pg.ddx.io pgsql-hackers@postgresql.org mailing list archive
help / color / mirror / Atom feedRe: AIO v2.2
16+ messages / 4 participants
[nested] [flat]
* Re: AIO v2.2
@ 2025-01-01 04:03 Andres Freund <andres@anarazel.de>
0 siblings, 3 replies; 16+ messages in thread
From: Andres Freund @ 2025-01-01 04:03 UTC (permalink / raw)
To: pgsql-hackers
Hi,
Attached is a new version of the AIO patchset.
The biggest changes are:
- The README has been extended with an overview of the API. I think it gives a
good overview of how the API fits together. I'd be very good to get
feedback from folks that aren't as familiar with AIO, I can't really see
what's easy/hard anymore.
- The read/write patches and the bounce buffer patches are split out, so that
there's no dependency between the first few AIO patches and the "don't dirty
while IO is going on" patcheset [1].
- Retries for partial IOs (i.e. short reads) are now implemented. Turned out
to take all of three lines and adding one missing variable initialization.
- I added quite a lot of function-header and file-header comments. There's
more to be done here, but see also the TODO section below.
- IO stats are now tracked. Specifically, the "time" for an IO is now the time
spent waiting for an IO, as discussed around [2]. I haven't updated the
docs yet.
- There now is a fastpath for executing AIO "synchronously", i.e. preparing an
IO and immediately submitting it.
- Previously one needed very large effective_io_concurrency values to get
sufficient asynchronous IO for sequential scans, as read_stream.c limited
max_pinned_buffers to effective_io_concurrency * 4. Unless
effective_io_concurrency was very high, that'd only allow a single IO to be
in-flight, due to io_combine_limit buffers getting merged into one IO.
Instead the pin limit is now capped by effective_io_concurrency *
io_combine_limit.
Right now that's part of one larger "hack up read_stream.c" commit, Thomas
said he'd take a look at how to do this properly. This is probably
something we could and should commit separately.
- io_method = sync has been made more similar to the way IO happens today. In
particular, we now continue to issue prefetch requests and the actual IO is
done only within WaitReadBuffers().
- When using buffered IO with io_uring, there previously was a small
regression, due to more IO happening in the process context with io_uring
(instead of in a kernel thread). While one could argue that it's better to
not increase CPU usage beyond one process, I don't find that sufficiently
convincing. To work around that I added a heuritic that tells IO uring to
execute IOs using it's worker infrastructure. That seems to have fixed this
problem entirely.
- IO worker infrastructure was cleaned up
- I pushed a few minor preliminary commits a while ago
- lots of other smaller stuff
The biggest TODOs are:
- Right now the API between bufmgr.c and read_stream.c kind of necessitates
that one StartReadBuffers() call actually can trigger multiple IOs, if
one of the buffers was read in by another backend, before "this" backend
called StartBufferIO().
I think Thomas and I figured out a way to evolve the interface so that this
isn't necessary anymore:
We allow StartReadBuffers() to memorize buffers it pinned but didn't
initiate IO on in the buffers[] argument. The next call to StartReadBuffers
then doesn't have to repin thse buffers. That doesn't just solve the
multiple-IOs for one "read operation" issue, it also make the - very common
- case of a bunch of "buffer misses" followed by a "buffer hit" cleaner, the
hit wouldn't be tracked in the same ReadBuffersOperation anymore.
- Right now bufmgr.h includes aio.h, because it needs to include a reference
to the AIO's result in ReadBuffersOperation. Requiring a dynamic allocation
would be noticeable overhead, so that's not an option. I think the best
option here would be to introduce something like aio_types.h, so fewer
things are included.
- There's no obvious way to tell "internal" function operating on an IO handle
apart from functions that are expected to be called by the issuer of an IO.
One way to deal with this would be to introduce a distinct "issuer IO
reference" type. I think that might be a good idea, it would also make it
clearer that a good number of the functions can only be called by the
issuer, before the IO is submitted.
This would also make it easier to order functions more sensibly in aio.c, as
all the issuer functions would be together.
The functions on AIO handles that everyone can call already have a distinct
type (PgAioHandleRef vs PgAioHandle*).
- While I've added a lot of comments, I only got so far adding them. More are
needed.
- The naming around PgAioReturn, PgAioResult, PgAioResultStatus needs to be
improved
- The debug logging functions are a bit of a mess, lots of very similar code
in lots of places. I think AIO needs a few ereport() wrappers to make this
easier.
- More tests are needed. None of our current test frameworks really makes this
easy :(.
- Several folks asked for pg_stat_aio to come back, in "v1" that showed the
set of currently in-flight AIOs. That's not particularly hard - except
that it doesn't really fit in the pg_stat_* namespace.
- I'm not sure that effective_io_concurrency as we have it right now really
makes sense, particularly not with the current default values. But that's a
mostly independent change.
Greetings,
Andres Freund
[1] https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf%40gcnactj4z56m
[2] https://postgr.es/m/tp63m6tcbi7mmsjlqgxd55sghhwvjxp3mkgeljffkbaujezvdl%40fvmdr3c6uhat
Attachments:
[text/x-diff] v2-0001-Ensure-a-resowner-exists-for-all-paths-that-may-p.patch (2.6K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/2-v2-0001-Ensure-a-resowner-exists-for-all-paths-that-may-p.patch)
download | inline diff:
From 42af1a44eadbfc3ac4e65ab23d280d6933b23284 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 8 Oct 2024 14:34:38 -0400
Subject: [PATCH v2 01/20] Ensure a resowner exists for all paths that may
perform AIO
Reviewed-by: Noah Misch <noah@leadboat.com>
Discussion: https://postgr.es/m/1f6b50a7-38ef-4d87-8246-786d39f46ab9@iki.fi
---
src/backend/bootstrap/bootstrap.c | 7 +++++++
src/backend/replication/logical/logical.c | 6 ++++++
src/backend/utils/init/postinit.c | 6 +++++-
3 files changed, 18 insertions(+), 1 deletion(-)
diff --git a/src/backend/bootstrap/bootstrap.c b/src/backend/bootstrap/bootstrap.c
index e0cb70ee9da..8ddcab0182a 100644
--- a/src/backend/bootstrap/bootstrap.c
+++ b/src/backend/bootstrap/bootstrap.c
@@ -361,8 +361,15 @@ BootstrapModeMain(int argc, char *argv[], bool check_only)
BaseInit();
bootstrap_signals();
+
+ /* need a resowner for IO during BootStrapXLOG() */
+ CreateAuxProcessResourceOwner();
+
BootStrapXLOG(bootstrap_data_checksum_version);
+ ReleaseAuxProcessResources(true);
+ CurrentResourceOwner = NULL;
+
/*
* To ensure that src/common/link-canary.c is linked into the backend, we
* must call it from somewhere. Here is as good as anywhere.
diff --git a/src/backend/replication/logical/logical.c b/src/backend/replication/logical/logical.c
index 4dc14fdb495..76fce6749a9 100644
--- a/src/backend/replication/logical/logical.c
+++ b/src/backend/replication/logical/logical.c
@@ -386,6 +386,12 @@ CreateInitDecodingContext(const char *plugin,
slot->data.plugin = plugin_name;
SpinLockRelease(&slot->mutex);
+ if (CurrentResourceOwner == NULL)
+ {
+ Assert(am_walsender);
+ CurrentResourceOwner = AuxProcessResourceOwner;
+ }
+
if (XLogRecPtrIsInvalid(restart_lsn))
ReplicationSlotReserveWal();
else
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 01c4016ced6..8a09c939eff 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -755,8 +755,12 @@ InitPostgres(const char *in_dbname, Oid dboid,
* We don't yet have an aux-process resource owner, but StartupXLOG
* and ShutdownXLOG will need one. Hence, create said resource owner
* (and register a callback to clean it up after ShutdownXLOG runs).
+ *
+ * In bootstrap mode CreateAuxProcessResourceOwner() was already
+ * called in BootstrapModeMain().
*/
- CreateAuxProcessResourceOwner();
+ if (!bootstrap)
+ CreateAuxProcessResourceOwner();
StartupXLOG();
/* Release (and warn about) any buffer pins leaked in StartupXLOG */
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0002-Allow-lwlocks-to-be-unowned.patch (5.0K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/3-v2-0002-Allow-lwlocks-to-be-unowned.patch)
download | inline diff:
From 5eff74f7f0bd0cf7102a04263a0dc9c0439123ed Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 5 Jan 2021 10:10:36 -0800
Subject: [PATCH v2 02/20] Allow lwlocks to be unowned
This is required for AIO so that the lock hold during a write can be released
in another backend. Which in turn is required to avoid the potential for
deadlocks.
---
src/include/storage/lwlock.h | 2 +
src/backend/storage/lmgr/lwlock.c | 110 ++++++++++++++++++++++--------
2 files changed, 82 insertions(+), 30 deletions(-)
diff --git a/src/include/storage/lwlock.h b/src/include/storage/lwlock.h
index d70e6d37e09..eabf813ce05 100644
--- a/src/include/storage/lwlock.h
+++ b/src/include/storage/lwlock.h
@@ -129,6 +129,8 @@ extern bool LWLockAcquireOrWait(LWLock *lock, LWLockMode mode);
extern void LWLockRelease(LWLock *lock);
extern void LWLockReleaseClearVar(LWLock *lock, pg_atomic_uint64 *valptr, uint64 val);
extern void LWLockReleaseAll(void);
+extern LWLockMode LWLockDisown(LWLock *l);
+extern void LWLockReleaseUnowned(LWLock *l, LWLockMode mode);
extern bool LWLockHeldByMe(LWLock *lock);
extern bool LWLockAnyHeldByMe(LWLock *lock, int nlocks, size_t stride);
extern bool LWLockHeldByMeInMode(LWLock *lock, LWLockMode mode);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 9cf3e4f4f3a..bc459dc5d2b 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -1773,52 +1773,36 @@ LWLockUpdateVar(LWLock *lock, pg_atomic_uint64 *valptr, uint64 val)
}
}
-
-/*
- * LWLockRelease - release a previously acquired lock
- */
-void
-LWLockRelease(LWLock *lock)
+static void
+LWLockReleaseInternal(LWLock *lock, LWLockMode mode)
{
- LWLockMode mode;
uint32 oldstate;
bool check_waiters;
- int i;
-
- /*
- * Remove lock from list of locks held. Usually, but not always, it will
- * be the latest-acquired lock; so search array backwards.
- */
- for (i = num_held_lwlocks; --i >= 0;)
- if (lock == held_lwlocks[i].lock)
- break;
-
- if (i < 0)
- elog(ERROR, "lock %s is not held", T_NAME(lock));
-
- mode = held_lwlocks[i].mode;
-
- num_held_lwlocks--;
- for (; i < num_held_lwlocks; i++)
- held_lwlocks[i] = held_lwlocks[i + 1];
-
- PRINT_LWDEBUG("LWLockRelease", lock, mode);
/*
* Release my hold on lock, after that it can immediately be acquired by
* others, even if we still have to wakeup other waiters.
*/
if (mode == LW_EXCLUSIVE)
- oldstate = pg_atomic_sub_fetch_u32(&lock->state, LW_VAL_EXCLUSIVE);
+ oldstate = pg_atomic_fetch_sub_u32(&lock->state, LW_VAL_EXCLUSIVE);
else
- oldstate = pg_atomic_sub_fetch_u32(&lock->state, LW_VAL_SHARED);
+ oldstate = pg_atomic_fetch_sub_u32(&lock->state, LW_VAL_SHARED);
/* nobody else can have that kind of lock */
- Assert(!(oldstate & LW_VAL_EXCLUSIVE));
+ if (mode == LW_EXCLUSIVE)
+ Assert((oldstate & LW_LOCK_MASK) == LW_VAL_EXCLUSIVE);
+ else
+ Assert((oldstate & LW_LOCK_MASK) < LW_VAL_EXCLUSIVE &&
+ (oldstate & LW_LOCK_MASK) >= LW_VAL_SHARED);
if (TRACE_POSTGRESQL_LWLOCK_RELEASE_ENABLED())
TRACE_POSTGRESQL_LWLOCK_RELEASE(T_NAME(lock));
+ if (mode == LW_EXCLUSIVE)
+ oldstate -= LW_VAL_EXCLUSIVE;
+ else
+ oldstate -= LW_VAL_SHARED;
+
/*
* We're still waiting for backends to get scheduled, don't wake them up
* again.
@@ -1841,6 +1825,72 @@ LWLockRelease(LWLock *lock)
LWLockWakeup(lock);
}
+ TRACE_POSTGRESQL_LWLOCK_RELEASE(T_NAME(lock));
+}
+
+void
+LWLockReleaseUnowned(LWLock *lock, LWLockMode mode)
+{
+ LWLockReleaseInternal(lock, mode);
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * After calling this function it's the callers responsibility to ensure that
+ * the lock gets released, even in case of an error. This only is desirable if
+ * the lock is going to be released in a different process than the process
+ * that acquired it.
+ *
+ * Returns the mode in which the lock was held by the current backend.
+ *
+ * NB: This will leave lock->owner pointing to the current backend (if
+ * LOCK_DEBUG is set). We could add a separate flag indicating that, but it
+ * doesn't really seem worth it.
+ *
+ * NB: This does not call RESUME_INTERRUPTS(), but leaves that responsibility
+ * of the caller.
+ */
+LWLockMode
+LWLockDisown(LWLock *lock)
+{
+ LWLockMode mode;
+ int i;
+
+ /*
+ * Remove lock from list of locks held. Usually, but not always, it will
+ * be the latest-acquired lock; so search array backwards.
+ */
+ for (i = num_held_lwlocks; --i >= 0;)
+ if (lock == held_lwlocks[i].lock)
+ break;
+
+ if (i < 0)
+ elog(ERROR, "lock %s is not held", T_NAME(lock));
+
+ mode = held_lwlocks[i].mode;
+
+ num_held_lwlocks--;
+ for (; i < num_held_lwlocks; i++)
+ held_lwlocks[i] = held_lwlocks[i + 1];
+
+ return mode;
+}
+
+/*
+ * LWLockRelease - release a previously acquired lock
+ */
+void
+LWLockRelease(LWLock *lock)
+{
+ LWLockMode mode;
+
+ mode = LWLockDisown(lock);
+
+ PRINT_LWDEBUG("LWLockRelease", lock, mode);
+
+ LWLockReleaseInternal(lock, mode);
+
/*
* Now okay to allow cancel/die interrupts.
*/
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0003-aio-Basic-subsystem-initialization.patch (9.8K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/4-v2-0003-aio-Basic-subsystem-initialization.patch)
download | inline diff:
From 93547a5a5b72fa0689b812ee6336b74c74eb95d7 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 10 Jun 2024 13:42:58 -0700
Subject: [PATCH v2 03/20] aio: Basic subsystem initialization
This is just separate to make it easier to review the tendrils into various
places.
---
src/include/storage/aio.h | 42 +++++++++++++++++++
src/include/storage/aio_init.h | 24 +++++++++++
src/backend/storage/aio/Makefile | 2 +
src/backend/storage/aio/aio.c | 33 +++++++++++++++
src/backend/storage/aio/aio_init.c | 41 ++++++++++++++++++
src/backend/storage/aio/meson.build | 2 +
src/backend/storage/ipc/ipci.c | 3 ++
src/backend/utils/init/postinit.c | 7 ++++
src/backend/utils/misc/guc_tables.c | 23 ++++++++++
src/backend/utils/misc/postgresql.conf.sample | 11 +++++
src/tools/pgindent/typedefs.list | 1 +
11 files changed, 189 insertions(+)
create mode 100644 src/include/storage/aio.h
create mode 100644 src/include/storage/aio_init.h
create mode 100644 src/backend/storage/aio/aio.c
create mode 100644 src/backend/storage/aio/aio_init.c
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
new file mode 100644
index 00000000000..0ee9d0043de
--- /dev/null
+++ b/src/include/storage/aio.h
@@ -0,0 +1,42 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio.h
+ * Main AIO interface
+ *
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/aio.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef AIO_H
+#define AIO_H
+
+
+#include "utils/guc_tables.h"
+
+
+/* GUC related */
+extern void assign_io_method(int newval, void *extra);
+
+
+/* Enum for io_method GUC. */
+typedef enum IoMethod
+{
+ IOMETHOD_SYNC = 0,
+} IoMethod;
+
+
+/* We'll default to synchronous execution. */
+#define DEFAULT_IO_METHOD IOMETHOD_SYNC
+
+
+/* GUCs */
+extern const struct config_enum_entry io_method_options[];
+extern int io_method;
+extern int io_max_concurrency;
+
+
+#endif /* AIO_H */
diff --git a/src/include/storage/aio_init.h b/src/include/storage/aio_init.h
new file mode 100644
index 00000000000..1c1d62baa79
--- /dev/null
+++ b/src/include/storage/aio_init.h
@@ -0,0 +1,24 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_init.h
+ * AIO initialization - kept separate as initialization sites don't need to
+ * know about AIO itself and AIO users don't need to know about initialization.
+ *
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/aio_init.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef AIO_INIT_H
+#define AIO_INIT_H
+
+
+extern Size AioShmemSize(void);
+extern void AioShmemInit(void);
+
+extern void pgaio_init_backend(void);
+
+#endif /* AIO_INIT_H */
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index 2f29a9ec4d1..eaeaeeee8e3 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -9,6 +9,8 @@ top_builddir = ../../../..
include $(top_builddir)/src/Makefile.global
OBJS = \
+ aio.o \
+ aio_init.o \
read_stream.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
new file mode 100644
index 00000000000..72110c0df3e
--- /dev/null
+++ b/src/backend/storage/aio/aio.c
@@ -0,0 +1,33 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio.c
+ * AIO - Core Logic
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/aio.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "storage/aio.h"
+
+
+/* Options for io_method. */
+const struct config_enum_entry io_method_options[] = {
+ {"sync", IOMETHOD_SYNC, false},
+ {NULL, 0, false}
+};
+
+int io_method = DEFAULT_IO_METHOD;
+int io_max_concurrency = -1;
+
+
+void
+assign_io_method(int newval, void *extra)
+{
+}
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
new file mode 100644
index 00000000000..84e0e37baae
--- /dev/null
+++ b/src/backend/storage/aio/aio_init.c
@@ -0,0 +1,41 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_init.c
+ * AIO - Subsystem Initialization
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/aio_init.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "storage/aio_init.h"
+
+
+Size
+AioShmemSize(void)
+{
+ Size sz = 0;
+
+ return sz;
+}
+
+void
+AioShmemInit(void)
+{
+}
+
+void
+pgaio_init_backend(void)
+{
+}
+
+void
+pgaio_postmaster_child_init_local(void)
+{
+}
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index 10e1aa3b20b..8d20759ebf8 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -1,5 +1,7 @@
# Copyright (c) 2024, PostgreSQL Global Development Group
backend_sources += files(
+ 'aio.c',
+ 'aio_init.c',
'read_stream.c',
)
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 7783ba854fc..c7703e5178e 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -37,6 +37,7 @@
#include "replication/slotsync.h"
#include "replication/walreceiver.h"
#include "replication/walsender.h"
+#include "storage/aio_init.h"
#include "storage/bufmgr.h"
#include "storage/dsm.h"
#include "storage/dsm_registry.h"
@@ -148,6 +149,7 @@ CalculateShmemSize(int *num_semaphores)
size = add_size(size, WaitEventCustomShmemSize());
size = add_size(size, InjectionPointShmemSize());
size = add_size(size, SlotSyncShmemSize());
+ size = add_size(size, AioShmemSize());
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -340,6 +342,7 @@ CreateOrAttachShmemStructs(void)
StatsShmemInit();
WaitEventCustomShmemInit();
InjectionPointShmemInit();
+ AioShmemInit();
}
/*
diff --git a/src/backend/utils/init/postinit.c b/src/backend/utils/init/postinit.c
index 8a09c939eff..9d1025e815b 100644
--- a/src/backend/utils/init/postinit.c
+++ b/src/backend/utils/init/postinit.c
@@ -43,6 +43,7 @@
#include "replication/slot.h"
#include "replication/slotsync.h"
#include "replication/walsender.h"
+#include "storage/aio_init.h"
#include "storage/bufmgr.h"
#include "storage/fd.h"
#include "storage/ipc.h"
@@ -626,6 +627,12 @@ BaseInit(void)
*/
pgstat_initialize();
+ /*
+ * Initialize AIO before infrastructure that might need to actually
+ * execute AIO.
+ */
+ pgaio_init_backend();
+
/* Do local initialization of storage and buffer managers */
InitSync();
smgrinit();
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 8cf1afbad20..6d4056c68b9 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -71,6 +71,7 @@
#include "replication/slot.h"
#include "replication/slotsync.h"
#include "replication/syncrep.h"
+#include "storage/aio.h"
#include "storage/bufmgr.h"
#include "storage/bufpage.h"
#include "storage/large_object.h"
@@ -3219,6 +3220,18 @@ struct config_int ConfigureNamesInt[] =
NULL, NULL, NULL
},
+ {
+ {"io_max_concurrency",
+ PGC_POSTMASTER,
+ RESOURCES_ASYNCHRONOUS,
+ gettext_noop("Number of IOs that may be in flight in one backend."),
+ NULL,
+ },
+ &io_max_concurrency,
+ -1, -1, 1024,
+ NULL, NULL, NULL
+ },
+
{
{"backend_flush_after", PGC_USERSET, RESOURCES_ASYNCHRONOUS,
gettext_noop("Number of pages after which previously performed writes are flushed to disk."),
@@ -5226,6 +5239,16 @@ struct config_enum ConfigureNamesEnum[] =
NULL, NULL, NULL
},
+ {
+ {"io_method", PGC_POSTMASTER, RESOURCES_MEM,
+ gettext_noop("Selects the method of asynchronous I/O to use."),
+ NULL
+ },
+ &io_method,
+ DEFAULT_IO_METHOD, io_method_options,
+ NULL, assign_io_method, NULL
+ },
+
/* End-of-list marker */
{
{NULL, 0, 0, NULL, NULL}, NULL, 0, NULL, NULL, NULL, NULL
diff --git a/src/backend/utils/misc/postgresql.conf.sample b/src/backend/utils/misc/postgresql.conf.sample
index a2ac7575ca7..c4c60da9845 100644
--- a/src/backend/utils/misc/postgresql.conf.sample
+++ b/src/backend/utils/misc/postgresql.conf.sample
@@ -838,6 +838,17 @@
#include = '...' # include file
+#------------------------------------------------------------------------------
+# WIP AIO GUC docs
+#------------------------------------------------------------------------------
+
+#io_method = sync # (change requires restart)
+
+#io_max_concurrency = 32 # Max number of IOs that may be in
+ # flight at the same time in one backend
+ # (change requires restart)
+
+
#------------------------------------------------------------------------------
# CUSTOMIZED OPTIONS
#------------------------------------------------------------------------------
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index e1c4f913f84..2586d1cf53f 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1262,6 +1262,7 @@ IntoClause
InvalMessageArray
InvalidationInfo
InvalidationMsgsGroup
+IoMethod
IpcMemoryId
IpcMemoryKey
IpcMemoryState
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0004-aio-Core-AIO-implementation.patch (63.5K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/5-v2-0004-aio-Core-AIO-implementation.patch)
download | inline diff:
From b64c247210c5a5067b5c76f6ab68c978606b0902 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 9 Dec 2024 14:14:13 -0500
Subject: [PATCH v2 04/20] aio: Core AIO implementation
At this point nothing can use AIO - this commit does not include any
implementation of aio subjects / callbacks. That will come in later commits.
Todo:
- lots of cleanup
---
src/include/storage/aio.h | 296 ++++++
src/include/storage/aio_internal.h | 244 +++++
src/include/storage/aio_ref.h | 24 +
src/include/utils/resowner.h | 5 +
src/backend/access/transam/xact.c | 9 +
src/backend/storage/aio/Makefile | 3 +
src/backend/storage/aio/aio.c | 906 ++++++++++++++++++
src/backend/storage/aio/aio_init.c | 186 +++-
src/backend/storage/aio/aio_io.c | 140 +++
src/backend/storage/aio/aio_subject.c | 231 +++++
src/backend/storage/aio/meson.build | 3 +
src/backend/storage/aio/method_sync.c | 45 +
.../utils/activity/wait_event_names.txt | 3 +
src/backend/utils/resowner/resowner.c | 30 +
src/tools/pgindent/typedefs.list | 18 +
15 files changed, 2139 insertions(+), 4 deletions(-)
create mode 100644 src/include/storage/aio_internal.h
create mode 100644 src/include/storage/aio_ref.h
create mode 100644 src/backend/storage/aio/aio_io.c
create mode 100644 src/backend/storage/aio/aio_subject.c
create mode 100644 src/backend/storage/aio/method_sync.c
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index 0ee9d0043de..b386dabc921 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -15,9 +15,305 @@
#define AIO_H
+#include "storage/aio_ref.h"
+#include "storage/procnumber.h"
#include "utils/guc_tables.h"
+typedef struct PgAioHandle PgAioHandle;
+
+typedef enum PgAioOp
+{
+ /* intentionally the zero value, to help catch zeroed memory etc */
+ PGAIO_OP_INVALID = 0,
+
+ PGAIO_OP_READV,
+ PGAIO_OP_WRITEV,
+
+ /**
+ * In the near term we'll need at least:
+ * - fsync / fdatasync
+ * - flush_range
+ *
+ * Eventually we'll additionally want at least:
+ * - send
+ * - recv
+ * - accept
+ **/
+} PgAioOp;
+
+#define PGAIO_OP_COUNT (PGAIO_OP_WRITEV + 1)
+
+
+/*
+ * On what is IO being performed.
+ *
+ * PgAioSharedCallback specific behaviour should be implemented in
+ * aio_subject.c.
+ */
+typedef enum PgAioSubjectID
+{
+ /* intentionally the zero value, to help catch zeroed memory etc */
+ ASI_INVALID = 0,
+} PgAioSubjectID;
+
+#define ASI_COUNT (ASI_INVALID + 1)
+
+/*
+ * Flags for an IO that can be set with pgaio_io_set_flag().
+ */
+typedef enum PgAioHandleFlags
+{
+ /* hint that IO will be executed synchronously */
+ AHF_SYNCHRONOUS = 1 << 0,
+
+ /* the IO references backend local memory */
+ AHF_REFERENCES_LOCAL = 1 << 1,
+
+ /*
+ * IO is using buffered IO, used to control heuristic in some IO
+ * methods. Advantageous to set, if applicable, but not required for
+ * correctness.
+ */
+ AHF_BUFFERED = 1 << 2,
+} PgAioHandleFlags;
+
+
+/*
+ * IDs for callbacks that can be registered on an IO.
+ *
+ * Callbacks are identified by an ID rather than a function pointer. There are
+ * two main reasons:
+
+ * 1) Memory within PgAioHandle is precious, due to the number of PgAioHandle
+ * structs in pre-allocated shared memory.
+
+ * 2) Due to EXEC_BACKEND function pointers are not necessarily stable between
+ * different backends, therefore function pointers cannot directly be in
+ * shared memory.
+ *
+ * Without 2), we could fairly easily allow to add new callbacks, by filling a
+ * ID->pointer mapping table on demand. In the presence of 2 that's still
+ * doable, but harder, because every process has to re-register the pointers
+ * so that a local ID->"backend local pointer" mapping can be maintained.
+ */
+typedef enum PgAioHandleSharedCallbackID
+{
+ ASC_INVALID,
+} PgAioHandleSharedCallbackID;
+
+
+/*
+ * Data necessary for basic IO types (PgAioOp).
+ *
+ * NB: Note that the FDs in here may *not* be relied upon for re-issuing
+ * requests (e.g. for partial reads/writes) - the FD might be from another
+ * process, or closed since. That's not a problem for IOs waiting to be issued
+ * only because the queue is flushed when closing an FD.
+ */
+typedef union
+{
+ struct
+ {
+ int fd;
+ uint16 iov_length;
+ uint64 offset;
+ } read;
+
+ struct
+ {
+ int fd;
+ uint16 iov_length;
+ uint64 offset;
+ } write;
+} PgAioOpData;
+
+
+/* XXX: Perhaps it's worth moving this to a dedicated file? */
+#include "storage/block.h"
+#include "storage/relfilelocator.h"
+
+typedef union PgAioSubjectData
+{
+ /* just as an example placeholder for later */
+ struct
+ {
+ uint32 queue_id;
+ } wal;
+} PgAioSubjectData;
+
+
+typedef enum PgAioResultStatus
+{
+ ARS_UNKNOWN, /* not yet completed / uninitialized */
+ ARS_OK,
+ ARS_PARTIAL, /* did not fully succeed, but no error */
+ ARS_ERROR,
+} PgAioResultStatus;
+
+typedef struct PgAioResult
+{
+ /*
+ * This is of type PgAioHandleSharedCallbackID, but can't use a bitfield
+ * of an enum, because some compilers treat enums as signed.
+ */
+ uint32 id:8;
+
+ /* of type PgAioResultStatus, see above */
+ uint32 status:2;
+
+ /* meaning defined by callback->error */
+ uint32 error_data:22;
+
+ int32 result;
+} PgAioResult;
+
+/*
+ * Result of IO operation, visible only to the initiator of IO.
+ */
+typedef struct PgAioReturn
+{
+ PgAioResult result;
+ PgAioSubjectData subject_data;
+} PgAioReturn;
+
+
+typedef struct PgAioSubjectInfo
+{
+ void (*reopen) (PgAioHandle *ioh);
+
+#ifdef NOT_YET
+ char *(*describe_identity) (PgAioHandle *ioh);
+#endif
+
+ const char *name;
+} PgAioSubjectInfo;
+
+
+typedef PgAioResult (*PgAioHandleSharedCallbackComplete) (PgAioHandle *ioh, PgAioResult prior_result);
+typedef void (*PgAioHandleSharedCallbackPrepare) (PgAioHandle *ioh);
+typedef void (*PgAioHandleSharedCallbackError) (PgAioResult result, const PgAioSubjectData *subject_data, int elevel);
+
+typedef struct PgAioHandleSharedCallbacks
+{
+ PgAioHandleSharedCallbackPrepare prepare;
+ PgAioHandleSharedCallbackComplete complete;
+ PgAioHandleSharedCallbackError error;
+} PgAioHandleSharedCallbacks;
+
+
+
+/*
+ * How many callbacks can be registered for one IO handle. Currently we only
+ * need two, but it's not hard to imagine needing a few more.
+ */
+#define AIO_MAX_SHARED_CALLBACKS 4
+
+
+
+/* AIO API */
+
+
+/* --------------------------------------------------------------------------------
+ * IO Handles
+ * --------------------------------------------------------------------------------
+ */
+
+struct ResourceOwnerData;
+extern PgAioHandle *pgaio_io_get(struct ResourceOwnerData *resowner, PgAioReturn *ret);
+extern PgAioHandle *pgaio_io_get_nb(struct ResourceOwnerData *resowner, PgAioReturn *ret);
+
+extern void pgaio_io_release(PgAioHandle *ioh);
+extern void pgaio_io_release_resowner(dlist_node *ioh_node, bool on_error);
+
+extern void pgaio_io_get_ref(PgAioHandle *ioh, PgAioHandleRef *ior);
+
+extern void pgaio_io_set_subject(PgAioHandle *ioh, PgAioSubjectID subjid);
+extern void pgaio_io_set_flag(PgAioHandle *ioh, PgAioHandleFlags flag);
+
+extern void pgaio_io_add_shared_cb(PgAioHandle *ioh, PgAioHandleSharedCallbackID cbid);
+
+extern void pgaio_io_set_io_data_32(PgAioHandle *ioh, uint32 *data, uint8 len);
+extern void pgaio_io_set_io_data_64(PgAioHandle *ioh, uint64 *data, uint8 len);
+extern uint64 *pgaio_io_get_io_data(PgAioHandle *ioh, uint8 *len);
+
+extern void pgaio_io_prepare(PgAioHandle *ioh, PgAioOp op);
+
+extern int pgaio_io_get_id(PgAioHandle *ioh);
+struct iovec;
+extern int pgaio_io_get_iovec(PgAioHandle *ioh, struct iovec **iov);
+extern bool pgaio_io_has_subject(PgAioHandle *ioh);
+
+extern PgAioSubjectData *pgaio_io_get_subject_data(PgAioHandle *ioh);
+extern PgAioOpData *pgaio_io_get_op_data(PgAioHandle *ioh);
+extern ProcNumber pgaio_io_get_owner(PgAioHandle *ioh);
+
+
+
+/* --------------------------------------------------------------------------------
+ * IO References
+ * --------------------------------------------------------------------------------
+ */
+
+extern void pgaio_io_ref_clear(PgAioHandleRef *ior);
+extern bool pgaio_io_ref_valid(PgAioHandleRef *ior);
+extern int pgaio_io_ref_get_id(PgAioHandleRef *ior);
+
+
+extern void pgaio_io_ref_wait(PgAioHandleRef *ior);
+extern bool pgaio_io_ref_check_done(PgAioHandleRef *ior);
+
+
+
+/* --------------------------------------------------------------------------------
+ * IO Result
+ * --------------------------------------------------------------------------------
+ */
+
+extern void pgaio_result_log(PgAioResult result, const PgAioSubjectData *subject_data,
+ int elevel);
+
+
+
+/* --------------------------------------------------------------------------------
+ * Actions on multiple IOs.
+ * --------------------------------------------------------------------------------
+ */
+
+extern void pgaio_submit_staged(void);
+extern bool pgaio_have_staged(void);
+
+
+
+/* --------------------------------------------------------------------------------
+ * Low level IO preparation routines
+ *
+ * These will often be called by code lowest level of initiating an
+ * IO. E.g. bufmgr.c may initiate IO for a buffer, but pgaio_io_prep_readv()
+ * will be called from within fd.c.
+ *
+ * Implemented in aio_io.c
+ * --------------------------------------------------------------------------------
+ */
+
+extern void pgaio_io_prep_readv(PgAioHandle *ioh,
+ int fd, int iovcnt, uint64 offset);
+
+extern void pgaio_io_prep_writev(PgAioHandle *ioh,
+ int fd, int iovcnt, uint64 offset);
+
+
+
+/* --------------------------------------------------------------------------------
+ * Other
+ * --------------------------------------------------------------------------------
+ */
+
+extern void pgaio_closing_fd(int fd);
+extern void pgaio_at_xact_end(bool is_subxact, bool is_commit);
+extern void pgaio_at_error(void);
+
+
/* GUC related */
extern void assign_io_method(int newval, void *extra);
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
new file mode 100644
index 00000000000..d600d45b4fd
--- /dev/null
+++ b/src/include/storage/aio_internal.h
@@ -0,0 +1,244 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_internal.h
+ * aio_internal
+ *
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/aio_internal.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef AIO_INTERNAL_H
+#define AIO_INTERNAL_H
+
+
+#include "lib/ilist.h"
+#include "port/pg_iovec.h"
+#include "storage/aio.h"
+#include "storage/condition_variable.h"
+
+
+#define PGAIO_VERBOSE
+
+
+/* AFIXME */
+#define PGAIO_SUBMIT_BATCH_SIZE 32
+
+
+
+typedef enum PgAioHandleState
+{
+ /* not in use */
+ AHS_IDLE = 0,
+
+ /* returned by pgaio_io_get() */
+ AHS_HANDED_OUT,
+
+ /* pgaio_io_start_*() has been called, but IO hasn't been submitted yet */
+ AHS_DEFINED,
+
+ /* subjects prepare() callback has been called */
+ AHS_PREPARED,
+
+ /* IO is being executed */
+ AHS_IN_FLIGHT,
+
+ /* IO finished, but result has not yet been processed */
+ AHS_REAPED,
+
+ /* IO completed, shared completion has been called */
+ AHS_COMPLETED_SHARED,
+
+ /* IO completed, local completion has been called */
+ AHS_COMPLETED_LOCAL,
+} PgAioHandleState;
+
+
+struct ResourceOwnerData;
+
+/* typedef is in public header */
+struct PgAioHandle
+{
+ /* all state updates should go through pgaio_io_update_state() */
+ PgAioHandleState state:8;
+
+ /* what are we operating on */
+ PgAioSubjectID subject:8;
+
+ /* which operation */
+ PgAioOp op:8;
+
+ /* bitfield of PgAioHandleFlags */
+ uint8 flags;
+
+ uint8 num_shared_callbacks;
+
+ /* using the proper type here would use more space */
+ uint8 shared_callbacks[AIO_MAX_SHARED_CALLBACKS];
+
+ uint8 iovec_data_len;
+
+ /* XXX: could be optimized out with some pointer math */
+ int32 owner_procno;
+
+ /* FIXME: remove in favor of distilled_result */
+ /* raw result of the IO operation */
+ int32 result;
+
+ /* index into PgAioCtl->iovecs */
+ uint32 iovec_off;
+
+ /**
+ * In which list the handle is registered, depends on the state:
+ * - IDLE, in per-backend list
+ * - HANDED_OUT - not in a list
+ * - DEFINED - in per-backend staged list
+ * - PREPARED - in per-backend staged list
+ * - IN_FLIGHT - in issuer's in_flight list
+ * - REAPED - in issuer's in_flight list
+ * - COMPLETED_SHARED - in issuer's in_flight list
+ * - COMPLETED_LOCAL - in issuer's in_flight list
+ *
+ * XXX: It probably make sense to optimize this out to save on per-io
+ * memory at the cost of per-backend memory.
+ **/
+ dlist_node node;
+
+ struct ResourceOwnerData *resowner;
+ dlist_node resowner_node;
+
+ /* incremented every time the IO handle is reused */
+ uint64 generation;
+
+ ConditionVariable cv;
+
+ /* result of shared callback, passed to issuer callback */
+ PgAioResult distilled_result;
+
+ PgAioReturn *report_return;
+
+ PgAioOpData op_data;
+
+ /*
+ * Data necessary for shared completions. Needs to be sufficient to allow
+ * another backend to retry an IO.
+ */
+ PgAioSubjectData scb_data;
+};
+
+
+typedef struct PgAioPerBackend
+{
+ /* index into PgAioCtl->io_handles */
+ uint32 io_handle_off;
+
+ /* IO Handles that currently are not used */
+ dclist_head idle_ios;
+
+ /*
+ * Only one IO may be returned by pgaio_io_get()/pgaio_io_get() without
+ * having been either defined (by actually associating it with IO) or by
+ * released (with pgaio_io_release()). This restriction is necessary to
+ * guarantee that we always can acquire an IO. ->handed_out_io is used to
+ * enforce that rule.
+ */
+ PgAioHandle *handed_out_io;
+
+ /*
+ * IOs that are defined, but not yet submitted.
+ */
+ uint16 num_staged_ios;
+ PgAioHandle *staged_ios[PGAIO_SUBMIT_BATCH_SIZE];
+
+ /*
+ * List of in-flight IOs. Also contains IOs that aren't strict speaking
+ * in-flight anymore, but have been waited-for and completed by another
+ * backend. Once this backend sees such an IO it'll be reclaimed.
+ *
+ * The list is ordered by submission time, with more recently submitted
+ * IOs being appended at the end.
+ */
+ dclist_head in_flight_ios;
+} PgAioPerBackend;
+
+
+typedef struct PgAioCtl
+{
+ int backend_state_count;
+ PgAioPerBackend *backend_state;
+
+ /*
+ * Array of iovec structs. Each iovec is owned by a specific backend. The
+ * allocation is in PgAioCtl to allow the maximum number of iovecs for
+ * individual IOs to be configurable with PGC_POSTMASTER GUC.
+ */
+ uint64 iovec_count;
+ struct iovec *iovecs;
+
+ /*
+ * For, e.g., an IO covering multiple buffers in shared / temp buffers, we
+ * need to get Buffer IDs during completion to be able to change the
+ * BufferDesc state accordingly. This space can be used to store e.g.
+ * Buffer IDs. Note that the actual iovec might be shorter than this,
+ * because we combine neighboring pages into one larger iovec entry.
+ */
+ uint64 *iovecs_data;
+
+ uint64 io_handle_count;
+ PgAioHandle *io_handles;
+} PgAioCtl;
+
+
+
+/*
+ * The set of callbacks that each IO method must implement.
+ */
+typedef struct IoMethodOps
+{
+ /* global initialization */
+ size_t (*shmem_size) (void);
+ void (*shmem_init) (bool first_time);
+
+ /* per-backend initialization */
+ void (*init_backend) (void);
+
+ /* handling of IOs */
+ bool (*needs_synchronous_execution) (PgAioHandle *ioh);
+ int (*submit) (uint16 num_staged_ios, PgAioHandle **staged_ios);
+
+ void (*wait_one) (PgAioHandle *ioh,
+ uint64 ref_generation);
+} IoMethodOps;
+
+
+extern bool pgaio_io_was_recycled(PgAioHandle *ioh, uint64 ref_generation, PgAioHandleState *state);
+
+extern void pgaio_io_prepare_subject(PgAioHandle *ioh);
+extern void pgaio_io_process_completion_subject(PgAioHandle *ioh);
+extern void pgaio_io_process_completion(PgAioHandle *ioh, int result);
+extern void pgaio_io_prepare_submit(PgAioHandle *ioh);
+
+extern bool pgaio_io_needs_synchronous_execution(PgAioHandle *ioh);
+extern void pgaio_io_perform_synchronously(PgAioHandle *ioh);
+
+extern bool pgaio_io_can_reopen(PgAioHandle *ioh);
+extern void pgaio_io_reopen(PgAioHandle *ioh);
+
+extern const char *pgaio_io_get_subject_name(PgAioHandle *ioh);
+extern const char *pgaio_io_get_op_name(PgAioHandle *ioh);
+extern const char *pgaio_io_get_state_name(PgAioHandle *ioh);
+
+
+/* Declarations for the tables of function pointers exposed by each IO method. */
+extern const IoMethodOps pgaio_sync_ops;
+
+extern const IoMethodOps *pgaio_impl;
+extern PgAioCtl *aio_ctl;
+extern PgAioPerBackend *my_aio;
+
+
+
+#endif /* AIO_INTERNAL_H */
diff --git a/src/include/storage/aio_ref.h b/src/include/storage/aio_ref.h
new file mode 100644
index 00000000000..ad7e9ad34f3
--- /dev/null
+++ b/src/include/storage/aio_ref.h
@@ -0,0 +1,24 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_ref.h Definition of PgAioHandleRef, which sometimes needs to be used in
+ * headers.
+ *
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/aio_ref.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef AIO_REF_H
+#define AIO_REF_H
+
+typedef struct PgAioHandleRef
+{
+ uint32 aio_index;
+ uint32 generation_upper;
+ uint32 generation_lower;
+} PgAioHandleRef;
+
+#endif /* AIO_REF_H */
diff --git a/src/include/utils/resowner.h b/src/include/utils/resowner.h
index 4e534bc3e70..2d55720a54c 100644
--- a/src/include/utils/resowner.h
+++ b/src/include/utils/resowner.h
@@ -164,4 +164,9 @@ struct LOCALLOCK;
extern void ResourceOwnerRememberLock(ResourceOwner owner, struct LOCALLOCK *locallock);
extern void ResourceOwnerForgetLock(ResourceOwner owner, struct LOCALLOCK *locallock);
+/* special support for AIO */
+struct dlist_node;
+extern void ResourceOwnerRememberAioHandle(ResourceOwner owner, struct dlist_node *ioh_node);
+extern void ResourceOwnerForgetAioHandle(ResourceOwner owner, struct dlist_node *ioh_node);
+
#endif /* RESOWNER_H */
diff --git a/src/backend/access/transam/xact.c b/src/backend/access/transam/xact.c
index 3ebd7c40418..0356552c499 100644
--- a/src/backend/access/transam/xact.c
+++ b/src/backend/access/transam/xact.c
@@ -51,6 +51,7 @@
#include "replication/origin.h"
#include "replication/snapbuild.h"
#include "replication/syncrep.h"
+#include "storage/aio.h"
#include "storage/condition_variable.h"
#include "storage/fd.h"
#include "storage/lmgr.h"
@@ -2475,6 +2476,8 @@ CommitTransaction(void)
AtEOXact_LogicalRepWorkers(true);
pgstat_report_xact_timestamp(0);
+ pgaio_at_xact_end( /* is_subxact = */ false, /* is_commit = */ true);
+
ResourceOwnerDelete(TopTransactionResourceOwner);
s->curTransactionOwner = NULL;
CurTransactionResourceOwner = NULL;
@@ -2988,6 +2991,8 @@ AbortTransaction(void)
pgstat_report_xact_timestamp(0);
}
+ pgaio_at_xact_end( /* is_subxact = */ false, /* is_commit = */ false);
+
/*
* State remains TRANS_ABORT until CleanupTransaction().
*/
@@ -5185,6 +5190,8 @@ CommitSubTransaction(void)
AtEOSubXact_PgStat(true, s->nestingLevel);
AtSubCommit_Snapshot(s->nestingLevel);
+ pgaio_at_xact_end( /* is_subxact = */ true, /* is_commit = */ true);
+
/*
* We need to restore the upper transaction's read-only state, in case the
* upper is read-write while the child is read-only; GUC will incorrectly
@@ -5351,6 +5358,8 @@ AbortSubTransaction(void)
AtSubAbort_Snapshot(s->nestingLevel);
}
+ pgaio_at_xact_end( /* is_subxact = */ true, /* is_commit = */ false);
+
/*
* Restore the upper transaction's read-only state, too. This should be
* redundant with GUC's cleanup but we may as well do it for consistency
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index eaeaeeee8e3..b253278f3c1 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -11,6 +11,9 @@ include $(top_builddir)/src/Makefile.global
OBJS = \
aio.o \
aio_init.o \
+ aio_io.o \
+ aio_subject.o \
+ method_sync.o \
read_stream.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index 72110c0df3e..3e2ff9718ca 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -3,6 +3,28 @@
* aio.c
* AIO - Core Logic
*
+ * For documentation about how AIO works on a higher level, including a
+ * schematic example, see README.md.
+ *
+ *
+ * AIO is a complicated subsystem. To keep things navigable it is split across
+ * a number of files:
+ *
+ * - aio.c - core AIO state handling
+ *
+ * - aio_init.c - initialization
+ *
+ * - aio_io.c - dealing with actual IO, including executing IOs synchronously
+ *
+ * - aio_subject.c - functionality related to executing IO for different
+ * subjects
+ *
+ * - method_*.c - different ways of executing AIO
+ *
+ * - read_stream.c - helper for accessing buffered relation data with
+ * look-ahead
+ *
+ *
* Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
* Portions Copyright (c) 1994, Regents of the University of California
*
@@ -14,7 +36,22 @@
#include "postgres.h"
+#include "miscadmin.h"
+#include "port/atomics.h"
#include "storage/aio.h"
+#include "storage/aio_internal.h"
+#include "storage/bufmgr.h"
+#include "utils/resowner.h"
+#include "utils/wait_event_types.h"
+
+
+
+static inline void pgaio_io_update_state(PgAioHandle *ioh, PgAioHandleState new_state);
+static void pgaio_io_reclaim(PgAioHandle *ioh);
+static void pgaio_io_resowner_register(PgAioHandle *ioh);
+static void pgaio_io_wait_for_free(void);
+static PgAioHandle *pgaio_io_from_ref(PgAioHandleRef *ior, uint64 *ref_generation);
+
/* Options for io_method. */
@@ -27,7 +64,876 @@ int io_method = DEFAULT_IO_METHOD;
int io_max_concurrency = -1;
+/* global control for AIO */
+PgAioCtl *aio_ctl;
+
+/* current backend's per-backend state */
+PgAioPerBackend *my_aio;
+
+
+static const IoMethodOps *pgaio_ops_table[] = {
+ [IOMETHOD_SYNC] = &pgaio_sync_ops,
+};
+
+
+const IoMethodOps *pgaio_impl;
+
+
+
+/* --------------------------------------------------------------------------------
+ * "Core" IO Api
+ * --------------------------------------------------------------------------------
+ */
+
+/*
+ * Acquire an AioHandle, waiting for IO completion if necessary.
+ *
+ * Each backend can only have one AIO handle that that has been "handed out"
+ * to code, but not yet submitted or released. This restriction is necessary
+ * to ensure that it is possible for code to wait for an unused handle by
+ * waiting for in-flight IO to complete. There is a limited number of handles
+ * in each backend, if multiple handles could be handed out without being
+ * submitted, waiting for all in-flight IO to complete would not guarantee
+ * that handles free up.
+ *
+ * It is cheap to acquire an IO handle, unless all handles are in use. In that
+ * case this function waits for the oldest IO to complete. In case that is not
+ * desirable, see pgaio_io_get_nb().
+ *
+ * If a handle was acquired but then does not turn out to be needed,
+ * e.g. because pgaio_io_get() is called before starting an IO in a critical
+ * section, the handle needs to be be released with pgaio_io_release().
+ *
+ *
+ * To react to the completion of the IO as soon as it is know to have
+ * completed, callbacks can be registered with pgaio_io_add_shared_cb().
+ *
+ * To actually execute IO using the returned handle, the pgaio_io_prep_*()
+ * family of functions is used. In many cases the pgaio_io_prep_*() call will
+ * not be done directly by code that acquired the handle, but by lower level
+ * code that gets passed the handle. E.g. if code in bufmgr.c wants to perform
+ * AIO, it typically will pass the handle to smgr., which will pass it on to
+ * md.c, on to fd.c, which then finally calls pgaio_io_prep_*(). This
+ * forwarding allows the various layers to react to the IO's completion by
+ * registering callbacks. These callbacks in turn can translate a lower
+ * layer's result into a result understandable by a higher layer.
+ *
+ * Once pgaio_io_prep_*() is called, the IO may be in the process of being
+ * executed and might even complete before the functions return. That is,
+ * however, not guaranteed, to allow IO submission to be batched. To guarantee
+ * IO submission pgaio_submit_staged() needs to be called.
+ *
+ * After pgaio_io_prep_*() the AioHandle is "consumed" and may not be
+ * referenced by the IO issuing code. To e.g. wait for IO, references to the
+ * IO can be established with pgaio_io_get_ref() *before* pgaio_io_prep_*() is
+ * called. pgaio_io_ref_wait() can be used to wait for the IO to complete.
+ *
+ *
+ * To know if the IO [partially] succeeded or failed, a PgAioReturn * can be
+ * passed to pgaio_io_get(). Once the issuing backend has called
+ * pgaio_io_ref_wait(), the PgAioReturn contains information about whether the
+ * operation succeeded and details about the first failure, if any. The error
+ * can be raised / logged with pgaio_result_log().
+ *
+ * The lifetime of the memory pointed to be *ret needs to be at least as long
+ * as the passed in resowner. If the resowner releases resources before the IO
+ * completes, the reference to *ret will be cleared.
+ */
+PgAioHandle *
+pgaio_io_get(struct ResourceOwnerData *resowner, PgAioReturn *ret)
+{
+ PgAioHandle *h;
+
+ while (true)
+ {
+ h = pgaio_io_get_nb(resowner, ret);
+
+ if (h != NULL)
+ return h;
+
+ /*
+ * Evidently all handles by this backend are in use. Just wait for
+ * some to complete.
+ */
+ pgaio_io_wait_for_free();
+ }
+}
+
+/*
+ * Acquire an AioHandle, returning NULL if no handles are free.
+ *
+ * See pgaio_io_get(). The only difference is that this function will return
+ * NULL if there are no idle handles, instead of blocking.
+ */
+PgAioHandle *
+pgaio_io_get_nb(struct ResourceOwnerData *resowner, PgAioReturn *ret)
+{
+ if (my_aio->num_staged_ios >= PGAIO_SUBMIT_BATCH_SIZE)
+ {
+ Assert(my_aio->num_staged_ios == PGAIO_SUBMIT_BATCH_SIZE);
+ pgaio_submit_staged();
+ }
+
+ if (my_aio->handed_out_io)
+ {
+ ereport(ERROR,
+ errmsg("API violation: Only one IO can be handed out"));
+ }
+
+ if (!dclist_is_empty(&my_aio->idle_ios))
+ {
+ dlist_node *ion = dclist_pop_head_node(&my_aio->idle_ios);
+ PgAioHandle *ioh = dclist_container(PgAioHandle, node, ion);
+
+ Assert(ioh->state == AHS_IDLE);
+ Assert(ioh->owner_procno == MyProcNumber);
+
+ pgaio_io_update_state(ioh, AHS_HANDED_OUT);
+ my_aio->handed_out_io = ioh;
+
+ if (resowner)
+ pgaio_io_resowner_register(ioh);
+
+ if (ret)
+ {
+ ioh->report_return = ret;
+ ret->result.status = ARS_UNKNOWN;
+ }
+
+ return ioh;
+ }
+
+ return NULL;
+}
+
+/*
+ * Release IO handle that turned out to not be required.
+ *
+ * See pgaio_io_get() for more details.
+ */
+void
+pgaio_io_release(PgAioHandle *ioh)
+{
+ if (ioh == my_aio->handed_out_io)
+ {
+ Assert(ioh->state == AHS_HANDED_OUT);
+ Assert(ioh->resowner);
+
+ my_aio->handed_out_io = NULL;
+ pgaio_io_reclaim(ioh);
+ }
+ else
+ {
+ elog(ERROR, "release in unexpected state");
+ }
+}
+
+/*
+ * Release IO handle during resource owner cleanup.
+ */
+void
+pgaio_io_release_resowner(dlist_node *ioh_node, bool on_error)
+{
+ PgAioHandle *ioh = dlist_container(PgAioHandle, resowner_node, ioh_node);
+
+ Assert(ioh->resowner);
+
+ ResourceOwnerForgetAioHandle(ioh->resowner, &ioh->resowner_node);
+ ioh->resowner = NULL;
+
+ switch (ioh->state)
+ {
+ case AHS_IDLE:
+ elog(ERROR, "unexpected");
+ break;
+ case AHS_HANDED_OUT:
+ Assert(ioh == my_aio->handed_out_io || my_aio->handed_out_io == NULL);
+
+ if (ioh == my_aio->handed_out_io)
+ {
+ my_aio->handed_out_io = NULL;
+ if (!on_error)
+ elog(WARNING, "leaked AIO handle");
+ }
+
+ pgaio_io_reclaim(ioh);
+ break;
+ case AHS_DEFINED:
+ case AHS_PREPARED:
+ /* XXX: Should we warn about this when is_commit? */
+ pgaio_submit_staged();
+ break;
+ case AHS_IN_FLIGHT:
+ case AHS_REAPED:
+ case AHS_COMPLETED_SHARED:
+ /* this is expected to happen */
+ break;
+ case AHS_COMPLETED_LOCAL:
+ /* XXX: unclear if this ought to be possible? */
+ pgaio_io_reclaim(ioh);
+ break;
+ }
+
+ /*
+ * Need to unregister the reporting of the IO's result, the memory it's
+ * referencing likely has gone away.
+ */
+ if (ioh->report_return)
+ ioh->report_return = NULL;
+}
+
+int
+pgaio_io_get_iovec(PgAioHandle *ioh, struct iovec **iov)
+{
+ Assert(ioh->state == AHS_HANDED_OUT);
+
+ *iov = &aio_ctl->iovecs[ioh->iovec_off];
+
+ /* AFIXME: Needs to be the value at startup time */
+ return io_combine_limit;
+}
+
+PgAioSubjectData *
+pgaio_io_get_subject_data(PgAioHandle *ioh)
+{
+ return &ioh->scb_data;
+}
+
+PgAioOpData *
+pgaio_io_get_op_data(PgAioHandle *ioh)
+{
+ return &ioh->op_data;
+}
+
+ProcNumber
+pgaio_io_get_owner(PgAioHandle *ioh)
+{
+ return ioh->owner_procno;
+}
+
+bool
+pgaio_io_has_subject(PgAioHandle *ioh)
+{
+ return ioh->subject != ASI_INVALID;
+}
+
+void
+pgaio_io_set_flag(PgAioHandle *ioh, PgAioHandleFlags flag)
+{
+ Assert(ioh->state == AHS_HANDED_OUT);
+
+ ioh->flags |= flag;
+}
+
+void
+pgaio_io_set_io_data_32(PgAioHandle *ioh, uint32 *data, uint8 len)
+{
+ Assert(ioh->state == AHS_HANDED_OUT);
+
+ for (int i = 0; i < len; i++)
+ aio_ctl->iovecs_data[ioh->iovec_off + i] = data[i];
+ ioh->iovec_data_len = len;
+}
+
+uint64 *
+pgaio_io_get_io_data(PgAioHandle *ioh, uint8 *len)
+{
+ Assert(ioh->iovec_data_len > 0);
+
+ *len = ioh->iovec_data_len;
+
+ return &aio_ctl->iovecs_data[ioh->iovec_off];
+}
+
+void
+pgaio_io_set_subject(PgAioHandle *ioh, PgAioSubjectID subjid)
+{
+ Assert(ioh->state == AHS_HANDED_OUT);
+
+ ioh->subject = subjid;
+
+ elog(DEBUG3, "io:%d, op %s, subject %s, set subject",
+ pgaio_io_get_id(ioh),
+ pgaio_io_get_op_name(ioh),
+ pgaio_io_get_subject_name(ioh));
+}
+
+void
+pgaio_io_get_ref(PgAioHandle *ioh, PgAioHandleRef *ior)
+{
+ Assert(ioh->state == AHS_HANDED_OUT ||
+ ioh->state == AHS_DEFINED ||
+ ioh->state == AHS_PREPARED);
+ Assert(ioh->generation != 0);
+
+ ior->aio_index = ioh - aio_ctl->io_handles;
+ ior->generation_upper = (uint32) (ioh->generation >> 32);
+ ior->generation_lower = (uint32) ioh->generation;
+}
+
+void
+pgaio_io_ref_clear(PgAioHandleRef *ior)
+{
+ ior->aio_index = PG_UINT32_MAX;
+}
+
+bool
+pgaio_io_ref_valid(PgAioHandleRef *ior)
+{
+ return ior->aio_index != PG_UINT32_MAX;
+}
+
+int
+pgaio_io_ref_get_id(PgAioHandleRef *ior)
+{
+ Assert(pgaio_io_ref_valid(ior));
+ return ior->aio_index;
+}
+
+bool
+pgaio_io_was_recycled(PgAioHandle *ioh, uint64 ref_generation, PgAioHandleState *state)
+{
+ *state = ioh->state;
+ pg_read_barrier();
+
+ return ioh->generation != ref_generation;
+}
+
+void
+pgaio_io_ref_wait(PgAioHandleRef *ior)
+{
+ uint64 ref_generation;
+ PgAioHandleState state;
+ bool am_owner;
+ PgAioHandle *ioh;
+
+ ioh = pgaio_io_from_ref(ior, &ref_generation);
+
+ am_owner = ioh->owner_procno == MyProcNumber;
+
+ if (pgaio_io_was_recycled(ioh, ref_generation, &state))
+ return;
+
+ if (am_owner)
+ {
+ if (state == AHS_DEFINED || state == AHS_PREPARED)
+ {
+ /* XXX: Arguably this should be prevented by callers? */
+ pgaio_submit_staged();
+ }
+ else if (state != AHS_IN_FLIGHT
+ && state != AHS_REAPED
+ && state != AHS_COMPLETED_SHARED
+ && state != AHS_COMPLETED_LOCAL)
+ {
+ elog(PANIC, "waiting for own IO in wrong state: %d",
+ state);
+ }
+
+ /*
+ * Somebody else completed the IO, need to execute issuer callback, so
+ * reclaim eagerly.
+ */
+ if (state == AHS_COMPLETED_LOCAL)
+ {
+ pgaio_io_reclaim(ioh);
+
+ return;
+ }
+ }
+
+ while (true)
+ {
+ if (pgaio_io_was_recycled(ioh, ref_generation, &state))
+ return;
+
+ switch (state)
+ {
+ case AHS_IDLE:
+ case AHS_HANDED_OUT:
+ elog(ERROR, "IO in wrong state: %d", state);
+ break;
+
+ case AHS_IN_FLIGHT:
+ /*
+ * If we need to wait via the IO method, do so now. Don't
+ * check via the IO method if the issuing backend is executing
+ * the IO synchronously.
+ */
+ if (pgaio_impl->wait_one && !(ioh->flags & AHF_SYNCHRONOUS))
+ {
+ pgaio_impl->wait_one(ioh, ref_generation);
+ continue;
+ }
+ /* fallthrough */
+
+ /* waiting for owner to submit */
+ case AHS_PREPARED:
+ case AHS_DEFINED:
+ /* waiting for reaper to complete */
+ /* fallthrough */
+ case AHS_REAPED:
+ /* shouldn't be able to hit this otherwise */
+ Assert(IsUnderPostmaster);
+ /* ensure we're going to get woken up */
+ ConditionVariablePrepareToSleep(&ioh->cv);
+
+ while (!pgaio_io_was_recycled(ioh, ref_generation, &state))
+ {
+ if (state != AHS_REAPED && state != AHS_DEFINED &&
+ state != AHS_IN_FLIGHT)
+ break;
+ ConditionVariableSleep(&ioh->cv, WAIT_EVENT_AIO_COMPLETION);
+ }
+
+ ConditionVariableCancelSleep();
+ break;
+
+ case AHS_COMPLETED_SHARED:
+ /* see above */
+ if (am_owner)
+ pgaio_io_reclaim(ioh);
+ return;
+ case AHS_COMPLETED_LOCAL:
+ return;
+ }
+ }
+}
+
+/*
+ * Check if the the referenced IO completed, without blocking.
+ */
+bool
+pgaio_io_ref_check_done(PgAioHandleRef *ior)
+{
+ uint64 ref_generation;
+ PgAioHandleState state;
+ bool am_owner;
+ PgAioHandle *ioh;
+
+ ioh = pgaio_io_from_ref(ior, &ref_generation);
+
+ if (pgaio_io_was_recycled(ioh, ref_generation, &state))
+ return true;
+
+
+ if (state == AHS_IDLE)
+ return true;
+
+ am_owner = ioh->owner_procno == MyProcNumber;
+
+ if (state == AHS_COMPLETED_SHARED || state == AHS_COMPLETED_LOCAL)
+ {
+ if (am_owner)
+ pgaio_io_reclaim(ioh);
+ return true;
+ }
+
+ return false;
+}
+
+int
+pgaio_io_get_id(PgAioHandle *ioh)
+{
+ Assert(ioh >= aio_ctl->io_handles &&
+ ioh <= (aio_ctl->io_handles + aio_ctl->io_handle_count));
+ return ioh - aio_ctl->io_handles;
+}
+
+const char *
+pgaio_io_get_state_name(PgAioHandle *ioh)
+{
+ switch (ioh->state)
+ {
+ case AHS_IDLE:
+ return "idle";
+ case AHS_HANDED_OUT:
+ return "handed_out";
+ case AHS_DEFINED:
+ return "DEFINED";
+ case AHS_PREPARED:
+ return "PREPARED";
+ case AHS_IN_FLIGHT:
+ return "IN_FLIGHT";
+ case AHS_REAPED:
+ return "REAPED";
+ case AHS_COMPLETED_SHARED:
+ return "COMPLETED_SHARED";
+ case AHS_COMPLETED_LOCAL:
+ return "COMPLETED_LOCAL";
+ }
+ pg_unreachable();
+}
+
+/*
+ * Internal, should only be called from pgaio_io_prep_*().
+ */
+void
+pgaio_io_prepare(PgAioHandle *ioh, PgAioOp op)
+{
+ bool needs_synchronous;
+
+ Assert(ioh->state == AHS_HANDED_OUT);
+ Assert(pgaio_io_has_subject(ioh));
+
+ ioh->op = op;
+ ioh->result = 0;
+
+ pgaio_io_update_state(ioh, AHS_DEFINED);
+
+ /* allow a new IO to be staged */
+ my_aio->handed_out_io = NULL;
+
+ pgaio_io_prepare_subject(ioh);
+
+ pgaio_io_update_state(ioh, AHS_PREPARED);
+
+ needs_synchronous = pgaio_io_needs_synchronous_execution(ioh);
+
+ elog(DEBUG3, "io:%d: prepared %s, executed synchronously: %d",
+ pgaio_io_get_id(ioh), pgaio_io_get_op_name(ioh),
+ needs_synchronous);
+
+ if (!needs_synchronous)
+ {
+ my_aio->staged_ios[my_aio->num_staged_ios++] = ioh;
+ Assert(my_aio->num_staged_ios <= PGAIO_SUBMIT_BATCH_SIZE);
+ }
+ else
+ {
+ pgaio_io_prepare_submit(ioh);
+ pgaio_io_perform_synchronously(ioh);
+ }
+}
+
+/*
+ * Handle IO getting completed by a method.
+ */
+void
+pgaio_io_process_completion(PgAioHandle *ioh, int result)
+{
+ Assert(ioh->state == AHS_IN_FLIGHT);
+
+ ioh->result = result;
+
+ pgaio_io_update_state(ioh, AHS_REAPED);
+
+ pgaio_io_process_completion_subject(ioh);
+
+ pgaio_io_update_state(ioh, AHS_COMPLETED_SHARED);
+
+ /* condition variable broadcast ensures state is visible before wakeup */
+ ConditionVariableBroadcast(&ioh->cv);
+
+ if (ioh->owner_procno == MyProcNumber)
+ pgaio_io_reclaim(ioh);
+}
+
+bool
+pgaio_io_needs_synchronous_execution(PgAioHandle *ioh)
+{
+ if (ioh->flags & AHF_SYNCHRONOUS)
+ {
+ /* XXX: should we also check if there are other IOs staged? */
+ return true;
+ }
+
+ if (pgaio_impl->needs_synchronous_execution)
+ return pgaio_impl->needs_synchronous_execution(ioh);
+ return false;
+}
+
+/*
+ * Handle IO being processed by IO method.
+ */
+void
+pgaio_io_prepare_submit(PgAioHandle *ioh)
+{
+ pgaio_io_update_state(ioh, AHS_IN_FLIGHT);
+
+ dclist_push_tail(&my_aio->in_flight_ios, &ioh->node);
+}
+
+static inline void
+pgaio_io_update_state(PgAioHandle *ioh, PgAioHandleState new_state)
+{
+ /*
+ * Ensure the changes signified by the new state are visible before the
+ * new state becomes visible.
+ */
+ pg_write_barrier();
+
+ ioh->state = new_state;
+}
+
+static PgAioHandle *
+pgaio_io_from_ref(PgAioHandleRef *ior, uint64 *ref_generation)
+{
+ PgAioHandle *ioh;
+
+ Assert(ior->aio_index < aio_ctl->io_handle_count);
+
+ ioh = &aio_ctl->io_handles[ior->aio_index];
+
+ *ref_generation = ((uint64) ior->generation_upper) << 32 |
+ ior->generation_lower;
+
+ Assert(*ref_generation != 0);
+
+ return ioh;
+}
+
+static void
+pgaio_io_resowner_register(PgAioHandle *ioh)
+{
+ Assert(!ioh->resowner);
+ Assert(CurrentResourceOwner);
+
+ ResourceOwnerRememberAioHandle(CurrentResourceOwner, &ioh->resowner_node);
+ ioh->resowner = CurrentResourceOwner;
+}
+
+static void
+pgaio_io_reclaim(PgAioHandle *ioh)
+{
+ /* This is only ok if it's our IO */
+ Assert(ioh->owner_procno == MyProcNumber);
+
+ ereport(DEBUG3,
+ errmsg("reclaiming io:%d, state: %s, op %s, subject %s, result: %d, distilled_result: AFIXME, report to: %p",
+ pgaio_io_get_id(ioh),
+ pgaio_io_get_state_name(ioh),
+ pgaio_io_get_op_name(ioh),
+ pgaio_io_get_subject_name(ioh),
+ ioh->result,
+ ioh->report_return
+ ),
+ errhidestmt(true), errhidecontext(true));
+
+ /* if the IO has been defined, we might need to do more work */
+ if (ioh->state != AHS_HANDED_OUT)
+ {
+ dclist_delete_from(&my_aio->in_flight_ios, &ioh->node);
+
+ if (ioh->report_return)
+ {
+ ioh->report_return->result = ioh->distilled_result;
+ ioh->report_return->subject_data = ioh->scb_data;
+ }
+ }
+
+ if (ioh->resowner)
+ {
+ ResourceOwnerForgetAioHandle(ioh->resowner, &ioh->resowner_node);
+ ioh->resowner = NULL;
+ }
+
+ Assert(!ioh->resowner);
+
+ ioh->num_shared_callbacks = 0;
+ ioh->iovec_data_len = 0;
+ ioh->report_return = NULL;
+ ioh->flags = 0;
+
+ /* XXX: the barrier is probably superfluous */
+ pg_write_barrier();
+ ioh->generation++;
+
+ pgaio_io_update_state(ioh, AHS_IDLE);
+
+ /*
+ * We push the IO to the head of the idle IO list, that seems more cache
+ * efficient in cases where only a few IOs are used.
+ */
+ dclist_push_head(&my_aio->idle_ios, &ioh->node);
+}
+
+static void
+pgaio_io_wait_for_free(void)
+{
+ int reclaimed = 0;
+
+ elog(DEBUG2,
+ "waiting for self: %d pending",
+ my_aio->num_staged_ios);
+
+ /*
+ * First check if any of our IOs actually have completed - when using
+ * worker, that'll often be the case. We could do so as part of the loop
+ * below, but that'd potentially lead us to wait for some IO submitted
+ * before.
+ */
+ for (int i = 0; i < io_max_concurrency; i++)
+ {
+ PgAioHandle *ioh = &aio_ctl->io_handles[my_aio->io_handle_off + i];
+
+ if (ioh->state == AHS_COMPLETED_SHARED)
+ {
+ pgaio_io_reclaim(ioh);
+ reclaimed++;
+ }
+ }
+
+ if (reclaimed > 0)
+ return;
+
+ /*
+ * If we have any unsubmitted IOs, submit them now. We'll start waiting in
+ * a second, so it's better they're in flight. This also addresses the
+ * edge-case that all IOs are unsubmitted.
+ */
+ if (my_aio->num_staged_ios > 0)
+ {
+ elog(DEBUG2, "submitting while acquiring free io");
+ pgaio_submit_staged();
+ }
+
+ /*
+ * It's possible that we recognized there were free IOs while submitting.
+ */
+ if (dclist_count(&my_aio->in_flight_ios) == 0)
+ {
+ elog(ERROR, "no free IOs despite no in-flight IOs");
+ }
+
+ /*
+ * Wait for the oldest in-flight IO to complete.
+ *
+ * XXX: Reusing the general IO wait is suboptimal, we don't need to wait
+ * for that specific IO to complete, we just need *any* IO to complete.
+ */
+ {
+ PgAioHandle *ioh = dclist_head_element(PgAioHandle, node, &my_aio->in_flight_ios);
+
+ switch (ioh->state)
+ {
+ /* should not be in in-flight list */
+ case AHS_IDLE:
+ case AHS_DEFINED:
+ case AHS_HANDED_OUT:
+ case AHS_PREPARED:
+ case AHS_COMPLETED_LOCAL:
+ elog(ERROR, "shouldn't get here with io:%d in state %d",
+ pgaio_io_get_id(ioh), ioh->state);
+ break;
+
+ case AHS_REAPED:
+ case AHS_IN_FLIGHT:
+ {
+ PgAioHandleRef ior;
+
+ ior.aio_index = ioh - aio_ctl->io_handles;
+ ior.generation_upper = (uint32) (ioh->generation >> 32);
+ ior.generation_lower = (uint32) ioh->generation;
+
+ pgaio_io_ref_wait(&ior);
+ elog(DEBUG2, "waited for io:%d",
+ pgaio_io_get_id(ioh));
+ }
+ break;
+ case AHS_COMPLETED_SHARED:
+ /* it's possible that another backend just finished this IO */
+ pgaio_io_reclaim(ioh);
+ break;
+ }
+
+ if (dclist_count(&my_aio->idle_ios) == 0)
+ elog(PANIC, "no idle IOs after waiting");
+ return;
+ }
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * Actions on multiple IOs.
+ * --------------------------------------------------------------------------------
+ */
+
+void
+pgaio_submit_staged(void)
+{
+ int total_submitted = 0;
+ int did_submit;
+
+ if (my_aio->num_staged_ios == 0)
+ return;
+
+
+ START_CRIT_SECTION();
+
+ did_submit = pgaio_impl->submit(my_aio->num_staged_ios, my_aio->staged_ios);
+
+ END_CRIT_SECTION();
+
+ total_submitted += did_submit;
+
+ Assert(total_submitted == did_submit);
+
+ my_aio->num_staged_ios = 0;
+
+#ifdef PGAIO_VERBOSE
+ ereport(DEBUG2,
+ errmsg("submitted %d",
+ total_submitted),
+ errhidestmt(true),
+ errhidecontext(true));
+#endif
+}
+
+bool
+pgaio_have_staged(void)
+{
+ return my_aio->num_staged_ios > 0;
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * Other
+ * --------------------------------------------------------------------------------
+ */
+
+/*
+ * Need to submit staged but not yet submitted IOs using the fd, otherwise
+ * the IO would end up targeting something bogus.
+ */
+void
+pgaio_closing_fd(int fd)
+{
+ /*
+ * Might be called before AIO is initialized or in a subprocess that
+ * doesn't use AIO.
+ */
+ if (!my_aio)
+ return;
+
+ /*
+ * For now just submit all staged IOs - we could be more selective, but
+ * it's probably not worth it.
+ */
+ pgaio_submit_staged();
+}
+
+void
+pgaio_at_xact_end(bool is_subxact, bool is_commit)
+{
+ Assert(!my_aio->handed_out_io);
+}
+
+/*
+ * Similar to pgaio_at_xact_end(..., is_commit = false), but for cases where
+ * errors happen outside of transactions.
+ */
+void
+pgaio_at_error(void)
+{
+ Assert(!my_aio->handed_out_io);
+}
+
+
void
assign_io_method(int newval, void *extra)
{
+ pgaio_impl = pgaio_ops_table[newval];
}
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index 84e0e37baae..b9bdf51680a 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -14,28 +14,206 @@
#include "postgres.h"
+#include "miscadmin.h"
+#include "storage/aio.h"
#include "storage/aio_init.h"
+#include "storage/aio_internal.h"
+#include "storage/bufmgr.h"
+#include "storage/proc.h"
+#include "storage/shmem.h"
+static Size
+AioCtlShmemSize(void)
+{
+ Size sz;
+
+ /* aio_ctl itself */
+ sz = offsetof(PgAioCtl, io_handles);
+
+ return sz;
+}
+
+static uint32
+AioProcs(void)
+{
+ return MaxBackends + NUM_AUXILIARY_PROCS;
+}
+
+static Size
+AioBackendShmemSize(void)
+{
+ return mul_size(AioProcs(), sizeof(PgAioPerBackend));
+}
+
+static Size
+AioHandleShmemSize(void)
+{
+ Size sz;
+
+ /* ios */
+ sz = mul_size(AioProcs(),
+ mul_size(io_max_concurrency, sizeof(PgAioHandle)));
+
+ return sz;
+}
+
+static Size
+AioIOVShmemSize(void)
+{
+ /* FIXME: io_combine_limit is USERSET */
+ return mul_size(sizeof(struct iovec),
+ mul_size(mul_size(io_combine_limit, AioProcs()),
+ io_max_concurrency));
+}
+
+static Size
+AioIOVDataShmemSize(void)
+{
+ /* FIXME: io_combine_limit is USERSET */
+ return mul_size(sizeof(uint64),
+ mul_size(mul_size(io_combine_limit, AioProcs()),
+ io_max_concurrency));
+}
+
+/*
+ * Choose a suitable value for io_max_concurrency.
+ *
+ * It's unlikely that we could have more IOs in flight than buffers that we
+ * would be allowed to pin.
+ *
+ * On the upper end, apply a cap too - just because shared_buffers is large,
+ * it doesn't make sense have millions of buffers undergo IO concurrently.
+ */
+static int
+AioChooseMaxConccurrency(void)
+{
+ uint32 max_backends;
+ int max_proportional_pins;
+
+ /* Similar logic to LimitAdditionalPins() */
+ max_backends = MaxBackends + NUM_AUXILIARY_PROCS;
+ max_proportional_pins = NBuffers / max_backends;
+
+ max_proportional_pins = Max(max_proportional_pins, 1);
+
+ /* apply upper limit */
+ return Min(max_proportional_pins, 64);
+}
+
Size
AioShmemSize(void)
{
Size sz = 0;
+ /*
+ * We prefer to report this value's source as PGC_S_DYNAMIC_DEFAULT.
+ * However, if the DBA explicitly set wal_buffers = -1 in the config file,
+ * then PGC_S_DYNAMIC_DEFAULT will fail to override that and we must force
+ *
+ */
+ if (io_max_concurrency == -1)
+ {
+ char buf[32];
+
+ snprintf(buf, sizeof(buf), "%d", AioChooseMaxConccurrency());
+ SetConfigOption("io_max_concurrency", buf, PGC_POSTMASTER,
+ PGC_S_DYNAMIC_DEFAULT);
+ if (io_max_concurrency == -1) /* failed to apply it? */
+ SetConfigOption("io_max_concurrency", buf, PGC_POSTMASTER,
+ PGC_S_OVERRIDE);
+ }
+
+ sz = add_size(sz, AioCtlShmemSize());
+ sz = add_size(sz, AioBackendShmemSize());
+ sz = add_size(sz, AioHandleShmemSize());
+ sz = add_size(sz, AioIOVShmemSize());
+ sz = add_size(sz, AioIOVDataShmemSize());
+
+ if (pgaio_impl->shmem_size)
+ sz = add_size(sz, pgaio_impl->shmem_size());
+
return sz;
}
void
AioShmemInit(void)
{
+ bool found;
+ uint32 io_handle_off = 0;
+ uint32 iovec_off = 0;
+ uint32 per_backend_iovecs = io_max_concurrency * io_combine_limit;
+
+ aio_ctl = (PgAioCtl *)
+ ShmemInitStruct("AioCtl", AioCtlShmemSize(), &found);
+
+ if (found)
+ goto out;
+
+ memset(aio_ctl, 0, AioCtlShmemSize());
+
+ aio_ctl->io_handle_count = AioProcs() * io_max_concurrency;
+ aio_ctl->iovec_count = AioProcs() * per_backend_iovecs;
+
+ aio_ctl->backend_state = (PgAioPerBackend *)
+ ShmemInitStruct("AioBackend", AioBackendShmemSize(), &found);
+
+ aio_ctl->io_handles = (PgAioHandle *)
+ ShmemInitStruct("AioHandle", AioHandleShmemSize(), &found);
+
+ aio_ctl->iovecs = ShmemInitStruct("AioIOV", AioIOVShmemSize(), &found);
+ aio_ctl->iovecs_data = ShmemInitStruct("AioIOVData", AioIOVDataShmemSize(), &found);
+
+ for (int procno = 0; procno < AioProcs(); procno++)
+ {
+ PgAioPerBackend *bs = &aio_ctl->backend_state[procno];
+
+ bs->io_handle_off = io_handle_off;
+ io_handle_off += io_max_concurrency;
+
+ dclist_init(&bs->idle_ios);
+ memset(bs->staged_ios, 0, sizeof(PgAioHandle *) * PGAIO_SUBMIT_BATCH_SIZE);
+ dclist_init(&bs->in_flight_ios);
+
+ /* initialize per-backend IOs */
+ for (int i = 0; i < io_max_concurrency; i++)
+ {
+ PgAioHandle *ioh = &aio_ctl->io_handles[bs->io_handle_off + i];
+
+ ioh->generation = 1;
+ ioh->owner_procno = procno;
+ ioh->iovec_off = iovec_off;
+ ioh->iovec_data_len = 0;
+ ioh->report_return = NULL;
+ ioh->resowner = NULL;
+ ioh->num_shared_callbacks = 0;
+ ioh->distilled_result.status = ARS_UNKNOWN;
+ ioh->flags = 0;
+
+ ConditionVariableInit(&ioh->cv);
+
+ dclist_push_tail(&bs->idle_ios, &ioh->node);
+ iovec_off += io_combine_limit;
+ }
+ }
+
+out:
+ /* Initialize IO method specific resources. */
+ if (pgaio_impl->shmem_init)
+ pgaio_impl->shmem_init(!found);
}
void
pgaio_init_backend(void)
{
-}
+ /* shouldn't be initialized twice */
+ Assert(!my_aio);
+
+ if (MyProc == NULL || MyProcNumber >= AioProcs())
+ elog(ERROR, "aio requires a normal PGPROC");
+
+ my_aio = &aio_ctl->backend_state[MyProcNumber];
-void
-pgaio_postmaster_child_init_local(void)
-{
+ if (pgaio_impl->init_backend)
+ pgaio_impl->init_backend();
}
diff --git a/src/backend/storage/aio/aio_io.c b/src/backend/storage/aio/aio_io.c
new file mode 100644
index 00000000000..3c255775833
--- /dev/null
+++ b/src/backend/storage/aio/aio_io.c
@@ -0,0 +1,140 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_io.c
+ * AIO - Low Level IO Handling
+ *
+ * Functions related to associating IO operations to IO Handles and IO-method
+ * independent support functions for actually performing IO.
+ *
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/aio_io.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
+#include "storage/fd.h"
+#include "utils/wait_event.h"
+
+
+static void pgaio_io_before_prep(PgAioHandle *ioh);
+
+
+
+/* --------------------------------------------------------------------------------
+ * "Preparation" routines for individual IO types
+ *
+ * These are called by place the place actually initiating an IO, to associate
+ * the IO specific data with an AIO handle.
+ *
+ * Each of the preparation routines first needs to call
+ * pgaio_io_before_prep(), then fill IO specific fields in the handle and then
+ * finally call pgaio_io_prepare().
+ * --------------------------------------------------------------------------------
+ */
+
+void
+pgaio_io_prep_readv(PgAioHandle *ioh,
+ int fd, int iovcnt, uint64 offset)
+{
+ pgaio_io_before_prep(ioh);
+
+ ioh->op_data.read.fd = fd;
+ ioh->op_data.read.offset = offset;
+ ioh->op_data.read.iov_length = iovcnt;
+
+ pgaio_io_prepare(ioh, PGAIO_OP_READV);
+}
+
+void
+pgaio_io_prep_writev(PgAioHandle *ioh,
+ int fd, int iovcnt, uint64 offset)
+{
+ pgaio_io_before_prep(ioh);
+
+ ioh->op_data.write.fd = fd;
+ ioh->op_data.write.offset = offset;
+ ioh->op_data.write.iov_length = iovcnt;
+
+ pgaio_io_prepare(ioh, PGAIO_OP_WRITEV);
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * Functions implementing IO handle operations that are directly related to IO
+ * operations.
+ * --------------------------------------------------------------------------------
+ */
+
+/*
+ * Execute IO operation synchronously. This is implemented here, not in
+ * method_sync.c, because other IO methods lso might use it / fall back to it.
+ */
+void
+pgaio_io_perform_synchronously(PgAioHandle *ioh)
+{
+ ssize_t result = 0;
+ struct iovec *iov = &aio_ctl->iovecs[ioh->iovec_off];
+
+ /* Perform IO. */
+ switch (ioh->op)
+ {
+ case PGAIO_OP_READV:
+ pgstat_report_wait_start(WAIT_EVENT_DATA_FILE_READ);
+ result = pg_preadv(ioh->op_data.read.fd, iov,
+ ioh->op_data.read.iov_length,
+ ioh->op_data.read.offset);
+ pgstat_report_wait_end();
+ break;
+ case PGAIO_OP_WRITEV:
+ pgstat_report_wait_start(WAIT_EVENT_DATA_FILE_WRITE);
+ result = pg_pwritev(ioh->op_data.write.fd, iov,
+ ioh->op_data.write.iov_length,
+ ioh->op_data.write.offset);
+ pgstat_report_wait_end();
+ break;
+ case PGAIO_OP_INVALID:
+ elog(ERROR, "trying to execute invalid IO operation");
+ }
+
+ ioh->result = result < 0 ? -errno : result;
+
+ pgaio_io_process_completion(ioh, ioh->result);
+}
+
+const char *
+pgaio_io_get_op_name(PgAioHandle *ioh)
+{
+ Assert(ioh->op >= 0 && ioh->op < PGAIO_OP_COUNT);
+
+ switch (ioh->op)
+ {
+ case PGAIO_OP_INVALID:
+ return "invalid";
+ case PGAIO_OP_READV:
+ return "read";
+ case PGAIO_OP_WRITEV:
+ return "write";
+ }
+
+ pg_unreachable();
+}
+
+/*
+ * Helper function to be called by IO operation preparation functions, before
+ * any data in the handle is set. Mostly to centralize assertions.
+ */
+static void
+pgaio_io_before_prep(PgAioHandle *ioh)
+{
+ Assert(ioh->state == AHS_HANDED_OUT);
+ Assert(pgaio_io_has_subject(ioh));
+}
diff --git a/src/backend/storage/aio/aio_subject.c b/src/backend/storage/aio/aio_subject.c
new file mode 100644
index 00000000000..8694cfafcd1
--- /dev/null
+++ b/src/backend/storage/aio/aio_subject.c
@@ -0,0 +1,231 @@
+/*-------------------------------------------------------------------------
+ *
+ * aio_subject.c
+ * AIO - Functionality related to executing IO for different subjects
+ *
+ * XXX Write me
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/aio_subject.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
+#include "storage/buf_internals.h"
+#include "storage/bufmgr.h"
+#include "storage/smgr.h"
+#include "utils/memutils.h"
+
+
+/*
+ * Registry for entities that can be the target of AIO.
+ *
+ * To support executing using worker processes, the file descriptor for an IO
+ * may need to be be reopened in a different process. This is done via the
+ * PgAioSubjectInfo.reopen callback.
+ */
+static const PgAioSubjectInfo *aio_subject_info[] = {
+ [ASI_INVALID] = &(PgAioSubjectInfo) {
+ .name = "invalid",
+ },
+};
+
+
+typedef struct PgAioHandleSharedCallbacksEntry
+{
+ const PgAioHandleSharedCallbacks *const cb;
+ const char *const name;
+} PgAioHandleSharedCallbacksEntry;
+
+static const PgAioHandleSharedCallbacksEntry aio_shared_cbs[] = {
+#define CALLBACK_ENTRY(id, callback) [id] = {.cb = &callback, .name = #callback}
+#undef CALLBACK_ENTRY
+};
+
+
+/*
+ * Register callback for the IO handle.
+ *
+ * Only a limited number (AIO_MAX_SHARED_CALLBACKS) of callbacks can be
+ * registered for each IO.
+ *
+ * Callbacks need to be registered before [indirectly] calling
+ * pgaio_io_prep_*(), as the IO may be executed immediately.
+ *
+ *
+ * Note that callbacks are executed in critical sections. This is necessary
+ * to be able to execute IO in critical sections (consider e.g. WAL
+ * logging). To perform AIO we first need to acquire a handle, which, if there
+ * are no free handles, requires waiting for IOs to complete and to execute
+ * their completion callbacks.
+ *
+ * Callbacks may be executed in the issuing backend but also in another
+ * backend (because that backend is waiting for the IO) or in IO workers (if
+ * io_method=worker is used).
+ *
+ *
+ * See PgAioHandleSharedCallbackID's definition for an explanation for why
+ * callbacks are not identified by a pointer.
+ */
+void
+pgaio_io_add_shared_cb(PgAioHandle *ioh, PgAioHandleSharedCallbackID cbid)
+{
+ const PgAioHandleSharedCallbacksEntry *ce = &aio_shared_cbs[cbid];
+
+ if (cbid >= lengthof(aio_shared_cbs))
+ elog(ERROR, "callback %d is out of range", cbid);
+ if (aio_shared_cbs[cbid].cb->complete == NULL)
+ elog(ERROR, "callback %d is undefined", cbid);
+ if (ioh->num_shared_callbacks >= AIO_MAX_SHARED_CALLBACKS)
+ elog(PANIC, "too many callbacks, the max is %d", AIO_MAX_SHARED_CALLBACKS);
+ ioh->shared_callbacks[ioh->num_shared_callbacks] = cbid;
+
+ elog(DEBUG3, "io:%d, op %s, subject %s, adding cb #%d, id %d/%s",
+ pgaio_io_get_id(ioh),
+ pgaio_io_get_op_name(ioh),
+ pgaio_io_get_subject_name(ioh),
+ ioh->num_shared_callbacks + 1,
+ cbid, ce->name);
+
+ ioh->num_shared_callbacks++;
+}
+
+/*
+ * Return the name for the subject associated with the IO. Mostly useful for
+ * debugging/logging.
+ */
+const char *
+pgaio_io_get_subject_name(PgAioHandle *ioh)
+{
+ Assert(ioh->subject >= 0 && ioh->subject < ASI_COUNT);
+
+ return aio_subject_info[ioh->subject]->name;
+}
+
+/*
+ * Internal function which invokes ->prepare for all the registered callbacks.
+ */
+void
+pgaio_io_prepare_subject(PgAioHandle *ioh)
+{
+ Assert(ioh->subject > ASI_INVALID && ioh->subject < ASI_COUNT);
+ Assert(ioh->op >= 0 && ioh->op < PGAIO_OP_COUNT);
+
+ for (int i = ioh->num_shared_callbacks; i > 0; i--)
+ {
+ PgAioHandleSharedCallbackID cbid = ioh->shared_callbacks[i - 1];
+ const PgAioHandleSharedCallbacksEntry *ce = &aio_shared_cbs[cbid];
+
+ if (!ce->cb->prepare)
+ continue;
+
+ elog(DEBUG3, "io:%d, op %s, subject %s, calling cb #%d %d/%s->prepare",
+ pgaio_io_get_id(ioh),
+ pgaio_io_get_op_name(ioh),
+ pgaio_io_get_subject_name(ioh),
+ i,
+ cbid, ce->name);
+ ce->cb->prepare(ioh);
+ }
+}
+
+/*
+ * Internal function which invokes ->complete for all the registered
+ * callbacks.
+ */
+void
+pgaio_io_process_completion_subject(PgAioHandle *ioh)
+{
+ PgAioResult result;
+
+ Assert(ioh->subject >= 0 && ioh->subject < ASI_COUNT);
+ Assert(ioh->op >= 0 && ioh->op < PGAIO_OP_COUNT);
+
+ result.status = ARS_OK; /* low level IO is always considered OK */
+ result.result = ioh->result;
+ result.id = ASC_INVALID;
+ result.error_data = 0;
+
+ for (int i = ioh->num_shared_callbacks; i > 0; i--)
+ {
+ PgAioHandleSharedCallbackID cbid = ioh->shared_callbacks[i - 1];
+ const PgAioHandleSharedCallbacksEntry *ce = &aio_shared_cbs[cbid];
+
+ elog(DEBUG3, "io:%d, op %s, subject %s, calling cb #%d, id %d/%s->complete with distilled result status %d, id %u, error_data: %d, result: %d",
+ pgaio_io_get_id(ioh),
+ pgaio_io_get_op_name(ioh),
+ pgaio_io_get_subject_name(ioh),
+ i,
+ cbid, ce->name,
+ result.status,
+ result.id,
+ result.error_data,
+ result.result);
+ result = ce->cb->complete(ioh, result);
+ }
+
+ ioh->distilled_result = result;
+
+ elog(DEBUG3, "io:%d, op %s, subject %s, distilled result status %d, id %u, error_data: %d, result: %d, raw_result %d",
+ pgaio_io_get_id(ioh),
+ pgaio_io_get_op_name(ioh),
+ pgaio_io_get_subject_name(ioh),
+ result.status,
+ result.id,
+ result.error_data,
+ result.result,
+ ioh->result);
+}
+
+/*
+ * Check if pgaio_io_reopen() is available for the IO.
+ */
+bool
+pgaio_io_can_reopen(PgAioHandle *ioh)
+{
+ return aio_subject_info[ioh->subject]->reopen != NULL;
+}
+
+/*
+ * Before executing an IO outside of the context of the process the IO has
+ * been prepared in, the file descriptor has to be reopened - any FD
+ * referenced in the IO itself, won't be valid in the separate process.
+ */
+void
+pgaio_io_reopen(PgAioHandle *ioh)
+{
+ Assert(ioh->subject >= 0 && ioh->subject < ASI_COUNT);
+ Assert(ioh->op >= 0 && ioh->op < PGAIO_OP_COUNT);
+
+ aio_subject_info[ioh->subject]->reopen(ioh);
+}
+
+
+
+/* --------------------------------------------------------------------------------
+ * IO Result
+ * --------------------------------------------------------------------------------
+ */
+
+void
+pgaio_result_log(PgAioResult result, const PgAioSubjectData *subject_data, int elevel)
+{
+ PgAioHandleSharedCallbackID cbid = result.id;
+ const PgAioHandleSharedCallbacksEntry *ce = &aio_shared_cbs[cbid];
+
+ Assert(result.status != ARS_UNKNOWN);
+ Assert(result.status != ARS_OK);
+
+ if (ce->cb->error == NULL)
+ elog(ERROR, "scb id %d/%s does not have an error callback",
+ result.id, ce->name);
+
+ ce->cb->error(result, subject_data, elevel);
+}
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index 8d20759ebf8..8339d473aae 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -3,5 +3,8 @@
backend_sources += files(
'aio.c',
'aio_init.c',
+ 'aio_io.c',
+ 'aio_subject.c',
+ 'method_sync.c',
'read_stream.c',
)
diff --git a/src/backend/storage/aio/method_sync.c b/src/backend/storage/aio/method_sync.c
new file mode 100644
index 00000000000..61fd06a277b
--- /dev/null
+++ b/src/backend/storage/aio/method_sync.c
@@ -0,0 +1,45 @@
+/*-------------------------------------------------------------------------
+ *
+ * method_sync.c
+ * AIO - perform "AIO" by executing it synchronously
+ *
+ * This method is mainly to check if AIO use causes regressions. Other IO
+ * methods might also fall back to the synchronous method for functionality
+ * they cannot provide.
+ *
+ * Portions Copyright (c) 1996-2021, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/method_sync.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
+
+static bool pgaio_sync_needs_synchronous_execution(PgAioHandle *ioh);
+static int pgaio_sync_submit(uint16 num_staged_ios, PgAioHandle **staged_ios);
+
+
+const IoMethodOps pgaio_sync_ops = {
+ .needs_synchronous_execution = pgaio_sync_needs_synchronous_execution,
+ .submit = pgaio_sync_submit,
+};
+
+static bool
+pgaio_sync_needs_synchronous_execution(PgAioHandle *ioh)
+{
+ return true;
+}
+
+static int
+pgaio_sync_submit(uint16 num_staged_ios, PgAioHandle **staged_ios)
+{
+ elog(ERROR, "should be unreachable");
+
+ return 0;
+}
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 16144c2b72d..7a2e2b4432e 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -190,6 +190,9 @@ ABI_compatibility:
Section: ClassName - WaitEventIO
+AIO_SUBMIT "Waiting for AIO submission."
+AIO_DRAIN "Waiting for IOs to finish."
+AIO_COMPLETION "Waiting for completion callback."
BASEBACKUP_READ "Waiting for base backup to read from a file."
BASEBACKUP_SYNC "Waiting for data written by a base backup to reach durable storage."
BASEBACKUP_WRITE "Waiting for base backup to write to a file."
diff --git a/src/backend/utils/resowner/resowner.c b/src/backend/utils/resowner/resowner.c
index 505534ee8d3..5cf14472ebd 100644
--- a/src/backend/utils/resowner/resowner.c
+++ b/src/backend/utils/resowner/resowner.c
@@ -47,6 +47,8 @@
#include "common/hashfn.h"
#include "common/int.h"
+#include "lib/ilist.h"
+#include "storage/aio.h"
#include "storage/ipc.h"
#include "storage/predicate.h"
#include "storage/proc.h"
@@ -155,6 +157,12 @@ struct ResourceOwnerData
/* The local locks cache. */
LOCALLOCK *locks[MAX_RESOWNER_LOCKS]; /* list of owned locks */
+
+ /*
+ * AIO handles need be registered in critical sections and therefore
+ * cannot use the normal ResoureElem mechanism.
+ */
+ dlist_head aio_handles;
};
@@ -425,6 +433,8 @@ ResourceOwnerCreate(ResourceOwner parent, const char *name)
parent->firstchild = owner;
}
+ dlist_init(&owner->aio_handles);
+
return owner;
}
@@ -725,6 +735,14 @@ ResourceOwnerReleaseInternal(ResourceOwner owner,
* so issue warnings. In the abort case, just clean up quietly.
*/
ResourceOwnerReleaseAll(owner, phase, isCommit);
+
+ /* XXX: Could probably be a later phase? */
+ while (!dlist_is_empty(&owner->aio_handles))
+ {
+ dlist_node *node = dlist_head_node(&owner->aio_handles);
+
+ pgaio_io_release_resowner(node, !isCommit);
+ }
}
else if (phase == RESOURCE_RELEASE_LOCKS)
{
@@ -1082,3 +1100,15 @@ ResourceOwnerForgetLock(ResourceOwner owner, LOCALLOCK *locallock)
elog(ERROR, "lock reference %p is not owned by resource owner %s",
locallock, owner->name);
}
+
+void
+ResourceOwnerRememberAioHandle(ResourceOwner owner, struct dlist_node *ioh_node)
+{
+ dlist_push_tail(&owner->aio_handles, ioh_node);
+}
+
+void
+ResourceOwnerForgetAioHandle(ResourceOwner owner, struct dlist_node *ioh_node)
+{
+ dlist_delete_from(&owner->aio_handles, ioh_node);
+}
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 2586d1cf53f..bc1acbb98ee 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1263,6 +1263,7 @@ InvalMessageArray
InvalidationInfo
InvalidationMsgsGroup
IoMethod
+IoMethodOps
IpcMemoryId
IpcMemoryKey
IpcMemoryState
@@ -2100,6 +2101,23 @@ Permutation
PermutationStep
PermutationStepBlocker
PermutationStepBlockerType
+PgAioCtl
+PgAioHandle
+PgAioHandleFlags
+PgAioHandleRef
+PgAioHandleSharedCallbackID
+PgAioHandleSharedCallbacks
+PgAioHandleSharedCallbacksEntry
+PgAioHandleState
+PgAioOp
+PgAioOpData
+PgAioPerBackend
+PgAioResultStatus
+PgAioResult
+PgAioReturn
+PgAioSubjectData
+PgAioSubjectID
+PgAioSubjectInfo
PgArchData
PgBackendGSSStatus
PgBackendSSLStatus
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0005-aio-Skeleton-IO-worker-infrastructure.patch (22.2K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/6-v2-0005-aio-Skeleton-IO-worker-infrastructure.patch)
download | inline diff:
From e6c7783183c0b36f94b9debfd9edde71e4d75bbc Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 25 Nov 2024 14:03:40 -0500
Subject: [PATCH v2 05/20] aio: Skeleton IO worker infrastructure
This doesn't do anything useful on its own, but the code that needs to be
touched is independent of other changes.
Remarks:
- should completely get rid of ID assignment logic in postmaster.c
- postmaster.c badly needs a refactoring.
- dynamic increase / decrease of workers based on IO load
Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
src/include/miscadmin.h | 2 +
src/include/postmaster/postmaster.h | 1 +
src/include/storage/aio_init.h | 2 +
src/include/storage/io_worker.h | 22 +++
src/include/storage/proc.h | 4 +-
src/backend/postmaster/launch_backend.c | 2 +
src/backend/postmaster/pmchild.c | 1 +
src/backend/postmaster/postmaster.c | 171 ++++++++++++++++--
src/backend/storage/aio/Makefile | 1 +
src/backend/storage/aio/aio_init.c | 7 +
src/backend/storage/aio/meson.build | 1 +
src/backend/storage/aio/method_worker.c | 86 +++++++++
src/backend/tcop/postgres.c | 2 +
src/backend/utils/activity/pgstat_backend.c | 1 +
src/backend/utils/activity/pgstat_io.c | 1 +
.../utils/activity/wait_event_names.txt | 1 +
src/backend/utils/init/miscinit.c | 3 +
src/backend/utils/misc/guc_tables.c | 13 ++
src/backend/utils/misc/postgresql.conf.sample | 1 +
19 files changed, 310 insertions(+), 12 deletions(-)
create mode 100644 src/include/storage/io_worker.h
create mode 100644 src/backend/storage/aio/method_worker.c
diff --git a/src/include/miscadmin.h b/src/include/miscadmin.h
index e4c0d1481e9..0afc57ebf27 100644
--- a/src/include/miscadmin.h
+++ b/src/include/miscadmin.h
@@ -360,6 +360,7 @@ typedef enum BackendType
B_ARCHIVER,
B_BG_WRITER,
B_CHECKPOINTER,
+ B_IO_WORKER,
B_STARTUP,
B_WAL_RECEIVER,
B_WAL_SUMMARIZER,
@@ -389,6 +390,7 @@ extern PGDLLIMPORT BackendType MyBackendType;
#define AmWalReceiverProcess() (MyBackendType == B_WAL_RECEIVER)
#define AmWalSummarizerProcess() (MyBackendType == B_WAL_SUMMARIZER)
#define AmWalWriterProcess() (MyBackendType == B_WAL_WRITER)
+#define AmIoWorkerProcess() (MyBackendType == B_IO_WORKER)
#define AmSpecialWorkerProcess() \
(AmAutoVacuumLauncherProcess() || \
diff --git a/src/include/postmaster/postmaster.h b/src/include/postmaster/postmaster.h
index 24d49a5439e..4d003b7f86d 100644
--- a/src/include/postmaster/postmaster.h
+++ b/src/include/postmaster/postmaster.h
@@ -98,6 +98,7 @@ extern void InitProcessGlobals(void);
extern int MaxLivePostmasterChildren(void);
extern bool PostmasterMarkPIDForWorkerNotify(int);
+extern void assign_io_workers(int newval, void *extra);
#ifdef WIN32
extern void pgwin32_register_deadchild_callback(HANDLE procHandle, DWORD procId);
diff --git a/src/include/storage/aio_init.h b/src/include/storage/aio_init.h
index 1c1d62baa79..70976791c93 100644
--- a/src/include/storage/aio_init.h
+++ b/src/include/storage/aio_init.h
@@ -21,4 +21,6 @@ extern void AioShmemInit(void);
extern void pgaio_init_backend(void);
+extern bool pgaio_workers_enabled(void);
+
#endif /* AIO_INIT_H */
diff --git a/src/include/storage/io_worker.h b/src/include/storage/io_worker.h
new file mode 100644
index 00000000000..ba5dcb9e6e4
--- /dev/null
+++ b/src/include/storage/io_worker.h
@@ -0,0 +1,22 @@
+/*-------------------------------------------------------------------------
+ *
+ * io_worker.h
+ * IO worker for implementing AIO "ourselves"
+ *
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/io.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef IO_WORKER_H
+#define IO_WORKER_H
+
+
+extern void IoWorkerMain(char *startup_data, size_t startup_data_len) pg_attribute_noreturn();
+
+extern int io_workers;
+
+#endif /* IO_WORKER_H */
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 0b1fa61310f..cafd0b334b9 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -461,7 +461,9 @@ extern PGDLLIMPORT PGPROC *PreparedXactProcs;
* 2 slots, but WAL writer is launched only after startup has exited, so we
* only need 6 slots.
*/
-#define NUM_AUXILIARY_PROCS 6
+#define MAX_IO_WORKERS 32
+#define NUM_AUXILIARY_PROCS (6 + MAX_IO_WORKERS)
+
/* configurable options */
extern PGDLLIMPORT int DeadlockTimeout;
diff --git a/src/backend/postmaster/launch_backend.c b/src/backend/postmaster/launch_backend.c
index 1f2d829ec5a..7399adfeae9 100644
--- a/src/backend/postmaster/launch_backend.c
+++ b/src/backend/postmaster/launch_backend.c
@@ -48,6 +48,7 @@
#include "replication/slotsync.h"
#include "replication/walreceiver.h"
#include "storage/dsm.h"
+#include "storage/io_worker.h"
#include "storage/pg_shmem.h"
#include "tcop/backend_startup.h"
#include "utils/memutils.h"
@@ -197,6 +198,7 @@ static child_process_kind child_process_kinds[] = {
[B_ARCHIVER] = {"archiver", PgArchiverMain, true},
[B_BG_WRITER] = {"bgwriter", BackgroundWriterMain, true},
[B_CHECKPOINTER] = {"checkpointer", CheckpointerMain, true},
+ [B_IO_WORKER] = {"io_worker", IoWorkerMain, true},
[B_STARTUP] = {"startup", StartupProcessMain, true},
[B_WAL_RECEIVER] = {"wal_receiver", WalReceiverMain, true},
[B_WAL_SUMMARIZER] = {"wal_summarizer", WalSummarizerMain, true},
diff --git a/src/backend/postmaster/pmchild.c b/src/backend/postmaster/pmchild.c
index 381cf005a9b..89ee626829d 100644
--- a/src/backend/postmaster/pmchild.c
+++ b/src/backend/postmaster/pmchild.c
@@ -101,6 +101,7 @@ InitPostmasterChildSlots(void)
pmchild_pools[B_AUTOVAC_WORKER].size = autovacuum_max_workers;
pmchild_pools[B_BG_WORKER].size = max_worker_processes;
+ pmchild_pools[B_IO_WORKER].size = MAX_IO_WORKERS;
/*
* There can be only one of each of these running at a time. They each
diff --git a/src/backend/postmaster/postmaster.c b/src/backend/postmaster/postmaster.c
index 6f849ffbcb5..8dab7072114 100644
--- a/src/backend/postmaster/postmaster.c
+++ b/src/backend/postmaster/postmaster.c
@@ -108,9 +108,12 @@
#include "replication/logicallauncher.h"
#include "replication/slotsync.h"
#include "replication/walsender.h"
+#include "storage/aio_init.h"
#include "storage/fd.h"
+#include "storage/io_worker.h"
#include "storage/ipc.h"
#include "storage/pmsignal.h"
+#include "storage/proc.h"
#include "tcop/backend_startup.h"
#include "tcop/tcopprot.h"
#include "utils/datetime.h"
@@ -172,6 +175,7 @@ btmask_all_except(BackendType t)
return mask;
}
+#ifdef NOT_USED
static inline BackendTypeMask
btmask_all_except2(BackendType t1, BackendType t2)
{
@@ -181,6 +185,18 @@ btmask_all_except2(BackendType t1, BackendType t2)
mask = btmask_del(mask, t2);
return mask;
}
+#endif
+
+static inline BackendTypeMask
+btmask_all_except3(BackendType t1, BackendType t2, BackendType t3)
+{
+ BackendTypeMask mask = BTYPE_MASK_ALL;
+
+ mask = btmask_del(mask, t1);
+ mask = btmask_del(mask, t2);
+ mask = btmask_del(mask, t3);
+ return mask;
+}
static inline bool
btmask_contains(BackendTypeMask mask, BackendType t)
@@ -329,6 +345,7 @@ typedef enum
* ckpt */
PM_SHUTDOWN_2, /* waiting for archiver and walsenders to
* finish */
+ PM_SHUTDOWN_IO, /* waiting for io workers to exit */
PM_WAIT_DEAD_END, /* waiting for dead-end children to exit */
PM_NO_CHILDREN, /* all important children have exited */
} PMState;
@@ -390,6 +407,10 @@ bool LoadedSSL = false;
static DNSServiceRef bonjour_sdref = NULL;
#endif
+/* State for IO worker management. */
+static int io_worker_count = 0;
+static PMChild *io_worker_children[MAX_IO_WORKERS];
+
/*
* postmaster.c - function prototypes
*/
@@ -424,6 +445,8 @@ static void TerminateChildren(int signal);
static int CountChildren(BackendTypeMask targetMask);
static void LaunchMissingBackgroundProcesses(void);
static void maybe_start_bgworkers(void);
+static bool maybe_reap_io_worker(int pid);
+static void maybe_adjust_io_workers(void);
static bool CreateOptsFile(int argc, char *argv[], char *fullprogname);
static PMChild *StartChildProcess(BackendType type);
static void StartSysLogger(void);
@@ -1351,6 +1374,11 @@ PostmasterMain(int argc, char *argv[])
*/
AddToDataDirLockFile(LOCK_FILE_LINE_PM_STATUS, PM_STATUS_STARTING);
+ pmState = PM_STARTUP;
+
+ /* Make sure we can perform I/O while starting up. */
+ maybe_adjust_io_workers();
+
/* Start bgwriter and checkpointer so they can help with recovery */
if (CheckpointerPMChild == NULL)
CheckpointerPMChild = StartChildProcess(B_CHECKPOINTER);
@@ -1363,7 +1391,6 @@ PostmasterMain(int argc, char *argv[])
StartupPMChild = StartChildProcess(B_STARTUP);
Assert(StartupPMChild != NULL);
StartupStatus = STARTUP_RUNNING;
- pmState = PM_STARTUP;
/* Some workers may be scheduled to start now */
maybe_start_bgworkers();
@@ -2503,6 +2530,16 @@ process_pm_child_exit(void)
continue;
}
+ /* Was it an IO worker? */
+ if (maybe_reap_io_worker(pid))
+ {
+ if (!EXIT_STATUS_0(exitstatus) && !EXIT_STATUS_1(exitstatus))
+ HandleChildCrash(pid, exitstatus, _("io worker"));
+
+ maybe_adjust_io_workers();
+ continue;
+ }
+
/*
* Was it a backend or a background worker?
*/
@@ -2867,10 +2904,10 @@ PostmasterStateMachine(void)
targetMask = btmask_add(targetMask, B_CHECKPOINTER);
/*
- * Walsenders and archiver will continue running; they will be
- * terminated later after writing the checkpoint record. We also let
- * dead-end children to keep running for now. The syslogger process
- * exits last.
+ * Walsenders, archiver and IO workers will continue running; they
+ * will be terminated later after writing the checkpoint record. We
+ * also let dead-end children to keep running for now. The syslogger
+ * process exits last.
*
* This assertion checks that we have covered all backend types,
* either by including them in targetMask, or by noting here that they
@@ -2882,6 +2919,7 @@ PostmasterStateMachine(void)
remainMask = btmask_add(remainMask, B_WAL_SENDER);
remainMask = btmask_add(remainMask, B_ARCHIVER);
+ remainMask = btmask_add(remainMask, B_IO_WORKER);
remainMask = btmask_add(remainMask, B_DEAD_END_BACKEND);
remainMask = btmask_add(remainMask, B_LOGGER);
@@ -2963,7 +3001,7 @@ PostmasterStateMachine(void)
pmState = PM_WAIT_DEAD_END;
ConfigurePostmasterWaitSet(false);
- /* Kill the walsenders and archiver too */
+ /* Kill walsenders, archiver and aio workers too */
SignalChildren(SIGQUIT, btmask_all_except(B_LOGGER));
}
}
@@ -2974,11 +3012,23 @@ PostmasterStateMachine(void)
{
/*
* PM_SHUTDOWN_2 state ends when there's no other children than
- * dead-end children left. There shouldn't be any regular backends
- * left by now anyway; what we're really waiting for is walsenders and
- * archiver.
+ * dead-end children and io workers left. There shouldn't be any
+ * regular backends left by now anyway; what we're really waiting for
+ * is walsenders and archiver.
*/
- if (CountChildren(btmask_all_except2(B_LOGGER, B_DEAD_END_BACKEND)) == 0)
+ if (CountChildren(btmask_all_except3(B_LOGGER, B_DEAD_END_BACKEND, B_IO_WORKER)) == 0)
+ {
+ pmState = PM_SHUTDOWN_IO;
+ SignalChildren(SIGUSR2, btmask(B_IO_WORKER));
+ }
+ }
+
+ if (pmState == PM_SHUTDOWN_IO)
+ {
+ /*
+ * PM_SHUTDOWN_IO state ends when there's only dead_end children left.
+ */
+ if (io_worker_count == 0)
{
pmState = PM_WAIT_DEAD_END;
ConfigurePostmasterWaitSet(false);
@@ -3094,10 +3144,14 @@ PostmasterStateMachine(void)
/* re-create shared memory and semaphores */
CreateSharedMemoryAndSemaphores();
+ pmState = PM_STARTUP;
+
+ /* Make sure we can perform I/O while starting up. */
+ maybe_adjust_io_workers();
+
StartupPMChild = StartChildProcess(B_STARTUP);
Assert(StartupPMChild != NULL);
StartupStatus = STARTUP_RUNNING;
- pmState = PM_STARTUP;
/* crash recovery started, reset SIGKILL flag */
AbortStartTime = 0;
@@ -3918,6 +3972,7 @@ bgworker_should_start_now(BgWorkerStartTime start_time)
{
case PM_NO_CHILDREN:
case PM_WAIT_DEAD_END:
+ case PM_SHUTDOWN_IO:
case PM_SHUTDOWN_2:
case PM_SHUTDOWN:
case PM_WAIT_BACKENDS:
@@ -4070,6 +4125,100 @@ maybe_start_bgworkers(void)
}
}
+static bool
+maybe_reap_io_worker(int pid)
+{
+ for (int id = 0; id < MAX_IO_WORKERS; ++id)
+ {
+ if (io_worker_children[id] &&
+ io_worker_children[id]->pid == pid)
+ {
+ ReleasePostmasterChildSlot(io_worker_children[id]);
+
+ --io_worker_count;
+ io_worker_children[id] = NULL;
+ return true;
+ }
+ }
+ return false;
+}
+
+static void
+maybe_adjust_io_workers(void)
+{
+ if (!pgaio_workers_enabled())
+ return;
+
+ /*
+ * If we're in final shutting down state, then we're just waiting for all
+ * processes to exit.
+ */
+ if (pmState >= PM_SHUTDOWN_IO)
+ return;
+
+ /* Don't start new workers during an immediate shutdown either. */
+ if (Shutdown >= ImmediateShutdown)
+ return;
+
+ /*
+ * Don't start new workers if we're in the shutdown phase of a crash
+ * restart. But we *do* need to start if we're already starting up again.
+ */
+ if (FatalError && pmState >= PM_STOP_BACKENDS)
+ return;
+
+ Assert(pmState < PM_SHUTDOWN_IO);
+
+ /* Not enough running? */
+ while (io_worker_count < io_workers)
+ {
+ PMChild *child;
+ int id;
+
+ /* find unused entry in io_worker_children array */
+ for (id = 0; id < MAX_IO_WORKERS; ++id)
+ {
+ if (io_worker_children[id] == NULL)
+ break;
+ }
+ if (id == MAX_IO_WORKERS)
+ elog(ERROR, "could not find a free IO worker ID");
+
+ /* Try to launch one. */
+ child = StartChildProcess(B_IO_WORKER);
+ if (child != NULL)
+ {
+ io_worker_children[id] = child;
+ ++io_worker_count;
+ }
+ else
+ break; /* XXX try again soon? */
+ }
+
+ /* Too many running? */
+ if (io_worker_count > io_workers)
+ {
+ /* ask the IO worker in the highest slot to exit */
+ for (int id = MAX_IO_WORKERS - 1; id >= 0; --id)
+ {
+ if (io_worker_children[id] != NULL)
+ {
+ kill(io_worker_children[id]->pid, SIGUSR2);
+ break;
+ }
+ }
+ }
+}
+
+void
+assign_io_workers(int newval, void *extra)
+{
+ io_workers = newval;
+ if (!IsUnderPostmaster && pmState > PM_INIT)
+ maybe_adjust_io_workers();
+}
+
+
/*
* When a backend asks to be notified about worker state changes, we
* set a flag in its backend entry. The background worker machinery needs
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index b253278f3c1..fa2a7e9e5df 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -14,6 +14,7 @@ OBJS = \
aio_io.o \
aio_subject.o \
method_sync.o \
+ method_worker.o \
read_stream.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index b9bdf51680a..0c2d77ec8ab 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -217,3 +217,10 @@ pgaio_init_backend(void)
if (pgaio_impl->init_backend)
pgaio_impl->init_backend();
}
+
+bool
+pgaio_workers_enabled(void)
+{
+ /* placeholder for future commit */
+ return false;
+}
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index 8339d473aae..62738ce1d14 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -6,5 +6,6 @@ backend_sources += files(
'aio_io.c',
'aio_subject.c',
'method_sync.c',
+ 'method_worker.c',
'read_stream.c',
)
diff --git a/src/backend/storage/aio/method_worker.c b/src/backend/storage/aio/method_worker.c
new file mode 100644
index 00000000000..0ea749a8ba8
--- /dev/null
+++ b/src/backend/storage/aio/method_worker.c
@@ -0,0 +1,86 @@
+/*-------------------------------------------------------------------------
+ *
+ * method_worker.c
+ * AIO implementation using workers
+ *
+ * Portions Copyright (c) 1996-2021, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/method_worker.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "libpq/pqsignal.h"
+#include "miscadmin.h"
+#include "postmaster/auxprocess.h"
+#include "postmaster/interrupt.h"
+#include "storage/io_worker.h"
+#include "storage/ipc.h"
+#include "storage/latch.h"
+#include "storage/proc.h"
+#include "tcop/tcopprot.h"
+#include "utils/wait_event.h"
+
+
+int io_workers = 3;
+
+
+void
+IoWorkerMain(char *startup_data, size_t startup_data_len)
+{
+ sigjmp_buf local_sigjmp_buf;
+
+ MyBackendType = B_IO_WORKER;
+ AuxiliaryProcessMainCommon();
+
+ /* TODO review all signals */
+ pqsignal(SIGHUP, SignalHandlerForConfigReload);
+ pqsignal(SIGINT, die); /* to allow manually triggering worker restart */
+
+ /*
+ * Ignore SIGTERM, will get explicit shutdown via SIGUSR2 later in the
+ * shutdown sequence, similar to checkpointer.
+ */
+ pqsignal(SIGTERM, SIG_IGN);
+ /* SIGQUIT handler was already set up by InitPostmasterChild */
+ pqsignal(SIGALRM, SIG_IGN);
+ pqsignal(SIGPIPE, SIG_IGN);
+ pqsignal(SIGUSR1, procsignal_sigusr1_handler);
+ pqsignal(SIGUSR2, SignalHandlerForShutdownRequest);
+ sigprocmask(SIG_SETMASK, &UnBlockSig, NULL);
+
+ /* see PostgresMain() */
+ if (sigsetjmp(local_sigjmp_buf, 1) != 0)
+ {
+ error_context_stack = NULL;
+ HOLD_INTERRUPTS();
+
+ /*
+ * We normally shouldn't get errors here. Need to do just enough error
+ * recovery so that we can mark the IO as failed and then exit.
+ */
+ LWLockReleaseAll();
+
+ /* TODO: recover from IO errors */
+
+ EmitErrorReport();
+ proc_exit(1);
+ }
+
+ /* We can now handle ereport(ERROR) */
+ PG_exception_stack = &local_sigjmp_buf;
+
+ while (!ShutdownRequestPending)
+ {
+ WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, -1,
+ WAIT_EVENT_IO_WORKER_MAIN);
+ ResetLatch(MyLatch);
+ CHECK_FOR_INTERRUPTS();
+ }
+
+ proc_exit(0);
+}
diff --git a/src/backend/tcop/postgres.c b/src/backend/tcop/postgres.c
index 85902788181..fcd3e1eb482 100644
--- a/src/backend/tcop/postgres.c
+++ b/src/backend/tcop/postgres.c
@@ -3313,6 +3313,8 @@ ProcessInterrupts(void)
(errcode(ERRCODE_ADMIN_SHUTDOWN),
errmsg("terminating background worker \"%s\" due to administrator command",
MyBgworkerEntry->bgw_type)));
+ else if (AmIoWorkerProcess())
+ proc_exit(0);
else
ereport(FATAL,
(errcode(ERRCODE_ADMIN_SHUTDOWN),
diff --git a/src/backend/utils/activity/pgstat_backend.c b/src/backend/utils/activity/pgstat_backend.c
index 6b2c9baa8c0..c48befef6a7 100644
--- a/src/backend/utils/activity/pgstat_backend.c
+++ b/src/backend/utils/activity/pgstat_backend.c
@@ -166,6 +166,7 @@ pgstat_tracks_backend_bktype(BackendType bktype)
case B_WAL_SUMMARIZER:
case B_BG_WRITER:
case B_CHECKPOINTER:
+ case B_IO_WORKER:
case B_STARTUP:
return false;
diff --git a/src/backend/utils/activity/pgstat_io.c b/src/backend/utils/activity/pgstat_io.c
index 011a3326dad..7869197dd1f 100644
--- a/src/backend/utils/activity/pgstat_io.c
+++ b/src/backend/utils/activity/pgstat_io.c
@@ -365,6 +365,7 @@ pgstat_tracks_io_bktype(BackendType bktype)
case B_INVALID:
case B_DEAD_END_BACKEND:
case B_ARCHIVER:
+ case B_IO_WORKER:
case B_LOGGER:
case B_WAL_RECEIVER:
case B_WAL_WRITER:
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 7a2e2b4432e..330a32a90ce 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -56,6 +56,7 @@ AUTOVACUUM_MAIN "Waiting in main loop of autovacuum launcher process."
BGWRITER_HIBERNATE "Waiting in background writer process, hibernating."
BGWRITER_MAIN "Waiting in main loop of background writer process."
CHECKPOINTER_MAIN "Waiting in main loop of checkpointer process."
+IO_WORKER_MAIN "Waiting in main loop of IO Worker process."
LOGICAL_APPLY_MAIN "Waiting in main loop of logical replication apply process."
LOGICAL_LAUNCHER_MAIN "Waiting in main loop of logical replication launcher process."
LOGICAL_PARALLEL_APPLY_MAIN "Waiting in main loop of logical replication parallel apply process."
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index 6349abb8fb6..56133cfdd08 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -293,6 +293,9 @@ GetBackendTypeDesc(BackendType backendType)
case B_CHECKPOINTER:
backendDesc = gettext_noop("checkpointer");
break;
+ case B_IO_WORKER:
+ backendDesc = "io worker";
+ break;
case B_LOGGER:
backendDesc = gettext_noop("logger");
break;
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 6d4056c68b9..b2999b86c24 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -74,6 +74,7 @@
#include "storage/aio.h"
#include "storage/bufmgr.h"
#include "storage/bufpage.h"
+#include "storage/io_worker.h"
#include "storage/large_object.h"
#include "storage/pg_shmem.h"
#include "storage/predicate.h"
@@ -3232,6 +3233,18 @@ struct config_int ConfigureNamesInt[] =
NULL, NULL, NULL
},
+ {
+ {"io_workers",
+ PGC_SIGHUP,
+ RESOURCES_ASYNCHRONOUS,
+ gettext_noop("Number of IO worker processes, for io_method=worker."),
+ NULL,
+ },
+ &io_workers,
+ 3, 1, MAX_IO_WORKERS,
+ NULL, assign_io_workers, NULL
+ },
+
{
{"backend_flush_after", PGC_USERSET, RESOURCES_ASYNCHRONOUS,
gettext_noop("Number of pages after which previously performed writes are flushed to disk."),
diff --git a/src/backend/utils/misc/postgresql.conf.sample b/src/backend/utils/misc/postgresql.conf.sample
index c4c60da9845..0f80a0680ec 100644
--- a/src/backend/utils/misc/postgresql.conf.sample
+++ b/src/backend/utils/misc/postgresql.conf.sample
@@ -843,6 +843,7 @@
#------------------------------------------------------------------------------
#io_method = sync # (change requires restart)
+#io_workers = 3 # 1-32;
#io_max_concurrency = 32 # Max number of IOs that may be in
# flight at the same time in one backend
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0006-aio-Add-worker-method.patch (17.6K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/7-v2-0006-aio-Add-worker-method.patch)
download | inline diff:
From 9c9bbb42fb561fb2cf7d6d5183db5359d37e004e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Fri, 8 Nov 2024 12:38:41 -0500
Subject: [PATCH v2 06/20] aio: Add worker method
---
src/include/storage/aio.h | 5 +-
src/include/storage/aio_internal.h | 1 +
src/include/storage/lwlocklist.h | 1 +
src/backend/storage/aio/aio.c | 2 +
src/backend/storage/aio/aio_init.c | 12 +-
src/backend/storage/aio/method_worker.c | 406 +++++++++++++++++-
.../utils/activity/wait_event_names.txt | 1 +
src/backend/utils/misc/postgresql.conf.sample | 2 +-
src/tools/pgindent/typedefs.list | 3 +
9 files changed, 423 insertions(+), 10 deletions(-)
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index b386dabc921..2e84abfea2d 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -322,11 +322,12 @@ extern void assign_io_method(int newval, void *extra);
typedef enum IoMethod
{
IOMETHOD_SYNC = 0,
+ IOMETHOD_WORKER,
} IoMethod;
-/* We'll default to synchronous execution. */
-#define DEFAULT_IO_METHOD IOMETHOD_SYNC
+/* We'll default to bgworker. */
+#define DEFAULT_IO_METHOD IOMETHOD_WORKER
/* GUCs */
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
index d600d45b4fd..f974c4accf5 100644
--- a/src/include/storage/aio_internal.h
+++ b/src/include/storage/aio_internal.h
@@ -234,6 +234,7 @@ extern const char *pgaio_io_get_state_name(PgAioHandle *ioh);
/* Declarations for the tables of function pointers exposed by each IO method. */
extern const IoMethodOps pgaio_sync_ops;
+extern const IoMethodOps pgaio_worker_ops;
extern const IoMethodOps *pgaio_impl;
extern PgAioCtl *aio_ctl;
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index 6a2f64c54fb..8d00d62e208 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -83,3 +83,4 @@ PG_LWLOCK(49, WALSummarizer)
PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
+PG_LWLOCK(53, AioWorkerSubmissionQueue)
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index 3e2ff9718ca..e4c9d439ddd 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -57,6 +57,7 @@ static PgAioHandle *pgaio_io_from_ref(PgAioHandleRef *ior, uint64 *ref_generatio
/* Options for io_method. */
const struct config_enum_entry io_method_options[] = {
{"sync", IOMETHOD_SYNC, false},
+ {"worker", IOMETHOD_WORKER, false},
{NULL, 0, false}
};
@@ -73,6 +74,7 @@ PgAioPerBackend *my_aio;
static const IoMethodOps *pgaio_ops_table[] = {
[IOMETHOD_SYNC] = &pgaio_sync_ops,
+ [IOMETHOD_WORKER] = &pgaio_worker_ops,
};
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index 0c2d77ec8ab..23adc5308e5 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -19,6 +19,7 @@
#include "storage/aio_init.h"
#include "storage/aio_internal.h"
#include "storage/bufmgr.h"
+#include "storage/io_worker.h"
#include "storage/proc.h"
#include "storage/shmem.h"
@@ -37,6 +38,11 @@ AioCtlShmemSize(void)
static uint32
AioProcs(void)
{
+ /*
+ * While AIO workers don't need their own AIO context, we can't currently
+ * guarantee nothing gets assigned to the a ProcNumber for an IO worker if
+ * we just subtracted MAX_IO_WORKERS.
+ */
return MaxBackends + NUM_AUXILIARY_PROCS;
}
@@ -209,6 +215,9 @@ pgaio_init_backend(void)
/* shouldn't be initialized twice */
Assert(!my_aio);
+ if (MyBackendType == B_IO_WORKER)
+ return;
+
if (MyProc == NULL || MyProcNumber >= AioProcs())
elog(ERROR, "aio requires a normal PGPROC");
@@ -221,6 +230,5 @@ pgaio_init_backend(void)
bool
pgaio_workers_enabled(void)
{
- /* placeholder for future commit */
- return false;
+ return io_method == IOMETHOD_WORKER;
}
diff --git a/src/backend/storage/aio/method_worker.c b/src/backend/storage/aio/method_worker.c
index 0ea749a8ba8..a508f53ebd4 100644
--- a/src/backend/storage/aio/method_worker.c
+++ b/src/backend/storage/aio/method_worker.c
@@ -1,7 +1,22 @@
/*-------------------------------------------------------------------------
*
* method_worker.c
- * AIO implementation using workers
+ * AIO - perform AIO using worker processes
+ *
+ * Worker processes consume IOs from a shared memory submission queue, run
+ * traditional synchronous system calls, and perform the shared completion
+ * handling immediately. Client code submits most requests by pushing IOs
+ * into the submission queue, and waits (if necessary) using condition
+ * variables. Some IOs cannot be performed in another process due to lack of
+ * infrastructure for reopening the file, and must processed synchronously by
+ * the client code when submitted.
+ *
+ * So that the submitter can make just one system call when submitting a batch
+ * of IOs, wakeups "fan out"; each woken backend can wake two more. XXX This
+ * could be improved by using futexes instead of latches to wake N waiters.
+ *
+ * This method of AIO is available in all builds on all operating systems, and
+ * is the default.
*
* Portions Copyright (c) 1996-2021, PostgreSQL Global Development Group
* Portions Copyright (c) 1994, Regents of the University of California
@@ -16,23 +31,323 @@
#include "libpq/pqsignal.h"
#include "miscadmin.h"
+#include "port/pg_bitutils.h"
#include "postmaster/auxprocess.h"
#include "postmaster/interrupt.h"
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
#include "storage/io_worker.h"
#include "storage/ipc.h"
#include "storage/latch.h"
#include "storage/proc.h"
#include "tcop/tcopprot.h"
+#include "utils/ps_status.h"
#include "utils/wait_event.h"
+/* How many workers should each worker wake up if needed? */
+#define IO_WORKER_WAKEUP_FANOUT 2
+
+
+typedef struct AioWorkerSubmissionQueue
+{
+ uint32 size;
+ uint32 mask;
+ uint32 head;
+ uint32 tail;
+ uint32 ios[FLEXIBLE_ARRAY_MEMBER];
+} AioWorkerSubmissionQueue;
+
+typedef struct AioWorkerSlot
+{
+ Latch *latch;
+ bool in_use;
+} AioWorkerSlot;
+
+typedef struct AioWorkerControl
+{
+ uint64 idle_worker_mask;
+ AioWorkerSlot workers[FLEXIBLE_ARRAY_MEMBER];
+} AioWorkerControl;
+
+
+static size_t pgaio_worker_shmem_size(void);
+static void pgaio_worker_shmem_init(bool first_time);
+
+static bool pgaio_worker_needs_synchronous_execution(PgAioHandle *ioh);
+static int pgaio_worker_submit(uint16 num_staged_ios, PgAioHandle **staged_ios);
+
+
+const IoMethodOps pgaio_worker_ops = {
+ .shmem_size = pgaio_worker_shmem_size,
+ .shmem_init = pgaio_worker_shmem_init,
+
+ .needs_synchronous_execution = pgaio_worker_needs_synchronous_execution,
+ .submit = pgaio_worker_submit,
+};
+
+
int io_workers = 3;
+static int io_worker_queue_size = 64;
+static int MyIoWorkerId;
+
+
+static AioWorkerSubmissionQueue *io_worker_submission_queue;
+static AioWorkerControl *io_worker_control;
+
+
+static size_t
+pgaio_worker_shmem_size(void)
+{
+ return
+ offsetof(AioWorkerSubmissionQueue, ios) +
+ sizeof(uint32) * io_worker_queue_size +
+ offsetof(AioWorkerControl, workers) +
+ sizeof(AioWorkerSlot) * io_workers;
+}
+
+static void
+pgaio_worker_shmem_init(bool first_time)
+{
+ bool found;
+ int size;
+
+ /* Round size up to next power of two so we can make a mask. */
+ size = pg_nextpower2_32(io_worker_queue_size);
+
+ io_worker_submission_queue =
+ ShmemInitStruct("AioWorkerSubmissionQueue",
+ offsetof(AioWorkerSubmissionQueue, ios) +
+ sizeof(uint32) * size,
+ &found);
+ if (!found)
+ {
+ io_worker_submission_queue->size = size;
+ io_worker_submission_queue->head = 0;
+ io_worker_submission_queue->tail = 0;
+ }
+
+ io_worker_control =
+ ShmemInitStruct("AioWorkerControl",
+ offsetof(AioWorkerControl, workers) +
+ sizeof(AioWorkerSlot) * io_workers,
+ &found);
+ if (!found)
+ {
+ io_worker_control->idle_worker_mask = 0;
+ for (int i = 0; i < io_workers; ++i)
+ {
+ io_worker_control->workers[i].latch = NULL;
+ io_worker_control->workers[i].in_use = false;
+ }
+ }
+}
+
+
+static int
+pgaio_choose_idle_worker(void)
+{
+ int worker;
+
+ if (io_worker_control->idle_worker_mask == 0)
+ return -1;
+
+ /* Find the lowest bit position, and clear it. */
+ worker = pg_rightmost_one_pos64(io_worker_control->idle_worker_mask);
+ io_worker_control->idle_worker_mask &= ~(UINT64_C(1) << worker);
+
+ return worker;
+}
+
+static bool
+pgaio_worker_submission_queue_insert(PgAioHandle *ioh)
+{
+ AioWorkerSubmissionQueue *queue;
+ uint32 new_head;
+
+ queue = io_worker_submission_queue;
+ new_head = (queue->head + 1) & (queue->size - 1);
+ if (new_head == queue->tail)
+ {
+ elog(DEBUG1, "full");
+ return false; /* full */
+ }
+
+ queue->ios[queue->head] = pgaio_io_get_id(ioh);
+ queue->head = new_head;
+
+ return true;
+}
+
+static uint32
+pgaio_worker_submission_queue_consume(void)
+{
+ AioWorkerSubmissionQueue *queue;
+ uint32 result;
+
+ queue = io_worker_submission_queue;
+ if (queue->tail == queue->head)
+ return UINT32_MAX; /* empty */
+
+ result = queue->ios[queue->tail];
+ queue->tail = (queue->tail + 1) & (queue->size - 1);
+
+ return result;
+}
+
+static uint32
+pgaio_worker_submission_queue_depth(void)
+{
+ uint32 head;
+ uint32 tail;
+
+ head = io_worker_submission_queue->head;
+ tail = io_worker_submission_queue->tail;
+
+ if (tail > head)
+ head += io_worker_submission_queue->size;
+
+ Assert(head >= tail);
+
+ return head - tail;
+}
+
+static void
+pgaio_worker_submit_internal(int nios, PgAioHandle *ios[])
+{
+ PgAioHandle *synchronous_ios[PGAIO_SUBMIT_BATCH_SIZE];
+ int nsync = 0;
+ Latch *wakeup = NULL;
+ int worker;
+
+ Assert(nios <= PGAIO_SUBMIT_BATCH_SIZE);
+
+ LWLockAcquire(AioWorkerSubmissionQueueLock, LW_EXCLUSIVE);
+ for (int i = 0; i < nios; ++i)
+ {
+ Assert(!pgaio_worker_needs_synchronous_execution(ios[i]));
+ if (!pgaio_worker_submission_queue_insert(ios[i]))
+ {
+ /*
+ * We'll do it synchronously, but only after we've sent as many as
+ * we can to workers, to maximize concurrency.
+ */
+ synchronous_ios[nsync++] = ios[i];
+ continue;
+ }
+
+ if (wakeup == NULL)
+ {
+ /* Choose an idle worker to wake up if we haven't already. */
+ worker = pgaio_choose_idle_worker();
+ if (worker >= 0)
+ wakeup = io_worker_control->workers[worker].latch;
+
+ ereport(DEBUG3,
+ errmsg("submission for io:%d choosing worker %d, latch %p",
+ pgaio_io_get_id(ios[i]), worker, wakeup),
+ errhidestmt(true), errhidecontext(true));
+ }
+ }
+ LWLockRelease(AioWorkerSubmissionQueueLock);
+
+ if (wakeup)
+ SetLatch(wakeup);
+
+ /* Run whatever is left synchronously. */
+ if (nsync > 0)
+ {
+ for (int i = 0; i < nsync; ++i)
+ {
+ pgaio_io_perform_synchronously(synchronous_ios[i]);
+ }
+ }
+}
+
+static bool
+pgaio_worker_needs_synchronous_execution(PgAioHandle *ioh)
+{
+ return
+ !IsUnderPostmaster
+ || ioh->flags & AHF_REFERENCES_LOCAL
+ || !pgaio_io_can_reopen(ioh);
+}
+
+static int
+pgaio_worker_submit(uint16 num_staged_ios, PgAioHandle **staged_ios)
+{
+ for (int i = 0; i < num_staged_ios; i++)
+ {
+ PgAioHandle *ioh = staged_ios[i];
+
+ pgaio_io_prepare_submit(ioh);
+ }
+
+ pgaio_worker_submit_internal(num_staged_ios, staged_ios);
+
+ return num_staged_ios;
+}
+
+/*
+ * shmem_exit() callback that releases the worker's slot in io_worker_control.
+ */
+static void
+pgaio_worker_die(int code, Datum arg)
+{
+ LWLockAcquire(AioWorkerSubmissionQueueLock, LW_EXCLUSIVE);
+ Assert(io_worker_control->workers[MyIoWorkerId].in_use);
+ Assert(io_worker_control->workers[MyIoWorkerId].latch == MyLatch);
+
+ io_worker_control->workers[MyIoWorkerId].in_use = false;
+ io_worker_control->workers[MyIoWorkerId].latch = NULL;
+ LWLockRelease(AioWorkerSubmissionQueueLock);
+}
+
+/*
+ * Register the worker in shared memory, assign MyWorkerId and register a
+ * shutdown callback to release registration.
+ */
+static void
+pgaio_worker_register(void)
+{
+ MyIoWorkerId = -1;
+
+ /*
+ * XXX: This could do with more fine-grained locking. But it's also not
+ * very common for the number of workers to change at the moment...
+ */
+ LWLockAcquire(AioWorkerSubmissionQueueLock, LW_EXCLUSIVE);
+
+ for (int i = 0; i < io_workers; ++i)
+ {
+ if (!io_worker_control->workers[i].in_use)
+ {
+ Assert(io_worker_control->workers[i].latch == NULL);
+ io_worker_control->workers[i].in_use = true;
+ MyIoWorkerId = i;
+ break;
+ }
+ else
+ Assert(io_worker_control->workers[i].latch != NULL);
+ }
+
+ if (MyIoWorkerId == -1)
+ elog(ERROR, "couldn't find a free worker slot");
+
+ io_worker_control->idle_worker_mask |= (UINT64_C(1) << MyIoWorkerId);
+ io_worker_control->workers[MyIoWorkerId].latch = MyLatch;
+ LWLockRelease(AioWorkerSubmissionQueueLock);
+
+ on_shmem_exit(pgaio_worker_die, 0);
+}
void
IoWorkerMain(char *startup_data, size_t startup_data_len)
{
sigjmp_buf local_sigjmp_buf;
+ volatile PgAioHandle *ioh = NULL;
+ char cmd[128];
MyBackendType = B_IO_WORKER;
AuxiliaryProcessMainCommon();
@@ -53,6 +368,11 @@ IoWorkerMain(char *startup_data, size_t startup_data_len)
pqsignal(SIGUSR2, SignalHandlerForShutdownRequest);
sigprocmask(SIG_SETMASK, &UnBlockSig, NULL);
+ pgaio_worker_register();
+
+ sprintf(cmd, "io worker: %d", MyIoWorkerId);
+ set_ps_display(cmd);
+
/* see PostgresMain() */
if (sigsetjmp(local_sigjmp_buf, 1) != 0)
{
@@ -66,8 +386,26 @@ IoWorkerMain(char *startup_data, size_t startup_data_len)
LWLockReleaseAll();
/* TODO: recover from IO errors */
+ if (ioh != NULL)
+ {
+#if 0
+ /* EINTR is treated as a retryable error */
+ pgaio_process_io_completion(unvolatize(PgAioInProgress *, io),
+ EINTR);
+#endif
+ }
EmitErrorReport();
+
+ /* FIXME: should probably be a before-shmem-exit instead */
+ LWLockAcquire(AioWorkerSubmissionQueueLock, LW_EXCLUSIVE);
+ Assert(io_worker_control->workers[MyIoWorkerId].in_use);
+ Assert(io_worker_control->workers[MyIoWorkerId].latch == MyLatch);
+
+ io_worker_control->workers[MyIoWorkerId].in_use = false;
+ io_worker_control->workers[MyIoWorkerId].latch = NULL;
+ LWLockRelease(AioWorkerSubmissionQueueLock);
+
proc_exit(1);
}
@@ -76,10 +414,68 @@ IoWorkerMain(char *startup_data, size_t startup_data_len)
while (!ShutdownRequestPending)
{
- WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, -1,
- WAIT_EVENT_IO_WORKER_MAIN);
- ResetLatch(MyLatch);
- CHECK_FOR_INTERRUPTS();
+ uint32 io_index;
+ Latch *latches[IO_WORKER_WAKEUP_FANOUT];
+ int nlatches = 0;
+ int nwakeups = 0;
+ int worker;
+
+ /* Try to get a job to do. */
+ LWLockAcquire(AioWorkerSubmissionQueueLock, LW_EXCLUSIVE);
+ if ((io_index = pgaio_worker_submission_queue_consume()) == UINT32_MAX)
+ {
+ /* Nothing to do. Mark self idle. */
+ /*
+ * XXX: Invent some kind of back pressure to reduce useless
+ * wakeups?
+ */
+ io_worker_control->idle_worker_mask |= (UINT64_C(1) << MyIoWorkerId);
+ }
+ else
+ {
+ /* Got one. Clear idle flag. */
+ io_worker_control->idle_worker_mask &= ~(UINT64_C(1) << MyIoWorkerId);
+
+ /* See if we can wake up some peers. */
+ nwakeups = Min(pgaio_worker_submission_queue_depth(),
+ IO_WORKER_WAKEUP_FANOUT);
+ for (int i = 0; i < nwakeups; ++i)
+ {
+ if ((worker = pgaio_choose_idle_worker()) < 0)
+ break;
+ latches[nlatches++] = io_worker_control->workers[worker].latch;
+ }
+#if 0
+ if (nwakeups > 0)
+ elog(LOG, "wake %d", nwakeups);
+#endif
+ }
+ LWLockRelease(AioWorkerSubmissionQueueLock);
+
+ for (int i = 0; i < nlatches; ++i)
+ SetLatch(latches[i]);
+
+ if (io_index != UINT32_MAX)
+ {
+ ioh = &aio_ctl->io_handles[io_index];
+
+ ereport(DEBUG3,
+ errmsg("worker processing io:%d",
+ pgaio_io_get_id(unvolatize(PgAioHandle *, ioh))),
+ errhidestmt(true), errhidecontext(true));
+
+ pgaio_io_reopen(unvolatize(PgAioHandle *, ioh));
+ pgaio_io_perform_synchronously(unvolatize(PgAioHandle *, ioh));
+
+ ioh = NULL;
+ }
+ else
+ {
+ WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, -1,
+ WAIT_EVENT_IO_WORKER_MAIN);
+ ResetLatch(MyLatch);
+ CHECK_FOR_INTERRUPTS();
+ }
}
proc_exit(0);
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 330a32a90ce..8c3aafd8a18 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -349,6 +349,7 @@ WALSummarizer "Waiting to read or update WAL summarization state."
DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
+AioWorkerSubmissionQueue "Waiting to access AIO worker submission queue."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/misc/postgresql.conf.sample b/src/backend/utils/misc/postgresql.conf.sample
index 0f80a0680ec..5893eb29228 100644
--- a/src/backend/utils/misc/postgresql.conf.sample
+++ b/src/backend/utils/misc/postgresql.conf.sample
@@ -842,7 +842,7 @@
# WIP AIO GUC docs
#------------------------------------------------------------------------------
-#io_method = sync # (change requires restart)
+#io_method = worker # (change requires restart)
#io_workers = 3 # 1-32;
#io_max_concurrency = 32 # Max number of IOs that may be in
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index bc1acbb98ee..9b9c8f0d1fc 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -54,6 +54,9 @@ AggStrategy
AggTransInfo
Aggref
AggregateInstrumentation
+AioWorkerControl
+AioWorkerSlot
+AioWorkerSubmissionQueue
AlenState
Alias
AllocBlock
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0007-aio-Add-liburing-dependency.patch (9.9K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/8-v2-0007-aio-Add-liburing-dependency.patch)
download | inline diff:
From 309863778a6051b0e18d949551961608dbf9d399 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 5 Jun 2024 19:37:25 -0700
Subject: [PATCH v2 07/20] aio: Add liburing dependency
Not yet used.
Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
meson.build | 14 ++++
meson_options.txt | 3 +
configure.ac | 11 +++
src/makefiles/meson.build | 3 +
src/include/pg_config.h.in | 3 +
configure | 138 +++++++++++++++++++++++++++++++++++++
src/Makefile.global.in | 4 ++
7 files changed, 176 insertions(+)
diff --git a/meson.build b/meson.build
index e5ce437a5c7..76c276437d7 100644
--- a/meson.build
+++ b/meson.build
@@ -854,6 +854,18 @@ endif
+###############################################################
+# Library: liburing
+###############################################################
+
+liburingopt = get_option('liburing')
+liburing = dependency('liburing', required: liburingopt)
+if liburing.found()
+ cdata.set('USE_LIBURING', 1)
+endif
+
+
+
###############################################################
# Library: libxml
###############################################################
@@ -3054,6 +3066,7 @@ backend_both_deps += [
icu_i18n,
ldap,
libintl,
+ liburing,
libxml,
lz4,
pam,
@@ -3698,6 +3711,7 @@ if meson.version().version_compare('>=0.57')
'gss': gssapi,
'icu': icu,
'ldap': ldap,
+ 'liburing': liburing,
'libxml': libxml,
'libxslt': libxslt,
'llvm': llvm,
diff --git a/meson_options.txt b/meson_options.txt
index 38935196394..6e8d376b3b2 100644
--- a/meson_options.txt
+++ b/meson_options.txt
@@ -103,6 +103,9 @@ option('ldap', type: 'feature', value: 'auto',
option('libedit_preferred', type: 'boolean', value: false,
description: 'Prefer BSD Libedit over GNU Readline')
+option('liburing', type : 'feature', value: 'auto',
+ description: 'Use liburing for async io')
+
option('libxml', type: 'feature', value: 'auto',
description: 'XML support')
diff --git a/configure.ac b/configure.ac
index 247ae97fa4c..dda296ee029 100644
--- a/configure.ac
+++ b/configure.ac
@@ -975,6 +975,14 @@ AC_SUBST(with_readline)
PGAC_ARG_BOOL(with, libedit-preferred, no,
[prefer BSD Libedit over GNU Readline])
+#
+# liburing
+#
+AC_MSG_CHECKING([whether to build with liburing support])
+PGAC_ARG_BOOL(with, liburing, no, [use liburing for async io],
+ [AC_DEFINE([USE_LIBURING], 1, [Define to build with io-uring support. (--with-liburing)])])
+AC_MSG_RESULT([$with_liburing])
+AC_SUBST(with_liburing)
#
# UUID library
@@ -1427,6 +1435,9 @@ elif test "$with_uuid" = ossp ; then
fi
AC_SUBST(UUID_LIBS)
+if test "$with_liburing" = yes; then
+ PKG_CHECK_MODULES(LIBURING, liburing)
+fi
##
## Header files
diff --git a/src/makefiles/meson.build b/src/makefiles/meson.build
index aba7411a1be..00613aebc79 100644
--- a/src/makefiles/meson.build
+++ b/src/makefiles/meson.build
@@ -199,6 +199,8 @@ pgxs_empty = [
'PTHREAD_CFLAGS', 'PTHREAD_LIBS',
'ICU_LIBS',
+
+ 'LIBURING_CFLAGS', 'LIBURING_LIBS',
]
if host_system == 'windows' and cc.get_argument_syntax() != 'msvc'
@@ -229,6 +231,7 @@ pgxs_deps = {
'gssapi': gssapi,
'icu': icu,
'ldap': ldap,
+ 'liburing': liburing,
'libxml': libxml,
'libxslt': libxslt,
'llvm': llvm,
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 07b2f798abd..6ab71a3dffe 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -663,6 +663,9 @@
/* Define to 1 to build with LDAP support. (--with-ldap) */
#undef USE_LDAP
+/* Define to build with io-uring support. (--with-liburing) */
+#undef USE_LIBURING
+
/* Define to 1 to build with XML support. (--with-libxml) */
#undef USE_LIBXML
diff --git a/configure b/configure
index 518c33b73a9..1c3fada9fe0 100755
--- a/configure
+++ b/configure
@@ -651,6 +651,8 @@ LIBOBJS
OPENSSL
ZSTD
LZ4
+LIBURING_LIBS
+LIBURING_CFLAGS
UUID_LIBS
LDAP_LIBS_BE
LDAP_LIBS_FE
@@ -709,6 +711,7 @@ XML2_CFLAGS
XML2_CONFIG
with_libxml
with_uuid
+with_liburing
with_readline
with_systemd
with_selinux
@@ -862,6 +865,7 @@ with_selinux
with_systemd
with_readline
with_libedit_preferred
+with_liburing
with_uuid
with_ossp_uuid
with_libxml
@@ -905,6 +909,8 @@ LDFLAGS_EX
LDFLAGS_SL
PERL
PYTHON
+LIBURING_CFLAGS
+LIBURING_LIBS
MSGFMT
TCLSH'
@@ -1572,6 +1578,7 @@ Optional Packages:
--without-readline do not use GNU Readline nor BSD Libedit for editing
--with-libedit-preferred
prefer BSD Libedit over GNU Readline
+ --with-liburing use liburing for async io
--with-uuid=LIB build contrib/uuid-ossp using LIB (bsd,e2fs,ossp)
--with-ossp-uuid obsolete spelling of --with-uuid=ossp
--with-libxml build with XML support
@@ -1618,6 +1625,10 @@ Some influential environment variables:
LDFLAGS_SL extra linker flags for linking shared libraries only
PERL Perl program
PYTHON Python program
+ LIBURING_CFLAGS
+ C compiler flags for LIBURING, overriding pkg-config
+ LIBURING_LIBS
+ linker flags for LIBURING, overriding pkg-config
MSGFMT msgfmt program for NLS
TCLSH Tcl interpreter program (tclsh)
@@ -8681,6 +8692,40 @@ fi
+#
+# liburing
+#
+{ $as_echo "$as_me:${as_lineno-$LINENO}: checking whether to build with liburing support" >&5
+$as_echo_n "checking whether to build with liburing support... " >&6; }
+
+
+
+# Check whether --with-liburing was given.
+if test "${with_liburing+set}" = set; then :
+ withval=$with_liburing;
+ case $withval in
+ yes)
+
+$as_echo "#define USE_LIBURING 1" >>confdefs.h
+
+ ;;
+ no)
+ :
+ ;;
+ *)
+ as_fn_error $? "no argument expected for --with-liburing option" "$LINENO" 5
+ ;;
+ esac
+
+else
+ with_liburing=no
+
+fi
+
+
+{ $as_echo "$as_me:${as_lineno-$LINENO}: result: $with_liburing" >&5
+$as_echo "$with_liburing" >&6; }
+
#
# UUID library
@@ -13222,6 +13267,99 @@ fi
fi
+if test "$with_liburing" = yes; then
+
+pkg_failed=no
+{ $as_echo "$as_me:${as_lineno-$LINENO}: checking for liburing" >&5
+$as_echo_n "checking for liburing... " >&6; }
+
+if test -n "$LIBURING_CFLAGS"; then
+ pkg_cv_LIBURING_CFLAGS="$LIBURING_CFLAGS"
+ elif test -n "$PKG_CONFIG"; then
+ if test -n "$PKG_CONFIG" && \
+ { { $as_echo "$as_me:${as_lineno-$LINENO}: \$PKG_CONFIG --exists --print-errors \"liburing\""; } >&5
+ ($PKG_CONFIG --exists --print-errors "liburing") 2>&5
+ ac_status=$?
+ $as_echo "$as_me:${as_lineno-$LINENO}: \$? = $ac_status" >&5
+ test $ac_status = 0; }; then
+ pkg_cv_LIBURING_CFLAGS=`$PKG_CONFIG --cflags "liburing" 2>/dev/null`
+ test "x$?" != "x0" && pkg_failed=yes
+else
+ pkg_failed=yes
+fi
+ else
+ pkg_failed=untried
+fi
+if test -n "$LIBURING_LIBS"; then
+ pkg_cv_LIBURING_LIBS="$LIBURING_LIBS"
+ elif test -n "$PKG_CONFIG"; then
+ if test -n "$PKG_CONFIG" && \
+ { { $as_echo "$as_me:${as_lineno-$LINENO}: \$PKG_CONFIG --exists --print-errors \"liburing\""; } >&5
+ ($PKG_CONFIG --exists --print-errors "liburing") 2>&5
+ ac_status=$?
+ $as_echo "$as_me:${as_lineno-$LINENO}: \$? = $ac_status" >&5
+ test $ac_status = 0; }; then
+ pkg_cv_LIBURING_LIBS=`$PKG_CONFIG --libs "liburing" 2>/dev/null`
+ test "x$?" != "x0" && pkg_failed=yes
+else
+ pkg_failed=yes
+fi
+ else
+ pkg_failed=untried
+fi
+
+
+
+if test $pkg_failed = yes; then
+ { $as_echo "$as_me:${as_lineno-$LINENO}: result: no" >&5
+$as_echo "no" >&6; }
+
+if $PKG_CONFIG --atleast-pkgconfig-version 0.20; then
+ _pkg_short_errors_supported=yes
+else
+ _pkg_short_errors_supported=no
+fi
+ if test $_pkg_short_errors_supported = yes; then
+ LIBURING_PKG_ERRORS=`$PKG_CONFIG --short-errors --print-errors --cflags --libs "liburing" 2>&1`
+ else
+ LIBURING_PKG_ERRORS=`$PKG_CONFIG --print-errors --cflags --libs "liburing" 2>&1`
+ fi
+ # Put the nasty error message in config.log where it belongs
+ echo "$LIBURING_PKG_ERRORS" >&5
+
+ as_fn_error $? "Package requirements (liburing) were not met:
+
+$LIBURING_PKG_ERRORS
+
+Consider adjusting the PKG_CONFIG_PATH environment variable if you
+installed software in a non-standard prefix.
+
+Alternatively, you may set the environment variables LIBURING_CFLAGS
+and LIBURING_LIBS to avoid the need to call pkg-config.
+See the pkg-config man page for more details." "$LINENO" 5
+elif test $pkg_failed = untried; then
+ { $as_echo "$as_me:${as_lineno-$LINENO}: result: no" >&5
+$as_echo "no" >&6; }
+ { { $as_echo "$as_me:${as_lineno-$LINENO}: error: in \`$ac_pwd':" >&5
+$as_echo "$as_me: error: in \`$ac_pwd':" >&2;}
+as_fn_error $? "The pkg-config script could not be found or is too old. Make sure it
+is in your PATH or set the PKG_CONFIG environment variable to the full
+path to pkg-config.
+
+Alternatively, you may set the environment variables LIBURING_CFLAGS
+and LIBURING_LIBS to avoid the need to call pkg-config.
+See the pkg-config man page for more details.
+
+To get pkg-config, see <http://pkg-config.freedesktop.org/>.
+See \`config.log' for more details" "$LINENO" 5; }
+else
+ LIBURING_CFLAGS=$pkg_cv_LIBURING_CFLAGS
+ LIBURING_LIBS=$pkg_cv_LIBURING_LIBS
+ { $as_echo "$as_me:${as_lineno-$LINENO}: result: yes" >&5
+$as_echo "yes" >&6; }
+
+fi
+fi
##
## Header files
diff --git a/src/Makefile.global.in b/src/Makefile.global.in
index eac3d001211..60393ed8fa4 100644
--- a/src/Makefile.global.in
+++ b/src/Makefile.global.in
@@ -190,6 +190,7 @@ with_systemd = @with_systemd@
with_gssapi = @with_gssapi@
with_krb_srvnam = @with_krb_srvnam@
with_ldap = @with_ldap@
+with_liburing = @with_liburing@
with_libxml = @with_libxml@
with_libxslt = @with_libxslt@
with_llvm = @with_llvm@
@@ -216,6 +217,9 @@ krb_srvtab = @krb_srvtab@
ICU_CFLAGS = @ICU_CFLAGS@
ICU_LIBS = @ICU_LIBS@
+LIBURING_CFLAGS = @LIBURING_CFLAGS@
+LIBURING_LIBS = @LIBURING_LIBS@
+
TCLSH = @TCLSH@
TCL_LIBS = @TCL_LIBS@
TCL_LIB_SPEC = @TCL_LIB_SPEC@
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0008-aio-Add-io_uring-method.patch (14.1K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/9-v2-0008-aio-Add-io_uring-method.patch)
download | inline diff:
From de57cec96e81a1867a9f1db4c44243cdc0072b20 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 4 Sep 2024 16:15:17 -0400
Subject: [PATCH v2 08/20] aio: Add io_uring method
---
src/include/storage/aio.h | 1 +
src/include/storage/aio_internal.h | 3 +
src/include/storage/lwlock.h | 1 +
src/backend/storage/aio/Makefile | 1 +
src/backend/storage/aio/aio.c | 6 +
src/backend/storage/aio/meson.build | 1 +
src/backend/storage/aio/method_io_uring.c | 386 ++++++++++++++++++++++
src/backend/storage/lmgr/lwlock.c | 1 +
src/tools/pgindent/typedefs.list | 1 +
9 files changed, 401 insertions(+)
create mode 100644 src/backend/storage/aio/method_io_uring.c
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index 2e84abfea2d..a1633a0ed3d 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -323,6 +323,7 @@ typedef enum IoMethod
{
IOMETHOD_SYNC = 0,
IOMETHOD_WORKER,
+ IOMETHOD_IO_URING,
} IoMethod;
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
index f974c4accf5..d2dc1516bdf 100644
--- a/src/include/storage/aio_internal.h
+++ b/src/include/storage/aio_internal.h
@@ -235,6 +235,9 @@ extern const char *pgaio_io_get_state_name(PgAioHandle *ioh);
/* Declarations for the tables of function pointers exposed by each IO method. */
extern const IoMethodOps pgaio_sync_ops;
extern const IoMethodOps pgaio_worker_ops;
+#ifdef USE_LIBURING
+extern const IoMethodOps pgaio_uring_ops;
+#endif
extern const IoMethodOps *pgaio_impl;
extern PgAioCtl *aio_ctl;
diff --git a/src/include/storage/lwlock.h b/src/include/storage/lwlock.h
index eabf813ce05..72f928b7602 100644
--- a/src/include/storage/lwlock.h
+++ b/src/include/storage/lwlock.h
@@ -217,6 +217,7 @@ typedef enum BuiltinTrancheIds
LWTRANCHE_SUBTRANS_SLRU,
LWTRANCHE_XACT_SLRU,
LWTRANCHE_PARALLEL_VACUUM_DSA,
+ LWTRANCHE_AIO_URING_COMPLETION,
LWTRANCHE_FIRST_USER_DEFINED,
} BuiltinTrancheIds;
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index fa2a7e9e5df..3bcb8a0b2ed 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -13,6 +13,7 @@ OBJS = \
aio_init.o \
aio_io.o \
aio_subject.o \
+ method_io_uring.o \
method_sync.o \
method_worker.o \
read_stream.o
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index e4c9d439ddd..701f06287d9 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -58,6 +58,9 @@ static PgAioHandle *pgaio_io_from_ref(PgAioHandleRef *ior, uint64 *ref_generatio
const struct config_enum_entry io_method_options[] = {
{"sync", IOMETHOD_SYNC, false},
{"worker", IOMETHOD_WORKER, false},
+#ifdef USE_LIBURING
+ {"io_uring", IOMETHOD_IO_URING, false},
+#endif
{NULL, 0, false}
};
@@ -75,6 +78,9 @@ PgAioPerBackend *my_aio;
static const IoMethodOps *pgaio_ops_table[] = {
[IOMETHOD_SYNC] = &pgaio_sync_ops,
[IOMETHOD_WORKER] = &pgaio_worker_ops,
+#ifdef USE_LIBURING
+ [IOMETHOD_IO_URING] = &pgaio_uring_ops,
+#endif
};
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index 62738ce1d14..537f23d446d 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -5,6 +5,7 @@ backend_sources += files(
'aio_init.c',
'aio_io.c',
'aio_subject.c',
+ 'method_io_uring.c',
'method_sync.c',
'method_worker.c',
'read_stream.c',
diff --git a/src/backend/storage/aio/method_io_uring.c b/src/backend/storage/aio/method_io_uring.c
new file mode 100644
index 00000000000..3f214e42767
--- /dev/null
+++ b/src/backend/storage/aio/method_io_uring.c
@@ -0,0 +1,386 @@
+/*-------------------------------------------------------------------------
+ *
+ * method_io_uring.c
+ * AIO - perform AIO using Linux' io_uring
+ *
+ * XXX Write me
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/method_io_uring.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#ifdef USE_LIBURING
+
+#include <liburing.h>
+
+#include "pgstat.h"
+#include "port/pg_iovec.h"
+#include "storage/aio_internal.h"
+#include "storage/fd.h"
+#include "storage/proc.h"
+#include "storage/shmem.h"
+
+
+/* Entry points for IoMethodOps. */
+static size_t pgaio_uring_shmem_size(void);
+static void pgaio_uring_shmem_init(bool first_time);
+static void pgaio_uring_init_backend(void);
+
+static int pgaio_uring_submit(uint16 num_staged_ios, PgAioHandle **staged_ios);
+static void pgaio_uring_wait_one(PgAioHandle *ioh, uint64 ref_generation);
+
+static void pgaio_uring_sq_from_io(PgAioHandle *ioh, struct io_uring_sqe *sqe);
+
+
+const IoMethodOps pgaio_uring_ops = {
+ .shmem_size = pgaio_uring_shmem_size,
+ .shmem_init = pgaio_uring_shmem_init,
+ .init_backend = pgaio_uring_init_backend,
+
+ .submit = pgaio_uring_submit,
+ .wait_one = pgaio_uring_wait_one,
+};
+
+typedef struct PgAioUringContext
+{
+ LWLock completion_lock;
+
+ struct io_uring io_uring_ring;
+ /* XXX: probably worth padding to a cacheline boundary here */
+} PgAioUringContext;
+
+
+static PgAioUringContext *aio_uring_contexts;
+static PgAioUringContext *my_shared_uring_context;
+
+/* io_uring local state */
+static struct io_uring local_ring;
+
+
+
+static Size
+AioContextShmemSize(void)
+{
+ uint32 TotalProcs = MaxBackends + NUM_AUXILIARY_PROCS - MAX_IO_WORKERS;
+
+ return mul_size(TotalProcs, sizeof(PgAioUringContext));
+}
+
+static size_t
+pgaio_uring_shmem_size(void)
+{
+ return AioContextShmemSize();
+}
+
+static void
+pgaio_uring_shmem_init(bool first_time)
+{
+ uint32 TotalProcs = MaxBackends + NUM_AUXILIARY_PROCS - MAX_IO_WORKERS;
+ bool found;
+
+ aio_uring_contexts = (PgAioUringContext *)
+ ShmemInitStruct("AioUring", pgaio_uring_shmem_size(), &found);
+
+ if (found)
+ return;
+
+ for (int contextno = 0; contextno < TotalProcs; contextno++)
+ {
+ PgAioUringContext *context = &aio_uring_contexts[contextno];
+ int ret;
+
+ /*
+ * XXX: Probably worth sharing the WQ between the different rings,
+ * when supported by the kernel. Could also cause additional
+ * contention, I guess?
+ */
+#if 0
+ if (!AcquireExternalFD())
+ elog(ERROR, "No external FD available");
+#endif
+ ret = io_uring_queue_init(io_max_concurrency, &context->io_uring_ring, 0);
+ if (ret < 0)
+ elog(ERROR, "io_uring_queue_init failed: %s", strerror(-ret));
+
+ LWLockInitialize(&context->completion_lock, LWTRANCHE_AIO_URING_COMPLETION);
+ }
+}
+
+static void
+pgaio_uring_init_backend(void)
+{
+ int ret;
+
+ my_shared_uring_context = &aio_uring_contexts[MyProcNumber];
+
+ ret = io_uring_queue_init(32, &local_ring, 0);
+ if (ret < 0)
+ elog(ERROR, "io_uring_queue_init failed: %s", strerror(-ret));
+}
+
+static int
+pgaio_uring_submit(uint16 num_staged_ios, PgAioHandle **staged_ios)
+{
+ struct io_uring *uring_instance = &my_shared_uring_context->io_uring_ring;
+ int in_flight_before = dclist_count(&my_aio->in_flight_ios);
+
+ Assert(num_staged_ios <= PGAIO_SUBMIT_BATCH_SIZE);
+
+ for (int i = 0; i < num_staged_ios; i++)
+ {
+ PgAioHandle *ioh = staged_ios[i];
+ struct io_uring_sqe *sqe;
+
+ sqe = io_uring_get_sqe(uring_instance);
+
+ if (!sqe)
+ elog(ERROR, "io_uring submission queue is unexpectedly full");
+
+ pgaio_io_prepare_submit(ioh);
+ pgaio_uring_sq_from_io(ioh, sqe);
+
+ /*
+ * io_uring executes IO in process context if possible. That's
+ * generally good, as it reduces context switching. When performing a
+ * lot of buffered IO that means that copying between page cache and
+ * userspace memory happens in the foreground, as it can't be
+ * offloaded to DMA hardware as is possible when using direct IO. When
+ * executing a lot of buffered IO this causes io_uring to be slower
+ * than worker mode, as worker mode parallelizes the copying.
+ * io_uring can be told to offload work to worker threads instead.
+ *
+ * If an IO is buffered IO and we already have IOs in flight or
+ * multiple IOs are being submitted, we thus tell io_uring to execute
+ * the IO in the background. We don't do so for the first few IOs
+ * being submitted as executing in this process' context has lower
+ * latency.
+ */
+ if (in_flight_before > 4 && (ioh->flags & AHF_BUFFERED))
+ io_uring_sqe_set_flags(sqe, IOSQE_ASYNC);
+
+ in_flight_before++;
+ }
+
+ while (true)
+ {
+ int ret;
+
+ pgstat_report_wait_start(WAIT_EVENT_AIO_SUBMIT);
+ ret = io_uring_submit(uring_instance);
+ pgstat_report_wait_end();
+
+ if (ret == -EINTR)
+ {
+ elog(DEBUG3, "submit EINTR, nios: %d", num_staged_ios);
+ continue;
+ }
+ if (ret < 0)
+ elog(PANIC, "failed: %d/%s",
+ ret, strerror(-ret));
+ else if (ret != num_staged_ios)
+ {
+ /* likely unreachable, but if it is, we would need to re-submit */
+ elog(PANIC, "submitted only %d of %d",
+ ret, num_staged_ios);
+ }
+ else
+ {
+ elog(DEBUG3, "submit nios: %d", num_staged_ios);
+ }
+ break;
+ }
+
+ return num_staged_ios;
+}
+
+
+#define PGAIO_MAX_LOCAL_REAPED 16
+
+static void
+pgaio_uring_drain_locked(PgAioUringContext *context)
+{
+ int ready;
+ int orig_ready;
+
+ /*
+ * Don't drain more events than available right now. Otherwise it's
+ * plausible that one backend could get stuck, for a while, receiving CQEs
+ * without actually processing them.
+ */
+ orig_ready = ready = io_uring_cq_ready(&context->io_uring_ring);
+
+ while (ready > 0)
+ {
+ struct io_uring_cqe *reaped_cqes[PGAIO_MAX_LOCAL_REAPED];
+ uint32 reaped;
+
+ START_CRIT_SECTION();
+ reaped =
+ io_uring_peek_batch_cqe(&context->io_uring_ring,
+ reaped_cqes,
+ Min(PGAIO_MAX_LOCAL_REAPED, ready));
+ Assert(reaped <= ready);
+
+ ready -= reaped;
+
+ for (int i = 0; i < reaped; i++)
+ {
+ struct io_uring_cqe *cqe = reaped_cqes[i];
+ PgAioHandle *ioh;
+
+ ioh = io_uring_cqe_get_data(cqe);
+ io_uring_cqe_seen(&context->io_uring_ring, cqe);
+
+ pgaio_io_process_completion(ioh, cqe->res);
+ }
+
+ END_CRIT_SECTION();
+
+ ereport(DEBUG3,
+ errmsg("drained %d/%d, now expecting %d",
+ reaped, orig_ready, io_uring_cq_ready(&context->io_uring_ring)),
+ errhidestmt(true),
+ errhidecontext(true));
+
+ }
+}
+
+static void
+pgaio_uring_wait_one(PgAioHandle *ioh, uint64 ref_generation)
+{
+ PgAioHandleState state;
+ ProcNumber owner_procno = ioh->owner_procno;
+ PgAioUringContext *owner_context = &aio_uring_contexts[owner_procno];
+ bool expect_cqe;
+ int waited = 0;
+
+ /*
+ * We ought to have a smarter locking scheme, nearly all the time the
+ * backend owning the ring will reap the completions, making the locking
+ * unnecessarily expensive.
+ */
+ LWLockAcquire(&owner_context->completion_lock, LW_EXCLUSIVE);
+
+ while (true)
+ {
+ ereport(DEBUG3,
+ errmsg("wait_one for io:%d io_gen: %llu, ref_gen: %llu, in state %s, cycle %d",
+ pgaio_io_get_id(ioh),
+ (long long unsigned) ref_generation,
+ (long long unsigned) ioh->generation,
+ pgaio_io_get_state_name(ioh), waited),
+ errhidestmt(true),
+ errhidecontext(true));
+
+ if (pgaio_io_was_recycled(ioh, ref_generation, &state) ||
+ state != AHS_IN_FLIGHT)
+ {
+ break;
+ }
+ else if (io_uring_cq_ready(&owner_context->io_uring_ring))
+ {
+ expect_cqe = true;
+ }
+ else
+ {
+ int ret;
+ struct io_uring_cqe *cqes;
+
+ pgstat_report_wait_start(WAIT_EVENT_AIO_DRAIN);
+ ret = io_uring_wait_cqes(&owner_context->io_uring_ring, &cqes, 1, NULL, NULL);
+ pgstat_report_wait_end();
+
+ if (ret == -EINTR)
+ {
+ continue;
+ }
+ else if (ret != 0)
+ {
+ elog(PANIC, "unexpected: %d/%s: %m", ret, strerror(-ret));
+ }
+ else
+ {
+ Assert(cqes != NULL);
+ expect_cqe = true;
+ waited++;
+ }
+ }
+
+ if (expect_cqe)
+ {
+ pgaio_uring_drain_locked(owner_context);
+ }
+ }
+
+ LWLockRelease(&owner_context->completion_lock);
+
+ ereport(DEBUG3,
+ errmsg("wait_one with %d sleeps",
+ waited),
+ errhidestmt(true),
+ errhidecontext(true));
+}
+
+static void
+pgaio_uring_sq_from_io(PgAioHandle *ioh, struct io_uring_sqe *sqe)
+{
+ struct iovec *iov;
+
+ switch (ioh->op)
+ {
+ case PGAIO_OP_READV:
+ iov = &aio_ctl->iovecs[ioh->iovec_off];
+ if (ioh->op_data.read.iov_length == 1)
+ {
+ io_uring_prep_read(sqe,
+ ioh->op_data.read.fd,
+ iov->iov_base,
+ iov->iov_len,
+ ioh->op_data.read.offset);
+ }
+ else
+ {
+ io_uring_prep_readv(sqe,
+ ioh->op_data.read.fd,
+ iov,
+ ioh->op_data.read.iov_length,
+ ioh->op_data.read.offset);
+
+ }
+ break;
+
+ case PGAIO_OP_WRITEV:
+ iov = &aio_ctl->iovecs[ioh->iovec_off];
+ if (ioh->op_data.write.iov_length == 1)
+ {
+ io_uring_prep_write(sqe,
+ ioh->op_data.write.fd,
+ iov->iov_base,
+ iov->iov_len,
+ ioh->op_data.write.offset);
+ }
+ else
+ {
+ io_uring_prep_writev(sqe,
+ ioh->op_data.write.fd,
+ iov,
+ ioh->op_data.write.iov_length,
+ ioh->op_data.write.offset);
+ }
+ break;
+
+ case PGAIO_OP_INVALID:
+ elog(ERROR, "trying to prepare invalid IO operation for execution");
+ }
+
+ io_uring_sqe_set_data(sqe, ioh);
+}
+
+#endif /* USE_LIBURING */
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index bc459dc5d2b..4fdcfb1df1b 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -166,6 +166,7 @@ static const char *const BuiltinTrancheNames[] = {
[LWTRANCHE_SUBTRANS_SLRU] = "SubtransSLRU",
[LWTRANCHE_XACT_SLRU] = "XactSLRU",
[LWTRANCHE_PARALLEL_VACUUM_DSA] = "ParallelVacuumDSA",
+ [LWTRANCHE_AIO_URING_COMPLETION] = "AioUringCompletion",
};
StaticAssertDecl(lengthof(BuiltinTrancheNames) ==
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 9b9c8f0d1fc..a5b12b48f99 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2121,6 +2121,7 @@ PgAioReturn
PgAioSubjectData
PgAioSubjectID
PgAioSubjectInfo
+PgAioUringContext
PgArchData
PgBackendGSSStatus
PgBackendSSLStatus
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0009-aio-Add-README.md-explaining-higher-level-design.patch (18.0K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/10-v2-0009-aio-Add-README.md-explaining-higher-level-design.patch)
download | inline diff:
From c95ba2c47ddc454f19703c4361f47690ff8ff05e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Fri, 6 Sep 2024 15:27:57 -0400
Subject: [PATCH v2 09/20] aio: Add README.md explaining higher level design
---
src/backend/storage/aio/README.md | 413 ++++++++++++++++++++++++++++++
src/backend/storage/aio/aio.c | 2 +
2 files changed, 415 insertions(+)
create mode 100644 src/backend/storage/aio/README.md
diff --git a/src/backend/storage/aio/README.md b/src/backend/storage/aio/README.md
new file mode 100644
index 00000000000..893f4ffe428
--- /dev/null
+++ b/src/backend/storage/aio/README.md
@@ -0,0 +1,413 @@
+# Asynchronous & Direct IO
+
+## AIO Usage Example
+
+In many cases code that can benefit from AIO does not directly have to
+interact with the AIO interface, but can use AIO via higher-level
+abstractions. See [Helpers](#helpers).
+
+In this example, a buffer will be read into shared buffers.
+
+```C
+/*
+ * Result of the operation, only to be accessed in this backend.
+ */
+PgAioReturn ioret;
+
+/*
+ * Acquire AIO Handle, ioret will get result upon completion.
+ */
+PgAioHandle *ioh = pgaio_io_get(CurrentResourceOwner, &ioret);
+
+/*
+ * Reference that can be used to wait for the IO we initiate below. This
+ * reference can reside in local or shared memory and waited upon by any
+ * process. An arbitrary number of references can be made for each IO.
+ */
+PgAioRef ior;
+
+pgaio_io_get_ref(ioh, &ior);
+
+/*
+ * Arrange for shared buffer completion callbacks to be called upon completion
+ * of the IO. This callback will update the buffer descriptors associated with
+ * the AioHandle, which e.g. allows other backends to access the buffer.
+ *
+ * Multiple completion callbacks can be registered for each handle.
+ */
+pgaio_io_add_shared_cb(ioh, ASC_SHARED_BUFFER_READ);
+
+/*
+ * The completion callback needs to know which buffers to update when the IO
+ * completes. As the AIO subsystem does not know about buffers, we have to
+ * associate this information with the AioHandle, for use by the completion
+ * callback registered above.
+ */
+pgaio_io_set_io_data_32(ioh, (uint32 *) buffer, 1);
+
+/*
+ * Hand AIO handle to lower-level function. When operating on the level of
+ * buffers, we don't know how exactly the IO is performed, that is the
+ * responsibility of the storage manager implementation.
+ *
+ * E.g. md.c needs to translate block numbers into offsets in segments.
+ *
+ * Once the IO handle has been handed of, it may not further be used, as the
+ * IO may immediately get executed below smgrstartreadv() and the handle reused
+ * for another IO.
+ */
+smgrstartreadv(ioh, operation->smgr, forknum, blkno,
+ BufferGetBlock(buffer), 1);
+
+/*
+ * As mentioned above, the IO might be initiated within smgrstartreadv(). That
+ * is however not guaranteed, to allow IO submission to be batched.
+ *
+ * Note that one needs to be careful while there may be unsubmitted IOs, as
+ * another backend may need to wait for one of the unsubmitted IOs. If this
+ * backend were to wait for the other backend, we'd have a deadlock. To avoid
+ * that, pending IOs need to be explicitly submitted before this backend
+ * might be blocked by a backend waiting for IO.
+ *
+ * Note that the IO might have immediately been submitted (e.g. due to reaching
+ * a limit on the number of unsubmitted IOs) and even completed during the
+ * smgrstartreadv() above.
+ *
+ * Once submitted, the IO is in-flight and can complete at any time.
+ */
+pgaio_submit_staged();
+
+/*
+ * To benefit from AIO, it is beneficial to perform other work, including
+ * submitting other IOs, before waiting for the IO to complete. Otherwise
+ * we could just have used synchronous, blocking IO.
+ */
+perform_other_work();
+
+/*
+ * We did some other work and now need the IO operation to have completed to
+ * continue.
+ */
+pgaio_io_ref_wait(&ior);
+
+/*
+ * At this point the IO has completed. We do not yet know whether it succeeded
+ * or failed, however. The buffer's state has been updated, which allows other
+ * backends to use the buffer (if the IO succeeded), or retry the IO (if it
+ * failed).
+ *
+ * Note that in case the IO has failed, a LOG message may have been emitted,
+ * but no ERROR has been raised. This is crucial, as another backend waiting
+ * for this IO should not see an ERROR.
+ *
+ * To check whether the operation succeeded, and to raise an ERROR, or if more
+ * appropriate LOG, the PgAioReturn we passed to pgaio_io_get() is used.
+ */
+if (ioret.result.status == ARS_ERROR)
+ pgaio_result_log(aio_ret.result, &aio_ret.subject_data, ERROR);
+
+/*
+ * Besides having succeeded completely, the IO could also have partially
+ * completed. If we e.g. tried to read many blocks at once, the read might have
+ * only succeeded for the first few blocks.
+ *
+ * If the IO partially succeeded and this backend needs all blocks to have
+ * completed, this backend needs to reissue the IO for the remaining buffers.
+ * The AIO subsystem cannot handle this retry transparently.
+ *
+ * As this example is already long, and we only read a single block, we'll just
+ * error out if there's a partial read.
+ */
+if (ioret.result.status == ARS_PARTIAL)
+ pgaio_result_log(aio_ret.result, &aio_ret.subject_data, ERROR);
+
+/*
+ * The IO succeeded, so we can use the buffer now.
+ */
+```
+
+
+## Design Criteria & Motivation
+
+### Why Asynchronous IO
+
+Until the introduction of asynchronous IO Postgres relied on the operating
+system to hide the cost of synchronous IO from Postgres. While this worked
+surprisingly well in a lot of workloads, it does not do as good a job on
+prefetching and controlled writeback as we would like.
+
+There are important expensive operations like `fdatasync()` where the operating
+system cannot hide the storage latency. This is particularly important for WAL
+writes, where the ability to asynchronously issue `fdatasync()` or O_DSYNC
+writes can yield significantly higher throughput.
+
+
+### Why Direct / unbuffered IO
+
+The main reason to want to use Direct IO are:
+
+- Lower CPU usage / higher throughput. Particularly on modern storage buffered
+ writes are bottlenecked by the operating system having to copy data from the
+ kernel's page cache to postgres buffer pool using the CPU. Whereas direct IO
+ can often move the data directly between the storage devices and postgres'
+ buffer cache, using DMA. While that transfer is ongoing, the CPU is free to
+ perform other work.
+- Reduced latency - Direct IO can have substantially lower latency than
+ buffered IO, which can be impactful for OLTP workloads bottlenecked by WAL
+ write latency.
+- Avoiding double buffering between operating system cache and postgres'
+ shared_buffers.
+- Better control over the timing and pace of dirty data writeback.
+
+
+The main reason *not* to use Direct IO are:
+
+- Without AIO, Direct IO is unusably slow for most purposes.
+- Even with AIO, many parts of postgres need to be modified to perform
+ explicit prefetching.
+- In situations where shared_buffers cannot be set appropriately large,
+ e.g. because there are many different postgres instances hosted on shared
+ hardware, performance will often be worse then when using buffered IO.
+
+
+### Deadlock and Starvation Dangers due to AIO
+
+Using AIO in a naive way can easily lead to deadlocks in an environment where
+the source/target of AIO are shared resources, like pages in postgres'
+shared_buffers.
+
+Consider one backend performing readahead on a table, initiating IO for a
+number of buffers ahead of the current "scan position". If that backend then
+performs some operation that blocks, or even just is slow, the IO completion
+for the asynchronously initiated read may not be processed.
+
+This AIO implementation solves this problem by requiring that AIO methods
+either allow AIO completions to be processed by any backend in the system
+(e.g. io_uring), or to guarantee that AIO processing will happen even when the
+issuing backend is blocked (e.g. worker mode, which offloads completion
+processing to the AIO workers).
+
+
+### IO can be started in critical sections
+
+Using AIO for WAL writes can reduce the overhead of WAL logging substantially:
+
+- AIO allows to start WAL writes eagerly, so they complete before needing to
+ wait
+- AIO allows to have multiple WAL flushes in progress at the same time
+- AIO makes it more realistic to use O\_DIRECT + O\_DSYNC, which can reduce
+ the number of roundtrips to storage on some OSs and storage HW (buffered IO
+ and direct IO without O_DSYNC needs to issue a write and after the writes
+ completion a cache cache flush, whereas O\_DIRECT + O\_DSYNC can use a
+ single FUA write).
+
+The need to be able to execute IO in critical sections has substantial design
+implication on the AIO subsystem. Mainly because completing IOs (see prior
+section) needs to be possible within a critical section, even if the
+to-be-completed IO itself was not issued in a critical section. Consider
+e.g. the case of a backend first starting a number of writes from shared
+buffers and then starting to flush the WAL. Because only a limited amount of
+IO can be in-progress at the same time, initiating the IO for flushing the WAL
+may require to first finish executing IO executed earlier.
+
+
+### State for AIO needs to live in shared memory
+
+Because postgres uses a process model and because AIOs need to be
+complete-able by any backend much of the state of the AIO subsystem needs to
+live in shared memory.
+
+In an `EXEC_BACKEND` build backends executable code and other process local
+state is not necessarily mapped to the same addresses in each process due to
+ASLR. This means that the shared memory cannot contain pointer to callbacks.
+
+
+## Design of the AIO Subsystem
+
+
+### AIO Methods
+
+To achieve portability and performance, multiple methods of performing AIO are
+implemented and others are likely worth adding in the future.
+
+
+#### Synchronous Mode
+
+`io_method=sync` does not actually perform AIO but allows to use the AIO API
+while performing synchronous IO. This can be useful for debugging. The code
+for the synchronous mode is also used as a fallback by e.g. the [worker
+mode](#worker) uses it to execute IO that cannot be executed by workers.
+
+
+#### Worker
+
+`io_method=worker` is available on every platform postgres runs on, and
+implements asynchronous IO - from the view of the issuing process - by
+dispatching the IO to one of several worker processes performing the IO in a
+synchronous manner.
+
+
+#### io_uring
+
+`io_method=io_uring` is available on Linux 5.1+. In contrast to worker mode it
+dispatches all IO from within the process, lowering context switch rate /
+latency.
+
+
+### AIO Handles
+
+The central API piece for postgres' AIO abstraction are AIO handles. To
+execute an IO one first has to acquire an IO handle (`pgaio_io_get()`) and
+then "defined", i.e. associate an IO operation with the handle.
+
+Often AIO handles are acquired on a higher level and then passed to a lower
+level to be fully defined. E.g., for IO to/from shared buffers, bufmgr.c
+routines acquire the handle, which is then passed through smgr.c, md.c to be
+finally fully defined in fd.c.
+
+The functions used at the lowest level to define the operation are
+`pgaio_io_prep_*()`.
+
+Because acquisition of an IO handle
+[must always succeed](#io-can-be-started-in-critical-sections)
+and the number of AIO Handles
+[has to be limited](#state-for-aio-needs-to-live-in-shared-memory)
+AIO handles can be reused as soon as they have completed. Obviously code needs
+to be able to react to IO completion. Shared state can be updated using
+[AIO Completion callbacks](#aio-callbacks)
+and the issuing backend can provide a backend local variable to receive the
+result of the IO, as described in
+[AIO Result](#aio-results)
+. An IO can be waited for, by both the issuing and any other backend, using
+[AIO References](#aio-references).
+
+
+Because an AIO Handle is not executable just after calling `pgaio_io_get()`
+and because `pgaio_io_get()` needs to be able to succeed, only a single AIO
+Handle may be acquired (i.e. returned by `pgaio_io_get()`) without causing the
+IO to have been defined (by, potentially indirectly, causing
+`pgaio_io_prep_*()` to have been called). Otherwise a backend could trivially
+self-deadlock by using up all AIO Handles without the ability to wait for some
+of the IOs to complete.
+
+If it turns out that an AIO Handle is not needed, e.g., because the handle was
+acquired before holding a contended lock, it can be released without being
+defined using `pgaio_io_release()`.
+
+
+### AIO Callbacks
+
+Commonly several layers need to react to completion of an IO. E.g. for a read
+md.c needs to check if the IO outright failed or was shorter than needed,
+bufmgr.c needs to verify the page looks valid and bufmgr.c needs to update the
+BufferDesc to update the buffer's state.
+
+The fact that several layers / subsystems need to react to IO completion poses
+a few challenges:
+
+- Upper layers should not need to know details of lower layers. E.g. bufmgr.c
+ should not assume the IO will pass through md.c. Therefore upper levels
+ cannot know what lower layers would consider an error.
+
+- Lower layers should not need to know about upper layers. E.g. smgr APIs are
+ used going through shared buffers but are also used bypassing shared
+ buffers. This means that e.g. md.c is not in a position to validate
+ checksums.
+
+- Having code in the AIO subsystem for every possible combination of layers
+ would lead to a lot of duplication.
+
+The "solution" to this the ability to associate multiple completion callbacks
+with a handle. E.g. bufmgr.c can have a callback to update the BufferDesc
+state and to verify the page and md.c. another callback to check if the IO
+operation was successful.
+
+As [mentioned](#state-for-aio-needs-to-live-in-shared-memory), shared memory
+currently cannot contain function pointers. Because of that completion
+callbacks are not directly identified by function pointers but by IDs
+(`PgAioHandleSharedCallbackID`). A substantial added benefit is that that
+allows callbacks to be identified by much smaller amount of memory (a single
+byte currently).
+
+In addition to completion, AIO callbacks also are called to "prepare" an
+IO. This is, e.g., used to acquire buffer pins owned by the AIO subsystem for
+IO to/from shared buffers, which is required to handle the case where the
+issuing backend errors out and releases its own pins.
+
+As [explained earlier](#io-can-be-started-in-critical-sections) IO completions
+need to be safe to execute in critical sections. To allow the backend that
+issued the IO to error out in case of failure [AIO Result](#aio-results) can
+be used.
+
+
+### AIO Subjects
+
+In addition to the completion callbacks describe above, each AIO Handle has
+exactly one "subject". Each subject has some space inside an AIO Handle with
+information specific to the subject and can provide callbacks to allow to
+reopen the underlying file (required for worker mode) and to describe the IO
+operation (used for debug logging and error messages).
+
+
+### AIO References
+
+As [described above](#aio-handles) can be reused immediately after completion
+and therefore cannot be used to wait for completion of the IO. Waiting is
+enabled using AIO references, which do not just identify an AIO Handle but
+also include the handles "generation".
+
+A reference to an AIO Handle can be acquired using `pgaio_io_get_ref()` and
+then waited upon using `pgaio_io_ref_wait()`.
+
+
+### AIO Results
+
+As AIO completion callbacks
+[are executed in critical sections](#io-can-be-started-in-critical-sections)
+and [may be executed by any backend](#deadlock-and-starvation-dangers-due-to-aio)
+completion callbacks cannot be used to, e.g., make the query that triggered an
+IO ERROR out.
+
+To allow to react to failing IOs the issuing backend can pass a pointer to a
+`PgAioReturn` in backend local memory. Before an AIO Handle is reused the
+`PgAioReturn` is filled with information about the IO. This includes
+information about whether the IO was successful (as a value of
+`PgAioResultStatus`) and enough information to raise an error in case of a
+failure (via `pgaio_result_log()`, with the error details encoded in
+`PgAioResult`).
+
+XXX: "return" vs "result" vs "result status" seems quite confusing. The naming
+should be improved.
+
+
+### AIO Errors
+
+It would be very convenient to have shared completion callbacks encode the
+details of errors as an `ErrorData` that could be raised at a later
+time. Unfortunately doing so would require allocating memory. While elog.c can
+guarantee (well, kinda) that logging a message will not run out of memory,
+that only works because a very limited number of messages are in the process
+of being logged. With AIO a large number of concurrently issued AIOs might
+fail.
+
+To avoid the need for preallocating a potentially large amount of memory (in
+shared memory no less!), completion callbacks instead have to encode errors in
+a more compact format that can be converted into an error message.
+
+
+## Helpers
+
+Using the low-level AIO API introduces too much complexity to do so all over
+the tree. Most uses of AIO should be done via reusable, higher-level,
+helpers.
+
+
+### Read Stream
+
+A common and very beneficial use of AIO are reads where a substantial number
+of to-be-read locations are known ahead of time. E.g., for a sequential scan
+the set of blocks that need to be read can be determined solely by knowing the
+current position and checking the buffer mapping table.
+
+The [Read Stream](../../../include/storage/read_stream.h) interface makes it
+comparatively easy to use AIO for such use cases.
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index 701f06287d9..2439ce3740d 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -24,6 +24,8 @@
* - read_stream.c - helper for accessing buffered relation data with
* look-ahead
*
+ * - README.md - higher-level overview over AIO
+ *
*
* Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
* Portions Copyright (c) 1994, Regents of the University of California
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0010-aio-Implement-smgr-md.c-aio-methods.patch (25.6K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/11-v2-0010-aio-Implement-smgr-md.c-aio-methods.patch)
download | inline diff:
From 45154f1e08ee325875673c14470479f019ef0461 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sun, 15 Dec 2024 12:36:32 -0500
Subject: [PATCH v2 10/20] aio: Implement smgr/md.c aio methods
---
src/include/storage/aio.h | 17 +-
src/include/storage/fd.h | 6 +
src/include/storage/md.h | 12 +
src/include/storage/smgr.h | 21 ++
src/backend/storage/aio/aio_subject.c | 4 +
src/backend/storage/file/fd.c | 68 ++++++
src/backend/storage/smgr/md.c | 314 ++++++++++++++++++++++++++
src/backend/storage/smgr/smgr.c | 91 ++++++++
8 files changed, 532 insertions(+), 1 deletion(-)
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index a1633a0ed3d..d693b0b0bd8 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -55,9 +55,10 @@ typedef enum PgAioSubjectID
{
/* intentionally the zero value, to help catch zeroed memory etc */
ASI_INVALID = 0,
+ ASI_SMGR,
} PgAioSubjectID;
-#define ASI_COUNT (ASI_INVALID + 1)
+#define ASI_COUNT (ASI_SMGR + 1)
/*
* Flags for an IO that can be set with pgaio_io_set_flag().
@@ -100,6 +101,9 @@ typedef enum PgAioHandleFlags
typedef enum PgAioHandleSharedCallbackID
{
ASC_INVALID,
+
+ ASC_MD_READV,
+ ASC_MD_WRITEV,
} PgAioHandleSharedCallbackID;
@@ -135,6 +139,17 @@ typedef union
typedef union PgAioSubjectData
{
+ struct
+ {
+ RelFileLocator rlocator; /* physical relation identifier */
+ BlockNumber blockNum; /* blknum relative to begin of reln */
+ int nblocks;
+ ForkNumber forkNum:8; /* don't waste 4 byte for four values */
+ bool is_temp; /* proc can be inferred by owning AIO */
+ bool release_lock;
+ int8 mode;
+ } smgr;
+
/* just as an example placeholder for later */
struct
{
diff --git a/src/include/storage/fd.h b/src/include/storage/fd.h
index 1456ab383a4..e993e1b671f 100644
--- a/src/include/storage/fd.h
+++ b/src/include/storage/fd.h
@@ -101,6 +101,8 @@ extern PGDLLIMPORT int max_safe_fds;
* prototypes for functions in fd.c
*/
+struct PgAioHandle;
+
/* Operations on virtual Files --- equivalent to Unix kernel file ops */
extern File PathNameOpenFile(const char *fileName, int fileFlags);
extern File PathNameOpenFilePerm(const char *fileName, int fileFlags, mode_t fileMode);
@@ -109,6 +111,10 @@ extern void FileClose(File file);
extern int FilePrefetch(File file, off_t offset, off_t amount, uint32 wait_event_info);
extern ssize_t FileReadV(File file, const struct iovec *iov, int iovcnt, off_t offset, uint32 wait_event_info);
extern ssize_t FileWriteV(File file, const struct iovec *iov, int iovcnt, off_t offset, uint32 wait_event_info);
+extern ssize_t FileReadV(File file, const struct iovec *iov, int iovcnt, off_t offset, uint32 wait_event_info);
+extern int FileStartReadV(struct PgAioHandle *ioh, File file, int iovcnt, off_t offset, uint32 wait_event_info);
+extern ssize_t FileWriteV(File file, const struct iovec *iov, int iovcnt, off_t offset, uint32 wait_event_info);
+extern int FileStartWriteV(struct PgAioHandle *ioh, File file, int iovcnt, off_t offset, uint32 wait_event_info);
extern int FileSync(File file, uint32 wait_event_info);
extern int FileZero(File file, off_t offset, off_t amount, uint32 wait_event_info);
extern int FileFallocate(File file, off_t offset, off_t amount, uint32 wait_event_info);
diff --git a/src/include/storage/md.h b/src/include/storage/md.h
index e7671dd6c18..c3a18465c6b 100644
--- a/src/include/storage/md.h
+++ b/src/include/storage/md.h
@@ -19,6 +19,10 @@
#include "storage/smgr.h"
#include "storage/sync.h"
+struct PgAioHandleSharedCallbacks;
+extern const struct PgAioHandleSharedCallbacks aio_md_readv_cb;
+extern const struct PgAioHandleSharedCallbacks aio_md_writev_cb;
+
/* md storage manager functionality */
extern void mdinit(void);
extern void mdopen(SMgrRelation reln);
@@ -36,9 +40,16 @@ extern uint32 mdmaxcombine(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum);
extern void mdreadv(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
void **buffers, BlockNumber nblocks);
+extern void mdstartreadv(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
+ void **buffers, BlockNumber nblocks);
extern void mdwritev(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum,
const void **buffers, BlockNumber nblocks, bool skipFsync);
+extern void mdstartwritev(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum,
+ BlockNumber blocknum,
+ const void **buffers, BlockNumber nblocks, bool skipFsync);
extern void mdwriteback(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum, BlockNumber nblocks);
extern BlockNumber mdnblocks(SMgrRelation reln, ForkNumber forknum);
@@ -46,6 +57,7 @@ extern void mdtruncate(SMgrRelation reln, ForkNumber forknum,
BlockNumber old_blocks, BlockNumber nblocks);
extern void mdimmedsync(SMgrRelation reln, ForkNumber forknum);
extern void mdregistersync(SMgrRelation reln, ForkNumber forknum);
+extern int mdfd(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum, uint32 *off);
extern void ForgetDatabaseSyncRequests(Oid dbid);
extern void DropRelationFiles(RelFileLocator *delrels, int ndelrels, bool isRedo);
diff --git a/src/include/storage/smgr.h b/src/include/storage/smgr.h
index 63a186bd346..fe23a7f744f 100644
--- a/src/include/storage/smgr.h
+++ b/src/include/storage/smgr.h
@@ -73,6 +73,11 @@ typedef SMgrRelationData *SMgrRelation;
#define SmgrIsTemp(smgr) \
RelFileLocatorBackendIsTemp((smgr)->smgr_rlocator)
+struct PgAioHandle;
+struct PgAioSubjectInfo;
+
+extern const struct PgAioSubjectInfo aio_smgr_subject_info;
+
extern void smgrinit(void);
extern SMgrRelation smgropen(RelFileLocator rlocator, ProcNumber backend);
extern bool smgrexists(SMgrRelation reln, ForkNumber forknum);
@@ -97,10 +102,19 @@ extern uint32 smgrmaxcombine(SMgrRelation reln, ForkNumber forknum,
extern void smgrreadv(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum,
void **buffers, BlockNumber nblocks);
+extern void smgrstartreadv(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum,
+ BlockNumber blocknum,
+ void **buffers, BlockNumber nblocks);
extern void smgrwritev(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum,
const void **buffers, BlockNumber nblocks,
bool skipFsync);
+extern void smgrstartwritev(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum,
+ BlockNumber blocknum,
+ const void **buffers, BlockNumber nblocks,
+ bool skipFsync);
extern void smgrwriteback(SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum, BlockNumber nblocks);
extern BlockNumber smgrnblocks(SMgrRelation reln, ForkNumber forknum);
@@ -110,6 +124,7 @@ extern void smgrtruncate(SMgrRelation reln, ForkNumber *forknum, int nforks,
BlockNumber *nblocks);
extern void smgrimmedsync(SMgrRelation reln, ForkNumber forknum);
extern void smgrregistersync(SMgrRelation reln, ForkNumber forknum);
+extern int smgrfd(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum, uint32 *off);
extern void AtEOXact_SMgr(void);
extern bool ProcessBarrierSmgrRelease(void);
@@ -127,4 +142,10 @@ smgrwrite(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
smgrwritev(reln, forknum, blocknum, &buffer, 1, skipFsync);
}
+extern void pgaio_io_set_subject_smgr(struct PgAioHandle *ioh,
+ SMgrRelationData *smgr,
+ ForkNumber forknum,
+ BlockNumber blocknum,
+ int nblocks);
+
#endif /* SMGR_H */
diff --git a/src/backend/storage/aio/aio_subject.c b/src/backend/storage/aio/aio_subject.c
index 8694cfafcd1..effb09c11c7 100644
--- a/src/backend/storage/aio/aio_subject.c
+++ b/src/backend/storage/aio/aio_subject.c
@@ -20,6 +20,7 @@
#include "storage/aio_internal.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
+#include "storage/md.h"
#include "storage/smgr.h"
#include "utils/memutils.h"
@@ -35,6 +36,7 @@ static const PgAioSubjectInfo *aio_subject_info[] = {
[ASI_INVALID] = &(PgAioSubjectInfo) {
.name = "invalid",
},
+ [ASI_SMGR] = &aio_smgr_subject_info,
};
@@ -46,6 +48,8 @@ typedef struct PgAioHandleSharedCallbacksEntry
static const PgAioHandleSharedCallbacksEntry aio_shared_cbs[] = {
#define CALLBACK_ENTRY(id, callback) [id] = {.cb = &callback, .name = #callback}
+ CALLBACK_ENTRY(ASC_MD_READV, aio_md_readv_cb),
+ CALLBACK_ENTRY(ASC_MD_WRITEV, aio_md_writev_cb),
#undef CALLBACK_ENTRY
};
diff --git a/src/backend/storage/file/fd.c b/src/backend/storage/file/fd.c
index 7c403fb360e..eeb6288a9b5 100644
--- a/src/backend/storage/file/fd.c
+++ b/src/backend/storage/file/fd.c
@@ -94,6 +94,7 @@
#include "miscadmin.h"
#include "pgstat.h"
#include "postmaster/startup.h"
+#include "storage/aio.h"
#include "storage/fd.h"
#include "storage/ipc.h"
#include "utils/guc.h"
@@ -1294,6 +1295,8 @@ LruDelete(File file)
vfdP = &VfdCache[file];
+ pgaio_closing_fd(vfdP->fd);
+
/*
* Close the file. We aren't expecting this to fail; if it does, better
* to leak the FD than to mess up our internal state.
@@ -1987,6 +1990,8 @@ FileClose(File file)
if (!FileIsNotOpen(file))
{
+ pgaio_closing_fd(vfdP->fd);
+
/* close the file */
if (close(vfdP->fd) != 0)
{
@@ -2210,6 +2215,32 @@ retry:
return returnCode;
}
+int
+FileStartReadV(struct PgAioHandle *ioh, File file,
+ int iovcnt, off_t offset,
+ uint32 wait_event_info)
+{
+ int returnCode;
+ Vfd *vfdP;
+
+ Assert(FileIsValid(file));
+
+ DO_DB(elog(LOG, "FileStartReadV: %d (%s) " INT64_FORMAT " %d",
+ file, VfdCache[file].fileName,
+ (int64) offset,
+ iovcnt));
+
+ returnCode = FileAccess(file);
+ if (returnCode < 0)
+ return returnCode;
+
+ vfdP = &VfdCache[file];
+
+ pgaio_io_prep_readv(ioh, vfdP->fd, iovcnt, offset);
+
+ return 0;
+}
+
ssize_t
FileWriteV(File file, const struct iovec *iov, int iovcnt, off_t offset,
uint32 wait_event_info)
@@ -2315,6 +2346,34 @@ retry:
return returnCode;
}
+int
+FileStartWriteV(struct PgAioHandle *ioh, File file,
+ int iovcnt, off_t offset,
+ uint32 wait_event_info)
+{
+ int returnCode;
+ Vfd *vfdP;
+
+ Assert(FileIsValid(file));
+
+ DO_DB(elog(LOG, "FileStartWriteV: %d (%s) " INT64_FORMAT " %d",
+ file, VfdCache[file].fileName,
+ (int64) offset,
+ iovcnt));
+
+ returnCode = FileAccess(file);
+ if (returnCode < 0)
+ return returnCode;
+
+ vfdP = &VfdCache[file];
+
+ /* FIXME: think about / reimplement temp_file_limit */
+
+ pgaio_io_prep_writev(ioh, vfdP->fd, iovcnt, offset);
+
+ return 0;
+}
+
int
FileSync(File file, uint32 wait_event_info)
{
@@ -2498,6 +2557,12 @@ FilePathName(File file)
int
FileGetRawDesc(File file)
{
+ int returnCode;
+
+ returnCode = FileAccess(file);
+ if (returnCode < 0)
+ return returnCode;
+
Assert(FileIsValid(file));
return VfdCache[file].fd;
}
@@ -2778,6 +2843,7 @@ FreeDesc(AllocateDesc *desc)
result = closedir(desc->desc.dir);
break;
case AllocateDescRawFD:
+ pgaio_closing_fd(desc->desc.fd);
result = close(desc->desc.fd);
break;
default:
@@ -2846,6 +2912,8 @@ CloseTransientFile(int fd)
/* Only get here if someone passes us a file not in allocatedDescs */
elog(WARNING, "fd passed to CloseTransientFile was not obtained from OpenTransientFile");
+ pgaio_closing_fd(fd);
+
return close(fd);
}
diff --git a/src/backend/storage/smgr/md.c b/src/backend/storage/smgr/md.c
index 11fccda475f..b1277ed97ae 100644
--- a/src/backend/storage/smgr/md.c
+++ b/src/backend/storage/smgr/md.c
@@ -31,6 +31,7 @@
#include "miscadmin.h"
#include "pg_trace.h"
#include "pgstat.h"
+#include "storage/aio.h"
#include "storage/bufmgr.h"
#include "storage/fd.h"
#include "storage/md.h"
@@ -132,6 +133,22 @@ static MdfdVec *_mdfd_getseg(SMgrRelation reln, ForkNumber forknum,
static BlockNumber _mdnblocks(SMgrRelation reln, ForkNumber forknum,
MdfdVec *seg);
+static PgAioResult md_readv_complete(PgAioHandle *ioh, PgAioResult prior_result);
+static void md_readv_error(PgAioResult result, const PgAioSubjectData *subject_data, int elevel);
+static PgAioResult md_writev_complete(PgAioHandle *ioh, PgAioResult prior_result);
+static void md_writev_error(PgAioResult result, const PgAioSubjectData *subject_data, int elevel);
+
+const struct PgAioHandleSharedCallbacks aio_md_readv_cb = {
+ .complete = md_readv_complete,
+ .error = md_readv_error,
+};
+
+const struct PgAioHandleSharedCallbacks aio_md_writev_cb = {
+ .complete = md_writev_complete,
+ .error = md_writev_error,
+};
+
+
static inline int
_mdfd_open_flags(void)
{
@@ -927,6 +944,52 @@ mdreadv(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
}
}
+void
+mdstartreadv(PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
+ void **buffers, BlockNumber nblocks)
+{
+ off_t seekpos;
+ MdfdVec *v;
+ BlockNumber nblocks_this_segment;
+ struct iovec *iov;
+ int iovcnt;
+
+ v = _mdfd_getseg(reln, forknum, blocknum, false,
+ EXTENSION_FAIL | EXTENSION_CREATE_RECOVERY);
+
+ seekpos = (off_t) BLCKSZ * (blocknum % ((BlockNumber) RELSEG_SIZE));
+
+ Assert(seekpos < (off_t) BLCKSZ * RELSEG_SIZE);
+
+ nblocks_this_segment =
+ Min(nblocks,
+ RELSEG_SIZE - (blocknum % ((BlockNumber) RELSEG_SIZE)));
+
+ if (nblocks_this_segment != nblocks)
+ elog(ERROR, "read crossing segment boundary");
+
+ iovcnt = pgaio_io_get_iovec(ioh, &iov);
+
+ Assert(nblocks <= iovcnt);
+
+ iovcnt = buffers_to_iovec(iov, buffers, nblocks_this_segment);
+
+ Assert(iovcnt <= nblocks_this_segment);
+
+ if (!(io_direct_flags & IO_DIRECT_DATA))
+ pgaio_io_set_flag(ioh, AHF_BUFFERED);
+
+ pgaio_io_set_subject_smgr(ioh,
+ reln,
+ forknum,
+ blocknum,
+ nblocks);
+ pgaio_io_add_shared_cb(ioh, ASC_MD_READV);
+
+ FileStartReadV(ioh, v->mdfd_vfd, iovcnt, seekpos, WAIT_EVENT_DATA_FILE_READ);
+}
+
/*
* mdwritev() -- Write the supplied blocks at the appropriate location.
*
@@ -1032,6 +1095,52 @@ mdwritev(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
}
}
+void
+mdstartwritev(PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
+ const void **buffers, BlockNumber nblocks, bool skipFsync)
+{
+ off_t seekpos;
+ MdfdVec *v;
+ BlockNumber nblocks_this_segment;
+ struct iovec *iov;
+ int iovcnt;
+
+ v = _mdfd_getseg(reln, forknum, blocknum, false,
+ EXTENSION_FAIL | EXTENSION_CREATE_RECOVERY);
+
+ seekpos = (off_t) BLCKSZ * (blocknum % ((BlockNumber) RELSEG_SIZE));
+
+ Assert(seekpos < (off_t) BLCKSZ * RELSEG_SIZE);
+
+ nblocks_this_segment =
+ Min(nblocks,
+ RELSEG_SIZE - (blocknum % ((BlockNumber) RELSEG_SIZE)));
+
+ if (nblocks_this_segment != nblocks)
+ elog(ERROR, "write crossing segment boundary");
+
+ iovcnt = pgaio_io_get_iovec(ioh, &iov);
+
+ Assert(nblocks <= iovcnt);
+
+ iovcnt = buffers_to_iovec(iov, unconstify(void **, buffers), nblocks_this_segment);
+
+ Assert(iovcnt <= nblocks_this_segment);
+
+ if (!(io_direct_flags & IO_DIRECT_DATA))
+ pgaio_io_set_flag(ioh, AHF_BUFFERED);
+
+ pgaio_io_set_subject_smgr(ioh,
+ reln,
+ forknum,
+ blocknum,
+ nblocks);
+ pgaio_io_add_shared_cb(ioh, ASC_MD_WRITEV);
+
+ FileStartWriteV(ioh, v->mdfd_vfd, iovcnt, seekpos, WAIT_EVENT_DATA_FILE_WRITE);
+}
+
/*
* mdwriteback() -- Tell the kernel to write pages back to storage.
@@ -1355,6 +1464,21 @@ mdimmedsync(SMgrRelation reln, ForkNumber forknum)
}
}
+int
+mdfd(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum, uint32 *off)
+{
+ MdfdVec *v = mdopenfork(reln, forknum, EXTENSION_FAIL);
+
+ v = _mdfd_getseg(reln, forknum, blocknum, false,
+ EXTENSION_FAIL);
+
+ *off = (off_t) BLCKSZ * (blocknum % ((BlockNumber) RELSEG_SIZE));
+
+ Assert(*off < (off_t) BLCKSZ * RELSEG_SIZE);
+
+ return FileGetRawDesc(v->mdfd_vfd);
+}
+
/*
* register_dirty_segment() -- Mark a relation segment as needing fsync
*
@@ -1838,3 +1962,193 @@ mdfiletagmatches(const FileTag *ftag, const FileTag *candidate)
*/
return ftag->rlocator.dbOid == candidate->rlocator.dbOid;
}
+
+/*
+ * AIO completion callback for mdstartreadv().
+ */
+static PgAioResult
+md_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
+{
+ PgAioSubjectData *sd = pgaio_io_get_subject_data(ioh);
+ PgAioResult result = prior_result;
+
+ if (prior_result.result < 0)
+ {
+ result.status = ARS_ERROR;
+ result.id = ASC_MD_READV;
+ /* For "hard" errors, track the error number in error_data */
+ result.error_data = -prior_result.result;
+ result.result = 0;
+
+ md_readv_error(result, sd, LOG);
+
+ return result;
+ }
+
+ result.result /= BLCKSZ;
+
+ if (result.result == 0)
+ {
+ /* consider 0 blocks read a failure */
+ result.status = ARS_ERROR;
+ result.id = ASC_MD_READV;
+ result.error_data = 0;
+
+ md_readv_error(result, sd, LOG);
+ }
+
+ if (result.status != ARS_ERROR &&
+ result.result < sd->smgr.nblocks)
+ {
+ /* partial reads should be retried at upper level */
+ result.status = ARS_PARTIAL;
+ result.id = ASC_MD_READV;
+ }
+
+ /* AFIXME: post-read portion of mdreadv() */
+
+ return result;
+}
+
+/*
+ * AIO error reporting callback for mdstartreadv().
+ */
+static void
+md_readv_error(PgAioResult result, const PgAioSubjectData *subject_data, int elevel)
+{
+ MemoryContext oldContext = CurrentMemoryContext;
+
+ /* AFIXME: */
+ oldContext = MemoryContextSwitchTo(ErrorContext);
+
+ if (result.error_data != 0)
+ {
+ errno = result.error_data; /* for errcode_for_file_access() */
+
+ ereport(elevel,
+ errcode_for_file_access(),
+ errmsg("could not read blocks %u..%u in file \"%s\": %m",
+ subject_data->smgr.blockNum,
+ subject_data->smgr.blockNum + subject_data->smgr.nblocks,
+ relpathperm(subject_data->smgr.rlocator, subject_data->smgr.forkNum)
+ )
+ );
+ }
+ else
+ {
+ /*
+ * NB: This will typically only be output in debug messages, while
+ * retrying a partial IO.
+ */
+ ereport(elevel,
+ errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("could not read blocks %u..%u in file \"%s\": read only %zu of %zu bytes",
+ subject_data->smgr.blockNum,
+ subject_data->smgr.blockNum + subject_data->smgr.nblocks - 1,
+ relpathperm(subject_data->smgr.rlocator, subject_data->smgr.forkNum),
+ result.result * (size_t) BLCKSZ,
+ subject_data->smgr.nblocks * (size_t) BLCKSZ
+ )
+ );
+ }
+
+ MemoryContextSwitchTo(oldContext);
+}
+
+/*
+ * AIO completion callback for mdstartwritev().
+ */
+static PgAioResult
+md_writev_complete(PgAioHandle *ioh, PgAioResult prior_result)
+{
+ PgAioSubjectData *sd = pgaio_io_get_subject_data(ioh);
+ PgAioResult result = prior_result;
+
+ if (prior_result.result < 0)
+ {
+ result.status = ARS_ERROR;
+ result.id = ASC_MD_WRITEV;
+ /* For "hard" errors, track the error number in error_data */
+ result.error_data = -prior_result.result;
+ result.result = 0;
+
+ md_writev_error(result, sd, LOG);
+
+ return result;
+ }
+
+ result.result /= BLCKSZ;
+
+ if (result.result == 0)
+ {
+ /* consider 0 blocks written a failure */
+ result.status = ARS_ERROR;
+ result.id = ASC_MD_WRITEV;
+ result.error_data = 0;
+
+ md_writev_error(result, sd, LOG);
+ }
+
+ if (result.status != ARS_ERROR &&
+ result.result < sd->smgr.nblocks)
+ {
+ /* partial writes should be retried at upper level */
+ result.status = ARS_PARTIAL;
+ result.id = ASC_MD_WRITEV;
+ }
+
+ if (prior_result.status == ARS_ERROR)
+ {
+ /* AFIXME: complain */
+ return prior_result;
+ }
+
+ prior_result.result /= BLCKSZ;
+
+ return prior_result;
+}
+
+/*
+ * AIO error reporting callback for mdstartwritev().
+ */
+static void
+md_writev_error(PgAioResult result, const PgAioSubjectData *subject_data, int elevel)
+{
+ MemoryContext oldContext = CurrentMemoryContext;
+
+ /* AFIXME: */
+ oldContext = MemoryContextSwitchTo(ErrorContext);
+
+ if (result.error_data != 0)
+ {
+ errno = result.error_data; /* for errcode_for_file_access() */
+
+ ereport(elevel,
+ errcode_for_file_access(),
+ errmsg("could not write blocks %u..%u in file \"%s\": %m",
+ subject_data->smgr.blockNum,
+ subject_data->smgr.blockNum + subject_data->smgr.nblocks,
+ relpathperm(subject_data->smgr.rlocator, subject_data->smgr.forkNum)
+ )
+ );
+ }
+ else
+ {
+ /*
+ * NB: This will typically only be output in debug messages, while
+ * retrying a partial IO.
+ */
+ ereport(elevel,
+ errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("could not write blocks %u..%u in file \"%s\": wrote only %zu of %zu bytes",
+ subject_data->smgr.blockNum,
+ subject_data->smgr.blockNum + subject_data->smgr.nblocks - 1,
+ relpathperm(subject_data->smgr.rlocator, subject_data->smgr.forkNum),
+ result.result * (size_t) BLCKSZ,
+ subject_data->smgr.nblocks * (size_t) BLCKSZ
+ )
+ );
+ }
+
+ MemoryContextSwitchTo(oldContext);
+}
diff --git a/src/backend/storage/smgr/smgr.c b/src/backend/storage/smgr/smgr.c
index 36ad34aa6ac..454ebe9c243 100644
--- a/src/backend/storage/smgr/smgr.c
+++ b/src/backend/storage/smgr/smgr.c
@@ -53,6 +53,7 @@
#include "access/xlogutils.h"
#include "lib/ilist.h"
+#include "storage/aio.h"
#include "storage/bufmgr.h"
#include "storage/ipc.h"
#include "storage/md.h"
@@ -93,10 +94,19 @@ typedef struct f_smgr
void (*smgr_readv) (SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum,
void **buffers, BlockNumber nblocks);
+ void (*smgr_startreadv) (struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum,
+ BlockNumber blocknum,
+ void **buffers, BlockNumber nblocks);
void (*smgr_writev) (SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum,
const void **buffers, BlockNumber nblocks,
bool skipFsync);
+ void (*smgr_startwritev) (struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum,
+ BlockNumber blocknum,
+ const void **buffers, BlockNumber nblocks,
+ bool skipFsync);
void (*smgr_writeback) (SMgrRelation reln, ForkNumber forknum,
BlockNumber blocknum, BlockNumber nblocks);
BlockNumber (*smgr_nblocks) (SMgrRelation reln, ForkNumber forknum);
@@ -104,6 +114,7 @@ typedef struct f_smgr
BlockNumber old_blocks, BlockNumber nblocks);
void (*smgr_immedsync) (SMgrRelation reln, ForkNumber forknum);
void (*smgr_registersync) (SMgrRelation reln, ForkNumber forknum);
+ int (*smgr_fd) (SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum, uint32 *off);
} f_smgr;
static const f_smgr smgrsw[] = {
@@ -121,12 +132,15 @@ static const f_smgr smgrsw[] = {
.smgr_prefetch = mdprefetch,
.smgr_maxcombine = mdmaxcombine,
.smgr_readv = mdreadv,
+ .smgr_startreadv = mdstartreadv,
.smgr_writev = mdwritev,
+ .smgr_startwritev = mdstartwritev,
.smgr_writeback = mdwriteback,
.smgr_nblocks = mdnblocks,
.smgr_truncate = mdtruncate,
.smgr_immedsync = mdimmedsync,
.smgr_registersync = mdregistersync,
+ .smgr_fd = mdfd,
}
};
@@ -145,6 +159,14 @@ static void smgrshutdown(int code, Datum arg);
static void smgrdestroy(SMgrRelation reln);
+static void smgr_aio_reopen(PgAioHandle *ioh);
+
+const struct PgAioSubjectInfo aio_smgr_subject_info = {
+ .name = "smgr",
+ .reopen = smgr_aio_reopen,
+};
+
+
/*
* smgrinit(), smgrshutdown() -- Initialize or shut down storage
* managers.
@@ -623,6 +645,19 @@ smgrreadv(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
nblocks);
}
+/*
+ * FILL ME IN
+ */
+void
+smgrstartreadv(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
+ void **buffers, BlockNumber nblocks)
+{
+ smgrsw[reln->smgr_which].smgr_startreadv(ioh,
+ reln, forknum, blocknum, buffers,
+ nblocks);
+}
+
/*
* smgrwritev() -- Write the supplied buffers out.
*
@@ -657,6 +692,16 @@ smgrwritev(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
buffers, nblocks, skipFsync);
}
+void
+smgrstartwritev(struct PgAioHandle *ioh,
+ SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum,
+ const void **buffers, BlockNumber nblocks, bool skipFsync)
+{
+ smgrsw[reln->smgr_which].smgr_startwritev(ioh,
+ reln, forknum, blocknum, buffers,
+ nblocks, skipFsync);
+}
+
/*
* smgrwriteback() -- Trigger kernel writeback for the supplied range of
* blocks.
@@ -819,6 +864,12 @@ smgrimmedsync(SMgrRelation reln, ForkNumber forknum)
smgrsw[reln->smgr_which].smgr_immedsync(reln, forknum);
}
+int
+smgrfd(SMgrRelation reln, ForkNumber forknum, BlockNumber blocknum, uint32 *off)
+{
+ return smgrsw[reln->smgr_which].smgr_fd(reln, forknum, blocknum, off);
+}
+
/*
* AtEOXact_SMgr
*
@@ -847,3 +898,43 @@ ProcessBarrierSmgrRelease(void)
smgrreleaseall();
return true;
}
+
+void
+pgaio_io_set_subject_smgr(PgAioHandle *ioh,
+ struct SMgrRelationData *smgr,
+ ForkNumber forknum,
+ BlockNumber blocknum,
+ int nblocks)
+{
+ PgAioSubjectData *sd = pgaio_io_get_subject_data(ioh);
+
+ pgaio_io_set_subject(ioh, ASI_SMGR);
+
+ /* backend is implied via IO owner */
+ sd->smgr.rlocator = smgr->smgr_rlocator.locator;
+ sd->smgr.forkNum = forknum;
+ sd->smgr.blockNum = blocknum;
+ sd->smgr.nblocks = nblocks;
+ sd->smgr.is_temp = SmgrIsTemp(smgr);
+ sd->smgr.release_lock = false;
+ sd->smgr.mode = RBM_NORMAL;
+}
+
+static void
+smgr_aio_reopen(PgAioHandle *ioh)
+{
+ PgAioSubjectData *sd = pgaio_io_get_subject_data(ioh);
+ PgAioOpData *od = pgaio_io_get_op_data(ioh);
+ SMgrRelation reln;
+ ProcNumber procno;
+ uint32 off;
+
+ if (sd->smgr.is_temp)
+ procno = pgaio_io_get_owner(ioh);
+ else
+ procno = INVALID_PROC_NUMBER;
+
+ reln = smgropen(sd->smgr.rlocator, procno);
+ od->read.fd = smgrfd(reln, sd->smgr.forkNum, sd->smgr.blockNum, &off);
+ Assert(off == od->read.offset);
+}
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0011-bufmgr-Implement-AIO-read-support.patch (19.2K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/12-v2-0011-bufmgr-Implement-AIO-read-support.patch)
download | inline diff:
From 7a42b48f7421f071dab6cff273e4cc5b1c3c755f Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sat, 31 Aug 2024 21:39:01 -0400
Subject: [PATCH v2 11/20] bufmgr: Implement AIO read support
As of this commit there are no users of these AIO facilities, that'll come in
later commits.
Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
src/include/storage/aio.h | 4 +
src/include/storage/buf_internals.h | 6 +
src/include/storage/bufmgr.h | 8 +
src/backend/storage/aio/aio_subject.c | 4 +
src/backend/storage/buffer/buf_init.c | 3 +
src/backend/storage/buffer/bufmgr.c | 364 +++++++++++++++++++++++++-
src/backend/storage/buffer/localbuf.c | 65 +++++
7 files changed, 447 insertions(+), 7 deletions(-)
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index d693b0b0bd8..ff44dac5bb2 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -104,6 +104,10 @@ typedef enum PgAioHandleSharedCallbackID
ASC_MD_READV,
ASC_MD_WRITEV,
+
+ ASC_SHARED_BUFFER_READ,
+
+ ASC_LOCAL_BUFFER_READ,
} PgAioHandleSharedCallbackID;
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index eda6c699212..37520890073 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -17,6 +17,7 @@
#include "pgstat.h"
#include "port/atomics.h"
+#include "storage/aio_ref.h"
#include "storage/buf.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
@@ -251,6 +252,8 @@ typedef struct BufferDesc
int wait_backend_pgprocno; /* backend of pin-count waiter */
int freeNext; /* link in freelist chain */
+
+ PgAioHandleRef io_in_progress;
LWLock content_lock; /* to lock access to buffer contents */
} BufferDesc;
@@ -464,4 +467,7 @@ extern void DropRelationLocalBuffers(RelFileLocator rlocator,
extern void DropRelationAllLocalBuffers(RelFileLocator rlocator);
extern void AtEOXact_LocalBuffers(bool isCommit);
+
+extern bool ReadBufferCompleteReadLocal(Buffer buffer, int mode, bool failed);
+
#endif /* BUFMGR_INTERNALS_H */
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index eb0fba4230b..ca8e8b51e68 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -177,6 +177,12 @@ extern PGDLLIMPORT int NLocBuffer;
extern PGDLLIMPORT Block *LocalBufferBlockPointers;
extern PGDLLIMPORT int32 *LocalRefCount;
+
+struct PgAioHandleSharedCallbacks;
+extern const struct PgAioHandleSharedCallbacks aio_shared_buffer_readv_cb;
+extern const struct PgAioHandleSharedCallbacks aio_local_buffer_readv_cb;
+
+
/* upper limit for effective_io_concurrency */
#define MAX_IO_CONCURRENCY 1000
@@ -194,6 +200,8 @@ extern PGDLLIMPORT int32 *LocalRefCount;
/*
* prototypes for functions in bufmgr.c
*/
+struct PgAioHandle;
+
extern PrefetchBufferResult PrefetchSharedBuffer(struct SMgrRelationData *smgr_reln,
ForkNumber forkNum,
BlockNumber blockNum);
diff --git a/src/backend/storage/aio/aio_subject.c b/src/backend/storage/aio/aio_subject.c
index effb09c11c7..21341aae425 100644
--- a/src/backend/storage/aio/aio_subject.c
+++ b/src/backend/storage/aio/aio_subject.c
@@ -50,6 +50,10 @@ static const PgAioHandleSharedCallbacksEntry aio_shared_cbs[] = {
#define CALLBACK_ENTRY(id, callback) [id] = {.cb = &callback, .name = #callback}
CALLBACK_ENTRY(ASC_MD_READV, aio_md_readv_cb),
CALLBACK_ENTRY(ASC_MD_WRITEV, aio_md_writev_cb),
+
+ CALLBACK_ENTRY(ASC_SHARED_BUFFER_READ, aio_shared_buffer_readv_cb),
+
+ CALLBACK_ENTRY(ASC_LOCAL_BUFFER_READ, aio_local_buffer_readv_cb),
#undef CALLBACK_ENTRY
};
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 56761a8eedc..7853b1877e0 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -14,6 +14,7 @@
*/
#include "postgres.h"
+#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
@@ -125,6 +126,8 @@ BufferManagerShmemInit(void)
buf->buf_id = i;
+ pgaio_io_ref_clear(&buf->io_in_progress);
+
/*
* Initially link all the buffers together as unused. Subsequent
* management of this list is done by freelist.c.
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 2622221809c..c0fb3028c95 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -48,6 +48,7 @@
#include "pg_trace.h"
#include "pgstat.h"
#include "postmaster/bgwriter.h"
+#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "storage/fd.h"
@@ -58,6 +59,7 @@
#include "storage/smgr.h"
#include "storage/standby.h"
#include "utils/memdebug.h"
+#include "utils/memutils.h"
#include "utils/ps_status.h"
#include "utils/rel.h"
#include "utils/resowner.h"
@@ -514,7 +516,8 @@ static int SyncOneBuffer(int buf_id, bool skip_recently_used,
static void WaitIO(BufferDesc *buf);
static bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
static void TerminateBufferIO(BufferDesc *buf, bool clear_dirty,
- uint32 set_flag_bits, bool forget_owner);
+ uint32 set_flag_bits, bool forget_owner,
+ bool syncio);
static void AbortBufferIO(Buffer buffer);
static void shared_buffer_write_error_callback(void *arg);
static void local_buffer_write_error_callback(void *arg);
@@ -1081,7 +1084,7 @@ ZeroAndLockBuffer(Buffer buffer, ReadBufferMode mode, bool already_valid)
else
{
/* Set BM_VALID, terminate IO, and wake up any waiters */
- TerminateBufferIO(bufHdr, false, BM_VALID, true);
+ TerminateBufferIO(bufHdr, false, BM_VALID, true, true);
}
}
else if (!isLocalBuf)
@@ -1566,7 +1569,7 @@ WaitReadBuffers(ReadBuffersOperation *operation)
else
{
/* Set BM_VALID, terminate IO, and wake up any waiters */
- TerminateBufferIO(bufHdr, false, BM_VALID, true);
+ TerminateBufferIO(bufHdr, false, BM_VALID, true, true);
}
/* Report I/Os as completing individually. */
@@ -2450,7 +2453,7 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
if (lock)
LWLockAcquire(BufferDescriptorGetContentLock(buf_hdr), LW_EXCLUSIVE);
- TerminateBufferIO(buf_hdr, false, BM_VALID, true);
+ TerminateBufferIO(buf_hdr, false, BM_VALID, true, true);
}
pgBufferUsage.shared_blks_written += extend_by;
@@ -3899,7 +3902,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
* Mark the buffer as clean (unless BM_JUST_DIRTIED has become set) and
* end the BM_IO_IN_PROGRESS state.
*/
- TerminateBufferIO(buf, true, 0, true);
+ TerminateBufferIO(buf, true, 0, true, true);
TRACE_POSTGRESQL_BUFFER_FLUSH_DONE(BufTagGetForkNum(&buf->tag),
buf->tag.blockNum,
@@ -5514,6 +5517,7 @@ WaitIO(BufferDesc *buf)
for (;;)
{
uint32 buf_state;
+ PgAioHandleRef ior;
/*
* It may not be necessary to acquire the spinlock to check the flag
@@ -5521,10 +5525,19 @@ WaitIO(BufferDesc *buf)
* play it safe.
*/
buf_state = LockBufHdr(buf);
+ ior = buf->io_in_progress;
UnlockBufHdr(buf, buf_state);
if (!(buf_state & BM_IO_IN_PROGRESS))
break;
+
+ if (pgaio_io_ref_valid(&ior))
+ {
+ pgaio_io_ref_wait(&ior);
+ ConditionVariablePrepareToSleep(cv);
+ continue;
+ }
+
ConditionVariableSleep(cv, WAIT_EVENT_BUFFER_IO);
}
ConditionVariableCancelSleep();
@@ -5613,7 +5626,7 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
*/
static void
TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
- bool forget_owner)
+ bool forget_owner, bool syncio)
{
uint32 buf_state;
@@ -5625,6 +5638,13 @@ TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
if (clear_dirty && !(buf_state & BM_JUST_DIRTIED))
buf_state &= ~(BM_DIRTY | BM_CHECKPOINT_NEEDED);
+ if (!syncio)
+ {
+ /* release ownership by the AIO subsystem */
+ buf_state -= BUF_REFCOUNT_ONE;
+ pgaio_io_ref_clear(&buf->io_in_progress);
+ }
+
buf_state |= set_flag_bits;
UnlockBufHdr(buf, buf_state);
@@ -5633,6 +5653,40 @@ TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
BufferDescriptorGetBuffer(buf));
ConditionVariableBroadcast(BufferDescriptorGetIOCV(buf));
+
+ /*
+ * If we just released a pin, need to do BM_PIN_COUNT_WAITER handling.
+ * Most of the time the current backend will hold another pin preventing
+ * that from happening, but that's e.g. not the case when completing an IO
+ * another backend started.
+ *
+ * AFIXME: Deduplicate with UnpinBufferNoOwner() or just replace
+ * BM_PIN_COUNT_WAITER with something saner.
+ */
+ /* Support LockBufferForCleanup() */
+ if (buf_state & BM_PIN_COUNT_WAITER)
+ {
+ /*
+ * Acquire the buffer header lock, re-check that there's a waiter.
+ * Another backend could have unpinned this buffer, and already woken
+ * up the waiter. There's no danger of the buffer being replaced
+ * after we unpinned it above, as it's pinned by the waiter.
+ */
+ buf_state = LockBufHdr(buf);
+
+ if ((buf_state & BM_PIN_COUNT_WAITER) &&
+ BUF_STATE_GET_REFCOUNT(buf_state) == 1)
+ {
+ /* we just released the last pin other than the waiter's */
+ int wait_backend_pgprocno = buf->wait_backend_pgprocno;
+
+ buf_state &= ~BM_PIN_COUNT_WAITER;
+ UnlockBufHdr(buf, buf_state);
+ ProcSendSignal(wait_backend_pgprocno);
+ }
+ else
+ UnlockBufHdr(buf, buf_state);
+ }
}
/*
@@ -5684,7 +5738,7 @@ AbortBufferIO(Buffer buffer)
}
}
- TerminateBufferIO(buf_hdr, false, BM_IO_ERROR, false);
+ TerminateBufferIO(buf_hdr, false, BM_IO_ERROR, false, true);
}
/*
@@ -6143,3 +6197,299 @@ EvictUnpinnedBuffer(Buffer buf)
return result;
}
+
+static bool
+ReadBufferCompleteReadShared(Buffer buffer, int mode, bool failed)
+{
+ BufferDesc *bufHdr = NULL;
+ BlockNumber blockno;
+ bool buf_failed = false;
+ char *bufdata = BufferGetBlock(buffer);
+
+ Assert(BufferIsValid(buffer));
+
+ bufHdr = GetBufferDescriptor(buffer - 1);
+ blockno = bufHdr->tag.blockNum;
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ uint32 buf_state = pg_atomic_read_u32(&bufHdr->state);
+
+ Assert(buf_state & BM_TAG_VALID);
+ Assert(!(buf_state & BM_VALID));
+ Assert(buf_state & BM_IO_IN_PROGRESS);
+ Assert(!(buf_state & BM_DIRTY));
+ }
+#endif
+
+ /* check for garbage data */
+ if (!failed &&
+ !PageIsVerifiedExtended((Page) bufdata, blockno,
+ PIV_LOG_WARNING | PIV_REPORT_STAT))
+ {
+ RelFileLocator rlocator = BufTagGetRelFileLocator(&bufHdr->tag);
+ BlockNumber forkNum = bufHdr->tag.forkNum;
+
+ /* AFIXME: relpathperm allocates memory */
+ MemoryContextSwitchTo(ErrorContext);
+ if (mode == READ_BUFFERS_ZERO_ON_ERROR || zero_damaged_pages)
+ {
+ ereport(LOG,
+ (errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("invalid page in block %u of relation %s; zeroing out page",
+ blockno,
+ relpathperm(rlocator, forkNum))));
+ memset(bufdata, 0, BLCKSZ);
+ }
+ else
+ {
+ ereport(LOG,
+ (errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("invalid page in block %u of relation %s",
+ blockno,
+ relpathperm(rlocator, forkNum))));
+ failed = true;
+ buf_failed = true;
+ }
+ }
+
+ /* Terminate I/O and set BM_VALID. */
+ TerminateBufferIO(bufHdr, false,
+ failed ? BM_IO_ERROR : BM_VALID,
+ false, false);
+
+ /* Report I/Os as completing individually. */
+
+ /* FIXME: Should we do TRACE_POSTGRESQL_BUFFER_READ_DONE here? */
+ return buf_failed;
+}
+
+/*
+ * Helper to prepare IO on shared buffers for execution, shared between reads
+ * and writes.
+ */
+static void
+shared_buffer_prepare_common(PgAioHandle *ioh, bool is_write)
+{
+ uint64 *io_data;
+ uint8 io_data_len;
+ PgAioHandleRef io_ref;
+ BufferTag first PG_USED_FOR_ASSERTS_ONLY = {0};
+
+ io_data = pgaio_io_get_io_data(ioh, &io_data_len);
+
+ pgaio_io_get_ref(ioh, &io_ref);
+
+ for (int i = 0; i < io_data_len; i++)
+ {
+ Buffer buf = (Buffer) io_data[i];
+ BufferDesc *bufHdr;
+ uint32 buf_state;
+
+ bufHdr = GetBufferDescriptor(buf - 1);
+
+ if (i == 0)
+ first = bufHdr->tag;
+ else
+ {
+ Assert(bufHdr->tag.relNumber == first.relNumber);
+ Assert(bufHdr->tag.blockNum == first.blockNum + i);
+ }
+
+
+ buf_state = LockBufHdr(bufHdr);
+
+ Assert(buf_state & BM_TAG_VALID);
+ if (is_write)
+ {
+ Assert(buf_state & BM_VALID);
+ Assert(buf_state & BM_DIRTY);
+ }
+ else
+ Assert(!(buf_state & BM_VALID));
+
+ Assert(buf_state & BM_IO_IN_PROGRESS);
+ Assert(BUF_STATE_GET_REFCOUNT(buf_state) >= 1);
+
+ buf_state += BUF_REFCOUNT_ONE;
+ bufHdr->io_in_progress = io_ref;
+
+ UnlockBufHdr(bufHdr, buf_state);
+
+ if (is_write)
+ {
+ LWLock *content_lock;
+
+ content_lock = BufferDescriptorGetContentLock(bufHdr);
+
+ Assert(LWLockHeldByMe(content_lock));
+
+ /*
+ * Lock is now owned by IO.
+ */
+ LWLockDisown(content_lock);
+ RESUME_INTERRUPTS();
+ }
+
+ /*
+ * Stop tracking this buffer via the resowner - the AIO system now
+ * keeps track.
+ */
+ ResourceOwnerForgetBufferIO(CurrentResourceOwner, buf);
+ }
+}
+
+static void
+shared_buffer_readv_prepare(PgAioHandle *ioh)
+{
+ shared_buffer_prepare_common(ioh, false);
+}
+
+static PgAioResult
+shared_buffer_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
+{
+ PgAioResult result = prior_result;
+ int mode = pgaio_io_get_subject_data(ioh)->smgr.mode;
+ uint64 *io_data;
+ uint8 io_data_len;
+
+ elog(DEBUG3, "%s: %d %d", __func__, prior_result.status, prior_result.result);
+
+ io_data = pgaio_io_get_io_data(ioh, &io_data_len);
+
+ for (int io_data_off = 0; io_data_off < io_data_len; io_data_off++)
+ {
+ Buffer buf = io_data[io_data_off];
+ bool buf_failed;
+ bool failed;
+
+ failed =
+ prior_result.status == ARS_ERROR
+ || prior_result.result <= io_data_off;
+
+ elog(DEBUG3, "calling rbcrs for buf %d with failed %d, error: %d, result: %d, data_off: %d",
+ buf, failed, prior_result.status, prior_result.result, io_data_off);
+
+ /*
+ * XXX: It might be better to not set BM_IO_ERROR (which is what
+ * failed = true leads to) when it's just a short read...
+ */
+ buf_failed = ReadBufferCompleteReadShared(buf,
+ mode,
+ failed);
+
+ if (result.status != ARS_ERROR && buf_failed)
+ {
+ result.status = ARS_ERROR;
+ result.id = ASC_SHARED_BUFFER_READ;
+ result.error_data = io_data_off + 1;
+ }
+ }
+
+ return result;
+}
+
+static void
+buffer_readv_error(PgAioResult result, const PgAioSubjectData *subject_data, int elevel)
+{
+ MemoryContext oldContext = CurrentMemoryContext;
+ ProcNumber errProc;
+
+ if (subject_data->smgr.is_temp)
+ errProc = MyProcNumber;
+ else
+ errProc = INVALID_PROC_NUMBER;
+
+ /* AFIXME: need infrastructure to allow memory allocation for error reporting */
+ oldContext = MemoryContextSwitchTo(ErrorContext);
+
+ ereport(elevel,
+ errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("invalid page in block %u of relation %s",
+ subject_data->smgr.blockNum + result.error_data,
+ relpathbackend(subject_data->smgr.rlocator, errProc, subject_data->smgr.forkNum)
+ )
+ );
+ MemoryContextSwitchTo(oldContext);
+}
+
+/*
+ * Helper to prepare IO on local buffers for execution, shared between reads
+ * and writes.
+ */
+static void
+local_buffer_readv_prepare(PgAioHandle *ioh)
+{
+ uint64 *io_data;
+ uint8 io_data_len;
+ PgAioHandleRef io_ref;
+
+ io_data = pgaio_io_get_io_data(ioh, &io_data_len);
+
+ pgaio_io_get_ref(ioh, &io_ref);
+
+ for (int i = 0; i < io_data_len; i++)
+ {
+ Buffer buf = (Buffer) io_data[i];
+ BufferDesc *bufHdr;
+ uint32 buf_state;
+
+ bufHdr = GetLocalBufferDescriptor(-buf - 1);
+
+ buf_state = pg_atomic_read_u32(&bufHdr->state);
+
+ bufHdr->io_in_progress = io_ref;
+ LocalRefCount[-buf - 1] += 1;
+
+ UnlockBufHdr(bufHdr, buf_state);
+ }
+}
+
+static PgAioResult
+local_buffer_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
+{
+ PgAioResult result = prior_result;
+ int mode = pgaio_io_get_subject_data(ioh)->smgr.mode;
+ uint64 *io_data;
+ uint8 io_data_len;
+
+ elog(DEBUG3, "%s: %d %d", __func__, prior_result.status, prior_result.result);
+
+ io_data = pgaio_io_get_io_data(ioh, &io_data_len);
+
+ for (int io_data_off = 0; io_data_off < io_data_len; io_data_off++)
+ {
+ Buffer buf = io_data[io_data_off];
+ bool buf_failed;
+ bool failed;
+
+ failed =
+ prior_result.status == ARS_ERROR
+ || prior_result.result <= io_data_off;
+
+ buf_failed = ReadBufferCompleteReadLocal(buf,
+ mode,
+ failed);
+
+ if (result.status != ARS_ERROR && buf_failed)
+ {
+ result.status = ARS_ERROR;
+ result.id = ASC_LOCAL_BUFFER_READ;
+ result.error_data = io_data_off + 1;
+ }
+ }
+
+ return result;
+}
+
+
+const struct PgAioHandleSharedCallbacks aio_shared_buffer_readv_cb = {
+ .prepare = shared_buffer_readv_prepare,
+ .complete = shared_buffer_readv_complete,
+ .error = buffer_readv_error,
+};
+const struct PgAioHandleSharedCallbacks aio_local_buffer_readv_cb = {
+ .prepare = local_buffer_readv_prepare,
+ .complete = local_buffer_readv_complete,
+ .error = buffer_readv_error,
+};
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 6fd1a6418d2..75c4d2570e0 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -18,6 +18,7 @@
#include "access/parallel.h"
#include "executor/instrument.h"
#include "pgstat.h"
+#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "storage/fd.h"
@@ -620,6 +621,8 @@ InitLocalBuffers(void)
*/
buf->buf_id = -i - 2;
+ pgaio_io_ref_clear(&buf->io_in_progress);
+
/*
* Intentionally do not initialize the buffer's atomic variable
* (besides zeroing the underlying memory above). That way we get
@@ -836,3 +839,65 @@ AtProcExit_LocalBuffers(void)
*/
CheckForLocalBufferLeaks();
}
+
+bool
+ReadBufferCompleteReadLocal(Buffer buffer, int mode, bool failed)
+{
+ BufferDesc *buf_hdr = NULL;
+ BlockNumber blockno;
+ bool buf_failed = false;
+ char *bufdata = BufferGetBlock(buffer);
+
+ Assert(BufferIsValid(buffer));
+
+ buf_hdr = GetLocalBufferDescriptor(-buffer - 1);
+ blockno = buf_hdr->tag.blockNum;
+
+ /* check for garbage data */
+ if (!failed &&
+ !PageIsVerifiedExtended((Page) bufdata, blockno,
+ PIV_LOG_WARNING | PIV_REPORT_STAT))
+ {
+ RelFileLocator rlocator = BufTagGetRelFileLocator(&buf_hdr->tag);
+ BlockNumber forkNum = buf_hdr->tag.forkNum;
+
+ MemoryContextSwitchTo(ErrorContext);
+
+ if (mode == READ_BUFFERS_ZERO_ON_ERROR || zero_damaged_pages)
+ {
+
+ ereport(WARNING,
+ (errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("invalid page in block %u of relation %s; zeroing out page",
+ blockno,
+ relpathperm(rlocator, forkNum))));
+ memset(bufdata, 0, BLCKSZ);
+ }
+ else
+ {
+ ereport(LOG,
+ (errcode(ERRCODE_DATA_CORRUPTED),
+ errmsg("invalid page in block %u of relation %s",
+ blockno,
+ relpathperm(rlocator, forkNum))));
+ failed = true;
+ buf_failed = true;
+ }
+ }
+
+ /* Terminate I/O and set BM_VALID. */
+ pgaio_io_ref_clear(&buf_hdr->io_in_progress);
+
+ {
+ uint32 buf_state;
+
+ buf_state = pg_atomic_read_u32(&buf_hdr->state);
+ buf_state |= BM_VALID;
+ pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+ }
+
+ /* release pin held by IO subsystem */
+ LocalRefCount[-buffer - 1] -= 1;
+
+ return buf_failed;
+}
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0012-bufmgr-Use-aio-for-StartReadBuffers.patch (19.0K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/13-v2-0012-bufmgr-Use-aio-for-StartReadBuffers.patch)
download | inline diff:
From e8a5a6318b0e386afb2c1ed2d7f4cc0372358ade Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sat, 31 Aug 2024 21:55:59 -0400
Subject: [PATCH v2 12/20] bufmgr: Use aio for StartReadBuffers()
Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
src/include/storage/bufmgr.h | 27 +-
src/backend/storage/buffer/bufmgr.c | 378 ++++++++++++++++++++--------
2 files changed, 300 insertions(+), 105 deletions(-)
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index ca8e8b51e68..7a12ef6e9be 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -15,6 +15,7 @@
#define BUFMGR_H
#include "port/pg_iovec.h"
+#include "storage/aio_ref.h"
#include "storage/block.h"
#include "storage/buf.h"
#include "storage/bufpage.h"
@@ -107,10 +108,23 @@ typedef struct BufferManagerRelation
#define BMR_REL(p_rel) ((BufferManagerRelation){.rel = p_rel})
#define BMR_SMGR(p_smgr, p_relpersistence) ((BufferManagerRelation){.smgr = p_smgr, .relpersistence = p_relpersistence})
+
+#define MAX_IO_COMBINE_LIMIT PG_IOV_MAX
+#define DEFAULT_IO_COMBINE_LIMIT Min(MAX_IO_COMBINE_LIMIT, (128 * 1024) / BLCKSZ)
+
+
/* Zero out page if reading fails. */
#define READ_BUFFERS_ZERO_ON_ERROR (1 << 0)
/* Call smgrprefetch() if I/O necessary. */
#define READ_BUFFERS_ISSUE_ADVICE (1 << 1)
+/* IO will immediately be waited for */
+#define READ_BUFFERS_SYNCHRONOUSLY (1 << 2)
+
+/*
+ * FIXME: PgAioReturn is defined in aio.h. It'd be much better if we didn't
+ * need to include that here. Perhaps this could live in a separate header?
+ */
+#include "storage/aio.h"
struct ReadBuffersOperation
{
@@ -131,6 +145,17 @@ struct ReadBuffersOperation
int flags;
int16 nblocks;
int16 io_buffers_len;
+
+ /*
+ * In some rare-ish cases one operation causes multiple reads (e.g. if a
+ * buffer was concurrently read by another backend). It'd be much better
+ * if we ensured that each ReadBuffersOperation covered only one IO - but
+ * that's not entirely trivial, due to having pinned victim buffers before
+ * starting IOs.
+ */
+ int16 nios;
+ PgAioHandleRef refs[MAX_IO_COMBINE_LIMIT];
+ PgAioReturn returns[MAX_IO_COMBINE_LIMIT];
};
typedef struct ReadBuffersOperation ReadBuffersOperation;
@@ -161,8 +186,6 @@ extern PGDLLIMPORT bool track_io_timing;
extern PGDLLIMPORT int effective_io_concurrency;
extern PGDLLIMPORT int maintenance_io_concurrency;
-#define MAX_IO_COMBINE_LIMIT PG_IOV_MAX
-#define DEFAULT_IO_COMBINE_LIMIT Min(MAX_IO_COMBINE_LIMIT, (128 * 1024) / BLCKSZ)
extern PGDLLIMPORT int io_combine_limit;
extern PGDLLIMPORT int checkpoint_flush_after;
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index c0fb3028c95..89cb7b41b03 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -1235,10 +1235,9 @@ ReadBuffer_common(Relation rel, SMgrRelation smgr, char smgr_persistence,
return buffer;
}
+ flags = READ_BUFFERS_SYNCHRONOUSLY;
if (mode == RBM_ZERO_ON_ERROR)
- flags = READ_BUFFERS_ZERO_ON_ERROR;
- else
- flags = 0;
+ flags |= READ_BUFFERS_ZERO_ON_ERROR;
operation.smgr = smgr;
operation.rel = rel;
operation.persistence = persistence;
@@ -1253,6 +1252,9 @@ ReadBuffer_common(Relation rel, SMgrRelation smgr, char smgr_persistence,
return buffer;
}
+static bool AsyncReadBuffers(ReadBuffersOperation *operation,
+ int nblocks);
+
static pg_attribute_always_inline bool
StartReadBuffersImpl(ReadBuffersOperation *operation,
Buffer *buffers,
@@ -1288,6 +1290,12 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
* so we stop here.
*/
actual_nblocks = i + 1;
+
+ ereport(DEBUG3,
+ errmsg("found buf %d, idx %i: %s, data %p",
+ buffers[i], i, DebugPrintBufferRefcount(buffers[i]),
+ BufferGetBlock(buffers[i])),
+ errhidestmt(true), errhidecontext(true));
break;
}
else
@@ -1324,28 +1332,51 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
operation->flags = flags;
operation->nblocks = actual_nblocks;
operation->io_buffers_len = io_buffers_len;
+ operation->nios = 0;
- if (flags & READ_BUFFERS_ISSUE_ADVICE)
+ /*
+ * When using AIO, start the IO in the background. If not, issue prefetch
+ * requests if desired by the caller.
+ *
+ * The reason we have a dedicated path for IOMETHOD_SYNC here is to derisk
+ * the introduction of AIO somewhat. It's a large architectural change,
+ * with lots of chances for unanticipated performance effects. Use of
+ * IOMETHOD_SYNC already leads to not actually performing IO
+ * asynchronously, but without the check here we'd execute IO earlier than
+ * we used to.
+ */
+ if (io_method != IOMETHOD_SYNC)
{
- /*
- * In theory we should only do this if PinBufferForBlock() had to
- * allocate new buffers above. That way, if two calls to
- * StartReadBuffers() were made for the same blocks before
- * WaitReadBuffers(), only the first would issue the advice. That'd be
- * a better simulation of true asynchronous I/O, which would only
- * start the I/O once, but isn't done here for simplicity. Note also
- * that the following call might actually issue two advice calls if we
- * cross a segment boundary; in a true asynchronous version we might
- * choose to process only one real I/O at a time in that case.
- */
- smgrprefetch(operation->smgr,
- operation->forknum,
- blockNum,
- operation->io_buffers_len);
+ /* initiate the IO asynchronously */
+ return AsyncReadBuffers(operation, io_buffers_len);
}
+ else
+ {
+ operation->flags |= READ_BUFFERS_SYNCHRONOUSLY;
+
+ if (flags & READ_BUFFERS_ISSUE_ADVICE)
+ {
+ /*
+ * In theory we should only do this if PinBufferForBlock() had to
+ * allocate new buffers above. That way, if two calls to
+ * StartReadBuffers() were made for the same blocks before
+ * WaitReadBuffers(), only the first would issue the
+ * advice. That'd be a better simulation of true asynchronous I/O,
+ * which would only start the I/O once, but isn't done here for
+ * simplicity. Note also that the following call might actually
+ * issue two advice calls if we cross a segment boundary; in a
+ * true asynchronous version we might choose to process only one
+ * real I/O at a time in that case.
+ */
+ smgrprefetch(operation->smgr,
+ operation->forknum,
+ blockNum,
+ operation->io_buffers_len);
+ }
- /* Indicate that WaitReadBuffers() should be called. */
- return true;
+ /* Indicate that WaitReadBuffers() should be called. */
+ return true;
+ }
}
/*
@@ -1397,12 +1428,31 @@ StartReadBuffer(ReadBuffersOperation *operation,
}
static inline bool
-WaitReadBuffersCanStartIO(Buffer buffer, bool nowait)
+ReadBuffersCanStartIO(Buffer buffer, bool nowait)
{
if (BufferIsLocal(buffer))
{
BufferDesc *bufHdr = GetLocalBufferDescriptor(-buffer - 1);
+ /*
+ * The buffer could have IO in progress by another scan. Right now
+ * localbuf.c doesn't use IO_IN_PROGRESS, which is why we need this
+ * hack.
+ *
+ * TODO: localbuf.c should use IO_IN_PROGRESS / have an equivalent of
+ * StartBufferIO().
+ */
+ if (pgaio_io_ref_valid(&bufHdr->io_in_progress))
+ {
+ PgAioHandleRef ior = bufHdr->io_in_progress;
+
+ ereport(DEBUG3,
+ errmsg("waiting for temp buffer IO in CSIO"),
+ errhidestmt(true), errhidecontext(true));
+ pgaio_io_ref_wait(&ior);
+ return false;
+ }
+
return (pg_atomic_read_u32(&bufHdr->state) & BM_VALID) == 0;
}
else
@@ -1412,13 +1462,38 @@ WaitReadBuffersCanStartIO(Buffer buffer, bool nowait)
void
WaitReadBuffers(ReadBuffersOperation *operation)
{
- Buffer *buffers;
+ IOContext io_context;
+ IOObject io_object;
int nblocks;
- BlockNumber blocknum;
- ForkNumber forknum;
- IOContext io_context;
- IOObject io_object;
- char persistence;
+ bool have_retryable_failure;
+
+ /*
+ * If we get here without any IO operations having been issued, the
+ * io_method == IOMETHOD_SYNC path must have been used. In that case, we
+ * start - as we used to before - the IO now, just before waiting.
+ */
+ if (operation->nios == 0)
+ {
+ Assert(io_method == IOMETHOD_SYNC);
+ if (!AsyncReadBuffers(operation, operation->io_buffers_len))
+ {
+ /* all blocks were already read in concurrently */
+ return;
+ }
+ }
+
+ if (operation->persistence == RELPERSISTENCE_TEMP)
+ {
+ io_context = IOCONTEXT_NORMAL;
+ io_object = IOOBJECT_TEMP_RELATION;
+ }
+ else
+ {
+ io_context = IOContextForStrategy(operation->strategy);
+ io_object = IOOBJECT_RELATION;
+ }
+
+restart:
/*
* Currently operations are only allowed to include a read of some range,
@@ -1433,15 +1508,101 @@ WaitReadBuffers(ReadBuffersOperation *operation)
if (nblocks == 0)
return; /* nothing to do */
- buffers = &operation->buffers[0];
- blocknum = operation->blocknum;
- forknum = operation->forknum;
- persistence = operation->persistence;
+ Assert(operation->nios > 0);
+ /*
+ * For IO timing we just count the time spent waiting for the IO.
+ *
+ * XXX: We probably should track the IO operation, rather than its time,
+ * separately, when initiating the IO. But right now that's not quite
+ * allowed by the interface.
+ */
+ have_retryable_failure = false;
+ for (int i = 0; i < operation->nios; i++)
+ {
+ PgAioReturn *aio_ret = &operation->returns[i];
+
+ /*
+ * Tracking a wait even if we don't actually need to wait a) is not
+ * cheap b) reports some time as waiting, even if we never waited.
+ */
+ if (aio_ret->result.status == ARS_UNKNOWN &&
+ !pgaio_io_ref_check_done(&operation->refs[i]))
+ {
+ instr_time io_start = pgstat_prepare_io_time(track_io_timing);
+
+ pgaio_io_ref_wait(&operation->refs[i]);
+
+ /*
+ * The IO operation itself was already counted earlier, in
+ * AsyncReadBuffers().
+ */
+ pgstat_count_io_op_time(io_object, io_context, IOOP_READ, io_start,
+ 0);
+ }
+ else
+ {
+ Assert(pgaio_io_ref_check_done(&operation->refs[i]));
+ }
+
+ if (aio_ret->result.status == ARS_PARTIAL)
+ {
+ /*
+ * We'll retry below, so we just emit a debug message the server
+ * log (or not even that in prod scenarios).
+ */
+ pgaio_result_log(aio_ret->result, &aio_ret->subject_data, DEBUG1);
+ have_retryable_failure = true;
+ }
+ else if (aio_ret->result.status != ARS_OK)
+ pgaio_result_log(aio_ret->result, &aio_ret->subject_data, ERROR);
+ }
+
+ /*
+ * If any of the associated IOs failed, try again to issue IOs. Buffers
+ * for which IO has completed successfully will be discovered as such and
+ * not retried.
+ */
+ if (have_retryable_failure)
+ {
+ nblocks = operation->io_buffers_len;
+
+ elog(DEBUG3, "retrying IO after partial failure");
+ CHECK_FOR_INTERRUPTS();
+ AsyncReadBuffers(operation, nblocks);
+ goto restart;
+ }
+
+ if (VacuumCostActive)
+ VacuumCostBalance += VacuumCostPageMiss * nblocks;
+
+ /* FIXME: READ_DONE tracepoint */
+}
+
+static bool
+AsyncReadBuffers(ReadBuffersOperation *operation,
+ int nblocks)
+{
+ int io_buffers_len = 0;
+ Buffer *buffers = &operation->buffers[0];
+ int flags = operation->flags;
+ BlockNumber blocknum = operation->blocknum;
+ ForkNumber forknum = operation->forknum;
+ IOContext io_context;
+ IOObject io_object;
+ char persistence;
+ bool did_start_io_overall = false;
+ PgAioHandle *ioh = NULL;
+ uint32 ioh_flags = 0;
+
+ persistence = operation->rel
+ ? operation->rel->rd_rel->relpersistence
+ : RELPERSISTENCE_PERMANENT;
if (persistence == RELPERSISTENCE_TEMP)
{
io_context = IOCONTEXT_NORMAL;
io_object = IOOBJECT_TEMP_RELATION;
+ ioh_flags |= AHF_REFERENCES_LOCAL;
}
else
{
@@ -1449,6 +1610,16 @@ WaitReadBuffers(ReadBuffersOperation *operation)
io_object = IOOBJECT_RELATION;
}
+ /*
+ * When this IO is executed synchronously, either because the caller will
+ * immediately block waiting for the IO or because IOMETHOD_SYNC is used,
+ * the AIO subsystem needs to know.
+ */
+ if (flags & READ_BUFFERS_SYNCHRONOUSLY)
+ ioh_flags |= AHF_SYNCHRONOUS;
+
+ operation->nios = 0;
+
/*
* We count all these blocks as read by this backend. This is traditional
* behavior, but might turn out to be not true if we find that someone
@@ -1464,19 +1635,38 @@ WaitReadBuffers(ReadBuffersOperation *operation)
for (int i = 0; i < nblocks; ++i)
{
- int io_buffers_len;
- Buffer io_buffers[MAX_IO_COMBINE_LIMIT];
void *io_pages[MAX_IO_COMBINE_LIMIT];
- instr_time io_start;
+ Buffer io_buffers[MAX_IO_COMBINE_LIMIT];
BlockNumber io_first_block;
+ bool did_start_io_this = false;
/*
- * Skip this block if someone else has already completed it. If an
- * I/O is already in progress in another backend, this will wait for
- * the outcome: either done, or something went wrong and we will
- * retry.
+ * Get IO before ReadBuffersCanStartIO, as pgaio_io_get() might block,
+ * which we don't want after setting IO_IN_PROGRESS.
+ *
+ * XXX: Should we attribute the time spent in here to the IO? If there
+ * already are a lot of IO operations in progress, getting an IO
+ * handle will block waiting for some other IO operation to finish.
+ *
+ * In most cases it'll be free to get the IO, so a timer would be
+ * overhead. Perhaps we should use pgaio_io_get_nb() and only account
+ * IO time when pgaio_io_get_nb() returned false?
*/
- if (!WaitReadBuffersCanStartIO(buffers[i], false))
+ if (likely(!ioh))
+ ioh = pgaio_io_get(CurrentResourceOwner, &operation->returns[operation->nios]);
+
+ /*
+ * Skip this block if someone else has already completed it.
+ *
+ * If an I/O is already in progress in another backend, this will wait
+ * for the outcome: either done, or something went wrong and we will
+ * retry. But don't wait if we have staged, but haven't issued,
+ * another IO.
+ *
+ * XXX: If we can't start IO due to unsubmitted IO, it might be worth
+ * to submit and then try to start IO again.
+ */
+ if (!ReadBuffersCanStartIO(buffers[i], did_start_io_overall))
{
/*
* Report this as a 'hit' for this backend, even though it must
@@ -1488,6 +1678,11 @@ WaitReadBuffers(ReadBuffersOperation *operation)
operation->smgr->smgr_rlocator.locator.relNumber,
operation->smgr->smgr_rlocator.backend,
true);
+
+ ereport(DEBUG3,
+ errmsg("can't start io for first buffer %u: %s",
+ buffers[i], DebugPrintBufferRefcount(buffers[i])),
+ errhidestmt(true), errhidecontext(true));
continue;
}
@@ -1497,6 +1692,11 @@ WaitReadBuffers(ReadBuffersOperation *operation)
io_first_block = blocknum + i;
io_buffers_len = 1;
+ ereport(DEBUG5,
+ errmsg("first prepped for io: %s, offset %d",
+ DebugPrintBufferRefcount(io_buffers[0]), i),
+ errhidestmt(true), errhidecontext(true));
+
/*
* How many neighboring-on-disk blocks can we scatter-read into other
* buffers at the same time? In this case we don't wait if we see an
@@ -1505,85 +1705,57 @@ WaitReadBuffers(ReadBuffersOperation *operation)
* We'll come back to this block again, above.
*/
while ((i + 1) < nblocks &&
- WaitReadBuffersCanStartIO(buffers[i + 1], true))
+ ReadBuffersCanStartIO(buffers[i + 1], true))
{
/* Must be consecutive block numbers. */
Assert(BufferGetBlockNumber(buffers[i + 1]) ==
BufferGetBlockNumber(buffers[i]) + 1);
+ ereport(DEBUG5,
+ errmsg("seq prepped for io: %s, offset %d",
+ DebugPrintBufferRefcount(buffers[i + 1]),
+ i + 1),
+ errhidestmt(true), errhidecontext(true));
+
io_buffers[io_buffers_len] = buffers[++i];
io_pages[io_buffers_len++] = BufferGetBlock(buffers[i]);
}
- io_start = pgstat_prepare_io_time(track_io_timing);
- smgrreadv(operation->smgr, forknum, io_first_block, io_pages, io_buffers_len);
- pgstat_count_io_op_time(io_object, io_context, IOOP_READ, io_start,
- io_buffers_len);
+ pgaio_io_get_ref(ioh, &operation->refs[operation->nios]);
- /* Verify each block we read, and terminate the I/O. */
- for (int j = 0; j < io_buffers_len; ++j)
- {
- BufferDesc *bufHdr;
- Block bufBlock;
+ pgaio_io_set_io_data_32(ioh, (uint32 *) io_buffers, io_buffers_len);
- if (persistence == RELPERSISTENCE_TEMP)
- {
- bufHdr = GetLocalBufferDescriptor(-io_buffers[j] - 1);
- bufBlock = LocalBufHdrGetBlock(bufHdr);
- }
- else
- {
- bufHdr = GetBufferDescriptor(io_buffers[j] - 1);
- bufBlock = BufHdrGetBlock(bufHdr);
- }
- /* check for garbage data */
- if (!PageIsVerifiedExtended((Page) bufBlock, io_first_block + j,
- PIV_LOG_WARNING | PIV_REPORT_STAT))
- {
- if ((operation->flags & READ_BUFFERS_ZERO_ON_ERROR) || zero_damaged_pages)
- {
- ereport(WARNING,
- (errcode(ERRCODE_DATA_CORRUPTED),
- errmsg("invalid page in block %u of relation %s; zeroing out page",
- io_first_block + j,
- relpath(operation->smgr->smgr_rlocator, forknum))));
- memset(bufBlock, 0, BLCKSZ);
- }
- else
- ereport(ERROR,
- (errcode(ERRCODE_DATA_CORRUPTED),
- errmsg("invalid page in block %u of relation %s",
- io_first_block + j,
- relpath(operation->smgr->smgr_rlocator, forknum))));
- }
+ if (persistence == RELPERSISTENCE_TEMP)
+ pgaio_io_add_shared_cb(ioh, ASC_LOCAL_BUFFER_READ);
+ else
+ pgaio_io_add_shared_cb(ioh, ASC_SHARED_BUFFER_READ);
- /* Terminate I/O and set BM_VALID. */
- if (persistence == RELPERSISTENCE_TEMP)
- {
- uint32 buf_state = pg_atomic_read_u32(&bufHdr->state);
+ pgaio_io_set_flag(ioh, ioh_flags);
- buf_state |= BM_VALID;
- pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
- }
- else
- {
- /* Set BM_VALID, terminate IO, and wake up any waiters */
- TerminateBufferIO(bufHdr, false, BM_VALID, true, true);
- }
+ did_start_io_overall = did_start_io_this = true;
+ smgrstartreadv(ioh, operation->smgr, forknum, io_first_block,
+ io_pages, io_buffers_len);
+ ioh = NULL;
+ operation->nios++;
- /* Report I/Os as completing individually. */
- TRACE_POSTGRESQL_BUFFER_READ_DONE(forknum, io_first_block + j,
- operation->smgr->smgr_rlocator.locator.spcOid,
- operation->smgr->smgr_rlocator.locator.dbOid,
- operation->smgr->smgr_rlocator.locator.relNumber,
- operation->smgr->smgr_rlocator.backend,
- false);
- }
+ /* not obvious what we'd use for time */
+ pgstat_count_io_op_n(io_object, io_context, IOOP_READ, io_buffers_len);
+ }
+
+ if (ioh)
+ {
+ pgaio_io_release(ioh);
+ ioh = NULL;
+ }
- if (VacuumCostActive)
- VacuumCostBalance += VacuumCostPageMiss * io_buffers_len;
+ if (did_start_io_overall)
+ {
+ pgaio_submit_staged();
+ return true;
}
+ else
+ return false;
}
/*
@@ -6367,7 +6539,7 @@ shared_buffer_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
prior_result.status == ARS_ERROR
|| prior_result.result <= io_data_off;
- elog(DEBUG3, "calling rbcrs for buf %d with failed %d, error: %d, result: %d, data_off: %d",
+ elog(DEBUG5, "calling rbcrs for buf %d with failed %d, error: %d, result: %d, data_off: %d",
buf, failed, prior_result.status, prior_result.result, io_data_off);
/*
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0013-aio-Very-WIP-read_stream.c-adjustments-for-real-A.patch (4.9K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/14-v2-0013-aio-Very-WIP-read_stream.c-adjustments-for-real-A.patch)
download | inline diff:
From b7123290712da81631ecfbfb2437b95eb42a8e9c Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sat, 31 Aug 2024 21:39:30 -0400
Subject: [PATCH v2 13/20] aio: Very-WIP: read_stream.c adjustments for real
AIO
---
src/include/storage/bufmgr.h | 2 ++
src/backend/storage/aio/read_stream.c | 31 +++++++++++++++++++++------
src/backend/storage/buffer/bufmgr.c | 3 ++-
3 files changed, 28 insertions(+), 8 deletions(-)
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 7a12ef6e9be..2a836cf98c6 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -119,6 +119,8 @@ typedef struct BufferManagerRelation
#define READ_BUFFERS_ISSUE_ADVICE (1 << 1)
/* IO will immediately be waited for */
#define READ_BUFFERS_SYNCHRONOUSLY (1 << 2)
+/* caller will issue more io, don't submit */
+#define READ_BUFFERS_MORE_MORE_MORE (1 << 3)
/*
* FIXME: PgAioReturn is defined in aio.h. It'd be much better if we didn't
diff --git a/src/backend/storage/aio/read_stream.c b/src/backend/storage/aio/read_stream.c
index 3d30e6224f7..5b5bae16c44 100644
--- a/src/backend/storage/aio/read_stream.c
+++ b/src/backend/storage/aio/read_stream.c
@@ -90,6 +90,7 @@
#include "postgres.h"
#include "miscadmin.h"
+#include "storage/aio.h"
#include "storage/fd.h"
#include "storage/smgr.h"
#include "storage/read_stream.h"
@@ -240,14 +241,18 @@ read_stream_start_pending_read(ReadStream *stream, bool suppress_advice)
/*
* If advice hasn't been suppressed, this system supports it, and this
* isn't a strictly sequential pattern, then we'll issue advice.
+ *
+ * XXX: Used to also check stream->pending_read_blocknum !=
+ * stream->seq_blocknum
*/
if (!suppress_advice &&
- stream->advice_enabled &&
- stream->pending_read_blocknum != stream->seq_blocknum)
+ stream->advice_enabled)
flags = READ_BUFFERS_ISSUE_ADVICE;
else
flags = 0;
+ flags |= READ_BUFFERS_MORE_MORE_MORE;
+
/* We say how many blocks we want to read, but may be smaller on return. */
buffer_index = stream->next_buffer_index;
io_index = stream->next_io_index;
@@ -306,6 +311,14 @@ read_stream_start_pending_read(ReadStream *stream, bool suppress_advice)
static void
read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
{
+ if (stream->distance > (io_combine_limit * 8))
+ {
+ if (stream->pinned_buffers + stream->pending_read_nblocks > ((stream->distance * 3) / 4))
+ {
+ return;
+ }
+ }
+
while (stream->ios_in_progress < stream->max_ios &&
stream->pinned_buffers + stream->pending_read_nblocks < stream->distance)
{
@@ -355,6 +368,7 @@ read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
{
/* And we've hit the limit. Rewind, and stop here. */
read_stream_unget_block(stream, blocknum);
+ pgaio_submit_staged();
return;
}
}
@@ -379,6 +393,8 @@ read_stream_look_ahead(ReadStream *stream, bool suppress_advice)
stream->distance == 0) &&
stream->ios_in_progress < stream->max_ios)
read_stream_start_pending_read(stream, suppress_advice);
+
+ pgaio_submit_staged();
}
/*
@@ -442,7 +458,7 @@ read_stream_begin_impl(int flags,
* overflow (even though that's not possible with the current GUC range
* limits), allowing also for the spare entry and the overflow space.
*/
- max_pinned_buffers = Max(max_ios * 4, io_combine_limit);
+ max_pinned_buffers = Max(max_ios * io_combine_limit, io_combine_limit);
max_pinned_buffers = Min(max_pinned_buffers,
PG_INT16_MAX - io_combine_limit - 1);
@@ -493,10 +509,11 @@ read_stream_begin_impl(int flags,
* direct I/O isn't enabled, the caller hasn't promised sequential access
* (overriding our detection heuristics), and max_ios hasn't been set to
* zero.
+ *
+ * FIXME: Used to also check (io_direct_flags & IO_DIRECT_DATA) == 0 &&
+ * (flags & READ_STREAM_SEQUENTIAL) == 0
*/
- if ((io_direct_flags & IO_DIRECT_DATA) == 0 &&
- (flags & READ_STREAM_SEQUENTIAL) == 0 &&
- max_ios > 0)
+ if (max_ios > 0)
stream->advice_enabled = true;
#endif
@@ -727,7 +744,7 @@ read_stream_next_buffer(ReadStream *stream, void **per_buffer_data)
if (++stream->oldest_io_index == stream->max_ios)
stream->oldest_io_index = 0;
- if (stream->ios[io_index].op.flags & READ_BUFFERS_ISSUE_ADVICE)
+ if (stream->ios[io_index].op.flags & (READ_BUFFERS_ISSUE_ADVICE | READ_BUFFERS_MORE_MORE_MORE))
{
/* Distance ramps up fast (behavior C). */
distance = stream->distance * 2;
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 89cb7b41b03..722e73eb7d0 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -1751,7 +1751,8 @@ AsyncReadBuffers(ReadBuffersOperation *operation,
if (did_start_io_overall)
{
- pgaio_submit_staged();
+ if (!(flags & READ_BUFFERS_MORE_MORE_MORE))
+ pgaio_submit_staged();
return true;
}
else
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0014-aio-Add-bounce-buffers.patch (20.6K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/15-v2-0014-aio-Add-bounce-buffers.patch)
download | inline diff:
From c1a5b7c868eb962a3e1e5348aa6309aa1005f4eb Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 25 Nov 2024 16:35:15 -0500
Subject: [PATCH v2 14/20] aio: Add bounce buffers
---
src/include/storage/aio.h | 18 ++
src/include/storage/aio_internal.h | 33 ++++
src/include/utils/resowner.h | 2 +
src/backend/storage/aio/README.md | 27 +++
src/backend/storage/aio/aio.c | 182 ++++++++++++++++++
src/backend/storage/aio/aio_init.c | 118 ++++++++++++
src/backend/utils/misc/guc_tables.c | 13 ++
src/backend/utils/misc/postgresql.conf.sample | 2 +
src/backend/utils/resowner/resowner.c | 25 ++-
src/tools/pgindent/typedefs.list | 1 +
10 files changed, 419 insertions(+), 2 deletions(-)
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index ff44dac5bb2..1bef475b0a9 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -222,6 +222,9 @@ typedef struct PgAioHandleSharedCallbacks
+typedef struct PgAioBounceBuffer PgAioBounceBuffer;
+
+
/*
* How many callbacks can be registered for one IO handle. Currently we only
* need two, but it's not hard to imagine needing a few more.
@@ -294,6 +297,20 @@ extern void pgaio_result_log(PgAioResult result, const PgAioSubjectData *subject
+/* --------------------------------------------------------------------------------
+ * Bounce Buffers
+ * --------------------------------------------------------------------------------
+ */
+
+extern PgAioBounceBuffer *pgaio_bounce_buffer_get(void);
+extern void pgaio_io_assoc_bounce_buffer(PgAioHandle *ioh, PgAioBounceBuffer *bb);
+extern uint32 pgaio_bounce_buffer_id(PgAioBounceBuffer *bb);
+extern void pgaio_bounce_buffer_release(PgAioBounceBuffer *bb);
+extern char *pgaio_bounce_buffer_buffer(PgAioBounceBuffer *bb);
+extern void pgaio_bounce_buffer_release_resowner(dlist_node *bb_node, bool on_error);
+
+
+
/* --------------------------------------------------------------------------------
* Actions on multiple IOs.
* --------------------------------------------------------------------------------
@@ -354,6 +371,7 @@ typedef enum IoMethod
extern const struct config_enum_entry io_method_options[];
extern int io_method;
extern int io_max_concurrency;
+extern int io_bounce_buffers;
#endif /* AIO_H */
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
index d2dc1516bdf..2065bde79c3 100644
--- a/src/include/storage/aio_internal.h
+++ b/src/include/storage/aio_internal.h
@@ -91,6 +91,12 @@ struct PgAioHandle
/* index into PgAioCtl->iovecs */
uint32 iovec_off;
+ /*
+ * List of bounce_buffers owned by IO. It would suffice to use an index
+ * based linked list here.
+ */
+ slist_head bounce_buffers;
+
/**
* In which list the handle is registered, depends on the state:
* - IDLE, in per-backend list
@@ -130,11 +136,23 @@ struct PgAioHandle
};
+struct PgAioBounceBuffer
+{
+ slist_node node;
+ struct ResourceOwnerData *resowner;
+ dlist_node resowner_node;
+ char *buffer;
+};
+
+
typedef struct PgAioPerBackend
{
/* index into PgAioCtl->io_handles */
uint32 io_handle_off;
+ /* index into PgAioCtl->bounce_buffers */
+ uint32 bounce_buffers_off;
+
/* IO Handles that currently are not used */
dclist_head idle_ios;
@@ -162,6 +180,12 @@ typedef struct PgAioPerBackend
* IOs being appended at the end.
*/
dclist_head in_flight_ios;
+
+ /* Bounce Buffers that currently are not used */
+ slist_head idle_bbs;
+
+ /* see handed_out_io */
+ PgAioBounceBuffer *handed_out_bb;
} PgAioPerBackend;
@@ -187,6 +211,15 @@ typedef struct PgAioCtl
*/
uint64 *iovecs_data;
+ /*
+ * To perform AIO on buffers that are not located in shared memory (either
+ * because they are not in shared memory or because we need to operate on
+ * a copy, as e.g. the case for writes when checksums are in use)
+ */
+ uint64 bounce_buffers_count;
+ PgAioBounceBuffer *bounce_buffers;
+ char *bounce_buffers_data;
+
uint64 io_handle_count;
PgAioHandle *io_handles;
} PgAioCtl;
diff --git a/src/include/utils/resowner.h b/src/include/utils/resowner.h
index 2d55720a54c..0cdd0c13ffb 100644
--- a/src/include/utils/resowner.h
+++ b/src/include/utils/resowner.h
@@ -168,5 +168,7 @@ extern void ResourceOwnerForgetLock(ResourceOwner owner, struct LOCALLOCK *local
struct dlist_node;
extern void ResourceOwnerRememberAioHandle(ResourceOwner owner, struct dlist_node *ioh_node);
extern void ResourceOwnerForgetAioHandle(ResourceOwner owner, struct dlist_node *ioh_node);
+extern void ResourceOwnerRememberAioBounceBuffer(ResourceOwner owner, struct dlist_node *bb_node);
+extern void ResourceOwnerForgetAioBounceBuffer(ResourceOwner owner, struct dlist_node *bb_node);
#endif /* RESOWNER_H */
diff --git a/src/backend/storage/aio/README.md b/src/backend/storage/aio/README.md
index 893f4ffe428..0076ea4aa10 100644
--- a/src/backend/storage/aio/README.md
+++ b/src/backend/storage/aio/README.md
@@ -395,6 +395,33 @@ shared memory no less!), completion callbacks instead have to encode errors in
a more compact format that can be converted into an error message.
+### AIO Bounce Buffers
+
+For some uses of AIO there is no convenient memory location as the source /
+destination of an AIO. E.g. when data checksums are enabled, writes from
+shared buffers currently cannot be done directly from shared buffers, as a
+shared buffer lock still allows some modification, e.g., for hint bits(see
+`FlushBuffer()`). If the write were done in-place, such modifications can
+cause the checksum to fail.
+
+For synchronous IO this is solved by copying the buffer to separate memory
+before computing the checksum and using that copy as the source buffer for the
+AIO.
+
+However, for AIO that is not a workable solution:
+- Instead of a single buffer many buffers are required, as many IOs might be
+ in flight
+- When using the [worker method](#worker), the source/target of IO needs to be
+ in shared memory, otherwise the workers won't be able to access the memory.
+
+The AIO subsystem addresses this by providing a limited number of bounce
+buffers that can be used as the source / target for IO. A bounce buffer be
+acquired with `pgaio_bounce_buffer_get()` and multiple bounce buffers can be
+associated with an AIO Handle with `pgaio_io_assoc_bounce_buffer()`.
+
+Bounce buffers are automatically released when the IO completes.
+
+
## Helpers
Using the low-level AIO API introduces too much complexity to do so all over
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index 2439ce3740d..e829e1752ca 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -54,6 +54,8 @@ static void pgaio_io_resowner_register(PgAioHandle *ioh);
static void pgaio_io_wait_for_free(void);
static PgAioHandle *pgaio_io_from_ref(PgAioHandleRef *ior, uint64 *ref_generation);
+static void pgaio_bounce_buffer_wait_for_free(void);
+
/* Options for io_method. */
@@ -68,6 +70,7 @@ const struct config_enum_entry io_method_options[] = {
int io_method = DEFAULT_IO_METHOD;
int io_max_concurrency = -1;
+int io_bounce_buffers = -1;
/* global control for AIO */
@@ -732,6 +735,21 @@ pgaio_io_reclaim(PgAioHandle *ioh)
}
}
+ /* reclaim all associated bounce buffers */
+ if (!slist_is_empty(&ioh->bounce_buffers))
+ {
+ slist_mutable_iter it;
+
+ slist_foreach_modify(it, &ioh->bounce_buffers)
+ {
+ PgAioBounceBuffer *bb = slist_container(PgAioBounceBuffer, node, it.cur);
+
+ slist_delete_current(&it);
+
+ slist_push_head(&my_aio->idle_bbs, &bb->node);
+ }
+ }
+
if (ioh->resowner)
{
ResourceOwnerForgetAioHandle(ioh->resowner, &ioh->resowner_node);
@@ -855,6 +873,168 @@ pgaio_io_wait_for_free(void)
+/* --------------------------------------------------------------------------------
+ * Bounce Buffers
+ * --------------------------------------------------------------------------------
+ */
+
+PgAioBounceBuffer *
+pgaio_bounce_buffer_get(void)
+{
+ PgAioBounceBuffer *bb = NULL;
+ slist_node *node;
+
+ if (my_aio->handed_out_bb != NULL)
+ elog(ERROR, "can only hand out one BB");
+
+ /*
+ * FIXME It probably is not correct to have bounce buffers be per backend,
+ * they use too much memory.
+ */
+ if (slist_is_empty(&my_aio->idle_bbs))
+ {
+ pgaio_bounce_buffer_wait_for_free();
+ }
+
+ node = slist_pop_head_node(&my_aio->idle_bbs);
+ bb = slist_container(PgAioBounceBuffer, node, node);
+
+ my_aio->handed_out_bb = bb;
+
+ bb->resowner = CurrentResourceOwner;
+ ResourceOwnerRememberAioBounceBuffer(bb->resowner, &bb->resowner_node);
+
+ return bb;
+}
+
+void
+pgaio_io_assoc_bounce_buffer(PgAioHandle *ioh, PgAioBounceBuffer *bb)
+{
+ if (my_aio->handed_out_bb != bb)
+ elog(ERROR, "can only assign handed out BB");
+ my_aio->handed_out_bb = NULL;
+
+ /*
+ * There can be many bounce buffers assigned in case of vectorized IOs.
+ */
+ slist_push_head(&ioh->bounce_buffers, &bb->node);
+
+ /* once associated with an IO, the IO has ownership */
+ ResourceOwnerForgetAioBounceBuffer(bb->resowner, &bb->resowner_node);
+ bb->resowner = NULL;
+}
+
+uint32
+pgaio_bounce_buffer_id(PgAioBounceBuffer *bb)
+{
+ return bb - aio_ctl->bounce_buffers;
+}
+
+void
+pgaio_bounce_buffer_release(PgAioBounceBuffer *bb)
+{
+ if (my_aio->handed_out_bb != bb)
+ elog(ERROR, "can only release handed out BB");
+
+ slist_push_head(&my_aio->idle_bbs, &bb->node);
+ my_aio->handed_out_bb = NULL;
+
+ ResourceOwnerForgetAioBounceBuffer(bb->resowner, &bb->resowner_node);
+ bb->resowner = NULL;
+}
+
+void
+pgaio_bounce_buffer_release_resowner(dlist_node *bb_node, bool on_error)
+{
+ PgAioBounceBuffer *bb = dlist_container(PgAioBounceBuffer, resowner_node, bb_node);
+
+ Assert(bb->resowner);
+
+ if (!on_error)
+ elog(WARNING, "leaked AIO bounce buffer");
+
+ pgaio_bounce_buffer_release(bb);
+}
+
+char *
+pgaio_bounce_buffer_buffer(PgAioBounceBuffer *bb)
+{
+ return bb->buffer;
+}
+
+static void
+pgaio_bounce_buffer_wait_for_free(void)
+{
+ static uint32 lastpos = 0;
+
+ if (my_aio->num_staged_ios > 0)
+ {
+ elog(DEBUG2, "submitting while acquiring free bb");
+ pgaio_submit_staged();
+ }
+
+ for (uint32 i = lastpos; i < lastpos + io_max_concurrency; i++)
+ {
+ uint32 thisoff = my_aio->io_handle_off + (i % io_max_concurrency);
+ PgAioHandle *ioh = &aio_ctl->io_handles[thisoff];
+
+ switch (ioh->state)
+ {
+ case AHS_IDLE:
+ case AHS_HANDED_OUT:
+ continue;
+ case AHS_DEFINED: /* should have been submitted above */
+ case AHS_PREPARED:
+ elog(ERROR, "shouldn't get here with io:%d in state %d",
+ pgaio_io_get_id(ioh), ioh->state);
+ break;
+ case AHS_REAPED:
+ case AHS_IN_FLIGHT:
+ if (!slist_is_empty(&ioh->bounce_buffers))
+ {
+ PgAioHandleRef ior;
+
+ ior.aio_index = ioh - aio_ctl->io_handles;
+ ior.generation_upper = (uint32) (ioh->generation >> 32);
+ ior.generation_lower = (uint32) ioh->generation;
+
+ pgaio_io_ref_wait(&ior);
+ elog(DEBUG2, "waited for io:%d to reclaim BB",
+ pgaio_io_get_id(ioh));
+
+ if (slist_is_empty(&my_aio->idle_bbs))
+ elog(WARNING, "empty after wait");
+
+ if (!slist_is_empty(&my_aio->idle_bbs))
+ {
+ lastpos = i;
+ return;
+ }
+ }
+ break;
+ case AHS_COMPLETED_SHARED:
+ case AHS_COMPLETED_LOCAL:
+ /* reclaim */
+ pgaio_io_reclaim(ioh);
+
+ if (!slist_is_empty(&my_aio->idle_bbs))
+ {
+ lastpos = i;
+ return;
+ }
+ break;
+ }
+ }
+
+ /*
+ * The submission above could have caused the IO to complete at any time.
+ */
+ if (slist_is_empty(&my_aio->idle_bbs))
+ elog(PANIC, "no more bbs");
+}
+
+
+
/* --------------------------------------------------------------------------------
* Actions on multiple IOs.
* --------------------------------------------------------------------------------
@@ -929,6 +1109,7 @@ void
pgaio_at_xact_end(bool is_subxact, bool is_commit)
{
Assert(!my_aio->handed_out_io);
+ Assert(!my_aio->handed_out_bb);
}
/*
@@ -939,6 +1120,7 @@ void
pgaio_at_error(void)
{
Assert(!my_aio->handed_out_io);
+ Assert(!my_aio->handed_out_bb);
}
diff --git a/src/backend/storage/aio/aio_init.c b/src/backend/storage/aio/aio_init.c
index 23adc5308e5..417526f3823 100644
--- a/src/backend/storage/aio/aio_init.c
+++ b/src/backend/storage/aio/aio_init.c
@@ -82,6 +82,32 @@ AioIOVDataShmemSize(void)
io_max_concurrency));
}
+static Size
+AioBounceBufferDescShmemSize(void)
+{
+ Size sz;
+
+ /* PgAioBounceBuffer itself */
+ sz = mul_size(sizeof(PgAioBounceBuffer),
+ mul_size(AioProcs(), io_bounce_buffers));
+
+ return sz;
+}
+
+static Size
+AioBounceBufferDataShmemSize(void)
+{
+ Size sz;
+
+ /* and the associated buffer */
+ sz = mul_size(BLCKSZ,
+ mul_size(io_bounce_buffers, AioProcs()));
+ /* memory for alignment */
+ sz += BLCKSZ;
+
+ return sz;
+}
+
/*
* Choose a suitable value for io_max_concurrency.
*
@@ -107,6 +133,33 @@ AioChooseMaxConccurrency(void)
return Min(max_proportional_pins, 64);
}
+/*
+ * Choose a suitable value for io_bounce_buffers.
+ *
+ * It's very unlikely to be useful to allocate more bounce buffers for each
+ * backend than the backend is allowed to pin. Additionally, bounce buffers
+ * currently are used for writes, it seems very uncommon for more than 10% of
+ * shared_buffers to be written out concurrently.
+ *
+ * XXX: This quickly can take up significant amounts of memory, the logic
+ * should probably fine tuned.
+ */
+static int
+AioChooseBounceBuffers(void)
+{
+ uint32 max_backends;
+ int max_proportional_pins;
+
+ /* Similar logic to LimitAdditionalPins() */
+ max_backends = MaxBackends + NUM_AUXILIARY_PROCS;
+ max_proportional_pins = (NBuffers / 10) / max_backends;
+
+ max_proportional_pins = Max(max_proportional_pins, 1);
+
+ /* apply upper limit */
+ return Min(max_proportional_pins, 256);
+}
+
Size
AioShmemSize(void)
{
@@ -130,11 +183,31 @@ AioShmemSize(void)
PGC_S_OVERRIDE);
}
+
+ /*
+ * If io_bounce_buffers is -1, we automatically choose a suitable value.
+ *
+ * See also comment above.
+ */
+ if (io_bounce_buffers == -1)
+ {
+ char buf[32];
+
+ snprintf(buf, sizeof(buf), "%d", AioChooseBounceBuffers());
+ SetConfigOption("io_bounce_buffers", buf, PGC_POSTMASTER,
+ PGC_S_DYNAMIC_DEFAULT);
+ if (io_bounce_buffers == -1) /* failed to apply it? */
+ SetConfigOption("io_bounce_buffers", buf, PGC_POSTMASTER,
+ PGC_S_OVERRIDE);
+ }
+
sz = add_size(sz, AioCtlShmemSize());
sz = add_size(sz, AioBackendShmemSize());
sz = add_size(sz, AioHandleShmemSize());
sz = add_size(sz, AioIOVShmemSize());
sz = add_size(sz, AioIOVDataShmemSize());
+ sz = add_size(sz, AioBounceBufferDescShmemSize());
+ sz = add_size(sz, AioBounceBufferDataShmemSize());
if (pgaio_impl->shmem_size)
sz = add_size(sz, pgaio_impl->shmem_size());
@@ -148,7 +221,10 @@ AioShmemInit(void)
bool found;
uint32 io_handle_off = 0;
uint32 iovec_off = 0;
+ uint32 bounce_buffers_off = 0;
uint32 per_backend_iovecs = io_max_concurrency * io_combine_limit;
+ uint32 per_backend_bb = io_bounce_buffers;
+ char *bounce_buffers_data;
aio_ctl = (PgAioCtl *)
ShmemInitStruct("AioCtl", AioCtlShmemSize(), &found);
@@ -160,6 +236,7 @@ AioShmemInit(void)
aio_ctl->io_handle_count = AioProcs() * io_max_concurrency;
aio_ctl->iovec_count = AioProcs() * per_backend_iovecs;
+ aio_ctl->bounce_buffers_count = AioProcs() * per_backend_bb;
aio_ctl->backend_state = (PgAioPerBackend *)
ShmemInitStruct("AioBackend", AioBackendShmemSize(), &found);
@@ -170,6 +247,35 @@ AioShmemInit(void)
aio_ctl->iovecs = ShmemInitStruct("AioIOV", AioIOVShmemSize(), &found);
aio_ctl->iovecs_data = ShmemInitStruct("AioIOVData", AioIOVDataShmemSize(), &found);
+ aio_ctl->bounce_buffers = ShmemInitStruct("AioBounceBufferDesc", AioBounceBufferDescShmemSize(), &found);
+
+ bounce_buffers_data = ShmemInitStruct("AioBounceBufferData", AioBounceBufferDataShmemSize(), &found);
+ bounce_buffers_data = (char *) TYPEALIGN(BLCKSZ, (uintptr_t) bounce_buffers_data);
+ aio_ctl->bounce_buffers_data = bounce_buffers_data;
+
+
+ /* Initialize IO handles. */
+ for (uint64 i = 0; i < aio_ctl->io_handle_count; i++)
+ {
+ PgAioHandle *ioh = &aio_ctl->io_handles[i];
+
+ ioh->op = PGAIO_OP_INVALID;
+ ioh->subject = ASI_INVALID;
+ ioh->state = AHS_IDLE;
+
+ slist_init(&ioh->bounce_buffers);
+ }
+
+ /* Initialize Bounce Buffers. */
+ for (uint64 i = 0; i < aio_ctl->bounce_buffers_count; i++)
+ {
+ PgAioBounceBuffer *bb = &aio_ctl->bounce_buffers[i];
+
+ bb->buffer = bounce_buffers_data;
+ bounce_buffers_data += BLCKSZ;
+ }
+
+
for (int procno = 0; procno < AioProcs(); procno++)
{
PgAioPerBackend *bs = &aio_ctl->backend_state[procno];
@@ -177,9 +283,13 @@ AioShmemInit(void)
bs->io_handle_off = io_handle_off;
io_handle_off += io_max_concurrency;
+ bs->bounce_buffers_off = bounce_buffers_off;
+ bounce_buffers_off += per_backend_bb;
+
dclist_init(&bs->idle_ios);
memset(bs->staged_ios, 0, sizeof(PgAioHandle *) * PGAIO_SUBMIT_BATCH_SIZE);
dclist_init(&bs->in_flight_ios);
+ slist_init(&bs->idle_bbs);
/* initialize per-backend IOs */
for (int i = 0; i < io_max_concurrency; i++)
@@ -201,6 +311,14 @@ AioShmemInit(void)
dclist_push_tail(&bs->idle_ios, &ioh->node);
iovec_off += io_combine_limit;
}
+
+ /* initialize per-backend bounce buffers */
+ for (int i = 0; i < per_backend_bb; i++)
+ {
+ PgAioBounceBuffer *bb = &aio_ctl->bounce_buffers[bs->bounce_buffers_off + i];
+
+ slist_push_head(&bs->idle_bbs, &bb->node);
+ }
}
out:
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index b2999b86c24..39e91ebd2a5 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -3233,6 +3233,19 @@ struct config_int ConfigureNamesInt[] =
NULL, NULL, NULL
},
+ {
+ {"io_bounce_buffers",
+ PGC_POSTMASTER,
+ RESOURCES_ASYNCHRONOUS,
+ gettext_noop("Number of IO Bounce Buffers reserved for each backend."),
+ NULL,
+ GUC_UNIT_BLOCKS
+ },
+ &io_bounce_buffers,
+ -1, -1, 4096,
+ NULL, NULL, NULL
+ },
+
{
{"io_workers",
PGC_SIGHUP,
diff --git a/src/backend/utils/misc/postgresql.conf.sample b/src/backend/utils/misc/postgresql.conf.sample
index 5893eb29228..da6e248a29e 100644
--- a/src/backend/utils/misc/postgresql.conf.sample
+++ b/src/backend/utils/misc/postgresql.conf.sample
@@ -848,6 +848,8 @@
#io_max_concurrency = 32 # Max number of IOs that may be in
# flight at the same time in one backend
# (change requires restart)
+#io_bounce_buffers = -1 # -1 sets based on shared_buffers
+ # (change requires restart)
#------------------------------------------------------------------------------
diff --git a/src/backend/utils/resowner/resowner.c b/src/backend/utils/resowner/resowner.c
index 5cf14472ebd..d1932b7393c 100644
--- a/src/backend/utils/resowner/resowner.c
+++ b/src/backend/utils/resowner/resowner.c
@@ -159,10 +159,11 @@ struct ResourceOwnerData
LOCALLOCK *locks[MAX_RESOWNER_LOCKS]; /* list of owned locks */
/*
- * AIO handles need be registered in critical sections and therefore
- * cannot use the normal ResoureElem mechanism.
+ * AIO handles & bounce buffers need be registered in critical sections
+ * and therefore cannot use the normal ResoureElem mechanism.
*/
dlist_head aio_handles;
+ dlist_head aio_bounce_buffers;
};
@@ -434,6 +435,7 @@ ResourceOwnerCreate(ResourceOwner parent, const char *name)
}
dlist_init(&owner->aio_handles);
+ dlist_init(&owner->aio_bounce_buffers);
return owner;
}
@@ -743,6 +745,13 @@ ResourceOwnerReleaseInternal(ResourceOwner owner,
pgaio_io_release_resowner(node, !isCommit);
}
+
+ while (!dlist_is_empty(&owner->aio_bounce_buffers))
+ {
+ dlist_node *node = dlist_head_node(&owner->aio_bounce_buffers);
+
+ pgaio_bounce_buffer_release_resowner(node, !isCommit);
+ }
}
else if (phase == RESOURCE_RELEASE_LOCKS)
{
@@ -1112,3 +1121,15 @@ ResourceOwnerForgetAioHandle(ResourceOwner owner, struct dlist_node *ioh_node)
{
dlist_delete_from(&owner->aio_handles, ioh_node);
}
+
+void
+ResourceOwnerRememberAioBounceBuffer(ResourceOwner owner, struct dlist_node *ioh_node)
+{
+ dlist_push_tail(&owner->aio_bounce_buffers, ioh_node);
+}
+
+void
+ResourceOwnerForgetAioBounceBuffer(ResourceOwner owner, struct dlist_node *ioh_node)
+{
+ dlist_delete_from(&owner->aio_bounce_buffers, ioh_node);
+}
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index a5b12b48f99..dc52d6165d4 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2104,6 +2104,7 @@ Permutation
PermutationStep
PermutationStepBlocker
PermutationStepBlockerType
+PgAioBounceBuffer
PgAioCtl
PgAioHandle
PgAioHandleFlags
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0015-bufmgr-Implement-AIO-write-support.patch (5.7K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/16-v2-0015-bufmgr-Implement-AIO-write-support.patch)
download | inline diff:
From 40e15609a95f6733a7fe0e202c5ec4add3044bad Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sat, 31 Aug 2024 21:39:01 -0400
Subject: [PATCH v2 15/20] bufmgr: Implement AIO write support
As of this commit there are no users of these AIO facilities, that'll come in
later commits.
Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
src/include/storage/aio.h | 2 +
src/include/storage/bufmgr.h | 2 +
src/backend/storage/aio/aio_subject.c | 2 +
src/backend/storage/buffer/bufmgr.c | 85 +++++++++++++++++++++++++++
4 files changed, 91 insertions(+)
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index 1bef475b0a9..caa52d2aaba 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -106,8 +106,10 @@ typedef enum PgAioHandleSharedCallbackID
ASC_MD_WRITEV,
ASC_SHARED_BUFFER_READ,
+ ASC_SHARED_BUFFER_WRITE,
ASC_LOCAL_BUFFER_READ,
+ ASC_LOCAL_BUFFER_WRITE,
} PgAioHandleSharedCallbackID;
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 2a836cf98c6..2e88b19619c 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -205,7 +205,9 @@ extern PGDLLIMPORT int32 *LocalRefCount;
struct PgAioHandleSharedCallbacks;
extern const struct PgAioHandleSharedCallbacks aio_shared_buffer_readv_cb;
+extern const struct PgAioHandleSharedCallbacks aio_shared_buffer_writev_cb;
extern const struct PgAioHandleSharedCallbacks aio_local_buffer_readv_cb;
+extern const struct PgAioHandleSharedCallbacks aio_local_buffer_writev_cb;
/* upper limit for effective_io_concurrency */
diff --git a/src/backend/storage/aio/aio_subject.c b/src/backend/storage/aio/aio_subject.c
index 21341aae425..b2bd0c235e7 100644
--- a/src/backend/storage/aio/aio_subject.c
+++ b/src/backend/storage/aio/aio_subject.c
@@ -52,8 +52,10 @@ static const PgAioHandleSharedCallbacksEntry aio_shared_cbs[] = {
CALLBACK_ENTRY(ASC_MD_WRITEV, aio_md_writev_cb),
CALLBACK_ENTRY(ASC_SHARED_BUFFER_READ, aio_shared_buffer_readv_cb),
+ CALLBACK_ENTRY(ASC_SHARED_BUFFER_WRITE, aio_shared_buffer_writev_cb),
CALLBACK_ENTRY(ASC_LOCAL_BUFFER_READ, aio_local_buffer_readv_cb),
+ CALLBACK_ENTRY(ASC_LOCAL_BUFFER_WRITE, aio_local_buffer_writev_cb),
#undef CALLBACK_ENTRY
};
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 722e73eb7d0..0f94db19f9d 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -6437,6 +6437,44 @@ ReadBufferCompleteReadShared(Buffer buffer, int mode, bool failed)
return buf_failed;
}
+static uint64
+ReadBufferCompleteWriteShared(Buffer buffer, bool release_lock, bool failed)
+{
+ BufferDesc *bufHdr;
+ bool result = false;
+
+ Assert(BufferIsValid(buffer));
+
+ bufHdr = GetBufferDescriptor(buffer - 1);
+
+#ifdef USE_ASSERT_CHECKING
+ {
+ uint32 buf_state = pg_atomic_read_u32(&bufHdr->state);
+
+ Assert(buf_state & BM_VALID);
+ Assert(buf_state & BM_TAG_VALID);
+ Assert(buf_state & BM_IO_IN_PROGRESS);
+ Assert(buf_state & BM_DIRTY);
+ }
+#endif
+
+ /* AFIXME: implement track_io_timing */
+
+ TerminateBufferIO(bufHdr, /* clear_dirty = */ true,
+ failed ? BM_IO_ERROR : 0,
+ /* forget_owner = */ false,
+ /* syncio = */ false);
+
+ /*
+ * The initiator of IO is not managing the lock (i.e. called
+ * LWLockDisown()), we are.
+ */
+ if (release_lock)
+ LWLockReleaseUnowned(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
+
+ return result;
+}
+
/*
* Helper to prepare IO on shared buffers for execution, shared between reads
* and writes.
@@ -6518,6 +6556,12 @@ shared_buffer_readv_prepare(PgAioHandle *ioh)
shared_buffer_prepare_common(ioh, false);
}
+static void
+shared_buffer_writev_prepare(PgAioHandle *ioh)
+{
+ shared_buffer_prepare_common(ioh, true);
+}
+
static PgAioResult
shared_buffer_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
{
@@ -6586,6 +6630,34 @@ buffer_readv_error(PgAioResult result, const PgAioSubjectData *subject_data, int
MemoryContextSwitchTo(oldContext);
}
+static PgAioResult
+shared_buffer_writev_complete(PgAioHandle *ioh, PgAioResult prior_result)
+{
+ PgAioResult result = prior_result;
+ uint64 *io_data;
+ uint8 io_data_len;
+
+ elog(DEBUG3, "%s: %d %d", __func__, prior_result.status, prior_result.result);
+
+ io_data = pgaio_io_get_io_data(ioh, &io_data_len);
+
+ /* FIXME: handle outright errors */
+
+ for (int io_data_off = 0; io_data_off < io_data_len; io_data_off++)
+ {
+ Buffer buf = io_data[io_data_off];
+
+ /* FIXME: handle short writes / failures */
+ /* FIXME: ioh->scb_data.shared_buffer.release_lock */
+ ReadBufferCompleteWriteShared(buf,
+ true,
+ false);
+
+ }
+
+ return result;
+}
+
/*
* Helper to prepare IO on local buffers for execution, shared between reads
* and writes.
@@ -6655,14 +6727,27 @@ local_buffer_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
return result;
}
+static void
+local_buffer_writev_prepare(PgAioHandle *ioh)
+{
+ elog(ERROR, "not yet");
+}
+
const struct PgAioHandleSharedCallbacks aio_shared_buffer_readv_cb = {
.prepare = shared_buffer_readv_prepare,
.complete = shared_buffer_readv_complete,
.error = buffer_readv_error,
};
+const struct PgAioHandleSharedCallbacks aio_shared_buffer_writev_cb = {
+ .prepare = shared_buffer_writev_prepare,
+ .complete = shared_buffer_writev_complete,
+};
const struct PgAioHandleSharedCallbacks aio_local_buffer_readv_cb = {
.prepare = local_buffer_readv_prepare,
.complete = local_buffer_readv_complete,
.error = buffer_readv_error,
};
+const struct PgAioHandleSharedCallbacks aio_local_buffer_writev_cb = {
+ .prepare = local_buffer_writev_prepare,
+};
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0016-aio-Add-IO-queue-helper.patch (7.2K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/17-v2-0016-aio-Add-IO-queue-helper.patch)
download | inline diff:
From 0d7dbde438633fbb7af0dd2f3efd3a2c6b587438 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 4 Sep 2024 16:15:42 -0400
Subject: [PATCH v2 16/20] aio: Add IO queue helper
This is likely never going to anywhere - Thomas Munro is working on something
more complete. But I needed a way to exercise aio for checkpointer / bgwriter.
---
src/include/storage/io_queue.h | 33 +++++
src/backend/storage/aio/Makefile | 1 +
src/backend/storage/aio/io_queue.c | 195 ++++++++++++++++++++++++++++
src/backend/storage/aio/meson.build | 1 +
src/tools/pgindent/typedefs.list | 2 +
5 files changed, 232 insertions(+)
create mode 100644 src/include/storage/io_queue.h
create mode 100644 src/backend/storage/aio/io_queue.c
diff --git a/src/include/storage/io_queue.h b/src/include/storage/io_queue.h
new file mode 100644
index 00000000000..28077158d6d
--- /dev/null
+++ b/src/include/storage/io_queue.h
@@ -0,0 +1,33 @@
+/*-------------------------------------------------------------------------
+ *
+ * io_queue.h
+ * Mechanism for tracking many IOs
+ *
+ *
+ * Portions Copyright (c) 1996-2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/storage/io_queue.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef IO_QUEUE_H
+#define IO_QUEUE_H
+
+#include "storage/bufmgr.h"
+
+struct IOQueue;
+typedef struct IOQueue IOQueue;
+
+struct PgAioHandleRef;
+
+extern IOQueue *io_queue_create(int depth, int flags);
+extern void io_queue_track(IOQueue *ioq, const struct PgAioHandleRef *ior);
+extern void io_queue_wait_one(IOQueue *ioq);
+extern void io_queue_wait_all(IOQueue *ioq);
+extern bool io_queue_is_empty(IOQueue *ioq);
+extern void io_queue_reserve(IOQueue *ioq);
+extern struct PgAioHandle *io_queue_get_io(IOQueue *ioq);
+extern void io_queue_free(IOQueue *ioq);
+
+#endif /* IO_QUEUE_H */
diff --git a/src/backend/storage/aio/Makefile b/src/backend/storage/aio/Makefile
index 3bcb8a0b2ed..f3a7f9e63d6 100644
--- a/src/backend/storage/aio/Makefile
+++ b/src/backend/storage/aio/Makefile
@@ -13,6 +13,7 @@ OBJS = \
aio_init.o \
aio_io.o \
aio_subject.o \
+ io_queue.o \
method_io_uring.o \
method_sync.o \
method_worker.o \
diff --git a/src/backend/storage/aio/io_queue.c b/src/backend/storage/aio/io_queue.c
new file mode 100644
index 00000000000..89ccfc2b9a7
--- /dev/null
+++ b/src/backend/storage/aio/io_queue.c
@@ -0,0 +1,195 @@
+/*-------------------------------------------------------------------------
+ *
+ * io_queue.c
+ * AIO - Mechanism for tracking many IOs
+ *
+ * Portions Copyright (c) 2024, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * IDENTIFICATION
+ * src/backend/storage/aio/io_queue.c
+ *
+ *-------------------------------------------------------------------------
+ */
+#include "postgres.h"
+
+#include "storage/io_queue.h"
+
+#include "storage/aio.h"
+
+
+typedef struct TrackedIO
+{
+ PgAioHandleRef ior;
+ dlist_node node;
+} TrackedIO;
+
+struct IOQueue
+{
+ int depth;
+ int unsubmitted;
+
+ bool has_reserved;
+
+ dclist_head idle;
+ dclist_head in_progress;
+
+ TrackedIO tracked_ios[FLEXIBLE_ARRAY_MEMBER];
+};
+
+
+IOQueue *
+io_queue_create(int depth, int flags)
+{
+ size_t sz;
+ IOQueue *ioq;
+
+ sz = offsetof(IOQueue, tracked_ios)
+ + sizeof(TrackedIO) * depth;
+
+ ioq = palloc0(sz);
+
+ ioq->depth = 0;
+
+ for (int i = 0; i < depth; i++)
+ {
+ TrackedIO *tio = &ioq->tracked_ios[i];
+
+ pgaio_io_ref_clear(&tio->ior);
+ dclist_push_tail(&ioq->idle, &tio->node);
+ }
+
+ return ioq;
+}
+
+void
+io_queue_wait_one(IOQueue *ioq)
+{
+ while (!dclist_is_empty(&ioq->in_progress))
+ {
+ /* FIXME: Should we really pop here already? */
+ dlist_node *node = dclist_pop_head_node(&ioq->in_progress);
+ TrackedIO *tio = dclist_container(TrackedIO, node, node);
+
+ pgaio_io_ref_wait(&tio->ior);
+ dclist_push_head(&ioq->idle, &tio->node);
+ }
+}
+
+void
+io_queue_reserve(IOQueue *ioq)
+{
+ if (ioq->has_reserved)
+ return;
+
+ if (dclist_is_empty(&ioq->idle))
+ io_queue_wait_one(ioq);
+
+ Assert(!dclist_is_empty(&ioq->idle));
+
+ ioq->has_reserved = true;
+}
+
+PgAioHandle *
+io_queue_get_io(IOQueue *ioq)
+{
+ PgAioHandle *ioh;
+
+ io_queue_reserve(ioq);
+
+ Assert(!dclist_is_empty(&ioq->idle));
+
+ if (!io_queue_is_empty(ioq))
+ {
+ ioh = pgaio_io_get_nb(CurrentResourceOwner, NULL);
+ if (ioh == NULL)
+ {
+ /*
+ * Need to wait for all IOs, blocking might not be legal in the
+ * context.
+ *
+ * XXX: This doesn't make a whole lot of sense, we're also
+ * blocking here. What was I smoking when I wrote the above?
+ */
+ io_queue_wait_all(ioq);
+ ioh = pgaio_io_get(CurrentResourceOwner, NULL);
+ }
+ }
+ else
+ {
+ ioh = pgaio_io_get(CurrentResourceOwner, NULL);
+ }
+
+ return ioh;
+}
+
+void
+io_queue_track(IOQueue *ioq, const struct PgAioHandleRef *ior)
+{
+ dlist_node *node;
+ TrackedIO *tio;
+
+ Assert(ioq->has_reserved);
+ ioq->has_reserved = false;
+
+ Assert(!dclist_is_empty(&ioq->idle));
+
+ node = dclist_pop_head_node(&ioq->idle);
+ tio = dclist_container(TrackedIO, node, node);
+
+ tio->ior = *ior;
+
+ dclist_push_tail(&ioq->in_progress, &tio->node);
+
+ ioq->unsubmitted++;
+
+ /*
+ * XXX: Should have some smarter logic here. We don't want to wait too
+ * long to submit, that'll mean we're more likely to block. But we also
+ * don't want to have the overhead of submitting every IO individually.
+ */
+ if (ioq->unsubmitted >= 4)
+ {
+ pgaio_submit_staged();
+ ioq->unsubmitted = 0;
+ }
+}
+
+void
+io_queue_wait_all(IOQueue *ioq)
+{
+ while (!dclist_is_empty(&ioq->in_progress))
+ {
+ /* wait for the last IO to minimize unnecessary wakeups */
+ dlist_node *node = dclist_tail_node(&ioq->in_progress);
+ TrackedIO *tio = dclist_container(TrackedIO, node, node);
+
+ if (!pgaio_io_ref_check_done(&tio->ior))
+ {
+ ereport(DEBUG3,
+ errmsg("io_queue_wait_all for io:%d",
+ pgaio_io_ref_get_id(&tio->ior)),
+ errhidestmt(true),
+ errhidecontext(true));
+
+ pgaio_io_ref_wait(&tio->ior);
+ }
+
+ dclist_delete_from(&ioq->in_progress, &tio->node);
+ dclist_push_head(&ioq->idle, &tio->node);
+ }
+}
+
+bool
+io_queue_is_empty(IOQueue *ioq)
+{
+ return dclist_is_empty(&ioq->in_progress);
+}
+
+void
+io_queue_free(IOQueue *ioq)
+{
+ io_queue_wait_all(ioq);
+
+ pfree(ioq);
+}
diff --git a/src/backend/storage/aio/meson.build b/src/backend/storage/aio/meson.build
index 537f23d446d..e8a88e615c0 100644
--- a/src/backend/storage/aio/meson.build
+++ b/src/backend/storage/aio/meson.build
@@ -5,6 +5,7 @@ backend_sources += files(
'aio_init.c',
'aio_io.c',
'aio_subject.c',
+ 'io_queue.c',
'method_io_uring.c',
'method_sync.c',
'method_worker.c',
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index dc52d6165d4..ca1e3427bc1 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1175,6 +1175,7 @@ IOContext
IOFuncSelector
IOObject
IOOp
+IOQueue
IO_STATUS_BLOCK
IPCompareMethod
ITEM
@@ -2974,6 +2975,7 @@ TocEntry
TokenAuxData
TokenizedAuthLine
TrackItem
+TrackedIO
TransApplyAction
TransInvalidationInfo
TransState
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0017-bufmgr-use-AIO-in-checkpointer-bgwriter.patch (31.6K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/18-v2-0017-bufmgr-use-AIO-in-checkpointer-bgwriter.patch)
download | inline diff:
From ffe8489a8b44bc0a0b11ad765d578aa12801925a Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 23 Jul 2024 10:01:23 -0700
Subject: [PATCH v2 17/20] bufmgr: use AIO in checkpointer, bgwriter
This is far from ready - just included to be able to exercise AIO writes and
get some preliminary numbers. In all likelihood this will instead be based
ontop of work by Thomas Munro instead of the preceding commit.
---
src/include/postmaster/bgwriter.h | 3 +-
src/include/storage/buf_internals.h | 2 +
src/include/storage/bufmgr.h | 3 +-
src/include/storage/bufpage.h | 1 +
src/backend/postmaster/bgwriter.c | 25 +-
src/backend/postmaster/checkpointer.c | 12 +-
src/backend/storage/buffer/bufmgr.c | 581 +++++++++++++++++++++++---
src/backend/storage/page/bufpage.c | 10 +
src/tools/pgindent/typedefs.list | 1 +
9 files changed, 580 insertions(+), 58 deletions(-)
diff --git a/src/include/postmaster/bgwriter.h b/src/include/postmaster/bgwriter.h
index 407f26e5302..01a936fbc0a 100644
--- a/src/include/postmaster/bgwriter.h
+++ b/src/include/postmaster/bgwriter.h
@@ -31,7 +31,8 @@ extern void BackgroundWriterMain(char *startup_data, size_t startup_data_len) pg
extern void CheckpointerMain(char *startup_data, size_t startup_data_len) pg_attribute_noreturn();
extern void RequestCheckpoint(int flags);
-extern void CheckpointWriteDelay(int flags, double progress);
+struct IOQueue;
+extern void CheckpointWriteDelay(struct IOQueue *ioq, int flags, double progress);
extern bool ForwardSyncRequest(const FileTag *ftag, SyncRequestType type);
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 37520890073..9d3123663b3 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -21,6 +21,8 @@
#include "storage/buf.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
+#include "storage/io_queue.h"
+#include "storage/latch.h"
#include "storage/lwlock.h"
#include "storage/shmem.h"
#include "storage/smgr.h"
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 2e88b19619c..455bbbcbfc4 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -327,7 +327,8 @@ extern bool ConditionalLockBufferForCleanup(Buffer buffer);
extern bool IsBufferCleanupOK(Buffer buffer);
extern bool HoldingBufferPinThatDelaysRecovery(void);
-extern bool BgBufferSync(struct WritebackContext *wb_context);
+struct IOQueue;
+extern bool BgBufferSync(struct IOQueue *ioq, struct WritebackContext *wb_context);
extern void LimitAdditionalPins(uint32 *additional_pins);
extern void LimitAdditionalLocalPins(uint32 *additional_pins);
diff --git a/src/include/storage/bufpage.h b/src/include/storage/bufpage.h
index 6222d46e535..6f8fe796da3 100644
--- a/src/include/storage/bufpage.h
+++ b/src/include/storage/bufpage.h
@@ -509,5 +509,6 @@ extern bool PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
Item newtup, Size newsize);
extern char *PageSetChecksumCopy(Page page, BlockNumber blkno);
extern void PageSetChecksumInplace(Page page, BlockNumber blkno);
+extern bool PageNeedsChecksumCopy(Page page);
#endif /* BUFPAGE_H */
diff --git a/src/backend/postmaster/bgwriter.c b/src/backend/postmaster/bgwriter.c
index 0f75548759a..71c08da45db 100644
--- a/src/backend/postmaster/bgwriter.c
+++ b/src/backend/postmaster/bgwriter.c
@@ -38,10 +38,12 @@
#include "postmaster/auxprocess.h"
#include "postmaster/bgwriter.h"
#include "postmaster/interrupt.h"
+#include "storage/aio.h"
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/fd.h"
+#include "storage/io_queue.h"
#include "storage/lwlock.h"
#include "storage/proc.h"
#include "storage/procsignal.h"
@@ -89,6 +91,7 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
sigjmp_buf local_sigjmp_buf;
MemoryContext bgwriter_context;
bool prev_hibernate;
+ IOQueue *ioq;
WritebackContext wb_context;
Assert(startup_data_len == 0);
@@ -130,6 +133,7 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
ALLOCSET_DEFAULT_SIZES);
MemoryContextSwitchTo(bgwriter_context);
+ ioq = io_queue_create(128, 0);
WritebackContextInit(&wb_context, &bgwriter_flush_after);
/*
@@ -167,6 +171,7 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
* about in bgwriter, but we do have LWLocks, buffers, and temp files.
*/
LWLockReleaseAll();
+ pgaio_at_error();
ConditionVariableCancelSleep();
UnlockBuffers();
ReleaseAuxProcessResources(false);
@@ -226,12 +231,27 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
/* Clear any already-pending wakeups */
ResetLatch(MyLatch);
+ /*
+ * XXX: Before exiting, wait for all IO to finish. That's only
+ * important to avoid spurious PrintBufferLeakWarning() /
+ * PrintAioIPLeakWarning() calls, triggered by
+ * ReleaseAuxProcessResources() being called with isCommit=true.
+ *
+ * FIXME: this is theoretically racy, but I didn't want to copy
+ * HandleMainLoopInterrupts() remaining body here.
+ */
+ if (ShutdownRequestPending)
+ {
+ io_queue_wait_all(ioq);
+ io_queue_free(ioq);
+ }
+
HandleMainLoopInterrupts();
/*
* Do one cycle of dirty-buffer writing.
*/
- can_hibernate = BgBufferSync(&wb_context);
+ can_hibernate = BgBufferSync(ioq, &wb_context);
/* Report pending statistics to the cumulative stats system */
pgstat_report_bgwriter();
@@ -248,6 +268,9 @@ BackgroundWriterMain(char *startup_data, size_t startup_data_len)
smgrdestroyall();
}
+ /* finish IO before sleeping, to avoid blocking other backends */
+ io_queue_wait_all(ioq);
+
/*
* Log a new xl_running_xacts every now and then so replication can
* get into a consistent state faster (think of suboverflowed
diff --git a/src/backend/postmaster/checkpointer.c b/src/backend/postmaster/checkpointer.c
index 982572a75db..0c08acd611f 100644
--- a/src/backend/postmaster/checkpointer.c
+++ b/src/backend/postmaster/checkpointer.c
@@ -46,9 +46,11 @@
#include "postmaster/bgwriter.h"
#include "postmaster/interrupt.h"
#include "replication/syncrep.h"
+#include "storage/aio.h"
#include "storage/bufmgr.h"
#include "storage/condition_variable.h"
#include "storage/fd.h"
+#include "storage/io_queue.h"
#include "storage/ipc.h"
#include "storage/lwlock.h"
#include "storage/proc.h"
@@ -266,6 +268,7 @@ CheckpointerMain(char *startup_data, size_t startup_data_len)
* files.
*/
LWLockReleaseAll();
+ pgaio_at_error();
ConditionVariableCancelSleep();
pgstat_report_wait_end();
UnlockBuffers();
@@ -719,7 +722,7 @@ ImmediateCheckpointRequested(void)
* fraction between 0.0 meaning none, and 1.0 meaning all done.
*/
void
-CheckpointWriteDelay(int flags, double progress)
+CheckpointWriteDelay(IOQueue *ioq, int flags, double progress)
{
static int absorb_counter = WRITES_PER_ABSORB;
@@ -752,6 +755,13 @@ CheckpointWriteDelay(int flags, double progress)
/* Report interim statistics to the cumulative stats system */
pgstat_report_checkpointer();
+ /*
+ * Ensure all pending IO is submitted to avoid unnecessary delays for
+ * other processes.
+ */
+ io_queue_wait_all(ioq);
+
+
/*
* This sleep used to be connected to bgwriter_delay, typically 200ms.
* That resulted in more frequent wakeups if not much work to do.
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 0f94db19f9d..863464f12da 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -52,6 +52,7 @@
#include "storage/buf_internals.h"
#include "storage/bufmgr.h"
#include "storage/fd.h"
+#include "storage/io_queue.h"
#include "storage/ipc.h"
#include "storage/lmgr.h"
#include "storage/proc.h"
@@ -77,6 +78,7 @@
/* Bits in SyncOneBuffer's return value */
#define BUF_WRITTEN 0x01
#define BUF_REUSABLE 0x02
+#define BUF_CANT_MERGE 0x04
#define RELS_BSEARCH_THRESHOLD 20
@@ -511,8 +513,6 @@ static void UnpinBuffer(BufferDesc *buf);
static void UnpinBufferNoOwner(BufferDesc *buf);
static void BufferSync(int flags);
static uint32 WaitBufHdrUnlocked(BufferDesc *buf);
-static int SyncOneBuffer(int buf_id, bool skip_recently_used,
- WritebackContext *wb_context);
static void WaitIO(BufferDesc *buf);
static bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
static void TerminateBufferIO(BufferDesc *buf, bool clear_dirty,
@@ -530,6 +530,7 @@ static inline BufferDesc *BufferAlloc(SMgrRelation smgr,
static Buffer GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context);
static void FlushBuffer(BufferDesc *buf, SMgrRelation reln,
IOObject io_object, IOContext io_context);
+
static void FindAndDropRelationBuffers(RelFileLocator rlocator,
ForkNumber forkNum,
BlockNumber nForkBlock,
@@ -3067,6 +3068,56 @@ UnpinBufferNoOwner(BufferDesc *buf)
}
}
+typedef struct BuffersToWrite
+{
+ int nbuffers;
+ BufferTag start_at_tag;
+ uint32 max_combine;
+
+ XLogRecPtr max_lsn;
+
+ PgAioHandle *ioh;
+ PgAioHandleRef ior;
+
+ uint64 total_writes;
+
+ Buffer buffers[IOV_MAX];
+ PgAioBounceBuffer *bounce_buffers[IOV_MAX];
+ const void *data_ptrs[IOV_MAX];
+} BuffersToWrite;
+
+static int PrepareToWriteBuffer(BuffersToWrite *to_write, Buffer buf,
+ bool skip_recently_used,
+ IOQueue *ioq, WritebackContext *wb_context);
+
+static void WriteBuffers(BuffersToWrite *to_write,
+ IOQueue *ioq, WritebackContext *wb_context);
+
+static void
+BuffersToWriteInit(BuffersToWrite *to_write,
+ IOQueue *ioq, WritebackContext *wb_context)
+{
+ to_write->total_writes = 0;
+ to_write->nbuffers = 0;
+ to_write->ioh = NULL;
+ pgaio_io_ref_clear(&to_write->ior);
+ to_write->max_lsn = InvalidXLogRecPtr;
+}
+
+static void
+BuffersToWriteEnd(BuffersToWrite *to_write)
+{
+ if (to_write->ioh != NULL)
+ {
+ pgaio_io_release(to_write->ioh);
+ to_write->ioh = NULL;
+ }
+
+ if (to_write->total_writes > 0)
+ pgaio_submit_staged();
+}
+
+
#define ST_SORT sort_checkpoint_bufferids
#define ST_ELEMENT_TYPE CkptSortItem
#define ST_COMPARE(a, b) ckpt_buforder_comparator(a, b)
@@ -3098,7 +3149,10 @@ BufferSync(int flags)
binaryheap *ts_heap;
int i;
int mask = BM_DIRTY;
+ IOQueue *ioq;
WritebackContext wb_context;
+ BuffersToWrite to_write;
+ int max_combine;
/*
* Unless this is a shutdown checkpoint or we have been explicitly told,
@@ -3160,7 +3214,9 @@ BufferSync(int flags)
if (num_to_scan == 0)
return; /* nothing to do */
+ ioq = io_queue_create(512, 0);
WritebackContextInit(&wb_context, &checkpoint_flush_after);
+ max_combine = Min(io_bounce_buffers, io_combine_limit);
TRACE_POSTGRESQL_BUFFER_SYNC_START(NBuffers, num_to_scan);
@@ -3268,48 +3324,91 @@ BufferSync(int flags)
*/
num_processed = 0;
num_written = 0;
+
+ BuffersToWriteInit(&to_write, ioq, &wb_context);
+
while (!binaryheap_empty(ts_heap))
{
BufferDesc *bufHdr = NULL;
CkptTsStatus *ts_stat = (CkptTsStatus *)
DatumGetPointer(binaryheap_first(ts_heap));
+ bool batch_continue = true;
- buf_id = CkptBufferIds[ts_stat->index].buf_id;
- Assert(buf_id != -1);
-
- bufHdr = GetBufferDescriptor(buf_id);
-
- num_processed++;
+ Assert(ts_stat->num_scanned <= ts_stat->num_to_scan);
/*
- * We don't need to acquire the lock here, because we're only looking
- * at a single bit. It's possible that someone else writes the buffer
- * and clears the flag right after we check, but that doesn't matter
- * since SyncOneBuffer will then do nothing. However, there is a
- * further race condition: it's conceivable that between the time we
- * examine the bit here and the time SyncOneBuffer acquires the lock,
- * someone else not only wrote the buffer but replaced it with another
- * page and dirtied it. In that improbable case, SyncOneBuffer will
- * write the buffer though we didn't need to. It doesn't seem worth
- * guarding against this, though.
+ * Collect a batch of buffers to write out from the current
+ * tablespace. That causes some imbalance between the tablespaces, but
+ * that's more than outweighed by the efficiency gain due to batching.
*/
- if (pg_atomic_read_u32(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
+ while (batch_continue &&
+ to_write.nbuffers < max_combine &&
+ ts_stat->num_scanned < ts_stat->num_to_scan)
{
- if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
+ buf_id = CkptBufferIds[ts_stat->index].buf_id;
+ Assert(buf_id != -1);
+
+ bufHdr = GetBufferDescriptor(buf_id);
+
+ num_processed++;
+
+ /*
+ * We don't need to acquire the lock here, because we're only
+ * looking at a single bit. It's possible that someone else writes
+ * the buffer and clears the flag right after we check, but that
+ * doesn't matter since SyncOneBuffer will then do nothing.
+ * However, there is a further race condition: it's conceivable
+ * that between the time we examine the bit here and the time
+ * SyncOneBuffer acquires the lock, someone else not only wrote
+ * the buffer but replaced it with another page and dirtied it. In
+ * that improbable case, SyncOneBuffer will write the buffer
+ * though we didn't need to. It doesn't seem worth guarding
+ * against this, though.
+ */
+ if (pg_atomic_read_u32(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
{
- TRACE_POSTGRESQL_BUFFER_SYNC_WRITTEN(buf_id);
- PendingCheckpointerStats.buffers_written++;
- num_written++;
+ int result = PrepareToWriteBuffer(&to_write, buf_id + 1, false,
+ ioq, &wb_context);
+
+ if (result & BUF_CANT_MERGE)
+ {
+ Assert(to_write.nbuffers > 0);
+ WriteBuffers(&to_write, ioq, &wb_context);
+
+ result = PrepareToWriteBuffer(&to_write, buf_id + 1, false,
+ ioq, &wb_context);
+ Assert(result != BUF_CANT_MERGE);
+ }
+
+ if (result & BUF_WRITTEN)
+ {
+ TRACE_POSTGRESQL_BUFFER_SYNC_WRITTEN(buf_id);
+ PendingCheckpointerStats.buffers_written++;
+ num_written++;
+ }
+ else
+ {
+ batch_continue = false;
+ }
}
+ else
+ {
+ if (to_write.nbuffers > 0)
+ WriteBuffers(&to_write, ioq, &wb_context);
+ }
+
+ /*
+ * Measure progress independent of actually having to flush the
+ * buffer - otherwise writing become unbalanced.
+ */
+ ts_stat->progress += ts_stat->progress_slice;
+ ts_stat->num_scanned++;
+ ts_stat->index++;
}
- /*
- * Measure progress independent of actually having to flush the buffer
- * - otherwise writing become unbalanced.
- */
- ts_stat->progress += ts_stat->progress_slice;
- ts_stat->num_scanned++;
- ts_stat->index++;
+ if (to_write.nbuffers > 0)
+ WriteBuffers(&to_write, ioq, &wb_context);
+
/* Have all the buffers from the tablespace been processed? */
if (ts_stat->num_scanned == ts_stat->num_to_scan)
@@ -3327,15 +3426,23 @@ BufferSync(int flags)
*
* (This will check for barrier events even if it doesn't sleep.)
*/
- CheckpointWriteDelay(flags, (double) num_processed / num_to_scan);
+ CheckpointWriteDelay(ioq, flags, (double) num_processed / num_to_scan);
}
+ Assert(to_write.nbuffers == 0);
+ io_queue_wait_all(ioq);
+
/*
* Issue all pending flushes. Only checkpointer calls BufferSync(), so
* IOContext will always be IOCONTEXT_NORMAL.
*/
IssuePendingWritebacks(&wb_context, IOCONTEXT_NORMAL);
+ io_queue_wait_all(ioq); /* IssuePendingWritebacks might have added
+ * more */
+ io_queue_free(ioq);
+ BuffersToWriteEnd(&to_write);
+
pfree(per_ts_stat);
per_ts_stat = NULL;
binaryheap_free(ts_heap);
@@ -3361,7 +3468,7 @@ BufferSync(int flags)
* bgwriter_lru_maxpages to 0.)
*/
bool
-BgBufferSync(WritebackContext *wb_context)
+BgBufferSync(IOQueue *ioq, WritebackContext *wb_context)
{
/* info obtained from freelist.c */
int strategy_buf_id;
@@ -3404,6 +3511,9 @@ BgBufferSync(WritebackContext *wb_context)
long new_strategy_delta;
uint32 new_recent_alloc;
+ BuffersToWrite to_write;
+ int max_combine;
+
/*
* Find out where the freelist clock sweep currently is, and how many
* buffer allocations have happened since our last call.
@@ -3424,6 +3534,8 @@ BgBufferSync(WritebackContext *wb_context)
return true;
}
+ max_combine = Min(io_bounce_buffers, io_combine_limit);
+
/*
* Compute strategy_delta = how many buffers have been scanned by the
* clock sweep since last time. If first time through, assume none. Then
@@ -3580,11 +3692,25 @@ BgBufferSync(WritebackContext *wb_context)
num_written = 0;
reusable_buffers = reusable_buffers_est;
+ BuffersToWriteInit(&to_write, ioq, wb_context);
+
/* Execute the LRU scan */
while (num_to_scan > 0 && reusable_buffers < upcoming_alloc_est)
{
- int sync_state = SyncOneBuffer(next_to_clean, true,
- wb_context);
+ int sync_state;
+
+ sync_state = PrepareToWriteBuffer(&to_write, next_to_clean + 1,
+ true, ioq, wb_context);
+ if (sync_state & BUF_CANT_MERGE)
+ {
+ Assert(to_write.nbuffers > 0);
+
+ WriteBuffers(&to_write, ioq, wb_context);
+
+ sync_state = PrepareToWriteBuffer(&to_write, next_to_clean + 1,
+ true, ioq, wb_context);
+ Assert(sync_state != BUF_CANT_MERGE);
+ }
if (++next_to_clean >= NBuffers)
{
@@ -3595,6 +3721,13 @@ BgBufferSync(WritebackContext *wb_context)
if (sync_state & BUF_WRITTEN)
{
+ Assert(sync_state & BUF_REUSABLE);
+
+ if (to_write.nbuffers == max_combine)
+ {
+ WriteBuffers(&to_write, ioq, wb_context);
+ }
+
reusable_buffers++;
if (++num_written >= bgwriter_lru_maxpages)
{
@@ -3606,6 +3739,11 @@ BgBufferSync(WritebackContext *wb_context)
reusable_buffers++;
}
+ if (to_write.nbuffers > 0)
+ WriteBuffers(&to_write, ioq, wb_context);
+
+ BuffersToWriteEnd(&to_write);
+
PendingBgWriterStats.buf_written_clean += num_written;
#ifdef BGW_DEBUG
@@ -3644,8 +3782,66 @@ BgBufferSync(WritebackContext *wb_context)
return (bufs_to_lap == 0 && recent_alloc == 0);
}
+static inline bool
+BufferTagsSameRel(const BufferTag *tag1, const BufferTag *tag2)
+{
+ return (tag1->spcOid == tag2->spcOid) &&
+ (tag1->dbOid == tag2->dbOid) &&
+ (tag1->relNumber == tag2->relNumber) &&
+ (tag1->forkNum == tag2->forkNum)
+ ;
+}
+
+static bool
+CanMergeWrite(BuffersToWrite *to_write, BufferDesc *cur_buf_hdr)
+{
+ BlockNumber cur_block = cur_buf_hdr->tag.blockNum;
+
+ Assert(to_write->nbuffers > 0); /* can't merge with nothing */
+ Assert(to_write->start_at_tag.relNumber != InvalidOid);
+ Assert(to_write->start_at_tag.blockNum != InvalidBlockNumber);
+
+ Assert(to_write->ioh != NULL);
+
+ /*
+ * First check if the blocknumber is one that we could actually merge,
+ * that's cheaper than checking the tablespace/db/relnumber/fork match.
+ */
+ if (to_write->start_at_tag.blockNum + to_write->nbuffers != cur_block)
+ return false;
+
+ if (!BufferTagsSameRel(&to_write->start_at_tag, &cur_buf_hdr->tag))
+ return false;
+
+ /*
+ * Need to check with smgr how large a write we're allowed to make. To
+ * reduce the overhead of the smgr check, only inquire once, when
+ * processing the first to-be-merged buffer. That avoids the overhead in
+ * the common case of writing out buffers that definitely not mergeable.
+ */
+ if (to_write->nbuffers == 1)
+ {
+ SMgrRelation smgr;
+
+ smgr = smgropen(BufTagGetRelFileLocator(&to_write->start_at_tag), INVALID_PROC_NUMBER);
+
+ to_write->max_combine = smgrmaxcombine(smgr,
+ to_write->start_at_tag.forkNum,
+ to_write->start_at_tag.blockNum);
+ }
+ else
+ {
+ Assert(to_write->max_combine > 0);
+ }
+
+ if (to_write->start_at_tag.blockNum + to_write->max_combine <= cur_block)
+ return false;
+
+ return true;
+}
+
/*
- * SyncOneBuffer -- process a single buffer during syncing.
+ * PrepareToWriteBuffer -- process a single buffer during syncing.
*
* If skip_recently_used is true, we don't write currently-pinned buffers, nor
* buffers marked recently used, as these are not replacement candidates.
@@ -3654,22 +3850,56 @@ BgBufferSync(WritebackContext *wb_context)
* BUF_WRITTEN: we wrote the buffer.
* BUF_REUSABLE: buffer is available for replacement, ie, it has
* pin count 0 and usage count 0.
+ * BUF_CANT_MERGE: can't combine this write with prior writes, caller needs
+ * to issue those first
*
* (BUF_WRITTEN could be set in error if FlushBuffer finds the buffer clean
* after locking it, but we don't care all that much.)
*/
static int
-SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
+PrepareToWriteBuffer(BuffersToWrite *to_write, Buffer buf,
+ bool skip_recently_used,
+ IOQueue *ioq, WritebackContext *wb_context)
{
- BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
+ BufferDesc *cur_buf_hdr = GetBufferDescriptor(buf - 1);
+ uint32 buf_state;
int result = 0;
- uint32 buf_state;
- BufferTag tag;
+ XLogRecPtr cur_buf_lsn;
+ LWLock *content_lock;
+ bool may_block;
+
+ /*
+ * Check if this buffer can be written out together with already prepared
+ * writes. We check before we have pinned the buffer, so the buffer can be
+ * written out and replaced between this check and us pinning the buffer -
+ * we'll recheck below. The reason for the pre-check is that we don't want
+ * to pin the buffer just to find out that we can't merge the IO.
+ */
+ if (to_write->nbuffers != 0)
+ {
+ if (!CanMergeWrite(to_write, cur_buf_hdr))
+ {
+ result |= BUF_CANT_MERGE;
+ return result;
+ }
+ }
+ else
+ {
+ if (to_write->ioh == NULL)
+ {
+ to_write->ioh = io_queue_get_io(ioq);
+ pgaio_io_get_ref(to_write->ioh, &to_write->ior);
+ }
+
+ to_write->start_at_tag = cur_buf_hdr->tag;
+ }
/* Make sure we can handle the pin */
ReservePrivateRefCountEntry();
ResourceOwnerEnlarge(CurrentResourceOwner);
+ /* XXX: Should also check if we are allowed to pin one more buffer */
+
/*
* Check whether buffer needs writing.
*
@@ -3679,7 +3909,7 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
* don't worry because our checkpoint.redo points before log record for
* upcoming changes and so we are not required to write such dirty buffer.
*/
- buf_state = LockBufHdr(bufHdr);
+ buf_state = LockBufHdr(cur_buf_hdr);
if (BUF_STATE_GET_REFCOUNT(buf_state) == 0 &&
BUF_STATE_GET_USAGECOUNT(buf_state) == 0)
@@ -3688,40 +3918,282 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
}
else if (skip_recently_used)
{
+#if 0
+ elog(LOG, "at block %d: skip recent with nbuffers %d",
+ cur_buf_hdr->tag.blockNum, to_write->nbuffers);
+#endif
/* Caller told us not to write recently-used buffers */
- UnlockBufHdr(bufHdr, buf_state);
+ UnlockBufHdr(cur_buf_hdr, buf_state);
return result;
}
if (!(buf_state & BM_VALID) || !(buf_state & BM_DIRTY))
{
/* It's clean, so nothing to do */
- UnlockBufHdr(bufHdr, buf_state);
+ UnlockBufHdr(cur_buf_hdr, buf_state);
return result;
}
+ /* pin the buffer, from now on its identity can't change anymore */
+ PinBuffer_Locked(cur_buf_hdr);
+
/*
- * Pin it, share-lock it, write it. (FlushBuffer will do nothing if the
- * buffer is clean by the time we've locked it.)
+ * If we are merging, check if the buffer's identity possibly changed
+ * while we hadn't yet pinned it.
+ *
+ * XXX: It might be worth checking if we still want to write the buffer
+ * out, e.g. it could have been replaced with a buffer that doesn't have
+ * BM_CHECKPOINT_NEEDED set.
*/
- PinBuffer_Locked(bufHdr);
- LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
+ if (to_write->nbuffers != 0)
+ {
+ if (!CanMergeWrite(to_write, cur_buf_hdr))
+ {
+ elog(LOG, "changed identity");
+ UnpinBuffer(cur_buf_hdr);
+
+ result |= BUF_CANT_MERGE;
+
+ return result;
+ }
+ }
+
+ may_block = to_write->nbuffers == 0
+ && !pgaio_have_staged()
+ && io_queue_is_empty(ioq)
+ ;
+ content_lock = BufferDescriptorGetContentLock(cur_buf_hdr);
+
+ if (!may_block)
+ {
+ if (LWLockConditionalAcquire(content_lock, LW_SHARED))
+ {
+ /* done */
+ }
+ else if (to_write->nbuffers == 0)
+ {
+ /*
+ * Need to wait for all prior IO to finish before blocking for
+ * lock acquisition, to avoid the risk a deadlock due to us
+ * waiting for another backend that is waiting for our unsubmitted
+ * IO to complete.
+ */
+ pgaio_submit_staged();
+ io_queue_wait_all(ioq);
+
+ elog(DEBUG2, "at block %u: can't block, nbuffers = 0",
+ cur_buf_hdr->tag.blockNum
+ );
+
+ may_block = to_write->nbuffers == 0
+ && !pgaio_have_staged()
+ && io_queue_is_empty(ioq)
+ ;
+ Assert(may_block);
+
+ LWLockAcquire(content_lock, LW_SHARED);
+ }
+ else
+ {
+ elog(DEBUG2, "at block %d: can't block nbuffers = %d",
+ cur_buf_hdr->tag.blockNum,
+ to_write->nbuffers);
+
+ UnpinBuffer(cur_buf_hdr);
+ result |= BUF_CANT_MERGE;
+ Assert(to_write->nbuffers > 0);
+
+ return result;
+ }
+ }
+ else
+ {
+ LWLockAcquire(content_lock, LW_SHARED);
+ }
+
+ if (!may_block)
+ {
+ if (!StartBufferIO(cur_buf_hdr, false, !may_block))
+ {
+ pgaio_submit_staged();
+ io_queue_wait_all(ioq);
- FlushBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
+ may_block = io_queue_is_empty(ioq) && to_write->nbuffers == 0 && !pgaio_have_staged();
- LWLockRelease(BufferDescriptorGetContentLock(bufHdr));
+ if (!StartBufferIO(cur_buf_hdr, false, !may_block))
+ {
+ elog(DEBUG2, "at block %d: non-waitable StartBufferIO returns false, %d",
+ cur_buf_hdr->tag.blockNum,
+ may_block);
- tag = bufHdr->tag;
+ /*
+ * FIXME: can't tell whether this is because the buffer has
+ * been cleaned
+ */
+ if (!may_block)
+ {
+ result |= BUF_CANT_MERGE;
+ Assert(to_write->nbuffers > 0);
+ }
+ LWLockRelease(content_lock);
+ UnpinBuffer(cur_buf_hdr);
- UnpinBuffer(bufHdr);
+ return result;
+ }
+ }
+ }
+ else
+ {
+ if (!StartBufferIO(cur_buf_hdr, false, false))
+ {
+ elog(DEBUG2, "waitable StartBufferIO returns false");
+ LWLockRelease(content_lock);
+ UnpinBuffer(cur_buf_hdr);
+
+ /*
+ * FIXME: Historically we returned BUF_WRITTEN in this case, which
+ * seems wrong
+ */
+ return result;
+ }
+ }
/*
- * SyncOneBuffer() is only called by checkpointer and bgwriter, so
- * IOContext will always be IOCONTEXT_NORMAL.
+ * Run PageGetLSN while holding header lock, since we don't have the
+ * buffer locked exclusively in all cases.
*/
- ScheduleBufferTagForWriteback(wb_context, IOCONTEXT_NORMAL, &tag);
+ buf_state = LockBufHdr(cur_buf_hdr);
+
+ cur_buf_lsn = BufferGetLSN(cur_buf_hdr);
+
+ /* To check if block content changes while flushing. - vadim 01/17/97 */
+ buf_state &= ~BM_JUST_DIRTIED;
+
+ UnlockBufHdr(cur_buf_hdr, buf_state);
+
+ to_write->buffers[to_write->nbuffers] = buf;
+ to_write->nbuffers++;
+
+ if (buf_state & BM_PERMANENT &&
+ (to_write->max_lsn == InvalidXLogRecPtr || to_write->max_lsn < cur_buf_lsn))
+ {
+ to_write->max_lsn = cur_buf_lsn;
+ }
+
+ result |= BUF_WRITTEN;
+
+ return result;
+}
+
+static void
+WriteBuffers(BuffersToWrite *to_write,
+ IOQueue *ioq, WritebackContext *wb_context)
+{
+ SMgrRelation smgr;
+ Buffer first_buf;
+ BufferDesc *first_buf_hdr;
+ bool needs_checksum;
+
+ Assert(to_write->nbuffers > 0 && to_write->nbuffers <= io_combine_limit);
+
+ first_buf = to_write->buffers[0];
+ first_buf_hdr = GetBufferDescriptor(first_buf - 1);
+
+ smgr = smgropen(BufTagGetRelFileLocator(&first_buf_hdr->tag), INVALID_PROC_NUMBER);
+
+ /*
+ * Force XLOG flush up to buffer's LSN. This implements the basic WAL
+ * rule that log updates must hit disk before any of the data-file changes
+ * they describe do.
+ *
+ * However, this rule does not apply to unlogged relations, which will be
+ * lost after a crash anyway. Most unlogged relation pages do not bear
+ * LSNs since we never emit WAL records for them, and therefore flushing
+ * up through the buffer LSN would be useless, but harmless. However,
+ * GiST indexes use LSNs internally to track page-splits, and therefore
+ * unlogged GiST pages bear "fake" LSNs generated by
+ * GetFakeLSNForUnloggedRel. It is unlikely but possible that the fake
+ * LSN counter could advance past the WAL insertion point; and if it did
+ * happen, attempting to flush WAL through that location would fail, with
+ * disastrous system-wide consequences. To make sure that can't happen,
+ * skip the flush if the buffer isn't permanent.
+ */
+ if (to_write->max_lsn != InvalidXLogRecPtr)
+ XLogFlush(to_write->max_lsn);
+
+ /*
+ * Now it's safe to write buffer to disk. Note that no one else should
+ * have been able to write it while we were busy with log flushing because
+ * only one process at a time can set the BM_IO_IN_PROGRESS bit.
+ */
+
+ for (int nbuf = 0; nbuf < to_write->nbuffers; nbuf++)
+ {
+ Buffer cur_buf = to_write->buffers[nbuf];
+ BufferDesc *cur_buf_hdr = GetBufferDescriptor(cur_buf - 1);
+ Block bufBlock;
+ char *bufToWrite;
+
+ bufBlock = BufHdrGetBlock(cur_buf_hdr);
+ needs_checksum = PageNeedsChecksumCopy((Page) bufBlock);
+
+ /*
+ * Update page checksum if desired. Since we have only shared lock on
+ * the buffer, other processes might be updating hint bits in it, so
+ * we must copy the page to a bounce buffer if we do checksumming.
+ */
+ if (needs_checksum)
+ {
+ PgAioBounceBuffer *bb = pgaio_bounce_buffer_get();
+
+ pgaio_io_assoc_bounce_buffer(to_write->ioh, bb);
+
+ bufToWrite = pgaio_bounce_buffer_buffer(bb);
+ memcpy(bufToWrite, bufBlock, BLCKSZ);
+ PageSetChecksumInplace((Page) bufToWrite, cur_buf_hdr->tag.blockNum);
+ }
+ else
+ {
+ bufToWrite = bufBlock;
+ }
+
+ to_write->data_ptrs[nbuf] = bufToWrite;
+ }
+
+ pgaio_io_set_io_data_32(to_write->ioh,
+ (uint32 *) to_write->buffers,
+ to_write->nbuffers);
+ pgaio_io_add_shared_cb(to_write->ioh, ASC_SHARED_BUFFER_WRITE);
+
+ smgrstartwritev(to_write->ioh, smgr,
+ BufTagGetForkNum(&first_buf_hdr->tag),
+ first_buf_hdr->tag.blockNum,
+ to_write->data_ptrs,
+ to_write->nbuffers,
+ false);
+ pgstat_count_io_op_n(IOOBJECT_RELATION, IOCONTEXT_NORMAL,
+ IOOP_WRITE, to_write->nbuffers);
+
+
+ for (int nbuf = 0; nbuf < to_write->nbuffers; nbuf++)
+ {
+ Buffer cur_buf = to_write->buffers[nbuf];
+ BufferDesc *cur_buf_hdr = GetBufferDescriptor(cur_buf - 1);
+
+ UnpinBuffer(cur_buf_hdr);
+ }
+
+ io_queue_track(ioq, &to_write->ior);
+ to_write->total_writes++;
- return result | BUF_WRITTEN;
+ /* clear state for next write */
+ to_write->nbuffers = 0;
+ to_write->start_at_tag.relNumber = InvalidOid;
+ to_write->start_at_tag.blockNum = InvalidBlockNumber;
+ to_write->max_combine = 0;
+ to_write->max_lsn = InvalidXLogRecPtr;
+ to_write->ioh = NULL;
+ pgaio_io_ref_clear(&to_write->ior);
}
/*
@@ -4087,6 +4559,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
error_context_stack = errcallback.previous;
}
+
/*
* RelationGetNumberOfBlocksInFork
* Determines the current number of pages in the specified relation fork.
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index aa264f61b9c..1f6b982c7e9 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -1480,6 +1480,16 @@ PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
return true;
}
+bool
+PageNeedsChecksumCopy(Page page)
+{
+ if (PageIsNew(page))
+ return false;
+
+ /* If we don't need a checksum, just return the passed-in data */
+ return DataChecksumsEnabled();
+}
+
/*
* Set checksum for a page in shared buffers.
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index ca1e3427bc1..cdfef5698e7 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -345,6 +345,7 @@ BufferManagerRelation
BufferStrategyControl
BufferTag
BufferUsage
+BuffersToWrite
BuildAccumulator
BuiltinScript
BulkInsertState
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0018-very-wip-test_aio-module.patch (42.2K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/19-v2-0018-very-wip-test_aio-module.patch)
download | inline diff:
From bdc7ed519ced00b6cc7fd7eb8137d5d79d846353 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sat, 31 Aug 2024 21:13:48 -0400
Subject: [PATCH v2 18/20] very-wip: test_aio module
Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
src/include/storage/aio_internal.h | 10 +
src/include/storage/buf_internals.h | 4 +
src/backend/storage/aio/aio.c | 38 ++
src/backend/storage/buffer/bufmgr.c | 3 +-
src/test/modules/Makefile | 1 +
src/test/modules/meson.build | 1 +
src/test/modules/test_aio/.gitignore | 6 +
src/test/modules/test_aio/Makefile | 34 ++
src/test/modules/test_aio/expected/inject.out | 295 ++++++++++
src/test/modules/test_aio/expected/io.out | 40 ++
.../modules/test_aio/expected/ownership.out | 148 +++++
src/test/modules/test_aio/expected/prep.out | 17 +
src/test/modules/test_aio/io_uring.conf | 5 +
src/test/modules/test_aio/meson.build | 78 +++
src/test/modules/test_aio/sql/inject.sql | 84 +++
src/test/modules/test_aio/sql/io.sql | 16 +
src/test/modules/test_aio/sql/ownership.sql | 65 +++
src/test/modules/test_aio/sql/prep.sql | 9 +
src/test/modules/test_aio/sync.conf | 5 +
src/test/modules/test_aio/test_aio--1.0.sql | 99 ++++
src/test/modules/test_aio/test_aio.c | 504 ++++++++++++++++++
src/test/modules/test_aio/test_aio.control | 3 +
src/test/modules/test_aio/worker.conf | 5 +
23 files changed, 1468 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_aio/.gitignore
create mode 100644 src/test/modules/test_aio/Makefile
create mode 100644 src/test/modules/test_aio/expected/inject.out
create mode 100644 src/test/modules/test_aio/expected/io.out
create mode 100644 src/test/modules/test_aio/expected/ownership.out
create mode 100644 src/test/modules/test_aio/expected/prep.out
create mode 100644 src/test/modules/test_aio/io_uring.conf
create mode 100644 src/test/modules/test_aio/meson.build
create mode 100644 src/test/modules/test_aio/sql/inject.sql
create mode 100644 src/test/modules/test_aio/sql/io.sql
create mode 100644 src/test/modules/test_aio/sql/ownership.sql
create mode 100644 src/test/modules/test_aio/sql/prep.sql
create mode 100644 src/test/modules/test_aio/sync.conf
create mode 100644 src/test/modules/test_aio/test_aio--1.0.sql
create mode 100644 src/test/modules/test_aio/test_aio.c
create mode 100644 src/test/modules/test_aio/test_aio.control
create mode 100644 src/test/modules/test_aio/worker.conf
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
index 2065bde79c3..f4c57438dd4 100644
--- a/src/include/storage/aio_internal.h
+++ b/src/include/storage/aio_internal.h
@@ -265,6 +265,16 @@ extern const char *pgaio_io_get_op_name(PgAioHandle *ioh);
extern const char *pgaio_io_get_state_name(PgAioHandle *ioh);
+
+/* These functions are just for use in tests, from within injection points */
+#ifdef USE_INJECTION_POINTS
+
+extern PgAioHandle *pgaio_inj_io_get(void);
+
+#endif
+
+
+
/* Declarations for the tables of function pointers exposed by each IO method. */
extern const IoMethodOps pgaio_sync_ops;
extern const IoMethodOps pgaio_worker_ops;
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 9d3123663b3..1b3329a25b4 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -423,6 +423,10 @@ extern void IssuePendingWritebacks(WritebackContext *wb_context, IOContext io_co
extern void ScheduleBufferTagForWriteback(WritebackContext *wb_context,
IOContext io_context, BufferTag *tag);
+/* solely to make it easier to write tests */
+extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
+
+
/* freelist.c */
extern IOContext IOContextForStrategy(BufferAccessStrategy strategy);
extern BufferDesc *StrategyGetBuffer(BufferAccessStrategy strategy,
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index e829e1752ca..261a752fb80 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -46,6 +46,9 @@
#include "utils/resowner.h"
#include "utils/wait_event_types.h"
+#ifdef USE_INJECTION_POINTS
+#include "utils/injection_point.h"
+#endif
static inline void pgaio_io_update_state(PgAioHandle *ioh, PgAioHandleState new_state);
@@ -92,6 +95,11 @@ static const IoMethodOps *pgaio_ops_table[] = {
const IoMethodOps *pgaio_impl;
+#ifdef USE_INJECTION_POINTS
+static PgAioHandle *inj_cur_handle;
+#endif
+
+
/* --------------------------------------------------------------------------------
* "Core" IO Api
@@ -631,6 +639,19 @@ pgaio_io_process_completion(PgAioHandle *ioh, int result)
pgaio_io_update_state(ioh, AHS_REAPED);
+#ifdef USE_INJECTION_POINTS
+ inj_cur_handle = ioh;
+
+ /*
+ * FIXME: This could be in a critical section - but it looks like we can't
+ * just InjectionPointLoad() at process start, as the injection point
+ * might not yet be defined.
+ */
+ InjectionPointCached("AIO_PROCESS_COMPLETION_BEFORE_SHARED");
+
+ inj_cur_handle = NULL;
+#endif
+
pgaio_io_process_completion_subject(ioh);
pgaio_io_update_state(ioh, AHS_COMPLETED_SHARED);
@@ -1129,3 +1150,20 @@ assign_io_method(int newval, void *extra)
{
pgaio_impl = pgaio_ops_table[newval];
}
+
+
+
+/* --------------------------------------------------------------------------------
+ * Injection point support
+ * --------------------------------------------------------------------------------
+ */
+
+#ifdef USE_INJECTION_POINTS
+
+PgAioHandle *
+pgaio_inj_io_get(void)
+{
+ return inj_cur_handle;
+}
+
+#endif
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 863464f12da..4a022440ada 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -514,7 +514,6 @@ static void UnpinBufferNoOwner(BufferDesc *buf);
static void BufferSync(int flags);
static uint32 WaitBufHdrUnlocked(BufferDesc *buf);
static void WaitIO(BufferDesc *buf);
-static bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
static void TerminateBufferIO(BufferDesc *buf, bool clear_dirty,
uint32 set_flag_bits, bool forget_owner,
bool syncio);
@@ -6213,7 +6212,7 @@ WaitIO(BufferDesc *buf)
* find out if they can perform the I/O as part of a larger operation, without
* waiting for the answer or distinguishing the reasons why not.
*/
-static bool
+bool
StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
{
uint32 buf_state;
diff --git a/src/test/modules/Makefile b/src/test/modules/Makefile
index c0d3cf0e14b..73ff9c55687 100644
--- a/src/test/modules/Makefile
+++ b/src/test/modules/Makefile
@@ -13,6 +13,7 @@ SUBDIRS = \
libpq_pipeline \
plsample \
spgist_name_ops \
+ test_aio \
test_bloomfilter \
test_copy_callbacks \
test_custom_rmgrs \
diff --git a/src/test/modules/meson.build b/src/test/modules/meson.build
index c829b619530..61c854a9b63 100644
--- a/src/test/modules/meson.build
+++ b/src/test/modules/meson.build
@@ -1,5 +1,6 @@
# Copyright (c) 2022-2024, PostgreSQL Global Development Group
+subdir('test_aio')
subdir('brin')
subdir('commit_ts')
subdir('delay_execution')
diff --git a/src/test/modules/test_aio/.gitignore b/src/test/modules/test_aio/.gitignore
new file mode 100644
index 00000000000..b4903eba657
--- /dev/null
+++ b/src/test/modules/test_aio/.gitignore
@@ -0,0 +1,6 @@
+# Generated subdirectories
+/log/
+/results/
+/output_iso/
+/tmp_check/
+/tmp_check_iso/
diff --git a/src/test/modules/test_aio/Makefile b/src/test/modules/test_aio/Makefile
new file mode 100644
index 00000000000..ae6d685835b
--- /dev/null
+++ b/src/test/modules/test_aio/Makefile
@@ -0,0 +1,34 @@
+# src/test/modules/delay_execution/Makefile
+
+PGFILEDESC = "test_aio - test code for AIO"
+
+MODULE_big = test_aio
+OBJS = \
+ $(WIN32RES) \
+ test_aio.o
+
+EXTENSION = test_aio
+DATA = test_aio--1.0.sql
+
+REGRESS = prep ownership io
+
+ifeq ($(enable_injection_points),yes)
+REGRESS += inject
+endif
+
+# FIXME: with meson this runs the tests once with worker and once - if
+# supported - with io_uring.
+
+# requires custom config
+NO_INSTALLCHECK = 1
+
+ifdef USE_PGXS
+PG_CONFIG = pg_config
+PGXS := $(shell $(PG_CONFIG) --pgxs)
+include $(PGXS)
+else
+subdir = src/test/modules/test_aio
+top_builddir = ../../../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+endif
diff --git a/src/test/modules/test_aio/expected/inject.out b/src/test/modules/test_aio/expected/inject.out
new file mode 100644
index 00000000000..e62e3718845
--- /dev/null
+++ b/src/test/modules/test_aio/expected/inject.out
@@ -0,0 +1,295 @@
+SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+ count
+-------
+ 1
+(1 row)
+
+-- injected what we'd expect
+SELECT inj_io_short_read_attach(8192);
+ inj_io_short_read_attach
+--------------------------
+
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 2);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+ count
+-------
+ 1
+(1 row)
+
+SELECT inj_io_short_read_detach();
+ inj_io_short_read_detach
+--------------------------
+
+(1 row)
+
+-- injected a read shorter than a single block, expecting error
+SELECT inj_io_short_read_attach(17);
+ inj_io_short_read_attach
+--------------------------
+
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 2);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT redact($$
+ SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+$$);
+NOTICE: wrapped error: could not read blocks 2..2 in file base/<redacted>: read only 0 of 8192 bytes
+ redact
+--------
+ f
+(1 row)
+
+SELECT inj_io_short_read_detach();
+ inj_io_short_read_detach
+--------------------------
+
+(1 row)
+
+-- shorten multi-block read to a single block, should retry
+SELECT count(*) FROM tbl_b; -- for comparison
+ count
+-------
+ 10000
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 0);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 1);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 2);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT inj_io_short_read_attach(8192);
+ inj_io_short_read_attach
+--------------------------
+
+(1 row)
+
+-- no need to redact, no messages to client
+SELECT count(*) FROM tbl_b;
+ count
+-------
+ 10000
+(1 row)
+
+SELECT inj_io_short_read_detach();
+ inj_io_short_read_detach
+--------------------------
+
+(1 row)
+
+-- shorten multi-block read to 1 1/2 blocks, should retry
+SELECT count(*) FROM tbl_b; -- for comparison
+ count
+-------
+ 10000
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 0);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 1);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 2);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT inj_io_short_read_attach(8192 + 4096);
+ inj_io_short_read_attach
+--------------------------
+
+(1 row)
+
+-- no need to redact, no messages to client
+SELECT count(*) FROM tbl_b;
+ count
+-------
+ 10000
+(1 row)
+
+SELECT inj_io_short_read_detach();
+ inj_io_short_read_detach
+--------------------------
+
+(1 row)
+
+-- shorten single-block read to read that block partially, we'll error out,
+-- because we assume we can read at least one block at a time.
+SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)'; -- for comparison
+ count
+-------
+ 1
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 2);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT inj_io_short_read_attach(4096);
+ inj_io_short_read_attach
+--------------------------
+
+(1 row)
+
+SELECT redact($$
+ SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+$$);
+NOTICE: wrapped error: could not read blocks 2..2 in file base/<redacted>: read only 0 of 8192 bytes
+ redact
+--------
+ f
+(1 row)
+
+SELECT inj_io_short_read_detach();
+ inj_io_short_read_detach
+--------------------------
+
+(1 row)
+
+-- shorten single-block read to read 0 bytes, expect that to error out
+SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)'; -- for comparison
+ count
+-------
+ 1
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 2);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT inj_io_short_read_attach(0);
+ inj_io_short_read_attach
+--------------------------
+
+(1 row)
+
+SELECT redact($$
+ SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+$$);
+NOTICE: wrapped error: could not read blocks 2..2 in file base/<redacted>: read only 0 of 8192 bytes
+ redact
+--------
+ f
+(1 row)
+
+SELECT inj_io_short_read_detach();
+ inj_io_short_read_detach
+--------------------------
+
+(1 row)
+
+-- verify that checksum errors are detected even as part of a shortened
+-- multi-block read
+-- (tbl_a, block 1 is corrupted)
+SELECT redact($$
+ SELECT count(*) FROM tbl_a WHERE ctid < '(2, 1)';
+$$);
+NOTICE: wrapped error: invalid page in block 2 of relation base/<redacted>
+ redact
+--------
+ f
+(1 row)
+
+SELECT inj_io_short_read_attach(8192);
+ inj_io_short_read_attach
+--------------------------
+
+(1 row)
+
+SELECT invalidate_rel_block('tbl_a', 0);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT invalidate_rel_block('tbl_a', 1);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT invalidate_rel_block('tbl_a', 2);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT redact($$
+ SELECT count(*) FROM tbl_a WHERE ctid < '(2, 1)';
+$$);
+NOTICE: wrapped error: invalid page in block 2 of relation base/<redacted>
+ redact
+--------
+ f
+(1 row)
+
+SELECT inj_io_short_read_detach();
+ inj_io_short_read_detach
+--------------------------
+
+(1 row)
+
+-- trigger a hard error, should error out
+SELECT inj_io_short_read_attach(-errno_from_string('EIO'));
+ inj_io_short_read_attach
+--------------------------
+
+(1 row)
+
+SELECT invalidate_rel_block('tbl_b', 2);
+ invalidate_rel_block
+----------------------
+
+(1 row)
+
+SELECT redact($$
+ SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+$$);
+NOTICE: wrapped error: could not read blocks 2..3 in file base/<redacted>: Input/output error
+ redact
+--------
+ f
+(1 row)
+
+SELECT inj_io_short_read_detach();
+ inj_io_short_read_detach
+--------------------------
+
+(1 row)
+
diff --git a/src/test/modules/test_aio/expected/io.out b/src/test/modules/test_aio/expected/io.out
new file mode 100644
index 00000000000..e46b582f290
--- /dev/null
+++ b/src/test/modules/test_aio/expected/io.out
@@ -0,0 +1,40 @@
+SELECT count(*) FROM tbl_a WHERE ctid = '(1, 1)';
+ count
+-------
+ 1
+(1 row)
+
+SELECT corrupt_rel_block('tbl_a', 1);
+ corrupt_rel_block
+-------------------
+
+(1 row)
+
+-- FIXME: Should report the error
+SELECT redact($$
+ SELECT read_corrupt_rel_block('tbl_a', 1);
+$$);
+ redact
+--------
+ t
+(1 row)
+
+-- verify the error is reported
+SELECT redact($$
+ SELECT count(*) FROM tbl_a WHERE ctid = '(1, 1)';
+$$);
+NOTICE: wrapped error: invalid page in block 2 of relation base/<redacted>
+ redact
+--------
+ f
+(1 row)
+
+SELECT redact($$
+ SELECT count(*) FROM tbl_a;
+$$);
+NOTICE: wrapped error: invalid page in block 2 of relation base/<redacted>
+ redact
+--------
+ f
+(1 row)
+
diff --git a/src/test/modules/test_aio/expected/ownership.out b/src/test/modules/test_aio/expected/ownership.out
new file mode 100644
index 00000000000..97fdad6c629
--- /dev/null
+++ b/src/test/modules/test_aio/expected/ownership.out
@@ -0,0 +1,148 @@
+-----
+-- IO handles
+----
+-- leak warning: implicit xact
+SELECT handle_get();
+WARNING: leaked AIO handle
+ handle_get
+------------
+
+(1 row)
+
+-- leak warning: explicit xact
+BEGIN; SELECT handle_get(); COMMIT;
+WARNING: leaked AIO handle
+ handle_get
+------------
+
+(1 row)
+
+-- leak warning + error: released in different command (thus resowner)
+BEGIN; SELECT handle_get(); SELECT handle_release_last(); COMMIT;
+WARNING: leaked AIO handle
+ handle_get
+------------
+
+(1 row)
+
+ERROR: release in unexpected state
+-- no leak, same command
+BEGIN; SELECT handle_get() UNION ALL SELECT handle_release_last(); COMMIT;
+ handle_get
+------------
+
+
+(2 rows)
+
+-- leak warning: subtrans
+BEGIN; SAVEPOINT foo; SELECT handle_get(); COMMIT;
+WARNING: leaked AIO handle
+ handle_get
+------------
+
+(1 row)
+
+-- normal handle use
+SELECT handle_get_release();
+ handle_get_release
+--------------------
+
+(1 row)
+
+-- should error out, API violation
+SELECT handle_get_twice();
+ERROR: API violation: Only one IO can be handed out
+-- recover after error in implicit xact
+SELECT handle_get_and_error(); SELECT handle_get_release();
+ERROR: as you command
+ handle_get_release
+--------------------
+
+(1 row)
+
+-- recover after error in explicit xact
+BEGIN; SELECT handle_get_and_error(); ROLLBACK; SELECT handle_get_release();
+ERROR: as you command
+ handle_get_release
+--------------------
+
+(1 row)
+
+-- recover after error in subtrans
+BEGIN; SAVEPOINT foo; SELECT handle_get_and_error(); ROLLBACK TO SAVEPOINT foo; SELECT handle_get_release(); ROLLBACK;
+ERROR: as you command
+ handle_get_release
+--------------------
+
+(1 row)
+
+-----
+-- Bounce Buffers handles
+----
+-- leak warning: implicit xact
+SELECT bb_get();
+WARNING: leaked AIO bounce buffer
+ bb_get
+--------
+
+(1 row)
+
+-- leak warning: explicit xact
+BEGIN; SELECT bb_get(); COMMIT;
+WARNING: leaked AIO bounce buffer
+ bb_get
+--------
+
+(1 row)
+
+-- missing leak warning: we should warn at command boundaries, not xact boundaries
+BEGIN; SELECT bb_get(); SELECT bb_release_last(); COMMIT;
+WARNING: leaked AIO bounce buffer
+ bb_get
+--------
+
+(1 row)
+
+ERROR: can only release handed out BB
+-- leak warning: subtrans
+BEGIN; SAVEPOINT foo; SELECT bb_get(); COMMIT;
+WARNING: leaked AIO bounce buffer
+ bb_get
+--------
+
+(1 row)
+
+-- normal bb use
+SELECT bb_get_release();
+ bb_get_release
+----------------
+
+(1 row)
+
+-- should error out, API violation
+SELECT bb_get_twice();
+ERROR: can only hand out one BB
+-- recover after error in implicit xact
+SELECT bb_get_and_error(); SELECT bb_get_release();
+ERROR: as you command
+ bb_get_release
+----------------
+
+(1 row)
+
+-- recover after error in explicit xact
+BEGIN; SELECT bb_get_and_error(); ROLLBACK; SELECT bb_get_release();
+ERROR: as you command
+ bb_get_release
+----------------
+
+(1 row)
+
+-- recover after error in subtrans
+BEGIN; SAVEPOINT foo; SELECT bb_get_and_error(); ROLLBACK TO SAVEPOINT foo; SELECT bb_get_release(); ROLLBACK;
+ERROR: as you command
+ bb_get_release
+----------------
+
+(1 row)
+
diff --git a/src/test/modules/test_aio/expected/prep.out b/src/test/modules/test_aio/expected/prep.out
new file mode 100644
index 00000000000..7fad6280db5
--- /dev/null
+++ b/src/test/modules/test_aio/expected/prep.out
@@ -0,0 +1,17 @@
+CREATE EXTENSION test_aio;
+CREATE TABLE tbl_a(data int not null);
+CREATE TABLE tbl_b(data int not null);
+INSERT INTO tbl_a SELECT generate_series(1, 10000);
+INSERT INTO tbl_b SELECT generate_series(1, 10000);
+SELECT grow_rel('tbl_a', 500);
+ grow_rel
+----------
+
+(1 row)
+
+SELECT grow_rel('tbl_b', 550);
+ grow_rel
+----------
+
+(1 row)
+
diff --git a/src/test/modules/test_aio/io_uring.conf b/src/test/modules/test_aio/io_uring.conf
new file mode 100644
index 00000000000..efd7ad143ff
--- /dev/null
+++ b/src/test/modules/test_aio/io_uring.conf
@@ -0,0 +1,5 @@
+shared_preload_libraries=test_aio
+io_method = 'io_uring'
+log_min_messages = 'DEBUG3'
+log_statement=all
+restart_after_crash=false
diff --git a/src/test/modules/test_aio/meson.build b/src/test/modules/test_aio/meson.build
new file mode 100644
index 00000000000..a4bef0ceeb0
--- /dev/null
+++ b/src/test/modules/test_aio/meson.build
@@ -0,0 +1,78 @@
+# Copyright (c) 2022-2024, PostgreSQL Global Development Group
+
+test_aio_sources = files(
+ 'test_aio.c',
+)
+
+if host_system == 'windows'
+ test_aio_sources += rc_lib_gen.process(win32ver_rc, extra_args: [
+ '--NAME', 'test_aio',
+ '--FILEDESC', 'test_aio - test code for AIO',])
+endif
+
+test_aio = shared_module('test_aio',
+ test_aio_sources,
+ kwargs: pg_test_mod_args,
+)
+test_install_libs += test_aio
+
+test_install_data += files(
+ 'test_aio.control',
+ 'test_aio--1.0.sql',
+)
+
+
+testfiles = [
+ 'prep',
+ 'ownership',
+ 'io',
+]
+
+if get_option('injection_points')
+ testfiles += 'inject'
+endif
+
+
+tests += {
+ 'name': 'test_aio_sync',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': testfiles,
+ 'regress_args': [
+ '--temp-config', files('sync.conf'),
+ ],
+ # requires custom config
+ 'runningcheck': false,
+ },
+}
+
+tests += {
+ 'name': 'test_aio_worker',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': testfiles,
+ 'regress_args': [
+ '--temp-config', files('worker.conf'),
+ ],
+ # requires custom config
+ 'runningcheck': false,
+ },
+}
+
+if liburing.found()
+ tests += {
+ 'name': 'test_aio_uring',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': testfiles,
+ 'regress_args': [
+ '--temp-config', files('io_uring.conf'),
+ ],
+ # requires custom config
+ 'runningcheck': false,
+ }
+ }
+endif
diff --git a/src/test/modules/test_aio/sql/inject.sql b/src/test/modules/test_aio/sql/inject.sql
new file mode 100644
index 00000000000..1190531f5ad
--- /dev/null
+++ b/src/test/modules/test_aio/sql/inject.sql
@@ -0,0 +1,84 @@
+SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+
+-- injected what we'd expect
+SELECT inj_io_short_read_attach(8192);
+SELECT invalidate_rel_block('tbl_b', 2);
+SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+SELECT inj_io_short_read_detach();
+
+
+-- injected a read shorter than a single block, expecting error
+SELECT inj_io_short_read_attach(17);
+SELECT invalidate_rel_block('tbl_b', 2);
+SELECT redact($$
+ SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+$$);
+SELECT inj_io_short_read_detach();
+
+
+-- shorten multi-block read to a single block, should retry
+SELECT count(*) FROM tbl_b; -- for comparison
+SELECT invalidate_rel_block('tbl_b', 0);
+SELECT invalidate_rel_block('tbl_b', 1);
+SELECT invalidate_rel_block('tbl_b', 2);
+SELECT inj_io_short_read_attach(8192);
+-- no need to redact, no messages to client
+SELECT count(*) FROM tbl_b;
+SELECT inj_io_short_read_detach();
+
+
+-- shorten multi-block read to 1 1/2 blocks, should retry
+SELECT count(*) FROM tbl_b; -- for comparison
+SELECT invalidate_rel_block('tbl_b', 0);
+SELECT invalidate_rel_block('tbl_b', 1);
+SELECT invalidate_rel_block('tbl_b', 2);
+SELECT inj_io_short_read_attach(8192 + 4096);
+-- no need to redact, no messages to client
+SELECT count(*) FROM tbl_b;
+SELECT inj_io_short_read_detach();
+
+
+-- shorten single-block read to read that block partially, we'll error out,
+-- because we assume we can read at least one block at a time.
+SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)'; -- for comparison
+SELECT invalidate_rel_block('tbl_b', 2);
+SELECT inj_io_short_read_attach(4096);
+SELECT redact($$
+ SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+$$);
+SELECT inj_io_short_read_detach();
+
+
+-- shorten single-block read to read 0 bytes, expect that to error out
+SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)'; -- for comparison
+SELECT invalidate_rel_block('tbl_b', 2);
+SELECT inj_io_short_read_attach(0);
+SELECT redact($$
+ SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+$$);
+SELECT inj_io_short_read_detach();
+
+
+-- verify that checksum errors are detected even as part of a shortened
+-- multi-block read
+-- (tbl_a, block 1 is corrupted)
+SELECT redact($$
+ SELECT count(*) FROM tbl_a WHERE ctid < '(2, 1)';
+$$);
+SELECT inj_io_short_read_attach(8192);
+SELECT invalidate_rel_block('tbl_a', 0);
+SELECT invalidate_rel_block('tbl_a', 1);
+SELECT invalidate_rel_block('tbl_a', 2);
+SELECT redact($$
+ SELECT count(*) FROM tbl_a WHERE ctid < '(2, 1)';
+$$);
+SELECT inj_io_short_read_detach();
+
+
+-- trigger a hard error, should error out
+SELECT inj_io_short_read_attach(-errno_from_string('EIO'));
+SELECT invalidate_rel_block('tbl_b', 2);
+SELECT redact($$
+ SELECT count(*) FROM tbl_b WHERE ctid = '(2, 1)';
+$$);
+SELECT inj_io_short_read_detach();
diff --git a/src/test/modules/test_aio/sql/io.sql b/src/test/modules/test_aio/sql/io.sql
new file mode 100644
index 00000000000..a29bb4eb15d
--- /dev/null
+++ b/src/test/modules/test_aio/sql/io.sql
@@ -0,0 +1,16 @@
+SELECT count(*) FROM tbl_a WHERE ctid = '(1, 1)';
+
+SELECT corrupt_rel_block('tbl_a', 1);
+
+-- FIXME: Should report the error
+SELECT redact($$
+ SELECT read_corrupt_rel_block('tbl_a', 1);
+$$);
+
+-- verify the error is reported
+SELECT redact($$
+ SELECT count(*) FROM tbl_a WHERE ctid = '(1, 1)';
+$$);
+SELECT redact($$
+ SELECT count(*) FROM tbl_a;
+$$);
diff --git a/src/test/modules/test_aio/sql/ownership.sql b/src/test/modules/test_aio/sql/ownership.sql
new file mode 100644
index 00000000000..63cf40c802a
--- /dev/null
+++ b/src/test/modules/test_aio/sql/ownership.sql
@@ -0,0 +1,65 @@
+-----
+-- IO handles
+----
+
+-- leak warning: implicit xact
+SELECT handle_get();
+
+-- leak warning: explicit xact
+BEGIN; SELECT handle_get(); COMMIT;
+
+-- leak warning + error: released in different command (thus resowner)
+BEGIN; SELECT handle_get(); SELECT handle_release_last(); COMMIT;
+
+-- no leak, same command
+BEGIN; SELECT handle_get() UNION ALL SELECT handle_release_last(); COMMIT;
+
+-- leak warning: subtrans
+BEGIN; SAVEPOINT foo; SELECT handle_get(); COMMIT;
+
+-- normal handle use
+SELECT handle_get_release();
+
+-- should error out, API violation
+SELECT handle_get_twice();
+
+-- recover after error in implicit xact
+SELECT handle_get_and_error(); SELECT handle_get_release();
+
+-- recover after error in explicit xact
+BEGIN; SELECT handle_get_and_error(); ROLLBACK; SELECT handle_get_release();
+
+-- recover after error in subtrans
+BEGIN; SAVEPOINT foo; SELECT handle_get_and_error(); ROLLBACK TO SAVEPOINT foo; SELECT handle_get_release(); ROLLBACK;
+
+
+-----
+-- Bounce Buffers handles
+----
+
+-- leak warning: implicit xact
+SELECT bb_get();
+
+-- leak warning: explicit xact
+BEGIN; SELECT bb_get(); COMMIT;
+
+-- missing leak warning: we should warn at command boundaries, not xact boundaries
+BEGIN; SELECT bb_get(); SELECT bb_release_last(); COMMIT;
+
+-- leak warning: subtrans
+BEGIN; SAVEPOINT foo; SELECT bb_get(); COMMIT;
+
+-- normal bb use
+SELECT bb_get_release();
+
+-- should error out, API violation
+SELECT bb_get_twice();
+
+-- recover after error in implicit xact
+SELECT bb_get_and_error(); SELECT bb_get_release();
+
+-- recover after error in explicit xact
+BEGIN; SELECT bb_get_and_error(); ROLLBACK; SELECT bb_get_release();
+
+-- recover after error in subtrans
+BEGIN; SAVEPOINT foo; SELECT bb_get_and_error(); ROLLBACK TO SAVEPOINT foo; SELECT bb_get_release(); ROLLBACK;
diff --git a/src/test/modules/test_aio/sql/prep.sql b/src/test/modules/test_aio/sql/prep.sql
new file mode 100644
index 00000000000..b8f225cbc98
--- /dev/null
+++ b/src/test/modules/test_aio/sql/prep.sql
@@ -0,0 +1,9 @@
+CREATE EXTENSION test_aio;
+
+CREATE TABLE tbl_a(data int not null);
+CREATE TABLE tbl_b(data int not null);
+
+INSERT INTO tbl_a SELECT generate_series(1, 10000);
+INSERT INTO tbl_b SELECT generate_series(1, 10000);
+SELECT grow_rel('tbl_a', 500);
+SELECT grow_rel('tbl_b', 550);
diff --git a/src/test/modules/test_aio/sync.conf b/src/test/modules/test_aio/sync.conf
new file mode 100644
index 00000000000..c480922d6cf
--- /dev/null
+++ b/src/test/modules/test_aio/sync.conf
@@ -0,0 +1,5 @@
+shared_preload_libraries=test_aio
+io_method = 'sync'
+log_min_messages = 'DEBUG3'
+log_statement=all
+restart_after_crash=false
diff --git a/src/test/modules/test_aio/test_aio--1.0.sql b/src/test/modules/test_aio/test_aio--1.0.sql
new file mode 100644
index 00000000000..e3d5ce29c60
--- /dev/null
+++ b/src/test/modules/test_aio/test_aio--1.0.sql
@@ -0,0 +1,99 @@
+/* src/test/modules/test_aio/test_aio--1.0.sql */
+
+-- complain if script is sourced in psql, rather than via CREATE EXTENSION
+\echo Use "CREATE EXTENSION test_aio" to load this file. \quit
+
+
+CREATE FUNCTION errno_from_string(sym text)
+RETURNS pg_catalog.int4 STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+CREATE FUNCTION grow_rel(rel regclass, nblocks int)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+CREATE FUNCTION corrupt_rel_block(rel regclass, blockno int)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION read_corrupt_rel_block(rel regclass, blockno int)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION invalidate_rel_block(rel regclass, blockno int)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION handle_get_and_error()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION handle_get_twice()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION handle_get()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION handle_get_release()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION handle_release_last()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+CREATE FUNCTION bb_get_and_error()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION bb_get_twice()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION bb_get()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION bb_get_release()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION bb_release_last()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+
+CREATE OR REPLACE FUNCTION redact(p_sql text)
+RETURNS bool
+LANGUAGE plpgsql
+AS $$
+ DECLARE
+ err_state text;
+ err_msg text;
+ BEGIN
+ EXECUTE p_sql;
+ RETURN true;
+ EXCEPTION WHEN OTHERS THEN
+ GET STACKED DIAGNOSTICS
+ err_state = RETURNED_SQLSTATE,
+ err_msg = MESSAGE_TEXT;
+ err_msg = regexp_replace(err_msg, '(file|relation) "?base/[0-9]+/[0-9]+"?', '\1 base/<redacted>');
+ RAISE NOTICE 'wrapped error: %', err_msg
+ USING ERRCODE = err_state;
+ RETURN false;
+ END;
+$$;
+
+
+CREATE FUNCTION inj_io_short_read_attach(result int)
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
+
+CREATE FUNCTION inj_io_short_read_detach()
+RETURNS pg_catalog.void STRICT
+AS 'MODULE_PATHNAME' LANGUAGE C;
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
new file mode 100644
index 00000000000..e495c5309b3
--- /dev/null
+++ b/src/test/modules/test_aio/test_aio.c
@@ -0,0 +1,504 @@
+/*-------------------------------------------------------------------------
+ *
+ * delay_execution.c
+ * Test module to allow delay between parsing and execution of a query.
+ *
+ * The delay is implemented by taking and immediately releasing a specified
+ * advisory lock. If another process has previously taken that lock, the
+ * current process will be blocked until the lock is released; otherwise,
+ * there's no effect. This allows an isolationtester script to reliably
+ * test behaviors where some specified action happens in another backend
+ * between parsing and execution of any desired query.
+ *
+ * Copyright (c) 2020-2024, PostgreSQL Global Development Group
+ *
+ * IDENTIFICATION
+ * src/test/modules/delay_execution/delay_execution.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "access/relation.h"
+#include "fmgr.h"
+#include "storage/aio.h"
+#include "storage/aio_internal.h"
+#include "storage/buf_internals.h"
+#include "storage/bufmgr.h"
+#include "storage/ipc.h"
+#include "storage/lwlock.h"
+#include "utils/builtins.h"
+#include "utils/injection_point.h"
+#include "utils/rel.h"
+
+
+PG_MODULE_MAGIC;
+
+
+typedef struct InjIoErrorState
+{
+ bool enabled;
+ bool result_set;
+ int result;
+} InjIoErrorState;
+
+static InjIoErrorState * inj_io_error_state;
+
+/* Shared memory init callbacks */
+static shmem_request_hook_type prev_shmem_request_hook = NULL;
+static shmem_startup_hook_type prev_shmem_startup_hook = NULL;
+
+
+static PgAioHandle *last_handle;
+static PgAioBounceBuffer *last_bb;
+
+
+
+static void
+test_aio_shmem_request(void)
+{
+ if (prev_shmem_request_hook)
+ prev_shmem_request_hook();
+
+ RequestAddinShmemSpace(sizeof(InjIoErrorState));
+}
+
+static void
+test_aio_shmem_startup(void)
+{
+ bool found;
+
+ if (prev_shmem_startup_hook)
+ prev_shmem_startup_hook();
+
+ /* Create or attach to the shared memory state */
+ LWLockAcquire(AddinShmemInitLock, LW_EXCLUSIVE);
+
+ inj_io_error_state = ShmemInitStruct("injection_points",
+ sizeof(InjIoErrorState),
+ &found);
+
+ if (!found)
+ {
+ /*
+ * First time through, so initialize. This is shared with the dynamic
+ * initialization using a DSM.
+ */
+ inj_io_error_state->enabled = false;
+
+#ifdef USE_INJECTION_POINTS
+ InjectionPointAttach("AIO_PROCESS_COMPLETION_BEFORE_SHARED",
+ "test_aio",
+ "inj_io_short_read",
+ NULL,
+ 0);
+ InjectionPointLoad("AIO_PROCESS_COMPLETION_BEFORE_SHARED");
+#endif
+ }
+ else
+ {
+#ifdef USE_INJECTION_POINTS
+ InjectionPointLoad("AIO_PROCESS_COMPLETION_BEFORE_SHARED");
+ elog(LOG, "injection point loaded");
+#endif
+ }
+
+ LWLockRelease(AddinShmemInitLock);
+}
+
+void
+_PG_init(void)
+{
+ if (!process_shared_preload_libraries_in_progress)
+ return;
+
+ /* Shared memory initialization */
+ prev_shmem_request_hook = shmem_request_hook;
+ shmem_request_hook = test_aio_shmem_request;
+ prev_shmem_startup_hook = shmem_startup_hook;
+ shmem_startup_hook = test_aio_shmem_startup;
+}
+
+
+PG_FUNCTION_INFO_V1(errno_from_string);
+Datum
+errno_from_string(PG_FUNCTION_ARGS)
+{
+ const char *sym = text_to_cstring(PG_GETARG_TEXT_PP(0));
+
+ if (strcmp(sym, "EIO") == 0)
+ PG_RETURN_INT32(EIO);
+ else if (strcmp(sym, "EAGAIN") == 0)
+ PG_RETURN_INT32(EAGAIN);
+ else if (strcmp(sym, "EINTR") == 0)
+ PG_RETURN_INT32(EINTR);
+ else if (strcmp(sym, "ENOSPC") == 0)
+ PG_RETURN_INT32(ENOSPC);
+ else if (strcmp(sym, "EROFS") == 0)
+ PG_RETURN_INT32(EROFS);
+
+ ereport(ERROR,
+ errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg_internal("%s is not a supported errno value", sym));
+ PG_RETURN_INT32(0);
+}
+
+
+PG_FUNCTION_INFO_V1(grow_rel);
+Datum
+grow_rel(PG_FUNCTION_ARGS)
+{
+ Oid relid = PG_GETARG_OID(0);
+ uint32 nblocks = PG_GETARG_UINT32(1);
+ Relation rel;
+#define MAX_BUFFERS_TO_EXTEND_BY 64
+ Buffer victim_buffers[MAX_BUFFERS_TO_EXTEND_BY];
+
+ rel = relation_open(relid, AccessExclusiveLock);
+
+ while (nblocks > 0)
+ {
+ uint32 extend_by_pages;
+
+ extend_by_pages = Min(nblocks, MAX_BUFFERS_TO_EXTEND_BY);
+
+ ExtendBufferedRelBy(BMR_REL(rel),
+ MAIN_FORKNUM,
+ NULL,
+ 0,
+ extend_by_pages,
+ victim_buffers,
+ &extend_by_pages);
+
+ nblocks -= extend_by_pages;
+
+ for (uint32 i = 0; i < extend_by_pages; i++)
+ {
+ ReleaseBuffer(victim_buffers[i]);
+ }
+ }
+
+ relation_close(rel, NoLock);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(corrupt_rel_block);
+Datum
+corrupt_rel_block(PG_FUNCTION_ARGS)
+{
+ Oid relid = PG_GETARG_OID(0);
+ uint32 block = PG_GETARG_UINT32(1);
+ Relation rel;
+ Buffer buf;
+ Page page;
+ PageHeader ph;
+
+ rel = relation_open(relid, AccessExclusiveLock);
+
+ buf = ReadBuffer(rel, block);
+ page = BufferGetPage(buf);
+
+ LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
+
+ MarkBufferDirty(buf);
+
+ PageInit(page, BufferGetPageSize(buf), 0);
+
+ ph = (PageHeader) page;
+ ph->pd_special = BLCKSZ + 1;
+
+ FlushOneBuffer(buf);
+
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+
+ ReleaseBuffer(buf);
+
+ EvictUnpinnedBuffer(buf);
+
+ relation_close(rel, NoLock);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(read_corrupt_rel_block);
+Datum
+read_corrupt_rel_block(PG_FUNCTION_ARGS)
+{
+ Oid relid = PG_GETARG_OID(0);
+ uint32 block = PG_GETARG_UINT32(1);
+ Relation rel;
+ Buffer buf;
+ BufferDesc *buf_hdr;
+ Page page;
+ PgAioHandle *ioh;
+ PgAioHandleRef ior;
+ SMgrRelation smgr;
+ uint32 buf_state;
+
+ rel = relation_open(relid, AccessExclusiveLock);
+
+ /* read buffer without erroring out */
+ buf = ReadBufferExtended(rel, MAIN_FORKNUM, block, RBM_ZERO_AND_LOCK, NULL);
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+
+ page = BufferGetBlock(buf);
+
+ ioh = pgaio_io_get(CurrentResourceOwner, NULL);
+ pgaio_io_get_ref(ioh, &ior);
+
+ buf_hdr = GetBufferDescriptor(buf - 1);
+ smgr = RelationGetSmgr(rel);
+
+ /* FIXME: even if just a test, we should verify nobody else uses this */
+ buf_state = LockBufHdr(buf_hdr);
+ buf_state &= ~(BM_VALID | BM_DIRTY);
+ UnlockBufHdr(buf_hdr, buf_state);
+
+ StartBufferIO(buf_hdr, true, false);
+
+ pgaio_io_set_io_data_32(ioh, (uint32 *) &buf, 1);
+ pgaio_io_add_shared_cb(ioh, ASC_SHARED_BUFFER_READ);
+
+ smgrstartreadv(ioh, smgr, MAIN_FORKNUM, block,
+ (void *) &page, 1);
+
+ ReleaseBuffer(buf);
+
+ pgaio_io_ref_wait(&ior);
+
+ relation_close(rel, NoLock);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(invalidate_rel_block);
+Datum
+invalidate_rel_block(PG_FUNCTION_ARGS)
+{
+ Oid relid = PG_GETARG_OID(0);
+ uint32 block = PG_GETARG_UINT32(1);
+ Relation rel;
+ PrefetchBufferResult pr;
+ Buffer buf;
+
+ rel = relation_open(relid, AccessExclusiveLock);
+
+ /* this is a gross hack, but there's no good API exposed */
+ pr = PrefetchBuffer(rel, MAIN_FORKNUM, block);
+ buf = pr.recent_buffer;
+ elog(LOG, "recent: %d", buf);
+ if (BufferIsValid(buf))
+ {
+ /* if the buffer contents aren't valid, this'll return false */
+ if (ReadRecentBuffer(rel->rd_locator, MAIN_FORKNUM, block, buf))
+ {
+ LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
+ FlushOneBuffer(buf);
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+ ReleaseBuffer(buf);
+
+ if (!EvictUnpinnedBuffer(buf))
+ elog(ERROR, "couldn't unpin");
+ }
+ }
+
+ relation_close(rel, AccessExclusiveLock);
+
+ PG_RETURN_VOID();
+}
+
+#if 0
+PG_FUNCTION_INFO_V1(test_unsubmitted_vs_close);
+Datum
+test_unsubmitted_vs_close(PG_FUNCTION_ARGS)
+{
+ Oid relid = PG_GETARG_OID(0);
+ uint32 block = PG_GETARG_UINT32(1);
+ Relation rel;
+ Buffer buf;
+ Page page;
+ PageHeader ph;
+
+ rel = relation_open(relid, AccessExclusiveLock);
+
+ buf = ReadBufferExtended(rel, MAIN_FORKNUM, block, RBM_ZERO_AND_LOCK, NULL);
+
+ buf = ReadBuffer(rel, block);
+ page = BufferGetPage(buf);
+
+ EvictUnpinnedBuffer(buf);
+
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+
+
+ MarkBufferDirty(buf);
+ ph->pd_special = BLCKSZ + 1;
+
+ /* last_handle = pgaio_io_get(); */
+
+ PG_RETURN_VOID();
+}
+#endif
+
+PG_FUNCTION_INFO_V1(handle_get);
+Datum
+handle_get(PG_FUNCTION_ARGS)
+{
+ last_handle = pgaio_io_get(CurrentResourceOwner, NULL);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(handle_release_last);
+Datum
+handle_release_last(PG_FUNCTION_ARGS)
+{
+ if (!last_handle)
+ elog(ERROR, "no handle");
+
+ pgaio_io_release(last_handle);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(handle_get_and_error);
+Datum
+handle_get_and_error(PG_FUNCTION_ARGS)
+{
+ pgaio_io_get(CurrentResourceOwner, NULL);
+
+ elog(ERROR, "as you command");
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(handle_get_twice);
+Datum
+handle_get_twice(PG_FUNCTION_ARGS)
+{
+ pgaio_io_get(CurrentResourceOwner, NULL);
+ pgaio_io_get(CurrentResourceOwner, NULL);
+
+ PG_RETURN_VOID();
+}
+
+
+PG_FUNCTION_INFO_V1(handle_get_release);
+Datum
+handle_get_release(PG_FUNCTION_ARGS)
+{
+ PgAioHandle *handle;
+
+ handle = pgaio_io_get(CurrentResourceOwner, NULL);
+ pgaio_io_release(handle);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(bb_get);
+Datum
+bb_get(PG_FUNCTION_ARGS)
+{
+ last_bb = pgaio_bounce_buffer_get();
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(bb_release_last);
+Datum
+bb_release_last(PG_FUNCTION_ARGS)
+{
+ if (!last_bb)
+ elog(ERROR, "no bb");
+
+ pgaio_bounce_buffer_release(last_bb);
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(bb_get_and_error);
+Datum
+bb_get_and_error(PG_FUNCTION_ARGS)
+{
+ pgaio_bounce_buffer_get();
+
+ elog(ERROR, "as you command");
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(bb_get_twice);
+Datum
+bb_get_twice(PG_FUNCTION_ARGS)
+{
+ pgaio_bounce_buffer_get();
+ pgaio_bounce_buffer_get();
+
+ PG_RETURN_VOID();
+}
+
+
+PG_FUNCTION_INFO_V1(bb_get_release);
+Datum
+bb_get_release(PG_FUNCTION_ARGS)
+{
+ PgAioBounceBuffer *bb;
+
+ bb = pgaio_bounce_buffer_get();
+ pgaio_bounce_buffer_release(bb);
+
+ PG_RETURN_VOID();
+}
+
+#ifdef USE_INJECTION_POINTS
+extern PGDLLEXPORT void inj_io_short_read(const char *name, const void *private_data);
+
+void
+inj_io_short_read(const char *name, const void *private_data)
+{
+ PgAioHandle *ioh;
+
+ elog(LOG, "short read called: %d", inj_io_error_state->enabled);
+
+ if (inj_io_error_state->enabled)
+ {
+ ioh = pgaio_inj_io_get();
+
+ if (inj_io_error_state->result_set)
+ {
+ elog(LOG, "short read, changing result from %d to %d",
+ ioh->result, inj_io_error_state->result);
+
+ ioh->result = inj_io_error_state->result;
+ }
+ }
+}
+#endif
+
+PG_FUNCTION_INFO_V1(inj_io_short_read_attach);
+Datum
+inj_io_short_read_attach(PG_FUNCTION_ARGS)
+{
+#ifdef USE_INJECTION_POINTS
+ inj_io_error_state->enabled = true;
+ inj_io_error_state->result_set = !PG_ARGISNULL(0);
+ if (inj_io_error_state->result_set)
+ inj_io_error_state->result = PG_GETARG_INT32(0);
+#else
+ elog(ERROR, "injection points not supported");
+#endif
+
+ PG_RETURN_VOID();
+}
+
+PG_FUNCTION_INFO_V1(inj_io_short_read_detach);
+Datum
+inj_io_short_read_detach(PG_FUNCTION_ARGS)
+{
+#ifdef USE_INJECTION_POINTS
+ inj_io_error_state->enabled = false;
+#else
+ elog(ERROR, "injection points not supported");
+#endif
+ PG_RETURN_VOID();
+}
diff --git a/src/test/modules/test_aio/test_aio.control b/src/test/modules/test_aio/test_aio.control
new file mode 100644
index 00000000000..cd91c3ed16b
--- /dev/null
+++ b/src/test/modules/test_aio/test_aio.control
@@ -0,0 +1,3 @@
+comment = 'Test code for AIO'
+default_version = '1.0'
+module_pathname = '$libdir/test_aio'
diff --git a/src/test/modules/test_aio/worker.conf b/src/test/modules/test_aio/worker.conf
new file mode 100644
index 00000000000..8104c201924
--- /dev/null
+++ b/src/test/modules/test_aio/worker.conf
@@ -0,0 +1,5 @@
+shared_preload_libraries=test_aio
+io_method = 'worker'
+log_min_messages = 'DEBUG3'
+log_statement=all
+restart_after_crash=false
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0019-Temporary-Increase-BAS_BULKREAD-size.patch (1.3K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/20-v2-0019-Temporary-Increase-BAS_BULKREAD-size.patch)
download | inline diff:
From 75c690243866d3f6b476ecfb9c249da8098122f0 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sun, 1 Sep 2024 00:42:27 -0400
Subject: [PATCH v2 19/20] Temporary: Increase BAS_BULKREAD size
Without this we only can execute very little AIO for sequential scans, as
there's just not enough buffers in the ring. This isn't the right fix, as
just increasing the ring size can have negative performance implications in
workloads where the kernel has all the data cached.
Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
src/backend/storage/buffer/freelist.c | 7 ++++++-
1 file changed, 6 insertions(+), 1 deletion(-)
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index dffdd57e9b5..f5795b509c7 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -555,7 +555,12 @@ GetAccessStrategy(BufferAccessStrategyType btype)
return NULL;
case BAS_BULKREAD:
- ring_size_kb = 256;
+
+ /*
+ * FIXME: Temporary increase to allow large enough streaming reads
+ * to actually benefit from AIO. This needs a better solution.
+ */
+ ring_size_kb = 2 * 1024;
break;
case BAS_BULKWRITE:
ring_size_kb = 16 * 1024;
--
2.45.2.746.g06e570c0df.dirty
[text/x-diff] v2-0020-WIP-Use-MAP_POPULATE.patch (1.1K, ../../bgixmidc73doecg7wskq3k76g3nqnglqub7irbrwp4ppjsx43j@fwre2x775mcl/21-v2-0020-WIP-Use-MAP_POPULATE.patch)
download | inline diff:
From e9c132e191cacc9fc946b611afc5f489762c4387 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 31 Dec 2024 13:25:56 -0500
Subject: [PATCH v2 20/20] WIP: Use MAP_POPULATE
For benchmarking it's quite annoying that the first time a memory is touched
has completely different perf characteristics than subsequent accesses. Using
MAP_POPULATE reduces that substantially.
---
src/backend/port/sysv_shmem.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/src/backend/port/sysv_shmem.c b/src/backend/port/sysv_shmem.c
index a5a4511f66d..2a45dffd5e0 100644
--- a/src/backend/port/sysv_shmem.c
+++ b/src/backend/port/sysv_shmem.c
@@ -620,7 +620,7 @@ CreateAnonymousSegment(Size *size)
allocsize += hugepagesize - (allocsize % hugepagesize);
ptr = mmap(NULL, allocsize, PROT_READ | PROT_WRITE,
- PG_MMAP_FLAGS | mmap_flags, -1, 0);
+ PG_MMAP_FLAGS | MAP_POPULATE | mmap_flags, -1, 0);
mmap_errno = errno;
if (huge_pages == HUGE_PAGES_TRY && ptr == MAP_FAILED)
elog(DEBUG1, "mmap(%zu) with MAP_HUGETLB failed, huge pages disabled: %m",
--
2.45.2.746.g06e570c0df.dirty
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-06 18:52 Noah Misch <noah@leadboat.com>
parent: Andres Freund <andres@anarazel.de>
2 siblings, 1 reply; 16+ messages in thread
From: Noah Misch @ 2025-01-06 18:52 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers
Patches 1 and 2 are still Ready for Committer.
On Tue, Dec 31, 2024 at 11:03:33PM -0500, Andres Freund wrote:
> - The README has been extended with an overview of the API. I think it gives a
> good overview of how the API fits together. I'd be very good to get
> feedback from folks that aren't as familiar with AIO, I can't really see
> what's easy/hard anymore.
That's a helpful addition. I've left inline comments on it, below.
> The biggest TODOs are:
>
> - Right now the API between bufmgr.c and read_stream.c kind of necessitates
> that one StartReadBuffers() call actually can trigger multiple IOs, if
> one of the buffers was read in by another backend, before "this" backend
> called StartBufferIO().
>
> I think Thomas and I figured out a way to evolve the interface so that this
> isn't necessary anymore:
>
> We allow StartReadBuffers() to memorize buffers it pinned but didn't
> initiate IO on in the buffers[] argument. The next call to StartReadBuffers
> then doesn't have to repin thse buffers. That doesn't just solve the
> multiple-IOs for one "read operation" issue, it also make the - very common
> - case of a bunch of "buffer misses" followed by a "buffer hit" cleaner, the
> hit wouldn't be tracked in the same ReadBuffersOperation anymore.
That sounds reasonable.
> - Right now bufmgr.h includes aio.h, because it needs to include a reference
> to the AIO's result in ReadBuffersOperation. Requiring a dynamic allocation
> would be noticeable overhead, so that's not an option. I think the best
> option here would be to introduce something like aio_types.h, so fewer
> things are included.
That sounds fine. Header splits aren't going to be perfect, so I'd pick
something (e.g. your proposal here) and move on.
> - There's no obvious way to tell "internal" function operating on an IO handle
> apart from functions that are expected to be called by the issuer of an IO.
>
> One way to deal with this would be to introduce a distinct "issuer IO
> reference" type. I think that might be a good idea, it would also make it
> clearer that a good number of the functions can only be called by the
> issuer, before the IO is submitted.
That's reasonable, albeit non-critical.
> - The naming around PgAioReturn, PgAioResult, PgAioResultStatus needs to be
> improved
POSIX uses the word "result" for the consequences of a function (e.g. the
result of unlink() is readdir() no longer finding the link). It uses the word
"return" for a memory value that describes a result. In that usage, the
struct currently called PgAioResult would be a Return. The struct currently
called PgAioReturn is PgAioResult plus the data to identify the IO. Possible
name changes:
PgAioResult -> PgAioReturn
PgAioReturn -> PgAioReturnIdentified | PgAioReturnID | PgAioReturnTagged [I don't love these]
PgAioResultStatus -> PgAioStatus | PgAioFill
That said, I don't dislike the existing names and would not have raised the
topic myself.
> - The debug logging functions are a bit of a mess, lots of very similar code
> in lots of places. I think AIO needs a few ereport() wrappers to make this
> easier.
May as well.
> - More tests are needed. None of our current test frameworks really makes this
> easy :(.
Which testing gap do you find most concerning? I'd be most interested in the
cases that would be undetected deadlocks under a naive design. An example
appeared at the end of postgr.es/m/20240916144349.74.nmisch@google.com
> - Several folks asked for pg_stat_aio to come back, in "v1" that showed the
> set of currently in-flight AIOs. That's not particularly hard - except
> that it doesn't really fit in the pg_stat_* namespace.
Later message
postgr.es/m/6vjl6jeaqvyhfbpgwziypwmhem2rwla4o5pgpuxwtg3o3o3jb5@evyzorb5meth is
considering the name pg_aios. Works for me.
> --- a/src/backend/storage/aio/aio.c
> +++ b/src/backend/storage/aio/aio.c
> @@ -3,6 +3,28 @@
> * aio.c
> * AIO - Core Logic
> *
> + * For documentation about how AIO works on a higher level, including a
> + * schematic example, see README.md.
> + *
> + *
> + * AIO is a complicated subsystem. To keep things navigable it is split across
> + * a number of files:
> + *
> + * - aio.c - core AIO state handling
> + *
> + * - aio_init.c - initialization
> + *
> + * - aio_io.c - dealing with actual IO, including executing IOs synchronously
> + *
> + * - aio_subject.c - functionality related to executing IO for different
> + * subjects
> + *
> + * - method_*.c - different ways of executing AIO
> + *
> + * - read_stream.c - helper for accessing buffered relation data with
> + * look-ahead
> + *
I felt like some list entries in this new header comment largely restated the
file name. Here's how I'd write them to avoid that:
* - method_*.c - different ways of executing AIO (e.g. worker process)
* - aio_io.c - method-independent code for specific IO ops (e.g. readv)
* - aio_subject.c - callbacks at IO operation lifecycle events
* - aio_init.c - per-fork and per-startup-process initialization
* - aio.c - all other topics
* - read_stream.c - helper for reading buffered relation data
> --- /dev/null
> +++ b/src/backend/storage/aio/README.md
> @@ -0,0 +1,413 @@
> +# Asynchronous & Direct IO
I would move "### Why Asynchronous IO" to here; that's good background before
getting into the example. I might also move "### Why Direct / unbuffered IO"
to here. For me as a reader, I'd benefit from seeing things in this order:
- "why"
- condensed usage example like manpage SYNOPSIS, comments and decls removed
- PgAioHandleState and discussion of valid transitions
- usage example as it is, with full comments
- the rest
In other words, like this:
# Asynchronous & Direct IO
## Motivation
### Why Asynchronous IO
[existing content moved from lower in the file]
## Synopsis
ioh = pgaio_io_get(CurrentResourceOwner, &ioret);
pgaio_io_get_ref(ioh, &ior);
pgaio_io_add_shared_cb(ioh, ASC_SHARED_BUFFER_READ);
pgaio_io_set_io_data_32(ioh, (uint32 *) buffer, 1);
smgrstartreadv(ioh, operation->smgr, forknum, blkno,
BufferGetBlock(buffer), 1);
pgaio_submit_staged();
pgaio_io_ref_wait(&ior);
if (ioret.result.status == ARS_ERROR)
pgaio_result_log(aio_ret.result, &aio_ret.subject_data, ERROR);
## I/O Operation States & Transitions
[PgAioHandleState and its transitions]
## AIO Usage Example
[your content:]
> +
> +## AIO Usage Example
> +
> +In many cases code that can benefit from AIO does not directly have to
> +interact with the AIO interface, but can use AIO via higher-level
> +abstractions. See [Helpers](#helpers).
> +
> +In this example, a buffer will be read into shared buffers.
> +
> +```C
> +/*
> + * Result of the operation, only to be accessed in this backend.
> + */
> +PgAioReturn ioret;
> +
> +/*
> + * Acquire AIO Handle, ioret will get result upon completion.
Consider adding: from here to pgaio_submit_staged(), don't do [description of
the kind of unacceptable blocking operations].
> + * Once the IO handle has been handed of, it may not further be used, as the
s/of/off/
> +### IO can be started in critical sections
...
> +The need to be able to execute IO in critical sections has substantial design
> +implication on the AIO subsystem. Mainly because completing IOs (see prior
> +section) needs to be possible within a critical section, even if the
> +to-be-completed IO itself was not issued in a critical section. Consider
> +e.g. the case of a backend first starting a number of writes from shared
> +buffers and then starting to flush the WAL. Because only a limited amount of
> +IO can be in-progress at the same time, initiating the IO for flushing the WAL
> +may require to first finish executing IO executed earlier.
The last line's two appearances of the word "execute" read awkwardly to me,
and it's an opportunity to use PgAioHandleState terms. Consider writing the
last line like "may first advance an existing IO from AHS_PREPARED to
AHS_COMPLETED_SHARED".
> +ASLR. This means that the shared memory cannot contain pointer to callbacks.
s/pointer/pointers/
> +### AIO Callbacks
...
> +In addition to completion, AIO callbacks also are called to "prepare" an
> +IO. This is, e.g., used to acquire buffer pins owned by the AIO subsystem for
> +IO to/from shared buffers, which is required to handle the case where the
> +issuing backend errors out and releases its own pins.
Reading this, it's not obvious to me how to reconcile "finishing an IO could
require pin acquisition" with "finishing an IO could happen in a critical
section". Pinning a buffer in a critical section sounds bad. I vaguely
recall understanding how it was okay as of my September review, but I've
already forgotten. Can this text have a sentence making that explicit?
> +### AIO Subjects
> +
> +In addition to the completion callbacks describe above, each AIO Handle has
> +exactly one "subject". Each subject has some space inside an AIO Handle with
> +information specific to the subject and can provide callbacks to allow to
> +reopen the underlying file (required for worker mode) and to describe the IO
> +operation (used for debug logging and error messages).
Can this say roughly how to decide when to add a new subject? Failing that,
can it give examples of what additional subjects might exist if certain
existing subsystems were to start using AIO?
> +### AIO Results
> +
> +As AIO completion callbacks
> +[are executed in critical sections](#io-can-be-started-in-critical-sections)
> +and [may be executed by any backend](#deadlock-and-starvation-dangers-due-to-aio)
> +completion callbacks cannot be used to, e.g., make the query that triggered an
> +IO ERROR out.
> +
> +To allow to react to failing IOs the issuing backend can pass a pointer to a
> +`PgAioReturn` in backend local memory. Before an AIO Handle is reused the
> +`PgAioReturn` is filled with information about the IO. This includes
> +information about whether the IO was successful (as a value of
> +`PgAioResultStatus`) and enough information to raise an error in case of a
> +failure (via `pgaio_result_log()`, with the error details encoded in
> +`PgAioResult`).
Can this have a sentence on how this fits in bounded shmem, given the lack of
guarantees about a backend's responsiveness? In other words, what makes it
okay to have requests take arbitrarily long to move from AHS_COMPLETED_SHARED
to AHS_COMPLETED_LOCAL?
Thanks,
nm
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-06 21:40 Andres Freund <andres@anarazel.de>
parent: Noah Misch <noah@leadboat.com>
0 siblings, 1 reply; 16+ messages in thread
From: Andres Freund @ 2025-01-06 21:40 UTC (permalink / raw)
To: Noah Misch <noah@leadboat.com>; +Cc: pgsql-hackers
Hi,
On 2025-01-06 10:52:20 -0800, Noah Misch wrote:
> Patches 1 and 2 are still Ready for Committer.
I feel somewhat weird about pushing 0002 without a user, but I guess it's
still exercised, so it's probably fine...
> On Tue, Dec 31, 2024 at 11:03:33PM -0500, Andres Freund wrote:
> > - The README has been extended with an overview of the API. I think it gives a
> > good overview of how the API fits together. I'd be very good to get
> > feedback from folks that aren't as familiar with AIO, I can't really see
> > what's easy/hard anymore.
>
> That's a helpful addition. I've left inline comments on it, below.
Cool!
> > - More tests are needed. None of our current test frameworks really makes this
> > easy :(.
>
> Which testing gap do you find most concerning?
Most of it isn't even AIO specific...
- temporary tables are rather poorly tested in general:
- e.g. trivial to exceed the number of buffers, but our tests don't reach that
- We have pretty no testing for IO errors. We have a bit of coverage due to
src/bin/pg_amcheck/t/003_check.pl, but that's for errors originating in
bufmgr.c itself.
- no real testing of StartBufferIO's etc wait paths
- no testing for BM_PIN_COUNT_WAITER
I e.g. just noticed that the error handling for AIO on temp tables was broken
- but our tests never reach that:
The bug exists due to temp tables not differentiating between "backend" pins
and a "global pincount" - which means that there's no real way for the AIO
subsystem to have a reference separate from the backend local pin -
CheckForLocalBufferLeaks() complains about any leftover pins. It seems to
works in non-assert mode, but with assertions transaction abort asserts out.
> I'd be most interested in the
> cases that would be undetected deadlocks under a naive design. An example
> appeared at the end of postgr.es/m/20240916144349.74.nmisch@google.com
That's a good one, yea.
I think I'll try to translate the regression tests I wrote into an isolation
test, I hope that'll make it a bit easier to cover more cases.
And then we'll need more injection points, I'm afraid :(.
> > - Several folks asked for pg_stat_aio to come back, in "v1" that showed the
> > set of currently in-flight AIOs. That's not particularly hard - except
> > that it doesn't really fit in the pg_stat_* namespace.
>
> Later message
> postgr.es/m/6vjl6jeaqvyhfbpgwziypwmhem2rwla4o5pgpuxwtg3o3o3jb5@evyzorb5meth is
> considering the name pg_aios. Works for me.
Cool.
> > --- a/src/backend/storage/aio/aio.c
> > +++ b/src/backend/storage/aio/aio.c
> > @@ -3,6 +3,28 @@
> > * aio.c
> > * AIO - Core Logic
> > *
> > + * For documentation about how AIO works on a higher level, including a
> > + * schematic example, see README.md.
> > + *
> > + *
> > + * AIO is a complicated subsystem. To keep things navigable it is split across
> > + * a number of files:
> > + *
> > + * - aio.c - core AIO state handling
> > + *
> > + * - aio_init.c - initialization
> > + *
> > + * - aio_io.c - dealing with actual IO, including executing IOs synchronously
> > + *
> > + * - aio_subject.c - functionality related to executing IO for different
> > + * subjects
> > + *
> > + * - method_*.c - different ways of executing AIO
> > + *
> > + * - read_stream.c - helper for accessing buffered relation data with
> > + * look-ahead
> > + *
>
> I felt like some list entries in this new header comment largely restated the
> file name. Here's how I'd write them to avoid that:
Thanks, adopting.
> * - method_*.c - different ways of executing AIO (e.g. worker process)
> * - aio_io.c - method-independent code for specific IO ops (e.g. readv)
> * - aio_subject.c - callbacks at IO operation lifecycle events
> * - aio_init.c - per-fork and per-startup-process initialization
I don't particularly like "per-startup-process", because "global
initialization" really is separate (and precedes) from startup processes
startup. Maybe "per-server and per-backend initialization"?
> * - aio.c - all other topics
> * - read_stream.c - helper for reading buffered relation data
Did the order you listed the files have a system to it? If so, what is it?
> > --- /dev/null
> > +++ b/src/backend/storage/aio/README.md
> > @@ -0,0 +1,413 @@
> > +# Asynchronous & Direct IO
>
> I would move "### Why Asynchronous IO" to here; that's good background before
> getting into the example.
I moved the example back and forth when writing because different readers
would benefit from a different order and I couldn't quite decide.
So I'm happy to adjust based on your feedback...
> I might also move "### Why Direct / unbuffered IO" to here. For me as a
> reader, I'd benefit from seeing things in this order:
>
> - "why"
> - condensed usage example like manpage SYNOPSIS, comments and decls removed
> - PgAioHandleState and discussion of valid transitions
Hm - why have PgAioHandleState and its states before the usage example? Seems
like it'd be harder to understand that way.
> - usage example as it is, with full comments
> - the rest
> ## Synopsis
>
> ioh = pgaio_io_get(CurrentResourceOwner, &ioret);
> pgaio_io_get_ref(ioh, &ior);
> pgaio_io_add_shared_cb(ioh, ASC_SHARED_BUFFER_READ);
> pgaio_io_set_io_data_32(ioh, (uint32 *) buffer, 1);
> smgrstartreadv(ioh, operation->smgr, forknum, blkno,
> BufferGetBlock(buffer), 1);
> pgaio_submit_staged();
> pgaio_io_ref_wait(&ior);
> if (ioret.result.status == ARS_ERROR)
> pgaio_result_log(aio_ret.result, &aio_ret.subject_data, ERROR);
Happy to add this, but I'm not entirely sure if that's really that useful to
have without commentary? The synopsis in manpages is helpful because it
provides the signature of various functions, but this wouldn't...
> > +
> > +## AIO Usage Example
> > +
> > +In many cases code that can benefit from AIO does not directly have to
> > +interact with the AIO interface, but can use AIO via higher-level
> > +abstractions. See [Helpers](#helpers).
> > +
> > +In this example, a buffer will be read into shared buffers.
> > +
> > +```C
> > +/*
> > + * Result of the operation, only to be accessed in this backend.
> > + */
> > +PgAioReturn ioret;
> > +
> > +/*
> > + * Acquire AIO Handle, ioret will get result upon completion.
>
> Consider adding: from here to pgaio_submit_staged(), don't do [description of
> the kind of unacceptable blocking operations].
Hm. Strictly speaking it's fine to block here, depending on whether
StartBufferIO() was already called. I'll clarify.
> > +### IO can be started in critical sections
> ...
> > +The need to be able to execute IO in critical sections has substantial design
> > +implication on the AIO subsystem. Mainly because completing IOs (see prior
> > +section) needs to be possible within a critical section, even if the
> > +to-be-completed IO itself was not issued in a critical section. Consider
> > +e.g. the case of a backend first starting a number of writes from shared
> > +buffers and then starting to flush the WAL. Because only a limited amount of
> > +IO can be in-progress at the same time, initiating the IO for flushing the WAL
> > +may require to first finish executing IO executed earlier.
>
> The last line's two appearances of the word "execute" read awkwardly to me,
> and it's an opportunity to use PgAioHandleState terms. Consider writing the
> last line like "may first advance an existing IO from AHS_PREPARED to
> AHS_COMPLETED_SHARED".
It is indeed awkward. I don't love referencing the state-constants here
though, somehow that feels like a reference-cycle ;). What about this:
> ... Consider
> e.g. the case of a backend first starting a number of writes from shared
> buffers and then starting to flush the WAL. Because only a limited amount of
> IO can be in-progress at the same time, initiating IO for flushing the WAL may
> require to first complete IO that was started earlier.
> > +### AIO Callbacks
> ...
> > +In addition to completion, AIO callbacks also are called to "prepare" an
> > +IO. This is, e.g., used to acquire buffer pins owned by the AIO subsystem for
> > +IO to/from shared buffers, which is required to handle the case where the
> > +issuing backend errors out and releases its own pins.
>
> Reading this, it's not obvious to me how to reconcile "finishing an IO could
> require pin acquisition" with "finishing an IO could happen in a critical
> section". Pinning a buffer in a critical section sounds bad. I vaguely
> recall understanding how it was okay as of my September review, but I've
> already forgotten. Can this text have a sentence making that explicit?
Ah, yes, that's easy to misunderstand. The answer basically is that we don't
newly pin a buffer, we just increment the reference count by 1. That should
never fail.
How about:
> In addition to completion, AIO callbacks also are called to "prepare" an
> IO. This is, e.g., used to increase buffer reference counts to account for the
> AIO subsystem referencing the buffer, which is required to handle the case
> where the issuing backend errors out and releases its own pins while the IO is
> still ongoing.
> > +### AIO Subjects
> > +
> > +In addition to the completion callbacks describe above, each AIO Handle has
> > +exactly one "subject". Each subject has some space inside an AIO Handle with
> > +information specific to the subject and can provide callbacks to allow to
> > +reopen the underlying file (required for worker mode) and to describe the IO
> > +operation (used for debug logging and error messages).
>
> Can this say roughly how to decide when to add a new subject?
Hm, there obviously is some fuzziness. I was trying to get to some of that by
mentioning that the subject needs to know how to [re-]open a file and describe
the target of the IO in terms that make sense to the user.
E.g. smgr seemed to make sense as a subject as the smgr layer knows how to
open a file by delegating that to the layer below and the layer above just
knows about smgr, not md.c (or other potential smgr implementations).
The reason to keep this separate from the callbacks was that smgr IO going
through shared buffers, bypassing shared buffers and different smgr
implemenentations all could share the same subject implementation, even if
callbacks would differ between these use cases.
How about:
> I.e., if two different uses of AIO can describe the identity of the file being
> operated on the same way, it likely makes sense to use the same
> subject. E.g. different smgr implementations can describe IO with
> RelFileLocator, ForkNumber and BlockNumber and can thus share a subject. In
> contrast, IO for a WAL file would be described with TimeLineID and XLogRecPtr
> and it would not make sense to use the same subject for smgr and WAL.
> Failing that, can it give examples of what additional subjects might exist
> if certain existing subsystems were to start using AIO?
I think the main ones I can think of are:
1) WAL logging
This was implemented in v1. I'd guess that "real" WAL logging and
initializing new WAL segments might use a different subject, but that's
probably a question of taste.
2) "raw" file IO, for things that don't use the smgr abstraction. I could
e.g. imagine using AIO in COPY to read / write the FROM/TO file or to
implement CREATE DATABASE ... STRATEGY file_copy with AIO.
This was used in v1, e.g. to implement the initial data directory sync
after a crash. We do that on a filesystem level, not going through smgr
etc.
3) FE/BE network IO
> > +### AIO Results
> > +
> > +As AIO completion callbacks
> > +[are executed in critical sections](#io-can-be-started-in-critical-sections)
> > +and [may be executed by any backend](#deadlock-and-starvation-dangers-due-to-aio)
> > +completion callbacks cannot be used to, e.g., make the query that triggered an
> > +IO ERROR out.
> > +
> > +To allow to react to failing IOs the issuing backend can pass a pointer to a
> > +`PgAioReturn` in backend local memory. Before an AIO Handle is reused the
> > +`PgAioReturn` is filled with information about the IO. This includes
> > +information about whether the IO was successful (as a value of
> > +`PgAioResultStatus`) and enough information to raise an error in case of a
> > +failure (via `pgaio_result_log()`, with the error details encoded in
> > +`PgAioResult`).
>
> Can this have a sentence on how this fits in bounded shmem, given the lack of
> guarantees about a backend's responsiveness? In other words, what makes it
> okay to have requests take arbitrarily long to move from AHS_COMPLETED_SHARED
> to AHS_COMPLETED_LOCAL?
I agree this should be explained somewhere - but not sure this is the best
place.
The reason it's ok is that each backend has a limited number of AIO handles
and if it runs out of IO handles we'll a) check if any IOs can be reclaimed b)
wait for the oldest IO to finish.
Thanks for the review!
Andres Freund
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-07 15:09 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Andres Freund <andres@anarazel.de>
2 siblings, 1 reply; 16+ messages in thread
From: Heikki Linnakangas @ 2025-01-07 15:09 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; pgsql-hackers
On 01/01/2025 06:03, Andres Freund wrote:
> Hi,
>
> Attached is a new version of the AIO patchset.
I haven't gone through it all yet, but some comments below.
> The biggest changes are:
>
> - The README has been extended with an overview of the API. I think it gives a
> good overview of how the API fits together. I'd be very good to get
> feedback from folks that aren't as familiar with AIO, I can't really see
> what's easy/hard anymore.
Thanks, the README is super helpful! I was overwhelmed by all the new
concepts before, now it all makes much more sense.
Now that it's all laid out more clearly, I see how many different
concepts and states there really are:
- For a single IO, there is an "IO handle", "IO references", and an "IO
return". You first allocate an IO handle (PgAioHandle), and then you get
a reference (PgAioHandleRef) and an "IO return" (PgAioReturn) struct for it.
- An IO handle has eight different states (PgAioHandleState).
I'm sure all those concepts exist for a reason. But still I wonder: can
we simplify?
pgaio_io_get() and pgaio_io_release() are a bit asymmetric, I'd suggest
pgaio_io_acquire() or similar. "get" also feels very innocent, even
though it may wait for previous IO to finish. Especially when
pgaio_io_get_ref() actually is innocent.
> typedef enum PgAioHandleState
> {
> /* not in use */
> AHS_IDLE = 0,
>
> /* returned by pgaio_io_get() */
> AHS_HANDED_OUT,
>
> /* pgaio_io_start_*() has been called, but IO hasn't been submitted yet */
> AHS_DEFINED,
>
> /* subject's prepare() callback has been called */
> AHS_PREPARED,
>
> /* IO has been submitted and is being executed */
> AHS_IN_FLIGHT,
>
> /* IO finished, but result has not yet been processed */
> AHS_REAPED,
>
> /* IO completed, shared completion has been called */
> AHS_COMPLETED_SHARED,
>
> /* IO completed, local completion has been called */
> AHS_COMPLETED_LOCAL,
> } PgAioHandleState;
Do we need to distinguish between DEFINED and PREPARED? At quick glance,
those states are treated the same. (The comment refers to
pgaio_io_start_*() functions, but there's no such thing)
I didn't quite understand the point of the prepare callbacks. For
example, when AsyncReadBuffers() calls smgrstartreadv(), the
shared_buffer_readv_prepare() callback will be called. Why doesn't
AsyncReadBuffers() do the "prepare" work itself directly; why does it
need to be in a callback? I assume it's somehow related to error
handling, but I didn't quite get it. Perhaps an "abort" callback that'd
be called on error, instead of a "prepare" callback, would be better?
There are some synonyms used in the code: I think "in-flight" and
"submitted" mean the same thing. And "prepared" and "staged". I'd
suggest picking just one term for each concept.
I didn't understand the COMPLETED_SHARED and COMPLETED_LOCAL states.
does a single IO go through both states, or are the mutually exclusive?
At quick glance, I don't actually see any code that would set the
COMPLETED_LOCAL state; is it dead code?
REAPED feels like a bad name. It sounds like a later stage than
COMPLETED, but it's actually vice versa.
I'm a little surprised that the term "IO request" isn't used anywhere. I
have no concrete suggestion, but perhaps that would be a useful term.
> - Retries for partial IOs (i.e. short reads) are now implemented. Turned out
> to take all of three lines and adding one missing variable initialization.
:-)
> - There's no obvious way to tell "internal" function operating on an IO handle
> apart from functions that are expected to be called by the issuer of an IO.
>
> One way to deal with this would be to introduce a distinct "issuer IO
> reference" type. I think that might be a good idea, it would also make it
> clearer that a good number of the functions can only be called by the
> issuer, before the IO is submitted.
>
> This would also make it easier to order functions more sensibly in aio.c, as
> all the issuer functions would be together.
>
> The functions on AIO handles that everyone can call already have a distinct
> type (PgAioHandleRef vs PgAioHandle*).
Hmm, yeah I think you might be onto something here.
Could pgaio_io_get() return an PgAioHandleRef directly, so that the
issuer would never see a raw PgAioHandle ?
Finally, attached are a couple of typos and other trivial suggestions.
--
Heikki Linnakangas
Neon (https://neon.tech)
Attachments:
[text/x-patch] aio-typos.patch (4.1K, ../../b7f6bd4a-d794-44d5-b6a7-314395939e2c@iki.fi/2-aio-typos.patch)
download | inline diff:
diff --git a/src/backend/storage/aio/README.md b/src/backend/storage/aio/README.md
index 0076ea4aa10..db3257c2705 100644
--- a/src/backend/storage/aio/README.md
+++ b/src/backend/storage/aio/README.md
@@ -15,7 +15,7 @@ In this example, a buffer will be read into shared buffers.
PgAioReturn ioret;
/*
- * Acquire AIO Handle, ioret will get result upon completion.
+ * Acquire an AIO Handle, ioret will get the result upon completion.
*/
PgAioHandle *ioh = pgaio_io_get(CurrentResourceOwner, &ioret);
@@ -46,15 +46,15 @@ pgaio_io_add_shared_cb(ioh, ASC_SHARED_BUFFER_READ);
pgaio_io_set_io_data_32(ioh, (uint32 *) buffer, 1);
/*
- * Hand AIO handle to lower-level function. When operating on the level of
+ * Pass the AIO handle to lower-level function. When operating on the level of
* buffers, we don't know how exactly the IO is performed, that is the
* responsibility of the storage manager implementation.
*
* E.g. md.c needs to translate block numbers into offsets in segments.
*
- * Once the IO handle has been handed of, it may not further be used, as the
- * IO may immediately get executed below smgrstartreadv() and the handle reused
- * for another IO.
+ * Once the IO handle has been handed off to smgrstartreadv(), it may not
+ * further be used, as the IO may immediately get executed in smgrstartreadv()
+ * and the handle reused for another IO.
*/
smgrstartreadv(ioh, operation->smgr, forknum, blkno,
BufferGetBlock(buffer), 1);
@@ -167,7 +167,7 @@ The main reason *not* to use Direct IO are:
explicit prefetching.
- In situations where shared_buffers cannot be set appropriately large,
e.g. because there are many different postgres instances hosted on shared
- hardware, performance will often be worse then when using buffered IO.
+ hardware, performance will often be worse than when using buffered IO.
### Deadlock and Starvation Dangers due to AIO
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index 261a752fb80..1cef6ef556b 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -123,10 +123,10 @@ static PgAioHandle *inj_cur_handle;
*
* If a handle was acquired but then does not turn out to be needed,
* e.g. because pgaio_io_get() is called before starting an IO in a critical
- * section, the handle needs to be be released with pgaio_io_release().
+ * section, the handle needs to be released with pgaio_io_release().
*
*
- * To react to the completion of the IO as soon as it is know to have
+ * To react to the completion of the IO as soon as it is known to have
* completed, callbacks can be registered with pgaio_io_add_shared_cb().
*
* To actually execute IO using the returned handle, the pgaio_io_prep_*()
diff --git a/src/backend/storage/aio/aio_io.c b/src/backend/storage/aio/aio_io.c
index 3c255775833..9e111c04b7e 100644
--- a/src/backend/storage/aio/aio_io.c
+++ b/src/backend/storage/aio/aio_io.c
@@ -31,7 +31,7 @@ static void pgaio_io_before_prep(PgAioHandle *ioh);
/* --------------------------------------------------------------------------------
* "Preparation" routines for individual IO types
*
- * These are called by place the place actually initiating an IO, to associate
+ * These are called by XXX place the place actually initiating an IO, to associate
* the IO specific data with an AIO handle.
*
* Each of the preparation routines first needs to call
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
index f4c57438dd4..7a81e211d48 100644
--- a/src/include/storage/aio_internal.h
+++ b/src/include/storage/aio_internal.h
@@ -38,12 +38,13 @@ typedef enum PgAioHandleState
AHS_HANDED_OUT,
/* pgaio_io_start_*() has been called, but IO hasn't been submitted yet */
+ /* XXX: there are no pgaio_io_start_*() functions */
AHS_DEFINED,
- /* subjects prepare() callback has been called */
+ /* subject's prepare() callback has been called */
AHS_PREPARED,
- /* IO is being executed */
+ /* IO has been submitted and is being executed */
AHS_IN_FLIGHT,
/* IO finished, but result has not yet been processed */
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-07 16:08 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Andres Freund <andres@anarazel.de>
2 siblings, 2 replies; 16+ messages in thread
From: Heikki Linnakangas @ 2025-01-07 16:08 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; pgsql-hackers
On LWLockDisown():
> +/*
> + * Stop treating lock as held by current backend.
> + *
> + * After calling this function it's the callers responsibility to ensure that
> + * the lock gets released, even in case of an error. This only is desirable if
> + * the lock is going to be released in a different process than the process
> + * that acquired it.
> + *
> + * Returns the mode in which the lock was held by the current backend.
Returning the lock mode feels a bit ad hoc..
> + * NB: This will leave lock->owner pointing to the current backend (if
> + * LOCK_DEBUG is set). We could add a separate flag indicating that, but it
> + * doesn't really seem worth it.
Hmm. I won't insist, but I feel it probably would be worth it. This is
only in LOCK_DEBUG mode so there's no performance penalty in non-debug
builds, and when you do compile with LOCK_DEBUG you probably appreciate
any extra information.
> + * NB: This does not call RESUME_INTERRUPTS(), but leaves that responsibility
> + * of the caller.
> + */
That feels weird. The only caller outside lwlock.c does call
RESUME_INTERRUPTS() immediately.
Perhaps it'd make for a better external interface if LWLockDisown() did
call RESUME_INTERRUPTS(), and there was a separate internal version that
didn't. And it might make more sense for the external version to return
'void' while we're at it. Returning a value that the caller ignores is
harmless, of course, but it feels a bit weird. It makes you wonder what
you're supposed to do with it.
> + {
> + {"io_method", PGC_POSTMASTER, RESOURCES_MEM,
> + gettext_noop("Selects the method of asynchronous I/O to use."),
> + NULL
> + },
> + &io_method,
> + DEFAULT_IO_METHOD, io_method_options,
> + NULL, assign_io_method, NULL
> + },
> +
The description is a bit funny because synchronous I/O is one of the
possible methods.
--
Heikki Linnakangas
Neon (https://neon.tech)
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-07 16:11 Andres Freund <andres@anarazel.de>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
0 siblings, 2 replies; 16+ messages in thread
From: Andres Freund @ 2025-01-07 16:11 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: pgsql-hackers
Hi,
On 2025-01-07 17:09:58 +0200, Heikki Linnakangas wrote:
> On 01/01/2025 06:03, Andres Freund wrote:
> > Hi,
> >
> > Attached is a new version of the AIO patchset.
>
> I haven't gone through it all yet, but some comments below.
Thanks!
> > The biggest changes are:
> >
> > - The README has been extended with an overview of the API. I think it gives a
> > good overview of how the API fits together. I'd be very good to get
> > feedback from folks that aren't as familiar with AIO, I can't really see
> > what's easy/hard anymore.
>
> Thanks, the README is super helpful! I was overwhelmed by all the new
> concepts before, now it all makes much more sense.
>
> Now that it's all laid out more clearly, I see how many different concepts
> and states there really are:
>
> - For a single IO, there is an "IO handle", "IO references", and an "IO
> return". You first allocate an IO handle (PgAioHandle), and then you get a
> reference (PgAioHandleRef) and an "IO return" (PgAioReturn) struct for it.
>
> - An IO handle has eight different states (PgAioHandleState).
>
> I'm sure all those concepts exist for a reason. But still I wonder: can we
> simplify?
Probably, but it's not exactly obvious to me where.
The difference between a handle and a reference is useful right now, to have
some separation between the functions that can be called by anyone (taking a
PgAioHandleRef) and only by the issuer (PgAioHandle). That might better be
solved by having a PgAioHandleIssuerRef ref or something.
Having PgAioReturn be separate from the AIO handle turns out to be rather
crucial, otherwise it's very hard to guarantee "forward progress",
i.e. guarantee that pgaio_io_get() will return something without blocking
forever.
> pgaio_io_get() and pgaio_io_release() are a bit asymmetric, I'd suggest
> pgaio_io_acquire() or similar. "get" also feels very innocent, even though
> it may wait for previous IO to finish. Especially when pgaio_io_get_ref()
> actually is innocent.
WFM.
> > typedef enum PgAioHandleState
> > {
> > /* not in use */
> > AHS_IDLE = 0,
> >
> > /* returned by pgaio_io_get() */
> > AHS_HANDED_OUT,
> >
> > /* pgaio_io_start_*() has been called, but IO hasn't been submitted yet */
> > AHS_DEFINED,
> >
> > /* subject's prepare() callback has been called */
> > AHS_PREPARED,
> >
> > /* IO has been submitted and is being executed */
> > AHS_IN_FLIGHT,
> >
> > /* IO finished, but result has not yet been processed */
> > AHS_REAPED,
> >
> > /* IO completed, shared completion has been called */
> > AHS_COMPLETED_SHARED,
> >
> > /* IO completed, local completion has been called */
> > AHS_COMPLETED_LOCAL,
> > } PgAioHandleState;
>
> Do we need to distinguish between DEFINED and PREPARED?
I found it to be rather confusing if it's not possible to tell if some action
(like the prepare callback) has already happened, or not. It's useful to be
able look at an IO in a backtrace or such and see exactly in what state it is
in.
In v1 I had several of the above states managed as separate boolean variables
- that turned out to be a huge mess, it's a lot easier to understand if
there's a single strictly monotonically increasing state.
> At quick glance, those states are treated the same. (The comment refers to
> pgaio_io_start_*() functions, but there's no such thing)
They're called pgaio_io_prep_{readv,writev} now, updated the comment.
> I didn't quite understand the point of the prepare callbacks. For example,
> when AsyncReadBuffers() calls smgrstartreadv(), the
> shared_buffer_readv_prepare() callback will be called. Why doesn't
> AsyncReadBuffers() do the "prepare" work itself directly; why does it need
> to be in a callback?
One big part of it is "ownership" - while the IO isn't completely "assembled",
we can release all buffer pins etc in case of an error. But if the error
happens just after the IO was staged, we can't - the buffer is still
referenced by the IO. For that the AIO subystem needs to take its own pins
etc. Initially the prepare callback didn't exist, the code in
AsyncReadBuffers() was a lot more complicated before it.
> I assume it's somehow related to error handling, but I didn't quite get
> it. Perhaps an "abort" callback that'd be called on error, instead of a
> "prepare" callback, would be better?
I don't think an error callback would be helpful - the whole thing is that we
basically need claim ownership of all IO related resources IFF the IO is
staged. Not before (because then the IO not getting staged would mean we have
a resource leak), not after (because we might error out and thus not keep
e.g. buffers pinned).
> There are some synonyms used in the code: I think "in-flight" and
> "submitted" mean the same thing.
Fair. I guess in my mind the process of moving an IO into flight is
"submitting" and the state of not having been submitted but not yet having
completed is being in flight. But that's probably not useful.
> And "prepared" and "staged". I'd suggest picking just one term for each
> concept.
Agreed.
> I didn't understand the COMPLETED_SHARED and COMPLETED_LOCAL states. does a
> single IO go through both states, or are the mutually exclusive? At quick
> glance, I don't actually see any code that would set the COMPLETED_LOCAL
> state; is it dead code?
It's dead code right now. I've made it dead and undead a couple times
:/. Unfortunately I think I need to revive it to make some corner cases with
temporary tables work (AIO for temp table is executed via IO uring, another
backend waits for *another* IO executed via that IO uring instance and reaps
the completion -> we can't update the local buffer state in the shared
completion callback).
> REAPED feels like a bad name. It sounds like a later stage than COMPLETED,
> but it's actually vice versa.
What would you call having gotten "completion notifications" from the kernel,
but not having processed them?
> > - There's no obvious way to tell "internal" function operating on an IO handle
> > apart from functions that are expected to be called by the issuer of an IO.
> >
> > One way to deal with this would be to introduce a distinct "issuer IO
> > reference" type. I think that might be a good idea, it would also make it
> > clearer that a good number of the functions can only be called by the
> > issuer, before the IO is submitted.
> >
> > This would also make it easier to order functions more sensibly in aio.c, as
> > all the issuer functions would be together.
> >
> > The functions on AIO handles that everyone can call already have a distinct
> > type (PgAioHandleRef vs PgAioHandle*).
>
> Hmm, yeah I think you might be onto something here.
I'll give it a try.
> Could pgaio_io_get() return an PgAioHandleRef directly, so that the issuer
> would never see a raw PgAioHandle ?
Don't think that would be helpful - that way there'd be no difference at all
anymore between what functions any backend can call and what the issuer can
do.
>
> Finally, attached are a couple of typos and other trivial suggestions.
Integrating...
Thanks!
Andres
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-07 16:32 Andres Freund <andres@anarazel.de>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
1 sibling, 0 replies; 16+ messages in thread
From: Andres Freund @ 2025-01-07 16:32 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: pgsql-hackers
Hi,
On 2025-01-07 18:08:51 +0200, Heikki Linnakangas wrote:
> On LWLockDisown():
>
> > +/*
> > + * Stop treating lock as held by current backend.
> > + *
> > + * After calling this function it's the callers responsibility to ensure that
> > + * the lock gets released, even in case of an error. This only is desirable if
> > + * the lock is going to be released in a different process than the process
> > + * that acquired it.
> > + *
> > + * Returns the mode in which the lock was held by the current backend.
>
> Returning the lock mode feels a bit ad hoc..
It seemed useful to me, that way callers could verify that the released lock
level is actually what it expected. What do we gain by hiding this information
anyway?
Orthogonal: I think it was a mistake that LWLockRelease() didn't require the
to-be-releaased lock mode to be passed in...
> > + * NB: This will leave lock->owner pointing to the current backend (if
> > + * LOCK_DEBUG is set). We could add a separate flag indicating that, but it
> > + * doesn't really seem worth it.
>
> Hmm. I won't insist, but I feel it probably would be worth it. This is only
> in LOCK_DEBUG mode so there's no performance penalty in non-debug builds,
> and when you do compile with LOCK_DEBUG you probably appreciate any extra
> information.
I actually thought it'd be more useful if it stays pointing to the 'original
owner'.
When you say "it" would be worth it, you mean resetting owner, or adding a
flag indicating that it's a disowned lock?
> > + * NB: This does not call RESUME_INTERRUPTS(), but leaves that responsibility
> > + * of the caller.
> > + */
>
> That feels weird. The only caller outside lwlock.c does call
> RESUME_INTERRUPTS() immediately.
Yea, I didn't feel happy with it either. It just seemed that the cure (a
separate function, or a parameter indicating whether interrupts should be
resumed) was as bad as the disease.
> Perhaps it'd make for a better external interface if LWLockDisown() did call
> RESUME_INTERRUPTS(), and there was a separate internal version that didn't.
Hm, that seems more complicated than it's worth. I'd either leave it as-is,
or add a parameter to LWLockDisown to indicate if interrupts should be
resumed.
> And it might make more sense for the external version to return 'void' while
> we're at it. Returning a value that the caller ignores is harmless, of
> course, but it feels a bit weird. It makes you wonder what you're supposed
> to do with it.
This one I disagree with, I think it makes a lot of sense to return the lock
mode of the lock you just disowned.
Doubtful it matters, but the compiler can trivially optimize that out for the
lwlock.c callers.
> > + {
> > + {"io_method", PGC_POSTMASTER, RESOURCES_MEM,
> > + gettext_noop("Selects the method of asynchronous I/O to use."),
> > + NULL
> > + },
> > + &io_method,
> > + DEFAULT_IO_METHOD, io_method_options,
> > + NULL, assign_io_method, NULL
> > + },
> > +
>
> The description is a bit funny because synchronous I/O is one of the
> possible methods.
Hah. How about:
"Selects the method of, potentially asynchronous, IO execution."?
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-07 19:10 Noah Misch <noah@leadboat.com>
parent: Andres Freund <andres@anarazel.de>
0 siblings, 0 replies; 16+ messages in thread
From: Noah Misch @ 2025-01-07 19:10 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers
On Mon, Jan 06, 2025 at 04:40:26PM -0500, Andres Freund wrote:
> On 2025-01-06 10:52:20 -0800, Noah Misch wrote:
> > On Tue, Dec 31, 2024 at 11:03:33PM -0500, Andres Freund wrote:
> - We have pretty no testing for IO errors.
Yes, that's remained a gap. I've wondered how much to address this via
targeted tests of specific sites vs. fuzzing, iterative fault injection, or
some other approach closer to brute force.
> > I'd be most interested in the
> > cases that would be undetected deadlocks under a naive design. An example
> > appeared at the end of postgr.es/m/20240916144349.74.nmisch@google.com
>
> That's a good one, yea.
>
> I think I'll try to translate the regression tests I wrote into an isolation
> test, I hope that'll make it a bit easier to cover more cases.
>
> And then we'll need more injection points, I'm afraid :(.
Sounds good.
> > * - method_*.c - different ways of executing AIO (e.g. worker process)
> > * - aio_io.c - method-independent code for specific IO ops (e.g. readv)
> > * - aio_subject.c - callbacks at IO operation lifecycle events
> > * - aio_init.c - per-fork and per-startup-process initialization
>
> I don't particularly like "per-startup-process", because "global
> initialization" really is separate (and precedes) from startup processes
> startup. Maybe "per-server and per-backend initialization"?
That works for me. I wrote "per-startup-process" because it can happen more
than once in a postmaster that reaches "all server processes terminated;
reinitializing". That said, there's little risk of "per-server" giving folks
a materially wrong idea.
> > * - aio.c - all other topics
> > * - read_stream.c - helper for reading buffered relation data
>
> Did the order you listed the files have a system to it? If so, what is it?
The rough idea was to avoid forward references:
* - method_*.c - different ways of executing AIO (e.g. worker process)
makes sense without other background
* - aio_io.c - method-independent code for specific IO ops (e.g. readv)
refers to methods, so listed after methods
* - aio_subject.c - callbacks at IO operation lifecycle events
refers to IO ops, so listed after aio_io.c
* - aio_init.c - per-fork and per-startup-process initialization
no surprise that this code will exist somewhere, so list it lower to deemphasize it
* - aio.c - all other topics
default route, hence last
* - read_stream.c - helper for reading buffered relation data
could just as easily come first, not last
could be under a distinct heading like "Recommended abstractions:"
> > I'd benefit from seeing things in this order:
> >
> > - "why"
> > - condensed usage example like manpage SYNOPSIS, comments and decls removed
> > - PgAioHandleState and discussion of valid transitions
>
> Hm - why have PgAioHandleState and its states before the usage example? Seems
> like it'd be harder to understand that way.
I usually look at the data structures before the code that manipulates them.
(Similarly, I look at the map before the directions.) I wouldn't mind it
appearing after the usage example, since order preferences do vary.
> > - usage example as it is, with full comments
> > - the rest
>
>
> > ## Synopsis
> >
> > ioh = pgaio_io_get(CurrentResourceOwner, &ioret);
> > pgaio_io_get_ref(ioh, &ior);
> > pgaio_io_add_shared_cb(ioh, ASC_SHARED_BUFFER_READ);
> > pgaio_io_set_io_data_32(ioh, (uint32 *) buffer, 1);
> > smgrstartreadv(ioh, operation->smgr, forknum, blkno,
> > BufferGetBlock(buffer), 1);
> > pgaio_submit_staged();
> > pgaio_io_ref_wait(&ior);
> > if (ioret.result.status == ARS_ERROR)
> > pgaio_result_log(aio_ret.result, &aio_ret.subject_data, ERROR);
>
> Happy to add this, but I'm not entirely sure if that's really that useful to
> have without commentary? The synopsis in manpages is helpful because it
> provides the signature of various functions, but this wouldn't...
I'm not sure either. Let's drop that idea.
> > > +### IO can be started in critical sections
> > ...
> > > +The need to be able to execute IO in critical sections has substantial design
> > > +implication on the AIO subsystem. Mainly because completing IOs (see prior
> > > +section) needs to be possible within a critical section, even if the
> > > +to-be-completed IO itself was not issued in a critical section. Consider
> > > +e.g. the case of a backend first starting a number of writes from shared
> > > +buffers and then starting to flush the WAL. Because only a limited amount of
> > > +IO can be in-progress at the same time, initiating the IO for flushing the WAL
> > > +may require to first finish executing IO executed earlier.
> >
> > The last line's two appearances of the word "execute" read awkwardly to me,
> > and it's an opportunity to use PgAioHandleState terms. Consider writing the
> > last line like "may first advance an existing IO from AHS_PREPARED to
> > AHS_COMPLETED_SHARED".
>
> It is indeed awkward. I don't love referencing the state-constants here
> though, somehow that feels like a reference-cycle ;). What about this:
>
> > ... Consider
> > e.g. the case of a backend first starting a number of writes from shared
> > buffers and then starting to flush the WAL. Because only a limited amount of
> > IO can be in-progress at the same time, initiating IO for flushing the WAL may
> > require to first complete IO that was started earlier.
That's non-awkward. I like specific state names here since "complete" could
mean AHS_COMPLETED_SHARED or AHS_COMPLETED_LOCAL, and it matters here. If the
state names changed so AHS_COMPLETED_LOCAL dropped the word "complete", that
too would solve it.
> > > +### AIO Callbacks
> > ...
> > > +In addition to completion, AIO callbacks also are called to "prepare" an
> > > +IO. This is, e.g., used to acquire buffer pins owned by the AIO subsystem for
> > > +IO to/from shared buffers, which is required to handle the case where the
> > > +issuing backend errors out and releases its own pins.
> >
> > Reading this, it's not obvious to me how to reconcile "finishing an IO could
> > require pin acquisition" with "finishing an IO could happen in a critical
> > section". Pinning a buffer in a critical section sounds bad. I vaguely
> > recall understanding how it was okay as of my September review, but I've
> > already forgotten. Can this text have a sentence making that explicit?
>
> Ah, yes, that's easy to misunderstand. The answer basically is that we don't
> newly pin a buffer, we just increment the reference count by 1. That should
> never fail.
>
> How about:
> > In addition to completion, AIO callbacks also are called to "prepare" an
> > IO. This is, e.g., used to increase buffer reference counts to account for the
> > AIO subsystem referencing the buffer, which is required to handle the case
> > where the issuing backend errors out and releases its own pins while the IO is
> > still ongoing.
Perfect.
> > > +### AIO Subjects
> > > +
> > > +In addition to the completion callbacks describe above, each AIO Handle has
> > > +exactly one "subject". Each subject has some space inside an AIO Handle with
> > > +information specific to the subject and can provide callbacks to allow to
> > > +reopen the underlying file (required for worker mode) and to describe the IO
> > > +operation (used for debug logging and error messages).
> >
> > Can this say roughly how to decide when to add a new subject?
>
> Hm, there obviously is some fuzziness. I was trying to get to some of that by
> mentioning that the subject needs to know how to [re-]open a file and describe
> the target of the IO in terms that make sense to the user.
>
> E.g. smgr seemed to make sense as a subject as the smgr layer knows how to
> open a file by delegating that to the layer below and the layer above just
> knows about smgr, not md.c (or other potential smgr implementations).
>
> The reason to keep this separate from the callbacks was that smgr IO going
> through shared buffers, bypassing shared buffers and different smgr
> implemenentations all could share the same subject implementation, even if
> callbacks would differ between these use cases.
>
>
> How about:
>
> > I.e., if two different uses of AIO can describe the identity of the file being
> > operated on the same way, it likely makes sense to use the same
> > subject. E.g. different smgr implementations can describe IO with
> > RelFileLocator, ForkNumber and BlockNumber and can thus share a subject. In
> > contrast, IO for a WAL file would be described with TimeLineID and XLogRecPtr
> > and it would not make sense to use the same subject for smgr and WAL.
Sounds good to include.
> > Can this have a sentence on how this fits in bounded shmem, given the lack of
> > guarantees about a backend's responsiveness? In other words, what makes it
> > okay to have requests take arbitrarily long to move from AHS_COMPLETED_SHARED
> > to AHS_COMPLETED_LOCAL?
>
> I agree this should be explained somewhere - but not sure this is the best
> place.
>
> The reason it's ok is that each backend has a limited number of AIO handles
> and if it runs out of IO handles we'll a) check if any IOs can be reclaimed b)
> wait for the oldest IO to finish.
Reading it again today, that topic may already have adequate coverage.
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-07 19:59 Robert Haas <robertmhaas@gmail.com>
parent: Andres Freund <andres@anarazel.de>
1 sibling, 1 reply; 16+ messages in thread
From: Robert Haas @ 2025-01-07 19:59 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; pgsql-hackers
On Tue, Jan 7, 2025 at 11:11 AM Andres Freund <andres@anarazel.de> wrote:
> The difference between a handle and a reference is useful right now, to have
> some separation between the functions that can be called by anyone (taking a
> PgAioHandleRef) and only by the issuer (PgAioHandle). That might better be
> solved by having a PgAioHandleIssuerRef ref or something.
To me, those names don't convey that. I would perhaps call the thing
that supports issuer-only operations a "PgAio" and the thing other
people can use a "PgAioHandle". Or "PgAioRequest" and "PgAioHandle" or
something like that. With PgAioHandleRef, IMHO you've got two words
that both imply a layer of indirection -- "handle" and "ref" -- which
doesn't seem quite as nice, because then the other thing --
"PgAioHandle" still sort of implies
one layer of indirection and the whole thing seems a bit less clear.
(I say all of this having looked at nothing, so feel free to ignore me
if that doesn't sound coherent.)
> > REAPED feels like a bad name. It sounds like a later stage than COMPLETED,
> > but it's actually vice versa.
>
> What would you call having gotten "completion notifications" from the kernel,
> but not having processed them?
The Linux kernel calls those zombie processes, so we could call it a
ZOMBIE state, but that seems like it might be a bit of inside
baseball. I do agree with Heikki that REAPED sounds later than
COMPLETED, because you reap zombie processes by collecting their exit
status. Maybe you could have AHS_COMPLETE or AHS_IO_COMPLETE for the
state where the I/O is done but there's still completion-related work
to be done, and then the other state could be AHS_DONE or AHS_FINISHED
or AHS_FINAL or AHS_REAPED or something.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-07 20:09 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Andres Freund <andres@anarazel.de>
1 sibling, 1 reply; 16+ messages in thread
From: Heikki Linnakangas @ 2025-01-07 20:09 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers
On 07/01/2025 18:11, Andres Freund wrote:
> The difference between a handle and a reference is useful right now, to have
> some separation between the functions that can be called by anyone (taking a
> PgAioHandleRef) and only by the issuer (PgAioHandle). That might better be
> solved by having a PgAioHandleIssuerRef ref or something.
>
> Having PgAioReturn be separate from the AIO handle turns out to be rather
> crucial, otherwise it's very hard to guarantee "forward progress",
> i.e. guarantee that pgaio_io_get() will return something without blocking
> forever.
Right, yeah, I can see that.
>>> typedef enum PgAioHandleState
>>> {
>>> /* not in use */
>>> AHS_IDLE = 0,
>>>
>>> /* returned by pgaio_io_get() */
>>> AHS_HANDED_OUT,
>>>
>>> /* pgaio_io_start_*() has been called, but IO hasn't been submitted yet */
>>> AHS_DEFINED,
>>>
>>> /* subject's prepare() callback has been called */
>>> AHS_PREPARED,
>>>
>>> /* IO has been submitted and is being executed */
>>> AHS_IN_FLIGHT,
>>>
>>> /* IO finished, but result has not yet been processed */
>>> AHS_REAPED,
>>>
>>> /* IO completed, shared completion has been called */
>>> AHS_COMPLETED_SHARED,
>>>
>>> /* IO completed, local completion has been called */
>>> AHS_COMPLETED_LOCAL,
>>> } PgAioHandleState;
>>
>> Do we need to distinguish between DEFINED and PREPARED?
>
> I found it to be rather confusing if it's not possible to tell if some action
> (like the prepare callback) has already happened, or not. It's useful to be
> able look at an IO in a backtrace or such and see exactly in what state it is
> in.
I see.
> In v1 I had several of the above states managed as separate boolean variables
> - that turned out to be a huge mess, it's a lot easier to understand if
> there's a single strictly monotonically increasing state.
Agreed on that
>> I didn't quite understand the point of the prepare callbacks. For example,
>> when AsyncReadBuffers() calls smgrstartreadv(), the
>> shared_buffer_readv_prepare() callback will be called. Why doesn't
>> AsyncReadBuffers() do the "prepare" work itself directly; why does it need
>> to be in a callback?
>
> One big part of it is "ownership" - while the IO isn't completely "assembled",
> we can release all buffer pins etc in case of an error. But if the error
> happens just after the IO was staged, we can't - the buffer is still
> referenced by the IO. For that the AIO subystem needs to take its own pins
> etc. Initially the prepare callback didn't exist, the code in
> AsyncReadBuffers() was a lot more complicated before it.
>
>
>> I assume it's somehow related to error handling, but I didn't quite get
>> it. Perhaps an "abort" callback that'd be called on error, instead of a
>> "prepare" callback, would be better?
>
> I don't think an error callback would be helpful - the whole thing is that we
> basically need claim ownership of all IO related resources IFF the IO is
> staged. Not before (because then the IO not getting staged would mean we have
> a resource leak), not after (because we might error out and thus not keep
> e.g. buffers pinned).
Hmm. The comments say that when you call smgrstartreadv(), the IO handle
may no longer be modified, as the IO may be executed immediately. What
if we changed that so that it never submits the IO, only adds the
necessary callbacks to it?
In that world, when smgrstartreadv() returns, the necessary details and
completion callbacks have been set in the IO handle, but the caller can
still do more preparation before the IO is submitted. The caller must
ensure that it gets submitted, however, so no erroring out in that state.
Currently the call stack looks like this:
AsyncReadBuffers()
-> smgrstartreadv()
-> mdstartreadv()
-> FileStartReadV()
-> pgaio_io_prep_readv()
-> shared_buffer_readv_prepare() (callback)
<- (return)
<- (return)
<- (return)
<- (return)
<- (return)
I'm thinking that the prepare work is done "on the way up" instead:
AsyncReadBuffers()
-> smgrstartreadv()
-> mdstartreadv()
-> FileStartReadV()
-> pgaio_io_prep_readv()
<- (return)
<- (return)
<- (return)
-> shared_buffer_readv_prepare()
<- (return)
Attached is a patch to demonstrate concretely what I mean.
This adds a new pgaio_io_stage() step to the issuer, and the issuer
needs to call the prepare functions explicitly, instead of having them
as callbacks. Nominally that's more steps, but IMHO it's better to be
explicit. The same actions were happening previously too, it was just
hidden in the callback. I updated the README to show that too.
I'm not wedded to this, but it feels a little better to me.
--
Heikki Linnakangas
Neon (https://neon.tech)
Attachments:
[text/x-patch] aio-remove-prepare-callback.patch (17.7K, ../../2cd4f43e-71fa-4d23-b316-06a0d451f8ef@iki.fi/2-aio-remove-prepare-callback.patch)
download | inline diff:
diff --git a/src/backend/storage/aio/README.md b/src/backend/storage/aio/README.md
index 0076ea4aa10..25b5f5d9529 100644
--- a/src/backend/storage/aio/README.md
+++ b/src/backend/storage/aio/README.md
@@ -60,7 +60,18 @@ smgrstartreadv(ioh, operation->smgr, forknum, blkno,
BufferGetBlock(buffer), 1);
/*
- * As mentioned above, the IO might be initiated within smgrstartreadv(). That
+ * After smgrstartreadv() has returned, we are committed to performing the IO.
+ * We may do more preparation or add more callbacks to the IO, but must
+ * *not* error out before calling pgaio_io_stage(). We don't have any such
+ * preparation to do here, so just call pgaio_io_stage() to indicate that we
+ * have completed building the IO request. It usually queues up the request
+ * for batching, but may submit it immediately if the batch is full or if
+ * the request needed to be processed synchronously.
+ */
+pgaio_io_stage(ioh);
+
+/*
+ * The IO might already have been initiated by pgaio_io_stage(). That
* is however not guaranteed, to allow IO submission to be batched.
*
* Note that one needs to be careful while there may be unsubmitted IOs, as
@@ -69,10 +80,6 @@ smgrstartreadv(ioh, operation->smgr, forknum, blkno,
* that, pending IOs need to be explicitly submitted before this backend
* might be blocked by a backend waiting for IO.
*
- * Note that the IO might have immediately been submitted (e.g. due to reaching
- * a limit on the number of unsubmitted IOs) and even completed during the
- * smgrstartreadv() above.
- *
* Once submitted, the IO is in-flight and can complete at any time.
*/
pgaio_submit_staged();
diff --git a/src/backend/storage/aio/aio.c b/src/backend/storage/aio/aio.c
index 261a752fb80..ed03fe03609 100644
--- a/src/backend/storage/aio/aio.c
+++ b/src/backend/storage/aio/aio.c
@@ -110,7 +110,7 @@ static PgAioHandle *inj_cur_handle;
* Acquire an AioHandle, waiting for IO completion if necessary.
*
* Each backend can only have one AIO handle that that has been "handed out"
- * to code, but not yet submitted or released. This restriction is necessary
+ * to code, but not yet staged or released. This restriction is necessary
* to ensure that it is possible for code to wait for an unused handle by
* waiting for in-flight IO to complete. There is a limited number of handles
* in each backend, if multiple handles could be handed out without being
@@ -249,6 +249,43 @@ pgaio_io_release(PgAioHandle *ioh)
}
}
+/*
+ * Finish building an IO request. Once a request has been staged, there's no
+ * going back; the IO subsystem will attempt to perform the IO. If the IO
+ * succeeds the completion callbacks will be called; on error, the error
+ * callbacks.
+ *
+ * This may add the IO to the current batch, or execute the request
+ * synchronously.
+ */
+void
+pgaio_io_stage(PgAioHandle *ioh)
+{
+ bool needs_synchronous;
+
+ /* allow a new IO to be staged */
+ my_aio->handed_out_io = NULL;
+
+ pgaio_io_update_state(ioh, AHS_PREPARED);
+
+ needs_synchronous = pgaio_io_needs_synchronous_execution(ioh);
+
+ elog(DEBUG3, "io:%d: staged %s, executed synchronously: %d",
+ pgaio_io_get_id(ioh), pgaio_io_get_op_name(ioh),
+ needs_synchronous);
+
+ if (!needs_synchronous)
+ {
+ my_aio->staged_ios[my_aio->num_staged_ios++] = ioh;
+ Assert(my_aio->num_staged_ios <= PGAIO_SUBMIT_BATCH_SIZE);
+ }
+ else
+ {
+ pgaio_io_prepare_submit(ioh);
+ pgaio_io_perform_synchronously(ioh);
+ }
+}
+
/*
* Release IO handle during resource owner cleanup.
*/
@@ -279,7 +316,7 @@ pgaio_io_release_resowner(dlist_node *ioh_node, bool on_error)
pgaio_io_reclaim(ioh);
break;
- case AHS_DEFINED:
+ case AHS_PREPARING:
case AHS_PREPARED:
/* XXX: Should we warn about this when is_commit? */
pgaio_submit_staged();
@@ -383,7 +420,7 @@ void
pgaio_io_get_ref(PgAioHandle *ioh, PgAioHandleRef *ior)
{
Assert(ioh->state == AHS_HANDED_OUT ||
- ioh->state == AHS_DEFINED ||
+ ioh->state == AHS_PREPARING ||
ioh->state == AHS_PREPARED);
Assert(ioh->generation != 0);
@@ -437,7 +474,7 @@ pgaio_io_ref_wait(PgAioHandleRef *ior)
if (am_owner)
{
- if (state == AHS_DEFINED || state == AHS_PREPARED)
+ if (state == AHS_PREPARING || state == AHS_PREPARED)
{
/* XXX: Arguably this should be prevented by callers? */
pgaio_submit_staged();
@@ -489,8 +526,8 @@ pgaio_io_ref_wait(PgAioHandleRef *ior)
/* fallthrough */
/* waiting for owner to submit */
+ case AHS_PREPARING:
case AHS_PREPARED:
- case AHS_DEFINED:
/* waiting for reaper to complete */
/* fallthrough */
case AHS_REAPED:
@@ -501,8 +538,7 @@ pgaio_io_ref_wait(PgAioHandleRef *ior)
while (!pgaio_io_was_recycled(ioh, ref_generation, &state))
{
- if (state != AHS_REAPED && state != AHS_DEFINED &&
- state != AHS_IN_FLIGHT)
+ if (state != AHS_REAPED && state != AHS_IN_FLIGHT)
break;
ConditionVariableSleep(&ioh->cv, WAIT_EVENT_AIO_COMPLETION);
}
@@ -570,8 +606,8 @@ pgaio_io_get_state_name(PgAioHandle *ioh)
return "idle";
case AHS_HANDED_OUT:
return "handed_out";
- case AHS_DEFINED:
- return "DEFINED";
+ case AHS_PREPARING:
+ return "PREPARING";
case AHS_PREPARED:
return "PREPARED";
case AHS_IN_FLIGHT:
@@ -588,43 +624,18 @@ pgaio_io_get_state_name(PgAioHandle *ioh)
/*
* Internal, should only be called from pgaio_io_prep_*().
+ *
+ * Switches the IO to PREPARING state.
*/
void
-pgaio_io_prepare(PgAioHandle *ioh, PgAioOp op)
+pgaio_io_start_staging(PgAioHandle *ioh)
{
- bool needs_synchronous;
-
Assert(ioh->state == AHS_HANDED_OUT);
Assert(pgaio_io_has_subject(ioh));
- ioh->op = op;
ioh->result = 0;
- pgaio_io_update_state(ioh, AHS_DEFINED);
-
- /* allow a new IO to be staged */
- my_aio->handed_out_io = NULL;
-
- pgaio_io_prepare_subject(ioh);
-
- pgaio_io_update_state(ioh, AHS_PREPARED);
-
- needs_synchronous = pgaio_io_needs_synchronous_execution(ioh);
-
- elog(DEBUG3, "io:%d: prepared %s, executed synchronously: %d",
- pgaio_io_get_id(ioh), pgaio_io_get_op_name(ioh),
- needs_synchronous);
-
- if (!needs_synchronous)
- {
- my_aio->staged_ios[my_aio->num_staged_ios++] = ioh;
- Assert(my_aio->num_staged_ios <= PGAIO_SUBMIT_BATCH_SIZE);
- }
- else
- {
- pgaio_io_prepare_submit(ioh);
- pgaio_io_perform_synchronously(ioh);
- }
+ pgaio_io_update_state(ioh, AHS_PREPARING);
}
/*
@@ -858,8 +869,8 @@ pgaio_io_wait_for_free(void)
{
/* should not be in in-flight list */
case AHS_IDLE:
- case AHS_DEFINED:
case AHS_HANDED_OUT:
+ case AHS_PREPARING:
case AHS_PREPARED:
case AHS_COMPLETED_LOCAL:
elog(ERROR, "shouldn't get here with io:%d in state %d",
@@ -1004,7 +1015,7 @@ pgaio_bounce_buffer_wait_for_free(void)
case AHS_IDLE:
case AHS_HANDED_OUT:
continue;
- case AHS_DEFINED: /* should have been submitted above */
+ case AHS_PREPARING: /* should have been submitted above */
case AHS_PREPARED:
elog(ERROR, "shouldn't get here with io:%d in state %d",
pgaio_io_get_id(ioh), ioh->state);
diff --git a/src/backend/storage/aio/aio_io.c b/src/backend/storage/aio/aio_io.c
index 3c255775833..e84b79d3f2e 100644
--- a/src/backend/storage/aio/aio_io.c
+++ b/src/backend/storage/aio/aio_io.c
@@ -46,11 +46,12 @@ pgaio_io_prep_readv(PgAioHandle *ioh,
{
pgaio_io_before_prep(ioh);
+ ioh->op = PGAIO_OP_READV;
ioh->op_data.read.fd = fd;
ioh->op_data.read.offset = offset;
ioh->op_data.read.iov_length = iovcnt;
- pgaio_io_prepare(ioh, PGAIO_OP_READV);
+ pgaio_io_start_staging(ioh);
}
void
@@ -59,11 +60,12 @@ pgaio_io_prep_writev(PgAioHandle *ioh,
{
pgaio_io_before_prep(ioh);
+ ioh->op = PGAIO_OP_WRITEV;
ioh->op_data.write.fd = fd;
ioh->op_data.write.offset = offset;
ioh->op_data.write.iov_length = iovcnt;
- pgaio_io_prepare(ioh, PGAIO_OP_WRITEV);
+ pgaio_io_start_staging(ioh);
}
diff --git a/src/backend/storage/aio/aio_subject.c b/src/backend/storage/aio/aio_subject.c
index b2bd0c235e7..321e1d8e975 100644
--- a/src/backend/storage/aio/aio_subject.c
+++ b/src/backend/storage/aio/aio_subject.c
@@ -119,33 +119,6 @@ pgaio_io_get_subject_name(PgAioHandle *ioh)
return aio_subject_info[ioh->subject]->name;
}
-/*
- * Internal function which invokes ->prepare for all the registered callbacks.
- */
-void
-pgaio_io_prepare_subject(PgAioHandle *ioh)
-{
- Assert(ioh->subject > ASI_INVALID && ioh->subject < ASI_COUNT);
- Assert(ioh->op >= 0 && ioh->op < PGAIO_OP_COUNT);
-
- for (int i = ioh->num_shared_callbacks; i > 0; i--)
- {
- PgAioHandleSharedCallbackID cbid = ioh->shared_callbacks[i - 1];
- const PgAioHandleSharedCallbacksEntry *ce = &aio_shared_cbs[cbid];
-
- if (!ce->cb->prepare)
- continue;
-
- elog(DEBUG3, "io:%d, op %s, subject %s, calling cb #%d %d/%s->prepare",
- pgaio_io_get_id(ioh),
- pgaio_io_get_op_name(ioh),
- pgaio_io_get_subject_name(ioh),
- i,
- cbid, ce->name);
- ce->cb->prepare(ioh);
- }
-}
-
/*
* Internal function which invokes ->complete for all the registered
* callbacks.
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 9bc0176a2ca..dd30856aca0 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -179,6 +179,9 @@ int backend_flush_after = DEFAULT_BACKEND_FLUSH_AFTER;
/* local state for LockBufferForCleanup */
static BufferDesc *PinCountWaitBuf = NULL;
+static void local_buffer_readv_prepare(PgAioHandle *ioh, Buffer *buffers, int nbuffers);
+static void shared_buffer_writev_prepare(PgAioHandle *ioh, Buffer *buffers, int nbuffers);
+
/*
* Backend-Private refcount management:
*
@@ -1725,7 +1728,6 @@ AsyncReadBuffers(ReadBuffersOperation *operation,
pgaio_io_set_io_data_32(ioh, (uint32 *) io_buffers, io_buffers_len);
-
if (persistence == RELPERSISTENCE_TEMP)
pgaio_io_add_shared_cb(ioh, ASC_LOCAL_BUFFER_READ);
else
@@ -1736,6 +1738,11 @@ AsyncReadBuffers(ReadBuffersOperation *operation,
did_start_io_overall = did_start_io_this = true;
smgrstartreadv(ioh, operation->smgr, forknum, io_first_block,
io_pages, io_buffers_len);
+ if (persistence == RELPERSISTENCE_TEMP)
+ local_buffer_readv_prepare(ioh, io_buffers, io_buffers_len);
+ else
+ shared_buffer_readv_prepare(ioh, io_buffers, io_buffers_len);
+ pgaio_io_stage(ioh);
ioh = NULL;
operation->nios++;
@@ -4170,10 +4177,11 @@ WriteBuffers(BuffersToWrite *to_write,
to_write->data_ptrs,
to_write->nbuffers,
false);
+ shared_buffer_writev_prepare(to_write->ioh, to_write->buffers, to_write->nbuffers);
+ pgaio_io_stage(to_write->ioh);
pgstat_count_io_op_n(IOOBJECT_RELATION, IOCONTEXT_NORMAL,
IOOP_WRITE, to_write->nbuffers);
-
for (int nbuf = 0; nbuf < to_write->nbuffers; nbuf++)
{
Buffer cur_buf = to_write->buffers[nbuf];
@@ -6952,20 +6960,16 @@ ReadBufferCompleteWriteShared(Buffer buffer, bool release_lock, bool failed)
* and writes.
*/
static void
-shared_buffer_prepare_common(PgAioHandle *ioh, bool is_write)
+shared_buffer_prepare_common(PgAioHandle *ioh, bool is_write, Buffer *buffers, int nbuffers)
{
- uint64 *io_data;
- uint8 io_data_len;
PgAioHandleRef io_ref;
BufferTag first PG_USED_FOR_ASSERTS_ONLY = {0};
- io_data = pgaio_io_get_io_data(ioh, &io_data_len);
-
pgaio_io_get_ref(ioh, &io_ref);
- for (int i = 0; i < io_data_len; i++)
+ for (int i = 0; i < nbuffers; i++)
{
- Buffer buf = (Buffer) io_data[i];
+ Buffer buf = buffers[i];
BufferDesc *bufHdr;
uint32 buf_state;
@@ -7022,16 +7026,16 @@ shared_buffer_prepare_common(PgAioHandle *ioh, bool is_write)
}
}
-static void
-shared_buffer_readv_prepare(PgAioHandle *ioh)
+void
+shared_buffer_readv_prepare(PgAioHandle *ioh, Buffer *buffers, int nbuffers)
{
- shared_buffer_prepare_common(ioh, false);
+ shared_buffer_prepare_common(ioh, false, buffers, nbuffers);
}
static void
-shared_buffer_writev_prepare(PgAioHandle *ioh)
+shared_buffer_writev_prepare(PgAioHandle *ioh, Buffer *buffers, int nbuffers)
{
- shared_buffer_prepare_common(ioh, true);
+ shared_buffer_prepare_common(ioh, true, buffers, nbuffers);
}
static PgAioResult
@@ -7135,19 +7139,15 @@ shared_buffer_writev_complete(PgAioHandle *ioh, PgAioResult prior_result)
* and writes.
*/
static void
-local_buffer_readv_prepare(PgAioHandle *ioh)
+local_buffer_readv_prepare(PgAioHandle *ioh, Buffer *buffers, int nbuffers)
{
- uint64 *io_data;
- uint8 io_data_len;
PgAioHandleRef io_ref;
- io_data = pgaio_io_get_io_data(ioh, &io_data_len);
-
pgaio_io_get_ref(ioh, &io_ref);
- for (int i = 0; i < io_data_len; i++)
+ for (int i = 0; i < nbuffers; i++)
{
- Buffer buf = (Buffer) io_data[i];
+ Buffer buf = buffers[i];
BufferDesc *bufHdr;
uint32 buf_state;
@@ -7199,27 +7199,17 @@ local_buffer_readv_complete(PgAioHandle *ioh, PgAioResult prior_result)
return result;
}
-static void
-local_buffer_writev_prepare(PgAioHandle *ioh)
-{
- elog(ERROR, "not yet");
-}
-
-
const struct PgAioHandleSharedCallbacks aio_shared_buffer_readv_cb = {
- .prepare = shared_buffer_readv_prepare,
.complete = shared_buffer_readv_complete,
.error = buffer_readv_error,
};
const struct PgAioHandleSharedCallbacks aio_shared_buffer_writev_cb = {
- .prepare = shared_buffer_writev_prepare,
.complete = shared_buffer_writev_complete,
};
const struct PgAioHandleSharedCallbacks aio_local_buffer_readv_cb = {
- .prepare = local_buffer_readv_prepare,
.complete = local_buffer_readv_complete,
.error = buffer_readv_error,
};
const struct PgAioHandleSharedCallbacks aio_local_buffer_writev_cb = {
- .prepare = local_buffer_writev_prepare,
+
};
diff --git a/src/backend/storage/smgr/md.c b/src/backend/storage/smgr/md.c
index d12225a9949..bf4522eeac6 100644
--- a/src/backend/storage/smgr/md.c
+++ b/src/backend/storage/smgr/md.c
@@ -985,9 +985,9 @@ mdstartreadv(PgAioHandle *ioh,
forknum,
blocknum,
nblocks);
- pgaio_io_add_shared_cb(ioh, ASC_MD_READV);
-
FileStartReadV(ioh, v->mdfd_vfd, iovcnt, seekpos, WAIT_EVENT_DATA_FILE_READ);
+
+ pgaio_io_add_shared_cb(ioh, ASC_MD_READV);
}
/*
@@ -1136,9 +1136,8 @@ mdstartwritev(PgAioHandle *ioh,
forknum,
blocknum,
nblocks);
- pgaio_io_add_shared_cb(ioh, ASC_MD_WRITEV);
-
FileStartWriteV(ioh, v->mdfd_vfd, iovcnt, seekpos, WAIT_EVENT_DATA_FILE_WRITE);
+ pgaio_io_add_shared_cb(ioh, ASC_MD_WRITEV);
}
diff --git a/src/include/storage/aio.h b/src/include/storage/aio.h
index caa52d2aaba..d126a10f9d4 100644
--- a/src/include/storage/aio.h
+++ b/src/include/storage/aio.h
@@ -212,12 +212,10 @@ typedef struct PgAioSubjectInfo
typedef PgAioResult (*PgAioHandleSharedCallbackComplete) (PgAioHandle *ioh, PgAioResult prior_result);
-typedef void (*PgAioHandleSharedCallbackPrepare) (PgAioHandle *ioh);
typedef void (*PgAioHandleSharedCallbackError) (PgAioResult result, const PgAioSubjectData *subject_data, int elevel);
typedef struct PgAioHandleSharedCallbacks
{
- PgAioHandleSharedCallbackPrepare prepare;
PgAioHandleSharedCallbackComplete complete;
PgAioHandleSharedCallbackError error;
} PgAioHandleSharedCallbacks;
@@ -247,6 +245,8 @@ struct ResourceOwnerData;
extern PgAioHandle *pgaio_io_get(struct ResourceOwnerData *resowner, PgAioReturn *ret);
extern PgAioHandle *pgaio_io_get_nb(struct ResourceOwnerData *resowner, PgAioReturn *ret);
+extern void pgaio_io_stage(PgAioHandle *ioh);
+
extern void pgaio_io_release(PgAioHandle *ioh);
extern void pgaio_io_release_resowner(dlist_node *ioh_node, bool on_error);
@@ -261,7 +261,7 @@ extern void pgaio_io_set_io_data_32(PgAioHandle *ioh, uint32 *data, uint8 len);
extern void pgaio_io_set_io_data_64(PgAioHandle *ioh, uint64 *data, uint8 len);
extern uint64 *pgaio_io_get_io_data(PgAioHandle *ioh, uint8 *len);
-extern void pgaio_io_prepare(PgAioHandle *ioh, PgAioOp op);
+extern void pgaio_io_start_staging(PgAioHandle *ioh);
extern int pgaio_io_get_id(PgAioHandle *ioh);
struct iovec;
diff --git a/src/include/storage/aio_internal.h b/src/include/storage/aio_internal.h
index f4c57438dd4..55677d7dc8c 100644
--- a/src/include/storage/aio_internal.h
+++ b/src/include/storage/aio_internal.h
@@ -37,10 +37,10 @@ typedef enum PgAioHandleState
/* returned by pgaio_io_get() */
AHS_HANDED_OUT,
- /* pgaio_io_start_*() has been called, but IO hasn't been submitted yet */
- AHS_DEFINED,
+ /* pgaio_io_start_staging() has been called, but IO hasn't been fully staged yet */
+ AHS_PREPARING,
- /* subjects prepare() callback has been called */
+ /* pgaio_io_stage() has been called, but the IO hasn't been submitted yet */
AHS_PREPARED,
/* IO is being executed */
@@ -249,7 +249,6 @@ typedef struct IoMethodOps
extern bool pgaio_io_was_recycled(PgAioHandle *ioh, uint64 ref_generation, PgAioHandleState *state);
-extern void pgaio_io_prepare_subject(PgAioHandle *ioh);
extern void pgaio_io_process_completion_subject(PgAioHandle *ioh);
extern void pgaio_io_process_completion(PgAioHandle *ioh, int result);
extern void pgaio_io_prepare_submit(PgAioHandle *ioh);
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 3523d8a3860..5c7d602d91b 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -425,6 +425,7 @@ extern void ScheduleBufferTagForWriteback(WritebackContext *wb_context,
/* solely to make it easier to write tests */
extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
+extern void shared_buffer_readv_prepare(PgAioHandle *ioh, Buffer *buffers, int nbuffers);
/* freelist.c */
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index e495c5309b3..446da4f0231 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -264,6 +264,8 @@ read_corrupt_rel_block(PG_FUNCTION_ARGS)
smgrstartreadv(ioh, smgr, MAIN_FORKNUM, block,
(void *) &page, 1);
+ shared_buffer_readv_prepare(ioh, &buf, 1);
+ pgaio_io_stage(ioh);
ReleaseBuffer(buf);
pgaio_io_ref_wait(&ior);
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-08 22:56 Andres Freund <andres@anarazel.de>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
0 siblings, 0 replies; 16+ messages in thread
From: Andres Freund @ 2025-01-08 22:56 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: pgsql-hackers
Hi,
On 2025-01-07 22:09:56 +0200, Heikki Linnakangas wrote:
> On 07/01/2025 18:11, Andres Freund wrote:
> > > I didn't quite understand the point of the prepare callbacks. For example,
> > > when AsyncReadBuffers() calls smgrstartreadv(), the
> > > shared_buffer_readv_prepare() callback will be called. Why doesn't
> > > AsyncReadBuffers() do the "prepare" work itself directly; why does it need
> > > to be in a callback?
> >
> > One big part of it is "ownership" - while the IO isn't completely "assembled",
> > we can release all buffer pins etc in case of an error. But if the error
> > happens just after the IO was staged, we can't - the buffer is still
> > referenced by the IO. For that the AIO subystem needs to take its own pins
> > etc. Initially the prepare callback didn't exist, the code in
> > AsyncReadBuffers() was a lot more complicated before it.
> >
> >
> > > I assume it's somehow related to error handling, but I didn't quite get
> > > it. Perhaps an "abort" callback that'd be called on error, instead of a
> > > "prepare" callback, would be better?
> >
> > I don't think an error callback would be helpful - the whole thing is that we
> > basically need claim ownership of all IO related resources IFF the IO is
> > staged. Not before (because then the IO not getting staged would mean we have
> > a resource leak), not after (because we might error out and thus not keep
> > e.g. buffers pinned).
>
> Hmm. The comments say that when you call smgrstartreadv(), the IO handle may
> no longer be modified, as the IO may be executed immediately. What if we
> changed that so that it never submits the IO, only adds the necessary
> callbacks to it?
> In that world, when smgrstartreadv() returns, the necessary details and
> completion callbacks have been set in the IO handle, but the caller can
> still do more preparation before the IO is submitted. The caller must ensure
> that it gets submitted, however, so no erroring out in that state.
>
> Currently the call stack looks like this:
>
> AsyncReadBuffers()
> -> smgrstartreadv()
> -> mdstartreadv()
> -> FileStartReadV()
> -> pgaio_io_prep_readv()
> -> shared_buffer_readv_prepare() (callback)
> <- (return)
> <- (return)
> <- (return)
> <- (return)
> <- (return)
>
> I'm thinking that the prepare work is done "on the way up" instead:
>
> AsyncReadBuffers()
> -> smgrstartreadv()
> -> mdstartreadv()
> -> FileStartReadV()
> -> pgaio_io_prep_readv()
> <- (return)
> <- (return)
> <- (return)
> -> shared_buffer_readv_prepare()
> <- (return)
>
> Attached is a patch to demonstrate concretely what I mean.
I think this would be somewhat limiting. Right now it's indeed just bufmgr.c
that needs to do a preparation (or "moving of ownership") step - but I don't
think it's necessarily going to stay that way.
Consider e.g. a hypothetical threaded future in which we have refcounted file
descriptors. While AIO is ongoing, the AIO subsystem would need to ensure that
the FD refcount is increased, otherwise you'd obviously run into trouble if
the issuing backend errored out and released its own reference as part of
resowner release.
I don't think the approach you suggest above would scale well for such a
situation - shared_buffer_readv_prepare() would again need to call to
smgr->md->fd. Whereas with the current approach md.c (or fd.c?) could just
define its own prepare callback that increased the refcount at the right
moment.
There's a few other scenarios I can think of:
- If somebody were - no idea what made me think of that - to write an smgr
implementation where storage is accessed over the network, one might need to
keep network buffers and sockets alive for the duration of the IO.
- It'd be rather useful to have support for asynchronously extending a
relation, that often requires filesystem journal IO and thus is slow. If
you're bulk loading, or the extension lock is contented, it'd be great if we
could start the next relation extension *before* it's needed and the
extension has to happen synchronously. To avoid deadlocks, such an
asynchronous extension would need to be able to release the lock in any
other backend, just like it's needed for the content locks when
asynchronously writing. Which in turn would require transferring ownership
of the relevant buffers *and* the extension lock. You could mash this
together, but it seems like a separate callback woul make it more
composable.
Does that make any sense to you?
> This adds a new pgaio_io_stage() step to the issuer, and the issuer needs to
> call the prepare functions explicitly, instead of having them as callbacks.
> Nominally that's more steps, but IMHO it's better to be explicit. The same
> actions were happening previously too, it was just hidden in the callback. I
> updated the README to show that too.
>
> I'm not wedded to this, but it feels a little better to me.
Right now the current approach seems to make more sense to me, but I'll think
about it more. I might also have missed something with my theorizing above.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-09 00:26 Andres Freund <andres@anarazel.de>
parent: Robert Haas <robertmhaas@gmail.com>
0 siblings, 1 reply; 16+ messages in thread
From: Andres Freund @ 2025-01-09 00:26 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; pgsql-hackers
Hi,
On 2025-01-07 14:59:58 -0500, Robert Haas wrote:
> On Tue, Jan 7, 2025 at 11:11 AM Andres Freund <andres@anarazel.de> wrote:
> > The difference between a handle and a reference is useful right now, to have
> > some separation between the functions that can be called by anyone (taking a
> > PgAioHandleRef) and only by the issuer (PgAioHandle). That might better be
> > solved by having a PgAioHandleIssuerRef ref or something.
>
> To me, those names don't convey that.
I'm certainly not wedded to these names - I went back and forth between
different names a fair bit, because I wasn't quite happy. I am however certain
that the current names are better than what it used to be (PgAioInProgress and
because that's long, a bunch of PgAioIP* names) :)
To make sure were talking about the same things, I am thinking of the
following "entities" needing names:
1) Shared memory representation of an IO, for the AIO subsystem internally
Currently: PgAioHandle
Because shared memory is limited, we need to reuse this entity. This reuse
needs to be possible "immediately" after completion, to avoid a bunch of
nasty scenarios.
To distinguish a reused PgAioHandle from its "prior" incarnation, each
PgAioHandle has a 64bit "generation counter.
In addition to being referenceable via pointer, it's also possible to
assign a 32bit integer to each PgAioHandle, as there is a fixed number of
them.
2) A way for the issuer of an IO to reference 1), to attach information to the
IO
Currently: PgAioHandle*
As long as the issuer hasn't yet staged the IO, it can't be
reused. Therefore it's OK to just point to the PgAioHandle.
One disadvantage of just using a pointer to PgAioHandle* is that it's
harder to distinguish subystem-internal functions that accept PgAioHandle*
from "public" functions that accept the "issuer reference".
3) A way for any backend to wait for a specific IO to complete
Currently: PgAioHandleRef
This references 1) using a 32 bit ID and the 64bit generation.
This is used to allow any backend to wait for a specific IO to
complete. E.g. by including it in the BufferDesc so that WaitIO can wait
for it.
Because it includes the generation it's trivial to detect whether the
PgAioHandle was reused.
> I would perhaps call the thing that supports issuer-only operations a
> "PgAio" and the thing other people can use a "PgAioHandle". Or
> "PgAioRequest" and "PgAioHandle" or something like that. With
> PgAioHandleRef, IMHO you've got two words that both imply a layer of
> indirection -- "handle" and "ref" -- which doesn't seem quite as nice,
> because then the other thing -- "PgAioHandle" still sort of implies one
> layer of indirection and the whole thing seems a bit less clear.
It's indirections all the way down. The PG representation of "one IO" in the
end is just an indirection for a kernel operation :)
I would like to split 1) and 2) above.
1) PgAio{Handle,Request,} (a large struct) - used internally by AIO subsystem,
"pointed to" by the following
2) PgAioIssuerRef (an ID or pointer) - used by the issuer to incrementally
define the IO
3) PgAioWaitRef - (an ID and generation) - used to wait for a specific IO to
complete, not affected by reuse of PgAioHandle
> > > REAPED feels like a bad name. It sounds like a later stage than COMPLETED,
> > > but it's actually vice versa.
> >
> > What would you call having gotten "completion notifications" from the kernel,
> > but not having processed them?
>
> The Linux kernel calls those zombie processes, so we could call it a ZOMBIE
> state, but that seems like it might be a bit of inside baseball.
ZOMBIE feels even later than REAPED to me :)
> I do agree with Heikki that REAPED sounds later than COMPLETED, because you
> reap zombie processes by collecting their exit status. Maybe you could have
> AHS_COMPLETE or AHS_IO_COMPLETE for the state where the I/O is done but
> there's still completion-related work to be done, and then the other state
> could be AHS_DONE or AHS_FINISHED or AHS_FINAL or AHS_REAPED or something.
How about
AHS_COMPLETE_KERNEL or AHS_COMPLETE_RAW - raw syscall completed
AHS_COMPLETE_SHARED_CB - shared callback completed
AHS_COMPLETE_LOCAL_CB - local callback completed
?
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-13 20:43 Robert Haas <robertmhaas@gmail.com>
parent: Andres Freund <andres@anarazel.de>
0 siblings, 1 reply; 16+ messages in thread
From: Robert Haas @ 2025-01-13 20:43 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; pgsql-hackers
On Wed, Jan 8, 2025 at 7:26 PM Andres Freund <andres@anarazel.de> wrote:
> 1) Shared memory representation of an IO, for the AIO subsystem internally
>
> Currently: PgAioHandle
>
> 2) A way for the issuer of an IO to reference 1), to attach information to the
> IO
>
> Currently: PgAioHandle*
>
> 3) A way for any backend to wait for a specific IO to complete
>
> Currently: PgAioHandleRef
With that additional information, I don't mind this naming too much,
but I still think PgAioHandle -> PgAio and PgAioHandleRef ->
PgAioHandle is worth considering. Compare BackgroundWorkerSlot and
BackgroundWorkerHandle, which suggests PgAioHandle -> PgAioSlot and
PgAioHandleRef -> PgAioHandle.
> ZOMBIE feels even later than REAPED to me :)
Makes logical sense, because you would assume that you die first and
then later become an undead creature, but the UNIX precedent is that
dying turns you into a zombie and someone then has to reap the exit
status for you to be just plain dead. :-)
> > I do agree with Heikki that REAPED sounds later than COMPLETED, because you
> > reap zombie processes by collecting their exit status. Maybe you could have
> > AHS_COMPLETE or AHS_IO_COMPLETE for the state where the I/O is done but
> > there's still completion-related work to be done, and then the other state
> > could be AHS_DONE or AHS_FINISHED or AHS_FINAL or AHS_REAPED or something.
>
> How about
>
> AHS_COMPLETE_KERNEL or AHS_COMPLETE_RAW - raw syscall completed
> AHS_COMPLETE_SHARED_CB - shared callback completed
> AHS_COMPLETE_LOCAL_CB - local callback completed
>
> ?
That's not bad. I like RAW better than KERNEL. I was hoping to use
different works like COMPLETE and DONE rather than, as you did it
here, COMPLETE and COMPLETE, but it's probably fine.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-13 21:46 Andres Freund <andres@anarazel.de>
parent: Robert Haas <robertmhaas@gmail.com>
0 siblings, 1 reply; 16+ messages in thread
From: Andres Freund @ 2025-01-13 21:46 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; pgsql-hackers
Hi,
On 2025-01-13 15:43:46 -0500, Robert Haas wrote:
> On Wed, Jan 8, 2025 at 7:26 PM Andres Freund <andres@anarazel.de> wrote:
> > 1) Shared memory representation of an IO, for the AIO subsystem internally
> >
> > Currently: PgAioHandle
> >
> > 2) A way for the issuer of an IO to reference 1), to attach information to the
> > IO
> >
> > Currently: PgAioHandle*
> >
> > 3) A way for any backend to wait for a specific IO to complete
> >
> > Currently: PgAioHandleRef
>
> With that additional information, I don't mind this naming too much,
> but I still think PgAioHandle -> PgAio and PgAioHandleRef ->
> PgAioHandle is worth considering. Compare BackgroundWorkerSlot and
> BackgroundWorkerHandle, which suggests PgAioHandle -> PgAioSlot and
> PgAioHandleRef -> PgAioHandle.
I don't love PgAioHandle -> PgAio as there are other things (e.g. per-backend
state) in the PgAio namespace...
> > > I do agree with Heikki that REAPED sounds later than COMPLETED, because you
> > > reap zombie processes by collecting their exit status. Maybe you could have
> > > AHS_COMPLETE or AHS_IO_COMPLETE for the state where the I/O is done but
> > > there's still completion-related work to be done, and then the other state
> > > could be AHS_DONE or AHS_FINISHED or AHS_FINAL or AHS_REAPED or something.
> >
> > How about
> >
> > AHS_COMPLETE_KERNEL or AHS_COMPLETE_RAW - raw syscall completed
> > AHS_COMPLETE_SHARED_CB - shared callback completed
> > AHS_COMPLETE_LOCAL_CB - local callback completed
> >
> > ?
>
> That's not bad. I like RAW better than KERNEL.
Cool.
> I was hoping to use different works like COMPLETE and DONE rather than, as
> you did it here, COMPLETE and COMPLETE, but it's probably fine.
Once the IO is really done, the handle is immediately recycled (and moved into
IDLE state, ready to be used again).
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-01-14 13:40 Robert Haas <robertmhaas@gmail.com>
parent: Andres Freund <andres@anarazel.de>
0 siblings, 0 replies; 16+ messages in thread
From: Robert Haas @ 2025-01-14 13:40 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; pgsql-hackers
On Mon, Jan 13, 2025 at 4:46 PM Andres Freund <andres@anarazel.de> wrote:
> Once the IO is really done, the handle is immediately recycled (and moved into
> IDLE state, ready to be used again).
OK, fair enough.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: AIO v2.2
@ 2025-02-21 19:31 Andres Freund <andres@anarazel.de>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
1 sibling, 0 replies; 16+ messages in thread
From: Andres Freund @ 2025-02-21 19:31 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: pgsql-hackers
Hi,
I was just going through comments about LWLockDisown() and was reminded of
this:
On 2025-01-07 18:08:51 +0200, Heikki Linnakangas wrote:
> On LWLockDisown():
> > + * NB: This will leave lock->owner pointing to the current backend (if
> > + * LOCK_DEBUG is set). We could add a separate flag indicating that, but it
> > + * doesn't really seem worth it.
>
> Hmm. I won't insist, but I feel it probably would be worth it. This is only
> in LOCK_DEBUG mode so there's no performance penalty in non-debug builds,
> and when you do compile with LOCK_DEBUG you probably appreciate any extra
> information.
I don't think that makes sense, as we, independent of this change, never clear
lock->owner. Not even when releasing a lock! The background to that, I think,
is that there were some cases where we forgot to wake up all backends due to
race conditions, and that for that it's really useful to know the last owner.
That could perhaps be evolved or documented better, but it's pretty much
independent of the patch at hand.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 16+ messages in thread
end of thread, other threads:[~2025-02-21 19:31 UTC | newest]
Thread overview: 16+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2025-01-01 04:03 Re: AIO v2.2 Andres Freund <andres@anarazel.de>
2025-01-06 18:52 ` Noah Misch <noah@leadboat.com>
2025-01-06 21:40 ` Andres Freund <andres@anarazel.de>
2025-01-07 19:10 ` Noah Misch <noah@leadboat.com>
2025-01-07 15:09 ` Heikki Linnakangas <hlinnaka@iki.fi>
2025-01-07 16:11 ` Andres Freund <andres@anarazel.de>
2025-01-07 19:59 ` Robert Haas <robertmhaas@gmail.com>
2025-01-09 00:26 ` Andres Freund <andres@anarazel.de>
2025-01-13 20:43 ` Robert Haas <robertmhaas@gmail.com>
2025-01-13 21:46 ` Andres Freund <andres@anarazel.de>
2025-01-14 13:40 ` Robert Haas <robertmhaas@gmail.com>
2025-01-07 20:09 ` Heikki Linnakangas <hlinnaka@iki.fi>
2025-01-08 22:56 ` Andres Freund <andres@anarazel.de>
2025-01-07 16:08 ` Heikki Linnakangas <hlinnaka@iki.fi>
2025-01-07 16:32 ` Andres Freund <andres@anarazel.de>
2025-02-21 19:31 ` Andres Freund <andres@anarazel.de>
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox