agora inbox for pgsql-hackers@postgresql.org  
help / color / mirror / Atom feed
Buffer locking is special (hints, checksums, AIO writes)
120+ messages / 18 participants
[nested] [flat]

* Buffer locking is special (hints, checksums, AIO writes)
@ 2025-08-22 19:44 Andres Freund <andres@anarazel.de>
  2025-08-23 10:31 ` Re: Buffer locking is special (hints, checksums, AIO writes) Mihail Nikalayeu <mihailnikalayeu@gmail.com>
  2025-08-26 20:21 ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 3 replies; 120+ messages in thread

From: Andres Freund @ 2025-08-22 19:44 UTC (permalink / raw)
  To: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>

Hi,


I'm working on making bufmgr.c ready for AIO writes. That requires some
infrastructure changes that I wanted to discuss... In this email I'm trying to
outline the problems and what I currently think we should do.


== Problem 1 - hint bits ==

Currently pages that are being written out are copied if checksums are
enabled. The copy is needed because hint bits can be set with just a share
lock. Writing out a buffer only requires a share lock. If we didn't copy the
buffer, the computed checksum can be falsified by the checksum.

Even without checksums hint bits being set while IO is ongoing is an issue,
e.g. btrfs assumes that pages do not change while being written out with
direct-io, and corrupts itself if they do [1].

The copy we make to avoid checksum failure are not at all cheap [2], but are
particularly problematic for AIO, where a lot of buffers can undergo AIO at
the same time. For AIO the issue is that the longer-lived bounce buffers that
I had initially prototyped actually end up using a lot of memory, making it
hard to figure out what defaults to set.

In [2] I had worked on an approach to avoid copying pages during writes. It
avoided the need to copy buffers by not allowing hint bits to be set while IO
is ongoing. It did so by skipping hint bit writes while IO is ongoing (plus a
fair bit of infrastructure to make that cheap enough).


In that thread Heikki was wondering whether we should instead not go for a
more fundamental solution to the problem, like introducing a separate lock
level that prevents concurrent modifications but allows share lockers. At the
time I was against that, because it seemed like a large project.

However, since then I realized there's a architecturally related second issue:


== Problem 2 - AIO writes vs exclusive locks ==

Separate from the hint bit issue, there is a second issue that I didn't have a
good answer for: Making acquiring an exclusive lock concurrency safe in the
presence of asynchronous writes:

The problem is that while a buffer is being written out, it obviously has to
be share locked. That's true even with AIO. With AIO the share lock is held
once the IO is completed. The problem is that if a backend wants to
exclusively lock a buffer undergoing AIO, it can't just wait for the content
lock as today, it might have to actually reap the IO completion from the
operating system. If one just were to wait for the content lock, there's no
forward progress guarantee.

The buffer's state "knows" that it's undergoing write IO (BM_VALID and
BM_IO_IN_PROGRESS are set). To ensure forward progress guarantee, an exclusive
locker needs to wait for the IO (pgaio_wref_wait(BufferDesc->->io_wref)). The
problem is that it's surprisingly hard to do so race free:

If a backend A were to just check if a buffer is undergoing IO before locking
it, a backend B could start IO on the buffer between A checking for
BM_IO_IN_PROGRESS and acquiring the content lock.  We obviously can't just
hold the buffer header spinlock across a blocking lwlock acquisition.

There potentially are ways to synchronize the buffer state and the content
lock, but it requires deep integration between bufmgr.c and lwlock.c.


== Problem 3 - Cacheline contention ==

This is unrelated to AIO, but might influence the architecture for potential
solutions.  It might make sense to skip over this section on a quicker
read-through.

The problem is that in some workloads the cacheline containing the BufferDesc
becomes very hot. The most common case are workloads with lots of index nested
loops, the root page and some of the other inner pages in the index can become
very contended.

I've seen cases where running the same workload in four separate copies in the
same postgres instance yields ~8x the throughput - on a machine with forty
cores, there's machines with many more these days [4]/

There are three main issues:

a) We manipulate the same cache line multiple times. E.g. btree always first
   pins the page and then locks the page, which performs two atomic operations
   on the same cacheline.

   I've experimented with putting the content lock on a separate cacheline,
   but that causes regressions in other cases.

   Leaving nontrivial implementation issues aside, we could release the lock
   and pin at the same time. With a bit increased difficulty we could do the
   same for pin and lock acquisition.  nbtree almost always does the two
   together, so it'd not be hard to make it benefit from such an optimization.


b) Many operations, like unpinning a buffer, need to use CAS, instead of an
   atomic subtraction.

   atomic add/sub scales a *lot* better than compare and swap, as there is no
   need to retry. With increased contention the number of retries increases
   further and further.

   The reason we need to use CAS in so many places is the following:

   Historically postgres' spinlocks don't use an atomic operation to release
   the spinlock on x86. When making buffer header locks their own thing, I
   carried that forward, to avoid performance regressions.  Because of that
   BufferDesc.state may not be modified while the buffer header spinlock is
   held - which is incompatible with using atomic-sub.

   However, since then we converted most of the "hot" uses of buffer header
   spinlocks into CAS loops (see e.g. PinBuffer()). That makes it feasible to
   use an atomic operation for the buffer header lock release, which in turn
   allows e.g. unpinning with an atomic sub.


c) Read accesses to the BufferDesc cause contention

   Some code, like nbtree, relies on functions like
   BufferGetBlockNumber(). Unfortunately that contends with concurrent
   modifications of the buffer descriptor (like pinning). Potential solutions
   are to rely less on functions like BufferGetBlockNumber() or to split out
   the memory for that into a separate (denser?) array.


d) Even after addressing all of the above, there's still a lot of contention

   I think the solution here would be something roughly to fastpath locks. If
   a buffer is very contended, we can mark it as super-pinned & share locked,
   avoiding any atomic operation on the buffer descriptor itself. Instead the
   current lock and pincount would be stored in each backends PGPROC.
   Obviously evicting or exclusively-locking such a buffer would be a lot more
   expensive.

   I've prototyped it and it helps a *lot*.  The reason I mention this here is
   that this seems impossible to do while using the generic lwlocks for the
   content lock.



== Solution ? ==

My conclusion from the above is that we ought to:


A) Make Buffer Locks something separate from lwlocks

   As part of that introduce a new lock level in-between share and exclusive
   locks (provisionally called share-exclusive, but I hate it). The new lock
   level allows other share lockers, but can only be held by one backend.

   This allows to change the rules so that:

   1) Share lockers are not allowed to modify buffers anymore
   2) Hint bits need the new lock mode (conditionally upgrading the lock in SetHintBits())
   3) Write IO needs to the new lock level

   This addresses 1) from above


B) Merge BufferDesc.state and the content lock

   This allows to address 2) from above, as we now atomically can check if IO
   was concurrently initiated.

   Obviously BufferDesc.state is not currently wide enough, therefore the
   buffer state has to be updated to a 64bit variable.


C) Allow some modifications of BufferDesc.state while holding spinlock

   Just naively doing the above two things reduces scalability, as the
   likelihood of CAS failures increases, due to increased number of
   modifications of the same atomic variable.

   However, by allowing unpinning while the buffer header spinlock is held,
   scalability considerably improves in my tests.

   Doing so requires changing all uses of LockBufHdr(), but by introducing a
   static inline helper the complexity can largely be encapsulated.



I've prototyped the above. The current state is pretty rough, but before I
spend the non-trivial time to make it into an understandable sequence of
changes, I wanted to get feedback.


Does this plan sound reasonable?


The hardest part about this change is that everything kind of depends on each
other. The changes are large enough that they clearly can't just be committed
at once, but doing them over time risks [temporary] performance regressions.




The order of changes I think makes the most sense is the following:

1) Allow some modifications while holding the buffer header spinlock

2) Reduce buffer pin with just an atomic-sub

   This needs to happen first, otherwise there are performance regressions
   during the later steps.

3) Widen BufferDesc.state to 64 bits

4) Implement buffer locking inside BufferDesc.state

5) Do IO while holding share-exclusive lock and require all buffer
   modifications to at least hold share exclusive lock

6) Wait for AIO when acquiring an exclusive content lock

(some of these will likely have parts of their own, but that's details)


Sane?


DOES ANYBODY HAVE A BETTER NAME THAN SHARE-EXCLUSIVE???!?


Greetings,

Andres Freund

[1] https://www.postgresql.org/message-id/CA%2BhUKGKSBaz78Fw3WTF3Q8ArqKCz1GgsTfRFiDPbu-j9OFz-jw%40mail.g...
[2] https://www.postgresql.org/message-id/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf%40gcnactj4z...
[3] In some cases it causes slowdowns for checkpointer close to 50%!
[4] https://anarazel.de/talks/2024-10-23-pgconf-eu-numa-vs-postgresql/numa-vs-postgresql.pdf - slide 19





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-08-23 10:31 ` Mihail Nikalayeu <mihailnikalayeu@gmail.com>
  2 siblings, 0 replies; 120+ messages in thread

From: Mihail Nikalayeu @ 2025-08-23 10:31 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>

Hello, Andres!

Andres Freund <andres@anarazel.de>:
>
>    As part of that introduce a new lock level in-between share and exclusive
>    locks (provisionally called share-exclusive, but I hate it). The new lock
>    level allows other share lockers, but can only be held by one backend.
>
>    This allows to change the rules so that:
>
>    1) Share lockers are not allowed to modify buffers anymore
>    2) Hint bits need the new lock mode (conditionally upgrading the lock in SetHintBits())
>    3) Write IO needs to the new lock level

IIUC, it may be mapped to existing locking system:

BUFFER_LOCK_SHARE          ---->      AccessShareLock
new lock mode                           ---->     ExclusiveLock
BUFFER_LOCK_EXCLUSIVE   ---->     AccessExclusiveLock

So, it all may be named:

BUFFER_LOCK_ACCESS_SHARE
BUFFER_LOCK_EXCLUSIVE
BUFFER_LOCK_ACCESS_EXCLUSIVE

being more consistent with table-level locking system.

Greetings,
Mikhail.





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-08-26 20:21 ` Robert Haas <robertmhaas@gmail.com>
  2025-08-26 21:00   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2 siblings, 1 reply; 120+ messages in thread

From: Robert Haas @ 2025-08-26 20:21 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>

On Fri, Aug 22, 2025 at 3:45 PM Andres Freund <andres@anarazel.de> wrote:
> My conclusion from the above is that we ought to:
>
> A) Make Buffer Locks something separate from lwlocks
> B) Merge BufferDesc.state and the content lock
> C) Allow some modifications of BufferDesc.state while holding spinlock

+1 to (A) and (B). No particular opinion on (C) but if it works well, great.

> The order of changes I think makes the most sense is the following:
>
> 1) Allow some modifications while holding the buffer header spinlock
> 2) Reduce buffer pin with just an atomic-sub
> 3) Widen BufferDesc.state to 64 bits
> 4) Implement buffer locking inside BufferDesc.state
> 5) Do IO while holding share-exclusive lock and require all buffer
>    modifications to at least hold share exclusive lock
> 6) Wait for AIO when acquiring an exclusive content lock

No strong objections. I certainly like getting to (5) and (6) and I
think those are in the right order. I'm not sure about the rest. I
thought (1) and (2) were the same change after reading your email; and
it surprises me a little bit that (2) is separate from (4). But I'm
sure you have a much better sense of this than I do.

> DOES ANYBODY HAVE A BETTER NAME THAN SHARE-EXCLUSIVE???!?

AFAIK "share exclusive" or "SX" is standard terminology. While I'm not
wholly hostile to the idea of coming up with something else, I don't
think our tendency to invent our own way to do everything is one of
our better tendencies as a project.

-- 
Robert Haas
EDB: http://www.enterprisedb.com





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-26 20:21 ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
@ 2025-08-26 21:00   ` Andres Freund <andres@anarazel.de>
  2025-08-27 00:14     ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-08-26 21:00 UTC (permalink / raw)
  To: Robert Haas <robertmhaas@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>

Hi,

On 2025-08-26 16:21:36 -0400, Robert Haas wrote:
> On Fri, Aug 22, 2025 at 3:45 PM Andres Freund <andres@anarazel.de> wrote:
> > My conclusion from the above is that we ought to:
> >
> > A) Make Buffer Locks something separate from lwlocks
> > B) Merge BufferDesc.state and the content lock
> > C) Allow some modifications of BufferDesc.state while holding spinlock
> 
> +1 to (A) and (B). No particular opinion on (C) but if it works well, great.

Without it I see performance regressions due to the increased rate of CAS
failures due to having more changes to one atomic variable :/



> > The order of changes I think makes the most sense is the following:
> >
> > 1) Allow some modifications while holding the buffer header spinlock
> > 2) Reduce buffer pin with just an atomic-sub
> > 3) Widen BufferDesc.state to 64 bits
> > 4) Implement buffer locking inside BufferDesc.state
> > 5) Do IO while holding share-exclusive lock and require all buffer
> >    modifications to at least hold share exclusive lock
> > 6) Wait for AIO when acquiring an exclusive content lock
> 
> No strong objections. I certainly like getting to (5) and (6) and I
> think those are in the right order. I'm not sure about the rest.


> I thought (1) and (2) were the same change after reading your email

They are certainly related. I thought it'd make sense to split them as
outlined above, as (1) is relatively verbose on its own, but far more
mechanical.


> and it surprises me a little bit that (2) is separate from (4).

Without doing 2) first, I see performance/scalability regressions doing
(4). Doing (3) without (2) also hurts...



> > DOES ANYBODY HAVE A BETTER NAME THAN SHARE-EXCLUSIVE???!?
> 
> AFAIK "share exclusive" or "SX" is standard terminology. While I'm not
> wholly hostile to the idea of coming up with something else, I don't
> think our tendency to invent our own way to do everything is one of
> our better tendencies as a project.

I guess it bothers me that we'd use share-exclusive to mean the buffer can't
be modified, but for real (vs share, which does allow some modifications). But
it's very well plausible that there's no meaningfully better name, in which
case we certainly shouldn't differ from what's somewhat commonly used.


Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-26 20:21 ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  2025-08-26 21:00   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-08-27 00:14     ` Noah Misch <noah@leadboat.com>
  2025-08-27 14:03       ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  2025-08-27 16:18       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 2 replies; 120+ messages in thread

From: Noah Misch @ 2025-08-27 00:14 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Robert Haas <robertmhaas@gmail.com>; pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>

On Fri, Aug 22, 2025 at 03:44:48PM -0400, Andres Freund wrote:
> I'm working on making bufmgr.c ready for AIO writes.

Nice!

> == Problem 2 - AIO writes vs exclusive locks ==
> 
> Separate from the hint bit issue, there is a second issue that I didn't have a
> good answer for: Making acquiring an exclusive lock concurrency safe in the
> presence of asynchronous writes:
> 
> The problem is that while a buffer is being written out, it obviously has to
> be share locked. That's true even with AIO. With AIO the share lock is held
> once the IO is completed. The problem is that if a backend wants to
> exclusively lock a buffer undergoing AIO, it can't just wait for the content
> lock as today, it might have to actually reap the IO completion from the
> operating system. If one just were to wait for the content lock, there's no
> forward progress guarantee.
> 
> The buffer's state "knows" that it's undergoing write IO (BM_VALID and
> BM_IO_IN_PROGRESS are set). To ensure forward progress guarantee, an exclusive
> locker needs to wait for the IO (pgaio_wref_wait(BufferDesc->->io_wref)). The
> problem is that it's surprisingly hard to do so race free:
> 
> If a backend A were to just check if a buffer is undergoing IO before locking
> it, a backend B could start IO on the buffer between A checking for
> BM_IO_IN_PROGRESS and acquiring the content lock.  We obviously can't just
> hold the buffer header spinlock across a blocking lwlock acquisition.
> 
> There potentially are ways to synchronize the buffer state and the content
> lock, but it requires deep integration between bufmgr.c and lwlock.c.

You may have considered and rejected simpler alternatives for (2) before
picking the approach you go on to outline.  Anything interesting?  For
example, I imagine these might work with varying degrees of inefficiency:

- Use LWLockConditionalAcquire() with some nonstandard waiting protocol when
  there's a non-I/O lock conflict.
- Take BM_IO_IN_PROGRESS before exclusive-locking, then release it.

> == Problem 3 - Cacheline contention ==

> c) Read accesses to the BufferDesc cause contention
> 
>    Some code, like nbtree, relies on functions like
>    BufferGetBlockNumber(). Unfortunately that contends with concurrent
>    modifications of the buffer descriptor (like pinning). Potential solutions
>    are to rely less on functions like BufferGetBlockNumber() or to split out
>    the memory for that into a separate (denser?) array.

Agreed.  BufferGetBlockNumber() could even use a new local (non-shmem) data
structure, since the buffer's mapping can't change until we unpin.

> d) Even after addressing all of the above, there's still a lot of contention
> 
>    I think the solution here would be something roughly to fastpath locks. If
>    a buffer is very contended, we can mark it as super-pinned & share locked,
>    avoiding any atomic operation on the buffer descriptor itself. Instead the
>    current lock and pincount would be stored in each backends PGPROC.
>    Obviously evicting or exclusively-locking such a buffer would be a lot more
>    expensive.
> 
>    I've prototyped it and it helps a *lot*.  The reason I mention this here is
>    that this seems impossible to do while using the generic lwlocks for the
>    content lock.

Nice.

On Tue, Aug 26, 2025 at 05:00:13PM -0400, Andres Freund wrote:
> On 2025-08-26 16:21:36 -0400, Robert Haas wrote:
> > On Fri, Aug 22, 2025 at 3:45 PM Andres Freund <andres@anarazel.de> wrote:
> > > The order of changes I think makes the most sense is the following:

No concerns so far.  I won't claim I can picture all the implications and be
sure this is the right thing, but it sounds promising.  I like your principle
of ordering changes to avoid performance regressions.

> > > DOES ANYBODY HAVE A BETTER NAME THAN SHARE-EXCLUSIVE???!?

I would consider {AccessShare, Exclusive, AccessExclusive}.  What the $SUBJECT
proposal calls SHARE-EXCLUSIVE would become Exclusive.  That has the same
conflict matrix as the corresponding heavyweight locks, which seems good.  I
don't love our mode names, particularly ShareRowExclusive being unsharable.
However, learning one special taxonomy is better than learning two.

> > AFAIK "share exclusive" or "SX" is standard terminology.

Can you say more about that?  I looked around at
https://google.com/search?q=share+exclusive+%22sx%22+lock but didn't find
anything well-aligned with the proposal:

https://dev.mysql.com/doc/dev/mysql-server/latest//PAGE_LOCK_ORDER.html looked
most relevant, but it doesn't give the big idea.
https://mysqlonarm.github.io/Understanding-InnoDB-rwlock-stats/ is less
authoritative but does articulate the big idea, as "Shared-Exclusive (SX):
offer write access to the resource with inconsistent read. (relaxed
exclusive)."  That differs from $SUBJECT semantics, in which SHARE-EXCLUSIVE
can't see inconsistent reads.

https://docs.oracle.com/en/database/oracle/oracle-database/19/arpls/DBMS_LOCK.html
has term SX = "sub exclusive".  I gather an SX lock on a table lets one do
SELECT FOR UPDATE on that table (each row is the "sub"component being locked).

https://man.freebsd.org/cgi/man.cgi?query=sx_slock&sektion=9&format=html uses
the term "SX", but it's more like our lwlocks.  One acquires S or X, not
blends of them.





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-26 20:21 ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  2025-08-26 21:00   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-27 00:14     ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
@ 2025-08-27 14:03       ` Robert Haas <robertmhaas@gmail.com>
  2025-09-01 02:03         ` Re: Buffer locking is special (hints, checksums, AIO writes) Michael Paquier <michael@paquier.xyz>
  1 sibling, 1 reply; 120+ messages in thread

From: Robert Haas @ 2025-08-27 14:03 UTC (permalink / raw)
  To: Noah Misch <noah@leadboat.com>; +Cc: Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>

On Tue, Aug 26, 2025 at 8:14 PM Noah Misch <noah@leadboat.com> wrote:
> > > AFAIK "share exclusive" or "SX" is standard terminology.
>
> Can you say more about that?

Looks like I was misremembering. I was thinking of Gray & Reuter,
Transaction Processing: Concepts and Techniques, 1993. However,
opening it up, I find that his vocabulary is slightly different. He
offers the following six lock modes: IS, IX, S, SIX, Update, X. "I"
means "intent" and acts as a modifier to the letter that follows.
Hence, SIX means "a course-granularity shared lock with intent to set
finer-granularity exclusive locks" (p. 408). His lock manager is
hierarchical, so taking a SIX lock on a table means that you are
allowed to read all the rows in the table and you are allowed to
exclusive-lock individual rows as desired and nobody is allowed to
exclusive-lock any rows in the table. It is compatible only with IS;
that is, it does not preclude other people from share-locking
individual rows (which might delay your exclusive locks on those
rows). Since we don't have intent-locking in PostgreSQL, I think my
brain mentally flattened this hierarchy down to S, X, SX, but that's
not what he actually wrote.

His "Update" locks are also somewhat interesting: an update lock is
exactly like an exclusive lock except that it permits PAST
share-locks. You take an update lock when you currently need a
share-lock but anticipate the possibility of needing an
exclusive-lock. This is a deadlock avoidance strategy: updaters will
take turns, and some of them will ultimately want exclusive locks and
others won't, but they can't deadlock against each other as long as
they all take "Update" locks initially and don't try to upgrade to
that level later. An updater's attempt to upgrade to an exclusive lock
can still be delayed by, or deadlock against, share lockers, but those
typically won't try to higher lock levels later.

If we were to use the existing PostgreSQL naming convention, I think
I'd probably argue that the nearest parallel to this level is
ShareUpdateExclusive: a self-exclusive lock level that permits
ordinary table access to continue while blocking exclusive locks, used
for an in-flight maintenance operation. But that's arguable, of
course.

-- 
Robert Haas
EDB: http://www.enterprisedb.com





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-26 20:21 ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  2025-08-26 21:00   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-27 00:14     ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2025-08-27 14:03       ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
@ 2025-09-01 02:03         ` Michael Paquier <michael@paquier.xyz>
  0 siblings, 0 replies; 120+ messages in thread

From: Michael Paquier @ 2025-09-01 02:03 UTC (permalink / raw)
  To: Robert Haas <robertmhaas@gmail.com>; +Cc: Noah Misch <noah@leadboat.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>

On Wed, Aug 27, 2025 at 10:03:08AM -0400, Robert Haas wrote:
> If we were to use the existing PostgreSQL naming convention, I think
> I'd probably argue that the nearest parallel to this level is
> ShareUpdateExclusive: a self-exclusive lock level that permits
> ordinary table access to continue while blocking exclusive locks, used
> for an in-flight maintenance operation. But that's arguable, of
> course.

ShareUpdateExclusive is a term that's been used for some time now and
relates to knowledge that's quite spread in the tree, so it feels like
a natural fit for the use-case described on this thread as we'd want a
self-conflicting lock.  share-exclusive did not sound that bad to me,
TBH, quite the contrary, when applied to buffer locking for aio.

"intent" is also a word I've bumped quite a lot into while looking at
some naming convention, but this is more related to the fact that a
lock is going to be taken, which we don't really have.  So that feels
off.
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../aLT_CzoHoRiGuLz9@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-26 20:21 ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  2025-08-26 21:00   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-27 00:14     ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
@ 2025-08-27 16:18       ` Andres Freund <andres@anarazel.de>
  2025-08-27 19:14         ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2025-08-27 19:22         ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  1 sibling, 2 replies; 120+ messages in thread

From: Andres Freund @ 2025-08-27 16:18 UTC (permalink / raw)
  To: Noah Misch <noah@leadboat.com>; +Cc: Robert Haas <robertmhaas@gmail.com>; pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>

Hi,

On 2025-08-26 17:14:49 -0700, Noah Misch wrote:
> On Fri, Aug 22, 2025 at 03:44:48PM -0400, Andres Freund wrote:
> > == Problem 2 - AIO writes vs exclusive locks ==
> >
> > Separate from the hint bit issue, there is a second issue that I didn't have a
> > good answer for: Making acquiring an exclusive lock concurrency safe in the
> > presence of asynchronous writes:
> >
> > The problem is that while a buffer is being written out, it obviously has to
> > be share locked. That's true even with AIO. With AIO the share lock is held
> > once the IO is completed. The problem is that if a backend wants to
> > exclusively lock a buffer undergoing AIO, it can't just wait for the content
> > lock as today, it might have to actually reap the IO completion from the
> > operating system. If one just were to wait for the content lock, there's no
> > forward progress guarantee.
> >
> > The buffer's state "knows" that it's undergoing write IO (BM_VALID and
> > BM_IO_IN_PROGRESS are set). To ensure forward progress guarantee, an exclusive
> > locker needs to wait for the IO (pgaio_wref_wait(BufferDesc->->io_wref)). The
> > problem is that it's surprisingly hard to do so race free:
> >
> > If a backend A were to just check if a buffer is undergoing IO before locking
> > it, a backend B could start IO on the buffer between A checking for
> > BM_IO_IN_PROGRESS and acquiring the content lock.  We obviously can't just
> > hold the buffer header spinlock across a blocking lwlock acquisition.
> >
> > There potentially are ways to synchronize the buffer state and the content
> > lock, but it requires deep integration between bufmgr.c and lwlock.c.
>
> You may have considered and rejected simpler alternatives for (2) before
> picking the approach you go on to outline.

I tried a few things...


> Anything interesting?

Not really.

The first one you propose is what I looked at for a while:

> For example, I imagine these might work with varying degrees of
> inefficiency:
>
> - Use LWLockConditionalAcquire() with some nonstandard waiting protocol when
>   there's a non-I/O lock conflict.

It's nontrivial to make this race free - the problem is the case where we *do*
have to wait for an exclusive content lock. It's possible for the lwlock to be
released by the owning backend and for IO to be started, after checking
whether IO is in progress (after LWLockConditionalAcquire() had failed).

I came up with a complicated scheme, where setting IO in progress would
afterwards wake up all lwlock waiters and all exclusive content lock waits
were done with LWLockAcquireOrWait().  I think that was working - but it's
also a slower and seems really fragile and ugly.


> - Take BM_IO_IN_PROGRESS before exclusive-locking, then release it.

That just seems expensive. We could make it cheaper by doing it only if a
LWLockConditionalAcquire() doesn't succeed. But it still seems not great.  And
it doesn't really help with addressing the 'setting hint bits while IO is in
progress" part...


> > == Problem 3 - Cacheline contention ==
>
> > c) Read accesses to the BufferDesc cause contention
> >
> >    Some code, like nbtree, relies on functions like
> >    BufferGetBlockNumber(). Unfortunately that contends with concurrent
> >    modifications of the buffer descriptor (like pinning). Potential solutions
> >    are to rely less on functions like BufferGetBlockNumber() or to split out
> >    the memory for that into a separate (denser?) array.
>
> Agreed.  BufferGetBlockNumber() could even use a new local (non-shmem) data
> structure, since the buffer's mapping can't change until we unpin.

Hm. I didn't think about a backend local datastructure for that, perhaps
because it seems not cheap to maintain (both from a runtime and a space
perspective).

If we store the read-only data for buffers separately from the read-write
data, we could access that from backends without a lock, since it can't change
with the buffer pinned.

One way to do that would be to maintain a back-pointer from the BufferDesc to
the BufferLookupEnt, since the latter *already* contains the BufferTag. We
probably don't want to add another indirection to the buffer mapping hash
table, otherwise we could deduplicate the other way round and just put padding
between the modified and read-only part of a buffer desc.



> On Tue, Aug 26, 2025 at 05:00:13PM -0400, Andres Freund wrote:
> > On 2025-08-26 16:21:36 -0400, Robert Haas wrote:
> > > On Fri, Aug 22, 2025 at 3:45 PM Andres Freund <andres@anarazel.de> wrote:
> > > > The order of changes I think makes the most sense is the following:
>
> No concerns so far.  I won't claim I can picture all the implications and be
> sure this is the right thing, but it sounds promising.  I like your principle
> of ordering changes to avoid performance regressions.

I suspect we'll have to merge this incrementally to stay sane, I don't want to
end up with a period of worse performance due to this, that could make it
harder to evaluate other work.


> > > > DOES ANYBODY HAVE A BETTER NAME THAN SHARE-EXCLUSIVE???!?

> I would consider {AccessShare, Exclusive, AccessExclusive}.

One thing I forgot to mention is that with the proposed re-architecture in
place, we could subsequently go further and make pinning just be a very
lightweight lock level, instead of that being a separate dedicated
infrstructure.  One nice outgrowth of that would be that that acquiring a
cleanup lock would just be a real lock acquisition, instead of the dedicated
limited machinery we have right now.

Which would leave us with:
- reference (pins today)
- share
- share-exclusive
- exclusive
- cleanup

This doesn't quite seem to map onto the heavyweight lock levels in a sensible
way...



> What the $SUBJECT proposal calls SHARE-EXCLUSIVE would become Exclusive.

There are a few hundred references to the lock levels though, seems painful to
rename them :(


> That has the same conflict matrix as the corresponding heavyweight locks,
> which seems good.

> I don't love our mode names, particularly ShareRowExclusive being
> unsharable.

I hate them with a passion :). Except for the most basic ones they just don't
stay in my head for more than a few hours.


> However, learning one special taxonomy is better than learning two.

But that's fair.


Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-26 20:21 ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  2025-08-26 21:00   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-27 00:14     ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2025-08-27 16:18       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-08-27 19:14         ` Noah Misch <noah@leadboat.com>
  2025-08-27 19:29           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  1 sibling, 1 reply; 120+ messages in thread

From: Noah Misch @ 2025-08-27 19:14 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Robert Haas <robertmhaas@gmail.com>; pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>

On Wed, Aug 27, 2025 at 12:18:27PM -0400, Andres Freund wrote:
> On 2025-08-26 17:14:49 -0700, Noah Misch wrote:
> > On Fri, Aug 22, 2025 at 03:44:48PM -0400, Andres Freund wrote:
> > > == Problem 3 - Cacheline contention ==
> >
> > > c) Read accesses to the BufferDesc cause contention
> > >
> > >    Some code, like nbtree, relies on functions like
> > >    BufferGetBlockNumber(). Unfortunately that contends with concurrent
> > >    modifications of the buffer descriptor (like pinning). Potential solutions
> > >    are to rely less on functions like BufferGetBlockNumber() or to split out
> > >    the memory for that into a separate (denser?) array.
> >
> > Agreed.  BufferGetBlockNumber() could even use a new local (non-shmem) data
> > structure, since the buffer's mapping can't change until we unpin.
> 
> Hm. I didn't think about a backend local datastructure for that, perhaps
> because it seems not cheap to maintain (both from a runtime and a space
> perspective).

Yes, paying off the cost of maintaining it could be tricky.  It could be the
kind of thing where the overhead loses at 10 cores and wins at 40 cores.  It
could also depend heavily on the workload's concurrent pins per backend.

> If we store the read-only data for buffers separately from the read-write
> data, we could access that from backends without a lock, since it can't change
> with the buffer pinned.

Good point.  That alone may be enough of a win.

> One way to do that would be to maintain a back-pointer from the BufferDesc to
> the BufferLookupEnt, since the latter *already* contains the BufferTag. We
> probably don't want to add another indirection to the buffer mapping hash
> table, otherwise we could deduplicate the other way round and just put padding
> between the modified and read-only part of a buffer desc.

I think you're saying clients would save the back-pointer once and dereference
it many times, with each dereference of a saved back-pointer avoiding a shmem
read of BufferDesc.tag.  Is that right?

> > On Tue, Aug 26, 2025 at 05:00:13PM -0400, Andres Freund wrote:
> > > On 2025-08-26 16:21:36 -0400, Robert Haas wrote:
> > > > On Fri, Aug 22, 2025 at 3:45 PM Andres Freund <andres@anarazel.de> wrote:
> > > > > DOES ANYBODY HAVE A BETTER NAME THAN SHARE-EXCLUSIVE???!?
> 
> > I would consider {AccessShare, Exclusive, AccessExclusive}.
> 
> One thing I forgot to mention is that with the proposed re-architecture in
> place, we could subsequently go further and make pinning just be a very
> lightweight lock level, instead of that being a separate dedicated
> infrstructure.  One nice outgrowth of that would be that that acquiring a
> cleanup lock would just be a real lock acquisition, instead of the dedicated
> limited machinery we have right now.
> 
> Which would leave us with:
> - reference (pins today)
> - share
> - share-exclusive
> - exclusive
> - cleanup
> 
> This doesn't quite seem to map onto the heavyweight lock levels in a sensible
> way...

Could map it like this:

AccessShare - pins today
RowShare - check tuple visibility (BUFFER_LOCK_SHARE today)
Share - set hint bits
ShareUpdateExclusive - clean/write out (borrowing Robert's idea)
Exclusive - add tuples, change xmax, etc. (BUFFER_LOCK_EXCLUSIVE today)
AccessExclusive - cleanup lock or evict the buffer

That has a separate level for hint bits vs. I/O, so multiple backends could
set hint bits.  I don't know whether the benchmarks would favor maintaining
that distinction.

> > What the $SUBJECT proposal calls SHARE-EXCLUSIVE would become Exclusive.
> 
> There are a few hundred references to the lock levels though, seems painful to
> rename them :(

Yes, especially in comments and extensions.  Likely more important than that
for the long-term, your latest proposal has the advantage of keeping short
names for the most-commonly-referenced lock types.  (We could keep
BUFFER_LOCK_SHARE with the lower layers translating that into RowShare, but
that weakens or eliminates the benefit of reducing what readers need to
learn.)  For what it's worth, 6 PGXN modules reference BUFFER_LOCK_SHARE
and/or BUFFER_LOCK_EXCLUSIVE.

Compared to share-exclusive, I think I'd prefer a name that describes the use
cases, "set-hints-or-write" (or separate "write" and "set-hints" levels).
What do you think of that?  I don't know whether that should win vs. names
like ShareUpdateExclusive, though.





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-26 20:21 ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  2025-08-26 21:00   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-27 00:14     ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2025-08-27 16:18       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-27 19:14         ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
@ 2025-08-27 19:29           ` Andres Freund <andres@anarazel.de>
  2025-08-27 23:04             ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-08-27 19:29 UTC (permalink / raw)
  To: Noah Misch <noah@leadboat.com>; +Cc: Robert Haas <robertmhaas@gmail.com>; pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>

Hi,

On 2025-08-27 12:14:41 -0700, Noah Misch wrote:
> On Wed, Aug 27, 2025 at 12:18:27PM -0400, Andres Freund wrote:
> > One way to do that would be to maintain a back-pointer from the BufferDesc to
> > the BufferLookupEnt, since the latter *already* contains the BufferTag. We
> > probably don't want to add another indirection to the buffer mapping hash
> > table, otherwise we could deduplicate the other way round and just put padding
> > between the modified and read-only part of a buffer desc.
> 
> I think you're saying clients would save the back-pointer once and dereference
> it many times, with each dereference of a saved back-pointer avoiding a shmem
> read of BufferDesc.tag.  Is that right?

I was thinking that we'd not have BufferDesc.tag, instead just storing it
solely in BufferLookupEnt. To get the tag of a BufferDesc, you'd every time
have to follow the back-reference.   But that's actually why it doesn't work -
reading the back-reference pointer would have the same issue as just reading
BufferDesc.tag...


> > > On Tue, Aug 26, 2025 at 05:00:13PM -0400, Andres Freund wrote:
> > > > On 2025-08-26 16:21:36 -0400, Robert Haas wrote:
> > > > > On Fri, Aug 22, 2025 at 3:45 PM Andres Freund <andres@anarazel.de> wrote:
> > > > > > DOES ANYBODY HAVE A BETTER NAME THAN SHARE-EXCLUSIVE???!?
> > 
> > > I would consider {AccessShare, Exclusive, AccessExclusive}.
> > 
> > One thing I forgot to mention is that with the proposed re-architecture in
> > place, we could subsequently go further and make pinning just be a very
> > lightweight lock level, instead of that being a separate dedicated
> > infrstructure.  One nice outgrowth of that would be that that acquiring a
> > cleanup lock would just be a real lock acquisition, instead of the dedicated
> > limited machinery we have right now.
> > 
> > Which would leave us with:
> > - reference (pins today)
> > - share
> > - share-exclusive
> > - exclusive
> > - cleanup
> > 
> > This doesn't quite seem to map onto the heavyweight lock levels in a sensible
> > way...
> 
> Could map it like this:
> 
> AccessShare - pins today
> RowShare - check tuple visibility (BUFFER_LOCK_SHARE today)
> Share - set hint bits
> ShareUpdateExclusive - clean/write out (borrowing Robert's idea)
> Exclusive - add tuples, change xmax, etc. (BUFFER_LOCK_EXCLUSIVE today)
> AccessExclusive - cleanup lock or evict the buffer

I tend think having things like RowShare for buffer locking is confusing
enough to actually make the similarity to the heavyweight locks to not be a
win...


> That has a separate level for hint bits vs. I/O, so multiple backends could
> set hint bits.  I don't know whether the benchmarks would favor maintaining
> that distinction.

I don't think it would - I actually found multiple backends setting the same
hint bits to *hurt* performance a bit.  But what's more important, we don't
have the space for it, I think.  Every lock that can be acquired multiple
times needs a lock count of 18 bits. And we need to store the buffer state
flags (10 bits). There's just not enough space in 64bit to have three 18bit
counters as well as flag bits etc.


> Compared to share-exclusive, I think I'd prefer a name that describes the use
> cases, "set-hints-or-write" (or separate "write" and "set-hints" levels).

I would too, I just couldn't come up with something that conveys the meanings
in a sufficiently concise way :)


> What do you think of that?  I don't know whether that should win vs. names
> like ShareUpdateExclusive, though.

I think it'd be a win compared to the heavyweight lock names...

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-26 20:21 ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  2025-08-26 21:00   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-27 00:14     ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2025-08-27 16:18       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-27 19:14         ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2025-08-27 19:29           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-08-27 23:04             ` Noah Misch <noah@leadboat.com>
  0 siblings, 0 replies; 120+ messages in thread

From: Noah Misch @ 2025-08-27 23:04 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Robert Haas <robertmhaas@gmail.com>; pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>

On Wed, Aug 27, 2025 at 03:29:02PM -0400, Andres Freund wrote:
> On 2025-08-27 12:14:41 -0700, Noah Misch wrote:
> > On Wed, Aug 27, 2025 at 12:18:27PM -0400, Andres Freund wrote:
> > > > On Tue, Aug 26, 2025 at 05:00:13PM -0400, Andres Freund wrote:
> > > > > On 2025-08-26 16:21:36 -0400, Robert Haas wrote:
> > > > > > On Fri, Aug 22, 2025 at 3:45 PM Andres Freund <andres@anarazel.de> wrote:
> > > > > > > DOES ANYBODY HAVE A BETTER NAME THAN SHARE-EXCLUSIVE???!?

> > > Which would leave us with:
> > > - reference (pins today)
> > > - share
> > > - share-exclusive
> > > - exclusive
> > > - cleanup

> > Compared to share-exclusive, I think I'd prefer a name that describes the use
> > cases, "set-hints-or-write" (or separate "write" and "set-hints" levels).

Another name idea is "self-exclusive", to contrast with "exclusive" excluding
all of (exclusive, self-exclusive, share).

Fortunately, not much code will acquire this lock type.  Hence, there's
relatively little damage if the name is less obvious than older lock types or
if the name changes later.





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-26 20:21 ` Re: Buffer locking is special (hints, checksums, AIO writes) Robert Haas <robertmhaas@gmail.com>
  2025-08-26 21:00   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-08-27 00:14     ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2025-08-27 16:18       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-08-27 19:22         ` Robert Haas <robertmhaas@gmail.com>
  1 sibling, 0 replies; 120+ messages in thread

From: Robert Haas @ 2025-08-27 19:22 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Noah Misch <noah@leadboat.com>; pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>

On Wed, Aug 27, 2025 at 12:18 PM Andres Freund <andres@anarazel.de> wrote:
> Which would leave us with:
> - reference (pins today)
> - share
> - share-exclusive
> - exclusive
> - cleanup
>
> This doesn't quite seem to map onto the heavyweight lock levels in a sensible
> way...

Could do: ACCESS SHARE, SHARE, SHARE UPDATE EXCLUSIVE, EXCLUSIVE,
ACCESS EXCLUSIVE.

I've always thought that a pin was a lot like an access share lock and
a cleanup lock was a lot like an access exclusive lock.

But then again, using the same terminology for two different things
might be confusing.

-- 
Robert Haas
EDB: http://www.enterprisedb.com





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-09-15 23:05 ` Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-09-15 23:05 UTC (permalink / raw)
  To: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-08-22 15:44:48 -0400, Andres Freund wrote:
> The hardest part about this change is that everything kind of depends on each
> other. The changes are large enough that they clearly can't just be committed
> at once, but doing them over time risks [temporary] performance regressions.
> 
> 
> 
> 
> The order of changes I think makes the most sense is the following:
> 
> 1) Allow some modifications while holding the buffer header spinlock
> 
> 2) Reduce buffer pin with just an atomic-sub
> 
>    This needs to happen first, otherwise there are performance regressions
>    during the later steps.

Here are the first few cleaned up patches implementing the above steps, as
well as some cleanups.  I included a commit from another thread, as it
conflicts with these changes, and we really should apply it - and it's
arguably required to make the changes viable, as it removes one more use of
PinBuffer_Locked().

Another change included is to not return the buffer with the spinlock held
from StrategyGetBuffer(), and instead pin the buffer in freelist.c. The reason
for that is to reduce the most common PinBuffer_locked() call. By definition
PinBuffer_locked() will become a bit slower due to 0003. But even without 0003
it 0002 is faster than master. And the previous approach also just seems
pretty unclean.   I don't love that it requires the new TrackNewBufferPin(),
but I don't really have a better idea.

I invite particular attention to the commit message for 0003 as well as the
comment changes in buf_internals.h within.

I'm not sure I like the TODOs added in 0003 and removed in 0004, but squashing
the changes doesn't really seem better either.

Greetings,

Andres Freund

Attachments:

  [text/x-diff] v3-0001-Improve-ReadRecentBuffer-scalability.patch (6.1K, ../../yivb2evcrj7fna5ymuunw3g5u5xxttwjbjxaa4ofkfkviystjv@4dfylftqxyxh/2-v3-0001-Improve-ReadRecentBuffer-scalability.patch)
  download | inline diff:
From 2267a08b45f1f040182376e9fc4372ab7ee740ad Mon Sep 17 00:00:00 2001
From: Thomas Munro <thomas.munro@gmail.com>
Date: Thu, 29 Jun 2023 10:52:56 +1200
Subject: [PATCH v3 1/6] Improve ReadRecentBuffer() scalability.

While testing a new potential use for ReadRecentBuffer(), Andres
reported that it scales badly when called concurrently for the same
buffer by many backends.  Instead of a naive (but wrong) coding with
PinBuffer(), it used the spinlock, so that it could be careful to pin
only if the buffer was valid and holding the expected block, to avoid
breaking invariants in eg GetVictimBuffer().  Unfortunately that made it
less scalable than PinBuffer(), which uses compare-exchange instead.

We can fix that by giving PinBuffer() a new skip_if_not_valid mode that
doesn't pin invalid buffers.  It might occasionally skip when it
shouldn't due to the unlocked read of the header flags, but that's
unlikely and perfectly acceptable for an opportunistic optimisation
routine, and it can only succeed when it really should due to the
compare-exchange loop.

XXX This also fixes ReadRecentBuffer()'s failure to bump the usage
count.  Fix separately or back-patch this?

Author: Thomas Munro <thomas.munro@gmail.com>
Reported-by: Andres Freund <andres@anarazel.de>
Reviewed-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/20230627020546.t6z4tntmj7wmjrfh%40awork3.anarazel.de
---
 src/backend/storage/buffer/bufmgr.c | 60 +++++++++++++----------------
 1 file changed, 26 insertions(+), 34 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index fe470de63f2..b5eebfb6990 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -512,7 +512,8 @@ static BlockNumber ExtendBufferedRelShared(BufferManagerRelation bmr,
 										   BlockNumber extend_upto,
 										   Buffer *buffers,
 										   uint32 *extended_by);
-static bool PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy);
+static bool PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
+					  bool skip_if_not_valid);
 static void PinBuffer_Locked(BufferDesc *buf);
 static void UnpinBuffer(BufferDesc *buf);
 static void UnpinBufferNoOwner(BufferDesc *buf);
@@ -685,7 +686,6 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 	BufferDesc *bufHdr;
 	BufferTag	tag;
 	uint32		buf_state;
-	bool		have_private_ref;
 
 	Assert(BufferIsValid(recent_buffer));
 
@@ -713,38 +713,24 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 	else
 	{
 		bufHdr = GetBufferDescriptor(recent_buffer - 1);
-		have_private_ref = GetPrivateRefCount(recent_buffer) > 0;
 
 		/*
-		 * Do we already have this buffer pinned with a private reference?  If
-		 * so, it must be valid and it is safe to check the tag without
-		 * locking.  If not, we have to lock the header first and then check.
+		 * Is it still valid and holding the right tag?  We do an unlocked tag
+		 * comparison first, to make it unlikely that we'll increment the
+		 * usage counter of the wrong buffer, if someone calls us with a very
+		 * out of date recent_buffer.  Then we'll check it again if we get the
+		 * pin.
 		 */
-		if (have_private_ref)
-			buf_state = pg_atomic_read_u32(&bufHdr->state);
-		else
-			buf_state = LockBufHdr(bufHdr);
-
-		if ((buf_state & BM_VALID) && BufferTagsEqual(&tag, &bufHdr->tag))
+		if (BufferTagsEqual(&tag, &bufHdr->tag) &&
+			PinBuffer(bufHdr, NULL, true))
 		{
-			/*
-			 * It's now safe to pin the buffer.  We can't pin first and ask
-			 * questions later, because it might confuse code paths like
-			 * InvalidateBuffer() if we pinned a random non-matching buffer.
-			 */
-			if (have_private_ref)
-				PinBuffer(bufHdr, NULL);	/* bump pin count */
-			else
-				PinBuffer_Locked(bufHdr);	/* pin for first time */
-
-			pgBufferUsage.shared_blks_hit++;
-
-			return true;
+			if (BufferTagsEqual(&tag, &bufHdr->tag))
+			{
+				pgBufferUsage.shared_blks_hit++;
+				return true;
+			}
+			UnpinBuffer(bufHdr);
 		}
-
-		/* If we locked the header above, now unlock. */
-		if (!have_private_ref)
-			UnlockBufHdr(bufHdr, buf_state);
 	}
 
 	return false;
@@ -2036,7 +2022,7 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 		 */
 		buf = GetBufferDescriptor(existing_buf_id);
 
-		valid = PinBuffer(buf, strategy);
+		valid = PinBuffer(buf, strategy, false);
 
 		/* Can release the mapping lock as soon as we've pinned it */
 		LWLockRelease(newPartitionLock);
@@ -2098,7 +2084,7 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 
 		existing_buf_hdr = GetBufferDescriptor(existing_buf_id);
 
-		valid = PinBuffer(existing_buf_hdr, strategy);
+		valid = PinBuffer(existing_buf_hdr, strategy, false);
 
 		/* Can release the mapping lock as soon as we've pinned it */
 		LWLockRelease(newPartitionLock);
@@ -2736,7 +2722,7 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 * Pin the existing buffer before releasing the partition lock,
 			 * preventing it from being evicted.
 			 */
-			valid = PinBuffer(existing_hdr, strategy);
+			valid = PinBuffer(existing_hdr, strategy, false);
 
 			LWLockRelease(partition_lock);
 			UnpinBuffer(victim_buf_hdr);
@@ -3035,10 +3021,13 @@ ReleaseAndReadBuffer(Buffer buffer,
  * must have been done already.
  *
  * Returns true if buffer is BM_VALID, else false.  This provision allows
- * some callers to avoid an extra spinlock cycle.
+ * some callers to avoid an extra spinlock cycle.  If skip_if_not_valid is
+ * true, then a false return value also indicates that the buffer was
+ * (recently) invalid and has not been pinned.
  */
 static bool
-PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy)
+PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
+		  bool skip_if_not_valid)
 {
 	Buffer		b = BufferDescriptorGetBuffer(buf);
 	bool		result;
@@ -3062,6 +3051,9 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy)
 			if (old_buf_state & BM_LOCKED)
 				old_buf_state = WaitBufHdrUnlocked(buf);
 
+			if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
+				return false;
+
 			buf_state = old_buf_state;
 
 			/* increase refcount */
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v3-0002-bufmgr-Don-t-lock-buffer-header-in-StrategyGetBuf.patch (10.0K, ../../yivb2evcrj7fna5ymuunw3g5u5xxttwjbjxaa4ofkfkviystjv@4dfylftqxyxh/3-v3-0002-bufmgr-Don-t-lock-buffer-header-in-StrategyGetBuf.patch)
  download | inline diff:
From 482f56744a72c168492eefac2af9dcf5ee153524 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 8 Sep 2025 17:16:01 -0400
Subject: [PATCH v3 2/6] bufmgr: Don't lock buffer header in
 StrategyGetBuffer()

This is a small improvement on its own, due to not needing to hold the buffer
spinlock as often. But the main reason for this is to reduce the use of
PinBuffer_Locked() and LockBufHdr(), which a future commit will make a bit
more expensive (to make more common paths faster).

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/buf_internals.h   |   4 +
 src/backend/storage/buffer/bufmgr.c   |  44 +++++------
 src/backend/storage/buffer/freelist.c | 110 +++++++++++++++++++-------
 3 files changed, 105 insertions(+), 53 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index dfd614f7ca4..c1206a46aba 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -371,6 +371,8 @@ UnlockBufHdr(BufferDesc *desc, uint32 buf_state)
 	pg_atomic_write_u32(&desc->state, buf_state & (~BM_LOCKED));
 }
 
+extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
+
 /* in bufmgr.c */
 
 /*
@@ -425,6 +427,8 @@ extern void IssuePendingWritebacks(WritebackContext *wb_context, IOContext io_co
 extern void ScheduleBufferTagForWriteback(WritebackContext *wb_context,
 										  IOContext io_context, BufferTag *tag);
 
+extern void TrackNewBufferPin(Buffer buf);
+
 /* solely to make it easier to write tests */
 extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
 extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index b5eebfb6990..0b841eb7838 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -518,7 +518,6 @@ static void PinBuffer_Locked(BufferDesc *buf);
 static void UnpinBuffer(BufferDesc *buf);
 static void UnpinBufferNoOwner(BufferDesc *buf);
 static void BufferSync(int flags);
-static uint32 WaitBufHdrUnlocked(BufferDesc *buf);
 static int	SyncOneBuffer(int buf_id, bool skip_recently_used,
 						  WritebackContext *wb_context);
 static void WaitIO(BufferDesc *buf);
@@ -2325,7 +2324,7 @@ GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context)
 
 	/*
 	 * Ensure, while the spinlock's not yet held, that there's a free refcount
-	 * entry, and a resource owner slot for the pin.
+	 * entry, and a resource owner slot for the pin. FIXME: Need to be updated
 	 */
 	ReservePrivateRefCountEntry();
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2334,17 +2333,11 @@ GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context)
 again:
 
 	/*
-	 * Select a victim buffer.  The buffer is returned with its header
-	 * spinlock still held!
+	 * Select a victim buffer.  The buffer is returned pinned by this backend.
 	 */
 	buf_hdr = StrategyGetBuffer(strategy, &buf_state, &from_ring);
 	buf = BufferDescriptorGetBuffer(buf_hdr);
 
-	Assert(BUF_STATE_GET_REFCOUNT(buf_state) == 0);
-
-	/* Pin the buffer and then release the buffer spinlock */
-	PinBuffer_Locked(buf_hdr);
-
 	/*
 	 * We shouldn't have any other pins for this buffer.
 	 */
@@ -3043,8 +3036,6 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		uint32		buf_state;
 		uint32		old_buf_state;
 
-		ref = NewPrivateRefCountEntry(b);
-
 		old_buf_state = pg_atomic_read_u32(&buf->state);
 		for (;;)
 		{
@@ -3091,6 +3082,8 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 				break;
 			}
 		}
+
+		TrackNewBufferPin(b);
 	}
 	else
 	{
@@ -3110,11 +3103,12 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		 * cannot meddle with that.
 		 */
 		result = (pg_atomic_read_u32(&buf->state) & BM_VALID) != 0;
+
+		Assert(ref->refcount > 0);
+		ref->refcount++;
+		ResourceOwnerRememberBuffer(CurrentResourceOwner, b);
 	}
 
-	ref->refcount++;
-	Assert(ref->refcount > 0);
-	ResourceOwnerRememberBuffer(CurrentResourceOwner, b);
 	return result;
 }
 
@@ -3143,8 +3137,6 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	Buffer		b;
-	PrivateRefCountEntry *ref;
 	uint32		buf_state;
 
 	/*
@@ -3169,12 +3161,7 @@ PinBuffer_Locked(BufferDesc *buf)
 	buf_state += BUF_REFCOUNT_ONE;
 	UnlockBufHdr(buf, buf_state);
 
-	b = BufferDescriptorGetBuffer(buf);
-
-	ref = NewPrivateRefCountEntry(b);
-	ref->refcount++;
-
-	ResourceOwnerRememberBuffer(CurrentResourceOwner, b);
+	TrackNewBufferPin(BufferDescriptorGetBuffer(buf));
 }
 
 /*
@@ -3289,6 +3276,17 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	}
 }
 
+inline void
+TrackNewBufferPin(Buffer buf)
+{
+	PrivateRefCountEntry *ref;
+
+	ref = NewPrivateRefCountEntry(buf);
+	ref->refcount++;
+
+	ResourceOwnerRememberBuffer(CurrentResourceOwner, buf);
+}
+
 #define ST_SORT sort_checkpoint_bufferids
 #define ST_ELEMENT_TYPE CkptSortItem
 #define ST_COMPARE(a, b) ckpt_buforder_comparator(a, b)
@@ -6242,7 +6240,7 @@ LockBufHdr(BufferDesc *desc)
  * Obviously the buffer could be locked by the time the value is returned, so
  * this is primarily useful in CAS style loops.
  */
-static uint32
+pg_noinline uint32
 WaitBufHdrUnlocked(BufferDesc *buf)
 {
 	SpinDelayStatus delayStatus;
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 7d59a92bd1a..9d9fb0471a0 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -164,8 +164,8 @@ ClockSweepTick(void)
  *
  *	strategy is a BufferAccessStrategy object, or NULL for default strategy.
  *
- *	To ensure that no one else can pin the buffer before we do, we must
- *	return the buffer with the buffer header spinlock still held.
+ *	The buffer is pinned before returning, but resource ownership is the
+ *	responsibility of the caller.
  */
 BufferDesc *
 StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_ring)
@@ -173,7 +173,6 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 	BufferDesc *buf;
 	int			bgwprocno;
 	int			trycounter;
-	uint32		local_buf_state;	/* to avoid repeated (de-)referencing */
 
 	*from_ring = false;
 
@@ -228,44 +227,74 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 	trycounter = NBuffers;
 	for (;;)
 	{
+		uint32		old_buf_state;
+		uint32		local_buf_state;
+
 		buf = GetBufferDescriptor(ClockSweepTick());
 
 		/*
 		 * If the buffer is pinned or has a nonzero usage_count, we cannot use
 		 * it; decrement the usage_count (unless pinned) and keep scanning.
 		 */
-		local_buf_state = LockBufHdr(buf);
 
-		if (BUF_STATE_GET_REFCOUNT(local_buf_state) == 0)
+		old_buf_state = pg_atomic_read_u32(&buf->state);
+
+		for (;;)
 		{
+			local_buf_state = old_buf_state;
+
+			if (BUF_STATE_GET_REFCOUNT(local_buf_state) != 0)
+			{
+				if (--trycounter == 0)
+				{
+					/*
+					 * We've scanned all the buffers without making any state
+					 * changes, so all the buffers are pinned (or were when we
+					 * looked at them). We could hope that someone will free
+					 * one eventually, but it's probably better to fail than
+					 * to risk getting stuck in an infinite loop.
+					 */
+					elog(ERROR, "no unpinned buffers available");
+				}
+				break;
+			}
+
+			if (unlikely(local_buf_state & BM_LOCKED))
+			{
+				old_buf_state = WaitBufHdrUnlocked(buf);
+				continue;
+			}
+
 			if (BUF_STATE_GET_USAGECOUNT(local_buf_state) != 0)
 			{
 				local_buf_state -= BUF_USAGECOUNT_ONE;
 
-				trycounter = NBuffers;
+				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+												   local_buf_state))
+				{
+					trycounter = NBuffers;
+					break;
+				}
 			}
 			else
 			{
-				/* Found a usable buffer */
-				if (strategy != NULL)
-					AddBufferToRing(strategy, buf);
-				*buf_state = local_buf_state;
-				return buf;
+				local_buf_state += BUF_REFCOUNT_ONE;
+
+				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+												   local_buf_state))
+				{
+					/* Found a usable buffer */
+					if (strategy != NULL)
+						AddBufferToRing(strategy, buf);
+					*buf_state = local_buf_state;
+
+					TrackNewBufferPin(BufferDescriptorGetBuffer(buf));
+
+					return buf;
+				}
 			}
+
 		}
-		else if (--trycounter == 0)
-		{
-			/*
-			 * We've scanned all the buffers without making any state changes,
-			 * so all the buffers are pinned (or were when we looked at them).
-			 * We could hope that someone will free one eventually, but it's
-			 * probably better to fail than to risk getting stuck in an
-			 * infinite loop.
-			 */
-			UnlockBufHdr(buf, local_buf_state);
-			elog(ERROR, "no unpinned buffers available");
-		}
-		UnlockBufHdr(buf, local_buf_state);
 	}
 }
 
@@ -621,6 +650,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 {
 	BufferDesc *buf;
 	Buffer		bufnum;
+	uint32		old_buf_state;
 	uint32		local_buf_state;	/* to avoid repeated (de-)referencing */
 
 
@@ -647,14 +677,34 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 	 * shouldn't re-use it.
 	 */
 	buf = GetBufferDescriptor(bufnum - 1);
-	local_buf_state = LockBufHdr(buf);
-	if (BUF_STATE_GET_REFCOUNT(local_buf_state) == 0
-		&& BUF_STATE_GET_USAGECOUNT(local_buf_state) <= 1)
+
+	old_buf_state = pg_atomic_read_u32(&buf->state);
+
+	for (;;)
 	{
-		*buf_state = local_buf_state;
-		return buf;
+		local_buf_state = old_buf_state;
+
+		if (BUF_STATE_GET_REFCOUNT(local_buf_state) != 0
+			|| BUF_STATE_GET_USAGECOUNT(local_buf_state) > 1)
+			break;
+
+		if (unlikely(local_buf_state & BM_LOCKED))
+		{
+			old_buf_state = WaitBufHdrUnlocked(buf);
+			continue;
+		}
+
+		local_buf_state += BUF_REFCOUNT_ONE;
+
+		if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+										   local_buf_state))
+		{
+			*buf_state = local_buf_state;
+
+			TrackNewBufferPin(BufferDescriptorGetBuffer(buf));
+			return buf;
+		}
 	}
-	UnlockBufHdr(buf, local_buf_state);
 
 	/*
 	 * Tell caller to allocate a new buffer with the normal allocation
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v3-0003-bufmgr-Allow-some-buffer-state-modifications-whil.patch (33.5K, ../../yivb2evcrj7fna5ymuunw3g5u5xxttwjbjxaa4ofkfkviystjv@4dfylftqxyxh/4-v3-0003-bufmgr-Allow-some-buffer-state-modifications-whil.patch)
  download | inline diff:
From ef9590e8928fcf4882efc6c03c844b3de156d064 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 15 Sep 2025 17:06:48 -0400
Subject: [PATCH v3 3/6] bufmgr: Allow some buffer state modifications while
 holding header lock

Until now BufferDesc.state was not allowed to be modified while the buffer
header spinlock was held. This meant that operations like unpinning buffers
needed to use a CAS loop, waiting for the buffer header spinlock to be
released before updating.

The benefit of that restriction is that it allowed us to unlock the buffer
header spinlock with just a write barrier and an unlocked write (instead of a
full atomic operation). That was important to avoid regressions in
48354581a49c. However, since then the hottest buffer header spinlock uses have
been replaced with atomic operations (in particular, the most common use of
PinBuffer_Locked(), in BufferAlloc(), has been removed in FIXME-TODO-FIXME).

This change will allow, in a subsequent commit, to release buffer pins with a
single atomic-sub operation. This previously was not possible while such
operations were not allowed while the buffer header spinlock was held, as an
atomic-sub would not have allowed a race-free check for the buffer header lock
being held.

Using atomic-sub to unpin buffers is a nice scalability win, however it is not
the primary motivation for this change (although it would be sufficient). The
primary motivation is that we would like to merge the buffer content lock into
BufferDesc.state, which will result in more frequent changes of the state
variable, which in some situations can cause a performance regression, due to
an increased CAS failure rate when unpinning buffers.  The regression entirely
vanishes when using atomic-sub.

Naively implementing this would require putting CAS loops in every place
modifying the buffer state while holding the buffer header lock. To avoid
that, introduce UnlockBufHdrExt(), which can set/add flags as well as the
refcount, together with releasing the lock.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/buf_internals.h           | 101 ++++--
 src/backend/storage/buffer/bufmgr.c           | 298 ++++++++++--------
 src/backend/storage/buffer/freelist.c         |   2 +
 contrib/pg_buffercache/pg_buffercache_pages.c |   7 +-
 contrib/pg_prewarm/autoprewarm.c              |   2 +-
 src/test/modules/test_aio/test_aio.c          |  10 +-
 6 files changed, 262 insertions(+), 158 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index c1206a46aba..e0d7fe56395 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -211,28 +211,29 @@ BufMappingPartitionLockByIndex(uint32 index)
 /*
  *	BufferDesc -- shared descriptor/state data for a single shared buffer.
  *
- * Note: Buffer header lock (BM_LOCKED flag) must be held to examine or change
- * tag, state or wait_backend_pgprocno fields.  In general, buffer header lock
- * is a spinlock which is combined with flags, refcount and usagecount into
- * single atomic variable.  This layout allow us to do some operations in a
- * single atomic operation, without actually acquiring and releasing spinlock;
- * for instance, increase or decrease refcount.  buf_id field never changes
- * after initialization, so does not need locking.  The LWLock can take care
- * of itself.  The buffer header lock is *not* used to control access to the
- * data in the buffer!
+ * The state of the buffer is controlled by the, drumroll, state variable. It
+ * only may be modified using atomic operations.  The state variable combines
+ * various flags, the buffer's refcount and usage count. See comment above
+ * BUF_REFCOUNT_BITS for details about the division.  This layout allow us to
+ * do some operations in a single atomic operation, without actually acquiring
+ * and releasing the spinlock; for instance, increasing or decreasing the
+ * refcount.
  *
- * It's assumed that nobody changes the state field while buffer header lock
- * is held.  Thus buffer header lock holder can do complex updates of the
- * state variable in single write, simultaneously with lock release (cleaning
- * BM_LOCKED flag).  On the other hand, updating of state without holding
- * buffer header lock is restricted to CAS, which ensures that BM_LOCKED flag
- * is not set.  Atomic increment/decrement, OR/AND etc. are not allowed.
+ * One of the aforementioned flags is BM_LOCKED, used to implement the buffer
+ * header lock. While the buffer header lock is held, the identity of the
+ * buffer cannot change (BufferDesc.tag) and no additional buffer pins may get
+ * established (we would like to relax the latter eventually). However,
+ * existing buffer references can be released while the buffer header spinlock
+ * is held.
  *
- * An exception is that if we have the buffer pinned, its tag can't change
- * underneath us, so we can examine the tag without locking the buffer header.
- * Also, in places we do one-time reads of the flags without bothering to
- * lock the buffer header; this is generally for situations where we don't
- * expect the flag bit being tested to be changing.
+ * The LWLock can take care of itself.  The buffer header lock is *not* used
+ * to control access to the data in the buffer!
+ *
+ * If we have the buffer pinned, its tag can't change underneath us, so we can
+ * examine the tag without locking the buffer header.  Also, in places we do
+ * one-time reads of the flags without bothering to lock the buffer header;
+ * this is generally for situations where we don't expect the flag bit being
+ * tested to be changing.
  *
  * We can't physically remove items from a disk page if another backend has
  * the buffer pinned.  Hence, a backend may need to wait for all other pins
@@ -257,12 +258,21 @@ BufMappingPartitionLockByIndex(uint32 index)
 typedef struct BufferDesc
 {
 	BufferTag	tag;			/* ID of page contained in buffer */
-	int			buf_id;			/* buffer's index number (from 0) */
+
+	/*
+	 * Buffer's index number (from 0). The field never changes after
+	 * initialization, so does not need locking.
+	 */
+	int			buf_id;
 
 	/* state of the tag, containing flags, refcount and usagecount */
 	pg_atomic_uint32 state;
 
-	int			wait_backend_pgprocno;	/* backend of pin-count waiter */
+	/*
+	 * Backend of pin-count waiter. The buffer header spinlock needs to be
+	 * held to modify this field.
+	 */
+	int			wait_backend_pgprocno;
 
 	PgAioWaitRef io_wref;		/* set iff AIO is in progress */
 	LWLock		content_lock;	/* to lock access to buffer contents */
@@ -364,11 +374,52 @@ BufferDescriptorGetContentLock(const BufferDesc *bdesc)
  */
 extern uint32 LockBufHdr(BufferDesc *desc);
 
+/*
+ * Unlock the buffer header.
+ *
+ * This can only be used if the caller did not modify BufferDesc.state. To
+ * set/unset flag bits or change the refcount use UnlockBufHdrExt().
+ */
 static inline void
-UnlockBufHdr(BufferDesc *desc, uint32 buf_state)
+UnlockBufHdr(BufferDesc *desc)
 {
-	pg_write_barrier();
-	pg_atomic_write_u32(&desc->state, buf_state & (~BM_LOCKED));
+	Assert(pg_atomic_read_u32(&desc->state) & BM_LOCKED);
+
+	pg_atomic_fetch_sub_u32(&desc->state, BM_LOCKED);
+}
+
+/*
+ * Unlock the buffer header, while atomically adding the flags in set_bits,
+ * unsetting the ones in unset_bits and changing the refcount by
+ * refcount_change.
+ *
+ * Note that this approach would not work for usagecount, since we need to cap
+ * the usagecount at BM_MAX_USAGE_COUNT.
+ */
+static inline uint32
+UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
+				uint32 set_bits, uint32 unset_bits,
+				int refcount_change)
+{
+	for (;;)
+	{
+		uint32		buf_state = old_buf_state;
+
+		Assert(buf_state & BM_LOCKED);
+
+		buf_state |= set_bits;
+		buf_state &= ~unset_bits;
+		buf_state &= ~BM_LOCKED;
+
+		if (refcount_change != 0)
+			buf_state += BUF_REFCOUNT_ONE * refcount_change;
+
+		if (pg_atomic_compare_exchange_u32(&desc->state, &old_buf_state,
+										   buf_state))
+		{
+			return old_buf_state;
+		}
+	}
 }
 
 extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 0b841eb7838..c498402e85b 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -1994,6 +1994,7 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	Buffer		victim_buffer;
 	BufferDesc *victim_buf_hdr;
 	uint32		victim_buf_state;
+	uint32		set_bits = 0;
 
 	/* Make sure we will have room to remember the buffer pin */
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2120,11 +2121,12 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	 * checkpoints, except for their "init" forks, which need to be treated
 	 * just like permanent relations.
 	 */
-	victim_buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
+	set_bits |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 	if (relpersistence == RELPERSISTENCE_PERMANENT || forkNum == INIT_FORKNUM)
-		victim_buf_state |= BM_PERMANENT;
+		set_bits |= BM_PERMANENT;
 
-	UnlockBufHdr(victim_buf_hdr, victim_buf_state);
+	UnlockBufHdrExt(victim_buf_hdr, victim_buf_state,
+					set_bits, 0, 0);
 
 	LWLockRelease(newPartitionLock);
 
@@ -2164,9 +2166,7 @@ InvalidateBuffer(BufferDesc *buf)
 	/* Save the original buffer tag before dropping the spinlock */
 	oldTag = buf->tag;
 
-	buf_state = pg_atomic_read_u32(&buf->state);
-	Assert(buf_state & BM_LOCKED);
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdr(buf);
 
 	/*
 	 * Need to compute the old tag's hashcode and partition lock ID. XXX is it
@@ -2190,7 +2190,7 @@ retry:
 	/* If it's changed while we were waiting for lock, do nothing */
 	if (!BufferTagsEqual(&buf->tag, &oldTag))
 	{
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		LWLockRelease(oldPartitionLock);
 		return;
 	}
@@ -2207,7 +2207,7 @@ retry:
 	 */
 	if (BUF_STATE_GET_REFCOUNT(buf_state) != 0)
 	{
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		LWLockRelease(oldPartitionLock);
 		/* safety check: should definitely not be our *own* pin */
 		if (GetPrivateRefCount(BufferDescriptorGetBuffer(buf)) > 0)
@@ -2222,8 +2222,11 @@ retry:
 	 */
 	oldFlags = buf_state & BUF_FLAG_MASK;
 	ClearBufferTag(&buf->tag);
-	buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
-	UnlockBufHdr(buf, buf_state);
+
+	UnlockBufHdrExt(buf, buf_state,
+					0,
+					BUF_FLAG_MASK | BUF_USAGECOUNT_MASK,
+					0);
 
 	/*
 	 * Remove the buffer from the lookup hashtable, if it was in there.
@@ -2283,7 +2286,7 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 	{
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 
-		UnlockBufHdr(buf_hdr, buf_state);
+		UnlockBufHdr(buf_hdr);
 		LWLockRelease(partition_lock);
 
 		return false;
@@ -2297,8 +2300,10 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 	 * tag (see e.g. FlushDatabaseBuffers()).
 	 */
 	ClearBufferTag(&buf_hdr->tag);
-	buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
-	UnlockBufHdr(buf_hdr, buf_state);
+	UnlockBufHdrExt(buf_hdr, buf_state,
+					0,
+					BUF_FLAG_MASK | BUF_USAGECOUNT_MASK,
+					0);
 
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 
@@ -2307,6 +2312,7 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 
 	LWLockRelease(partition_lock);
 
+	buf_state = pg_atomic_read_u32(&buf_hdr->state);
 	Assert(!(buf_state & (BM_DIRTY | BM_VALID | BM_TAG_VALID)));
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u32(&buf_hdr->state)) > 0);
@@ -2396,7 +2402,7 @@ again:
 			/* Read the LSN while holding buffer header lock */
 			buf_state = LockBufHdr(buf_hdr);
 			lsn = BufferGetLSN(buf_hdr);
-			UnlockBufHdr(buf_hdr, buf_state);
+			UnlockBufHdr(buf_hdr);
 
 			if (XLogNeedsFlush(lsn)
 				&& StrategyRejectBuffer(strategy, buf_hdr, from_ring))
@@ -2741,15 +2747,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 */
 			do
 			{
-				uint32		buf_state = LockBufHdr(existing_hdr);
-
-				buf_state &= ~BM_VALID;
-				UnlockBufHdr(existing_hdr, buf_state);
+				pg_atomic_fetch_and_u32(&existing_hdr->state, ~BM_VALID);
 			} while (!StartBufferIO(existing_hdr, true, false));
 		}
 		else
 		{
 			uint32		buf_state;
+			uint32		set_bits = 0;
 
 			buf_state = LockBufHdr(victim_buf_hdr);
 
@@ -2759,11 +2763,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 
 			victim_buf_hdr->tag = tag;
 
-			buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
+			set_bits |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 			if (bmr.relpersistence == RELPERSISTENCE_PERMANENT || fork == INIT_FORKNUM)
-				buf_state |= BM_PERMANENT;
+				set_bits |= BM_PERMANENT;
 
-			UnlockBufHdr(victim_buf_hdr, buf_state);
+			UnlockBufHdrExt(victim_buf_hdr, buf_state,
+							set_bits, 0,
+							0);
 
 			LWLockRelease(partition_lock);
 
@@ -2918,6 +2924,11 @@ MarkBufferDirty(Buffer buffer)
 	Assert(LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
 								LW_EXCLUSIVE));
 
+	/*
+	 * TODO: A future commit will replace this loop with a single atomic
+	 * operation, we do not need to wait for the buffer header spinlock to be
+	 * released anymore.
+	 */
 	old_buf_state = pg_atomic_read_u32(&bufHdr->state);
 	for (;;)
 	{
@@ -3039,6 +3050,10 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		old_buf_state = pg_atomic_read_u32(&buf->state);
 		for (;;)
 		{
+			/*
+			 * We're not allowed to increase the refcount while the buffer
+			 * header spinlock is held. Wait for the lock to be released.
+			 */
 			if (old_buf_state & BM_LOCKED)
 				old_buf_state = WaitBufHdrUnlocked(buf);
 
@@ -3137,7 +3152,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	uint32		buf_state;
+	uint32		old_buf_state;
 
 	/*
 	 * As explained, We don't expect any preexisting pins. That allows us to
@@ -3156,10 +3171,10 @@ PinBuffer_Locked(BufferDesc *buf)
 	 * Since we hold the buffer spinlock, we can update the buffer state and
 	 * release the lock in one operation.
 	 */
-	buf_state = pg_atomic_read_u32(&buf->state);
-	Assert(buf_state & BM_LOCKED);
-	buf_state += BUF_REFCOUNT_ONE;
-	UnlockBufHdr(buf, buf_state);
+	old_buf_state = pg_atomic_read_u32(&buf->state);
+
+	UnlockBufHdrExt(buf, old_buf_state,
+					0, 0, 1);
 
 	TrackNewBufferPin(BufferDescriptorGetBuffer(buf));
 }
@@ -3194,12 +3209,13 @@ WakePinCountWaiter(BufferDesc *buf)
 		/* we just released the last pin other than the waiter's */
 		int			wait_backend_pgprocno = buf->wait_backend_pgprocno;
 
-		buf_state &= ~BM_PIN_COUNT_WAITER;
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdrExt(buf, buf_state,
+						0, BM_PIN_COUNT_WAITER,
+						0);
 		ProcSendSignal(wait_backend_pgprocno);
 	}
 	else
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 }
 
 /*
@@ -3252,6 +3268,10 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 *
 		 * Since buffer spinlock holder can update status using just write,
 		 * it's not safe to use atomic decrement here; thus use a CAS loop.
+		 *
+		 * TODO: The above requirement does not hold anymore, in a future
+		 * commit this will be rewritten to release the pin in a single atomic
+		 * operation.
 		 */
 		old_buf_state = pg_atomic_read_u32(&buf->state);
 		for (;;)
@@ -3349,6 +3369,7 @@ BufferSync(int flags)
 	for (buf_id = 0; buf_id < NBuffers; buf_id++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
+		uint32		set_bits = 0;
 
 		/*
 		 * Header spinlock is enough to examine BM_DIRTY, see comment in
@@ -3360,7 +3381,7 @@ BufferSync(int flags)
 		{
 			CkptSortItem *item;
 
-			buf_state |= BM_CHECKPOINT_NEEDED;
+			set_bits = BM_CHECKPOINT_NEEDED;
 
 			item = &CkptBufferIds[num_to_scan++];
 			item->buf_id = buf_id;
@@ -3370,7 +3391,9 @@ BufferSync(int flags)
 			item->blockNum = bufHdr->tag.blockNum;
 		}
 
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdrExt(bufHdr, buf_state,
+						set_bits, 0,
+						0);
 
 		/* Check for barrier events in case NBuffers is large. */
 		if (ProcSignalBarrierPending)
@@ -3909,14 +3932,14 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	else if (skip_recently_used)
 	{
 		/* Caller told us not to write recently-used buffers */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return result;
 	}
 
 	if (!(buf_state & BM_VALID) || !(buf_state & BM_DIRTY))
 	{
 		/* It's clean, so nothing to do */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return result;
 	}
 
@@ -4288,8 +4311,9 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	recptr = BufferGetLSN(buf);
 
 	/* To check if block content changes while flushing. - vadim 01/17/97 */
-	buf_state &= ~BM_JUST_DIRTIED;
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdrExt(buf, buf_state,
+					0, BM_JUST_DIRTIED,
+					0);
 
 	/*
 	 * Force XLOG flush up to buffer's LSN.  This implements the basic WAL
@@ -4452,7 +4476,6 @@ BufferGetLSNAtomic(Buffer buffer)
 	char	   *page = BufferGetPage(buffer);
 	BufferDesc *bufHdr;
 	XLogRecPtr	lsn;
-	uint32		buf_state;
 
 	/*
 	 * If we don't need locking for correctness, fastpath out.
@@ -4465,9 +4488,9 @@ BufferGetLSNAtomic(Buffer buffer)
 	Assert(BufferIsPinned(buffer));
 
 	bufHdr = GetBufferDescriptor(buffer - 1);
-	buf_state = LockBufHdr(bufHdr);
+	LockBufHdr(bufHdr);
 	lsn = PageGetLSN(page);
-	UnlockBufHdr(bufHdr, buf_state);
+	UnlockBufHdr(bufHdr);
 
 	return lsn;
 }
@@ -4568,7 +4591,6 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
 	for (i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
 
 		/*
 		 * We can make this a tad faster by prechecking the buffer tag before
@@ -4589,7 +4611,7 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
 		if (!BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator.locator))
 			continue;
 
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 
 		for (j = 0; j < nforks; j++)
 		{
@@ -4602,7 +4624,7 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
 			}
 		}
 		if (j >= nforks)
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -4731,7 +4753,6 @@ DropRelationsAllBuffers(SMgrRelation *smgr_reln, int nlocators)
 	{
 		RelFileLocator *rlocator = NULL;
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -4765,11 +4786,11 @@ DropRelationsAllBuffers(SMgrRelation *smgr_reln, int nlocators)
 		if (rlocator == NULL)
 			continue;
 
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 		if (BufTagMatchesRelFileLocator(&bufHdr->tag, rlocator))
 			InvalidateBuffer(bufHdr);	/* releases spinlock */
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 
 	pfree(locators);
@@ -4799,7 +4820,6 @@ FindAndDropRelationBuffers(RelFileLocator rlocator, ForkNumber forkNum,
 		LWLock	   *bufPartitionLock;	/* buffer partition lock for it */
 		int			buf_id;
 		BufferDesc *bufHdr;
-		uint32		buf_state;
 
 		/* create a tag so we can lookup the buffer */
 		InitBufferTag(&bufTag, &rlocator, forkNum, curBlock);
@@ -4824,14 +4844,14 @@ FindAndDropRelationBuffers(RelFileLocator rlocator, ForkNumber forkNum,
 		 * evicted by some other backend loading blocks for a different
 		 * relation after we release lock on the BufMapping table.
 		 */
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 
 		if (BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator) &&
 			BufTagGetForkNum(&bufHdr->tag) == forkNum &&
 			bufHdr->tag.blockNum >= firstDelBlock)
 			InvalidateBuffer(bufHdr);	/* releases spinlock */
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -4859,7 +4879,6 @@ DropDatabaseBuffers(Oid dbid)
 	for (i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -4868,11 +4887,11 @@ DropDatabaseBuffers(Oid dbid)
 		if (bufHdr->tag.dbOid != dbid)
 			continue;
 
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 		if (bufHdr->tag.dbOid == dbid)
 			InvalidateBuffer(bufHdr);	/* releases spinlock */
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -4971,7 +4990,7 @@ FlushRelationBuffers(Relation rel)
 			UnpinBuffer(bufHdr);
 		}
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -5068,7 +5087,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 			UnpinBuffer(bufHdr);
 		}
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 
 	pfree(srels);
@@ -5296,7 +5315,7 @@ FlushDatabaseBuffers(Oid dbid)
 			UnpinBuffer(bufHdr);
 		}
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -5506,8 +5525,9 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 				PageSetLSN(page, lsn);
 		}
 
-		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdrExt(bufHdr, buf_state,
+						BM_DIRTY | BM_JUST_DIRTIED,
+						0, 0);
 
 		if (delayChkptFlags)
 			MyProc->delayChkptFlags &= ~DELAY_CHKPT_START;
@@ -5538,6 +5558,7 @@ UnlockBuffers(void)
 	if (buf)
 	{
 		uint32		buf_state;
+		uint32		unset_bits = 0;
 
 		buf_state = LockBufHdr(buf);
 
@@ -5547,9 +5568,11 @@ UnlockBuffers(void)
 		 */
 		if ((buf_state & BM_PIN_COUNT_WAITER) != 0 &&
 			buf->wait_backend_pgprocno == MyProcNumber)
-			buf_state &= ~BM_PIN_COUNT_WAITER;
+			unset_bits = BM_PIN_COUNT_WAITER;
 
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdrExt(buf, buf_state,
+						0, unset_bits,
+						0);
 
 		PinCountWaitBuf = NULL;
 	}
@@ -5667,6 +5690,7 @@ LockBufferForCleanup(Buffer buffer)
 	for (;;)
 	{
 		uint32		buf_state;
+		uint32		unset_bits = 0;
 
 		/* Try to acquire lock */
 		LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
@@ -5676,7 +5700,7 @@ LockBufferForCleanup(Buffer buffer)
 		if (BUF_STATE_GET_REFCOUNT(buf_state) == 1)
 		{
 			/* Successfully acquired exclusive lock with pincount 1 */
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 
 			/*
 			 * Emit the log message if recovery conflict on buffer pin was
@@ -5699,14 +5723,15 @@ LockBufferForCleanup(Buffer buffer)
 		/* Failed, so mark myself as waiting for pincount 1 */
 		if (buf_state & BM_PIN_COUNT_WAITER)
 		{
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 			LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 			elog(ERROR, "multiple backends attempting to wait for pincount 1");
 		}
 		bufHdr->wait_backend_pgprocno = MyProcNumber;
 		PinCountWaitBuf = bufHdr;
-		buf_state |= BM_PIN_COUNT_WAITER;
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdrExt(bufHdr, buf_state,
+						BM_PIN_COUNT_WAITER, 0,
+						0);
 		LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 
 		/* Wait to be signaled by UnpinBuffer() */
@@ -5768,8 +5793,11 @@ LockBufferForCleanup(Buffer buffer)
 		buf_state = LockBufHdr(bufHdr);
 		if ((buf_state & BM_PIN_COUNT_WAITER) != 0 &&
 			bufHdr->wait_backend_pgprocno == MyProcNumber)
-			buf_state &= ~BM_PIN_COUNT_WAITER;
-		UnlockBufHdr(bufHdr, buf_state);
+			unset_bits |= BM_PIN_COUNT_WAITER;
+
+		UnlockBufHdrExt(bufHdr, buf_state,
+						0, unset_bits,
+						0);
 
 		PinCountWaitBuf = NULL;
 		/* Loop back and try again */
@@ -5846,12 +5874,12 @@ ConditionalLockBufferForCleanup(Buffer buffer)
 	if (refcount == 1)
 	{
 		/* Successfully acquired exclusive lock with pincount 1 */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return true;
 	}
 
 	/* Failed, so release the lock */
-	UnlockBufHdr(bufHdr, buf_state);
+	UnlockBufHdr(bufHdr);
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 	return false;
 }
@@ -5899,11 +5927,11 @@ IsBufferCleanupOK(Buffer buffer)
 	if (BUF_STATE_GET_REFCOUNT(buf_state) == 1)
 	{
 		/* pincount is OK. */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return true;
 	}
 
-	UnlockBufHdr(bufHdr, buf_state);
+	UnlockBufHdr(bufHdr);
 	return false;
 }
 
@@ -5941,7 +5969,7 @@ WaitIO(BufferDesc *buf)
 		 * clearing the wref while it's being read.
 		 */
 		iow = buf->io_wref;
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 
 		/* no IO in progress, we don't need to wait */
 		if (!(buf_state & BM_IO_IN_PROGRESS))
@@ -6009,7 +6037,7 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 
 		if (!(buf_state & BM_IO_IN_PROGRESS))
 			break;
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		if (nowait)
 			return false;
 		WaitIO(buf);
@@ -6020,12 +6048,13 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 	/* Check if someone else already did the I/O */
 	if (forInput ? (buf_state & BM_VALID) : !(buf_state & BM_DIRTY))
 	{
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		return false;
 	}
 
-	buf_state |= BM_IO_IN_PROGRESS;
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdrExt(buf, buf_state,
+					BM_IO_IN_PROGRESS, 0,
+					0);
 
 	ResourceOwnerRememberBufferIO(CurrentResourceOwner,
 								  BufferDescriptorGetBuffer(buf));
@@ -6057,29 +6086,35 @@ void
 TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
 				  bool forget_owner, bool release_aio)
 {
-	uint32		buf_state;
+	uint32		old_buf_state;
+	uint32		unset_flag_bits = 0;
+	int			refcount_change = 0;
 
-	buf_state = LockBufHdr(buf);
+	old_buf_state = LockBufHdr(buf);
 
-	Assert(buf_state & BM_IO_IN_PROGRESS);
-	buf_state &= ~BM_IO_IN_PROGRESS;
+	if (release_aio)
+		pgaio_wref_clear(&buf->io_wref);
+
+	Assert(old_buf_state & BM_IO_IN_PROGRESS);
+
+	unset_flag_bits |= BM_IO_IN_PROGRESS;
 
 	/* Clear earlier errors, if this IO failed, it'll be marked again */
-	buf_state &= ~BM_IO_ERROR;
+	unset_flag_bits |= BM_IO_ERROR;
 
-	if (clear_dirty && !(buf_state & BM_JUST_DIRTIED))
-		buf_state &= ~(BM_DIRTY | BM_CHECKPOINT_NEEDED);
+	if (clear_dirty && !(old_buf_state & BM_JUST_DIRTIED))
+		unset_flag_bits |= BM_DIRTY | BM_CHECKPOINT_NEEDED;
 
 	if (release_aio)
 	{
 		/* release ownership by the AIO subsystem */
-		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-		buf_state -= BUF_REFCOUNT_ONE;
-		pgaio_wref_clear(&buf->io_wref);
+		Assert(BUF_STATE_GET_REFCOUNT(old_buf_state) > 0);
+		refcount_change = -1;
 	}
 
-	buf_state |= set_flag_bits;
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdrExt(buf, old_buf_state,
+					set_flag_bits, unset_flag_bits,
+					refcount_change);
 
 	if (forget_owner)
 		ResourceOwnerForgetBufferIO(CurrentResourceOwner,
@@ -6095,7 +6130,7 @@ TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
 	 * example, this backend is completing an IO issued by another backend, it
 	 * may be time to wake the waiter.
 	 */
-	if (release_aio && (buf_state & BM_PIN_COUNT_WAITER))
+	if (release_aio && (old_buf_state & BM_PIN_COUNT_WAITER))
 		WakePinCountWaiter(buf);
 }
 
@@ -6124,12 +6159,12 @@ AbortBufferIO(Buffer buffer)
 	if (!(buf_state & BM_VALID))
 	{
 		Assert(!(buf_state & BM_DIRTY));
-		UnlockBufHdr(buf_hdr, buf_state);
+		UnlockBufHdr(buf_hdr);
 	}
 	else
 	{
 		Assert(buf_state & BM_DIRTY);
-		UnlockBufHdr(buf_hdr, buf_state);
+		UnlockBufHdr(buf_hdr);
 
 		/* Issue notice if this is not the first failure... */
 		if (buf_state & BM_IO_ERROR)
@@ -6551,14 +6586,14 @@ EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 
 	if ((buf_state & BM_VALID) == 0)
 	{
-		UnlockBufHdr(desc, buf_state);
+		UnlockBufHdr(desc);
 		return false;
 	}
 
 	/* Check that it's not pinned already. */
 	if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
 	{
-		UnlockBufHdr(desc, buf_state);
+		UnlockBufHdr(desc);
 		return false;
 	}
 
@@ -6710,7 +6745,7 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
 		if ((buf_state & BM_VALID) == 0 ||
 			!BufTagMatchesRelFileLocator(&desc->tag, &rel->rd_locator))
 		{
-			UnlockBufHdr(desc, buf_state);
+			UnlockBufHdr(desc);
 			continue;
 		}
 
@@ -6757,7 +6792,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		BufferDesc *buf_hdr = is_temp ?
 			GetLocalBufferDescriptor(-buffer - 1)
 			: GetBufferDescriptor(buffer - 1);
-		uint32		buf_state;
+		uint32		old_buf_state;
 
 		/*
 		 * Check that all the buffers are actually ones that could conceivably
@@ -6774,47 +6809,58 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 			Assert(buf_hdr->tag.blockNum == first.blockNum + i);
 		}
 
-		if (is_temp)
-			buf_state = pg_atomic_read_u32(&buf_hdr->state);
-		else
-			buf_state = LockBufHdr(buf_hdr);
+		old_buf_state = pg_atomic_read_u32(&buf_hdr->state);
 
-		/* verify the buffer is in the expected state */
-		Assert(buf_state & BM_TAG_VALID);
-		if (is_write)
+		for (;;)
 		{
-			Assert(buf_state & BM_VALID);
-			Assert(buf_state & BM_DIRTY);
-		}
-		else
-		{
-			Assert(!(buf_state & BM_VALID));
-			Assert(!(buf_state & BM_DIRTY));
-		}
+			uint32		buf_state = old_buf_state;
 
-		/* temp buffers don't use BM_IO_IN_PROGRESS */
-		if (!is_temp)
-			Assert(buf_state & BM_IO_IN_PROGRESS);
+			/* verify the buffer is in the expected state */
+			Assert(buf_state & BM_TAG_VALID);
+			if (is_write)
+			{
+				Assert(buf_state & BM_VALID);
+				Assert(buf_state & BM_DIRTY);
+			}
+			else
+			{
+				Assert(!(buf_state & BM_VALID));
+				Assert(!(buf_state & BM_DIRTY));
+			}
 
-		Assert(BUF_STATE_GET_REFCOUNT(buf_state) >= 1);
+			/* temp buffers don't use BM_IO_IN_PROGRESS */
+			if (!is_temp)
+				Assert(buf_state & BM_IO_IN_PROGRESS);
 
-		/*
-		 * Reflect that the buffer is now owned by the AIO subsystem.
-		 *
-		 * For local buffers: This can't be done just via LocalRefCount, as
-		 * one might initially think, as this backend could error out while
-		 * AIO is still in progress, releasing all the pins by the backend
-		 * itself.
-		 *
-		 * This pin is released again in TerminateBufferIO().
-		 */
-		buf_state += BUF_REFCOUNT_ONE;
-		buf_hdr->io_wref = io_ref;
+			Assert(BUF_STATE_GET_REFCOUNT(buf_state) >= 1);
 
-		if (is_temp)
-			pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
-		else
-			UnlockBufHdr(buf_hdr, buf_state);
+			/*
+			 * Reflect that the buffer is now owned by the AIO subsystem.
+			 *
+			 * For local buffers: This can't be done just via LocalRefCount,
+			 * as one might initially think, as this backend could error out
+			 * while AIO is still in progress, releasing all the pins by the
+			 * backend itself.
+			 *
+			 * This pin is released again in TerminateBufferIO().
+			 */
+			buf_state += BUF_REFCOUNT_ONE;
+			buf_hdr->io_wref = io_ref;
+
+			if (is_temp)
+			{
+				pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+				break;
+			}
+			else
+			{
+				if (pg_atomic_compare_exchange_u32(&buf_hdr->state, &old_buf_state,
+												   buf_state))
+				{
+					break;
+				}
+			}
+		}
 
 		/*
 		 * Ensure the content lock that prevents buffer modifications while
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 9d9fb0471a0..c98080a7d0b 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -259,6 +259,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 				break;
 			}
 
+			/* See equivalent code in PinBuffer() */
 			if (unlikely(local_buf_state & BM_LOCKED))
 			{
 				old_buf_state = WaitBufHdrUnlocked(buf);
@@ -688,6 +689,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 			|| BUF_STATE_GET_USAGECOUNT(local_buf_state) > 1)
 			break;
 
+		/* See equivalent code in PinBuffer() */
 		if (unlikely(local_buf_state & BM_LOCKED))
 		{
 			old_buf_state = WaitBufHdrUnlocked(buf);
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 3df04c98959..ab790533ff6 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -220,7 +220,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
 			else
 				fctx->record[i].isvalid = false;
 
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 		}
 	}
 
@@ -460,7 +460,6 @@ pg_buffercache_numa_pages(PG_FUNCTION_ARGS)
 		{
 			char	   *buffptr = (char *) BufferGetBlock(i + 1);
 			BufferDesc *bufHdr;
-			uint32		buf_state;
 			uint32		bufferid;
 			int32		page_num;
 			char	   *startptr_buff,
@@ -471,9 +470,9 @@ pg_buffercache_numa_pages(PG_FUNCTION_ARGS)
 			bufHdr = GetBufferDescriptor(i);
 
 			/* Lock each buffer header before inspecting. */
-			buf_state = LockBufHdr(bufHdr);
+			LockBufHdr(bufHdr);
 			bufferid = BufferDescriptorGetBuffer(bufHdr);
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 
 			/* start of the first page of this buffer */
 			startptr_buff = (char *) TYPEALIGN_DOWN(os_page_size, buffptr);
diff --git a/contrib/pg_prewarm/autoprewarm.c b/contrib/pg_prewarm/autoprewarm.c
index 8b68dafc261..5ba1240d51f 100644
--- a/contrib/pg_prewarm/autoprewarm.c
+++ b/contrib/pg_prewarm/autoprewarm.c
@@ -730,7 +730,7 @@ apw_dump_now(bool is_bgworker, bool dump_unlogged)
 			++num_blocks;
 		}
 
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 	}
 
 	snprintf(transient_dump_file_path, MAXPGPATH, "%s.tmp", AUTOPREWARM_FILE);
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index c55cf6c0aac..d7eadeab256 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -310,6 +310,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	BufferDesc *buf_hdr;
 	uint32		buf_state;
 	bool		was_pinned = false;
+	uint32		unset_bits = 0;
 
 	/* place buffer in shared buffers without erroring out */
 	buf = ReadBufferExtended(rel, MAIN_FORKNUM, blkno, RBM_ZERO_AND_LOCK, NULL);
@@ -334,12 +335,17 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (BUF_STATE_GET_REFCOUNT(buf_state) > 1)
 		was_pinned = true;
 	else
-		buf_state &= ~(BM_VALID | BM_DIRTY);
+		unset_bits |= BM_VALID | BM_DIRTY;
 
 	if (RelationUsesLocalBuffers(rel))
+	{
+		buf_state &= ~unset_bits;
 		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+	}
 	else
-		UnlockBufHdr(buf_hdr, buf_state);
+	{
+		UnlockBufHdrExt(buf_hdr, buf_state, 0, unset_bits, 0);
+	}
 
 	if (was_pinned)
 		elog(ERROR, "toy buffer %d was already pinned",
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v3-0004-bufmgr-Use-atomic-sub-or-for-unpinning-and-markin.patch (3.4K, ../../yivb2evcrj7fna5ymuunw3g5u5xxttwjbjxaa4ofkfkviystjv@4dfylftqxyxh/5-v3-0004-bufmgr-Use-atomic-sub-or-for-unpinning-and-markin.patch)
  download | inline diff:
From 054647bcc53e57a5f30fecd9c7de1d8e74570417 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 15 Sep 2025 17:14:12 -0400
Subject: [PATCH v3 4/6] bufmgr: Use atomic sub/or for unpinning and marking
 buffers dirty

The prior commit made it legal to modify BufferDesc.state while the buffer
header spinlock is held. This allows us to replace a few CAS loops with atomic
sub/or.

Particularly the change in UnpinBufferNoOwner() improves scalability
significantly. See the prior commit for more background.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/storage/buffer/bufmgr.c | 53 ++++-------------------------
 1 file changed, 6 insertions(+), 47 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index c498402e85b..ec8fad2d72e 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2906,8 +2906,7 @@ void
 MarkBufferDirty(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
-	uint32		old_buf_state;
+	uint32		old_buf_state PG_USED_FOR_ASSERTS_ONLY;
 
 	if (!BufferIsValid(buffer))
 		elog(ERROR, "bad buffer ID: %d", buffer);
@@ -2924,26 +2923,9 @@ MarkBufferDirty(Buffer buffer)
 	Assert(LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
 								LW_EXCLUSIVE));
 
-	/*
-	 * TODO: A future commit will replace this loop with a single atomic
-	 * operation, we do not need to wait for the buffer header spinlock to be
-	 * released anymore.
-	 */
-	old_buf_state = pg_atomic_read_u32(&bufHdr->state);
-	for (;;)
-	{
-		if (old_buf_state & BM_LOCKED)
-			old_buf_state = WaitBufHdrUnlocked(bufHdr);
+	old_buf_state = pg_atomic_fetch_or_u32(&bufHdr->state, BM_DIRTY | BM_JUST_DIRTIED);
 
-		buf_state = old_buf_state;
-
-		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
-
-		if (pg_atomic_compare_exchange_u32(&bufHdr->state, &old_buf_state,
-										   buf_state))
-			break;
-	}
+	Assert(BUF_STATE_GET_REFCOUNT(old_buf_state) > 0);
 
 	/*
 	 * If the buffer was not dirty already, do vacuum accounting.
@@ -3248,7 +3230,6 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	ref->refcount--;
 	if (ref->refcount == 0)
 	{
-		uint32		buf_state;
 		uint32		old_buf_state;
 
 		/*
@@ -3263,33 +3244,11 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		/* I'd better not still hold the buffer content lock */
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
-		/*
-		 * Decrement the shared reference count.
-		 *
-		 * Since buffer spinlock holder can update status using just write,
-		 * it's not safe to use atomic decrement here; thus use a CAS loop.
-		 *
-		 * TODO: The above requirement does not hold anymore, in a future
-		 * commit this will be rewritten to release the pin in a single atomic
-		 * operation.
-		 */
-		old_buf_state = pg_atomic_read_u32(&buf->state);
-		for (;;)
-		{
-			if (old_buf_state & BM_LOCKED)
-				old_buf_state = WaitBufHdrUnlocked(buf);
-
-			buf_state = old_buf_state;
-
-			buf_state -= BUF_REFCOUNT_ONE;
-
-			if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
-											   buf_state))
-				break;
-		}
+		/* decrement the shared reference count */
+		old_buf_state = pg_atomic_fetch_sub_u32(&buf->state, BUF_REFCOUNT_ONE);
 
 		/* Support LockBufferForCleanup() */
-		if (buf_state & BM_PIN_COUNT_WAITER)
+		if (old_buf_state & BM_PIN_COUNT_WAITER)
 			WakePinCountWaiter(buf);
 
 		ForgetPrivateRefCountEntry(ref);
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v3-0005-bufmgr-fewer-calls-to-BufferDescriptorGetContentL.patch (7.2K, ../../yivb2evcrj7fna5ymuunw3g5u5xxttwjbjxaa4ofkfkviystjv@4dfylftqxyxh/6-v3-0005-bufmgr-fewer-calls-to-BufferDescriptorGetContentL.patch)
  download | inline diff:
From 87e1624829b02d7f0d52960e23fe0d28c3af6bc2 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 30 Jun 2025 16:54:39 -0400
Subject: [PATCH v3 5/6] bufmgr: fewer calls to BufferDescriptorGetContentLock

We're planning to merge buffer content locks into BufferDesc.state. To reduce
the size of that patch, centralize BufferDescriptorGetContentLock().

The biggest part of the change is in assertions, by introducing
BufferIsLockedByMe[InMode]() (and removing BufferIsExclusiveLocked()). This
seems like an improvement even without aforementioned plans.

Additionally replace some direct calls to LWLockAcquire() with calls to
LockBuffer().

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/bufmgr.h            |  3 +-
 src/backend/access/heap/visibilitymap.c |  3 +-
 src/backend/access/transam/xloginsert.c |  3 +-
 src/backend/storage/buffer/bufmgr.c     | 70 +++++++++++++++++++------
 4 files changed, 61 insertions(+), 18 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 41fdc1e7693..7ca17cd9f6c 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -230,7 +230,8 @@ extern void WaitReadBuffers(ReadBuffersOperation *operation);
 
 extern void ReleaseBuffer(Buffer buffer);
 extern void UnlockReleaseBuffer(Buffer buffer);
-extern bool BufferIsExclusiveLocked(Buffer buffer);
+extern bool BufferIsLockedByMe(Buffer buffer);
+extern bool BufferIsLockedByMeInMode(Buffer buffer, int mode);
 extern bool BufferIsDirty(Buffer buffer);
 extern void MarkBufferDirty(Buffer buffer);
 extern void IncrBufferRefCount(Buffer buffer);
diff --git a/src/backend/access/heap/visibilitymap.c b/src/backend/access/heap/visibilitymap.c
index 7306c16f05c..0414ce1945c 100644
--- a/src/backend/access/heap/visibilitymap.c
+++ b/src/backend/access/heap/visibilitymap.c
@@ -270,7 +270,8 @@ visibilitymap_set(Relation rel, BlockNumber heapBlk, Buffer heapBuf,
 	if (BufferIsValid(heapBuf) && BufferGetBlockNumber(heapBuf) != heapBlk)
 		elog(ERROR, "wrong heap buffer passed to visibilitymap_set");
 
-	Assert(!BufferIsValid(heapBuf) || BufferIsExclusiveLocked(heapBuf));
+	Assert(!BufferIsValid(heapBuf) ||
+		   BufferIsLockedByMeInMode(heapBuf, BUFFER_LOCK_EXCLUSIVE));
 
 	/* Check that we have the right VM page pinned */
 	if (!BufferIsValid(vmBuf) || BufferGetBlockNumber(vmBuf) != mapBlock)
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index c7571429e8e..496e0fa4ac6 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -258,7 +258,8 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	 */
 #ifdef USE_ASSERT_CHECKING
 	if (!(flags & REGBUF_NO_CHANGE))
-		Assert(BufferIsExclusiveLocked(buffer) && BufferIsDirty(buffer));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) &&
+			   BufferIsDirty(buffer));
 #endif
 
 	if (block_id >= max_registered_block_id)
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index ec8fad2d72e..ba6c9bbf7c2 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -1065,7 +1065,7 @@ ZeroAndLockBuffer(Buffer buffer, ReadBufferMode mode, bool already_valid)
 		 * already valid.)
 		 */
 		if (!isLocalBuf)
-			LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_EXCLUSIVE);
+			LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
 
 		/* Set BM_VALID, terminate IO, and wake up any waiters */
 		if (isLocalBuf)
@@ -2822,7 +2822,7 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 		}
 
 		if (lock)
-			LWLockAcquire(BufferDescriptorGetContentLock(buf_hdr), LW_EXCLUSIVE);
+			LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
 		TerminateBufferIO(buf_hdr, false, BM_VALID, true, false);
 	}
@@ -2835,14 +2835,14 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 }
 
 /*
- * BufferIsExclusiveLocked
+ * BufferIsLockedByMe
  *
- *      Checks if buffer is exclusive-locked.
+ *      Checks if this backend has the buffer locked in any mode.
  *
  * Buffer must be pinned.
  */
 bool
-BufferIsExclusiveLocked(Buffer buffer)
+BufferIsLockedByMe(Buffer buffer)
 {
 	BufferDesc *bufHdr;
 
@@ -2855,9 +2855,49 @@ BufferIsExclusiveLocked(Buffer buffer)
 	}
 	else
 	{
+		bufHdr = GetBufferDescriptor(buffer - 1);
+		return LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr));
+	}
+}
+
+/*
+ * BufferIsLockedByMeInMode
+ *
+ *      Checks if this backend has the buffer locked in the specified mode.
+ *
+ * Buffer must be pinned.
+ */
+bool
+BufferIsLockedByMeInMode(Buffer buffer, int mode)
+{
+	BufferDesc *bufHdr;
+
+	Assert(BufferIsPinned(buffer));
+
+	if (BufferIsLocal(buffer))
+	{
+		/* Content locks are not maintained for local buffers. */
+		return true;
+	}
+	else
+	{
+		LWLockMode	lw_mode;
+
+		switch (mode)
+		{
+			case BUFFER_LOCK_EXCLUSIVE:
+				lw_mode = LW_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE:
+				lw_mode = LW_SHARED;
+				break;
+			default:
+				pg_unreachable();
+		}
+
 		bufHdr = GetBufferDescriptor(buffer - 1);
 		return LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-									LW_EXCLUSIVE);
+									lw_mode);
 	}
 }
 
@@ -2886,8 +2926,7 @@ BufferIsDirty(Buffer buffer)
 	else
 	{
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		Assert(LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-									LW_EXCLUSIVE));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 	}
 
 	return pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY;
@@ -2920,8 +2959,7 @@ MarkBufferDirty(Buffer buffer)
 	bufHdr = GetBufferDescriptor(buffer - 1);
 
 	Assert(BufferIsPinned(buffer));
-	Assert(LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-								LW_EXCLUSIVE));
+	Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 
 	old_buf_state = pg_atomic_fetch_or_u32(&bufHdr->state, BM_DIRTY | BM_JUST_DIRTIED);
 
@@ -3241,7 +3279,10 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 */
 		VALGRIND_MAKE_MEM_NOACCESS(BufHdrGetBlock(buf), BLCKSZ);
 
-		/* I'd better not still hold the buffer content lock */
+		/*
+		 * I'd better not still hold the buffer content lock. Can't use
+		 * BufferIsLockedByMe(), as that asserts the buffer is pinned.
+		 */
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
 		/* decrement the shared reference count */
@@ -5294,7 +5335,7 @@ FlushOneBuffer(Buffer buffer)
 
 	bufHdr = GetBufferDescriptor(buffer - 1);
 
-	Assert(LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr)));
+	Assert(BufferIsLockedByMe(buffer));
 
 	FlushBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 }
@@ -5385,7 +5426,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 
 	Assert(GetPrivateRefCount(buffer) > 0);
 	/* here, either share or exclusive lock is OK */
-	Assert(LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr)));
+	Assert(BufferIsLockedByMe(buffer));
 
 	/*
 	 * This routine might get called many times on the same page, if we are
@@ -5877,8 +5918,7 @@ IsBufferCleanupOK(Buffer buffer)
 	bufHdr = GetBufferDescriptor(buffer - 1);
 
 	/* caller must hold exclusive lock on buffer */
-	Assert(LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-								LW_EXCLUSIVE));
+	Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 
 	buf_state = LockBufHdr(bufHdr);
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v3-0006-bufmgr-Introduce-FlushUnlockedBuffer.patch (4.5K, ../../yivb2evcrj7fna5ymuunw3g5u5xxttwjbjxaa4ofkfkviystjv@4dfylftqxyxh/7-v3-0006-bufmgr-Introduce-FlushUnlockedBuffer.patch)
  download | inline diff:
From 9ad42cd32c3dd98c1d90f1001ce8ffce6ba96616 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 30 Jun 2025 14:30:38 -0400
Subject: [PATCH v3 6/6] bufmgr: Introduce FlushUnlockedBuffer

There were several copies of code locking a buffer, flushing its contents, and
unlocking the buffer. It seems worth centralizing that into a helper function.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/storage/buffer/bufmgr.c | 36 ++++++++++++++++-------------
 1 file changed, 20 insertions(+), 16 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index ba6c9bbf7c2..4f70cf62a9f 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -533,6 +533,8 @@ static inline BufferDesc *BufferAlloc(SMgrRelation smgr,
 static bool AsyncReadBuffers(ReadBuffersOperation *operation, int *nblocks_progress);
 static void CheckReadBuffersOperation(ReadBuffersOperation *operation, bool is_complete);
 static Buffer GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context);
+static void FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
+								IOObject io_object, IOContext io_context);
 static void FlushBuffer(BufferDesc *buf, SMgrRelation reln,
 						IOObject io_object, IOContext io_context);
 static void FindAndDropRelationBuffers(RelFileLocator rlocator,
@@ -3948,11 +3950,8 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	 * buffer is clean by the time we've locked it.)
 	 */
 	PinBuffer_Locked(bufHdr);
-	LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
 
-	FlushBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-
-	LWLockRelease(BufferDescriptorGetContentLock(bufHdr));
+	FlushUnlockedBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 
 	tag = bufHdr->tag;
 
@@ -4400,6 +4399,19 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	error_context_stack = errcallback.previous;
 }
 
+/*
+ * Convenience wrapper around FlushBuffer() that locks/unlocks the buffer
+ * before/after calling FlushBuffer().
+ */
+static void
+FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
+					IOObject io_object, IOContext io_context)
+{
+	LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
+	LWLockRelease(BufferDescriptorGetContentLock(buf));
+}
+
 /*
  * RelationGetNumberOfBlocksInFork
  *		Determines the current number of pages in the specified relation fork.
@@ -4984,9 +4996,7 @@ FlushRelationBuffers(Relation rel)
 			(buf_state & (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 		{
 			PinBuffer_Locked(bufHdr);
-			LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
-			FlushBuffer(bufHdr, srel, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-			LWLockRelease(BufferDescriptorGetContentLock(bufHdr));
+			FlushUnlockedBuffer(bufHdr, srel, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 			UnpinBuffer(bufHdr);
 		}
 		else
@@ -5081,9 +5091,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 			(buf_state & (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 		{
 			PinBuffer_Locked(bufHdr);
-			LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
-			FlushBuffer(bufHdr, srelent->srel, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-			LWLockRelease(BufferDescriptorGetContentLock(bufHdr));
+			FlushUnlockedBuffer(bufHdr, srelent->srel, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 			UnpinBuffer(bufHdr);
 		}
 		else
@@ -5309,9 +5317,7 @@ FlushDatabaseBuffers(Oid dbid)
 			(buf_state & (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 		{
 			PinBuffer_Locked(bufHdr);
-			LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
-			FlushBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-			LWLockRelease(BufferDescriptorGetContentLock(bufHdr));
+			FlushUnlockedBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 			UnpinBuffer(bufHdr);
 		}
 		else
@@ -6601,10 +6607,8 @@ EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 	/* If it was dirty, try to clean it once. */
 	if (buf_state & BM_DIRTY)
 	{
-		LWLockAcquire(BufferDescriptorGetContentLock(desc), LW_SHARED);
-		FlushBuffer(desc, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
+		FlushUnlockedBuffer(desc, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 		*buffer_flushed = true;
-		LWLockRelease(BufferDescriptorGetContentLock(desc));
 	}
 
 	/* This will return false if it becomes dirty or someone else pins it. */
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-09-22 22:14   ` Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-09-22 22:14 UTC (permalink / raw)
  To: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-09-15 19:05:37 -0400, Andres Freund wrote:
> On 2025-08-22 15:44:48 -0400, Andres Freund wrote:
> > The hardest part about this change is that everything kind of depends on each
> > other. The changes are large enough that they clearly can't just be committed
> > at once, but doing them over time risks [temporary] performance regressions.
> >
> >
> >
> >
> > The order of changes I think makes the most sense is the following:
> >
> > 1) Allow some modifications while holding the buffer header spinlock
> >
> > 2) Reduce buffer pin with just an atomic-sub
> >
> >    This needs to happen first, otherwise there are performance regressions
> >    during the later steps.
>
> Here are the first few cleaned up patches implementing the above steps, as
> well as some cleanups.  I included a commit from another thread, as it
> conflicts with these changes, and we really should apply it - and it's
> arguably required to make the changes viable, as it removes one more use of
> PinBuffer_Locked().
>
> Another change included is to not return the buffer with the spinlock held
> from StrategyGetBuffer(), and instead pin the buffer in freelist.c. The reason
> for that is to reduce the most common PinBuffer_locked() call. By definition
> PinBuffer_locked() will become a bit slower due to 0003. But even without 0003
> it 0002 is faster than master. And the previous approach also just seems
> pretty unclean.   I don't love that it requires the new TrackNewBufferPin(),
> but I don't really have a better idea.
>
> I invite particular attention to the commit message for 0003 as well as the
> comment changes in buf_internals.h within.

Robert looked at the patches while we were chatting, and I addressed his
feedback in this new version.

Changes:

- Updated patch description for 0002, giving a lot more background

- Improved BufferDesc comments a fair bit more in 0003

- Reduced the size of 0003 a bit, by using UnlockBufHdrExt() instead of a CAS
  loop in buffer_stage_common() and reordering some things in
  TerminateBufferIO()

- Made 0004 only use the non-looping atomic op in UnpinBufferNoOwner(), not
  MarkBufferDirty(). I realized the latter would take additional complexity to
  make safe (a CAS loop in TerminateBufferIO()).  I am somewhat doubtful that
  there are workloads where it matters...


Greetings,

Andres Freund

Attachments:

  [text/x-diff] v4-0001-Improve-ReadRecentBuffer-scalability.patch (6.1K, ../../6kmid26do57ykqfpvq6iieniy4djsymhrypkjccazq5g4bbe6a@2y6owwv7qpex/2-v4-0001-Improve-ReadRecentBuffer-scalability.patch)
  download | inline diff:
From 5829480f2ee93a8ee438756ebc79d54e82e3dc88 Mon Sep 17 00:00:00 2001
From: Thomas Munro <thomas.munro@gmail.com>
Date: Thu, 29 Jun 2023 10:52:56 +1200
Subject: [PATCH v4 1/6] Improve ReadRecentBuffer() scalability.

While testing a new potential use for ReadRecentBuffer(), Andres
reported that it scales badly when called concurrently for the same
buffer by many backends.  Instead of a naive (but wrong) coding with
PinBuffer(), it used the spinlock, so that it could be careful to pin
only if the buffer was valid and holding the expected block, to avoid
breaking invariants in eg GetVictimBuffer().  Unfortunately that made it
less scalable than PinBuffer(), which uses compare-exchange instead.

We can fix that by giving PinBuffer() a new skip_if_not_valid mode that
doesn't pin invalid buffers.  It might occasionally skip when it
shouldn't due to the unlocked read of the header flags, but that's
unlikely and perfectly acceptable for an opportunistic optimisation
routine, and it can only succeed when it really should due to the
compare-exchange loop.

XXX This also fixes ReadRecentBuffer()'s failure to bump the usage
count.  Fix separately or back-patch this?

Author: Thomas Munro <thomas.munro@gmail.com>
Reported-by: Andres Freund <andres@anarazel.de>
Reviewed-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/20230627020546.t6z4tntmj7wmjrfh%40awork3.anarazel.de
---
 src/backend/storage/buffer/bufmgr.c | 60 +++++++++++++----------------
 1 file changed, 26 insertions(+), 34 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index fe470de63f2..b5eebfb6990 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -512,7 +512,8 @@ static BlockNumber ExtendBufferedRelShared(BufferManagerRelation bmr,
 										   BlockNumber extend_upto,
 										   Buffer *buffers,
 										   uint32 *extended_by);
-static bool PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy);
+static bool PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
+					  bool skip_if_not_valid);
 static void PinBuffer_Locked(BufferDesc *buf);
 static void UnpinBuffer(BufferDesc *buf);
 static void UnpinBufferNoOwner(BufferDesc *buf);
@@ -685,7 +686,6 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 	BufferDesc *bufHdr;
 	BufferTag	tag;
 	uint32		buf_state;
-	bool		have_private_ref;
 
 	Assert(BufferIsValid(recent_buffer));
 
@@ -713,38 +713,24 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 	else
 	{
 		bufHdr = GetBufferDescriptor(recent_buffer - 1);
-		have_private_ref = GetPrivateRefCount(recent_buffer) > 0;
 
 		/*
-		 * Do we already have this buffer pinned with a private reference?  If
-		 * so, it must be valid and it is safe to check the tag without
-		 * locking.  If not, we have to lock the header first and then check.
+		 * Is it still valid and holding the right tag?  We do an unlocked tag
+		 * comparison first, to make it unlikely that we'll increment the
+		 * usage counter of the wrong buffer, if someone calls us with a very
+		 * out of date recent_buffer.  Then we'll check it again if we get the
+		 * pin.
 		 */
-		if (have_private_ref)
-			buf_state = pg_atomic_read_u32(&bufHdr->state);
-		else
-			buf_state = LockBufHdr(bufHdr);
-
-		if ((buf_state & BM_VALID) && BufferTagsEqual(&tag, &bufHdr->tag))
+		if (BufferTagsEqual(&tag, &bufHdr->tag) &&
+			PinBuffer(bufHdr, NULL, true))
 		{
-			/*
-			 * It's now safe to pin the buffer.  We can't pin first and ask
-			 * questions later, because it might confuse code paths like
-			 * InvalidateBuffer() if we pinned a random non-matching buffer.
-			 */
-			if (have_private_ref)
-				PinBuffer(bufHdr, NULL);	/* bump pin count */
-			else
-				PinBuffer_Locked(bufHdr);	/* pin for first time */
-
-			pgBufferUsage.shared_blks_hit++;
-
-			return true;
+			if (BufferTagsEqual(&tag, &bufHdr->tag))
+			{
+				pgBufferUsage.shared_blks_hit++;
+				return true;
+			}
+			UnpinBuffer(bufHdr);
 		}
-
-		/* If we locked the header above, now unlock. */
-		if (!have_private_ref)
-			UnlockBufHdr(bufHdr, buf_state);
 	}
 
 	return false;
@@ -2036,7 +2022,7 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 		 */
 		buf = GetBufferDescriptor(existing_buf_id);
 
-		valid = PinBuffer(buf, strategy);
+		valid = PinBuffer(buf, strategy, false);
 
 		/* Can release the mapping lock as soon as we've pinned it */
 		LWLockRelease(newPartitionLock);
@@ -2098,7 +2084,7 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 
 		existing_buf_hdr = GetBufferDescriptor(existing_buf_id);
 
-		valid = PinBuffer(existing_buf_hdr, strategy);
+		valid = PinBuffer(existing_buf_hdr, strategy, false);
 
 		/* Can release the mapping lock as soon as we've pinned it */
 		LWLockRelease(newPartitionLock);
@@ -2736,7 +2722,7 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 * Pin the existing buffer before releasing the partition lock,
 			 * preventing it from being evicted.
 			 */
-			valid = PinBuffer(existing_hdr, strategy);
+			valid = PinBuffer(existing_hdr, strategy, false);
 
 			LWLockRelease(partition_lock);
 			UnpinBuffer(victim_buf_hdr);
@@ -3035,10 +3021,13 @@ ReleaseAndReadBuffer(Buffer buffer,
  * must have been done already.
  *
  * Returns true if buffer is BM_VALID, else false.  This provision allows
- * some callers to avoid an extra spinlock cycle.
+ * some callers to avoid an extra spinlock cycle.  If skip_if_not_valid is
+ * true, then a false return value also indicates that the buffer was
+ * (recently) invalid and has not been pinned.
  */
 static bool
-PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy)
+PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
+		  bool skip_if_not_valid)
 {
 	Buffer		b = BufferDescriptorGetBuffer(buf);
 	bool		result;
@@ -3062,6 +3051,9 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy)
 			if (old_buf_state & BM_LOCKED)
 				old_buf_state = WaitBufHdrUnlocked(buf);
 
+			if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
+				return false;
+
 			buf_state = old_buf_state;
 
 			/* increase refcount */
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v4-0002-bufmgr-Don-t-lock-buffer-header-in-StrategyGetBuf.patch (11.1K, ../../6kmid26do57ykqfpvq6iieniy4djsymhrypkjccazq5g4bbe6a@2y6owwv7qpex/3-v4-0002-bufmgr-Don-t-lock-buffer-header-in-StrategyGetBuf.patch)
  download | inline diff:
From f2f8ce7f495c6ac0e84847e532fcba2d6914b37d Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 8 Sep 2025 17:16:01 -0400
Subject: [PATCH v4 2/6] bufmgr: Don't lock buffer header in
 StrategyGetBuffer()

Previously StrategyGetBuffer() acquired the buffer header spinlock for every
buffer, whether it was reusable or not. If reusable, it'd be returned, with
the lock held, to GetVictimBuffer(), which then would pin the buffer with
PinBuffer_Locked(). That's somewhat violating the spirit of the guidelines for
holding spinlocks (i.e. that they are only held for a few lines of consecutive
code) and necessitates using PinBuffer_Locked(), which scales worse than
PinBuffer() due to holding the spinlock.  This alone makes it worth changing
the code.

However, the main reason to change this is that a future commit will make
PinBuffer_Locked() slower (due to making UnlockBufHdr() slower), to gain
scalability for the much more common case of pinning a pre-existing buffer. By
pinning the buffer with a single atomic operation, iff the buffer is reusable,
we avoid any potential regression for miss-heavy workloads. There strictly are
fewer atomic operations for each potential buffer after this change.

The price for this improvement is that freelist.c needs two CAS loops and
needs to be able to set up the resource accounting for pinned buffers. The
latter is achieved by exposing a new function for that purpose from bufmgr.c,
that seems better than exposing the entire private refcount infrastructure.
The improvement seems worth the complexity.

Reviewed-by: Robert Haas <robertmhaas@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h   |   4 +
 src/backend/storage/buffer/bufmgr.c   |  44 +++++------
 src/backend/storage/buffer/freelist.c | 110 +++++++++++++++++++-------
 3 files changed, 105 insertions(+), 53 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index dfd614f7ca4..c1206a46aba 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -371,6 +371,8 @@ UnlockBufHdr(BufferDesc *desc, uint32 buf_state)
 	pg_atomic_write_u32(&desc->state, buf_state & (~BM_LOCKED));
 }
 
+extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
+
 /* in bufmgr.c */
 
 /*
@@ -425,6 +427,8 @@ extern void IssuePendingWritebacks(WritebackContext *wb_context, IOContext io_co
 extern void ScheduleBufferTagForWriteback(WritebackContext *wb_context,
 										  IOContext io_context, BufferTag *tag);
 
+extern void TrackNewBufferPin(Buffer buf);
+
 /* solely to make it easier to write tests */
 extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
 extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index b5eebfb6990..0b841eb7838 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -518,7 +518,6 @@ static void PinBuffer_Locked(BufferDesc *buf);
 static void UnpinBuffer(BufferDesc *buf);
 static void UnpinBufferNoOwner(BufferDesc *buf);
 static void BufferSync(int flags);
-static uint32 WaitBufHdrUnlocked(BufferDesc *buf);
 static int	SyncOneBuffer(int buf_id, bool skip_recently_used,
 						  WritebackContext *wb_context);
 static void WaitIO(BufferDesc *buf);
@@ -2325,7 +2324,7 @@ GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context)
 
 	/*
 	 * Ensure, while the spinlock's not yet held, that there's a free refcount
-	 * entry, and a resource owner slot for the pin.
+	 * entry, and a resource owner slot for the pin. FIXME: Need to be updated
 	 */
 	ReservePrivateRefCountEntry();
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2334,17 +2333,11 @@ GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context)
 again:
 
 	/*
-	 * Select a victim buffer.  The buffer is returned with its header
-	 * spinlock still held!
+	 * Select a victim buffer.  The buffer is returned pinned by this backend.
 	 */
 	buf_hdr = StrategyGetBuffer(strategy, &buf_state, &from_ring);
 	buf = BufferDescriptorGetBuffer(buf_hdr);
 
-	Assert(BUF_STATE_GET_REFCOUNT(buf_state) == 0);
-
-	/* Pin the buffer and then release the buffer spinlock */
-	PinBuffer_Locked(buf_hdr);
-
 	/*
 	 * We shouldn't have any other pins for this buffer.
 	 */
@@ -3043,8 +3036,6 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		uint32		buf_state;
 		uint32		old_buf_state;
 
-		ref = NewPrivateRefCountEntry(b);
-
 		old_buf_state = pg_atomic_read_u32(&buf->state);
 		for (;;)
 		{
@@ -3091,6 +3082,8 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 				break;
 			}
 		}
+
+		TrackNewBufferPin(b);
 	}
 	else
 	{
@@ -3110,11 +3103,12 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		 * cannot meddle with that.
 		 */
 		result = (pg_atomic_read_u32(&buf->state) & BM_VALID) != 0;
+
+		Assert(ref->refcount > 0);
+		ref->refcount++;
+		ResourceOwnerRememberBuffer(CurrentResourceOwner, b);
 	}
 
-	ref->refcount++;
-	Assert(ref->refcount > 0);
-	ResourceOwnerRememberBuffer(CurrentResourceOwner, b);
 	return result;
 }
 
@@ -3143,8 +3137,6 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	Buffer		b;
-	PrivateRefCountEntry *ref;
 	uint32		buf_state;
 
 	/*
@@ -3169,12 +3161,7 @@ PinBuffer_Locked(BufferDesc *buf)
 	buf_state += BUF_REFCOUNT_ONE;
 	UnlockBufHdr(buf, buf_state);
 
-	b = BufferDescriptorGetBuffer(buf);
-
-	ref = NewPrivateRefCountEntry(b);
-	ref->refcount++;
-
-	ResourceOwnerRememberBuffer(CurrentResourceOwner, b);
+	TrackNewBufferPin(BufferDescriptorGetBuffer(buf));
 }
 
 /*
@@ -3289,6 +3276,17 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	}
 }
 
+inline void
+TrackNewBufferPin(Buffer buf)
+{
+	PrivateRefCountEntry *ref;
+
+	ref = NewPrivateRefCountEntry(buf);
+	ref->refcount++;
+
+	ResourceOwnerRememberBuffer(CurrentResourceOwner, buf);
+}
+
 #define ST_SORT sort_checkpoint_bufferids
 #define ST_ELEMENT_TYPE CkptSortItem
 #define ST_COMPARE(a, b) ckpt_buforder_comparator(a, b)
@@ -6242,7 +6240,7 @@ LockBufHdr(BufferDesc *desc)
  * Obviously the buffer could be locked by the time the value is returned, so
  * this is primarily useful in CAS style loops.
  */
-static uint32
+pg_noinline uint32
 WaitBufHdrUnlocked(BufferDesc *buf)
 {
 	SpinDelayStatus delayStatus;
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 7d59a92bd1a..9d9fb0471a0 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -164,8 +164,8 @@ ClockSweepTick(void)
  *
  *	strategy is a BufferAccessStrategy object, or NULL for default strategy.
  *
- *	To ensure that no one else can pin the buffer before we do, we must
- *	return the buffer with the buffer header spinlock still held.
+ *	The buffer is pinned before returning, but resource ownership is the
+ *	responsibility of the caller.
  */
 BufferDesc *
 StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_ring)
@@ -173,7 +173,6 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 	BufferDesc *buf;
 	int			bgwprocno;
 	int			trycounter;
-	uint32		local_buf_state;	/* to avoid repeated (de-)referencing */
 
 	*from_ring = false;
 
@@ -228,44 +227,74 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 	trycounter = NBuffers;
 	for (;;)
 	{
+		uint32		old_buf_state;
+		uint32		local_buf_state;
+
 		buf = GetBufferDescriptor(ClockSweepTick());
 
 		/*
 		 * If the buffer is pinned or has a nonzero usage_count, we cannot use
 		 * it; decrement the usage_count (unless pinned) and keep scanning.
 		 */
-		local_buf_state = LockBufHdr(buf);
 
-		if (BUF_STATE_GET_REFCOUNT(local_buf_state) == 0)
+		old_buf_state = pg_atomic_read_u32(&buf->state);
+
+		for (;;)
 		{
+			local_buf_state = old_buf_state;
+
+			if (BUF_STATE_GET_REFCOUNT(local_buf_state) != 0)
+			{
+				if (--trycounter == 0)
+				{
+					/*
+					 * We've scanned all the buffers without making any state
+					 * changes, so all the buffers are pinned (or were when we
+					 * looked at them). We could hope that someone will free
+					 * one eventually, but it's probably better to fail than
+					 * to risk getting stuck in an infinite loop.
+					 */
+					elog(ERROR, "no unpinned buffers available");
+				}
+				break;
+			}
+
+			if (unlikely(local_buf_state & BM_LOCKED))
+			{
+				old_buf_state = WaitBufHdrUnlocked(buf);
+				continue;
+			}
+
 			if (BUF_STATE_GET_USAGECOUNT(local_buf_state) != 0)
 			{
 				local_buf_state -= BUF_USAGECOUNT_ONE;
 
-				trycounter = NBuffers;
+				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+												   local_buf_state))
+				{
+					trycounter = NBuffers;
+					break;
+				}
 			}
 			else
 			{
-				/* Found a usable buffer */
-				if (strategy != NULL)
-					AddBufferToRing(strategy, buf);
-				*buf_state = local_buf_state;
-				return buf;
+				local_buf_state += BUF_REFCOUNT_ONE;
+
+				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+												   local_buf_state))
+				{
+					/* Found a usable buffer */
+					if (strategy != NULL)
+						AddBufferToRing(strategy, buf);
+					*buf_state = local_buf_state;
+
+					TrackNewBufferPin(BufferDescriptorGetBuffer(buf));
+
+					return buf;
+				}
 			}
+
 		}
-		else if (--trycounter == 0)
-		{
-			/*
-			 * We've scanned all the buffers without making any state changes,
-			 * so all the buffers are pinned (or were when we looked at them).
-			 * We could hope that someone will free one eventually, but it's
-			 * probably better to fail than to risk getting stuck in an
-			 * infinite loop.
-			 */
-			UnlockBufHdr(buf, local_buf_state);
-			elog(ERROR, "no unpinned buffers available");
-		}
-		UnlockBufHdr(buf, local_buf_state);
 	}
 }
 
@@ -621,6 +650,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 {
 	BufferDesc *buf;
 	Buffer		bufnum;
+	uint32		old_buf_state;
 	uint32		local_buf_state;	/* to avoid repeated (de-)referencing */
 
 
@@ -647,14 +677,34 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 	 * shouldn't re-use it.
 	 */
 	buf = GetBufferDescriptor(bufnum - 1);
-	local_buf_state = LockBufHdr(buf);
-	if (BUF_STATE_GET_REFCOUNT(local_buf_state) == 0
-		&& BUF_STATE_GET_USAGECOUNT(local_buf_state) <= 1)
+
+	old_buf_state = pg_atomic_read_u32(&buf->state);
+
+	for (;;)
 	{
-		*buf_state = local_buf_state;
-		return buf;
+		local_buf_state = old_buf_state;
+
+		if (BUF_STATE_GET_REFCOUNT(local_buf_state) != 0
+			|| BUF_STATE_GET_USAGECOUNT(local_buf_state) > 1)
+			break;
+
+		if (unlikely(local_buf_state & BM_LOCKED))
+		{
+			old_buf_state = WaitBufHdrUnlocked(buf);
+			continue;
+		}
+
+		local_buf_state += BUF_REFCOUNT_ONE;
+
+		if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+										   local_buf_state))
+		{
+			*buf_state = local_buf_state;
+
+			TrackNewBufferPin(BufferDescriptorGetBuffer(buf));
+			return buf;
+		}
 	}
-	UnlockBufHdr(buf, local_buf_state);
 
 	/*
 	 * Tell caller to allocate a new buffer with the normal allocation
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v4-0003-bufmgr-Allow-some-buffer-state-modifications-whil.patch (31.2K, ../../6kmid26do57ykqfpvq6iieniy4djsymhrypkjccazq5g4bbe6a@2y6owwv7qpex/4-v4-0003-bufmgr-Allow-some-buffer-state-modifications-whil.patch)
  download | inline diff:
From 27d90068f2e4841434f35e3f5ecb6fe7dceeca50 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 15 Sep 2025 17:06:48 -0400
Subject: [PATCH v4 3/6] bufmgr: Allow some buffer state modifications while
 holding header lock

Until now BufferDesc.state was not allowed to be modified while the buffer
header spinlock was held. This meant that operations like unpinning buffers
needed to use a CAS loop, waiting for the buffer header spinlock to be
released before updating.

The benefit of that restriction is that it allowed us to unlock the buffer
header spinlock with just a write barrier and an unlocked write (instead of a
full atomic operation). That was important to avoid regressions in
48354581a49c. However, since then the hottest buffer header spinlock uses have
been replaced with atomic operations (in particular, the most common use of
PinBuffer_Locked(), in GetVictimBuffer() (formerly in BufferAlloc()), has been
removed in FIXME-TODO-FIXME).

This change will allow, in a subsequent commit, to release buffer pins with a
single atomic-sub operation. This previously was not possible while such
operations were not allowed while the buffer header spinlock was held, as an
atomic-sub would not have allowed a race-free check for the buffer header lock
being held.

Using atomic-sub to unpin buffers is a nice scalability win, however it is not
the primary motivation for this change (although it would be sufficient). The
primary motivation is that we would like to merge the buffer content lock into
BufferDesc.state, which will result in more frequent changes of the state
variable, which in some situations can cause a performance regression, due to
an increased CAS failure rate when unpinning buffers.  The regression entirely
vanishes when using atomic-sub.

Naively implementing this would require putting CAS loops in every place
modifying the buffer state while holding the buffer header lock. To avoid
that, introduce UnlockBufHdrExt(), which can set/add flags as well as the
refcount, together with releasing the lock.

Reviewed-by: Robert Haas <robertmhaas@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           | 119 +++++++---
 src/backend/storage/buffer/bufmgr.c           | 203 ++++++++++--------
 src/backend/storage/buffer/freelist.c         |   2 +
 contrib/pg_buffercache/pg_buffercache_pages.c |   7 +-
 contrib/pg_prewarm/autoprewarm.c              |   2 +-
 src/test/modules/test_aio/test_aio.c          |  10 +-
 6 files changed, 224 insertions(+), 119 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index c1206a46aba..e25b1ece3f9 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -211,28 +211,36 @@ BufMappingPartitionLockByIndex(uint32 index)
 /*
  *	BufferDesc -- shared descriptor/state data for a single shared buffer.
  *
- * Note: Buffer header lock (BM_LOCKED flag) must be held to examine or change
- * tag, state or wait_backend_pgprocno fields.  In general, buffer header lock
- * is a spinlock which is combined with flags, refcount and usagecount into
- * single atomic variable.  This layout allow us to do some operations in a
- * single atomic operation, without actually acquiring and releasing spinlock;
- * for instance, increase or decrease refcount.  buf_id field never changes
- * after initialization, so does not need locking.  The LWLock can take care
- * of itself.  The buffer header lock is *not* used to control access to the
- * data in the buffer!
+ * The state of the buffer is controlled by the, drumroll, state variable. It
+ * only may be modified using atomic operations.  The state variable combines
+ * various flags, the buffer's refcount and usage count. See comment above
+ * BUF_REFCOUNT_BITS for details about the division.  This layout allow us to
+ * do some operations in a single atomic operation, without actually acquiring
+ * and releasing the spinlock; for instance, increasing or decreasing the
+ * refcount.
  *
- * It's assumed that nobody changes the state field while buffer header lock
- * is held.  Thus buffer header lock holder can do complex updates of the
- * state variable in single write, simultaneously with lock release (cleaning
- * BM_LOCKED flag).  On the other hand, updating of state without holding
- * buffer header lock is restricted to CAS, which ensures that BM_LOCKED flag
- * is not set.  Atomic increment/decrement, OR/AND etc. are not allowed.
+ * One of the aforementioned flags is BM_LOCKED, used to implement the buffer
+ * header lock. See the following paragraphs, as well as the documentation for
+ * individual fields for more details.
  *
- * An exception is that if we have the buffer pinned, its tag can't change
- * underneath us, so we can examine the tag without locking the buffer header.
- * Also, in places we do one-time reads of the flags without bothering to
- * lock the buffer header; this is generally for situations where we don't
- * expect the flag bit being tested to be changing.
+ * The identity of the buffer (BufferDesc.tag) can only be changed by the
+ * backend holding the buffer header lock.
+ *
+ * If the lock is held by another backend, neither additional buffer pins may
+ * be established (we would like to relax this eventually), nor can flags be
+ * set/cleared. These operations either need to acquire the buffer header
+ * spinlock, or need to use a CAS loop, waiting for the lock to be released if
+ * it is held.  However, existing buffer pins may be released while the buffer
+ * header spinlock is held, using an atomic subtraction.
+ *
+ * The LWLock can take care of itself.  The buffer header lock is *not* used
+ * to control access to the data in the buffer!
+ *
+ * If we have the buffer pinned, its tag can't change underneath us, so we can
+ * examine the tag without locking the buffer header.  Also, in places we do
+ * one-time reads of the flags without bothering to lock the buffer header;
+ * this is generally for situations where we don't expect the flag bit being
+ * tested to be changing.
  *
  * We can't physically remove items from a disk page if another backend has
  * the buffer pinned.  Hence, a backend may need to wait for all other pins
@@ -256,13 +264,29 @@ BufMappingPartitionLockByIndex(uint32 index)
  */
 typedef struct BufferDesc
 {
-	BufferTag	tag;			/* ID of page contained in buffer */
-	int			buf_id;			/* buffer's index number (from 0) */
+	/*
+	 * ID of page contained in buffer. The buffer header spinlock needs to be
+	 * held to modify this field.
+	 */
+	BufferTag	tag;
 
-	/* state of the tag, containing flags, refcount and usagecount */
+	/*
+	 * Buffer's index number (from 0). The field never changes after
+	 * initialization, so does not need locking.
+	 */
+	int			buf_id;
+
+	/*
+	 * State of the buffer, containing flags, refcount and usagecount. See
+	 * BUF_* and BM_* defines at the top of this file.
+	 */
 	pg_atomic_uint32 state;
 
-	int			wait_backend_pgprocno;	/* backend of pin-count waiter */
+	/*
+	 * Backend of pin-count waiter. The buffer header spinlock needs to be
+	 * held to modify this field.
+	 */
+	int			wait_backend_pgprocno;
 
 	PgAioWaitRef io_wref;		/* set iff AIO is in progress */
 	LWLock		content_lock;	/* to lock access to buffer contents */
@@ -364,11 +388,52 @@ BufferDescriptorGetContentLock(const BufferDesc *bdesc)
  */
 extern uint32 LockBufHdr(BufferDesc *desc);
 
+/*
+ * Unlock the buffer header.
+ *
+ * This can only be used if the caller did not modify BufferDesc.state. To
+ * set/unset flag bits or change the refcount use UnlockBufHdrExt().
+ */
 static inline void
-UnlockBufHdr(BufferDesc *desc, uint32 buf_state)
+UnlockBufHdr(BufferDesc *desc)
 {
-	pg_write_barrier();
-	pg_atomic_write_u32(&desc->state, buf_state & (~BM_LOCKED));
+	Assert(pg_atomic_read_u32(&desc->state) & BM_LOCKED);
+
+	pg_atomic_fetch_sub_u32(&desc->state, BM_LOCKED);
+}
+
+/*
+ * Unlock the buffer header, while atomically adding the flags in set_bits,
+ * unsetting the ones in unset_bits and changing the refcount by
+ * refcount_change.
+ *
+ * Note that this approach would not work for usagecount, since we need to cap
+ * the usagecount at BM_MAX_USAGE_COUNT.
+ */
+static inline uint32
+UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
+				uint32 set_bits, uint32 unset_bits,
+				int refcount_change)
+{
+	for (;;)
+	{
+		uint32		buf_state = old_buf_state;
+
+		Assert(buf_state & BM_LOCKED);
+
+		buf_state |= set_bits;
+		buf_state &= ~unset_bits;
+		buf_state &= ~BM_LOCKED;
+
+		if (refcount_change != 0)
+			buf_state += BUF_REFCOUNT_ONE * refcount_change;
+
+		if (pg_atomic_compare_exchange_u32(&desc->state, &old_buf_state,
+										   buf_state))
+		{
+			return old_buf_state;
+		}
+	}
 }
 
 extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 0b841eb7838..0ce4a736b58 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -1994,6 +1994,7 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	Buffer		victim_buffer;
 	BufferDesc *victim_buf_hdr;
 	uint32		victim_buf_state;
+	uint32		set_bits = 0;
 
 	/* Make sure we will have room to remember the buffer pin */
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2120,11 +2121,12 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	 * checkpoints, except for their "init" forks, which need to be treated
 	 * just like permanent relations.
 	 */
-	victim_buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
+	set_bits |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 	if (relpersistence == RELPERSISTENCE_PERMANENT || forkNum == INIT_FORKNUM)
-		victim_buf_state |= BM_PERMANENT;
+		set_bits |= BM_PERMANENT;
 
-	UnlockBufHdr(victim_buf_hdr, victim_buf_state);
+	UnlockBufHdrExt(victim_buf_hdr, victim_buf_state,
+					set_bits, 0, 0);
 
 	LWLockRelease(newPartitionLock);
 
@@ -2164,9 +2166,7 @@ InvalidateBuffer(BufferDesc *buf)
 	/* Save the original buffer tag before dropping the spinlock */
 	oldTag = buf->tag;
 
-	buf_state = pg_atomic_read_u32(&buf->state);
-	Assert(buf_state & BM_LOCKED);
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdr(buf);
 
 	/*
 	 * Need to compute the old tag's hashcode and partition lock ID. XXX is it
@@ -2190,7 +2190,7 @@ retry:
 	/* If it's changed while we were waiting for lock, do nothing */
 	if (!BufferTagsEqual(&buf->tag, &oldTag))
 	{
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		LWLockRelease(oldPartitionLock);
 		return;
 	}
@@ -2207,7 +2207,7 @@ retry:
 	 */
 	if (BUF_STATE_GET_REFCOUNT(buf_state) != 0)
 	{
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		LWLockRelease(oldPartitionLock);
 		/* safety check: should definitely not be our *own* pin */
 		if (GetPrivateRefCount(BufferDescriptorGetBuffer(buf)) > 0)
@@ -2222,8 +2222,11 @@ retry:
 	 */
 	oldFlags = buf_state & BUF_FLAG_MASK;
 	ClearBufferTag(&buf->tag);
-	buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
-	UnlockBufHdr(buf, buf_state);
+
+	UnlockBufHdrExt(buf, buf_state,
+					0,
+					BUF_FLAG_MASK | BUF_USAGECOUNT_MASK,
+					0);
 
 	/*
 	 * Remove the buffer from the lookup hashtable, if it was in there.
@@ -2283,7 +2286,7 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 	{
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 
-		UnlockBufHdr(buf_hdr, buf_state);
+		UnlockBufHdr(buf_hdr);
 		LWLockRelease(partition_lock);
 
 		return false;
@@ -2297,8 +2300,10 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 	 * tag (see e.g. FlushDatabaseBuffers()).
 	 */
 	ClearBufferTag(&buf_hdr->tag);
-	buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
-	UnlockBufHdr(buf_hdr, buf_state);
+	UnlockBufHdrExt(buf_hdr, buf_state,
+					0,
+					BUF_FLAG_MASK | BUF_USAGECOUNT_MASK,
+					0);
 
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 
@@ -2307,6 +2312,7 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 
 	LWLockRelease(partition_lock);
 
+	buf_state = pg_atomic_read_u32(&buf_hdr->state);
 	Assert(!(buf_state & (BM_DIRTY | BM_VALID | BM_TAG_VALID)));
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u32(&buf_hdr->state)) > 0);
@@ -2396,7 +2402,7 @@ again:
 			/* Read the LSN while holding buffer header lock */
 			buf_state = LockBufHdr(buf_hdr);
 			lsn = BufferGetLSN(buf_hdr);
-			UnlockBufHdr(buf_hdr, buf_state);
+			UnlockBufHdr(buf_hdr);
 
 			if (XLogNeedsFlush(lsn)
 				&& StrategyRejectBuffer(strategy, buf_hdr, from_ring))
@@ -2741,15 +2747,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 */
 			do
 			{
-				uint32		buf_state = LockBufHdr(existing_hdr);
-
-				buf_state &= ~BM_VALID;
-				UnlockBufHdr(existing_hdr, buf_state);
+				pg_atomic_fetch_and_u32(&existing_hdr->state, ~BM_VALID);
 			} while (!StartBufferIO(existing_hdr, true, false));
 		}
 		else
 		{
 			uint32		buf_state;
+			uint32		set_bits = 0;
 
 			buf_state = LockBufHdr(victim_buf_hdr);
 
@@ -2759,11 +2763,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 
 			victim_buf_hdr->tag = tag;
 
-			buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
+			set_bits |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 			if (bmr.relpersistence == RELPERSISTENCE_PERMANENT || fork == INIT_FORKNUM)
-				buf_state |= BM_PERMANENT;
+				set_bits |= BM_PERMANENT;
 
-			UnlockBufHdr(victim_buf_hdr, buf_state);
+			UnlockBufHdrExt(victim_buf_hdr, buf_state,
+							set_bits, 0,
+							0);
 
 			LWLockRelease(partition_lock);
 
@@ -2918,6 +2924,10 @@ MarkBufferDirty(Buffer buffer)
 	Assert(LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
 								LW_EXCLUSIVE));
 
+	/*
+	 * NB: We have to wait for the buffer header spinlock to be not held, as
+	 * TerminateBufferIO() relies on the spinlock.
+	 */
 	old_buf_state = pg_atomic_read_u32(&bufHdr->state);
 	for (;;)
 	{
@@ -3039,6 +3049,10 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		old_buf_state = pg_atomic_read_u32(&buf->state);
 		for (;;)
 		{
+			/*
+			 * We're not allowed to increase the refcount while the buffer
+			 * header spinlock is held. Wait for the lock to be released.
+			 */
 			if (old_buf_state & BM_LOCKED)
 				old_buf_state = WaitBufHdrUnlocked(buf);
 
@@ -3137,7 +3151,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	uint32		buf_state;
+	uint32		old_buf_state;
 
 	/*
 	 * As explained, We don't expect any preexisting pins. That allows us to
@@ -3156,10 +3170,10 @@ PinBuffer_Locked(BufferDesc *buf)
 	 * Since we hold the buffer spinlock, we can update the buffer state and
 	 * release the lock in one operation.
 	 */
-	buf_state = pg_atomic_read_u32(&buf->state);
-	Assert(buf_state & BM_LOCKED);
-	buf_state += BUF_REFCOUNT_ONE;
-	UnlockBufHdr(buf, buf_state);
+	old_buf_state = pg_atomic_read_u32(&buf->state);
+
+	UnlockBufHdrExt(buf, old_buf_state,
+					0, 0, 1);
 
 	TrackNewBufferPin(BufferDescriptorGetBuffer(buf));
 }
@@ -3194,12 +3208,13 @@ WakePinCountWaiter(BufferDesc *buf)
 		/* we just released the last pin other than the waiter's */
 		int			wait_backend_pgprocno = buf->wait_backend_pgprocno;
 
-		buf_state &= ~BM_PIN_COUNT_WAITER;
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdrExt(buf, buf_state,
+						0, BM_PIN_COUNT_WAITER,
+						0);
 		ProcSendSignal(wait_backend_pgprocno);
 	}
 	else
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 }
 
 /*
@@ -3252,6 +3267,10 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 *
 		 * Since buffer spinlock holder can update status using just write,
 		 * it's not safe to use atomic decrement here; thus use a CAS loop.
+		 *
+		 * TODO: The above requirement does not hold anymore, in a future
+		 * commit this will be rewritten to release the pin in a single atomic
+		 * operation.
 		 */
 		old_buf_state = pg_atomic_read_u32(&buf->state);
 		for (;;)
@@ -3349,6 +3368,7 @@ BufferSync(int flags)
 	for (buf_id = 0; buf_id < NBuffers; buf_id++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
+		uint32		set_bits = 0;
 
 		/*
 		 * Header spinlock is enough to examine BM_DIRTY, see comment in
@@ -3360,7 +3380,7 @@ BufferSync(int flags)
 		{
 			CkptSortItem *item;
 
-			buf_state |= BM_CHECKPOINT_NEEDED;
+			set_bits = BM_CHECKPOINT_NEEDED;
 
 			item = &CkptBufferIds[num_to_scan++];
 			item->buf_id = buf_id;
@@ -3370,7 +3390,9 @@ BufferSync(int flags)
 			item->blockNum = bufHdr->tag.blockNum;
 		}
 
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdrExt(bufHdr, buf_state,
+						set_bits, 0,
+						0);
 
 		/* Check for barrier events in case NBuffers is large. */
 		if (ProcSignalBarrierPending)
@@ -3909,14 +3931,14 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	else if (skip_recently_used)
 	{
 		/* Caller told us not to write recently-used buffers */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return result;
 	}
 
 	if (!(buf_state & BM_VALID) || !(buf_state & BM_DIRTY))
 	{
 		/* It's clean, so nothing to do */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return result;
 	}
 
@@ -4288,8 +4310,9 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	recptr = BufferGetLSN(buf);
 
 	/* To check if block content changes while flushing. - vadim 01/17/97 */
-	buf_state &= ~BM_JUST_DIRTIED;
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdrExt(buf, buf_state,
+					0, BM_JUST_DIRTIED,
+					0);
 
 	/*
 	 * Force XLOG flush up to buffer's LSN.  This implements the basic WAL
@@ -4452,7 +4475,6 @@ BufferGetLSNAtomic(Buffer buffer)
 	char	   *page = BufferGetPage(buffer);
 	BufferDesc *bufHdr;
 	XLogRecPtr	lsn;
-	uint32		buf_state;
 
 	/*
 	 * If we don't need locking for correctness, fastpath out.
@@ -4465,9 +4487,9 @@ BufferGetLSNAtomic(Buffer buffer)
 	Assert(BufferIsPinned(buffer));
 
 	bufHdr = GetBufferDescriptor(buffer - 1);
-	buf_state = LockBufHdr(bufHdr);
+	LockBufHdr(bufHdr);
 	lsn = PageGetLSN(page);
-	UnlockBufHdr(bufHdr, buf_state);
+	UnlockBufHdr(bufHdr);
 
 	return lsn;
 }
@@ -4568,7 +4590,6 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
 	for (i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
 
 		/*
 		 * We can make this a tad faster by prechecking the buffer tag before
@@ -4589,7 +4610,7 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
 		if (!BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator.locator))
 			continue;
 
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 
 		for (j = 0; j < nforks; j++)
 		{
@@ -4602,7 +4623,7 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
 			}
 		}
 		if (j >= nforks)
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -4731,7 +4752,6 @@ DropRelationsAllBuffers(SMgrRelation *smgr_reln, int nlocators)
 	{
 		RelFileLocator *rlocator = NULL;
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -4765,11 +4785,11 @@ DropRelationsAllBuffers(SMgrRelation *smgr_reln, int nlocators)
 		if (rlocator == NULL)
 			continue;
 
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 		if (BufTagMatchesRelFileLocator(&bufHdr->tag, rlocator))
 			InvalidateBuffer(bufHdr);	/* releases spinlock */
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 
 	pfree(locators);
@@ -4799,7 +4819,6 @@ FindAndDropRelationBuffers(RelFileLocator rlocator, ForkNumber forkNum,
 		LWLock	   *bufPartitionLock;	/* buffer partition lock for it */
 		int			buf_id;
 		BufferDesc *bufHdr;
-		uint32		buf_state;
 
 		/* create a tag so we can lookup the buffer */
 		InitBufferTag(&bufTag, &rlocator, forkNum, curBlock);
@@ -4824,14 +4843,14 @@ FindAndDropRelationBuffers(RelFileLocator rlocator, ForkNumber forkNum,
 		 * evicted by some other backend loading blocks for a different
 		 * relation after we release lock on the BufMapping table.
 		 */
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 
 		if (BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator) &&
 			BufTagGetForkNum(&bufHdr->tag) == forkNum &&
 			bufHdr->tag.blockNum >= firstDelBlock)
 			InvalidateBuffer(bufHdr);	/* releases spinlock */
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -4859,7 +4878,6 @@ DropDatabaseBuffers(Oid dbid)
 	for (i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -4868,11 +4886,11 @@ DropDatabaseBuffers(Oid dbid)
 		if (bufHdr->tag.dbOid != dbid)
 			continue;
 
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 		if (bufHdr->tag.dbOid == dbid)
 			InvalidateBuffer(bufHdr);	/* releases spinlock */
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -4971,7 +4989,7 @@ FlushRelationBuffers(Relation rel)
 			UnpinBuffer(bufHdr);
 		}
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -5068,7 +5086,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 			UnpinBuffer(bufHdr);
 		}
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 
 	pfree(srels);
@@ -5296,7 +5314,7 @@ FlushDatabaseBuffers(Oid dbid)
 			UnpinBuffer(bufHdr);
 		}
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -5506,8 +5524,9 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 				PageSetLSN(page, lsn);
 		}
 
-		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdrExt(bufHdr, buf_state,
+						BM_DIRTY | BM_JUST_DIRTIED,
+						0, 0);
 
 		if (delayChkptFlags)
 			MyProc->delayChkptFlags &= ~DELAY_CHKPT_START;
@@ -5538,6 +5557,7 @@ UnlockBuffers(void)
 	if (buf)
 	{
 		uint32		buf_state;
+		uint32		unset_bits = 0;
 
 		buf_state = LockBufHdr(buf);
 
@@ -5547,9 +5567,11 @@ UnlockBuffers(void)
 		 */
 		if ((buf_state & BM_PIN_COUNT_WAITER) != 0 &&
 			buf->wait_backend_pgprocno == MyProcNumber)
-			buf_state &= ~BM_PIN_COUNT_WAITER;
+			unset_bits = BM_PIN_COUNT_WAITER;
 
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdrExt(buf, buf_state,
+						0, unset_bits,
+						0);
 
 		PinCountWaitBuf = NULL;
 	}
@@ -5667,6 +5689,7 @@ LockBufferForCleanup(Buffer buffer)
 	for (;;)
 	{
 		uint32		buf_state;
+		uint32		unset_bits = 0;
 
 		/* Try to acquire lock */
 		LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
@@ -5676,7 +5699,7 @@ LockBufferForCleanup(Buffer buffer)
 		if (BUF_STATE_GET_REFCOUNT(buf_state) == 1)
 		{
 			/* Successfully acquired exclusive lock with pincount 1 */
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 
 			/*
 			 * Emit the log message if recovery conflict on buffer pin was
@@ -5699,14 +5722,15 @@ LockBufferForCleanup(Buffer buffer)
 		/* Failed, so mark myself as waiting for pincount 1 */
 		if (buf_state & BM_PIN_COUNT_WAITER)
 		{
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 			LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 			elog(ERROR, "multiple backends attempting to wait for pincount 1");
 		}
 		bufHdr->wait_backend_pgprocno = MyProcNumber;
 		PinCountWaitBuf = bufHdr;
-		buf_state |= BM_PIN_COUNT_WAITER;
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdrExt(bufHdr, buf_state,
+						BM_PIN_COUNT_WAITER, 0,
+						0);
 		LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 
 		/* Wait to be signaled by UnpinBuffer() */
@@ -5768,8 +5792,11 @@ LockBufferForCleanup(Buffer buffer)
 		buf_state = LockBufHdr(bufHdr);
 		if ((buf_state & BM_PIN_COUNT_WAITER) != 0 &&
 			bufHdr->wait_backend_pgprocno == MyProcNumber)
-			buf_state &= ~BM_PIN_COUNT_WAITER;
-		UnlockBufHdr(bufHdr, buf_state);
+			unset_bits |= BM_PIN_COUNT_WAITER;
+
+		UnlockBufHdrExt(bufHdr, buf_state,
+						0, unset_bits,
+						0);
 
 		PinCountWaitBuf = NULL;
 		/* Loop back and try again */
@@ -5846,12 +5873,12 @@ ConditionalLockBufferForCleanup(Buffer buffer)
 	if (refcount == 1)
 	{
 		/* Successfully acquired exclusive lock with pincount 1 */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return true;
 	}
 
 	/* Failed, so release the lock */
-	UnlockBufHdr(bufHdr, buf_state);
+	UnlockBufHdr(bufHdr);
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 	return false;
 }
@@ -5899,11 +5926,11 @@ IsBufferCleanupOK(Buffer buffer)
 	if (BUF_STATE_GET_REFCOUNT(buf_state) == 1)
 	{
 		/* pincount is OK. */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return true;
 	}
 
-	UnlockBufHdr(bufHdr, buf_state);
+	UnlockBufHdr(bufHdr);
 	return false;
 }
 
@@ -5941,7 +5968,7 @@ WaitIO(BufferDesc *buf)
 		 * clearing the wref while it's being read.
 		 */
 		iow = buf->io_wref;
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 
 		/* no IO in progress, we don't need to wait */
 		if (!(buf_state & BM_IO_IN_PROGRESS))
@@ -6009,7 +6036,7 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 
 		if (!(buf_state & BM_IO_IN_PROGRESS))
 			break;
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		if (nowait)
 			return false;
 		WaitIO(buf);
@@ -6020,12 +6047,13 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 	/* Check if someone else already did the I/O */
 	if (forInput ? (buf_state & BM_VALID) : !(buf_state & BM_DIRTY))
 	{
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		return false;
 	}
 
-	buf_state |= BM_IO_IN_PROGRESS;
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdrExt(buf, buf_state,
+					BM_IO_IN_PROGRESS, 0,
+					0);
 
 	ResourceOwnerRememberBufferIO(CurrentResourceOwner,
 								  BufferDescriptorGetBuffer(buf));
@@ -6058,28 +6086,31 @@ TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
 				  bool forget_owner, bool release_aio)
 {
 	uint32		buf_state;
+	uint32		unset_flag_bits = 0;
+	int			refcount_change = 0;
 
 	buf_state = LockBufHdr(buf);
 
 	Assert(buf_state & BM_IO_IN_PROGRESS);
-	buf_state &= ~BM_IO_IN_PROGRESS;
+	unset_flag_bits |= BM_IO_IN_PROGRESS;
 
 	/* Clear earlier errors, if this IO failed, it'll be marked again */
-	buf_state &= ~BM_IO_ERROR;
+	unset_flag_bits |= BM_IO_ERROR;
 
 	if (clear_dirty && !(buf_state & BM_JUST_DIRTIED))
-		buf_state &= ~(BM_DIRTY | BM_CHECKPOINT_NEEDED);
+		unset_flag_bits |= BM_DIRTY | BM_CHECKPOINT_NEEDED;
 
 	if (release_aio)
 	{
 		/* release ownership by the AIO subsystem */
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-		buf_state -= BUF_REFCOUNT_ONE;
+		refcount_change = -1;
 		pgaio_wref_clear(&buf->io_wref);
 	}
 
-	buf_state |= set_flag_bits;
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdrExt(buf, buf_state,
+					set_flag_bits, unset_flag_bits,
+					refcount_change);
 
 	if (forget_owner)
 		ResourceOwnerForgetBufferIO(CurrentResourceOwner,
@@ -6124,12 +6155,12 @@ AbortBufferIO(Buffer buffer)
 	if (!(buf_state & BM_VALID))
 	{
 		Assert(!(buf_state & BM_DIRTY));
-		UnlockBufHdr(buf_hdr, buf_state);
+		UnlockBufHdr(buf_hdr);
 	}
 	else
 	{
 		Assert(buf_state & BM_DIRTY);
-		UnlockBufHdr(buf_hdr, buf_state);
+		UnlockBufHdr(buf_hdr);
 
 		/* Issue notice if this is not the first failure... */
 		if (buf_state & BM_IO_ERROR)
@@ -6551,14 +6582,14 @@ EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 
 	if ((buf_state & BM_VALID) == 0)
 	{
-		UnlockBufHdr(desc, buf_state);
+		UnlockBufHdr(desc);
 		return false;
 	}
 
 	/* Check that it's not pinned already. */
 	if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
 	{
-		UnlockBufHdr(desc, buf_state);
+		UnlockBufHdr(desc);
 		return false;
 	}
 
@@ -6710,7 +6741,7 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
 		if ((buf_state & BM_VALID) == 0 ||
 			!BufTagMatchesRelFileLocator(&desc->tag, &rel->rd_locator))
 		{
-			UnlockBufHdr(desc, buf_state);
+			UnlockBufHdr(desc);
 			continue;
 		}
 
@@ -6808,13 +6839,15 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		 *
 		 * This pin is released again in TerminateBufferIO().
 		 */
-		buf_state += BUF_REFCOUNT_ONE;
 		buf_hdr->io_wref = io_ref;
 
 		if (is_temp)
+		{
+			buf_state += BUF_REFCOUNT_ONE;
 			pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		}
 		else
-			UnlockBufHdr(buf_hdr, buf_state);
+			UnlockBufHdrExt(buf_hdr, buf_state, 0, 0, 1);
 
 		/*
 		 * Ensure the content lock that prevents buffer modifications while
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 9d9fb0471a0..c98080a7d0b 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -259,6 +259,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 				break;
 			}
 
+			/* See equivalent code in PinBuffer() */
 			if (unlikely(local_buf_state & BM_LOCKED))
 			{
 				old_buf_state = WaitBufHdrUnlocked(buf);
@@ -688,6 +689,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 			|| BUF_STATE_GET_USAGECOUNT(local_buf_state) > 1)
 			break;
 
+		/* See equivalent code in PinBuffer() */
 		if (unlikely(local_buf_state & BM_LOCKED))
 		{
 			old_buf_state = WaitBufHdrUnlocked(buf);
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 3df04c98959..ab790533ff6 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -220,7 +220,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
 			else
 				fctx->record[i].isvalid = false;
 
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 		}
 	}
 
@@ -460,7 +460,6 @@ pg_buffercache_numa_pages(PG_FUNCTION_ARGS)
 		{
 			char	   *buffptr = (char *) BufferGetBlock(i + 1);
 			BufferDesc *bufHdr;
-			uint32		buf_state;
 			uint32		bufferid;
 			int32		page_num;
 			char	   *startptr_buff,
@@ -471,9 +470,9 @@ pg_buffercache_numa_pages(PG_FUNCTION_ARGS)
 			bufHdr = GetBufferDescriptor(i);
 
 			/* Lock each buffer header before inspecting. */
-			buf_state = LockBufHdr(bufHdr);
+			LockBufHdr(bufHdr);
 			bufferid = BufferDescriptorGetBuffer(bufHdr);
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 
 			/* start of the first page of this buffer */
 			startptr_buff = (char *) TYPEALIGN_DOWN(os_page_size, buffptr);
diff --git a/contrib/pg_prewarm/autoprewarm.c b/contrib/pg_prewarm/autoprewarm.c
index 8b68dafc261..5ba1240d51f 100644
--- a/contrib/pg_prewarm/autoprewarm.c
+++ b/contrib/pg_prewarm/autoprewarm.c
@@ -730,7 +730,7 @@ apw_dump_now(bool is_bgworker, bool dump_unlogged)
 			++num_blocks;
 		}
 
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 	}
 
 	snprintf(transient_dump_file_path, MAXPGPATH, "%s.tmp", AUTOPREWARM_FILE);
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index c55cf6c0aac..d7eadeab256 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -310,6 +310,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	BufferDesc *buf_hdr;
 	uint32		buf_state;
 	bool		was_pinned = false;
+	uint32		unset_bits = 0;
 
 	/* place buffer in shared buffers without erroring out */
 	buf = ReadBufferExtended(rel, MAIN_FORKNUM, blkno, RBM_ZERO_AND_LOCK, NULL);
@@ -334,12 +335,17 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (BUF_STATE_GET_REFCOUNT(buf_state) > 1)
 		was_pinned = true;
 	else
-		buf_state &= ~(BM_VALID | BM_DIRTY);
+		unset_bits |= BM_VALID | BM_DIRTY;
 
 	if (RelationUsesLocalBuffers(rel))
+	{
+		buf_state &= ~unset_bits;
 		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+	}
 	else
-		UnlockBufHdr(buf_hdr, buf_state);
+	{
+		UnlockBufHdrExt(buf_hdr, buf_state, 0, unset_bits, 0);
+	}
 
 	if (was_pinned)
 		elog(ERROR, "toy buffer %d was already pinned",
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v4-0004-bufmgr-Use-atomic-sub-for-unpinning-buffers.patch (2.3K, ../../6kmid26do57ykqfpvq6iieniy4djsymhrypkjccazq5g4bbe6a@2y6owwv7qpex/5-v4-0004-bufmgr-Use-atomic-sub-for-unpinning-buffers.patch)
  download | inline diff:
From e763e0073f1b6c82b77be77fdd101ce3e54f650f Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 22 Sep 2025 14:44:52 -0400
Subject: [PATCH v4 4/6] bufmgr: Use atomic sub for unpinning buffers

The prior commit made it legal to modify BufferDesc.state while the buffer
header spinlock is held. This allows us to replace the CAS loop
inUnpinBufferNoOwner() with an atomic sub. This improves scalability
significantly. See the prior commits for more background.

Reviewed-by: Robert Haas <robertmhaas@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/backend/storage/buffer/bufmgr.c | 29 +++--------------------------
 1 file changed, 3 insertions(+), 26 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 0ce4a736b58..594a8e5d512 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3247,7 +3247,6 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	ref->refcount--;
 	if (ref->refcount == 0)
 	{
-		uint32		buf_state;
 		uint32		old_buf_state;
 
 		/*
@@ -3262,33 +3261,11 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		/* I'd better not still hold the buffer content lock */
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
-		/*
-		 * Decrement the shared reference count.
-		 *
-		 * Since buffer spinlock holder can update status using just write,
-		 * it's not safe to use atomic decrement here; thus use a CAS loop.
-		 *
-		 * TODO: The above requirement does not hold anymore, in a future
-		 * commit this will be rewritten to release the pin in a single atomic
-		 * operation.
-		 */
-		old_buf_state = pg_atomic_read_u32(&buf->state);
-		for (;;)
-		{
-			if (old_buf_state & BM_LOCKED)
-				old_buf_state = WaitBufHdrUnlocked(buf);
-
-			buf_state = old_buf_state;
-
-			buf_state -= BUF_REFCOUNT_ONE;
-
-			if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
-											   buf_state))
-				break;
-		}
+		/* decrement the shared reference count */
+		old_buf_state = pg_atomic_fetch_sub_u32(&buf->state, BUF_REFCOUNT_ONE);
 
 		/* Support LockBufferForCleanup() */
-		if (buf_state & BM_PIN_COUNT_WAITER)
+		if (old_buf_state & BM_PIN_COUNT_WAITER)
 			WakePinCountWaiter(buf);
 
 		ForgetPrivateRefCountEntry(ref);
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v4-0005-bufmgr-fewer-calls-to-BufferDescriptorGetContentL.patch (7.2K, ../../6kmid26do57ykqfpvq6iieniy4djsymhrypkjccazq5g4bbe6a@2y6owwv7qpex/6-v4-0005-bufmgr-fewer-calls-to-BufferDescriptorGetContentL.patch)
  download | inline diff:
From c044083cde67f216c12c1d6a481a8b4390a5882f Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 30 Jun 2025 16:54:39 -0400
Subject: [PATCH v4 5/6] bufmgr: fewer calls to BufferDescriptorGetContentLock

We're planning to merge buffer content locks into BufferDesc.state. To reduce
the size of that patch, centralize BufferDescriptorGetContentLock().

The biggest part of the change is in assertions, by introducing
BufferIsLockedByMe[InMode]() (and removing BufferIsExclusiveLocked()). This
seems like an improvement even without aforementioned plans.

Additionally replace some direct calls to LWLockAcquire() with calls to
LockBuffer().

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/bufmgr.h            |  3 +-
 src/backend/access/heap/visibilitymap.c |  3 +-
 src/backend/access/transam/xloginsert.c |  3 +-
 src/backend/storage/buffer/bufmgr.c     | 70 +++++++++++++++++++------
 4 files changed, 61 insertions(+), 18 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 47360a3d3d8..3f37b294af6 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -230,7 +230,8 @@ extern void WaitReadBuffers(ReadBuffersOperation *operation);
 
 extern void ReleaseBuffer(Buffer buffer);
 extern void UnlockReleaseBuffer(Buffer buffer);
-extern bool BufferIsExclusiveLocked(Buffer buffer);
+extern bool BufferIsLockedByMe(Buffer buffer);
+extern bool BufferIsLockedByMeInMode(Buffer buffer, int mode);
 extern bool BufferIsDirty(Buffer buffer);
 extern void MarkBufferDirty(Buffer buffer);
 extern void IncrBufferRefCount(Buffer buffer);
diff --git a/src/backend/access/heap/visibilitymap.c b/src/backend/access/heap/visibilitymap.c
index 7306c16f05c..0414ce1945c 100644
--- a/src/backend/access/heap/visibilitymap.c
+++ b/src/backend/access/heap/visibilitymap.c
@@ -270,7 +270,8 @@ visibilitymap_set(Relation rel, BlockNumber heapBlk, Buffer heapBuf,
 	if (BufferIsValid(heapBuf) && BufferGetBlockNumber(heapBuf) != heapBlk)
 		elog(ERROR, "wrong heap buffer passed to visibilitymap_set");
 
-	Assert(!BufferIsValid(heapBuf) || BufferIsExclusiveLocked(heapBuf));
+	Assert(!BufferIsValid(heapBuf) ||
+		   BufferIsLockedByMeInMode(heapBuf, BUFFER_LOCK_EXCLUSIVE));
 
 	/* Check that we have the right VM page pinned */
 	if (!BufferIsValid(vmBuf) || BufferGetBlockNumber(vmBuf) != mapBlock)
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index c7571429e8e..496e0fa4ac6 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -258,7 +258,8 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	 */
 #ifdef USE_ASSERT_CHECKING
 	if (!(flags & REGBUF_NO_CHANGE))
-		Assert(BufferIsExclusiveLocked(buffer) && BufferIsDirty(buffer));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) &&
+			   BufferIsDirty(buffer));
 #endif
 
 	if (block_id >= max_registered_block_id)
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 594a8e5d512..350b35362ce 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -1065,7 +1065,7 @@ ZeroAndLockBuffer(Buffer buffer, ReadBufferMode mode, bool already_valid)
 		 * already valid.)
 		 */
 		if (!isLocalBuf)
-			LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_EXCLUSIVE);
+			LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
 
 		/* Set BM_VALID, terminate IO, and wake up any waiters */
 		if (isLocalBuf)
@@ -2822,7 +2822,7 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 		}
 
 		if (lock)
-			LWLockAcquire(BufferDescriptorGetContentLock(buf_hdr), LW_EXCLUSIVE);
+			LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
 		TerminateBufferIO(buf_hdr, false, BM_VALID, true, false);
 	}
@@ -2835,14 +2835,14 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 }
 
 /*
- * BufferIsExclusiveLocked
+ * BufferIsLockedByMe
  *
- *      Checks if buffer is exclusive-locked.
+ *      Checks if this backend has the buffer locked in any mode.
  *
  * Buffer must be pinned.
  */
 bool
-BufferIsExclusiveLocked(Buffer buffer)
+BufferIsLockedByMe(Buffer buffer)
 {
 	BufferDesc *bufHdr;
 
@@ -2855,9 +2855,49 @@ BufferIsExclusiveLocked(Buffer buffer)
 	}
 	else
 	{
+		bufHdr = GetBufferDescriptor(buffer - 1);
+		return LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr));
+	}
+}
+
+/*
+ * BufferIsLockedByMeInMode
+ *
+ *      Checks if this backend has the buffer locked in the specified mode.
+ *
+ * Buffer must be pinned.
+ */
+bool
+BufferIsLockedByMeInMode(Buffer buffer, int mode)
+{
+	BufferDesc *bufHdr;
+
+	Assert(BufferIsPinned(buffer));
+
+	if (BufferIsLocal(buffer))
+	{
+		/* Content locks are not maintained for local buffers. */
+		return true;
+	}
+	else
+	{
+		LWLockMode	lw_mode;
+
+		switch (mode)
+		{
+			case BUFFER_LOCK_EXCLUSIVE:
+				lw_mode = LW_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE:
+				lw_mode = LW_SHARED;
+				break;
+			default:
+				pg_unreachable();
+		}
+
 		bufHdr = GetBufferDescriptor(buffer - 1);
 		return LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-									LW_EXCLUSIVE);
+									lw_mode);
 	}
 }
 
@@ -2886,8 +2926,7 @@ BufferIsDirty(Buffer buffer)
 	else
 	{
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		Assert(LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-									LW_EXCLUSIVE));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 	}
 
 	return pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY;
@@ -2921,8 +2960,7 @@ MarkBufferDirty(Buffer buffer)
 	bufHdr = GetBufferDescriptor(buffer - 1);
 
 	Assert(BufferIsPinned(buffer));
-	Assert(LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-								LW_EXCLUSIVE));
+	Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 
 	/*
 	 * NB: We have to wait for the buffer header spinlock to be not held, as
@@ -3258,7 +3296,10 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 */
 		VALGRIND_MAKE_MEM_NOACCESS(BufHdrGetBlock(buf), BLCKSZ);
 
-		/* I'd better not still hold the buffer content lock */
+		/*
+		 * I'd better not still hold the buffer content lock. Can't use
+		 * BufferIsLockedByMe(), as that asserts the buffer is pinned.
+		 */
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
 		/* decrement the shared reference count */
@@ -5311,7 +5352,7 @@ FlushOneBuffer(Buffer buffer)
 
 	bufHdr = GetBufferDescriptor(buffer - 1);
 
-	Assert(LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr)));
+	Assert(BufferIsLockedByMe(buffer));
 
 	FlushBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 }
@@ -5402,7 +5443,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 
 	Assert(GetPrivateRefCount(buffer) > 0);
 	/* here, either share or exclusive lock is OK */
-	Assert(LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr)));
+	Assert(BufferIsLockedByMe(buffer));
 
 	/*
 	 * This routine might get called many times on the same page, if we are
@@ -5894,8 +5935,7 @@ IsBufferCleanupOK(Buffer buffer)
 	bufHdr = GetBufferDescriptor(buffer - 1);
 
 	/* caller must hold exclusive lock on buffer */
-	Assert(LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-								LW_EXCLUSIVE));
+	Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 
 	buf_state = LockBufHdr(bufHdr);
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v4-0006-bufmgr-Introduce-FlushUnlockedBuffer.patch (4.5K, ../../6kmid26do57ykqfpvq6iieniy4djsymhrypkjccazq5g4bbe6a@2y6owwv7qpex/7-v4-0006-bufmgr-Introduce-FlushUnlockedBuffer.patch)
  download | inline diff:
From 3a70f842d5b5dfd5e8e477bc9cea918efbb1b8ec Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 30 Jun 2025 14:30:38 -0400
Subject: [PATCH v4 6/6] bufmgr: Introduce FlushUnlockedBuffer

There were several copies of code locking a buffer, flushing its contents, and
unlocking the buffer. It seems worth centralizing that into a helper function.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/storage/buffer/bufmgr.c | 36 ++++++++++++++++-------------
 1 file changed, 20 insertions(+), 16 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 350b35362ce..ff961ec46d9 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -533,6 +533,8 @@ static inline BufferDesc *BufferAlloc(SMgrRelation smgr,
 static bool AsyncReadBuffers(ReadBuffersOperation *operation, int *nblocks_progress);
 static void CheckReadBuffersOperation(ReadBuffersOperation *operation, bool is_complete);
 static Buffer GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context);
+static void FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
+								IOObject io_object, IOContext io_context);
 static void FlushBuffer(BufferDesc *buf, SMgrRelation reln,
 						IOObject io_object, IOContext io_context);
 static void FindAndDropRelationBuffers(RelFileLocator rlocator,
@@ -3965,11 +3967,8 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	 * buffer is clean by the time we've locked it.)
 	 */
 	PinBuffer_Locked(bufHdr);
-	LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
 
-	FlushBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-
-	LWLockRelease(BufferDescriptorGetContentLock(bufHdr));
+	FlushUnlockedBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 
 	tag = bufHdr->tag;
 
@@ -4417,6 +4416,19 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	error_context_stack = errcallback.previous;
 }
 
+/*
+ * Convenience wrapper around FlushBuffer() that locks/unlocks the buffer
+ * before/after calling FlushBuffer().
+ */
+static void
+FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
+					IOObject io_object, IOContext io_context)
+{
+	LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
+	LWLockRelease(BufferDescriptorGetContentLock(buf));
+}
+
 /*
  * RelationGetNumberOfBlocksInFork
  *		Determines the current number of pages in the specified relation fork.
@@ -5001,9 +5013,7 @@ FlushRelationBuffers(Relation rel)
 			(buf_state & (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 		{
 			PinBuffer_Locked(bufHdr);
-			LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
-			FlushBuffer(bufHdr, srel, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-			LWLockRelease(BufferDescriptorGetContentLock(bufHdr));
+			FlushUnlockedBuffer(bufHdr, srel, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 			UnpinBuffer(bufHdr);
 		}
 		else
@@ -5098,9 +5108,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 			(buf_state & (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 		{
 			PinBuffer_Locked(bufHdr);
-			LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
-			FlushBuffer(bufHdr, srelent->srel, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-			LWLockRelease(BufferDescriptorGetContentLock(bufHdr));
+			FlushUnlockedBuffer(bufHdr, srelent->srel, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 			UnpinBuffer(bufHdr);
 		}
 		else
@@ -5326,9 +5334,7 @@ FlushDatabaseBuffers(Oid dbid)
 			(buf_state & (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 		{
 			PinBuffer_Locked(bufHdr);
-			LWLockAcquire(BufferDescriptorGetContentLock(bufHdr), LW_SHARED);
-			FlushBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-			LWLockRelease(BufferDescriptorGetContentLock(bufHdr));
+			FlushUnlockedBuffer(bufHdr, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 			UnpinBuffer(bufHdr);
 		}
 		else
@@ -6615,10 +6621,8 @@ EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 	/* If it was dirty, try to clean it once. */
 	if (buf_state & BM_DIRTY)
 	{
-		LWLockAcquire(BufferDescriptorGetContentLock(desc), LW_SHARED);
-		FlushBuffer(desc, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
+		FlushUnlockedBuffer(desc, NULL, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 		*buffer_flushed = true;
-		LWLockRelease(BufferDescriptorGetContentLock(desc));
 	}
 
 	/* This will return false if it becomes dirty or someone else pins it. */
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-10-04 07:05     ` Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Matthias van de Meent @ 2025-10-04 07:05 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Tue, 23 Sept 2025 at 00:14, Andres Freund <andres@anarazel.de> wrote:
> On 2025-09-15 19:05:37 -0400, Andres Freund wrote:
> > Here are the first few cleaned up patches implementing the above steps, as
> > well as some cleanups.  I included a commit from another thread, as it
> > conflicts with these changes, and we really should apply it - and it's
> > arguably required to make the changes viable, as it removes one more use of
> > PinBuffer_Locked().
> >
> > Another change included is to not return the buffer with the spinlock held
> > from StrategyGetBuffer(), and instead pin the buffer in freelist.c. The reason
> > for that is to reduce the most common PinBuffer_locked() call. By definition
> > PinBuffer_locked() will become a bit slower due to 0003. But even without 0003
> > it 0002 is faster than master. And the previous approach also just seems
> > pretty unclean.   I don't love that it requires the new TrackNewBufferPin(),
> > but I don't really have a better idea.
> >
> > I invite particular attention to the commit message for 0003 as well as the
> > comment changes in buf_internals.h within.
>
> Robert looked at the patches while we were chatting, and I addressed his
> feedback in this new version.

I like these changes, and have some minor comments:

0001 ensures that ReadRecentBuffer increments the usage counter, which
someone who uses an access strategy may want to prevent. I know this
isn't exactly new behaviour, but something I noticed anyway. Apart
from that observation, LGTM

0002 has a FIXME in a comment in GetVictimBuffer. Assuming it's about
the comment itself needing updates, how about:

+     * Ensure, before we pin a victim buffer, that there's a free refcount
+     * entry, and a resource owner slot for the pin.

Again, LGTM.

0003's UnlockBufHdrExt:
This is implemented with CAS, even when we only want to change bits we
know the state of (or could know, if we spent the effort).
Given its inline nature, wouldn't it be better to use atomic_sub
instructions? Or is this to handle cases where the bits we want to
(un)set might be (un)set by a concurrent process?
If the latter, could we specialize this to do a single atomic_sub
whenever we want to change state bits that we know can be only changed
whilst holding the spinlock?

0004: LGTM

0005: LGTM

0006: LGTM

Kind regards,

Matthias van de Meent





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
@ 2025-10-06 22:55       ` Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-10-06 22:55 UTC (permalink / raw)
  To: Matthias van de Meent <boekewurm+postgres@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-10-04 09:05:45 +0200, Matthias van de Meent wrote:
> On Tue, 23 Sept 2025 at 00:14, Andres Freund <andres@anarazel.de> wrote:
> > On 2025-09-15 19:05:37 -0400, Andres Freund wrote:
> > > Here are the first few cleaned up patches implementing the above steps, as
> > > well as some cleanups.  I included a commit from another thread, as it
> > > conflicts with these changes, and we really should apply it - and it's
> > > arguably required to make the changes viable, as it removes one more use of
> > > PinBuffer_Locked().
> > >
> > > Another change included is to not return the buffer with the spinlock held
> > > from StrategyGetBuffer(), and instead pin the buffer in freelist.c. The reason
> > > for that is to reduce the most common PinBuffer_locked() call. By definition
> > > PinBuffer_locked() will become a bit slower due to 0003. But even without 0003
> > > it 0002 is faster than master. And the previous approach also just seems
> > > pretty unclean.   I don't love that it requires the new TrackNewBufferPin(),
> > > but I don't really have a better idea.
> > >
> > > I invite particular attention to the commit message for 0003 as well as the
> > > comment changes in buf_internals.h within.
> >
> > Robert looked at the patches while we were chatting, and I addressed his
> > feedback in this new version.
> 
> I like these changes, and have some minor comments:

Thank for reviewing!


> 0001 ensures that ReadRecentBuffer increments the usage counter, which
> someone who uses an access strategy may want to prevent. I know this
> isn't exactly new behaviour, but something I noticed anyway. Apart
> from that observation, LGTM

Are you proposing to change behaviour? Right now ReadRecentBuffer doesn't even
accept a strategy, so I don't really see this as something that needs to be
tackled at this point.

I'm not sure I see any real use cases for ReadRecentBuffer() that would
benefit from a strategy, but I very well might just not be thinking wide
enough.


> 0002 has a FIXME in a comment in GetVictimBuffer. Assuming it's about
> the comment itself needing updates

Indeed.


> , how about:
> 
> +     * Ensure, before we pin a victim buffer, that there's a free refcount
> +     * entry, and a resource owner slot for the pin.
> 
> Again, LGTM.

WFM.


> 0003's UnlockBufHdrExt:
> This is implemented with CAS, even when we only want to change bits we
> know the state of (or could know, if we spent the effort).
> Given its inline nature, wouldn't it be better to use atomic_sub
> instructions? Or is this to handle cases where the bits we want to
> (un)set might be (un)set by a concurrent process?

Yes, it's to handle concurrent changes to the buffer state.

> If the latter, could we specialize this to do a single atomic_sub
> whenever we want to change state bits that we know can be only changed
> whilst holding the spinlock?

We probably could optimize some cases as an atomic-sub, some others as an
atomic-and and others again as an atomic-or. The latter to however are
implemented as a CAS on x86 anyway...

After 0004 I don't think any of the paths using this are actually particularly
hot, so I'm somewhat doubtful it's worth to try to optimize this too much. If
there are hot paths, we really should try to work towards not even needing the
buffer header spinlock, that has a bigger impact that improving the code for
unlocking the buffer header...

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-10-07 16:40         ` Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Matthias van de Meent @ 2025-10-07 16:40 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On Tue, 7 Oct 2025 at 00:55, Andres Freund <andres@anarazel.de> wrote:
> On 2025-10-04 09:05:45 +0200, Matthias van de Meent wrote:
> > 0001 ensures that ReadRecentBuffer increments the usage counter, which
> > someone who uses an access strategy may want to prevent. I know this
> > isn't exactly new behaviour, but something I noticed anyway. Apart
> > from that observation, LGTM
>
> Are you proposing to change behaviour?

Eventually, yes, but not necessarily now or with this patchset.

> Right now ReadRecentBuffer doesn't even
> accept a strategy, so I don't really see this as something that needs to be
> tackled at this point.
>
> I'm not sure I see any real use cases for ReadRecentBuffer() that would
> benefit from a strategy, but I very well might just not be thinking wide
> enough.

I think it's rather strange that there is no Extended variant of
ReadRecentBuffer, like how there is a ReadBufferExtended for
ReadBuffer. Yes, ReadRecentBuffer has more arguments to fill and so
has a smaller difference versus ReadBufferExtended, but it's no
complete replacement for ReadBuffer[Ext] when you're aware of a recent
buffer of the page.

I'm not saying it will definitely happen, but I could see that e.g.
amcheck might want to keep track of buffer IDs of recent heap pages it
accessed to verify index's results, without holding a pin on all the
pages; and instead using ReadRecentBuffer[Extended] with a
BufferAccessStrategy to allow re-acquiring the buffer pin without
blowing out shared buffers or making parts of the pool take forever to
evict again.

> > 0003's UnlockBufHdrExt:
> > This is implemented with CAS, even when we only want to change bits we
> > know the state of (or could know, if we spent the effort).
> > Given its inline nature, wouldn't it be better to use atomic_sub
> > instructions? Or is this to handle cases where the bits we want to
> > (un)set might be (un)set by a concurrent process?
>
> Yes, it's to handle concurrent changes to the buffer state.
>
> > If the latter, could we specialize this to do a single atomic_sub
> > whenever we want to change state bits that we know can be only changed
> > whilst holding the spinlock?
>
> We probably could optimize some cases as an atomic-sub, some others as an
> atomic-and and others again as an atomic-or. The latter to however are
> implemented as a CAS on x86 anyway...
>
> After 0004 I don't think any of the paths using this are actually particularly
> hot, so I'm somewhat doubtful it's worth to try to optimize this too much. If
> there are hot paths, we really should try to work towards not even needing the
> buffer header spinlock, that has a bigger impact that improving the code for
> unlocking the buffer header...

Fair enough; I guess we'll see if further optimization would have much
impact once this all has been committed.

Kind regards,

Matthias van de Meent
Databricks





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
@ 2025-10-09 20:35           ` Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-10-09 20:35 UTC (permalink / raw)
  To: Matthias van de Meent <boekewurm+postgres@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

I pushed a few commits from this patchset after Matthias' review
(thanks!). Unfortunately in 5e899859287 I missed that the valgrind annotations
would not be done anymore for the buffers returned by
StrategyGetBuffer(). Which turned skink red.

The attached 0001 patch centralizes the valgrind initialization in
TrackNewBufferPin(), which 5e899859287 had added. The nice side effect of that
is that there are fewer VALGRIND_MAKE_MEM_DEFINED() calls than before. The
naming isn't the perfect match, but it seems fine to me.


0002 is a new version of "Allow some buffer state modifications while holding
header lock", mainly with a fair bit of comment polishing around BufferDesc
and one small oversight fixed (didn't update a buffer_state variable in one
place).

Greetings,

Andres Freund

Attachments:

  [text/x-diff] v5-0001-bufmgr-Fix-valgrind-checking-for-buffers-pinned-i.patch (3.1K, ../../3je3ahgf7rrmmurxo6hnlhg5d3ffwfrtjwjxd6jm5srlv5iebp@vxqk5qtgmowr/2-v5-0001-bufmgr-Fix-valgrind-checking-for-buffers-pinned-i.patch)
  download | inline diff:
From 353b4d30f9b8cf014d2a2b851ae989301fd894ab Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 9 Oct 2025 16:27:08 -0400
Subject: [PATCH v5 1/3] bufmgr: Fix valgrind checking for buffers pinned in
 StrategyGetBuffer()

In 5e899859287 I made StrategyGetBuffer() pin buffers with a single CAS,
instead of using PinBuffer_Locked(). Unfortunately I missed that
PinBuffer_Locked() marked the page as defined for valgrind.

Fix this oversight by centralizing the valgrind initialization into
TrackNewBufferPin(), which also allows us to reduce the number of places doing
VALGRIND_MAKE_MEM_DEFINED.

Per buildfarm animal skink and Amit Langote.

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
Discussion: https://postgr.es/m/CA+HiwqGKJ6nEXEPQW7EpykVsEtzxp5-up_xhtcUAkWFtATVQvQ@mail.gmail.com
---
 src/backend/storage/buffer/bufmgr.c | 32 ++++++++++++++---------------
 1 file changed, 16 insertions(+), 16 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index ca62b9cdcf0..edf17ce3ea1 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3113,15 +3113,6 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 				result = (buf_state & BM_VALID) != 0;
 
 				TrackNewBufferPin(b);
-
-				/*
-				 * Assume that we acquired a buffer pin for the purposes of
-				 * Valgrind buffer client checks (even in !result case) to
-				 * keep things simple.  Buffers that are unsafe to access are
-				 * not generally guaranteed to be marked undefined or
-				 * non-accessible in any case.
-				 */
-				VALGRIND_MAKE_MEM_DEFINED(BufHdrGetBlock(buf), BLCKSZ);
 				break;
 			}
 		}
@@ -3186,13 +3177,6 @@ PinBuffer_Locked(BufferDesc *buf)
 	 */
 	Assert(GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf), false) == NULL);
 
-	/*
-	 * Buffer can't have a preexisting pin, so mark its page as defined to
-	 * Valgrind (this is similar to the PinBuffer() case where the backend
-	 * doesn't already have a buffer pin)
-	 */
-	VALGRIND_MAKE_MEM_DEFINED(BufHdrGetBlock(buf), BLCKSZ);
-
 	/*
 	 * Since we hold the buffer spinlock, we can update the buffer state and
 	 * release the lock in one operation.
@@ -3320,6 +3304,10 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	}
 }
 
+/*
+ * Set up backend-local tracking of a buffer pinned the first time by this
+ * backend.
+ */
 inline void
 TrackNewBufferPin(Buffer buf)
 {
@@ -3329,6 +3317,18 @@ TrackNewBufferPin(Buffer buf)
 	ref->refcount++;
 
 	ResourceOwnerRememberBuffer(CurrentResourceOwner, buf);
+
+	/*
+	 * This is the first pin for this page by this backend, mark its page as
+	 * defined to valgrind. While the page contents might not actually be
+	 * valid yet, we don't currently guarantee that such pages are marked
+	 * undefined or non-accessible.
+	 *
+	 * It's not necessarily the prettiest to do this here, but otherwise we'd
+	 * need this block of code in multiple places.
+	 */
+	VALGRIND_MAKE_MEM_DEFINED(BufHdrGetBlock(GetBufferDescriptor(buf - 1)),
+							  BLCKSZ);
 }
 
 #define ST_SORT sort_checkpoint_bufferids
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v5-0002-bufmgr-Allow-some-buffer-state-modifications-whil.patch (31.3K, ../../3je3ahgf7rrmmurxo6hnlhg5d3ffwfrtjwjxd6jm5srlv5iebp@vxqk5qtgmowr/3-v5-0002-bufmgr-Allow-some-buffer-state-modifications-whil.patch)
  download | inline diff:
From d10f731d7e74968c25e90233a60fc4261c445a66 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 15 Sep 2025 17:06:48 -0400
Subject: [PATCH v5 2/3] bufmgr: Allow some buffer state modifications while
 holding header lock

Until now BufferDesc.state was not allowed to be modified while the buffer
header spinlock was held. This meant that operations like unpinning buffers
needed to use a CAS loop, waiting for the buffer header spinlock to be
released before updating.

The benefit of that restriction is that it allowed us to unlock the buffer
header spinlock with just a write barrier and an unlocked write (instead of a
full atomic operation). That was important to avoid regressions in
48354581a49c. However, since then the hottest buffer header spinlock uses have
been replaced with atomic operations (in particular, the most common use of
PinBuffer_Locked(), in GetVictimBuffer() (formerly in BufferAlloc()), has been
removed in 5e899859287).

This change will allow, in a subsequent commit, to release buffer pins with a
single atomic-sub operation. This previously was not possible while such
operations were not allowed while the buffer header spinlock was held, as an
atomic-sub would not have allowed a race-free check for the buffer header lock
being held.

Using atomic-sub to unpin buffers is a nice scalability win, however it is not
the primary motivation for this change (although it would be sufficient). The
primary motivation is that we would like to merge the buffer content lock into
BufferDesc.state, which will result in more frequent changes of the state
variable, which in some situations can cause a performance regression, due to
an increased CAS failure rate when unpinning buffers.  The regression entirely
vanishes when using atomic-sub.

Naively implementing this would require putting CAS loops in every place
modifying the buffer state while holding the buffer header lock. To avoid
that, introduce UnlockBufHdrExt(), which can set/add flags as well as the
refcount, together with releasing the lock.

Reviewed-by: Robert Haas <robertmhaas@gmail.com>
Reviewed-by: Matthias van de Meent <boekewurm+postgres@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           | 119 +++++++---
 src/backend/storage/buffer/bufmgr.c           | 203 ++++++++++--------
 src/backend/storage/buffer/freelist.c         |   2 +
 contrib/pg_buffercache/pg_buffercache_pages.c |   7 +-
 contrib/pg_prewarm/autoprewarm.c              |   2 +-
 src/test/modules/test_aio/test_aio.c          |  10 +-
 6 files changed, 224 insertions(+), 119 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index c1206a46aba..5400c56a965 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -211,28 +211,36 @@ BufMappingPartitionLockByIndex(uint32 index)
 /*
  *	BufferDesc -- shared descriptor/state data for a single shared buffer.
  *
- * Note: Buffer header lock (BM_LOCKED flag) must be held to examine or change
- * tag, state or wait_backend_pgprocno fields.  In general, buffer header lock
- * is a spinlock which is combined with flags, refcount and usagecount into
- * single atomic variable.  This layout allow us to do some operations in a
- * single atomic operation, without actually acquiring and releasing spinlock;
- * for instance, increase or decrease refcount.  buf_id field never changes
- * after initialization, so does not need locking.  The LWLock can take care
- * of itself.  The buffer header lock is *not* used to control access to the
- * data in the buffer!
+ * The state of the buffer is controlled by the, drumroll, state variable. It
+ * only may be modified using atomic operations.  The state variable combines
+ * various flags, the buffer's refcount and usage count. See comment above
+ * BUF_REFCOUNT_BITS for details about the division.  This layout allow us to
+ * do some operations in a single atomic operation, without actually acquiring
+ * and releasing the spinlock; for instance, increasing or decreasing the
+ * refcount.
  *
- * It's assumed that nobody changes the state field while buffer header lock
- * is held.  Thus buffer header lock holder can do complex updates of the
- * state variable in single write, simultaneously with lock release (cleaning
- * BM_LOCKED flag).  On the other hand, updating of state without holding
- * buffer header lock is restricted to CAS, which ensures that BM_LOCKED flag
- * is not set.  Atomic increment/decrement, OR/AND etc. are not allowed.
+ * One of the aforementioned flags is BM_LOCKED, used to implement the buffer
+ * header lock. See the following paragraphs, as well as the documentation for
+ * individual fields, for more details.
  *
- * An exception is that if we have the buffer pinned, its tag can't change
- * underneath us, so we can examine the tag without locking the buffer header.
- * Also, in places we do one-time reads of the flags without bothering to
- * lock the buffer header; this is generally for situations where we don't
- * expect the flag bit being tested to be changing.
+ * The identity of the buffer (BufferDesc.tag) can only be changed by the
+ * backend holding the buffer header lock.
+ *
+ * If the lock is held by another backend, neither additional buffer pins may
+ * be established (we would like to relax this eventually), nor can flags be
+ * set/cleared. These operations either need to acquire the buffer header
+ * spinlock, or need to use a CAS loop, waiting for the lock to be released if
+ * it is held.  However, existing buffer pins may be released while the buffer
+ * header spinlock is held, using an atomic subtraction.
+ *
+ * The LWLock can take care of itself.  The buffer header lock is *not* used
+ * to control access to the data in the buffer!
+ *
+ * If we have the buffer pinned, its tag can't change underneath us, so we can
+ * examine the tag without locking the buffer header.  Also, in places we do
+ * one-time reads of the flags without bothering to lock the buffer header;
+ * this is generally for situations where we don't expect the flag bit being
+ * tested to be changing.
  *
  * We can't physically remove items from a disk page if another backend has
  * the buffer pinned.  Hence, a backend may need to wait for all other pins
@@ -256,13 +264,29 @@ BufMappingPartitionLockByIndex(uint32 index)
  */
 typedef struct BufferDesc
 {
-	BufferTag	tag;			/* ID of page contained in buffer */
-	int			buf_id;			/* buffer's index number (from 0) */
+	/*
+	 * ID of page contained in buffer. The buffer header spinlock needs to be
+	 * held to modify this field.
+	 */
+	BufferTag	tag;
 
-	/* state of the tag, containing flags, refcount and usagecount */
+	/*
+	 * Buffer's index number (from 0). The field never changes after
+	 * initialization, so does not need locking.
+	 */
+	int			buf_id;
+
+	/*
+	 * State of the buffer, containing flags, refcount and usagecount. See
+	 * BUF_* and BM_* defines at the top of this file.
+	 */
 	pg_atomic_uint32 state;
 
-	int			wait_backend_pgprocno;	/* backend of pin-count waiter */
+	/*
+	 * Backend of pin-count waiter. The buffer header spinlock needs to be
+	 * held to modify this field.
+	 */
+	int			wait_backend_pgprocno;
 
 	PgAioWaitRef io_wref;		/* set iff AIO is in progress */
 	LWLock		content_lock;	/* to lock access to buffer contents */
@@ -364,11 +388,52 @@ BufferDescriptorGetContentLock(const BufferDesc *bdesc)
  */
 extern uint32 LockBufHdr(BufferDesc *desc);
 
+/*
+ * Unlock the buffer header.
+ *
+ * This can only be used if the caller did not modify BufferDesc.state. To
+ * set/unset flag bits or change the refcount use UnlockBufHdrExt().
+ */
 static inline void
-UnlockBufHdr(BufferDesc *desc, uint32 buf_state)
+UnlockBufHdr(BufferDesc *desc)
 {
-	pg_write_barrier();
-	pg_atomic_write_u32(&desc->state, buf_state & (~BM_LOCKED));
+	Assert(pg_atomic_read_u32(&desc->state) & BM_LOCKED);
+
+	pg_atomic_fetch_sub_u32(&desc->state, BM_LOCKED);
+}
+
+/*
+ * Unlock the buffer header, while atomically adding the flags in set_bits,
+ * unsetting the ones in unset_bits and changing the refcount by
+ * refcount_change.
+ *
+ * Note that this approach would not work for usagecount, since we need to cap
+ * the usagecount at BM_MAX_USAGE_COUNT.
+ */
+static inline uint32
+UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
+				uint32 set_bits, uint32 unset_bits,
+				int refcount_change)
+{
+	for (;;)
+	{
+		uint32		buf_state = old_buf_state;
+
+		Assert(buf_state & BM_LOCKED);
+
+		buf_state |= set_bits;
+		buf_state &= ~unset_bits;
+		buf_state &= ~BM_LOCKED;
+
+		if (refcount_change != 0)
+			buf_state += BUF_REFCOUNT_ONE * refcount_change;
+
+		if (pg_atomic_compare_exchange_u32(&desc->state, &old_buf_state,
+										   buf_state))
+		{
+			return old_buf_state;
+		}
+	}
 }
 
 extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index edf17ce3ea1..42e1f50c7c8 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -1996,6 +1996,7 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	Buffer		victim_buffer;
 	BufferDesc *victim_buf_hdr;
 	uint32		victim_buf_state;
+	uint32		set_bits = 0;
 
 	/* Make sure we will have room to remember the buffer pin */
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2122,11 +2123,12 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	 * checkpoints, except for their "init" forks, which need to be treated
 	 * just like permanent relations.
 	 */
-	victim_buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
+	set_bits |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 	if (relpersistence == RELPERSISTENCE_PERMANENT || forkNum == INIT_FORKNUM)
-		victim_buf_state |= BM_PERMANENT;
+		set_bits |= BM_PERMANENT;
 
-	UnlockBufHdr(victim_buf_hdr, victim_buf_state);
+	UnlockBufHdrExt(victim_buf_hdr, victim_buf_state,
+					set_bits, 0, 0);
 
 	LWLockRelease(newPartitionLock);
 
@@ -2166,9 +2168,7 @@ InvalidateBuffer(BufferDesc *buf)
 	/* Save the original buffer tag before dropping the spinlock */
 	oldTag = buf->tag;
 
-	buf_state = pg_atomic_read_u32(&buf->state);
-	Assert(buf_state & BM_LOCKED);
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdr(buf);
 
 	/*
 	 * Need to compute the old tag's hashcode and partition lock ID. XXX is it
@@ -2192,7 +2192,7 @@ retry:
 	/* If it's changed while we were waiting for lock, do nothing */
 	if (!BufferTagsEqual(&buf->tag, &oldTag))
 	{
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		LWLockRelease(oldPartitionLock);
 		return;
 	}
@@ -2209,7 +2209,7 @@ retry:
 	 */
 	if (BUF_STATE_GET_REFCOUNT(buf_state) != 0)
 	{
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		LWLockRelease(oldPartitionLock);
 		/* safety check: should definitely not be our *own* pin */
 		if (GetPrivateRefCount(BufferDescriptorGetBuffer(buf)) > 0)
@@ -2224,8 +2224,11 @@ retry:
 	 */
 	oldFlags = buf_state & BUF_FLAG_MASK;
 	ClearBufferTag(&buf->tag);
-	buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
-	UnlockBufHdr(buf, buf_state);
+
+	UnlockBufHdrExt(buf, buf_state,
+					0,
+					BUF_FLAG_MASK | BUF_USAGECOUNT_MASK,
+					0);
 
 	/*
 	 * Remove the buffer from the lookup hashtable, if it was in there.
@@ -2285,7 +2288,7 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 	{
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 
-		UnlockBufHdr(buf_hdr, buf_state);
+		UnlockBufHdr(buf_hdr);
 		LWLockRelease(partition_lock);
 
 		return false;
@@ -2299,8 +2302,10 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 	 * tag (see e.g. FlushDatabaseBuffers()).
 	 */
 	ClearBufferTag(&buf_hdr->tag);
-	buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
-	UnlockBufHdr(buf_hdr, buf_state);
+	UnlockBufHdrExt(buf_hdr, buf_state,
+					0,
+					BUF_FLAG_MASK | BUF_USAGECOUNT_MASK,
+					0);
 
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 
@@ -2309,6 +2314,7 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 
 	LWLockRelease(partition_lock);
 
+	buf_state = pg_atomic_read_u32(&buf_hdr->state);
 	Assert(!(buf_state & (BM_DIRTY | BM_VALID | BM_TAG_VALID)));
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u32(&buf_hdr->state)) > 0);
@@ -2399,7 +2405,7 @@ again:
 			/* Read the LSN while holding buffer header lock */
 			buf_state = LockBufHdr(buf_hdr);
 			lsn = BufferGetLSN(buf_hdr);
-			UnlockBufHdr(buf_hdr, buf_state);
+			UnlockBufHdr(buf_hdr);
 
 			if (XLogNeedsFlush(lsn)
 				&& StrategyRejectBuffer(strategy, buf_hdr, from_ring))
@@ -2744,15 +2750,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 */
 			do
 			{
-				uint32		buf_state = LockBufHdr(existing_hdr);
-
-				buf_state &= ~BM_VALID;
-				UnlockBufHdr(existing_hdr, buf_state);
+				pg_atomic_fetch_and_u32(&existing_hdr->state, ~BM_VALID);
 			} while (!StartBufferIO(existing_hdr, true, false));
 		}
 		else
 		{
 			uint32		buf_state;
+			uint32		set_bits = 0;
 
 			buf_state = LockBufHdr(victim_buf_hdr);
 
@@ -2762,11 +2766,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 
 			victim_buf_hdr->tag = tag;
 
-			buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
+			set_bits |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 			if (bmr.relpersistence == RELPERSISTENCE_PERMANENT || fork == INIT_FORKNUM)
-				buf_state |= BM_PERMANENT;
+				set_bits |= BM_PERMANENT;
 
-			UnlockBufHdr(victim_buf_hdr, buf_state);
+			UnlockBufHdrExt(victim_buf_hdr, buf_state,
+							set_bits, 0,
+							0);
 
 			LWLockRelease(partition_lock);
 
@@ -2959,6 +2965,10 @@ MarkBufferDirty(Buffer buffer)
 	Assert(BufferIsPinned(buffer));
 	Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 
+	/*
+	 * NB: We have to wait for the buffer header spinlock to be not held, as
+	 * TerminateBufferIO() relies on the spinlock.
+	 */
 	old_buf_state = pg_atomic_read_u32(&bufHdr->state);
 	for (;;)
 	{
@@ -3083,6 +3093,10 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 			if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
 				return false;
 
+			/*
+			 * We're not allowed to increase the refcount while the buffer
+			 * header spinlock is held. Wait for the lock to be released.
+			 */
 			if (old_buf_state & BM_LOCKED)
 				old_buf_state = WaitBufHdrUnlocked(buf);
 
@@ -3169,7 +3183,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	uint32		buf_state;
+	uint32		old_buf_state;
 
 	/*
 	 * As explained, We don't expect any preexisting pins. That allows us to
@@ -3181,10 +3195,10 @@ PinBuffer_Locked(BufferDesc *buf)
 	 * Since we hold the buffer spinlock, we can update the buffer state and
 	 * release the lock in one operation.
 	 */
-	buf_state = pg_atomic_read_u32(&buf->state);
-	Assert(buf_state & BM_LOCKED);
-	buf_state += BUF_REFCOUNT_ONE;
-	UnlockBufHdr(buf, buf_state);
+	old_buf_state = pg_atomic_read_u32(&buf->state);
+
+	UnlockBufHdrExt(buf, old_buf_state,
+					0, 0, 1);
 
 	TrackNewBufferPin(BufferDescriptorGetBuffer(buf));
 }
@@ -3219,12 +3233,13 @@ WakePinCountWaiter(BufferDesc *buf)
 		/* we just released the last pin other than the waiter's */
 		int			wait_backend_pgprocno = buf->wait_backend_pgprocno;
 
-		buf_state &= ~BM_PIN_COUNT_WAITER;
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdrExt(buf, buf_state,
+						0, BM_PIN_COUNT_WAITER,
+						0);
 		ProcSendSignal(wait_backend_pgprocno);
 	}
 	else
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 }
 
 /*
@@ -3280,6 +3295,10 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 *
 		 * Since buffer spinlock holder can update status using just write,
 		 * it's not safe to use atomic decrement here; thus use a CAS loop.
+		 *
+		 * TODO: The above requirement does not hold anymore, in a future
+		 * commit this will be rewritten to release the pin in a single atomic
+		 * operation.
 		 */
 		old_buf_state = pg_atomic_read_u32(&buf->state);
 		for (;;)
@@ -3393,6 +3412,7 @@ BufferSync(int flags)
 	for (buf_id = 0; buf_id < NBuffers; buf_id++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
+		uint32		set_bits = 0;
 
 		/*
 		 * Header spinlock is enough to examine BM_DIRTY, see comment in
@@ -3404,7 +3424,7 @@ BufferSync(int flags)
 		{
 			CkptSortItem *item;
 
-			buf_state |= BM_CHECKPOINT_NEEDED;
+			set_bits = BM_CHECKPOINT_NEEDED;
 
 			item = &CkptBufferIds[num_to_scan++];
 			item->buf_id = buf_id;
@@ -3414,7 +3434,9 @@ BufferSync(int flags)
 			item->blockNum = bufHdr->tag.blockNum;
 		}
 
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdrExt(bufHdr, buf_state,
+						set_bits, 0,
+						0);
 
 		/* Check for barrier events in case NBuffers is large. */
 		if (ProcSignalBarrierPending)
@@ -3953,14 +3975,14 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	else if (skip_recently_used)
 	{
 		/* Caller told us not to write recently-used buffers */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return result;
 	}
 
 	if (!(buf_state & BM_VALID) || !(buf_state & BM_DIRTY))
 	{
 		/* It's clean, so nothing to do */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return result;
 	}
 
@@ -4329,8 +4351,9 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	recptr = BufferGetLSN(buf);
 
 	/* To check if block content changes while flushing. - vadim 01/17/97 */
-	buf_state &= ~BM_JUST_DIRTIED;
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdrExt(buf, buf_state,
+					0, BM_JUST_DIRTIED,
+					0);
 
 	/*
 	 * Force XLOG flush up to buffer's LSN.  This implements the basic WAL
@@ -4506,7 +4529,6 @@ BufferGetLSNAtomic(Buffer buffer)
 	char	   *page = BufferGetPage(buffer);
 	BufferDesc *bufHdr;
 	XLogRecPtr	lsn;
-	uint32		buf_state;
 
 	/*
 	 * If we don't need locking for correctness, fastpath out.
@@ -4519,9 +4541,9 @@ BufferGetLSNAtomic(Buffer buffer)
 	Assert(BufferIsPinned(buffer));
 
 	bufHdr = GetBufferDescriptor(buffer - 1);
-	buf_state = LockBufHdr(bufHdr);
+	LockBufHdr(bufHdr);
 	lsn = PageGetLSN(page);
-	UnlockBufHdr(bufHdr, buf_state);
+	UnlockBufHdr(bufHdr);
 
 	return lsn;
 }
@@ -4622,7 +4644,6 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
 	for (i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
 
 		/*
 		 * We can make this a tad faster by prechecking the buffer tag before
@@ -4643,7 +4664,7 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
 		if (!BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator.locator))
 			continue;
 
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 
 		for (j = 0; j < nforks; j++)
 		{
@@ -4656,7 +4677,7 @@ DropRelationBuffers(SMgrRelation smgr_reln, ForkNumber *forkNum,
 			}
 		}
 		if (j >= nforks)
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -4785,7 +4806,6 @@ DropRelationsAllBuffers(SMgrRelation *smgr_reln, int nlocators)
 	{
 		RelFileLocator *rlocator = NULL;
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -4819,11 +4839,11 @@ DropRelationsAllBuffers(SMgrRelation *smgr_reln, int nlocators)
 		if (rlocator == NULL)
 			continue;
 
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 		if (BufTagMatchesRelFileLocator(&bufHdr->tag, rlocator))
 			InvalidateBuffer(bufHdr);	/* releases spinlock */
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 
 	pfree(locators);
@@ -4853,7 +4873,6 @@ FindAndDropRelationBuffers(RelFileLocator rlocator, ForkNumber forkNum,
 		LWLock	   *bufPartitionLock;	/* buffer partition lock for it */
 		int			buf_id;
 		BufferDesc *bufHdr;
-		uint32		buf_state;
 
 		/* create a tag so we can lookup the buffer */
 		InitBufferTag(&bufTag, &rlocator, forkNum, curBlock);
@@ -4878,14 +4897,14 @@ FindAndDropRelationBuffers(RelFileLocator rlocator, ForkNumber forkNum,
 		 * evicted by some other backend loading blocks for a different
 		 * relation after we release lock on the BufMapping table.
 		 */
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 
 		if (BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator) &&
 			BufTagGetForkNum(&bufHdr->tag) == forkNum &&
 			bufHdr->tag.blockNum >= firstDelBlock)
 			InvalidateBuffer(bufHdr);	/* releases spinlock */
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -4913,7 +4932,6 @@ DropDatabaseBuffers(Oid dbid)
 	for (i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -4922,11 +4940,11 @@ DropDatabaseBuffers(Oid dbid)
 		if (bufHdr->tag.dbOid != dbid)
 			continue;
 
-		buf_state = LockBufHdr(bufHdr);
+		LockBufHdr(bufHdr);
 		if (bufHdr->tag.dbOid == dbid)
 			InvalidateBuffer(bufHdr);	/* releases spinlock */
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -5023,7 +5041,7 @@ FlushRelationBuffers(Relation rel)
 			UnpinBuffer(bufHdr);
 		}
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -5118,7 +5136,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 			UnpinBuffer(bufHdr);
 		}
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 
 	pfree(srels);
@@ -5344,7 +5362,7 @@ FlushDatabaseBuffers(Oid dbid)
 			UnpinBuffer(bufHdr);
 		}
 		else
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 	}
 }
 
@@ -5554,8 +5572,9 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 				PageSetLSN(page, lsn);
 		}
 
-		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdrExt(bufHdr, buf_state,
+						BM_DIRTY | BM_JUST_DIRTIED,
+						0, 0);
 
 		if (delayChkptFlags)
 			MyProc->delayChkptFlags &= ~DELAY_CHKPT_START;
@@ -5586,6 +5605,7 @@ UnlockBuffers(void)
 	if (buf)
 	{
 		uint32		buf_state;
+		uint32		unset_bits = 0;
 
 		buf_state = LockBufHdr(buf);
 
@@ -5595,9 +5615,11 @@ UnlockBuffers(void)
 		 */
 		if ((buf_state & BM_PIN_COUNT_WAITER) != 0 &&
 			buf->wait_backend_pgprocno == MyProcNumber)
-			buf_state &= ~BM_PIN_COUNT_WAITER;
+			unset_bits = BM_PIN_COUNT_WAITER;
 
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdrExt(buf, buf_state,
+						0, unset_bits,
+						0);
 
 		PinCountWaitBuf = NULL;
 	}
@@ -5715,6 +5737,7 @@ LockBufferForCleanup(Buffer buffer)
 	for (;;)
 	{
 		uint32		buf_state;
+		uint32		unset_bits = 0;
 
 		/* Try to acquire lock */
 		LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
@@ -5724,7 +5747,7 @@ LockBufferForCleanup(Buffer buffer)
 		if (BUF_STATE_GET_REFCOUNT(buf_state) == 1)
 		{
 			/* Successfully acquired exclusive lock with pincount 1 */
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 
 			/*
 			 * Emit the log message if recovery conflict on buffer pin was
@@ -5747,14 +5770,15 @@ LockBufferForCleanup(Buffer buffer)
 		/* Failed, so mark myself as waiting for pincount 1 */
 		if (buf_state & BM_PIN_COUNT_WAITER)
 		{
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 			LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 			elog(ERROR, "multiple backends attempting to wait for pincount 1");
 		}
 		bufHdr->wait_backend_pgprocno = MyProcNumber;
 		PinCountWaitBuf = bufHdr;
-		buf_state |= BM_PIN_COUNT_WAITER;
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdrExt(bufHdr, buf_state,
+						BM_PIN_COUNT_WAITER, 0,
+						0);
 		LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 
 		/* Wait to be signaled by UnpinBuffer() */
@@ -5816,8 +5840,11 @@ LockBufferForCleanup(Buffer buffer)
 		buf_state = LockBufHdr(bufHdr);
 		if ((buf_state & BM_PIN_COUNT_WAITER) != 0 &&
 			bufHdr->wait_backend_pgprocno == MyProcNumber)
-			buf_state &= ~BM_PIN_COUNT_WAITER;
-		UnlockBufHdr(bufHdr, buf_state);
+			unset_bits |= BM_PIN_COUNT_WAITER;
+
+		UnlockBufHdrExt(bufHdr, buf_state,
+						0, unset_bits,
+						0);
 
 		PinCountWaitBuf = NULL;
 		/* Loop back and try again */
@@ -5894,12 +5921,12 @@ ConditionalLockBufferForCleanup(Buffer buffer)
 	if (refcount == 1)
 	{
 		/* Successfully acquired exclusive lock with pincount 1 */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return true;
 	}
 
 	/* Failed, so release the lock */
-	UnlockBufHdr(bufHdr, buf_state);
+	UnlockBufHdr(bufHdr);
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 	return false;
 }
@@ -5946,11 +5973,11 @@ IsBufferCleanupOK(Buffer buffer)
 	if (BUF_STATE_GET_REFCOUNT(buf_state) == 1)
 	{
 		/* pincount is OK. */
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 		return true;
 	}
 
-	UnlockBufHdr(bufHdr, buf_state);
+	UnlockBufHdr(bufHdr);
 	return false;
 }
 
@@ -5988,7 +6015,7 @@ WaitIO(BufferDesc *buf)
 		 * clearing the wref while it's being read.
 		 */
 		iow = buf->io_wref;
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 
 		/* no IO in progress, we don't need to wait */
 		if (!(buf_state & BM_IO_IN_PROGRESS))
@@ -6056,7 +6083,7 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 
 		if (!(buf_state & BM_IO_IN_PROGRESS))
 			break;
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		if (nowait)
 			return false;
 		WaitIO(buf);
@@ -6067,12 +6094,13 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 	/* Check if someone else already did the I/O */
 	if (forInput ? (buf_state & BM_VALID) : !(buf_state & BM_DIRTY))
 	{
-		UnlockBufHdr(buf, buf_state);
+		UnlockBufHdr(buf);
 		return false;
 	}
 
-	buf_state |= BM_IO_IN_PROGRESS;
-	UnlockBufHdr(buf, buf_state);
+	UnlockBufHdrExt(buf, buf_state,
+					BM_IO_IN_PROGRESS, 0,
+					0);
 
 	ResourceOwnerRememberBufferIO(CurrentResourceOwner,
 								  BufferDescriptorGetBuffer(buf));
@@ -6105,28 +6133,31 @@ TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
 				  bool forget_owner, bool release_aio)
 {
 	uint32		buf_state;
+	uint32		unset_flag_bits = 0;
+	int			refcount_change = 0;
 
 	buf_state = LockBufHdr(buf);
 
 	Assert(buf_state & BM_IO_IN_PROGRESS);
-	buf_state &= ~BM_IO_IN_PROGRESS;
+	unset_flag_bits |= BM_IO_IN_PROGRESS;
 
 	/* Clear earlier errors, if this IO failed, it'll be marked again */
-	buf_state &= ~BM_IO_ERROR;
+	unset_flag_bits |= BM_IO_ERROR;
 
 	if (clear_dirty && !(buf_state & BM_JUST_DIRTIED))
-		buf_state &= ~(BM_DIRTY | BM_CHECKPOINT_NEEDED);
+		unset_flag_bits |= BM_DIRTY | BM_CHECKPOINT_NEEDED;
 
 	if (release_aio)
 	{
 		/* release ownership by the AIO subsystem */
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-		buf_state -= BUF_REFCOUNT_ONE;
+		refcount_change = -1;
 		pgaio_wref_clear(&buf->io_wref);
 	}
 
-	buf_state |= set_flag_bits;
-	UnlockBufHdr(buf, buf_state);
+	buf_state = UnlockBufHdrExt(buf, buf_state,
+								set_flag_bits, unset_flag_bits,
+								refcount_change);
 
 	if (forget_owner)
 		ResourceOwnerForgetBufferIO(CurrentResourceOwner,
@@ -6171,12 +6202,12 @@ AbortBufferIO(Buffer buffer)
 	if (!(buf_state & BM_VALID))
 	{
 		Assert(!(buf_state & BM_DIRTY));
-		UnlockBufHdr(buf_hdr, buf_state);
+		UnlockBufHdr(buf_hdr);
 	}
 	else
 	{
 		Assert(buf_state & BM_DIRTY);
-		UnlockBufHdr(buf_hdr, buf_state);
+		UnlockBufHdr(buf_hdr);
 
 		/* Issue notice if this is not the first failure... */
 		if (buf_state & BM_IO_ERROR)
@@ -6598,14 +6629,14 @@ EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 
 	if ((buf_state & BM_VALID) == 0)
 	{
-		UnlockBufHdr(desc, buf_state);
+		UnlockBufHdr(desc);
 		return false;
 	}
 
 	/* Check that it's not pinned already. */
 	if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
 	{
-		UnlockBufHdr(desc, buf_state);
+		UnlockBufHdr(desc);
 		return false;
 	}
 
@@ -6755,7 +6786,7 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
 		if ((buf_state & BM_VALID) == 0 ||
 			!BufTagMatchesRelFileLocator(&desc->tag, &rel->rd_locator))
 		{
-			UnlockBufHdr(desc, buf_state);
+			UnlockBufHdr(desc);
 			continue;
 		}
 
@@ -6853,13 +6884,15 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		 *
 		 * This pin is released again in TerminateBufferIO().
 		 */
-		buf_state += BUF_REFCOUNT_ONE;
 		buf_hdr->io_wref = io_ref;
 
 		if (is_temp)
+		{
+			buf_state += BUF_REFCOUNT_ONE;
 			pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		}
 		else
-			UnlockBufHdr(buf_hdr, buf_state);
+			UnlockBufHdrExt(buf_hdr, buf_state, 0, 0, 1);
 
 		/*
 		 * Ensure the content lock that prevents buffer modifications while
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 7fe34d3ef4c..53668b92400 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -266,6 +266,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 				break;
 			}
 
+			/* See equivalent code in PinBuffer() */
 			if (unlikely(local_buf_state & BM_LOCKED))
 			{
 				old_buf_state = WaitBufHdrUnlocked(buf);
@@ -700,6 +701,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 			|| BUF_STATE_GET_USAGECOUNT(local_buf_state) > 1)
 			break;
 
+		/* See equivalent code in PinBuffer() */
 		if (unlikely(local_buf_state & BM_LOCKED))
 		{
 			old_buf_state = WaitBufHdrUnlocked(buf);
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 3df04c98959..ab790533ff6 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -220,7 +220,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
 			else
 				fctx->record[i].isvalid = false;
 
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 		}
 	}
 
@@ -460,7 +460,6 @@ pg_buffercache_numa_pages(PG_FUNCTION_ARGS)
 		{
 			char	   *buffptr = (char *) BufferGetBlock(i + 1);
 			BufferDesc *bufHdr;
-			uint32		buf_state;
 			uint32		bufferid;
 			int32		page_num;
 			char	   *startptr_buff,
@@ -471,9 +470,9 @@ pg_buffercache_numa_pages(PG_FUNCTION_ARGS)
 			bufHdr = GetBufferDescriptor(i);
 
 			/* Lock each buffer header before inspecting. */
-			buf_state = LockBufHdr(bufHdr);
+			LockBufHdr(bufHdr);
 			bufferid = BufferDescriptorGetBuffer(bufHdr);
-			UnlockBufHdr(bufHdr, buf_state);
+			UnlockBufHdr(bufHdr);
 
 			/* start of the first page of this buffer */
 			startptr_buff = (char *) TYPEALIGN_DOWN(os_page_size, buffptr);
diff --git a/contrib/pg_prewarm/autoprewarm.c b/contrib/pg_prewarm/autoprewarm.c
index 8b68dafc261..5ba1240d51f 100644
--- a/contrib/pg_prewarm/autoprewarm.c
+++ b/contrib/pg_prewarm/autoprewarm.c
@@ -730,7 +730,7 @@ apw_dump_now(bool is_bgworker, bool dump_unlogged)
 			++num_blocks;
 		}
 
-		UnlockBufHdr(bufHdr, buf_state);
+		UnlockBufHdr(bufHdr);
 	}
 
 	snprintf(transient_dump_file_path, MAXPGPATH, "%s.tmp", AUTOPREWARM_FILE);
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index c55cf6c0aac..d7eadeab256 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -310,6 +310,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	BufferDesc *buf_hdr;
 	uint32		buf_state;
 	bool		was_pinned = false;
+	uint32		unset_bits = 0;
 
 	/* place buffer in shared buffers without erroring out */
 	buf = ReadBufferExtended(rel, MAIN_FORKNUM, blkno, RBM_ZERO_AND_LOCK, NULL);
@@ -334,12 +335,17 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (BUF_STATE_GET_REFCOUNT(buf_state) > 1)
 		was_pinned = true;
 	else
-		buf_state &= ~(BM_VALID | BM_DIRTY);
+		unset_bits |= BM_VALID | BM_DIRTY;
 
 	if (RelationUsesLocalBuffers(rel))
+	{
+		buf_state &= ~unset_bits;
 		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+	}
 	else
-		UnlockBufHdr(buf_hdr, buf_state);
+	{
+		UnlockBufHdrExt(buf_hdr, buf_state, 0, unset_bits, 0);
+	}
 
 	if (was_pinned)
 		elog(ERROR, "toy buffer %d was already pinned",
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v5-0003-bufmgr-Use-atomic-sub-for-unpinning-buffers.patch (2.3K, ../../3je3ahgf7rrmmurxo6hnlhg5d3ffwfrtjwjxd6jm5srlv5iebp@vxqk5qtgmowr/4-v5-0003-bufmgr-Use-atomic-sub-for-unpinning-buffers.patch)
  download | inline diff:
From 39e605e3e224e14108bea363df86049c96b66f1d Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 22 Sep 2025 14:44:52 -0400
Subject: [PATCH v5 3/3] bufmgr: Use atomic sub for unpinning buffers

The prior commit made it legal to modify BufferDesc.state while the buffer
header spinlock is held. This allows us to replace the CAS loop
inUnpinBufferNoOwner() with an atomic sub. This improves scalability
significantly. See the prior commits for more background.

Reviewed-by: Matthias van de Meent <boekewurm+postgres@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/backend/storage/buffer/bufmgr.c | 29 +++--------------------------
 1 file changed, 3 insertions(+), 26 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 42e1f50c7c8..3528a355eff 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -3272,7 +3272,6 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	ref->refcount--;
 	if (ref->refcount == 0)
 	{
-		uint32		buf_state;
 		uint32		old_buf_state;
 
 		/*
@@ -3290,33 +3289,11 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 */
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
-		/*
-		 * Decrement the shared reference count.
-		 *
-		 * Since buffer spinlock holder can update status using just write,
-		 * it's not safe to use atomic decrement here; thus use a CAS loop.
-		 *
-		 * TODO: The above requirement does not hold anymore, in a future
-		 * commit this will be rewritten to release the pin in a single atomic
-		 * operation.
-		 */
-		old_buf_state = pg_atomic_read_u32(&buf->state);
-		for (;;)
-		{
-			if (old_buf_state & BM_LOCKED)
-				old_buf_state = WaitBufHdrUnlocked(buf);
-
-			buf_state = old_buf_state;
-
-			buf_state -= BUF_REFCOUNT_ONE;
-
-			if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
-											   buf_state))
-				break;
-		}
+		/* decrement the shared reference count */
+		old_buf_state = pg_atomic_fetch_sub_u32(&buf->state, BUF_REFCOUNT_ONE);
 
 		/* Support LockBufferForCleanup() */
-		if (buf_state & BM_PIN_COUNT_WAITER)
+		if (old_buf_state & BM_PIN_COUNT_WAITER)
 			WakePinCountWaiter(buf);
 
 		ForgetPrivateRefCountEntry(ref);
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-10-09 21:16             ` Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-10-09 21:16 UTC (permalink / raw)
  To: Matthias van de Meent <boekewurm+postgres@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On 2025-10-09 16:35:44 -0400, Andres Freund wrote:
> I pushed a few commits from this patchset after Matthias' review
> (thanks!). Unfortunately in 5e899859287 I missed that the valgrind annotations
> would not be done anymore for the buffers returned by
> StrategyGetBuffer(). Which turned skink red.
> 
> The attached 0001 patch centralizes the valgrind initialization in
> TrackNewBufferPin(), which 5e899859287 had added. The nice side effect of that
> is that there are fewer VALGRIND_MAKE_MEM_DEFINED() calls than before. The
> naming isn't the perfect match, but it seems fine to me.

Forgot to say: I'll push this patch soon, to get skink back to green. Unless
somebody says something.  We can adjust this later, if the comment and/or
placement of VALGRIND_MAKE_MEM_DEFINED() isn't to everyones liking.





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-11-20 02:47               ` Andres Freund <andres@anarazel.de>
  2025-11-20 19:08                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Greg Burd <greg@burd.me>
  2025-11-21 17:52                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-24 20:57                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  0 siblings, 4 replies; 120+ messages in thread

From: Andres Freund @ 2025-11-20 02:47 UTC (permalink / raw)
  To: Matthias van de Meent <boekewurm+postgres@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-10-09 17:16:49 -0400, Andres Freund wrote:
> On 2025-10-09 16:35:44 -0400, Andres Freund wrote:
> > I pushed a few commits from this patchset after Matthias' review
> > (thanks!). Unfortunately in 5e899859287 I missed that the valgrind annotations
> > would not be done anymore for the buffers returned by
> > StrategyGetBuffer(). Which turned skink red.
> > 
> > The attached 0001 patch centralizes the valgrind initialization in
> > TrackNewBufferPin(), which 5e899859287 had added. The nice side effect of that
> > is that there are fewer VALGRIND_MAKE_MEM_DEFINED() calls than before. The
> > naming isn't the perfect match, but it seems fine to me.
> 
> Forgot to say: I'll push this patch soon, to get skink back to green. Unless
> somebody says something.  We can adjust this later, if the comment and/or
> placement of VALGRIND_MAKE_MEM_DEFINED() isn't to everyones liking.

I have pushed that fix as well as the subsequent buffer header locking changes
a while ago.


Attached is a patchset that actually implements the buffer content locks in
bufmgr.c. This isn't that close to a committable shape yet, but it seemed
useful to get it out there.  The first few patches seem closer, so it'll also
be useful to narrow this down.

0001: A straight-up bugfix in lwlock.c - albeit for a bug that seems currently
	effectively harmless.

0002: Not really required, but seems like an improvement to me

0003: A prerequisite to 0004, pretty boring itself

0004: Use 64bit atomics for BufferDesc.state - at this point nothing uses the
additional bits yet, though.  Some annoying reformatting required to avoid
long lines.

0005: There already was a wait event class for BUFFERPIN. It seems better to
make that more general than to implement them separately.

0006+0007: This is preparatory work for 0008, but also worthwhile on its
own. The private refcount stuff does show up in profiles. The reason it's
related is that without these changes the added information in 0008 makes that
worse.

0008: The main change. Implements buffer content locking independently from
lwlock.c. There's obviously a lot of similarity between lwlock.c code and
this, but I've not found a good way to reduce the duplication without giving
up too much.  This patch does immediately introduce share-exclusive as a new
lock level, mostly because it was too painful to do separately.

0009+0010+0011: Preparatory work for 0012.

0012: One of the main goals of this patchset - use the new share-exclusive
lock level to only allow hint bits to be set while no IO is going on.

0013: Prototype of making UnlockReleaseBuffer() faster and of using it more
widely in nbtree.c

0014: Now that hint bits can't be done while IO is going on, we don't need to
copy pages anymore.  This needs a fair bit more work, as denoted by the FIXMEs
in the code.

I've tried to add detail to the more important commit messages, at least until
0012.


I want to again emphasize that the important commits (i.e. 0008, 0012, 0014)
aren't close to being mergeable. But I think they're in a stage that they
could benefit from "lenient" high-level review.

Greetings,

Andres Freund

Attachments:

  [text/x-diff] v6-0001-lwlock-Fix-currently-harmless-bug-in-LWLockWakeup.patch (1.5K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/2-v6-0001-lwlock-Fix-currently-harmless-bug-in-LWLockWakeup.patch)
  download | inline diff:
From cf5f78299faf99d42c31acc795617fc2b9046844 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Fri, 7 Nov 2025 16:47:47 -0500
Subject: [PATCH v6 01/14] lwlock: Fix, currently harmless, bug in
 LWLockWakeup()

Accidentally the code in LWLockWakeup() checked the list of to-be-woken up
processes to see if LW_FLAG_HAS_WAITERS should be unset. That means that
HAS_WAITERS would not get unset immediately, but only during the next,
unnecessary, call to LWLockWakeup().

Luckily, as the code stands, this is just a small efficiency issue.

However, if there were (as in a patch of mine) a case in which LWLockWakeup()
would not find any backend to wake, despite the wait list not being empty,
we'd wrongly unset LW_FLAG_HAS_WAITERS, leading to potentially hanging.

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/backend/storage/lmgr/lwlock.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index b017880f5e4..255cfa8fa95 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -998,7 +998,7 @@ LWLockWakeup(LWLock *lock)
 			else
 				desired_state &= ~LW_FLAG_RELEASE_OK;
 
-			if (proclist_is_empty(&wakeup))
+			if (proclist_is_empty(&lock->waiters))
 				desired_state &= ~LW_FLAG_HAS_WAITERS;
 
 			desired_state &= ~LW_FLAG_LOCKED;	/* release lock */
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0002-bufmgr-Turn-BUFFER_LOCK_-into-an-enum.patch (3.1K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/3-v6-0002-bufmgr-Turn-BUFFER_LOCK_-into-an-enum.patch)
  download | inline diff:
From 70d0457b7c9527bf1304921b367a34e49d26fd6e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Fri, 7 Nov 2025 16:51:52 -0500
Subject: [PATCH v6 02/14] bufmgr: Turn BUFFER_LOCK_* into an enum

This way we will be able to benefit from compiler-warnings for code using a
switch() over all lock modes.

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/bufmgr.h        | 13 ++++++++-----
 src/backend/storage/buffer/bufmgr.c |  4 ++--
 src/tools/pgindent/typedefs.list    |  1 +
 3 files changed, 11 insertions(+), 7 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index b5f8f3c5d42..5fc3de20abc 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -200,9 +200,12 @@ extern PGDLLIMPORT int32 *LocalRefCount;
 /*
  * Buffer content lock modes (mode argument for LockBuffer())
  */
-#define BUFFER_LOCK_UNLOCK		0
-#define BUFFER_LOCK_SHARE		1
-#define BUFFER_LOCK_EXCLUSIVE	2
+typedef enum BufferLockMode
+{
+	BUFFER_LOCK_UNLOCK,
+	BUFFER_LOCK_SHARE,
+	BUFFER_LOCK_EXCLUSIVE,
+} BufferLockMode;
 
 
 /*
@@ -238,7 +241,7 @@ extern void WaitReadBuffers(ReadBuffersOperation *operation);
 extern void ReleaseBuffer(Buffer buffer);
 extern void UnlockReleaseBuffer(Buffer buffer);
 extern bool BufferIsLockedByMe(Buffer buffer);
-extern bool BufferIsLockedByMeInMode(Buffer buffer, int mode);
+extern bool BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode);
 extern bool BufferIsDirty(Buffer buffer);
 extern void MarkBufferDirty(Buffer buffer);
 extern void IncrBufferRefCount(Buffer buffer);
@@ -299,7 +302,7 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
 extern void UnlockBuffers(void);
-extern void LockBuffer(Buffer buffer, int mode);
+extern void LockBuffer(Buffer buffer, BufferLockMode mode);
 extern bool ConditionalLockBuffer(Buffer buffer);
 extern void LockBufferForCleanup(Buffer buffer);
 extern bool ConditionalLockBufferForCleanup(Buffer buffer);
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 327ddb7adc8..b682878b1fb 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2866,7 +2866,7 @@ BufferIsLockedByMe(Buffer buffer)
  * Buffer must be pinned.
  */
 bool
-BufferIsLockedByMeInMode(Buffer buffer, int mode)
+BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode)
 {
 	BufferDesc *bufHdr;
 
@@ -5601,7 +5601,7 @@ UnlockBuffers(void)
  * Acquire or release the content_lock for the buffer.
  */
 void
-LockBuffer(Buffer buffer, int mode)
+LockBuffer(Buffer buffer, BufferLockMode mode)
 {
 	BufferDesc *buf;
 
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 57f2a9ccdc5..5769227e41c 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -345,6 +345,7 @@ BufferCachePagesRec
 BufferDesc
 BufferDescPadded
 BufferHeapTupleTableSlot
+BufferLockMode
 BufferLookupEnt
 BufferManagerRelation
 BufferStrategyControl
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0003-Add-pg_atomic_unlocked_write_u64.patch (1.8K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/4-v6-0003-Add-pg_atomic_unlocked_write_u64.patch)
  download | inline diff:
From e854fb8fb8737c4cb972f5b307f37a84dde32a4f Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 5 Nov 2025 19:12:37 -0500
Subject: [PATCH v6 03/14] Add pg_atomic_unlocked_write_u64

The 64bit equivalent of pg_atomic_unlocked_write_u32(), to be used in an
upcoming patch converting BufferDesc.state into a 64bit atomic.
---
 src/include/port/atomics.h         | 10 ++++++++++
 src/include/port/atomics/generic.h |  9 +++++++++
 2 files changed, 19 insertions(+)

diff --git a/src/include/port/atomics.h b/src/include/port/atomics.h
index 96f1858da97..830ea5c7c52 100644
--- a/src/include/port/atomics.h
+++ b/src/include/port/atomics.h
@@ -488,6 +488,16 @@ pg_atomic_write_u64(volatile pg_atomic_uint64 *ptr, uint64 val)
 	pg_atomic_write_u64_impl(ptr, val);
 }
 
+static inline void
+pg_atomic_unlocked_write_u64(volatile pg_atomic_uint64 *ptr, uint64 val)
+{
+#ifndef PG_HAVE_ATOMIC_U64_SIMULATION
+	AssertPointerAlignment(ptr, 8);
+#endif
+
+	pg_atomic_unlocked_write_u64_impl(ptr, val);
+}
+
 static inline void
 pg_atomic_write_membarrier_u64(volatile pg_atomic_uint64 *ptr, uint64 val)
 {
diff --git a/src/include/port/atomics/generic.h b/src/include/port/atomics/generic.h
index 6b61a7b5416..00aa152f908 100644
--- a/src/include/port/atomics/generic.h
+++ b/src/include/port/atomics/generic.h
@@ -297,6 +297,15 @@ pg_atomic_write_u64_impl(volatile pg_atomic_uint64 *ptr, uint64 val)
 #endif /* PG_HAVE_8BYTE_SINGLE_COPY_ATOMICITY && !PG_HAVE_ATOMIC_U64_SIMULATION */
 #endif /* PG_HAVE_ATOMIC_WRITE_U64 */
 
+#ifndef PG_HAVE_ATOMIC_UNLOCKED_WRITE_U64
+#define PG_HAVE_ATOMIC_UNLOCKED_WRITE_U64
+static inline void
+pg_atomic_unlocked_write_u64_impl(volatile pg_atomic_uint64 *ptr, uint64 val)
+{
+	ptr->value = val;
+}
+#endif
+
 #ifndef PG_HAVE_ATOMIC_READ_U64
 #define PG_HAVE_ATOMIC_READ_U64
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0004-bufmgr-Change-BufferDesc.state-to-be-a-64bit-atom.patch (43.6K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/5-v6-0004-bufmgr-Change-BufferDesc.state-to-be-a-64bit-atom.patch)
  download | inline diff:
From 90697d38319c839ce2533dd4425fcce7058b07ca Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 5 Nov 2025 19:22:15 -0500
Subject: [PATCH v6 04/14] bufmgr: Change BufferDesc.state to be a 64bit atomic

This is motivated by wanting to merge buffer content locks into
BufferDesc.state in a future commit, rather than having a separate lwlock (see
commit c75ebc657ff more details). As this change is rather mechanical, it
seems to make sense to split it out into a separate commit, for easier review.

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  88 ++++++----
 src/backend/storage/buffer/buf_init.c         |   2 +-
 src/backend/storage/buffer/bufmgr.c           | 158 +++++++++---------
 src/backend/storage/buffer/freelist.c         |  24 +--
 src/backend/storage/buffer/localbuf.c         |  72 ++++----
 contrib/pg_buffercache/pg_buffercache_pages.c |   8 +-
 src/test/modules/test_aio/test_aio.c          |  12 +-
 7 files changed, 192 insertions(+), 172 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 5400c56a965..28519ad2813 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -30,7 +30,7 @@
 #include "utils/resowner.h"
 
 /*
- * Buffer state is a single 32-bit variable where following data is combined.
+ * Buffer state is a single 64-bit variable where following data is combined.
  *
  * - 18 bits refcount
  * - 4 bits usage count
@@ -39,6 +39,9 @@
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
+ * NB: A future commit will use a significant portion of the remaining bits to
+ * implement buffer locking as part of the state variable.
+ *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
@@ -49,15 +52,21 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 				 "parts of buffer state space need to equal 32");
 
 #define BUF_REFCOUNT_ONE 1
-#define BUF_REFCOUNT_MASK ((1U << BUF_REFCOUNT_BITS) - 1)
-#define BUF_USAGECOUNT_MASK (((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_REFCOUNT_BITS))
-#define BUF_USAGECOUNT_ONE (1U << BUF_REFCOUNT_BITS)
+#define BUF_REFCOUNT_MASK \
+	((UINT64CONST(1) << BUF_REFCOUNT_BITS) - 1)
+#define BUF_USAGECOUNT_MASK \
+	(((UINT64CONST(1) << BUF_USAGECOUNT_BITS) - 1) << (BUF_REFCOUNT_BITS))
+#define BUF_USAGECOUNT_ONE \
+	(UINT64CONST(1) << BUF_REFCOUNT_BITS)
 #define BUF_USAGECOUNT_SHIFT BUF_REFCOUNT_BITS
-#define BUF_FLAG_MASK (((1U << BUF_FLAG_BITS) - 1) << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS))
+#define BUF_FLAG_MASK \
+	(((UINT64CONST(1) << BUF_FLAG_BITS) - 1) << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS))
 
 /* Get refcount and usagecount from buffer state */
-#define BUF_STATE_GET_REFCOUNT(state) ((state) & BUF_REFCOUNT_MASK)
-#define BUF_STATE_GET_USAGECOUNT(state) (((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
+#define BUF_STATE_GET_REFCOUNT(state) \
+	((uint32)((state) & BUF_REFCOUNT_MASK))
+#define BUF_STATE_GET_USAGECOUNT(state) \
+	((uint32)(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT))
 
 /*
  * Flags for buffer descriptors
@@ -65,17 +74,28 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  * Note: BM_TAG_VALID essentially means that there is a buffer hashtable
  * entry associated with the buffer's tag.
  */
-#define BM_LOCKED				(1U << 22)	/* buffer header is locked */
-#define BM_DIRTY				(1U << 23)	/* data needs writing */
-#define BM_VALID				(1U << 24)	/* data is valid */
-#define BM_TAG_VALID			(1U << 25)	/* tag is assigned */
-#define BM_IO_IN_PROGRESS		(1U << 26)	/* read or write in progress */
-#define BM_IO_ERROR				(1U << 27)	/* previous I/O failed */
-#define BM_JUST_DIRTIED			(1U << 28)	/* dirtied since write started */
-#define BM_PIN_COUNT_WAITER		(1U << 29)	/* have waiter for sole pin */
-#define BM_CHECKPOINT_NEEDED	(1U << 30)	/* must write for checkpoint */
-#define BM_PERMANENT			(1U << 31)	/* permanent buffer (not unlogged,
-											 * or init fork) */
+
+/* buffer header is locked */
+#define BM_LOCKED				(UINT64CONST(1) << 22)
+/* data needs writing */
+#define BM_DIRTY				(UINT64CONST(1) << 23)
+/* data is valid */
+#define BM_VALID				(UINT64CONST(1) << 24)
+/* tag is assigned */
+#define BM_TAG_VALID			(UINT64CONST(1) << 25)
+/* read or write in progress */
+#define BM_IO_IN_PROGRESS		(UINT64CONST(1) << 26)
+/* previous I/O failed */
+#define BM_IO_ERROR				(UINT64CONST(1) << 27)
+/* dirtied since write started */
+#define BM_JUST_DIRTIED			(UINT64CONST(1) << 28)
+/* have waiter for sole pin */
+#define BM_PIN_COUNT_WAITER		(UINT64CONST(1) << 29)
+/* must write for checkpoint */
+#define BM_CHECKPOINT_NEEDED	(UINT64CONST(1) << 30)
+/* permanent buffer (not unlogged, or init fork) */
+#define BM_PERMANENT			(UINT64CONST(1) << 31)
+
 /*
  * The maximum allowed value of usage_count represents a tradeoff between
  * accuracy and speed of the clock-sweep buffer management algorithm.  A
@@ -86,7 +106,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  */
 #define BM_MAX_USAGE_COUNT	5
 
-StaticAssertDecl(BM_MAX_USAGE_COUNT < (1 << BUF_USAGECOUNT_BITS),
+StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 				 "BM_MAX_USAGE_COUNT doesn't fit in BUF_USAGECOUNT_BITS bits");
 StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
 				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
@@ -251,8 +271,8 @@ BufMappingPartitionLockByIndex(uint32 index)
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
- * atomic operations (i.e. only pg_atomic_read_u32() and
- * pg_atomic_unlocked_write_u32()).
+ * atomic operations (i.e. only pg_atomic_read_u64() and
+ * pg_atomic_unlocked_write_u64()).
  *
  * Be careful to avoid increasing the size of the struct when adding or
  * reordering members.  Keeping it below 64 bytes (the most common CPU
@@ -280,7 +300,7 @@ typedef struct BufferDesc
 	 * State of the buffer, containing flags, refcount and usagecount. See
 	 * BUF_* and BM_* defines at the top of this file.
 	 */
-	pg_atomic_uint32 state;
+	pg_atomic_uint64 state;
 
 	/*
 	 * Backend of pin-count waiter. The buffer header spinlock needs to be
@@ -386,7 +406,7 @@ BufferDescriptorGetContentLock(const BufferDesc *bdesc)
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
  */
-extern uint32 LockBufHdr(BufferDesc *desc);
+extern uint64 LockBufHdr(BufferDesc *desc);
 
 /*
  * Unlock the buffer header.
@@ -397,9 +417,9 @@ extern uint32 LockBufHdr(BufferDesc *desc);
 static inline void
 UnlockBufHdr(BufferDesc *desc)
 {
-	Assert(pg_atomic_read_u32(&desc->state) & BM_LOCKED);
+	Assert(pg_atomic_read_u64(&desc->state) & BM_LOCKED);
 
-	pg_atomic_fetch_sub_u32(&desc->state, BM_LOCKED);
+	pg_atomic_fetch_sub_u64(&desc->state, BM_LOCKED);
 }
 
 /*
@@ -410,14 +430,14 @@ UnlockBufHdr(BufferDesc *desc)
  * Note that this approach would not work for usagecount, since we need to cap
  * the usagecount at BM_MAX_USAGE_COUNT.
  */
-static inline uint32
-UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
-				uint32 set_bits, uint32 unset_bits,
+static inline uint64
+UnlockBufHdrExt(BufferDesc *desc, uint64 old_buf_state,
+				uint64 set_bits, uint64 unset_bits,
 				int refcount_change)
 {
 	for (;;)
 	{
-		uint32		buf_state = old_buf_state;
+		uint64		buf_state = old_buf_state;
 
 		Assert(buf_state & BM_LOCKED);
 
@@ -428,7 +448,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 		if (refcount_change != 0)
 			buf_state += BUF_REFCOUNT_ONE * refcount_change;
 
-		if (pg_atomic_compare_exchange_u32(&desc->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&desc->state, &old_buf_state,
 										   buf_state))
 		{
 			return old_buf_state;
@@ -436,7 +456,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 	}
 }
 
-extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
+extern uint64 WaitBufHdrUnlocked(BufferDesc *buf);
 
 /* in bufmgr.c */
 
@@ -496,14 +516,14 @@ extern void TrackNewBufferPin(Buffer buf);
 
 /* solely to make it easier to write tests */
 extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
-extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 							  bool forget_owner, bool release_aio);
 
 
 /* freelist.c */
 extern IOContext IOContextForStrategy(BufferAccessStrategy strategy);
 extern BufferDesc *StrategyGetBuffer(BufferAccessStrategy strategy,
-									 uint32 *buf_state, bool *from_ring);
+									 uint64 *buf_state, bool *from_ring);
 extern bool StrategyRejectBuffer(BufferAccessStrategy strategy,
 								 BufferDesc *buf, bool from_ring);
 
@@ -539,7 +559,7 @@ extern BlockNumber ExtendBufferedRelLocal(BufferManagerRelation bmr,
 										  uint32 *extended_by);
 extern void MarkLocalBufferDirty(Buffer buffer);
 extern void TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty,
-								   uint32 set_flag_bits, bool release_aio);
+								   uint64 set_flag_bits, bool release_aio);
 extern bool StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait);
 extern void FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln);
 extern void InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced);
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6fd3a6bbac5..25f71191ec3 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -121,7 +121,7 @@ BufferManagerShmemInit(void)
 
 			ClearBufferTag(&buf->tag);
 
-			pg_atomic_init_u32(&buf->state, 0);
+			pg_atomic_init_u64(&buf->state, 0);
 			buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
 
 			buf->buf_id = i;
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index b682878b1fb..e33fa0cbfec 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -686,7 +686,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 {
 	BufferDesc *bufHdr;
 	BufferTag	tag;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(recent_buffer));
 
@@ -699,7 +699,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 		int			b = -recent_buffer - 1;
 
 		bufHdr = GetLocalBufferDescriptor(b);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		/* Is it still valid and holding the right tag? */
 		if ((buf_state & BM_VALID) && BufferTagsEqual(&tag, &bufHdr->tag))
@@ -1292,8 +1292,8 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
 				bufHdr = GetLocalBufferDescriptor(-buffers[i] - 1);
 			else
 				bufHdr = GetBufferDescriptor(buffers[i] - 1);
-			Assert(pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID);
-			found = pg_atomic_read_u32(&bufHdr->state) & BM_VALID;
+			Assert(pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID);
+			found = pg_atomic_read_u64(&bufHdr->state) & BM_VALID;
 		}
 		else
 		{
@@ -1519,10 +1519,10 @@ CheckReadBuffersOperation(ReadBuffersOperation *operation, bool is_complete)
 			GetBufferDescriptor(buffer - 1);
 
 		Assert(BufferGetBlockNumber(buffer) == operation->blocknum + i);
-		Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_TAG_VALID);
+		Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_TAG_VALID);
 
 		if (i < operation->nblocks_done)
-			Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_VALID);
+			Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_VALID);
 	}
 #endif
 }
@@ -1989,8 +1989,8 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	int			existing_buf_id;
 	Buffer		victim_buffer;
 	BufferDesc *victim_buf_hdr;
-	uint32		victim_buf_state;
-	uint32		set_bits = 0;
+	uint64		victim_buf_state;
+	uint64		set_bits = 0;
 
 	/* Make sure we will have room to remember the buffer pin */
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2157,7 +2157,7 @@ InvalidateBuffer(BufferDesc *buf)
 	uint32		oldHash;		/* hash value for oldTag */
 	LWLock	   *oldPartitionLock;	/* buffer partition lock for it */
 	uint32		oldFlags;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/* Save the original buffer tag before dropping the spinlock */
 	oldTag = buf->tag;
@@ -2248,7 +2248,7 @@ retry:
 static bool
 InvalidateVictimBuffer(BufferDesc *buf_hdr)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	uint32		hash;
 	LWLock	   *partition_lock;
 	BufferTag	tag;
@@ -2308,10 +2308,10 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 
 	LWLockRelease(partition_lock);
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	Assert(!(buf_state & (BM_DIRTY | BM_VALID | BM_TAG_VALID)));
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u32(&buf_hdr->state)) > 0);
+	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u64(&buf_hdr->state)) > 0);
 
 	return true;
 }
@@ -2321,7 +2321,7 @@ GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context)
 {
 	BufferDesc *buf_hdr;
 	Buffer		buf;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		from_ring;
 
 	/*
@@ -2454,7 +2454,7 @@ again:
 
 	/* a final set of sanity checks */
 #ifdef USE_ASSERT_CHECKING
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) == 1);
 	Assert(!(buf_state & (BM_TAG_VALID | BM_VALID | BM_DIRTY)));
@@ -2745,13 +2745,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 */
 			do
 			{
-				pg_atomic_fetch_and_u32(&existing_hdr->state, ~BM_VALID);
+				pg_atomic_fetch_and_u64(&existing_hdr->state, ~BM_VALID);
 			} while (!StartBufferIO(existing_hdr, true, false));
 		}
 		else
 		{
-			uint32		buf_state;
-			uint32		set_bits = 0;
+			uint64		buf_state;
+			uint64		set_bits = 0;
 
 			buf_state = LockBufHdr(victim_buf_hdr);
 
@@ -2927,7 +2927,7 @@ BufferIsDirty(Buffer buffer)
 		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 	}
 
-	return pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY;
+	return pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY;
 }
 
 /*
@@ -2943,8 +2943,8 @@ void
 MarkBufferDirty(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
-	uint32		old_buf_state;
+	uint64		buf_state;
+	uint64		old_buf_state;
 
 	if (!BufferIsValid(buffer))
 		elog(ERROR, "bad buffer ID: %d", buffer);
@@ -2964,7 +2964,7 @@ MarkBufferDirty(Buffer buffer)
 	 * NB: We have to wait for the buffer header spinlock to be not held, as
 	 * TerminateBufferIO() relies on the spinlock.
 	 */
-	old_buf_state = pg_atomic_read_u32(&bufHdr->state);
+	old_buf_state = pg_atomic_read_u64(&bufHdr->state);
 	for (;;)
 	{
 		if (old_buf_state & BM_LOCKED)
@@ -2975,7 +2975,7 @@ MarkBufferDirty(Buffer buffer)
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
 
-		if (pg_atomic_compare_exchange_u32(&bufHdr->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&bufHdr->state, &old_buf_state,
 										   buf_state))
 			break;
 	}
@@ -3079,10 +3079,10 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 
 	if (ref == NULL)
 	{
-		uint32		buf_state;
-		uint32		old_buf_state;
+		uint64		buf_state;
+		uint64		old_buf_state;
 
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
@@ -3116,7 +3116,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 					buf_state += BUF_USAGECOUNT_ONE;
 			}
 
-			if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+			if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 											   buf_state))
 			{
 				result = (buf_state & BM_VALID) != 0;
@@ -3143,7 +3143,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		 * that the buffer page is legitimately non-accessible here.  We
 		 * cannot meddle with that.
 		 */
-		result = (pg_atomic_read_u32(&buf->state) & BM_VALID) != 0;
+		result = (pg_atomic_read_u64(&buf->state) & BM_VALID) != 0;
 
 		Assert(ref->refcount > 0);
 		ref->refcount++;
@@ -3178,7 +3178,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	/*
 	 * As explained, We don't expect any preexisting pins. That allows us to
@@ -3190,7 +3190,7 @@ PinBuffer_Locked(BufferDesc *buf)
 	 * Since we hold the buffer spinlock, we can update the buffer state and
 	 * release the lock in one operation.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 
 	UnlockBufHdrExt(buf, old_buf_state,
 					0, 0, 1);
@@ -3220,7 +3220,7 @@ WakePinCountWaiter(BufferDesc *buf)
 	 * BM_PIN_COUNT_WAITER if it stops waiting for a reason other than this
 	 * backend waking it up.
 	 */
-	uint32		buf_state = LockBufHdr(buf);
+	uint64		buf_state = LockBufHdr(buf);
 
 	if ((buf_state & BM_PIN_COUNT_WAITER) &&
 		BUF_STATE_GET_REFCOUNT(buf_state) == 1)
@@ -3267,7 +3267,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	ref->refcount--;
 	if (ref->refcount == 0)
 	{
-		uint32		old_buf_state;
+		uint64		old_buf_state;
 
 		/*
 		 * Mark buffer non-accessible to Valgrind.
@@ -3285,7 +3285,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
 		/* decrement the shared reference count */
-		old_buf_state = pg_atomic_fetch_sub_u32(&buf->state, BUF_REFCOUNT_ONE);
+		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
 
 		/* Support LockBufferForCleanup() */
 		if (old_buf_state & BM_PIN_COUNT_WAITER)
@@ -3342,7 +3342,7 @@ TrackNewBufferPin(Buffer buf)
 static void
 BufferSync(int flags)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	int			buf_id;
 	int			num_to_scan;
 	int			num_spaces;
@@ -3352,7 +3352,7 @@ BufferSync(int flags)
 	Oid			last_tsid;
 	binaryheap *ts_heap;
 	int			i;
-	uint32		mask = BM_DIRTY;
+	uint64		mask = BM_DIRTY;
 	WritebackContext wb_context;
 
 	/*
@@ -3384,7 +3384,7 @@ BufferSync(int flags)
 	for (buf_id = 0; buf_id < NBuffers; buf_id++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
-		uint32		set_bits = 0;
+		uint64		set_bits = 0;
 
 		/*
 		 * Header spinlock is enough to examine BM_DIRTY, see comment in
@@ -3551,7 +3551,7 @@ BufferSync(int flags)
 		 * write the buffer though we didn't need to.  It doesn't seem worth
 		 * guarding against this, though.
 		 */
-		if (pg_atomic_read_u32(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
+		if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
 		{
 			if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
 			{
@@ -3921,7 +3921,7 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 {
 	BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
 	int			result = 0;
-	uint32		buf_state;
+	uint64		buf_state;
 	BufferTag	tag;
 
 	/* Make sure we can handle the pin */
@@ -4169,7 +4169,7 @@ DebugPrintBufferRefcount(Buffer buffer)
 	int32		loccount;
 	char	   *result;
 	ProcNumber	backend;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 	if (BufferIsLocal(buffer))
@@ -4186,9 +4186,9 @@ DebugPrintBufferRefcount(Buffer buffer)
 	}
 
 	/* theoretically we should lock the bufhdr here */
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
-	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%x, refcount=%u %d)",
+	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%" PRIx64 ", refcount=%u %d)",
 					  buffer,
 					  relpathbackend(BufTagGetRelFileLocator(&buf->tag), backend,
 									 BufTagGetForkNum(&buf->tag)).str,
@@ -4288,7 +4288,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	instr_time	io_start;
 	Block		bufBlock;
 	char	   *bufToWrite;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
@@ -4486,7 +4486,7 @@ BufferIsPermanent(Buffer buffer)
 	 * not random garbage.
 	 */
 	bufHdr = GetBufferDescriptor(buffer - 1);
-	return (pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT) != 0;
+	return (pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT) != 0;
 }
 
 /*
@@ -4949,11 +4949,11 @@ FlushRelationBuffers(Relation rel)
 	{
 		for (i = 0; i < NLocBuffer; i++)
 		{
-			uint32		buf_state;
+			uint64		buf_state;
 
 			bufHdr = GetLocalBufferDescriptor(i);
 			if (BufTagMatchesRelFileLocator(&bufHdr->tag, &rel->rd_locator) &&
-				((buf_state = pg_atomic_read_u32(&bufHdr->state)) &
+				((buf_state = pg_atomic_read_u64(&bufHdr->state)) &
 				 (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 			{
 				ErrorContextCallback errcallback;
@@ -4989,7 +4989,7 @@ FlushRelationBuffers(Relation rel)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5061,7 +5061,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 	{
 		SMgrSortArray *srelent = NULL;
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -5310,7 +5310,7 @@ FlushDatabaseBuffers(Oid dbid)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5458,13 +5458,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u32(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
 		(BM_DIRTY | BM_JUST_DIRTIED))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
 		bool		delayChkptFlags = false;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * If we need to protect hint bit updates from torn writes, WAL-log a
@@ -5476,7 +5476,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
 		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT))
+			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5576,8 +5576,8 @@ UnlockBuffers(void)
 
 	if (buf)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		buf_state = LockBufHdr(buf);
 
@@ -5708,8 +5708,8 @@ LockBufferForCleanup(Buffer buffer)
 
 	for (;;)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		/* Try to acquire lock */
 		LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
@@ -5857,7 +5857,7 @@ bool
 ConditionalLockBufferForCleanup(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state,
+	uint64		buf_state,
 				refcount;
 
 	Assert(BufferIsValid(buffer));
@@ -5915,7 +5915,7 @@ bool
 IsBufferCleanupOK(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 
@@ -5971,7 +5971,7 @@ WaitIO(BufferDesc *buf)
 	ConditionVariablePrepareToSleep(cv);
 	for (;;)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 		PgAioWaitRef iow;
 
 		/*
@@ -6045,7 +6045,7 @@ WaitIO(BufferDesc *buf)
 bool
 StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	ResourceOwnerEnlarge(CurrentResourceOwner);
 
@@ -6101,11 +6101,11 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
  * is being released)
  */
 void
-TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 				  bool forget_owner, bool release_aio)
 {
-	uint32		buf_state;
-	uint32		unset_flag_bits = 0;
+	uint64		buf_state;
+	uint64		unset_flag_bits = 0;
 	int			refcount_change = 0;
 
 	buf_state = LockBufHdr(buf);
@@ -6166,7 +6166,7 @@ static void
 AbortBufferIO(Buffer buffer)
 {
 	BufferDesc *buf_hdr = GetBufferDescriptor(buffer - 1);
-	uint32		buf_state;
+	uint64		buf_state;
 
 	buf_state = LockBufHdr(buf_hdr);
 	Assert(buf_state & (BM_IO_IN_PROGRESS | BM_TAG_VALID));
@@ -6260,11 +6260,11 @@ rlocator_comparator(const void *p1, const void *p2)
 /*
  * Lock buffer header - set BM_LOCKED in buffer state.
  */
-uint32
+uint64
 LockBufHdr(BufferDesc *desc)
 {
 	SpinDelayStatus delayStatus;
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	Assert(!BufferIsLocal(BufferDescriptorGetBuffer(desc)));
 
@@ -6273,7 +6273,7 @@ LockBufHdr(BufferDesc *desc)
 	while (true)
 	{
 		/* set BM_LOCKED flag */
-		old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
+		old_buf_state = pg_atomic_fetch_or_u64(&desc->state, BM_LOCKED);
 		/* if it wasn't set before we're OK */
 		if (!(old_buf_state & BM_LOCKED))
 			break;
@@ -6290,20 +6290,20 @@ LockBufHdr(BufferDesc *desc)
  * Obviously the buffer could be locked by the time the value is returned, so
  * this is primarily useful in CAS style loops.
  */
-pg_noinline uint32
+pg_noinline uint64
 WaitBufHdrUnlocked(BufferDesc *buf)
 {
 	SpinDelayStatus delayStatus;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	init_local_spin_delay(&delayStatus);
 
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
 	while (buf_state & BM_LOCKED)
 	{
 		perform_spin_delay(&delayStatus);
-		buf_state = pg_atomic_read_u32(&buf->state);
+		buf_state = pg_atomic_read_u64(&buf->state);
 	}
 
 	finish_spin_delay(&delayStatus);
@@ -6591,12 +6591,12 @@ ResOwnerPrintBufferPin(Datum res)
 static bool
 EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result;
 
 	*buffer_flushed = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -6690,12 +6690,12 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -6742,7 +6742,7 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
@@ -6809,7 +6809,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		BufferDesc *buf_hdr = is_temp ?
 			GetLocalBufferDescriptor(-buffer - 1)
 			: GetBufferDescriptor(buffer - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * Check that all the buffers are actually ones that could conceivably
@@ -6827,7 +6827,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		}
 
 		if (is_temp)
-			buf_state = pg_atomic_read_u32(&buf_hdr->state);
+			buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		else
 			buf_state = LockBufHdr(buf_hdr);
 
@@ -6865,7 +6865,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		if (is_temp)
 		{
 			buf_state += BUF_REFCOUNT_ONE;
-			pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 		}
 		else
 			UnlockBufHdrExt(buf_hdr, buf_state, 0, 0, 1);
@@ -7051,13 +7051,13 @@ buffer_readv_complete_one(PgAioTargetData *td, uint8 buf_off, Buffer buffer,
 		: GetBufferDescriptor(buffer - 1);
 	BufferTag	tag = buf_hdr->tag;
 	char	   *bufdata = BufferGetBlock(buffer);
-	uint32		set_flag_bits;
+	uint64		set_flag_bits;
 	int			piv_flags;
 
 	/* check that the buffer is in the expected state for a read */
 #ifdef USE_ASSERT_CHECKING
 	{
-		uint32		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(!(buf_state & BM_VALID));
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 28d952b3534..1d4f19a9afd 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -86,7 +86,7 @@ typedef struct BufferAccessStrategyData
 
 /* Prototypes for internal functions */
 static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
-									 uint32 *buf_state);
+									 uint64 *buf_state);
 static void AddBufferToRing(BufferAccessStrategy strategy,
 							BufferDesc *buf);
 
@@ -171,7 +171,7 @@ ClockSweepTick(void)
  *	before returning.
  */
 BufferDesc *
-StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_ring)
+StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_ring)
 {
 	BufferDesc *buf;
 	int			bgwprocno;
@@ -230,8 +230,8 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 	trycounter = NBuffers;
 	for (;;)
 	{
-		uint32		old_buf_state;
-		uint32		local_buf_state;
+		uint64		old_buf_state;
+		uint64		local_buf_state;
 
 		buf = GetBufferDescriptor(ClockSweepTick());
 
@@ -239,7 +239,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 		 * Check whether the buffer can be used and pin it if so. Do this
 		 * using a CAS loop, to avoid having to lock the buffer header.
 		 */
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			local_buf_state = old_buf_state;
@@ -277,7 +277,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 			{
 				local_buf_state -= BUF_USAGECOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					trycounter = NBuffers;
@@ -289,7 +289,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 				/* pin the buffer if the CAS succeeds */
 				local_buf_state += BUF_REFCOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					/* Found a usable buffer */
@@ -655,12 +655,12 @@ FreeAccessStrategy(BufferAccessStrategy strategy)
  * returning.
  */
 static BufferDesc *
-GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
+GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
 {
 	BufferDesc *buf;
 	Buffer		bufnum;
-	uint32		old_buf_state;
-	uint32		local_buf_state;	/* to avoid repeated (de-)referencing */
+	uint64		old_buf_state;
+	uint64		local_buf_state;	/* to avoid repeated (de-)referencing */
 
 
 	/* Advance to next ring slot */
@@ -682,7 +682,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 	 * Check whether the buffer can be used and pin it if so. Do this using a
 	 * CAS loop, to avoid having to lock the buffer header.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 	for (;;)
 	{
 		local_buf_state = old_buf_state;
@@ -710,7 +710,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 		/* pin the buffer if the CAS succeeds */
 		local_buf_state += BUF_REFCOUNT_ONE;
 
-		if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 										   local_buf_state))
 		{
 			*buf_state = local_buf_state;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 15aac7d1c9f..a41a5facd3a 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -148,7 +148,7 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 	}
 	else
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		victim_buffer = GetLocalVictimBuffer();
 		bufid = -victim_buffer - 1;
@@ -165,10 +165,10 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 		 */
 		bufHdr->tag = newTag;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
 		buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 		*foundPtr = false;
 	}
@@ -245,12 +245,12 @@ GetLocalVictimBuffer(void)
 
 		if (LocalRefCount[victim_bufid] == 0)
 		{
-			uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 			if (BUF_STATE_GET_USAGECOUNT(buf_state) > 0)
 			{
 				buf_state -= BUF_USAGECOUNT_ONE;
-				pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+				pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 				trycounter = NLocBuffer;
 			}
 			else if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
@@ -286,13 +286,13 @@ GetLocalVictimBuffer(void)
 	 * this buffer is not referenced but it might still be dirty. if that's
 	 * the case, write it out before reusing it!
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY)
 		FlushLocalBuffer(bufHdr, NULL);
 
 	/*
 	 * Remove the victim buffer from the hashtable and mark as invalid.
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID)
 	{
 		InvalidateLocalBuffer(bufHdr, false);
 
@@ -417,7 +417,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 		if (found)
 		{
 			BufferDesc *existing_hdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			UnpinLocalBuffer(BufferDescriptorGetBuffer(victim_buf_hdr));
 
@@ -428,18 +428,18 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 			/*
 			 * Clear the BM_VALID bit, do StartLocalBufferIO() and proceed.
 			 */
-			buf_state = pg_atomic_read_u32(&existing_hdr->state);
+			buf_state = pg_atomic_read_u64(&existing_hdr->state);
 			Assert(buf_state & BM_TAG_VALID);
 			Assert(!(buf_state & BM_DIRTY));
 			buf_state &= ~BM_VALID;
-			pg_atomic_unlocked_write_u32(&existing_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&existing_hdr->state, buf_state);
 
 			/* no need to loop for local buffers */
 			StartLocalBufferIO(existing_hdr, true, false);
 		}
 		else
 		{
-			uint32		buf_state = pg_atomic_read_u32(&victim_buf_hdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&victim_buf_hdr->state);
 
 			Assert(!(buf_state & (BM_VALID | BM_TAG_VALID | BM_DIRTY | BM_JUST_DIRTIED)));
 
@@ -447,7 +447,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 
 			buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 
-			pg_atomic_unlocked_write_u32(&victim_buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&victim_buf_hdr->state, buf_state);
 
 			hresult->id = victim_buf_id;
 
@@ -467,13 +467,13 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 	{
 		Buffer		buf = buffers[i];
 		BufferDesc *buf_hdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		buf_state |= BM_VALID;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 
 	*extended_by = extend_by;
@@ -492,7 +492,7 @@ MarkLocalBufferDirty(Buffer buffer)
 {
 	int			bufid;
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsLocal(buffer));
 
@@ -506,14 +506,14 @@ MarkLocalBufferDirty(Buffer buffer)
 
 	bufHdr = GetLocalBufferDescriptor(bufid);
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	if (!(buf_state & BM_DIRTY))
 		pgBufferUsage.local_blks_dirtied++;
 
 	buf_state |= BM_DIRTY;
 
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -522,7 +522,7 @@ MarkLocalBufferDirty(Buffer buffer)
 bool
 StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * With AIO the buffer could have IO in progress, e.g. when there are two
@@ -542,7 +542,7 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 	/* Once we get here, there is definitely no I/O active on this buffer */
 
 	/* Check if someone else already did the I/O */
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 	if (forInput ? (buf_state & BM_VALID) : !(buf_state & BM_DIRTY))
 	{
 		return false;
@@ -559,11 +559,11 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
  * Like TerminateBufferIO, but for local buffers
  */
 void
-TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bits,
+TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint64 set_flag_bits,
 					   bool release_aio)
 {
 	/* Only need to adjust flags */
-	uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+	uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/* BM_IO_IN_PROGRESS isn't currently used for local buffers */
 
@@ -582,7 +582,7 @@ TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bit
 	}
 
 	buf_state |= set_flag_bits;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 	/* local buffers don't track IO using resowners */
 
@@ -606,7 +606,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(bufHdr);
 	int			bufid = -buffer - 1;
-	uint32		buf_state;
+	uint64		buf_state;
 	LocalBufferLookupEnt *hresult;
 
 	/*
@@ -622,7 +622,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 		Assert(!pgaio_wref_valid(&bufHdr->io_wref));
 	}
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/*
 	 * We need to test not just LocalRefCount[bufid] but also the BufferDesc
@@ -647,7 +647,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 	ClearBufferTag(&bufHdr->tag);
 	buf_state &= ~BUF_FLAG_MASK;
 	buf_state &= ~BUF_USAGECOUNT_MASK;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -671,9 +671,9 @@ DropRelationLocalBuffers(RelFileLocator rlocator, ForkNumber *forkNum,
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (!(buf_state & BM_TAG_VALID) ||
 			!BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -706,9 +706,9 @@ DropRelationAllLocalBuffers(RelFileLocator rlocator)
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if ((buf_state & BM_TAG_VALID) &&
 			BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -804,11 +804,11 @@ InitLocalBuffers(void)
 bool
 PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	Buffer		buffer = BufferDescriptorGetBuffer(buf_hdr);
 	int			bufid = -buffer - 1;
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	if (LocalRefCount[bufid] == 0)
 	{
@@ -819,7 +819,7 @@ PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 		{
 			buf_state += BUF_USAGECOUNT_ONE;
 		}
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/*
 		 * See comment in PinBuffer().
@@ -856,14 +856,14 @@ UnpinLocalBufferNoOwner(Buffer buffer)
 	if (--LocalRefCount[buffid] == 0)
 	{
 		BufferDesc *buf_hdr = GetLocalBufferDescriptor(buffid);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		NLocalPinnedBuffers--;
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state -= BUF_REFCOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/* see comment in UnpinBufferNoOwner */
 		VALGRIND_MAKE_MEM_NOACCESS(LocalBufHdrGetBlock(buf_hdr), BLCKSZ);
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index c29b784dfa1..32bd8aa784a 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -192,7 +192,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
 		for (i = 0; i < NBuffers; i++)
 		{
 			BufferDesc *bufHdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			CHECK_FOR_INTERRUPTS();
 
@@ -559,7 +559,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -570,7 +570,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 		 * noticeably increase the cost of the function.
 		 */
 		bufHdr = GetBufferDescriptor(i);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (buf_state & BM_VALID)
 		{
@@ -620,7 +620,7 @@ pg_buffercache_usage_counts(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		int			usage_count;
 
 		CHECK_FOR_INTERRUPTS();
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index d7eadeab256..488d98e7e66 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -308,9 +308,9 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 {
 	Buffer		buf;
 	BufferDesc *buf_hdr;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		was_pinned = false;
-	uint32		unset_bits = 0;
+	uint64		unset_bits = 0;
 
 	/* place buffer in shared buffers without erroring out */
 	buf = ReadBufferExtended(rel, MAIN_FORKNUM, blkno, RBM_ZERO_AND_LOCK, NULL);
@@ -319,7 +319,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	}
 	else
 	{
@@ -340,7 +340,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_state &= ~unset_bits;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 	else
 	{
@@ -489,7 +489,7 @@ invalidate_rel_block(PG_FUNCTION_ARGS)
 
 			LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
-			if (pg_atomic_read_u32(&buf_hdr->state) & BM_DIRTY)
+			if (pg_atomic_read_u64(&buf_hdr->state) & BM_DIRTY)
 			{
 				if (BufferIsLocal(buf))
 					FlushLocalBuffer(buf_hdr, NULL);
@@ -572,7 +572,7 @@ buffer_call_terminate_io(PG_FUNCTION_ARGS)
 	bool		io_error = PG_GETARG_BOOL(3);
 	bool		release_aio = PG_GETARG_BOOL(4);
 	bool		clear_dirty = false;
-	uint32		set_flag_bits = 0;
+	uint64		set_flag_bits = 0;
 
 	if (io_error)
 		set_flag_bits |= BM_IO_ERROR;
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0005-Rename-BUFFERPIN-wait-event-class-to-BUFFER.patch (6.5K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/6-v6-0005-Rename-BUFFERPIN-wait-event-class-to-BUFFER.patch)
  download | inline diff:
From 38bf7740795793403d5fba883be29a74d05ac91e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 6 Nov 2025 09:15:18 -0500
Subject: [PATCH v6 05/14] Rename BUFFERPIN wait event class to BUFFER

In an upcoming patch more wait events will be added to the wait event
class (for buffer locking), making the current name too
specific. Alternatively we could introduce a dedicated wait event class for
those, but it seems somewhat confusing to have a BUFFERPIN and a BUFFER wait
event class.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/utils/wait_classes.h                |  2 +-
 src/backend/storage/buffer/bufmgr.c             |  2 +-
 src/backend/storage/ipc/standby.c               |  2 +-
 src/backend/utils/activity/wait_event.c         | 12 ++++++------
 src/backend/utils/activity/wait_event_names.txt |  6 +++---
 doc/src/sgml/monitoring.sgml                    |  8 +++-----
 src/test/recovery/t/048_vacuum_horizon_floor.pl |  2 +-
 src/test/regress/expected/sysviews.out          |  2 +-
 8 files changed, 17 insertions(+), 19 deletions(-)

diff --git a/src/include/utils/wait_classes.h b/src/include/utils/wait_classes.h
index 51ee68397d5..57888aa62f7 100644
--- a/src/include/utils/wait_classes.h
+++ b/src/include/utils/wait_classes.h
@@ -17,7 +17,7 @@
  */
 #define PG_WAIT_LWLOCK				0x01000000U
 #define PG_WAIT_LOCK				0x03000000U
-#define PG_WAIT_BUFFERPIN			0x04000000U
+#define PG_WAIT_BUFFER				0x04000000U
 #define PG_WAIT_ACTIVITY			0x05000000U
 #define PG_WAIT_CLIENT				0x06000000U
 #define PG_WAIT_EXTENSION			0x07000000U
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index e33fa0cbfec..d4235ca7939 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5799,7 +5799,7 @@ LockBufferForCleanup(Buffer buffer)
 			SetStartupBufferPinWaitBufId(-1);
 		}
 		else
-			ProcWaitForSignal(WAIT_EVENT_BUFFER_PIN);
+			ProcWaitForSignal(WAIT_EVENT_BUFFER_CLEANUP);
 
 		/*
 		 * Remove flag marking us as waiter. Normally this will not be set
diff --git a/src/backend/storage/ipc/standby.c b/src/backend/storage/ipc/standby.c
index 4222bdab078..fc45d72c79b 100644
--- a/src/backend/storage/ipc/standby.c
+++ b/src/backend/storage/ipc/standby.c
@@ -840,7 +840,7 @@ ResolveRecoveryConflictWithBufferPin(void)
 	 * SIGHUP signal handler, etc cannot do that because it uses the different
 	 * latch from that ProcWaitForSignal() waits on.
 	 */
-	ProcWaitForSignal(WAIT_EVENT_BUFFER_PIN);
+	ProcWaitForSignal(WAIT_EVENT_BUFFER_CLEANUP);
 
 	if (got_standby_delay_timeout)
 		SendRecoveryConflictWithBufferPin(PROCSIG_RECOVERY_CONFLICT_BUFFERPIN);
diff --git a/src/backend/utils/activity/wait_event.c b/src/backend/utils/activity/wait_event.c
index d9b8f34a355..96d61f77f6e 100644
--- a/src/backend/utils/activity/wait_event.c
+++ b/src/backend/utils/activity/wait_event.c
@@ -29,7 +29,7 @@
 
 
 static const char *pgstat_get_wait_activity(WaitEventActivity w);
-static const char *pgstat_get_wait_bufferpin(WaitEventBufferPin w);
+static const char *pgstat_get_wait_buffer(WaitEventBuffer w);
 static const char *pgstat_get_wait_client(WaitEventClient w);
 static const char *pgstat_get_wait_ipc(WaitEventIPC w);
 static const char *pgstat_get_wait_timeout(WaitEventTimeout w);
@@ -389,8 +389,8 @@ pgstat_get_wait_event_type(uint32 wait_event_info)
 		case PG_WAIT_LOCK:
 			event_type = "Lock";
 			break;
-		case PG_WAIT_BUFFERPIN:
-			event_type = "BufferPin";
+		case PG_WAIT_BUFFER:
+			event_type = "Buffer";
 			break;
 		case PG_WAIT_ACTIVITY:
 			event_type = "Activity";
@@ -453,11 +453,11 @@ pgstat_get_wait_event(uint32 wait_event_info)
 		case PG_WAIT_INJECTIONPOINT:
 			event_name = GetWaitEventCustomIdentifier(wait_event_info);
 			break;
-		case PG_WAIT_BUFFERPIN:
+		case PG_WAIT_BUFFER:
 			{
-				WaitEventBufferPin w = (WaitEventBufferPin) wait_event_info;
+				WaitEventBuffer w = (WaitEventBuffer) wait_event_info;
 
-				event_name = pgstat_get_wait_bufferpin(w);
+				event_name = pgstat_get_wait_buffer(w);
 				break;
 			}
 		case PG_WAIT_ACTIVITY:
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index c1ac71ff7f2..1e5e368a5dc 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -279,12 +279,12 @@ WAL_WRITE	"Waiting for a write to a WAL file."
 ABI_compatibility:
 
 #
-# Wait Events - Buffer Pin
+# Wait Events - Buffer
 #
 
-Section: ClassName - WaitEventBufferPin
+Section: ClassName - WaitEventBuffer
 
-BUFFER_PIN	"Waiting to acquire an exclusive pin on a buffer."
+BUFFER_CLEANUP	"Waiting to acquire an exclusive pin on a buffer. Buffer pin waits can be protracted if another process holds an open cursor that last read data from the buffer in question."
 
 ABI_compatibility:
 
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 436ef0e8bd0..97ae34a92aa 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -1053,11 +1053,9 @@ postgres   27093  0.0  0.0  30096  2752 ?        Ss   11:34   0:00 postgres: ser
       </entry>
      </row>
      <row>
-      <entry><literal>BufferPin</literal></entry>
-      <entry>The server process is waiting for exclusive access to
-       a data buffer.  Buffer pin waits can be protracted if
-       another process holds an open cursor that last read data from the
-       buffer in question. See <xref linkend="wait-event-bufferpin-table"/>.
+      <entry><literal>Buffer</literal></entry>
+      <entry>The server process is waiting for access to a data buffer.
+      See <xref linkend="wait-event-buffer-table"/>.
       </entry>
      </row>
      <row>
diff --git a/src/test/recovery/t/048_vacuum_horizon_floor.pl b/src/test/recovery/t/048_vacuum_horizon_floor.pl
index 668eedd71b2..9cdf6cee8a7 100644
--- a/src/test/recovery/t/048_vacuum_horizon_floor.pl
+++ b/src/test/recovery/t/048_vacuum_horizon_floor.pl
@@ -194,7 +194,7 @@ $node_primary->poll_query_until(
 	qq[
 	SELECT count(*) >= 1 FROM pg_stat_activity
 		WHERE pid = $vacuum_pid
-		AND wait_event = 'BufferPin';
+		AND wait_event = 'BufferCleanup';
 	],
 	't');
 
diff --git a/src/test/regress/expected/sysviews.out b/src/test/regress/expected/sysviews.out
index 3b37fafa65b..0411db832f1 100644
--- a/src/test/regress/expected/sysviews.out
+++ b/src/test/regress/expected/sysviews.out
@@ -182,7 +182,7 @@ select type, count(*) > 0 as ok FROM pg_wait_events
    type    | ok 
 -----------+----
  Activity  | t
- BufferPin | t
+ Buffer    | t
  Client    | t
  Extension | t
  IO        | t
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0006-bufmgr-Separate-keys-for-private-refcount-infrast.patch (10.0K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/7-v6-0006-bufmgr-Separate-keys-for-private-refcount-infrast.patch)
  download | inline diff:
From f2e8d9de5bd2dae9b16264cb77d92bcc8e5ec8df Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 12 Nov 2025 12:50:52 -0500
Subject: [PATCH v6 06/14] bufmgr: Separate keys for private refcount
 infrastructure

This makes lookups faster, due to allowing auto-vectorized lookups. It is also
beneficial for an upcoming patch, independent of auto-vectorization, as the
upcoming patch wants to track more information for each pinned buffer, making
the existing loop, iterating over an array of PrivateRefCountEntry, more
expensive due to increasing its size.

Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/storage/buffer/bufmgr.c | 123 ++++++++++++++++++----------
 src/tools/pgindent/typedefs.list    |   1 +
 2 files changed, 79 insertions(+), 45 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index d4235ca7939..972eae1fc7b 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -90,10 +90,18 @@
  */
 #define BUF_DROP_FULL_SCAN_THRESHOLD		(uint64) (NBuffers / 32)
 
-typedef struct PrivateRefCountEntry
+typedef struct PrivateRefCountData
 {
-	Buffer		buffer;
+	/*
+	 * How many times has the buffer been pinned by this backend.
+	 */
 	int32		refcount;
+} PrivateRefCountData;
+
+typedef struct PrivateRefCountEntry
+{
+	Buffer		buffer;
+	PrivateRefCountData data;
 } PrivateRefCountEntry;
 
 /* 64 bytes, about the size of a cache line on common systems */
@@ -212,11 +220,12 @@ static BufferDesc *PinCountWaitBuf = NULL;
  * memory allocations in NewPrivateRefCountEntry() which can be important
  * because in some scenarios it's called with a spinlock held...
  */
+static Buffer PrivateRefCountArrayKeys[REFCOUNT_ARRAY_ENTRIES];
 static struct PrivateRefCountEntry PrivateRefCountArray[REFCOUNT_ARRAY_ENTRIES];
 static HTAB *PrivateRefCountHash = NULL;
 static int32 PrivateRefCountOverflowed = 0;
 static uint32 PrivateRefCountClock = 0;
-static PrivateRefCountEntry *ReservedRefCountEntry = NULL;
+static int	ReservedRefCountSlot = -1;
 
 static uint32 MaxProportionalPins;
 
@@ -259,7 +268,7 @@ static void
 ReservePrivateRefCountEntry(void)
 {
 	/* Already reserved (or freed), nothing to do */
-	if (ReservedRefCountEntry != NULL)
+	if (ReservedRefCountSlot != -1)
 		return;
 
 	/*
@@ -271,16 +280,19 @@ ReservePrivateRefCountEntry(void)
 
 		for (i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
 		{
-			PrivateRefCountEntry *res;
-
-			res = &PrivateRefCountArray[i];
-
-			if (res->buffer == InvalidBuffer)
+			if (PrivateRefCountArrayKeys[i] == InvalidBuffer)
 			{
-				ReservedRefCountEntry = res;
-				return;
+				ReservedRefCountSlot = i;
+
+				/*
+				 * We could return immediately, but iterating till the end of
+				 * the array allows compiler-autovectorization.
+				 */
 			}
 		}
+
+		if (ReservedRefCountSlot != -1)
+			return;
 	}
 
 	/*
@@ -292,27 +304,34 @@ ReservePrivateRefCountEntry(void)
 		 * Move entry from the current clock position in the array into the
 		 * hashtable. Use that slot.
 		 */
+		int			victim_slot;
+		PrivateRefCountEntry *victim_entry;
 		PrivateRefCountEntry *hashent;
 		bool		found;
 
 		/* select victim slot */
-		ReservedRefCountEntry =
-			&PrivateRefCountArray[PrivateRefCountClock++ % REFCOUNT_ARRAY_ENTRIES];
+		victim_slot = PrivateRefCountClock++ % REFCOUNT_ARRAY_ENTRIES;
+		victim_entry = &PrivateRefCountArray[victim_slot];
+		ReservedRefCountSlot = victim_slot;
 
 		/* Better be used, otherwise we shouldn't get here. */
-		Assert(ReservedRefCountEntry->buffer != InvalidBuffer);
+		Assert(PrivateRefCountArrayKeys[victim_slot] != InvalidBuffer);
+		Assert(PrivateRefCountArray[victim_slot].buffer != InvalidBuffer);
+		Assert(PrivateRefCountArrayKeys[victim_slot] == PrivateRefCountArray[victim_slot].buffer);
 
 		/* enter victim array entry into hashtable */
 		hashent = hash_search(PrivateRefCountHash,
-							  &(ReservedRefCountEntry->buffer),
+							  &PrivateRefCountArrayKeys[victim_slot],
 							  HASH_ENTER,
 							  &found);
 		Assert(!found);
-		hashent->refcount = ReservedRefCountEntry->refcount;
+		hashent->data = victim_entry->data;
 
 		/* clear the now free array slot */
-		ReservedRefCountEntry->buffer = InvalidBuffer;
-		ReservedRefCountEntry->refcount = 0;
+		PrivateRefCountArrayKeys[victim_slot] = InvalidBuffer;
+		victim_entry->buffer = InvalidBuffer;
+		memset(&victim_entry->data, 0, sizeof(victim_entry->data));
+		victim_entry->data.refcount = 0;
 
 		PrivateRefCountOverflowed++;
 	}
@@ -327,15 +346,17 @@ NewPrivateRefCountEntry(Buffer buffer)
 	PrivateRefCountEntry *res;
 
 	/* only allowed to be called when a reservation has been made */
-	Assert(ReservedRefCountEntry != NULL);
+	Assert(ReservedRefCountSlot != -1);
 
 	/* use up the reserved entry */
-	res = ReservedRefCountEntry;
-	ReservedRefCountEntry = NULL;
+	res = &PrivateRefCountArray[ReservedRefCountSlot];
 
 	/* and fill it */
+	PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
 	res->buffer = buffer;
-	res->refcount = 0;
+	res->data.refcount = 0;
+
+	ReservedRefCountSlot = -1;
 
 	return res;
 }
@@ -347,10 +368,11 @@ NewPrivateRefCountEntry(Buffer buffer)
  * do_move is true, and the entry resides in the hashtable the entry is
  * optimized for frequent access by moving it to the array.
  */
-static PrivateRefCountEntry *
+static inline PrivateRefCountEntry *
 GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 {
 	PrivateRefCountEntry *res;
+	int			match = -1;
 	int			i;
 
 	Assert(BufferIsValid(buffer));
@@ -362,12 +384,16 @@ GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 	 */
 	for (i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
 	{
-		res = &PrivateRefCountArray[i];
-
-		if (res->buffer == buffer)
-			return res;
+		if (PrivateRefCountArrayKeys[i] == buffer)
+		{
+			match = i;
+			/* see ReservePrivateRefCountEntry() for why we don't return */
+		}
 	}
 
+	if (match != -1)
+		return &PrivateRefCountArray[match];
+
 	/*
 	 * By here we know that the buffer, if already pinned, isn't residing in
 	 * the array.
@@ -397,14 +423,18 @@ GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 		ReservePrivateRefCountEntry();
 
 		/* Use up the reserved slot */
-		Assert(ReservedRefCountEntry != NULL);
-		free = ReservedRefCountEntry;
-		ReservedRefCountEntry = NULL;
+		Assert(ReservedRefCountSlot != -1);
+		free = &PrivateRefCountArray[ReservedRefCountSlot];
+		Assert(PrivateRefCountArrayKeys[ReservedRefCountSlot] == free->buffer);
 		Assert(free->buffer == InvalidBuffer);
 
 		/* and fill it */
 		free->buffer = buffer;
-		free->refcount = res->refcount;
+		free->data = res->data;
+		PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
+
+		ReservedRefCountSlot = -1;
+
 
 		/* delete from hashtable */
 		hash_search(PrivateRefCountHash, &buffer, HASH_REMOVE, &found);
@@ -437,7 +467,7 @@ GetPrivateRefCount(Buffer buffer)
 
 	if (ref == NULL)
 		return 0;
-	return ref->refcount;
+	return ref->data.refcount;
 }
 
 /*
@@ -447,19 +477,21 @@ GetPrivateRefCount(Buffer buffer)
 static void
 ForgetPrivateRefCountEntry(PrivateRefCountEntry *ref)
 {
-	Assert(ref->refcount == 0);
+	Assert(ref->data.refcount == 0);
 
 	if (ref >= &PrivateRefCountArray[0] &&
 		ref < &PrivateRefCountArray[REFCOUNT_ARRAY_ENTRIES])
 	{
 		ref->buffer = InvalidBuffer;
+		PrivateRefCountArrayKeys[ref - PrivateRefCountArray] = InvalidBuffer;
+
 
 		/*
 		 * Mark the just used entry as reserved - in many scenarios that
 		 * allows us to avoid ever having to search the array/hash for free
 		 * entries.
 		 */
-		ReservedRefCountEntry = ref;
+		ReservedRefCountSlot = ref - PrivateRefCountArray;
 	}
 	else
 	{
@@ -3073,7 +3105,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 	PrivateRefCountEntry *ref;
 
 	Assert(!BufferIsLocal(b));
-	Assert(ReservedRefCountEntry != NULL);
+	Assert(ReservedRefCountSlot != -1);
 
 	ref = GetPrivateRefCountEntry(b, true);
 
@@ -3145,8 +3177,8 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		 */
 		result = (pg_atomic_read_u64(&buf->state) & BM_VALID) != 0;
 
-		Assert(ref->refcount > 0);
-		ref->refcount++;
+		Assert(ref->data.refcount > 0);
+		ref->data.refcount++;
 		ResourceOwnerRememberBuffer(CurrentResourceOwner, b);
 	}
 
@@ -3263,9 +3295,9 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	/* not moving as we're likely deleting it soon anyway */
 	ref = GetPrivateRefCountEntry(b, false);
 	Assert(ref != NULL);
-	Assert(ref->refcount > 0);
-	ref->refcount--;
-	if (ref->refcount == 0)
+	Assert(ref->data.refcount > 0);
+	ref->data.refcount--;
+	if (ref->data.refcount == 0)
 	{
 		uint64		old_buf_state;
 
@@ -3305,7 +3337,7 @@ TrackNewBufferPin(Buffer buf)
 	PrivateRefCountEntry *ref;
 
 	ref = NewPrivateRefCountEntry(buf);
-	ref->refcount++;
+	ref->data.refcount++;
 
 	ResourceOwnerRememberBuffer(CurrentResourceOwner, buf);
 
@@ -4018,6 +4050,7 @@ InitBufferManagerAccess(void)
 	MaxProportionalPins = NBuffers / (MaxBackends + NUM_AUXILIARY_PROCS);
 
 	memset(&PrivateRefCountArray, 0, sizeof(PrivateRefCountArray));
+	memset(&PrivateRefCountArrayKeys, 0, sizeof(Buffer));
 
 	hash_ctl.keysize = sizeof(int32);
 	hash_ctl.entrysize = sizeof(PrivateRefCountEntry);
@@ -4067,10 +4100,10 @@ CheckForBufferLeaks(void)
 	/* check the array */
 	for (i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
 	{
-		res = &PrivateRefCountArray[i];
-
-		if (res->buffer != InvalidBuffer)
+		if (PrivateRefCountArrayKeys[i] != InvalidBuffer)
 		{
+			res = &PrivateRefCountArray[i];
+
 			s = DebugPrintBufferRefcount(res->buffer);
 			elog(WARNING, "buffer refcount leak: %s", s);
 			pfree(s);
@@ -5407,7 +5440,7 @@ IncrBufferRefCount(Buffer buffer)
 
 		ref = GetPrivateRefCountEntry(buffer, true);
 		Assert(ref != NULL);
-		ref->refcount++;
+		ref->data.refcount++;
 	}
 	ResourceOwnerRememberBuffer(CurrentResourceOwner, buffer);
 }
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 5769227e41c..9a89e68c59c 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2327,6 +2327,7 @@ PrintfArgValue
 PrintfTarget
 PrinttupAttrInfo
 PrivTarget
+PrivateRefCountData
 PrivateRefCountEntry
 ProcArrayStruct
 ProcLangInfo
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0007-bufmgr-Add-one-entry-cache-for-private-refcount.patch (1.9K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/8-v6-0007-bufmgr-Add-one-entry-cache-for-private-refcount.patch)
  download | inline diff:
From b503f94124620433d1e8276341079c70d037d360 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 18 Nov 2025 09:39:59 -0500
Subject: [PATCH v6 07/14] bufmgr: Add one-entry cache for private refcount

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/storage/buffer/bufmgr.c | 17 +++++++++++++++++
 1 file changed, 17 insertions(+)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 972eae1fc7b..b7ce4bafdea 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -226,6 +226,8 @@ static HTAB *PrivateRefCountHash = NULL;
 static int32 PrivateRefCountOverflowed = 0;
 static uint32 PrivateRefCountClock = 0;
 static int	ReservedRefCountSlot = -1;
+static int	PrivateRefcountEntryLast = -1;
+
 
 static uint32 MaxProportionalPins;
 
@@ -356,6 +358,8 @@ NewPrivateRefCountEntry(Buffer buffer)
 	res->buffer = buffer;
 	res->data.refcount = 0;
 
+	PrivateRefcountEntryLast = ReservedRefCountSlot;
+
 	ReservedRefCountSlot = -1;
 
 	return res;
@@ -378,6 +382,16 @@ GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 	Assert(BufferIsValid(buffer));
 	Assert(!BufferIsLocal(buffer));
 
+	/*
+	 * It's very common to look up the same buffer repeatedly. To make that
+	 * fast, we have a one-entry cache.
+	 */
+	if (likely(PrivateRefcountEntryLast != -1) &&
+		likely(PrivateRefCountArrayKeys[PrivateRefcountEntryLast] == buffer))
+	{
+		return &PrivateRefCountArray[PrivateRefcountEntryLast];
+	}
+
 	/*
 	 * First search for references in the array, that'll be sufficient in the
 	 * majority of cases.
@@ -392,7 +406,10 @@ GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 	}
 
 	if (match != -1)
+	{
+		PrivateRefcountEntryLast = match;
 		return &PrivateRefCountArray[match];
+	}
 
 	/*
 	 * By here we know that the buffer, if already pinned, isn't residing in
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0008-bufmgr-Implement-buffer-content-locks-independent.patch (39.9K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/9-v6-0008-bufmgr-Implement-buffer-content-locks-independent.patch)
  download | inline diff:
From 905c67a8387286bbd39ebca49c2b127403de4e80 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 19 Nov 2025 16:37:26 -0500
Subject: [PATCH v6 08/14] bufmgr: Implement buffer content locks independently
 of lwlocks

Until now buffer content locks were implemented using lwlocks. That has the
obvious advantage of not needing a separate efficient implementation of
locks. However, the time for a dedicated buffer content lock implementation
has come:

1) Hint bits are currently set while holding only a share lock. This leads to
   having to copy pages while they are being written out if checksums are
   enabled, which is not cheap. We would like to add AIO writes, however once
   many buffers can be written out at the same time, it gets a lot more
   expensive to copy them, particularly because that copy needs to reside in
   shared buffers (for worker mode to have access to the buffer).

   In addition, modifying buffers while they are being written out can cause
   issues with unbuffered/direct-IO, as some filesystems (like btrfs) do not
   like that, due to filesystem internal checksums getting corrupted.

   The solution to this is to require a new share-exclusive lock-level to set
   hint bits and to write out buffers, making those operations mutually
   exclusive. We could introduce such a lock level into the generic lwlock
   implementation, however it does not look like there would be other users,
   and it does add some overhead into important codepaths.

2) For AIO writes we need to be able to race-freely check whether a buffer is
   undergoing IO and whether an exclusive lock on the page can be acquired. That
   is rather hard to do efficiently when the buffer state and the lock state
   are separate atomic variables. This is a major hindrance to allowing writes
   to be done asynchronously.

3) Buffer locks are by far the most frequently taken locks. Optimizing them
   specifically for their use case is worth the effort. E.g. by merging
   content locks into buffer locks we will be able to release a buffer lock
   and pin in one atomic operation.

4) There are more complicated optimizations, like long-lived "super pinned &
   locked" pages, that cannot realistically be implemented with the generic
   lwlock implementation.

Therefore implement content locks inside bufmgr.c. The lockstate is stored as
part of BufferDesc.state. The implementation of buffer content locks is fairly
similar to lwlocks, with a few important differences:

1) An additional lock-level share-exclusive has been added. This lock level
   conflicts with exclusive locks and itself, but not share locks.

2) Error recovery for content locks is implemented as part of the already
   existing private-refcount tracking mechanism in combination with resowners,
   instead of a bespoke mechanism as the case for lwlocks. This means we do
   not need to add dedicated error-recovery codepaths to release all content
   locks (like done with LWLockReleaseAll() for content locks).

3) The lock state is embedded in BufferDesc.state instead of having its own
   struct.

4) The wakeup logic is a tad more complicated due to needing to support the
   additional lock level

This commit unfortunately introduces some code that is very similar to the
code in lwlock.c, however the code is not equivalent enough to easily merge
it. The future wins that this commit makes possible seem worth the cost.

As of this commit nothing uses the new share-exclusive lock mode. It will be
documented and used in a future commit. It seemed too complicated to introduce
the lock-level in a separate commit.

TODO:
- Address FIXMEs

- Perhaps move the locking code into a buffer_locking.h or such? Needs to be
  inline functions for efficiency unfortunately.

- reflow some comments that I didn't reflow to make the diff more readable

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  55 +-
 src/include/storage/bufmgr.h                  |  20 +-
 src/include/storage/proc.h                    |   8 +-
 src/backend/postmaster/auxprocess.c           |   1 +
 src/backend/storage/buffer/buf_init.c         |   7 +-
 src/backend/storage/buffer/bufmgr.c           | 776 ++++++++++++++++--
 .../utils/activity/wait_event_names.txt       |   3 +
 7 files changed, 804 insertions(+), 66 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 28519ad2813..0a145d95024 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -23,6 +23,7 @@
 #include "storage/condition_variable.h"
 #include "storage/lwlock.h"
 #include "storage/procnumber.h"
+#include "storage/proclist_types.h"
 #include "storage/shmem.h"
 #include "storage/smgr.h"
 #include "storage/spin.h"
@@ -32,22 +33,29 @@
 /*
  * Buffer state is a single 64-bit variable where following data is combined.
  *
+ * State of the buffer itself:
  * - 18 bits refcount
  * - 4 bits usage count
  * - 10 bits of flags
  *
+ * State of the content lock:
+ * - 1 bit has_waiter
+ * - 1 bit release_ok
+ * - 1 bit lock state locked
+ * - 1 bit exclusively locked
+ * - 1 bit share exclusively locked
+ * - 18 bits share lock count
+ *
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
- * NB: A future commit will use a significant portion of the remaining bits to
- * implement buffer locking as part of the state variable.
- *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
 #define BUF_USAGECOUNT_BITS 4
 #define BUF_FLAG_BITS 10
 
+/* FIXME: Also assert lock state size, just not yet sure how */
 StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 				 "parts of buffer state space need to equal 32");
 
@@ -69,7 +77,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 	((uint32)(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT))
 
 /*
- * Flags for buffer descriptors
+ * Flags for buffer descriptor state
  *
  * Note: BM_TAG_VALID essentially means that there is a buffer hashtable
  * entry associated with the buffer's tag.
@@ -111,6 +119,20 @@ StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
 				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
 
+
+/*
+ * Definitions related to buffer content locks
+ */
+#define BM_LOCK_HAS_WAITERS         (UINT64CONST(1) << 63)
+#define BM_LOCK_RELEASE_OK          (UINT64CONST(1) << 62)
+
+#define BM_LOCK_VAL_SHARED          (UINT64CONST(1) << 32)
+#define BM_LOCK_VAL_SHARE_EXCLUSIVE (UINT64CONST(1) << (32 + MAX_BACKENDS_BITS))
+#define BM_LOCK_VAL_EXCLUSIVE       (UINT64CONST(1) << (32 + 1 + MAX_BACKENDS_BITS))
+
+#define BM_LOCK_MASK                (((uint64)MAX_BACKENDS << 32) | BM_LOCK_VAL_SHARE_EXCLUSIVE | BM_LOCK_VAL_EXCLUSIVE)
+
+
 /*
  * Buffer tag identifies which disk block the buffer contains.
  *
@@ -253,9 +275,6 @@ BufMappingPartitionLockByIndex(uint32 index)
  * it is held.  However, existing buffer pins may be released while the buffer
  * header spinlock is held, using an atomic subtraction.
  *
- * The LWLock can take care of itself.  The buffer header lock is *not* used
- * to control access to the data in the buffer!
- *
  * If we have the buffer pinned, its tag can't change underneath us, so we can
  * examine the tag without locking the buffer header.  Also, in places we do
  * one-time reads of the flags without bothering to lock the buffer header;
@@ -268,6 +287,15 @@ BufMappingPartitionLockByIndex(uint32 index)
  * wait_backend_pgprocno and setting flag bit BM_PIN_COUNT_WAITER.  At present,
  * there can be only one such waiter per buffer.
  *
+ * The content of buffers is protected via the buffer content lock,
+ * implemented as part buffer state. Note that the buffer header lock is *not*
+ * used to control access to the data in the buffer! We used to use an LWLock
+ * to implement the content lock, but having a dedicated implementation of
+ * content locks allows to implement some otherwise hard things (e.g.
+ * race-freely checking if AIO is in progress before locking a buffer
+ * exclusively) and makes otherwise impossible optimizations possible
+ * (e.g. unlocking and unpinning a buffer in one atomic operation).
+ *
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
@@ -309,7 +337,12 @@ typedef struct BufferDesc
 	int			wait_backend_pgprocno;
 
 	PgAioWaitRef io_wref;		/* set iff AIO is in progress */
-	LWLock		content_lock;	/* to lock access to buffer contents */
+
+	/*
+	 * List of PGPROCs waiting for the buffer content lock. Protected by the
+	 * buffer header spinlock.
+	 */
+	proclist_head lock_waiters;
 } BufferDesc;
 
 /*
@@ -396,12 +429,6 @@ BufferDescriptorGetIOCV(const BufferDesc *bdesc)
 	return &(BufferIOCVArray[bdesc->buf_id]).cv;
 }
 
-static inline LWLock *
-BufferDescriptorGetContentLock(const BufferDesc *bdesc)
-{
-	return (LWLock *) (&bdesc->content_lock);
-}
-
 /*
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 5fc3de20abc..8e442492d4d 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -204,6 +204,7 @@ typedef enum BufferLockMode
 {
 	BUFFER_LOCK_UNLOCK,
 	BUFFER_LOCK_SHARE,
+	BUFFER_LOCK_SHARE_EXCLUSIVE,
 	BUFFER_LOCK_EXCLUSIVE,
 } BufferLockMode;
 
@@ -302,7 +303,24 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
 extern void UnlockBuffers(void);
-extern void LockBuffer(Buffer buffer, BufferLockMode mode);
+extern void UnlockBuffer(Buffer buffer);
+extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
+
+/*
+ * Handling BUFFER_LOCK_UNLOCK in bufmgr.c leads to sufficiently worse branch
+ * prediction to impact performance. Therefore handle that switch here, were
+ * most of the time `mode` will be a constant and thus can be optimized out by
+ * the compiler.
+ */
+static inline void
+LockBuffer(Buffer buffer, BufferLockMode mode)
+{
+	if (mode == BUFFER_LOCK_UNLOCK)
+		UnlockBuffer(buffer);
+	else
+		LockBufferInternal(buffer, mode);
+}
+
 extern bool ConditionalLockBuffer(Buffer buffer);
 extern void LockBufferForCleanup(Buffer buffer);
 extern bool ConditionalLockBufferForCleanup(Buffer buffer);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index c6f5ebceefd..d1f6d314d57 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -242,7 +242,13 @@ struct PGPROC
 	 */
 	bool		recoveryConflictPending;
 
-	/* Info about LWLock the process is currently waiting for, if any. */
+	/*
+	 * Info about LWLock the process is currently waiting for, if any.
+	 *
+	 * This is currently used both for lwlocks and buffer content locks, which
+	 * is acceptable, although not pretty, because a backend can't wait for
+	 * both types of locks at the same time.
+	 */
 	uint8		lwWaiting;		/* see LWLockWaitState */
 	uint8		lwWaitMode;		/* lwlock mode being waited for */
 	proclist_node lwWaitLink;	/* position in LW lock wait list */
diff --git a/src/backend/postmaster/auxprocess.c b/src/backend/postmaster/auxprocess.c
index a6d3630398f..2012f959fd9 100644
--- a/src/backend/postmaster/auxprocess.c
+++ b/src/backend/postmaster/auxprocess.c
@@ -18,6 +18,7 @@
 #include "miscadmin.h"
 #include "pgstat.h"
 #include "postmaster/auxprocess.h"
+#include "storage/bufmgr.h"
 #include "storage/condition_variable.h"
 #include "storage/ipc.h"
 #include "storage/proc.h"
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 25f71191ec3..f3224c793c4 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
 #include "storage/aio.h"
 #include "storage/buf_internals.h"
 #include "storage/bufmgr.h"
+#include "storage/proclist.h"
 
 BufferDescPadded *BufferDescriptors;
 char	   *BufferBlocks;
@@ -121,16 +122,14 @@ BufferManagerShmemInit(void)
 
 			ClearBufferTag(&buf->tag);
 
-			pg_atomic_init_u64(&buf->state, 0);
+			pg_atomic_init_u64(&buf->state, BM_LOCK_RELEASE_OK);
 			buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
 
 			buf->buf_id = i;
 
 			pgaio_wref_clear(&buf->io_wref);
 
-			LWLockInitialize(BufferDescriptorGetContentLock(buf),
-							 LWTRANCHE_BUFFER_CONTENT);
-
+			proclist_init(&buf->lock_waiters);
 			ConditionVariableInit(BufferDescriptorGetIOCV(buf));
 		}
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index b7ce4bafdea..da83b775d0b 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -58,6 +58,7 @@
 #include "storage/ipc.h"
 #include "storage/lmgr.h"
 #include "storage/proc.h"
+#include "storage/proclist.h"
 #include "storage/read_stream.h"
 #include "storage/smgr.h"
 #include "storage/standby.h"
@@ -96,6 +97,12 @@ typedef struct PrivateRefCountData
 	 * How many times has the buffer been pinned by this backend.
 	 */
 	int32		refcount;
+
+	/*
+	 * Is the buffer locked by this backend? BUFFER_LOCK_UNLOCK indicates that
+	 * the buffer is not locked.
+	 */
+	BufferLockMode lockmode;
 } PrivateRefCountData;
 
 typedef struct PrivateRefCountEntry
@@ -334,6 +341,7 @@ ReservePrivateRefCountEntry(void)
 		victim_entry->buffer = InvalidBuffer;
 		memset(&victim_entry->data, 0, sizeof(victim_entry->data));
 		victim_entry->data.refcount = 0;
+		victim_entry->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 		PrivateRefCountOverflowed++;
 	}
@@ -357,6 +365,7 @@ NewPrivateRefCountEntry(Buffer buffer)
 	PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
 	res->buffer = buffer;
 	res->data.refcount = 0;
+	res->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 	PrivateRefcountEntryLast = ReservedRefCountSlot;
 
@@ -495,6 +504,7 @@ static void
 ForgetPrivateRefCountEntry(PrivateRefCountEntry *ref)
 {
 	Assert(ref->data.refcount == 0);
+	Assert(ref->data.lockmode == BUFFER_LOCK_UNLOCK);
 
 	if (ref >= &PrivateRefCountArray[0] &&
 		ref < &PrivateRefCountArray[REFCOUNT_ARRAY_ENTRIES])
@@ -604,6 +614,20 @@ static inline int buffertag_comparator(const BufferTag *ba, const BufferTag *bb)
 static inline int ckpt_buforder_comparator(const CkptSortItem *a, const CkptSortItem *b);
 static int	ts_ckpt_progress_comparator(Datum a, Datum b, void *arg);
 
+static void BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr);
+static bool BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMe(BufferDesc *buf_hdr);
+static inline void BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr);
+static inline int BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr);
+static inline bool BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockDequeueSelf(BufferDesc *buf_hdr);
+static void BufferLockWakeup(BufferDesc *buf_hdr, bool unlocked);
+static void BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate);
+static inline uint64 BufferLockReleaseSub(BufferLockMode mode);
+
 
 /*
  * Implementation of PrefetchBuffer() for shared buffers.
@@ -2404,8 +2428,6 @@ again:
 	 */
 	if (buf_state & BM_DIRTY)
 	{
-		LWLock	   *content_lock;
-
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(buf_state & BM_VALID);
 
@@ -2423,8 +2445,7 @@ again:
 		 * one just happens to be trying to split the page the first one got
 		 * from StrategyGetBuffer.)
 		 */
-		content_lock = BufferDescriptorGetContentLock(buf_hdr);
-		if (!LWLockConditionalAcquire(content_lock, LW_SHARED))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -2453,7 +2474,7 @@ again:
 			if (XLogNeedsFlush(lsn)
 				&& StrategyRejectBuffer(strategy, buf_hdr, from_ring))
 			{
-				LWLockRelease(content_lock);
+				LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 				UnpinBuffer(buf_hdr);
 				goto again;
 			}
@@ -2461,7 +2482,7 @@ again:
 
 		/* OK, do the I/O */
 		FlushBuffer(buf_hdr, NULL, IOOBJECT_RELATION, io_context);
-		LWLockRelease(content_lock);
+		LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 		ScheduleBufferTagForWriteback(&BackendWritebackContext, io_context,
 									  &buf_hdr->tag);
@@ -2903,7 +2924,7 @@ BufferIsLockedByMe(Buffer buffer)
 	else
 	{
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr));
+		return BufferLockHeldByMe(bufHdr);
 	}
 }
 
@@ -2928,23 +2949,8 @@ BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode)
 	}
 	else
 	{
-		LWLockMode	lw_mode;
-
-		switch (mode)
-		{
-			case BUFFER_LOCK_EXCLUSIVE:
-				lw_mode = LW_EXCLUSIVE;
-				break;
-			case BUFFER_LOCK_SHARE:
-				lw_mode = LW_SHARED;
-				break;
-			default:
-				pg_unreachable();
-		}
-
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-									lw_mode);
+		return BufferLockHeldByMeInMode(bufHdr, mode);
 	}
 }
 
@@ -3331,7 +3337,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 * I'd better not still hold the buffer content lock. Can't use
 		 * BufferIsLockedByMe(), as that asserts the buffer is pinned.
 		 */
-		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
+		Assert(!BufferLockHeldByMe(buf));
 
 		/* decrement the shared reference count */
 		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
@@ -4176,6 +4182,8 @@ static void
 AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
 						   void *unused_context)
 {
+	/* FIXME */
+#ifdef NOT_YET
 	BufferDesc *bufHdr;
 	BufferTag	tag;
 	Oid			relid;
@@ -4205,6 +4213,7 @@ AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
 		return;
 
 	Assert(!IsCatalogRelationOid(relid));
+#endif
 }
 #endif
 
@@ -4470,9 +4479,11 @@ static void
 FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 					IOObject io_object, IOContext io_context)
 {
-	LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	Buffer		buffer = BufferDescriptorGetBuffer(buf);
+
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-	LWLockRelease(BufferDescriptorGetContentLock(buf));
+	BufferLockUnlock(buffer, buf);
 }
 
 /*
@@ -5615,9 +5626,10 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
  *
  * Used to clean up after errors.
  *
- * Currently, we can expect that lwlock.c's LWLockReleaseAll() took care
- * of releasing buffer content locks per se; the only thing we need to deal
- * with here is clearing any PIN_COUNT request that was in progress.
+ * Currently, we can expect that resource owner cleanup, via
+ * ResOwnerReleaseBufferPin(), took care releasing buffer content locks per
+ * se; the only thing we need to deal with here is clearing any PIN_COUNT
+ * request that was in progress.
  */
 void
 UnlockBuffers(void)
@@ -5648,25 +5660,684 @@ UnlockBuffers(void)
 }
 
 /*
- * Acquire or release the content_lock for the buffer.
+ * Acquire the buffer content lock in the specified mode
+ *
+ * If the lock is not available, sleep until it is.
+ *
+ * Side effect: cancel/die interrupts are held off until lock release.
+ *
+ * This uses almost the same locking approach as lwlock.c's
+ * LWLockAcquire(). See documentation atop of lwlock.c for a more detailed
+ * discussion.
+ *
+ * The reason that this, and most of the other BufferLock* functions, get both
+ * the Buffer and BufferDesc* as parameters, is that looking up one from the
+ * other repeatedly shows up noticeably in profiles.
+ */
+static inline void
+BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry;
+	int			extraWaits = 0;
+
+	/*
+	 * Get reference to the refcount entry before we hold the lock, it seems
+	 * better to do before holding the lock.
+	 */
+	entry = GetPrivateRefCountEntry(buffer, true);
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	for (;;)
+	{
+		bool		mustwait;
+		uint32		wait_event;
+
+		/*
+		 * Try to grab the lock the first time, we're not in the waitqueue
+		 * yet/anymore.
+		 */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		if (likely(!mustwait))
+		{
+			break;
+		}
+
+		/*
+		 * Ok, at this point we couldn't grab the lock on the first try. We
+		 * cannot simply queue ourselves to the end of the list and wait to be
+		 * woken up because by now the lock could long have been released.
+		 * Instead add us to the queue and try to grab the lock again. If we
+		 * succeed we need to revert the queuing and be happy, otherwise we
+		 * recheck the lock. If we still couldn't grab it, we know that the
+		 * other locker will see our queue entries when releasing since they
+		 * existed before we checked for the lock.
+		 */
+
+		/* add to the queue */
+		BufferLockQueueSelf(buf_hdr, mode);
+
+		/* we're now guaranteed to be woken up if necessary */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		/* ok, grabbed the lock the second time round, need to undo queueing */
+		if (!mustwait)
+		{
+			BufferLockDequeueSelf(buf_hdr);
+			break;
+		}
+
+		switch (mode)
+		{
+			case BUFFER_LOCK_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_SHARE_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE:
+				wait_event = WAIT_EVENT_BUFFER_SHARED;
+				break;
+			case BUFFER_LOCK_UNLOCK:
+				pg_unreachable();
+
+		}
+		pgstat_report_wait_start(wait_event);
+
+		/*
+		 * Wait until awakened.
+		 *
+		 * It is possible that we get awakened for a reason other than being
+		 * signaled by LWLockRelease.  If so, loop back and wait again.  Once
+		 * we've gotten the LWLock, re-increment the sema by the number of
+		 * additional signals received.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		pgstat_report_wait_end();
+
+		/* Retrying, allow BufferLockRelease to release waiters again. */
+		pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_RELEASE_OK);
+	}
+
+	/* Remember that we now hold this lock */
+	entry->data.lockmode = mode;
+
+	/*
+	 * Fix the process wait semaphore's count for any absorbed wakeups.
+	 */
+	while (unlikely(extraWaits-- > 0))
+		PGSemaphoreUnlock(MyProc->sem);
+}
+
+/*
+ * Release a previously acquired buffer content lock.
+ */
+static void
+BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	uint64		oldstate;
+	uint64		sub;
+
+	mode = BufferLockDisownInternal(buffer, buf_hdr);
+
+	/*
+	 * Release my hold on lock, after that it can immediately be acquired by
+	 * others, even if we still have to wakeup other waiters.
+	 */
+	sub = BufferLockReleaseSub(mode);
+
+	oldstate = pg_atomic_sub_fetch_u64(&buf_hdr->state, sub);
+
+	BufferLockProcessRelease(buf_hdr, mode, oldstate);
+
+	/*
+	 * Now okay to allow cancel/die interrupts.
+	 */
+	RESUME_INTERRUPTS();
+}
+
+
+/*
+ * Acquire the content lock for the buffer, but only if we don't have to wait.
+ */
+static bool
+BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	bool		mustwait;
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	/* Check for the lock */
+	mustwait = BufferLockAttempt(buf_hdr, mode);
+
+	if (mustwait)
+	{
+		/* Failed to get lock, so release interrupt holdoff */
+		RESUME_INTERRUPTS();
+	}
+	else
+	{
+		PrivateRefCountEntry *entry =
+			GetPrivateRefCountEntry(buffer, true);
+
+		entry->data.lockmode = mode;
+	}
+
+	return !mustwait;
+}
+
+/*
+ * Internal function that tries to atomically acquire the content lock in the
+ * passed in mode.
+ *
+ * This function will not block waiting for a lock to become free - that's the
+ * caller's job.
+ *
+ * Similar to LWLockAttemptLock().
+ */
+static inline bool
+BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	uint64		old_state;
+
+	/*
+	 * Read once outside the loop, later iterations will get the newer value
+	 * via compare & exchange.
+	 */
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+
+	/* loop until we've determined whether we could acquire the lock or not */
+	while (true)
+	{
+		uint64		desired_state;
+		bool		lock_free;
+
+		desired_state = old_state;
+
+		if (mode == BUFFER_LOCK_EXCLUSIVE)
+		{
+			lock_free = (old_state & BM_LOCK_MASK) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_EXCLUSIVE;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			lock_free = (old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+		}
+		else
+		{
+			lock_free = (old_state & BM_LOCK_VAL_EXCLUSIVE) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARED;
+		}
+
+		/*
+		 * Attempt to swap in the state we are expecting. If we didn't see
+		 * lock to be free, that's just the old value. If we saw it as free,
+		 * we'll attempt to mark it acquired. The reason that we always swap
+		 * in the value is that this doubles as a memory barrier. We could try
+		 * to be smarter and only swap in values if we saw the lock as free,
+		 * but benchmark haven't shown it as beneficial so far.
+		 *
+		 * Retry if the value changed since we last looked at it.
+		 */
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			if (lock_free)
+			{
+				/* Great! Got the lock. */
+				return false;
+			}
+			else
+				return true;	/* somebody else has the lock */
+		}
+	}
+
+	pg_unreachable();
+}
+
+/*
+ * Add ourselves to the end of the content lock's wait queue.
+ */
+static void
+BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	/*
+	 * If we don't have a PGPROC structure, there's no way to wait. This
+	 * should never occur, since MyProc should only be null during shared
+	 * memory initialization.
+	 */
+	if (MyProc == NULL)
+		elog(PANIC, "cannot wait without a PGPROC structure");
+
+	if (MyProc->lwWaiting != LW_WS_NOT_WAITING)
+		elog(PANIC, "queueing for lock while waiting on another one");
+
+	LockBufHdr(buf_hdr);
+
+	/* setting the flag is protected by the spinlock */
+	pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_HAS_WAITERS);
+
+	/*
+	 * FIXME: This is reusing the lwlock fields. That's not a correctness
+	 * issue, a backend can't wait for both an lwlock and a buffer content
+	 * lock at the same time. However, it seems pretty ugly, particularly
+	 * given that the field names have an lw* prefix. But duplicating the
+	 * fields also seems somewhat superfluous.
+	 */
+	MyProc->lwWaiting = LW_WS_WAITING;
+	MyProc->lwWaitMode = mode;
+
+	proclist_push_tail(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	/* Can release the mutex now */
+	UnlockBufHdr(buf_hdr);
+}
+
+/*
+ * Remove ourselves from the waitlist.
+ *
+ * This is used if we queued ourselves because we thought we needed to sleep
+ * but, after further checking, we discovered that we don't actually need to
+ * do so.
+ */
+static void
+BufferLockDequeueSelf(BufferDesc *buf_hdr)
+{
+	bool		on_waitlist;
+
+	LockBufHdr(buf_hdr);
+
+	on_waitlist = MyProc->lwWaiting == LW_WS_WAITING;
+	if (on_waitlist)
+		proclist_delete(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	if (proclist_is_empty(&buf_hdr->lock_waiters) &&
+		(pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS) != 0)
+	{
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_HAS_WAITERS);
+	}
+
+	/* XXX: combine with fetch_and above? */
+	UnlockBufHdr(buf_hdr);
+
+	/* clear waiting state again, nice for debugging */
+	if (on_waitlist)
+		MyProc->lwWaiting = LW_WS_NOT_WAITING;
+	else
+	{
+		int			extraWaits = 0;
+
+
+		/*
+		 * Somebody else dequeued us and has or will wake us up. Deal with the
+		 * superfluous absorption of a wakeup.
+		 */
+
+		/*
+		 * Reset RELEASE_OK flag if somebody woke us before we removed
+		 * ourselves - they'll have set it to false.
+		 */
+		pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_RELEASE_OK);
+
+		/*
+		 * Now wait for the scheduled wakeup, otherwise our ->lwWaiting would
+		 * get reset at some inconvenient point later. Most of the time this
+		 * will immediately return.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		/*
+		 * Fix the process wait semaphore's count for any absorbed wakeups.
+		 */
+		while (extraWaits-- > 0)
+			PGSemaphoreUnlock(MyProc->sem);
+	}
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * After calling this function it's the callers responsibility to ensure that
+ * the lock gets released, even in case of an error. This only is desirable if
+ * the lock is going to be released in a different process than the process
+ * that acquired it.
+ */
+static inline void
+BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockDisownInternal(buffer, buf_hdr);
+	RESUME_INTERRUPTS();
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * This is the code that can be shared between actually releasing a lock
+ * (BufferLockUnlock()) and just not tracking ownership of the lock anymore
+ * without releasing the lock (BufferLockDisown()).
+ */
+static inline int
+BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	PrivateRefCountEntry *ref;
+
+	ref = GetPrivateRefCountEntry(buffer, false);
+	if (ref == NULL)
+		elog(ERROR, "lock %d is not held", buffer);
+	mode = ref->data.lockmode;
+	ref->data.lockmode = BUFFER_LOCK_UNLOCK;
+
+	return mode;
+}
+
+/*
+ * Wakeup all the lockers that currently have a chance to acquire the lock.
+ */
+static void
+BufferLockWakeup(BufferDesc *buf_hdr, bool unlocked)
+{
+	bool		new_release_ok;
+	bool		wake_exclusive = unlocked;
+	bool		wake_share_exclusive = true;
+	proclist_head wakeup;
+	proclist_mutable_iter iter;
+
+	proclist_init(&wakeup);
+
+	new_release_ok = true;
+
+	/* lock wait list while collecting backends to wake up */
+	LockBufHdr(buf_hdr);
+
+	proclist_foreach_modify(iter, &buf_hdr->lock_waiters, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		if (!wake_exclusive && waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+			continue;
+
+		if (!wake_share_exclusive && waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+			continue;
+
+		proclist_delete(&buf_hdr->lock_waiters, iter.cur, lwWaitLink);
+		proclist_push_tail(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Prevent additional wakeups until retryer gets to run. Backends that
+		 * are just waiting for the lock to become free don't retry
+		 * automatically.
+		 */
+		new_release_ok = false;
+
+		/*
+		 * Don't wakeup further share-exclusive/exclusive lock waiters after
+		 * waking a conflicting waiter.
+		 */
+		if (waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+		{
+			wake_exclusive = false;
+			wake_share_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+			wake_share_exclusive = false;
+
+		/*
+		 * Signal that the process isn't on the wait list anymore. This allows
+		 * BufferLockDequeueSelf() to remove itself of the waitlist with a
+		 * proclist_delete(), rather than having to check if it has been
+		 * removed from the list.
+		 */
+		Assert(waiter->lwWaiting == LW_WS_WAITING);
+		waiter->lwWaiting = LW_WS_PENDING_WAKEUP;
+
+		/*
+		 * Once we've woken up an exclusive lock, there's no point in waking
+		 * up anybody else.
+		 */
+		if (waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+			break;
+	}
+
+	Assert(proclist_is_empty(&wakeup) || pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS);
+
+	/* unset required flags, and release lock, in one fell swoop */
+	{
+		uint64		old_state;
+		uint64		desired_state;
+
+		old_state = pg_atomic_read_u64(&buf_hdr->state);
+		while (true)
+		{
+			desired_state = old_state;
+
+			/* compute desired flags */
+
+			if (new_release_ok)
+				desired_state |= BM_LOCK_RELEASE_OK;
+			else
+				desired_state &= ~BM_LOCK_RELEASE_OK;
+
+			if (proclist_is_empty(&buf_hdr->lock_waiters))
+				desired_state &= ~BM_LOCK_HAS_WAITERS;
+
+			desired_state &= ~BM_LOCKED;	/* release lock */
+
+			if (pg_atomic_compare_exchange_u64(&buf_hdr->state, &old_state,
+											   desired_state))
+				break;
+		}
+	}
+
+	/* Awaken any waiters I removed from the queue. */
+	proclist_foreach_modify(iter, &wakeup, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		proclist_delete(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Guarantee that lwWaiting being unset only becomes visible once the
+		 * unlink from the link has completed. Otherwise the target backend
+		 * could be woken up for other reason and enqueue for a new lock - if
+		 * that happens before the list unlink happens, the list would end up
+		 * being corrupted.
+		 *
+		 * The barrier pairs with the LWLockWaitListLock() when enqueuing for
+		 * another lock.
+		 */
+		pg_write_barrier();
+		waiter->lwWaiting = LW_WS_NOT_WAITING;
+		PGSemaphoreUnlock(waiter->sem);
+	}
+}
+
+/*
+ * Compute subtraction from buffer state for a release of a held lock in
+ * `mode`.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places.
+ */
+static inline uint64
+BufferLockReleaseSub(BufferLockMode mode)
+{
+
+	/*
+	 * Turns out that a switch() leads gcc to generate sufficiently worse code
+	 * for this to show up in profiles...
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE)
+		return BM_LOCK_VAL_EXCLUSIVE;
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		return BM_LOCK_VAL_SHARE_EXCLUSIVE;
+	else
+	{
+		Assert(mode == BUFFER_LOCK_SHARE);
+		return BM_LOCK_VAL_SHARED;
+	}
+
+	return 0;
+}
+
+/*
+ * Handle work that needs to be done after releasing a lock that was held in
+ * `mode`, where `lockstate` is the result of the atomic operation modifying
+ * the state variable.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places.
+ */
+static void
+BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate)
+{
+	bool		check_waiters = false;
+	bool		unlocked = false;
+
+	/* nobody else can have that kind of lock */
+	Assert(!(lockstate & BM_LOCK_VAL_EXCLUSIVE));
+
+	/*
+	 * We're still waiting for backends to get scheduled, don't wake them up
+	 * again.
+	 */
+	if ((lockstate & (BM_LOCK_HAS_WAITERS | BM_LOCK_RELEASE_OK)) ==
+		(BM_LOCK_HAS_WAITERS | BM_LOCK_RELEASE_OK))
+	{
+		if ((lockstate & BM_LOCK_MASK) == 0)
+		{
+			check_waiters = true;
+			unlocked = true;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			check_waiters = true;
+			unlocked = false;
+		}
+	}
+
+	/*
+	 * As waking up waiters requires the spinlock to be acquired, only do so
+	 * if necessary.
+	 */
+	if (check_waiters)
+		BufferLockWakeup(buf_hdr, unlocked);
+}
+
+/*
+ * BufferLockHeldByMeInMode - test whether my process holds the content lock
+ * in the specified mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode == mode;
+
+}
+
+/*
+ * BufferLockHeldByMe - test whether my process holds the content lock in any
+ * mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMe(BufferDesc *buf_hdr)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode != BUFFER_LOCK_UNLOCK;
+}
+
+/*
+ * Release the content lock for the buffer.
+ */
+void
+UnlockBuffer(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+
+	Assert(BufferIsPinned(buffer));
+	if (BufferIsLocal(buffer))
+		return;					/* local buffers need no lock */
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+	BufferLockUnlock(buffer, buf_hdr);
+}
+
+/*
+ * Acquire the content_lock for the buffer.
  */
 void
-LockBuffer(Buffer buffer, BufferLockMode mode)
+LockBufferInternal(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *buf;
+	BufferDesc *buf_hdr;
+
+	/*
+	 * We can't wait if we haven't got a PGPROC.  This should only occur
+	 * during bootstrap or shared memory initialization.  Put an Assert here
+	 * to catch unsafe coding practices.
+	 */
+	Assert(!(MyProc == NULL && IsUnderPostmaster));
+
+	/* handled in LockBuffer() wrapper */
+	Assert(mode != BUFFER_LOCK_UNLOCK);
 
 	Assert(BufferIsPinned(buffer));
 	if (BufferIsLocal(buffer))
 		return;					/* local buffers need no lock */
 
-	buf = GetBufferDescriptor(buffer - 1);
+	buf_hdr = GetBufferDescriptor(buffer - 1);
 
-	if (mode == BUFFER_LOCK_UNLOCK)
-		LWLockRelease(BufferDescriptorGetContentLock(buf));
-	else if (mode == BUFFER_LOCK_SHARE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	if (mode == BUFFER_LOCK_SHARE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE);
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	else if (mode == BUFFER_LOCK_EXCLUSIVE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_EXCLUSIVE);
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_EXCLUSIVE);
 	else
 		elog(ERROR, "unrecognized buffer lock mode: %d", mode);
 }
@@ -5687,8 +6358,7 @@ ConditionalLockBuffer(Buffer buffer)
 
 	buf = GetBufferDescriptor(buffer - 1);
 
-	return LWLockConditionalAcquire(BufferDescriptorGetContentLock(buf),
-									LW_EXCLUSIVE);
+	return BufferLockConditional(buffer, buf, BUFFER_LOCK_EXCLUSIVE);
 }
 
 /*
@@ -6625,7 +7295,25 @@ ResOwnerReleaseBufferPin(Datum res)
 	if (BufferIsLocal(buffer))
 		UnpinLocalBufferNoOwner(buffer);
 	else
+	{
+		PrivateRefCountEntry *ref;
+
+		ref = GetPrivateRefCountEntry(buffer, false);
+
+		/*
+		 * If the buffer was locked at the time of the resowner release,
+		 * release the lock now. This should only happen after errors.
+		 */
+		if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
+		{
+			BufferDesc *buf = GetBufferDescriptor(buffer - 1);
+
+			HOLD_INTERRUPTS();	/* match the upcoming RESUME_INTERRUPTS */
+			BufferLockUnlock(buffer, buf);
+		}
+
 		UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
+	}
 }
 
 static char *
@@ -6927,16 +7615,12 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		 */
 		if (is_write && !is_temp)
 		{
-			LWLock	   *content_lock;
-
-			content_lock = BufferDescriptorGetContentLock(buf_hdr);
-
-			Assert(LWLockHeldByMe(content_lock));
+			Assert(BufferLockHeldByMe(buf_hdr));
 
 			/*
 			 * Lock is now owned by AIO subsystem.
 			 */
-			LWLockDisown(content_lock);
+			BufferLockDisown(buffer, buf_hdr);
 		}
 
 		/*
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 1e5e368a5dc..39ae7cdf856 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -285,6 +285,9 @@ ABI_compatibility:
 Section: ClassName - WaitEventBuffer
 
 BUFFER_CLEANUP	"Waiting to acquire an exclusive pin on a buffer. Buffer pin waits can be protracted if another process holds an open cursor that last read data from the buffer in question."
+BUFFER_SHARED	"Waiting to acquire shared lock on a buffer."
+BUFFER_SHARE_EXCLUSIVE	"Waiting to acquire share exclusive lock on a buffer."
+BUFFER_EXCLUSIVE	"Waiting to acquire exclusive lock on a buffer."
 
 ABI_compatibility:
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0009-heapam-Move-logic-to-handle-HEAP_MOVED-into-a-hel.patch (11.5K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/10-v6-0009-heapam-Move-logic-to-handle-HEAP_MOVED-into-a-hel.patch)
  download | inline diff:
From 5244aaf57a4dba00139058173e1be476ace0ee8e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 23 Sep 2024 12:23:33 -0400
Subject: [PATCH v6 09/14] heapam: Move logic to handle HEAP_MOVED into a
 helper function

Before we dealt with this in 6 near identical and one very similar copy.

The helper function errors out when encountering a
HEAP_MOVED_IN/HEAP_MOVED_OUT tuple with xvac considered current or
in-progress. It'd be preferrable to do that change separately, but otherwise
it'd not be possible to deduplicate the handling in
HeapTupleSatisfiesVacuum().

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/heap/heapam_visibility.c | 307 ++++----------------
 1 file changed, 61 insertions(+), 246 deletions(-)

diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 05f6946fe60..4fefcbca5f5 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -144,6 +144,55 @@ HeapTupleSetHintBits(HeapTupleHeader tuple, Buffer buffer,
 	SetHintBits(tuple, buffer, infomask, xid);
 }
 
+/*
+ * If HEAP_MOVED_OFF or HEAP_MOVED_IN are set on the tuple, remove them and
+ * adjust hint bits. See the comment for SetHintBits() for more background.
+ *
+ * This helper returns false if the row ought to be invisible, true otherwise.
+ */
+static inline bool
+HeapTupleCleanMoved(HeapTupleHeader tuple, Buffer buffer)
+{
+	TransactionId xvac;
+
+	/* only used by pre-9.0 binary upgrades */
+	if (likely(!(tuple->t_infomask & (HEAP_MOVED_OFF | HEAP_MOVED_IN))))
+		return true;
+
+	xvac = HeapTupleHeaderGetXvac(tuple);
+
+	if (TransactionIdIsCurrentTransactionId(xvac))
+		elog(ERROR, "encountered tuple with HEAP_MOVED considered current");
+
+	if (TransactionIdIsInProgress(xvac))
+		elog(ERROR, "encountered tuple with HEAP_MOVED considered in-progress");
+
+	if (tuple->t_infomask & HEAP_MOVED_OFF)
+	{
+		if (TransactionIdDidCommit(xvac))
+		{
+			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
+						InvalidTransactionId);
+			return false;
+		}
+		SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
+					InvalidTransactionId);
+	}
+	else if (tuple->t_infomask & HEAP_MOVED_IN)
+	{
+		if (TransactionIdDidCommit(xvac))
+			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
+						InvalidTransactionId);
+		else
+		{
+			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
+						InvalidTransactionId);
+			return false;
+		}
+	}
+
+	return true;
+}
 
 /*
  * HeapTupleSatisfiesSelf
@@ -179,45 +228,8 @@ HeapTupleSatisfiesSelf(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (tuple->t_infomask & HEAP_XMAX_INVALID)	/* xid invalid */
@@ -372,45 +384,8 @@ HeapTupleSatisfiesToast(HeapTuple htup, Snapshot snapshot,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 
 		/*
 		 * An invalid Xmin can be left behind by a speculative insertion that
@@ -468,45 +443,8 @@ HeapTupleSatisfiesUpdate(HeapTuple htup, CommandId curcid,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return TM_Invisible;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return TM_Invisible;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return TM_Invisible;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return TM_Invisible;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return TM_Invisible;
-				}
-			}
-		}
+		else if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (HeapTupleHeaderGetCmin(tuple) >= curcid)
@@ -756,45 +694,8 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (tuple->t_infomask & HEAP_XMAX_INVALID)	/* xid invalid */
@@ -979,45 +880,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!XidInMVCCSnapshot(xvac, snapshot))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (XidInMVCCSnapshot(xvac, snapshot))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (HeapTupleHeaderGetCmin(tuple) >= snapshot->curcid)
@@ -1222,57 +1086,8 @@ HeapTupleSatisfiesVacuumHorizon(HeapTuple htup, Buffer buffer, TransactionId *de
 	{
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return HEAPTUPLE_DEAD;
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return HEAPTUPLE_DELETE_IN_PROGRESS;
-			if (TransactionIdIsInProgress(xvac))
-				return HEAPTUPLE_DELETE_IN_PROGRESS;
-			if (TransactionIdDidCommit(xvac))
-			{
-				SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-							InvalidTransactionId);
-				return HEAPTUPLE_DEAD;
-			}
-			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-						InvalidTransactionId);
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return HEAPTUPLE_INSERT_IN_PROGRESS;
-			if (TransactionIdIsInProgress(xvac))
-				return HEAPTUPLE_INSERT_IN_PROGRESS;
-			if (TransactionIdDidCommit(xvac))
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			else
-			{
-				SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-							InvalidTransactionId);
-				return HEAPTUPLE_DEAD;
-			}
-		}
-		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
-		{
-			if (tuple->t_infomask & HEAP_XMAX_INVALID)	/* xid invalid */
-				return HEAPTUPLE_INSERT_IN_PROGRESS;
-			/* only locked? run infomask-only check first, for performance */
-			if (HEAP_XMAX_IS_LOCKED_ONLY(tuple->t_infomask) ||
-				HeapTupleHeaderIsOnlyLocked(tuple))
-				return HEAPTUPLE_INSERT_IN_PROGRESS;
-			/* inserted and then deleted by same xact */
-			if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetUpdateXid(tuple)))
-				return HEAPTUPLE_DELETE_IN_PROGRESS;
-			/* deleting subtransaction must have aborted */
-			return HEAPTUPLE_INSERT_IN_PROGRESS;
-		}
+		else if (!HeapTupleCleanMoved(tuple, buffer))
+			return HEAPTUPLE_DEAD;
 		else if (TransactionIdIsInProgress(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			/*
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0010-heapam-Use-exclusive-lock-on-old-page-in-CLUSTER.patch (2.7K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/11-v6-0010-heapam-Use-exclusive-lock-on-old-page-in-CLUSTER.patch)
  download | inline diff:
From 1c1f8f7b7d5a4baffe6cb282cdd4089671361b22 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sun, 26 Jan 2025 15:18:46 -0500
Subject: [PATCH v6 10/14] heapam: Use exclusive lock on old page in CLUSTER

To be able to guarantee that we can set the hint bit, acquire an exclusive
lock on the old buffer. We need the hint bits to be set as otherwise
reform_and_rewrite_tuple() -> rewrite_heap_tuple() -> heap_freeze_tuple() will
get confused.

It'd be better if we somehow could avoid setting hint bits on the old page. A
commonreason to use VACUUM FULL are very bloated tables - rewriting most of
the old table before during VACUUM FULL doesn't exactly help.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/heap/heapam_handler.c    | 13 ++++++++++++-
 src/backend/access/heap/heapam_visibility.c |  7 +++++++
 2 files changed, 19 insertions(+), 1 deletion(-)

diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index bcbac844bb6..f84254f0737 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -837,7 +837,18 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		tuple = ExecFetchSlotHeapTuple(slot, false, NULL);
 		buf = hslot->buffer;
 
-		LockBuffer(buf, BUFFER_LOCK_SHARE);
+		/*
+		 * To be able to guarantee that we can set the hint bit, acquire an
+		 * exclusive lock on the old buffer. We need the hint bits to be set
+		 * as otherwise reform_and_rewrite_tuple() -> rewrite_heap_tuple() ->
+		 * heap_freeze_tuple() will get confused.
+		 *
+		 * It'd be better if we somehow could avoid setting hint bits on the
+		 * old page. One reason to use VACUUM FULL are very bloated tables -
+		 * rewriting most of the old table before during VACUUM FULL doesn't
+		 * exactly help...
+		 */
+		LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
 		switch (HeapTupleSatisfiesVacuum(tuple, OldestXmin, buf))
 		{
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 4fefcbca5f5..762538a2040 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -141,6 +141,13 @@ void
 HeapTupleSetHintBits(HeapTupleHeader tuple, Buffer buffer,
 					 uint16 infomask, TransactionId xid)
 {
+	/*
+	 * The uses from heapam.c rely on being able to perform the hint bit
+	 * updates, which can only be guaranteed if we are holding an exclusive
+	 * lock on the buffer - which all callers are doing.
+	 */
+	Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+
 	SetHintBits(tuple, buffer, infomask, xid);
 }
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0011-heapam-Add-batch-mode-mvcc-check-and-use-it-in-pa.patch (7.6K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/12-v6-0011-heapam-Add-batch-mode-mvcc-check-and-use-it-in-pa.patch)
  download | inline diff:
From c93a15d76539720a8564de1b0a1100c4734389a7 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 17 Oct 2024 13:16:36 -0400
Subject: [PATCH v6 11/14] heapam: Add batch mode mvcc check and use it in page
 mode

There are two reasons for doing so:

1) It is generally faster to perform checks in a batched fashion and making
   sequential scans faster is nice.

2) We would like to stop setting hint bits while pages are being written
   out. The necessary locking becomes visible for page mode scans if done for
   every tuple. With batching the overhead can be amortized to only happen
   once per page.

There are substantial further optimization opportunities along these
lines:

- Right now HeapTupleSatisfiesMVCCBatch() simply uses the single-tuple
  HeapTupleSatisfiesMVCC(), relying on the compiler to inline it. We could
  instead write an explicitly optimized version that avoids repeated xid
  tests.

- Introduce batched version of the serializability test

- Introduce batched version of HeapTupleSatisfiesVacuum

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/access/heapam.h                 | 28 +++++++
 src/backend/access/heap/heapam.c            | 91 ++++++++++++++++-----
 src/backend/access/heap/heapam_visibility.c | 47 +++++++++++
 src/tools/pgindent/typedefs.list            |  1 +
 4 files changed, 147 insertions(+), 20 deletions(-)

diff --git a/src/include/access/heapam.h b/src/include/access/heapam.h
index 909db73b7bb..13e4a4096c3 100644
--- a/src/include/access/heapam.h
+++ b/src/include/access/heapam.h
@@ -410,6 +410,34 @@ extern bool HeapTupleHeaderIsOnlyLocked(HeapTupleHeader tuple);
 extern bool HeapTupleIsSurelyDead(HeapTuple htup,
 								  GlobalVisState *vistest);
 
+/*
+ * FIXME: define to be removed
+ *
+ * Without this I see worse performance. But it's a bit ugly, so I thought
+ * it'd be useful to leave a way in for others to experiment with this.
+ */
+#define BATCHMVCC_FEWER_ARGS
+
+#ifdef BATCHMVCC_FEWER_ARGS
+typedef struct BatchMVCCState
+{
+	HeapTupleData tuples[MaxHeapTuplesPerPage];
+	bool		visible[MaxHeapTuplesPerPage];
+} BatchMVCCState;
+#endif
+
+extern int	HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
+										int ntups,
+#ifdef BATCHMVCC_FEWER_ARGS
+										BatchMVCCState *batchmvcc,
+#else
+										HeapTupleData *tuples,
+										bool *visible,
+#endif
+										OffsetNumber *vistuples_dense);
+
+
+
 /*
  * To avoid leaking too much knowledge about reorderbuffer implementation
  * details this is implemented in reorderbuffer.c not heapam_visibility.c
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 4b0c49f4bb0..ddabd1a3ec3 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -504,42 +504,93 @@ page_collect_tuples(HeapScanDesc scan, Snapshot snapshot,
 					BlockNumber block, int lines,
 					bool all_visible, bool check_serializable)
 {
+	Oid			relid = RelationGetRelid(scan->rs_base.rs_rd);
+#ifdef BATCHMVCC_FEWER_ARGS
+	BatchMVCCState batchmvcc;
+	HeapTupleData *tuples = batchmvcc.tuples;
+	bool	   *visible = batchmvcc.visible;
+#else
+	HeapTupleData tuples[MaxHeapTuplesPerPage];
+	bool		visible[MaxHeapTuplesPerPage];
+#endif
 	int			ntup = 0;
-	OffsetNumber lineoff;
+	int			nvis = 0;
 
-	for (lineoff = FirstOffsetNumber; lineoff <= lines; lineoff++)
+	/* page at a time should have been disabled otherwise */
+	Assert(IsMVCCSnapshot(snapshot));
+
+	/* first find all tuples on the page */
+	for (OffsetNumber lineoff = FirstOffsetNumber; lineoff <= lines; lineoff++)
 	{
 		ItemId		lpp = PageGetItemId(page, lineoff);
-		HeapTupleData loctup;
-		bool		valid;
+		HeapTuple	tup;
 
-		if (!ItemIdIsNormal(lpp))
+		if (unlikely(!ItemIdIsNormal(lpp)))
 			continue;
 
-		loctup.t_data = (HeapTupleHeader) PageGetItem(page, lpp);
-		loctup.t_len = ItemIdGetLength(lpp);
-		loctup.t_tableOid = RelationGetRelid(scan->rs_base.rs_rd);
-		ItemPointerSet(&(loctup.t_self), block, lineoff);
+		/*
+		 * If the page is not all-visible or we need to check serializability,
+		 * maintain enough state to be able to refind the tuple efficiently,
+		 * without again needing to extract it from the page.
+		 */
+		if (!all_visible || check_serializable)
+		{
+			tup = &tuples[ntup];
 
+			tup->t_data = (HeapTupleHeader) PageGetItem(page, lpp);
+			tup->t_len = ItemIdGetLength(lpp);
+			tup->t_tableOid = relid;
+			ItemPointerSet(&(tup->t_self), block, lineoff);
+		}
+
+		/*
+		 * If the page is all visible, these fields won'otherwise wont be
+		 * populated in loop below.
+		 */
 		if (all_visible)
-			valid = true;
-		else
-			valid = HeapTupleSatisfiesVisibility(&loctup, snapshot, buffer);
-
-		if (check_serializable)
-			HeapCheckForSerializableConflictOut(valid, scan->rs_base.rs_rd,
-												&loctup, buffer, snapshot);
-
-		if (valid)
 		{
+			if (check_serializable)
+			{
+				visible[ntup] = true;
+			}
 			scan->rs_vistuples[ntup] = lineoff;
-			ntup++;
 		}
+
+		ntup++;
 	}
 
 	Assert(ntup <= MaxHeapTuplesPerPage);
 
-	return ntup;
+	/* unless the page is all visible, test visibility for all tuples one go */
+	if (all_visible)
+		nvis = ntup;
+	else
+		nvis = HeapTupleSatisfiesMVCCBatch(snapshot, buffer,
+										   ntup,
+#ifdef BATCHMVCC_FEWER_ARGS
+										   &batchmvcc,
+#else
+										   tuples, visible,
+#endif
+										   scan->rs_vistuples
+			);
+
+	/*
+	 * So far we don't have batch API for testing serializabilty, so do so
+	 * one-by-one.
+	 */
+	if (check_serializable)
+	{
+		for (int i = 0; i < ntup; i++)
+		{
+			HeapCheckForSerializableConflictOut(visible[i],
+												scan->rs_base.rs_rd,
+												&tuples[i],
+												buffer, snapshot);
+		}
+	}
+
+	return nvis;
 }
 
 /*
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 762538a2040..5645cfd8a49 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -1584,6 +1584,53 @@ HeapTupleSatisfiesHistoricMVCC(HeapTuple htup, Snapshot snapshot,
 		return true;
 }
 
+/*
+ * Perform HeaptupleSatisfiesMVCC() on each passed in tuple. This is more
+ * efficient than doing HeapTupleSatisfiesMVCC() one-by-one.
+ *
+ * To be checked tuples are passed via BatchMVCCState->tuples. Each tuple's
+ * visibility is set in batchmvcc->visible[]. In addition, ->vistuples_dense
+ * is set to contain the offsets of visible tuples.
+ *
+ * Returns the number of visible tuples.
+ */
+int
+HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
+							int ntups,
+#ifdef BATCHMVCC_FEWER_ARGS
+							BatchMVCCState *batchmvcc,
+#else
+							HeapTupleData *tuples,
+							bool *visible,
+#endif
+							OffsetNumber *vistuples_dense)
+{
+	int			nvis = 0;
+#ifdef BATCHMVCC_FEWER_ARGS
+	HeapTupleData *tuples = batchmvcc->tuples;
+	bool	   *visible = batchmvcc->visible;
+#endif
+
+	Assert(IsMVCCSnapshot(snapshot));
+
+	for (int i = 0; i < ntups; i++)
+	{
+		bool		valid;
+		HeapTuple	tup = &tuples[i];
+
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		visible[i] = valid;
+
+		if (likely(valid))
+		{
+			vistuples_dense[nvis] = tup->t_self.ip_posid;
+			nvis++;
+		}
+	}
+
+	return nvis;
+}
+
 /*
  * HeapTupleSatisfiesVisibility
  *		True iff heap tuple satisfies a time qual.
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 9a89e68c59c..9d14239b4c4 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -249,6 +249,7 @@ Barrier
 BaseBackupCmd
 BaseBackupTargetHandle
 BaseBackupTargetType
+BatchMVCCState
 BeginDirectModify_function
 BeginForeignInsert_function
 BeginForeignModify_function
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0012-Require-share-exclusive-lock-to-set-hint-bits.patch (32.7K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/13-v6-0012-Require-share-exclusive-lock-to-set-hint-bits.patch)
  download | inline diff:
From 6e61b1b2d2202c23674f27ba80cd50ff596fb95b Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 18 Nov 2025 09:22:28 -0500
Subject: [PATCH v6 12/14] Require share-exclusive lock to set hint bits

At the moment hint bits can be set with just a share lock on a page (and in
one place even without any lock). Because of this we need to copy pages while
writing them out, as otherwise the checksum could be corrupted.

The need to copy the page is problematic to implement AIO writes:

1) Instead of just needing a single buffer for a copied page we need one for
   each page that's potentially undergoing IO
2) To be able to use the "worker" AIO implementation the copied page needs to
   reside in shared memory.

It also causes problems for using unbuffered/direct-IO, independent of AIO:
Some filesystems, raid implementations, ... do not tolerate the data being
written out to change during the write. E.g. they may compute internal
checksums that can be invalidated by concurrent modifications, leading e.g. to
filesystem errors (as the case with btrfs).

It also just is plain odd to allow modifications of buffers that are just
share locked.

To address these issue, this commit changes the rules so that modifications to
pages are not allowed anymore while holding a share lock. Instead the new
share-exclusive lock (introduced in FIXME XXXX TODO) allows at most one
backend to modify a buffer while other backends have the same page share
locked. An existing share-lock can be upgraded to a share-exclusive lock, if
there are no conflicting locks. For that
BufferBeginSetHintBits()/BufferBeginSetHintBits() and BufferSetHintBits16()
have been introduced.

The biggest change to adapt to this is in heapam. To avoid performance
regressions for sequential scans that need to set a lot of hint bits, we need
to amortize the cost of BufferBeginSetHintBits() for cases where hint bits are
set at a high frequency, HeapTupleSatisfiesMVCCBatch() uses the new
SetHintBitsExt() which defers BufferFinishSetHintBits() until all hint bits on
a page have been set.  Conversely, to avoid regressions in cases where we
can't set hint bits in bulk (because we're looking only at individual tuples),
use BufferSetHintBits16() when setting hint bits without batching.

Several other places also need to be adapted, but those changes are
comparatively simpler.

After this we do not need to copy buffers to write them out anymore. That
change is done separately however.

TODO:
- Address FIXMEs
- reflow parts of storage/buffer/README that I didn't reindent to make the
  diff more readable

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf%40gcnactj4z56m
---
 src/include/storage/bufmgr.h                |   4 +
 src/backend/access/gist/gistget.c           |  19 +-
 src/backend/access/hash/hashutil.c          |  10 +-
 src/backend/access/heap/heapam_visibility.c | 124 ++++++++--
 src/backend/access/nbtree/nbtinsert.c       |  28 ++-
 src/backend/access/nbtree/nbtutils.c        |  16 +-
 src/backend/storage/buffer/README           |  32 ++-
 src/backend/storage/buffer/bufmgr.c         | 241 +++++++++++++++-----
 src/backend/storage/freespace/freespace.c   |  20 +-
 src/backend/storage/freespace/fsmpage.c     |  11 +-
 src/tools/pgindent/typedefs.list            |   1 +
 11 files changed, 392 insertions(+), 114 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 8e442492d4d..afa16afffc9 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -302,6 +302,10 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
+extern bool BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer);
+extern bool BufferBeginSetHintBits(Buffer buffer);
+extern void BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std);
+
 extern void UnlockBuffers(void);
 extern void UnlockBuffer(Buffer buffer);
 extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
diff --git a/src/backend/access/gist/gistget.c b/src/backend/access/gist/gistget.c
index 9ba45acfff3..956ece6bed5 100644
--- a/src/backend/access/gist/gistget.c
+++ b/src/backend/access/gist/gistget.c
@@ -63,11 +63,7 @@ gistkillitems(IndexScanDesc scan)
 	 * safe.
 	 */
 	if (BufferGetLSNAtomic(buffer) != so->curPageLSN)
-	{
-		UnlockReleaseBuffer(buffer);
-		so->numKilled = 0;		/* reset counter */
-		return;
-	}
+		goto unlock;
 
 	Assert(GistPageIsLeaf(page));
 
@@ -77,6 +73,16 @@ gistkillitems(IndexScanDesc scan)
 	 */
 	for (i = 0; i < so->numKilled; i++)
 	{
+		if (!killedsomething)
+		{
+			/*
+			 * Use hint bit infrastructure to be allowed to modify the page
+			 * without holding an exclusive lock.
+			 */
+			if (!BufferBeginSetHintBits(buffer))
+				goto unlock;
+		}
+
 		offnum = so->killedItems[i];
 		iid = PageGetItemId(page, offnum);
 		ItemIdMarkDead(iid);
@@ -86,9 +92,10 @@ gistkillitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		GistMarkPageHasGarbage(page);
-		MarkBufferDirtyHint(buffer, true);
+		BufferFinishSetHintBits(buffer, true, true);
 	}
 
+unlock:
 	UnlockReleaseBuffer(buffer);
 
 	/*
diff --git a/src/backend/access/hash/hashutil.c b/src/backend/access/hash/hashutil.c
index f41233fcd07..d1d603770b2 100644
--- a/src/backend/access/hash/hashutil.c
+++ b/src/backend/access/hash/hashutil.c
@@ -593,6 +593,13 @@ _hash_kill_items(IndexScanDesc scan)
 
 			if (ItemPointerEquals(&ituple->t_tid, &currItem->heapTid))
 			{
+				/*
+				 * Use hint bit infrastructure to be allowed to modify the
+				 * page without holding an exclusive lock.
+				 */
+				if (!BufferBeginSetHintBits(so->currPos.buf))
+					goto unlock_page;
+
 				/* found the item */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -610,9 +617,10 @@ _hash_kill_items(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->hasho_flag |= LH_PAGE_HAS_DEAD_TUPLES;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(so->currPos.buf, true, true);
 	}
 
+unlock_page:
 	if (so->hashso_bucket_buf == so->currPos.buf ||
 		havePin)
 		LockBuffer(so->currPos.buf, BUFFER_LOCK_UNLOCK);
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 5645cfd8a49..630ba7df167 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -80,10 +80,38 @@
 
 
 /*
- * SetHintBits()
+ * To be allowed to set hint bits, SetHintBits() needs to call
+ * BufferBeginSetHintBits(). However, that's not free, and some callsites call
+ * SetHintBits() on many tuples in a row. For those it makes sense to amortize
+ * the cost of BufferBeginSetHintBits(). Additionally it's desirable to defer
+ * the cost of BufferBeginSetHintBits() until a hint bit needs to actually be
+ * set. This enum serves as the necessary state space passed to
+ * SetHintbitsExt().
+ */
+typedef enum SetHintBitsState
+{
+	/* not yet checked if hint bits may be set */
+	SHB_INITIAL,
+	/* failed to get permission to set hint bits, don't check again */
+	SHB_DISABLED,
+	/* allowed to set hint bits */
+	SHB_ENABLED,
+} SetHintBitsState;
+
+/*
+ * SetHintBitsExt()
  *
  * Set commit/abort hint bits on a tuple, if appropriate at this time.
  *
+ * To be allowed to set a hint bit on a tuple, the page must not be undergoing
+ * IO at this time (otherwise we e.g. could corrupt PG's page checksum or even
+ * the filesystem's, as is known to happen with btrfs).
+ *
+ * The right to set a hint bit can be acquired on a page level with
+ * BufferBeginSetHintBits(). Only a single backend gets the right to set hint
+ * bits at a time.  Alternatively, if called with a NULL SetHintBitsState*,
+ * hint bits are set with BufferSetHintBits16().
+ *
  * It is only safe to set a transaction-committed hint bit if we know the
  * transaction's commit record is guaranteed to be flushed to disk before the
  * buffer, or if the table is temporary or unlogged and will be obliterated by
@@ -111,24 +139,68 @@
  * InvalidTransactionId if no check is needed.
  */
 static inline void
-SetHintBits(HeapTupleHeader tuple, Buffer buffer,
-			uint16 infomask, TransactionId xid)
+SetHintBitsExt(HeapTupleHeader tuple, Buffer buffer,
+			   uint16 infomask, TransactionId xid, SetHintBitsState *state)
 {
 	if (TransactionIdIsValid(xid))
 	{
-		/* NB: xid must be known committed here! */
-		XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+		if (BufferIsPermanent(buffer))
+		{
+			/* NB: xid must be known committed here! */
+			XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+
+			if (XLogNeedsFlush(commitLSN) &&
+				BufferGetLSNAtomic(buffer) < commitLSN)
+			{
+				/* not flushed and no LSN interlock, so don't set hint */
+				return;
+			}
+		}
+	}
+
+	/*
+	 * If we're not operating in batch mode, use BufferSetHintBits16 to mark
+	 * the page dirty, that's cheaper than
+	 * BufferBeginSetHintBits()/BufferFinishSetHintBits(). That's important
+	 * for cases where we set a lot of hint bits on a page individually.
+	 */
+	if (!state)
+	{
+		BufferSetHintBits16(&tuple->t_infomask, tuple->t_infomask | infomask, buffer);
+		return;
+	}
+
+	/*
+	 * In batched mode and we previously did not get permission to set hint
+	 * bits. Don't try again, in all likelihood IO is still going on.
+	 */
+	if (*state == SHB_DISABLED)
+		return;
 
-		if (BufferIsPermanent(buffer) && XLogNeedsFlush(commitLSN) &&
-			BufferGetLSNAtomic(buffer) < commitLSN)
+	if (*state == SHB_INITIAL)
+	{
+		if (!BufferBeginSetHintBits(buffer))
 		{
-			/* not flushed and no LSN interlock, so don't set hint */
+			*state = SHB_DISABLED;
 			return;
 		}
+
+		if (state)
+			*state = SHB_ENABLED;
+
 	}
-
 	tuple->t_infomask |= infomask;
-	MarkBufferDirtyHint(buffer, true);
+}
+
+/*
+ * Simple wrapper around SetHintBitExt(), use when operating on a single
+ * tuple.
+ */
+static inline void
+SetHintBits(HeapTupleHeader tuple, Buffer buffer,
+			uint16 infomask, TransactionId xid)
+{
+	SetHintBitsExt(tuple, buffer, infomask, xid, NULL);
 }
 
 /*
@@ -864,9 +936,9 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
  * inserting/deleting transaction was still running --- which was more cycles
  * and more contention on ProcArrayLock.
  */
-static bool
+static inline bool
 HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
-					   Buffer buffer)
+					   Buffer buffer, SetHintBitsState *state)
 {
 	HeapTupleHeader tuple = htup->t_data;
 
@@ -921,8 +993,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 			if (!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmax(tuple)))
 			{
 				/* deleting subtransaction must have aborted */
-				SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-							InvalidTransactionId);
+				SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+							   InvalidTransactionId, state);
 				return true;
 			}
 
@@ -934,13 +1006,13 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		else if (XidInMVCCSnapshot(HeapTupleHeaderGetRawXmin(tuple), snapshot))
 			return false;
 		else if (TransactionIdDidCommit(HeapTupleHeaderGetRawXmin(tuple)))
-			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-						HeapTupleHeaderGetRawXmin(tuple));
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_COMMITTED,
+						   HeapTupleHeaderGetRawXmin(tuple), state);
 		else
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_INVALID,
+						   InvalidTransactionId, state);
 			return false;
 		}
 	}
@@ -1003,14 +1075,14 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		if (!TransactionIdDidCommit(HeapTupleHeaderGetRawXmax(tuple)))
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+						   InvalidTransactionId, state);
 			return true;
 		}
 
 		/* xmax transaction committed */
-		SetHintBits(tuple, buffer, HEAP_XMAX_COMMITTED,
-					HeapTupleHeaderGetRawXmax(tuple));
+		SetHintBitsExt(tuple, buffer, HEAP_XMAX_COMMITTED,
+					   HeapTupleHeaderGetRawXmax(tuple), state);
 	}
 	else
 	{
@@ -1606,6 +1678,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 							OffsetNumber *vistuples_dense)
 {
 	int			nvis = 0;
+	SetHintBitsState state = SHB_INITIAL;
 #ifdef BATCHMVCC_FEWER_ARGS
 	HeapTupleData *tuples = batchmvcc->tuples;
 	bool	   *visible = batchmvcc->visible;
@@ -1618,7 +1691,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		bool		valid;
 		HeapTuple	tup = &tuples[i];
 
-		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer, &state);
 		visible[i] = valid;
 
 		if (likely(valid))
@@ -1628,6 +1701,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		}
 	}
 
+	if (state == SHB_ENABLED)
+		BufferFinishSetHintBits(buffer, true, true);
+
 	return nvis;
 }
 
@@ -1647,7 +1723,7 @@ HeapTupleSatisfiesVisibility(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 	switch (snapshot->snapshot_type)
 	{
 		case SNAPSHOT_MVCC:
-			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer);
+			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer, NULL);
 		case SNAPSHOT_SELF:
 			return HeapTupleSatisfiesSelf(htup, snapshot, buffer);
 		case SNAPSHOT_ANY:
diff --git a/src/backend/access/nbtree/nbtinsert.c b/src/backend/access/nbtree/nbtinsert.c
index 7c113c007e5..545e1d7d9e0 100644
--- a/src/backend/access/nbtree/nbtinsert.c
+++ b/src/backend/access/nbtree/nbtinsert.c
@@ -680,20 +680,28 @@ _bt_check_unique(Relation rel, BTInsertState insertstate, Relation heapRel,
 				{
 					/*
 					 * The conflicting tuple (or all HOT chains pointed to by
-					 * all posting list TIDs) is dead to everyone, so mark the
-					 * index entry killed.
+					 * all posting list TIDs) is dead to everyone, so try to
+					 * mark the index entry killed. It's ok if we're not
+					 * allowed to, this isn't required for correctness.
 					 */
-					ItemIdMarkDead(curitemid);
-					opaque->btpo_flags |= BTP_HAS_GARBAGE;
+					Buffer		buf;
 
-					/*
-					 * Mark buffer with a dirty hint, since state is not
-					 * crucial. Be sure to mark the proper buffer dirty.
-					 */
+					/* Be sure to operate on the proper buffer */
 					if (nbuf != InvalidBuffer)
-						MarkBufferDirtyHint(nbuf, true);
+						buf = nbuf;
 					else
-						MarkBufferDirtyHint(insertstate->buf, true);
+						buf = insertstate->buf;
+
+					/*
+					 * Can't use BufferSetHintBits16() here as we update two
+					 * different locations.
+					 */
+					if (BufferBeginSetHintBits(buf))
+					{
+						ItemIdMarkDead(curitemid);
+						opaque->btpo_flags |= BTP_HAS_GARBAGE;
+						BufferFinishSetHintBits(buf, true, true);
+					}
 				}
 
 				/*
diff --git a/src/backend/access/nbtree/nbtutils.c b/src/backend/access/nbtree/nbtutils.c
index ab0f98b0287..34e548b9930 100644
--- a/src/backend/access/nbtree/nbtutils.c
+++ b/src/backend/access/nbtree/nbtutils.c
@@ -3542,10 +3542,19 @@ _bt_killitems(IndexScanDesc scan)
 			 * it's possible that multiple processes attempt to do this
 			 * simultaneously, leading to multiple full-page images being sent
 			 * to WAL (if wal_log_hints or data checksums are enabled), which
-			 * is undesirable.
+			 * is undesirable.  We need to use the hint bit infrastructure to
+			 * update the page while just holding a share lock.
 			 */
 			if (killtuple && !ItemIdIsDead(iid))
 			{
+				/*
+				 * If we're not able to set hint bits, there's no point
+				 * continuing.
+				 */
+				if (!killedsomething &&
+					!BufferBeginSetHintBits(buf))
+					goto unlock_page;
+
 				/* found the item/all posting list items */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -3556,8 +3565,6 @@ _bt_killitems(IndexScanDesc scan)
 	}
 
 	/*
-	 * Since this can be redone later if needed, mark as dirty hint.
-	 *
 	 * Whenever we mark anything LP_DEAD, we also set the page's
 	 * BTP_HAS_GARBAGE flag, which is likewise just a hint.  (Note that we
 	 * only rely on the page-level flag in !heapkeyspace indexes.)
@@ -3565,9 +3572,10 @@ _bt_killitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->btpo_flags |= BTP_HAS_GARBAGE;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(buf, true, true);
 	}
 
+unlock_page:
 	if (!so->dropPin)
 		_bt_unlockbuf(rel, buf);
 	else
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index 119f31b5d65..9a4dc101c26 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -25,14 +25,20 @@ that might need to do such a wait is instead handled by waiting to obtain
 the relation-level lock, which is why you'd better hold one first.)  Pins
 may not be held across transaction boundaries, however.
 
-Buffer content locks: there are two kinds of buffer lock, shared and exclusive,
-which act just as you'd expect: multiple backends can hold shared locks on
-the same buffer, but an exclusive lock prevents anyone else from holding
-either shared or exclusive lock.  (These can alternatively be called READ
-and WRITE locks.)  These locks are intended to be short-term: they should not
-be held for long.  Buffer locks are acquired and released by LockBuffer().
-It will *not* work for a single backend to try to acquire multiple locks on
-the same buffer.  One must pin a buffer before trying to lock it.
+Buffer content locks: there three kinds of buffer lock, shared,
+share-exclusive and exclusive:
+a) multiple backends can hold shared locks on the same buffer
+   (alternatively called a READ lock)
+b) one backend can hold an share-exclusive lock on a buffer while multiple
+   backends can hold a share lock
+c) an exclusive lock prevents anyone else from holding either shared or
+   exclusive lock.
+   (alternatively called a WRITE lock)
+
+These locks are intended to be short-term: they should not be held for long.
+Buffer locks are acquired and released by LockBuffer().  It will *not* work
+for a single backend to try to acquire multiple locks on the same buffer.  One
+must pin a buffer before trying to lock it.
 
 Buffer access rules:
 
@@ -55,8 +61,14 @@ one must hold a pin and an exclusive content lock on the containing buffer.
 This ensures that no one else might see a partially-updated state of the
 tuple while they are doing visibility checks.
 
-4. It is considered OK to update tuple commit status bits (ie, OR the
-values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
+4. Non-critical information on a page ("hint bits") may be modified while
+holding only a share-exclusive lock and pin on the page. To do so in cases
+where only a share lock is already held, use BufferBeginSetHintBits() &
+BufferFinishSetHintBits() (if multiple hint bits are to be set) or
+BufferSetHintBits16() (if a single hit bit is set).
+
+E.g. for heapam, a share-exclusive lock allows to update tuple commit status
+bits (ie, OR the values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
 HEAP_XMAX_INVALID into t_infomask) while holding only a shared lock and
 pin on a buffer.  This is OK because another backend looking at the tuple
 at about the same time would OR the same bits into the field, so there
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index da83b775d0b..9ed7a368d74 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2422,9 +2422,8 @@ again:
 	/*
 	 * If the buffer was dirty, try to write it out.  There is a race
 	 * condition here, in that someone might dirty it after we released the
-	 * buffer header lock above, or even while we are writing it out (since
-	 * our share-lock won't prevent hint-bit updates).  We will recheck the
-	 * dirty bit after re-locking the buffer header.
+	 * buffer header lock above.  We will recheck the dirty bit after
+	 * re-locking the buffer header.
 	 */
 	if (buf_state & BM_DIRTY)
 	{
@@ -2432,12 +2431,12 @@ again:
 		Assert(buf_state & BM_VALID);
 
 		/*
-		 * We need a share-lock on the buffer contents to write it out (else
+		 * We need a share-exclusive lock on the buffer contents to write it out (else
 		 * we might write invalid data, eg because someone else is compacting
 		 * the page contents while we write).  We must use a conditional lock
 		 * acquisition here to avoid deadlock.  Even though the buffer was not
 		 * pinned (and therefore surely not locked) when StrategyGetBuffer
-		 * returned it, someone else could have pinned and exclusive-locked it
+		 * returned it, someone else could have pinned and (share-)exclusive-locked it
 		 * by the time we get here. If we try to get the lock unconditionally,
 		 * we'd block waiting for them; if they later block waiting for us,
 		 * deadlock ensues. (This has been observed to happen when two
@@ -2445,7 +2444,7 @@ again:
 		 * one just happens to be trying to split the page the first one got
 		 * from StrategyGetBuffer.)
 		 */
-		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -4014,8 +4013,8 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	}
 
 	/*
-	 * Pin it, share-lock it, write it.  (FlushBuffer will do nothing if the
-	 * buffer is clean by the time we've locked it.)
+	 * Pin it, share-exclusive-lock it, write it.  (FlushBuffer will do
+	 * nothing if the buffer is clean by the time we've locked it.)
 	 */
 	PinBuffer_Locked(bufHdr);
 
@@ -4329,11 +4328,8 @@ BufferGetTag(Buffer buffer, RelFileLocator *rlocator, ForkNumber *forknum,
  * However, we will need to force the changes to disk via fsync before
  * we can checkpoint WAL.
  *
- * The caller must hold a pin on the buffer and have share-locked the
- * buffer contents.  (Note: a share-lock does not prevent updates of
- * hint bits in the buffer, so the page could change while the write
- * is in progress, but we assume that that will not invalidate the data
- * written.)
+ * The caller must hold a pin on the buffer and have
+ * (share-)exclusively-locked the buffer contents.
  *
  * If the caller has an smgr reference for the buffer's relation, pass it
  * as the second parameter.  If not, pass NULL.
@@ -4349,6 +4345,9 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	char	   *bufToWrite;
 	uint64		buf_state;
 
+	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(buf, BUFFER_LOCK_SHARE_EXCLUSIVE));
+
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
 	 * someone else flushed the buffer before we could, so we need not do
@@ -4481,7 +4480,7 @@ FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(buf);
 
-	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 	BufferLockUnlock(buffer, buf);
 }
@@ -5400,8 +5399,8 @@ FlushDatabaseBuffers(Oid dbid)
 }
 
 /*
- * Flush a previously, shared or exclusively, locked and pinned buffer to the
- * OS.
+ * Flush a previously, share-exclusively or exclusively, locked and pinned
+ * buffer to the OS.
  */
 void
 FlushOneBuffer(Buffer buffer)
@@ -5474,39 +5473,23 @@ IncrBufferRefCount(Buffer buffer)
 }
 
 /*
- * MarkBufferDirtyHint
+ * Shared-buffer only helper for MarkBufferDirtyHint() and
+ * BufferSetHintBits16().
  *
- *	Mark a buffer dirty for non-critical changes.
- *
- * This is essentially the same as MarkBufferDirty, except:
- *
- * 1. The caller does not write WAL; so if checksums are enabled, we may need
- *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
- * 2. The caller might have only share-lock instead of exclusive-lock on the
- *	  buffer's content lock.
- * 3. This function does not guarantee that the buffer is always marked dirty
- *	  (due to a race condition), so it cannot be used for important changes.
+ * This is separated out because it turns out that the repeated checks for
+ * local buffers, repeated GetBufferDescriptor() and repeated reading of the
+ * buffer's state sufficiently hurts the performance of BufferSetHintBits16().
  */
-void
-MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+static inline void
+MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate, bool buffer_std)
 {
-	BufferDesc *bufHdr;
 	Page		page = BufferGetPage(buffer);
 
-	if (!BufferIsValid(buffer))
-		elog(ERROR, "bad buffer ID: %d", buffer);
-
-	if (BufferIsLocal(buffer))
-	{
-		MarkLocalBufferDirty(buffer);
-		return;
-	}
-
-	bufHdr = GetBufferDescriptor(buffer - 1);
-
 	Assert(GetPrivateRefCount(buffer) > 0);
-	/* here, either share or exclusive lock is OK */
-	Assert(BufferIsLockedByMe(buffer));
+
+	/* here, either share-exclusive or exclusive lock is OK */
+	Assert(BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_SHARE_EXCLUSIVE));
 
 	/*
 	 * This routine might get called many times on the same page, if we are
@@ -5519,8 +5502,8 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
-		(BM_DIRTY | BM_JUST_DIRTIED))
+	if (unlikely((lockstate & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+				 (BM_DIRTY | BM_JUST_DIRTIED)))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
@@ -5589,13 +5572,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 			dirtied = true;		/* Means "will be dirtied by this action" */
 
 			/*
-			 * Set the page LSN if we wrote a backup block. We aren't supposed
-			 * to set this when only holding a share lock but as long as we
-			 * serialise it somehow we're OK. We choose to set LSN while
-			 * holding the buffer header lock, which causes any reader of an
-			 * LSN who holds only a share lock to also obtain a buffer header
-			 * lock before using PageGetLSN(), which is enforced in
-			 * BufferGetLSNAtomic().
+			 * Set the page LSN if we wrote a backup block. To allow backends
+			 * that only hold a share lock on the buffer to read the LSN in a
+			 * tear-free manner, we set the page LSN while holding the buffer
+			 * header lock. This allows any reader of an LSN who holds only a
+			 * share lock to also obtain a buffer header lock before using
+			 * PageGetLSN() to read the LSN in a tear free way. This is done
+			 * in BufferGetLSNAtomic().
 			 *
 			 * If checksums are enabled, you might think we should reset the
 			 * checksum here. That will happen when the page is written
@@ -5621,6 +5604,40 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	}
 }
 
+/*
+ * MarkBufferDirtyHint
+ *
+ *	Mark a buffer dirty for non-critical changes.
+ *
+ * This is essentially the same as MarkBufferDirty, except:
+ *
+ * 1. The caller does not write WAL; so if checksums are enabled, we may need
+ *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
+ * 2. The caller might have only share-exclusive-lock instead of
+ *	  exclusive-lock on the buffer's content lock.
+ * 3. This function does not guarantee that the buffer is always marked dirty
+ *	  (due to a race condition), so it cannot be used for important changes.
+ */
+inline void
+MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+{
+	BufferDesc *bufHdr;
+
+	bufHdr = GetBufferDescriptor(buffer - 1);
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		MarkLocalBufferDirty(buffer);
+		return;
+	}
+
+	MarkSharedBufferDirtyHint(buffer, bufHdr, pg_atomic_read_u64(&bufHdr->state),
+							  buffer_std);
+}
+
 /*
  * Release buffer content locks for shared buffers.
  *
@@ -6673,6 +6690,126 @@ IsBufferCleanupOK(Buffer buffer)
 	return false;
 }
 
+static inline bool
+SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
+{
+	uint64		old_state;
+	PrivateRefCountEntry *ref;
+	BufferLockMode mode;
+
+	ref = GetPrivateRefCountEntry(buffer, true);
+
+	if (ref == NULL)
+		elog(ERROR, "lock is not held");
+
+	mode = ref->data.lockmode;
+	if (mode == BUFFER_LOCK_UNLOCK)
+		elog(ERROR, "buffer is not locked");
+
+	/*
+	 * Already am holding the required lock level.
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE || mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+	{
+		*lockstate = pg_atomic_read_u64(&buf_hdr->state);
+		return true;
+	}
+
+	/*
+	 * Only holding a share lock right now, try to upgrade to SHARE_EXCLUSIVE.
+	 */
+	Assert(mode == BUFFER_LOCK_SHARE);
+
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+	while (true)
+	{
+		uint64		desired_state;
+
+		desired_state = old_state;
+
+		/*
+		 * Can't upgrade if somebody else holds the lock in exlusive or
+		 * share-exclusive mode.
+		 */
+		if (unlikely((old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) != 0))
+		{
+			return false;
+		}
+
+		/* currently held lock state */
+		desired_state -= BM_LOCK_VAL_SHARED;
+
+		/* new lock level */
+		desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			ref->data.lockmode = BUFFER_LOCK_SHARE_EXCLUSIVE;
+			*lockstate = desired_state;
+
+			return true;
+		}
+	}
+
+}
+
+bool
+BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+
+	if (BufferIsLocal(buffer))
+	{
+		*ptr = val;
+
+		MarkLocalBufferDirty(buffer);
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	if (SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate))
+	{
+		*ptr = val;
+
+		MarkSharedBufferDirtyHint(buffer, buf_hdr, lockstate, true);
+
+		return true;
+	}
+
+	return false;
+}
+
+bool
+BufferBeginSetHintBits(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+
+	if (BufferIsLocal(buffer))
+	{
+		/*
+		 * TODO: will need to check for write IO once that's done
+		 * asynchronously.
+		 */
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	return SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate);
+}
+
+void
+BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std)
+{
+	if (mark_dirty)
+		MarkBufferDirtyHint(buffer, buffer_std);
+}
 
 /*
  *	Functions for buffer I/O handling
diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index 4773a9cc65e..6cdfbbeb260 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -904,14 +904,22 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 	max_avail = fsm_get_max_avail(page);
 
 	/*
-	 * Reset the next slot pointer. This encourages the use of low-numbered
-	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation.  We don't bother with a lock here, nor with marking the page
-	 * dirty if it wasn't already, since this is just a hint.
+	 * Try to reset the next slot pointer. This encourages the use of
+	 * low-numbered pages, increasing the chances that a later vacuum can
+	 * truncate the relation.  We don't bother with a lock here, nor with
+	 * marking the page dirty if it wasn't already, since this is just a hint.
+	 *
+	 * To be allowed to update the page without an exclusive lock, we have to
+	 * use the hint bit infrastructure.
 	 */
-	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	LockBuffer(buf, BUFFER_LOCK_SHARE);
+	if (BufferBeginSetHintBits(buf))
+	{
+		((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+		BufferFinishSetHintBits(buf, false, false);
+	}
 
-	ReleaseBuffer(buf);
+	UnlockReleaseBuffer(buf);
 
 	return max_avail;
 }
diff --git a/src/backend/storage/freespace/fsmpage.c b/src/backend/storage/freespace/fsmpage.c
index 66a5c80b5a6..a59696b6484 100644
--- a/src/backend/storage/freespace/fsmpage.c
+++ b/src/backend/storage/freespace/fsmpage.c
@@ -298,9 +298,18 @@ restart:
 	 * lock and get a garbled next pointer every now and then, than take the
 	 * concurrency hit of an exclusive lock.
 	 *
+	 * Without an exclusive lock, we need to use the hint bit infrastructure
+	 * to be allowed to modify the page.
+	 *
 	 * Wrap-around is handled at the beginning of this function.
 	 */
-	fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+	if (exclusive_lock_held || BufferBeginSetHintBits(buf))
+	{
+		fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+
+		if (!exclusive_lock_held)
+			BufferFinishSetHintBits(buf, false, true);
+	}
 
 	return slot;
 }
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 9d14239b4c4..66e8f2e9fa6 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2731,6 +2731,7 @@ SetConstraintStateData
 SetConstraintTriggerData
 SetExprState
 SetFunctionReturnMode
+SetHintBitsState
 SetOp
 SetOpCmd
 SetOpPath
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0013-WIP-Make-UnlockReleaseBuffer-more-efficient.patch (3.5K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/14-v6-0013-WIP-Make-UnlockReleaseBuffer-more-efficient.patch)
  download | inline diff:
From 1c24af4fedb372b9b89d0635895860889da86b47 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 19 Nov 2025 15:32:20 -0500
Subject: [PATCH v6 13/14] WIP: Make UnlockReleaseBuffer() more efficient

Now that the buffer content lock is implemented as part of BufferDesc.state,
releasing the lock and unpinning the buffer can be implemented as a single
atomic operation.

Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/nbtree/nbtpage.c | 22 +++++++++++-
 src/backend/storage/buffer/bufmgr.c | 52 ++++++++++++++++++++++++++++-
 2 files changed, 72 insertions(+), 2 deletions(-)

diff --git a/src/backend/access/nbtree/nbtpage.c b/src/backend/access/nbtree/nbtpage.c
index 30b43a4dd18..2fd8141854c 100644
--- a/src/backend/access/nbtree/nbtpage.c
+++ b/src/backend/access/nbtree/nbtpage.c
@@ -1006,11 +1006,18 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 
 	Assert(BlockNumberIsValid(blkno));
 	if (BufferIsValid(obuf))
+	{
+		_bt_relbuf(rel, obuf);
+#if 0
+		Assert(BufferGetBlockNumber(obuf) != blkno);
 		_bt_unlockbuf(rel, obuf);
-	buf = ReleaseAndReadBuffer(obuf, rel, blkno);
+#endif
+	}
+	buf = ReadBuffer(rel, blkno);
 	_bt_lockbuf(rel, buf, access);
 
 	_bt_checkpage(rel, buf);
+
 	return buf;
 }
 
@@ -1022,8 +1029,21 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 void
 _bt_relbuf(Relation rel, Buffer buf)
 {
+#if 0
 	_bt_unlockbuf(rel, buf);
 	ReleaseBuffer(buf);
+#else
+	/*
+	 * Buffer is pinned and locked, which means that it is expected to be
+	 * defined and addressable.  Check that proactively.
+	 */
+	VALGRIND_CHECK_MEM_IS_DEFINED(BufferGetPage(buf), BLCKSZ);
+
+	UnlockReleaseBuffer(buf);
+
+	if (!RelationUsesLocalBuffers(rel))
+		VALGRIND_MAKE_MEM_NOACCESS(BufferGetPage(buf), BLCKSZ);
+#endif
 }
 
 /*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 9ed7a368d74..584c3b2ee75 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5437,13 +5437,63 @@ ReleaseBuffer(Buffer buffer)
 /*
  * UnlockReleaseBuffer -- release the content lock and pin on a buffer
  *
- * This is just a shorthand for a common combination.
+ * This is just a, more efficient, shorthand for a common combination.
  */
 void
 UnlockReleaseBuffer(Buffer buffer)
 {
+#if 1
+	int			mode;
+	BufferDesc *buf;
+	PrivateRefCountEntry *ref;
+	uint64		sub;
+	uint64		lockstate;
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		UnpinLocalBuffer(buffer);
+		return;
+	}
+
+	ResourceOwnerForgetBuffer(CurrentResourceOwner, buffer);
+
+	buf = GetBufferDescriptor(buffer - 1);
+
+	mode = BufferLockDisownInternal(buffer, buf);
+
+	/* compute state modification for lock release */
+	sub = BufferLockReleaseSub(mode);
+
+	/* compute state modification for pin release */
+	ref = GetPrivateRefCountEntry(buffer, false);
+	Assert(ref != NULL);
+	Assert(ref->data.refcount > 0);
+	ref->data.refcount--;
+
+	if (ref->data.refcount == 0)
+	{
+		sub |= BUF_REFCOUNT_ONE;
+		ForgetPrivateRefCountEntry(ref);
+	}
+
+	/* perform the lock and pin release in one atomic op */
+	lockstate = pg_atomic_sub_fetch_u64(&buf->state, sub);
+
+	/* wake up waiters etc */
+	BufferLockProcessRelease(buf, mode, lockstate);
+
+	if (lockstate & BM_PIN_COUNT_WAITER)
+		WakePinCountWaiter(buf);
+
+	RESUME_INTERRUPTS();
+
+#else
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 	ReleaseBuffer(buffer);
+#endif
 }
 
 /*
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v6-0014-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch (11.6K, ../../6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar/15-v6-0014-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch)
  download | inline diff:
From 8b73c9143118487ee3e7663b4adb050db99b86d8 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 17 Oct 2024 14:14:35 -0400
Subject: [PATCH v6 14/14] WIP: bufmgr: Don't copy pages while writing out

After the series of preceding commits introducing and using
BufferBeginSetHintBits()/BufferSetHintBits16() hint bits are not set
anymore while IO is going on. Therefore we do not need to copy pages while
they are being written out anymore.

TODO: Update comments

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/bufpage.h           |  3 +-
 src/backend/access/hash/hashpage.c      |  2 +-
 src/backend/access/transam/xloginsert.c | 43 ++++++----------------
 src/backend/storage/buffer/bufmgr.c     | 21 +++++------
 src/backend/storage/buffer/localbuf.c   |  2 +-
 src/backend/storage/page/bufpage.c      | 48 ++++---------------------
 src/backend/storage/smgr/bulk_write.c   |  2 +-
 src/test/modules/test_aio/test_aio.c    |  2 +-
 8 files changed, 33 insertions(+), 90 deletions(-)

diff --git a/src/include/storage/bufpage.h b/src/include/storage/bufpage.h
index abc2cf2a020..f8f621446c4 100644
--- a/src/include/storage/bufpage.h
+++ b/src/include/storage/bufpage.h
@@ -504,7 +504,6 @@ extern void PageIndexMultiDelete(Page page, OffsetNumber *itemnos, int nitems);
 extern void PageIndexTupleDeleteNoCompact(Page page, OffsetNumber offnum);
 extern bool PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 									const void *newtup, Size newsize);
-extern char *PageSetChecksumCopy(Page page, BlockNumber blkno);
-extern void PageSetChecksumInplace(Page page, BlockNumber blkno);
+extern void PageSetChecksum(Page page, BlockNumber blkno);
 
 #endif							/* BUFPAGE_H */
diff --git a/src/backend/access/hash/hashpage.c b/src/backend/access/hash/hashpage.c
index b8e5bd005e5..dd17eff59d1 100644
--- a/src/backend/access/hash/hashpage.c
+++ b/src/backend/access/hash/hashpage.c
@@ -1029,7 +1029,7 @@ _hash_alloc_buckets(Relation rel, BlockNumber firstblock, uint32 nblocks)
 					zerobuf.data,
 					true);
 
-	PageSetChecksumInplace(page, lastblock);
+	PageSetChecksum(page, lastblock);
 	smgrextend(RelationGetSmgr(rel), MAIN_FORKNUM, lastblock, zerobuf.data,
 			   false);
 
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index a56d5a55282..0af148e9496 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -261,8 +261,11 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	 */
 #ifdef USE_ASSERT_CHECKING
 	if (!(flags & REGBUF_NO_CHANGE))
-		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) &&
-			   BufferIsDirty(buffer));
+	{
+		Assert(BufferIsDirty(buffer));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE));
+	}
 #endif
 
 	if (block_id >= max_registered_block_id)
@@ -1066,7 +1069,7 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * Write a backup block if needed when we are setting a hint. Note that
  * this may be called for a variety of page types, not just heaps.
  *
- * Callable while holding just share lock on the buffer content.
+ * Callable while holding just share-exclusive lock on the buffer content.
  *
  * We can't use the plain backup block mechanism since that relies on the
  * Buffer being exclusively locked. Since some modifications (setting LSN, hint
@@ -1074,6 +1077,8 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * failures. So instead we copy the page and insert the copied data as normal
  * record data.
  *
+ * FIXME: outdated
+ *
  * We only need to do something if page has not yet been full page written in
  * this checkpoint round. The LSN of the inserted wal record is returned if we
  * had to write, InvalidXLogRecPtr otherwise.
@@ -1102,46 +1107,20 @@ XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
 
 	/*
 	 * We assume page LSN is first data on *every* page that can be passed to
-	 * XLogInsert, whether it has the standard page layout or not. Since we're
-	 * only holding a share-lock on the page, we must take the buffer header
-	 * lock when we look at the LSN.
+	 * XLogInsert, whether it has the standard page layout or not.
 	 */
 	lsn = BufferGetLSNAtomic(buffer);
 
 	if (lsn <= RedoRecPtr)
 	{
-		int			flags = 0;
-		PGAlignedBlock copied_buffer;
-		char	   *origdata = (char *) BufferGetBlock(buffer);
-		RelFileLocator rlocator;
-		ForkNumber	forkno;
-		BlockNumber blkno;
-
-		/*
-		 * Copy buffer so we don't have to worry about concurrent hint bit or
-		 * lsn updates. We assume pd_lower/upper cannot be changed without an
-		 * exclusive lock, so the contents bkp are not racy.
-		 */
-		if (buffer_std)
-		{
-			/* Assume we can omit data between pd_lower and pd_upper */
-			Page		page = BufferGetPage(buffer);
-			uint16		lower = ((PageHeader) page)->pd_lower;
-			uint16		upper = ((PageHeader) page)->pd_upper;
-
-			memcpy(copied_buffer.data, origdata, lower);
-			memcpy(copied_buffer.data + upper, origdata + upper, BLCKSZ - upper);
-		}
-		else
-			memcpy(copied_buffer.data, origdata, BLCKSZ);
+		int			flags = REGBUF_NO_CHANGE;
 
 		XLogBeginInsert();
 
 		if (buffer_std)
 			flags |= REGBUF_STANDARD;
 
-		BufferGetTag(buffer, &rlocator, &forkno, &blkno);
-		XLogRegisterBlock(0, &rlocator, forkno, blkno, copied_buffer.data, flags);
+		XLogRegisterBuffer(0, buffer, flags);
 
 		recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI_FOR_HINT);
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 584c3b2ee75..d6a638613ae 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -4342,7 +4342,6 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	ErrorContextCallback errcallback;
 	instr_time	io_start;
 	Block		bufBlock;
-	char	   *bufToWrite;
 	uint64		buf_state;
 
 	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
@@ -4413,12 +4412,8 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	 */
 	bufBlock = BufHdrGetBlock(buf);
 
-	/*
-	 * Update page checksum if desired.  Since we have only shared lock on the
-	 * buffer, other processes might be updating hint bits in it, so we must
-	 * copy the page to private storage if we do checksumming.
-	 */
-	bufToWrite = PageSetChecksumCopy((Page) bufBlock, buf->tag.blockNum);
+	/* Update page checksum if desired. */
+	PageSetChecksum((Page) bufBlock, buf->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
@@ -4428,7 +4423,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	smgrwrite(reln,
 			  BufTagGetForkNum(&buf->tag),
 			  buf->tag.blockNum,
-			  bufToWrite,
+			  bufBlock,
 			  false);
 
 	/*
@@ -4552,8 +4547,8 @@ BufferIsPermanent(Buffer buffer)
 /*
  * BufferGetLSNAtomic
  *		Retrieves the LSN of the buffer atomically using a buffer header lock.
- *		This is necessary for some callers who may not have an exclusive lock
- *		on the buffer.
+ *		This is necessary for some callers who may not have a (share-)exclusive
+ *		lock on the buffer.
  */
 XLogRecPtr
 BufferGetLSNAtomic(Buffer buffer)
@@ -5606,6 +5601,12 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate, b
 			 * It's possible we may enter here without an xid, so it is
 			 * essential that CreateCheckPoint waits for virtual transactions
 			 * rather than full transactionids.
+			 *
+			 * FIXME: I think we now should simply mark the page dirty before
+			 * WAL logging the hint bit - afaikt it then should work just like
+			 * any other buffer write (due to SyncBuffers()/SyncOneBuffer()
+			 * seeing the dirty bit and trying to lock the page
+			 * share-exclusive, and thus having to wait).
 			 */
 			Assert((MyProc->delayChkptFlags & DELAY_CHKPT_START) == 0);
 			MyProc->delayChkptFlags |= DELAY_CHKPT_START;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index a41a5facd3a..5826d4b54c6 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -199,7 +199,7 @@ FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln)
 		reln = smgropen(BufTagGetRelFileLocator(&bufHdr->tag),
 						MyProcNumber);
 
-	PageSetChecksumInplace(localpage, bufHdr->tag.blockNum);
+	PageSetChecksum(localpage, bufHdr->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index aac6e695954..c8cbdd1f7a6 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -1494,51 +1494,15 @@ PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 /*
  * Set checksum for a page in shared buffers.
  *
- * If checksums are disabled, or if the page is not initialized, just return
- * the input.  Otherwise, we must make a copy of the page before calculating
- * the checksum, to prevent concurrent modifications (e.g. setting hint bits)
- * from making the final checksum invalid.  It doesn't matter if we include or
- * exclude hints during the copy, as long as we write a valid page and
- * associated checksum.
+ * If checksums are disabled, or if the page is not initialized, just
+ * return. Otherwise compute and set the checksum.
  *
- * Returns a pointer to the block-sized data that needs to be written. Uses
- * statically-allocated memory, so the caller must immediately write the
- * returned page and not refer to it again.
- */
-char *
-PageSetChecksumCopy(Page page, BlockNumber blkno)
-{
-	static char *pageCopy = NULL;
-
-	/* If we don't need a checksum, just return the passed-in data */
-	if (PageIsNew(page) || !DataChecksumsEnabled())
-		return page;
-
-	/*
-	 * We allocate the copy space once and use it over on each subsequent
-	 * call.  The point of palloc'ing here, rather than having a static char
-	 * array, is first to ensure adequate alignment for the checksumming code
-	 * and second to avoid wasting space in processes that never call this.
-	 */
-	if (pageCopy == NULL)
-		pageCopy = MemoryContextAllocAligned(TopMemoryContext,
-											 BLCKSZ,
-											 PG_IO_ALIGN_SIZE,
-											 0);
-
-	memcpy(pageCopy, page, BLCKSZ);
-	((PageHeader) pageCopy)->pd_checksum = pg_checksum_page(pageCopy, blkno);
-	return pageCopy;
-}
-
-/*
- * Set checksum for a page in private memory.
- *
- * This must only be used when we know that no other process can be modifying
- * the page buffer.
+ * In the past this needed to be done on a copy of the page, due to the
+ * possibility of e.g. hint bits being set concurrently. However, this is not
+ * necessary anymore as hint bits won't be set while IO is going on.
  */
 void
-PageSetChecksumInplace(Page page, BlockNumber blkno)
+PageSetChecksum(Page page, BlockNumber blkno)
 {
 	/* If we don't need a checksum, just return */
 	if (PageIsNew(page) || !DataChecksumsEnabled())
diff --git a/src/backend/storage/smgr/bulk_write.c b/src/backend/storage/smgr/bulk_write.c
index b958be15716..f4d07543365 100644
--- a/src/backend/storage/smgr/bulk_write.c
+++ b/src/backend/storage/smgr/bulk_write.c
@@ -279,7 +279,7 @@ smgr_bulk_flush(BulkWriteState *bulkstate)
 		BlockNumber blkno = pending_writes[i].blkno;
 		Page		page = pending_writes[i].buf->data;
 
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 
 		if (blkno >= bulkstate->relsize)
 		{
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index 488d98e7e66..e5fc7642dc2 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -288,7 +288,7 @@ modify_rel_block(PG_FUNCTION_ARGS)
 	}
 	else
 	{
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 	}
 
 	smgrwrite(RelationGetSmgr(rel),
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-11-20 19:08                 ` Greg Burd <greg@burd.me>
  2025-11-20 20:51                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  3 siblings, 1 reply; 120+ messages in thread

From: Greg Burd @ 2025-11-20 19:08 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers; Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>


On Nov 19 2025, at 9:47 pm, Andres Freund <andres@anarazel.de> wrote:

> Hi,
> 
> On 2025-10-09 17:16:49 -0400, Andres Freund wrote:
>> On 2025-10-09 16:35:44 -0400, Andres Freund wrote:
>> > I pushed a few commits from this patchset after Matthias' review
>> > (thanks!). Unfortunately in 5e899859287 I missed that the valgrind annotations
>> > would not be done anymore for the buffers returned by
>> > StrategyGetBuffer(). Which turned skink red.
>> > 
>> > The attached 0001 patch centralizes the valgrind initialization in
>> > TrackNewBufferPin(), which 5e899859287 had added. The nice side
>> effect of that
>> > is that there are fewer VALGRIND_MAKE_MEM_DEFINED() calls than
>> before. The
>> > naming isn't the perfect match, but it seems fine to me.
>> 
>> Forgot to say: I'll push this patch soon, to get skink back to green. Unless
>> somebody says something.  We can adjust this later, if the comment and/or
>> placement of VALGRIND_MAKE_MEM_DEFINED() isn't to everyones liking.
> 
> I have pushed that fix as well as the subsequent buffer header locking changes
> a while ago.

Hello Andres,

After talking to you about these ideas at PGConf in NYC I've been
anxiously awaiting this patch set. Thanks for dedicating the energy and
time to get it to this stage.

High level feedback after reading the patches/email/commit messages is
that it looks to get you to where you wanted to be, unblocking AIO
writes. I think they'll end up being faster than what's in place now,
even before you get to the AIO piece.  Certainly removing the copy of
each page to do a checksum will help.  Opening the door to a future
where we can have super-pinned/locked pages is also a net win.

Everything before/after 0008 was rather easy to digest and understand
and I found nothing really to call out at this stage

0008 is understandable too, it's just sizable. While it is large, I find
it well laid out and more readable than before.  I gave the locking code
a good look, it seems correct AFAICT.

Keep going, I'll be happy to dedicate time to testing and digging into
the commits as you get this into a final state.  I look forward to
extending/enhancing this code once integrated.

> I want to again emphasize that the important commits (i.e. 0008, 0012, 0014)
> aren't close to being mergeable. But I think they're in a stage that they
> could benefit from "lenient" high-level review.
> 
> Greetings,
> 
> Andres Freund

cheers.

-greg





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 19:08                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Greg Burd <greg@burd.me>
@ 2025-11-20 20:51                   ` Andres Freund <andres@anarazel.de>
  2025-11-25 00:09                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Thomas Munro <thomas.munro@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-11-20 20:51 UTC (permalink / raw)
  To: Greg Burd <greg@burd.me>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers; Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-11-20 14:08:57 -0500, Greg Burd wrote:
> After talking to you about these ideas at PGConf in NYC I've been
> anxiously awaiting this patch set. Thanks for dedicating the energy and
> time to get it to this stage.

Thanks!


> High level feedback after reading the patches/email/commit messages is
> that it looks to get you to where you wanted to be, unblocking AIO
> writes. I think they'll end up being faster than what's in place now,
> even before you get to the AIO piece.  Certainly removing the copy of
> each page to do a checksum will help.  Opening the door to a future
> where we can have super-pinned/locked pages is also a net win.

I actually had planned to write about the performance effects, in particular
of 0012, a bit more:

It's worth pointing out that the new way of setting hint bits is inherently
more expensive than what we did before - upgrading a lock to a different lock
level isn't free, compared to doing, well, nothing.

For paths that set the hint bits of a whole page, like a seqscan, that cost is
more than amortized by the batched approach introduced in 0011. Those get
faster with the patch, both when already hinted and when not.

However, there are paths that aren't easily amenable to that approach, like
e.g. an ordered index scan referencing unhinted tuples. There we only ever
access a single tuple and release the upgraded lock after every tuple. If the
index scan is perfectly correlated with the table and every tuple is unhinted,
that's a decent amount of additional work.

I've spent a lot of time micro-optimizing that workload, to avoid any
significiant regressions. An extreme stress-test started out being about 20%
slower than today, as of my current local version, it's a bit faster (~1%) on
one of my machines and a bit slower (~2%) on another. Partially that was
achieved by optimizing the hint-bit-lock-upgrade code more (e.g. having a fast
path for updating a single hint bit, avoiding redundant reads of the lock
state by having MarkSharedBufferDirtyHint(), ...), partially by optimizing the
locking code.  The latter is a bit of a cheat though - things would be even
faster if we went with the old way of setting hint bits, but with the
independent optimizations applied.

I think that's ok though:

1) the old way of setting hint bits is a pretty dirty hack that causes issues
   in quite a few places.

2) by definition, having to set hint bits is an ephemeral state, once the hint
   bits are set, the difference vanishes

3) no normal workload shows the difference - my stress test does
   SELECT * FROM manyrows_idx ORDER BY i OFFSET 10000000;
   on a perfectly correlated table with very narrow rows, i.e. an index scan
   of the whole table, where none of the scan results are ever used. Once one
   actually uses the resulting rows, the performance difference completely
   vanishes.

4) as part of the index prefetching work, we might get the infrastructure to
   actually batch the hint-bit setting in this case too.


I see some mild performance gains in workloads like pgbench [-S], but nothing
to write home about. Which I think is about what I would expect - there's a
minor efficiency gain due to the private refcount changes and not having state
tracking for two error recovery mechanisms (lwlocks' and private refcount).

Non-all-visible seqscans do see some performance gain due to 0011, whether the
table is hinted or not. But it's again something that's mostly noticeable in
microbenchmarks, as e.g. tuple deforming or qual evaluation has a much bigger
impact.


> Everything before/after 0008 was rather easy to digest and understand
> and I found nothing really to call out at this stage
>
> 0008 is understandable too, it's just sizable. While it is large, I find
> it well laid out and more readable than before.  I gave the locking code
> a good look, it seems correct AFAICT.

I hope so :), the locking logic it's largely the same as lwlock.c, with some
exceptions due to the added lock level and differences in error handling
state.

I don't really see a good way to split 0008 unfortunately...  I previously had
split 0012 into four patches (core changes, heapam changes, the rest of the
adaptions to setting hint bits in various places, adding assertions relying on
the different lock levels), but I found it pretty unwieldly from a "comment
management" perspective, because comments that need to be rewritten are
temporarily wrong, or would need to be modified in yet another path.  But I'm
open to going back to that approach.

I guess I could pull out the addition of UnlockBuffer() and the "redirection"
to it from LockBuffer() into a separate patch.


> Keep going, I'll be happy to dedicate time to testing and digging into
> the commits as you get this into a final state.  I look forward to
> extending/enhancing this code once integrated.

Cool!

I think 0001, 0002, 0003, 0005 and 0009 should be mergeable pretty soon.  I've
some further polishing to do for 0006 and 0007, but I think they could go in
well ahead of the rest.

For 0008, in addition to what's noted in the commit message, I think there
needs to be an additional section in src/backend/storage/buffer/README (or
such), explaining that buffer content locks used to be lwlocks but aren't
anymore for xyz reasons.  I suspect it'd also be good to have a few references
from lwlock.c to bufmgr.c to make sure the code is co-evolved.

For 0012, I think it might make sense to pull out some of the changes to
fsm_vacuum_page() out into a separate commit. Basically changing the code to
acquire a lock on the page - I still can't quite believe that somebody thought
it's sane to update the page without even bothering with a share lock.

0014 needs a lot more polishing, there's references to the hint bit / locking
interactions all over. Pretty hard to find all the references :(.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 19:08                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Greg Burd <greg@burd.me>
  2025-11-20 20:51                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-11-25 00:09                     ` Thomas Munro <thomas.munro@gmail.com>
  0 siblings, 0 replies; 120+ messages in thread

From: Thomas Munro @ 2025-11-25 00:09 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Greg Burd <greg@burd.me>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers; Melanie Plageman <melanieplageman@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Fri, Nov 21, 2025 at 9:51 AM Andres Freund <andres@anarazel.de> wrote:
> It's worth pointing out that the new way of setting hint bits is inherently
> more expensive than what we did before - upgrading a lock to a different lock
> level isn't free, compared to doing, well, nothing.
>
> For paths that set the hint bits of a whole page, like a seqscan, that cost is
> more than amortized by the batched approach introduced in 0011. Those get
> faster with the patch, both when already hinted and when not.

Nice work!

> However, there are paths that aren't easily amenable to that approach, like
> e.g. an ordered index scan referencing unhinted tuples. There we only ever
> access a single tuple and release the upgraded lock after every tuple. If the
> index scan is perfectly correlated with the table and every tuple is unhinted,
> that's a decent amount of additional work.

Yeah, but it was only faster because it was cheating.  It presumably
doesn't happen when you bulk load and then create index.  It
presumably does happen when you insert a lot of data in order, on
first correlated index scan.  Seems like an inherent limitation of the
current tuple-at-a-time architecture when combined with the *required*
interlocking, and not a blocker for this work.

+ Some filesystems, raid implementations, ... do not tolerate the data being

I was aware of BTRFS (EIO on read) and ZFS 2.4 (EIO on read or write
depending on configuration option), but hadn't thought about RAID.
Ugh, right, non-matching RAID1 mirrors (and I guess also b0rked RAID5
parity bits?).  Fun.

https://bugzilla.kernel.org/show_bug.cgi?id=99171

> I've spent a lot of time micro-optimizing that workload, to avoid any
> significiant regressions. An extreme stress-test started out being about 20%
> slower than today, as of my current local version, it's a bit faster (~1%) on
> one of my machines and a bit slower (~2%) on another. Partially that was
> achieved by optimizing the hint-bit-lock-upgrade code more (e.g. having a fast
> path for updating a single hint bit, avoiding redundant reads of the lock
> state by having MarkSharedBufferDirtyHint(), ...), partially by optimizing the
> locking code.  The latter is a bit of a cheat though - things would be even
> faster if we went with the old way of setting hint bits, but with the
> independent optimizations applied.
>
> I think that's ok though:
>
> 1) the old way of setting hint bits is a pretty dirty hack that causes issues
>    in quite a few places.
>
> 2) by definition, having to set hint bits is an ephemeral state, once the hint
>    bits are set, the difference vanishes
>
> 3) no normal workload shows the difference - my stress test does
>    SELECT * FROM manyrows_idx ORDER BY i OFFSET 10000000;
>    on a perfectly correlated table with very narrow rows, i.e. an index scan
>    of the whole table, where none of the scan results are ever used. Once one
>    actually uses the resulting rows, the performance difference completely
>    vanishes.
>
> 4) as part of the index prefetching work, we might get the infrastructure to
>    actually batch the hint-bit setting in this case too.

Yeah.  Was just thinking the same.  Both the streaming and batching
projects have opportunities to figure out an amortisation scheme.  I
have a few vague ideas about stream-based approaches already, hmm...

+1, I think this is OK for now.





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-11-21 17:52                 ` Melanie Plageman <melanieplageman@gmail.com>
  2025-12-01 20:28                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  3 siblings, 1 reply; 120+ messages in thread

From: Melanie Plageman @ 2025-11-21 17:52 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

This email is just a review for some (specified below) of the patches 0001-0007

On Wed, Nov 19, 2025 at 9:47 PM Andres Freund <andres@anarazel.de> wrote:
>
> 0002: Not really required, but seems like an improvement to me

The commit message says the point is to get compiler warnings for
switch cases that should be exhaustive, but as soon as I looked for a
switch case like that, I see BufferIsLockedByMeInMode()

        switch (mode)
        {
            case BUFFER_LOCK_EXCLUSIVE:
                lw_mode = LW_EXCLUSIVE;
                break;
            case BUFFER_LOCK_SHARE:
                lw_mode = LW_SHARED;
                break;
            default:
                pg_unreachable();
        }

Which makes it impossible to get such a warning. When I add a lock
mode, it specifically doesn't warn when compiling.
However, I'm a big fan of using enums instead of macros when
appropriate, so I have no issue with this change. I just think the
commit message is a bit confusing.

> 0003: A prerequisite to 0004, pretty boring itself

LGTM.

> 0004: Use 64bit atomics for BufferDesc.state - at this point nothing uses the
> additional bits yet, though.  Some annoying reformatting required to avoid
> long lines.

I noticed that the BUF_STATE_GET_REFCOUNT and BUF_STATE_GET_USAGECOUNT
macros cast the return value to a uint32. We won't use the extra bits
but we did bother to keep the macro result sized to the field width
before so keeping it uint32 is probably more confusing now that state
is 64 bit.

Not related to this patch, but I noticed GetBufferDescriptor() calls
for a uint32 and all the callers pretty much pass a signed int —
wonder why it calls for uint32.

> 0005: There already was a wait event class for BUFFERPIN. It seems better to
> make that more general than to implement them separately.

I reviewed and see no issues with the code, but I don't have an
opinion on this wait event naming so maybe you better _wait_ for some
other review ;)

> 0006+0007: This is preparatory work for 0008, but also worthwhile on its
> own. The private refcount stuff does show up in profiles. The reason it's
> related is that without these changes the added information in 0008 makes that
> worse.

I found it slightly confusing that this commit appears to
unnecessarily add the PrivateRefCountData struct (given that it
doesn't need it to do the new parallel array thing). You could wait
until you need it in 0008, but 0008 is big as it is, so it probably is
fine where it is.

in InitBufferManagerAccess(), why do you have

memset(&PrivateRefCountArrayKeys, 0, sizeof(Buffer));
seems like it should be
memset(PrivateRefCountArrayKeys, 0, sizeof(PrivateRefCountArrayKeys));

I wonder how easy it will be to keep the Buffer in sync between
PrivateRefCountArrayKeys and the PrivateRefCountEntry — would a helper
function help?

ForgetPrivateRefCountEntry doesn’t clear the data member — but maybe
it doesn’t matter...

in ReservePrivateRefCountEntry() there is a superfluous clear

memset(&victim_entry->data, 0, sizeof(victim_entry->data));
victim_entry->data.refcount = 0;

0007
needs a commit message. overall seems fine though.
You should probably capitalize the "c" of "count" in
PrivateRefcountEntryLast to be consistent with the other names.

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-21 17:52                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
@ 2025-12-01 20:28                   ` Andres Freund <andres@anarazel.de>
  2025-12-01 20:41                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-12-01 20:28 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-11-21 12:52:38 -0500, Melanie Plageman wrote:
> > 0004: Use 64bit atomics for BufferDesc.state - at this point nothing uses the
> > additional bits yet, though.  Some annoying reformatting required to avoid
> > long lines.
>
> I noticed that the BUF_STATE_GET_REFCOUNT and BUF_STATE_GET_USAGECOUNT
> macros cast the return value to a uint32. We won't use the extra bits
> but we did bother to keep the macro result sized to the field width
> before so keeping it uint32 is probably more confusing now that state
> is 64 bit.

I can't really follow - why would we want to return a 64bit value if none of
the values ever can get anywhere near that big?


> Not related to this patch, but I noticed GetBufferDescriptor() calls
> for a uint32 and all the callers pretty much pass a signed int —
> wonder why it calls for uint32.

It's not strictly required - no Buffers exist bet INT32_MAX and
UINT32_MAX. However, GetBufferDescriptor() cannot be used for local buffers
(which would have a negative buffer id), therefore a uint32 is fine. Many of
the callers have dedicated branches to deal with local buffers and therefore
couldn't use a uint32.


> > 0006+0007: This is preparatory work for 0008, but also worthwhile on its
> > own. The private refcount stuff does show up in profiles. The reason it's
> > related is that without these changes the added information in 0008 makes that
> > worse.
>
> I found it slightly confusing that this commit appears to
> unnecessarily add the PrivateRefCountData struct (given that it
> doesn't need it to do the new parallel array thing). You could wait
> until you need it in 0008, but 0008 is big as it is, so it probably is
> fine where it is.

It seemed too annoying to whack the code around multiple times...


> in InitBufferManagerAccess(), why do you have
>
> memset(&PrivateRefCountArrayKeys, 0, sizeof(Buffer));
> seems like it should be
> memset(PrivateRefCountArrayKeys, 0, sizeof(PrivateRefCountArrayKeys));

Ugh, indeed.


> I wonder how easy it will be to keep the Buffer in sync between
> PrivateRefCountArrayKeys and the PrivateRefCountEntry — would a helper
> function help?

I don't think we really need it, the existing helper functions are where it
should be manipulated.


> ForgetPrivateRefCountEntry doesn’t clear the data member — but maybe
> it doesn’t matter...

It asserts that the fields are reset, which seems to suffice.


> in ReservePrivateRefCountEntry() there is a superfluous clear
>
> memset(&victim_entry->data, 0, sizeof(victim_entry->data));
> victim_entry->data.refcount = 0;

It's indeed superfluous today. I guess I put it in as a belt and suspenders
approach to future members... The compiler is easily be able to optimize that
redundancy away.


> 0007
> needs a commit message. overall seems fine though.

This is what I've since written:

    bufmgr: Add one-entry cache for private refcount

    The private refcount entry for a buffer is often looked up repeatedly for the
    same buffer, e.g. to pin and then unpin a buffer. Benchmarking shows that it's
    worthwhile to have a one-entry cache for that case. With that cache in place,
    it's worth splitting GetPrivateRefCountEntry() into a small inline
    portion (for the cache hit case) and an out-of-line helper for the rest.

    This is helpful for some workloads today, but becomes more important in an
    upcoming patch that will utilize the private refcount infrastructure to also
    store whether the buffer is currently locked, as that increases the rate of
    lookups substantially.



> You should probably capitalize the "c" of "count" in
> PrivateRefcountEntryLast to be consistent with the other names.

Ooops, yes.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-21 17:52                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-12-01 20:28                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-12-01 20:41                     ` Melanie Plageman <melanieplageman@gmail.com>
  0 siblings, 0 replies; 120+ messages in thread

From: Melanie Plageman @ 2025-12-01 20:41 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Mon, Dec 1, 2025 at 3:28 PM Andres Freund <andres@anarazel.de> wrote:
>
> > I noticed that the BUF_STATE_GET_REFCOUNT and BUF_STATE_GET_USAGECOUNT
> > macros cast the return value to a uint32. We won't use the extra bits
> > but we did bother to keep the macro result sized to the field width
> > before so keeping it uint32 is probably more confusing now that state
> > is 64 bit.
>
> I can't really follow - why would we want to return a 64bit value if none of
> the values ever can get anywhere near that big?

Well, usagecount could never have reached anything close to a uint32
either, so I thought that this cast was meant to match the datatype. I
found it rather confusing otherwise.

> > 0007
> > needs a commit message. overall seems fine though.
>
> This is what I've since written:
>
>     bufmgr: Add one-entry cache for private refcount
>
>     The private refcount entry for a buffer is often looked up repeatedly for the
>     same buffer, e.g. to pin and then unpin a buffer. Benchmarking shows that it's
>     worthwhile to have a one-entry cache for that case. With that cache in place,
>     it's worth splitting GetPrivateRefCountEntry() into a small inline
>     portion (for the cache hit case) and an out-of-line helper for the rest.
>
>     This is helpful for some workloads today, but becomes more important in an
>     upcoming patch that will utilize the private refcount infrastructure to also
>     store whether the buffer is currently locked, as that increases the rate of
>     lookups substantially.

Sounds good based on what I recall of the patch.

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-11-24 20:57                 ` Andres Freund <andres@anarazel.de>
  2025-11-24 21:04                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  3 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-11-24 20:57 UTC (permalink / raw)
  To: Matthias van de Meent <boekewurm+postgres@gmail.com>; +Cc: pgsql-hackers@postgresql.org, Melanie Plageman <melanieplageman@gmail.com>; Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-11-19 21:47:49 -0500, Andres Freund wrote:
> 0001: A straight-up bugfix in lwlock.c - albeit for a bug that seems currently
> 	effectively harmless.

Does anybody have opinions about whether to backpatch this fix? Given that it
has no real consequences I'm mildly inclined not to, but maybe there are cases
where the additional wait list lock cycle matters?

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-24 20:57                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-11-24 21:04                   ` Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 00:17                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Melanie Plageman @ 2025-11-24 21:04 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Mon, Nov 24, 2025 at 3:58 PM Andres Freund <andres@anarazel.de> wrote:
>
> Hi,
>
> On 2025-11-19 21:47:49 -0500, Andres Freund wrote:
> > 0001: A straight-up bugfix in lwlock.c - albeit for a bug that seems currently
> >       effectively harmless.
>
> Does anybody have opinions about whether to backpatch this fix? Given that it
> has no real consequences I'm mildly inclined not to, but maybe there are cases
> where the additional wait list lock cycle matters?

Since it is a mistake, I am mildly in favor of backporting to avoid
confusion for future developers. It's pretty weird that LWLockWakeup()
has to be called again to actually unset LW_FLAG_HAS_WAITERS. But
since it's not really harmful, this is a very mild opinion.

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-24 20:57                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-24 21:04                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
@ 2025-11-25 00:17                     ` Andres Freund <andres@anarazel.de>
  0 siblings, 0 replies; 120+ messages in thread

From: Andres Freund @ 2025-11-25 00:17 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-11-24 16:04:41 -0500, Melanie Plageman wrote:
> On Mon, Nov 24, 2025 at 3:58 PM Andres Freund <andres@anarazel.de> wrote:
> > On 2025-11-19 21:47:49 -0500, Andres Freund wrote:
> > > 0001: A straight-up bugfix in lwlock.c - albeit for a bug that seems currently
> > >       effectively harmless.
> >
> > Does anybody have opinions about whether to backpatch this fix? Given that it
> > has no real consequences I'm mildly inclined not to, but maybe there are cases
> > where the additional wait list lock cycle matters?
> 
> Since it is a mistake, I am mildly in favor of backporting to avoid
> confusion for future developers. It's pretty weird that LWLockWakeup()
> has to be called again to actually unset LW_FLAG_HAS_WAITERS. But
> since it's not really harmful, this is a very mild opinion.

Thanks for chiming in, done.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-11-25 15:44                 ` Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  3 siblings, 1 reply; 120+ messages in thread

From: Melanie Plageman @ 2025-11-25 15:44 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Wed, Nov 19, 2025 at 9:47 PM Andres Freund <andres@anarazel.de> wrote:
>
> 0008: The main change. Implements buffer content locking independently from
> lwlock.c. There's obviously a lot of similarity between lwlock.c code and
> this, but I've not found a good way to reduce the duplication without giving
> up too much.  This patch does immediately introduce share-exclusive as a new
> lock level, mostly because it was too painful to do separately.
>
> 0009+0010+0011: Preparatory work for 0012.
>
> 0012: One of the main goals of this patchset - use the new share-exclusive
> lock level to only allow hint bits to be set while no IO is going on.

Below is a review of 0008-0012
I skipped 0013 and 0014 after seeing "#if 1" in 0013 :)

0008:
-------
> [PATCH v6 08/14] bufmgr: Implement buffer content locks independently
 of lwlocks
...
> This commit unfortunately introduces some code that is very similar to the
> code in lwlock.c, however the code is not equivalent enough to easily merge
> it. The future wins that this commit makes possible seem worth the cost.

It is a truly unfortunate amount of duplication. I tried some
refactoring myself just to convince myself it wasn't a good idea.
However, ISTM there is no reason for the lwlock and buffer locking
implementations to have to stay in sync. So they can diverge as
features are added and perhaps the duplication won't be as conspicuous
in the future.

> As of this commit nothing uses the new share-exclusive lock mode. It will be
>documented and used in a future commit. It seemed too complicated to introduce
> the lock-level in a separate commit.

I would have liked this mentioned earlier in the commit message. Also,
I don't know how I feel about it being "documented" in a future
commit...perhaps just don't say that.

diff --git a/src/include/storage/buf_internals.h
b/src/include/storage/buf_internals.h
index 28519ad2813..0a145d95024 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -32,22 +33,29 @@
 /*
  * Buffer state is a single 64-bit variable where following data is combined.
  *
+ * State of the buffer itself:
  * - 18 bits refcount
  * - 4 bits usage count
  * - 10 bits of flags
  *
+ * State of the content lock:
+ * - 1 bit has_waiter
+ * - 1 bit release_ok
+ * - 1 bit lock state locked

Somewhere you should clearly explain the scenarios in which you still
need to take the buffer header lock with LockBufHdr() vs when you can
just use atomic operations/CAS loops.

Separately, you should clearly explain what BM_LOCKED (I presume what
"lock state locked" refers to) protects.

+/*
+ * Definitions related to buffer content locks
+ */
+#define BM_LOCK_HAS_WAITERS         (UINT64CONST(1) << 63)
+#define BM_LOCK_RELEASE_OK          (UINT64CONST(1) << 62)
+
+#define BM_LOCK_VAL_SHARED          (UINT64CONST(1) << 32)
+#define BM_LOCK_VAL_SHARE_EXCLUSIVE (UINT64CONST(1) << (32 +
MAX_BACKENDS_BITS))
+#define BM_LOCK_VAL_EXCLUSIVE       (UINT64CONST(1) << (32 + 1 +
MAX_BACKENDS_BITS))
+
+#define BM_LOCK_MASK                (((uint64)MAX_BACKENDS << 32) |
BM_LOCK_VAL_SHARE_EXCLUSIVE | BM_LOCK_VAL_EXCLUSIVE)

Not the fault of this patch, but I always found it confusing that the
BM flags were for buffer manager. I always think BM stands for buffer
mapping.

On a more actionable note: you probably want to update the comment in
procnumber.h in your earlier patch to take out the part about how the
limitation could be lifted using a 64 bit state -- since you made the
state 64 bit and didn't lift the limitation (see below for the
specific comment).

/*
 * Note: MAX_BACKENDS_BITS is 18 as that is the space available for buffer
 * refcounts in buf_internals.h.  This limitation could be lifted by using a
 * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
 * currently realistic configurations. Even if that limitation were removed,
 * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
 * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
 * 4*MaxBackends without any overflow check.  We check that the configured
 * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
 */

@@ -268,6 +287,15 @@ BufMappingPartitionLockByIndex(uint32 index)
  * wait_backend_pgprocno and setting flag bit BM_PIN_COUNT_WAITER.  At present,
  * there can be only one such waiter per buffer.
  *
+ * The content of buffers is protected via the buffer content lock,
+ * implemented as part buffer state. Note that the buffer header lock is *not*
+ * used to control access to the data in the buffer! We used to use an LWLock
+ * to implement the content lock, but having a dedicated implementation of
+ * content locks allows to implement some otherwise hard things (e.g.
+ * race-freely checking if AIO is in progress before locking a buffer
+ * exclusively) and makes otherwise impossible optimizations possible
+ * (e.g. unlocking and unpinning a buffer in one atomic operation).
+ *

I don't really like that this includes a description of how it used to
work. It makes the description more confusing when trying to
understand the current state. And comments like that make most sense
when the current state is confusing because of a long annoying
history.

@@ -302,7 +303,24 @@ extern void BufferGetTag(Buffer buffer,
RelFileLocator *rlocator,
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 extern void UnlockBuffers(void);
-extern void LockBuffer(Buffer buffer, BufferLockMode mode);
+extern void UnlockBuffer(Buffer buffer);
+extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
+
+/*
+ * Handling BUFFER_LOCK_UNLOCK in bufmgr.c leads to sufficiently worse branch
+ * prediction to impact performance. Therefore handle that switch here, were
+ * most of the time `mode` will be a constant and thus can be optimized out by
+ * the compiler.

Typo: were -> where

+ */
+static inline void
+LockBuffer(Buffer buffer, BufferLockMode mode)
+{
+    if (mode == BUFFER_LOCK_UNLOCK)
+        UnlockBuffer(buffer);
+    else
+        LockBufferInternal(buffer, mode);
+}
+

Is there a reason to stick with the LockBuffer(buf,
BUFFER_LOCK_UNLOCK) interface here? It seems like a cgood time to
start using UnlockBuffer() which I thought most people find more
intuitive.

 /*
- * Acquire or release the content_lock for the buffer.
+ * Acquire the buffer content lock in the specified mode
+ *
+ * If the lock is not available, sleep until it is.
+ *
+ * Side effect: cancel/die interrupts are held off until lock release.
+ *
+ * This uses almost the same locking approach as lwlock.c's
+ * LWLockAcquire(). See documentation atop of lwlock.c for a more detailed
+ * discussion.
+ *
+ * The reason that this, and most of the other BufferLock* functions, get both
+ * the Buffer and BufferDesc* as parameters, is that looking up one from the
+ * other repeatedly shows up noticeably in profiles.
+ */
+static inline void
+BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)

Either here or over the lock mode enum, I'd spend a few sentences
describing how the buffer lock modes interact -- what conflicts with
what. You spend more time saying how this is like LWLock than
explaining what it actually is.

+/*
+ * Remove ourselves from the waitlist.
+ *
+ * This is used if we queued ourselves because we thought we needed to sleep
+ * but, after further checking, we discovered that we don't actually need to
+ * do so.
+ */
+static void
+BufferLockDequeueSelf(BufferDesc *buf_hdr)
+{
+    bool        on_waitlist;
+
+    LockBufHdr(buf_hdr);
+
...
+    /* XXX: combine with fetch_and above? */
+    UnlockBufHdr(buf_hdr);

I know you just ported this comment from the LWLock implementation but
I can't see how it could be worth the effort. It won't be in the
hottest path as you must satisfy a few conditions to get here.


+ * Wakeup all the lockers that currently have a chance to acquire the lock.
+ */
+static void
+BufferLockWakeup(BufferDesc *buf_hdr, bool unlocked)

This should have a more explanatory comment. You should especially
document the unlocked parameter. I would expect something somewhere in
or above this function that basically says (maybe less verbosely)

- Before waking anything: allow all
- After waking SHARE:
    wake SHARE only
- After waking SHARE_EXCLUSIVE:
    wake SHARE only
    (no more SHARE_EXCLUSIVE)
- After waking EXCLUSIVE:
    wake none

+{
+    bool        new_release_ok;
+    bool        wake_exclusive = unlocked;
+    bool        wake_share_exclusive = true;
+    proclist_head wakeup;
+    proclist_mutable_iter iter;
+
+    proclist_init(&wakeup);
+
+    new_release_ok = true;
+
+    /* lock wait list while collecting backends to wake up */
+    LockBufHdr(buf_hdr);
+
+    proclist_foreach_modify(iter, &buf_hdr->lock_waiters, lwWaitLink)
+    {
+        PGPROC       *waiter = GetPGProcByNumber(iter.cur);
+
+        if (!wake_exclusive && waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+            continue;
+
+        if (!wake_share_exclusive && waiter->lwWaitMode ==
BUFFER_LOCK_SHARE_EXCLUSIVE)
+            continue;
+

It seems there is a difference between LWLockWakeup() and
BufferLockWakeup() if the queue contains a share lock waiter followed
by an exclusive lock waiter.

In LWLockWakeup(), it will wake up the share waiter, set
wokeup_somebody to true, and then not wake up the exclusive lock
waiter because wokeup_somebody is true.

In BufferLockWakeup() when unlocked is true, it will wake up the share
lock waiter and then wake up the exclusive lock waiter because
wake_exclusive and wake_share_exclusive are both still true.

This might be the intended behavior, but it is a difference from
LWLockWakeup(), so I think it is worth documenting why it is okay.


+         * Signal that the process isn't on the wait list anymore. This allows
+         * BufferLockDequeueSelf() to remove itself of the waitlist with a
+         * proclist_delete(), rather than having to check if it has been

I know this comment was ported over, but the "remove itself of the
waitlist" -- the "of" is confusing in the original comment and it is
confusing here.

+ * Compute subtraction from buffer state for a release of a held lock in
+ * `mode`.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places.
+ */
+static inline uint64
+BufferLockReleaseSub(BufferLockMode mode)

I don't understand why this is a separate function even with your comment.

+ * BufferLockHeldByMe - test whether my process holds the content lock in any
+ * mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMe(BufferDesc *buf_hdr)
+{
+    PrivateRefCountEntry *entry =
+        GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+    if (!entry)
+        return false;
+    else
+        return entry->data.lockmode != BUFFER_LOCK_UNLOCK;
+}

Previously, if I called LockBuffer(buf, BUFFER_LOCK_SHARE) again after
calling it once, say in RelationCopyStorageUsingBuffer(), I would trip
an assert when unpinning the buffer about how I needed to have
released the lock.

Now, though, it doesn't trip the assert because we don't track
multiple locks. When I call LockBuffer(UNLOCK), it sets lockmode to
BUFFER_LOCK_UNLOCK and BufferLockHeldByMe() only checks the private
refcount entry. Whereas LWLockHeldByMe() checked held_lwlocks which
had multiple entries and would report that I still held a lock. Now,
if I call LockBuffer(SHARE) twice, I'll only find out later when I
have a hang because someone is trying to get an exclusive lock and the
actual BufferDesc->state still has a lock set.

Maybe you should add an assert to the lock acquisition path that the
prevate ref count entry mode is UNLOCK?

 void
-LockBuffer(Buffer buffer, BufferLockMode mode)
+LockBufferInternal(Buffer buffer, BufferLockMode mode)
 {
-    buf = GetBufferDescriptor(buffer - 1);
+    buf_hdr = GetBufferDescriptor(buffer - 1);

-    if (mode == BUFFER_LOCK_UNLOCK)
-        LWLockRelease(BufferDescriptorGetContentLock(buf));
-    else if (mode == BUFFER_LOCK_SHARE)
-        LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+    if (mode == BUFFER_LOCK_SHARE)
+        BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE);
+    else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+        BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE);
     else if (mode == BUFFER_LOCK_EXCLUSIVE)
-        LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_EXCLUSIVE);
+        BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_EXCLUSIVE);
     else
         elog(ERROR, "unrecognized buffer lock mode: %d", mode);

I presume you've stuck with this if statement structure because that
is what LockBuffer() used. Even though now the BufferLockMode passed
through to BufferLockAcquire is the exact thing you are testing. It
caught my eye and I had to check multiple times if the mode being
passed in is the same. Basically I found it a bit distracting.


@@ -6625,7 +7295,25 @@ ResOwnerReleaseBufferPin(Datum res)
     if (BufferIsLocal(buffer))
         UnpinLocalBufferNoOwner(buffer);
     else
+    {
+        PrivateRefCountEntry *ref;
+
+        ref = GetPrivateRefCountEntry(buffer, false);
+
+        /*
+         * If the buffer was locked at the time of the resowner release,
+         * release the lock now. This should only happen after errors.
+         */
+        if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
+        {
+            BufferDesc *buf = GetBufferDescriptor(buffer - 1);
+
+            HOLD_INTERRUPTS();    /* match the upcoming RESUME_INTERRUPTS */
+            BufferLockUnlock(buffer, buf);
+        }
+
         UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
+    }
 }

 Bit confusing that ResOwnerReleaseBufferBin() now releases locks as well.

0009 -- I didn't look closely at 0009

0010:
--------
> [PATCH v6 10/14] heapam: Use exclusive lock on old page in CLUSTER

> To be able to guarantee that we can set the hint bit, acquire an exclusive
> lock on the old buffer. We need the hint bits to be set as otherwise
> reform_and_rewrite_tuple() -> rewrite_heap_tuple() -> heap_freeze_tuple() will
> get confused.

So, this is an active bug? And what exactly do you mean
heap_freeze_tuple() gets confused? I thought somewhere in there we
would check the clog.

0011:
--------
> [PATCH v6 11/14] heapam: Add batch mode mvcc check and use it in page mode

> 2) We would like to stop setting hint bits while pages are being written
> out. The necessary locking becomes visible for page mode scans if done for
> every tuple. With batching the overhead can be amortized to only happen
> once per page.

I don't understand the above point. What does this patch have to do
with not setting hint bits while pages are being written out (that
happens in the next patch)? And I presume you mean don't set hint bits
on a buffer that is being flushed by someone else -- but it sounds
like you mean not to set hint bits as part of flushing a buffer.

diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 4b0c49f4bb0..ddabd1a3ec3 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -504,42 +504,93 @@ page_collect_tuples(HeapScanDesc scan, Snapshot snapshot,
                     BlockNumber block, int lines,
                     bool all_visible, bool check_serializable)
 {
+    Oid            relid = RelationGetRelid(scan->rs_base.rs_rd);
+#ifdef BATCHMVCC_FEWER_ARGS
+    BatchMVCCState batchmvcc;
+    HeapTupleData *tuples = batchmvcc.tuples;
+    bool       *visible = batchmvcc.visible;
+#else
+    HeapTupleData tuples[MaxHeapTuplesPerPage];
+    bool        visible[MaxHeapTuplesPerPage];
+#endif
     int            ntup = 0;

It's pretty confusing when visible is an output vs input parameter and
who fills it in when. (i.e. it's filled in in page_collect_tuples() if
check_serializable and all_visible are true, otherwise it's filled in
in HeapTupleSatisifiesMVCCBatch())

Personally, I think I'd almost prefer an all-visible and
not-all-visible version of page_collect_tuples() (or helper functions
containing the loop) that separate this. (I haven't tried it though)

BATCHMVCC_FEWER_ARGS definitely doesn't make it any easier to read --
which I assume you are removing.

+        /*
+         * If the page is not all-visible or we need to check serializability,
+         * maintain enough state to be able to refind the tuple efficiently,
+         * without again needing to extract it from the page.
+         */
+        if (!all_visible || check_serializable)
+        {

"enough state" is pretty vague here. Maybe mention what it wouldn't be
valid to do with the tuples array?

/*
* If the page is all visible, these fields won'otherwise wont be
* populated in loop below.
*/

Some spelling issues with "won'" and "wont"

+ * visibility is set in batchmvcc->visible[]. In addition, ->vistuples_dense
+ * is set to contain the offsets of visible tuples.
+ *
+ * Returns the number of visible tuples.
+ */
+int
+HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
+                            int ntups,

I don't really get why this (the patch in general) would be
substantially faster.  You still call HeapTupleSatisfiesMVCC() for
each tuple in a loop. The difference is that you've got pointers into
the array of tuples instead of doing PageGetItem(), then calling
HeapTupleSatisfiesMVCC() for each tuple.


0012:
-------

> [PATCH v6 12/14] Require share-exclusive lock to set hint bits
> To address these issue, this commit changes the rules so that modifications to
> pages are not allowed anymore while holding a share lock. Instead the new

In the commit message, you make it sound like you only change the lock
level for setting the hint bits. But that wouldn't solve any problems
if FlushBuffer() could still happen with a share lock. I would try to
make it clear that you change the lock level both for setting the hint
bits and flushing the buffer.

@@ -77,6 +73,16 @@ gistkillitems(IndexScanDesc scan)
      */
     for (i = 0; i < so->numKilled; i++)
     {
+        if (!killedsomething)
+        {
+            /*
+             * Use hint bit infrastructure to be allowed to modify the page
+             * without holding an exclusive lock.
+             */
+            if (!BufferBeginSetHintBits(buffer))
+                goto unlock;
+        }
+

I don't understand why this is in the loop. Clearly you want to call
BufferBeginSetHintBits() once, but why would you do it in the loop?

In the comment, I might also note that the lock level will be upgraded
as needed or something since we only have a share lock, it is
confusing at first

- * SetHintBits()
+ * To be allowed to set hint bits, SetHintBits() needs to call
+ * BufferBeginSetHintBits(). However, that's not free, and some callsites call
+ * SetHintBits() on many tuples in a row. For those it makes sense to amortize
+ * the cost of BufferBeginSetHintBits(). Additionally it's desirable to defer
+ * the cost of BufferBeginSetHintBits() until a hint bit needs to actually be
+ * set. This enum serves as the necessary state space passed to
+ * SetHintbitsExt().
+ */
+typedef enum SetHintBitsState
+{
+   /* not yet checked if hint bits may be set */
+   SHB_INITIAL,
+   /* failed to get permission to set hint bits, don't check again */
+   SHB_DISABLED,
+   /* allowed to set hint bits */
+   SHB_ENABLED,
+} SetHintBitsState;

I dislike the SHB prefix. Perhaps something involving the word hint?
And should the enum name itself (SetHintBitsState) include the word
"batch"? I know that would make it long. At least the comment should
explain that these are needed when batch setting hint bits.

+SetHintBitsExt(HeapTupleHeader tuple, Buffer buffer,
+               uint16 infomask, TransactionId xid, SetHintBitsState *state)
 {
     if (TransactionIdIsValid(xid))
     {
-        /* NB: xid must be known committed here! */
-        XLogRecPtr    commitLSN = TransactionIdGetCommitLSN(xid);
+        if (BufferIsPermanent(buffer))
+        {

I really wish there was a way to better pull apart the batch and
non-batch cases in a way that could allow the below block to be in a
helper do_set_hint() (or whatever) which SetHintBitsExt() and
SetHintBits() called. And then you inlined BufferSetHintBits16().

I presume you didn't do this because HeapTupleSatisifiesMVCC() for
SNAPSHOT_MVCC calls the non-batch version (and, of course,
HeapTupleSatisifiesVisibility() is the much more common case).

if (TransactionIdIsValid(xid))
{
        if (BufferIsPermanent(buffer))
        {
                /* NB: xid must be known committed here! */
                XLogRecPtr    commitLSN = TransactionIdGetCommitLSN(xid);

                if (XLogNeedsFlush(commitLSN) &&
                        BufferGetLSNAtomic(buffer) < commitLSN)
                {
                        /* not flushed and no LSN interlock, so don't
set hint */
                        return; false;
                }
        }
}

Separately, I was thinking, should we assert here about having the
right lock type?

 * It is only safe to set a transaction-committed hint bit if we know the
 * transaction's commit record is guaranteed to be flushed to disk before the
 * buffer, or if the table is temporary or unlogged and will be obliterated by
 * a crash anyway.  We cannot change the LSN of the page here, because we may
 * hold only a share lock on the buffer, so we can only use the LSN to
 * interlock this if the buffer's LSN already is newer than the commit LSN;
 * otherwise we have to just refrain from setting the hint bit until some
 * future re-examination of the tuple.
 *

Should this say "we may hold only a share exclusive lock on the
buffer". Also what is "this" in "only use the LSN to interlock this"?

@@ -1628,6 +1701,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot,
Buffer buffer,
+    if (state == SHB_ENABLED)
+        BufferFinishSetHintBits(buffer, true, true);
+
     return nvis;
 }

I wondered if it would be more natural for BufferBeginSetHintBits()
and BufferFinishSetHintBits() to set SHB_INITIAL and SHB_DISABLED
instead of having callers do it. But, I guess you don't do this
because of gist and hash indexes using this for doing their own
modifications. It seems like it also would help make it clear that
BufferBegin and BufferFinish are for batch mode.

I can't help but feel a bit of awkwardness in the whole API between
this and the way you've supported non-batch mode in SetHintBitsExt().
But it's easy for me to criticize without providing concrete ideas of
how to reorganize it.

+static inline bool
+SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr,
uint64 *locksta>
+{
+   uint64      old_state;
+   PrivateRefCountEntry *ref;
+   BufferLockMode mode;

These functions probably ought to have comments. And I wonder if you
should say anything (or even assert anything) about how IO better not
be in progress (though it is enforced by the locks).

+        /*
+         * Can't upgrade if somebody else holds the lock in exlusive or
+         * share-exclusive mode.
+         */

typo: exlusive -> exclusive

-4. It is considered OK to update tuple commit status bits (ie, OR the
-values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
+4. Non-critical information on a page ("hint bits") may be modified while
+holding only a share-exclusive lock and pin on the page. To do so in cases
+where only a share lock is already held, use BufferBeginSetHintBits() &
+BufferFinishSetHintBits() (if multiple hint bits are to be set) or
+BufferSetHintBits16() (if a single hit bit is set).

typo: single hit -> single hint
Also, you should probably say that you use BufferBeginSetHintBits() to
actually upgrade the lock.

I also don't see anywhere in the README where flushing a buffer is
mentioned -- and what lock level is needed for that in different
situations. It kind of feels like it is worth mentioning.

-   if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
-       (BM_DIRTY | BM_JUST_DIRTIED))
+   if (unlikely((lockstate & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+                (BM_DIRTY | BM_JUST_DIRTIED)))

I don't quite understand why you do the atomic read in the call to
MarkSharedBufferDirtyHint() form MarkBufferDirtyHint() instead of at
the top of MarkSharedBufferDirtyHint() -- and then you wouldn't have
to pass the lockstate as a parameter, right?

--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -904,14 +904,22 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
    max_avail = fsm_get_max_avail(page);
    /*
-    * Reset the next slot pointer. This encourages the use of low-numbered
-    * pages, increasing the chances that a later vacuum can truncate the
-    * relation.  We don't bother with a lock here, nor with marking the page
-    * dirty if it wasn't already, since this is just a hint.
+    * Try to reset the next slot pointer. This encourages the use of
+    * low-numbered pages, increasing the chances that a later vacuum can
+    * truncate the relation.  We don't bother with a lock here, nor with
+    * marking the page dirty if it wasn't already, since this is just a hint.
+    *
+    * To be allowed to update the page without an exclusive lock, we have to
+    * use the hint bit infrastructure.
     */

What the heck? This didn't even take a share lock before...

diff --git a/src/backend/storage/freespace/fsmpage.c
b/src/backend/storage/freesp>
index 66a5c80b5a6..a59696b6484 100644
--- a/src/backend/storage/freespace/fsmpage.c
+++ b/src/backend/storage/freespace/fsmpage.c
@@ -298,9 +298,18 @@ restart:
     * lock and get a garbled next pointer every now and then, than take the
     * concurrency hit of an exclusive lock.
     *
+    * Without an exclusive lock, we need to use the hint bit infrastructure
+    * to be allowed to modify the page.
+    *

Is the sentence above this still correct?

/*
* Update the next-target pointer. Note that we do this even if we're only
* holding a shared lock, on the grounds that it's better to use a shared
* lock and get a garbled next pointer every now and then, than take the
* concurrency hit of an exclusive lock.

We appear to avoid the garbling now?

In general on 0012, I didn't spend much time checking if you caught
all the places where we mention our hint bit hackery (only taking the
share lock). But those can always be caught later as we inevitably
encounter them.

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
@ 2025-11-25 16:54                   ` Andres Freund <andres@anarazel.de>
  2025-11-25 20:02                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 2 replies; 120+ messages in thread

From: Andres Freund @ 2025-11-25 16:54 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

Thanks a lot for that detailed review!  A few questions and comments, before I
try to address the comments in the next version.


On 2025-11-25 10:44:31 -0500, Melanie Plageman wrote:
> On Wed, Nov 19, 2025 at 9:47 PM Andres Freund <andres@anarazel.de> wrote:
> >
> > 0008: The main change. Implements buffer content locking independently from
> > lwlock.c. There's obviously a lot of similarity between lwlock.c code and
> > this, but I've not found a good way to reduce the duplication without giving
> > up too much.  This patch does immediately introduce share-exclusive as a new
> > lock level, mostly because it was too painful to do separately.
> >
> > 0009+0010+0011: Preparatory work for 0012.
> >
> > 0012: One of the main goals of this patchset - use the new share-exclusive
> > lock level to only allow hint bits to be set while no IO is going on.
> 
> Below is a review of 0008-0012
> I skipped 0013 and 0014 after seeing "#if 1" in 0013 :)

I left that in there to make the comparison easier. But clearly that's not to
be committed...


> 0008:
> -------
> > [PATCH v6 08/14] bufmgr: Implement buffer content locks independently
>  of lwlocks
> ...
> > This commit unfortunately introduces some code that is very similar to the
> > code in lwlock.c, however the code is not equivalent enough to easily merge
> > it. The future wins that this commit makes possible seem worth the cost.
> 
> It is a truly unfortunate amount of duplication. I tried some
> refactoring myself just to convince myself it wasn't a good idea.

Without success, I guess?


> However, ISTM there is no reason for the lwlock and buffer locking
> implementations to have to stay in sync. So they can diverge as
> features are added and perhaps the duplication won't be as conspicuous
> in the future.

Right - I'd expect further divergence.


> > As of this commit nothing uses the new share-exclusive lock mode. It will be
> > documented and used in a future commit. It seemed too complicated to introduce
> > the lock-level in a separate commit.
> 
> I would have liked this mentioned earlier in the commit message.

Hm, ok.  I could split share-exclusive out again, but it was somewhat painful,
because it doesn't just lead to adding code, but to changing code.


> Also, I don't know how I feel about it being "documented" in a future
> commit...perhaps just don't say that.

Hm, why?


> diff --git a/src/include/storage/buf_internals.h
> b/src/include/storage/buf_internals.h
> index 28519ad2813..0a145d95024 100644
> --- a/src/include/storage/buf_internals.h
> +++ b/src/include/storage/buf_internals.h
> @@ -32,22 +33,29 @@
>  /*
>   * Buffer state is a single 64-bit variable where following data is combined.
>   *
> + * State of the buffer itself:
>   * - 18 bits refcount
>   * - 4 bits usage count
>   * - 10 bits of flags
>   *
> + * State of the content lock:
> + * - 1 bit has_waiter
> + * - 1 bit release_ok
> + * - 1 bit lock state locked
> 
> Somewhere you should clearly explain the scenarios in which you still
> need to take the buffer header lock with LockBufHdr() vs when you can
> just use atomic operations/CAS loops.

There is an explanation of that, or at least my attempt at it ;). See the
dcoumentation for BufferDesc.


> @@ -268,6 +287,15 @@ BufMappingPartitionLockByIndex(uint32 index)
>   * wait_backend_pgprocno and setting flag bit BM_PIN_COUNT_WAITER.  At present,
>   * there can be only one such waiter per buffer.
>   *
> + * The content of buffers is protected via the buffer content lock,
> + * implemented as part buffer state. Note that the buffer header lock is *not*
> + * used to control access to the data in the buffer! We used to use an LWLock
> + * to implement the content lock, but having a dedicated implementation of
> + * content locks allows to implement some otherwise hard things (e.g.
> + * race-freely checking if AIO is in progress before locking a buffer
> + * exclusively) and makes otherwise impossible optimizations possible
> + * (e.g. unlocking and unpinning a buffer in one atomic operation).
> + *
> 
> I don't really like that this includes a description of how it used to
> work. It makes the description more confusing when trying to
> understand the current state. And comments like that make most sense
> when the current state is confusing because of a long annoying
> history.

I generally agree with that - however, in this case it seemed like folks, in
the future, might actually be wondering why this isn't using lwlocks.


> + */
> +static inline void
> +LockBuffer(Buffer buffer, BufferLockMode mode)
> +{
> +    if (mode == BUFFER_LOCK_UNLOCK)
> +        UnlockBuffer(buffer);
> +    else
> +        LockBufferInternal(buffer, mode);
> +}
> +
> 
> Is there a reason to stick with the LockBuffer(buf,
> BUFFER_LOCK_UNLOCK) interface here? It seems like a cgood time to
> start using UnlockBuffer() which I thought most people find more
> intuitive.

There are like 200 uses of BUFFER_LOCK_UNLOCK|GIN_UNLOCK|GIST_UNLOCK. And
probably a lot of external ones.  I'm not against using UnlockBuffer()
directly in the future, but changing all that code as part of this change
doesn't quite seem to make sense.


>  /*
> - * Acquire or release the content_lock for the buffer.
> + * Acquire the buffer content lock in the specified mode
> + *
> + * If the lock is not available, sleep until it is.
> + *
> + * Side effect: cancel/die interrupts are held off until lock release.
> + *
> + * This uses almost the same locking approach as lwlock.c's
> + * LWLockAcquire(). See documentation atop of lwlock.c for a more detailed
> + * discussion.
> + *
> + * The reason that this, and most of the other BufferLock* functions, get both
> + * the Buffer and BufferDesc* as parameters, is that looking up one from the
> + * other repeatedly shows up noticeably in profiles.
> + */
> +static inline void
> +BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
> 
> Either here or over the lock mode enum, I'd spend a few sentences
> describing how the buffer lock modes interact -- what conflicts with
> what. You spend more time saying how this is like LWLock than
> explaining what it actually is.

Historically they're just explained in the README. But yea, it makes sense to
add to the enum. I like adding it to the enum better than to the function, as
there are multiple ways to acquire a lock.


> +{
> +    bool        new_release_ok;
> +    bool        wake_exclusive = unlocked;
> +    bool        wake_share_exclusive = true;
> +    proclist_head wakeup;
> +    proclist_mutable_iter iter;
> +
> +    proclist_init(&wakeup);
> +
> +    new_release_ok = true;
> +
> +    /* lock wait list while collecting backends to wake up */
> +    LockBufHdr(buf_hdr);
> +
> +    proclist_foreach_modify(iter, &buf_hdr->lock_waiters, lwWaitLink)
> +    {
> +        PGPROC       *waiter = GetPGProcByNumber(iter.cur);
> +
> +        if (!wake_exclusive && waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
> +            continue;
> +
> +        if (!wake_share_exclusive && waiter->lwWaitMode ==
> BUFFER_LOCK_SHARE_EXCLUSIVE)
> +            continue;
> +
> 
> It seems there is a difference between LWLockWakeup() and
> BufferLockWakeup() if the queue contains a share lock waiter followed
> by an exclusive lock waiter.
> 
> In LWLockWakeup(), it will wake up the share waiter, set
> wokeup_somebody to true, and then not wake up the exclusive lock
> waiter because wokeup_somebody is true.
> 
> In BufferLockWakeup() when unlocked is true, it will wake up the share
> lock waiter and then wake up the exclusive lock waiter because
> wake_exclusive and wake_share_exclusive are both still true.
> 
> This might be the intended behavior, but it is a difference from
> LWLockWakeup(), so I think it is worth documenting why it is okay.

Yea, that doesn't look quite right. I think I just whacked the code around too
much and somewhere lost that branch.


> 
> +         * Signal that the process isn't on the wait list anymore. This allows
> +         * BufferLockDequeueSelf() to remove itself of the waitlist with a
> +         * proclist_delete(), rather than having to check if it has been
> 
> I know this comment was ported over, but the "remove itself of the
> waitlist" -- the "of" is confusing in the original comment and it is
> confusing here.

As in s/of/from/?


> + * Compute subtraction from buffer state for a release of a held lock in
> + * `mode`.
> + *
> + * This is separated from BufferLockUnlock() as we want to combine the lock
> + * release with other atomic operations when possible, leading to the lock
> + * release being done in multiple places.
> + */
> +static inline uint64
> +BufferLockReleaseSub(BufferLockMode mode)
> 
> I don't understand why this is a separate function even with your comment.

Because there are operations where we want to unlock the buffer as well as do
something else. E.g. in 0013, UnlockReleaseBuffer() we want to unlock the
buffer and decrease the refcount in one atomic operation. For that we need to
know what to subtract from the state variable for the lock portion - hence
BufferLockReleaseSub().
 

> + * BufferLockHeldByMe - test whether my process holds the content lock in any
> + * mode
> + *
> + * This is meant as debug support only.
> + */
> +static bool
> +BufferLockHeldByMe(BufferDesc *buf_hdr)
> +{
> +    PrivateRefCountEntry *entry =
> +        GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
> +
> +    if (!entry)
> +        return false;
> +    else
> +        return entry->data.lockmode != BUFFER_LOCK_UNLOCK;
> +}
> 
> Previously, if I called LockBuffer(buf, BUFFER_LOCK_SHARE) again after
> calling it once, say in RelationCopyStorageUsingBuffer(), I would trip
> an assert when unpinning the buffer about how I needed to have
> released the lock.

Right, that's important to keep that way.


> Now, though, it doesn't trip the assert because we don't track
> multiple locks. When I call LockBuffer(UNLOCK), it sets lockmode to
> BUFFER_LOCK_UNLOCK and BufferLockHeldByMe() only checks the private
> refcount entry.

That part seems right.


> Whereas LWLockHeldByMe() checked held_lwlocks which had multiple entries and
> would report that I still held a lock.

It'd be much better if we had detected that redundant lock acquisition, that's
not legal...


> Now, if I call LockBuffer(SHARE) twice, I'll only find out later when I have
> a hang because someone is trying to get an exclusive lock and the actual
> BufferDesc->state still has a lock set.
> 
> Maybe you should add an assert to the lock acquisition path that the
> prevate ref count entry mode is UNLOCK?

Yes, we should...

It'd be nice if we had a decent way to test things that we except to
crash. Like repeated buffer lock acquisitions...

>  void
> -LockBuffer(Buffer buffer, BufferLockMode mode)
> +LockBufferInternal(Buffer buffer, BufferLockMode mode)
>  {
> -    buf = GetBufferDescriptor(buffer - 1);
> +    buf_hdr = GetBufferDescriptor(buffer - 1);
> 
> -    if (mode == BUFFER_LOCK_UNLOCK)
> -        LWLockRelease(BufferDescriptorGetContentLock(buf));
> -    else if (mode == BUFFER_LOCK_SHARE)
> -        LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
> +    if (mode == BUFFER_LOCK_SHARE)
> +        BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE);
> +    else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
> +        BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE);
>      else if (mode == BUFFER_LOCK_EXCLUSIVE)
> -        LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_EXCLUSIVE);
> +        BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_EXCLUSIVE);
>      else
>          elog(ERROR, "unrecognized buffer lock mode: %d", mode);
> 
> I presume you've stuck with this if statement structure because that
> is what LockBuffer() used.
>
> Even though now the BufferLockMode passed
> through to BufferLockAcquire is the exact thing you are testing.

Ah. Using explicit constants allows the compiler to do constant propagation
(so e.g. BufferLockAttempt() doesn't have runtime branches on mode), which
turns out to generate more efficient code.  I'll add a comment to that effect.


> @@ -6625,7 +7295,25 @@ ResOwnerReleaseBufferPin(Datum res)
>      if (BufferIsLocal(buffer))
>          UnpinLocalBufferNoOwner(buffer);
>      else
> +    {
> +        PrivateRefCountEntry *ref;
> +
> +        ref = GetPrivateRefCountEntry(buffer, false);
> +
> +        /*
> +         * If the buffer was locked at the time of the resowner release,
> +         * release the lock now. This should only happen after errors.
> +         */
> +        if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
> +        {
> +            BufferDesc *buf = GetBufferDescriptor(buffer - 1);
> +
> +            HOLD_INTERRUPTS();    /* match the upcoming RESUME_INTERRUPTS */
> +            BufferLockUnlock(buffer, buf);
> +        }
> +
>          UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
> +    }
>  }
> 
>  Bit confusing that ResOwnerReleaseBufferBin() now releases locks as well.

Do you have a better suggestion? I'll add a comment that makes that explicit,
but other than that I don't have a great idea. Renaming the whole buffer pin
mechanism seems pretty noisy.


> 0009 -- I didn't look closely at 0009

That's hopefully just a boring code move...


> 0010:
> --------
> > [PATCH v6 10/14] heapam: Use exclusive lock on old page in CLUSTER
> 
> > To be able to guarantee that we can set the hint bit, acquire an exclusive
> > lock on the old buffer. We need the hint bits to be set as otherwise
> > reform_and_rewrite_tuple() -> rewrite_heap_tuple() -> heap_freeze_tuple() will
> > get confused.
> 
> So, this is an active bug?

I don't think so - we have an AEL on the relation at that point, so nobody
else can access the buffer, aside from checkpointer writing it out. Which
doesn't currrently block hint bits from being set.


> And what exactly do you mean heap_freeze_tuple() gets confused? I thought
> somewhere in there we would check the clog.

heap_freeze_tuple()->heap_prepare_freeze_tuple() assumes that:
 * It is assumed that the caller has checked the tuple with
 * HeapTupleSatisfiesVacuum() and determined that it is not HEAPTUPLE_DEAD
 * (else we should be removing the tuple, not freezing it).

In the new world, if HTSV were not allowed to set the hint bit, e.g. because
the page is in the process of being written out, there wouldn't be a guarantee
that hints were set.  Which then leads to confusion, because some of the code
assumes that hint bits were already set.  TBH, I don't quite remember the
precise details, it's been a while (this was already part of an earlier
attempt at dealing with the hint bit stuff).  I'll try to reconstruct.


> 0011:
> --------
> > [PATCH v6 11/14] heapam: Add batch mode mvcc check and use it in page mode
> 
> > 2) We would like to stop setting hint bits while pages are being written
> > out. The necessary locking becomes visible for page mode scans if done for
> > every tuple. With batching the overhead can be amortized to only happen
> > once per page.
> 
> I don't understand the above point. What does this patch have to do
> with not setting hint bits while pages are being written out (that
> happens in the next patch)?

It's just an explanation for why we want to use batch mode. Without the
changed locking model the reason for that is a lot less clear.

> And I presume you mean don't set hint bits on a buffer that is being flushed
> by someone else -- but it sounds like you mean not to set hint bits as part
> of flushing a buffer.

Right, I do mean the former.


> diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
> index 4b0c49f4bb0..ddabd1a3ec3 100644
> --- a/src/backend/access/heap/heapam.c
> +++ b/src/backend/access/heap/heapam.c
> @@ -504,42 +504,93 @@ page_collect_tuples(HeapScanDesc scan, Snapshot snapshot,
>                      BlockNumber block, int lines,
>                      bool all_visible, bool check_serializable)
>  {
> +    Oid            relid = RelationGetRelid(scan->rs_base.rs_rd);
> +#ifdef BATCHMVCC_FEWER_ARGS
> +    BatchMVCCState batchmvcc;
> +    HeapTupleData *tuples = batchmvcc.tuples;
> +    bool       *visible = batchmvcc.visible;
> +#else
> +    HeapTupleData tuples[MaxHeapTuplesPerPage];
> +    bool        visible[MaxHeapTuplesPerPage];
> +#endif
>      int            ntup = 0;
> 
> It's pretty confusing when visible is an output vs input parameter and
> who fills it in when. (i.e. it's filled in in page_collect_tuples() if
> check_serializable and all_visible are true, otherwise it's filled in
> in HeapTupleSatisifiesMVCCBatch())

I don't really see a good alternative.


> Personally, I think I'd almost prefer an all-visible and
> not-all-visible version of page_collect_tuples() (or helper functions
> containing the loop) that separate this. (I haven't tried it though)

I have a hard time seeing how that comes out better. Where would you make the
switch between the different functions and how would that lead to easier to
understand code?


> BATCHMVCC_FEWER_ARGS definitely doesn't make it any easier to read --
> which I assume you are removing.

Yea, I only left it in so others perhaps can reproduce the performance effect
of needing BATCHMVCC_FEWER_ARGS.


> +        /*
> +         * If the page is not all-visible or we need to check serializability,
> +         * maintain enough state to be able to refind the tuple efficiently,
> +         * without again needing to extract it from the page.
> +         */
> +        if (!all_visible || check_serializable)
> +        {
> 
> "enough state" is pretty vague here.

I don't really follow. Wouldn't going into more detail just restate the code
in a comment?


> Maybe mention what it wouldn't be valid to do with the tuples array?

Hm? The tuples array *is* that state?



> + * visibility is set in batchmvcc->visible[]. In addition, ->vistuples_dense
> + * is set to contain the offsets of visible tuples.
> + *
> + * Returns the number of visible tuples.
> + */
> +int
> +HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
> +                            int ntups,
> 
> I don't really get why this (the patch in general) would be
> substantially faster.  You still call HeapTupleSatisfiesMVCC() for
> each tuple in a loop. The difference is that you've got pointers into
> the array of tuples instead of doing PageGetItem(), then calling
> HeapTupleSatisfiesMVCC() for each tuple.

There are a few reasons:

1) Most importantly, in the batched world, we only need to check if we can set
   hint bits once. That check is far from free. We also only need to mark the
   buffer dirty once.

2) HeapTupleSatisfiesMVCCBatch() can inline HeapTupleSatisfiesMVCC() and avoid
   some redundant work, e.g. it doesn't have to set up a stack frame each
   time.

3) A loop over the tuples in heapam.c needs an external function call to
   HeapTupleSatisfiesMVCC(), as that's in heapam_visibility.c. That function
   call quickly shows up.


> 
> 0012:
> -------
> 
> > [PATCH v6 12/14] Require share-exclusive lock to set hint bits
> > To address these issue, this commit changes the rules so that modifications to
> > pages are not allowed anymore while holding a share lock. Instead the new
> 
> In the commit message, you make it sound like you only change the lock
> level for setting the hint bits. But that wouldn't solve any problems
> if FlushBuffer() could still happen with a share lock. I would try to
> make it clear that you change the lock level both for setting the hint
> bits and flushing the buffer.

Good point.


> @@ -77,6 +73,16 @@ gistkillitems(IndexScanDesc scan)
>       */
>      for (i = 0; i < so->numKilled; i++)
>      {
> +        if (!killedsomething)
> +        {
> +            /*
> +             * Use hint bit infrastructure to be allowed to modify the page
> +             * without holding an exclusive lock.
> +             */
> +            if (!BufferBeginSetHintBits(buffer))
> +                goto unlock;
> +        }
> +
> 
> I don't understand why this is in the loop. Clearly you want to call
> BufferBeginSetHintBits() once, but why would you do it in the loop?

Why would we want to continue if we can't set hint bits?


> In the comment, I might also note that the lock level will be upgraded
> as needed or something since we only have a share lock, it is
> confusing at first

I don't think copying the way this works into all the callers of
BufferBeginSetHintBits() is a good idea. That just makes it harder to adjust
going down the road.


> - * SetHintBits()
> + * To be allowed to set hint bits, SetHintBits() needs to call
> + * BufferBeginSetHintBits(). However, that's not free, and some callsites call
> + * SetHintBits() on many tuples in a row. For those it makes sense to amortize
> + * the cost of BufferBeginSetHintBits(). Additionally it's desirable to defer
> + * the cost of BufferBeginSetHintBits() until a hint bit needs to actually be
> + * set. This enum serves as the necessary state space passed to
> + * SetHintbitsExt().
> + */
> +typedef enum SetHintBitsState
> +{
> +   /* not yet checked if hint bits may be set */
> +   SHB_INITIAL,
> +   /* failed to get permission to set hint bits, don't check again */
> +   SHB_DISABLED,
> +   /* allowed to set hint bits */
> +   SHB_ENABLED,
> +} SetHintBitsState;
> 
> I dislike the SHB prefix. Perhaps something involving the word hint?

Shrug. It's a very locally used enum, I don't think it matters terribly.


> And should the enum name itself (SetHintBitsState) include the word
> "batch"? I know that would make it

> long. At least the comment should explain that these are needed when batch
> setting hint bits.

Isn't that what the comment does, explaining that we want to amortize the cost
of BufferBeginSetHintBits()?


> +SetHintBitsExt(HeapTupleHeader tuple, Buffer buffer,
> +               uint16 infomask, TransactionId xid, SetHintBitsState *state)
>  {
>      if (TransactionIdIsValid(xid))
>      {
> -        /* NB: xid must be known committed here! */
> -        XLogRecPtr    commitLSN = TransactionIdGetCommitLSN(xid);
> +        if (BufferIsPermanent(buffer))
> +        {
> 
> I really wish there was a way to better pull apart the batch and
> non-batch cases in a way that could allow the below block to be in a
> helper do_set_hint() (or whatever) which SetHintBitsExt() and
> SetHintBits() called.

I couldn't see a way that didn't lead to substantially more code duplication.


> And then you inlined BufferSetHintBits16().

Hm?


> I presume you didn't do this because HeapTupleSatisifiesMVCC() for
> SNAPSHOT_MVCC calls the non-batch version (and, of course,
> HeapTupleSatisifiesVisibility() is the much more common case).
> 
> if (TransactionIdIsValid(xid))
> {
>         if (BufferIsPermanent(buffer))
>         {
>                 /* NB: xid must be known committed here! */
>                 XLogRecPtr    commitLSN = TransactionIdGetCommitLSN(xid);
> 
>                 if (XLogNeedsFlush(commitLSN) &&
>                         BufferGetLSNAtomic(buffer) < commitLSN)
>                 {
>                         /* not flushed and no LSN interlock, so don't
> set hint */
>                         return; false;
>                 }
>         }
> }
> 
> Separately, I was thinking, should we assert here about having the
> right lock type?

Not sure I get what assert of what locktype where?


>  * It is only safe to set a transaction-committed hint bit if we know the
>  * transaction's commit record is guaranteed to be flushed to disk before the
>  * buffer, or if the table is temporary or unlogged and will be obliterated by
>  * a crash anyway.  We cannot change the LSN of the page here, because we may
>  * hold only a share lock on the buffer, so we can only use the LSN to
>  * interlock this if the buffer's LSN already is newer than the commit LSN;
>  * otherwise we have to just refrain from setting the hint bit until some
>  * future re-examination of the tuple.
>  *
> 
> Should this say "we may hold only a share exclusive lock on the
> buffer". Also what is "this" in "only use the LSN to interlock this"?

That's a pre-existing comment, right?


> @@ -1628,6 +1701,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot,
> Buffer buffer,
> +    if (state == SHB_ENABLED)
> +        BufferFinishSetHintBits(buffer, true, true);
> +
>      return nvis;
>  }
> 
> I wondered if it would be more natural for BufferBeginSetHintBits()
> and BufferFinishSetHintBits() to set SHB_INITIAL and SHB_DISABLED
> instead of having callers do it. But, I guess you don't do this
> because of gist and hash indexes using this for doing their own
> modifications.

I thought about it. But yea, the different callers seemed to make that not
really useful. It'd also mean that we'd do an external function call for every
tuple, which I think would be prohibitively expensive.



> --- a/src/backend/storage/freespace/freespace.c
> +++ b/src/backend/storage/freespace/freespace.c
> @@ -904,14 +904,22 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
>     max_avail = fsm_get_max_avail(page);
>     /*
> -    * Reset the next slot pointer. This encourages the use of low-numbered
> -    * pages, increasing the chances that a later vacuum can truncate the
> -    * relation.  We don't bother with a lock here, nor with marking the page
> -    * dirty if it wasn't already, since this is just a hint.
> +    * Try to reset the next slot pointer. This encourages the use of
> +    * low-numbered pages, increasing the chances that a later vacuum can
> +    * truncate the relation.  We don't bother with a lock here, nor with
> +    * marking the page dirty if it wasn't already, since this is just a hint.
> +    *
> +    * To be allowed to update the page without an exclusive lock, we have to
> +    * use the hint bit infrastructure.
>      */
> 
> What the heck? This didn't even take a share lock before...

It is insane. I do not understand how anybody thought this was ok.  I think I
probably should split this out into a separate commit, since it's so insane.


> diff --git a/src/backend/storage/freespace/fsmpage.c
> b/src/backend/storage/freesp>
> index 66a5c80b5a6..a59696b6484 100644
> --- a/src/backend/storage/freespace/fsmpage.c
> +++ b/src/backend/storage/freespace/fsmpage.c
> @@ -298,9 +298,18 @@ restart:
>      * lock and get a garbled next pointer every now and then, than take the
>      * concurrency hit of an exclusive lock.
>      *
> +    * Without an exclusive lock, we need to use the hint bit infrastructure
> +    * to be allowed to modify the page.
> +    *
> 
> Is the sentence above this still correct?

Seems ok enough to me.


> /*
> * Update the next-target pointer. Note that we do this even if we're only
> * holding a shared lock, on the grounds that it's better to use a shared
> * lock and get a garbled next pointer every now and then, than take the
> * concurrency hit of an exclusive lock.
> 
> We appear to avoid the garbling now?

I don't think so. Two backends concurrently can do fsm_search_avail() and
one backend might set a hint to a page that is already used up by the other
one. At least I think so?


> In general on 0012, I didn't spend much time checking if you caught
> all the places where we mention our hint bit hackery (only taking the
> share lock). But those can always be caught later as we inevitably
> encounter them.

There's definitely some more search needed as part of polishing that
commit. But I think it's pretty much inevitable that we'll miss some comment
somewhere :(

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-11-25 20:02                     ` Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 20:46                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  1 sibling, 1 reply; 120+ messages in thread

From: Melanie Plageman @ 2025-11-25 20:02 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Tue, Nov 25, 2025 at 11:54 AM Andres Freund <andres@anarazel.de> wrote:
>
> > I skipped 0013 and 0014 after seeing "#if 1" in 0013 :)
>
> I left that in there to make the comparison easier. But clearly that's not to
> be committed...

Eh, I mostly just ran out of steam.

> > > [PATCH v6 08/14] bufmgr: Implement buffer content locks independently
> >  of lwlocks
> > ...
> > > This commit unfortunately introduces some code that is very similar to the
> > > code in lwlock.c, however the code is not equivalent enough to easily merge
> > > it. The future wins that this commit makes possible seem worth the cost.
> >
> > It is a truly unfortunate amount of duplication. I tried some
> > refactoring myself just to convince myself it wasn't a good idea.
>
> Without success, I guess?

Nothing that seemed better...and not worse :)

> > > As of this commit nothing uses the new share-exclusive lock mode. It will be
> > > documented and used in a future commit. It seemed too complicated to introduce
> > > the lock-level in a separate commit.
> >
> > I would have liked this mentioned earlier in the commit message.
>
> Hm, ok.  I could split share-exclusive out again, but it was somewhat painful,
> because it doesn't just lead to adding code, but to changing code.

I don't think you need to introduce share-exclusive in a separate
commit. I just mean it would be good to mention in the commit message
before you get into so many other details that this commit doesn't use
the share-exclusive level yet.

> > Also, I don't know how I feel about it being "documented" in a future
> > commit...perhaps just don't say that.
>
> Hm, why?

I don't really see you documenting share-exclusive lock semantics in a
later commit. Documentation for what the lock level is should be in
the same commit that introduces it -- which I think you mostly have
done. I feel like it is weird to say you will document that lock level
later when 1) you don't really do that and 2) this commit mostly
(besides some requests I had for elaboration) already does that.

> > diff --git a/src/include/storage/buf_internals.h
> > b/src/include/storage/buf_internals.h
> > index 28519ad2813..0a145d95024 100644
> > --- a/src/include/storage/buf_internals.h
> > +++ b/src/include/storage/buf_internals.h
> > @@ -32,22 +33,29 @@
> >  /*
> >   * Buffer state is a single 64-bit variable where following data is combined.
> >   *
> > + * State of the buffer itself:
> >   * - 18 bits refcount
> >   * - 4 bits usage count
> >   * - 10 bits of flags
> >   *
> > + * State of the content lock:
> > + * - 1 bit has_waiter
> > + * - 1 bit release_ok
> > + * - 1 bit lock state locked
> >
> > Somewhere you should clearly explain the scenarios in which you still
> > need to take the buffer header lock with LockBufHdr() vs when you can
> > just use atomic operations/CAS loops.
>
> There is an explanation of that, or at least my attempt at it ;). See the
> dcoumentation for BufferDesc.

Hmm. I suppose. I was imagining something above the state member since
you have added to it, but okay.

Separately, I'll admit I don't quite understand when I have to use
LockBufHdr() and when I should use a CAS loop to update the
bufferdesc->state.

> >
> > +         * Signal that the process isn't on the wait list anymore. This allows
> > +         * BufferLockDequeueSelf() to remove itself of the waitlist with a
> > +         * proclist_delete(), rather than having to check if it has been
> >
> > I know this comment was ported over, but the "remove itself of the
> > waitlist" -- the "of" is confusing in the original comment and it is
> > confusing here.
>
> As in s/of/from/?

Yes, I wasn't sure if it meant "from" or "off" -- but I believe "from"
is more grammatically correct.

> > +static inline uint64
> > +BufferLockReleaseSub(BufferLockMode mode)
> >
> > I don't understand why this is a separate function even with your comment.
>
> Because there are operations where we want to unlock the buffer as well as do
> something else. E.g. in 0013, UnlockReleaseBuffer() we want to unlock the
> buffer and decrease the refcount in one atomic operation. For that we need to
> know what to subtract from the state variable for the lock portion - hence
> BufferLockReleaseSub().

Hmm. Okay well maybe save it for 0013? I don't care that much, though.

> > Maybe you should add an assert to the lock acquisition path that the
> > prevate ref count entry mode is UNLOCK?
>
> Yes, we should...
>
> It'd be nice if we had a decent way to test things that we except to
> crash. Like repeated buffer lock acquisitions...

Indeed.

> > @@ -6625,7 +7295,25 @@ ResOwnerReleaseBufferPin(Datum res)
> >      if (BufferIsLocal(buffer))
> >          UnpinLocalBufferNoOwner(buffer);
> >      else
> > +    {
> > +        PrivateRefCountEntry *ref;
> > +
> > +        ref = GetPrivateRefCountEntry(buffer, false);
> > +
> > +        /*
> > +         * If the buffer was locked at the time of the resowner release,
> > +         * release the lock now. This should only happen after errors.
> > +         */
> > +        if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
> > +        {
> > +            BufferDesc *buf = GetBufferDescriptor(buffer - 1);
> > +
> > +            HOLD_INTERRUPTS();    /* match the upcoming RESUME_INTERRUPTS */
> > +            BufferLockUnlock(buffer, buf);
> > +        }
> > +
> >          UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
> > +    }
> >  }
> >
> >  Bit confusing that ResOwnerReleaseBufferBin() now releases locks as well.
>
> Do you have a better suggestion? I'll add a comment that makes that explicit,
> but other than that I don't have a great idea. Renaming the whole buffer pin
> mechanism seems pretty noisy.

ResOwnerReleaseBuffer()?

What do you mean renaming the whole buffer pin mechanism?

> > 0011:
> > --------
> > > [PATCH v6 11/14] heapam: Add batch mode mvcc check and use it in page mode
> >
> > > 2) We would like to stop setting hint bits while pages are being written
> > > out. The necessary locking becomes visible for page mode scans if done for
> > > every tuple. With batching the overhead can be amortized to only happen
> > > once per page.
>
> > And I presume you mean don't set hint bits on a buffer that is being flushed
> > by someone else -- but it sounds like you mean not to set hint bits as part
> > of flushing a buffer.
>
> Right, I do mean the former.

I would try and state it more clearly then.

> > +        /*
> > +         * If the page is not all-visible or we need to check serializability,
> > +         * maintain enough state to be able to refind the tuple efficiently,
> > +         * without again needing to extract it from the page.
> > +         */
> > +        if (!all_visible || check_serializable)
> > +        {
> >
> > "enough state" is pretty vague here.
>
> I don't really follow. Wouldn't going into more detail just restate the code
> in a comment?

I guess maybe something like

If the page is not all-visible or we need to check serializability,
keep track of the tuples so we can examine them later without the
overhead of extracting them from the page again.

> > 0012:
> > -------
>
> > @@ -77,6 +73,16 @@ gistkillitems(IndexScanDesc scan)
> >       */
> >      for (i = 0; i < so->numKilled; i++)
> >      {
> > +        if (!killedsomething)
> > +        {
> > +            /*
> > +             * Use hint bit infrastructure to be allowed to modify the page
> > +             * without holding an exclusive lock.
> > +             */
> > +            if (!BufferBeginSetHintBits(buffer))
> > +                goto unlock;
> > +        }
> > +
> >
> > I don't understand why this is in the loop. Clearly you want to call
> > BufferBeginSetHintBits() once, but why would you do it in the loop?
>
> Why would we want to continue if we can't set hint bits?

No, I'm suggesting that you move BufferBeginSetHintBits() outside the
loop so it is more obvious that it is happening once at the beginning.
Like this:

    if (so->numKilled > 0 && !BufferBeginSetHintBits(buffer))
        goto unlock;

    for (i = 0; i < so->numKilled; i++)
    {
        offnum = so->killedItems[i];
        iid = PageGetItemId(page, offnum);
        ItemIdMarkDead(iid);
        killedsomething = true;
    }

    if (killedsomething)
    {
        GistMarkPageHasGarbage(page);
        BufferFinishSetHintBits(buffer, true, true);
    }

unlock:
    UnlockReleaseBuffer(buffer);

> > And should the enum name itself (SetHintBitsState) include the word
> > "batch"? I know that would make it
> > long. At least the comment should explain that these are needed when batch
> > setting hint bits.
>
> Isn't that what the comment does, explaining that we want to amortize the cost
> of BufferBeginSetHintBits()?

well, amortize is a pretty fancy word and you don't say "batch" anywhere.

> > And then you inlined BufferSetHintBits16().

This was part of my wishlist for separating the batch and non-batch versions.

> > I presume you didn't do this because HeapTupleSatisifiesMVCC() for
> > SNAPSHOT_MVCC calls the non-batch version (and, of course,
> > HeapTupleSatisifiesVisibility() is the much more common case).
> >
> > if (TransactionIdIsValid(xid))
> > {
> >         if (BufferIsPermanent(buffer))
> >         {
> >                 /* NB: xid must be known committed here! */
> >                 XLogRecPtr    commitLSN = TransactionIdGetCommitLSN(xid);
> >
> >                 if (XLogNeedsFlush(commitLSN) &&
> >                         BufferGetLSNAtomic(buffer) < commitLSN)
> >                 {
> >                         /* not flushed and no LSN interlock, so don't
> > set hint */
> >                         return; false;
> >                 }
> >         }
> > }
> >
> > Separately, I was thinking, should we assert here about having the
> > right lock type?
>
> Not sure I get what assert of what locktype where?

In SetHintBitsExt() that we have share-exclusive or above. Or are
there still callers with only a share lock?

> >  * It is only safe to set a transaction-committed hint bit if we know the
> >  * transaction's commit record is guaranteed to be flushed to disk before the
> >  * buffer, or if the table is temporary or unlogged and will be obliterated by
> >  * a crash anyway.  We cannot change the LSN of the page here, because we may
> >  * hold only a share lock on the buffer, so we can only use the LSN to
> >  * interlock this if the buffer's LSN already is newer than the commit LSN;
> >  * otherwise we have to just refrain from setting the hint bit until some
> >  * future re-examination of the tuple.
> >  *
> >
> > Should this say "we may hold only a share exclusive lock on the
> > buffer". Also what is "this" in "only use the LSN to interlock this"?
>
> That's a pre-existing comment, right?

Right, but are there callers that will only have a share lock after your change?

> > @@ -1628,6 +1701,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot,
> > Buffer buffer,
> > +    if (state == SHB_ENABLED)
> > +        BufferFinishSetHintBits(buffer, true, true);
> > +
> >      return nvis;
> >  }
> >
> > I wondered if it would be more natural for BufferBeginSetHintBits()
> > and BufferFinishSetHintBits() to set SHB_INITIAL and SHB_DISABLED
> > instead of having callers do it. But, I guess you don't do this
> > because of gist and hash indexes using this for doing their own
> > modifications.
>
> I thought about it. But yea, the different callers seemed to make that not
> really useful. It'd also mean that we'd do an external function call for every
> tuple, which I think would be prohibitively expensive.

I'm confused why this would mean more external function calls.

> > --- a/src/backend/storage/freespace/freespace.c
> > +++ b/src/backend/storage/freespace/freespace.c
> > @@ -904,14 +904,22 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
> >     max_avail = fsm_get_max_avail(page);
> >     /*
> > -    * Reset the next slot pointer. This encourages the use of low-numbered
> > -    * pages, increasing the chances that a later vacuum can truncate the
> > -    * relation.  We don't bother with a lock here, nor with marking the page
> > -    * dirty if it wasn't already, since this is just a hint.
> > +    * Try to reset the next slot pointer. This encourages the use of
> > +    * low-numbered pages, increasing the chances that a later vacuum can
> > +    * truncate the relation.  We don't bother with a lock here, nor with
> > +    * marking the page dirty if it wasn't already, since this is just a hint.
> > +    *
> > +    * To be allowed to update the page without an exclusive lock, we have to
> > +    * use the hint bit infrastructure.
> >      */
> >
> > What the heck? This didn't even take a share lock before...
>
> It is insane. I do not understand how anybody thought this was ok.  I think I
> probably should split this out into a separate commit, since it's so insane.

Aren't you going to backport having it take a lock? Then it can be
separate (e.g. have it take a share lock in the "fix" commit and then
this commit bumps it to share-exclusive).

> > diff --git a/src/backend/storage/freespace/fsmpage.c
> > b/src/backend/storage/freesp>
> > /*
> > * Update the next-target pointer. Note that we do this even if we're only
> > * holding a shared lock, on the grounds that it's better to use a shared
> > * lock and get a garbled next pointer every now and then, than take the
> > * concurrency hit of an exclusive lock.
> >
> > We appear to avoid the garbling now?
>
> I don't think so. Two backends concurrently can do fsm_search_avail() and
> one backend might set a hint to a page that is already used up by the other
> one. At least I think so?

Maybe I don't know what it meant by garbled, but I thought it was
talking about two backends each trying to set fp_next_slot. If they
now have to have a share-exclusive lock and they can't both have a
share-exclusive lock at the same time, then it seems like that
wouldn't be a problem. It sounds like you may be talking about a
backend taking up the freespace of a page that is referred to by the
fp_next_slot?

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 20:02                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
@ 2025-11-25 20:46                       ` Andres Freund <andres@anarazel.de>
  2025-11-25 21:23                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-12-02 08:01                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  0 siblings, 2 replies; 120+ messages in thread

From: Andres Freund @ 2025-11-25 20:46 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-11-25 15:02:02 -0500, Melanie Plageman wrote:
> On Tue, Nov 25, 2025 at 11:54 AM Andres Freund <andres@anarazel.de> wrote:
> > > > As of this commit nothing uses the new share-exclusive lock mode. It will be
> > > > documented and used in a future commit. It seemed too complicated to introduce
> > > > the lock-level in a separate commit.
> > >
> > > I would have liked this mentioned earlier in the commit message.
> >
> > Hm, ok.  I could split share-exclusive out again, but it was somewhat painful,
> > because it doesn't just lead to adding code, but to changing code.
> 
> I don't think you need to introduce share-exclusive in a separate
> commit. I just mean it would be good to mention in the commit message
> before you get into so many other details that this commit doesn't use
> the share-exclusive level yet.
> 
> > > Also, I don't know how I feel about it being "documented" in a future
> > > commit...perhaps just don't say that.
> >
> > Hm, why?
> 
> I don't really see you documenting share-exclusive lock semantics in a
> later commit. Documentation for what the lock level is should be in
> the same commit that introduces it -- which I think you mostly have
> done. I feel like it is weird to say you will document that lock level
> later when 1) you don't really do that and 2) this commit mostly
> (besides some requests I had for elaboration) already does that.

The problem I faced was that the explanation for the lock level depends on
later changes - how do you document why share-exclusive is useful without
explaining the hint bit thing, which isn't yet implemented as of that commit.


> > > diff --git a/src/include/storage/buf_internals.h
> > > b/src/include/storage/buf_internals.h
> > > index 28519ad2813..0a145d95024 100644
> > > --- a/src/include/storage/buf_internals.h
> > > +++ b/src/include/storage/buf_internals.h
> > > @@ -32,22 +33,29 @@
> > >  /*
> > >   * Buffer state is a single 64-bit variable where following data is combined.
> > >   *
> > > + * State of the buffer itself:
> > >   * - 18 bits refcount
> > >   * - 4 bits usage count
> > >   * - 10 bits of flags
> > >   *
> > > + * State of the content lock:
> > > + * - 1 bit has_waiter
> > > + * - 1 bit release_ok
> > > + * - 1 bit lock state locked
> > >
> > > Somewhere you should clearly explain the scenarios in which you still
> > > need to take the buffer header lock with LockBufHdr() vs when you can
> > > just use atomic operations/CAS loops.
> >
> > There is an explanation of that, or at least my attempt at it ;). See the
> > dcoumentation for BufferDesc.
> 
> Hmm. I suppose. I was imagining something above the state member since
> you have added to it, but okay.
> 
> Separately, I'll admit I don't quite understand when I have to use
> LockBufHdr() and when I should use a CAS loop to update the
> bufferdesc->state.

It's definitely subtle :(. Not sure what to do about that, other than to work
on eventually just making everything doable with just a CAS. But that's a fair
bit of future work away.


> > > @@ -6625,7 +7295,25 @@ ResOwnerReleaseBufferPin(Datum res)
> > >      if (BufferIsLocal(buffer))
> > >          UnpinLocalBufferNoOwner(buffer);
> > >      else
> > > +    {
> > > +        PrivateRefCountEntry *ref;
> > > +
> > > +        ref = GetPrivateRefCountEntry(buffer, false);
> > > +
> > > +        /*
> > > +         * If the buffer was locked at the time of the resowner release,
> > > +         * release the lock now. This should only happen after errors.
> > > +         */
> > > +        if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
> > > +        {
> > > +            BufferDesc *buf = GetBufferDescriptor(buffer - 1);
> > > +
> > > +            HOLD_INTERRUPTS();    /* match the upcoming RESUME_INTERRUPTS */
> > > +            BufferLockUnlock(buffer, buf);
> > > +        }
> > > +
> > >          UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
> > > +    }
> > >  }
> > >
> > >  Bit confusing that ResOwnerReleaseBufferBin() now releases locks as well.
> >
> > Do you have a better suggestion? I'll add a comment that makes that explicit,
> > but other than that I don't have a great idea. Renaming the whole buffer pin
> > mechanism seems pretty noisy.
> 
> ResOwnerReleaseBuffer()?
> 
> What do you mean renaming the whole buffer pin mechanism?

All the *PrivateRefCount* stuff arguably aught to be renamed...



> > > I presume you didn't do this because HeapTupleSatisifiesMVCC() for
> > > SNAPSHOT_MVCC calls the non-batch version (and, of course,
> > > HeapTupleSatisifiesVisibility() is the much more common case).
> > >
> > > if (TransactionIdIsValid(xid))
> > > {
> > >         if (BufferIsPermanent(buffer))
> > >         {
> > >                 /* NB: xid must be known committed here! */
> > >                 XLogRecPtr    commitLSN = TransactionIdGetCommitLSN(xid);
> > >
> > >                 if (XLogNeedsFlush(commitLSN) &&
> > >                         BufferGetLSNAtomic(buffer) < commitLSN)
> > >                 {
> > >                         /* not flushed and no LSN interlock, so don't
> > > set hint */
> > >                         return; false;
> > >                 }
> > >         }
> > > }
> > >
> > > Separately, I was thinking, should we assert here about having the
> > > right lock type?
> >
> > Not sure I get what assert of what locktype where?
> 
> In SetHintBitsExt() that we have share-exclusive or above.

We *don't* necessarily hold that though, it'll just be acquired by
BufferBeginSetHintBits().  Or do you mean in the SHB_ENABLED case?


> Or are there still callers with only a share lock?

Almost all of them, I think? We can't just call this with an unconditional
share-exclusive lock, since that'd destroy concurrency. Instead we just try to
upgrade to share-exclusive to set the hint bit. If we can't get
share-exclusive, we don't need to block, we just haven't set the hint bit.


> > > @@ -1628,6 +1701,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot,
> > > Buffer buffer,
> > > +    if (state == SHB_ENABLED)
> > > +        BufferFinishSetHintBits(buffer, true, true);
> > > +
> > >      return nvis;
> > >  }
> > >
> > > I wondered if it would be more natural for BufferBeginSetHintBits()
> > > and BufferFinishSetHintBits() to set SHB_INITIAL and SHB_DISABLED
> > > instead of having callers do it. But, I guess you don't do this
> > > because of gist and hash indexes using this for doing their own
> > > modifications.
> >
> > I thought about it. But yea, the different callers seemed to make that not
> > really useful. It'd also mean that we'd do an external function call for every
> > tuple, which I think would be prohibitively expensive.
> 
> I'm confused why this would mean more external function calls.

The visibility logic is in heapam_visibility.c, the locking for the buffer in
bufmgr.c? If the knowledge about SetHintBitsState lives in
BufferBeginSetHintBits(), we need to call it to do that check.


> > > --- a/src/backend/storage/freespace/freespace.c
> > > +++ b/src/backend/storage/freespace/freespace.c
> > > @@ -904,14 +904,22 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
> > >     max_avail = fsm_get_max_avail(page);
> > >     /*
> > > -    * Reset the next slot pointer. This encourages the use of low-numbered
> > > -    * pages, increasing the chances that a later vacuum can truncate the
> > > -    * relation.  We don't bother with a lock here, nor with marking the page
> > > -    * dirty if it wasn't already, since this is just a hint.
> > > +    * Try to reset the next slot pointer. This encourages the use of
> > > +    * low-numbered pages, increasing the chances that a later vacuum can
> > > +    * truncate the relation.  We don't bother with a lock here, nor with
> > > +    * marking the page dirty if it wasn't already, since this is just a hint.
> > > +    *
> > > +    * To be allowed to update the page without an exclusive lock, we have to
> > > +    * use the hint bit infrastructure.
> > >      */
> > >
> > > What the heck? This didn't even take a share lock before...
> >
> > It is insane. I do not understand how anybody thought this was ok.  I think I
> > probably should split this out into a separate commit, since it's so insane.
> 
> Aren't you going to backport having it take a lock? Then it can be
> separate (e.g. have it take a share lock in the "fix" commit and then
> this commit bumps it to share-exclusive).

I wasn't thinking we should backport that. It's been this way for ages,
without having caused known issues....


> > > diff --git a/src/backend/storage/freespace/fsmpage.c
> > > b/src/backend/storage/freesp>
> > > /*
> > > * Update the next-target pointer. Note that we do this even if we're only
> > > * holding a shared lock, on the grounds that it's better to use a shared
> > > * lock and get a garbled next pointer every now and then, than take the
> > > * concurrency hit of an exclusive lock.
> > >
> > > We appear to avoid the garbling now?
> >
> > I don't think so. Two backends concurrently can do fsm_search_avail() and
> > one backend might set a hint to a page that is already used up by the other
> > one. At least I think so?
> 
> Maybe I don't know what it meant by garbled, but I thought it was
> talking about two backends each trying to set fp_next_slot. If they
> now have to have a share-exclusive lock and they can't both have a
> share-exclusive lock at the same time, then it seems like that
> wouldn't be a problem. It sounds like you may be talking about a
> backend taking up the freespace of a page that is referred to by the
> fp_next_slot?

Yes, a version of the latter. The value that fp_next_slot will be set to can
be outdated by the time we actually set it, unless we do all of
fsm_search_avail() under some form of exclusive lock - clearly not something
desirable.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 20:02                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 20:46                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-11-25 21:23                         ` Melanie Plageman <melanieplageman@gmail.com>
  1 sibling, 0 replies; 120+ messages in thread

From: Melanie Plageman @ 2025-11-25 21:23 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Tue, Nov 25, 2025 at 3:46 PM Andres Freund <andres@anarazel.de> wrote:
>
>
> > > > I presume you didn't do this because HeapTupleSatisifiesMVCC() for
> > > > SNAPSHOT_MVCC calls the non-batch version (and, of course,
> > > > HeapTupleSatisifiesVisibility() is the much more common case).
> > > >
> > > > if (TransactionIdIsValid(xid))
> > > > {
> > > >         if (BufferIsPermanent(buffer))
> > > >         {
> > > >                 /* NB: xid must be known committed here! */
> > > >                 XLogRecPtr    commitLSN = TransactionIdGetCommitLSN(xid);
> > > >
> > > >                 if (XLogNeedsFlush(commitLSN) &&
> > > >                         BufferGetLSNAtomic(buffer) < commitLSN)
> > > >                 {
> > > >                         /* not flushed and no LSN interlock, so don't
> > > > set hint */
> > > >                         return; false;
> > > >                 }
> > > >         }
> > > > }
> > > >
> > > > Separately, I was thinking, should we assert here about having the
> > > > right lock type?
> > >
> > > Not sure I get what assert of what locktype where?
> >
> > In SetHintBitsExt() that we have share-exclusive or above.
>
> We *don't* necessarily hold that though, it'll just be acquired by
> BufferBeginSetHintBits().  Or do you mean in the SHB_ENABLED case?

Yea, in the enabled case. Also, can't we skip the whole

    if (TransactionIdIsValid(xid))
    {
        if (BufferIsPermanent(buffer))
        {
            /* NB: xid must be known committed here! */
            XLogRecPtr    commitLSN = TransactionIdGetCommitLSN(xid);

            if (XLogNeedsFlush(commitLSN) &&
                BufferGetLSNAtomic(buffer) < commitLSN)
            {
                /* not flushed and no LSN interlock, so don't set hint */
                return;
            }
        }
    }

part if state is SHB_DISABLED?

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 20:02                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 20:46                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-12-02 08:01                         ` Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-02 13:20                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  1 sibling, 1 reply; 120+ messages in thread

From: Heikki Linnakangas @ 2025-12-02 08:01 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; Melanie Plageman <melanieplageman@gmail.com>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On 25/11/2025 22:46, Andres Freund wrote:
> On 2025-11-25 15:02:02 -0500, Melanie Plageman wrote:
>> On Tue, Nov 25, 2025 at 11:54 AM Andres Freund <andres@anarazel.de> wrote:
>>>> --- a/src/backend/storage/freespace/freespace.c
>>>> +++ b/src/backend/storage/freespace/freespace.c
>>>> @@ -904,14 +904,22 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
>>>>      max_avail = fsm_get_max_avail(page);
>>>>      /*
>>>> -    * Reset the next slot pointer. This encourages the use of low-numbered
>>>> -    * pages, increasing the chances that a later vacuum can truncate the
>>>> -    * relation.  We don't bother with a lock here, nor with marking the page
>>>> -    * dirty if it wasn't already, since this is just a hint.
>>>> +    * Try to reset the next slot pointer. This encourages the use of
>>>> +    * low-numbered pages, increasing the chances that a later vacuum can
>>>> +    * truncate the relation.  We don't bother with a lock here, nor with
>>>> +    * marking the page dirty if it wasn't already, since this is just a hint.
>>>> +    *
>>>> +    * To be allowed to update the page without an exclusive lock, we have to
>>>> +    * use the hint bit infrastructure.
>>>>       */
>>>>
>>>> What the heck? This didn't even take a share lock before...
>>>
>>> It is insane. I do not understand how anybody thought this was ok.  I think I
>>> probably should split this out into a separate commit, since it's so insane.
>>
>> Aren't you going to backport having it take a lock? Then it can be
>> separate (e.g. have it take a share lock in the "fix" commit and then
>> this commit bumps it to share-exclusive).
> 
> I wasn't thinking we should backport that. It's been this way for ages,
> without having caused known issues....

I wrote that originally :-). The FSM code always treats the FSM page 
contents as potentially corrupted garbage. FSM page updates are not 
WAL-logged, for starters. And as the comment says, the next slot pointer 
is just a hint for where within the page to start looking for free space.

Page checksums were added later, and now we know about the problems of 
modifying a page a write() is in progress, on some filesystems with 
filesystem-level checksums. So it's a good idea to tighten it up, but it 
was fine back then.

>>>> diff --git a/src/backend/storage/freespace/fsmpage.c
>>>> b/src/backend/storage/freesp>
>>>> /*
>>>> * Update the next-target pointer. Note that we do this even if we're only
>>>> * holding a shared lock, on the grounds that it's better to use a shared
>>>> * lock and get a garbled next pointer every now and then, than take the
>>>> * concurrency hit of an exclusive lock.
>>>>
>>>> We appear to avoid the garbling now?
>>>
>>> I don't think so. Two backends concurrently can do fsm_search_avail() and
>>> one backend might set a hint to a page that is already used up by the other
>>> one. At least I think so?
>>
>> Maybe I don't know what it meant by garbled, but I thought it was
>> talking about two backends each trying to set fp_next_slot. If they
>> now have to have a share-exclusive lock and they can't both have a
>> share-exclusive lock at the same time, then it seems like that
>> wouldn't be a problem. It sounds like you may be talking about a
>> backend taking up the freespace of a page that is referred to by the
>> fp_next_slot?
> 
> Yes, a version of the latter. The value that fp_next_slot will be set to can
> be outdated by the time we actually set it, unless we do all of
> fsm_search_avail() under some form of exclusive lock - clearly not something
> desirable.

I'm pretty sure the "garbled" in the comment means the former, not the 
latter. I.e. it means that the pointer itself might become garbage. 
Would be good to update the comment if that's no longer possible.

But speaking of that: why do we not allow two processes to concurrently 
set hint bits on a page anymore?

- Heikki






^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 20:02                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 20:46                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-02 08:01                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
@ 2025-12-02 13:20                           ` Andres Freund <andres@anarazel.de>
  2025-12-02 13:38                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-12-02 13:20 UTC (permalink / raw)
  To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-12-02 10:01:06 +0200, Heikki Linnakangas wrote:
> On 25/11/2025 22:46, Andres Freund wrote:
> > > > > diff --git a/src/backend/storage/freespace/fsmpage.c
> > > > > b/src/backend/storage/freesp>
> > > > > /*
> > > > > * Update the next-target pointer. Note that we do this even if we're only
> > > > > * holding a shared lock, on the grounds that it's better to use a shared
> > > > > * lock and get a garbled next pointer every now and then, than take the
> > > > > * concurrency hit of an exclusive lock.
> > > > > 
> > > > > We appear to avoid the garbling now?
> > > > 
> > > > I don't think so. Two backends concurrently can do fsm_search_avail() and
> > > > one backend might set a hint to a page that is already used up by the other
> > > > one. At least I think so?
> > > 
> > > Maybe I don't know what it meant by garbled, but I thought it was
> > > talking about two backends each trying to set fp_next_slot. If they
> > > now have to have a share-exclusive lock and they can't both have a
> > > share-exclusive lock at the same time, then it seems like that
> > > wouldn't be a problem. It sounds like you may be talking about a
> > > backend taking up the freespace of a page that is referred to by the
> > > fp_next_slot?
> > 
> > Yes, a version of the latter. The value that fp_next_slot will be set to can
> > be outdated by the time we actually set it, unless we do all of
> > fsm_search_avail() under some form of exclusive lock - clearly not something
> > desirable.
> 
> I'm pretty sure the "garbled" in the comment means the former, not the
> latter. I.e. it means that the pointer itself might become garbage. Would be
> good to update the comment if that's no longer possible.

Hm. I thought we had always assumed that 4byte values can be read/written
tear-free. Hence thinking that garbled couldn't refer to reading entire
garbage due to a concurrent write.


> But speaking of that: why do we not allow two processes to concurrently set
> hint bits on a page anymore?

It'd make the locking a lot more complicated without much of a benefit.

The new share-exclusive lock mode only requires one additional bit of lock
state, for the single allowed holder. If we wanted a new lockmode that
prevented the page from being written out concurrently, but could be held
multiple times, we'd need at least MAX_BACKENDS bits for the lock level
allowing hint bits to be set and another lock level to acquire while writing
out the buffer.

At the same time, there seems to be little benefit in setting hint bits on a
page concurrently. A very common case is that the same hint bit(s) would be
set by multiple backends, we don't gain anything from that. And in the cases
where hint bits were intended to be set for different tuples, the window in
which that is not allowed is very narrow, and the cost of not setting right in
that moment is pretty small and the cost of not setting the hint bit right
then and there isn't high.

Makes sense?

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 20:02                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 20:46                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-02 08:01                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-02 13:20                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-12-02 13:38                             ` Heikki Linnakangas <hlinnaka@iki.fi>
  0 siblings, 0 replies; 120+ messages in thread

From: Heikki Linnakangas @ 2025-12-02 13:38 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On 02/12/2025 15:20, Andres Freund wrote:
> On 2025-12-02 10:01:06 +0200, Heikki Linnakangas wrote:
>> But speaking of that: why do we not allow two processes to concurrently set
>> hint bits on a page anymore?
> 
> It'd make the locking a lot more complicated without much of a benefit.
> 
> The new share-exclusive lock mode only requires one additional bit of lock
> state, for the single allowed holder. If we wanted a new lockmode that
> prevented the page from being written out concurrently, but could be held
> multiple times, we'd need at least MAX_BACKENDS bits for the lock level
> allowing hint bits to be set and another lock level to acquire while writing
> out the buffer.
> 
> At the same time, there seems to be little benefit in setting hint bits on a
> page concurrently. A very common case is that the same hint bit(s) would be
> set by multiple backends, we don't gain anything from that. And in the cases
> where hint bits were intended to be set for different tuples, the window in
> which that is not allowed is very narrow, and the cost of not setting right in
> that moment is pretty small and the cost of not setting the hint bit right
> then and there isn't high.
> 
> Makes sense?

Yep, makes sense. Would be good to put that in a comment somewhere, and 
in the commit message. "We could allow multiple backends to set hint 
bits concurrently, but it'd make the lock implementation more complicated"

- Heikki






^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-12-03 00:47                     ` Andres Freund <andres@anarazel.de>
  2025-12-03 01:12                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
  2025-12-03 16:03                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  1 sibling, 3 replies; 120+ messages in thread

From: Andres Freund @ 2025-12-03 00:47 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-11-25 11:54:00 -0500, Andres Freund wrote:
> Thanks a lot for that detailed review!  A few questions and comments, before I
> try to address the comments in the next version.

Here's that new new version, with the following changes

- Some more micro-optimizations, most importantly adding a commit that doesn't
  initialize the delay in LockBufHdr() unless needed. With those I don't see a
  consistent slowdown anymore (slight speedup on one workstation, slight
  slowdown on another, in an absurdly adverse workload)

- Tried to address Melanie's feedback, with some exceptions (some noted below,
  but I also need to make another pass through the reviews)

- re-implemented AssertNotCatalogBufferLock() in the new world

- Substantially expanded comments around setting hint bits (in buffer/README,
  heapam_visibility.c and bufmgr.c)

- split out the change to fsm_vacuum_page() to start to lock the page into is
  own commit

- reordered patch series so that smaller changes are before the 64bit-state
  and "Implement buffer content locks independently of" commits, so they can
  be committed while we finish cleaning the later changes

- I didn't invest much in cleaning up the later patches ("Don't copy pages
  while writing out" and "Make UnlockReleaseBuffer() more efficient") yet,
  wanted to focus on the earlier patches first


Todo:

- still need to rename ResOwnerReleaseBufferPin(). Wondering about what to
  rename ResourceOwnerDesc.name to. "buffer ownership" maybe? Not great...

- gistkillitems() complaint by Melanie

- amortize vs batch vs SetHintBits comment + SHB_* names

- for the next version I'll remove the BATCHMVCC_FEWER_ARGS conditionals from
  0010. I don't love needing BatchMVCCState but I don't really see an
  alternative, the performance difference is pretty persistent.


Questions:
- ForEachLWLockHeldByMe() and LWLockDisown() aren't used anymore, should we
  remove them?


Greetings,

Andres Freund

Attachments:

  [text/x-diff] v7-0001-bufmgr-Turn-BUFFER_LOCK_-into-an-enum.patch (3.2K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/2-v7-0001-bufmgr-Turn-BUFFER_LOCK_-into-an-enum.patch)
  download | inline diff:
From 2c020be4cbb3c810c8ac4b5b9e5aef03791d25bd Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Fri, 7 Nov 2025 16:51:52 -0500
Subject: [PATCH v7 01/15] bufmgr: Turn BUFFER_LOCK_* into an enum

It seems cleaner to use an enum to tie the different values together. It also
helps to have a more descriptive type in the argument to various functions.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/bufmgr.h        | 13 ++++++++-----
 src/backend/storage/buffer/bufmgr.c |  4 ++--
 src/tools/pgindent/typedefs.list    |  1 +
 3 files changed, 11 insertions(+), 7 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 9f6785910e0..97c1124c12a 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -200,9 +200,12 @@ extern PGDLLIMPORT int32 *LocalRefCount;
 /*
  * Buffer content lock modes (mode argument for LockBuffer())
  */
-#define BUFFER_LOCK_UNLOCK		0
-#define BUFFER_LOCK_SHARE		1
-#define BUFFER_LOCK_EXCLUSIVE	2
+typedef enum BufferLockMode
+{
+	BUFFER_LOCK_UNLOCK,
+	BUFFER_LOCK_SHARE,
+	BUFFER_LOCK_EXCLUSIVE,
+} BufferLockMode;
 
 
 /*
@@ -238,7 +241,7 @@ extern void WaitReadBuffers(ReadBuffersOperation *operation);
 extern void ReleaseBuffer(Buffer buffer);
 extern void UnlockReleaseBuffer(Buffer buffer);
 extern bool BufferIsLockedByMe(Buffer buffer);
-extern bool BufferIsLockedByMeInMode(Buffer buffer, int mode);
+extern bool BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode);
 extern bool BufferIsDirty(Buffer buffer);
 extern void MarkBufferDirty(Buffer buffer);
 extern void IncrBufferRefCount(Buffer buffer);
@@ -299,7 +302,7 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
 extern void UnlockBuffers(void);
-extern void LockBuffer(Buffer buffer, int mode);
+extern void LockBuffer(Buffer buffer, BufferLockMode mode);
 extern bool ConditionalLockBuffer(Buffer buffer);
 extern void LockBufferForCleanup(Buffer buffer);
 extern bool ConditionalLockBufferForCleanup(Buffer buffer);
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index f373cead95f..00d9a23b675 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2866,7 +2866,7 @@ BufferIsLockedByMe(Buffer buffer)
  * Buffer must be pinned.
  */
 bool
-BufferIsLockedByMeInMode(Buffer buffer, int mode)
+BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode)
 {
 	BufferDesc *bufHdr;
 
@@ -5601,7 +5601,7 @@ UnlockBuffers(void)
  * Acquire or release the content_lock for the buffer.
  */
 void
-LockBuffer(Buffer buffer, int mode)
+LockBuffer(Buffer buffer, BufferLockMode mode)
 {
 	BufferDesc *buf;
 
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index cf3f6a7dafd..abcb3c4c4cc 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -345,6 +345,7 @@ BufferCachePagesRec
 BufferDesc
 BufferDescPadded
 BufferHeapTupleTableSlot
+BufferLockMode
 BufferLookupEnt
 BufferManagerRelation
 BufferStrategyControl
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0002-Add-pg_atomic_unlocked_write_u64.patch (2.0K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/3-v7-0002-Add-pg_atomic_unlocked_write_u64.patch)
  download | inline diff:
From 1d8fee659fc3cb08df9191f29550bc9528ea7250 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 5 Nov 2025 19:12:37 -0500
Subject: [PATCH v7 02/15] Add pg_atomic_unlocked_write_u64

The 64bit equivalent of pg_atomic_unlocked_write_u32(), to be used in an
upcoming patch converting BufferDesc.state into a 64bit atomic.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/port/atomics.h         | 10 ++++++++++
 src/include/port/atomics/generic.h |  9 +++++++++
 2 files changed, 19 insertions(+)

diff --git a/src/include/port/atomics.h b/src/include/port/atomics.h
index 96f1858da97..830ea5c7c52 100644
--- a/src/include/port/atomics.h
+++ b/src/include/port/atomics.h
@@ -488,6 +488,16 @@ pg_atomic_write_u64(volatile pg_atomic_uint64 *ptr, uint64 val)
 	pg_atomic_write_u64_impl(ptr, val);
 }
 
+static inline void
+pg_atomic_unlocked_write_u64(volatile pg_atomic_uint64 *ptr, uint64 val)
+{
+#ifndef PG_HAVE_ATOMIC_U64_SIMULATION
+	AssertPointerAlignment(ptr, 8);
+#endif
+
+	pg_atomic_unlocked_write_u64_impl(ptr, val);
+}
+
 static inline void
 pg_atomic_write_membarrier_u64(volatile pg_atomic_uint64 *ptr, uint64 val)
 {
diff --git a/src/include/port/atomics/generic.h b/src/include/port/atomics/generic.h
index 6b61a7b5416..00aa152f908 100644
--- a/src/include/port/atomics/generic.h
+++ b/src/include/port/atomics/generic.h
@@ -297,6 +297,15 @@ pg_atomic_write_u64_impl(volatile pg_atomic_uint64 *ptr, uint64 val)
 #endif /* PG_HAVE_8BYTE_SINGLE_COPY_ATOMICITY && !PG_HAVE_ATOMIC_U64_SIMULATION */
 #endif /* PG_HAVE_ATOMIC_WRITE_U64 */
 
+#ifndef PG_HAVE_ATOMIC_UNLOCKED_WRITE_U64
+#define PG_HAVE_ATOMIC_UNLOCKED_WRITE_U64
+static inline void
+pg_atomic_unlocked_write_u64_impl(volatile pg_atomic_uint64 *ptr, uint64 val)
+{
+	ptr->value = val;
+}
+#endif
+
 #ifndef PG_HAVE_ATOMIC_READ_U64
 #define PG_HAVE_ATOMIC_READ_U64
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0003-Rename-BUFFERPIN-wait-event-class-to-BUFFER.patch (6.6K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/4-v7-0003-Rename-BUFFERPIN-wait-event-class-to-BUFFER.patch)
  download | inline diff:
From 464fd5a3cd99c4a9afe5331226d6e7443c52e606 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 6 Nov 2025 09:15:18 -0500
Subject: [PATCH v7 03/15] Rename BUFFERPIN wait event class to BUFFER

In an upcoming patch more wait events will be added to the wait event
class (for buffer locking), making the current name too
specific. Alternatively we could introduce a dedicated wait event class for
those, but it seems somewhat confusing to have a BUFFERPIN and a BUFFER wait
event class.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/utils/wait_classes.h                |  2 +-
 src/backend/storage/buffer/bufmgr.c             |  2 +-
 src/backend/storage/ipc/standby.c               |  2 +-
 src/backend/utils/activity/wait_event.c         | 12 ++++++------
 src/backend/utils/activity/wait_event_names.txt |  6 +++---
 doc/src/sgml/monitoring.sgml                    |  8 +++-----
 src/test/recovery/t/048_vacuum_horizon_floor.pl |  2 +-
 src/test/regress/expected/sysviews.out          |  2 +-
 8 files changed, 17 insertions(+), 19 deletions(-)

diff --git a/src/include/utils/wait_classes.h b/src/include/utils/wait_classes.h
index 51ee68397d5..57888aa62f7 100644
--- a/src/include/utils/wait_classes.h
+++ b/src/include/utils/wait_classes.h
@@ -17,7 +17,7 @@
  */
 #define PG_WAIT_LWLOCK				0x01000000U
 #define PG_WAIT_LOCK				0x03000000U
-#define PG_WAIT_BUFFERPIN			0x04000000U
+#define PG_WAIT_BUFFER				0x04000000U
 #define PG_WAIT_ACTIVITY			0x05000000U
 #define PG_WAIT_CLIENT				0x06000000U
 #define PG_WAIT_EXTENSION			0x07000000U
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 00d9a23b675..62f420dd344 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5799,7 +5799,7 @@ LockBufferForCleanup(Buffer buffer)
 			SetStartupBufferPinWaitBufId(-1);
 		}
 		else
-			ProcWaitForSignal(WAIT_EVENT_BUFFER_PIN);
+			ProcWaitForSignal(WAIT_EVENT_BUFFER_CLEANUP);
 
 		/*
 		 * Remove flag marking us as waiter. Normally this will not be set
diff --git a/src/backend/storage/ipc/standby.c b/src/backend/storage/ipc/standby.c
index 4222bdab078..fc45d72c79b 100644
--- a/src/backend/storage/ipc/standby.c
+++ b/src/backend/storage/ipc/standby.c
@@ -840,7 +840,7 @@ ResolveRecoveryConflictWithBufferPin(void)
 	 * SIGHUP signal handler, etc cannot do that because it uses the different
 	 * latch from that ProcWaitForSignal() waits on.
 	 */
-	ProcWaitForSignal(WAIT_EVENT_BUFFER_PIN);
+	ProcWaitForSignal(WAIT_EVENT_BUFFER_CLEANUP);
 
 	if (got_standby_delay_timeout)
 		SendRecoveryConflictWithBufferPin(PROCSIG_RECOVERY_CONFLICT_BUFFERPIN);
diff --git a/src/backend/utils/activity/wait_event.c b/src/backend/utils/activity/wait_event.c
index d9b8f34a355..96d61f77f6e 100644
--- a/src/backend/utils/activity/wait_event.c
+++ b/src/backend/utils/activity/wait_event.c
@@ -29,7 +29,7 @@
 
 
 static const char *pgstat_get_wait_activity(WaitEventActivity w);
-static const char *pgstat_get_wait_bufferpin(WaitEventBufferPin w);
+static const char *pgstat_get_wait_buffer(WaitEventBuffer w);
 static const char *pgstat_get_wait_client(WaitEventClient w);
 static const char *pgstat_get_wait_ipc(WaitEventIPC w);
 static const char *pgstat_get_wait_timeout(WaitEventTimeout w);
@@ -389,8 +389,8 @@ pgstat_get_wait_event_type(uint32 wait_event_info)
 		case PG_WAIT_LOCK:
 			event_type = "Lock";
 			break;
-		case PG_WAIT_BUFFERPIN:
-			event_type = "BufferPin";
+		case PG_WAIT_BUFFER:
+			event_type = "Buffer";
 			break;
 		case PG_WAIT_ACTIVITY:
 			event_type = "Activity";
@@ -453,11 +453,11 @@ pgstat_get_wait_event(uint32 wait_event_info)
 		case PG_WAIT_INJECTIONPOINT:
 			event_name = GetWaitEventCustomIdentifier(wait_event_info);
 			break;
-		case PG_WAIT_BUFFERPIN:
+		case PG_WAIT_BUFFER:
 			{
-				WaitEventBufferPin w = (WaitEventBufferPin) wait_event_info;
+				WaitEventBuffer w = (WaitEventBuffer) wait_event_info;
 
-				event_name = pgstat_get_wait_bufferpin(w);
+				event_name = pgstat_get_wait_buffer(w);
 				break;
 			}
 		case PG_WAIT_ACTIVITY:
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index c1ac71ff7f2..1e5e368a5dc 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -279,12 +279,12 @@ WAL_WRITE	"Waiting for a write to a WAL file."
 ABI_compatibility:
 
 #
-# Wait Events - Buffer Pin
+# Wait Events - Buffer
 #
 
-Section: ClassName - WaitEventBufferPin
+Section: ClassName - WaitEventBuffer
 
-BUFFER_PIN	"Waiting to acquire an exclusive pin on a buffer."
+BUFFER_CLEANUP	"Waiting to acquire an exclusive pin on a buffer. Buffer pin waits can be protracted if another process holds an open cursor that last read data from the buffer in question."
 
 ABI_compatibility:
 
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index e0556b6baac..039d73691be 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -1053,11 +1053,9 @@ postgres   27093  0.0  0.0  30096  2752 ?        Ss   11:34   0:00 postgres: ser
       </entry>
      </row>
      <row>
-      <entry><literal>BufferPin</literal></entry>
-      <entry>The server process is waiting for exclusive access to
-       a data buffer.  Buffer pin waits can be protracted if
-       another process holds an open cursor that last read data from the
-       buffer in question. See <xref linkend="wait-event-bufferpin-table"/>.
+      <entry><literal>Buffer</literal></entry>
+      <entry>The server process is waiting for access to a data buffer.
+      See <xref linkend="wait-event-buffer-table"/>.
       </entry>
      </row>
      <row>
diff --git a/src/test/recovery/t/048_vacuum_horizon_floor.pl b/src/test/recovery/t/048_vacuum_horizon_floor.pl
index 668eedd71b2..9cdf6cee8a7 100644
--- a/src/test/recovery/t/048_vacuum_horizon_floor.pl
+++ b/src/test/recovery/t/048_vacuum_horizon_floor.pl
@@ -194,7 +194,7 @@ $node_primary->poll_query_until(
 	qq[
 	SELECT count(*) >= 1 FROM pg_stat_activity
 		WHERE pid = $vacuum_pid
-		AND wait_event = 'BufferPin';
+		AND wait_event = 'BufferCleanup';
 	],
 	't');
 
diff --git a/src/test/regress/expected/sysviews.out b/src/test/regress/expected/sysviews.out
index 3b37fafa65b..0411db832f1 100644
--- a/src/test/regress/expected/sysviews.out
+++ b/src/test/regress/expected/sysviews.out
@@ -182,7 +182,7 @@ select type, count(*) > 0 as ok FROM pg_wait_events
    type    | ok 
 -----------+----
  Activity  | t
- BufferPin | t
+ Buffer    | t
  Client    | t
  Extension | t
  IO        | t
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0004-bufmgr-Optimize-LockBufHdr-by-delaying-spin-delay.patch (2.3K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/5-v7-0004-bufmgr-Optimize-LockBufHdr-by-delaying-spin-delay.patch)
  download | inline diff:
From 8c9488e9b467e0f405c59dd04143e19cbc1bd34d Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 2 Dec 2025 18:44:54 -0500
Subject: [PATCH v7 04/15] bufmgr: Optimize LockBufHdr() by delaying spin-delay
 setup

Previously we always initialized the SpinDelayStatus. That is sufficiently
expensive / buffer header lock acquisitions are sufficiently frequent to make
it worthwhile to instead have a fastpath that does not initialize the
SpinDelayStatus.

While this is a small gain on its own, it mainly is aimed at preventing a
regression after a future commit, which requires additional locking to set
hint bits.

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/backend/storage/buffer/bufmgr.c | 32 ++++++++++++++++++++---------
 1 file changed, 22 insertions(+), 10 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 62f420dd344..0ffad2cd735 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -6263,23 +6263,35 @@ rlocator_comparator(const void *p1, const void *p2)
 uint32
 LockBufHdr(BufferDesc *desc)
 {
-	SpinDelayStatus delayStatus;
 	uint32		old_buf_state;
 
 	Assert(!BufferIsLocal(BufferDescriptorGetBuffer(desc)));
 
-	init_local_spin_delay(&delayStatus);
+	/*
+	 * Try to acquire the lock once, without setting up the spin-delay
+	 * infrastructure. The work necessary for that shows up in profiles and is
+	 * rarely necessary.
+	 */
+	old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
 
-	while (true)
+	if (unlikely(old_buf_state & BM_LOCKED))
 	{
-		/* set BM_LOCKED flag */
-		old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
-		/* if it wasn't set before we're OK */
-		if (!(old_buf_state & BM_LOCKED))
-			break;
-		perform_spin_delay(&delayStatus);
+		SpinDelayStatus delayStatus;
+
+		init_local_spin_delay(&delayStatus);
+
+		while (true)
+		{
+			/* set BM_LOCKED flag */
+			old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
+			/* if it wasn't set before we're OK */
+			if (!(old_buf_state & BM_LOCKED))
+				break;
+			perform_spin_delay(&delayStatus);
+		}
+		finish_spin_delay(&delayStatus);
 	}
-	finish_spin_delay(&delayStatus);
+
 	return old_buf_state | BM_LOCKED;
 }
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0005-bufmgr-Separate-keys-for-private-refcount-infrast.patch (11.2K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/6-v7-0005-bufmgr-Separate-keys-for-private-refcount-infrast.patch)
  download | inline diff:
From 9009631b0890e5e8729715e1de192e82c75a504e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 12 Nov 2025 12:50:52 -0500
Subject: [PATCH v7 05/15] bufmgr: Separate keys for private refcount
 infrastructure

This makes lookups faster, due to allowing auto-vectorized lookups. It is also
beneficial for an upcoming patch, independent of auto-vectorization, as the
upcoming patch wants to track more information for each pinned buffer, making
the existing loop, iterating over an array of PrivateRefCountEntry, more
expensive due to increasing its size.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/backend/storage/buffer/bufmgr.c | 138 ++++++++++++++++++----------
 src/tools/pgindent/typedefs.list    |   1 +
 2 files changed, 93 insertions(+), 46 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 0ffad2cd735..4e147d477c7 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -90,10 +90,28 @@
  */
 #define BUF_DROP_FULL_SCAN_THRESHOLD		(uint64) (NBuffers / 32)
 
-typedef struct PrivateRefCountEntry
+typedef struct PrivateRefCountData
 {
-	Buffer		buffer;
+	/*
+	 * How many times has the buffer been pinned by this backend.
+	 */
 	int32		refcount;
+} PrivateRefCountData;
+
+typedef struct PrivateRefCountEntry
+{
+	/*
+	 * Note that this needs to be same as the entry's corresponding
+	 * PrivateRefCountArrayKeys[i], if the entry is stored in the array. We
+	 * store it in both places as this is used for the hashtable key and
+	 * because it is more convenient (passing around a PrivateRefCountEntry
+	 * suffices to identify the buffer) and faster (checking the keys array is
+	 * faster when checking many entries, checking the entry is faster if just
+	 * checking a single entry).
+	 */
+	Buffer		buffer;
+
+	PrivateRefCountData data;
 } PrivateRefCountEntry;
 
 /* 64 bytes, about the size of a cache line on common systems */
@@ -194,7 +212,8 @@ static BufferDesc *PinCountWaitBuf = NULL;
  *
  * To avoid - as we used to - requiring an array with NBuffers entries to keep
  * track of local buffers, we use a small sequentially searched array
- * (PrivateRefCountArray) and an overflow hash table (PrivateRefCountHash) to
+ * (PrivateRefCountArrayKeys, with the corresponding data stored in
+ * PrivateRefCountArray) and an overflow hash table (PrivateRefCountHash) to
  * keep track of backend local pins.
  *
  * Until no more than REFCOUNT_ARRAY_ENTRIES buffers are pinned at once, all
@@ -212,11 +231,12 @@ static BufferDesc *PinCountWaitBuf = NULL;
  * memory allocations in NewPrivateRefCountEntry() which can be important
  * because in some scenarios it's called with a spinlock held...
  */
+static Buffer PrivateRefCountArrayKeys[REFCOUNT_ARRAY_ENTRIES];
 static struct PrivateRefCountEntry PrivateRefCountArray[REFCOUNT_ARRAY_ENTRIES];
 static HTAB *PrivateRefCountHash = NULL;
 static int32 PrivateRefCountOverflowed = 0;
 static uint32 PrivateRefCountClock = 0;
-static PrivateRefCountEntry *ReservedRefCountEntry = NULL;
+static int	ReservedRefCountSlot = -1;
 
 static uint32 MaxProportionalPins;
 
@@ -259,7 +279,7 @@ static void
 ReservePrivateRefCountEntry(void)
 {
 	/* Already reserved (or freed), nothing to do */
-	if (ReservedRefCountEntry != NULL)
+	if (ReservedRefCountSlot != -1)
 		return;
 
 	/*
@@ -271,16 +291,19 @@ ReservePrivateRefCountEntry(void)
 
 		for (i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
 		{
-			PrivateRefCountEntry *res;
-
-			res = &PrivateRefCountArray[i];
-
-			if (res->buffer == InvalidBuffer)
+			if (PrivateRefCountArrayKeys[i] == InvalidBuffer)
 			{
-				ReservedRefCountEntry = res;
-				return;
+				ReservedRefCountSlot = i;
+
+				/*
+				 * We could return immediately, but iterating till the end of
+				 * the array allows compiler-autovectorization.
+				 */
 			}
 		}
+
+		if (ReservedRefCountSlot != -1)
+			return;
 	}
 
 	/*
@@ -292,27 +315,36 @@ ReservePrivateRefCountEntry(void)
 		 * Move entry from the current clock position in the array into the
 		 * hashtable. Use that slot.
 		 */
+		int			victim_slot;
+		PrivateRefCountEntry *victim_entry;
 		PrivateRefCountEntry *hashent;
 		bool		found;
 
 		/* select victim slot */
-		ReservedRefCountEntry =
-			&PrivateRefCountArray[PrivateRefCountClock++ % REFCOUNT_ARRAY_ENTRIES];
+		victim_slot = PrivateRefCountClock++ % REFCOUNT_ARRAY_ENTRIES;
+		victim_entry = &PrivateRefCountArray[victim_slot];
+		ReservedRefCountSlot = victim_slot;
 
 		/* Better be used, otherwise we shouldn't get here. */
-		Assert(ReservedRefCountEntry->buffer != InvalidBuffer);
+		Assert(PrivateRefCountArrayKeys[victim_slot] != InvalidBuffer);
+		Assert(PrivateRefCountArray[victim_slot].buffer != InvalidBuffer);
+		Assert(PrivateRefCountArrayKeys[victim_slot] == PrivateRefCountArray[victim_slot].buffer);
 
 		/* enter victim array entry into hashtable */
 		hashent = hash_search(PrivateRefCountHash,
-							  &(ReservedRefCountEntry->buffer),
+							  &PrivateRefCountArrayKeys[victim_slot],
 							  HASH_ENTER,
 							  &found);
 		Assert(!found);
-		hashent->refcount = ReservedRefCountEntry->refcount;
+		hashent->data = victim_entry->data;
 
 		/* clear the now free array slot */
-		ReservedRefCountEntry->buffer = InvalidBuffer;
-		ReservedRefCountEntry->refcount = 0;
+		PrivateRefCountArrayKeys[victim_slot] = InvalidBuffer;
+		victim_entry->buffer = InvalidBuffer;
+
+		/* clear the whole data member, just for future proofing */
+		memset(&victim_entry->data, 0, sizeof(victim_entry->data));
+		victim_entry->data.refcount = 0;
 
 		PrivateRefCountOverflowed++;
 	}
@@ -327,15 +359,17 @@ NewPrivateRefCountEntry(Buffer buffer)
 	PrivateRefCountEntry *res;
 
 	/* only allowed to be called when a reservation has been made */
-	Assert(ReservedRefCountEntry != NULL);
+	Assert(ReservedRefCountSlot != -1);
 
 	/* use up the reserved entry */
-	res = ReservedRefCountEntry;
-	ReservedRefCountEntry = NULL;
+	res = &PrivateRefCountArray[ReservedRefCountSlot];
 
 	/* and fill it */
+	PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
 	res->buffer = buffer;
-	res->refcount = 0;
+	res->data.refcount = 0;
+
+	ReservedRefCountSlot = -1;
 
 	return res;
 }
@@ -347,10 +381,11 @@ NewPrivateRefCountEntry(Buffer buffer)
  * do_move is true, and the entry resides in the hashtable the entry is
  * optimized for frequent access by moving it to the array.
  */
-static PrivateRefCountEntry *
+static inline PrivateRefCountEntry *
 GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 {
 	PrivateRefCountEntry *res;
+	int			match = -1;
 	int			i;
 
 	Assert(BufferIsValid(buffer));
@@ -362,12 +397,16 @@ GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 	 */
 	for (i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
 	{
-		res = &PrivateRefCountArray[i];
-
-		if (res->buffer == buffer)
-			return res;
+		if (PrivateRefCountArrayKeys[i] == buffer)
+		{
+			match = i;
+			/* see ReservePrivateRefCountEntry() for why we don't return */
+		}
 	}
 
+	if (match != -1)
+		return &PrivateRefCountArray[match];
+
 	/*
 	 * By here we know that the buffer, if already pinned, isn't residing in
 	 * the array.
@@ -397,14 +436,18 @@ GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 		ReservePrivateRefCountEntry();
 
 		/* Use up the reserved slot */
-		Assert(ReservedRefCountEntry != NULL);
-		free = ReservedRefCountEntry;
-		ReservedRefCountEntry = NULL;
+		Assert(ReservedRefCountSlot != -1);
+		free = &PrivateRefCountArray[ReservedRefCountSlot];
+		Assert(PrivateRefCountArrayKeys[ReservedRefCountSlot] == free->buffer);
 		Assert(free->buffer == InvalidBuffer);
 
 		/* and fill it */
 		free->buffer = buffer;
-		free->refcount = res->refcount;
+		free->data = res->data;
+		PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
+
+		ReservedRefCountSlot = -1;
+
 
 		/* delete from hashtable */
 		hash_search(PrivateRefCountHash, &buffer, HASH_REMOVE, &found);
@@ -437,7 +480,7 @@ GetPrivateRefCount(Buffer buffer)
 
 	if (ref == NULL)
 		return 0;
-	return ref->refcount;
+	return ref->data.refcount;
 }
 
 /*
@@ -447,19 +490,21 @@ GetPrivateRefCount(Buffer buffer)
 static void
 ForgetPrivateRefCountEntry(PrivateRefCountEntry *ref)
 {
-	Assert(ref->refcount == 0);
+	Assert(ref->data.refcount == 0);
 
 	if (ref >= &PrivateRefCountArray[0] &&
 		ref < &PrivateRefCountArray[REFCOUNT_ARRAY_ENTRIES])
 	{
 		ref->buffer = InvalidBuffer;
+		PrivateRefCountArrayKeys[ref - PrivateRefCountArray] = InvalidBuffer;
+
 
 		/*
 		 * Mark the just used entry as reserved - in many scenarios that
 		 * allows us to avoid ever having to search the array/hash for free
 		 * entries.
 		 */
-		ReservedRefCountEntry = ref;
+		ReservedRefCountSlot = ref - PrivateRefCountArray;
 	}
 	else
 	{
@@ -3073,7 +3118,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 	PrivateRefCountEntry *ref;
 
 	Assert(!BufferIsLocal(b));
-	Assert(ReservedRefCountEntry != NULL);
+	Assert(ReservedRefCountSlot != -1);
 
 	ref = GetPrivateRefCountEntry(b, true);
 
@@ -3145,8 +3190,8 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		 */
 		result = (pg_atomic_read_u32(&buf->state) & BM_VALID) != 0;
 
-		Assert(ref->refcount > 0);
-		ref->refcount++;
+		Assert(ref->data.refcount > 0);
+		ref->data.refcount++;
 		ResourceOwnerRememberBuffer(CurrentResourceOwner, b);
 	}
 
@@ -3263,9 +3308,9 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	/* not moving as we're likely deleting it soon anyway */
 	ref = GetPrivateRefCountEntry(b, false);
 	Assert(ref != NULL);
-	Assert(ref->refcount > 0);
-	ref->refcount--;
-	if (ref->refcount == 0)
+	Assert(ref->data.refcount > 0);
+	ref->data.refcount--;
+	if (ref->data.refcount == 0)
 	{
 		uint32		old_buf_state;
 
@@ -3305,7 +3350,7 @@ TrackNewBufferPin(Buffer buf)
 	PrivateRefCountEntry *ref;
 
 	ref = NewPrivateRefCountEntry(buf);
-	ref->refcount++;
+	ref->data.refcount++;
 
 	ResourceOwnerRememberBuffer(CurrentResourceOwner, buf);
 
@@ -4018,6 +4063,7 @@ InitBufferManagerAccess(void)
 	MaxProportionalPins = NBuffers / (MaxBackends + NUM_AUXILIARY_PROCS);
 
 	memset(&PrivateRefCountArray, 0, sizeof(PrivateRefCountArray));
+	memset(&PrivateRefCountArrayKeys, 0, sizeof(PrivateRefCountArrayKeys));
 
 	hash_ctl.keysize = sizeof(int32);
 	hash_ctl.entrysize = sizeof(PrivateRefCountEntry);
@@ -4067,10 +4113,10 @@ CheckForBufferLeaks(void)
 	/* check the array */
 	for (i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
 	{
-		res = &PrivateRefCountArray[i];
-
-		if (res->buffer != InvalidBuffer)
+		if (PrivateRefCountArrayKeys[i] != InvalidBuffer)
 		{
+			res = &PrivateRefCountArray[i];
+
 			s = DebugPrintBufferRefcount(res->buffer);
 			elog(WARNING, "buffer refcount leak: %s", s);
 			pfree(s);
@@ -5407,7 +5453,7 @@ IncrBufferRefCount(Buffer buffer)
 
 		ref = GetPrivateRefCountEntry(buffer, true);
 		Assert(ref != NULL);
-		ref->refcount++;
+		ref->data.refcount++;
 	}
 	ResourceOwnerRememberBuffer(CurrentResourceOwner, buffer);
 }
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index abcb3c4c4cc..c4656bfe858 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2331,6 +2331,7 @@ PrintfArgValue
 PrintfTarget
 PrinttupAttrInfo
 PrivTarget
+PrivateRefCountData
 PrivateRefCountEntry
 ProcArrayStruct
 ProcLangInfo
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0006-bufmgr-Add-one-entry-cache-for-private-refcount.patch (4.9K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/7-v7-0006-bufmgr-Add-one-entry-cache-for-private-refcount.patch)
  download | inline diff:
From 847d3ed8f906a20558aa3b458d86eaf6f99cbb5f Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 18 Nov 2025 09:39:59 -0500
Subject: [PATCH v7 06/15] bufmgr: Add one-entry cache for private refcount

The private refcount entry for a buffer is often looked up repeatedly for the
same buffer, e.g. to pin and then unpin a buffer. Benchmarking shows that it's
worthwhile to have a one-entry cache for that case. With that cache in place,
it's worth splitting GetPrivateRefCountEntry() into a small inline
portion (for the cache hit case) and an out-of-line helper for the rest.

This is helpful for some workloads today, but becomes more important in an
upcoming patch that will utilize the private refcount infrastructure to also
store whether the buffer is currently locked, as that increases the rate of
lookups substantially.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar
---
 src/backend/storage/buffer/bufmgr.c | 66 ++++++++++++++++++++++++-----
 1 file changed, 55 insertions(+), 11 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 4e147d477c7..be32bd596f6 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -237,6 +237,7 @@ static HTAB *PrivateRefCountHash = NULL;
 static int32 PrivateRefCountOverflowed = 0;
 static uint32 PrivateRefCountClock = 0;
 static int	ReservedRefCountSlot = -1;
+static int	PrivateRefCountEntryLast = -1;
 
 static uint32 MaxProportionalPins;
 
@@ -369,28 +370,27 @@ NewPrivateRefCountEntry(Buffer buffer)
 	res->buffer = buffer;
 	res->data.refcount = 0;
 
+	/* update cache for the next lookup */
+	PrivateRefCountEntryLast = ReservedRefCountSlot;
+
 	ReservedRefCountSlot = -1;
 
 	return res;
 }
 
 /*
- * Return the PrivateRefCount entry for the passed buffer.
- *
- * Returns NULL if a buffer doesn't have a refcount entry. Otherwise, if
- * do_move is true, and the entry resides in the hashtable the entry is
- * optimized for frequent access by moving it to the array.
+ * Slow-path for GetPrivateRefCountEntry(). This is big enough to not be worth
+ * inlining. This particularly seems to be true if the compiler is capable of
+ * auto-vectorizing the code, as that imposes additional stack-alignment
+ * requirements etc.
  */
-static inline PrivateRefCountEntry *
-GetPrivateRefCountEntry(Buffer buffer, bool do_move)
+static pg_noinline PrivateRefCountEntry *
+GetPrivateRefCountEntrySlow(Buffer buffer, bool do_move)
 {
 	PrivateRefCountEntry *res;
 	int			match = -1;
 	int			i;
 
-	Assert(BufferIsValid(buffer));
-	Assert(!BufferIsLocal(buffer));
-
 	/*
 	 * First search for references in the array, that'll be sufficient in the
 	 * majority of cases.
@@ -404,8 +404,13 @@ GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 		}
 	}
 
-	if (match != -1)
+	if (likely(match != -1))
+	{
+		/* update cache for the next lookup */
+		PrivateRefCountEntryLast = match;
+
 		return &PrivateRefCountArray[match];
+	}
 
 	/*
 	 * By here we know that the buffer, if already pinned, isn't residing in
@@ -445,6 +450,8 @@ GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 		free->buffer = buffer;
 		free->data = res->data;
 		PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
+		/* update cache for the next lookup */
+		PrivateRefCountEntryLast = match;
 
 		ReservedRefCountSlot = -1;
 
@@ -459,6 +466,43 @@ GetPrivateRefCountEntry(Buffer buffer, bool do_move)
 	}
 }
 
+/*
+ * Return the PrivateRefCount entry for the passed buffer.
+ *
+ * Returns NULL if a buffer doesn't have a refcount entry. Otherwise, if
+ * do_move is true, and the entry resides in the hashtable the entry is
+ * optimized for frequent access by moving it to the array.
+ */
+static inline PrivateRefCountEntry *
+GetPrivateRefCountEntry(Buffer buffer, bool do_move)
+{
+	Assert(BufferIsValid(buffer));
+	Assert(!BufferIsLocal(buffer));
+
+	/*
+	 * It's very common to look up the same buffer repeatedly. To make that
+	 * fast, we have a one-entry cache.
+	 *
+	 * In contrast to the loop below, here it faster to check
+	 * PrivateRefCountArray[].buffer, as in the case of a hit, as fewer
+	 * addresses are computed and fewer cachelines are accessed. Whereas in
+	 * the loop case below, checking PrivateRefCountArrayKeys saves a lot of
+	 * memory accesses.
+	 */
+	if (likely(PrivateRefCountEntryLast != -1) &&
+		likely(PrivateRefCountArray[PrivateRefCountEntryLast].buffer == buffer))
+	{
+		return &PrivateRefCountArray[PrivateRefCountEntryLast];
+	}
+
+	/*
+	 * The code for the cached lookup is small enough to be worth inlining
+	 * into the caller. In the miss case however, that empirically doesn't
+	 * seem worth it.
+	 */
+	return GetPrivateRefCountEntrySlow(buffer, do_move);
+}
+
 /*
  * Returns how many times the passed buffer is pinned by this backend.
  *
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0007-freespace-Don-t-modify-page-without-any-lock.patch (2.0K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/8-v7-0007-freespace-Don-t-modify-page-without-any-lock.patch)
  download | inline diff:
From 02a92ff1d7ef354b44fd5fc1ee3f9016ce88f905 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 1 Dec 2025 22:31:44 -0500
Subject: [PATCH v7 07/15] freespace: Don't modify page without any lock

Before this commit fsm_vacuum_page() modified the page without any lock on the
page. Historically that was kind of ok, as we didn't rely on the freespace to
really stay consistent and we did not have checksums. But these days pages are
checksummed and there are ways for FSM pages to be included in WAL records,
even if the FSM itself is still not WAL logged. If a FSM page ever were
modified while a WAL record referenced that page, we'd be in trouble, as the
WAL CRC could end up getting corrupted.

The reason to address this right now is a series of patches with the goal to
only allow modifications of pages with an appropriate lock level. Obviously
not having any lock is not appropriate :)

Discussion: https://postgr.es/m/4wggb7purufpto6x35fd2kwhasehnzfdy3zdcu47qryubs2hdz@fa5kannykekr
Discussion: https://postgr.es/m/e6a8f734-2198-4958-a028-aba863d4a204@iki.fi
---
 src/backend/storage/freespace/freespace.c | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index 4773a9cc65e..48ac15d3487 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -906,10 +906,12 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 	/*
 	 * Reset the next slot pointer. This encourages the use of low-numbered
 	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation.  We don't bother with a lock here, nor with marking the page
-	 * dirty if it wasn't already, since this is just a hint.
+	 * relation. We don't bother with marking the page dirty if it wasn't
+	 * already, since this is just a hint.
 	 */
+	LockBuffer(buf, BUFFER_LOCK_SHARE);
 	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 	ReleaseBuffer(buf);
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0008-heapam-Move-logic-to-handle-HEAP_MOVED-into-a-hel.patch (11.5K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/9-v7-0008-heapam-Move-logic-to-handle-HEAP_MOVED-into-a-hel.patch)
  download | inline diff:
From d74eda121de588dc1cc7be453ea2c979e04e6d27 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 23 Sep 2024 12:23:33 -0400
Subject: [PATCH v7 08/15] heapam: Move logic to handle HEAP_MOVED into a
 helper function

Before we dealt with this in 6 near identical and one very similar copy.

The helper function errors out when encountering a
HEAP_MOVED_IN/HEAP_MOVED_OUT tuple with xvac considered current or
in-progress. It'd be preferrable to do that change separately, but otherwise
it'd not be possible to deduplicate the handling in
HeapTupleSatisfiesVacuum().

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/heap/heapam_visibility.c | 307 ++++----------------
 1 file changed, 61 insertions(+), 246 deletions(-)

diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 05f6946fe60..4fefcbca5f5 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -144,6 +144,55 @@ HeapTupleSetHintBits(HeapTupleHeader tuple, Buffer buffer,
 	SetHintBits(tuple, buffer, infomask, xid);
 }
 
+/*
+ * If HEAP_MOVED_OFF or HEAP_MOVED_IN are set on the tuple, remove them and
+ * adjust hint bits. See the comment for SetHintBits() for more background.
+ *
+ * This helper returns false if the row ought to be invisible, true otherwise.
+ */
+static inline bool
+HeapTupleCleanMoved(HeapTupleHeader tuple, Buffer buffer)
+{
+	TransactionId xvac;
+
+	/* only used by pre-9.0 binary upgrades */
+	if (likely(!(tuple->t_infomask & (HEAP_MOVED_OFF | HEAP_MOVED_IN))))
+		return true;
+
+	xvac = HeapTupleHeaderGetXvac(tuple);
+
+	if (TransactionIdIsCurrentTransactionId(xvac))
+		elog(ERROR, "encountered tuple with HEAP_MOVED considered current");
+
+	if (TransactionIdIsInProgress(xvac))
+		elog(ERROR, "encountered tuple with HEAP_MOVED considered in-progress");
+
+	if (tuple->t_infomask & HEAP_MOVED_OFF)
+	{
+		if (TransactionIdDidCommit(xvac))
+		{
+			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
+						InvalidTransactionId);
+			return false;
+		}
+		SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
+					InvalidTransactionId);
+	}
+	else if (tuple->t_infomask & HEAP_MOVED_IN)
+	{
+		if (TransactionIdDidCommit(xvac))
+			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
+						InvalidTransactionId);
+		else
+		{
+			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
+						InvalidTransactionId);
+			return false;
+		}
+	}
+
+	return true;
+}
 
 /*
  * HeapTupleSatisfiesSelf
@@ -179,45 +228,8 @@ HeapTupleSatisfiesSelf(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (tuple->t_infomask & HEAP_XMAX_INVALID)	/* xid invalid */
@@ -372,45 +384,8 @@ HeapTupleSatisfiesToast(HeapTuple htup, Snapshot snapshot,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 
 		/*
 		 * An invalid Xmin can be left behind by a speculative insertion that
@@ -468,45 +443,8 @@ HeapTupleSatisfiesUpdate(HeapTuple htup, CommandId curcid,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return TM_Invisible;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return TM_Invisible;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return TM_Invisible;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return TM_Invisible;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return TM_Invisible;
-				}
-			}
-		}
+		else if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (HeapTupleHeaderGetCmin(tuple) >= curcid)
@@ -756,45 +694,8 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (tuple->t_infomask & HEAP_XMAX_INVALID)	/* xid invalid */
@@ -979,45 +880,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!XidInMVCCSnapshot(xvac, snapshot))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (XidInMVCCSnapshot(xvac, snapshot))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (HeapTupleHeaderGetCmin(tuple) >= snapshot->curcid)
@@ -1222,57 +1086,8 @@ HeapTupleSatisfiesVacuumHorizon(HeapTuple htup, Buffer buffer, TransactionId *de
 	{
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return HEAPTUPLE_DEAD;
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return HEAPTUPLE_DELETE_IN_PROGRESS;
-			if (TransactionIdIsInProgress(xvac))
-				return HEAPTUPLE_DELETE_IN_PROGRESS;
-			if (TransactionIdDidCommit(xvac))
-			{
-				SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-							InvalidTransactionId);
-				return HEAPTUPLE_DEAD;
-			}
-			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-						InvalidTransactionId);
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return HEAPTUPLE_INSERT_IN_PROGRESS;
-			if (TransactionIdIsInProgress(xvac))
-				return HEAPTUPLE_INSERT_IN_PROGRESS;
-			if (TransactionIdDidCommit(xvac))
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			else
-			{
-				SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-							InvalidTransactionId);
-				return HEAPTUPLE_DEAD;
-			}
-		}
-		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
-		{
-			if (tuple->t_infomask & HEAP_XMAX_INVALID)	/* xid invalid */
-				return HEAPTUPLE_INSERT_IN_PROGRESS;
-			/* only locked? run infomask-only check first, for performance */
-			if (HEAP_XMAX_IS_LOCKED_ONLY(tuple->t_infomask) ||
-				HeapTupleHeaderIsOnlyLocked(tuple))
-				return HEAPTUPLE_INSERT_IN_PROGRESS;
-			/* inserted and then deleted by same xact */
-			if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetUpdateXid(tuple)))
-				return HEAPTUPLE_DELETE_IN_PROGRESS;
-			/* deleting subtransaction must have aborted */
-			return HEAPTUPLE_INSERT_IN_PROGRESS;
-		}
+		else if (!HeapTupleCleanMoved(tuple, buffer))
+			return HEAPTUPLE_DEAD;
 		else if (TransactionIdIsInProgress(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			/*
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0009-heapam-Use-exclusive-lock-on-old-page-in-CLUSTER.patch (2.7K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/10-v7-0009-heapam-Use-exclusive-lock-on-old-page-in-CLUSTER.patch)
  download | inline diff:
From fb82b00d1e5c595222bd7d5754d0bd41b2631bbc Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sun, 26 Jan 2025 15:18:46 -0500
Subject: [PATCH v7 09/15] heapam: Use exclusive lock on old page in CLUSTER

To be able to guarantee that we can set the hint bit, acquire an exclusive
lock on the old buffer. We need the hint bits to be set as otherwise
reform_and_rewrite_tuple() -> rewrite_heap_tuple() -> heap_freeze_tuple() will
get confused.

It'd be better if we somehow could avoid setting hint bits on the old page. A
commonreason to use VACUUM FULL are very bloated tables - rewriting most of
the old table before during VACUUM FULL doesn't exactly help.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/heap/heapam_handler.c    | 13 ++++++++++++-
 src/backend/access/heap/heapam_visibility.c |  7 +++++++
 2 files changed, 19 insertions(+), 1 deletion(-)

diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index bcbac844bb6..f84254f0737 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -837,7 +837,18 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		tuple = ExecFetchSlotHeapTuple(slot, false, NULL);
 		buf = hslot->buffer;
 
-		LockBuffer(buf, BUFFER_LOCK_SHARE);
+		/*
+		 * To be able to guarantee that we can set the hint bit, acquire an
+		 * exclusive lock on the old buffer. We need the hint bits to be set
+		 * as otherwise reform_and_rewrite_tuple() -> rewrite_heap_tuple() ->
+		 * heap_freeze_tuple() will get confused.
+		 *
+		 * It'd be better if we somehow could avoid setting hint bits on the
+		 * old page. One reason to use VACUUM FULL are very bloated tables -
+		 * rewriting most of the old table before during VACUUM FULL doesn't
+		 * exactly help...
+		 */
+		LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
 		switch (HeapTupleSatisfiesVacuum(tuple, OldestXmin, buf))
 		{
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 4fefcbca5f5..762538a2040 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -141,6 +141,13 @@ void
 HeapTupleSetHintBits(HeapTupleHeader tuple, Buffer buffer,
 					 uint16 infomask, TransactionId xid)
 {
+	/*
+	 * The uses from heapam.c rely on being able to perform the hint bit
+	 * updates, which can only be guaranteed if we are holding an exclusive
+	 * lock on the buffer - which all callers are doing.
+	 */
+	Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+
 	SetHintBits(tuple, buffer, infomask, xid);
 }
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0010-heapam-Add-batch-mode-mvcc-check-and-use-it-in-pa.patch (7.6K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/11-v7-0010-heapam-Add-batch-mode-mvcc-check-and-use-it-in-pa.patch)
  download | inline diff:
From 2c75985744a2772e9532800bf7af43b05d899892 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 17 Oct 2024 13:16:36 -0400
Subject: [PATCH v7 10/15] heapam: Add batch mode mvcc check and use it in page
 mode

There are two reasons for doing so:

1) It is generally faster to perform checks in a batched fashion and making
   sequential scans faster is nice.

2) We would like to stop setting hint bits while pages are being written
   out. The necessary locking becomes visible for page mode scans if done for
   every tuple. With batching the overhead can be amortized to only happen
   once per page.

There are substantial further optimization opportunities along these
lines:

- Right now HeapTupleSatisfiesMVCCBatch() simply uses the single-tuple
  HeapTupleSatisfiesMVCC(), relying on the compiler to inline it. We could
  instead write an explicitly optimized version that avoids repeated xid
  tests.

- Introduce batched version of the serializability test

- Introduce batched version of HeapTupleSatisfiesVacuum

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/access/heapam.h                 | 28 +++++++
 src/backend/access/heap/heapam.c            | 91 ++++++++++++++++-----
 src/backend/access/heap/heapam_visibility.c | 47 +++++++++++
 src/tools/pgindent/typedefs.list            |  1 +
 4 files changed, 147 insertions(+), 20 deletions(-)

diff --git a/src/include/access/heapam.h b/src/include/access/heapam.h
index 632c4332a8c..3528e6e4b59 100644
--- a/src/include/access/heapam.h
+++ b/src/include/access/heapam.h
@@ -449,6 +449,34 @@ extern bool HeapTupleHeaderIsOnlyLocked(HeapTupleHeader tuple);
 extern bool HeapTupleIsSurelyDead(HeapTuple htup,
 								  GlobalVisState *vistest);
 
+/*
+ * FIXME: define to be removed
+ *
+ * Without this I see worse performance. But it's a bit ugly, so I thought
+ * it'd be useful to leave a way in for others to experiment with this.
+ */
+#define BATCHMVCC_FEWER_ARGS
+
+#ifdef BATCHMVCC_FEWER_ARGS
+typedef struct BatchMVCCState
+{
+	HeapTupleData tuples[MaxHeapTuplesPerPage];
+	bool		visible[MaxHeapTuplesPerPage];
+} BatchMVCCState;
+#endif
+
+extern int	HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
+										int ntups,
+#ifdef BATCHMVCC_FEWER_ARGS
+										BatchMVCCState *batchmvcc,
+#else
+										HeapTupleData *tuples,
+										bool *visible,
+#endif
+										OffsetNumber *vistuples_dense);
+
+
+
 /*
  * To avoid leaking too much knowledge about reorderbuffer implementation
  * details this is implemented in reorderbuffer.c not heapam_visibility.c
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 4d382a04338..7e56da9b196 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -519,42 +519,93 @@ page_collect_tuples(HeapScanDesc scan, Snapshot snapshot,
 					BlockNumber block, int lines,
 					bool all_visible, bool check_serializable)
 {
+	Oid			relid = RelationGetRelid(scan->rs_base.rs_rd);
+#ifdef BATCHMVCC_FEWER_ARGS
+	BatchMVCCState batchmvcc;
+	HeapTupleData *tuples = batchmvcc.tuples;
+	bool	   *visible = batchmvcc.visible;
+#else
+	HeapTupleData tuples[MaxHeapTuplesPerPage];
+	bool		visible[MaxHeapTuplesPerPage];
+#endif
 	int			ntup = 0;
-	OffsetNumber lineoff;
+	int			nvis = 0;
 
-	for (lineoff = FirstOffsetNumber; lineoff <= lines; lineoff++)
+	/* page at a time should have been disabled otherwise */
+	Assert(IsMVCCSnapshot(snapshot));
+
+	/* first find all tuples on the page */
+	for (OffsetNumber lineoff = FirstOffsetNumber; lineoff <= lines; lineoff++)
 	{
 		ItemId		lpp = PageGetItemId(page, lineoff);
-		HeapTupleData loctup;
-		bool		valid;
+		HeapTuple	tup;
 
-		if (!ItemIdIsNormal(lpp))
+		if (unlikely(!ItemIdIsNormal(lpp)))
 			continue;
 
-		loctup.t_data = (HeapTupleHeader) PageGetItem(page, lpp);
-		loctup.t_len = ItemIdGetLength(lpp);
-		loctup.t_tableOid = RelationGetRelid(scan->rs_base.rs_rd);
-		ItemPointerSet(&(loctup.t_self), block, lineoff);
+		/*
+		 * If the page is not all-visible or we need to check serializability,
+		 * maintain enough state to be able to refind the tuple efficiently,
+		 * without again needing to extract it from the page.
+		 */
+		if (!all_visible || check_serializable)
+		{
+			tup = &tuples[ntup];
 
+			tup->t_data = (HeapTupleHeader) PageGetItem(page, lpp);
+			tup->t_len = ItemIdGetLength(lpp);
+			tup->t_tableOid = relid;
+			ItemPointerSet(&(tup->t_self), block, lineoff);
+		}
+
+		/*
+		 * If the page is all visible, these fields otherwise won't be
+		 * populated in loop below.
+		 */
 		if (all_visible)
-			valid = true;
-		else
-			valid = HeapTupleSatisfiesVisibility(&loctup, snapshot, buffer);
-
-		if (check_serializable)
-			HeapCheckForSerializableConflictOut(valid, scan->rs_base.rs_rd,
-												&loctup, buffer, snapshot);
-
-		if (valid)
 		{
+			if (check_serializable)
+			{
+				visible[ntup] = true;
+			}
 			scan->rs_vistuples[ntup] = lineoff;
-			ntup++;
 		}
+
+		ntup++;
 	}
 
 	Assert(ntup <= MaxHeapTuplesPerPage);
 
-	return ntup;
+	/* unless the page is all visible, test visibility for all tuples one go */
+	if (all_visible)
+		nvis = ntup;
+	else
+		nvis = HeapTupleSatisfiesMVCCBatch(snapshot, buffer,
+										   ntup,
+#ifdef BATCHMVCC_FEWER_ARGS
+										   &batchmvcc,
+#else
+										   tuples, visible,
+#endif
+										   scan->rs_vistuples
+			);
+
+	/*
+	 * So far we don't have batch API for testing serializabilty, so do so
+	 * one-by-one.
+	 */
+	if (check_serializable)
+	{
+		for (int i = 0; i < ntup; i++)
+		{
+			HeapCheckForSerializableConflictOut(visible[i],
+												scan->rs_base.rs_rd,
+												&tuples[i],
+												buffer, snapshot);
+		}
+	}
+
+	return nvis;
 }
 
 /*
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 762538a2040..5645cfd8a49 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -1584,6 +1584,53 @@ HeapTupleSatisfiesHistoricMVCC(HeapTuple htup, Snapshot snapshot,
 		return true;
 }
 
+/*
+ * Perform HeaptupleSatisfiesMVCC() on each passed in tuple. This is more
+ * efficient than doing HeapTupleSatisfiesMVCC() one-by-one.
+ *
+ * To be checked tuples are passed via BatchMVCCState->tuples. Each tuple's
+ * visibility is set in batchmvcc->visible[]. In addition, ->vistuples_dense
+ * is set to contain the offsets of visible tuples.
+ *
+ * Returns the number of visible tuples.
+ */
+int
+HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
+							int ntups,
+#ifdef BATCHMVCC_FEWER_ARGS
+							BatchMVCCState *batchmvcc,
+#else
+							HeapTupleData *tuples,
+							bool *visible,
+#endif
+							OffsetNumber *vistuples_dense)
+{
+	int			nvis = 0;
+#ifdef BATCHMVCC_FEWER_ARGS
+	HeapTupleData *tuples = batchmvcc->tuples;
+	bool	   *visible = batchmvcc->visible;
+#endif
+
+	Assert(IsMVCCSnapshot(snapshot));
+
+	for (int i = 0; i < ntups; i++)
+	{
+		bool		valid;
+		HeapTuple	tup = &tuples[i];
+
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		visible[i] = valid;
+
+		if (likely(valid))
+		{
+			vistuples_dense[nvis] = tup->t_self.ip_posid;
+			nvis++;
+		}
+	}
+
+	return nvis;
+}
+
 /*
  * HeapTupleSatisfiesVisibility
  *		True iff heap tuple satisfies a time qual.
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index c4656bfe858..6641a988ae8 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -249,6 +249,7 @@ Barrier
 BaseBackupCmd
 BaseBackupTargetHandle
 BaseBackupTargetType
+BatchMVCCState
 BeginDirectModify_function
 BeginForeignInsert_function
 BeginForeignModify_function
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0011-bufmgr-Change-BufferDesc.state-to-be-a-64bit-atom.patch (46.6K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/12-v7-0011-bufmgr-Change-BufferDesc.state-to-be-a-64bit-atom.patch)
  download | inline diff:
From 187dc02e6cce69dcccc0b30771ba55e5805b14c5 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 2 Dec 2025 18:46:58 -0500
Subject: [PATCH v7 11/15] bufmgr: Change BufferDesc.state to be a 64bit atomic

This is motivated by wanting to merge buffer content locks into
BufferDesc.state in a future commit, rather than having a separate lwlock (see
commit c75ebc657ff more details). As this change is rather mechanical, it
seems to make sense to split it out into a separate commit, for easier review.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  88 +++++----
 src/include/storage/procnumber.h              |  14 +-
 src/backend/storage/buffer/buf_init.c         |   2 +-
 src/backend/storage/buffer/bufmgr.c           | 170 +++++++++---------
 src/backend/storage/buffer/freelist.c         |  24 +--
 src/backend/storage/buffer/localbuf.c         |  72 ++++----
 contrib/pg_buffercache/pg_buffercache_pages.c |   8 +-
 src/test/modules/test_aio/test_aio.c          |  12 +-
 8 files changed, 205 insertions(+), 185 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 5400c56a965..28519ad2813 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -30,7 +30,7 @@
 #include "utils/resowner.h"
 
 /*
- * Buffer state is a single 32-bit variable where following data is combined.
+ * Buffer state is a single 64-bit variable where following data is combined.
  *
  * - 18 bits refcount
  * - 4 bits usage count
@@ -39,6 +39,9 @@
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
+ * NB: A future commit will use a significant portion of the remaining bits to
+ * implement buffer locking as part of the state variable.
+ *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
@@ -49,15 +52,21 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 				 "parts of buffer state space need to equal 32");
 
 #define BUF_REFCOUNT_ONE 1
-#define BUF_REFCOUNT_MASK ((1U << BUF_REFCOUNT_BITS) - 1)
-#define BUF_USAGECOUNT_MASK (((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_REFCOUNT_BITS))
-#define BUF_USAGECOUNT_ONE (1U << BUF_REFCOUNT_BITS)
+#define BUF_REFCOUNT_MASK \
+	((UINT64CONST(1) << BUF_REFCOUNT_BITS) - 1)
+#define BUF_USAGECOUNT_MASK \
+	(((UINT64CONST(1) << BUF_USAGECOUNT_BITS) - 1) << (BUF_REFCOUNT_BITS))
+#define BUF_USAGECOUNT_ONE \
+	(UINT64CONST(1) << BUF_REFCOUNT_BITS)
 #define BUF_USAGECOUNT_SHIFT BUF_REFCOUNT_BITS
-#define BUF_FLAG_MASK (((1U << BUF_FLAG_BITS) - 1) << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS))
+#define BUF_FLAG_MASK \
+	(((UINT64CONST(1) << BUF_FLAG_BITS) - 1) << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS))
 
 /* Get refcount and usagecount from buffer state */
-#define BUF_STATE_GET_REFCOUNT(state) ((state) & BUF_REFCOUNT_MASK)
-#define BUF_STATE_GET_USAGECOUNT(state) (((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
+#define BUF_STATE_GET_REFCOUNT(state) \
+	((uint32)((state) & BUF_REFCOUNT_MASK))
+#define BUF_STATE_GET_USAGECOUNT(state) \
+	((uint32)(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT))
 
 /*
  * Flags for buffer descriptors
@@ -65,17 +74,28 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  * Note: BM_TAG_VALID essentially means that there is a buffer hashtable
  * entry associated with the buffer's tag.
  */
-#define BM_LOCKED				(1U << 22)	/* buffer header is locked */
-#define BM_DIRTY				(1U << 23)	/* data needs writing */
-#define BM_VALID				(1U << 24)	/* data is valid */
-#define BM_TAG_VALID			(1U << 25)	/* tag is assigned */
-#define BM_IO_IN_PROGRESS		(1U << 26)	/* read or write in progress */
-#define BM_IO_ERROR				(1U << 27)	/* previous I/O failed */
-#define BM_JUST_DIRTIED			(1U << 28)	/* dirtied since write started */
-#define BM_PIN_COUNT_WAITER		(1U << 29)	/* have waiter for sole pin */
-#define BM_CHECKPOINT_NEEDED	(1U << 30)	/* must write for checkpoint */
-#define BM_PERMANENT			(1U << 31)	/* permanent buffer (not unlogged,
-											 * or init fork) */
+
+/* buffer header is locked */
+#define BM_LOCKED				(UINT64CONST(1) << 22)
+/* data needs writing */
+#define BM_DIRTY				(UINT64CONST(1) << 23)
+/* data is valid */
+#define BM_VALID				(UINT64CONST(1) << 24)
+/* tag is assigned */
+#define BM_TAG_VALID			(UINT64CONST(1) << 25)
+/* read or write in progress */
+#define BM_IO_IN_PROGRESS		(UINT64CONST(1) << 26)
+/* previous I/O failed */
+#define BM_IO_ERROR				(UINT64CONST(1) << 27)
+/* dirtied since write started */
+#define BM_JUST_DIRTIED			(UINT64CONST(1) << 28)
+/* have waiter for sole pin */
+#define BM_PIN_COUNT_WAITER		(UINT64CONST(1) << 29)
+/* must write for checkpoint */
+#define BM_CHECKPOINT_NEEDED	(UINT64CONST(1) << 30)
+/* permanent buffer (not unlogged, or init fork) */
+#define BM_PERMANENT			(UINT64CONST(1) << 31)
+
 /*
  * The maximum allowed value of usage_count represents a tradeoff between
  * accuracy and speed of the clock-sweep buffer management algorithm.  A
@@ -86,7 +106,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  */
 #define BM_MAX_USAGE_COUNT	5
 
-StaticAssertDecl(BM_MAX_USAGE_COUNT < (1 << BUF_USAGECOUNT_BITS),
+StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 				 "BM_MAX_USAGE_COUNT doesn't fit in BUF_USAGECOUNT_BITS bits");
 StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
 				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
@@ -251,8 +271,8 @@ BufMappingPartitionLockByIndex(uint32 index)
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
- * atomic operations (i.e. only pg_atomic_read_u32() and
- * pg_atomic_unlocked_write_u32()).
+ * atomic operations (i.e. only pg_atomic_read_u64() and
+ * pg_atomic_unlocked_write_u64()).
  *
  * Be careful to avoid increasing the size of the struct when adding or
  * reordering members.  Keeping it below 64 bytes (the most common CPU
@@ -280,7 +300,7 @@ typedef struct BufferDesc
 	 * State of the buffer, containing flags, refcount and usagecount. See
 	 * BUF_* and BM_* defines at the top of this file.
 	 */
-	pg_atomic_uint32 state;
+	pg_atomic_uint64 state;
 
 	/*
 	 * Backend of pin-count waiter. The buffer header spinlock needs to be
@@ -386,7 +406,7 @@ BufferDescriptorGetContentLock(const BufferDesc *bdesc)
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
  */
-extern uint32 LockBufHdr(BufferDesc *desc);
+extern uint64 LockBufHdr(BufferDesc *desc);
 
 /*
  * Unlock the buffer header.
@@ -397,9 +417,9 @@ extern uint32 LockBufHdr(BufferDesc *desc);
 static inline void
 UnlockBufHdr(BufferDesc *desc)
 {
-	Assert(pg_atomic_read_u32(&desc->state) & BM_LOCKED);
+	Assert(pg_atomic_read_u64(&desc->state) & BM_LOCKED);
 
-	pg_atomic_fetch_sub_u32(&desc->state, BM_LOCKED);
+	pg_atomic_fetch_sub_u64(&desc->state, BM_LOCKED);
 }
 
 /*
@@ -410,14 +430,14 @@ UnlockBufHdr(BufferDesc *desc)
  * Note that this approach would not work for usagecount, since we need to cap
  * the usagecount at BM_MAX_USAGE_COUNT.
  */
-static inline uint32
-UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
-				uint32 set_bits, uint32 unset_bits,
+static inline uint64
+UnlockBufHdrExt(BufferDesc *desc, uint64 old_buf_state,
+				uint64 set_bits, uint64 unset_bits,
 				int refcount_change)
 {
 	for (;;)
 	{
-		uint32		buf_state = old_buf_state;
+		uint64		buf_state = old_buf_state;
 
 		Assert(buf_state & BM_LOCKED);
 
@@ -428,7 +448,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 		if (refcount_change != 0)
 			buf_state += BUF_REFCOUNT_ONE * refcount_change;
 
-		if (pg_atomic_compare_exchange_u32(&desc->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&desc->state, &old_buf_state,
 										   buf_state))
 		{
 			return old_buf_state;
@@ -436,7 +456,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 	}
 }
 
-extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
+extern uint64 WaitBufHdrUnlocked(BufferDesc *buf);
 
 /* in bufmgr.c */
 
@@ -496,14 +516,14 @@ extern void TrackNewBufferPin(Buffer buf);
 
 /* solely to make it easier to write tests */
 extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
-extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 							  bool forget_owner, bool release_aio);
 
 
 /* freelist.c */
 extern IOContext IOContextForStrategy(BufferAccessStrategy strategy);
 extern BufferDesc *StrategyGetBuffer(BufferAccessStrategy strategy,
-									 uint32 *buf_state, bool *from_ring);
+									 uint64 *buf_state, bool *from_ring);
 extern bool StrategyRejectBuffer(BufferAccessStrategy strategy,
 								 BufferDesc *buf, bool from_ring);
 
@@ -539,7 +559,7 @@ extern BlockNumber ExtendBufferedRelLocal(BufferManagerRelation bmr,
 										  uint32 *extended_by);
 extern void MarkLocalBufferDirty(Buffer buffer);
 extern void TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty,
-								   uint32 set_flag_bits, bool release_aio);
+								   uint64 set_flag_bits, bool release_aio);
 extern bool StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait);
 extern void FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln);
 extern void InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced);
diff --git a/src/include/storage/procnumber.h b/src/include/storage/procnumber.h
index 2ddaaf0c646..6baac7c77f1 100644
--- a/src/include/storage/procnumber.h
+++ b/src/include/storage/procnumber.h
@@ -27,13 +27,13 @@ typedef int ProcNumber;
 
 /*
  * Note: MAX_BACKENDS_BITS is 18 as that is the space available for buffer
- * refcounts in buf_internals.h.  This limitation could be lifted by using a
- * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
- * currently realistic configurations. Even if that limitation were removed,
- * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
- * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
- * 4*MaxBackends without any overflow check.  We check that the configured
- * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
+ * refcounts in buf_internals.h.  This limitation could be lifted, but it's
+ * unlikely to be worthwhile as 2^18-1 backends exceed currently realistic
+ * configurations. Even if that limitation were removed, we still could not a)
+ * exceed 2^23-1 because inval.c stores the ProcNumber as a 3-byte signed
+ * integer, b) INT_MAX/4 because some places compute 4*MaxBackends without any
+ * overflow check.  We check that the configured number of backends does not
+ * exceed MAX_BACKENDS in InitializeMaxBackends().
  */
 #define MAX_BACKENDS_BITS		18
 #define MAX_BACKENDS			((1U << MAX_BACKENDS_BITS)-1)
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6fd3a6bbac5..25f71191ec3 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -121,7 +121,7 @@ BufferManagerShmemInit(void)
 
 			ClearBufferTag(&buf->tag);
 
-			pg_atomic_init_u32(&buf->state, 0);
+			pg_atomic_init_u64(&buf->state, 0);
 			buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
 
 			buf->buf_id = i;
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index be32bd596f6..d0b8f8d20eb 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -775,7 +775,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 {
 	BufferDesc *bufHdr;
 	BufferTag	tag;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(recent_buffer));
 
@@ -788,7 +788,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 		int			b = -recent_buffer - 1;
 
 		bufHdr = GetLocalBufferDescriptor(b);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		/* Is it still valid and holding the right tag? */
 		if ((buf_state & BM_VALID) && BufferTagsEqual(&tag, &bufHdr->tag))
@@ -1381,8 +1381,8 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
 				bufHdr = GetLocalBufferDescriptor(-buffers[i] - 1);
 			else
 				bufHdr = GetBufferDescriptor(buffers[i] - 1);
-			Assert(pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID);
-			found = pg_atomic_read_u32(&bufHdr->state) & BM_VALID;
+			Assert(pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID);
+			found = pg_atomic_read_u64(&bufHdr->state) & BM_VALID;
 		}
 		else
 		{
@@ -1608,10 +1608,10 @@ CheckReadBuffersOperation(ReadBuffersOperation *operation, bool is_complete)
 			GetBufferDescriptor(buffer - 1);
 
 		Assert(BufferGetBlockNumber(buffer) == operation->blocknum + i);
-		Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_TAG_VALID);
+		Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_TAG_VALID);
 
 		if (i < operation->nblocks_done)
-			Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_VALID);
+			Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_VALID);
 	}
 #endif
 }
@@ -2078,8 +2078,8 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	int			existing_buf_id;
 	Buffer		victim_buffer;
 	BufferDesc *victim_buf_hdr;
-	uint32		victim_buf_state;
-	uint32		set_bits = 0;
+	uint64		victim_buf_state;
+	uint64		set_bits = 0;
 
 	/* Make sure we will have room to remember the buffer pin */
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2246,7 +2246,7 @@ InvalidateBuffer(BufferDesc *buf)
 	uint32		oldHash;		/* hash value for oldTag */
 	LWLock	   *oldPartitionLock;	/* buffer partition lock for it */
 	uint32		oldFlags;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/* Save the original buffer tag before dropping the spinlock */
 	oldTag = buf->tag;
@@ -2337,7 +2337,7 @@ retry:
 static bool
 InvalidateVictimBuffer(BufferDesc *buf_hdr)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	uint32		hash;
 	LWLock	   *partition_lock;
 	BufferTag	tag;
@@ -2397,10 +2397,10 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 
 	LWLockRelease(partition_lock);
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	Assert(!(buf_state & (BM_DIRTY | BM_VALID | BM_TAG_VALID)));
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u32(&buf_hdr->state)) > 0);
+	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u64(&buf_hdr->state)) > 0);
 
 	return true;
 }
@@ -2410,7 +2410,7 @@ GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context)
 {
 	BufferDesc *buf_hdr;
 	Buffer		buf;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		from_ring;
 
 	/*
@@ -2543,7 +2543,7 @@ again:
 
 	/* a final set of sanity checks */
 #ifdef USE_ASSERT_CHECKING
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) == 1);
 	Assert(!(buf_state & (BM_TAG_VALID | BM_VALID | BM_DIRTY)));
@@ -2834,13 +2834,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 */
 			do
 			{
-				pg_atomic_fetch_and_u32(&existing_hdr->state, ~BM_VALID);
+				pg_atomic_fetch_and_u64(&existing_hdr->state, ~BM_VALID);
 			} while (!StartBufferIO(existing_hdr, true, false));
 		}
 		else
 		{
-			uint32		buf_state;
-			uint32		set_bits = 0;
+			uint64		buf_state;
+			uint64		set_bits = 0;
 
 			buf_state = LockBufHdr(victim_buf_hdr);
 
@@ -3016,7 +3016,7 @@ BufferIsDirty(Buffer buffer)
 		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 	}
 
-	return pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY;
+	return pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY;
 }
 
 /*
@@ -3032,8 +3032,8 @@ void
 MarkBufferDirty(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
-	uint32		old_buf_state;
+	uint64		buf_state;
+	uint64		old_buf_state;
 
 	if (!BufferIsValid(buffer))
 		elog(ERROR, "bad buffer ID: %d", buffer);
@@ -3053,7 +3053,7 @@ MarkBufferDirty(Buffer buffer)
 	 * NB: We have to wait for the buffer header spinlock to be not held, as
 	 * TerminateBufferIO() relies on the spinlock.
 	 */
-	old_buf_state = pg_atomic_read_u32(&bufHdr->state);
+	old_buf_state = pg_atomic_read_u64(&bufHdr->state);
 	for (;;)
 	{
 		if (old_buf_state & BM_LOCKED)
@@ -3064,7 +3064,7 @@ MarkBufferDirty(Buffer buffer)
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
 
-		if (pg_atomic_compare_exchange_u32(&bufHdr->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&bufHdr->state, &old_buf_state,
 										   buf_state))
 			break;
 	}
@@ -3168,10 +3168,10 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 
 	if (ref == NULL)
 	{
-		uint32		buf_state;
-		uint32		old_buf_state;
+		uint64		buf_state;
+		uint64		old_buf_state;
 
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
@@ -3205,7 +3205,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 					buf_state += BUF_USAGECOUNT_ONE;
 			}
 
-			if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+			if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 											   buf_state))
 			{
 				result = (buf_state & BM_VALID) != 0;
@@ -3232,7 +3232,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		 * that the buffer page is legitimately non-accessible here.  We
 		 * cannot meddle with that.
 		 */
-		result = (pg_atomic_read_u32(&buf->state) & BM_VALID) != 0;
+		result = (pg_atomic_read_u64(&buf->state) & BM_VALID) != 0;
 
 		Assert(ref->data.refcount > 0);
 		ref->data.refcount++;
@@ -3267,7 +3267,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	/*
 	 * As explained, We don't expect any preexisting pins. That allows us to
@@ -3279,7 +3279,7 @@ PinBuffer_Locked(BufferDesc *buf)
 	 * Since we hold the buffer spinlock, we can update the buffer state and
 	 * release the lock in one operation.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 
 	UnlockBufHdrExt(buf, old_buf_state,
 					0, 0, 1);
@@ -3309,7 +3309,7 @@ WakePinCountWaiter(BufferDesc *buf)
 	 * BM_PIN_COUNT_WAITER if it stops waiting for a reason other than this
 	 * backend waking it up.
 	 */
-	uint32		buf_state = LockBufHdr(buf);
+	uint64		buf_state = LockBufHdr(buf);
 
 	if ((buf_state & BM_PIN_COUNT_WAITER) &&
 		BUF_STATE_GET_REFCOUNT(buf_state) == 1)
@@ -3356,7 +3356,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	ref->data.refcount--;
 	if (ref->data.refcount == 0)
 	{
-		uint32		old_buf_state;
+		uint64		old_buf_state;
 
 		/*
 		 * Mark buffer non-accessible to Valgrind.
@@ -3374,7 +3374,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
 		/* decrement the shared reference count */
-		old_buf_state = pg_atomic_fetch_sub_u32(&buf->state, BUF_REFCOUNT_ONE);
+		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
 
 		/* Support LockBufferForCleanup() */
 		if (old_buf_state & BM_PIN_COUNT_WAITER)
@@ -3431,7 +3431,7 @@ TrackNewBufferPin(Buffer buf)
 static void
 BufferSync(int flags)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	int			buf_id;
 	int			num_to_scan;
 	int			num_spaces;
@@ -3441,7 +3441,7 @@ BufferSync(int flags)
 	Oid			last_tsid;
 	binaryheap *ts_heap;
 	int			i;
-	uint32		mask = BM_DIRTY;
+	uint64		mask = BM_DIRTY;
 	WritebackContext wb_context;
 
 	/*
@@ -3473,7 +3473,7 @@ BufferSync(int flags)
 	for (buf_id = 0; buf_id < NBuffers; buf_id++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
-		uint32		set_bits = 0;
+		uint64		set_bits = 0;
 
 		/*
 		 * Header spinlock is enough to examine BM_DIRTY, see comment in
@@ -3640,7 +3640,7 @@ BufferSync(int flags)
 		 * write the buffer though we didn't need to.  It doesn't seem worth
 		 * guarding against this, though.
 		 */
-		if (pg_atomic_read_u32(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
+		if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
 		{
 			if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
 			{
@@ -4010,7 +4010,7 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 {
 	BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
 	int			result = 0;
-	uint32		buf_state;
+	uint64		buf_state;
 	BufferTag	tag;
 
 	/* Make sure we can handle the pin */
@@ -4259,7 +4259,7 @@ DebugPrintBufferRefcount(Buffer buffer)
 	int32		loccount;
 	char	   *result;
 	ProcNumber	backend;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 	if (BufferIsLocal(buffer))
@@ -4276,9 +4276,9 @@ DebugPrintBufferRefcount(Buffer buffer)
 	}
 
 	/* theoretically we should lock the bufhdr here */
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
-	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%x, refcount=%u %d)",
+	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%" PRIx64 ", refcount=%u %d)",
 					  buffer,
 					  relpathbackend(BufTagGetRelFileLocator(&buf->tag), backend,
 									 BufTagGetForkNum(&buf->tag)).str,
@@ -4378,7 +4378,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	instr_time	io_start;
 	Block		bufBlock;
 	char	   *bufToWrite;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
@@ -4576,7 +4576,7 @@ BufferIsPermanent(Buffer buffer)
 	 * not random garbage.
 	 */
 	bufHdr = GetBufferDescriptor(buffer - 1);
-	return (pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT) != 0;
+	return (pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT) != 0;
 }
 
 /*
@@ -5039,11 +5039,11 @@ FlushRelationBuffers(Relation rel)
 	{
 		for (i = 0; i < NLocBuffer; i++)
 		{
-			uint32		buf_state;
+			uint64		buf_state;
 
 			bufHdr = GetLocalBufferDescriptor(i);
 			if (BufTagMatchesRelFileLocator(&bufHdr->tag, &rel->rd_locator) &&
-				((buf_state = pg_atomic_read_u32(&bufHdr->state)) &
+				((buf_state = pg_atomic_read_u64(&bufHdr->state)) &
 				 (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 			{
 				ErrorContextCallback errcallback;
@@ -5079,7 +5079,7 @@ FlushRelationBuffers(Relation rel)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5151,7 +5151,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 	{
 		SMgrSortArray *srelent = NULL;
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -5400,7 +5400,7 @@ FlushDatabaseBuffers(Oid dbid)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5548,13 +5548,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u32(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
 		(BM_DIRTY | BM_JUST_DIRTIED))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
 		bool		delayChkptFlags = false;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * If we need to protect hint bit updates from torn writes, WAL-log a
@@ -5566,7 +5566,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
 		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT))
+			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5666,8 +5666,8 @@ UnlockBuffers(void)
 
 	if (buf)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		buf_state = LockBufHdr(buf);
 
@@ -5798,8 +5798,8 @@ LockBufferForCleanup(Buffer buffer)
 
 	for (;;)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		/* Try to acquire lock */
 		LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
@@ -5947,7 +5947,7 @@ bool
 ConditionalLockBufferForCleanup(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state,
+	uint64		buf_state,
 				refcount;
 
 	Assert(BufferIsValid(buffer));
@@ -6005,7 +6005,7 @@ bool
 IsBufferCleanupOK(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 
@@ -6061,7 +6061,7 @@ WaitIO(BufferDesc *buf)
 	ConditionVariablePrepareToSleep(cv);
 	for (;;)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 		PgAioWaitRef iow;
 
 		/*
@@ -6135,7 +6135,7 @@ WaitIO(BufferDesc *buf)
 bool
 StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	ResourceOwnerEnlarge(CurrentResourceOwner);
 
@@ -6191,11 +6191,11 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
  * is being released)
  */
 void
-TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 				  bool forget_owner, bool release_aio)
 {
-	uint32		buf_state;
-	uint32		unset_flag_bits = 0;
+	uint64		buf_state;
+	uint64		unset_flag_bits = 0;
 	int			refcount_change = 0;
 
 	buf_state = LockBufHdr(buf);
@@ -6256,7 +6256,7 @@ static void
 AbortBufferIO(Buffer buffer)
 {
 	BufferDesc *buf_hdr = GetBufferDescriptor(buffer - 1);
-	uint32		buf_state;
+	uint64		buf_state;
 
 	buf_state = LockBufHdr(buf_hdr);
 	Assert(buf_state & (BM_IO_IN_PROGRESS | BM_TAG_VALID));
@@ -6350,10 +6350,10 @@ rlocator_comparator(const void *p1, const void *p2)
 /*
  * Lock buffer header - set BM_LOCKED in buffer state.
  */
-uint32
+uint64
 LockBufHdr(BufferDesc *desc)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	Assert(!BufferIsLocal(BufferDescriptorGetBuffer(desc)));
 
@@ -6362,7 +6362,7 @@ LockBufHdr(BufferDesc *desc)
 	 * infrastructure. The work necessary for that shows up in profiles and is
 	 * rarely necessary.
 	 */
-	old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
+	old_buf_state = pg_atomic_fetch_or_u64(&desc->state, BM_LOCKED);
 
 	if (unlikely(old_buf_state & BM_LOCKED))
 	{
@@ -6373,7 +6373,7 @@ LockBufHdr(BufferDesc *desc)
 		while (true)
 		{
 			/* set BM_LOCKED flag */
-			old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
+			old_buf_state = pg_atomic_fetch_or_u64(&desc->state, BM_LOCKED);
 			/* if it wasn't set before we're OK */
 			if (!(old_buf_state & BM_LOCKED))
 				break;
@@ -6392,20 +6392,20 @@ LockBufHdr(BufferDesc *desc)
  * Obviously the buffer could be locked by the time the value is returned, so
  * this is primarily useful in CAS style loops.
  */
-pg_noinline uint32
+pg_noinline uint64
 WaitBufHdrUnlocked(BufferDesc *buf)
 {
 	SpinDelayStatus delayStatus;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	init_local_spin_delay(&delayStatus);
 
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
 	while (buf_state & BM_LOCKED)
 	{
 		perform_spin_delay(&delayStatus);
-		buf_state = pg_atomic_read_u32(&buf->state);
+		buf_state = pg_atomic_read_u64(&buf->state);
 	}
 
 	finish_spin_delay(&delayStatus);
@@ -6693,12 +6693,12 @@ ResOwnerPrintBufferPin(Datum res)
 static bool
 EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result;
 
 	*buffer_flushed = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -6792,12 +6792,12 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -6844,7 +6844,7 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
@@ -6886,12 +6886,12 @@ static bool
 MarkDirtyUnpinnedBufferInternal(Buffer buf, BufferDesc *desc,
 								bool *buffer_already_dirty)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result = false;
 
 	*buffer_already_dirty = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -6989,7 +6989,7 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_already_dirty;
 
 		CHECK_FOR_INTERRUPTS();
@@ -7043,12 +7043,12 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_already_dirty;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -7099,7 +7099,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		BufferDesc *buf_hdr = is_temp ?
 			GetLocalBufferDescriptor(-buffer - 1)
 			: GetBufferDescriptor(buffer - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * Check that all the buffers are actually ones that could conceivably
@@ -7117,7 +7117,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		}
 
 		if (is_temp)
-			buf_state = pg_atomic_read_u32(&buf_hdr->state);
+			buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		else
 			buf_state = LockBufHdr(buf_hdr);
 
@@ -7155,7 +7155,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		if (is_temp)
 		{
 			buf_state += BUF_REFCOUNT_ONE;
-			pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 		}
 		else
 			UnlockBufHdrExt(buf_hdr, buf_state, 0, 0, 1);
@@ -7341,13 +7341,13 @@ buffer_readv_complete_one(PgAioTargetData *td, uint8 buf_off, Buffer buffer,
 		: GetBufferDescriptor(buffer - 1);
 	BufferTag	tag = buf_hdr->tag;
 	char	   *bufdata = BufferGetBlock(buffer);
-	uint32		set_flag_bits;
+	uint64		set_flag_bits;
 	int			piv_flags;
 
 	/* check that the buffer is in the expected state for a read */
 #ifdef USE_ASSERT_CHECKING
 	{
-		uint32		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(!(buf_state & BM_VALID));
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 28d952b3534..1d4f19a9afd 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -86,7 +86,7 @@ typedef struct BufferAccessStrategyData
 
 /* Prototypes for internal functions */
 static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
-									 uint32 *buf_state);
+									 uint64 *buf_state);
 static void AddBufferToRing(BufferAccessStrategy strategy,
 							BufferDesc *buf);
 
@@ -171,7 +171,7 @@ ClockSweepTick(void)
  *	before returning.
  */
 BufferDesc *
-StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_ring)
+StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_ring)
 {
 	BufferDesc *buf;
 	int			bgwprocno;
@@ -230,8 +230,8 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 	trycounter = NBuffers;
 	for (;;)
 	{
-		uint32		old_buf_state;
-		uint32		local_buf_state;
+		uint64		old_buf_state;
+		uint64		local_buf_state;
 
 		buf = GetBufferDescriptor(ClockSweepTick());
 
@@ -239,7 +239,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 		 * Check whether the buffer can be used and pin it if so. Do this
 		 * using a CAS loop, to avoid having to lock the buffer header.
 		 */
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			local_buf_state = old_buf_state;
@@ -277,7 +277,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 			{
 				local_buf_state -= BUF_USAGECOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					trycounter = NBuffers;
@@ -289,7 +289,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 				/* pin the buffer if the CAS succeeds */
 				local_buf_state += BUF_REFCOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					/* Found a usable buffer */
@@ -655,12 +655,12 @@ FreeAccessStrategy(BufferAccessStrategy strategy)
  * returning.
  */
 static BufferDesc *
-GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
+GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
 {
 	BufferDesc *buf;
 	Buffer		bufnum;
-	uint32		old_buf_state;
-	uint32		local_buf_state;	/* to avoid repeated (de-)referencing */
+	uint64		old_buf_state;
+	uint64		local_buf_state;	/* to avoid repeated (de-)referencing */
 
 
 	/* Advance to next ring slot */
@@ -682,7 +682,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 	 * Check whether the buffer can be used and pin it if so. Do this using a
 	 * CAS loop, to avoid having to lock the buffer header.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 	for (;;)
 	{
 		local_buf_state = old_buf_state;
@@ -710,7 +710,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 		/* pin the buffer if the CAS succeeds */
 		local_buf_state += BUF_REFCOUNT_ONE;
 
-		if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 										   local_buf_state))
 		{
 			*buf_state = local_buf_state;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 15aac7d1c9f..a41a5facd3a 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -148,7 +148,7 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 	}
 	else
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		victim_buffer = GetLocalVictimBuffer();
 		bufid = -victim_buffer - 1;
@@ -165,10 +165,10 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 		 */
 		bufHdr->tag = newTag;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
 		buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 		*foundPtr = false;
 	}
@@ -245,12 +245,12 @@ GetLocalVictimBuffer(void)
 
 		if (LocalRefCount[victim_bufid] == 0)
 		{
-			uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 			if (BUF_STATE_GET_USAGECOUNT(buf_state) > 0)
 			{
 				buf_state -= BUF_USAGECOUNT_ONE;
-				pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+				pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 				trycounter = NLocBuffer;
 			}
 			else if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
@@ -286,13 +286,13 @@ GetLocalVictimBuffer(void)
 	 * this buffer is not referenced but it might still be dirty. if that's
 	 * the case, write it out before reusing it!
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY)
 		FlushLocalBuffer(bufHdr, NULL);
 
 	/*
 	 * Remove the victim buffer from the hashtable and mark as invalid.
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID)
 	{
 		InvalidateLocalBuffer(bufHdr, false);
 
@@ -417,7 +417,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 		if (found)
 		{
 			BufferDesc *existing_hdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			UnpinLocalBuffer(BufferDescriptorGetBuffer(victim_buf_hdr));
 
@@ -428,18 +428,18 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 			/*
 			 * Clear the BM_VALID bit, do StartLocalBufferIO() and proceed.
 			 */
-			buf_state = pg_atomic_read_u32(&existing_hdr->state);
+			buf_state = pg_atomic_read_u64(&existing_hdr->state);
 			Assert(buf_state & BM_TAG_VALID);
 			Assert(!(buf_state & BM_DIRTY));
 			buf_state &= ~BM_VALID;
-			pg_atomic_unlocked_write_u32(&existing_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&existing_hdr->state, buf_state);
 
 			/* no need to loop for local buffers */
 			StartLocalBufferIO(existing_hdr, true, false);
 		}
 		else
 		{
-			uint32		buf_state = pg_atomic_read_u32(&victim_buf_hdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&victim_buf_hdr->state);
 
 			Assert(!(buf_state & (BM_VALID | BM_TAG_VALID | BM_DIRTY | BM_JUST_DIRTIED)));
 
@@ -447,7 +447,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 
 			buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 
-			pg_atomic_unlocked_write_u32(&victim_buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&victim_buf_hdr->state, buf_state);
 
 			hresult->id = victim_buf_id;
 
@@ -467,13 +467,13 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 	{
 		Buffer		buf = buffers[i];
 		BufferDesc *buf_hdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		buf_state |= BM_VALID;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 
 	*extended_by = extend_by;
@@ -492,7 +492,7 @@ MarkLocalBufferDirty(Buffer buffer)
 {
 	int			bufid;
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsLocal(buffer));
 
@@ -506,14 +506,14 @@ MarkLocalBufferDirty(Buffer buffer)
 
 	bufHdr = GetLocalBufferDescriptor(bufid);
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	if (!(buf_state & BM_DIRTY))
 		pgBufferUsage.local_blks_dirtied++;
 
 	buf_state |= BM_DIRTY;
 
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -522,7 +522,7 @@ MarkLocalBufferDirty(Buffer buffer)
 bool
 StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * With AIO the buffer could have IO in progress, e.g. when there are two
@@ -542,7 +542,7 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 	/* Once we get here, there is definitely no I/O active on this buffer */
 
 	/* Check if someone else already did the I/O */
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 	if (forInput ? (buf_state & BM_VALID) : !(buf_state & BM_DIRTY))
 	{
 		return false;
@@ -559,11 +559,11 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
  * Like TerminateBufferIO, but for local buffers
  */
 void
-TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bits,
+TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint64 set_flag_bits,
 					   bool release_aio)
 {
 	/* Only need to adjust flags */
-	uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+	uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/* BM_IO_IN_PROGRESS isn't currently used for local buffers */
 
@@ -582,7 +582,7 @@ TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bit
 	}
 
 	buf_state |= set_flag_bits;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 	/* local buffers don't track IO using resowners */
 
@@ -606,7 +606,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(bufHdr);
 	int			bufid = -buffer - 1;
-	uint32		buf_state;
+	uint64		buf_state;
 	LocalBufferLookupEnt *hresult;
 
 	/*
@@ -622,7 +622,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 		Assert(!pgaio_wref_valid(&bufHdr->io_wref));
 	}
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/*
 	 * We need to test not just LocalRefCount[bufid] but also the BufferDesc
@@ -647,7 +647,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 	ClearBufferTag(&bufHdr->tag);
 	buf_state &= ~BUF_FLAG_MASK;
 	buf_state &= ~BUF_USAGECOUNT_MASK;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -671,9 +671,9 @@ DropRelationLocalBuffers(RelFileLocator rlocator, ForkNumber *forkNum,
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (!(buf_state & BM_TAG_VALID) ||
 			!BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -706,9 +706,9 @@ DropRelationAllLocalBuffers(RelFileLocator rlocator)
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if ((buf_state & BM_TAG_VALID) &&
 			BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -804,11 +804,11 @@ InitLocalBuffers(void)
 bool
 PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	Buffer		buffer = BufferDescriptorGetBuffer(buf_hdr);
 	int			bufid = -buffer - 1;
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	if (LocalRefCount[bufid] == 0)
 	{
@@ -819,7 +819,7 @@ PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 		{
 			buf_state += BUF_USAGECOUNT_ONE;
 		}
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/*
 		 * See comment in PinBuffer().
@@ -856,14 +856,14 @@ UnpinLocalBufferNoOwner(Buffer buffer)
 	if (--LocalRefCount[buffid] == 0)
 	{
 		BufferDesc *buf_hdr = GetLocalBufferDescriptor(buffid);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		NLocalPinnedBuffers--;
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state -= BUF_REFCOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/* see comment in UnpinBufferNoOwner */
 		VALGRIND_MAKE_MEM_NOACCESS(LocalBufHdrGetBlock(buf_hdr), BLCKSZ);
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 702307a49e2..294caf7a1eb 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -199,7 +199,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
 		for (i = 0; i < NBuffers; i++)
 		{
 			BufferDesc *bufHdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			CHECK_FOR_INTERRUPTS();
 
@@ -615,7 +615,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -626,7 +626,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 		 * noticeably increase the cost of the function.
 		 */
 		bufHdr = GetBufferDescriptor(i);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (buf_state & BM_VALID)
 		{
@@ -676,7 +676,7 @@ pg_buffercache_usage_counts(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		int			usage_count;
 
 		CHECK_FOR_INTERRUPTS();
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index d7eadeab256..488d98e7e66 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -308,9 +308,9 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 {
 	Buffer		buf;
 	BufferDesc *buf_hdr;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		was_pinned = false;
-	uint32		unset_bits = 0;
+	uint64		unset_bits = 0;
 
 	/* place buffer in shared buffers without erroring out */
 	buf = ReadBufferExtended(rel, MAIN_FORKNUM, blkno, RBM_ZERO_AND_LOCK, NULL);
@@ -319,7 +319,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	}
 	else
 	{
@@ -340,7 +340,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_state &= ~unset_bits;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 	else
 	{
@@ -489,7 +489,7 @@ invalidate_rel_block(PG_FUNCTION_ARGS)
 
 			LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
-			if (pg_atomic_read_u32(&buf_hdr->state) & BM_DIRTY)
+			if (pg_atomic_read_u64(&buf_hdr->state) & BM_DIRTY)
 			{
 				if (BufferIsLocal(buf))
 					FlushLocalBuffer(buf_hdr, NULL);
@@ -572,7 +572,7 @@ buffer_call_terminate_io(PG_FUNCTION_ARGS)
 	bool		io_error = PG_GETARG_BOOL(3);
 	bool		release_aio = PG_GETARG_BOOL(4);
 	bool		clear_dirty = false;
-	uint32		set_flag_bits = 0;
+	uint64		set_flag_bits = 0;
 
 	if (io_error)
 		set_flag_bits |= BM_IO_ERROR;
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0012-bufmgr-Implement-buffer-content-locks-independent.patch (44.6K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/13-v7-0012-bufmgr-Implement-buffer-content-locks-independent.patch)
  download | inline diff:
From c8c5958a23253664ccaf4f9b8b4f79b79918724f Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 19 Nov 2025 16:37:26 -0500
Subject: [PATCH v7 12/15] bufmgr: Implement buffer content locks independently
 of lwlocks

Until now buffer content locks were implemented using lwlocks. That has the
obvious advantage of not needing a separate efficient implementation of
locks. However, the time for a dedicated buffer content lock implementation
has come:

1) Hint bits are currently set while holding only a share lock. This leads to
   having to copy pages while they are being written out if checksums are
   enabled, which is not cheap. We would like to add AIO writes, however once
   many buffers can be written out at the same time, it gets a lot more
   expensive to copy them, particularly because that copy needs to reside in
   shared buffers (for worker mode to have access to the buffer).

   In addition, modifying buffers while they are being written out can cause
   issues with unbuffered/direct-IO, as some filesystems (like btrfs) do not
   like that, due to filesystem internal checksums getting corrupted.

   The solution to this is to require a new share-exclusive lock-level to set
   hint bits and to write out buffers, making those operations mutually
   exclusive. We could introduce such a lock level into the generic lwlock
   implementation, however it does not look like there would be other users,
   and it does add some overhead into important codepaths.

2) For AIO writes we need to be able to race-freely check whether a buffer is
   undergoing IO and whether an exclusive lock on the page can be acquired. That
   is rather hard to do efficiently when the buffer state and the lock state
   are separate atomic variables. This is a major hindrance to allowing writes
   to be done asynchronously.

3) Buffer locks are by far the most frequently taken locks. Optimizing them
   specifically for their use case is worth the effort. E.g. by merging
   content locks into buffer locks we will be able to release a buffer lock
   and pin in one atomic operation.

4) There are more complicated optimizations, like long-lived "super pinned &
   locked" pages, that cannot realistically be implemented with the generic
   lwlock implementation.

Therefore implement content locks inside bufmgr.c. The lockstate is stored as
part of BufferDesc.state. The implementation of buffer content locks is fairly
similar to lwlocks, with a few important differences:

1) An additional lock-level share-exclusive has been added. This lock level
   conflicts with exclusive locks and itself, but not share locks.

2) Error recovery for content locks is implemented as part of the already
   existing private-refcount tracking mechanism in combination with resowners,
   instead of a bespoke mechanism as the case for lwlocks. This means we do
   not need to add dedicated error-recovery codepaths to release all content
   locks (like done with LWLockReleaseAll() for content locks).

3) The lock state is embedded in BufferDesc.state instead of having its own
   struct.

4) The wakeup logic is a tad more complicated due to needing to support the
   additional lock level

This commit unfortunately introduces some code that is very similar to the
code in lwlock.c, however the code is not equivalent enough to easily merge
it. The future wins that this commit makes possible seem worth the cost.

As of this commit nothing uses the new share-exclusive lock mode. It will be
used in a future commit. It seemed too complicated to introduce the lock-level
in a separate commit.

TODO:
- Address FIXMEs

- Perhaps move the locking code into a buffer_locking.h or such? Needs to be
  inline functions for efficiency unfortunately.

- reflow some comments that I didn't reflow to make the diff more readable

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Reviewed-by: Greg Burd <greg@burd.me>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  55 +-
 src/include/storage/bufmgr.h                  |  32 +-
 src/include/storage/proc.h                    |   8 +-
 src/backend/storage/buffer/buf_init.c         |   7 +-
 src/backend/storage/buffer/bufmgr.c           | 872 ++++++++++++++++--
 .../utils/activity/wait_event_names.txt       |   3 +
 6 files changed, 893 insertions(+), 84 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 28519ad2813..0a145d95024 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -23,6 +23,7 @@
 #include "storage/condition_variable.h"
 #include "storage/lwlock.h"
 #include "storage/procnumber.h"
+#include "storage/proclist_types.h"
 #include "storage/shmem.h"
 #include "storage/smgr.h"
 #include "storage/spin.h"
@@ -32,22 +33,29 @@
 /*
  * Buffer state is a single 64-bit variable where following data is combined.
  *
+ * State of the buffer itself:
  * - 18 bits refcount
  * - 4 bits usage count
  * - 10 bits of flags
  *
+ * State of the content lock:
+ * - 1 bit has_waiter
+ * - 1 bit release_ok
+ * - 1 bit lock state locked
+ * - 1 bit exclusively locked
+ * - 1 bit share exclusively locked
+ * - 18 bits share lock count
+ *
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
- * NB: A future commit will use a significant portion of the remaining bits to
- * implement buffer locking as part of the state variable.
- *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
 #define BUF_USAGECOUNT_BITS 4
 #define BUF_FLAG_BITS 10
 
+/* FIXME: Also assert lock state size, just not yet sure how */
 StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 				 "parts of buffer state space need to equal 32");
 
@@ -69,7 +77,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 	((uint32)(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT))
 
 /*
- * Flags for buffer descriptors
+ * Flags for buffer descriptor state
  *
  * Note: BM_TAG_VALID essentially means that there is a buffer hashtable
  * entry associated with the buffer's tag.
@@ -111,6 +119,20 @@ StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
 				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
 
+
+/*
+ * Definitions related to buffer content locks
+ */
+#define BM_LOCK_HAS_WAITERS         (UINT64CONST(1) << 63)
+#define BM_LOCK_RELEASE_OK          (UINT64CONST(1) << 62)
+
+#define BM_LOCK_VAL_SHARED          (UINT64CONST(1) << 32)
+#define BM_LOCK_VAL_SHARE_EXCLUSIVE (UINT64CONST(1) << (32 + MAX_BACKENDS_BITS))
+#define BM_LOCK_VAL_EXCLUSIVE       (UINT64CONST(1) << (32 + 1 + MAX_BACKENDS_BITS))
+
+#define BM_LOCK_MASK                (((uint64)MAX_BACKENDS << 32) | BM_LOCK_VAL_SHARE_EXCLUSIVE | BM_LOCK_VAL_EXCLUSIVE)
+
+
 /*
  * Buffer tag identifies which disk block the buffer contains.
  *
@@ -253,9 +275,6 @@ BufMappingPartitionLockByIndex(uint32 index)
  * it is held.  However, existing buffer pins may be released while the buffer
  * header spinlock is held, using an atomic subtraction.
  *
- * The LWLock can take care of itself.  The buffer header lock is *not* used
- * to control access to the data in the buffer!
- *
  * If we have the buffer pinned, its tag can't change underneath us, so we can
  * examine the tag without locking the buffer header.  Also, in places we do
  * one-time reads of the flags without bothering to lock the buffer header;
@@ -268,6 +287,15 @@ BufMappingPartitionLockByIndex(uint32 index)
  * wait_backend_pgprocno and setting flag bit BM_PIN_COUNT_WAITER.  At present,
  * there can be only one such waiter per buffer.
  *
+ * The content of buffers is protected via the buffer content lock,
+ * implemented as part buffer state. Note that the buffer header lock is *not*
+ * used to control access to the data in the buffer! We used to use an LWLock
+ * to implement the content lock, but having a dedicated implementation of
+ * content locks allows to implement some otherwise hard things (e.g.
+ * race-freely checking if AIO is in progress before locking a buffer
+ * exclusively) and makes otherwise impossible optimizations possible
+ * (e.g. unlocking and unpinning a buffer in one atomic operation).
+ *
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
@@ -309,7 +337,12 @@ typedef struct BufferDesc
 	int			wait_backend_pgprocno;
 
 	PgAioWaitRef io_wref;		/* set iff AIO is in progress */
-	LWLock		content_lock;	/* to lock access to buffer contents */
+
+	/*
+	 * List of PGPROCs waiting for the buffer content lock. Protected by the
+	 * buffer header spinlock.
+	 */
+	proclist_head lock_waiters;
 } BufferDesc;
 
 /*
@@ -396,12 +429,6 @@ BufferDescriptorGetIOCV(const BufferDesc *bdesc)
 	return &(BufferIOCVArray[bdesc->buf_id]).cv;
 }
 
-static inline LWLock *
-BufferDescriptorGetContentLock(const BufferDesc *bdesc)
-{
-	return (LWLock *) (&bdesc->content_lock);
-}
-
 /*
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 97c1124c12a..df170fe9553 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -203,7 +203,20 @@ extern PGDLLIMPORT int32 *LocalRefCount;
 typedef enum BufferLockMode
 {
 	BUFFER_LOCK_UNLOCK,
+
+	/*
+	 * A share lock conflicts with exclusive locks.
+	 */
 	BUFFER_LOCK_SHARE,
+
+	/*
+	 * A share-exclusive lock conflicts with itself and exclusive locks.
+	 */
+	BUFFER_LOCK_SHARE_EXCLUSIVE,
+
+	/*
+	 * An exclusive lock conflicts with every other lock type.
+	 */
 	BUFFER_LOCK_EXCLUSIVE,
 } BufferLockMode;
 
@@ -302,7 +315,24 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
 extern void UnlockBuffers(void);
-extern void LockBuffer(Buffer buffer, BufferLockMode mode);
+extern void UnlockBuffer(Buffer buffer);
+extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
+
+/*
+ * Handling BUFFER_LOCK_UNLOCK in bufmgr.c leads to sufficiently worse branch
+ * prediction to impact performance. Therefore handle that switch here, where
+ * most of the time `mode` will be a constant and thus can be optimized out by
+ * the compiler.
+ */
+static inline void
+LockBuffer(Buffer buffer, BufferLockMode mode)
+{
+	if (mode == BUFFER_LOCK_UNLOCK)
+		UnlockBuffer(buffer);
+	else
+		LockBufferInternal(buffer, mode);
+}
+
 extern bool ConditionalLockBuffer(Buffer buffer);
 extern void LockBufferForCleanup(Buffer buffer);
 extern bool ConditionalLockBufferForCleanup(Buffer buffer);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index c6f5ebceefd..d1f6d314d57 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -242,7 +242,13 @@ struct PGPROC
 	 */
 	bool		recoveryConflictPending;
 
-	/* Info about LWLock the process is currently waiting for, if any. */
+	/*
+	 * Info about LWLock the process is currently waiting for, if any.
+	 *
+	 * This is currently used both for lwlocks and buffer content locks, which
+	 * is acceptable, although not pretty, because a backend can't wait for
+	 * both types of locks at the same time.
+	 */
 	uint8		lwWaiting;		/* see LWLockWaitState */
 	uint8		lwWaitMode;		/* lwlock mode being waited for */
 	proclist_node lwWaitLink;	/* position in LW lock wait list */
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 25f71191ec3..f3224c793c4 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
 #include "storage/aio.h"
 #include "storage/buf_internals.h"
 #include "storage/bufmgr.h"
+#include "storage/proclist.h"
 
 BufferDescPadded *BufferDescriptors;
 char	   *BufferBlocks;
@@ -121,16 +122,14 @@ BufferManagerShmemInit(void)
 
 			ClearBufferTag(&buf->tag);
 
-			pg_atomic_init_u64(&buf->state, 0);
+			pg_atomic_init_u64(&buf->state, BM_LOCK_RELEASE_OK);
 			buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
 
 			buf->buf_id = i;
 
 			pgaio_wref_clear(&buf->io_wref);
 
-			LWLockInitialize(BufferDescriptorGetContentLock(buf),
-							 LWTRANCHE_BUFFER_CONTENT);
-
+			proclist_init(&buf->lock_waiters);
 			ConditionVariableInit(BufferDescriptorGetIOCV(buf));
 		}
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index d0b8f8d20eb..32634858514 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -58,6 +58,7 @@
 #include "storage/ipc.h"
 #include "storage/lmgr.h"
 #include "storage/proc.h"
+#include "storage/proclist.h"
 #include "storage/read_stream.h"
 #include "storage/smgr.h"
 #include "storage/standby.h"
@@ -96,6 +97,12 @@ typedef struct PrivateRefCountData
 	 * How many times has the buffer been pinned by this backend.
 	 */
 	int32		refcount;
+
+	/*
+	 * Is the buffer locked by this backend? BUFFER_LOCK_UNLOCK indicates that
+	 * the buffer is not locked.
+	 */
+	BufferLockMode lockmode;
 } PrivateRefCountData;
 
 typedef struct PrivateRefCountEntry
@@ -206,8 +213,10 @@ static BufferDesc *PinCountWaitBuf = NULL;
  * Each buffer also has a private refcount that keeps track of the number of
  * times the buffer is pinned in the current process.  This is so that the
  * shared refcount needs to be modified only once if a buffer is pinned more
- * than once by an individual backend.  It's also used to check that no buffers
- * are still pinned at the end of transactions and when exiting.
+ * than once by an individual backend.  It's also used to check that no
+ * buffers are still pinned at the end of transactions and when exiting. We
+ * also use this mechanism to track whether this backend has a buffer locked,
+ * and, if so, in what mode.
  *
  *
  * To avoid - as we used to - requiring an array with NBuffers entries to keep
@@ -346,6 +355,7 @@ ReservePrivateRefCountEntry(void)
 		/* clear the whole data member, just for future proofing */
 		memset(&victim_entry->data, 0, sizeof(victim_entry->data));
 		victim_entry->data.refcount = 0;
+		victim_entry->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 		PrivateRefCountOverflowed++;
 	}
@@ -369,6 +379,7 @@ NewPrivateRefCountEntry(Buffer buffer)
 	PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
 	res->buffer = buffer;
 	res->data.refcount = 0;
+	res->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 	/* update cache for the next lookup */
 	PrivateRefCountEntryLast = ReservedRefCountSlot;
@@ -535,6 +546,7 @@ static void
 ForgetPrivateRefCountEntry(PrivateRefCountEntry *ref)
 {
 	Assert(ref->data.refcount == 0);
+	Assert(ref->data.lockmode == BUFFER_LOCK_UNLOCK);
 
 	if (ref >= &PrivateRefCountArray[0] &&
 		ref < &PrivateRefCountArray[REFCOUNT_ARRAY_ENTRIES])
@@ -636,14 +648,27 @@ static void RelationCopyStorageUsingBuffer(RelFileLocator srclocator,
 static void AtProcExit_Buffers(int code, Datum arg);
 static void CheckForBufferLeaks(void);
 #ifdef USE_ASSERT_CHECKING
-static void AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
-									   void *unused_context);
+static void AssertNotCatalogBufferLock(Buffer buffer, BufferLockMode mode);
 #endif
 static int	rlocator_comparator(const void *p1, const void *p2);
 static inline int buffertag_comparator(const BufferTag *ba, const BufferTag *bb);
 static inline int ckpt_buforder_comparator(const CkptSortItem *a, const CkptSortItem *b);
 static int	ts_ckpt_progress_comparator(Datum a, Datum b, void *arg);
 
+static void BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr);
+static bool BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMe(BufferDesc *buf_hdr);
+static inline void BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr);
+static inline int BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr);
+static inline bool BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockDequeueSelf(BufferDesc *buf_hdr);
+static void BufferLockWakeup(BufferDesc *buf_hdr, bool unlocked);
+static void BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate);
+static inline uint64 BufferLockReleaseSub(BufferLockMode mode);
+
 
 /*
  * Implementation of PrefetchBuffer() for shared buffers.
@@ -2444,8 +2469,6 @@ again:
 	 */
 	if (buf_state & BM_DIRTY)
 	{
-		LWLock	   *content_lock;
-
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(buf_state & BM_VALID);
 
@@ -2463,8 +2486,7 @@ again:
 		 * one just happens to be trying to split the page the first one got
 		 * from StrategyGetBuffer.)
 		 */
-		content_lock = BufferDescriptorGetContentLock(buf_hdr);
-		if (!LWLockConditionalAcquire(content_lock, LW_SHARED))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -2493,7 +2515,7 @@ again:
 			if (XLogNeedsFlush(lsn)
 				&& StrategyRejectBuffer(strategy, buf_hdr, from_ring))
 			{
-				LWLockRelease(content_lock);
+				LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 				UnpinBuffer(buf_hdr);
 				goto again;
 			}
@@ -2501,7 +2523,7 @@ again:
 
 		/* OK, do the I/O */
 		FlushBuffer(buf_hdr, NULL, IOOBJECT_RELATION, io_context);
-		LWLockRelease(content_lock);
+		LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 		ScheduleBufferTagForWriteback(&BackendWritebackContext, io_context,
 									  &buf_hdr->tag);
@@ -2943,7 +2965,7 @@ BufferIsLockedByMe(Buffer buffer)
 	else
 	{
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr));
+		return BufferLockHeldByMe(bufHdr);
 	}
 }
 
@@ -2968,23 +2990,8 @@ BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode)
 	}
 	else
 	{
-		LWLockMode	lw_mode;
-
-		switch (mode)
-		{
-			case BUFFER_LOCK_EXCLUSIVE:
-				lw_mode = LW_EXCLUSIVE;
-				break;
-			case BUFFER_LOCK_SHARE:
-				lw_mode = LW_SHARED;
-				break;
-			default:
-				pg_unreachable();
-		}
-
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-									lw_mode);
+		return BufferLockHeldByMeInMode(bufHdr, mode);
 	}
 }
 
@@ -3371,7 +3378,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 * I'd better not still hold the buffer content lock. Can't use
 		 * BufferIsLockedByMe(), as that asserts the buffer is pinned.
 		 */
-		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
+		Assert(!BufferLockHeldByMe(buf));
 
 		/* decrement the shared reference count */
 		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
@@ -4193,7 +4200,7 @@ CheckForBufferLeaks(void)
  * Check for exclusive-locked catalog buffers.  This is the core of
  * AssertCouldGetRelation().
  *
- * A backend would self-deadlock on LWLocks if the catalog scan read the
+ * A backend would self-deadlock on the content if the catalog scan read the
  * exclusive-locked buffer.  The main threat is exclusive-locked buffers of
  * catalogs used in relcache, because a catcache search on any catalog may
  * build that catalog's relcache entry.  We don't have an inventory of
@@ -4209,26 +4216,45 @@ CheckForBufferLeaks(void)
 void
 AssertBufferLocksPermitCatalogRead(void)
 {
-	ForEachLWLockHeldByMe(AssertNotCatalogBufferLock, NULL);
+	PrivateRefCountEntry *res;
+
+	/* check the array */
+	for (int i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
+	{
+		if (PrivateRefCountArrayKeys[i] != InvalidBuffer)
+		{
+			res = &PrivateRefCountArray[i];
+
+			if (res->buffer == InvalidBuffer)
+				continue;
+
+			AssertNotCatalogBufferLock(res->buffer, res->data.lockmode);
+		}
+	}
+
+	/* if necessary search the hash */
+	if (PrivateRefCountOverflowed)
+	{
+		HASH_SEQ_STATUS hstat;
+
+		hash_seq_init(&hstat, PrivateRefCountHash);
+		while ((res = (PrivateRefCountEntry *) hash_seq_search(&hstat)) != NULL)
+		{
+			AssertNotCatalogBufferLock(res->buffer, res->data.lockmode);
+		}
+	}
 }
 
 static void
-AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
-						   void *unused_context)
+AssertNotCatalogBufferLock(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *bufHdr;
+	BufferDesc *bufHdr = GetBufferDescriptor(buffer - 1);
 	BufferTag	tag;
 	Oid			relid;
 
-	if (mode != LW_EXCLUSIVE)
+	if (mode != BUFFER_LOCK_EXCLUSIVE)
 		return;
 
-	if (!((BufferDescPadded *) lock > BufferDescriptors &&
-		  (BufferDescPadded *) lock < BufferDescriptors + NBuffers))
-		return;					/* not a buffer lock */
-
-	bufHdr = (BufferDesc *)
-		((char *) lock - offsetof(BufferDesc, content_lock));
 	tag = bufHdr->tag;
 
 	/*
@@ -4510,9 +4536,11 @@ static void
 FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 					IOObject io_object, IOContext io_context)
 {
-	LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	Buffer		buffer = BufferDescriptorGetBuffer(buf);
+
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-	LWLockRelease(BufferDescriptorGetContentLock(buf));
+	BufferLockUnlock(buffer, buf);
 }
 
 /*
@@ -5655,9 +5683,10 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
  *
  * Used to clean up after errors.
  *
- * Currently, we can expect that lwlock.c's LWLockReleaseAll() took care
- * of releasing buffer content locks per se; the only thing we need to deal
- * with here is clearing any PIN_COUNT request that was in progress.
+ * Currently, we can expect that resource owner cleanup, via
+ * ResOwnerReleaseBufferPin(), took care releasing buffer content locks per
+ * se; the only thing we need to deal with here is clearing any PIN_COUNT
+ * request that was in progress.
  */
 void
 UnlockBuffers(void)
@@ -5688,25 +5717,727 @@ UnlockBuffers(void)
 }
 
 /*
- * Acquire or release the content_lock for the buffer.
+ * Acquire the buffer content lock in the specified mode
+ *
+ * If the lock is not available, sleep until it is.
+ *
+ * Side effect: cancel/die interrupts are held off until lock release.
+ *
+ * This uses almost the same locking approach as lwlock.c's
+ * LWLockAcquire(). See documentation atop of lwlock.c for a more detailed
+ * discussion.
+ *
+ * The reason that this, and most of the other BufferLock* functions, get both
+ * the Buffer and BufferDesc* as parameters, is that looking up one from the
+ * other repeatedly shows up noticeably in profiles.
+ *
+ * Callers should provide a constant for mode, for more efficient code
+ * generation.
+ */
+static inline void
+BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry;
+	int			extraWaits = 0;
+
+	/*
+	 * Get reference to the refcount entry before we hold the lock, it seems
+	 * better to do before holding the lock.
+	 */
+	entry = GetPrivateRefCountEntry(buffer, true);
+
+	/*
+	 * We better not already hold a lock on the buffer.
+	 */
+	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	for (;;)
+	{
+		bool		mustwait;
+		uint32		wait_event;
+
+		/*
+		 * Try to grab the lock the first time, we're not in the waitqueue
+		 * yet/anymore.
+		 */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		if (likely(!mustwait))
+		{
+			break;
+		}
+
+		/*
+		 * Ok, at this point we couldn't grab the lock on the first try. We
+		 * cannot simply queue ourselves to the end of the list and wait to be
+		 * woken up because by now the lock could long have been released.
+		 * Instead add us to the queue and try to grab the lock again. If we
+		 * succeed we need to revert the queuing and be happy, otherwise we
+		 * recheck the lock. If we still couldn't grab it, we know that the
+		 * other locker will see our queue entries when releasing since they
+		 * existed before we checked for the lock.
+		 */
+
+		/* add to the queue */
+		BufferLockQueueSelf(buf_hdr, mode);
+
+		/* we're now guaranteed to be woken up if necessary */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		/* ok, grabbed the lock the second time round, need to undo queueing */
+		if (!mustwait)
+		{
+			BufferLockDequeueSelf(buf_hdr);
+			break;
+		}
+
+		switch (mode)
+		{
+			case BUFFER_LOCK_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_SHARE_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE:
+				wait_event = WAIT_EVENT_BUFFER_SHARED;
+				break;
+			case BUFFER_LOCK_UNLOCK:
+				pg_unreachable();
+
+		}
+		pgstat_report_wait_start(wait_event);
+
+		/*
+		 * Wait until awakened.
+		 *
+		 * It is possible that we get awakened for a reason other than being
+		 * signaled by LWLockRelease.  If so, loop back and wait again.  Once
+		 * we've gotten the LWLock, re-increment the sema by the number of
+		 * additional signals received.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		pgstat_report_wait_end();
+
+		/* Retrying, allow BufferLockRelease to release waiters again. */
+		pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_RELEASE_OK);
+	}
+
+	/* Remember that we now hold this lock */
+	entry->data.lockmode = mode;
+
+	/*
+	 * Fix the process wait semaphore's count for any absorbed wakeups.
+	 */
+	while (unlikely(extraWaits-- > 0))
+		PGSemaphoreUnlock(MyProc->sem);
+}
+
+/*
+ * Release a previously acquired buffer content lock.
+ */
+static void
+BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	uint64		oldstate;
+	uint64		sub;
+
+	mode = BufferLockDisownInternal(buffer, buf_hdr);
+
+	/*
+	 * Release my hold on lock, after that it can immediately be acquired by
+	 * others, even if we still have to wakeup other waiters.
+	 */
+	sub = BufferLockReleaseSub(mode);
+
+	oldstate = pg_atomic_sub_fetch_u64(&buf_hdr->state, sub);
+
+	BufferLockProcessRelease(buf_hdr, mode, oldstate);
+
+	/*
+	 * Now okay to allow cancel/die interrupts.
+	 */
+	RESUME_INTERRUPTS();
+}
+
+
+/*
+ * Acquire the content lock for the buffer, but only if we don't have to wait.
+ */
+static bool
+BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	bool		mustwait;
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	/* Check for the lock */
+	mustwait = BufferLockAttempt(buf_hdr, mode);
+
+	if (mustwait)
+	{
+		/* Failed to get lock, so release interrupt holdoff */
+		RESUME_INTERRUPTS();
+	}
+	else
+	{
+		PrivateRefCountEntry *entry =
+			GetPrivateRefCountEntry(buffer, true);
+
+		entry->data.lockmode = mode;
+	}
+
+	return !mustwait;
+}
+
+/*
+ * Internal function that tries to atomically acquire the content lock in the
+ * passed in mode.
+ *
+ * This function will not block waiting for a lock to become free - that's the
+ * caller's job.
+ *
+ * Similar to LWLockAttemptLock().
+ */
+static inline bool
+BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	uint64		old_state;
+
+	/*
+	 * Read once outside the loop, later iterations will get the newer value
+	 * via compare & exchange.
+	 */
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+
+	/* loop until we've determined whether we could acquire the lock or not */
+	while (true)
+	{
+		uint64		desired_state;
+		bool		lock_free;
+
+		desired_state = old_state;
+
+		if (mode == BUFFER_LOCK_EXCLUSIVE)
+		{
+			lock_free = (old_state & BM_LOCK_MASK) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_EXCLUSIVE;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			lock_free = (old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+		}
+		else
+		{
+			lock_free = (old_state & BM_LOCK_VAL_EXCLUSIVE) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARED;
+		}
+
+		/*
+		 * Attempt to swap in the state we are expecting. If we didn't see
+		 * lock to be free, that's just the old value. If we saw it as free,
+		 * we'll attempt to mark it acquired. The reason that we always swap
+		 * in the value is that this doubles as a memory barrier. We could try
+		 * to be smarter and only swap in values if we saw the lock as free,
+		 * but benchmark haven't shown it as beneficial so far.
+		 *
+		 * Retry if the value changed since we last looked at it.
+		 */
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			if (lock_free)
+			{
+				/* Great! Got the lock. */
+				return false;
+			}
+			else
+				return true;	/* somebody else has the lock */
+		}
+	}
+
+	pg_unreachable();
+}
+
+/*
+ * Add ourselves to the end of the content lock's wait queue.
+ */
+static void
+BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	/*
+	 * If we don't have a PGPROC structure, there's no way to wait. This
+	 * should never occur, since MyProc should only be null during shared
+	 * memory initialization.
+	 */
+	if (MyProc == NULL)
+		elog(PANIC, "cannot wait without a PGPROC structure");
+
+	if (MyProc->lwWaiting != LW_WS_NOT_WAITING)
+		elog(PANIC, "queueing for lock while waiting on another one");
+
+	LockBufHdr(buf_hdr);
+
+	/* setting the flag is protected by the spinlock */
+	pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_HAS_WAITERS);
+
+	/*
+	 * FIXME: This is reusing the lwlock fields. That's not a correctness
+	 * issue, a backend can't wait for both an lwlock and a buffer content
+	 * lock at the same time. However, it seems pretty ugly, particularly
+	 * given that the field names have an lw* prefix. But duplicating the
+	 * fields also seems somewhat superfluous.
+	 */
+	MyProc->lwWaiting = LW_WS_WAITING;
+	MyProc->lwWaitMode = mode;
+
+	proclist_push_tail(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	/* Can release the mutex now */
+	UnlockBufHdr(buf_hdr);
+}
+
+/*
+ * Remove ourselves from the waitlist.
+ *
+ * This is used if we queued ourselves because we thought we needed to sleep
+ * but, after further checking, we discovered that we don't actually need to
+ * do so.
+ */
+static void
+BufferLockDequeueSelf(BufferDesc *buf_hdr)
+{
+	bool		on_waitlist;
+
+	LockBufHdr(buf_hdr);
+
+	on_waitlist = MyProc->lwWaiting == LW_WS_WAITING;
+	if (on_waitlist)
+		proclist_delete(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	if (proclist_is_empty(&buf_hdr->lock_waiters) &&
+		(pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS) != 0)
+	{
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_HAS_WAITERS);
+	}
+
+	/* XXX: combine with fetch_and above? */
+	UnlockBufHdr(buf_hdr);
+
+	/* clear waiting state again, nice for debugging */
+	if (on_waitlist)
+		MyProc->lwWaiting = LW_WS_NOT_WAITING;
+	else
+	{
+		int			extraWaits = 0;
+
+
+		/*
+		 * Somebody else dequeued us and has or will wake us up. Deal with the
+		 * superfluous absorption of a wakeup.
+		 */
+
+		/*
+		 * Reset RELEASE_OK flag if somebody woke us before we removed
+		 * ourselves - they'll have set it to false.
+		 */
+		pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_RELEASE_OK);
+
+		/*
+		 * Now wait for the scheduled wakeup, otherwise our ->lwWaiting would
+		 * get reset at some inconvenient point later. Most of the time this
+		 * will immediately return.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		/*
+		 * Fix the process wait semaphore's count for any absorbed wakeups.
+		 */
+		while (extraWaits-- > 0)
+			PGSemaphoreUnlock(MyProc->sem);
+	}
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * After calling this function it's the callers responsibility to ensure that
+ * the lock gets released, even in case of an error. This only is desirable if
+ * the lock is going to be released in a different process than the process
+ * that acquired it.
+ */
+static inline void
+BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockDisownInternal(buffer, buf_hdr);
+	RESUME_INTERRUPTS();
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * This is the code that can be shared between actually releasing a lock
+ * (BufferLockUnlock()) and just not tracking ownership of the lock anymore
+ * without releasing the lock (BufferLockDisown()).
+ */
+static inline int
+BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	PrivateRefCountEntry *ref;
+
+	ref = GetPrivateRefCountEntry(buffer, false);
+	if (ref == NULL)
+		elog(ERROR, "lock %d is not held", buffer);
+	mode = ref->data.lockmode;
+	ref->data.lockmode = BUFFER_LOCK_UNLOCK;
+
+	return mode;
+}
+
+/*
+ * Wakeup all the lockers that currently have a chance to acquire the lock.
+ *
+ * wake_exclusive indicates whether exlusive lock waiters should be woken up.
+ */
+static void
+BufferLockWakeup(BufferDesc *buf_hdr, bool wake_exclusive)
+{
+	bool		new_release_ok;
+	bool		wake_share_exclusive = true;
+	proclist_head wakeup;
+	proclist_mutable_iter iter;
+
+	proclist_init(&wakeup);
+
+	new_release_ok = true;
+
+	/* lock wait list while collecting backends to wake up */
+	LockBufHdr(buf_hdr);
+
+	proclist_foreach_modify(iter, &buf_hdr->lock_waiters, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		/*
+		 * Already woke up a conflicting lock, so skip over this wait list
+		 * entry.
+		 */
+		if (!wake_exclusive && waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+			continue;
+		if (!wake_share_exclusive && waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+			continue;
+
+		proclist_delete(&buf_hdr->lock_waiters, iter.cur, lwWaitLink);
+		proclist_push_tail(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Prevent additional wakeups until retryer gets to run. Backends that
+		 * are just waiting for the lock to become free don't retry
+		 * automatically.
+		 */
+		new_release_ok = false;
+
+		/*
+		 * Signal that the process isn't on the wait list anymore. This allows
+		 * BufferLockDequeueSelf() to remove itself from the waitlist with a
+		 * proclist_delete(), rather than having to check if it has been
+		 * removed from the list.
+		 */
+		Assert(waiter->lwWaiting == LW_WS_WAITING);
+		waiter->lwWaiting = LW_WS_PENDING_WAKEUP;
+
+		/*
+		 * Don't wakeup further waiters after waking a conflicting waiter.
+		 */
+		if (waiter->lwWaitMode == BUFFER_LOCK_SHARE)
+		{
+			/*
+			 * Share locks conflict with exclusive locks.
+			 */
+			wake_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			/*
+			 * Share-exclusive locks conflict with share-exclusive eand
+			 * exclusive locks.
+			 */
+			wake_exclusive = false;
+			wake_share_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+		{
+
+			/*
+			 * Exclusive locks conflict with all other locks, there's no point
+			 * in waking up anybody else.
+			 */
+			break;
+		}
+	}
+
+	Assert(proclist_is_empty(&wakeup) || pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS);
+
+	/* unset required flags, and release lock, in one fell swoop */
+	{
+		uint64		old_state;
+		uint64		desired_state;
+
+		old_state = pg_atomic_read_u64(&buf_hdr->state);
+		while (true)
+		{
+			desired_state = old_state;
+
+			/* compute desired flags */
+
+			if (new_release_ok)
+				desired_state |= BM_LOCK_RELEASE_OK;
+			else
+				desired_state &= ~BM_LOCK_RELEASE_OK;
+
+			if (proclist_is_empty(&buf_hdr->lock_waiters))
+				desired_state &= ~BM_LOCK_HAS_WAITERS;
+
+			desired_state &= ~BM_LOCKED;	/* release lock */
+
+			if (pg_atomic_compare_exchange_u64(&buf_hdr->state, &old_state,
+											   desired_state))
+				break;
+		}
+	}
+
+	/* Awaken any waiters I removed from the queue. */
+	proclist_foreach_modify(iter, &wakeup, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		proclist_delete(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Guarantee that lwWaiting being unset only becomes visible once the
+		 * unlink from the link has completed. Otherwise the target backend
+		 * could be woken up for other reason and enqueue for a new lock - if
+		 * that happens before the list unlink happens, the list would end up
+		 * being corrupted.
+		 *
+		 * The barrier pairs with the LWLockWaitListLock() when enqueuing for
+		 * another lock.
+		 */
+		pg_write_barrier();
+		waiter->lwWaiting = LW_WS_NOT_WAITING;
+		PGSemaphoreUnlock(waiter->sem);
+	}
+}
+
+/*
+ * Compute subtraction from buffer state for a release of a held lock in
+ * `mode`.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places, each needing to compute what to
+ * subtract from the lock state.
+ */
+static inline uint64
+BufferLockReleaseSub(BufferLockMode mode)
+{
+
+	/*
+	 * Turns out that a switch() leads gcc to generate sufficiently worse code
+	 * for this to show up in profiles...
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE)
+		return BM_LOCK_VAL_EXCLUSIVE;
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		return BM_LOCK_VAL_SHARE_EXCLUSIVE;
+	else
+	{
+		Assert(mode == BUFFER_LOCK_SHARE);
+		return BM_LOCK_VAL_SHARED;
+	}
+
+	return 0;
+}
+
+/*
+ * Handle work that needs to be done after releasing a lock that was held in
+ * `mode`, where `lockstate` is the result of the atomic operation modifying
+ * the state variable.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places.
+ */
+static void
+BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate)
+{
+	bool		check_waiters = false;
+	bool		wake_exclusive = false;
+
+	/* nobody else can have that kind of lock */
+	Assert(!(lockstate & BM_LOCK_VAL_EXCLUSIVE));
+
+	/*
+	 * If we're still waiting for backends to get scheduled, don't wake them
+	 * up again. Otherwise check if we need to look through the waitqueue to
+	 * wake other backends.
+	 */
+	if ((lockstate & (BM_LOCK_HAS_WAITERS | BM_LOCK_RELEASE_OK)) ==
+		(BM_LOCK_HAS_WAITERS | BM_LOCK_RELEASE_OK))
+	{
+		if ((lockstate & BM_LOCK_MASK) == 0)
+		{
+			/*
+			 * We released a lock and the lock was, in that moment, free. We
+			 * therefore can wake waiters for any kind of lock.
+			 */
+			check_waiters = true;
+			wake_exclusive = true;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			/*
+			 * We released the lock, but another backend still holds a lock.
+			 * We can't have released an exclusive lock, as there couldn't
+			 * have been other lock holders. If we released a share lock, no
+			 * waiters need to be woken up, as there must be other share
+			 * lockers. However, if we held a share-exclusive lock, another
+			 * backend now could acquire a share-exclusive lock.
+			 */
+			check_waiters = true;
+			wake_exclusive = false;
+		}
+	}
+
+	/*
+	 * As waking up waiters requires the spinlock to be acquired, only do so
+	 * if necessary.
+	 */
+	if (check_waiters)
+		BufferLockWakeup(buf_hdr, wake_exclusive);
+}
+
+/*
+ * BufferLockHeldByMeInMode - test whether my process holds the content lock
+ * in the specified mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode == mode;
+
+}
+
+/*
+ * BufferLockHeldByMe - test whether my process holds the content lock in any
+ * mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMe(BufferDesc *buf_hdr)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode != BUFFER_LOCK_UNLOCK;
+}
+
+/*
+ * Release the content lock for the buffer.
+ */
+void
+UnlockBuffer(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+
+	Assert(BufferIsPinned(buffer));
+	if (BufferIsLocal(buffer))
+		return;					/* local buffers need no lock */
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+	BufferLockUnlock(buffer, buf_hdr);
+}
+
+/*
+ * Acquire the content_lock for the buffer.
  */
 void
-LockBuffer(Buffer buffer, BufferLockMode mode)
+LockBufferInternal(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *buf;
+	BufferDesc *buf_hdr;
+
+	/*
+	 * We can't wait if we haven't got a PGPROC.  This should only occur
+	 * during bootstrap or shared memory initialization.  Put an Assert here
+	 * to catch unsafe coding practices.
+	 */
+	Assert(!(MyProc == NULL && IsUnderPostmaster));
+
+	/* handled in LockBuffer() wrapper */
+	Assert(mode != BUFFER_LOCK_UNLOCK);
 
 	Assert(BufferIsPinned(buffer));
 	if (BufferIsLocal(buffer))
 		return;					/* local buffers need no lock */
 
-	buf = GetBufferDescriptor(buffer - 1);
+	buf_hdr = GetBufferDescriptor(buffer - 1);
 
-	if (mode == BUFFER_LOCK_UNLOCK)
-		LWLockRelease(BufferDescriptorGetContentLock(buf));
-	else if (mode == BUFFER_LOCK_SHARE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	/*
+	 * Test the most frequent lock modes first. While a switch (mode) would be
+	 * nice, at least gcc generates considerably worse code for it.
+	 *
+	 * Call BufferLockAcquire() with a constant argument for mode, to generate
+	 * more efficient code for the different lock modes.
+	 */
+	if (mode == BUFFER_LOCK_SHARE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE);
 	else if (mode == BUFFER_LOCK_EXCLUSIVE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_EXCLUSIVE);
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_EXCLUSIVE);
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	else
 		elog(ERROR, "unrecognized buffer lock mode: %d", mode);
 }
@@ -5727,8 +6458,7 @@ ConditionalLockBuffer(Buffer buffer)
 
 	buf = GetBufferDescriptor(buffer - 1);
 
-	return LWLockConditionalAcquire(BufferDescriptorGetContentLock(buf),
-									LW_EXCLUSIVE);
+	return BufferLockConditional(buffer, buf, BUFFER_LOCK_EXCLUSIVE);
 }
 
 /*
@@ -6677,7 +7407,25 @@ ResOwnerReleaseBufferPin(Datum res)
 	if (BufferIsLocal(buffer))
 		UnpinLocalBufferNoOwner(buffer);
 	else
+	{
+		PrivateRefCountEntry *ref;
+
+		ref = GetPrivateRefCountEntry(buffer, false);
+
+		/*
+		 * If the buffer was locked at the time of the resowner release,
+		 * release the lock now. This should only happen after errors.
+		 */
+		if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
+		{
+			BufferDesc *buf = GetBufferDescriptor(buffer - 1);
+
+			HOLD_INTERRUPTS();	/* match the upcoming RESUME_INTERRUPTS */
+			BufferLockUnlock(buffer, buf);
+		}
+
 		UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
+	}
 }
 
 static char *
@@ -6913,10 +7661,10 @@ MarkDirtyUnpinnedBufferInternal(Buffer buf, BufferDesc *desc,
 	/* If it was not already dirty, mark it as dirty. */
 	if (!(buf_state & BM_DIRTY))
 	{
-		LWLockAcquire(BufferDescriptorGetContentLock(desc), LW_EXCLUSIVE);
+		BufferLockAcquire(buf, desc, BUFFER_LOCK_EXCLUSIVE);
 		MarkBufferDirty(buf);
 		result = true;
-		LWLockRelease(BufferDescriptorGetContentLock(desc));
+		BufferLockUnlock(buf, desc);
 	}
 	else
 		*buffer_already_dirty = true;
@@ -7167,16 +7915,12 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		 */
 		if (is_write && !is_temp)
 		{
-			LWLock	   *content_lock;
-
-			content_lock = BufferDescriptorGetContentLock(buf_hdr);
-
-			Assert(LWLockHeldByMe(content_lock));
+			Assert(BufferLockHeldByMe(buf_hdr));
 
 			/*
 			 * Lock is now owned by AIO subsystem.
 			 */
-			LWLockDisown(content_lock);
+			BufferLockDisown(buffer, buf_hdr);
 		}
 
 		/*
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 1e5e368a5dc..39ae7cdf856 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -285,6 +285,9 @@ ABI_compatibility:
 Section: ClassName - WaitEventBuffer
 
 BUFFER_CLEANUP	"Waiting to acquire an exclusive pin on a buffer. Buffer pin waits can be protracted if another process holds an open cursor that last read data from the buffer in question."
+BUFFER_SHARED	"Waiting to acquire shared lock on a buffer."
+BUFFER_SHARE_EXCLUSIVE	"Waiting to acquire share exclusive lock on a buffer."
+BUFFER_EXCLUSIVE	"Waiting to acquire exclusive lock on a buffer."
 
 ABI_compatibility:
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0013-Require-share-exclusive-lock-to-set-hint-bits-and.patch (37.9K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/14-v7-0013-Require-share-exclusive-lock-to-set-hint-bits-and.patch)
  download | inline diff:
From 9fbc92b06d8f2a4d5164dc7ff8b234634b96fd49 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 18 Nov 2025 09:22:28 -0500
Subject: [PATCH v7 13/15] Require share-exclusive lock to set hint bits and to
 flush

At the moment hint bits can be set with just a share lock on a page (and in
one place even without any lock). Because of this we need to copy pages while
writing them out, as otherwise the checksum could be corrupted.

The need to copy the page is problematic to implement AIO writes:

1) Instead of just needing a single buffer for a copied page we need one for
   each page that's potentially undergoing IO
2) To be able to use the "worker" AIO implementation the copied page needs to
   reside in shared memory

It also causes problems for using unbuffered/direct-IO, independent of AIO:
Some filesystems, raid implementations, ... do not tolerate the data being
written out to change during the write. E.g. they may compute internal
checksums that can be invalidated by concurrent modifications, leading e.g. to
filesystem errors (as the case with btrfs).

It also just is plain odd to allow modifications of buffers that are just
share locked.

To address these issue, this commit changes the rules so that modifications to
pages are not allowed anymore while holding a share lock. Instead the new
share-exclusive lock (introduced in FIXME XXXX TODO) allows at most one
backend to modify a buffer while other backends have the same page share
locked. An existing share-lock can be upgraded to a share-exclusive lock, if
there are no conflicting locks. For that
BufferBeginSetHintBits()/BufferBeginSetHintBits() and BufferSetHintBits16()
have been introduced.

To prevent hint bits from being set while the buffer is being written out,
writing out buffers now requires a share-exclusive lock.

The use of share-exclusive to gate setting hint bits means that from now on
only one backend can set hint bits at a time. To allow multiple backends
setting hint bits would require more complicated locking, for setting hint
bits we'd need to store the count of backends currently setting hint bits and
we would need another lock-level for I/O conflicting with the lock-level to
set hint bits. Given that the share-exclusive lock for setting hint bits is
only held for a short time, that often backends would just set the same hint
bits and that the cost of occasionally not setting hint bits in hotly accessed
pages is fairly low, this seems like an acceptable tradeoff.

The biggest change to adapt to this is in heapam. To avoid performance
regressions for sequential scans that need to set a lot of hint bits, we need
to amortize the cost of BufferBeginSetHintBits() for cases where hint bits are
set at a high frequency, HeapTupleSatisfiesMVCCBatch() uses the new
SetHintBitsExt() which defers BufferFinishSetHintBits() until all hint bits on
a page have been set.  Conversely, to avoid regressions in cases where we
can't set hint bits in bulk (because we're looking only at individual tuples),
use BufferSetHintBits16() when setting hint bits without batching.

Several other places also need to be adapted, but those changes are
comparatively simpler.

After this we do not need to copy buffers to write them out anymore. That
change is done separately however.

TODO:
- Update commit reference above
- reflow parts of storage/buffer/README that I didn't reindent to make the
  diff more readable

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf%40gcnactj4z56m
---
 src/include/storage/bufmgr.h                |   4 +
 src/backend/access/gist/gistget.c           |  19 +-
 src/backend/access/hash/hashutil.c          |  10 +-
 src/backend/access/heap/heapam_visibility.c | 125 ++++++--
 src/backend/access/nbtree/nbtinsert.c       |  28 +-
 src/backend/access/nbtree/nbtutils.c        |  16 +-
 src/backend/storage/buffer/README           |  44 ++-
 src/backend/storage/buffer/bufmgr.c         | 307 ++++++++++++++++----
 src/backend/storage/freespace/freespace.c   |  14 +-
 src/backend/storage/freespace/fsmpage.c     |  11 +-
 src/tools/pgindent/typedefs.list            |   1 +
 11 files changed, 459 insertions(+), 120 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index df170fe9553..6e7fec7e723 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -314,6 +314,10 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
+extern bool BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer);
+extern bool BufferBeginSetHintBits(Buffer buffer);
+extern void BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std);
+
 extern void UnlockBuffers(void);
 extern void UnlockBuffer(Buffer buffer);
 extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
diff --git a/src/backend/access/gist/gistget.c b/src/backend/access/gist/gistget.c
index 9ba45acfff3..956ece6bed5 100644
--- a/src/backend/access/gist/gistget.c
+++ b/src/backend/access/gist/gistget.c
@@ -63,11 +63,7 @@ gistkillitems(IndexScanDesc scan)
 	 * safe.
 	 */
 	if (BufferGetLSNAtomic(buffer) != so->curPageLSN)
-	{
-		UnlockReleaseBuffer(buffer);
-		so->numKilled = 0;		/* reset counter */
-		return;
-	}
+		goto unlock;
 
 	Assert(GistPageIsLeaf(page));
 
@@ -77,6 +73,16 @@ gistkillitems(IndexScanDesc scan)
 	 */
 	for (i = 0; i < so->numKilled; i++)
 	{
+		if (!killedsomething)
+		{
+			/*
+			 * Use hint bit infrastructure to be allowed to modify the page
+			 * without holding an exclusive lock.
+			 */
+			if (!BufferBeginSetHintBits(buffer))
+				goto unlock;
+		}
+
 		offnum = so->killedItems[i];
 		iid = PageGetItemId(page, offnum);
 		ItemIdMarkDead(iid);
@@ -86,9 +92,10 @@ gistkillitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		GistMarkPageHasGarbage(page);
-		MarkBufferDirtyHint(buffer, true);
+		BufferFinishSetHintBits(buffer, true, true);
 	}
 
+unlock:
 	UnlockReleaseBuffer(buffer);
 
 	/*
diff --git a/src/backend/access/hash/hashutil.c b/src/backend/access/hash/hashutil.c
index f41233fcd07..d1d603770b2 100644
--- a/src/backend/access/hash/hashutil.c
+++ b/src/backend/access/hash/hashutil.c
@@ -593,6 +593,13 @@ _hash_kill_items(IndexScanDesc scan)
 
 			if (ItemPointerEquals(&ituple->t_tid, &currItem->heapTid))
 			{
+				/*
+				 * Use hint bit infrastructure to be allowed to modify the
+				 * page without holding an exclusive lock.
+				 */
+				if (!BufferBeginSetHintBits(so->currPos.buf))
+					goto unlock_page;
+
 				/* found the item */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -610,9 +617,10 @@ _hash_kill_items(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->hasho_flag |= LH_PAGE_HAS_DEAD_TUPLES;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(so->currPos.buf, true, true);
 	}
 
+unlock_page:
 	if (so->hashso_bucket_buf == so->currPos.buf ||
 		havePin)
 		LockBuffer(so->currPos.buf, BUFFER_LOCK_UNLOCK);
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 5645cfd8a49..1a961517281 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -80,10 +80,38 @@
 
 
 /*
- * SetHintBits()
+ * To be allowed to set hint bits, SetHintBits() needs to call
+ * BufferBeginSetHintBits(). However, that's not free, and some callsites call
+ * SetHintBits() on many tuples in a row. For those it makes sense to amortize
+ * the cost of BufferBeginSetHintBits(). Additionally it's desirable to defer
+ * the cost of BufferBeginSetHintBits() until a hint bit needs to actually be
+ * set. This enum serves as the necessary state space passed to
+ * SetHintbitsExt().
+ */
+typedef enum SetHintBitsState
+{
+	/* not yet checked if hint bits may be set */
+	SHB_INITIAL,
+	/* failed to get permission to set hint bits, don't check again */
+	SHB_DISABLED,
+	/* allowed to set hint bits */
+	SHB_ENABLED,
+} SetHintBitsState;
+
+/*
+ * SetHintBitsExt()
  *
  * Set commit/abort hint bits on a tuple, if appropriate at this time.
  *
+ * To be allowed to set a hint bit on a tuple, the page must not be undergoing
+ * IO at this time (otherwise we e.g. could corrupt PG's page checksum or even
+ * the filesystem's, as is known to happen with btrfs).
+ *
+ * The right to set a hint bit can be acquired on a page level with
+ * BufferBeginSetHintBits(). Only a single backend gets the right to set hint
+ * bits at a time.  Alternatively, if called with a NULL SetHintBitsState*,
+ * hint bits are set with BufferSetHintBits16().
+ *
  * It is only safe to set a transaction-committed hint bit if we know the
  * transaction's commit record is guaranteed to be flushed to disk before the
  * buffer, or if the table is temporary or unlogged and will be obliterated by
@@ -111,24 +139,69 @@
  * InvalidTransactionId if no check is needed.
  */
 static inline void
-SetHintBits(HeapTupleHeader tuple, Buffer buffer,
-			uint16 infomask, TransactionId xid)
+SetHintBitsExt(HeapTupleHeader tuple, Buffer buffer,
+			   uint16 infomask, TransactionId xid, SetHintBitsState *state)
 {
+	/*
+	 * In batched mode and we previously did not get permission to set hint
+	 * bits. Don't try again, in all likelihood IO is still going on.
+	 */
+	if (state && *state == SHB_DISABLED)
+		return;
+
 	if (TransactionIdIsValid(xid))
 	{
-		/* NB: xid must be known committed here! */
-		XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+		if (BufferIsPermanent(buffer))
+		{
+			/* NB: xid must be known committed here! */
+			XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+
+			if (XLogNeedsFlush(commitLSN) &&
+				BufferGetLSNAtomic(buffer) < commitLSN)
+			{
+				/* not flushed and no LSN interlock, so don't set hint */
+				return;
+			}
+		}
+	}
+
+	/*
+	 * If we're not operating in batch mode, use BufferSetHintBits16() to mark
+	 * the page dirty, that's cheaper than
+	 * BufferBeginSetHintBits()/BufferFinishSetHintBits(). That's important
+	 * for cases where we set a lot of hint bits on a page individually.
+	 */
+	if (!state)
+	{
+		BufferSetHintBits16(&tuple->t_infomask,
+							tuple->t_infomask | infomask, buffer);
+		return;
+	}
 
-		if (BufferIsPermanent(buffer) && XLogNeedsFlush(commitLSN) &&
-			BufferGetLSNAtomic(buffer) < commitLSN)
+	if (*state == SHB_INITIAL)
+	{
+		if (!BufferBeginSetHintBits(buffer))
 		{
-			/* not flushed and no LSN interlock, so don't set hint */
+			*state = SHB_DISABLED;
 			return;
 		}
+
+		if (state)
+			*state = SHB_ENABLED;
+
 	}
-
 	tuple->t_infomask |= infomask;
-	MarkBufferDirtyHint(buffer, true);
+}
+
+/*
+ * Simple wrapper around SetHintBitExt(), use when operating on a single
+ * tuple.
+ */
+static inline void
+SetHintBits(HeapTupleHeader tuple, Buffer buffer,
+			uint16 infomask, TransactionId xid)
+{
+	SetHintBitsExt(tuple, buffer, infomask, xid, NULL);
 }
 
 /*
@@ -864,9 +937,9 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
  * inserting/deleting transaction was still running --- which was more cycles
  * and more contention on ProcArrayLock.
  */
-static bool
+static inline bool
 HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
-					   Buffer buffer)
+					   Buffer buffer, SetHintBitsState *state)
 {
 	HeapTupleHeader tuple = htup->t_data;
 
@@ -921,8 +994,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 			if (!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmax(tuple)))
 			{
 				/* deleting subtransaction must have aborted */
-				SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-							InvalidTransactionId);
+				SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+							   InvalidTransactionId, state);
 				return true;
 			}
 
@@ -934,13 +1007,13 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		else if (XidInMVCCSnapshot(HeapTupleHeaderGetRawXmin(tuple), snapshot))
 			return false;
 		else if (TransactionIdDidCommit(HeapTupleHeaderGetRawXmin(tuple)))
-			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-						HeapTupleHeaderGetRawXmin(tuple));
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_COMMITTED,
+						   HeapTupleHeaderGetRawXmin(tuple), state);
 		else
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_INVALID,
+						   InvalidTransactionId, state);
 			return false;
 		}
 	}
@@ -1003,14 +1076,14 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		if (!TransactionIdDidCommit(HeapTupleHeaderGetRawXmax(tuple)))
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+						   InvalidTransactionId, state);
 			return true;
 		}
 
 		/* xmax transaction committed */
-		SetHintBits(tuple, buffer, HEAP_XMAX_COMMITTED,
-					HeapTupleHeaderGetRawXmax(tuple));
+		SetHintBitsExt(tuple, buffer, HEAP_XMAX_COMMITTED,
+					   HeapTupleHeaderGetRawXmax(tuple), state);
 	}
 	else
 	{
@@ -1606,6 +1679,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 							OffsetNumber *vistuples_dense)
 {
 	int			nvis = 0;
+	SetHintBitsState state = SHB_INITIAL;
 #ifdef BATCHMVCC_FEWER_ARGS
 	HeapTupleData *tuples = batchmvcc->tuples;
 	bool	   *visible = batchmvcc->visible;
@@ -1618,7 +1692,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		bool		valid;
 		HeapTuple	tup = &tuples[i];
 
-		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer, &state);
 		visible[i] = valid;
 
 		if (likely(valid))
@@ -1628,6 +1702,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		}
 	}
 
+	if (state == SHB_ENABLED)
+		BufferFinishSetHintBits(buffer, true, true);
+
 	return nvis;
 }
 
@@ -1647,7 +1724,7 @@ HeapTupleSatisfiesVisibility(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 	switch (snapshot->snapshot_type)
 	{
 		case SNAPSHOT_MVCC:
-			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer);
+			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer, NULL);
 		case SNAPSHOT_SELF:
 			return HeapTupleSatisfiesSelf(htup, snapshot, buffer);
 		case SNAPSHOT_ANY:
diff --git a/src/backend/access/nbtree/nbtinsert.c b/src/backend/access/nbtree/nbtinsert.c
index 7c113c007e5..545e1d7d9e0 100644
--- a/src/backend/access/nbtree/nbtinsert.c
+++ b/src/backend/access/nbtree/nbtinsert.c
@@ -680,20 +680,28 @@ _bt_check_unique(Relation rel, BTInsertState insertstate, Relation heapRel,
 				{
 					/*
 					 * The conflicting tuple (or all HOT chains pointed to by
-					 * all posting list TIDs) is dead to everyone, so mark the
-					 * index entry killed.
+					 * all posting list TIDs) is dead to everyone, so try to
+					 * mark the index entry killed. It's ok if we're not
+					 * allowed to, this isn't required for correctness.
 					 */
-					ItemIdMarkDead(curitemid);
-					opaque->btpo_flags |= BTP_HAS_GARBAGE;
+					Buffer		buf;
 
-					/*
-					 * Mark buffer with a dirty hint, since state is not
-					 * crucial. Be sure to mark the proper buffer dirty.
-					 */
+					/* Be sure to operate on the proper buffer */
 					if (nbuf != InvalidBuffer)
-						MarkBufferDirtyHint(nbuf, true);
+						buf = nbuf;
 					else
-						MarkBufferDirtyHint(insertstate->buf, true);
+						buf = insertstate->buf;
+
+					/*
+					 * Can't use BufferSetHintBits16() here as we update two
+					 * different locations.
+					 */
+					if (BufferBeginSetHintBits(buf))
+					{
+						ItemIdMarkDead(curitemid);
+						opaque->btpo_flags |= BTP_HAS_GARBAGE;
+						BufferFinishSetHintBits(buf, true, true);
+					}
 				}
 
 				/*
diff --git a/src/backend/access/nbtree/nbtutils.c b/src/backend/access/nbtree/nbtutils.c
index ab0f98b0287..34e548b9930 100644
--- a/src/backend/access/nbtree/nbtutils.c
+++ b/src/backend/access/nbtree/nbtutils.c
@@ -3542,10 +3542,19 @@ _bt_killitems(IndexScanDesc scan)
 			 * it's possible that multiple processes attempt to do this
 			 * simultaneously, leading to multiple full-page images being sent
 			 * to WAL (if wal_log_hints or data checksums are enabled), which
-			 * is undesirable.
+			 * is undesirable.  We need to use the hint bit infrastructure to
+			 * update the page while just holding a share lock.
 			 */
 			if (killtuple && !ItemIdIsDead(iid))
 			{
+				/*
+				 * If we're not able to set hint bits, there's no point
+				 * continuing.
+				 */
+				if (!killedsomething &&
+					!BufferBeginSetHintBits(buf))
+					goto unlock_page;
+
 				/* found the item/all posting list items */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -3556,8 +3565,6 @@ _bt_killitems(IndexScanDesc scan)
 	}
 
 	/*
-	 * Since this can be redone later if needed, mark as dirty hint.
-	 *
 	 * Whenever we mark anything LP_DEAD, we also set the page's
 	 * BTP_HAS_GARBAGE flag, which is likewise just a hint.  (Note that we
 	 * only rely on the page-level flag in !heapkeyspace indexes.)
@@ -3565,9 +3572,10 @@ _bt_killitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->btpo_flags |= BTP_HAS_GARBAGE;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(buf, true, true);
 	}
 
+unlock_page:
 	if (!so->dropPin)
 		_bt_unlockbuf(rel, buf);
 	else
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index 119f31b5d65..5e6735edd23 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -25,21 +25,26 @@ that might need to do such a wait is instead handled by waiting to obtain
 the relation-level lock, which is why you'd better hold one first.)  Pins
 may not be held across transaction boundaries, however.
 
-Buffer content locks: there are two kinds of buffer lock, shared and exclusive,
-which act just as you'd expect: multiple backends can hold shared locks on
-the same buffer, but an exclusive lock prevents anyone else from holding
-either shared or exclusive lock.  (These can alternatively be called READ
-and WRITE locks.)  These locks are intended to be short-term: they should not
-be held for long.  Buffer locks are acquired and released by LockBuffer().
-It will *not* work for a single backend to try to acquire multiple locks on
-the same buffer.  One must pin a buffer before trying to lock it.
+Buffer content locks: there three kinds of buffer lock, shared,
+share-exclusive and exclusive:
+a) multiple backends can hold shared locks on the same buffer
+   (alternatively called a READ lock)
+b) one backend can hold an share-exclusive lock on a buffer while multiple
+   backends can hold a share lock
+c) an exclusive lock prevents anyone else from holding shared, share-exclusive
+   or exclusive lock.
+   (alternatively called a WRITE lock)
+
+These locks are intended to be short-term: they should not be held for long.
+Buffer locks are acquired and released by LockBuffer().  It will *not* work
+for a single backend to try to acquire multiple locks on the same buffer.  One
+must pin a buffer before trying to lock it.
 
 Buffer access rules:
 
-1. To scan a page for tuples, one must hold a pin and either shared or
-exclusive content lock.  To examine the commit status (XIDs and status bits)
-of a tuple in a shared buffer, one must likewise hold a pin and either shared
-or exclusive lock.
+1. To scan a page for tuples, one must hold a pin and at least a share lock.
+To examine the commit status (XIDs and status bits) of a tuple in a shared
+buffer, one must likewise hold a pin and at least a share lock.
 
 2. Once one has determined that a tuple is interesting (visible to the
 current transaction) one may drop the content lock, yet continue to access
@@ -55,8 +60,14 @@ one must hold a pin and an exclusive content lock on the containing buffer.
 This ensures that no one else might see a partially-updated state of the
 tuple while they are doing visibility checks.
 
-4. It is considered OK to update tuple commit status bits (ie, OR the
-values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
+4. Non-critical information on a page ("hint bits") may be modified while
+holding only a share-exclusive lock and pin on the page. To do so in cases
+where only a share lock is already held, use BufferBeginSetHintBits() &
+BufferFinishSetHintBits() (if multiple hint bits are to be set) or
+BufferSetHintBits16() (if a single hit bit is set).
+
+E.g. for heapam, a share-exclusive lock allows to update tuple commit status
+bits (ie, OR the values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
 HEAP_XMAX_INVALID into t_infomask) while holding only a shared lock and
 pin on a buffer.  This is OK because another backend looking at the tuple
 at about the same time would OR the same bits into the field, so there
@@ -80,7 +91,6 @@ buffer (increment the refcount) while one is performing the cleanup, but
 it won't be able to actually examine the page until it acquires shared
 or exclusive content lock.
 
-
 Obtaining the lock needed under rule #5 is done by the bufmgr routines
 LockBufferForCleanup() or ConditionalLockBufferForCleanup().  They first get
 an exclusive lock and then check to see if the shared pin count is currently
@@ -96,6 +106,10 @@ VACUUM's use, since we don't allow multiple VACUUMs concurrently on a single
 relation anyway.  Anyone wishing to obtain a cleanup lock outside of recovery
 or a VACUUM must use the conditional variant of the function.
 
+6. To write out a buffer, a share-exclusive lock needs to be held. This
+prevents the buffer from being modified while written out, which could corrupt
+checksums and cause issues on the OS or device level when direct-IO is used.
+
 
 Buffer Manager's Internal Locking
 ---------------------------------
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 32634858514..7409c4e7e42 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2463,9 +2463,8 @@ again:
 	/*
 	 * If the buffer was dirty, try to write it out.  There is a race
 	 * condition here, in that someone might dirty it after we released the
-	 * buffer header lock above, or even while we are writing it out (since
-	 * our share-lock won't prevent hint-bit updates).  We will recheck the
-	 * dirty bit after re-locking the buffer header.
+	 * buffer header lock above.  We will recheck the dirty bit after
+	 * re-locking the buffer header.
 	 */
 	if (buf_state & BM_DIRTY)
 	{
@@ -2473,12 +2472,12 @@ again:
 		Assert(buf_state & BM_VALID);
 
 		/*
-		 * We need a share-lock on the buffer contents to write it out (else
+		 * We need a share-exclusive lock on the buffer contents to write it out (else
 		 * we might write invalid data, eg because someone else is compacting
 		 * the page contents while we write).  We must use a conditional lock
 		 * acquisition here to avoid deadlock.  Even though the buffer was not
 		 * pinned (and therefore surely not locked) when StrategyGetBuffer
-		 * returned it, someone else could have pinned and exclusive-locked it
+		 * returned it, someone else could have pinned and (share-)exclusive-locked it
 		 * by the time we get here. If we try to get the lock unconditionally,
 		 * we'd block waiting for them; if they later block waiting for us,
 		 * deadlock ensues. (This has been observed to happen when two
@@ -2486,7 +2485,7 @@ again:
 		 * one just happens to be trying to split the page the first one got
 		 * from StrategyGetBuffer.)
 		 */
-		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -4055,8 +4054,8 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	}
 
 	/*
-	 * Pin it, share-lock it, write it.  (FlushBuffer will do nothing if the
-	 * buffer is clean by the time we've locked it.)
+	 * Pin it, share-exclusive-lock it, write it.  (FlushBuffer will do
+	 * nothing if the buffer is clean by the time we've locked it.)
 	 */
 	PinBuffer_Locked(bufHdr);
 
@@ -4386,11 +4385,8 @@ BufferGetTag(Buffer buffer, RelFileLocator *rlocator, ForkNumber *forknum,
  * However, we will need to force the changes to disk via fsync before
  * we can checkpoint WAL.
  *
- * The caller must hold a pin on the buffer and have share-locked the
- * buffer contents.  (Note: a share-lock does not prevent updates of
- * hint bits in the buffer, so the page could change while the write
- * is in progress, but we assume that that will not invalidate the data
- * written.)
+ * The caller must hold a pin on the buffer and have
+ * (share-)exclusively-locked the buffer contents.
  *
  * If the caller has an smgr reference for the buffer's relation, pass it
  * as the second parameter.  If not, pass NULL.
@@ -4406,6 +4402,9 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	char	   *bufToWrite;
 	uint64		buf_state;
 
+	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(buf, BUFFER_LOCK_SHARE_EXCLUSIVE));
+
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
 	 * someone else flushed the buffer before we could, so we need not do
@@ -4538,7 +4537,7 @@ FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(buf);
 
-	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 	BufferLockUnlock(buffer, buf);
 }
@@ -5457,8 +5456,8 @@ FlushDatabaseBuffers(Oid dbid)
 }
 
 /*
- * Flush a previously, shared or exclusively, locked and pinned buffer to the
- * OS.
+ * Flush a previously, share-exclusively or exclusively, locked and pinned
+ * buffer to the OS.
  */
 void
 FlushOneBuffer(Buffer buffer)
@@ -5531,39 +5530,23 @@ IncrBufferRefCount(Buffer buffer)
 }
 
 /*
- * MarkBufferDirtyHint
+ * Shared-buffer only helper for MarkBufferDirtyHint() and
+ * BufferSetHintBits16().
  *
- *	Mark a buffer dirty for non-critical changes.
- *
- * This is essentially the same as MarkBufferDirty, except:
- *
- * 1. The caller does not write WAL; so if checksums are enabled, we may need
- *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
- * 2. The caller might have only share-lock instead of exclusive-lock on the
- *	  buffer's content lock.
- * 3. This function does not guarantee that the buffer is always marked dirty
- *	  (due to a race condition), so it cannot be used for important changes.
+ * This is separated out because it turns out that the repeated checks for
+ * local buffers, repeated GetBufferDescriptor() and repeated reading of the
+ * buffer's state sufficiently hurts the performance of BufferSetHintBits16().
  */
-void
-MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+static inline void
+MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate, bool buffer_std)
 {
-	BufferDesc *bufHdr;
 	Page		page = BufferGetPage(buffer);
 
-	if (!BufferIsValid(buffer))
-		elog(ERROR, "bad buffer ID: %d", buffer);
-
-	if (BufferIsLocal(buffer))
-	{
-		MarkLocalBufferDirty(buffer);
-		return;
-	}
-
-	bufHdr = GetBufferDescriptor(buffer - 1);
-
 	Assert(GetPrivateRefCount(buffer) > 0);
-	/* here, either share or exclusive lock is OK */
-	Assert(BufferIsLockedByMe(buffer));
+
+	/* here, either share-exclusive or exclusive lock is OK */
+	Assert(BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_SHARE_EXCLUSIVE));
 
 	/*
 	 * This routine might get called many times on the same page, if we are
@@ -5576,8 +5559,8 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
-		(BM_DIRTY | BM_JUST_DIRTIED))
+	if (unlikely((lockstate & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+				 (BM_DIRTY | BM_JUST_DIRTIED)))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
@@ -5593,8 +5576,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * We don't check full_page_writes here because that logic is included
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
-		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
+		if (XLogHintBitIsNeeded() && (lockstate & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5646,13 +5628,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 			dirtied = true;		/* Means "will be dirtied by this action" */
 
 			/*
-			 * Set the page LSN if we wrote a backup block. We aren't supposed
-			 * to set this when only holding a share lock but as long as we
-			 * serialise it somehow we're OK. We choose to set LSN while
-			 * holding the buffer header lock, which causes any reader of an
-			 * LSN who holds only a share lock to also obtain a buffer header
-			 * lock before using PageGetLSN(), which is enforced in
-			 * BufferGetLSNAtomic().
+			 * Set the page LSN if we wrote a backup block. To allow backends
+			 * that only hold a share lock on the buffer to read the LSN in a
+			 * tear-free manner, we set the page LSN while holding the buffer
+			 * header lock. This allows any reader of an LSN who holds only a
+			 * share lock to also obtain a buffer header lock before using
+			 * PageGetLSN() to read the LSN in a tear free way. This is done
+			 * in BufferGetLSNAtomic().
 			 *
 			 * If checksums are enabled, you might think we should reset the
 			 * checksum here. That will happen when the page is written
@@ -5678,6 +5660,41 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	}
 }
 
+/*
+ * MarkBufferDirtyHint
+ *
+ *	Mark a buffer dirty for non-critical changes.
+ *
+ * This is essentially the same as MarkBufferDirty, except:
+ *
+ * 1. The caller does not write WAL; so if checksums are enabled, we may need
+ *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
+ * 2. The caller might have only share-exclusive-lock instead of
+ *	  exclusive-lock on the buffer's content lock.
+ * 3. This function does not guarantee that the buffer is always marked dirty
+ *	  (due to a race condition), so it cannot be used for important changes.
+ */
+inline void
+MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+{
+	BufferDesc *bufHdr;
+
+	bufHdr = GetBufferDescriptor(buffer - 1);
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		MarkLocalBufferDirty(buffer);
+		return;
+	}
+
+	MarkSharedBufferDirtyHint(buffer, bufHdr,
+							  pg_atomic_read_u64(&bufHdr->state),
+							  buffer_std);
+}
+
 /*
  * Release buffer content locks for shared buffers.
  *
@@ -6773,6 +6790,188 @@ IsBufferCleanupOK(Buffer buffer)
 	return false;
 }
 
+/*
+ * Helper for BufferBeginSetHintBits() and BufferSetHintBits16().
+ *
+ * This checks if the current lock mode already suffices to allow hint bits
+ * being set and, if not, whether the current lock can be upgraded.
+ */
+static inline bool
+SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
+{
+	uint64		old_state;
+	PrivateRefCountEntry *ref;
+	BufferLockMode mode;
+
+	ref = GetPrivateRefCountEntry(buffer, true);
+
+	if (ref == NULL)
+		elog(ERROR, "lock is not held");
+
+	mode = ref->data.lockmode;
+	if (mode == BUFFER_LOCK_UNLOCK)
+		elog(ERROR, "buffer is not locked");
+
+	/*
+	 * Already am holding a sufficient lock level.
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE || mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+	{
+		*lockstate = pg_atomic_read_u64(&buf_hdr->state);
+		return true;
+	}
+
+	/*
+	 * Only holding a share lock right now, try to upgrade to SHARE_EXCLUSIVE.
+	 */
+	Assert(mode == BUFFER_LOCK_SHARE);
+
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+	while (true)
+	{
+		uint64		desired_state;
+
+		desired_state = old_state;
+
+		/*
+		 * Can't upgrade if somebody else holds the lock in exlusive or
+		 * share-exclusive mode.
+		 */
+		if (unlikely((old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) != 0))
+		{
+			return false;
+		}
+
+		/* currently held lock state */
+		desired_state -= BM_LOCK_VAL_SHARED;
+
+		/* new lock level */
+		desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			ref->data.lockmode = BUFFER_LOCK_SHARE_EXCLUSIVE;
+			*lockstate = desired_state;
+
+			return true;
+		}
+	}
+
+}
+
+/*
+ * Try to acquire the right to set hint bits on the buffer.
+ *
+ * To be allowed to set hint bits, this backend needs to hold either a
+ * share-exclusive or an exclusive lock. In case this backend only holds a
+ * share lock, this function will try to upgrade the lock to
+ * share-exclusive. The caller is only allowed to set hint bits if true is
+ * returned.
+ *
+ * Once BufferBeginSetHintBits() has returned true, hint bits may be set
+ * without further calls to BufferBeginSetHintBits(), until the buffer is
+ * unlocked.
+ *
+ *
+ * Requiring a share-exclusive lock to set hint bits prevents setting hint
+ * bits on buffers that are currently being written out, which could corrupt
+ * the checksum on the page. Flushing buffers also requires a share-exclusive
+ * lock.
+ *
+ * Due to a lock >= share-exclusive being required to set hint bits, only one
+ * backend can set hint bits at a time. To allow multiple backends setting
+ * hint bits would require more complicated locking, for setting hint bits
+ * we'd need to store the count of backends currently setting hint bits and we
+ * would need another lock-level for I/O conflicting with the lock-level to
+ * set hint bits. Given that the share-exclusive lock for setting hint bits is
+ * only held for a short time, that often backends would just set the same
+ * hint bits and that the cost of occasionally not setting hint bits in hotly
+ * accessed pages is fairly low, this seems like an acceptable tradeoff.
+ */
+bool
+BufferBeginSetHintBits(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+
+	if (BufferIsLocal(buffer))
+	{
+		/*
+		 * TODO: will need to check for write IO once that's done
+		 * asynchronously.
+		 */
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	return SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate);
+}
+
+/*
+ * End a phase of setting hint bits on this buffer, started with
+ * BufferBeginSetHintBits().
+ *
+ * This would strictly speaking not be required (i.e. the caller could do
+ * MarkBufferDirtyHint() if so desired), but allows us to perform some sanity
+ * checks.
+ */
+void
+BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std)
+{
+	if (!BufferIsLocal(buffer))
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+
+	if (mark_dirty)
+		MarkBufferDirtyHint(buffer, buffer_std);
+}
+
+/*
+ * Ty to set a single hint bit in a buffer.
+ *
+ * This is a bit faster than BufferBeginSetHintBits() /
+ * BufferFinishSetHintBits() when setting a single hint bit, but slower than
+ * the former when setting several hint bits.
+ */
+bool
+BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+#ifdef USE_ASSERT_CHECKING
+	char	   *page;
+
+	/* verify that the address is on the page */
+	page = BufferGetPage(buffer);
+	Assert((char *) ptr >= page && (char *) ptr < (page + BLCKSZ));
+#endif
+
+	if (BufferIsLocal(buffer))
+	{
+		*ptr = val;
+
+		MarkLocalBufferDirty(buffer);
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	if (SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate))
+	{
+		*ptr = val;
+
+		MarkSharedBufferDirtyHint(buffer, buf_hdr, lockstate, true);
+
+		return true;
+	}
+
+	return false;
+}
+
 
 /*
  *	Functions for buffer I/O handling
diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index 48ac15d3487..bd4a2cff3a4 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -904,13 +904,17 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 	max_avail = fsm_get_max_avail(page);
 
 	/*
-	 * Reset the next slot pointer. This encourages the use of low-numbered
-	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation. We don't bother with marking the page dirty if it wasn't
-	 * already, since this is just a hint.
+	 * Try to reset the next slot pointer. This encourages the use of
+	 * low-numbered pages, increasing the chances that a later vacuum can
+	 * truncate the relation. We don't bother with marking the page dirty if
+	 * it wasn't already, since this is just a hint.
 	 */
 	LockBuffer(buf, BUFFER_LOCK_SHARE);
-	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	if (BufferBeginSetHintBits(buf))
+	{
+		((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+		BufferFinishSetHintBits(buf, false, false);
+	}
 	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 	ReleaseBuffer(buf);
diff --git a/src/backend/storage/freespace/fsmpage.c b/src/backend/storage/freespace/fsmpage.c
index 66a5c80b5a6..a59696b6484 100644
--- a/src/backend/storage/freespace/fsmpage.c
+++ b/src/backend/storage/freespace/fsmpage.c
@@ -298,9 +298,18 @@ restart:
 	 * lock and get a garbled next pointer every now and then, than take the
 	 * concurrency hit of an exclusive lock.
 	 *
+	 * Without an exclusive lock, we need to use the hint bit infrastructure
+	 * to be allowed to modify the page.
+	 *
 	 * Wrap-around is handled at the beginning of this function.
 	 */
-	fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+	if (exclusive_lock_held || BufferBeginSetHintBits(buf))
+	{
+		fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+
+		if (!exclusive_lock_held)
+			BufferFinishSetHintBits(buf, false, true);
+	}
 
 	return slot;
 }
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 6641a988ae8..f76a67ed256 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2736,6 +2736,7 @@ SetConstraintStateData
 SetConstraintTriggerData
 SetExprState
 SetFunctionReturnMode
+SetHintBitsState
 SetOp
 SetOpCmd
 SetOpPath
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0014-WIP-Make-UnlockReleaseBuffer-more-efficient.patch (3.5K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/15-v7-0014-WIP-Make-UnlockReleaseBuffer-more-efficient.patch)
  download | inline diff:
From 3a8c7fe452116946c4050c6bab9a936b33374c6f Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 19 Nov 2025 15:32:20 -0500
Subject: [PATCH v7 14/15] WIP: Make UnlockReleaseBuffer() more efficient

Now that the buffer content lock is implemented as part of BufferDesc.state,
releasing the lock and unpinning the buffer can be implemented as a single
atomic operation.

Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/nbtree/nbtpage.c | 22 +++++++++++-
 src/backend/storage/buffer/bufmgr.c | 52 ++++++++++++++++++++++++++++-
 2 files changed, 72 insertions(+), 2 deletions(-)

diff --git a/src/backend/access/nbtree/nbtpage.c b/src/backend/access/nbtree/nbtpage.c
index 30b43a4dd18..2fd8141854c 100644
--- a/src/backend/access/nbtree/nbtpage.c
+++ b/src/backend/access/nbtree/nbtpage.c
@@ -1006,11 +1006,18 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 
 	Assert(BlockNumberIsValid(blkno));
 	if (BufferIsValid(obuf))
+	{
+		_bt_relbuf(rel, obuf);
+#if 0
+		Assert(BufferGetBlockNumber(obuf) != blkno);
 		_bt_unlockbuf(rel, obuf);
-	buf = ReleaseAndReadBuffer(obuf, rel, blkno);
+#endif
+	}
+	buf = ReadBuffer(rel, blkno);
 	_bt_lockbuf(rel, buf, access);
 
 	_bt_checkpage(rel, buf);
+
 	return buf;
 }
 
@@ -1022,8 +1029,21 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 void
 _bt_relbuf(Relation rel, Buffer buf)
 {
+#if 0
 	_bt_unlockbuf(rel, buf);
 	ReleaseBuffer(buf);
+#else
+	/*
+	 * Buffer is pinned and locked, which means that it is expected to be
+	 * defined and addressable.  Check that proactively.
+	 */
+	VALGRIND_CHECK_MEM_IS_DEFINED(BufferGetPage(buf), BLCKSZ);
+
+	UnlockReleaseBuffer(buf);
+
+	if (!RelationUsesLocalBuffers(rel))
+		VALGRIND_MAKE_MEM_NOACCESS(BufferGetPage(buf), BLCKSZ);
+#endif
 }
 
 /*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 7409c4e7e42..3437854e766 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5494,13 +5494,63 @@ ReleaseBuffer(Buffer buffer)
 /*
  * UnlockReleaseBuffer -- release the content lock and pin on a buffer
  *
- * This is just a shorthand for a common combination.
+ * This is just a, more efficient, shorthand for a common combination.
  */
 void
 UnlockReleaseBuffer(Buffer buffer)
 {
+#if 1
+	int			mode;
+	BufferDesc *buf;
+	PrivateRefCountEntry *ref;
+	uint64		sub;
+	uint64		lockstate;
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		UnpinLocalBuffer(buffer);
+		return;
+	}
+
+	ResourceOwnerForgetBuffer(CurrentResourceOwner, buffer);
+
+	buf = GetBufferDescriptor(buffer - 1);
+
+	mode = BufferLockDisownInternal(buffer, buf);
+
+	/* compute state modification for lock release */
+	sub = BufferLockReleaseSub(mode);
+
+	/* compute state modification for pin release */
+	ref = GetPrivateRefCountEntry(buffer, false);
+	Assert(ref != NULL);
+	Assert(ref->data.refcount > 0);
+	ref->data.refcount--;
+
+	if (ref->data.refcount == 0)
+	{
+		sub |= BUF_REFCOUNT_ONE;
+		ForgetPrivateRefCountEntry(ref);
+	}
+
+	/* perform the lock and pin release in one atomic op */
+	lockstate = pg_atomic_sub_fetch_u64(&buf->state, sub);
+
+	/* wake up waiters etc */
+	BufferLockProcessRelease(buf, mode, lockstate);
+
+	if (lockstate & BM_PIN_COUNT_WAITER)
+		WakePinCountWaiter(buf);
+
+	RESUME_INTERRUPTS();
+
+#else
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 	ReleaseBuffer(buffer);
+#endif
 }
 
 /*
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v7-0015-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch (11.6K, ../../lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg@noioxkquezdw/16-v7-0015-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch)
  download | inline diff:
From c834c7c2e297000a23dcbdbd12735413cd467827 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 17 Oct 2024 14:14:35 -0400
Subject: [PATCH v7 15/15] WIP: bufmgr: Don't copy pages while writing out

After the series of preceding commits introducing and using
BufferBeginSetHintBits()/BufferSetHintBits16() hint bits are not set
anymore while IO is going on. Therefore we do not need to copy pages while
they are being written out anymore.

TODO: Update comments

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/bufpage.h           |  3 +-
 src/backend/access/hash/hashpage.c      |  2 +-
 src/backend/access/transam/xloginsert.c | 43 ++++++----------------
 src/backend/storage/buffer/bufmgr.c     | 21 +++++------
 src/backend/storage/buffer/localbuf.c   |  2 +-
 src/backend/storage/page/bufpage.c      | 48 ++++---------------------
 src/backend/storage/smgr/bulk_write.c   |  2 +-
 src/test/modules/test_aio/test_aio.c    |  2 +-
 8 files changed, 33 insertions(+), 90 deletions(-)

diff --git a/src/include/storage/bufpage.h b/src/include/storage/bufpage.h
index abc2cf2a020..f8f621446c4 100644
--- a/src/include/storage/bufpage.h
+++ b/src/include/storage/bufpage.h
@@ -504,7 +504,6 @@ extern void PageIndexMultiDelete(Page page, OffsetNumber *itemnos, int nitems);
 extern void PageIndexTupleDeleteNoCompact(Page page, OffsetNumber offnum);
 extern bool PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 									const void *newtup, Size newsize);
-extern char *PageSetChecksumCopy(Page page, BlockNumber blkno);
-extern void PageSetChecksumInplace(Page page, BlockNumber blkno);
+extern void PageSetChecksum(Page page, BlockNumber blkno);
 
 #endif							/* BUFPAGE_H */
diff --git a/src/backend/access/hash/hashpage.c b/src/backend/access/hash/hashpage.c
index b8e5bd005e5..dd17eff59d1 100644
--- a/src/backend/access/hash/hashpage.c
+++ b/src/backend/access/hash/hashpage.c
@@ -1029,7 +1029,7 @@ _hash_alloc_buckets(Relation rel, BlockNumber firstblock, uint32 nblocks)
 					zerobuf.data,
 					true);
 
-	PageSetChecksumInplace(page, lastblock);
+	PageSetChecksum(page, lastblock);
 	smgrextend(RelationGetSmgr(rel), MAIN_FORKNUM, lastblock, zerobuf.data,
 			   false);
 
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index a56d5a55282..0af148e9496 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -261,8 +261,11 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	 */
 #ifdef USE_ASSERT_CHECKING
 	if (!(flags & REGBUF_NO_CHANGE))
-		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) &&
-			   BufferIsDirty(buffer));
+	{
+		Assert(BufferIsDirty(buffer));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE));
+	}
 #endif
 
 	if (block_id >= max_registered_block_id)
@@ -1066,7 +1069,7 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * Write a backup block if needed when we are setting a hint. Note that
  * this may be called for a variety of page types, not just heaps.
  *
- * Callable while holding just share lock on the buffer content.
+ * Callable while holding just share-exclusive lock on the buffer content.
  *
  * We can't use the plain backup block mechanism since that relies on the
  * Buffer being exclusively locked. Since some modifications (setting LSN, hint
@@ -1074,6 +1077,8 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * failures. So instead we copy the page and insert the copied data as normal
  * record data.
  *
+ * FIXME: outdated
+ *
  * We only need to do something if page has not yet been full page written in
  * this checkpoint round. The LSN of the inserted wal record is returned if we
  * had to write, InvalidXLogRecPtr otherwise.
@@ -1102,46 +1107,20 @@ XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
 
 	/*
 	 * We assume page LSN is first data on *every* page that can be passed to
-	 * XLogInsert, whether it has the standard page layout or not. Since we're
-	 * only holding a share-lock on the page, we must take the buffer header
-	 * lock when we look at the LSN.
+	 * XLogInsert, whether it has the standard page layout or not.
 	 */
 	lsn = BufferGetLSNAtomic(buffer);
 
 	if (lsn <= RedoRecPtr)
 	{
-		int			flags = 0;
-		PGAlignedBlock copied_buffer;
-		char	   *origdata = (char *) BufferGetBlock(buffer);
-		RelFileLocator rlocator;
-		ForkNumber	forkno;
-		BlockNumber blkno;
-
-		/*
-		 * Copy buffer so we don't have to worry about concurrent hint bit or
-		 * lsn updates. We assume pd_lower/upper cannot be changed without an
-		 * exclusive lock, so the contents bkp are not racy.
-		 */
-		if (buffer_std)
-		{
-			/* Assume we can omit data between pd_lower and pd_upper */
-			Page		page = BufferGetPage(buffer);
-			uint16		lower = ((PageHeader) page)->pd_lower;
-			uint16		upper = ((PageHeader) page)->pd_upper;
-
-			memcpy(copied_buffer.data, origdata, lower);
-			memcpy(copied_buffer.data + upper, origdata + upper, BLCKSZ - upper);
-		}
-		else
-			memcpy(copied_buffer.data, origdata, BLCKSZ);
+		int			flags = REGBUF_NO_CHANGE;
 
 		XLogBeginInsert();
 
 		if (buffer_std)
 			flags |= REGBUF_STANDARD;
 
-		BufferGetTag(buffer, &rlocator, &forkno, &blkno);
-		XLogRegisterBlock(0, &rlocator, forkno, blkno, copied_buffer.data, flags);
+		XLogRegisterBuffer(0, buffer, flags);
 
 		recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI_FOR_HINT);
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 3437854e766..83d21ca16bd 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -4399,7 +4399,6 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	ErrorContextCallback errcallback;
 	instr_time	io_start;
 	Block		bufBlock;
-	char	   *bufToWrite;
 	uint64		buf_state;
 
 	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
@@ -4470,12 +4469,8 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	 */
 	bufBlock = BufHdrGetBlock(buf);
 
-	/*
-	 * Update page checksum if desired.  Since we have only shared lock on the
-	 * buffer, other processes might be updating hint bits in it, so we must
-	 * copy the page to private storage if we do checksumming.
-	 */
-	bufToWrite = PageSetChecksumCopy((Page) bufBlock, buf->tag.blockNum);
+	/* Update page checksum if desired. */
+	PageSetChecksum((Page) bufBlock, buf->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
@@ -4485,7 +4480,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	smgrwrite(reln,
 			  BufTagGetForkNum(&buf->tag),
 			  buf->tag.blockNum,
-			  bufToWrite,
+			  bufBlock,
 			  false);
 
 	/*
@@ -4609,8 +4604,8 @@ BufferIsPermanent(Buffer buffer)
 /*
  * BufferGetLSNAtomic
  *		Retrieves the LSN of the buffer atomically using a buffer header lock.
- *		This is necessary for some callers who may not have an exclusive lock
- *		on the buffer.
+ *		This is necessary for some callers who may not have a (share-)exclusive
+ *		lock on the buffer.
  */
 XLogRecPtr
 BufferGetLSNAtomic(Buffer buffer)
@@ -5662,6 +5657,12 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate, b
 			 * It's possible we may enter here without an xid, so it is
 			 * essential that CreateCheckPoint waits for virtual transactions
 			 * rather than full transactionids.
+			 *
+			 * FIXME: I think we now should simply mark the page dirty before
+			 * WAL logging the hint bit - afaikt it then should work just like
+			 * any other buffer write (due to SyncBuffers()/SyncOneBuffer()
+			 * seeing the dirty bit and trying to lock the page
+			 * share-exclusive, and thus having to wait).
 			 */
 			Assert((MyProc->delayChkptFlags & DELAY_CHKPT_START) == 0);
 			MyProc->delayChkptFlags |= DELAY_CHKPT_START;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index a41a5facd3a..5826d4b54c6 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -199,7 +199,7 @@ FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln)
 		reln = smgropen(BufTagGetRelFileLocator(&bufHdr->tag),
 						MyProcNumber);
 
-	PageSetChecksumInplace(localpage, bufHdr->tag.blockNum);
+	PageSetChecksum(localpage, bufHdr->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index aac6e695954..c8cbdd1f7a6 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -1494,51 +1494,15 @@ PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 /*
  * Set checksum for a page in shared buffers.
  *
- * If checksums are disabled, or if the page is not initialized, just return
- * the input.  Otherwise, we must make a copy of the page before calculating
- * the checksum, to prevent concurrent modifications (e.g. setting hint bits)
- * from making the final checksum invalid.  It doesn't matter if we include or
- * exclude hints during the copy, as long as we write a valid page and
- * associated checksum.
+ * If checksums are disabled, or if the page is not initialized, just
+ * return. Otherwise compute and set the checksum.
  *
- * Returns a pointer to the block-sized data that needs to be written. Uses
- * statically-allocated memory, so the caller must immediately write the
- * returned page and not refer to it again.
- */
-char *
-PageSetChecksumCopy(Page page, BlockNumber blkno)
-{
-	static char *pageCopy = NULL;
-
-	/* If we don't need a checksum, just return the passed-in data */
-	if (PageIsNew(page) || !DataChecksumsEnabled())
-		return page;
-
-	/*
-	 * We allocate the copy space once and use it over on each subsequent
-	 * call.  The point of palloc'ing here, rather than having a static char
-	 * array, is first to ensure adequate alignment for the checksumming code
-	 * and second to avoid wasting space in processes that never call this.
-	 */
-	if (pageCopy == NULL)
-		pageCopy = MemoryContextAllocAligned(TopMemoryContext,
-											 BLCKSZ,
-											 PG_IO_ALIGN_SIZE,
-											 0);
-
-	memcpy(pageCopy, page, BLCKSZ);
-	((PageHeader) pageCopy)->pd_checksum = pg_checksum_page(pageCopy, blkno);
-	return pageCopy;
-}
-
-/*
- * Set checksum for a page in private memory.
- *
- * This must only be used when we know that no other process can be modifying
- * the page buffer.
+ * In the past this needed to be done on a copy of the page, due to the
+ * possibility of e.g. hint bits being set concurrently. However, this is not
+ * necessary anymore as hint bits won't be set while IO is going on.
  */
 void
-PageSetChecksumInplace(Page page, BlockNumber blkno)
+PageSetChecksum(Page page, BlockNumber blkno)
 {
 	/* If we don't need a checksum, just return */
 	if (PageIsNew(page) || !DataChecksumsEnabled())
diff --git a/src/backend/storage/smgr/bulk_write.c b/src/backend/storage/smgr/bulk_write.c
index b958be15716..f4d07543365 100644
--- a/src/backend/storage/smgr/bulk_write.c
+++ b/src/backend/storage/smgr/bulk_write.c
@@ -279,7 +279,7 @@ smgr_bulk_flush(BulkWriteState *bulkstate)
 		BlockNumber blkno = pending_writes[i].blkno;
 		Page		page = pending_writes[i].buf->data;
 
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 
 		if (blkno >= bulkstate->relsize)
 		{
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index 488d98e7e66..e5fc7642dc2 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -288,7 +288,7 @@ modify_rel_block(PG_FUNCTION_ARGS)
 	}
 	else
 	{
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 	}
 
 	smgrwrite(RelationGetSmgr(rel),
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-12-03 01:12                       ` Peter Geoghegan <pg@bowt.ie>
  2025-12-03 01:18                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2 siblings, 1 reply; 120+ messages in thread

From: Peter Geoghegan @ 2025-12-03 01:12 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Tue, Dec 2, 2025 at 7:47 PM Andres Freund <andres@anarazel.de> wrote:
> On 2025-11-25 11:54:00 -0500, Andres Freund wrote:
> > Thanks a lot for that detailed review!  A few questions and comments, before I
> > try to address the comments in the next version.
>
> Here's that new new version, with the following changes

_bt_check_unique will hold an exclusive buffer lock on the page being
LP_DEAD-set in the vast majority of cases. Should we expect your
changes to have no effect at all in that common case?

The BTP_HAS_GARBAGE flag is deprecated these days; we basically don't
use it anymore. How much value might there be in avoiding setting
BTP_HAS_GARBAGE as a way of being able to use BufferSetHintBits16 more
often in nbtree?

-- 
Peter Geoghegan





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 01:12                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
@ 2025-12-03 01:18                         ` Andres Freund <andres@anarazel.de>
  0 siblings, 0 replies; 120+ messages in thread

From: Andres Freund @ 2025-12-03 01:18 UTC (permalink / raw)
  To: Peter Geoghegan <pg@bowt.ie>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-12-02 20:12:14 -0500, Peter Geoghegan wrote:
> On Tue, Dec 2, 2025 at 7:47 PM Andres Freund <andres@anarazel.de> wrote:
> > On 2025-11-25 11:54:00 -0500, Andres Freund wrote:
> > > Thanks a lot for that detailed review!  A few questions and comments, before I
> > > try to address the comments in the next version.
> >
> > Here's that new new version, with the following changes
> 
> _bt_check_unique will hold an exclusive buffer lock on the page being
> LP_DEAD-set in the vast majority of cases. Should we expect your
> changes to have no effect at all in that common case?

If we already have an exclusive lock, BufferBeginSetHintBits() will quickly
return true and won't ever return false.


> The BTP_HAS_GARBAGE flag is deprecated these days; we basically don't
> use it anymore. How much value might there be in avoiding setting
> BTP_HAS_GARBAGE as a way of being able to use BufferSetHintBits16 more
> often in nbtree?

None of the MarkBufferDirtyHint() cases in nbtree that had to be modified
looked like they would benefit from BufferSetHintBits16(), since they will
typically modify the page multiple times.  But maybe I'm just misunderstanding
what you mean?

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-12-03 16:03                       ` Andres Freund <andres@anarazel.de>
  2 siblings, 0 replies; 120+ messages in thread

From: Andres Freund @ 2025-12-03 16:03 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-12-02 19:47:35 -0500, Andres Freund wrote:
> On 2025-11-25 11:54:00 -0500, Andres Freund wrote:
> > Thanks a lot for that detailed review!  A few questions and comments, before I
> > try to address the comments in the next version.
> 
> Here's that new new version, with the following changes
> 
> - Some more micro-optimizations, most importantly adding a commit that doesn't
>   initialize the delay in LockBufHdr() unless needed. With those I don't see a
>   consistent slowdown anymore (slight speedup on one workstation, slight
>   slowdown on another, in an absurdly adverse workload)
> 
> - Tried to address Melanie's feedback, with some exceptions (some noted below,
>   but I also need to make another pass through the reviews)
> 
> - re-implemented AssertNotCatalogBufferLock() in the new world
> 
> - Substantially expanded comments around setting hint bits (in buffer/README,
>   heapam_visibility.c and bufmgr.c)
> 
> - split out the change to fsm_vacuum_page() to start to lock the page into is
>   own commit
> 
> - reordered patch series so that smaller changes are before the 64bit-state
>   and "Implement buffer content locks independently of" commits, so they can
>   be committed while we finish cleaning the later changes
> 
> - I didn't invest much in cleaning up the later patches ("Don't copy pages
>   while writing out" and "Make UnlockReleaseBuffer() more efficient") yet,
>   wanted to focus on the earlier patches first
> 
> 
> Todo:
> 
> - still need to rename ResOwnerReleaseBufferPin(). Wondering about what to
>   rename ResourceOwnerDesc.name to. "buffer ownership" maybe? Not great...
> 
> - gistkillitems() complaint by Melanie
> 
> - amortize vs batch vs SetHintBits comment + SHB_* names
> 
> - for the next version I'll remove the BATCHMVCC_FEWER_ARGS conditionals from
>   0010. I don't love needing BatchMVCCState but I don't really see an
>   alternative, the performance difference is pretty persistent.
> 
> 
> Questions:
> - ForEachLWLockHeldByMe() and LWLockDisown() aren't used anymore, should we
>   remove them?

I'm planning to work on committing 0001, 0002, 0003, 0008 soon-ish, unless
somebody sees a reason to hold off on that.  After that I think 0005, 0006
would be next.  I think 0004 is a clear improvement, but nobody has looked at
it yet...

For 0007, I wished ConditionalLockBuffer() accepted the lock level, there's no
point in waiting for the lock in the use case. I'm on the fence about whether
it's worth changing the ~12 users of ConditionalLockBuffer()...

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-12-17 09:25                       ` Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2 siblings, 1 reply; 120+ messages in thread

From: Heikki Linnakangas @ 2025-12-17 09:25 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; Melanie Plageman <melanieplageman@gmail.com>; +Cc: Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On 03/12/2025 02:47, Andres Freund wrote:
> On 2025-11-25 11:54:00 -0500, Andres Freund wrote:
>> Thanks a lot for that detailed review!  A few questions and comments, before I
>> try to address the comments in the next version.
> 
> Here's that new new version, with the following changes
> 
> - Some more micro-optimizations, most importantly adding a commit that doesn't
>    initialize the delay in LockBufHdr() unless needed. With those I don't see a
>    consistent slowdown anymore (slight speedup on one workstation, slight
>    slowdown on another, in an absurdly adverse workload)

+1

I'm comparing the patched LockBufHdr() with LWLockWaitListLock(), which 
does pretty much the same thing, and LWLockWaitListLock() already did 
the initialization of the delay that way. But there are some small 
differences:

- LockBufHdr() uses unlikely() in the initial attempt, 
LWLockWaitListLock() does not
- LWLockWaitListLock() uses pg_atomic_read_u32() after spinning, 
LockBufHdr() retries directly with pg_atomic_fetch_or_u32().

Are there reasons for the differences, or is it just that they were 
developed separately and ended up looking slightly different?

- Heikki






^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
@ 2025-12-17 14:54                         ` Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-12-17 14:54 UTC (permalink / raw)
  To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-12-17 11:25:50 +0200, Heikki Linnakangas wrote:
> On 03/12/2025 02:47, Andres Freund wrote:
> > On 2025-11-25 11:54:00 -0500, Andres Freund wrote:
> > > Thanks a lot for that detailed review!  A few questions and comments, before I
> > > try to address the comments in the next version.
> > 
> > Here's that new new version, with the following changes
> > 
> > - Some more micro-optimizations, most importantly adding a commit that doesn't
> >    initialize the delay in LockBufHdr() unless needed. With those I don't see a
> >    consistent slowdown anymore (slight speedup on one workstation, slight
> >    slowdown on another, in an absurdly adverse workload)
> 
> +1
> 
> I'm comparing the patched LockBufHdr() with LWLockWaitListLock(), which does
> pretty much the same thing, and LWLockWaitListLock() already did the
> initialization of the delay that way. But there are some small differences:
> 
> - LockBufHdr() uses unlikely() in the initial attempt, LWLockWaitListLock()
> does not

I think we probably ought to do that in LWLockWaitListLock() too.


> - LWLockWaitListLock() uses pg_atomic_read_u32() after spinning,
> LockBufHdr() retries directly with pg_atomic_fetch_or_u32().

I think here LWLockWaitListLock() is likely right - but it seems like a change
to LockBufHdr() that I would probably make in a separate commit?


> Are there reasons for the differences, or is it just that they were
> developed separately and ended up looking slightly different?

I think it's just the latter...


Thanks for reviewing,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-12-18 17:03                           ` Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-12-18 17:03 UTC (permalink / raw)
  To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-12-17 09:54:32 -0500, Andres Freund wrote:
> On 2025-12-17 11:25:50 +0200, Heikki Linnakangas wrote:
> > - LWLockWaitListLock() uses pg_atomic_read_u32() after spinning,
> > LockBufHdr() retries directly with pg_atomic_fetch_or_u32().
> 
> I think here LWLockWaitListLock() is likely right - but it seems like a change
> to LockBufHdr() that I would probably make in a separate commit?

FWIW, I couldn't come up with a scenario where it makes a performance
difference - exclusive content locks just aren't *that* frequent. And because
of that the wait list lock doesn't have similar contention as some non-content
lwlocks (like XidGenLock). The most extreme workload I could think of was
pgbench hammering a single sequence across many sessions. While the exclusive
locks show up in wait events, the buffer header spinlock itself doesn't..

So I'm inclined to not change anything about this for now.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-12-18 17:20                             ` Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Heikki Linnakangas @ 2025-12-18 17:20 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On 18/12/2025 19:03, Andres Freund wrote:
> Hi,
> 
> On 2025-12-17 09:54:32 -0500, Andres Freund wrote:
>> On 2025-12-17 11:25:50 +0200, Heikki Linnakangas wrote:
>>> - LWLockWaitListLock() uses pg_atomic_read_u32() after spinning,
>>> LockBufHdr() retries directly with pg_atomic_fetch_or_u32().
>>
>> I think here LWLockWaitListLock() is likely right - but it seems like a change
>> to LockBufHdr() that I would probably make in a separate commit?
> 
> FWIW, I couldn't come up with a scenario where it makes a performance
> difference - exclusive content locks just aren't *that* frequent. And because
> of that the wait list lock doesn't have similar contention as some non-content
> lwlocks (like XidGenLock). The most extreme workload I could think of was
> pgbench hammering a single sequence across many sessions. While the exclusive
> locks show up in wait events, the buffer header spinlock itself doesn't..
> 
> So I'm inclined to not change anything about this for now.

Ok. My thinking was just that LockBufHdr() and LWLockWaitListLock() 
should be consistent with each other. Otherwise anyone reading the code 
will ask the question "why are they different?". They're the only two 
things using the spin delay mechanism in our codebase, in addition to 
actual spinlocks.

BTW, I wonder if it would be worthwhile to have an inlineable fast-path 
of LockBufHdr() for the common case that the lock is free? I see that 
UnlockBufHdr() is already a static inline function.

- Heikki






^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
@ 2025-12-18 22:06                               ` Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-12-18 22:06 UTC (permalink / raw)
  To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-12-18 19:20:38 +0200, Heikki Linnakangas wrote:
> On 18/12/2025 19:03, Andres Freund wrote:
> > On 2025-12-17 09:54:32 -0500, Andres Freund wrote:
> > > On 2025-12-17 11:25:50 +0200, Heikki Linnakangas wrote:
> > > > - LWLockWaitListLock() uses pg_atomic_read_u32() after spinning,
> > > > LockBufHdr() retries directly with pg_atomic_fetch_or_u32().
> > >
> > > I think here LWLockWaitListLock() is likely right - but it seems like a change
> > > to LockBufHdr() that I would probably make in a separate commit?
> >
> > FWIW, I couldn't come up with a scenario where it makes a performance
> > difference - exclusive content locks just aren't *that* frequent. And because
> > of that the wait list lock doesn't have similar contention as some non-content
> > lwlocks (like XidGenLock). The most extreme workload I could think of was
> > pgbench hammering a single sequence across many sessions. While the exclusive
> > locks show up in wait events, the buffer header spinlock itself doesn't..
> >
> > So I'm inclined to not change anything about this for now.
>
> Ok. My thinking was just that LockBufHdr() and LWLockWaitListLock() should
> be consistent with each other. Otherwise anyone reading the code will ask
> the question "why are they different?". They're the only two things using
> the spin delay mechanism in our codebase, in addition to actual spinlocks.

I guess for me it didn't really seem like this patch's job to fix
that. Regardless of that, here's a version that tries to make them more
similar.

I did check, adding a likely() to LWLockWaitListLock()'s break does improve
code generation (verified by looking at the generated code) and seems to
improve performance in some very extreme workloads (e.g. [1]) a bit.

I'll try to come up with a combined patch that applies the optimizations in
LWLockWaitListLock() and LockBufHdr() to each other.


> BTW, I wonder if it would be worthwhile to have an inlineable fast-path of
> LockBufHdr() for the common case that the lock is free? I see that
> UnlockBufHdr() is already a static inline function.

I tried that a while ago and couldn't see any improvement, I think because all
the performance relevant callers are in bufmgr.c and thus can already inline
[parts of] the implementation.  I guess you could make the generated code a
bit smaller if you use pg_noinline on the slowpath, but that seems like a
separate project / effort to me.

Greetings,

Andres Freund


[1] Many connections doing
DO $do$
    BEGIN
        FOR i IN 1 .. 1000 LOOP
            PERFORM txid_current();
	    COMMIT;
	END LOOP;
     END;
$do$;





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2025-12-18 23:39                                 ` Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2025-12-18 23:39 UTC (permalink / raw)
  To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2025-12-18 17:06:43 -0500, Andres Freund wrote:
> I'll try to come up with a combined patch that applies the optimizations in
> LWLockWaitListLock() and LockBufHdr() to each other.

Attached is a rebased version of this patch series, with the above patch as
0001.

The later patches have a lot, but not yet all, of Melanie's feedback
addressed. I mainly am sending them because cfbot wanted a rebase anyway.

I'm mainly not yet happy with how 0007 changes buf_internals.h, some more
comment work is also needed. 0009 & 0010 still are more for POC than ready to
be in-depth reviewed.

Greetings,

Andres Freund

Attachments:

  [text/x-diff] v8-0001-bufmgr-Optimize-harmonize-LockBufHdr-LWLockWaitLi.patch (3.7K, ../../ossv2eistssmubfsir6xjll76tynvxv5lup4zkrfzjkud7fycw@rf5vii6l6cha/2-v8-0001-bufmgr-Optimize-harmonize-LockBufHdr-LWLockWaitLi.patch)
  download | inline diff:
From 74bbf209e18ce7ca9b135b94abe0077d0e0f8dfa Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 2 Dec 2025 18:44:54 -0500
Subject: [PATCH v8 01/10] bufmgr: Optimize & harmonize LockBufHdr(),
 LWLockWaitListLock()

The main optimization is for LockBufHdr() to delay initializing
SpinDelayStatus, similar to what LWLockWaitListLock already did. The
initialization is sufficiently expensive & buffer header lock acquisitions are
sufficiently frequent, to make it worthwhile to instead have a fastpath (via a
likely() branch) that does not initialize the SpinDelayStatus.

While LWLockWaitListLock() already the aforementioned optimization, it did not
use likely(), and inspection of the assembly shows that this indeed leads to
worse code generation (also observed in a microbenchmark). Fix that by adding
the likely().

While the LockBufHdr() improvement is a small gain on its own, it mainly is
aimed at preventing a regression after a future commit, which requires
additional locking to set hint bits.

While touching both, also make the comments more similar to each other.

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/backend/storage/buffer/bufmgr.c | 36 +++++++++++++++++++++--------
 src/backend/storage/lmgr/lwlock.c   |  8 +++++--
 2 files changed, 33 insertions(+), 11 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index a768fb129ae..eb55102b0d7 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -6358,23 +6358,41 @@ rlocator_comparator(const void *p1, const void *p2)
 uint32
 LockBufHdr(BufferDesc *desc)
 {
-	SpinDelayStatus delayStatus;
 	uint32		old_buf_state;
 
 	Assert(!BufferIsLocal(BufferDescriptorGetBuffer(desc)));
 
-	init_local_spin_delay(&delayStatus);
-
 	while (true)
 	{
-		/* set BM_LOCKED flag */
+		/*
+		 * Always try once to acquire the lock directly, without setting up
+		 * the spin-delay infrastructure. The work necessary for that shows up
+		 * in profiles and is rarely necessary.
+		 */
 		old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
-		/* if it wasn't set before we're OK */
-		if (!(old_buf_state & BM_LOCKED))
-			break;
-		perform_spin_delay(&delayStatus);
+		if (likely(!(old_buf_state & BM_LOCKED)))
+			break;				/* got lock */
+
+		/* and then spin without atomic operations until lock is released */
+		{
+			SpinDelayStatus delayStatus;
+
+			init_local_spin_delay(&delayStatus);
+
+			while (old_buf_state & BM_LOCKED)
+			{
+				perform_spin_delay(&delayStatus);
+				old_buf_state = pg_atomic_read_u32(&desc->state);
+			}
+			finish_spin_delay(&delayStatus);
+		}
+
+		/*
+		 * Retry. The lock might obviously already be re-acquired by the time
+		 * we're attempting to get it again.
+		 */
 	}
-	finish_spin_delay(&delayStatus);
+
 	return old_buf_state | BM_LOCKED;
 }
 
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 255cfa8fa95..b839ace57cb 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -870,9 +870,13 @@ LWLockWaitListLock(LWLock *lock)
 
 	while (true)
 	{
-		/* always try once to acquire lock directly */
+		/*
+		 * Always try once to acquire the lock directly, without setting up
+		 * the spin-delay infrastructure. The work necessary for that shows up
+		 * in profiles and is rarely necessary.
+		 */
 		old_state = pg_atomic_fetch_or_u32(&lock->state, LW_FLAG_LOCKED);
-		if (!(old_state & LW_FLAG_LOCKED))
+		if (likely(!(old_state & LW_FLAG_LOCKED)))
 			break;				/* got lock */
 
 		/* and then spin without atomic operations until lock is released */
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v8-0002-heapam-Move-logic-to-handle-HEAP_MOVED-into-a-hel.patch (11.1K, ../../ossv2eistssmubfsir6xjll76tynvxv5lup4zkrfzjkud7fycw@rf5vii6l6cha/3-v8-0002-heapam-Move-logic-to-handle-HEAP_MOVED-into-a-hel.patch)
  download | inline diff:
From d9143031e6b2d3ed1c2611b2aa0fbf0a6b9a4bff Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 23 Sep 2024 12:23:33 -0400
Subject: [PATCH v8 02/10] heapam: Move logic to handle HEAP_MOVED into a
 helper function

Before we dealt with this in 6 near identical and one very similar copy.

The helper function errors out when encountering a
HEAP_MOVED_IN/HEAP_MOVED_OUT tuple with xvac considered current or
in-progress. It'd be preferrable to do that change separately, but otherwise
it'd not be possible to deduplicate the handling in
HeapTupleSatisfiesVacuum().

Reviewed-by: Heikki Linnakangas <hlinnaka@iki.fi>
Discussion: https://postgr.es/m/lxzj26ga6ippdeunz6kuncectr5gfuugmm2ry22qu6hcx6oid6@lzx3sjsqhmt6
Discussion: https://postgr.es/m/6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar
---
 src/backend/access/heap/heapam_visibility.c | 293 ++++----------------
 1 file changed, 61 insertions(+), 232 deletions(-)

diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 05f6946fe60..bf899c2d2c6 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -144,6 +144,55 @@ HeapTupleSetHintBits(HeapTupleHeader tuple, Buffer buffer,
 	SetHintBits(tuple, buffer, infomask, xid);
 }
 
+/*
+ * If HEAP_MOVED_OFF or HEAP_MOVED_IN are set on the tuple, remove them and
+ * adjust hint bits. See the comment for SetHintBits() for more background.
+ *
+ * This helper returns false if the row ought to be invisible, true otherwise.
+ */
+static inline bool
+HeapTupleCleanMoved(HeapTupleHeader tuple, Buffer buffer)
+{
+	TransactionId xvac;
+
+	/* only used by pre-9.0 binary upgrades */
+	if (likely(!(tuple->t_infomask & (HEAP_MOVED_OFF | HEAP_MOVED_IN))))
+		return true;
+
+	xvac = HeapTupleHeaderGetXvac(tuple);
+
+	if (TransactionIdIsCurrentTransactionId(xvac))
+		elog(ERROR, "encountered tuple with HEAP_MOVED considered current");
+
+	if (TransactionIdIsInProgress(xvac))
+		elog(ERROR, "encountered tuple with HEAP_MOVED considered in-progress");
+
+	if (tuple->t_infomask & HEAP_MOVED_OFF)
+	{
+		if (TransactionIdDidCommit(xvac))
+		{
+			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
+						InvalidTransactionId);
+			return false;
+		}
+		SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
+					InvalidTransactionId);
+	}
+	else if (tuple->t_infomask & HEAP_MOVED_IN)
+	{
+		if (TransactionIdDidCommit(xvac))
+			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
+						InvalidTransactionId);
+		else
+		{
+			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
+						InvalidTransactionId);
+			return false;
+		}
+	}
+
+	return true;
+}
 
 /*
  * HeapTupleSatisfiesSelf
@@ -179,45 +228,8 @@ HeapTupleSatisfiesSelf(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (tuple->t_infomask & HEAP_XMAX_INVALID)	/* xid invalid */
@@ -372,45 +384,8 @@ HeapTupleSatisfiesToast(HeapTuple htup, Snapshot snapshot,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 
 		/*
 		 * An invalid Xmin can be left behind by a speculative insertion that
@@ -468,45 +443,8 @@ HeapTupleSatisfiesUpdate(HeapTuple htup, CommandId curcid,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return TM_Invisible;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return TM_Invisible;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return TM_Invisible;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return TM_Invisible;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return TM_Invisible;
-				}
-			}
-		}
+		else if (!HeapTupleCleanMoved(tuple, buffer))
+			return TM_Invisible;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (HeapTupleHeaderGetCmin(tuple) >= curcid)
@@ -756,45 +694,8 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!TransactionIdIsInProgress(xvac))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (TransactionIdIsInProgress(xvac))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (tuple->t_infomask & HEAP_XMAX_INVALID)	/* xid invalid */
@@ -979,45 +880,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return false;
 
-		/* Used by pre-9.0 binary upgrades */
-		if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return false;
-			if (!XidInMVCCSnapshot(xvac, snapshot))
-			{
-				if (TransactionIdDidCommit(xvac))
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			}
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (!TransactionIdIsCurrentTransactionId(xvac))
-			{
-				if (XidInMVCCSnapshot(xvac, snapshot))
-					return false;
-				if (TransactionIdDidCommit(xvac))
-					SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-								InvalidTransactionId);
-				else
-				{
-					SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-								InvalidTransactionId);
-					return false;
-				}
-			}
-		}
+		if (!HeapTupleCleanMoved(tuple, buffer))
+			return false;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (HeapTupleHeaderGetCmin(tuple) >= snapshot->curcid)
@@ -1222,43 +1086,8 @@ HeapTupleSatisfiesVacuumHorizon(HeapTuple htup, Buffer buffer, TransactionId *de
 	{
 		if (HeapTupleHeaderXminInvalid(tuple))
 			return HEAPTUPLE_DEAD;
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_OFF)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return HEAPTUPLE_DELETE_IN_PROGRESS;
-			if (TransactionIdIsInProgress(xvac))
-				return HEAPTUPLE_DELETE_IN_PROGRESS;
-			if (TransactionIdDidCommit(xvac))
-			{
-				SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-							InvalidTransactionId);
-				return HEAPTUPLE_DEAD;
-			}
-			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-						InvalidTransactionId);
-		}
-		/* Used by pre-9.0 binary upgrades */
-		else if (tuple->t_infomask & HEAP_MOVED_IN)
-		{
-			TransactionId xvac = HeapTupleHeaderGetXvac(tuple);
-
-			if (TransactionIdIsCurrentTransactionId(xvac))
-				return HEAPTUPLE_INSERT_IN_PROGRESS;
-			if (TransactionIdIsInProgress(xvac))
-				return HEAPTUPLE_INSERT_IN_PROGRESS;
-			if (TransactionIdDidCommit(xvac))
-				SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-							InvalidTransactionId);
-			else
-			{
-				SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-							InvalidTransactionId);
-				return HEAPTUPLE_DEAD;
-			}
-		}
+		else if (!HeapTupleCleanMoved(tuple, buffer))
+			return HEAPTUPLE_DEAD;
 		else if (TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmin(tuple)))
 		{
 			if (tuple->t_infomask & HEAP_XMAX_INVALID)	/* xid invalid */
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v8-0003-freespace-Don-t-modify-page-without-any-lock.patch (2.0K, ../../ossv2eistssmubfsir6xjll76tynvxv5lup4zkrfzjkud7fycw@rf5vii6l6cha/4-v8-0003-freespace-Don-t-modify-page-without-any-lock.patch)
  download | inline diff:
From d7a10d2b8390e0cf835ba67786ea11e014aaf105 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 1 Dec 2025 22:31:44 -0500
Subject: [PATCH v8 03/10] freespace: Don't modify page without any lock

Before this commit fsm_vacuum_page() modified the page without any lock on the
page. Historically that was kind of ok, as we didn't rely on the freespace to
really stay consistent and we did not have checksums. But these days pages are
checksummed and there are ways for FSM pages to be included in WAL records,
even if the FSM itself is still not WAL logged. If a FSM page ever were
modified while a WAL record referenced that page, we'd be in trouble, as the
WAL CRC could end up getting corrupted.

The reason to address this right now is a series of patches with the goal to
only allow modifications of pages with an appropriate lock level. Obviously
not having any lock is not appropriate :)

Discussion: https://postgr.es/m/4wggb7purufpto6x35fd2kwhasehnzfdy3zdcu47qryubs2hdz@fa5kannykekr
Discussion: https://postgr.es/m/e6a8f734-2198-4958-a028-aba863d4a204@iki.fi
---
 src/backend/storage/freespace/freespace.c | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index 4773a9cc65e..48ac15d3487 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -906,10 +906,12 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 	/*
 	 * Reset the next slot pointer. This encourages the use of low-numbered
 	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation.  We don't bother with a lock here, nor with marking the page
-	 * dirty if it wasn't already, since this is just a hint.
+	 * relation. We don't bother with marking the page dirty if it wasn't
+	 * already, since this is just a hint.
 	 */
+	LockBuffer(buf, BUFFER_LOCK_SHARE);
 	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 	ReleaseBuffer(buf);
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v8-0004-heapam-Use-exclusive-lock-on-old-page-in-CLUSTER.patch (2.7K, ../../ossv2eistssmubfsir6xjll76tynvxv5lup4zkrfzjkud7fycw@rf5vii6l6cha/5-v8-0004-heapam-Use-exclusive-lock-on-old-page-in-CLUSTER.patch)
  download | inline diff:
From a7095841a82576e6984b31ffd04ae851fec91993 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sun, 26 Jan 2025 15:18:46 -0500
Subject: [PATCH v8 04/10] heapam: Use exclusive lock on old page in CLUSTER

To be able to guarantee that we can set the hint bit, acquire an exclusive
lock on the old buffer. We need the hint bits to be set as otherwise
reform_and_rewrite_tuple() -> rewrite_heap_tuple() -> heap_freeze_tuple() will
get confused.

It'd be better if we somehow could avoid setting hint bits on the old page. A
commonreason to use VACUUM FULL are very bloated tables - rewriting most of
the old table before during VACUUM FULL doesn't exactly help.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/heap/heapam_handler.c    | 13 ++++++++++++-
 src/backend/access/heap/heapam_visibility.c |  7 +++++++
 2 files changed, 19 insertions(+), 1 deletion(-)

diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index dd4fe6bf62f..0c684786382 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -837,7 +837,18 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		tuple = ExecFetchSlotHeapTuple(slot, false, NULL);
 		buf = hslot->buffer;
 
-		LockBuffer(buf, BUFFER_LOCK_SHARE);
+		/*
+		 * To be able to guarantee that we can set the hint bit, acquire an
+		 * exclusive lock on the old buffer. We need the hint bits to be set
+		 * as otherwise reform_and_rewrite_tuple() -> rewrite_heap_tuple() ->
+		 * heap_freeze_tuple() will get confused.
+		 *
+		 * It'd be better if we somehow could avoid setting hint bits on the
+		 * old page. One reason to use VACUUM FULL are very bloated tables -
+		 * rewriting most of the old table before during VACUUM FULL doesn't
+		 * exactly help...
+		 */
+		LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
 		switch (HeapTupleSatisfiesVacuum(tuple, OldestXmin, buf))
 		{
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index bf899c2d2c6..debf5d56b95 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -141,6 +141,13 @@ void
 HeapTupleSetHintBits(HeapTupleHeader tuple, Buffer buffer,
 					 uint16 infomask, TransactionId xid)
 {
+	/*
+	 * The uses from heapam.c rely on being able to perform the hint bit
+	 * updates, which can only be guaranteed if we are holding an exclusive
+	 * lock on the buffer - which all callers are doing.
+	 */
+	Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+
 	SetHintBits(tuple, buffer, infomask, xid);
 }
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v8-0005-heapam-Add-batch-mode-mvcc-check-and-use-it-in-pa.patch (7.5K, ../../ossv2eistssmubfsir6xjll76tynvxv5lup4zkrfzjkud7fycw@rf5vii6l6cha/6-v8-0005-heapam-Add-batch-mode-mvcc-check-and-use-it-in-pa.patch)
  download | inline diff:
From f17b3654daf9219fd9ed3229e9ba82330345e0e2 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 17 Oct 2024 13:16:36 -0400
Subject: [PATCH v8 05/10] heapam: Add batch mode mvcc check and use it in page
 mode

There are two reasons for doing so:

1) It is generally faster to perform checks in a batched fashion and making
   sequential scans faster is nice.

2) We would like to stop setting hint bits while pages are being written
   out. The necessary locking becomes visible for page mode scans if done for
   every tuple. With batching the overhead can be amortized to only happen
   once per page.

There are substantial further optimization opportunities along these
lines:

- Right now HeapTupleSatisfiesMVCCBatch() simply uses the single-tuple
  HeapTupleSatisfiesMVCC(), relying on the compiler to inline it. We could
  instead write an explicitly optimized version that avoids repeated xid
  tests.

- Introduce batched version of the serializability test

- Introduce batched version of HeapTupleSatisfiesVacuum

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar
---
 src/include/access/heapam.h                 | 17 +++++
 src/backend/access/heap/heapam.c            | 84 ++++++++++++++++-----
 src/backend/access/heap/heapam_visibility.c | 42 +++++++++++
 src/tools/pgindent/typedefs.list            |  1 +
 4 files changed, 124 insertions(+), 20 deletions(-)

diff --git a/src/include/access/heapam.h b/src/include/access/heapam.h
index f7e4ae3843c..cf4cd3e4dbd 100644
--- a/src/include/access/heapam.h
+++ b/src/include/access/heapam.h
@@ -449,6 +449,23 @@ extern bool HeapTupleHeaderIsOnlyLocked(HeapTupleHeader tuple);
 extern bool HeapTupleIsSurelyDead(HeapTuple htup,
 								  GlobalVisState *vistest);
 
+/*
+ * The output of HeapTupleSatisfiesMVCCBatch() is passed via this struct, as
+ * otherwise the increased number of arguments to
+ * HeapTupleSatisfiesMVCCBatch() leads to on-stack argument passing on x86-64,
+ * which causes a small regression.
+ */
+typedef struct BatchMVCCState
+{
+	HeapTupleData tuples[MaxHeapTuplesPerPage];
+	bool		visible[MaxHeapTuplesPerPage];
+} BatchMVCCState;
+
+extern int	HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
+										int ntups,
+										BatchMVCCState *batchmvcc,
+										OffsetNumber *vistuples_dense);
+
 /*
  * To avoid leaking too much knowledge about reorderbuffer implementation
  * details this is implemented in reorderbuffer.c not heapam_visibility.c
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 6daf4a87dec..513a9b275a2 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -519,42 +519,86 @@ page_collect_tuples(HeapScanDesc scan, Snapshot snapshot,
 					BlockNumber block, int lines,
 					bool all_visible, bool check_serializable)
 {
+	Oid			relid = RelationGetRelid(scan->rs_base.rs_rd);
 	int			ntup = 0;
-	OffsetNumber lineoff;
+	int			nvis = 0;
+	BatchMVCCState batchmvcc;
 
-	for (lineoff = FirstOffsetNumber; lineoff <= lines; lineoff++)
+	/* page at a time should have been disabled otherwise */
+	Assert(IsMVCCSnapshot(snapshot));
+
+	/* first find all tuples on the page */
+	for (OffsetNumber lineoff = FirstOffsetNumber; lineoff <= lines; lineoff++)
 	{
 		ItemId		lpp = PageGetItemId(page, lineoff);
-		HeapTupleData loctup;
-		bool		valid;
+		HeapTuple	tup;
 
-		if (!ItemIdIsNormal(lpp))
+		if (unlikely(!ItemIdIsNormal(lpp)))
 			continue;
 
-		loctup.t_data = (HeapTupleHeader) PageGetItem(page, lpp);
-		loctup.t_len = ItemIdGetLength(lpp);
-		loctup.t_tableOid = RelationGetRelid(scan->rs_base.rs_rd);
-		ItemPointerSet(&(loctup.t_self), block, lineoff);
+		/*
+		 * If the page is not all-visible or we need to check serializability,
+		 * maintain enough state to be able to refind the tuple efficiently,
+		 * without again first needing to fetch the item and then via that the
+		 * tuple.
+		 */
+		if (!all_visible || check_serializable)
+		{
+			tup = &batchmvcc.tuples[ntup];
 
+			tup->t_data = (HeapTupleHeader) PageGetItem(page, lpp);
+			tup->t_len = ItemIdGetLength(lpp);
+			tup->t_tableOid = relid;
+			ItemPointerSet(&(tup->t_self), block, lineoff);
+		}
+
+		/*
+		 * If the page is all visible, these fields otherwise won't be
+		 * populated in loop below.
+		 */
 		if (all_visible)
-			valid = true;
-		else
-			valid = HeapTupleSatisfiesVisibility(&loctup, snapshot, buffer);
-
-		if (check_serializable)
-			HeapCheckForSerializableConflictOut(valid, scan->rs_base.rs_rd,
-												&loctup, buffer, snapshot);
-
-		if (valid)
 		{
+			if (check_serializable)
+			{
+				batchmvcc.visible[ntup] = true;
+			}
 			scan->rs_vistuples[ntup] = lineoff;
-			ntup++;
 		}
+
+		ntup++;
 	}
 
 	Assert(ntup <= MaxHeapTuplesPerPage);
 
-	return ntup;
+	/*
+	 * Unless the page is all visible, test visibility for all tuples one
+	 * go. That is considerably more efficient than calling
+	 * HeapTupleSatisfiesMVCC() one-by-one.
+	 */
+	if (all_visible)
+		nvis = ntup;
+	else
+		nvis = HeapTupleSatisfiesMVCCBatch(snapshot, buffer,
+										   ntup,
+										   &batchmvcc,
+										   scan->rs_vistuples);
+
+	/*
+	 * So far we don't have batch API for testing serializabilty, so do so
+	 * one-by-one.
+	 */
+	if (check_serializable)
+	{
+		for (int i = 0; i < ntup; i++)
+		{
+			HeapCheckForSerializableConflictOut(batchmvcc.visible[i],
+												scan->rs_base.rs_rd,
+												&batchmvcc.tuples[i],
+												buffer, snapshot);
+		}
+	}
+
+	return nvis;
 }
 
 /*
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index debf5d56b95..04284c4b2eb 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -1598,6 +1598,48 @@ HeapTupleSatisfiesHistoricMVCC(HeapTuple htup, Snapshot snapshot,
 		return true;
 }
 
+/*
+ * Perform HeaptupleSatisfiesMVCC() on each passed in tuple. This is more
+ * efficient than doing HeapTupleSatisfiesMVCC() one-by-one.
+ *
+ * To be checked tuples are passed via BatchMVCCState->tuples. Each tuple's
+ * visibility is stored in batchmvcc->visible[]. In addition,
+ * ->vistuples_dense is set to contain the offsets of visible tuples.
+ *
+ * The reason this is more efficient than HeapTupleSatisfiesMVCC() is that it
+ * avoids a cross-translation-unit function call for each tuple. In the future
+ * it will also allow more efficient setting of hint bits.
+ *
+ * Returns the number of visible tuples.
+ */
+int
+HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
+							int ntups,
+							BatchMVCCState *batchmvcc,
+							OffsetNumber *vistuples_dense)
+{
+	int			nvis = 0;
+
+	Assert(IsMVCCSnapshot(snapshot));
+
+	for (int i = 0; i < ntups; i++)
+	{
+		bool		valid;
+		HeapTuple	tup = &batchmvcc->tuples[i];
+
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		batchmvcc->visible[i] = valid;
+
+		if (likely(valid))
+		{
+			vistuples_dense[nvis] = tup->t_self.ip_posid;
+			nvis++;
+		}
+	}
+
+	return nvis;
+}
+
 /*
  * HeapTupleSatisfiesVisibility
  *		True iff heap tuple satisfies a time qual.
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 04845d5e680..2ffdf364386 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -255,6 +255,7 @@ Barrier
 BaseBackupCmd
 BaseBackupTargetHandle
 BaseBackupTargetType
+BatchMVCCState
 BeginDirectModify_function
 BeginForeignInsert_function
 BeginForeignModify_function
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v8-0006-bufmgr-Change-BufferDesc.state-to-be-a-64bit-atom.patch (46.6K, ../../ossv2eistssmubfsir6xjll76tynvxv5lup4zkrfzjkud7fycw@rf5vii6l6cha/7-v8-0006-bufmgr-Change-BufferDesc.state-to-be-a-64bit-atom.patch)
  download | inline diff:
From f9bfa4638c961015c447be6bbbe9f5b29bf3f17b Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 18 Dec 2025 18:28:50 -0500
Subject: [PATCH v8 06/10] bufmgr: Change BufferDesc.state to be a 64bit atomic

This is motivated by wanting to merge buffer content locks into
BufferDesc.state in a future commit, rather than having a separate lwlock (see
commit c75ebc657ff more details). As this change is rather mechanical, it
seems to make sense to split it out into a separate commit, for easier review.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  88 +++++----
 src/include/storage/procnumber.h              |  14 +-
 src/backend/storage/buffer/buf_init.c         |   2 +-
 src/backend/storage/buffer/bufmgr.c           | 170 +++++++++---------
 src/backend/storage/buffer/freelist.c         |  24 +--
 src/backend/storage/buffer/localbuf.c         |  72 ++++----
 contrib/pg_buffercache/pg_buffercache_pages.c |   8 +-
 src/test/modules/test_aio/test_aio.c          |  12 +-
 8 files changed, 205 insertions(+), 185 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 5400c56a965..28519ad2813 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -30,7 +30,7 @@
 #include "utils/resowner.h"
 
 /*
- * Buffer state is a single 32-bit variable where following data is combined.
+ * Buffer state is a single 64-bit variable where following data is combined.
  *
  * - 18 bits refcount
  * - 4 bits usage count
@@ -39,6 +39,9 @@
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
+ * NB: A future commit will use a significant portion of the remaining bits to
+ * implement buffer locking as part of the state variable.
+ *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
@@ -49,15 +52,21 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 				 "parts of buffer state space need to equal 32");
 
 #define BUF_REFCOUNT_ONE 1
-#define BUF_REFCOUNT_MASK ((1U << BUF_REFCOUNT_BITS) - 1)
-#define BUF_USAGECOUNT_MASK (((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_REFCOUNT_BITS))
-#define BUF_USAGECOUNT_ONE (1U << BUF_REFCOUNT_BITS)
+#define BUF_REFCOUNT_MASK \
+	((UINT64CONST(1) << BUF_REFCOUNT_BITS) - 1)
+#define BUF_USAGECOUNT_MASK \
+	(((UINT64CONST(1) << BUF_USAGECOUNT_BITS) - 1) << (BUF_REFCOUNT_BITS))
+#define BUF_USAGECOUNT_ONE \
+	(UINT64CONST(1) << BUF_REFCOUNT_BITS)
 #define BUF_USAGECOUNT_SHIFT BUF_REFCOUNT_BITS
-#define BUF_FLAG_MASK (((1U << BUF_FLAG_BITS) - 1) << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS))
+#define BUF_FLAG_MASK \
+	(((UINT64CONST(1) << BUF_FLAG_BITS) - 1) << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS))
 
 /* Get refcount and usagecount from buffer state */
-#define BUF_STATE_GET_REFCOUNT(state) ((state) & BUF_REFCOUNT_MASK)
-#define BUF_STATE_GET_USAGECOUNT(state) (((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
+#define BUF_STATE_GET_REFCOUNT(state) \
+	((uint32)((state) & BUF_REFCOUNT_MASK))
+#define BUF_STATE_GET_USAGECOUNT(state) \
+	((uint32)(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT))
 
 /*
  * Flags for buffer descriptors
@@ -65,17 +74,28 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  * Note: BM_TAG_VALID essentially means that there is a buffer hashtable
  * entry associated with the buffer's tag.
  */
-#define BM_LOCKED				(1U << 22)	/* buffer header is locked */
-#define BM_DIRTY				(1U << 23)	/* data needs writing */
-#define BM_VALID				(1U << 24)	/* data is valid */
-#define BM_TAG_VALID			(1U << 25)	/* tag is assigned */
-#define BM_IO_IN_PROGRESS		(1U << 26)	/* read or write in progress */
-#define BM_IO_ERROR				(1U << 27)	/* previous I/O failed */
-#define BM_JUST_DIRTIED			(1U << 28)	/* dirtied since write started */
-#define BM_PIN_COUNT_WAITER		(1U << 29)	/* have waiter for sole pin */
-#define BM_CHECKPOINT_NEEDED	(1U << 30)	/* must write for checkpoint */
-#define BM_PERMANENT			(1U << 31)	/* permanent buffer (not unlogged,
-											 * or init fork) */
+
+/* buffer header is locked */
+#define BM_LOCKED				(UINT64CONST(1) << 22)
+/* data needs writing */
+#define BM_DIRTY				(UINT64CONST(1) << 23)
+/* data is valid */
+#define BM_VALID				(UINT64CONST(1) << 24)
+/* tag is assigned */
+#define BM_TAG_VALID			(UINT64CONST(1) << 25)
+/* read or write in progress */
+#define BM_IO_IN_PROGRESS		(UINT64CONST(1) << 26)
+/* previous I/O failed */
+#define BM_IO_ERROR				(UINT64CONST(1) << 27)
+/* dirtied since write started */
+#define BM_JUST_DIRTIED			(UINT64CONST(1) << 28)
+/* have waiter for sole pin */
+#define BM_PIN_COUNT_WAITER		(UINT64CONST(1) << 29)
+/* must write for checkpoint */
+#define BM_CHECKPOINT_NEEDED	(UINT64CONST(1) << 30)
+/* permanent buffer (not unlogged, or init fork) */
+#define BM_PERMANENT			(UINT64CONST(1) << 31)
+
 /*
  * The maximum allowed value of usage_count represents a tradeoff between
  * accuracy and speed of the clock-sweep buffer management algorithm.  A
@@ -86,7 +106,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  */
 #define BM_MAX_USAGE_COUNT	5
 
-StaticAssertDecl(BM_MAX_USAGE_COUNT < (1 << BUF_USAGECOUNT_BITS),
+StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 				 "BM_MAX_USAGE_COUNT doesn't fit in BUF_USAGECOUNT_BITS bits");
 StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
 				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
@@ -251,8 +271,8 @@ BufMappingPartitionLockByIndex(uint32 index)
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
- * atomic operations (i.e. only pg_atomic_read_u32() and
- * pg_atomic_unlocked_write_u32()).
+ * atomic operations (i.e. only pg_atomic_read_u64() and
+ * pg_atomic_unlocked_write_u64()).
  *
  * Be careful to avoid increasing the size of the struct when adding or
  * reordering members.  Keeping it below 64 bytes (the most common CPU
@@ -280,7 +300,7 @@ typedef struct BufferDesc
 	 * State of the buffer, containing flags, refcount and usagecount. See
 	 * BUF_* and BM_* defines at the top of this file.
 	 */
-	pg_atomic_uint32 state;
+	pg_atomic_uint64 state;
 
 	/*
 	 * Backend of pin-count waiter. The buffer header spinlock needs to be
@@ -386,7 +406,7 @@ BufferDescriptorGetContentLock(const BufferDesc *bdesc)
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
  */
-extern uint32 LockBufHdr(BufferDesc *desc);
+extern uint64 LockBufHdr(BufferDesc *desc);
 
 /*
  * Unlock the buffer header.
@@ -397,9 +417,9 @@ extern uint32 LockBufHdr(BufferDesc *desc);
 static inline void
 UnlockBufHdr(BufferDesc *desc)
 {
-	Assert(pg_atomic_read_u32(&desc->state) & BM_LOCKED);
+	Assert(pg_atomic_read_u64(&desc->state) & BM_LOCKED);
 
-	pg_atomic_fetch_sub_u32(&desc->state, BM_LOCKED);
+	pg_atomic_fetch_sub_u64(&desc->state, BM_LOCKED);
 }
 
 /*
@@ -410,14 +430,14 @@ UnlockBufHdr(BufferDesc *desc)
  * Note that this approach would not work for usagecount, since we need to cap
  * the usagecount at BM_MAX_USAGE_COUNT.
  */
-static inline uint32
-UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
-				uint32 set_bits, uint32 unset_bits,
+static inline uint64
+UnlockBufHdrExt(BufferDesc *desc, uint64 old_buf_state,
+				uint64 set_bits, uint64 unset_bits,
 				int refcount_change)
 {
 	for (;;)
 	{
-		uint32		buf_state = old_buf_state;
+		uint64		buf_state = old_buf_state;
 
 		Assert(buf_state & BM_LOCKED);
 
@@ -428,7 +448,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 		if (refcount_change != 0)
 			buf_state += BUF_REFCOUNT_ONE * refcount_change;
 
-		if (pg_atomic_compare_exchange_u32(&desc->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&desc->state, &old_buf_state,
 										   buf_state))
 		{
 			return old_buf_state;
@@ -436,7 +456,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 	}
 }
 
-extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
+extern uint64 WaitBufHdrUnlocked(BufferDesc *buf);
 
 /* in bufmgr.c */
 
@@ -496,14 +516,14 @@ extern void TrackNewBufferPin(Buffer buf);
 
 /* solely to make it easier to write tests */
 extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
-extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 							  bool forget_owner, bool release_aio);
 
 
 /* freelist.c */
 extern IOContext IOContextForStrategy(BufferAccessStrategy strategy);
 extern BufferDesc *StrategyGetBuffer(BufferAccessStrategy strategy,
-									 uint32 *buf_state, bool *from_ring);
+									 uint64 *buf_state, bool *from_ring);
 extern bool StrategyRejectBuffer(BufferAccessStrategy strategy,
 								 BufferDesc *buf, bool from_ring);
 
@@ -539,7 +559,7 @@ extern BlockNumber ExtendBufferedRelLocal(BufferManagerRelation bmr,
 										  uint32 *extended_by);
 extern void MarkLocalBufferDirty(Buffer buffer);
 extern void TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty,
-								   uint32 set_flag_bits, bool release_aio);
+								   uint64 set_flag_bits, bool release_aio);
 extern bool StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait);
 extern void FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln);
 extern void InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced);
diff --git a/src/include/storage/procnumber.h b/src/include/storage/procnumber.h
index 2ddaaf0c646..6baac7c77f1 100644
--- a/src/include/storage/procnumber.h
+++ b/src/include/storage/procnumber.h
@@ -27,13 +27,13 @@ typedef int ProcNumber;
 
 /*
  * Note: MAX_BACKENDS_BITS is 18 as that is the space available for buffer
- * refcounts in buf_internals.h.  This limitation could be lifted by using a
- * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
- * currently realistic configurations. Even if that limitation were removed,
- * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
- * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
- * 4*MaxBackends without any overflow check.  We check that the configured
- * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
+ * refcounts in buf_internals.h.  This limitation could be lifted, but it's
+ * unlikely to be worthwhile as 2^18-1 backends exceed currently realistic
+ * configurations. Even if that limitation were removed, we still could not a)
+ * exceed 2^23-1 because inval.c stores the ProcNumber as a 3-byte signed
+ * integer, b) INT_MAX/4 because some places compute 4*MaxBackends without any
+ * overflow check.  We check that the configured number of backends does not
+ * exceed MAX_BACKENDS in InitializeMaxBackends().
  */
 #define MAX_BACKENDS_BITS		18
 #define MAX_BACKENDS			((1U << MAX_BACKENDS_BITS)-1)
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 6fd3a6bbac5..25f71191ec3 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -121,7 +121,7 @@ BufferManagerShmemInit(void)
 
 			ClearBufferTag(&buf->tag);
 
-			pg_atomic_init_u32(&buf->state, 0);
+			pg_atomic_init_u64(&buf->state, 0);
 			buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
 
 			buf->buf_id = i;
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index eb55102b0d7..03d99d294d5 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -780,7 +780,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 {
 	BufferDesc *bufHdr;
 	BufferTag	tag;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(recent_buffer));
 
@@ -793,7 +793,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 		int			b = -recent_buffer - 1;
 
 		bufHdr = GetLocalBufferDescriptor(b);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		/* Is it still valid and holding the right tag? */
 		if ((buf_state & BM_VALID) && BufferTagsEqual(&tag, &bufHdr->tag))
@@ -1386,8 +1386,8 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
 				bufHdr = GetLocalBufferDescriptor(-buffers[i] - 1);
 			else
 				bufHdr = GetBufferDescriptor(buffers[i] - 1);
-			Assert(pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID);
-			found = pg_atomic_read_u32(&bufHdr->state) & BM_VALID;
+			Assert(pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID);
+			found = pg_atomic_read_u64(&bufHdr->state) & BM_VALID;
 		}
 		else
 		{
@@ -1613,10 +1613,10 @@ CheckReadBuffersOperation(ReadBuffersOperation *operation, bool is_complete)
 			GetBufferDescriptor(buffer - 1);
 
 		Assert(BufferGetBlockNumber(buffer) == operation->blocknum + i);
-		Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_TAG_VALID);
+		Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_TAG_VALID);
 
 		if (i < operation->nblocks_done)
-			Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_VALID);
+			Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_VALID);
 	}
 #endif
 }
@@ -2083,8 +2083,8 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	int			existing_buf_id;
 	Buffer		victim_buffer;
 	BufferDesc *victim_buf_hdr;
-	uint32		victim_buf_state;
-	uint32		set_bits = 0;
+	uint64		victim_buf_state;
+	uint64		set_bits = 0;
 
 	/* Make sure we will have room to remember the buffer pin */
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2251,7 +2251,7 @@ InvalidateBuffer(BufferDesc *buf)
 	uint32		oldHash;		/* hash value for oldTag */
 	LWLock	   *oldPartitionLock;	/* buffer partition lock for it */
 	uint32		oldFlags;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/* Save the original buffer tag before dropping the spinlock */
 	oldTag = buf->tag;
@@ -2342,7 +2342,7 @@ retry:
 static bool
 InvalidateVictimBuffer(BufferDesc *buf_hdr)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	uint32		hash;
 	LWLock	   *partition_lock;
 	BufferTag	tag;
@@ -2402,10 +2402,10 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 
 	LWLockRelease(partition_lock);
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	Assert(!(buf_state & (BM_DIRTY | BM_VALID | BM_TAG_VALID)));
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u32(&buf_hdr->state)) > 0);
+	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u64(&buf_hdr->state)) > 0);
 
 	return true;
 }
@@ -2415,7 +2415,7 @@ GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context)
 {
 	BufferDesc *buf_hdr;
 	Buffer		buf;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		from_ring;
 
 	/*
@@ -2548,7 +2548,7 @@ again:
 
 	/* a final set of sanity checks */
 #ifdef USE_ASSERT_CHECKING
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) == 1);
 	Assert(!(buf_state & (BM_TAG_VALID | BM_VALID | BM_DIRTY)));
@@ -2839,13 +2839,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 */
 			do
 			{
-				pg_atomic_fetch_and_u32(&existing_hdr->state, ~BM_VALID);
+				pg_atomic_fetch_and_u64(&existing_hdr->state, ~BM_VALID);
 			} while (!StartBufferIO(existing_hdr, true, false));
 		}
 		else
 		{
-			uint32		buf_state;
-			uint32		set_bits = 0;
+			uint64		buf_state;
+			uint64		set_bits = 0;
 
 			buf_state = LockBufHdr(victim_buf_hdr);
 
@@ -3021,7 +3021,7 @@ BufferIsDirty(Buffer buffer)
 		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 	}
 
-	return pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY;
+	return pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY;
 }
 
 /*
@@ -3037,8 +3037,8 @@ void
 MarkBufferDirty(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
-	uint32		old_buf_state;
+	uint64		buf_state;
+	uint64		old_buf_state;
 
 	if (!BufferIsValid(buffer))
 		elog(ERROR, "bad buffer ID: %d", buffer);
@@ -3058,7 +3058,7 @@ MarkBufferDirty(Buffer buffer)
 	 * NB: We have to wait for the buffer header spinlock to be not held, as
 	 * TerminateBufferIO() relies on the spinlock.
 	 */
-	old_buf_state = pg_atomic_read_u32(&bufHdr->state);
+	old_buf_state = pg_atomic_read_u64(&bufHdr->state);
 	for (;;)
 	{
 		if (old_buf_state & BM_LOCKED)
@@ -3069,7 +3069,7 @@ MarkBufferDirty(Buffer buffer)
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
 
-		if (pg_atomic_compare_exchange_u32(&bufHdr->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&bufHdr->state, &old_buf_state,
 										   buf_state))
 			break;
 	}
@@ -3173,10 +3173,10 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 
 	if (ref == NULL)
 	{
-		uint32		buf_state;
-		uint32		old_buf_state;
+		uint64		buf_state;
+		uint64		old_buf_state;
 
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
@@ -3210,7 +3210,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 					buf_state += BUF_USAGECOUNT_ONE;
 			}
 
-			if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+			if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 											   buf_state))
 			{
 				result = (buf_state & BM_VALID) != 0;
@@ -3237,7 +3237,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		 * that the buffer page is legitimately non-accessible here.  We
 		 * cannot meddle with that.
 		 */
-		result = (pg_atomic_read_u32(&buf->state) & BM_VALID) != 0;
+		result = (pg_atomic_read_u64(&buf->state) & BM_VALID) != 0;
 
 		Assert(ref->data.refcount > 0);
 		ref->data.refcount++;
@@ -3272,7 +3272,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	/*
 	 * As explained, We don't expect any preexisting pins. That allows us to
@@ -3284,7 +3284,7 @@ PinBuffer_Locked(BufferDesc *buf)
 	 * Since we hold the buffer spinlock, we can update the buffer state and
 	 * release the lock in one operation.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 
 	UnlockBufHdrExt(buf, old_buf_state,
 					0, 0, 1);
@@ -3314,7 +3314,7 @@ WakePinCountWaiter(BufferDesc *buf)
 	 * BM_PIN_COUNT_WAITER if it stops waiting for a reason other than this
 	 * backend waking it up.
 	 */
-	uint32		buf_state = LockBufHdr(buf);
+	uint64		buf_state = LockBufHdr(buf);
 
 	if ((buf_state & BM_PIN_COUNT_WAITER) &&
 		BUF_STATE_GET_REFCOUNT(buf_state) == 1)
@@ -3361,7 +3361,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	ref->data.refcount--;
 	if (ref->data.refcount == 0)
 	{
-		uint32		old_buf_state;
+		uint64		old_buf_state;
 
 		/*
 		 * Mark buffer non-accessible to Valgrind.
@@ -3379,7 +3379,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
 		/* decrement the shared reference count */
-		old_buf_state = pg_atomic_fetch_sub_u32(&buf->state, BUF_REFCOUNT_ONE);
+		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
 
 		/* Support LockBufferForCleanup() */
 		if (old_buf_state & BM_PIN_COUNT_WAITER)
@@ -3436,7 +3436,7 @@ TrackNewBufferPin(Buffer buf)
 static void
 BufferSync(int flags)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	int			buf_id;
 	int			num_to_scan;
 	int			num_spaces;
@@ -3446,7 +3446,7 @@ BufferSync(int flags)
 	Oid			last_tsid;
 	binaryheap *ts_heap;
 	int			i;
-	uint32		mask = BM_DIRTY;
+	uint64		mask = BM_DIRTY;
 	WritebackContext wb_context;
 
 	/*
@@ -3478,7 +3478,7 @@ BufferSync(int flags)
 	for (buf_id = 0; buf_id < NBuffers; buf_id++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
-		uint32		set_bits = 0;
+		uint64		set_bits = 0;
 
 		/*
 		 * Header spinlock is enough to examine BM_DIRTY, see comment in
@@ -3645,7 +3645,7 @@ BufferSync(int flags)
 		 * write the buffer though we didn't need to.  It doesn't seem worth
 		 * guarding against this, though.
 		 */
-		if (pg_atomic_read_u32(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
+		if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
 		{
 			if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
 			{
@@ -4015,7 +4015,7 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 {
 	BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
 	int			result = 0;
-	uint32		buf_state;
+	uint64		buf_state;
 	BufferTag	tag;
 
 	/* Make sure we can handle the pin */
@@ -4264,7 +4264,7 @@ DebugPrintBufferRefcount(Buffer buffer)
 	int32		loccount;
 	char	   *result;
 	ProcNumber	backend;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 	if (BufferIsLocal(buffer))
@@ -4281,9 +4281,9 @@ DebugPrintBufferRefcount(Buffer buffer)
 	}
 
 	/* theoretically we should lock the bufhdr here */
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
-	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%x, refcount=%u %d)",
+	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%" PRIx64 ", refcount=%u %d)",
 					  buffer,
 					  relpathbackend(BufTagGetRelFileLocator(&buf->tag), backend,
 									 BufTagGetForkNum(&buf->tag)).str,
@@ -4383,7 +4383,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	instr_time	io_start;
 	Block		bufBlock;
 	char	   *bufToWrite;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
@@ -4581,7 +4581,7 @@ BufferIsPermanent(Buffer buffer)
 	 * not random garbage.
 	 */
 	bufHdr = GetBufferDescriptor(buffer - 1);
-	return (pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT) != 0;
+	return (pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT) != 0;
 }
 
 /*
@@ -5044,11 +5044,11 @@ FlushRelationBuffers(Relation rel)
 	{
 		for (i = 0; i < NLocBuffer; i++)
 		{
-			uint32		buf_state;
+			uint64		buf_state;
 
 			bufHdr = GetLocalBufferDescriptor(i);
 			if (BufTagMatchesRelFileLocator(&bufHdr->tag, &rel->rd_locator) &&
-				((buf_state = pg_atomic_read_u32(&bufHdr->state)) &
+				((buf_state = pg_atomic_read_u64(&bufHdr->state)) &
 				 (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 			{
 				ErrorContextCallback errcallback;
@@ -5084,7 +5084,7 @@ FlushRelationBuffers(Relation rel)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5156,7 +5156,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 	{
 		SMgrSortArray *srelent = NULL;
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -5405,7 +5405,7 @@ FlushDatabaseBuffers(Oid dbid)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5553,13 +5553,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u32(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
 		(BM_DIRTY | BM_JUST_DIRTIED))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
 		bool		delayChkptFlags = false;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * If we need to protect hint bit updates from torn writes, WAL-log a
@@ -5571,7 +5571,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
 		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT))
+			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5671,8 +5671,8 @@ UnlockBuffers(void)
 
 	if (buf)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		buf_state = LockBufHdr(buf);
 
@@ -5803,8 +5803,8 @@ LockBufferForCleanup(Buffer buffer)
 
 	for (;;)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		/* Try to acquire lock */
 		LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
@@ -5952,7 +5952,7 @@ bool
 ConditionalLockBufferForCleanup(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state,
+	uint64		buf_state,
 				refcount;
 
 	Assert(BufferIsValid(buffer));
@@ -6010,7 +6010,7 @@ bool
 IsBufferCleanupOK(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 
@@ -6066,7 +6066,7 @@ WaitIO(BufferDesc *buf)
 	ConditionVariablePrepareToSleep(cv);
 	for (;;)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 		PgAioWaitRef iow;
 
 		/*
@@ -6140,7 +6140,7 @@ WaitIO(BufferDesc *buf)
 bool
 StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	ResourceOwnerEnlarge(CurrentResourceOwner);
 
@@ -6196,11 +6196,11 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
  * is being released)
  */
 void
-TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 				  bool forget_owner, bool release_aio)
 {
-	uint32		buf_state;
-	uint32		unset_flag_bits = 0;
+	uint64		buf_state;
+	uint64		unset_flag_bits = 0;
 	int			refcount_change = 0;
 
 	buf_state = LockBufHdr(buf);
@@ -6261,7 +6261,7 @@ static void
 AbortBufferIO(Buffer buffer)
 {
 	BufferDesc *buf_hdr = GetBufferDescriptor(buffer - 1);
-	uint32		buf_state;
+	uint64		buf_state;
 
 	buf_state = LockBufHdr(buf_hdr);
 	Assert(buf_state & (BM_IO_IN_PROGRESS | BM_TAG_VALID));
@@ -6355,10 +6355,10 @@ rlocator_comparator(const void *p1, const void *p2)
 /*
  * Lock buffer header - set BM_LOCKED in buffer state.
  */
-uint32
+uint64
 LockBufHdr(BufferDesc *desc)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	Assert(!BufferIsLocal(BufferDescriptorGetBuffer(desc)));
 
@@ -6369,7 +6369,7 @@ LockBufHdr(BufferDesc *desc)
 		 * the spin-delay infrastructure. The work necessary for that shows up
 		 * in profiles and is rarely necessary.
 		 */
-		old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
+		old_buf_state = pg_atomic_fetch_or_u64(&desc->state, BM_LOCKED);
 		if (likely(!(old_buf_state & BM_LOCKED)))
 			break;				/* got lock */
 
@@ -6382,7 +6382,7 @@ LockBufHdr(BufferDesc *desc)
 			while (old_buf_state & BM_LOCKED)
 			{
 				perform_spin_delay(&delayStatus);
-				old_buf_state = pg_atomic_read_u32(&desc->state);
+				old_buf_state = pg_atomic_read_u64(&desc->state);
 			}
 			finish_spin_delay(&delayStatus);
 		}
@@ -6403,20 +6403,20 @@ LockBufHdr(BufferDesc *desc)
  * Obviously the buffer could be locked by the time the value is returned, so
  * this is primarily useful in CAS style loops.
  */
-pg_noinline uint32
+pg_noinline uint64
 WaitBufHdrUnlocked(BufferDesc *buf)
 {
 	SpinDelayStatus delayStatus;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	init_local_spin_delay(&delayStatus);
 
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
 	while (buf_state & BM_LOCKED)
 	{
 		perform_spin_delay(&delayStatus);
-		buf_state = pg_atomic_read_u32(&buf->state);
+		buf_state = pg_atomic_read_u64(&buf->state);
 	}
 
 	finish_spin_delay(&delayStatus);
@@ -6704,12 +6704,12 @@ ResOwnerPrintBufferPin(Datum res)
 static bool
 EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result;
 
 	*buffer_flushed = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -6803,12 +6803,12 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -6855,7 +6855,7 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
@@ -6897,12 +6897,12 @@ static bool
 MarkDirtyUnpinnedBufferInternal(Buffer buf, BufferDesc *desc,
 								bool *buffer_already_dirty)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result = false;
 
 	*buffer_already_dirty = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -7000,7 +7000,7 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_already_dirty;
 
 		CHECK_FOR_INTERRUPTS();
@@ -7054,12 +7054,12 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_already_dirty;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -7110,7 +7110,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		BufferDesc *buf_hdr = is_temp ?
 			GetLocalBufferDescriptor(-buffer - 1)
 			: GetBufferDescriptor(buffer - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * Check that all the buffers are actually ones that could conceivably
@@ -7128,7 +7128,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		}
 
 		if (is_temp)
-			buf_state = pg_atomic_read_u32(&buf_hdr->state);
+			buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		else
 			buf_state = LockBufHdr(buf_hdr);
 
@@ -7166,7 +7166,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		if (is_temp)
 		{
 			buf_state += BUF_REFCOUNT_ONE;
-			pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 		}
 		else
 			UnlockBufHdrExt(buf_hdr, buf_state, 0, 0, 1);
@@ -7352,13 +7352,13 @@ buffer_readv_complete_one(PgAioTargetData *td, uint8 buf_off, Buffer buffer,
 		: GetBufferDescriptor(buffer - 1);
 	BufferTag	tag = buf_hdr->tag;
 	char	   *bufdata = BufferGetBlock(buffer);
-	uint32		set_flag_bits;
+	uint64		set_flag_bits;
 	int			piv_flags;
 
 	/* check that the buffer is in the expected state for a read */
 #ifdef USE_ASSERT_CHECKING
 	{
-		uint32		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(!(buf_state & BM_VALID));
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 28d952b3534..1d4f19a9afd 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -86,7 +86,7 @@ typedef struct BufferAccessStrategyData
 
 /* Prototypes for internal functions */
 static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
-									 uint32 *buf_state);
+									 uint64 *buf_state);
 static void AddBufferToRing(BufferAccessStrategy strategy,
 							BufferDesc *buf);
 
@@ -171,7 +171,7 @@ ClockSweepTick(void)
  *	before returning.
  */
 BufferDesc *
-StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_ring)
+StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_ring)
 {
 	BufferDesc *buf;
 	int			bgwprocno;
@@ -230,8 +230,8 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 	trycounter = NBuffers;
 	for (;;)
 	{
-		uint32		old_buf_state;
-		uint32		local_buf_state;
+		uint64		old_buf_state;
+		uint64		local_buf_state;
 
 		buf = GetBufferDescriptor(ClockSweepTick());
 
@@ -239,7 +239,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 		 * Check whether the buffer can be used and pin it if so. Do this
 		 * using a CAS loop, to avoid having to lock the buffer header.
 		 */
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			local_buf_state = old_buf_state;
@@ -277,7 +277,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 			{
 				local_buf_state -= BUF_USAGECOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					trycounter = NBuffers;
@@ -289,7 +289,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 				/* pin the buffer if the CAS succeeds */
 				local_buf_state += BUF_REFCOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					/* Found a usable buffer */
@@ -655,12 +655,12 @@ FreeAccessStrategy(BufferAccessStrategy strategy)
  * returning.
  */
 static BufferDesc *
-GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
+GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
 {
 	BufferDesc *buf;
 	Buffer		bufnum;
-	uint32		old_buf_state;
-	uint32		local_buf_state;	/* to avoid repeated (de-)referencing */
+	uint64		old_buf_state;
+	uint64		local_buf_state;	/* to avoid repeated (de-)referencing */
 
 
 	/* Advance to next ring slot */
@@ -682,7 +682,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 	 * Check whether the buffer can be used and pin it if so. Do this using a
 	 * CAS loop, to avoid having to lock the buffer header.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 	for (;;)
 	{
 		local_buf_state = old_buf_state;
@@ -710,7 +710,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 		/* pin the buffer if the CAS succeeds */
 		local_buf_state += BUF_REFCOUNT_ONE;
 
-		if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 										   local_buf_state))
 		{
 			*buf_state = local_buf_state;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 15aac7d1c9f..a41a5facd3a 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -148,7 +148,7 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 	}
 	else
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		victim_buffer = GetLocalVictimBuffer();
 		bufid = -victim_buffer - 1;
@@ -165,10 +165,10 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 		 */
 		bufHdr->tag = newTag;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
 		buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 		*foundPtr = false;
 	}
@@ -245,12 +245,12 @@ GetLocalVictimBuffer(void)
 
 		if (LocalRefCount[victim_bufid] == 0)
 		{
-			uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 			if (BUF_STATE_GET_USAGECOUNT(buf_state) > 0)
 			{
 				buf_state -= BUF_USAGECOUNT_ONE;
-				pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+				pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 				trycounter = NLocBuffer;
 			}
 			else if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
@@ -286,13 +286,13 @@ GetLocalVictimBuffer(void)
 	 * this buffer is not referenced but it might still be dirty. if that's
 	 * the case, write it out before reusing it!
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY)
 		FlushLocalBuffer(bufHdr, NULL);
 
 	/*
 	 * Remove the victim buffer from the hashtable and mark as invalid.
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID)
 	{
 		InvalidateLocalBuffer(bufHdr, false);
 
@@ -417,7 +417,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 		if (found)
 		{
 			BufferDesc *existing_hdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			UnpinLocalBuffer(BufferDescriptorGetBuffer(victim_buf_hdr));
 
@@ -428,18 +428,18 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 			/*
 			 * Clear the BM_VALID bit, do StartLocalBufferIO() and proceed.
 			 */
-			buf_state = pg_atomic_read_u32(&existing_hdr->state);
+			buf_state = pg_atomic_read_u64(&existing_hdr->state);
 			Assert(buf_state & BM_TAG_VALID);
 			Assert(!(buf_state & BM_DIRTY));
 			buf_state &= ~BM_VALID;
-			pg_atomic_unlocked_write_u32(&existing_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&existing_hdr->state, buf_state);
 
 			/* no need to loop for local buffers */
 			StartLocalBufferIO(existing_hdr, true, false);
 		}
 		else
 		{
-			uint32		buf_state = pg_atomic_read_u32(&victim_buf_hdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&victim_buf_hdr->state);
 
 			Assert(!(buf_state & (BM_VALID | BM_TAG_VALID | BM_DIRTY | BM_JUST_DIRTIED)));
 
@@ -447,7 +447,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 
 			buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 
-			pg_atomic_unlocked_write_u32(&victim_buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&victim_buf_hdr->state, buf_state);
 
 			hresult->id = victim_buf_id;
 
@@ -467,13 +467,13 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 	{
 		Buffer		buf = buffers[i];
 		BufferDesc *buf_hdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		buf_state |= BM_VALID;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 
 	*extended_by = extend_by;
@@ -492,7 +492,7 @@ MarkLocalBufferDirty(Buffer buffer)
 {
 	int			bufid;
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsLocal(buffer));
 
@@ -506,14 +506,14 @@ MarkLocalBufferDirty(Buffer buffer)
 
 	bufHdr = GetLocalBufferDescriptor(bufid);
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	if (!(buf_state & BM_DIRTY))
 		pgBufferUsage.local_blks_dirtied++;
 
 	buf_state |= BM_DIRTY;
 
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -522,7 +522,7 @@ MarkLocalBufferDirty(Buffer buffer)
 bool
 StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * With AIO the buffer could have IO in progress, e.g. when there are two
@@ -542,7 +542,7 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 	/* Once we get here, there is definitely no I/O active on this buffer */
 
 	/* Check if someone else already did the I/O */
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 	if (forInput ? (buf_state & BM_VALID) : !(buf_state & BM_DIRTY))
 	{
 		return false;
@@ -559,11 +559,11 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
  * Like TerminateBufferIO, but for local buffers
  */
 void
-TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bits,
+TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint64 set_flag_bits,
 					   bool release_aio)
 {
 	/* Only need to adjust flags */
-	uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+	uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/* BM_IO_IN_PROGRESS isn't currently used for local buffers */
 
@@ -582,7 +582,7 @@ TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bit
 	}
 
 	buf_state |= set_flag_bits;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 	/* local buffers don't track IO using resowners */
 
@@ -606,7 +606,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(bufHdr);
 	int			bufid = -buffer - 1;
-	uint32		buf_state;
+	uint64		buf_state;
 	LocalBufferLookupEnt *hresult;
 
 	/*
@@ -622,7 +622,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 		Assert(!pgaio_wref_valid(&bufHdr->io_wref));
 	}
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/*
 	 * We need to test not just LocalRefCount[bufid] but also the BufferDesc
@@ -647,7 +647,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 	ClearBufferTag(&bufHdr->tag);
 	buf_state &= ~BUF_FLAG_MASK;
 	buf_state &= ~BUF_USAGECOUNT_MASK;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -671,9 +671,9 @@ DropRelationLocalBuffers(RelFileLocator rlocator, ForkNumber *forkNum,
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (!(buf_state & BM_TAG_VALID) ||
 			!BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -706,9 +706,9 @@ DropRelationAllLocalBuffers(RelFileLocator rlocator)
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if ((buf_state & BM_TAG_VALID) &&
 			BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -804,11 +804,11 @@ InitLocalBuffers(void)
 bool
 PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	Buffer		buffer = BufferDescriptorGetBuffer(buf_hdr);
 	int			bufid = -buffer - 1;
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	if (LocalRefCount[bufid] == 0)
 	{
@@ -819,7 +819,7 @@ PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 		{
 			buf_state += BUF_USAGECOUNT_ONE;
 		}
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/*
 		 * See comment in PinBuffer().
@@ -856,14 +856,14 @@ UnpinLocalBufferNoOwner(Buffer buffer)
 	if (--LocalRefCount[buffid] == 0)
 	{
 		BufferDesc *buf_hdr = GetLocalBufferDescriptor(buffid);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		NLocalPinnedBuffers--;
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state -= BUF_REFCOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/* see comment in UnpinBufferNoOwner */
 		VALGRIND_MAKE_MEM_NOACCESS(LocalBufHdrGetBlock(buf_hdr), BLCKSZ);
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 0c58e4b265c..529803346ce 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -199,7 +199,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
 		for (i = 0; i < NBuffers; i++)
 		{
 			BufferDesc *bufHdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			CHECK_FOR_INTERRUPTS();
 
@@ -615,7 +615,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -626,7 +626,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 		 * noticeably increase the cost of the function.
 		 */
 		bufHdr = GetBufferDescriptor(i);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (buf_state & BM_VALID)
 		{
@@ -676,7 +676,7 @@ pg_buffercache_usage_counts(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		int			usage_count;
 
 		CHECK_FOR_INTERRUPTS();
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index d7eadeab256..488d98e7e66 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -308,9 +308,9 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 {
 	Buffer		buf;
 	BufferDesc *buf_hdr;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		was_pinned = false;
-	uint32		unset_bits = 0;
+	uint64		unset_bits = 0;
 
 	/* place buffer in shared buffers without erroring out */
 	buf = ReadBufferExtended(rel, MAIN_FORKNUM, blkno, RBM_ZERO_AND_LOCK, NULL);
@@ -319,7 +319,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	}
 	else
 	{
@@ -340,7 +340,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_state &= ~unset_bits;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 	else
 	{
@@ -489,7 +489,7 @@ invalidate_rel_block(PG_FUNCTION_ARGS)
 
 			LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
-			if (pg_atomic_read_u32(&buf_hdr->state) & BM_DIRTY)
+			if (pg_atomic_read_u64(&buf_hdr->state) & BM_DIRTY)
 			{
 				if (BufferIsLocal(buf))
 					FlushLocalBuffer(buf_hdr, NULL);
@@ -572,7 +572,7 @@ buffer_call_terminate_io(PG_FUNCTION_ARGS)
 	bool		io_error = PG_GETARG_BOOL(3);
 	bool		release_aio = PG_GETARG_BOOL(4);
 	bool		clear_dirty = false;
-	uint32		set_flag_bits = 0;
+	uint64		set_flag_bits = 0;
 
 	if (io_error)
 		set_flag_bits |= BM_IO_ERROR;
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v8-0007-bufmgr-Implement-buffer-content-locks-independent.patch (44.6K, ../../ossv2eistssmubfsir6xjll76tynvxv5lup4zkrfzjkud7fycw@rf5vii6l6cha/8-v8-0007-bufmgr-Implement-buffer-content-locks-independent.patch)
  download | inline diff:
From 63d6aa0317a779efaf4542dbcb047bfd5b7ab130 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 19 Nov 2025 16:37:26 -0500
Subject: [PATCH v8 07/10] bufmgr: Implement buffer content locks independently
 of lwlocks

Until now buffer content locks were implemented using lwlocks. That has the
obvious advantage of not needing a separate efficient implementation of
locks. However, the time for a dedicated buffer content lock implementation
has come:

1) Hint bits are currently set while holding only a share lock. This leads to
   having to copy pages while they are being written out if checksums are
   enabled, which is not cheap. We would like to add AIO writes, however once
   many buffers can be written out at the same time, it gets a lot more
   expensive to copy them, particularly because that copy needs to reside in
   shared buffers (for worker mode to have access to the buffer).

   In addition, modifying buffers while they are being written out can cause
   issues with unbuffered/direct-IO, as some filesystems (like btrfs) do not
   like that, due to filesystem internal checksums getting corrupted.

   The solution to this is to require a new share-exclusive lock-level to set
   hint bits and to write out buffers, making those operations mutually
   exclusive. We could introduce such a lock level into the generic lwlock
   implementation, however it does not look like there would be other users,
   and it does add some overhead into important codepaths.

2) For AIO writes we need to be able to race-freely check whether a buffer is
   undergoing IO and whether an exclusive lock on the page can be acquired. That
   is rather hard to do efficiently when the buffer state and the lock state
   are separate atomic variables. This is a major hindrance to allowing writes
   to be done asynchronously.

3) Buffer locks are by far the most frequently taken locks. Optimizing them
   specifically for their use case is worth the effort. E.g. by merging
   content locks into buffer locks we will be able to release a buffer lock
   and pin in one atomic operation.

4) There are more complicated optimizations, like long-lived "super pinned &
   locked" pages, that cannot realistically be implemented with the generic
   lwlock implementation.

Therefore implement content locks inside bufmgr.c. The lockstate is stored as
part of BufferDesc.state. The implementation of buffer content locks is fairly
similar to lwlocks, with a few important differences:

1) An additional lock-level share-exclusive has been added. This lock level
   conflicts with exclusive locks and itself, but not share locks.

2) Error recovery for content locks is implemented as part of the already
   existing private-refcount tracking mechanism in combination with resowners,
   instead of a bespoke mechanism as the case for lwlocks. This means we do
   not need to add dedicated error-recovery codepaths to release all content
   locks (like done with LWLockReleaseAll() for content locks).

3) The lock state is embedded in BufferDesc.state instead of having its own
   struct.

4) The wakeup logic is a tad more complicated due to needing to support the
   additional lock level

This commit unfortunately introduces some code that is very similar to the
code in lwlock.c, however the code is not equivalent enough to easily merge
it. The future wins that this commit makes possible seem worth the cost.

As of this commit nothing uses the new share-exclusive lock mode. It will be
used in a future commit. It seemed too complicated to introduce the lock-level
in a separate commit.

TODO:
- Address FIXMEs

- Perhaps move the locking code into a buffer_locking.h or such? Needs to be
  inline functions for efficiency unfortunately.

- reflow some comments that I didn't reflow to make the diff more readable

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Reviewed-by: Greg Burd <greg@burd.me>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  55 +-
 src/include/storage/bufmgr.h                  |  32 +-
 src/include/storage/proc.h                    |   8 +-
 src/backend/storage/buffer/buf_init.c         |   7 +-
 src/backend/storage/buffer/bufmgr.c           | 872 ++++++++++++++++--
 .../utils/activity/wait_event_names.txt       |   3 +
 6 files changed, 893 insertions(+), 84 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 28519ad2813..0a145d95024 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -23,6 +23,7 @@
 #include "storage/condition_variable.h"
 #include "storage/lwlock.h"
 #include "storage/procnumber.h"
+#include "storage/proclist_types.h"
 #include "storage/shmem.h"
 #include "storage/smgr.h"
 #include "storage/spin.h"
@@ -32,22 +33,29 @@
 /*
  * Buffer state is a single 64-bit variable where following data is combined.
  *
+ * State of the buffer itself:
  * - 18 bits refcount
  * - 4 bits usage count
  * - 10 bits of flags
  *
+ * State of the content lock:
+ * - 1 bit has_waiter
+ * - 1 bit release_ok
+ * - 1 bit lock state locked
+ * - 1 bit exclusively locked
+ * - 1 bit share exclusively locked
+ * - 18 bits share lock count
+ *
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
- * NB: A future commit will use a significant portion of the remaining bits to
- * implement buffer locking as part of the state variable.
- *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
 #define BUF_USAGECOUNT_BITS 4
 #define BUF_FLAG_BITS 10
 
+/* FIXME: Also assert lock state size, just not yet sure how */
 StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 				 "parts of buffer state space need to equal 32");
 
@@ -69,7 +77,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 	((uint32)(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT))
 
 /*
- * Flags for buffer descriptors
+ * Flags for buffer descriptor state
  *
  * Note: BM_TAG_VALID essentially means that there is a buffer hashtable
  * entry associated with the buffer's tag.
@@ -111,6 +119,20 @@ StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
 				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
 
+
+/*
+ * Definitions related to buffer content locks
+ */
+#define BM_LOCK_HAS_WAITERS         (UINT64CONST(1) << 63)
+#define BM_LOCK_RELEASE_OK          (UINT64CONST(1) << 62)
+
+#define BM_LOCK_VAL_SHARED          (UINT64CONST(1) << 32)
+#define BM_LOCK_VAL_SHARE_EXCLUSIVE (UINT64CONST(1) << (32 + MAX_BACKENDS_BITS))
+#define BM_LOCK_VAL_EXCLUSIVE       (UINT64CONST(1) << (32 + 1 + MAX_BACKENDS_BITS))
+
+#define BM_LOCK_MASK                (((uint64)MAX_BACKENDS << 32) | BM_LOCK_VAL_SHARE_EXCLUSIVE | BM_LOCK_VAL_EXCLUSIVE)
+
+
 /*
  * Buffer tag identifies which disk block the buffer contains.
  *
@@ -253,9 +275,6 @@ BufMappingPartitionLockByIndex(uint32 index)
  * it is held.  However, existing buffer pins may be released while the buffer
  * header spinlock is held, using an atomic subtraction.
  *
- * The LWLock can take care of itself.  The buffer header lock is *not* used
- * to control access to the data in the buffer!
- *
  * If we have the buffer pinned, its tag can't change underneath us, so we can
  * examine the tag without locking the buffer header.  Also, in places we do
  * one-time reads of the flags without bothering to lock the buffer header;
@@ -268,6 +287,15 @@ BufMappingPartitionLockByIndex(uint32 index)
  * wait_backend_pgprocno and setting flag bit BM_PIN_COUNT_WAITER.  At present,
  * there can be only one such waiter per buffer.
  *
+ * The content of buffers is protected via the buffer content lock,
+ * implemented as part buffer state. Note that the buffer header lock is *not*
+ * used to control access to the data in the buffer! We used to use an LWLock
+ * to implement the content lock, but having a dedicated implementation of
+ * content locks allows to implement some otherwise hard things (e.g.
+ * race-freely checking if AIO is in progress before locking a buffer
+ * exclusively) and makes otherwise impossible optimizations possible
+ * (e.g. unlocking and unpinning a buffer in one atomic operation).
+ *
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
@@ -309,7 +337,12 @@ typedef struct BufferDesc
 	int			wait_backend_pgprocno;
 
 	PgAioWaitRef io_wref;		/* set iff AIO is in progress */
-	LWLock		content_lock;	/* to lock access to buffer contents */
+
+	/*
+	 * List of PGPROCs waiting for the buffer content lock. Protected by the
+	 * buffer header spinlock.
+	 */
+	proclist_head lock_waiters;
 } BufferDesc;
 
 /*
@@ -396,12 +429,6 @@ BufferDescriptorGetIOCV(const BufferDesc *bdesc)
 	return &(BufferIOCVArray[bdesc->buf_id]).cv;
 }
 
-static inline LWLock *
-BufferDescriptorGetContentLock(const BufferDesc *bdesc)
-{
-	return (LWLock *) (&bdesc->content_lock);
-}
-
 /*
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 97c1124c12a..df170fe9553 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -203,7 +203,20 @@ extern PGDLLIMPORT int32 *LocalRefCount;
 typedef enum BufferLockMode
 {
 	BUFFER_LOCK_UNLOCK,
+
+	/*
+	 * A share lock conflicts with exclusive locks.
+	 */
 	BUFFER_LOCK_SHARE,
+
+	/*
+	 * A share-exclusive lock conflicts with itself and exclusive locks.
+	 */
+	BUFFER_LOCK_SHARE_EXCLUSIVE,
+
+	/*
+	 * An exclusive lock conflicts with every other lock type.
+	 */
 	BUFFER_LOCK_EXCLUSIVE,
 } BufferLockMode;
 
@@ -302,7 +315,24 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
 extern void UnlockBuffers(void);
-extern void LockBuffer(Buffer buffer, BufferLockMode mode);
+extern void UnlockBuffer(Buffer buffer);
+extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
+
+/*
+ * Handling BUFFER_LOCK_UNLOCK in bufmgr.c leads to sufficiently worse branch
+ * prediction to impact performance. Therefore handle that switch here, where
+ * most of the time `mode` will be a constant and thus can be optimized out by
+ * the compiler.
+ */
+static inline void
+LockBuffer(Buffer buffer, BufferLockMode mode)
+{
+	if (mode == BUFFER_LOCK_UNLOCK)
+		UnlockBuffer(buffer);
+	else
+		LockBufferInternal(buffer, mode);
+}
+
 extern bool ConditionalLockBuffer(Buffer buffer);
 extern void LockBufferForCleanup(Buffer buffer);
 extern bool ConditionalLockBufferForCleanup(Buffer buffer);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index c6f5ebceefd..d1f6d314d57 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -242,7 +242,13 @@ struct PGPROC
 	 */
 	bool		recoveryConflictPending;
 
-	/* Info about LWLock the process is currently waiting for, if any. */
+	/*
+	 * Info about LWLock the process is currently waiting for, if any.
+	 *
+	 * This is currently used both for lwlocks and buffer content locks, which
+	 * is acceptable, although not pretty, because a backend can't wait for
+	 * both types of locks at the same time.
+	 */
 	uint8		lwWaiting;		/* see LWLockWaitState */
 	uint8		lwWaitMode;		/* lwlock mode being waited for */
 	proclist_node lwWaitLink;	/* position in LW lock wait list */
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 25f71191ec3..f3224c793c4 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
 #include "storage/aio.h"
 #include "storage/buf_internals.h"
 #include "storage/bufmgr.h"
+#include "storage/proclist.h"
 
 BufferDescPadded *BufferDescriptors;
 char	   *BufferBlocks;
@@ -121,16 +122,14 @@ BufferManagerShmemInit(void)
 
 			ClearBufferTag(&buf->tag);
 
-			pg_atomic_init_u64(&buf->state, 0);
+			pg_atomic_init_u64(&buf->state, BM_LOCK_RELEASE_OK);
 			buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
 
 			buf->buf_id = i;
 
 			pgaio_wref_clear(&buf->io_wref);
 
-			LWLockInitialize(BufferDescriptorGetContentLock(buf),
-							 LWTRANCHE_BUFFER_CONTENT);
-
+			proclist_init(&buf->lock_waiters);
 			ConditionVariableInit(BufferDescriptorGetIOCV(buf));
 		}
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 03d99d294d5..a9cafdc8ff1 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -58,6 +58,7 @@
 #include "storage/ipc.h"
 #include "storage/lmgr.h"
 #include "storage/proc.h"
+#include "storage/proclist.h"
 #include "storage/read_stream.h"
 #include "storage/smgr.h"
 #include "storage/standby.h"
@@ -100,6 +101,12 @@ typedef struct PrivateRefCountData
 	 * How many times has the buffer been pinned by this backend.
 	 */
 	int32		refcount;
+
+	/*
+	 * Is the buffer locked by this backend? BUFFER_LOCK_UNLOCK indicates that
+	 * the buffer is not locked.
+	 */
+	BufferLockMode lockmode;
 } PrivateRefCountData;
 
 typedef struct PrivateRefCountEntry
@@ -210,8 +217,10 @@ static BufferDesc *PinCountWaitBuf = NULL;
  * Each buffer also has a private refcount that keeps track of the number of
  * times the buffer is pinned in the current process.  This is so that the
  * shared refcount needs to be modified only once if a buffer is pinned more
- * than once by an individual backend.  It's also used to check that no buffers
- * are still pinned at the end of transactions and when exiting.
+ * than once by an individual backend.  It's also used to check that no
+ * buffers are still pinned at the end of transactions and when exiting. We
+ * also use this mechanism to track whether this backend has a buffer locked,
+ * and, if so, in what mode.
  *
  *
  * To avoid - as we used to - requiring an array with NBuffers entries to keep
@@ -351,6 +360,7 @@ ReservePrivateRefCountEntry(void)
 		/* clear the whole data member, just for future proofing */
 		memset(&victim_entry->data, 0, sizeof(victim_entry->data));
 		victim_entry->data.refcount = 0;
+		victim_entry->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 		PrivateRefCountOverflowed++;
 	}
@@ -374,6 +384,7 @@ NewPrivateRefCountEntry(Buffer buffer)
 	PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
 	res->buffer = buffer;
 	res->data.refcount = 0;
+	res->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 	/* update cache for the next lookup */
 	PrivateRefCountEntryLast = ReservedRefCountSlot;
@@ -540,6 +551,7 @@ static void
 ForgetPrivateRefCountEntry(PrivateRefCountEntry *ref)
 {
 	Assert(ref->data.refcount == 0);
+	Assert(ref->data.lockmode == BUFFER_LOCK_UNLOCK);
 
 	if (ref >= &PrivateRefCountArray[0] &&
 		ref < &PrivateRefCountArray[REFCOUNT_ARRAY_ENTRIES])
@@ -641,14 +653,27 @@ static void RelationCopyStorageUsingBuffer(RelFileLocator srclocator,
 static void AtProcExit_Buffers(int code, Datum arg);
 static void CheckForBufferLeaks(void);
 #ifdef USE_ASSERT_CHECKING
-static void AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
-									   void *unused_context);
+static void AssertNotCatalogBufferLock(Buffer buffer, BufferLockMode mode);
 #endif
 static int	rlocator_comparator(const void *p1, const void *p2);
 static inline int buffertag_comparator(const BufferTag *ba, const BufferTag *bb);
 static inline int ckpt_buforder_comparator(const CkptSortItem *a, const CkptSortItem *b);
 static int	ts_ckpt_progress_comparator(Datum a, Datum b, void *arg);
 
+static void BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr);
+static bool BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMe(BufferDesc *buf_hdr);
+static inline void BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr);
+static inline int BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr);
+static inline bool BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockDequeueSelf(BufferDesc *buf_hdr);
+static void BufferLockWakeup(BufferDesc *buf_hdr, bool unlocked);
+static void BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate);
+static inline uint64 BufferLockReleaseSub(BufferLockMode mode);
+
 
 /*
  * Implementation of PrefetchBuffer() for shared buffers.
@@ -2449,8 +2474,6 @@ again:
 	 */
 	if (buf_state & BM_DIRTY)
 	{
-		LWLock	   *content_lock;
-
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(buf_state & BM_VALID);
 
@@ -2468,8 +2491,7 @@ again:
 		 * one just happens to be trying to split the page the first one got
 		 * from StrategyGetBuffer.)
 		 */
-		content_lock = BufferDescriptorGetContentLock(buf_hdr);
-		if (!LWLockConditionalAcquire(content_lock, LW_SHARED))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -2498,7 +2520,7 @@ again:
 			if (XLogNeedsFlush(lsn)
 				&& StrategyRejectBuffer(strategy, buf_hdr, from_ring))
 			{
-				LWLockRelease(content_lock);
+				LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 				UnpinBuffer(buf_hdr);
 				goto again;
 			}
@@ -2506,7 +2528,7 @@ again:
 
 		/* OK, do the I/O */
 		FlushBuffer(buf_hdr, NULL, IOOBJECT_RELATION, io_context);
-		LWLockRelease(content_lock);
+		LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 		ScheduleBufferTagForWriteback(&BackendWritebackContext, io_context,
 									  &buf_hdr->tag);
@@ -2948,7 +2970,7 @@ BufferIsLockedByMe(Buffer buffer)
 	else
 	{
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr));
+		return BufferLockHeldByMe(bufHdr);
 	}
 }
 
@@ -2973,23 +2995,8 @@ BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode)
 	}
 	else
 	{
-		LWLockMode	lw_mode;
-
-		switch (mode)
-		{
-			case BUFFER_LOCK_EXCLUSIVE:
-				lw_mode = LW_EXCLUSIVE;
-				break;
-			case BUFFER_LOCK_SHARE:
-				lw_mode = LW_SHARED;
-				break;
-			default:
-				pg_unreachable();
-		}
-
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-									lw_mode);
+		return BufferLockHeldByMeInMode(bufHdr, mode);
 	}
 }
 
@@ -3376,7 +3383,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 * I'd better not still hold the buffer content lock. Can't use
 		 * BufferIsLockedByMe(), as that asserts the buffer is pinned.
 		 */
-		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
+		Assert(!BufferLockHeldByMe(buf));
 
 		/* decrement the shared reference count */
 		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
@@ -4198,7 +4205,7 @@ CheckForBufferLeaks(void)
  * Check for exclusive-locked catalog buffers.  This is the core of
  * AssertCouldGetRelation().
  *
- * A backend would self-deadlock on LWLocks if the catalog scan read the
+ * A backend would self-deadlock on the content if the catalog scan read the
  * exclusive-locked buffer.  The main threat is exclusive-locked buffers of
  * catalogs used in relcache, because a catcache search on any catalog may
  * build that catalog's relcache entry.  We don't have an inventory of
@@ -4214,26 +4221,45 @@ CheckForBufferLeaks(void)
 void
 AssertBufferLocksPermitCatalogRead(void)
 {
-	ForEachLWLockHeldByMe(AssertNotCatalogBufferLock, NULL);
+	PrivateRefCountEntry *res;
+
+	/* check the array */
+	for (int i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
+	{
+		if (PrivateRefCountArrayKeys[i] != InvalidBuffer)
+		{
+			res = &PrivateRefCountArray[i];
+
+			if (res->buffer == InvalidBuffer)
+				continue;
+
+			AssertNotCatalogBufferLock(res->buffer, res->data.lockmode);
+		}
+	}
+
+	/* if necessary search the hash */
+	if (PrivateRefCountOverflowed)
+	{
+		HASH_SEQ_STATUS hstat;
+
+		hash_seq_init(&hstat, PrivateRefCountHash);
+		while ((res = (PrivateRefCountEntry *) hash_seq_search(&hstat)) != NULL)
+		{
+			AssertNotCatalogBufferLock(res->buffer, res->data.lockmode);
+		}
+	}
 }
 
 static void
-AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
-						   void *unused_context)
+AssertNotCatalogBufferLock(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *bufHdr;
+	BufferDesc *bufHdr = GetBufferDescriptor(buffer - 1);
 	BufferTag	tag;
 	Oid			relid;
 
-	if (mode != LW_EXCLUSIVE)
+	if (mode != BUFFER_LOCK_EXCLUSIVE)
 		return;
 
-	if (!((BufferDescPadded *) lock > BufferDescriptors &&
-		  (BufferDescPadded *) lock < BufferDescriptors + NBuffers))
-		return;					/* not a buffer lock */
-
-	bufHdr = (BufferDesc *)
-		((char *) lock - offsetof(BufferDesc, content_lock));
 	tag = bufHdr->tag;
 
 	/*
@@ -4515,9 +4541,11 @@ static void
 FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 					IOObject io_object, IOContext io_context)
 {
-	LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	Buffer		buffer = BufferDescriptorGetBuffer(buf);
+
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-	LWLockRelease(BufferDescriptorGetContentLock(buf));
+	BufferLockUnlock(buffer, buf);
 }
 
 /*
@@ -5660,9 +5688,10 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
  *
  * Used to clean up after errors.
  *
- * Currently, we can expect that lwlock.c's LWLockReleaseAll() took care
- * of releasing buffer content locks per se; the only thing we need to deal
- * with here is clearing any PIN_COUNT request that was in progress.
+ * Currently, we can expect that resource owner cleanup, via
+ * ResOwnerReleaseBufferPin(), took care releasing buffer content locks per
+ * se; the only thing we need to deal with here is clearing any PIN_COUNT
+ * request that was in progress.
  */
 void
 UnlockBuffers(void)
@@ -5693,25 +5722,727 @@ UnlockBuffers(void)
 }
 
 /*
- * Acquire or release the content_lock for the buffer.
+ * Acquire the buffer content lock in the specified mode
+ *
+ * If the lock is not available, sleep until it is.
+ *
+ * Side effect: cancel/die interrupts are held off until lock release.
+ *
+ * This uses almost the same locking approach as lwlock.c's
+ * LWLockAcquire(). See documentation atop of lwlock.c for a more detailed
+ * discussion.
+ *
+ * The reason that this, and most of the other BufferLock* functions, get both
+ * the Buffer and BufferDesc* as parameters, is that looking up one from the
+ * other repeatedly shows up noticeably in profiles.
+ *
+ * Callers should provide a constant for mode, for more efficient code
+ * generation.
+ */
+static inline void
+BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry;
+	int			extraWaits = 0;
+
+	/*
+	 * Get reference to the refcount entry before we hold the lock, it seems
+	 * better to do before holding the lock.
+	 */
+	entry = GetPrivateRefCountEntry(buffer, true);
+
+	/*
+	 * We better not already hold a lock on the buffer.
+	 */
+	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	for (;;)
+	{
+		bool		mustwait;
+		uint32		wait_event;
+
+		/*
+		 * Try to grab the lock the first time, we're not in the waitqueue
+		 * yet/anymore.
+		 */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		if (likely(!mustwait))
+		{
+			break;
+		}
+
+		/*
+		 * Ok, at this point we couldn't grab the lock on the first try. We
+		 * cannot simply queue ourselves to the end of the list and wait to be
+		 * woken up because by now the lock could long have been released.
+		 * Instead add us to the queue and try to grab the lock again. If we
+		 * succeed we need to revert the queuing and be happy, otherwise we
+		 * recheck the lock. If we still couldn't grab it, we know that the
+		 * other locker will see our queue entries when releasing since they
+		 * existed before we checked for the lock.
+		 */
+
+		/* add to the queue */
+		BufferLockQueueSelf(buf_hdr, mode);
+
+		/* we're now guaranteed to be woken up if necessary */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		/* ok, grabbed the lock the second time round, need to undo queueing */
+		if (!mustwait)
+		{
+			BufferLockDequeueSelf(buf_hdr);
+			break;
+		}
+
+		switch (mode)
+		{
+			case BUFFER_LOCK_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_SHARE_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE:
+				wait_event = WAIT_EVENT_BUFFER_SHARED;
+				break;
+			case BUFFER_LOCK_UNLOCK:
+				pg_unreachable();
+
+		}
+		pgstat_report_wait_start(wait_event);
+
+		/*
+		 * Wait until awakened.
+		 *
+		 * It is possible that we get awakened for a reason other than being
+		 * signaled by LWLockRelease.  If so, loop back and wait again.  Once
+		 * we've gotten the LWLock, re-increment the sema by the number of
+		 * additional signals received.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		pgstat_report_wait_end();
+
+		/* Retrying, allow BufferLockRelease to release waiters again. */
+		pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_RELEASE_OK);
+	}
+
+	/* Remember that we now hold this lock */
+	entry->data.lockmode = mode;
+
+	/*
+	 * Fix the process wait semaphore's count for any absorbed wakeups.
+	 */
+	while (unlikely(extraWaits-- > 0))
+		PGSemaphoreUnlock(MyProc->sem);
+}
+
+/*
+ * Release a previously acquired buffer content lock.
+ */
+static void
+BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	uint64		oldstate;
+	uint64		sub;
+
+	mode = BufferLockDisownInternal(buffer, buf_hdr);
+
+	/*
+	 * Release my hold on lock, after that it can immediately be acquired by
+	 * others, even if we still have to wakeup other waiters.
+	 */
+	sub = BufferLockReleaseSub(mode);
+
+	oldstate = pg_atomic_sub_fetch_u64(&buf_hdr->state, sub);
+
+	BufferLockProcessRelease(buf_hdr, mode, oldstate);
+
+	/*
+	 * Now okay to allow cancel/die interrupts.
+	 */
+	RESUME_INTERRUPTS();
+}
+
+
+/*
+ * Acquire the content lock for the buffer, but only if we don't have to wait.
+ */
+static bool
+BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	bool		mustwait;
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	/* Check for the lock */
+	mustwait = BufferLockAttempt(buf_hdr, mode);
+
+	if (mustwait)
+	{
+		/* Failed to get lock, so release interrupt holdoff */
+		RESUME_INTERRUPTS();
+	}
+	else
+	{
+		PrivateRefCountEntry *entry =
+			GetPrivateRefCountEntry(buffer, true);
+
+		entry->data.lockmode = mode;
+	}
+
+	return !mustwait;
+}
+
+/*
+ * Internal function that tries to atomically acquire the content lock in the
+ * passed in mode.
+ *
+ * This function will not block waiting for a lock to become free - that's the
+ * caller's job.
+ *
+ * Similar to LWLockAttemptLock().
+ */
+static inline bool
+BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	uint64		old_state;
+
+	/*
+	 * Read once outside the loop, later iterations will get the newer value
+	 * via compare & exchange.
+	 */
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+
+	/* loop until we've determined whether we could acquire the lock or not */
+	while (true)
+	{
+		uint64		desired_state;
+		bool		lock_free;
+
+		desired_state = old_state;
+
+		if (mode == BUFFER_LOCK_EXCLUSIVE)
+		{
+			lock_free = (old_state & BM_LOCK_MASK) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_EXCLUSIVE;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			lock_free = (old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+		}
+		else
+		{
+			lock_free = (old_state & BM_LOCK_VAL_EXCLUSIVE) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARED;
+		}
+
+		/*
+		 * Attempt to swap in the state we are expecting. If we didn't see
+		 * lock to be free, that's just the old value. If we saw it as free,
+		 * we'll attempt to mark it acquired. The reason that we always swap
+		 * in the value is that this doubles as a memory barrier. We could try
+		 * to be smarter and only swap in values if we saw the lock as free,
+		 * but benchmark haven't shown it as beneficial so far.
+		 *
+		 * Retry if the value changed since we last looked at it.
+		 */
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			if (lock_free)
+			{
+				/* Great! Got the lock. */
+				return false;
+			}
+			else
+				return true;	/* somebody else has the lock */
+		}
+	}
+
+	pg_unreachable();
+}
+
+/*
+ * Add ourselves to the end of the content lock's wait queue.
+ */
+static void
+BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	/*
+	 * If we don't have a PGPROC structure, there's no way to wait. This
+	 * should never occur, since MyProc should only be null during shared
+	 * memory initialization.
+	 */
+	if (MyProc == NULL)
+		elog(PANIC, "cannot wait without a PGPROC structure");
+
+	if (MyProc->lwWaiting != LW_WS_NOT_WAITING)
+		elog(PANIC, "queueing for lock while waiting on another one");
+
+	LockBufHdr(buf_hdr);
+
+	/* setting the flag is protected by the spinlock */
+	pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_HAS_WAITERS);
+
+	/*
+	 * FIXME: This is reusing the lwlock fields. That's not a correctness
+	 * issue, a backend can't wait for both an lwlock and a buffer content
+	 * lock at the same time. However, it seems pretty ugly, particularly
+	 * given that the field names have an lw* prefix. But duplicating the
+	 * fields also seems somewhat superfluous.
+	 */
+	MyProc->lwWaiting = LW_WS_WAITING;
+	MyProc->lwWaitMode = mode;
+
+	proclist_push_tail(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	/* Can release the mutex now */
+	UnlockBufHdr(buf_hdr);
+}
+
+/*
+ * Remove ourselves from the waitlist.
+ *
+ * This is used if we queued ourselves because we thought we needed to sleep
+ * but, after further checking, we discovered that we don't actually need to
+ * do so.
+ */
+static void
+BufferLockDequeueSelf(BufferDesc *buf_hdr)
+{
+	bool		on_waitlist;
+
+	LockBufHdr(buf_hdr);
+
+	on_waitlist = MyProc->lwWaiting == LW_WS_WAITING;
+	if (on_waitlist)
+		proclist_delete(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	if (proclist_is_empty(&buf_hdr->lock_waiters) &&
+		(pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS) != 0)
+	{
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_HAS_WAITERS);
+	}
+
+	/* XXX: combine with fetch_and above? */
+	UnlockBufHdr(buf_hdr);
+
+	/* clear waiting state again, nice for debugging */
+	if (on_waitlist)
+		MyProc->lwWaiting = LW_WS_NOT_WAITING;
+	else
+	{
+		int			extraWaits = 0;
+
+
+		/*
+		 * Somebody else dequeued us and has or will wake us up. Deal with the
+		 * superfluous absorption of a wakeup.
+		 */
+
+		/*
+		 * Reset RELEASE_OK flag if somebody woke us before we removed
+		 * ourselves - they'll have set it to false.
+		 */
+		pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_RELEASE_OK);
+
+		/*
+		 * Now wait for the scheduled wakeup, otherwise our ->lwWaiting would
+		 * get reset at some inconvenient point later. Most of the time this
+		 * will immediately return.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		/*
+		 * Fix the process wait semaphore's count for any absorbed wakeups.
+		 */
+		while (extraWaits-- > 0)
+			PGSemaphoreUnlock(MyProc->sem);
+	}
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * After calling this function it's the callers responsibility to ensure that
+ * the lock gets released, even in case of an error. This only is desirable if
+ * the lock is going to be released in a different process than the process
+ * that acquired it.
+ */
+static inline void
+BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockDisownInternal(buffer, buf_hdr);
+	RESUME_INTERRUPTS();
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * This is the code that can be shared between actually releasing a lock
+ * (BufferLockUnlock()) and just not tracking ownership of the lock anymore
+ * without releasing the lock (BufferLockDisown()).
+ */
+static inline int
+BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	PrivateRefCountEntry *ref;
+
+	ref = GetPrivateRefCountEntry(buffer, false);
+	if (ref == NULL)
+		elog(ERROR, "lock %d is not held", buffer);
+	mode = ref->data.lockmode;
+	ref->data.lockmode = BUFFER_LOCK_UNLOCK;
+
+	return mode;
+}
+
+/*
+ * Wakeup all the lockers that currently have a chance to acquire the lock.
+ *
+ * wake_exclusive indicates whether exlusive lock waiters should be woken up.
+ */
+static void
+BufferLockWakeup(BufferDesc *buf_hdr, bool wake_exclusive)
+{
+	bool		new_release_ok;
+	bool		wake_share_exclusive = true;
+	proclist_head wakeup;
+	proclist_mutable_iter iter;
+
+	proclist_init(&wakeup);
+
+	new_release_ok = true;
+
+	/* lock wait list while collecting backends to wake up */
+	LockBufHdr(buf_hdr);
+
+	proclist_foreach_modify(iter, &buf_hdr->lock_waiters, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		/*
+		 * Already woke up a conflicting lock, so skip over this wait list
+		 * entry.
+		 */
+		if (!wake_exclusive && waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+			continue;
+		if (!wake_share_exclusive && waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+			continue;
+
+		proclist_delete(&buf_hdr->lock_waiters, iter.cur, lwWaitLink);
+		proclist_push_tail(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Prevent additional wakeups until retryer gets to run. Backends that
+		 * are just waiting for the lock to become free don't retry
+		 * automatically.
+		 */
+		new_release_ok = false;
+
+		/*
+		 * Signal that the process isn't on the wait list anymore. This allows
+		 * BufferLockDequeueSelf() to remove itself from the waitlist with a
+		 * proclist_delete(), rather than having to check if it has been
+		 * removed from the list.
+		 */
+		Assert(waiter->lwWaiting == LW_WS_WAITING);
+		waiter->lwWaiting = LW_WS_PENDING_WAKEUP;
+
+		/*
+		 * Don't wakeup further waiters after waking a conflicting waiter.
+		 */
+		if (waiter->lwWaitMode == BUFFER_LOCK_SHARE)
+		{
+			/*
+			 * Share locks conflict with exclusive locks.
+			 */
+			wake_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			/*
+			 * Share-exclusive locks conflict with share-exclusive eand
+			 * exclusive locks.
+			 */
+			wake_exclusive = false;
+			wake_share_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+		{
+
+			/*
+			 * Exclusive locks conflict with all other locks, there's no point
+			 * in waking up anybody else.
+			 */
+			break;
+		}
+	}
+
+	Assert(proclist_is_empty(&wakeup) || pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS);
+
+	/* unset required flags, and release lock, in one fell swoop */
+	{
+		uint64		old_state;
+		uint64		desired_state;
+
+		old_state = pg_atomic_read_u64(&buf_hdr->state);
+		while (true)
+		{
+			desired_state = old_state;
+
+			/* compute desired flags */
+
+			if (new_release_ok)
+				desired_state |= BM_LOCK_RELEASE_OK;
+			else
+				desired_state &= ~BM_LOCK_RELEASE_OK;
+
+			if (proclist_is_empty(&buf_hdr->lock_waiters))
+				desired_state &= ~BM_LOCK_HAS_WAITERS;
+
+			desired_state &= ~BM_LOCKED;	/* release lock */
+
+			if (pg_atomic_compare_exchange_u64(&buf_hdr->state, &old_state,
+											   desired_state))
+				break;
+		}
+	}
+
+	/* Awaken any waiters I removed from the queue. */
+	proclist_foreach_modify(iter, &wakeup, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		proclist_delete(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Guarantee that lwWaiting being unset only becomes visible once the
+		 * unlink from the link has completed. Otherwise the target backend
+		 * could be woken up for other reason and enqueue for a new lock - if
+		 * that happens before the list unlink happens, the list would end up
+		 * being corrupted.
+		 *
+		 * The barrier pairs with the LWLockWaitListLock() when enqueuing for
+		 * another lock.
+		 */
+		pg_write_barrier();
+		waiter->lwWaiting = LW_WS_NOT_WAITING;
+		PGSemaphoreUnlock(waiter->sem);
+	}
+}
+
+/*
+ * Compute subtraction from buffer state for a release of a held lock in
+ * `mode`.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places, each needing to compute what to
+ * subtract from the lock state.
+ */
+static inline uint64
+BufferLockReleaseSub(BufferLockMode mode)
+{
+
+	/*
+	 * Turns out that a switch() leads gcc to generate sufficiently worse code
+	 * for this to show up in profiles...
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE)
+		return BM_LOCK_VAL_EXCLUSIVE;
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		return BM_LOCK_VAL_SHARE_EXCLUSIVE;
+	else
+	{
+		Assert(mode == BUFFER_LOCK_SHARE);
+		return BM_LOCK_VAL_SHARED;
+	}
+
+	return 0;
+}
+
+/*
+ * Handle work that needs to be done after releasing a lock that was held in
+ * `mode`, where `lockstate` is the result of the atomic operation modifying
+ * the state variable.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places.
+ */
+static void
+BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate)
+{
+	bool		check_waiters = false;
+	bool		wake_exclusive = false;
+
+	/* nobody else can have that kind of lock */
+	Assert(!(lockstate & BM_LOCK_VAL_EXCLUSIVE));
+
+	/*
+	 * If we're still waiting for backends to get scheduled, don't wake them
+	 * up again. Otherwise check if we need to look through the waitqueue to
+	 * wake other backends.
+	 */
+	if ((lockstate & (BM_LOCK_HAS_WAITERS | BM_LOCK_RELEASE_OK)) ==
+		(BM_LOCK_HAS_WAITERS | BM_LOCK_RELEASE_OK))
+	{
+		if ((lockstate & BM_LOCK_MASK) == 0)
+		{
+			/*
+			 * We released a lock and the lock was, in that moment, free. We
+			 * therefore can wake waiters for any kind of lock.
+			 */
+			check_waiters = true;
+			wake_exclusive = true;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			/*
+			 * We released the lock, but another backend still holds a lock.
+			 * We can't have released an exclusive lock, as there couldn't
+			 * have been other lock holders. If we released a share lock, no
+			 * waiters need to be woken up, as there must be other share
+			 * lockers. However, if we held a share-exclusive lock, another
+			 * backend now could acquire a share-exclusive lock.
+			 */
+			check_waiters = true;
+			wake_exclusive = false;
+		}
+	}
+
+	/*
+	 * As waking up waiters requires the spinlock to be acquired, only do so
+	 * if necessary.
+	 */
+	if (check_waiters)
+		BufferLockWakeup(buf_hdr, wake_exclusive);
+}
+
+/*
+ * BufferLockHeldByMeInMode - test whether my process holds the content lock
+ * in the specified mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode == mode;
+
+}
+
+/*
+ * BufferLockHeldByMe - test whether my process holds the content lock in any
+ * mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMe(BufferDesc *buf_hdr)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode != BUFFER_LOCK_UNLOCK;
+}
+
+/*
+ * Release the content lock for the buffer.
+ */
+void
+UnlockBuffer(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+
+	Assert(BufferIsPinned(buffer));
+	if (BufferIsLocal(buffer))
+		return;					/* local buffers need no lock */
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+	BufferLockUnlock(buffer, buf_hdr);
+}
+
+/*
+ * Acquire the content_lock for the buffer.
  */
 void
-LockBuffer(Buffer buffer, BufferLockMode mode)
+LockBufferInternal(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *buf;
+	BufferDesc *buf_hdr;
+
+	/*
+	 * We can't wait if we haven't got a PGPROC.  This should only occur
+	 * during bootstrap or shared memory initialization.  Put an Assert here
+	 * to catch unsafe coding practices.
+	 */
+	Assert(!(MyProc == NULL && IsUnderPostmaster));
+
+	/* handled in LockBuffer() wrapper */
+	Assert(mode != BUFFER_LOCK_UNLOCK);
 
 	Assert(BufferIsPinned(buffer));
 	if (BufferIsLocal(buffer))
 		return;					/* local buffers need no lock */
 
-	buf = GetBufferDescriptor(buffer - 1);
+	buf_hdr = GetBufferDescriptor(buffer - 1);
 
-	if (mode == BUFFER_LOCK_UNLOCK)
-		LWLockRelease(BufferDescriptorGetContentLock(buf));
-	else if (mode == BUFFER_LOCK_SHARE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	/*
+	 * Test the most frequent lock modes first. While a switch (mode) would be
+	 * nice, at least gcc generates considerably worse code for it.
+	 *
+	 * Call BufferLockAcquire() with a constant argument for mode, to generate
+	 * more efficient code for the different lock modes.
+	 */
+	if (mode == BUFFER_LOCK_SHARE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE);
 	else if (mode == BUFFER_LOCK_EXCLUSIVE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_EXCLUSIVE);
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_EXCLUSIVE);
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	else
 		elog(ERROR, "unrecognized buffer lock mode: %d", mode);
 }
@@ -5732,8 +6463,7 @@ ConditionalLockBuffer(Buffer buffer)
 
 	buf = GetBufferDescriptor(buffer - 1);
 
-	return LWLockConditionalAcquire(BufferDescriptorGetContentLock(buf),
-									LW_EXCLUSIVE);
+	return BufferLockConditional(buffer, buf, BUFFER_LOCK_EXCLUSIVE);
 }
 
 /*
@@ -6688,7 +7418,25 @@ ResOwnerReleaseBufferPin(Datum res)
 	if (BufferIsLocal(buffer))
 		UnpinLocalBufferNoOwner(buffer);
 	else
+	{
+		PrivateRefCountEntry *ref;
+
+		ref = GetPrivateRefCountEntry(buffer, false);
+
+		/*
+		 * If the buffer was locked at the time of the resowner release,
+		 * release the lock now. This should only happen after errors.
+		 */
+		if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
+		{
+			BufferDesc *buf = GetBufferDescriptor(buffer - 1);
+
+			HOLD_INTERRUPTS();	/* match the upcoming RESUME_INTERRUPTS */
+			BufferLockUnlock(buffer, buf);
+		}
+
 		UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
+	}
 }
 
 static char *
@@ -6924,10 +7672,10 @@ MarkDirtyUnpinnedBufferInternal(Buffer buf, BufferDesc *desc,
 	/* If it was not already dirty, mark it as dirty. */
 	if (!(buf_state & BM_DIRTY))
 	{
-		LWLockAcquire(BufferDescriptorGetContentLock(desc), LW_EXCLUSIVE);
+		BufferLockAcquire(buf, desc, BUFFER_LOCK_EXCLUSIVE);
 		MarkBufferDirty(buf);
 		result = true;
-		LWLockRelease(BufferDescriptorGetContentLock(desc));
+		BufferLockUnlock(buf, desc);
 	}
 	else
 		*buffer_already_dirty = true;
@@ -7178,16 +7926,12 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		 */
 		if (is_write && !is_temp)
 		{
-			LWLock	   *content_lock;
-
-			content_lock = BufferDescriptorGetContentLock(buf_hdr);
-
-			Assert(LWLockHeldByMe(content_lock));
+			Assert(BufferLockHeldByMe(buf_hdr));
 
 			/*
 			 * Lock is now owned by AIO subsystem.
 			 */
-			LWLockDisown(content_lock);
+			BufferLockDisown(buffer, buf_hdr);
 		}
 
 		/*
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index c0632bf901a..d236c87294f 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -286,6 +286,9 @@ ABI_compatibility:
 Section: ClassName - WaitEventBuffer
 
 BUFFER_CLEANUP	"Waiting to acquire an exclusive pin on a buffer. Buffer pin waits can be protracted if another process holds an open cursor that last read data from the buffer in question."
+BUFFER_SHARED	"Waiting to acquire shared lock on a buffer."
+BUFFER_SHARE_EXCLUSIVE	"Waiting to acquire share exclusive lock on a buffer."
+BUFFER_EXCLUSIVE	"Waiting to acquire exclusive lock on a buffer."
 
 ABI_compatibility:
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v8-0008-Require-share-exclusive-lock-to-set-hint-bits-and.patch (37.8K, ../../ossv2eistssmubfsir6xjll76tynvxv5lup4zkrfzjkud7fycw@rf5vii6l6cha/9-v8-0008-Require-share-exclusive-lock-to-set-hint-bits-and.patch)
  download | inline diff:
From 4d8ef9702f05f3ca215a2da90ed02c213d2d8da5 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Fri, 12 Dec 2025 15:31:01 -0500
Subject: [PATCH v8 08/10] Require share-exclusive lock to set hint bits and to
 flush

At the moment hint bits can be set with just a share lock on a page (and in
one place even without any lock). Because of this we need to copy pages while
writing them out, as otherwise the checksum could be corrupted.

The need to copy the page is problematic to implement AIO writes:

1) Instead of just needing a single buffer for a copied page we need one for
   each page that's potentially undergoing IO
2) To be able to use the "worker" AIO implementation the copied page needs to
   reside in shared memory

It also causes problems for using unbuffered/direct-IO, independent of AIO:
Some filesystems, raid implementations, ... do not tolerate the data being
written out to change during the write. E.g. they may compute internal
checksums that can be invalidated by concurrent modifications, leading e.g. to
filesystem errors (as the case with btrfs).

It also just is plain odd to allow modifications of buffers that are just
share locked.

To address these issue, this commit changes the rules so that modifications to
pages are not allowed anymore while holding a share lock. Instead the new
share-exclusive lock (introduced in FIXME XXXX TODO) allows at most one
backend to modify a buffer while other backends have the same page share
locked. An existing share-lock can be upgraded to a share-exclusive lock, if
there are no conflicting locks. For that
BufferBeginSetHintBits()/BufferBeginSetHintBits() and BufferSetHintBits16()
have been introduced.

To prevent hint bits from being set while the buffer is being written out,
writing out buffers now requires a share-exclusive lock.

The use of share-exclusive to gate setting hint bits means that from now on
only one backend can set hint bits at a time. To allow multiple backends
setting hint bits would require more complicated locking, for setting hint
bits we'd need to store the count of backends currently setting hint bits and
we would need another lock-level for I/O conflicting with the lock-level to
set hint bits. Given that the share-exclusive lock for setting hint bits is
only held for a short time, that often backends would just set the same hint
bits and that the cost of occasionally not setting hint bits in hotly accessed
pages is fairly low, this seems like an acceptable tradeoff.

The biggest change to adapt to this is in heapam. To avoid performance
regressions for sequential scans that need to set a lot of hint bits, we need
to amortize the cost of BufferBeginSetHintBits() for cases where hint bits are
set at a high frequency, HeapTupleSatisfiesMVCCBatch() uses the new
SetHintBitsExt() which defers BufferFinishSetHintBits() until all hint bits on
a page have been set.  Conversely, to avoid regressions in cases where we
can't set hint bits in bulk (because we're looking only at individual tuples),
use BufferSetHintBits16() when setting hint bits without batching.

Several other places also need to be adapted, but those changes are
comparatively simpler.

After this we do not need to copy buffers to write them out anymore. That
change is done separately however.

TODO:
- Update commit reference above
- reflow parts of storage/buffer/README that I didn't reindent to make the
  diff more readable

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf%40gcnactj4z56m
---
 src/include/storage/bufmgr.h                |   4 +
 src/backend/access/gist/gistget.c           |  19 +-
 src/backend/access/hash/hashutil.c          |  10 +-
 src/backend/access/heap/heapam_visibility.c | 125 ++++++--
 src/backend/access/nbtree/nbtinsert.c       |  28 +-
 src/backend/access/nbtree/nbtutils.c        |  16 +-
 src/backend/storage/buffer/README           |  44 ++-
 src/backend/storage/buffer/bufmgr.c         | 307 ++++++++++++++++----
 src/backend/storage/freespace/freespace.c   |  14 +-
 src/backend/storage/freespace/fsmpage.c     |  11 +-
 src/tools/pgindent/typedefs.list            |   1 +
 11 files changed, 459 insertions(+), 120 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index df170fe9553..6e7fec7e723 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -314,6 +314,10 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
+extern bool BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer);
+extern bool BufferBeginSetHintBits(Buffer buffer);
+extern void BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std);
+
 extern void UnlockBuffers(void);
 extern void UnlockBuffer(Buffer buffer);
 extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
diff --git a/src/backend/access/gist/gistget.c b/src/backend/access/gist/gistget.c
index 9ba45acfff3..956ece6bed5 100644
--- a/src/backend/access/gist/gistget.c
+++ b/src/backend/access/gist/gistget.c
@@ -63,11 +63,7 @@ gistkillitems(IndexScanDesc scan)
 	 * safe.
 	 */
 	if (BufferGetLSNAtomic(buffer) != so->curPageLSN)
-	{
-		UnlockReleaseBuffer(buffer);
-		so->numKilled = 0;		/* reset counter */
-		return;
-	}
+		goto unlock;
 
 	Assert(GistPageIsLeaf(page));
 
@@ -77,6 +73,16 @@ gistkillitems(IndexScanDesc scan)
 	 */
 	for (i = 0; i < so->numKilled; i++)
 	{
+		if (!killedsomething)
+		{
+			/*
+			 * Use hint bit infrastructure to be allowed to modify the page
+			 * without holding an exclusive lock.
+			 */
+			if (!BufferBeginSetHintBits(buffer))
+				goto unlock;
+		}
+
 		offnum = so->killedItems[i];
 		iid = PageGetItemId(page, offnum);
 		ItemIdMarkDead(iid);
@@ -86,9 +92,10 @@ gistkillitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		GistMarkPageHasGarbage(page);
-		MarkBufferDirtyHint(buffer, true);
+		BufferFinishSetHintBits(buffer, true, true);
 	}
 
+unlock:
 	UnlockReleaseBuffer(buffer);
 
 	/*
diff --git a/src/backend/access/hash/hashutil.c b/src/backend/access/hash/hashutil.c
index f41233fcd07..d1d603770b2 100644
--- a/src/backend/access/hash/hashutil.c
+++ b/src/backend/access/hash/hashutil.c
@@ -593,6 +593,13 @@ _hash_kill_items(IndexScanDesc scan)
 
 			if (ItemPointerEquals(&ituple->t_tid, &currItem->heapTid))
 			{
+				/*
+				 * Use hint bit infrastructure to be allowed to modify the
+				 * page without holding an exclusive lock.
+				 */
+				if (!BufferBeginSetHintBits(so->currPos.buf))
+					goto unlock_page;
+
 				/* found the item */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -610,9 +617,10 @@ _hash_kill_items(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->hasho_flag |= LH_PAGE_HAS_DEAD_TUPLES;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(so->currPos.buf, true, true);
 	}
 
+unlock_page:
 	if (so->hashso_bucket_buf == so->currPos.buf ||
 		havePin)
 		LockBuffer(so->currPos.buf, BUFFER_LOCK_UNLOCK);
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 04284c4b2eb..b932139bf55 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -80,10 +80,38 @@
 
 
 /*
- * SetHintBits()
+ * To be allowed to set hint bits, SetHintBits() needs to call
+ * BufferBeginSetHintBits(). However, that's not free, and some callsites call
+ * SetHintBits() on many tuples in a row. For those it makes sense to amortize
+ * the cost of BufferBeginSetHintBits(). Additionally it's desirable to defer
+ * the cost of BufferBeginSetHintBits() until a hint bit needs to actually be
+ * set. This enum serves as the necessary state space passed to
+ * SetHintbitsExt().
+ */
+typedef enum SetHintBitsState
+{
+	/* not yet checked if hint bits may be set */
+	SHB_INITIAL,
+	/* failed to get permission to set hint bits, don't check again */
+	SHB_DISABLED,
+	/* allowed to set hint bits */
+	SHB_ENABLED,
+} SetHintBitsState;
+
+/*
+ * SetHintBitsExt()
  *
  * Set commit/abort hint bits on a tuple, if appropriate at this time.
  *
+ * To be allowed to set a hint bit on a tuple, the page must not be undergoing
+ * IO at this time (otherwise we e.g. could corrupt PG's page checksum or even
+ * the filesystem's, as is known to happen with btrfs).
+ *
+ * The right to set a hint bit can be acquired on a page level with
+ * BufferBeginSetHintBits(). Only a single backend gets the right to set hint
+ * bits at a time.  Alternatively, if called with a NULL SetHintBitsState*,
+ * hint bits are set with BufferSetHintBits16().
+ *
  * It is only safe to set a transaction-committed hint bit if we know the
  * transaction's commit record is guaranteed to be flushed to disk before the
  * buffer, or if the table is temporary or unlogged and will be obliterated by
@@ -111,24 +139,69 @@
  * InvalidTransactionId if no check is needed.
  */
 static inline void
-SetHintBits(HeapTupleHeader tuple, Buffer buffer,
-			uint16 infomask, TransactionId xid)
+SetHintBitsExt(HeapTupleHeader tuple, Buffer buffer,
+			   uint16 infomask, TransactionId xid, SetHintBitsState *state)
 {
+	/*
+	 * In batched mode and we previously did not get permission to set hint
+	 * bits. Don't try again, in all likelihood IO is still going on.
+	 */
+	if (state && *state == SHB_DISABLED)
+		return;
+
 	if (TransactionIdIsValid(xid))
 	{
-		/* NB: xid must be known committed here! */
-		XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+		if (BufferIsPermanent(buffer))
+		{
+			/* NB: xid must be known committed here! */
+			XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+
+			if (XLogNeedsFlush(commitLSN) &&
+				BufferGetLSNAtomic(buffer) < commitLSN)
+			{
+				/* not flushed and no LSN interlock, so don't set hint */
+				return;
+			}
+		}
+	}
+
+	/*
+	 * If we're not operating in batch mode, use BufferSetHintBits16() to mark
+	 * the page dirty, that's cheaper than
+	 * BufferBeginSetHintBits()/BufferFinishSetHintBits(). That's important
+	 * for cases where we set a lot of hint bits on a page individually.
+	 */
+	if (!state)
+	{
+		BufferSetHintBits16(&tuple->t_infomask,
+							tuple->t_infomask | infomask, buffer);
+		return;
+	}
 
-		if (BufferIsPermanent(buffer) && XLogNeedsFlush(commitLSN) &&
-			BufferGetLSNAtomic(buffer) < commitLSN)
+	if (*state == SHB_INITIAL)
+	{
+		if (!BufferBeginSetHintBits(buffer))
 		{
-			/* not flushed and no LSN interlock, so don't set hint */
+			*state = SHB_DISABLED;
 			return;
 		}
+
+		if (state)
+			*state = SHB_ENABLED;
+
 	}
-
 	tuple->t_infomask |= infomask;
-	MarkBufferDirtyHint(buffer, true);
+}
+
+/*
+ * Simple wrapper around SetHintBitExt(), use when operating on a single
+ * tuple.
+ */
+static inline void
+SetHintBits(HeapTupleHeader tuple, Buffer buffer,
+			uint16 infomask, TransactionId xid)
+{
+	SetHintBitsExt(tuple, buffer, infomask, xid, NULL);
 }
 
 /*
@@ -864,9 +937,9 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
  * inserting/deleting transaction was still running --- which was more cycles
  * and more contention on ProcArrayLock.
  */
-static bool
+static inline bool
 HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
-					   Buffer buffer)
+					   Buffer buffer, SetHintBitsState *state)
 {
 	HeapTupleHeader tuple = htup->t_data;
 
@@ -921,8 +994,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 			if (!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmax(tuple)))
 			{
 				/* deleting subtransaction must have aborted */
-				SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-							InvalidTransactionId);
+				SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+							   InvalidTransactionId, state);
 				return true;
 			}
 
@@ -934,13 +1007,13 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		else if (XidInMVCCSnapshot(HeapTupleHeaderGetRawXmin(tuple), snapshot))
 			return false;
 		else if (TransactionIdDidCommit(HeapTupleHeaderGetRawXmin(tuple)))
-			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-						HeapTupleHeaderGetRawXmin(tuple));
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_COMMITTED,
+						   HeapTupleHeaderGetRawXmin(tuple), state);
 		else
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_INVALID,
+						   InvalidTransactionId, state);
 			return false;
 		}
 	}
@@ -1003,14 +1076,14 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		if (!TransactionIdDidCommit(HeapTupleHeaderGetRawXmax(tuple)))
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+						   InvalidTransactionId, state);
 			return true;
 		}
 
 		/* xmax transaction committed */
-		SetHintBits(tuple, buffer, HEAP_XMAX_COMMITTED,
-					HeapTupleHeaderGetRawXmax(tuple));
+		SetHintBitsExt(tuple, buffer, HEAP_XMAX_COMMITTED,
+					   HeapTupleHeaderGetRawXmax(tuple), state);
 	}
 	else
 	{
@@ -1619,6 +1692,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 							OffsetNumber *vistuples_dense)
 {
 	int			nvis = 0;
+	SetHintBitsState state = SHB_INITIAL;
 
 	Assert(IsMVCCSnapshot(snapshot));
 
@@ -1627,7 +1701,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		bool		valid;
 		HeapTuple	tup = &batchmvcc->tuples[i];
 
-		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer, &state);
 		batchmvcc->visible[i] = valid;
 
 		if (likely(valid))
@@ -1637,6 +1711,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		}
 	}
 
+	if (state == SHB_ENABLED)
+		BufferFinishSetHintBits(buffer, true, true);
+
 	return nvis;
 }
 
@@ -1656,7 +1733,7 @@ HeapTupleSatisfiesVisibility(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 	switch (snapshot->snapshot_type)
 	{
 		case SNAPSHOT_MVCC:
-			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer);
+			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer, NULL);
 		case SNAPSHOT_SELF:
 			return HeapTupleSatisfiesSelf(htup, snapshot, buffer);
 		case SNAPSHOT_ANY:
diff --git a/src/backend/access/nbtree/nbtinsert.c b/src/backend/access/nbtree/nbtinsert.c
index 031eb76ba8c..f7cc56afc88 100644
--- a/src/backend/access/nbtree/nbtinsert.c
+++ b/src/backend/access/nbtree/nbtinsert.c
@@ -681,20 +681,28 @@ _bt_check_unique(Relation rel, BTInsertState insertstate, Relation heapRel,
 				{
 					/*
 					 * The conflicting tuple (or all HOT chains pointed to by
-					 * all posting list TIDs) is dead to everyone, so mark the
-					 * index entry killed.
+					 * all posting list TIDs) is dead to everyone, so try to
+					 * mark the index entry killed. It's ok if we're not
+					 * allowed to, this isn't required for correctness.
 					 */
-					ItemIdMarkDead(curitemid);
-					opaque->btpo_flags |= BTP_HAS_GARBAGE;
+					Buffer		buf;
 
-					/*
-					 * Mark buffer with a dirty hint, since state is not
-					 * crucial. Be sure to mark the proper buffer dirty.
-					 */
+					/* Be sure to operate on the proper buffer */
 					if (nbuf != InvalidBuffer)
-						MarkBufferDirtyHint(nbuf, true);
+						buf = nbuf;
 					else
-						MarkBufferDirtyHint(insertstate->buf, true);
+						buf = insertstate->buf;
+
+					/*
+					 * Can't use BufferSetHintBits16() here as we update two
+					 * different locations.
+					 */
+					if (BufferBeginSetHintBits(buf))
+					{
+						ItemIdMarkDead(curitemid);
+						opaque->btpo_flags |= BTP_HAS_GARBAGE;
+						BufferFinishSetHintBits(buf, true, true);
+					}
 				}
 
 				/*
diff --git a/src/backend/access/nbtree/nbtutils.c b/src/backend/access/nbtree/nbtutils.c
index a451d48e11e..5fb7c421adb 100644
--- a/src/backend/access/nbtree/nbtutils.c
+++ b/src/backend/access/nbtree/nbtutils.c
@@ -357,10 +357,19 @@ _bt_killitems(IndexScanDesc scan)
 			 * it's possible that multiple processes attempt to do this
 			 * simultaneously, leading to multiple full-page images being sent
 			 * to WAL (if wal_log_hints or data checksums are enabled), which
-			 * is undesirable.
+			 * is undesirable.  We need to use the hint bit infrastructure to
+			 * update the page while just holding a share lock.
 			 */
 			if (killtuple && !ItemIdIsDead(iid))
 			{
+				/*
+				 * If we're not able to set hint bits, there's no point
+				 * continuing.
+				 */
+				if (!killedsomething &&
+					!BufferBeginSetHintBits(buf))
+					goto unlock_page;
+
 				/* found the item/all posting list items */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -371,8 +380,6 @@ _bt_killitems(IndexScanDesc scan)
 	}
 
 	/*
-	 * Since this can be redone later if needed, mark as dirty hint.
-	 *
 	 * Whenever we mark anything LP_DEAD, we also set the page's
 	 * BTP_HAS_GARBAGE flag, which is likewise just a hint.  (Note that we
 	 * only rely on the page-level flag in !heapkeyspace indexes.)
@@ -380,9 +387,10 @@ _bt_killitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->btpo_flags |= BTP_HAS_GARBAGE;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(buf, true, true);
 	}
 
+unlock_page:
 	if (!so->dropPin)
 		_bt_unlockbuf(rel, buf);
 	else
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index 119f31b5d65..5e6735edd23 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -25,21 +25,26 @@ that might need to do such a wait is instead handled by waiting to obtain
 the relation-level lock, which is why you'd better hold one first.)  Pins
 may not be held across transaction boundaries, however.
 
-Buffer content locks: there are two kinds of buffer lock, shared and exclusive,
-which act just as you'd expect: multiple backends can hold shared locks on
-the same buffer, but an exclusive lock prevents anyone else from holding
-either shared or exclusive lock.  (These can alternatively be called READ
-and WRITE locks.)  These locks are intended to be short-term: they should not
-be held for long.  Buffer locks are acquired and released by LockBuffer().
-It will *not* work for a single backend to try to acquire multiple locks on
-the same buffer.  One must pin a buffer before trying to lock it.
+Buffer content locks: there three kinds of buffer lock, shared,
+share-exclusive and exclusive:
+a) multiple backends can hold shared locks on the same buffer
+   (alternatively called a READ lock)
+b) one backend can hold an share-exclusive lock on a buffer while multiple
+   backends can hold a share lock
+c) an exclusive lock prevents anyone else from holding shared, share-exclusive
+   or exclusive lock.
+   (alternatively called a WRITE lock)
+
+These locks are intended to be short-term: they should not be held for long.
+Buffer locks are acquired and released by LockBuffer().  It will *not* work
+for a single backend to try to acquire multiple locks on the same buffer.  One
+must pin a buffer before trying to lock it.
 
 Buffer access rules:
 
-1. To scan a page for tuples, one must hold a pin and either shared or
-exclusive content lock.  To examine the commit status (XIDs and status bits)
-of a tuple in a shared buffer, one must likewise hold a pin and either shared
-or exclusive lock.
+1. To scan a page for tuples, one must hold a pin and at least a share lock.
+To examine the commit status (XIDs and status bits) of a tuple in a shared
+buffer, one must likewise hold a pin and at least a share lock.
 
 2. Once one has determined that a tuple is interesting (visible to the
 current transaction) one may drop the content lock, yet continue to access
@@ -55,8 +60,14 @@ one must hold a pin and an exclusive content lock on the containing buffer.
 This ensures that no one else might see a partially-updated state of the
 tuple while they are doing visibility checks.
 
-4. It is considered OK to update tuple commit status bits (ie, OR the
-values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
+4. Non-critical information on a page ("hint bits") may be modified while
+holding only a share-exclusive lock and pin on the page. To do so in cases
+where only a share lock is already held, use BufferBeginSetHintBits() &
+BufferFinishSetHintBits() (if multiple hint bits are to be set) or
+BufferSetHintBits16() (if a single hit bit is set).
+
+E.g. for heapam, a share-exclusive lock allows to update tuple commit status
+bits (ie, OR the values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
 HEAP_XMAX_INVALID into t_infomask) while holding only a shared lock and
 pin on a buffer.  This is OK because another backend looking at the tuple
 at about the same time would OR the same bits into the field, so there
@@ -80,7 +91,6 @@ buffer (increment the refcount) while one is performing the cleanup, but
 it won't be able to actually examine the page until it acquires shared
 or exclusive content lock.
 
-
 Obtaining the lock needed under rule #5 is done by the bufmgr routines
 LockBufferForCleanup() or ConditionalLockBufferForCleanup().  They first get
 an exclusive lock and then check to see if the shared pin count is currently
@@ -96,6 +106,10 @@ VACUUM's use, since we don't allow multiple VACUUMs concurrently on a single
 relation anyway.  Anyone wishing to obtain a cleanup lock outside of recovery
 or a VACUUM must use the conditional variant of the function.
 
+6. To write out a buffer, a share-exclusive lock needs to be held. This
+prevents the buffer from being modified while written out, which could corrupt
+checksums and cause issues on the OS or device level when direct-IO is used.
+
 
 Buffer Manager's Internal Locking
 ---------------------------------
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index a9cafdc8ff1..7c90d3413ac 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2468,9 +2468,8 @@ again:
 	/*
 	 * If the buffer was dirty, try to write it out.  There is a race
 	 * condition here, in that someone might dirty it after we released the
-	 * buffer header lock above, or even while we are writing it out (since
-	 * our share-lock won't prevent hint-bit updates).  We will recheck the
-	 * dirty bit after re-locking the buffer header.
+	 * buffer header lock above.  We will recheck the dirty bit after
+	 * re-locking the buffer header.
 	 */
 	if (buf_state & BM_DIRTY)
 	{
@@ -2478,12 +2477,12 @@ again:
 		Assert(buf_state & BM_VALID);
 
 		/*
-		 * We need a share-lock on the buffer contents to write it out (else
+		 * We need a share-exclusive lock on the buffer contents to write it out (else
 		 * we might write invalid data, eg because someone else is compacting
 		 * the page contents while we write).  We must use a conditional lock
 		 * acquisition here to avoid deadlock.  Even though the buffer was not
 		 * pinned (and therefore surely not locked) when StrategyGetBuffer
-		 * returned it, someone else could have pinned and exclusive-locked it
+		 * returned it, someone else could have pinned and (share-)exclusive-locked it
 		 * by the time we get here. If we try to get the lock unconditionally,
 		 * we'd block waiting for them; if they later block waiting for us,
 		 * deadlock ensues. (This has been observed to happen when two
@@ -2491,7 +2490,7 @@ again:
 		 * one just happens to be trying to split the page the first one got
 		 * from StrategyGetBuffer.)
 		 */
-		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -4060,8 +4059,8 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	}
 
 	/*
-	 * Pin it, share-lock it, write it.  (FlushBuffer will do nothing if the
-	 * buffer is clean by the time we've locked it.)
+	 * Pin it, share-exclusive-lock it, write it.  (FlushBuffer will do
+	 * nothing if the buffer is clean by the time we've locked it.)
 	 */
 	PinBuffer_Locked(bufHdr);
 
@@ -4391,11 +4390,8 @@ BufferGetTag(Buffer buffer, RelFileLocator *rlocator, ForkNumber *forknum,
  * However, we will need to force the changes to disk via fsync before
  * we can checkpoint WAL.
  *
- * The caller must hold a pin on the buffer and have share-locked the
- * buffer contents.  (Note: a share-lock does not prevent updates of
- * hint bits in the buffer, so the page could change while the write
- * is in progress, but we assume that that will not invalidate the data
- * written.)
+ * The caller must hold a pin on the buffer and have
+ * (share-)exclusively-locked the buffer contents.
  *
  * If the caller has an smgr reference for the buffer's relation, pass it
  * as the second parameter.  If not, pass NULL.
@@ -4411,6 +4407,9 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	char	   *bufToWrite;
 	uint64		buf_state;
 
+	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(buf, BUFFER_LOCK_SHARE_EXCLUSIVE));
+
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
 	 * someone else flushed the buffer before we could, so we need not do
@@ -4543,7 +4542,7 @@ FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(buf);
 
-	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 	BufferLockUnlock(buffer, buf);
 }
@@ -5462,8 +5461,8 @@ FlushDatabaseBuffers(Oid dbid)
 }
 
 /*
- * Flush a previously, shared or exclusively, locked and pinned buffer to the
- * OS.
+ * Flush a previously, share-exclusively or exclusively, locked and pinned
+ * buffer to the OS.
  */
 void
 FlushOneBuffer(Buffer buffer)
@@ -5536,39 +5535,23 @@ IncrBufferRefCount(Buffer buffer)
 }
 
 /*
- * MarkBufferDirtyHint
+ * Shared-buffer only helper for MarkBufferDirtyHint() and
+ * BufferSetHintBits16().
  *
- *	Mark a buffer dirty for non-critical changes.
- *
- * This is essentially the same as MarkBufferDirty, except:
- *
- * 1. The caller does not write WAL; so if checksums are enabled, we may need
- *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
- * 2. The caller might have only share-lock instead of exclusive-lock on the
- *	  buffer's content lock.
- * 3. This function does not guarantee that the buffer is always marked dirty
- *	  (due to a race condition), so it cannot be used for important changes.
+ * This is separated out because it turns out that the repeated checks for
+ * local buffers, repeated GetBufferDescriptor() and repeated reading of the
+ * buffer's state sufficiently hurts the performance of BufferSetHintBits16().
  */
-void
-MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+static inline void
+MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate, bool buffer_std)
 {
-	BufferDesc *bufHdr;
 	Page		page = BufferGetPage(buffer);
 
-	if (!BufferIsValid(buffer))
-		elog(ERROR, "bad buffer ID: %d", buffer);
-
-	if (BufferIsLocal(buffer))
-	{
-		MarkLocalBufferDirty(buffer);
-		return;
-	}
-
-	bufHdr = GetBufferDescriptor(buffer - 1);
-
 	Assert(GetPrivateRefCount(buffer) > 0);
-	/* here, either share or exclusive lock is OK */
-	Assert(BufferIsLockedByMe(buffer));
+
+	/* here, either share-exclusive or exclusive lock is OK */
+	Assert(BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_SHARE_EXCLUSIVE));
 
 	/*
 	 * This routine might get called many times on the same page, if we are
@@ -5581,8 +5564,8 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
-		(BM_DIRTY | BM_JUST_DIRTIED))
+	if (unlikely((lockstate & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+				 (BM_DIRTY | BM_JUST_DIRTIED)))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
@@ -5598,8 +5581,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * We don't check full_page_writes here because that logic is included
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
-		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
+		if (XLogHintBitIsNeeded() && (lockstate & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5651,13 +5633,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 			dirtied = true;		/* Means "will be dirtied by this action" */
 
 			/*
-			 * Set the page LSN if we wrote a backup block. We aren't supposed
-			 * to set this when only holding a share lock but as long as we
-			 * serialise it somehow we're OK. We choose to set LSN while
-			 * holding the buffer header lock, which causes any reader of an
-			 * LSN who holds only a share lock to also obtain a buffer header
-			 * lock before using PageGetLSN(), which is enforced in
-			 * BufferGetLSNAtomic().
+			 * Set the page LSN if we wrote a backup block. To allow backends
+			 * that only hold a share lock on the buffer to read the LSN in a
+			 * tear-free manner, we set the page LSN while holding the buffer
+			 * header lock. This allows any reader of an LSN who holds only a
+			 * share lock to also obtain a buffer header lock before using
+			 * PageGetLSN() to read the LSN in a tear free way. This is done
+			 * in BufferGetLSNAtomic().
 			 *
 			 * If checksums are enabled, you might think we should reset the
 			 * checksum here. That will happen when the page is written
@@ -5683,6 +5665,41 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	}
 }
 
+/*
+ * MarkBufferDirtyHint
+ *
+ *	Mark a buffer dirty for non-critical changes.
+ *
+ * This is essentially the same as MarkBufferDirty, except:
+ *
+ * 1. The caller does not write WAL; so if checksums are enabled, we may need
+ *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
+ * 2. The caller might have only share-exclusive-lock instead of
+ *	  exclusive-lock on the buffer's content lock.
+ * 3. This function does not guarantee that the buffer is always marked dirty
+ *	  (due to a race condition), so it cannot be used for important changes.
+ */
+inline void
+MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+{
+	BufferDesc *bufHdr;
+
+	bufHdr = GetBufferDescriptor(buffer - 1);
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		MarkLocalBufferDirty(buffer);
+		return;
+	}
+
+	MarkSharedBufferDirtyHint(buffer, bufHdr,
+							  pg_atomic_read_u64(&bufHdr->state),
+							  buffer_std);
+}
+
 /*
  * Release buffer content locks for shared buffers.
  *
@@ -6778,6 +6795,188 @@ IsBufferCleanupOK(Buffer buffer)
 	return false;
 }
 
+/*
+ * Helper for BufferBeginSetHintBits() and BufferSetHintBits16().
+ *
+ * This checks if the current lock mode already suffices to allow hint bits
+ * being set and, if not, whether the current lock can be upgraded.
+ */
+static inline bool
+SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
+{
+	uint64		old_state;
+	PrivateRefCountEntry *ref;
+	BufferLockMode mode;
+
+	ref = GetPrivateRefCountEntry(buffer, true);
+
+	if (ref == NULL)
+		elog(ERROR, "lock is not held");
+
+	mode = ref->data.lockmode;
+	if (mode == BUFFER_LOCK_UNLOCK)
+		elog(ERROR, "buffer is not locked");
+
+	/*
+	 * Already am holding a sufficient lock level.
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE || mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+	{
+		*lockstate = pg_atomic_read_u64(&buf_hdr->state);
+		return true;
+	}
+
+	/*
+	 * Only holding a share lock right now, try to upgrade to SHARE_EXCLUSIVE.
+	 */
+	Assert(mode == BUFFER_LOCK_SHARE);
+
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+	while (true)
+	{
+		uint64		desired_state;
+
+		desired_state = old_state;
+
+		/*
+		 * Can't upgrade if somebody else holds the lock in exlusive or
+		 * share-exclusive mode.
+		 */
+		if (unlikely((old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) != 0))
+		{
+			return false;
+		}
+
+		/* currently held lock state */
+		desired_state -= BM_LOCK_VAL_SHARED;
+
+		/* new lock level */
+		desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			ref->data.lockmode = BUFFER_LOCK_SHARE_EXCLUSIVE;
+			*lockstate = desired_state;
+
+			return true;
+		}
+	}
+
+}
+
+/*
+ * Try to acquire the right to set hint bits on the buffer.
+ *
+ * To be allowed to set hint bits, this backend needs to hold either a
+ * share-exclusive or an exclusive lock. In case this backend only holds a
+ * share lock, this function will try to upgrade the lock to
+ * share-exclusive. The caller is only allowed to set hint bits if true is
+ * returned.
+ *
+ * Once BufferBeginSetHintBits() has returned true, hint bits may be set
+ * without further calls to BufferBeginSetHintBits(), until the buffer is
+ * unlocked.
+ *
+ *
+ * Requiring a share-exclusive lock to set hint bits prevents setting hint
+ * bits on buffers that are currently being written out, which could corrupt
+ * the checksum on the page. Flushing buffers also requires a share-exclusive
+ * lock.
+ *
+ * Due to a lock >= share-exclusive being required to set hint bits, only one
+ * backend can set hint bits at a time. To allow multiple backends setting
+ * hint bits would require more complicated locking, for setting hint bits
+ * we'd need to store the count of backends currently setting hint bits and we
+ * would need another lock-level for I/O conflicting with the lock-level to
+ * set hint bits. Given that the share-exclusive lock for setting hint bits is
+ * only held for a short time, that often backends would just set the same
+ * hint bits and that the cost of occasionally not setting hint bits in hotly
+ * accessed pages is fairly low, this seems like an acceptable tradeoff.
+ */
+bool
+BufferBeginSetHintBits(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+
+	if (BufferIsLocal(buffer))
+	{
+		/*
+		 * TODO: will need to check for write IO once that's done
+		 * asynchronously.
+		 */
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	return SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate);
+}
+
+/*
+ * End a phase of setting hint bits on this buffer, started with
+ * BufferBeginSetHintBits().
+ *
+ * This would strictly speaking not be required (i.e. the caller could do
+ * MarkBufferDirtyHint() if so desired), but allows us to perform some sanity
+ * checks.
+ */
+void
+BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std)
+{
+	if (!BufferIsLocal(buffer))
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+
+	if (mark_dirty)
+		MarkBufferDirtyHint(buffer, buffer_std);
+}
+
+/*
+ * Ty to set a single hint bit in a buffer.
+ *
+ * This is a bit faster than BufferBeginSetHintBits() /
+ * BufferFinishSetHintBits() when setting a single hint bit, but slower than
+ * the former when setting several hint bits.
+ */
+bool
+BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+#ifdef USE_ASSERT_CHECKING
+	char	   *page;
+
+	/* verify that the address is on the page */
+	page = BufferGetPage(buffer);
+	Assert((char *) ptr >= page && (char *) ptr < (page + BLCKSZ));
+#endif
+
+	if (BufferIsLocal(buffer))
+	{
+		*ptr = val;
+
+		MarkLocalBufferDirty(buffer);
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	if (SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate))
+	{
+		*ptr = val;
+
+		MarkSharedBufferDirtyHint(buffer, buf_hdr, lockstate, true);
+
+		return true;
+	}
+
+	return false;
+}
+
 
 /*
  *	Functions for buffer I/O handling
diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index 48ac15d3487..bd4a2cff3a4 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -904,13 +904,17 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 	max_avail = fsm_get_max_avail(page);
 
 	/*
-	 * Reset the next slot pointer. This encourages the use of low-numbered
-	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation. We don't bother with marking the page dirty if it wasn't
-	 * already, since this is just a hint.
+	 * Try to reset the next slot pointer. This encourages the use of
+	 * low-numbered pages, increasing the chances that a later vacuum can
+	 * truncate the relation. We don't bother with marking the page dirty if
+	 * it wasn't already, since this is just a hint.
 	 */
 	LockBuffer(buf, BUFFER_LOCK_SHARE);
-	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	if (BufferBeginSetHintBits(buf))
+	{
+		((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+		BufferFinishSetHintBits(buf, false, false);
+	}
 	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 	ReleaseBuffer(buf);
diff --git a/src/backend/storage/freespace/fsmpage.c b/src/backend/storage/freespace/fsmpage.c
index 66a5c80b5a6..a59696b6484 100644
--- a/src/backend/storage/freespace/fsmpage.c
+++ b/src/backend/storage/freespace/fsmpage.c
@@ -298,9 +298,18 @@ restart:
 	 * lock and get a garbled next pointer every now and then, than take the
 	 * concurrency hit of an exclusive lock.
 	 *
+	 * Without an exclusive lock, we need to use the hint bit infrastructure
+	 * to be allowed to modify the page.
+	 *
 	 * Wrap-around is handled at the beginning of this function.
 	 */
-	fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+	if (exclusive_lock_held || BufferBeginSetHintBits(buf))
+	{
+		fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+
+		if (!exclusive_lock_held)
+			BufferFinishSetHintBits(buf, false, true);
+	}
 
 	return slot;
 }
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 2ffdf364386..55de9f66de4 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2750,6 +2750,7 @@ SetConstraintStateData
 SetConstraintTriggerData
 SetExprState
 SetFunctionReturnMode
+SetHintBitsState
 SetOp
 SetOpCmd
 SetOpPath
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v8-0009-WIP-Make-UnlockReleaseBuffer-more-efficient.patch (3.5K, ../../ossv2eistssmubfsir6xjll76tynvxv5lup4zkrfzjkud7fycw@rf5vii6l6cha/10-v8-0009-WIP-Make-UnlockReleaseBuffer-more-efficient.patch)
  download | inline diff:
From d23ab6033ebc36fe57c20c333f3be3dded4cf374 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 19 Nov 2025 15:32:20 -0500
Subject: [PATCH v8 09/10] WIP: Make UnlockReleaseBuffer() more efficient

Now that the buffer content lock is implemented as part of BufferDesc.state,
releasing the lock and unpinning the buffer can be implemented as a single
atomic operation.

Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/nbtree/nbtpage.c | 22 +++++++++++-
 src/backend/storage/buffer/bufmgr.c | 52 ++++++++++++++++++++++++++++-
 2 files changed, 72 insertions(+), 2 deletions(-)

diff --git a/src/backend/access/nbtree/nbtpage.c b/src/backend/access/nbtree/nbtpage.c
index cfb07b2bca9..9f604751dc7 100644
--- a/src/backend/access/nbtree/nbtpage.c
+++ b/src/backend/access/nbtree/nbtpage.c
@@ -1007,11 +1007,18 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 
 	Assert(BlockNumberIsValid(blkno));
 	if (BufferIsValid(obuf))
+	{
+		_bt_relbuf(rel, obuf);
+#if 0
+		Assert(BufferGetBlockNumber(obuf) != blkno);
 		_bt_unlockbuf(rel, obuf);
-	buf = ReleaseAndReadBuffer(obuf, rel, blkno);
+#endif
+	}
+	buf = ReadBuffer(rel, blkno);
 	_bt_lockbuf(rel, buf, access);
 
 	_bt_checkpage(rel, buf);
+
 	return buf;
 }
 
@@ -1023,8 +1030,21 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 void
 _bt_relbuf(Relation rel, Buffer buf)
 {
+#if 0
 	_bt_unlockbuf(rel, buf);
 	ReleaseBuffer(buf);
+#else
+	/*
+	 * Buffer is pinned and locked, which means that it is expected to be
+	 * defined and addressable.  Check that proactively.
+	 */
+	VALGRIND_CHECK_MEM_IS_DEFINED(BufferGetPage(buf), BLCKSZ);
+
+	UnlockReleaseBuffer(buf);
+
+	if (!RelationUsesLocalBuffers(rel))
+		VALGRIND_MAKE_MEM_NOACCESS(BufferGetPage(buf), BLCKSZ);
+#endif
 }
 
 /*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 7c90d3413ac..7d8799bd9a1 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5499,13 +5499,63 @@ ReleaseBuffer(Buffer buffer)
 /*
  * UnlockReleaseBuffer -- release the content lock and pin on a buffer
  *
- * This is just a shorthand for a common combination.
+ * This is just a, more efficient, shorthand for a common combination.
  */
 void
 UnlockReleaseBuffer(Buffer buffer)
 {
+#if 1
+	int			mode;
+	BufferDesc *buf;
+	PrivateRefCountEntry *ref;
+	uint64		sub;
+	uint64		lockstate;
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		UnpinLocalBuffer(buffer);
+		return;
+	}
+
+	ResourceOwnerForgetBuffer(CurrentResourceOwner, buffer);
+
+	buf = GetBufferDescriptor(buffer - 1);
+
+	mode = BufferLockDisownInternal(buffer, buf);
+
+	/* compute state modification for lock release */
+	sub = BufferLockReleaseSub(mode);
+
+	/* compute state modification for pin release */
+	ref = GetPrivateRefCountEntry(buffer, false);
+	Assert(ref != NULL);
+	Assert(ref->data.refcount > 0);
+	ref->data.refcount--;
+
+	if (ref->data.refcount == 0)
+	{
+		sub |= BUF_REFCOUNT_ONE;
+		ForgetPrivateRefCountEntry(ref);
+	}
+
+	/* perform the lock and pin release in one atomic op */
+	lockstate = pg_atomic_sub_fetch_u64(&buf->state, sub);
+
+	/* wake up waiters etc */
+	BufferLockProcessRelease(buf, mode, lockstate);
+
+	if (lockstate & BM_PIN_COUNT_WAITER)
+		WakePinCountWaiter(buf);
+
+	RESUME_INTERRUPTS();
+
+#else
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 	ReleaseBuffer(buffer);
+#endif
 }
 
 /*
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v8-0010-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch (11.6K, ../../ossv2eistssmubfsir6xjll76tynvxv5lup4zkrfzjkud7fycw@rf5vii6l6cha/11-v8-0010-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch)
  download | inline diff:
From 92b2b6a5e18bfc22f104538fee08346adef0f0d1 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 17 Oct 2024 14:14:35 -0400
Subject: [PATCH v8 10/10] WIP: bufmgr: Don't copy pages while writing out

After the series of preceding commits introducing and using
BufferBeginSetHintBits()/BufferSetHintBits16() hint bits are not set
anymore while IO is going on. Therefore we do not need to copy pages while
they are being written out anymore.

TODO: Update comments

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/bufpage.h           |  3 +-
 src/backend/access/hash/hashpage.c      |  2 +-
 src/backend/access/transam/xloginsert.c | 43 ++++++----------------
 src/backend/storage/buffer/bufmgr.c     | 21 +++++------
 src/backend/storage/buffer/localbuf.c   |  2 +-
 src/backend/storage/page/bufpage.c      | 48 ++++---------------------
 src/backend/storage/smgr/bulk_write.c   |  2 +-
 src/test/modules/test_aio/test_aio.c    |  2 +-
 8 files changed, 33 insertions(+), 90 deletions(-)

diff --git a/src/include/storage/bufpage.h b/src/include/storage/bufpage.h
index abc2cf2a020..f8f621446c4 100644
--- a/src/include/storage/bufpage.h
+++ b/src/include/storage/bufpage.h
@@ -504,7 +504,6 @@ extern void PageIndexMultiDelete(Page page, OffsetNumber *itemnos, int nitems);
 extern void PageIndexTupleDeleteNoCompact(Page page, OffsetNumber offnum);
 extern bool PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 									const void *newtup, Size newsize);
-extern char *PageSetChecksumCopy(Page page, BlockNumber blkno);
-extern void PageSetChecksumInplace(Page page, BlockNumber blkno);
+extern void PageSetChecksum(Page page, BlockNumber blkno);
 
 #endif							/* BUFPAGE_H */
diff --git a/src/backend/access/hash/hashpage.c b/src/backend/access/hash/hashpage.c
index b8e5bd005e5..dd17eff59d1 100644
--- a/src/backend/access/hash/hashpage.c
+++ b/src/backend/access/hash/hashpage.c
@@ -1029,7 +1029,7 @@ _hash_alloc_buckets(Relation rel, BlockNumber firstblock, uint32 nblocks)
 					zerobuf.data,
 					true);
 
-	PageSetChecksumInplace(page, lastblock);
+	PageSetChecksum(page, lastblock);
 	smgrextend(RelationGetSmgr(rel), MAIN_FORKNUM, lastblock, zerobuf.data,
 			   false);
 
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index a56d5a55282..0af148e9496 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -261,8 +261,11 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	 */
 #ifdef USE_ASSERT_CHECKING
 	if (!(flags & REGBUF_NO_CHANGE))
-		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) &&
-			   BufferIsDirty(buffer));
+	{
+		Assert(BufferIsDirty(buffer));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE));
+	}
 #endif
 
 	if (block_id >= max_registered_block_id)
@@ -1066,7 +1069,7 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * Write a backup block if needed when we are setting a hint. Note that
  * this may be called for a variety of page types, not just heaps.
  *
- * Callable while holding just share lock on the buffer content.
+ * Callable while holding just share-exclusive lock on the buffer content.
  *
  * We can't use the plain backup block mechanism since that relies on the
  * Buffer being exclusively locked. Since some modifications (setting LSN, hint
@@ -1074,6 +1077,8 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * failures. So instead we copy the page and insert the copied data as normal
  * record data.
  *
+ * FIXME: outdated
+ *
  * We only need to do something if page has not yet been full page written in
  * this checkpoint round. The LSN of the inserted wal record is returned if we
  * had to write, InvalidXLogRecPtr otherwise.
@@ -1102,46 +1107,20 @@ XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
 
 	/*
 	 * We assume page LSN is first data on *every* page that can be passed to
-	 * XLogInsert, whether it has the standard page layout or not. Since we're
-	 * only holding a share-lock on the page, we must take the buffer header
-	 * lock when we look at the LSN.
+	 * XLogInsert, whether it has the standard page layout or not.
 	 */
 	lsn = BufferGetLSNAtomic(buffer);
 
 	if (lsn <= RedoRecPtr)
 	{
-		int			flags = 0;
-		PGAlignedBlock copied_buffer;
-		char	   *origdata = (char *) BufferGetBlock(buffer);
-		RelFileLocator rlocator;
-		ForkNumber	forkno;
-		BlockNumber blkno;
-
-		/*
-		 * Copy buffer so we don't have to worry about concurrent hint bit or
-		 * lsn updates. We assume pd_lower/upper cannot be changed without an
-		 * exclusive lock, so the contents bkp are not racy.
-		 */
-		if (buffer_std)
-		{
-			/* Assume we can omit data between pd_lower and pd_upper */
-			Page		page = BufferGetPage(buffer);
-			uint16		lower = ((PageHeader) page)->pd_lower;
-			uint16		upper = ((PageHeader) page)->pd_upper;
-
-			memcpy(copied_buffer.data, origdata, lower);
-			memcpy(copied_buffer.data + upper, origdata + upper, BLCKSZ - upper);
-		}
-		else
-			memcpy(copied_buffer.data, origdata, BLCKSZ);
+		int			flags = REGBUF_NO_CHANGE;
 
 		XLogBeginInsert();
 
 		if (buffer_std)
 			flags |= REGBUF_STANDARD;
 
-		BufferGetTag(buffer, &rlocator, &forkno, &blkno);
-		XLogRegisterBlock(0, &rlocator, forkno, blkno, copied_buffer.data, flags);
+		XLogRegisterBuffer(0, buffer, flags);
 
 		recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI_FOR_HINT);
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 7d8799bd9a1..1dfef0c69b8 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -4404,7 +4404,6 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	ErrorContextCallback errcallback;
 	instr_time	io_start;
 	Block		bufBlock;
-	char	   *bufToWrite;
 	uint64		buf_state;
 
 	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
@@ -4475,12 +4474,8 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	 */
 	bufBlock = BufHdrGetBlock(buf);
 
-	/*
-	 * Update page checksum if desired.  Since we have only shared lock on the
-	 * buffer, other processes might be updating hint bits in it, so we must
-	 * copy the page to private storage if we do checksumming.
-	 */
-	bufToWrite = PageSetChecksumCopy((Page) bufBlock, buf->tag.blockNum);
+	/* Update page checksum if desired. */
+	PageSetChecksum((Page) bufBlock, buf->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
@@ -4490,7 +4485,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	smgrwrite(reln,
 			  BufTagGetForkNum(&buf->tag),
 			  buf->tag.blockNum,
-			  bufToWrite,
+			  bufBlock,
 			  false);
 
 	/*
@@ -4614,8 +4609,8 @@ BufferIsPermanent(Buffer buffer)
 /*
  * BufferGetLSNAtomic
  *		Retrieves the LSN of the buffer atomically using a buffer header lock.
- *		This is necessary for some callers who may not have an exclusive lock
- *		on the buffer.
+ *		This is necessary for some callers who may not have a (share-)exclusive
+ *		lock on the buffer.
  */
 XLogRecPtr
 BufferGetLSNAtomic(Buffer buffer)
@@ -5667,6 +5662,12 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate, b
 			 * It's possible we may enter here without an xid, so it is
 			 * essential that CreateCheckPoint waits for virtual transactions
 			 * rather than full transactionids.
+			 *
+			 * FIXME: I think we now should simply mark the page dirty before
+			 * WAL logging the hint bit - afaikt it then should work just like
+			 * any other buffer write (due to SyncBuffers()/SyncOneBuffer()
+			 * seeing the dirty bit and trying to lock the page
+			 * share-exclusive, and thus having to wait).
 			 */
 			Assert((MyProc->delayChkptFlags & DELAY_CHKPT_START) == 0);
 			MyProc->delayChkptFlags |= DELAY_CHKPT_START;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index a41a5facd3a..5826d4b54c6 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -199,7 +199,7 @@ FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln)
 		reln = smgropen(BufTagGetRelFileLocator(&bufHdr->tag),
 						MyProcNumber);
 
-	PageSetChecksumInplace(localpage, bufHdr->tag.blockNum);
+	PageSetChecksum(localpage, bufHdr->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index 05376431ef2..64af5bd7b5e 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -1494,51 +1494,15 @@ PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 /*
  * Set checksum for a page in shared buffers.
  *
- * If checksums are disabled, or if the page is not initialized, just return
- * the input.  Otherwise, we must make a copy of the page before calculating
- * the checksum, to prevent concurrent modifications (e.g. setting hint bits)
- * from making the final checksum invalid.  It doesn't matter if we include or
- * exclude hints during the copy, as long as we write a valid page and
- * associated checksum.
+ * If checksums are disabled, or if the page is not initialized, just
+ * return. Otherwise compute and set the checksum.
  *
- * Returns a pointer to the block-sized data that needs to be written. Uses
- * statically-allocated memory, so the caller must immediately write the
- * returned page and not refer to it again.
- */
-char *
-PageSetChecksumCopy(Page page, BlockNumber blkno)
-{
-	static char *pageCopy = NULL;
-
-	/* If we don't need a checksum, just return the passed-in data */
-	if (PageIsNew(page) || !DataChecksumsEnabled())
-		return page;
-
-	/*
-	 * We allocate the copy space once and use it over on each subsequent
-	 * call.  The point of palloc'ing here, rather than having a static char
-	 * array, is first to ensure adequate alignment for the checksumming code
-	 * and second to avoid wasting space in processes that never call this.
-	 */
-	if (pageCopy == NULL)
-		pageCopy = MemoryContextAllocAligned(TopMemoryContext,
-											 BLCKSZ,
-											 PG_IO_ALIGN_SIZE,
-											 0);
-
-	memcpy(pageCopy, page, BLCKSZ);
-	((PageHeader) pageCopy)->pd_checksum = pg_checksum_page(pageCopy, blkno);
-	return pageCopy;
-}
-
-/*
- * Set checksum for a page in private memory.
- *
- * This must only be used when we know that no other process can be modifying
- * the page buffer.
+ * In the past this needed to be done on a copy of the page, due to the
+ * possibility of e.g. hint bits being set concurrently. However, this is not
+ * necessary anymore as hint bits won't be set while IO is going on.
  */
 void
-PageSetChecksumInplace(Page page, BlockNumber blkno)
+PageSetChecksum(Page page, BlockNumber blkno)
 {
 	/* If we don't need a checksum, just return */
 	if (PageIsNew(page) || !DataChecksumsEnabled())
diff --git a/src/backend/storage/smgr/bulk_write.c b/src/backend/storage/smgr/bulk_write.c
index d43c30da48e..48448de356a 100644
--- a/src/backend/storage/smgr/bulk_write.c
+++ b/src/backend/storage/smgr/bulk_write.c
@@ -279,7 +279,7 @@ smgr_bulk_flush(BulkWriteState *bulkstate)
 		BlockNumber blkno = pending_writes[i].blkno;
 		Page		page = pending_writes[i].buf->data;
 
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 
 		if (blkno >= bulkstate->relsize)
 		{
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index 488d98e7e66..e5fc7642dc2 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -288,7 +288,7 @@ modify_rel_block(PG_FUNCTION_ARGS)
 	}
 	else
 	{
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 	}
 
 	smgrwrite(RelationGetSmgr(rel),
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-09 00:29                                   ` Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-01-09 00:29 UTC (permalink / raw)
  To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

I pushed what was 0001, 0002 in v8. Attached is an updated set of patches for
the rest.

Main changes:

- the explanation in "heapam: Use exclusive lock on old page in CLUSTER" as to
  why it's problematic to not set hint bits wasn't quite right. I've updated
  it:

    heapam: Use exclusive lock on old page in CLUSTER

    To be able to guarantee that we can set the hint bit, acquire an exclusive
    lock on the old buffer. This is required as a future commit will only allow
    hint bits to be set with a new lock level, which is acquired as-needed in a
    non-blocking fashion.

    We need the hint bits, set in heapam_relation_copy_for_cluster() ->
    HeapTupleSatisfiesVacuum(), to be set, as otherwise reform_and_rewrite_tuple()
    -> rewrite_heap_tuple() will get confused. Specifically, rewrite_heap_tuple()
    checks for HEAP_XMAX_INVALID in the old tuple to determine whether to check
    the old-to-new mapping hash table.

- I added a patch that inverts the meaning of LW_FLAG_RELEASE_OK, to make the
  equivalent code for content locks easier. For buffer content locks we reset
  the flags when invalidating, and otherwise we'd either need to not have the
  equivalent of LW_FLAG_RELEASE_OK in the flag mask or explicitly add it after
  making the buffer valid.

  I think it's also nicer this way round, because we e.g. can assert that
  there are no pending wakeups when invalidating a buffer.

- I added a patch to reorganize some of the flags stuff in buf_internals.h, to
  make the later patches cleaner. In particular flags are now defined with a
  macro so that changing at which offset flag bits are doesn't require
  touching every single flag value.

- For the main commit, the reorganized flag stuff removed one of the remaining
  FIXMEs.

- I removed the performance instrumentation stuff from the batch visibility
  commit.


I think 0001, 0002, 0003 can be committed. 0004, 0005 are new and probably
could use a sanity check.  0006 hasn't changed much and is imo pretty much
ready, but should be pushed together with 0007.  0007 is getting close, I
think.  0008-0010 need a bit more work, but I think that can wait until 0007
has been pushed.

For 0008, it'd be nice if somebody could look at the way buf_internals.h now
looks.

The only remaining FIXME in 0008 is about the the reuse of
PGPROC->{lwWaiting,lwWaitMode,lwWaitLink}. I think reusing them for content
locks isn't pretty, but it's probably not worth duplicating them. Thoughts?

Greetings,

Andres Freund

Attachments:

  [text/x-diff] v9-0001-freespace-Don-t-modify-page-without-any-lock.patch (2.0K, ../../4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf/2-v9-0001-freespace-Don-t-modify-page-without-any-lock.patch)
  download | inline diff:
From 00e407beda9b6e2d9190b8b07d442115dd306568 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 1 Dec 2025 22:31:44 -0500
Subject: [PATCH v9 01/10] freespace: Don't modify page without any lock

Before this commit fsm_vacuum_page() modified the page without any lock on the
page. Historically that was kind of ok, as we didn't rely on the freespace to
really stay consistent and we did not have checksums. But these days pages are
checksummed and there are ways for FSM pages to be included in WAL records,
even if the FSM itself is still not WAL logged. If a FSM page ever were
modified while a WAL record referenced that page, we'd be in trouble, as the
WAL CRC could end up getting corrupted.

The reason to address this right now is a series of patches with the goal to
only allow modifications of pages with an appropriate lock level. Obviously
not having any lock is not appropriate :)

Discussion: https://postgr.es/m/4wggb7purufpto6x35fd2kwhasehnzfdy3zdcu47qryubs2hdz@fa5kannykekr
Discussion: https://postgr.es/m/e6a8f734-2198-4958-a028-aba863d4a204@iki.fi
---
 src/backend/storage/freespace/freespace.c | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index 19cf4155263..ad337c00871 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -906,10 +906,12 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 	/*
 	 * Reset the next slot pointer. This encourages the use of low-numbered
 	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation.  We don't bother with a lock here, nor with marking the page
-	 * dirty if it wasn't already, since this is just a hint.
+	 * relation. We don't bother with marking the page dirty if it wasn't
+	 * already, since this is just a hint.
 	 */
+	LockBuffer(buf, BUFFER_LOCK_SHARE);
 	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 	ReleaseBuffer(buf);
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v9-0002-heapam-Use-exclusive-lock-on-old-page-in-CLUSTER.patch (3.9K, ../../4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf/3-v9-0002-heapam-Use-exclusive-lock-on-old-page-in-CLUSTER.patch)
  download | inline diff:
From dcd2c7887b26d313351b7616e7b24a8e199c9445 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Sun, 26 Jan 2025 15:18:46 -0500
Subject: [PATCH v9 02/10] heapam: Use exclusive lock on old page in CLUSTER

To be able to guarantee that we can set the hint bit, acquire an exclusive
lock on the old buffer. This is required as a future commit will only allow
hint bits to be set with a new lock level, which is acquired as-needed in a
non-blocking fashion.

We need the hint bits, set in heapam_relation_copy_for_cluster() ->
HeapTupleSatisfiesVacuum(), to be set, as otherwise reform_and_rewrite_tuple()
-> rewrite_heap_tuple() will get confused. Specifically, rewrite_heap_tuple()
checks for HEAP_XMAX_INVALID in the old tuple to determine whether to check
the old-to-new mapping hash table.

It'd be better if we somehow could avoid setting hint bits on the old page. A
common reason to use VACUUM FULL are very bloated tables - rewriting most of
the old table before during VACUUM FULL doesn't exactly help.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/heap/heapam_handler.c    | 16 +++++++++++++++-
 src/backend/access/heap/heapam_visibility.c |  7 +++++++
 src/backend/access/heap/rewriteheap.c       |  3 +++
 3 files changed, 25 insertions(+), 1 deletion(-)

diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index 09a456e9966..505faaaf9d2 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -837,7 +837,21 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		tuple = ExecFetchSlotHeapTuple(slot, false, NULL);
 		buf = hslot->buffer;
 
-		LockBuffer(buf, BUFFER_LOCK_SHARE);
+		/*
+		 * To be able to guarantee that we can set the hint bit, acquire an
+		 * exclusive lock on the old buffer. We need the hint bits, set in
+		 * heapam_relation_copy_for_cluster() -> HeapTupleSatisfiesVacuum(),
+		 * to be set, as otherwise reform_and_rewrite_tuple() ->
+		 * rewrite_heap_tuple() will get confused. Specifically,
+		 * rewrite_heap_tuple() checks for HEAP_XMAX_INVALID in the old tuple
+		 * to determine whether to check the old-to-new mapping hash table.
+		 *
+		 * It'd be better if we somehow could avoid setting hint bits on the
+		 * old page. One reason to use VACUUM FULL are very bloated tables -
+		 * rewriting most of the old table before during VACUUM FULL doesn't
+		 * exactly help...
+		 */
+		LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
 		switch (HeapTupleSatisfiesVacuum(tuple, OldestXmin, buf))
 		{
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 05e70b7d92a..9a034d5c9e8 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -141,6 +141,13 @@ void
 HeapTupleSetHintBits(HeapTupleHeader tuple, Buffer buffer,
 					 uint16 infomask, TransactionId xid)
 {
+	/*
+	 * The uses from heapam.c rely on being able to perform the hint bit
+	 * updates, which can only be guaranteed if we are holding an exclusive
+	 * lock on the buffer - which all callers are doing.
+	 */
+	Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+
 	SetHintBits(tuple, buffer, infomask, xid);
 }
 
diff --git a/src/backend/access/heap/rewriteheap.c b/src/backend/access/heap/rewriteheap.c
index bae3a2da77a..77fd48eb59e 100644
--- a/src/backend/access/heap/rewriteheap.c
+++ b/src/backend/access/heap/rewriteheap.c
@@ -382,6 +382,9 @@ rewrite_heap_tuple(RewriteState state,
 
 	/*
 	 * If the tuple has been updated, check the old-to-new mapping hash table.
+	 *
+	 * Note that this check relies on the HeapTupleSatisfiesVacuum() in
+	 * heapam_relation_copy_for_cluster() to have set hint bits.
 	 */
 	if (!((old_tuple->t_data->t_infomask & HEAP_XMAX_INVALID) ||
 		  HeapTupleHeaderIsOnlyLocked(old_tuple->t_data)) &&
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v9-0003-heapam-Add-batch-mode-mvcc-check-and-use-it-in-pa.patch (7.5K, ../../4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf/4-v9-0003-heapam-Add-batch-mode-mvcc-check-and-use-it-in-pa.patch)
  download | inline diff:
From 9f7947fb309b7aba4756b77d95ea9065dd531ada Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 17 Oct 2024 13:16:36 -0400
Subject: [PATCH v9 03/10] heapam: Add batch mode mvcc check and use it in page
 mode

There are two reasons for doing so:

1) It is generally faster to perform checks in a batched fashion and making
   sequential scans faster is nice.

2) We would like to stop setting hint bits while pages are being written
   out. The necessary locking becomes visible for page mode scans if done for
   every tuple. With batching the overhead can be amortized to only happen
   once per page.

There are substantial further optimization opportunities along these
lines:

- Right now HeapTupleSatisfiesMVCCBatch() simply uses the single-tuple
  HeapTupleSatisfiesMVCC(), relying on the compiler to inline it. We could
  instead write an explicitly optimized version that avoids repeated xid
  tests.

- Introduce batched version of the serializability test

- Introduce batched version of HeapTupleSatisfiesVacuum

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/6rgb2nvhyvnszz4ul3wfzlf5rheb2kkwrglthnna7qhe24onwr@vw27225tkyar
---
 src/include/access/heapam.h                 | 17 +++++
 src/backend/access/heap/heapam.c            | 84 ++++++++++++++++-----
 src/backend/access/heap/heapam_visibility.c | 42 +++++++++++
 src/tools/pgindent/typedefs.list            |  1 +
 4 files changed, 124 insertions(+), 20 deletions(-)

diff --git a/src/include/access/heapam.h b/src/include/access/heapam.h
index ce48fac42ba..177e1b525a8 100644
--- a/src/include/access/heapam.h
+++ b/src/include/access/heapam.h
@@ -449,6 +449,23 @@ extern bool HeapTupleHeaderIsOnlyLocked(HeapTupleHeader tuple);
 extern bool HeapTupleIsSurelyDead(HeapTuple htup,
 								  GlobalVisState *vistest);
 
+/*
+ * The output of HeapTupleSatisfiesMVCCBatch() is passed via this struct, as
+ * otherwise the increased number of arguments to
+ * HeapTupleSatisfiesMVCCBatch() leads to on-stack argument passing on x86-64,
+ * which causes a small regression.
+ */
+typedef struct BatchMVCCState
+{
+	HeapTupleData tuples[MaxHeapTuplesPerPage];
+	bool		visible[MaxHeapTuplesPerPage];
+} BatchMVCCState;
+
+extern int	HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
+										int ntups,
+										BatchMVCCState *batchmvcc,
+										OffsetNumber *vistuples_dense);
+
 /*
  * To avoid leaking too much knowledge about reorderbuffer implementation
  * details this is implemented in reorderbuffer.c not heapam_visibility.c
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index ad9d6338ec2..f30a56ecf55 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -522,42 +522,86 @@ page_collect_tuples(HeapScanDesc scan, Snapshot snapshot,
 					BlockNumber block, int lines,
 					bool all_visible, bool check_serializable)
 {
+	Oid			relid = RelationGetRelid(scan->rs_base.rs_rd);
 	int			ntup = 0;
-	OffsetNumber lineoff;
+	int			nvis = 0;
+	BatchMVCCState batchmvcc;
 
-	for (lineoff = FirstOffsetNumber; lineoff <= lines; lineoff++)
+	/* page at a time should have been disabled otherwise */
+	Assert(IsMVCCSnapshot(snapshot));
+
+	/* first find all tuples on the page */
+	for (OffsetNumber lineoff = FirstOffsetNumber; lineoff <= lines; lineoff++)
 	{
 		ItemId		lpp = PageGetItemId(page, lineoff);
-		HeapTupleData loctup;
-		bool		valid;
+		HeapTuple	tup;
 
-		if (!ItemIdIsNormal(lpp))
+		if (unlikely(!ItemIdIsNormal(lpp)))
 			continue;
 
-		loctup.t_data = (HeapTupleHeader) PageGetItem(page, lpp);
-		loctup.t_len = ItemIdGetLength(lpp);
-		loctup.t_tableOid = RelationGetRelid(scan->rs_base.rs_rd);
-		ItemPointerSet(&(loctup.t_self), block, lineoff);
+		/*
+		 * If the page is not all-visible or we need to check serializability,
+		 * maintain enough state to be able to refind the tuple efficiently,
+		 * without again first needing to fetch the item and then via that the
+		 * tuple.
+		 */
+		if (!all_visible || check_serializable)
+		{
+			tup = &batchmvcc.tuples[ntup];
 
+			tup->t_data = (HeapTupleHeader) PageGetItem(page, lpp);
+			tup->t_len = ItemIdGetLength(lpp);
+			tup->t_tableOid = relid;
+			ItemPointerSet(&(tup->t_self), block, lineoff);
+		}
+
+		/*
+		 * If the page is all visible, these fields otherwise won't be
+		 * populated in loop below.
+		 */
 		if (all_visible)
-			valid = true;
-		else
-			valid = HeapTupleSatisfiesVisibility(&loctup, snapshot, buffer);
-
-		if (check_serializable)
-			HeapCheckForSerializableConflictOut(valid, scan->rs_base.rs_rd,
-												&loctup, buffer, snapshot);
-
-		if (valid)
 		{
+			if (check_serializable)
+			{
+				batchmvcc.visible[ntup] = true;
+			}
 			scan->rs_vistuples[ntup] = lineoff;
-			ntup++;
 		}
+
+		ntup++;
 	}
 
 	Assert(ntup <= MaxHeapTuplesPerPage);
 
-	return ntup;
+	/*
+	 * Unless the page is all visible, test visibility for all tuples one go.
+	 * That is considerably more efficient than calling
+	 * HeapTupleSatisfiesMVCC() one-by-one.
+	 */
+	if (all_visible)
+		nvis = ntup;
+	else
+		nvis = HeapTupleSatisfiesMVCCBatch(snapshot, buffer,
+										   ntup,
+										   &batchmvcc,
+										   scan->rs_vistuples);
+
+	/*
+	 * So far we don't have batch API for testing serializabilty, so do so
+	 * one-by-one.
+	 */
+	if (check_serializable)
+	{
+		for (int i = 0; i < ntup; i++)
+		{
+			HeapCheckForSerializableConflictOut(batchmvcc.visible[i],
+												scan->rs_base.rs_rd,
+												&batchmvcc.tuples[i],
+												buffer, snapshot);
+		}
+	}
+
+	return nvis;
 }
 
 /*
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 9a034d5c9e8..5d56c6e5075 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -1598,6 +1598,48 @@ HeapTupleSatisfiesHistoricMVCC(HeapTuple htup, Snapshot snapshot,
 		return true;
 }
 
+/*
+ * Perform HeaptupleSatisfiesMVCC() on each passed in tuple. This is more
+ * efficient than doing HeapTupleSatisfiesMVCC() one-by-one.
+ *
+ * To be checked tuples are passed via BatchMVCCState->tuples. Each tuple's
+ * visibility is stored in batchmvcc->visible[]. In addition,
+ * ->vistuples_dense is set to contain the offsets of visible tuples.
+ *
+ * The reason this is more efficient than HeapTupleSatisfiesMVCC() is that it
+ * avoids a cross-translation-unit function call for each tuple. In the future
+ * it will also allow more efficient setting of hint bits.
+ *
+ * Returns the number of visible tuples.
+ */
+int
+HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
+							int ntups,
+							BatchMVCCState *batchmvcc,
+							OffsetNumber *vistuples_dense)
+{
+	int			nvis = 0;
+
+	Assert(IsMVCCSnapshot(snapshot));
+
+	for (int i = 0; i < ntups; i++)
+	{
+		bool		valid;
+		HeapTuple	tup = &batchmvcc->tuples[i];
+
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		batchmvcc->visible[i] = valid;
+
+		if (likely(valid))
+		{
+			vistuples_dense[nvis] = tup->t_self.ip_posid;
+			nvis++;
+		}
+	}
+
+	return nvis;
+}
+
 /*
  * HeapTupleSatisfiesVisibility
  *		True iff heap tuple satisfies a time qual.
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 09e7f1d420e..14dec2d49c1 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -255,6 +255,7 @@ Barrier
 BaseBackupCmd
 BaseBackupTargetHandle
 BaseBackupTargetType
+BatchMVCCState
 BeginDirectModify_function
 BeginForeignInsert_function
 BeginForeignModify_function
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v9-0004-lwlock-Invert-meaning-of-LW_FLAG_RELEASE_OK.patch (5.5K, ../../4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf/5-v9-0004-lwlock-Invert-meaning-of-LW_FLAG_RELEASE_OK.patch)
  download | inline diff:
From aee0f0645919d3e7b46fcea9a55f8747b635bf7e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 5 Jan 2026 20:40:38 -0500
Subject: [PATCH v9 04/10] lwlock: Invert meaning of LW_FLAG_RELEASE_OK

Instead of setting a flag whenever a lock release is supposed to wake up
waiters - the majority of the time - set a flag whenever wakeups are
inhibited. The motivation for this that in an upcoming commit, buffer content
locks are implemented separately from lwlocks, and for buffer content locks it
is useful to be able to reset all buffer flags when a buffer invalidated,
alternatively we would have to set the release-ok flag when making a buffer
valid.  It seems good to keep the implementation of lwlocks and buffer content
locks as similar as reasonably possible.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/storage/lmgr/lwlock.c | 42 +++++++++++++++----------------
 1 file changed, 20 insertions(+), 22 deletions(-)

diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 6a9f86d5025..148309cc186 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -92,7 +92,7 @@
 
 
 #define LW_FLAG_HAS_WAITERS			((uint32) 1 << 31)
-#define LW_FLAG_RELEASE_OK			((uint32) 1 << 30)
+#define LW_FLAG_WAKE_IN_PROGRESS	((uint32) 1 << 30)
 #define LW_FLAG_LOCKED				((uint32) 1 << 29)
 #define LW_FLAG_BITS				3
 #define LW_FLAG_MASK				(((1<<LW_FLAG_BITS)-1)<<(32-LW_FLAG_BITS))
@@ -246,14 +246,14 @@ PRINT_LWDEBUG(const char *where, LWLock *lock, LWLockMode mode)
 		ereport(LOG,
 				(errhidestmt(true),
 				 errhidecontext(true),
-				 errmsg_internal("%d: %s(%s %p): excl %u shared %u haswaiters %u waiters %u rOK %d",
+				 errmsg_internal("%d: %s(%s %p): excl %u shared %u haswaiters %u waiters %u waking %d",
 								 MyProcPid,
 								 where, T_NAME(lock), lock,
 								 (state & LW_VAL_EXCLUSIVE) != 0,
 								 state & LW_SHARED_MASK,
 								 (state & LW_FLAG_HAS_WAITERS) != 0,
 								 pg_atomic_read_u32(&lock->nwaiters),
-								 (state & LW_FLAG_RELEASE_OK) != 0)));
+								 (state & LW_FLAG_WAKE_IN_PROGRESS) != 0)));
 	}
 }
 
@@ -700,7 +700,7 @@ LWLockInitialize(LWLock *lock, int tranche_id)
 	/* verify the tranche_id is valid */
 	(void) GetLWTrancheName(tranche_id);
 
-	pg_atomic_init_u32(&lock->state, LW_FLAG_RELEASE_OK);
+	pg_atomic_init_u32(&lock->state, 0);
 #ifdef LOCK_DEBUG
 	pg_atomic_init_u32(&lock->nwaiters, 0);
 #endif
@@ -929,15 +929,13 @@ LWLockWaitListUnlock(LWLock *lock)
 static void
 LWLockWakeup(LWLock *lock)
 {
-	bool		new_release_ok;
+	bool		new_release_in_progress = false;
 	bool		wokeup_somebody = false;
 	proclist_head wakeup;
 	proclist_mutable_iter iter;
 
 	proclist_init(&wakeup);
 
-	new_release_ok = true;
-
 	/* lock wait list while collecting backends to wake up */
 	LWLockWaitListLock(lock);
 
@@ -958,7 +956,7 @@ LWLockWakeup(LWLock *lock)
 			 * that are just waiting for the lock to become free don't retry
 			 * automatically.
 			 */
-			new_release_ok = false;
+			new_release_in_progress = true;
 
 			/*
 			 * Don't wakeup (further) exclusive locks.
@@ -997,10 +995,10 @@ LWLockWakeup(LWLock *lock)
 
 			/* compute desired flags */
 
-			if (new_release_ok)
-				desired_state |= LW_FLAG_RELEASE_OK;
+			if (new_release_in_progress)
+				desired_state |= LW_FLAG_WAKE_IN_PROGRESS;
 			else
-				desired_state &= ~LW_FLAG_RELEASE_OK;
+				desired_state &= ~LW_FLAG_WAKE_IN_PROGRESS;
 
 			if (proclist_is_empty(&lock->waiters))
 				desired_state &= ~LW_FLAG_HAS_WAITERS;
@@ -1131,10 +1129,10 @@ LWLockDequeueSelf(LWLock *lock)
 		 */
 
 		/*
-		 * Reset RELEASE_OK flag if somebody woke us before we removed
-		 * ourselves - they'll have set it to false.
+		 * Clear LW_FLAG_WAKE_IN_PROGRESS if somebody woke us before we
+		 * removed ourselves - they'll have set it.
 		 */
-		pg_atomic_fetch_or_u32(&lock->state, LW_FLAG_RELEASE_OK);
+		pg_atomic_fetch_and_u32(&lock->state, ~LW_FLAG_WAKE_IN_PROGRESS);
 
 		/*
 		 * Now wait for the scheduled wakeup, otherwise our ->lwWaiting would
@@ -1301,7 +1299,7 @@ LWLockAcquire(LWLock *lock, LWLockMode mode)
 		}
 
 		/* Retrying, allow LWLockRelease to release waiters again. */
-		pg_atomic_fetch_or_u32(&lock->state, LW_FLAG_RELEASE_OK);
+		pg_atomic_fetch_and_u32(&lock->state, ~LW_FLAG_WAKE_IN_PROGRESS);
 
 #ifdef LOCK_DEBUG
 		{
@@ -1636,10 +1634,10 @@ LWLockWaitForVar(LWLock *lock, pg_atomic_uint64 *valptr, uint64 oldval,
 		LWLockQueueSelf(lock, LW_WAIT_UNTIL_FREE);
 
 		/*
-		 * Set RELEASE_OK flag, to make sure we get woken up as soon as the
-		 * lock is released.
+		 * Clear LW_FLAG_WAKE_IN_PROGRESS flag, to make sure we get woken up
+		 * as soon as the lock is released.
 		 */
-		pg_atomic_fetch_or_u32(&lock->state, LW_FLAG_RELEASE_OK);
+		pg_atomic_fetch_and_u32(&lock->state, ~LW_FLAG_WAKE_IN_PROGRESS);
 
 		/*
 		 * We're now guaranteed to be woken up if necessary. Recheck the lock
@@ -1852,11 +1850,11 @@ LWLockReleaseInternal(LWLock *lock, LWLockMode mode)
 		TRACE_POSTGRESQL_LWLOCK_RELEASE(T_NAME(lock));
 
 	/*
-	 * We're still waiting for backends to get scheduled, don't wake them up
-	 * again.
+	 * Check if we're still waiting for backends to get scheduled, if so,
+	 * don't wake them up again.
 	 */
-	if ((oldstate & (LW_FLAG_HAS_WAITERS | LW_FLAG_RELEASE_OK)) ==
-		(LW_FLAG_HAS_WAITERS | LW_FLAG_RELEASE_OK) &&
+	if ((oldstate & LW_FLAG_HAS_WAITERS) &&
+		!(oldstate & LW_FLAG_WAKE_IN_PROGRESS) &&
 		(oldstate & LW_LOCK_MASK) == 0)
 		check_waiters = true;
 	else
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v9-0005-bufmgr-Make-definitions-related-to-buffer-descrip.patch (4.5K, ../../4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf/6-v9-0005-bufmgr-Make-definitions-related-to-buffer-descrip.patch)
  download | inline diff:
From d2204f78871fed341e55c4b346f3417441e1ad75 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 7 Jan 2026 17:21:48 -0500
Subject: [PATCH v9 05/10] bufmgr: Make definitions related to buffer
 descriptor easier to modify

This is in preparation to widening the buffer state to 64 bits, which in turn
is preparation for implementing content locks in bufmgr. This commit aims to
make the subsequent commits a bit easier to review, by separating out
reformatting etc from the actual changes.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/buf_internals.h | 65 +++++++++++++++++++++--------
 1 file changed, 47 insertions(+), 18 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index fa43cf4458d..2f607ea2ac5 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -32,6 +32,7 @@
 /*
  * Buffer state is a single 32-bit variable where following data is combined.
  *
+ * State of the buffer itself (in order):
  * - 18 bits refcount
  * - 4 bits usage count
  * - 10 bits of flags
@@ -48,16 +49,30 @@
 StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 				 "parts of buffer state space need to equal 32");
 
+/* refcount related definitions */
 #define BUF_REFCOUNT_ONE 1
-#define BUF_REFCOUNT_MASK ((1U << BUF_REFCOUNT_BITS) - 1)
-#define BUF_USAGECOUNT_MASK (((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_REFCOUNT_BITS))
-#define BUF_USAGECOUNT_ONE (1U << BUF_REFCOUNT_BITS)
-#define BUF_USAGECOUNT_SHIFT BUF_REFCOUNT_BITS
-#define BUF_FLAG_MASK (((1U << BUF_FLAG_BITS) - 1) << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS))
+#define BUF_REFCOUNT_MASK \
+	((1U << BUF_REFCOUNT_BITS) - 1)
+
+/* usage count related definitions */
+#define BUF_USAGECOUNT_SHIFT \
+	BUF_REFCOUNT_BITS
+#define BUF_USAGECOUNT_MASK \
+	(((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_USAGECOUNT_SHIFT))
+#define BUF_USAGECOUNT_ONE \
+	(1U << BUF_REFCOUNT_BITS)
+
+/* flags related definitions */
+#define BUF_FLAG_SHIFT \
+	(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS)
+#define BUF_FLAG_MASK \
+	(((1U << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
 
 /* Get refcount and usagecount from buffer state */
-#define BUF_STATE_GET_REFCOUNT(state) ((state) & BUF_REFCOUNT_MASK)
-#define BUF_STATE_GET_USAGECOUNT(state) (((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
+#define BUF_STATE_GET_REFCOUNT(state) \
+	((state) & BUF_REFCOUNT_MASK)
+#define BUF_STATE_GET_USAGECOUNT(state) \
+	(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
 
 /*
  * Flags for buffer descriptors
@@ -65,17 +80,31 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  * Note: BM_TAG_VALID essentially means that there is a buffer hashtable
  * entry associated with the buffer's tag.
  */
-#define BM_LOCKED				(1U << 22)	/* buffer header is locked */
-#define BM_DIRTY				(1U << 23)	/* data needs writing */
-#define BM_VALID				(1U << 24)	/* data is valid */
-#define BM_TAG_VALID			(1U << 25)	/* tag is assigned */
-#define BM_IO_IN_PROGRESS		(1U << 26)	/* read or write in progress */
-#define BM_IO_ERROR				(1U << 27)	/* previous I/O failed */
-#define BM_JUST_DIRTIED			(1U << 28)	/* dirtied since write started */
-#define BM_PIN_COUNT_WAITER		(1U << 29)	/* have waiter for sole pin */
-#define BM_CHECKPOINT_NEEDED	(1U << 30)	/* must write for checkpoint */
-#define BM_PERMANENT			(1U << 31)	/* permanent buffer (not unlogged,
-											 * or init fork) */
+
+#define BUF_DEFINE_FLAG(flagno)	\
+	(1U << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + (flagno)))
+
+/* buffer header is locked */
+#define BM_LOCKED					BUF_DEFINE_FLAG( 0)
+/* data needs writing */
+#define BM_DIRTY					BUF_DEFINE_FLAG( 1)
+/* data is valid */
+#define BM_VALID					BUF_DEFINE_FLAG( 2)
+/* tag is assigned */
+#define BM_TAG_VALID				BUF_DEFINE_FLAG( 3)
+/* read or write in progress */
+#define BM_IO_IN_PROGRESS			BUF_DEFINE_FLAG( 4)
+/* previous I/O failed */
+#define BM_IO_ERROR					BUF_DEFINE_FLAG( 5)
+/* dirtied since write started */
+#define BM_JUST_DIRTIED				BUF_DEFINE_FLAG( 6)
+/* have waiter for sole pin */
+#define BM_PIN_COUNT_WAITER			BUF_DEFINE_FLAG( 7)
+/* must write for checkpoint */
+#define BM_CHECKPOINT_NEEDED		BUF_DEFINE_FLAG( 8)
+/* permanent buffer (not unlogged, or init fork) */
+#define BM_PERMANENT				BUF_DEFINE_FLAG( 9)
+
 /*
  * The maximum allowed value of usage_count represents a tradeoff between
  * accuracy and speed of the clock-sweep buffer management algorithm.  A
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v9-0006-bufmgr-Change-BufferDesc.state-to-be-a-64bit-atom.patch (45.1K, ../../4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf/7-v9-0006-bufmgr-Change-BufferDesc.state-to-be-a-64bit-atom.patch)
  download | inline diff:
From 60fd51df183262ff6634c314e03685bc9922043b Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 7 Jan 2026 17:26:25 -0500
Subject: [PATCH v9 06/10] bufmgr: Change BufferDesc.state to be a 64bit atomic

This is motivated by wanting to merge buffer content locks into
BufferDesc.state in a future commit, rather than having a separate lwlock (see
commit c75ebc657ff more details). As this change is rather mechanical, it
seems to make sense to split it out into a separate commit, for easier review.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  51 +++---
 src/include/storage/procnumber.h              |  14 +-
 src/backend/storage/buffer/buf_init.c         |   2 +-
 src/backend/storage/buffer/bufmgr.c           | 170 +++++++++---------
 src/backend/storage/buffer/freelist.c         |  24 +--
 src/backend/storage/buffer/localbuf.c         |  72 ++++----
 contrib/pg_buffercache/pg_buffercache_pages.c |   8 +-
 src/test/modules/test_aio/test_aio.c          |  12 +-
 8 files changed, 178 insertions(+), 175 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 2f607ea2ac5..a4d36e9ca01 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -30,7 +30,7 @@
 #include "utils/resowner.h"
 
 /*
- * Buffer state is a single 32-bit variable where following data is combined.
+ * Buffer state is a single 64-bit variable where following data is combined.
  *
  * State of the buffer itself (in order):
  * - 18 bits refcount
@@ -40,6 +40,9 @@
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
+ * NB: A future commit will use a significant portion of the remaining bits to
+ * implement buffer locking as part of the state variable.
+ *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
@@ -52,27 +55,27 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 /* refcount related definitions */
 #define BUF_REFCOUNT_ONE 1
 #define BUF_REFCOUNT_MASK \
-	((1U << BUF_REFCOUNT_BITS) - 1)
+	((UINT64CONST(1) << BUF_REFCOUNT_BITS) - 1)
 
 /* usage count related definitions */
 #define BUF_USAGECOUNT_SHIFT \
 	BUF_REFCOUNT_BITS
 #define BUF_USAGECOUNT_MASK \
-	(((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_USAGECOUNT_SHIFT))
+	(((UINT64CONST(1) << BUF_USAGECOUNT_BITS) - 1) << (BUF_USAGECOUNT_SHIFT))
 #define BUF_USAGECOUNT_ONE \
-	(1U << BUF_REFCOUNT_BITS)
+	(UINT64CONST(1) << BUF_REFCOUNT_BITS)
 
 /* flags related definitions */
 #define BUF_FLAG_SHIFT \
 	(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS)
 #define BUF_FLAG_MASK \
-	(((1U << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
+	(((UINT64CONST(1) << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
 
 /* Get refcount and usagecount from buffer state */
 #define BUF_STATE_GET_REFCOUNT(state) \
-	((state) & BUF_REFCOUNT_MASK)
+	((uint32)((state) & BUF_REFCOUNT_MASK))
 #define BUF_STATE_GET_USAGECOUNT(state) \
-	(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
+	((uint32)(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT))
 
 /*
  * Flags for buffer descriptors
@@ -82,7 +85,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  */
 
 #define BUF_DEFINE_FLAG(flagno)	\
-	(1U << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + (flagno)))
+	(UINT64CONST(1) << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + (flagno)))
 
 /* buffer header is locked */
 #define BM_LOCKED					BUF_DEFINE_FLAG( 0)
@@ -115,7 +118,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  */
 #define BM_MAX_USAGE_COUNT	5
 
-StaticAssertDecl(BM_MAX_USAGE_COUNT < (1 << BUF_USAGECOUNT_BITS),
+StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 				 "BM_MAX_USAGE_COUNT doesn't fit in BUF_USAGECOUNT_BITS bits");
 StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
 				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
@@ -280,8 +283,8 @@ BufMappingPartitionLockByIndex(uint32 index)
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
- * atomic operations (i.e. only pg_atomic_read_u32() and
- * pg_atomic_unlocked_write_u32()).
+ * atomic operations (i.e. only pg_atomic_read_u64() and
+ * pg_atomic_unlocked_write_u64()).
  *
  * Be careful to avoid increasing the size of the struct when adding or
  * reordering members.  Keeping it below 64 bytes (the most common CPU
@@ -309,7 +312,7 @@ typedef struct BufferDesc
 	 * State of the buffer, containing flags, refcount and usagecount. See
 	 * BUF_* and BM_* defines at the top of this file.
 	 */
-	pg_atomic_uint32 state;
+	pg_atomic_uint64 state;
 
 	/*
 	 * Backend of pin-count waiter. The buffer header spinlock needs to be
@@ -415,7 +418,7 @@ BufferDescriptorGetContentLock(const BufferDesc *bdesc)
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
  */
-extern uint32 LockBufHdr(BufferDesc *desc);
+extern uint64 LockBufHdr(BufferDesc *desc);
 
 /*
  * Unlock the buffer header.
@@ -426,9 +429,9 @@ extern uint32 LockBufHdr(BufferDesc *desc);
 static inline void
 UnlockBufHdr(BufferDesc *desc)
 {
-	Assert(pg_atomic_read_u32(&desc->state) & BM_LOCKED);
+	Assert(pg_atomic_read_u64(&desc->state) & BM_LOCKED);
 
-	pg_atomic_fetch_sub_u32(&desc->state, BM_LOCKED);
+	pg_atomic_fetch_sub_u64(&desc->state, BM_LOCKED);
 }
 
 /*
@@ -439,14 +442,14 @@ UnlockBufHdr(BufferDesc *desc)
  * Note that this approach would not work for usagecount, since we need to cap
  * the usagecount at BM_MAX_USAGE_COUNT.
  */
-static inline uint32
-UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
-				uint32 set_bits, uint32 unset_bits,
+static inline uint64
+UnlockBufHdrExt(BufferDesc *desc, uint64 old_buf_state,
+				uint64 set_bits, uint64 unset_bits,
 				int refcount_change)
 {
 	for (;;)
 	{
-		uint32		buf_state = old_buf_state;
+		uint64		buf_state = old_buf_state;
 
 		Assert(buf_state & BM_LOCKED);
 
@@ -457,7 +460,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 		if (refcount_change != 0)
 			buf_state += BUF_REFCOUNT_ONE * refcount_change;
 
-		if (pg_atomic_compare_exchange_u32(&desc->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&desc->state, &old_buf_state,
 										   buf_state))
 		{
 			return old_buf_state;
@@ -465,7 +468,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 	}
 }
 
-extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
+extern uint64 WaitBufHdrUnlocked(BufferDesc *buf);
 
 /* in bufmgr.c */
 
@@ -525,14 +528,14 @@ extern void TrackNewBufferPin(Buffer buf);
 
 /* solely to make it easier to write tests */
 extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
-extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 							  bool forget_owner, bool release_aio);
 
 
 /* freelist.c */
 extern IOContext IOContextForStrategy(BufferAccessStrategy strategy);
 extern BufferDesc *StrategyGetBuffer(BufferAccessStrategy strategy,
-									 uint32 *buf_state, bool *from_ring);
+									 uint64 *buf_state, bool *from_ring);
 extern bool StrategyRejectBuffer(BufferAccessStrategy strategy,
 								 BufferDesc *buf, bool from_ring);
 
@@ -568,7 +571,7 @@ extern BlockNumber ExtendBufferedRelLocal(BufferManagerRelation bmr,
 										  uint32 *extended_by);
 extern void MarkLocalBufferDirty(Buffer buffer);
 extern void TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty,
-								   uint32 set_flag_bits, bool release_aio);
+								   uint64 set_flag_bits, bool release_aio);
 extern bool StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait);
 extern void FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln);
 extern void InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced);
diff --git a/src/include/storage/procnumber.h b/src/include/storage/procnumber.h
index 30c360ad350..bd9cb3891cc 100644
--- a/src/include/storage/procnumber.h
+++ b/src/include/storage/procnumber.h
@@ -27,13 +27,13 @@ typedef int ProcNumber;
 
 /*
  * Note: MAX_BACKENDS_BITS is 18 as that is the space available for buffer
- * refcounts in buf_internals.h.  This limitation could be lifted by using a
- * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
- * currently realistic configurations. Even if that limitation were removed,
- * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
- * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
- * 4*MaxBackends without any overflow check.  We check that the configured
- * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
+ * refcounts in buf_internals.h.  This limitation could be lifted, but it's
+ * unlikely to be worthwhile as 2^18-1 backends exceed currently realistic
+ * configurations. Even if that limitation were removed, we still could not a)
+ * exceed 2^23-1 because inval.c stores the ProcNumber as a 3-byte signed
+ * integer, b) INT_MAX/4 because some places compute 4*MaxBackends without any
+ * overflow check.  We check that the configured number of backends does not
+ * exceed MAX_BACKENDS in InitializeMaxBackends().
  */
 #define MAX_BACKENDS_BITS		18
 #define MAX_BACKENDS			((1U << MAX_BACKENDS_BITS)-1)
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 9a312bcc7b3..7d894522526 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -121,7 +121,7 @@ BufferManagerShmemInit(void)
 
 			ClearBufferTag(&buf->tag);
 
-			pg_atomic_init_u32(&buf->state, 0);
+			pg_atomic_init_u64(&buf->state, 0);
 			buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
 
 			buf->buf_id = i;
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index a036c2aa275..b0de8e45d4d 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -780,7 +780,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 {
 	BufferDesc *bufHdr;
 	BufferTag	tag;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(recent_buffer));
 
@@ -793,7 +793,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 		int			b = -recent_buffer - 1;
 
 		bufHdr = GetLocalBufferDescriptor(b);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		/* Is it still valid and holding the right tag? */
 		if ((buf_state & BM_VALID) && BufferTagsEqual(&tag, &bufHdr->tag))
@@ -1386,8 +1386,8 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
 				bufHdr = GetLocalBufferDescriptor(-buffers[i] - 1);
 			else
 				bufHdr = GetBufferDescriptor(buffers[i] - 1);
-			Assert(pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID);
-			found = pg_atomic_read_u32(&bufHdr->state) & BM_VALID;
+			Assert(pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID);
+			found = pg_atomic_read_u64(&bufHdr->state) & BM_VALID;
 		}
 		else
 		{
@@ -1613,10 +1613,10 @@ CheckReadBuffersOperation(ReadBuffersOperation *operation, bool is_complete)
 			GetBufferDescriptor(buffer - 1);
 
 		Assert(BufferGetBlockNumber(buffer) == operation->blocknum + i);
-		Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_TAG_VALID);
+		Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_TAG_VALID);
 
 		if (i < operation->nblocks_done)
-			Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_VALID);
+			Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_VALID);
 	}
 #endif
 }
@@ -2083,8 +2083,8 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	int			existing_buf_id;
 	Buffer		victim_buffer;
 	BufferDesc *victim_buf_hdr;
-	uint32		victim_buf_state;
-	uint32		set_bits = 0;
+	uint64		victim_buf_state;
+	uint64		set_bits = 0;
 
 	/* Make sure we will have room to remember the buffer pin */
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2251,7 +2251,7 @@ InvalidateBuffer(BufferDesc *buf)
 	uint32		oldHash;		/* hash value for oldTag */
 	LWLock	   *oldPartitionLock;	/* buffer partition lock for it */
 	uint32		oldFlags;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/* Save the original buffer tag before dropping the spinlock */
 	oldTag = buf->tag;
@@ -2342,7 +2342,7 @@ retry:
 static bool
 InvalidateVictimBuffer(BufferDesc *buf_hdr)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	uint32		hash;
 	LWLock	   *partition_lock;
 	BufferTag	tag;
@@ -2402,10 +2402,10 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 
 	LWLockRelease(partition_lock);
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	Assert(!(buf_state & (BM_DIRTY | BM_VALID | BM_TAG_VALID)));
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u32(&buf_hdr->state)) > 0);
+	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u64(&buf_hdr->state)) > 0);
 
 	return true;
 }
@@ -2415,7 +2415,7 @@ GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context)
 {
 	BufferDesc *buf_hdr;
 	Buffer		buf;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		from_ring;
 
 	/*
@@ -2548,7 +2548,7 @@ again:
 
 	/* a final set of sanity checks */
 #ifdef USE_ASSERT_CHECKING
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) == 1);
 	Assert(!(buf_state & (BM_TAG_VALID | BM_VALID | BM_DIRTY)));
@@ -2839,13 +2839,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 */
 			do
 			{
-				pg_atomic_fetch_and_u32(&existing_hdr->state, ~BM_VALID);
+				pg_atomic_fetch_and_u64(&existing_hdr->state, ~BM_VALID);
 			} while (!StartBufferIO(existing_hdr, true, false));
 		}
 		else
 		{
-			uint32		buf_state;
-			uint32		set_bits = 0;
+			uint64		buf_state;
+			uint64		set_bits = 0;
 
 			buf_state = LockBufHdr(victim_buf_hdr);
 
@@ -3021,7 +3021,7 @@ BufferIsDirty(Buffer buffer)
 		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 	}
 
-	return pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY;
+	return pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY;
 }
 
 /*
@@ -3037,8 +3037,8 @@ void
 MarkBufferDirty(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
-	uint32		old_buf_state;
+	uint64		buf_state;
+	uint64		old_buf_state;
 
 	if (!BufferIsValid(buffer))
 		elog(ERROR, "bad buffer ID: %d", buffer);
@@ -3058,7 +3058,7 @@ MarkBufferDirty(Buffer buffer)
 	 * NB: We have to wait for the buffer header spinlock to be not held, as
 	 * TerminateBufferIO() relies on the spinlock.
 	 */
-	old_buf_state = pg_atomic_read_u32(&bufHdr->state);
+	old_buf_state = pg_atomic_read_u64(&bufHdr->state);
 	for (;;)
 	{
 		if (old_buf_state & BM_LOCKED)
@@ -3069,7 +3069,7 @@ MarkBufferDirty(Buffer buffer)
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
 
-		if (pg_atomic_compare_exchange_u32(&bufHdr->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&bufHdr->state, &old_buf_state,
 										   buf_state))
 			break;
 	}
@@ -3173,10 +3173,10 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 
 	if (ref == NULL)
 	{
-		uint32		buf_state;
-		uint32		old_buf_state;
+		uint64		buf_state;
+		uint64		old_buf_state;
 
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
@@ -3210,7 +3210,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 					buf_state += BUF_USAGECOUNT_ONE;
 			}
 
-			if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+			if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 											   buf_state))
 			{
 				result = (buf_state & BM_VALID) != 0;
@@ -3237,7 +3237,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		 * that the buffer page is legitimately non-accessible here.  We
 		 * cannot meddle with that.
 		 */
-		result = (pg_atomic_read_u32(&buf->state) & BM_VALID) != 0;
+		result = (pg_atomic_read_u64(&buf->state) & BM_VALID) != 0;
 
 		Assert(ref->data.refcount > 0);
 		ref->data.refcount++;
@@ -3272,7 +3272,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	/*
 	 * As explained, We don't expect any preexisting pins. That allows us to
@@ -3284,7 +3284,7 @@ PinBuffer_Locked(BufferDesc *buf)
 	 * Since we hold the buffer spinlock, we can update the buffer state and
 	 * release the lock in one operation.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 
 	UnlockBufHdrExt(buf, old_buf_state,
 					0, 0, 1);
@@ -3314,7 +3314,7 @@ WakePinCountWaiter(BufferDesc *buf)
 	 * BM_PIN_COUNT_WAITER if it stops waiting for a reason other than this
 	 * backend waking it up.
 	 */
-	uint32		buf_state = LockBufHdr(buf);
+	uint64		buf_state = LockBufHdr(buf);
 
 	if ((buf_state & BM_PIN_COUNT_WAITER) &&
 		BUF_STATE_GET_REFCOUNT(buf_state) == 1)
@@ -3361,7 +3361,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	ref->data.refcount--;
 	if (ref->data.refcount == 0)
 	{
-		uint32		old_buf_state;
+		uint64		old_buf_state;
 
 		/*
 		 * Mark buffer non-accessible to Valgrind.
@@ -3379,7 +3379,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
 		/* decrement the shared reference count */
-		old_buf_state = pg_atomic_fetch_sub_u32(&buf->state, BUF_REFCOUNT_ONE);
+		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
 
 		/* Support LockBufferForCleanup() */
 		if (old_buf_state & BM_PIN_COUNT_WAITER)
@@ -3436,7 +3436,7 @@ TrackNewBufferPin(Buffer buf)
 static void
 BufferSync(int flags)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	int			buf_id;
 	int			num_to_scan;
 	int			num_spaces;
@@ -3446,7 +3446,7 @@ BufferSync(int flags)
 	Oid			last_tsid;
 	binaryheap *ts_heap;
 	int			i;
-	uint32		mask = BM_DIRTY;
+	uint64		mask = BM_DIRTY;
 	WritebackContext wb_context;
 
 	/*
@@ -3478,7 +3478,7 @@ BufferSync(int flags)
 	for (buf_id = 0; buf_id < NBuffers; buf_id++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
-		uint32		set_bits = 0;
+		uint64		set_bits = 0;
 
 		/*
 		 * Header spinlock is enough to examine BM_DIRTY, see comment in
@@ -3645,7 +3645,7 @@ BufferSync(int flags)
 		 * write the buffer though we didn't need to.  It doesn't seem worth
 		 * guarding against this, though.
 		 */
-		if (pg_atomic_read_u32(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
+		if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
 		{
 			if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
 			{
@@ -4015,7 +4015,7 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 {
 	BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
 	int			result = 0;
-	uint32		buf_state;
+	uint64		buf_state;
 	BufferTag	tag;
 
 	/* Make sure we can handle the pin */
@@ -4264,7 +4264,7 @@ DebugPrintBufferRefcount(Buffer buffer)
 	int32		loccount;
 	char	   *result;
 	ProcNumber	backend;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 	if (BufferIsLocal(buffer))
@@ -4281,9 +4281,9 @@ DebugPrintBufferRefcount(Buffer buffer)
 	}
 
 	/* theoretically we should lock the bufHdr here */
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
-	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%x, refcount=%u %d)",
+	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%" PRIx64 ", refcount=%u %d)",
 					  buffer,
 					  relpathbackend(BufTagGetRelFileLocator(&buf->tag), backend,
 									 BufTagGetForkNum(&buf->tag)).str,
@@ -4383,7 +4383,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	instr_time	io_start;
 	Block		bufBlock;
 	char	   *bufToWrite;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
@@ -4581,7 +4581,7 @@ BufferIsPermanent(Buffer buffer)
 	 * not random garbage.
 	 */
 	bufHdr = GetBufferDescriptor(buffer - 1);
-	return (pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT) != 0;
+	return (pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT) != 0;
 }
 
 /*
@@ -5044,11 +5044,11 @@ FlushRelationBuffers(Relation rel)
 	{
 		for (i = 0; i < NLocBuffer; i++)
 		{
-			uint32		buf_state;
+			uint64		buf_state;
 
 			bufHdr = GetLocalBufferDescriptor(i);
 			if (BufTagMatchesRelFileLocator(&bufHdr->tag, &rel->rd_locator) &&
-				((buf_state = pg_atomic_read_u32(&bufHdr->state)) &
+				((buf_state = pg_atomic_read_u64(&bufHdr->state)) &
 				 (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 			{
 				ErrorContextCallback errcallback;
@@ -5084,7 +5084,7 @@ FlushRelationBuffers(Relation rel)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5156,7 +5156,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 	{
 		SMgrSortArray *srelent = NULL;
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -5405,7 +5405,7 @@ FlushDatabaseBuffers(Oid dbid)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5553,13 +5553,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u32(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
 		(BM_DIRTY | BM_JUST_DIRTIED))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
 		bool		delayChkptFlags = false;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * If we need to protect hint bit updates from torn writes, WAL-log a
@@ -5571,7 +5571,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
 		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT))
+			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5671,8 +5671,8 @@ UnlockBuffers(void)
 
 	if (buf)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		buf_state = LockBufHdr(buf);
 
@@ -5803,8 +5803,8 @@ LockBufferForCleanup(Buffer buffer)
 
 	for (;;)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		/* Try to acquire lock */
 		LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
@@ -5952,7 +5952,7 @@ bool
 ConditionalLockBufferForCleanup(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state,
+	uint64		buf_state,
 				refcount;
 
 	Assert(BufferIsValid(buffer));
@@ -6010,7 +6010,7 @@ bool
 IsBufferCleanupOK(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 
@@ -6066,7 +6066,7 @@ WaitIO(BufferDesc *buf)
 	ConditionVariablePrepareToSleep(cv);
 	for (;;)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 		PgAioWaitRef iow;
 
 		/*
@@ -6140,7 +6140,7 @@ WaitIO(BufferDesc *buf)
 bool
 StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	ResourceOwnerEnlarge(CurrentResourceOwner);
 
@@ -6196,11 +6196,11 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
  * is being released)
  */
 void
-TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 				  bool forget_owner, bool release_aio)
 {
-	uint32		buf_state;
-	uint32		unset_flag_bits = 0;
+	uint64		buf_state;
+	uint64		unset_flag_bits = 0;
 	int			refcount_change = 0;
 
 	buf_state = LockBufHdr(buf);
@@ -6261,7 +6261,7 @@ static void
 AbortBufferIO(Buffer buffer)
 {
 	BufferDesc *buf_hdr = GetBufferDescriptor(buffer - 1);
-	uint32		buf_state;
+	uint64		buf_state;
 
 	buf_state = LockBufHdr(buf_hdr);
 	Assert(buf_state & (BM_IO_IN_PROGRESS | BM_TAG_VALID));
@@ -6355,10 +6355,10 @@ rlocator_comparator(const void *p1, const void *p2)
 /*
  * Lock buffer header - set BM_LOCKED in buffer state.
  */
-uint32
+uint64
 LockBufHdr(BufferDesc *desc)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	Assert(!BufferIsLocal(BufferDescriptorGetBuffer(desc)));
 
@@ -6369,7 +6369,7 @@ LockBufHdr(BufferDesc *desc)
 		 * the spin-delay infrastructure. The work necessary for that shows up
 		 * in profiles and is rarely necessary.
 		 */
-		old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
+		old_buf_state = pg_atomic_fetch_or_u64(&desc->state, BM_LOCKED);
 		if (likely(!(old_buf_state & BM_LOCKED)))
 			break;				/* got lock */
 
@@ -6382,7 +6382,7 @@ LockBufHdr(BufferDesc *desc)
 			while (old_buf_state & BM_LOCKED)
 			{
 				perform_spin_delay(&delayStatus);
-				old_buf_state = pg_atomic_read_u32(&desc->state);
+				old_buf_state = pg_atomic_read_u64(&desc->state);
 			}
 			finish_spin_delay(&delayStatus);
 		}
@@ -6403,20 +6403,20 @@ LockBufHdr(BufferDesc *desc)
  * Obviously the buffer could be locked by the time the value is returned, so
  * this is primarily useful in CAS style loops.
  */
-pg_noinline uint32
+pg_noinline uint64
 WaitBufHdrUnlocked(BufferDesc *buf)
 {
 	SpinDelayStatus delayStatus;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	init_local_spin_delay(&delayStatus);
 
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
 	while (buf_state & BM_LOCKED)
 	{
 		perform_spin_delay(&delayStatus);
-		buf_state = pg_atomic_read_u32(&buf->state);
+		buf_state = pg_atomic_read_u64(&buf->state);
 	}
 
 	finish_spin_delay(&delayStatus);
@@ -6704,12 +6704,12 @@ ResOwnerPrintBufferPin(Datum res)
 static bool
 EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result;
 
 	*buffer_flushed = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -6803,12 +6803,12 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -6855,7 +6855,7 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
@@ -6897,12 +6897,12 @@ static bool
 MarkDirtyUnpinnedBufferInternal(Buffer buf, BufferDesc *desc,
 								bool *buffer_already_dirty)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result = false;
 
 	*buffer_already_dirty = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -7000,7 +7000,7 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_already_dirty;
 
 		CHECK_FOR_INTERRUPTS();
@@ -7054,12 +7054,12 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_already_dirty;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -7110,7 +7110,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		BufferDesc *buf_hdr = is_temp ?
 			GetLocalBufferDescriptor(-buffer - 1)
 			: GetBufferDescriptor(buffer - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * Check that all the buffers are actually ones that could conceivably
@@ -7128,7 +7128,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		}
 
 		if (is_temp)
-			buf_state = pg_atomic_read_u32(&buf_hdr->state);
+			buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		else
 			buf_state = LockBufHdr(buf_hdr);
 
@@ -7166,7 +7166,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		if (is_temp)
 		{
 			buf_state += BUF_REFCOUNT_ONE;
-			pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 		}
 		else
 			UnlockBufHdrExt(buf_hdr, buf_state, 0, 0, 1);
@@ -7352,13 +7352,13 @@ buffer_readv_complete_one(PgAioTargetData *td, uint8 buf_off, Buffer buffer,
 		: GetBufferDescriptor(buffer - 1);
 	BufferTag	tag = buf_hdr->tag;
 	char	   *bufdata = BufferGetBlock(buffer);
-	uint32		set_flag_bits;
+	uint64		set_flag_bits;
 	int			piv_flags;
 
 	/* check that the buffer is in the expected state for a read */
 #ifdef USE_ASSERT_CHECKING
 	{
-		uint32		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(!(buf_state & BM_VALID));
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 9a93fb335fc..b7687836188 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -86,7 +86,7 @@ typedef struct BufferAccessStrategyData
 
 /* Prototypes for internal functions */
 static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
-									 uint32 *buf_state);
+									 uint64 *buf_state);
 static void AddBufferToRing(BufferAccessStrategy strategy,
 							BufferDesc *buf);
 
@@ -171,7 +171,7 @@ ClockSweepTick(void)
  *	before returning.
  */
 BufferDesc *
-StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_ring)
+StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_ring)
 {
 	BufferDesc *buf;
 	int			bgwprocno;
@@ -230,8 +230,8 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 	trycounter = NBuffers;
 	for (;;)
 	{
-		uint32		old_buf_state;
-		uint32		local_buf_state;
+		uint64		old_buf_state;
+		uint64		local_buf_state;
 
 		buf = GetBufferDescriptor(ClockSweepTick());
 
@@ -239,7 +239,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 		 * Check whether the buffer can be used and pin it if so. Do this
 		 * using a CAS loop, to avoid having to lock the buffer header.
 		 */
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			local_buf_state = old_buf_state;
@@ -277,7 +277,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 			{
 				local_buf_state -= BUF_USAGECOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					trycounter = NBuffers;
@@ -289,7 +289,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 				/* pin the buffer if the CAS succeeds */
 				local_buf_state += BUF_REFCOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					/* Found a usable buffer */
@@ -655,12 +655,12 @@ FreeAccessStrategy(BufferAccessStrategy strategy)
  * returning.
  */
 static BufferDesc *
-GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
+GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
 {
 	BufferDesc *buf;
 	Buffer		bufnum;
-	uint32		old_buf_state;
-	uint32		local_buf_state;	/* to avoid repeated (de-)referencing */
+	uint64		old_buf_state;
+	uint64		local_buf_state;	/* to avoid repeated (de-)referencing */
 
 
 	/* Advance to next ring slot */
@@ -682,7 +682,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 	 * Check whether the buffer can be used and pin it if so. Do this using a
 	 * CAS loop, to avoid having to lock the buffer header.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 	for (;;)
 	{
 		local_buf_state = old_buf_state;
@@ -710,7 +710,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 		/* pin the buffer if the CAS succeeds */
 		local_buf_state += BUF_REFCOUNT_ONE;
 
-		if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 										   local_buf_state))
 		{
 			*buf_state = local_buf_state;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index f6e2b1aa288..04a540379a2 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -148,7 +148,7 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 	}
 	else
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		victim_buffer = GetLocalVictimBuffer();
 		bufid = -victim_buffer - 1;
@@ -165,10 +165,10 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 		 */
 		bufHdr->tag = newTag;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
 		buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 		*foundPtr = false;
 	}
@@ -245,12 +245,12 @@ GetLocalVictimBuffer(void)
 
 		if (LocalRefCount[victim_bufid] == 0)
 		{
-			uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 			if (BUF_STATE_GET_USAGECOUNT(buf_state) > 0)
 			{
 				buf_state -= BUF_USAGECOUNT_ONE;
-				pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+				pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 				trycounter = NLocBuffer;
 			}
 			else if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
@@ -286,13 +286,13 @@ GetLocalVictimBuffer(void)
 	 * this buffer is not referenced but it might still be dirty. if that's
 	 * the case, write it out before reusing it!
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY)
 		FlushLocalBuffer(bufHdr, NULL);
 
 	/*
 	 * Remove the victim buffer from the hashtable and mark as invalid.
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID)
 	{
 		InvalidateLocalBuffer(bufHdr, false);
 
@@ -417,7 +417,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 		if (found)
 		{
 			BufferDesc *existing_hdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			UnpinLocalBuffer(BufferDescriptorGetBuffer(victim_buf_hdr));
 
@@ -428,18 +428,18 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 			/*
 			 * Clear the BM_VALID bit, do StartLocalBufferIO() and proceed.
 			 */
-			buf_state = pg_atomic_read_u32(&existing_hdr->state);
+			buf_state = pg_atomic_read_u64(&existing_hdr->state);
 			Assert(buf_state & BM_TAG_VALID);
 			Assert(!(buf_state & BM_DIRTY));
 			buf_state &= ~BM_VALID;
-			pg_atomic_unlocked_write_u32(&existing_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&existing_hdr->state, buf_state);
 
 			/* no need to loop for local buffers */
 			StartLocalBufferIO(existing_hdr, true, false);
 		}
 		else
 		{
-			uint32		buf_state = pg_atomic_read_u32(&victim_buf_hdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&victim_buf_hdr->state);
 
 			Assert(!(buf_state & (BM_VALID | BM_TAG_VALID | BM_DIRTY | BM_JUST_DIRTIED)));
 
@@ -447,7 +447,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 
 			buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 
-			pg_atomic_unlocked_write_u32(&victim_buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&victim_buf_hdr->state, buf_state);
 
 			hresult->id = victim_buf_id;
 
@@ -467,13 +467,13 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 	{
 		Buffer		buf = buffers[i];
 		BufferDesc *buf_hdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		buf_state |= BM_VALID;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 
 	*extended_by = extend_by;
@@ -492,7 +492,7 @@ MarkLocalBufferDirty(Buffer buffer)
 {
 	int			bufid;
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsLocal(buffer));
 
@@ -506,14 +506,14 @@ MarkLocalBufferDirty(Buffer buffer)
 
 	bufHdr = GetLocalBufferDescriptor(bufid);
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	if (!(buf_state & BM_DIRTY))
 		pgBufferUsage.local_blks_dirtied++;
 
 	buf_state |= BM_DIRTY;
 
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -522,7 +522,7 @@ MarkLocalBufferDirty(Buffer buffer)
 bool
 StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * With AIO the buffer could have IO in progress, e.g. when there are two
@@ -542,7 +542,7 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 	/* Once we get here, there is definitely no I/O active on this buffer */
 
 	/* Check if someone else already did the I/O */
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 	if (forInput ? (buf_state & BM_VALID) : !(buf_state & BM_DIRTY))
 	{
 		return false;
@@ -559,11 +559,11 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
  * Like TerminateBufferIO, but for local buffers
  */
 void
-TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bits,
+TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint64 set_flag_bits,
 					   bool release_aio)
 {
 	/* Only need to adjust flags */
-	uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+	uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/* BM_IO_IN_PROGRESS isn't currently used for local buffers */
 
@@ -582,7 +582,7 @@ TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bit
 	}
 
 	buf_state |= set_flag_bits;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 	/* local buffers don't track IO using resowners */
 
@@ -606,7 +606,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(bufHdr);
 	int			bufid = -buffer - 1;
-	uint32		buf_state;
+	uint64		buf_state;
 	LocalBufferLookupEnt *hresult;
 
 	/*
@@ -622,7 +622,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 		Assert(!pgaio_wref_valid(&bufHdr->io_wref));
 	}
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/*
 	 * We need to test not just LocalRefCount[bufid] but also the BufferDesc
@@ -647,7 +647,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 	ClearBufferTag(&bufHdr->tag);
 	buf_state &= ~BUF_FLAG_MASK;
 	buf_state &= ~BUF_USAGECOUNT_MASK;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -671,9 +671,9 @@ DropRelationLocalBuffers(RelFileLocator rlocator, ForkNumber *forkNum,
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (!(buf_state & BM_TAG_VALID) ||
 			!BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -706,9 +706,9 @@ DropRelationAllLocalBuffers(RelFileLocator rlocator)
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if ((buf_state & BM_TAG_VALID) &&
 			BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -804,11 +804,11 @@ InitLocalBuffers(void)
 bool
 PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	Buffer		buffer = BufferDescriptorGetBuffer(buf_hdr);
 	int			bufid = -buffer - 1;
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	if (LocalRefCount[bufid] == 0)
 	{
@@ -819,7 +819,7 @@ PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 		{
 			buf_state += BUF_USAGECOUNT_ONE;
 		}
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/*
 		 * See comment in PinBuffer().
@@ -856,14 +856,14 @@ UnpinLocalBufferNoOwner(Buffer buffer)
 	if (--LocalRefCount[buffid] == 0)
 	{
 		BufferDesc *buf_hdr = GetLocalBufferDescriptor(buffid);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		NLocalPinnedBuffers--;
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state -= BUF_REFCOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/* see comment in UnpinBufferNoOwner */
 		VALGRIND_MAKE_MEM_NOACCESS(LocalBufHdrGetBlock(buf_hdr), BLCKSZ);
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 0c58e4b265c..529803346ce 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -199,7 +199,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
 		for (i = 0; i < NBuffers; i++)
 		{
 			BufferDesc *bufHdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			CHECK_FOR_INTERRUPTS();
 
@@ -615,7 +615,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -626,7 +626,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 		 * noticeably increase the cost of the function.
 		 */
 		bufHdr = GetBufferDescriptor(i);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (buf_state & BM_VALID)
 		{
@@ -676,7 +676,7 @@ pg_buffercache_usage_counts(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		int			usage_count;
 
 		CHECK_FOR_INTERRUPTS();
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index e046b08f3d5..b1aa8af9ec0 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -308,9 +308,9 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 {
 	Buffer		buf;
 	BufferDesc *buf_hdr;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		was_pinned = false;
-	uint32		unset_bits = 0;
+	uint64		unset_bits = 0;
 
 	/* place buffer in shared buffers without erroring out */
 	buf = ReadBufferExtended(rel, MAIN_FORKNUM, blkno, RBM_ZERO_AND_LOCK, NULL);
@@ -319,7 +319,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	}
 	else
 	{
@@ -340,7 +340,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_state &= ~unset_bits;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 	else
 	{
@@ -489,7 +489,7 @@ invalidate_rel_block(PG_FUNCTION_ARGS)
 
 			LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
-			if (pg_atomic_read_u32(&buf_hdr->state) & BM_DIRTY)
+			if (pg_atomic_read_u64(&buf_hdr->state) & BM_DIRTY)
 			{
 				if (BufferIsLocal(buf))
 					FlushLocalBuffer(buf_hdr, NULL);
@@ -572,7 +572,7 @@ buffer_call_terminate_io(PG_FUNCTION_ARGS)
 	bool		io_error = PG_GETARG_BOOL(3);
 	bool		release_aio = PG_GETARG_BOOL(4);
 	bool		clear_dirty = false;
-	uint32		set_flag_bits = 0;
+	uint64		set_flag_bits = 0;
 
 	if (io_error)
 		set_flag_bits |= BM_IO_ERROR;
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v9-0007-bufmgr-Implement-buffer-content-locks-independent.patch (46.0K, ../../4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf/8-v9-0007-bufmgr-Implement-buffer-content-locks-independent.patch)
  download | inline diff:
From ec495a985d0d8172bddc617bf20b2e5fbe37c20a Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 19 Nov 2025 16:37:26 -0500
Subject: [PATCH v9 07/10] bufmgr: Implement buffer content locks independently
 of lwlocks

Until now buffer content locks were implemented using lwlocks. That has the
obvious advantage of not needing a separate efficient implementation of
locks. However, the time for a dedicated buffer content lock implementation
has come:

1) Hint bits are currently set while holding only a share lock. This leads to
   having to copy pages while they are being written out if checksums are
   enabled, which is not cheap. We would like to add AIO writes, however once
   many buffers can be written out at the same time, it gets a lot more
   expensive to copy them, particularly because that copy needs to reside in
   shared buffers (for worker mode to have access to the buffer).

   In addition, modifying buffers while they are being written out can cause
   issues with unbuffered/direct-IO, as some filesystems (like btrfs) do not
   like that, due to filesystem internal checksums getting corrupted.

   The solution to this is to require a new share-exclusive lock-level to set
   hint bits and to write out buffers, making those operations mutually
   exclusive. We could introduce such a lock level into the generic lwlock
   implementation, however it does not look like there would be other users,
   and it does add some overhead into important codepaths.

2) For AIO writes we need to be able to race-freely check whether a buffer is
   undergoing IO and whether an exclusive lock on the page can be acquired. That
   is rather hard to do efficiently when the buffer state and the lock state
   are separate atomic variables. This is a major hindrance to allowing writes
   to be done asynchronously.

3) Buffer locks are by far the most frequently taken locks. Optimizing them
   specifically for their use case is worth the effort. E.g. by merging
   content locks into buffer locks we will be able to release a buffer lock
   and pin in one atomic operation.

4) There are more complicated optimizations, like long-lived "super pinned &
   locked" pages, that cannot realistically be implemented with the generic
   lwlock implementation.

Therefore implement content locks inside bufmgr.c. The lockstate is stored as
part of BufferDesc.state. The implementation of buffer content locks is fairly
similar to lwlocks, with a few important differences:

1) An additional lock-level share-exclusive has been added. This lock level
   conflicts with exclusive locks and itself, but not share locks.

2) Error recovery for content locks is implemented as part of the already
   existing private-refcount tracking mechanism in combination with resowners,
   instead of a bespoke mechanism as the case for lwlocks. This means we do
   not need to add dedicated error-recovery codepaths to release all content
   locks (like done with LWLockReleaseAll() for content locks).

3) The lock state is embedded in BufferDesc.state instead of having its own
   struct.

4) The wakeup logic is a tad more complicated due to needing to support the
   additional lock level

This commit unfortunately introduces some code that is very similar to the
code in lwlock.c, however the code is not equivalent enough to easily merge
it. The future wins that this commit makes possible seem worth the cost.

As of this commit nothing uses the new share-exclusive lock mode. It will be
used in a future commit. It seemed too complicated to introduce the lock-level
in a separate commit.

TODO:
- Address FIXMEs

- Perhaps move the locking code into a buffer_locking.h or such? Needs to be
  inline functions for efficiency unfortunately.

- reflow some comments that I didn't reflow to make the diff more readable

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Reviewed-by: Greg Burd <greg@burd.me>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  67 +-
 src/include/storage/bufmgr.h                  |  32 +-
 src/include/storage/proc.h                    |   8 +-
 src/backend/storage/buffer/buf_init.c         |   5 +-
 src/backend/storage/buffer/bufmgr.c           | 882 ++++++++++++++++--
 .../utils/activity/wait_event_names.txt       |   3 +
 6 files changed, 909 insertions(+), 88 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index a4d36e9ca01..3ea198404b1 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -23,6 +23,7 @@
 #include "storage/condition_variable.h"
 #include "storage/lwlock.h"
 #include "storage/procnumber.h"
+#include "storage/proclist_types.h"
 #include "storage/shmem.h"
 #include "storage/smgr.h"
 #include "storage/spin.h"
@@ -35,22 +36,23 @@
  * State of the buffer itself (in order):
  * - 18 bits refcount
  * - 4 bits usage count
- * - 10 bits of flags
+ * - 12 bits of flags
+ * - 18 bits share-lock-count
+ * - 1 bit share-exclusive locked
+ * - 1 bit exclusive locked
  *
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
- * NB: A future commit will use a significant portion of the remaining bits to
- * implement buffer locking as part of the state variable.
- *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
 #define BUF_USAGECOUNT_BITS 4
-#define BUF_FLAG_BITS 10
+#define BUF_FLAG_BITS 12
+#define BUF_LOCK_BITS (18+2)
 
-StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
-				 "parts of buffer state space need to equal 32");
+StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS + BUF_LOCK_BITS <= 64,
+				 "parts of buffer state space need to be <= 64");
 
 /* refcount related definitions */
 #define BUF_REFCOUNT_ONE 1
@@ -71,6 +73,19 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 #define BUF_FLAG_MASK \
 	(((UINT64CONST(1) << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
 
+/* lock state related definitions */
+#define BM_LOCK_SHIFT \
+	(BUF_FLAG_SHIFT + BUF_FLAG_BITS)
+#define BM_LOCK_VAL_SHARED \
+	(UINT64CONST(1) << (BM_LOCK_SHIFT))
+#define BM_LOCK_VAL_SHARE_EXCLUSIVE \
+	(UINT64CONST(1) << (BM_LOCK_SHIFT + MAX_BACKENDS_BITS))
+#define BM_LOCK_VAL_EXCLUSIVE \
+	(UINT64CONST(1) << (BM_LOCK_SHIFT + MAX_BACKENDS_BITS + 1))
+#define BM_LOCK_MASK \
+	((((uint64) MAX_BACKENDS) << BM_LOCK_SHIFT) | BM_LOCK_VAL_SHARE_EXCLUSIVE | BM_LOCK_VAL_EXCLUSIVE)
+
+
 /* Get refcount and usagecount from buffer state */
 #define BUF_STATE_GET_REFCOUNT(state) \
 	((uint32)((state) & BUF_REFCOUNT_MASK))
@@ -107,6 +122,17 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 #define BM_CHECKPOINT_NEEDED		BUF_DEFINE_FLAG( 8)
 /* permanent buffer (not unlogged, or init fork) */
 #define BM_PERMANENT				BUF_DEFINE_FLAG( 9)
+/* content lock has waiters */
+#define BM_LOCK_HAS_WAITERS			BUF_DEFINE_FLAG(10)
+/* waiter for content lock has been signalled but not yet run */
+#define BM_LOCK_WAKE_IN_PROGRESS	BUF_DEFINE_FLAG(11)
+
+
+StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
+				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
+StaticAssertDecl(MAX_BACKENDS_BITS <= (BUF_LOCK_BITS - 2),
+				 "MAX_BACKENDS_BITS needs to be <= BUF_LOCK_BITS - 2");
+
 
 /*
  * The maximum allowed value of usage_count represents a tradeoff between
@@ -120,8 +146,6 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 
 StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 				 "BM_MAX_USAGE_COUNT doesn't fit in BUF_USAGECOUNT_BITS bits");
-StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
-				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
 
 /*
  * Buffer tag identifies which disk block the buffer contains.
@@ -265,9 +289,6 @@ BufMappingPartitionLockByIndex(uint32 index)
  * it is held.  However, existing buffer pins may be released while the buffer
  * header spinlock is held, using an atomic subtraction.
  *
- * The LWLock can take care of itself.  The buffer header lock is *not* used
- * to control access to the data in the buffer!
- *
  * If we have the buffer pinned, its tag can't change underneath us, so we can
  * examine the tag without locking the buffer header.  Also, in places we do
  * one-time reads of the flags without bothering to lock the buffer header;
@@ -280,6 +301,15 @@ BufMappingPartitionLockByIndex(uint32 index)
  * wait_backend_pgprocno and setting flag bit BM_PIN_COUNT_WAITER.  At present,
  * there can be only one such waiter per buffer.
  *
+ * The content of buffers is protected via the buffer content lock,
+ * implemented as part buffer state. Note that the buffer header lock is *not*
+ * used to control access to the data in the buffer! We used to use an LWLock
+ * to implement the content lock, but having a dedicated implementation of
+ * content locks allows to implement some otherwise hard things (e.g.
+ * race-freely checking if AIO is in progress before locking a buffer
+ * exclusively) and makes otherwise impossible optimizations possible
+ * (e.g. unlocking and unpinning a buffer in one atomic operation).
+ *
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
@@ -321,7 +351,12 @@ typedef struct BufferDesc
 	int			wait_backend_pgprocno;
 
 	PgAioWaitRef io_wref;		/* set iff AIO is in progress */
-	LWLock		content_lock;	/* to lock access to buffer contents */
+
+	/*
+	 * List of PGPROCs waiting for the buffer content lock. Protected by the
+	 * buffer header spinlock.
+	 */
+	proclist_head lock_waiters;
 } BufferDesc;
 
 /*
@@ -408,12 +443,6 @@ BufferDescriptorGetIOCV(const BufferDesc *bdesc)
 	return &(BufferIOCVArray[bdesc->buf_id]).cv;
 }
 
-static inline LWLock *
-BufferDescriptorGetContentLock(const BufferDesc *bdesc)
-{
-	return (LWLock *) (&bdesc->content_lock);
-}
-
 /*
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 715ae96f0f0..a40adf6b2a8 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -203,7 +203,20 @@ extern PGDLLIMPORT int32 *LocalRefCount;
 typedef enum BufferLockMode
 {
 	BUFFER_LOCK_UNLOCK,
+
+	/*
+	 * A share lock conflicts with exclusive locks.
+	 */
 	BUFFER_LOCK_SHARE,
+
+	/*
+	 * A share-exclusive lock conflicts with itself and exclusive locks.
+	 */
+	BUFFER_LOCK_SHARE_EXCLUSIVE,
+
+	/*
+	 * An exclusive lock conflicts with every other lock type.
+	 */
 	BUFFER_LOCK_EXCLUSIVE,
 } BufferLockMode;
 
@@ -302,7 +315,24 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
 extern void UnlockBuffers(void);
-extern void LockBuffer(Buffer buffer, BufferLockMode mode);
+extern void UnlockBuffer(Buffer buffer);
+extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
+
+/*
+ * Handling BUFFER_LOCK_UNLOCK in bufmgr.c leads to sufficiently worse branch
+ * prediction to impact performance. Therefore handle that switch here, where
+ * most of the time `mode` will be a constant and thus can be optimized out by
+ * the compiler.
+ */
+static inline void
+LockBuffer(Buffer buffer, BufferLockMode mode)
+{
+	if (mode == BUFFER_LOCK_UNLOCK)
+		UnlockBuffer(buffer);
+	else
+		LockBufferInternal(buffer, mode);
+}
+
 extern bool ConditionalLockBuffer(Buffer buffer);
 extern void LockBufferForCleanup(Buffer buffer);
 extern bool ConditionalLockBufferForCleanup(Buffer buffer);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index de7b2e0bd2c..039bc8353be 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -242,7 +242,13 @@ struct PGPROC
 	 */
 	bool		recoveryConflictPending;
 
-	/* Info about LWLock the process is currently waiting for, if any. */
+	/*
+	 * Info about LWLock the process is currently waiting for, if any.
+	 *
+	 * This is currently used both for lwlocks and buffer content locks, which
+	 * is acceptable, although not pretty, because a backend can't wait for
+	 * both types of locks at the same time.
+	 */
 	uint8		lwWaiting;		/* see LWLockWaitState */
 	uint8		lwWaitMode;		/* lwlock mode being waited for */
 	proclist_node lwWaitLink;	/* position in LW lock wait list */
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 7d894522526..c0c223b2e32 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
 #include "storage/aio.h"
 #include "storage/buf_internals.h"
 #include "storage/bufmgr.h"
+#include "storage/proclist.h"
 
 BufferDescPadded *BufferDescriptors;
 char	   *BufferBlocks;
@@ -128,9 +129,7 @@ BufferManagerShmemInit(void)
 
 			pgaio_wref_clear(&buf->io_wref);
 
-			LWLockInitialize(BufferDescriptorGetContentLock(buf),
-							 LWTRANCHE_BUFFER_CONTENT);
-
+			proclist_init(&buf->lock_waiters);
 			ConditionVariableInit(BufferDescriptorGetIOCV(buf));
 		}
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index b0de8e45d4d..30852e41862 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -58,6 +58,7 @@
 #include "storage/ipc.h"
 #include "storage/lmgr.h"
 #include "storage/proc.h"
+#include "storage/proclist.h"
 #include "storage/read_stream.h"
 #include "storage/smgr.h"
 #include "storage/standby.h"
@@ -100,6 +101,12 @@ typedef struct PrivateRefCountData
 	 * How many times has the buffer been pinned by this backend.
 	 */
 	int32		refcount;
+
+	/*
+	 * Is the buffer locked by this backend? BUFFER_LOCK_UNLOCK indicates that
+	 * the buffer is not locked.
+	 */
+	BufferLockMode lockmode;
 } PrivateRefCountData;
 
 typedef struct PrivateRefCountEntry
@@ -210,8 +217,10 @@ static BufferDesc *PinCountWaitBuf = NULL;
  * Each buffer also has a private refcount that keeps track of the number of
  * times the buffer is pinned in the current process.  This is so that the
  * shared refcount needs to be modified only once if a buffer is pinned more
- * than once by an individual backend.  It's also used to check that no buffers
- * are still pinned at the end of transactions and when exiting.
+ * than once by an individual backend.  It's also used to check that no
+ * buffers are still pinned at the end of transactions and when exiting. We
+ * also use this mechanism to track whether this backend has a buffer locked,
+ * and, if so, in what mode.
  *
  *
  * To avoid - as we used to - requiring an array with NBuffers entries to keep
@@ -351,6 +360,7 @@ ReservePrivateRefCountEntry(void)
 		/* clear the whole data member, just for future proofing */
 		memset(&victim_entry->data, 0, sizeof(victim_entry->data));
 		victim_entry->data.refcount = 0;
+		victim_entry->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 		PrivateRefCountOverflowed++;
 	}
@@ -374,6 +384,7 @@ NewPrivateRefCountEntry(Buffer buffer)
 	PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
 	res->buffer = buffer;
 	res->data.refcount = 0;
+	res->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 	/* update cache for the next lookup */
 	PrivateRefCountEntryLast = ReservedRefCountSlot;
@@ -540,6 +551,7 @@ static void
 ForgetPrivateRefCountEntry(PrivateRefCountEntry *ref)
 {
 	Assert(ref->data.refcount == 0);
+	Assert(ref->data.lockmode == BUFFER_LOCK_UNLOCK);
 
 	if (ref >= &PrivateRefCountArray[0] &&
 		ref < &PrivateRefCountArray[REFCOUNT_ARRAY_ENTRIES])
@@ -641,14 +653,27 @@ static void RelationCopyStorageUsingBuffer(RelFileLocator srclocator,
 static void AtProcExit_Buffers(int code, Datum arg);
 static void CheckForBufferLeaks(void);
 #ifdef USE_ASSERT_CHECKING
-static void AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
-									   void *unused_context);
+static void AssertNotCatalogBufferLock(Buffer buffer, BufferLockMode mode);
 #endif
 static int	rlocator_comparator(const void *p1, const void *p2);
 static inline int buffertag_comparator(const BufferTag *ba, const BufferTag *bb);
 static inline int ckpt_buforder_comparator(const CkptSortItem *a, const CkptSortItem *b);
 static int	ts_ckpt_progress_comparator(Datum a, Datum b, void *arg);
 
+static void BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr);
+static bool BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMe(BufferDesc *buf_hdr);
+static inline void BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr);
+static inline int BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr);
+static inline bool BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockDequeueSelf(BufferDesc *buf_hdr);
+static void BufferLockWakeup(BufferDesc *buf_hdr, bool unlocked);
+static void BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate);
+static inline uint64 BufferLockReleaseSub(BufferLockMode mode);
+
 
 /*
  * Implementation of PrefetchBuffer() for shared buffers.
@@ -2306,6 +2331,12 @@ retry:
 		goto retry;
 	}
 
+	/*
+	 * An invalidated buffer should not have any backends waiting to lock the
+	 * buffer, therefore BM_LOCK_WAKE_IN_PROGRESS should not be set.
+	 */
+	Assert(!(buf_state & BM_LOCK_WAKE_IN_PROGRESS));
+
 	/*
 	 * Clear out the buffer's tag and flags.  We must do this to ensure that
 	 * linear scans of the buffer array don't think the buffer is valid.
@@ -2382,6 +2413,12 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 		return false;
 	}
 
+	/*
+	 * An invalidated buffer should not have any backends waiting to lock the
+	 * buffer, therefore BM_LOCK_WAKE_IN_PROGRESS should not be set.
+	 */
+	Assert(!(buf_state & BM_LOCK_WAKE_IN_PROGRESS));
+
 	/*
 	 * Clear out the buffer's tag and flags and usagecount.  This is not
 	 * strictly required, as BM_TAG_VALID/BM_VALID needs to be checked before
@@ -2449,8 +2486,6 @@ again:
 	 */
 	if (buf_state & BM_DIRTY)
 	{
-		LWLock	   *content_lock;
-
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(buf_state & BM_VALID);
 
@@ -2468,8 +2503,7 @@ again:
 		 * one just happens to be trying to split the page the first one got
 		 * from StrategyGetBuffer.)
 		 */
-		content_lock = BufferDescriptorGetContentLock(buf_hdr);
-		if (!LWLockConditionalAcquire(content_lock, LW_SHARED))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -2498,7 +2532,7 @@ again:
 			if (XLogNeedsFlush(lsn)
 				&& StrategyRejectBuffer(strategy, buf_hdr, from_ring))
 			{
-				LWLockRelease(content_lock);
+				LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 				UnpinBuffer(buf_hdr);
 				goto again;
 			}
@@ -2506,7 +2540,7 @@ again:
 
 		/* OK, do the I/O */
 		FlushBuffer(buf_hdr, NULL, IOOBJECT_RELATION, io_context);
-		LWLockRelease(content_lock);
+		LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 		ScheduleBufferTagForWriteback(&BackendWritebackContext, io_context,
 									  &buf_hdr->tag);
@@ -2948,7 +2982,7 @@ BufferIsLockedByMe(Buffer buffer)
 	else
 	{
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr));
+		return BufferLockHeldByMe(bufHdr);
 	}
 }
 
@@ -2973,23 +3007,8 @@ BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode)
 	}
 	else
 	{
-		LWLockMode	lw_mode;
-
-		switch (mode)
-		{
-			case BUFFER_LOCK_EXCLUSIVE:
-				lw_mode = LW_EXCLUSIVE;
-				break;
-			case BUFFER_LOCK_SHARE:
-				lw_mode = LW_SHARED;
-				break;
-			default:
-				pg_unreachable();
-		}
-
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-									lw_mode);
+		return BufferLockHeldByMeInMode(bufHdr, mode);
 	}
 }
 
@@ -3376,7 +3395,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 * I'd better not still hold the buffer content lock. Can't use
 		 * BufferIsLockedByMe(), as that asserts the buffer is pinned.
 		 */
-		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
+		Assert(!BufferLockHeldByMe(buf));
 
 		/* decrement the shared reference count */
 		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
@@ -4198,7 +4217,7 @@ CheckForBufferLeaks(void)
  * Check for exclusive-locked catalog buffers.  This is the core of
  * AssertCouldGetRelation().
  *
- * A backend would self-deadlock on LWLocks if the catalog scan read the
+ * A backend would self-deadlock on the content if the catalog scan read the
  * exclusive-locked buffer.  The main threat is exclusive-locked buffers of
  * catalogs used in relcache, because a catcache search on any catalog may
  * build that catalog's relcache entry.  We don't have an inventory of
@@ -4214,26 +4233,45 @@ CheckForBufferLeaks(void)
 void
 AssertBufferLocksPermitCatalogRead(void)
 {
-	ForEachLWLockHeldByMe(AssertNotCatalogBufferLock, NULL);
+	PrivateRefCountEntry *res;
+
+	/* check the array */
+	for (int i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
+	{
+		if (PrivateRefCountArrayKeys[i] != InvalidBuffer)
+		{
+			res = &PrivateRefCountArray[i];
+
+			if (res->buffer == InvalidBuffer)
+				continue;
+
+			AssertNotCatalogBufferLock(res->buffer, res->data.lockmode);
+		}
+	}
+
+	/* if necessary search the hash */
+	if (PrivateRefCountOverflowed)
+	{
+		HASH_SEQ_STATUS hstat;
+
+		hash_seq_init(&hstat, PrivateRefCountHash);
+		while ((res = (PrivateRefCountEntry *) hash_seq_search(&hstat)) != NULL)
+		{
+			AssertNotCatalogBufferLock(res->buffer, res->data.lockmode);
+		}
+	}
 }
 
 static void
-AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
-						   void *unused_context)
+AssertNotCatalogBufferLock(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *bufHdr;
+	BufferDesc *bufHdr = GetBufferDescriptor(buffer - 1);
 	BufferTag	tag;
 	Oid			relid;
 
-	if (mode != LW_EXCLUSIVE)
+	if (mode != BUFFER_LOCK_EXCLUSIVE)
 		return;
 
-	if (!((BufferDescPadded *) lock > BufferDescriptors &&
-		  (BufferDescPadded *) lock < BufferDescriptors + NBuffers))
-		return;					/* not a buffer lock */
-
-	bufHdr = (BufferDesc *)
-		((char *) lock - offsetof(BufferDesc, content_lock));
 	tag = bufHdr->tag;
 
 	/*
@@ -4515,9 +4553,11 @@ static void
 FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 					IOObject io_object, IOContext io_context)
 {
-	LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	Buffer		buffer = BufferDescriptorGetBuffer(buf);
+
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-	LWLockRelease(BufferDescriptorGetContentLock(buf));
+	BufferLockUnlock(buffer, buf);
 }
 
 /*
@@ -5660,9 +5700,10 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
  *
  * Used to clean up after errors.
  *
- * Currently, we can expect that lwlock.c's LWLockReleaseAll() took care
- * of releasing buffer content locks per se; the only thing we need to deal
- * with here is clearing any PIN_COUNT request that was in progress.
+ * Currently, we can expect that resource owner cleanup, via
+ * ResOwnerReleaseBufferPin(), took care releasing buffer content locks per
+ * se; the only thing we need to deal with here is clearing any PIN_COUNT
+ * request that was in progress.
  */
 void
 UnlockBuffers(void)
@@ -5693,25 +5734,725 @@ UnlockBuffers(void)
 }
 
 /*
- * Acquire or release the content_lock for the buffer.
+ * Acquire the buffer content lock in the specified mode
+ *
+ * If the lock is not available, sleep until it is.
+ *
+ * Side effect: cancel/die interrupts are held off until lock release.
+ *
+ * This uses almost the same locking approach as lwlock.c's
+ * LWLockAcquire(). See documentation atop of lwlock.c for a more detailed
+ * discussion.
+ *
+ * The reason that this, and most of the other BufferLock* functions, get both
+ * the Buffer and BufferDesc* as parameters, is that looking up one from the
+ * other repeatedly shows up noticeably in profiles.
+ *
+ * Callers should provide a constant for mode, for more efficient code
+ * generation.
+ */
+static inline void
+BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry;
+	int			extraWaits = 0;
+
+	/*
+	 * Get reference to the refcount entry before we hold the lock, it seems
+	 * better to do before holding the lock.
+	 */
+	entry = GetPrivateRefCountEntry(buffer, true);
+
+	/*
+	 * We better not already hold a lock on the buffer.
+	 */
+	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	for (;;)
+	{
+		bool		mustwait;
+		uint32		wait_event;
+
+		/*
+		 * Try to grab the lock the first time, we're not in the waitqueue
+		 * yet/anymore.
+		 */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		if (likely(!mustwait))
+		{
+			break;
+		}
+
+		/*
+		 * Ok, at this point we couldn't grab the lock on the first try. We
+		 * cannot simply queue ourselves to the end of the list and wait to be
+		 * woken up because by now the lock could long have been released.
+		 * Instead add us to the queue and try to grab the lock again. If we
+		 * succeed we need to revert the queuing and be happy, otherwise we
+		 * recheck the lock. If we still couldn't grab it, we know that the
+		 * other locker will see our queue entries when releasing since they
+		 * existed before we checked for the lock.
+		 */
+
+		/* add to the queue */
+		BufferLockQueueSelf(buf_hdr, mode);
+
+		/* we're now guaranteed to be woken up if necessary */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		/* ok, grabbed the lock the second time round, need to undo queueing */
+		if (!mustwait)
+		{
+			BufferLockDequeueSelf(buf_hdr);
+			break;
+		}
+
+		switch (mode)
+		{
+			case BUFFER_LOCK_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_SHARE_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE:
+				wait_event = WAIT_EVENT_BUFFER_SHARED;
+				break;
+			case BUFFER_LOCK_UNLOCK:
+				pg_unreachable();
+
+		}
+		pgstat_report_wait_start(wait_event);
+
+		/*
+		 * Wait until awakened.
+		 *
+		 * It is possible that we get awakened for a reason other than being
+		 * signaled by LWLockRelease.  If so, loop back and wait again.  Once
+		 * we've gotten the LWLock, re-increment the sema by the number of
+		 * additional signals received.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		pgstat_report_wait_end();
+
+		/* Retrying, allow BufferLockRelease to release waiters again. */
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_WAKE_IN_PROGRESS);
+	}
+
+	/* Remember that we now hold this lock */
+	entry->data.lockmode = mode;
+
+	/*
+	 * Fix the process wait semaphore's count for any absorbed wakeups.
+	 */
+	while (unlikely(extraWaits-- > 0))
+		PGSemaphoreUnlock(MyProc->sem);
+}
+
+/*
+ * Release a previously acquired buffer content lock.
+ */
+static void
+BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	uint64		oldstate;
+	uint64		sub;
+
+	mode = BufferLockDisownInternal(buffer, buf_hdr);
+
+	/*
+	 * Release my hold on lock, after that it can immediately be acquired by
+	 * others, even if we still have to wakeup other waiters.
+	 */
+	sub = BufferLockReleaseSub(mode);
+
+	oldstate = pg_atomic_sub_fetch_u64(&buf_hdr->state, sub);
+
+	BufferLockProcessRelease(buf_hdr, mode, oldstate);
+
+	/*
+	 * Now okay to allow cancel/die interrupts.
+	 */
+	RESUME_INTERRUPTS();
+}
+
+
+/*
+ * Acquire the content lock for the buffer, but only if we don't have to wait.
+ */
+static bool
+BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	bool		mustwait;
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	/* Check for the lock */
+	mustwait = BufferLockAttempt(buf_hdr, mode);
+
+	if (mustwait)
+	{
+		/* Failed to get lock, so release interrupt holdoff */
+		RESUME_INTERRUPTS();
+	}
+	else
+	{
+		PrivateRefCountEntry *entry =
+			GetPrivateRefCountEntry(buffer, true);
+
+		entry->data.lockmode = mode;
+	}
+
+	return !mustwait;
+}
+
+/*
+ * Internal function that tries to atomically acquire the content lock in the
+ * passed in mode.
+ *
+ * This function will not block waiting for a lock to become free - that's the
+ * caller's job.
+ *
+ * Similar to LWLockAttemptLock().
+ */
+static inline bool
+BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	uint64		old_state;
+
+	/*
+	 * Read once outside the loop, later iterations will get the newer value
+	 * via compare & exchange.
+	 */
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+
+	/* loop until we've determined whether we could acquire the lock or not */
+	while (true)
+	{
+		uint64		desired_state;
+		bool		lock_free;
+
+		desired_state = old_state;
+
+		if (mode == BUFFER_LOCK_EXCLUSIVE)
+		{
+			lock_free = (old_state & BM_LOCK_MASK) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_EXCLUSIVE;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			lock_free = (old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+		}
+		else
+		{
+			lock_free = (old_state & BM_LOCK_VAL_EXCLUSIVE) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARED;
+		}
+
+		/*
+		 * Attempt to swap in the state we are expecting. If we didn't see
+		 * lock to be free, that's just the old value. If we saw it as free,
+		 * we'll attempt to mark it acquired. The reason that we always swap
+		 * in the value is that this doubles as a memory barrier. We could try
+		 * to be smarter and only swap in values if we saw the lock as free,
+		 * but benchmark haven't shown it as beneficial so far.
+		 *
+		 * Retry if the value changed since we last looked at it.
+		 */
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			if (lock_free)
+			{
+				/* Great! Got the lock. */
+				return false;
+			}
+			else
+				return true;	/* somebody else has the lock */
+		}
+	}
+
+	pg_unreachable();
+}
+
+/*
+ * Add ourselves to the end of the content lock's wait queue.
+ */
+static void
+BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	/*
+	 * If we don't have a PGPROC structure, there's no way to wait. This
+	 * should never occur, since MyProc should only be null during shared
+	 * memory initialization.
+	 */
+	if (MyProc == NULL)
+		elog(PANIC, "cannot wait without a PGPROC structure");
+
+	if (MyProc->lwWaiting != LW_WS_NOT_WAITING)
+		elog(PANIC, "queueing for lock while waiting on another one");
+
+	LockBufHdr(buf_hdr);
+
+	/* setting the flag is protected by the spinlock */
+	pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_HAS_WAITERS);
+
+	/*
+	 * FIXME: This is reusing the lwlock fields. That's not a correctness
+	 * issue, a backend can't wait for both an lwlock and a buffer content
+	 * lock at the same time. However, it seems pretty ugly, particularly
+	 * given that the field names have an lw* prefix. But duplicating the
+	 * fields also seems somewhat superfluous.
+	 */
+	MyProc->lwWaiting = LW_WS_WAITING;
+	MyProc->lwWaitMode = mode;
+
+	proclist_push_tail(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	/* Can release the mutex now */
+	UnlockBufHdr(buf_hdr);
+}
+
+/*
+ * Remove ourselves from the waitlist.
+ *
+ * This is used if we queued ourselves because we thought we needed to sleep
+ * but, after further checking, we discovered that we don't actually need to
+ * do so.
+ */
+static void
+BufferLockDequeueSelf(BufferDesc *buf_hdr)
+{
+	bool		on_waitlist;
+
+	LockBufHdr(buf_hdr);
+
+	on_waitlist = MyProc->lwWaiting == LW_WS_WAITING;
+	if (on_waitlist)
+		proclist_delete(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	if (proclist_is_empty(&buf_hdr->lock_waiters) &&
+		(pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS) != 0)
+	{
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_HAS_WAITERS);
+	}
+
+	/* XXX: combine with fetch_and above? */
+	UnlockBufHdr(buf_hdr);
+
+	/* clear waiting state again, nice for debugging */
+	if (on_waitlist)
+		MyProc->lwWaiting = LW_WS_NOT_WAITING;
+	else
+	{
+		int			extraWaits = 0;
+
+
+		/*
+		 * Somebody else dequeued us and has or will wake us up. Deal with the
+		 * superfluous absorption of a wakeup.
+		 */
+
+		/*
+		 * Reset RELEASE_OK flag if somebody woke us before we removed
+		 * ourselves - they'll have set it to false.
+		 */
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_WAKE_IN_PROGRESS);
+
+		/*
+		 * Now wait for the scheduled wakeup, otherwise our ->lwWaiting would
+		 * get reset at some inconvenient point later. Most of the time this
+		 * will immediately return.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		/*
+		 * Fix the process wait semaphore's count for any absorbed wakeups.
+		 */
+		while (extraWaits-- > 0)
+			PGSemaphoreUnlock(MyProc->sem);
+	}
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * After calling this function it's the callers responsibility to ensure that
+ * the lock gets released, even in case of an error. This only is desirable if
+ * the lock is going to be released in a different process than the process
+ * that acquired it.
+ */
+static inline void
+BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockDisownInternal(buffer, buf_hdr);
+	RESUME_INTERRUPTS();
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * This is the code that can be shared between actually releasing a lock
+ * (BufferLockUnlock()) and just not tracking ownership of the lock anymore
+ * without releasing the lock (BufferLockDisown()).
+ */
+static inline int
+BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	PrivateRefCountEntry *ref;
+
+	ref = GetPrivateRefCountEntry(buffer, false);
+	if (ref == NULL)
+		elog(ERROR, "lock %d is not held", buffer);
+	mode = ref->data.lockmode;
+	ref->data.lockmode = BUFFER_LOCK_UNLOCK;
+
+	return mode;
+}
+
+/*
+ * Wakeup all the lockers that currently have a chance to acquire the lock.
+ *
+ * wake_exclusive indicates whether exlusive lock waiters should be woken up.
+ */
+static void
+BufferLockWakeup(BufferDesc *buf_hdr, bool wake_exclusive)
+{
+	bool		new_wake_in_progress = false;
+	bool		wake_share_exclusive = true;
+	proclist_head wakeup;
+	proclist_mutable_iter iter;
+
+	proclist_init(&wakeup);
+
+	/* lock wait list while collecting backends to wake up */
+	LockBufHdr(buf_hdr);
+
+	proclist_foreach_modify(iter, &buf_hdr->lock_waiters, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		/*
+		 * Already woke up a conflicting lock, so skip over this wait list
+		 * entry.
+		 */
+		if (!wake_exclusive && waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+			continue;
+		if (!wake_share_exclusive && waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+			continue;
+
+		proclist_delete(&buf_hdr->lock_waiters, iter.cur, lwWaitLink);
+		proclist_push_tail(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Prevent additional wakeups until retryer gets to run. Backends that
+		 * are just waiting for the lock to become free don't retry
+		 * automatically.
+		 */
+		new_wake_in_progress = true;
+
+		/*
+		 * Signal that the process isn't on the wait list anymore. This allows
+		 * BufferLockDequeueSelf() to remove itself from the waitlist with a
+		 * proclist_delete(), rather than having to check if it has been
+		 * removed from the list.
+		 */
+		Assert(waiter->lwWaiting == LW_WS_WAITING);
+		waiter->lwWaiting = LW_WS_PENDING_WAKEUP;
+
+		/*
+		 * Don't wakeup further waiters after waking a conflicting waiter.
+		 */
+		if (waiter->lwWaitMode == BUFFER_LOCK_SHARE)
+		{
+			/*
+			 * Share locks conflict with exclusive locks.
+			 */
+			wake_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			/*
+			 * Share-exclusive locks conflict with share-exclusive eand
+			 * exclusive locks.
+			 */
+			wake_exclusive = false;
+			wake_share_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+		{
+
+			/*
+			 * Exclusive locks conflict with all other locks, there's no point
+			 * in waking up anybody else.
+			 */
+			break;
+		}
+	}
+
+	Assert(proclist_is_empty(&wakeup) || pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS);
+
+	/* unset required flags, and release lock, in one fell swoop */
+	{
+		uint64		old_state;
+		uint64		desired_state;
+
+		old_state = pg_atomic_read_u64(&buf_hdr->state);
+		while (true)
+		{
+			desired_state = old_state;
+
+			/* compute desired flags */
+
+			if (new_wake_in_progress)
+				desired_state |= BM_LOCK_WAKE_IN_PROGRESS;
+			else
+				desired_state &= ~BM_LOCK_WAKE_IN_PROGRESS;
+
+			if (proclist_is_empty(&buf_hdr->lock_waiters))
+				desired_state &= ~BM_LOCK_HAS_WAITERS;
+
+			desired_state &= ~BM_LOCKED;	/* release lock */
+
+			if (pg_atomic_compare_exchange_u64(&buf_hdr->state, &old_state,
+											   desired_state))
+				break;
+		}
+	}
+
+	/* Awaken any waiters I removed from the queue. */
+	proclist_foreach_modify(iter, &wakeup, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		proclist_delete(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Guarantee that lwWaiting being unset only becomes visible once the
+		 * unlink from the link has completed. Otherwise the target backend
+		 * could be woken up for other reason and enqueue for a new lock - if
+		 * that happens before the list unlink happens, the list would end up
+		 * being corrupted.
+		 *
+		 * The barrier pairs with the LWLockWaitListLock() when enqueuing for
+		 * another lock.
+		 */
+		pg_write_barrier();
+		waiter->lwWaiting = LW_WS_NOT_WAITING;
+		PGSemaphoreUnlock(waiter->sem);
+	}
+}
+
+/*
+ * Compute subtraction from buffer state for a release of a held lock in
+ * `mode`.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places, each needing to compute what to
+ * subtract from the lock state.
+ */
+static inline uint64
+BufferLockReleaseSub(BufferLockMode mode)
+{
+
+	/*
+	 * Turns out that a switch() leads gcc to generate sufficiently worse code
+	 * for this to show up in profiles...
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE)
+		return BM_LOCK_VAL_EXCLUSIVE;
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		return BM_LOCK_VAL_SHARE_EXCLUSIVE;
+	else
+	{
+		Assert(mode == BUFFER_LOCK_SHARE);
+		return BM_LOCK_VAL_SHARED;
+	}
+
+	return 0;
+}
+
+/*
+ * Handle work that needs to be done after releasing a lock that was held in
+ * `mode`, where `lockstate` is the result of the atomic operation modifying
+ * the state variable.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places.
+ */
+static void
+BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate)
+{
+	bool		check_waiters = false;
+	bool		wake_exclusive = false;
+
+	/* nobody else can have that kind of lock */
+	Assert(!(lockstate & BM_LOCK_VAL_EXCLUSIVE));
+
+	/*
+	 * If we're still waiting for backends to get scheduled, don't wake them
+	 * up again. Otherwise check if we need to look through the waitqueue to
+	 * wake other backends.
+	 */
+	if ((lockstate & BM_LOCK_HAS_WAITERS) &&
+		!(lockstate & BM_LOCK_WAKE_IN_PROGRESS))
+	{
+		if ((lockstate & BM_LOCK_MASK) == 0)
+		{
+			/*
+			 * We released a lock and the lock was, in that moment, free. We
+			 * therefore can wake waiters for any kind of lock.
+			 */
+			check_waiters = true;
+			wake_exclusive = true;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			/*
+			 * We released the lock, but another backend still holds a lock.
+			 * We can't have released an exclusive lock, as there couldn't
+			 * have been other lock holders. If we released a share lock, no
+			 * waiters need to be woken up, as there must be other share
+			 * lockers. However, if we held a share-exclusive lock, another
+			 * backend now could acquire a share-exclusive lock.
+			 */
+			check_waiters = true;
+			wake_exclusive = false;
+		}
+	}
+
+	/*
+	 * As waking up waiters requires the spinlock to be acquired, only do so
+	 * if necessary.
+	 */
+	if (check_waiters)
+		BufferLockWakeup(buf_hdr, wake_exclusive);
+}
+
+/*
+ * BufferLockHeldByMeInMode - test whether my process holds the content lock
+ * in the specified mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode == mode;
+
+}
+
+/*
+ * BufferLockHeldByMe - test whether my process holds the content lock in any
+ * mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMe(BufferDesc *buf_hdr)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode != BUFFER_LOCK_UNLOCK;
+}
+
+/*
+ * Release the content lock for the buffer.
+ */
+void
+UnlockBuffer(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+
+	Assert(BufferIsPinned(buffer));
+	if (BufferIsLocal(buffer))
+		return;					/* local buffers need no lock */
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+	BufferLockUnlock(buffer, buf_hdr);
+}
+
+/*
+ * Acquire the content_lock for the buffer.
  */
 void
-LockBuffer(Buffer buffer, BufferLockMode mode)
+LockBufferInternal(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *buf;
+	BufferDesc *buf_hdr;
+
+	/*
+	 * We can't wait if we haven't got a PGPROC.  This should only occur
+	 * during bootstrap or shared memory initialization.  Put an Assert here
+	 * to catch unsafe coding practices.
+	 */
+	Assert(!(MyProc == NULL && IsUnderPostmaster));
+
+	/* handled in LockBuffer() wrapper */
+	Assert(mode != BUFFER_LOCK_UNLOCK);
 
 	Assert(BufferIsPinned(buffer));
 	if (BufferIsLocal(buffer))
 		return;					/* local buffers need no lock */
 
-	buf = GetBufferDescriptor(buffer - 1);
+	buf_hdr = GetBufferDescriptor(buffer - 1);
 
-	if (mode == BUFFER_LOCK_UNLOCK)
-		LWLockRelease(BufferDescriptorGetContentLock(buf));
-	else if (mode == BUFFER_LOCK_SHARE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	/*
+	 * Test the most frequent lock modes first. While a switch (mode) would be
+	 * nice, at least gcc generates considerably worse code for it.
+	 *
+	 * Call BufferLockAcquire() with a constant argument for mode, to generate
+	 * more efficient code for the different lock modes.
+	 */
+	if (mode == BUFFER_LOCK_SHARE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE);
 	else if (mode == BUFFER_LOCK_EXCLUSIVE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_EXCLUSIVE);
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_EXCLUSIVE);
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	else
 		elog(ERROR, "unrecognized buffer lock mode: %d", mode);
 }
@@ -5732,8 +6473,7 @@ ConditionalLockBuffer(Buffer buffer)
 
 	buf = GetBufferDescriptor(buffer - 1);
 
-	return LWLockConditionalAcquire(BufferDescriptorGetContentLock(buf),
-									LW_EXCLUSIVE);
+	return BufferLockConditional(buffer, buf, BUFFER_LOCK_EXCLUSIVE);
 }
 
 /*
@@ -6688,7 +7428,25 @@ ResOwnerReleaseBufferPin(Datum res)
 	if (BufferIsLocal(buffer))
 		UnpinLocalBufferNoOwner(buffer);
 	else
+	{
+		PrivateRefCountEntry *ref;
+
+		ref = GetPrivateRefCountEntry(buffer, false);
+
+		/*
+		 * If the buffer was locked at the time of the resowner release,
+		 * release the lock now. This should only happen after errors.
+		 */
+		if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
+		{
+			BufferDesc *buf = GetBufferDescriptor(buffer - 1);
+
+			HOLD_INTERRUPTS();	/* match the upcoming RESUME_INTERRUPTS */
+			BufferLockUnlock(buffer, buf);
+		}
+
 		UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
+	}
 }
 
 static char *
@@ -6924,10 +7682,10 @@ MarkDirtyUnpinnedBufferInternal(Buffer buf, BufferDesc *desc,
 	/* If it was not already dirty, mark it as dirty. */
 	if (!(buf_state & BM_DIRTY))
 	{
-		LWLockAcquire(BufferDescriptorGetContentLock(desc), LW_EXCLUSIVE);
+		BufferLockAcquire(buf, desc, BUFFER_LOCK_EXCLUSIVE);
 		MarkBufferDirty(buf);
 		result = true;
-		LWLockRelease(BufferDescriptorGetContentLock(desc));
+		BufferLockUnlock(buf, desc);
 	}
 	else
 		*buffer_already_dirty = true;
@@ -7178,16 +7936,12 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		 */
 		if (is_write && !is_temp)
 		{
-			LWLock	   *content_lock;
-
-			content_lock = BufferDescriptorGetContentLock(buf_hdr);
-
-			Assert(LWLockHeldByMe(content_lock));
+			Assert(BufferLockHeldByMe(buf_hdr));
 
 			/*
 			 * Lock is now owned by AIO subsystem.
 			 */
-			LWLockDisown(content_lock);
+			BufferLockDisown(buffer, buf_hdr);
 		}
 
 		/*
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 3299de23bb3..a2785492bf7 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -287,6 +287,9 @@ ABI_compatibility:
 Section: ClassName - WaitEventBuffer
 
 BUFFER_CLEANUP	"Waiting to acquire an exclusive pin on a buffer. Buffer pin waits can be protracted if another process holds an open cursor that last read data from the buffer in question."
+BUFFER_SHARED	"Waiting to acquire shared lock on a buffer."
+BUFFER_SHARE_EXCLUSIVE	"Waiting to acquire share exclusive lock on a buffer."
+BUFFER_EXCLUSIVE	"Waiting to acquire exclusive lock on a buffer."
 
 ABI_compatibility:
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v9-0008-Require-share-exclusive-lock-to-set-hint-bits-and.patch (37.8K, ../../4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf/9-v9-0008-Require-share-exclusive-lock-to-set-hint-bits-and.patch)
  download | inline diff:
From 0de571fe89b196917884ffb0e95cc9f3121fbb45 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Fri, 12 Dec 2025 15:31:01 -0500
Subject: [PATCH v9 08/10] Require share-exclusive lock to set hint bits and to
 flush

At the moment hint bits can be set with just a share lock on a page (and in
one place even without any lock). Because of this we need to copy pages while
writing them out, as otherwise the checksum could be corrupted.

The need to copy the page is problematic to implement AIO writes:

1) Instead of just needing a single buffer for a copied page we need one for
   each page that's potentially undergoing IO
2) To be able to use the "worker" AIO implementation the copied page needs to
   reside in shared memory

It also causes problems for using unbuffered/direct-IO, independent of AIO:
Some filesystems, raid implementations, ... do not tolerate the data being
written out to change during the write. E.g. they may compute internal
checksums that can be invalidated by concurrent modifications, leading e.g. to
filesystem errors (as the case with btrfs).

It also just is plain odd to allow modifications of buffers that are just
share locked.

To address these issue, this commit changes the rules so that modifications to
pages are not allowed anymore while holding a share lock. Instead the new
share-exclusive lock (introduced in FIXME XXXX TODO) allows at most one
backend to modify a buffer while other backends have the same page share
locked. An existing share-lock can be upgraded to a share-exclusive lock, if
there are no conflicting locks. For that
BufferBeginSetHintBits()/BufferBeginSetHintBits() and BufferSetHintBits16()
have been introduced.

To prevent hint bits from being set while the buffer is being written out,
writing out buffers now requires a share-exclusive lock.

The use of share-exclusive to gate setting hint bits means that from now on
only one backend can set hint bits at a time. To allow multiple backends
setting hint bits would require more complicated locking, for setting hint
bits we'd need to store the count of backends currently setting hint bits and
we would need another lock-level for I/O conflicting with the lock-level to
set hint bits. Given that the share-exclusive lock for setting hint bits is
only held for a short time, that often backends would just set the same hint
bits and that the cost of occasionally not setting hint bits in hotly accessed
pages is fairly low, this seems like an acceptable tradeoff.

The biggest change to adapt to this is in heapam. To avoid performance
regressions for sequential scans that need to set a lot of hint bits, we need
to amortize the cost of BufferBeginSetHintBits() for cases where hint bits are
set at a high frequency, HeapTupleSatisfiesMVCCBatch() uses the new
SetHintBitsExt() which defers BufferFinishSetHintBits() until all hint bits on
a page have been set.  Conversely, to avoid regressions in cases where we
can't set hint bits in bulk (because we're looking only at individual tuples),
use BufferSetHintBits16() when setting hint bits without batching.

Several other places also need to be adapted, but those changes are
comparatively simpler.

After this we do not need to copy buffers to write them out anymore. That
change is done separately however.

TODO:
- Update commit reference above
- reflow parts of storage/buffer/README that I didn't reindent to make the
  diff more readable

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf%40gcnactj4z56m
---
 src/include/storage/bufmgr.h                |   4 +
 src/backend/access/gist/gistget.c           |  19 +-
 src/backend/access/hash/hashutil.c          |  10 +-
 src/backend/access/heap/heapam_visibility.c | 125 ++++++--
 src/backend/access/nbtree/nbtinsert.c       |  28 +-
 src/backend/access/nbtree/nbtutils.c        |  16 +-
 src/backend/storage/buffer/README           |  44 ++-
 src/backend/storage/buffer/bufmgr.c         | 307 ++++++++++++++++----
 src/backend/storage/freespace/freespace.c   |  14 +-
 src/backend/storage/freespace/fsmpage.c     |  11 +-
 src/tools/pgindent/typedefs.list            |   1 +
 11 files changed, 459 insertions(+), 120 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index a40adf6b2a8..4017896f951 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -314,6 +314,10 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
+extern bool BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer);
+extern bool BufferBeginSetHintBits(Buffer buffer);
+extern void BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std);
+
 extern void UnlockBuffers(void);
 extern void UnlockBuffer(Buffer buffer);
 extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
diff --git a/src/backend/access/gist/gistget.c b/src/backend/access/gist/gistget.c
index 6d05a5fdc34..05972265250 100644
--- a/src/backend/access/gist/gistget.c
+++ b/src/backend/access/gist/gistget.c
@@ -63,11 +63,7 @@ gistkillitems(IndexScanDesc scan)
 	 * safe.
 	 */
 	if (BufferGetLSNAtomic(buffer) != so->curPageLSN)
-	{
-		UnlockReleaseBuffer(buffer);
-		so->numKilled = 0;		/* reset counter */
-		return;
-	}
+		goto unlock;
 
 	Assert(GistPageIsLeaf(page));
 
@@ -77,6 +73,16 @@ gistkillitems(IndexScanDesc scan)
 	 */
 	for (i = 0; i < so->numKilled; i++)
 	{
+		if (!killedsomething)
+		{
+			/*
+			 * Use hint bit infrastructure to be allowed to modify the page
+			 * without holding an exclusive lock.
+			 */
+			if (!BufferBeginSetHintBits(buffer))
+				goto unlock;
+		}
+
 		offnum = so->killedItems[i];
 		iid = PageGetItemId(page, offnum);
 		ItemIdMarkDead(iid);
@@ -86,9 +92,10 @@ gistkillitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		GistMarkPageHasGarbage(page);
-		MarkBufferDirtyHint(buffer, true);
+		BufferFinishSetHintBits(buffer, true, true);
 	}
 
+unlock:
 	UnlockReleaseBuffer(buffer);
 
 	/*
diff --git a/src/backend/access/hash/hashutil.c b/src/backend/access/hash/hashutil.c
index cf7f0b90176..b917c97321a 100644
--- a/src/backend/access/hash/hashutil.c
+++ b/src/backend/access/hash/hashutil.c
@@ -593,6 +593,13 @@ _hash_kill_items(IndexScanDesc scan)
 
 			if (ItemPointerEquals(&ituple->t_tid, &currItem->heapTid))
 			{
+				/*
+				 * Use hint bit infrastructure to be allowed to modify the
+				 * page without holding an exclusive lock.
+				 */
+				if (!BufferBeginSetHintBits(so->currPos.buf))
+					goto unlock_page;
+
 				/* found the item */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -610,9 +617,10 @@ _hash_kill_items(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->hasho_flag |= LH_PAGE_HAS_DEAD_TUPLES;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(so->currPos.buf, true, true);
 	}
 
+unlock_page:
 	if (so->hashso_bucket_buf == so->currPos.buf ||
 		havePin)
 		LockBuffer(so->currPos.buf, BUFFER_LOCK_UNLOCK);
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 5d56c6e5075..9e9ec45e7c5 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -80,10 +80,38 @@
 
 
 /*
- * SetHintBits()
+ * To be allowed to set hint bits, SetHintBits() needs to call
+ * BufferBeginSetHintBits(). However, that's not free, and some callsites call
+ * SetHintBits() on many tuples in a row. For those it makes sense to amortize
+ * the cost of BufferBeginSetHintBits(). Additionally it's desirable to defer
+ * the cost of BufferBeginSetHintBits() until a hint bit needs to actually be
+ * set. This enum serves as the necessary state space passed to
+ * SetHintbitsExt().
+ */
+typedef enum SetHintBitsState
+{
+	/* not yet checked if hint bits may be set */
+	SHB_INITIAL,
+	/* failed to get permission to set hint bits, don't check again */
+	SHB_DISABLED,
+	/* allowed to set hint bits */
+	SHB_ENABLED,
+} SetHintBitsState;
+
+/*
+ * SetHintBitsExt()
  *
  * Set commit/abort hint bits on a tuple, if appropriate at this time.
  *
+ * To be allowed to set a hint bit on a tuple, the page must not be undergoing
+ * IO at this time (otherwise we e.g. could corrupt PG's page checksum or even
+ * the filesystem's, as is known to happen with btrfs).
+ *
+ * The right to set a hint bit can be acquired on a page level with
+ * BufferBeginSetHintBits(). Only a single backend gets the right to set hint
+ * bits at a time.  Alternatively, if called with a NULL SetHintBitsState*,
+ * hint bits are set with BufferSetHintBits16().
+ *
  * It is only safe to set a transaction-committed hint bit if we know the
  * transaction's commit record is guaranteed to be flushed to disk before the
  * buffer, or if the table is temporary or unlogged and will be obliterated by
@@ -111,24 +139,69 @@
  * InvalidTransactionId if no check is needed.
  */
 static inline void
-SetHintBits(HeapTupleHeader tuple, Buffer buffer,
-			uint16 infomask, TransactionId xid)
+SetHintBitsExt(HeapTupleHeader tuple, Buffer buffer,
+			   uint16 infomask, TransactionId xid, SetHintBitsState *state)
 {
+	/*
+	 * In batched mode and we previously did not get permission to set hint
+	 * bits. Don't try again, in all likelihood IO is still going on.
+	 */
+	if (state && *state == SHB_DISABLED)
+		return;
+
 	if (TransactionIdIsValid(xid))
 	{
-		/* NB: xid must be known committed here! */
-		XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+		if (BufferIsPermanent(buffer))
+		{
+			/* NB: xid must be known committed here! */
+			XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+
+			if (XLogNeedsFlush(commitLSN) &&
+				BufferGetLSNAtomic(buffer) < commitLSN)
+			{
+				/* not flushed and no LSN interlock, so don't set hint */
+				return;
+			}
+		}
+	}
+
+	/*
+	 * If we're not operating in batch mode, use BufferSetHintBits16() to mark
+	 * the page dirty, that's cheaper than
+	 * BufferBeginSetHintBits()/BufferFinishSetHintBits(). That's important
+	 * for cases where we set a lot of hint bits on a page individually.
+	 */
+	if (!state)
+	{
+		BufferSetHintBits16(&tuple->t_infomask,
+							tuple->t_infomask | infomask, buffer);
+		return;
+	}
 
-		if (BufferIsPermanent(buffer) && XLogNeedsFlush(commitLSN) &&
-			BufferGetLSNAtomic(buffer) < commitLSN)
+	if (*state == SHB_INITIAL)
+	{
+		if (!BufferBeginSetHintBits(buffer))
 		{
-			/* not flushed and no LSN interlock, so don't set hint */
+			*state = SHB_DISABLED;
 			return;
 		}
+
+		if (state)
+			*state = SHB_ENABLED;
+
 	}
-
 	tuple->t_infomask |= infomask;
-	MarkBufferDirtyHint(buffer, true);
+}
+
+/*
+ * Simple wrapper around SetHintBitExt(), use when operating on a single
+ * tuple.
+ */
+static inline void
+SetHintBits(HeapTupleHeader tuple, Buffer buffer,
+			uint16 infomask, TransactionId xid)
+{
+	SetHintBitsExt(tuple, buffer, infomask, xid, NULL);
 }
 
 /*
@@ -864,9 +937,9 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
  * inserting/deleting transaction was still running --- which was more cycles
  * and more contention on ProcArrayLock.
  */
-static bool
+static inline bool
 HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
-					   Buffer buffer)
+					   Buffer buffer, SetHintBitsState *state)
 {
 	HeapTupleHeader tuple = htup->t_data;
 
@@ -921,8 +994,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 			if (!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmax(tuple)))
 			{
 				/* deleting subtransaction must have aborted */
-				SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-							InvalidTransactionId);
+				SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+							   InvalidTransactionId, state);
 				return true;
 			}
 
@@ -934,13 +1007,13 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		else if (XidInMVCCSnapshot(HeapTupleHeaderGetRawXmin(tuple), snapshot))
 			return false;
 		else if (TransactionIdDidCommit(HeapTupleHeaderGetRawXmin(tuple)))
-			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-						HeapTupleHeaderGetRawXmin(tuple));
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_COMMITTED,
+						   HeapTupleHeaderGetRawXmin(tuple), state);
 		else
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_INVALID,
+						   InvalidTransactionId, state);
 			return false;
 		}
 	}
@@ -1003,14 +1076,14 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		if (!TransactionIdDidCommit(HeapTupleHeaderGetRawXmax(tuple)))
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+						   InvalidTransactionId, state);
 			return true;
 		}
 
 		/* xmax transaction committed */
-		SetHintBits(tuple, buffer, HEAP_XMAX_COMMITTED,
-					HeapTupleHeaderGetRawXmax(tuple));
+		SetHintBitsExt(tuple, buffer, HEAP_XMAX_COMMITTED,
+					   HeapTupleHeaderGetRawXmax(tuple), state);
 	}
 	else
 	{
@@ -1619,6 +1692,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 							OffsetNumber *vistuples_dense)
 {
 	int			nvis = 0;
+	SetHintBitsState state = SHB_INITIAL;
 
 	Assert(IsMVCCSnapshot(snapshot));
 
@@ -1627,7 +1701,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		bool		valid;
 		HeapTuple	tup = &batchmvcc->tuples[i];
 
-		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer, &state);
 		batchmvcc->visible[i] = valid;
 
 		if (likely(valid))
@@ -1637,6 +1711,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		}
 	}
 
+	if (state == SHB_ENABLED)
+		BufferFinishSetHintBits(buffer, true, true);
+
 	return nvis;
 }
 
@@ -1656,7 +1733,7 @@ HeapTupleSatisfiesVisibility(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 	switch (snapshot->snapshot_type)
 	{
 		case SNAPSHOT_MVCC:
-			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer);
+			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer, NULL);
 		case SNAPSHOT_SELF:
 			return HeapTupleSatisfiesSelf(htup, snapshot, buffer);
 		case SNAPSHOT_ANY:
diff --git a/src/backend/access/nbtree/nbtinsert.c b/src/backend/access/nbtree/nbtinsert.c
index 63eda08f7a2..da43af3ec96 100644
--- a/src/backend/access/nbtree/nbtinsert.c
+++ b/src/backend/access/nbtree/nbtinsert.c
@@ -681,20 +681,28 @@ _bt_check_unique(Relation rel, BTInsertState insertstate, Relation heapRel,
 				{
 					/*
 					 * The conflicting tuple (or all HOT chains pointed to by
-					 * all posting list TIDs) is dead to everyone, so mark the
-					 * index entry killed.
+					 * all posting list TIDs) is dead to everyone, so try to
+					 * mark the index entry killed. It's ok if we're not
+					 * allowed to, this isn't required for correctness.
 					 */
-					ItemIdMarkDead(curitemid);
-					opaque->btpo_flags |= BTP_HAS_GARBAGE;
+					Buffer		buf;
 
-					/*
-					 * Mark buffer with a dirty hint, since state is not
-					 * crucial. Be sure to mark the proper buffer dirty.
-					 */
+					/* Be sure to operate on the proper buffer */
 					if (nbuf != InvalidBuffer)
-						MarkBufferDirtyHint(nbuf, true);
+						buf = nbuf;
 					else
-						MarkBufferDirtyHint(insertstate->buf, true);
+						buf = insertstate->buf;
+
+					/*
+					 * Can't use BufferSetHintBits16() here as we update two
+					 * different locations.
+					 */
+					if (BufferBeginSetHintBits(buf))
+					{
+						ItemIdMarkDead(curitemid);
+						opaque->btpo_flags |= BTP_HAS_GARBAGE;
+						BufferFinishSetHintBits(buf, true, true);
+					}
 				}
 
 				/*
diff --git a/src/backend/access/nbtree/nbtutils.c b/src/backend/access/nbtree/nbtutils.c
index 5c50f0dd1bd..a76d90f2d8e 100644
--- a/src/backend/access/nbtree/nbtutils.c
+++ b/src/backend/access/nbtree/nbtutils.c
@@ -357,10 +357,19 @@ _bt_killitems(IndexScanDesc scan)
 			 * it's possible that multiple processes attempt to do this
 			 * simultaneously, leading to multiple full-page images being sent
 			 * to WAL (if wal_log_hints or data checksums are enabled), which
-			 * is undesirable.
+			 * is undesirable.  We need to use the hint bit infrastructure to
+			 * update the page while just holding a share lock.
 			 */
 			if (killtuple && !ItemIdIsDead(iid))
 			{
+				/*
+				 * If we're not able to set hint bits, there's no point
+				 * continuing.
+				 */
+				if (!killedsomething &&
+					!BufferBeginSetHintBits(buf))
+					goto unlock_page;
+
 				/* found the item/all posting list items */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -371,8 +380,6 @@ _bt_killitems(IndexScanDesc scan)
 	}
 
 	/*
-	 * Since this can be redone later if needed, mark as dirty hint.
-	 *
 	 * Whenever we mark anything LP_DEAD, we also set the page's
 	 * BTP_HAS_GARBAGE flag, which is likewise just a hint.  (Note that we
 	 * only rely on the page-level flag in !heapkeyspace indexes.)
@@ -380,9 +387,10 @@ _bt_killitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->btpo_flags |= BTP_HAS_GARBAGE;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(buf, true, true);
 	}
 
+unlock_page:
 	if (!so->dropPin)
 		_bt_unlockbuf(rel, buf);
 	else
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index 119f31b5d65..5e6735edd23 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -25,21 +25,26 @@ that might need to do such a wait is instead handled by waiting to obtain
 the relation-level lock, which is why you'd better hold one first.)  Pins
 may not be held across transaction boundaries, however.
 
-Buffer content locks: there are two kinds of buffer lock, shared and exclusive,
-which act just as you'd expect: multiple backends can hold shared locks on
-the same buffer, but an exclusive lock prevents anyone else from holding
-either shared or exclusive lock.  (These can alternatively be called READ
-and WRITE locks.)  These locks are intended to be short-term: they should not
-be held for long.  Buffer locks are acquired and released by LockBuffer().
-It will *not* work for a single backend to try to acquire multiple locks on
-the same buffer.  One must pin a buffer before trying to lock it.
+Buffer content locks: there three kinds of buffer lock, shared,
+share-exclusive and exclusive:
+a) multiple backends can hold shared locks on the same buffer
+   (alternatively called a READ lock)
+b) one backend can hold an share-exclusive lock on a buffer while multiple
+   backends can hold a share lock
+c) an exclusive lock prevents anyone else from holding shared, share-exclusive
+   or exclusive lock.
+   (alternatively called a WRITE lock)
+
+These locks are intended to be short-term: they should not be held for long.
+Buffer locks are acquired and released by LockBuffer().  It will *not* work
+for a single backend to try to acquire multiple locks on the same buffer.  One
+must pin a buffer before trying to lock it.
 
 Buffer access rules:
 
-1. To scan a page for tuples, one must hold a pin and either shared or
-exclusive content lock.  To examine the commit status (XIDs and status bits)
-of a tuple in a shared buffer, one must likewise hold a pin and either shared
-or exclusive lock.
+1. To scan a page for tuples, one must hold a pin and at least a share lock.
+To examine the commit status (XIDs and status bits) of a tuple in a shared
+buffer, one must likewise hold a pin and at least a share lock.
 
 2. Once one has determined that a tuple is interesting (visible to the
 current transaction) one may drop the content lock, yet continue to access
@@ -55,8 +60,14 @@ one must hold a pin and an exclusive content lock on the containing buffer.
 This ensures that no one else might see a partially-updated state of the
 tuple while they are doing visibility checks.
 
-4. It is considered OK to update tuple commit status bits (ie, OR the
-values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
+4. Non-critical information on a page ("hint bits") may be modified while
+holding only a share-exclusive lock and pin on the page. To do so in cases
+where only a share lock is already held, use BufferBeginSetHintBits() &
+BufferFinishSetHintBits() (if multiple hint bits are to be set) or
+BufferSetHintBits16() (if a single hit bit is set).
+
+E.g. for heapam, a share-exclusive lock allows to update tuple commit status
+bits (ie, OR the values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
 HEAP_XMAX_INVALID into t_infomask) while holding only a shared lock and
 pin on a buffer.  This is OK because another backend looking at the tuple
 at about the same time would OR the same bits into the field, so there
@@ -80,7 +91,6 @@ buffer (increment the refcount) while one is performing the cleanup, but
 it won't be able to actually examine the page until it acquires shared
 or exclusive content lock.
 
-
 Obtaining the lock needed under rule #5 is done by the bufmgr routines
 LockBufferForCleanup() or ConditionalLockBufferForCleanup().  They first get
 an exclusive lock and then check to see if the shared pin count is currently
@@ -96,6 +106,10 @@ VACUUM's use, since we don't allow multiple VACUUMs concurrently on a single
 relation anyway.  Anyone wishing to obtain a cleanup lock outside of recovery
 or a VACUUM must use the conditional variant of the function.
 
+6. To write out a buffer, a share-exclusive lock needs to be held. This
+prevents the buffer from being modified while written out, which could corrupt
+checksums and cause issues on the OS or device level when direct-IO is used.
+
 
 Buffer Manager's Internal Locking
 ---------------------------------
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 30852e41862..e944c09fb33 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2480,9 +2480,8 @@ again:
 	/*
 	 * If the buffer was dirty, try to write it out.  There is a race
 	 * condition here, in that someone might dirty it after we released the
-	 * buffer header lock above, or even while we are writing it out (since
-	 * our share-lock won't prevent hint-bit updates).  We will recheck the
-	 * dirty bit after re-locking the buffer header.
+	 * buffer header lock above.  We will recheck the dirty bit after
+	 * re-locking the buffer header.
 	 */
 	if (buf_state & BM_DIRTY)
 	{
@@ -2490,12 +2489,12 @@ again:
 		Assert(buf_state & BM_VALID);
 
 		/*
-		 * We need a share-lock on the buffer contents to write it out (else
+		 * We need a share-exclusive lock on the buffer contents to write it out (else
 		 * we might write invalid data, eg because someone else is compacting
 		 * the page contents while we write).  We must use a conditional lock
 		 * acquisition here to avoid deadlock.  Even though the buffer was not
 		 * pinned (and therefore surely not locked) when StrategyGetBuffer
-		 * returned it, someone else could have pinned and exclusive-locked it
+		 * returned it, someone else could have pinned and (share-)exclusive-locked it
 		 * by the time we get here. If we try to get the lock unconditionally,
 		 * we'd block waiting for them; if they later block waiting for us,
 		 * deadlock ensues. (This has been observed to happen when two
@@ -2503,7 +2502,7 @@ again:
 		 * one just happens to be trying to split the page the first one got
 		 * from StrategyGetBuffer.)
 		 */
-		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -4072,8 +4071,8 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	}
 
 	/*
-	 * Pin it, share-lock it, write it.  (FlushBuffer will do nothing if the
-	 * buffer is clean by the time we've locked it.)
+	 * Pin it, share-exclusive-lock it, write it.  (FlushBuffer will do
+	 * nothing if the buffer is clean by the time we've locked it.)
 	 */
 	PinBuffer_Locked(bufHdr);
 
@@ -4403,11 +4402,8 @@ BufferGetTag(Buffer buffer, RelFileLocator *rlocator, ForkNumber *forknum,
  * However, we will need to force the changes to disk via fsync before
  * we can checkpoint WAL.
  *
- * The caller must hold a pin on the buffer and have share-locked the
- * buffer contents.  (Note: a share-lock does not prevent updates of
- * hint bits in the buffer, so the page could change while the write
- * is in progress, but we assume that that will not invalidate the data
- * written.)
+ * The caller must hold a pin on the buffer and have
+ * (share-)exclusively-locked the buffer contents.
  *
  * If the caller has an smgr reference for the buffer's relation, pass it
  * as the second parameter.  If not, pass NULL.
@@ -4423,6 +4419,9 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	char	   *bufToWrite;
 	uint64		buf_state;
 
+	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(buf, BUFFER_LOCK_SHARE_EXCLUSIVE));
+
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
 	 * someone else flushed the buffer before we could, so we need not do
@@ -4555,7 +4554,7 @@ FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(buf);
 
-	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 	BufferLockUnlock(buffer, buf);
 }
@@ -5474,8 +5473,8 @@ FlushDatabaseBuffers(Oid dbid)
 }
 
 /*
- * Flush a previously, shared or exclusively, locked and pinned buffer to the
- * OS.
+ * Flush a previously, share-exclusively or exclusively, locked and pinned
+ * buffer to the OS.
  */
 void
 FlushOneBuffer(Buffer buffer)
@@ -5548,39 +5547,23 @@ IncrBufferRefCount(Buffer buffer)
 }
 
 /*
- * MarkBufferDirtyHint
+ * Shared-buffer only helper for MarkBufferDirtyHint() and
+ * BufferSetHintBits16().
  *
- *	Mark a buffer dirty for non-critical changes.
- *
- * This is essentially the same as MarkBufferDirty, except:
- *
- * 1. The caller does not write WAL; so if checksums are enabled, we may need
- *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
- * 2. The caller might have only share-lock instead of exclusive-lock on the
- *	  buffer's content lock.
- * 3. This function does not guarantee that the buffer is always marked dirty
- *	  (due to a race condition), so it cannot be used for important changes.
+ * This is separated out because it turns out that the repeated checks for
+ * local buffers, repeated GetBufferDescriptor() and repeated reading of the
+ * buffer's state sufficiently hurts the performance of BufferSetHintBits16().
  */
-void
-MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+static inline void
+MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate, bool buffer_std)
 {
-	BufferDesc *bufHdr;
 	Page		page = BufferGetPage(buffer);
 
-	if (!BufferIsValid(buffer))
-		elog(ERROR, "bad buffer ID: %d", buffer);
-
-	if (BufferIsLocal(buffer))
-	{
-		MarkLocalBufferDirty(buffer);
-		return;
-	}
-
-	bufHdr = GetBufferDescriptor(buffer - 1);
-
 	Assert(GetPrivateRefCount(buffer) > 0);
-	/* here, either share or exclusive lock is OK */
-	Assert(BufferIsLockedByMe(buffer));
+
+	/* here, either share-exclusive or exclusive lock is OK */
+	Assert(BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_SHARE_EXCLUSIVE));
 
 	/*
 	 * This routine might get called many times on the same page, if we are
@@ -5593,8 +5576,8 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
-		(BM_DIRTY | BM_JUST_DIRTIED))
+	if (unlikely((lockstate & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+				 (BM_DIRTY | BM_JUST_DIRTIED)))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
@@ -5610,8 +5593,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * We don't check full_page_writes here because that logic is included
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
-		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
+		if (XLogHintBitIsNeeded() && (lockstate & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5663,13 +5645,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 			dirtied = true;		/* Means "will be dirtied by this action" */
 
 			/*
-			 * Set the page LSN if we wrote a backup block. We aren't supposed
-			 * to set this when only holding a share lock but as long as we
-			 * serialise it somehow we're OK. We choose to set LSN while
-			 * holding the buffer header lock, which causes any reader of an
-			 * LSN who holds only a share lock to also obtain a buffer header
-			 * lock before using PageGetLSN(), which is enforced in
-			 * BufferGetLSNAtomic().
+			 * Set the page LSN if we wrote a backup block. To allow backends
+			 * that only hold a share lock on the buffer to read the LSN in a
+			 * tear-free manner, we set the page LSN while holding the buffer
+			 * header lock. This allows any reader of an LSN who holds only a
+			 * share lock to also obtain a buffer header lock before using
+			 * PageGetLSN() to read the LSN in a tear free way. This is done
+			 * in BufferGetLSNAtomic().
 			 *
 			 * If checksums are enabled, you might think we should reset the
 			 * checksum here. That will happen when the page is written
@@ -5695,6 +5677,41 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	}
 }
 
+/*
+ * MarkBufferDirtyHint
+ *
+ *	Mark a buffer dirty for non-critical changes.
+ *
+ * This is essentially the same as MarkBufferDirty, except:
+ *
+ * 1. The caller does not write WAL; so if checksums are enabled, we may need
+ *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
+ * 2. The caller might have only share-exclusive-lock instead of
+ *	  exclusive-lock on the buffer's content lock.
+ * 3. This function does not guarantee that the buffer is always marked dirty
+ *	  (due to a race condition), so it cannot be used for important changes.
+ */
+inline void
+MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+{
+	BufferDesc *bufHdr;
+
+	bufHdr = GetBufferDescriptor(buffer - 1);
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		MarkLocalBufferDirty(buffer);
+		return;
+	}
+
+	MarkSharedBufferDirtyHint(buffer, bufHdr,
+							  pg_atomic_read_u64(&bufHdr->state),
+							  buffer_std);
+}
+
 /*
  * Release buffer content locks for shared buffers.
  *
@@ -6788,6 +6805,188 @@ IsBufferCleanupOK(Buffer buffer)
 	return false;
 }
 
+/*
+ * Helper for BufferBeginSetHintBits() and BufferSetHintBits16().
+ *
+ * This checks if the current lock mode already suffices to allow hint bits
+ * being set and, if not, whether the current lock can be upgraded.
+ */
+static inline bool
+SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
+{
+	uint64		old_state;
+	PrivateRefCountEntry *ref;
+	BufferLockMode mode;
+
+	ref = GetPrivateRefCountEntry(buffer, true);
+
+	if (ref == NULL)
+		elog(ERROR, "lock is not held");
+
+	mode = ref->data.lockmode;
+	if (mode == BUFFER_LOCK_UNLOCK)
+		elog(ERROR, "buffer is not locked");
+
+	/*
+	 * Already am holding a sufficient lock level.
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE || mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+	{
+		*lockstate = pg_atomic_read_u64(&buf_hdr->state);
+		return true;
+	}
+
+	/*
+	 * Only holding a share lock right now, try to upgrade to SHARE_EXCLUSIVE.
+	 */
+	Assert(mode == BUFFER_LOCK_SHARE);
+
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+	while (true)
+	{
+		uint64		desired_state;
+
+		desired_state = old_state;
+
+		/*
+		 * Can't upgrade if somebody else holds the lock in exlusive or
+		 * share-exclusive mode.
+		 */
+		if (unlikely((old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) != 0))
+		{
+			return false;
+		}
+
+		/* currently held lock state */
+		desired_state -= BM_LOCK_VAL_SHARED;
+
+		/* new lock level */
+		desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			ref->data.lockmode = BUFFER_LOCK_SHARE_EXCLUSIVE;
+			*lockstate = desired_state;
+
+			return true;
+		}
+	}
+
+}
+
+/*
+ * Try to acquire the right to set hint bits on the buffer.
+ *
+ * To be allowed to set hint bits, this backend needs to hold either a
+ * share-exclusive or an exclusive lock. In case this backend only holds a
+ * share lock, this function will try to upgrade the lock to
+ * share-exclusive. The caller is only allowed to set hint bits if true is
+ * returned.
+ *
+ * Once BufferBeginSetHintBits() has returned true, hint bits may be set
+ * without further calls to BufferBeginSetHintBits(), until the buffer is
+ * unlocked.
+ *
+ *
+ * Requiring a share-exclusive lock to set hint bits prevents setting hint
+ * bits on buffers that are currently being written out, which could corrupt
+ * the checksum on the page. Flushing buffers also requires a share-exclusive
+ * lock.
+ *
+ * Due to a lock >= share-exclusive being required to set hint bits, only one
+ * backend can set hint bits at a time. To allow multiple backends setting
+ * hint bits would require more complicated locking, for setting hint bits
+ * we'd need to store the count of backends currently setting hint bits and we
+ * would need another lock-level for I/O conflicting with the lock-level to
+ * set hint bits. Given that the share-exclusive lock for setting hint bits is
+ * only held for a short time, that often backends would just set the same
+ * hint bits and that the cost of occasionally not setting hint bits in hotly
+ * accessed pages is fairly low, this seems like an acceptable tradeoff.
+ */
+bool
+BufferBeginSetHintBits(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+
+	if (BufferIsLocal(buffer))
+	{
+		/*
+		 * TODO: will need to check for write IO once that's done
+		 * asynchronously.
+		 */
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	return SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate);
+}
+
+/*
+ * End a phase of setting hint bits on this buffer, started with
+ * BufferBeginSetHintBits().
+ *
+ * This would strictly speaking not be required (i.e. the caller could do
+ * MarkBufferDirtyHint() if so desired), but allows us to perform some sanity
+ * checks.
+ */
+void
+BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std)
+{
+	if (!BufferIsLocal(buffer))
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+
+	if (mark_dirty)
+		MarkBufferDirtyHint(buffer, buffer_std);
+}
+
+/*
+ * Ty to set a single hint bit in a buffer.
+ *
+ * This is a bit faster than BufferBeginSetHintBits() /
+ * BufferFinishSetHintBits() when setting a single hint bit, but slower than
+ * the former when setting several hint bits.
+ */
+bool
+BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+#ifdef USE_ASSERT_CHECKING
+	char	   *page;
+
+	/* verify that the address is on the page */
+	page = BufferGetPage(buffer);
+	Assert((char *) ptr >= page && (char *) ptr < (page + BLCKSZ));
+#endif
+
+	if (BufferIsLocal(buffer))
+	{
+		*ptr = val;
+
+		MarkLocalBufferDirty(buffer);
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	if (SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate))
+	{
+		*ptr = val;
+
+		MarkSharedBufferDirtyHint(buffer, buf_hdr, lockstate, true);
+
+		return true;
+	}
+
+	return false;
+}
+
 
 /*
  *	Functions for buffer I/O handling
diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index ad337c00871..b9a8f368a63 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -904,13 +904,17 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 	max_avail = fsm_get_max_avail(page);
 
 	/*
-	 * Reset the next slot pointer. This encourages the use of low-numbered
-	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation. We don't bother with marking the page dirty if it wasn't
-	 * already, since this is just a hint.
+	 * Try to reset the next slot pointer. This encourages the use of
+	 * low-numbered pages, increasing the chances that a later vacuum can
+	 * truncate the relation. We don't bother with marking the page dirty if
+	 * it wasn't already, since this is just a hint.
 	 */
 	LockBuffer(buf, BUFFER_LOCK_SHARE);
-	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	if (BufferBeginSetHintBits(buf))
+	{
+		((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+		BufferFinishSetHintBits(buf, false, false);
+	}
 	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 	ReleaseBuffer(buf);
diff --git a/src/backend/storage/freespace/fsmpage.c b/src/backend/storage/freespace/fsmpage.c
index 33ee825529c..e46bf2631fc 100644
--- a/src/backend/storage/freespace/fsmpage.c
+++ b/src/backend/storage/freespace/fsmpage.c
@@ -298,9 +298,18 @@ restart:
 	 * lock and get a garbled next pointer every now and then, than take the
 	 * concurrency hit of an exclusive lock.
 	 *
+	 * Without an exclusive lock, we need to use the hint bit infrastructure
+	 * to be allowed to modify the page.
+	 *
 	 * Wrap-around is handled at the beginning of this function.
 	 */
-	fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+	if (exclusive_lock_held || BufferBeginSetHintBits(buf))
+	{
+		fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+
+		if (!exclusive_lock_held)
+			BufferFinishSetHintBits(buf, false, true);
+	}
 
 	return slot;
 }
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 14dec2d49c1..efea48fcef7 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2750,6 +2750,7 @@ SetConstraintStateData
 SetConstraintTriggerData
 SetExprState
 SetFunctionReturnMode
+SetHintBitsState
 SetOp
 SetOpCmd
 SetOpPath
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v9-0009-WIP-Make-UnlockReleaseBuffer-more-efficient.patch (3.5K, ../../4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf/10-v9-0009-WIP-Make-UnlockReleaseBuffer-more-efficient.patch)
  download | inline diff:
From eb3f2965e3909a675c96e73de0dbe6d99482db23 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 19 Nov 2025 15:32:20 -0500
Subject: [PATCH v9 09/10] WIP: Make UnlockReleaseBuffer() more efficient

Now that the buffer content lock is implemented as part of BufferDesc.state,
releasing the lock and unpinning the buffer can be implemented as a single
atomic operation.

Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/nbtree/nbtpage.c | 22 +++++++++++-
 src/backend/storage/buffer/bufmgr.c | 52 ++++++++++++++++++++++++++++-
 2 files changed, 72 insertions(+), 2 deletions(-)

diff --git a/src/backend/access/nbtree/nbtpage.c b/src/backend/access/nbtree/nbtpage.c
index 2ff0085b96f..073f5a05f81 100644
--- a/src/backend/access/nbtree/nbtpage.c
+++ b/src/backend/access/nbtree/nbtpage.c
@@ -1007,11 +1007,18 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 
 	Assert(BlockNumberIsValid(blkno));
 	if (BufferIsValid(obuf))
+	{
+		_bt_relbuf(rel, obuf);
+#if 0
+		Assert(BufferGetBlockNumber(obuf) != blkno);
 		_bt_unlockbuf(rel, obuf);
-	buf = ReleaseAndReadBuffer(obuf, rel, blkno);
+#endif
+	}
+	buf = ReadBuffer(rel, blkno);
 	_bt_lockbuf(rel, buf, access);
 
 	_bt_checkpage(rel, buf);
+
 	return buf;
 }
 
@@ -1023,8 +1030,21 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 void
 _bt_relbuf(Relation rel, Buffer buf)
 {
+#if 0
 	_bt_unlockbuf(rel, buf);
 	ReleaseBuffer(buf);
+#else
+	/*
+	 * Buffer is pinned and locked, which means that it is expected to be
+	 * defined and addressable.  Check that proactively.
+	 */
+	VALGRIND_CHECK_MEM_IS_DEFINED(BufferGetPage(buf), BLCKSZ);
+
+	UnlockReleaseBuffer(buf);
+
+	if (!RelationUsesLocalBuffers(rel))
+		VALGRIND_MAKE_MEM_NOACCESS(BufferGetPage(buf), BLCKSZ);
+#endif
 }
 
 /*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index e944c09fb33..647af7a167c 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5511,13 +5511,63 @@ ReleaseBuffer(Buffer buffer)
 /*
  * UnlockReleaseBuffer -- release the content lock and pin on a buffer
  *
- * This is just a shorthand for a common combination.
+ * This is just a, more efficient, shorthand for a common combination.
  */
 void
 UnlockReleaseBuffer(Buffer buffer)
 {
+#if 1
+	int			mode;
+	BufferDesc *buf;
+	PrivateRefCountEntry *ref;
+	uint64		sub;
+	uint64		lockstate;
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		UnpinLocalBuffer(buffer);
+		return;
+	}
+
+	ResourceOwnerForgetBuffer(CurrentResourceOwner, buffer);
+
+	buf = GetBufferDescriptor(buffer - 1);
+
+	mode = BufferLockDisownInternal(buffer, buf);
+
+	/* compute state modification for lock release */
+	sub = BufferLockReleaseSub(mode);
+
+	/* compute state modification for pin release */
+	ref = GetPrivateRefCountEntry(buffer, false);
+	Assert(ref != NULL);
+	Assert(ref->data.refcount > 0);
+	ref->data.refcount--;
+
+	if (ref->data.refcount == 0)
+	{
+		sub |= BUF_REFCOUNT_ONE;
+		ForgetPrivateRefCountEntry(ref);
+	}
+
+	/* perform the lock and pin release in one atomic op */
+	lockstate = pg_atomic_sub_fetch_u64(&buf->state, sub);
+
+	/* wake up waiters etc */
+	BufferLockProcessRelease(buf, mode, lockstate);
+
+	if (lockstate & BM_PIN_COUNT_WAITER)
+		WakePinCountWaiter(buf);
+
+	RESUME_INTERRUPTS();
+
+#else
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 	ReleaseBuffer(buffer);
+#endif
 }
 
 /*
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v9-0010-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch (11.6K, ../../4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf/11-v9-0010-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch)
  download | inline diff:
From 275d8a35fd0a2553c293ac9a8dff76a535bb081e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 17 Oct 2024 14:14:35 -0400
Subject: [PATCH v9 10/10] WIP: bufmgr: Don't copy pages while writing out

After the series of preceding commits introducing and using
BufferBeginSetHintBits()/BufferSetHintBits16() hint bits are not set
anymore while IO is going on. Therefore we do not need to copy pages while
they are being written out anymore.

TODO: Update comments

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/bufpage.h           |  3 +-
 src/backend/access/hash/hashpage.c      |  2 +-
 src/backend/access/transam/xloginsert.c | 43 ++++++----------------
 src/backend/storage/buffer/bufmgr.c     | 21 +++++------
 src/backend/storage/buffer/localbuf.c   |  2 +-
 src/backend/storage/page/bufpage.c      | 48 ++++---------------------
 src/backend/storage/smgr/bulk_write.c   |  2 +-
 src/test/modules/test_aio/test_aio.c    |  2 +-
 8 files changed, 33 insertions(+), 90 deletions(-)

diff --git a/src/include/storage/bufpage.h b/src/include/storage/bufpage.h
index ae3725b3b81..31ec9a8a047 100644
--- a/src/include/storage/bufpage.h
+++ b/src/include/storage/bufpage.h
@@ -504,7 +504,6 @@ extern void PageIndexMultiDelete(Page page, OffsetNumber *itemnos, int nitems);
 extern void PageIndexTupleDeleteNoCompact(Page page, OffsetNumber offnum);
 extern bool PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 									const void *newtup, Size newsize);
-extern char *PageSetChecksumCopy(Page page, BlockNumber blkno);
-extern void PageSetChecksumInplace(Page page, BlockNumber blkno);
+extern void PageSetChecksum(Page page, BlockNumber blkno);
 
 #endif							/* BUFPAGE_H */
diff --git a/src/backend/access/hash/hashpage.c b/src/backend/access/hash/hashpage.c
index 8e220a3ae16..52c20208c66 100644
--- a/src/backend/access/hash/hashpage.c
+++ b/src/backend/access/hash/hashpage.c
@@ -1029,7 +1029,7 @@ _hash_alloc_buckets(Relation rel, BlockNumber firstblock, uint32 nblocks)
 					zerobuf.data,
 					true);
 
-	PageSetChecksumInplace(page, lastblock);
+	PageSetChecksum(page, lastblock);
 	smgrextend(RelationGetSmgr(rel), MAIN_FORKNUM, lastblock, zerobuf.data,
 			   false);
 
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index 92c48e768c3..4bab484fd5d 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -261,8 +261,11 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	 */
 #ifdef USE_ASSERT_CHECKING
 	if (!(flags & REGBUF_NO_CHANGE))
-		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) &&
-			   BufferIsDirty(buffer));
+	{
+		Assert(BufferIsDirty(buffer));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE));
+	}
 #endif
 
 	if (block_id >= max_registered_block_id)
@@ -1066,7 +1069,7 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * Write a backup block if needed when we are setting a hint. Note that
  * this may be called for a variety of page types, not just heaps.
  *
- * Callable while holding just share lock on the buffer content.
+ * Callable while holding just share-exclusive lock on the buffer content.
  *
  * We can't use the plain backup block mechanism since that relies on the
  * Buffer being exclusively locked. Since some modifications (setting LSN, hint
@@ -1074,6 +1077,8 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * failures. So instead we copy the page and insert the copied data as normal
  * record data.
  *
+ * FIXME: outdated
+ *
  * We only need to do something if page has not yet been full page written in
  * this checkpoint round. The LSN of the inserted wal record is returned if we
  * had to write, InvalidXLogRecPtr otherwise.
@@ -1102,46 +1107,20 @@ XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
 
 	/*
 	 * We assume page LSN is first data on *every* page that can be passed to
-	 * XLogInsert, whether it has the standard page layout or not. Since we're
-	 * only holding a share-lock on the page, we must take the buffer header
-	 * lock when we look at the LSN.
+	 * XLogInsert, whether it has the standard page layout or not.
 	 */
 	lsn = BufferGetLSNAtomic(buffer);
 
 	if (lsn <= RedoRecPtr)
 	{
-		int			flags = 0;
-		PGAlignedBlock copied_buffer;
-		char	   *origdata = (char *) BufferGetBlock(buffer);
-		RelFileLocator rlocator;
-		ForkNumber	forkno;
-		BlockNumber blkno;
-
-		/*
-		 * Copy buffer so we don't have to worry about concurrent hint bit or
-		 * lsn updates. We assume pd_lower/upper cannot be changed without an
-		 * exclusive lock, so the contents bkp are not racy.
-		 */
-		if (buffer_std)
-		{
-			/* Assume we can omit data between pd_lower and pd_upper */
-			Page		page = BufferGetPage(buffer);
-			uint16		lower = ((PageHeader) page)->pd_lower;
-			uint16		upper = ((PageHeader) page)->pd_upper;
-
-			memcpy(copied_buffer.data, origdata, lower);
-			memcpy(copied_buffer.data + upper, origdata + upper, BLCKSZ - upper);
-		}
-		else
-			memcpy(copied_buffer.data, origdata, BLCKSZ);
+		int			flags = REGBUF_NO_CHANGE;
 
 		XLogBeginInsert();
 
 		if (buffer_std)
 			flags |= REGBUF_STANDARD;
 
-		BufferGetTag(buffer, &rlocator, &forkno, &blkno);
-		XLogRegisterBlock(0, &rlocator, forkno, blkno, copied_buffer.data, flags);
+		XLogRegisterBuffer(0, buffer, flags);
 
 		recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI_FOR_HINT);
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 647af7a167c..b24f25ce405 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -4416,7 +4416,6 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	ErrorContextCallback errcallback;
 	instr_time	io_start;
 	Block		bufBlock;
-	char	   *bufToWrite;
 	uint64		buf_state;
 
 	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
@@ -4487,12 +4486,8 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	 */
 	bufBlock = BufHdrGetBlock(buf);
 
-	/*
-	 * Update page checksum if desired.  Since we have only shared lock on the
-	 * buffer, other processes might be updating hint bits in it, so we must
-	 * copy the page to private storage if we do checksumming.
-	 */
-	bufToWrite = PageSetChecksumCopy((Page) bufBlock, buf->tag.blockNum);
+	/* Update page checksum if desired. */
+	PageSetChecksum((Page) bufBlock, buf->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
@@ -4502,7 +4497,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	smgrwrite(reln,
 			  BufTagGetForkNum(&buf->tag),
 			  buf->tag.blockNum,
-			  bufToWrite,
+			  bufBlock,
 			  false);
 
 	/*
@@ -4626,8 +4621,8 @@ BufferIsPermanent(Buffer buffer)
 /*
  * BufferGetLSNAtomic
  *		Retrieves the LSN of the buffer atomically using a buffer header lock.
- *		This is necessary for some callers who may not have an exclusive lock
- *		on the buffer.
+ *		This is necessary for some callers who may not have a (share-)exclusive
+ *		lock on the buffer.
  */
 XLogRecPtr
 BufferGetLSNAtomic(Buffer buffer)
@@ -5679,6 +5674,12 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate, b
 			 * It's possible we may enter here without an xid, so it is
 			 * essential that CreateCheckPoint waits for virtual transactions
 			 * rather than full transactionids.
+			 *
+			 * FIXME: I think we now should simply mark the page dirty before
+			 * WAL logging the hint bit - afaikt it then should work just like
+			 * any other buffer write (due to SyncBuffers()/SyncOneBuffer()
+			 * seeing the dirty bit and trying to lock the page
+			 * share-exclusive, and thus having to wait).
 			 */
 			Assert((MyProc->delayChkptFlags & DELAY_CHKPT_START) == 0);
 			MyProc->delayChkptFlags |= DELAY_CHKPT_START;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 04a540379a2..55e17e03acb 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -199,7 +199,7 @@ FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln)
 		reln = smgropen(BufTagGetRelFileLocator(&bufHdr->tag),
 						MyProcNumber);
 
-	PageSetChecksumInplace(localpage, bufHdr->tag.blockNum);
+	PageSetChecksum(localpage, bufHdr->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index de85911e3ac..2072bb1c72c 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -1494,51 +1494,15 @@ PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 /*
  * Set checksum for a page in shared buffers.
  *
- * If checksums are disabled, or if the page is not initialized, just return
- * the input.  Otherwise, we must make a copy of the page before calculating
- * the checksum, to prevent concurrent modifications (e.g. setting hint bits)
- * from making the final checksum invalid.  It doesn't matter if we include or
- * exclude hints during the copy, as long as we write a valid page and
- * associated checksum.
+ * If checksums are disabled, or if the page is not initialized, just
+ * return. Otherwise compute and set the checksum.
  *
- * Returns a pointer to the block-sized data that needs to be written. Uses
- * statically-allocated memory, so the caller must immediately write the
- * returned page and not refer to it again.
- */
-char *
-PageSetChecksumCopy(Page page, BlockNumber blkno)
-{
-	static char *pageCopy = NULL;
-
-	/* If we don't need a checksum, just return the passed-in data */
-	if (PageIsNew(page) || !DataChecksumsEnabled())
-		return page;
-
-	/*
-	 * We allocate the copy space once and use it over on each subsequent
-	 * call.  The point of palloc'ing here, rather than having a static char
-	 * array, is first to ensure adequate alignment for the checksumming code
-	 * and second to avoid wasting space in processes that never call this.
-	 */
-	if (pageCopy == NULL)
-		pageCopy = MemoryContextAllocAligned(TopMemoryContext,
-											 BLCKSZ,
-											 PG_IO_ALIGN_SIZE,
-											 0);
-
-	memcpy(pageCopy, page, BLCKSZ);
-	((PageHeader) pageCopy)->pd_checksum = pg_checksum_page(pageCopy, blkno);
-	return pageCopy;
-}
-
-/*
- * Set checksum for a page in private memory.
- *
- * This must only be used when we know that no other process can be modifying
- * the page buffer.
+ * In the past this needed to be done on a copy of the page, due to the
+ * possibility of e.g. hint bits being set concurrently. However, this is not
+ * necessary anymore as hint bits won't be set while IO is going on.
  */
 void
-PageSetChecksumInplace(Page page, BlockNumber blkno)
+PageSetChecksum(Page page, BlockNumber blkno)
 {
 	/* If we don't need a checksum, just return */
 	if (PageIsNew(page) || !DataChecksumsEnabled())
diff --git a/src/backend/storage/smgr/bulk_write.c b/src/backend/storage/smgr/bulk_write.c
index 36b28824ec8..f3c24082a69 100644
--- a/src/backend/storage/smgr/bulk_write.c
+++ b/src/backend/storage/smgr/bulk_write.c
@@ -279,7 +279,7 @@ smgr_bulk_flush(BulkWriteState *bulkstate)
 		BlockNumber blkno = pending_writes[i].blkno;
 		Page		page = pending_writes[i].buf->data;
 
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 
 		if (blkno >= bulkstate->relsize)
 		{
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index b1aa8af9ec0..2ae4a559fab 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -288,7 +288,7 @@ modify_rel_block(PG_FUNCTION_ARGS)
 	}
 	else
 	{
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 	}
 
 	smgrwrite(RelationGetSmgr(rel),
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-09 08:08                                     ` Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Kirill Reshke @ 2026-01-09 08:08 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi!

On Fri, 9 Jan 2026 at 05:29, Andres Freund <andres@anarazel.de> wrote:
>
> I think 0001, 0002, 0003 can be committed. 0004, 0005 are new and probably

0001 LGTM.

I also did look at 0002, looks sane.

Other patches are out of my comprehension for now, I did not review them .



-- 
Best regards,
Kirill Reshke





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
@ 2026-01-12 17:45                                       ` Andres Freund <andres@anarazel.de>
  2026-01-12 22:27                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 2 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-12 17:45 UTC (permalink / raw)
  To: Kirill Reshke <reshkekirill@gmail.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-09 13:08:43 +0500, Kirill Reshke wrote:
> On Fri, 9 Jan 2026 at 05:29, Andres Freund <andres@anarazel.de> wrote:
> >
> > I think 0001, 0002, 0003 can be committed. 0004, 0005 are new and probably
> 
> 0001 LGTM.
> 
> I also did look at 0002, looks sane.

Thanks for looking!

I've pushed 0001/0002 now.  I fixed a typo or two since the last published
version.

I'm doing another pass through 0003 and will push that if I don't find
anything significant.

Also working on doing comment polishing of the later patches, found a few
things, but not quite enough to be worth reposting yet.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-12 22:27                                         ` Melanie Plageman <melanieplageman@gmail.com>
  2026-01-12 23:22                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  1 sibling, 1 reply; 120+ messages in thread

From: Melanie Plageman @ 2026-01-12 22:27 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Mon, Jan 12, 2026 at 12:45 PM Andres Freund <andres@anarazel.de> wrote:
>
> Also working on doing comment polishing of the later patches, found a few
> things, but not quite enough to be worth reposting yet.

I looked at 0004 and 0005 and re-looked at 0006 - 0007.

For 0004, I think you should clarify the commit message a bit. I had
trouble understanding when WAKE_IN_PROGRESS is set. So, before
RELEASE_OK was set all the time except when a process woke up and
hadn't run yet. Now, WAKE_IN_PROGRESS is only set when a process is
woken up but hasn't run yet. Personally, I just needed a bit more
specificity (maybe even a bit more formality and grammatically
correctness) from the commit message to get it.

I agree that separating 0005 is helpful.

0007 looks basically fine to me. I'd comb through it with an AI tool
to catch a few nits I saw like an outdated reference to RELEASE_OK and
a missing word in the commit message, etc.

Otherwise, I mostly looked to see if the wakeup semantics seemed right
and if anything jumped out at me while skimming (i.e. I didn't go
through every line with a fine-toothed comb).

The two things I came up with were:

I wondered why this was needed (i.e. why it wasn't needed before)

@@ -6688,7 +7428,25 @@ ResOwnerReleaseBufferPin(Datum res)
     if (BufferIsLocal(buffer))
         UnpinLocalBufferNoOwner(buffer);
     else
+    {
+        PrivateRefCountEntry *ref;
+
+        ref = GetPrivateRefCountEntry(buffer, false);
+
+        /*
+         * If the buffer was locked at the time of the resowner release,
+         * release the lock now. This should only happen after errors.
+         */
+        if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
+        {
+            BufferDesc *buf = GetBufferDescriptor(buffer - 1);
+
+            HOLD_INTERRUPTS();    /* match the upcoming RESUME_INTERRUPTS */
+            BufferLockUnlock(buffer, buf);
+        }
+
         UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
+    }
 }

 is it related to your comment in the commit message

2) Error recovery for content locks is implemented as part of the
already existing private-refcount tracking mechanism in combination
with resowners?

As for your FIXMEs,
+    /*
+     * FIXME: This is reusing the lwlock fields. That's not a correctness
+     * issue, a backend can't wait for both an lwlock and a buffer content
+     * lock at the same time. However, it seems pretty ugly, particularly
+     * given that the field names have an lw* prefix. But duplicating the
+     * fields also seems somewhat superfluous.
+     */

personally I can live with reusing the lwlock fields now that it's
fairly well documented.

+    /* XXX: combine with fetch_and above? */
+    UnlockBufHdr(buf_hdr);

Are you thinking about adding a helper that stops waiting and unlocks?

> Perhaps move the locking code into a buffer_locking.h or such? Needs to be inline functions for efficiency unfortunately.

So you mean put all of the static buffer locking functions you added
to bufmgr.c inline into a header file?

bufmgr.c is super long anyway, so it's not like making it separate
makes the file manageable. On the other hand, it's probably better to
not keep making it worse. For example, I find it really annoying that
the helper function prototypes for res owner and ref count related
functions are grouped before their implementations and then below that
there is another seemingly arbitrary group of prototypes and then
their implementations. Like, what is the logic there?

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-12 22:27                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
@ 2026-01-12 23:22                                           ` Andres Freund <andres@anarazel.de>
  2026-01-13 14:59                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-01-12 23:22 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-12 17:27:30 -0500, Melanie Plageman wrote:
> On Mon, Jan 12, 2026 at 12:45 PM Andres Freund <andres@anarazel.de> wrote:
> >
> > Also working on doing comment polishing of the later patches, found a few
> > things, but not quite enough to be worth reposting yet.
>
> I looked at 0004 and 0005 and re-looked at 0006 - 0007.
>
> For 0004, I think you should clarify the commit message a bit. I had
> trouble understanding when WAKE_IN_PROGRESS is set. So, before
> RELEASE_OK was set all the time except when a process woke up and
> hadn't run yet. Now, WAKE_IN_PROGRESS is only set when a process is
> woken up but hasn't run yet. Personally, I just needed a bit more
> specificity (maybe even a bit more formality and grammatically
> correctness) from the commit message to get it.

Is this better?
    lwlock: Invert meaning of LW_FLAG_RELEASE_OK

    Previously, a flag was set to indicate that a lock release should wake up
    waiters. Since waking waiters is the default behavior in the majority of
    cases, this logic has been inverted. The new LW_FLAG_WAKE_IN_PROGRESS flag is
    now set iff wakeups are explicitly inhibited.

    The motivation for this change is that in an upcoming commit, content locks
    will be implemented independently of lwlocks, with the lock state stored as
    part of BufferDesc.state. As all of a buffer's flags are cleared when the
    buffer is invalidated, without this change we would have to re-add the
    RELEASE_OK flag after clearing the flags; otherwise, the next lock release
    would not wake waiters.

    It seems good to keep the implementation of lwlocks and buffer content locks
    as similar as reasonably possible.

    Discussion: https://postgr.es/m/4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf



> I agree that separating 0005 is helpful.

Kewl.


> 0007 looks basically fine to me. I'd comb through it with an AI tool
> to catch a few nits I saw like an outdated reference to RELEASE_OK and
> a missing word in the commit message, etc.

Found a few that way and with some manual searching.


> Otherwise, I mostly looked to see if the wakeup semantics seemed right
> and if anything jumped out at me while skimming (i.e. I didn't go
> through every line with a fine-toothed comb).
>
> The two things I came up with were:
>
> I wondered why this was needed (i.e. why it wasn't needed before)

> @@ -6688,7 +7428,25 @@ ResOwnerReleaseBufferPin(Datum res)
>      if (BufferIsLocal(buffer))
>          UnpinLocalBufferNoOwner(buffer);
>      else
> +    {
> +        PrivateRefCountEntry *ref;
> +
> +        ref = GetPrivateRefCountEntry(buffer, false);
> +
> +        /*
> +         * If the buffer was locked at the time of the resowner release,
> +         * release the lock now. This should only happen after errors.
> +         */
> +        if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
> +        {
> +            BufferDesc *buf = GetBufferDescriptor(buffer - 1);
> +
> +            HOLD_INTERRUPTS();    /* match the upcoming RESUME_INTERRUPTS */
> +            BufferLockUnlock(buffer, buf);
> +        }
> +
>          UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
> +    }
>  }

It's needed because previously content locks were released as part of the
LWLockReleaseAll() that are sprinkled across various error recovery paths. Now
that content locks aren't implemented via lwlocks anymore, something new is needed.


>  is it related to your comment in the commit message

> 2) Error recovery for content locks is implemented as part of the
> already existing private-refcount tracking mechanism in combination
> with resowners?

Yes.


> As for your FIXMEs,
> +    /*
> +     * FIXME: This is reusing the lwlock fields. That's not a correctness
> +     * issue, a backend can't wait for both an lwlock and a buffer content
> +     * lock at the same time. However, it seems pretty ugly, particularly
> +     * given that the field names have an lw* prefix. But duplicating the
> +     * fields also seems somewhat superfluous.
> +     */
>
> personally I can live with reusing the lwlock fields now that it's
> fairly well documented.

Cool. That's the conclusion I also came to. So unless somebody pipes up soon,
I'll remove the FIXME from the commit message and code.


> +    /* XXX: combine with fetch_and above? */
> +    UnlockBufHdr(buf_hdr);
>
> Are you thinking about adding a helper that stops waiting and unlocks?

I'm not sure what you mean by that? Just whether I plan to implement the
FIXME?



> > Perhaps move the locking code into a buffer_locking.h or such? Needs to be inline functions for efficiency unfortunately.
>
> So you mean put all of the static buffer locking functions you added
> to bufmgr.c inline into a header file?

Yes, that's what I was wondering about.


> bufmgr.c is super long anyway, so it's not like making it separate
> makes the file manageable. On the other hand, it's probably better to
> not keep making it worse.

Yea. OTOH I don't know if a header that's just included by one file is really
an improvement :/


> For example, I find it really annoying that the helper function prototypes
> for res owner and ref count related functions are grouped before their
> implementations and then below that there is another seemingly arbitrary
> group of prototypes and then their implementations. Like, what is the logic
> there?

I agree it's pretty awful this way. I don't know how the hell that happened,
despite probably being the party to blame (4b4b680c3d6d). Nobody in that
thread commented upon it, it was that way starting in the first version. Odd.
I guess I should propose fixing that :/

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-12 22:27                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2026-01-12 23:22                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-13 14:59                                             ` Melanie Plageman <melanieplageman@gmail.com>
  0 siblings, 0 replies; 120+ messages in thread

From: Melanie Plageman @ 2026-01-13 14:59 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Mon, Jan 12, 2026 at 6:22 PM Andres Freund <andres@anarazel.de> wrote:
>
> Is this better?
>     lwlock: Invert meaning of LW_FLAG_RELEASE_OK
>
>     Previously, a flag was set to indicate that a lock release should wake up
>     waiters. Since waking waiters is the default behavior in the majority of
>     cases, this logic has been inverted. The new LW_FLAG_WAKE_IN_PROGRESS flag is
>     now set iff wakeups are explicitly inhibited.

I think what you have would work for most people. The key thing for me
is that the wakeups are inhibited _because_ someone else is already
awake. So, you don't have to wake anyone up when you release the lock
because there is already someone awake. Having you explain that
off-list was necessary for me to bridge the gap between RELEASE_NOT_OK
and WAKE_IN_PROGRESS. And I do agree that WAKE_IN_PROGRESS is more
descriptive of when the flag is actually set. RELEASE_NOT_OK doesn't
explain the state or who/when it should be set.

> > I wondered why this was needed (i.e. why it wasn't needed before)
>
> > @@ -6688,7 +7428,25 @@ ResOwnerReleaseBufferPin(Datum res)
> >      if (BufferIsLocal(buffer))
> >          UnpinLocalBufferNoOwner(buffer);
> >      else
> > +    {
> > +        PrivateRefCountEntry *ref;
> > +
> > +        ref = GetPrivateRefCountEntry(buffer, false);
> > +
> > +        /*
> > +         * If the buffer was locked at the time of the resowner release,
> > +         * release the lock now. This should only happen after errors.
> > +         */
> > +        if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
> > +        {
> > +            BufferDesc *buf = GetBufferDescriptor(buffer - 1);
> > +
> > +            HOLD_INTERRUPTS();    /* match the upcoming RESUME_INTERRUPTS */
> > +            BufferLockUnlock(buffer, buf);
> > +        }
> > +
> >          UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
> > +    }
> >  }
>
> It's needed because previously content locks were released as part of the
> LWLockReleaseAll() that are sprinkled across various error recovery paths. Now
> that content locks aren't implemented via lwlocks anymore, something new is needed.

And all those LWLockReleaseAll()s are still needed because we might
hold other LWLocks even though we won't hold them for buffer content
access?

> > +    /* XXX: combine with fetch_and above? */
> > +    UnlockBufHdr(buf_hdr);
> >
> > Are you thinking about adding a helper that stops waiting and unlocks?
>
> I'm not sure what you mean by that? Just whether I plan to implement the
> FIXME?

I was trying to figure out why you left it as a FIXME and didn't just
do it or not do it. I thought maybe it was because you weren't sure if
you wanted to add another helper in addition to UnlockBufHdr().

> > bufmgr.c is super long anyway, so it's not like making it separate
> > makes the file manageable. On the other hand, it's probably better to
> > not keep making it worse.
>
> Yea. OTOH I don't know if a header that's just included by one file is really
> an improvement :/

Yea, I suppose that is a bit odd. Though it could be a pattern you
start for organizing gigantic files. I'm overall a +0.7 unless you
explain some other downsides than oddity.

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-13 00:33                                         ` Andres Freund <andres@anarazel.de>
  2026-01-13 15:05                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2026-01-14 02:26                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-09 11:42                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  1 sibling, 5 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-13 00:33 UTC (permalink / raw)
  To: Kirill Reshke <reshkekirill@gmail.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-12 12:45:03 -0500, Andres Freund wrote:
> I'm doing another pass through 0003 and will push that if I don't find
> anything significant.

Done, after adjust two comments in minor ways.


> Also working on doing comment polishing of the later patches, found a few
> things, but not quite enough to be worth reposting yet.

Here are the remaining commits, with a bit of polish:

- fixed references to old names in some places (lwlocks, release_ok)

- Aded an assert that we don't already hold a lock in BufferLockConditional()

- typo and grammar fixes

- updated the commit message of the LW_FLAG_RELEASE_OK, as "requested" by
  Melanie. I hope this explains the situation better.

- added a commit that renames ResOwnerReleaseBufferPin to
  ResOwnerReleaseBuffer (et al), as it now also releases content locks if held

  I kept this separate as I'm not yet sure about the new name, partially due
  to there also being a "buffer io" resowner.  I tried "buffer ownership" for
  the resowner that tracks pins and locks, but that was long and not clearly
  better.

Greetings,

Andres Freund

Attachments:

  [text/x-diff] v10-0001-lwlock-Invert-meaning-of-LW_FLAG_RELEASE_OK.patch (5.7K, ../../jtg5cu4n6h5lib3kzx66ju4yhh6kmviaud7oq6dtut6c4q4rdi@xwsfoagt3c2b/2-v10-0001-lwlock-Invert-meaning-of-LW_FLAG_RELEASE_OK.patch)
  download | inline diff:
From ea4bfffc90bf14c0a0f7cd1e1fe29ebca1430414 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 5 Jan 2026 20:40:38 -0500
Subject: [PATCH v10 1/8] lwlock: Invert meaning of LW_FLAG_RELEASE_OK

Previously, a flag was set to indicate that a lock release should wake up
waiters. Since waking waiters is the default behavior in the majority of
cases, this logic has been inverted. The new LW_FLAG_WAKE_IN_PROGRESS flag is
now set iff wakeups are explicitly inhibited.

The motivation for this change is that in an upcoming commit, content locks
will be implemented independently of lwlocks, with the lock state stored as
part of BufferDesc.state. As all of a buffer's flags are cleared when the
buffer is invalidated, without this change we would have to re-add the
RELEASE_OK flag after clearing the flags; otherwise, the next lock release
would not wake waiters.

It seems good to keep the implementation of lwlocks and buffer content locks
as similar as reasonably possible.

Discussion: https://postgr.es/m/4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf
---
 src/backend/storage/lmgr/lwlock.c | 42 +++++++++++++++----------------
 1 file changed, 20 insertions(+), 22 deletions(-)

diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 6a9f86d5025..148309cc186 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -92,7 +92,7 @@
 
 
 #define LW_FLAG_HAS_WAITERS			((uint32) 1 << 31)
-#define LW_FLAG_RELEASE_OK			((uint32) 1 << 30)
+#define LW_FLAG_WAKE_IN_PROGRESS	((uint32) 1 << 30)
 #define LW_FLAG_LOCKED				((uint32) 1 << 29)
 #define LW_FLAG_BITS				3
 #define LW_FLAG_MASK				(((1<<LW_FLAG_BITS)-1)<<(32-LW_FLAG_BITS))
@@ -246,14 +246,14 @@ PRINT_LWDEBUG(const char *where, LWLock *lock, LWLockMode mode)
 		ereport(LOG,
 				(errhidestmt(true),
 				 errhidecontext(true),
-				 errmsg_internal("%d: %s(%s %p): excl %u shared %u haswaiters %u waiters %u rOK %d",
+				 errmsg_internal("%d: %s(%s %p): excl %u shared %u haswaiters %u waiters %u waking %d",
 								 MyProcPid,
 								 where, T_NAME(lock), lock,
 								 (state & LW_VAL_EXCLUSIVE) != 0,
 								 state & LW_SHARED_MASK,
 								 (state & LW_FLAG_HAS_WAITERS) != 0,
 								 pg_atomic_read_u32(&lock->nwaiters),
-								 (state & LW_FLAG_RELEASE_OK) != 0)));
+								 (state & LW_FLAG_WAKE_IN_PROGRESS) != 0)));
 	}
 }
 
@@ -700,7 +700,7 @@ LWLockInitialize(LWLock *lock, int tranche_id)
 	/* verify the tranche_id is valid */
 	(void) GetLWTrancheName(tranche_id);
 
-	pg_atomic_init_u32(&lock->state, LW_FLAG_RELEASE_OK);
+	pg_atomic_init_u32(&lock->state, 0);
 #ifdef LOCK_DEBUG
 	pg_atomic_init_u32(&lock->nwaiters, 0);
 #endif
@@ -929,15 +929,13 @@ LWLockWaitListUnlock(LWLock *lock)
 static void
 LWLockWakeup(LWLock *lock)
 {
-	bool		new_release_ok;
+	bool		new_release_in_progress = false;
 	bool		wokeup_somebody = false;
 	proclist_head wakeup;
 	proclist_mutable_iter iter;
 
 	proclist_init(&wakeup);
 
-	new_release_ok = true;
-
 	/* lock wait list while collecting backends to wake up */
 	LWLockWaitListLock(lock);
 
@@ -958,7 +956,7 @@ LWLockWakeup(LWLock *lock)
 			 * that are just waiting for the lock to become free don't retry
 			 * automatically.
 			 */
-			new_release_ok = false;
+			new_release_in_progress = true;
 
 			/*
 			 * Don't wakeup (further) exclusive locks.
@@ -997,10 +995,10 @@ LWLockWakeup(LWLock *lock)
 
 			/* compute desired flags */
 
-			if (new_release_ok)
-				desired_state |= LW_FLAG_RELEASE_OK;
+			if (new_release_in_progress)
+				desired_state |= LW_FLAG_WAKE_IN_PROGRESS;
 			else
-				desired_state &= ~LW_FLAG_RELEASE_OK;
+				desired_state &= ~LW_FLAG_WAKE_IN_PROGRESS;
 
 			if (proclist_is_empty(&lock->waiters))
 				desired_state &= ~LW_FLAG_HAS_WAITERS;
@@ -1131,10 +1129,10 @@ LWLockDequeueSelf(LWLock *lock)
 		 */
 
 		/*
-		 * Reset RELEASE_OK flag if somebody woke us before we removed
-		 * ourselves - they'll have set it to false.
+		 * Clear LW_FLAG_WAKE_IN_PROGRESS if somebody woke us before we
+		 * removed ourselves - they'll have set it.
 		 */
-		pg_atomic_fetch_or_u32(&lock->state, LW_FLAG_RELEASE_OK);
+		pg_atomic_fetch_and_u32(&lock->state, ~LW_FLAG_WAKE_IN_PROGRESS);
 
 		/*
 		 * Now wait for the scheduled wakeup, otherwise our ->lwWaiting would
@@ -1301,7 +1299,7 @@ LWLockAcquire(LWLock *lock, LWLockMode mode)
 		}
 
 		/* Retrying, allow LWLockRelease to release waiters again. */
-		pg_atomic_fetch_or_u32(&lock->state, LW_FLAG_RELEASE_OK);
+		pg_atomic_fetch_and_u32(&lock->state, ~LW_FLAG_WAKE_IN_PROGRESS);
 
 #ifdef LOCK_DEBUG
 		{
@@ -1636,10 +1634,10 @@ LWLockWaitForVar(LWLock *lock, pg_atomic_uint64 *valptr, uint64 oldval,
 		LWLockQueueSelf(lock, LW_WAIT_UNTIL_FREE);
 
 		/*
-		 * Set RELEASE_OK flag, to make sure we get woken up as soon as the
-		 * lock is released.
+		 * Clear LW_FLAG_WAKE_IN_PROGRESS flag, to make sure we get woken up
+		 * as soon as the lock is released.
 		 */
-		pg_atomic_fetch_or_u32(&lock->state, LW_FLAG_RELEASE_OK);
+		pg_atomic_fetch_and_u32(&lock->state, ~LW_FLAG_WAKE_IN_PROGRESS);
 
 		/*
 		 * We're now guaranteed to be woken up if necessary. Recheck the lock
@@ -1852,11 +1850,11 @@ LWLockReleaseInternal(LWLock *lock, LWLockMode mode)
 		TRACE_POSTGRESQL_LWLOCK_RELEASE(T_NAME(lock));
 
 	/*
-	 * We're still waiting for backends to get scheduled, don't wake them up
-	 * again.
+	 * Check if we're still waiting for backends to get scheduled, if so,
+	 * don't wake them up again.
 	 */
-	if ((oldstate & (LW_FLAG_HAS_WAITERS | LW_FLAG_RELEASE_OK)) ==
-		(LW_FLAG_HAS_WAITERS | LW_FLAG_RELEASE_OK) &&
+	if ((oldstate & LW_FLAG_HAS_WAITERS) &&
+		!(oldstate & LW_FLAG_WAKE_IN_PROGRESS) &&
 		(oldstate & LW_LOCK_MASK) == 0)
 		check_waiters = true;
 	else
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v10-0002-bufmgr-Make-definitions-related-to-buffer-descri.patch (4.5K, ../../jtg5cu4n6h5lib3kzx66ju4yhh6kmviaud7oq6dtut6c4q4rdi@xwsfoagt3c2b/3-v10-0002-bufmgr-Make-definitions-related-to-buffer-descri.patch)
  download | inline diff:
From 2829cdad54bf2878c0cdc2d9e90596edcfb3ad09 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 7 Jan 2026 17:21:48 -0500
Subject: [PATCH v10 2/8] bufmgr: Make definitions related to buffer descriptor
 easier to modify

This is in preparation to widening the buffer state to 64 bits, which in turn
is preparation for implementing content locks in bufmgr. This commit aims to
make the subsequent commits a bit easier to review, by separating out
reformatting etc from the actual changes.

Discussion: https://postgr.es/m/4csodkvvfbfloxxjlkgsnl2lgfv2mtzdl7phqzd4jxjadxm4o5@usw7feyb5bzf
---
 src/include/storage/buf_internals.h | 65 +++++++++++++++++++++--------
 1 file changed, 47 insertions(+), 18 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index fa43cf4458d..2f607ea2ac5 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -32,6 +32,7 @@
 /*
  * Buffer state is a single 32-bit variable where following data is combined.
  *
+ * State of the buffer itself (in order):
  * - 18 bits refcount
  * - 4 bits usage count
  * - 10 bits of flags
@@ -48,16 +49,30 @@
 StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 				 "parts of buffer state space need to equal 32");
 
+/* refcount related definitions */
 #define BUF_REFCOUNT_ONE 1
-#define BUF_REFCOUNT_MASK ((1U << BUF_REFCOUNT_BITS) - 1)
-#define BUF_USAGECOUNT_MASK (((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_REFCOUNT_BITS))
-#define BUF_USAGECOUNT_ONE (1U << BUF_REFCOUNT_BITS)
-#define BUF_USAGECOUNT_SHIFT BUF_REFCOUNT_BITS
-#define BUF_FLAG_MASK (((1U << BUF_FLAG_BITS) - 1) << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS))
+#define BUF_REFCOUNT_MASK \
+	((1U << BUF_REFCOUNT_BITS) - 1)
+
+/* usage count related definitions */
+#define BUF_USAGECOUNT_SHIFT \
+	BUF_REFCOUNT_BITS
+#define BUF_USAGECOUNT_MASK \
+	(((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_USAGECOUNT_SHIFT))
+#define BUF_USAGECOUNT_ONE \
+	(1U << BUF_REFCOUNT_BITS)
+
+/* flags related definitions */
+#define BUF_FLAG_SHIFT \
+	(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS)
+#define BUF_FLAG_MASK \
+	(((1U << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
 
 /* Get refcount and usagecount from buffer state */
-#define BUF_STATE_GET_REFCOUNT(state) ((state) & BUF_REFCOUNT_MASK)
-#define BUF_STATE_GET_USAGECOUNT(state) (((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
+#define BUF_STATE_GET_REFCOUNT(state) \
+	((state) & BUF_REFCOUNT_MASK)
+#define BUF_STATE_GET_USAGECOUNT(state) \
+	(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
 
 /*
  * Flags for buffer descriptors
@@ -65,17 +80,31 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  * Note: BM_TAG_VALID essentially means that there is a buffer hashtable
  * entry associated with the buffer's tag.
  */
-#define BM_LOCKED				(1U << 22)	/* buffer header is locked */
-#define BM_DIRTY				(1U << 23)	/* data needs writing */
-#define BM_VALID				(1U << 24)	/* data is valid */
-#define BM_TAG_VALID			(1U << 25)	/* tag is assigned */
-#define BM_IO_IN_PROGRESS		(1U << 26)	/* read or write in progress */
-#define BM_IO_ERROR				(1U << 27)	/* previous I/O failed */
-#define BM_JUST_DIRTIED			(1U << 28)	/* dirtied since write started */
-#define BM_PIN_COUNT_WAITER		(1U << 29)	/* have waiter for sole pin */
-#define BM_CHECKPOINT_NEEDED	(1U << 30)	/* must write for checkpoint */
-#define BM_PERMANENT			(1U << 31)	/* permanent buffer (not unlogged,
-											 * or init fork) */
+
+#define BUF_DEFINE_FLAG(flagno)	\
+	(1U << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + (flagno)))
+
+/* buffer header is locked */
+#define BM_LOCKED					BUF_DEFINE_FLAG( 0)
+/* data needs writing */
+#define BM_DIRTY					BUF_DEFINE_FLAG( 1)
+/* data is valid */
+#define BM_VALID					BUF_DEFINE_FLAG( 2)
+/* tag is assigned */
+#define BM_TAG_VALID				BUF_DEFINE_FLAG( 3)
+/* read or write in progress */
+#define BM_IO_IN_PROGRESS			BUF_DEFINE_FLAG( 4)
+/* previous I/O failed */
+#define BM_IO_ERROR					BUF_DEFINE_FLAG( 5)
+/* dirtied since write started */
+#define BM_JUST_DIRTIED				BUF_DEFINE_FLAG( 6)
+/* have waiter for sole pin */
+#define BM_PIN_COUNT_WAITER			BUF_DEFINE_FLAG( 7)
+/* must write for checkpoint */
+#define BM_CHECKPOINT_NEEDED		BUF_DEFINE_FLAG( 8)
+/* permanent buffer (not unlogged, or init fork) */
+#define BM_PERMANENT				BUF_DEFINE_FLAG( 9)
+
 /*
  * The maximum allowed value of usage_count represents a tradeoff between
  * accuracy and speed of the clock-sweep buffer management algorithm.  A
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v10-0003-bufmgr-Change-BufferDesc.state-to-be-a-64-bit-at.patch (45.1K, ../../jtg5cu4n6h5lib3kzx66ju4yhh6kmviaud7oq6dtut6c4q4rdi@xwsfoagt3c2b/4-v10-0003-bufmgr-Change-BufferDesc.state-to-be-a-64-bit-at.patch)
  download | inline diff:
From 2ed035719c043173f5fbfc6961758de26a19bd90 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 7 Jan 2026 17:26:25 -0500
Subject: [PATCH v10 3/8] bufmgr: Change BufferDesc.state to be a 64-bit atomic

This is motivated by wanting to merge buffer content locks into
BufferDesc.state in a future commit, rather than having a separate lwlock (see
commit c75ebc657ff for more details). As this change is rather mechanical, it
seems to make sense to split it out into a separate commit, for easier review.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  51 +++---
 src/include/storage/procnumber.h              |  14 +-
 src/backend/storage/buffer/buf_init.c         |   2 +-
 src/backend/storage/buffer/bufmgr.c           | 170 +++++++++---------
 src/backend/storage/buffer/freelist.c         |  24 +--
 src/backend/storage/buffer/localbuf.c         |  72 ++++----
 contrib/pg_buffercache/pg_buffercache_pages.c |   8 +-
 src/test/modules/test_aio/test_aio.c          |  12 +-
 8 files changed, 178 insertions(+), 175 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 2f607ea2ac5..a4d36e9ca01 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -30,7 +30,7 @@
 #include "utils/resowner.h"
 
 /*
- * Buffer state is a single 32-bit variable where following data is combined.
+ * Buffer state is a single 64-bit variable where following data is combined.
  *
  * State of the buffer itself (in order):
  * - 18 bits refcount
@@ -40,6 +40,9 @@
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
+ * NB: A future commit will use a significant portion of the remaining bits to
+ * implement buffer locking as part of the state variable.
+ *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
@@ -52,27 +55,27 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 /* refcount related definitions */
 #define BUF_REFCOUNT_ONE 1
 #define BUF_REFCOUNT_MASK \
-	((1U << BUF_REFCOUNT_BITS) - 1)
+	((UINT64CONST(1) << BUF_REFCOUNT_BITS) - 1)
 
 /* usage count related definitions */
 #define BUF_USAGECOUNT_SHIFT \
 	BUF_REFCOUNT_BITS
 #define BUF_USAGECOUNT_MASK \
-	(((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_USAGECOUNT_SHIFT))
+	(((UINT64CONST(1) << BUF_USAGECOUNT_BITS) - 1) << (BUF_USAGECOUNT_SHIFT))
 #define BUF_USAGECOUNT_ONE \
-	(1U << BUF_REFCOUNT_BITS)
+	(UINT64CONST(1) << BUF_REFCOUNT_BITS)
 
 /* flags related definitions */
 #define BUF_FLAG_SHIFT \
 	(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS)
 #define BUF_FLAG_MASK \
-	(((1U << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
+	(((UINT64CONST(1) << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
 
 /* Get refcount and usagecount from buffer state */
 #define BUF_STATE_GET_REFCOUNT(state) \
-	((state) & BUF_REFCOUNT_MASK)
+	((uint32)((state) & BUF_REFCOUNT_MASK))
 #define BUF_STATE_GET_USAGECOUNT(state) \
-	(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
+	((uint32)(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT))
 
 /*
  * Flags for buffer descriptors
@@ -82,7 +85,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  */
 
 #define BUF_DEFINE_FLAG(flagno)	\
-	(1U << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + (flagno)))
+	(UINT64CONST(1) << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + (flagno)))
 
 /* buffer header is locked */
 #define BM_LOCKED					BUF_DEFINE_FLAG( 0)
@@ -115,7 +118,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  */
 #define BM_MAX_USAGE_COUNT	5
 
-StaticAssertDecl(BM_MAX_USAGE_COUNT < (1 << BUF_USAGECOUNT_BITS),
+StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 				 "BM_MAX_USAGE_COUNT doesn't fit in BUF_USAGECOUNT_BITS bits");
 StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
 				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
@@ -280,8 +283,8 @@ BufMappingPartitionLockByIndex(uint32 index)
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
- * atomic operations (i.e. only pg_atomic_read_u32() and
- * pg_atomic_unlocked_write_u32()).
+ * atomic operations (i.e. only pg_atomic_read_u64() and
+ * pg_atomic_unlocked_write_u64()).
  *
  * Be careful to avoid increasing the size of the struct when adding or
  * reordering members.  Keeping it below 64 bytes (the most common CPU
@@ -309,7 +312,7 @@ typedef struct BufferDesc
 	 * State of the buffer, containing flags, refcount and usagecount. See
 	 * BUF_* and BM_* defines at the top of this file.
 	 */
-	pg_atomic_uint32 state;
+	pg_atomic_uint64 state;
 
 	/*
 	 * Backend of pin-count waiter. The buffer header spinlock needs to be
@@ -415,7 +418,7 @@ BufferDescriptorGetContentLock(const BufferDesc *bdesc)
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
  */
-extern uint32 LockBufHdr(BufferDesc *desc);
+extern uint64 LockBufHdr(BufferDesc *desc);
 
 /*
  * Unlock the buffer header.
@@ -426,9 +429,9 @@ extern uint32 LockBufHdr(BufferDesc *desc);
 static inline void
 UnlockBufHdr(BufferDesc *desc)
 {
-	Assert(pg_atomic_read_u32(&desc->state) & BM_LOCKED);
+	Assert(pg_atomic_read_u64(&desc->state) & BM_LOCKED);
 
-	pg_atomic_fetch_sub_u32(&desc->state, BM_LOCKED);
+	pg_atomic_fetch_sub_u64(&desc->state, BM_LOCKED);
 }
 
 /*
@@ -439,14 +442,14 @@ UnlockBufHdr(BufferDesc *desc)
  * Note that this approach would not work for usagecount, since we need to cap
  * the usagecount at BM_MAX_USAGE_COUNT.
  */
-static inline uint32
-UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
-				uint32 set_bits, uint32 unset_bits,
+static inline uint64
+UnlockBufHdrExt(BufferDesc *desc, uint64 old_buf_state,
+				uint64 set_bits, uint64 unset_bits,
 				int refcount_change)
 {
 	for (;;)
 	{
-		uint32		buf_state = old_buf_state;
+		uint64		buf_state = old_buf_state;
 
 		Assert(buf_state & BM_LOCKED);
 
@@ -457,7 +460,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 		if (refcount_change != 0)
 			buf_state += BUF_REFCOUNT_ONE * refcount_change;
 
-		if (pg_atomic_compare_exchange_u32(&desc->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&desc->state, &old_buf_state,
 										   buf_state))
 		{
 			return old_buf_state;
@@ -465,7 +468,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 	}
 }
 
-extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
+extern uint64 WaitBufHdrUnlocked(BufferDesc *buf);
 
 /* in bufmgr.c */
 
@@ -525,14 +528,14 @@ extern void TrackNewBufferPin(Buffer buf);
 
 /* solely to make it easier to write tests */
 extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
-extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 							  bool forget_owner, bool release_aio);
 
 
 /* freelist.c */
 extern IOContext IOContextForStrategy(BufferAccessStrategy strategy);
 extern BufferDesc *StrategyGetBuffer(BufferAccessStrategy strategy,
-									 uint32 *buf_state, bool *from_ring);
+									 uint64 *buf_state, bool *from_ring);
 extern bool StrategyRejectBuffer(BufferAccessStrategy strategy,
 								 BufferDesc *buf, bool from_ring);
 
@@ -568,7 +571,7 @@ extern BlockNumber ExtendBufferedRelLocal(BufferManagerRelation bmr,
 										  uint32 *extended_by);
 extern void MarkLocalBufferDirty(Buffer buffer);
 extern void TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty,
-								   uint32 set_flag_bits, bool release_aio);
+								   uint64 set_flag_bits, bool release_aio);
 extern bool StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait);
 extern void FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln);
 extern void InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced);
diff --git a/src/include/storage/procnumber.h b/src/include/storage/procnumber.h
index 30c360ad350..bd9cb3891cc 100644
--- a/src/include/storage/procnumber.h
+++ b/src/include/storage/procnumber.h
@@ -27,13 +27,13 @@ typedef int ProcNumber;
 
 /*
  * Note: MAX_BACKENDS_BITS is 18 as that is the space available for buffer
- * refcounts in buf_internals.h.  This limitation could be lifted by using a
- * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
- * currently realistic configurations. Even if that limitation were removed,
- * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
- * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
- * 4*MaxBackends without any overflow check.  We check that the configured
- * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
+ * refcounts in buf_internals.h.  This limitation could be lifted, but it's
+ * unlikely to be worthwhile as 2^18-1 backends exceed currently realistic
+ * configurations. Even if that limitation were removed, we still could not a)
+ * exceed 2^23-1 because inval.c stores the ProcNumber as a 3-byte signed
+ * integer, b) INT_MAX/4 because some places compute 4*MaxBackends without any
+ * overflow check.  We check that the configured number of backends does not
+ * exceed MAX_BACKENDS in InitializeMaxBackends().
  */
 #define MAX_BACKENDS_BITS		18
 #define MAX_BACKENDS			((1U << MAX_BACKENDS_BITS)-1)
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 9a312bcc7b3..7d894522526 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -121,7 +121,7 @@ BufferManagerShmemInit(void)
 
 			ClearBufferTag(&buf->tag);
 
-			pg_atomic_init_u32(&buf->state, 0);
+			pg_atomic_init_u64(&buf->state, 0);
 			buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
 
 			buf->buf_id = i;
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index a036c2aa275..b0de8e45d4d 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -780,7 +780,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 {
 	BufferDesc *bufHdr;
 	BufferTag	tag;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(recent_buffer));
 
@@ -793,7 +793,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 		int			b = -recent_buffer - 1;
 
 		bufHdr = GetLocalBufferDescriptor(b);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		/* Is it still valid and holding the right tag? */
 		if ((buf_state & BM_VALID) && BufferTagsEqual(&tag, &bufHdr->tag))
@@ -1386,8 +1386,8 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
 				bufHdr = GetLocalBufferDescriptor(-buffers[i] - 1);
 			else
 				bufHdr = GetBufferDescriptor(buffers[i] - 1);
-			Assert(pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID);
-			found = pg_atomic_read_u32(&bufHdr->state) & BM_VALID;
+			Assert(pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID);
+			found = pg_atomic_read_u64(&bufHdr->state) & BM_VALID;
 		}
 		else
 		{
@@ -1613,10 +1613,10 @@ CheckReadBuffersOperation(ReadBuffersOperation *operation, bool is_complete)
 			GetBufferDescriptor(buffer - 1);
 
 		Assert(BufferGetBlockNumber(buffer) == operation->blocknum + i);
-		Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_TAG_VALID);
+		Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_TAG_VALID);
 
 		if (i < operation->nblocks_done)
-			Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_VALID);
+			Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_VALID);
 	}
 #endif
 }
@@ -2083,8 +2083,8 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	int			existing_buf_id;
 	Buffer		victim_buffer;
 	BufferDesc *victim_buf_hdr;
-	uint32		victim_buf_state;
-	uint32		set_bits = 0;
+	uint64		victim_buf_state;
+	uint64		set_bits = 0;
 
 	/* Make sure we will have room to remember the buffer pin */
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2251,7 +2251,7 @@ InvalidateBuffer(BufferDesc *buf)
 	uint32		oldHash;		/* hash value for oldTag */
 	LWLock	   *oldPartitionLock;	/* buffer partition lock for it */
 	uint32		oldFlags;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/* Save the original buffer tag before dropping the spinlock */
 	oldTag = buf->tag;
@@ -2342,7 +2342,7 @@ retry:
 static bool
 InvalidateVictimBuffer(BufferDesc *buf_hdr)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	uint32		hash;
 	LWLock	   *partition_lock;
 	BufferTag	tag;
@@ -2402,10 +2402,10 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 
 	LWLockRelease(partition_lock);
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	Assert(!(buf_state & (BM_DIRTY | BM_VALID | BM_TAG_VALID)));
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u32(&buf_hdr->state)) > 0);
+	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u64(&buf_hdr->state)) > 0);
 
 	return true;
 }
@@ -2415,7 +2415,7 @@ GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context)
 {
 	BufferDesc *buf_hdr;
 	Buffer		buf;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		from_ring;
 
 	/*
@@ -2548,7 +2548,7 @@ again:
 
 	/* a final set of sanity checks */
 #ifdef USE_ASSERT_CHECKING
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) == 1);
 	Assert(!(buf_state & (BM_TAG_VALID | BM_VALID | BM_DIRTY)));
@@ -2839,13 +2839,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 */
 			do
 			{
-				pg_atomic_fetch_and_u32(&existing_hdr->state, ~BM_VALID);
+				pg_atomic_fetch_and_u64(&existing_hdr->state, ~BM_VALID);
 			} while (!StartBufferIO(existing_hdr, true, false));
 		}
 		else
 		{
-			uint32		buf_state;
-			uint32		set_bits = 0;
+			uint64		buf_state;
+			uint64		set_bits = 0;
 
 			buf_state = LockBufHdr(victim_buf_hdr);
 
@@ -3021,7 +3021,7 @@ BufferIsDirty(Buffer buffer)
 		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 	}
 
-	return pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY;
+	return pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY;
 }
 
 /*
@@ -3037,8 +3037,8 @@ void
 MarkBufferDirty(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
-	uint32		old_buf_state;
+	uint64		buf_state;
+	uint64		old_buf_state;
 
 	if (!BufferIsValid(buffer))
 		elog(ERROR, "bad buffer ID: %d", buffer);
@@ -3058,7 +3058,7 @@ MarkBufferDirty(Buffer buffer)
 	 * NB: We have to wait for the buffer header spinlock to be not held, as
 	 * TerminateBufferIO() relies on the spinlock.
 	 */
-	old_buf_state = pg_atomic_read_u32(&bufHdr->state);
+	old_buf_state = pg_atomic_read_u64(&bufHdr->state);
 	for (;;)
 	{
 		if (old_buf_state & BM_LOCKED)
@@ -3069,7 +3069,7 @@ MarkBufferDirty(Buffer buffer)
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
 
-		if (pg_atomic_compare_exchange_u32(&bufHdr->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&bufHdr->state, &old_buf_state,
 										   buf_state))
 			break;
 	}
@@ -3173,10 +3173,10 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 
 	if (ref == NULL)
 	{
-		uint32		buf_state;
-		uint32		old_buf_state;
+		uint64		buf_state;
+		uint64		old_buf_state;
 
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
@@ -3210,7 +3210,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 					buf_state += BUF_USAGECOUNT_ONE;
 			}
 
-			if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+			if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 											   buf_state))
 			{
 				result = (buf_state & BM_VALID) != 0;
@@ -3237,7 +3237,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		 * that the buffer page is legitimately non-accessible here.  We
 		 * cannot meddle with that.
 		 */
-		result = (pg_atomic_read_u32(&buf->state) & BM_VALID) != 0;
+		result = (pg_atomic_read_u64(&buf->state) & BM_VALID) != 0;
 
 		Assert(ref->data.refcount > 0);
 		ref->data.refcount++;
@@ -3272,7 +3272,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	/*
 	 * As explained, We don't expect any preexisting pins. That allows us to
@@ -3284,7 +3284,7 @@ PinBuffer_Locked(BufferDesc *buf)
 	 * Since we hold the buffer spinlock, we can update the buffer state and
 	 * release the lock in one operation.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 
 	UnlockBufHdrExt(buf, old_buf_state,
 					0, 0, 1);
@@ -3314,7 +3314,7 @@ WakePinCountWaiter(BufferDesc *buf)
 	 * BM_PIN_COUNT_WAITER if it stops waiting for a reason other than this
 	 * backend waking it up.
 	 */
-	uint32		buf_state = LockBufHdr(buf);
+	uint64		buf_state = LockBufHdr(buf);
 
 	if ((buf_state & BM_PIN_COUNT_WAITER) &&
 		BUF_STATE_GET_REFCOUNT(buf_state) == 1)
@@ -3361,7 +3361,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	ref->data.refcount--;
 	if (ref->data.refcount == 0)
 	{
-		uint32		old_buf_state;
+		uint64		old_buf_state;
 
 		/*
 		 * Mark buffer non-accessible to Valgrind.
@@ -3379,7 +3379,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
 		/* decrement the shared reference count */
-		old_buf_state = pg_atomic_fetch_sub_u32(&buf->state, BUF_REFCOUNT_ONE);
+		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
 
 		/* Support LockBufferForCleanup() */
 		if (old_buf_state & BM_PIN_COUNT_WAITER)
@@ -3436,7 +3436,7 @@ TrackNewBufferPin(Buffer buf)
 static void
 BufferSync(int flags)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	int			buf_id;
 	int			num_to_scan;
 	int			num_spaces;
@@ -3446,7 +3446,7 @@ BufferSync(int flags)
 	Oid			last_tsid;
 	binaryheap *ts_heap;
 	int			i;
-	uint32		mask = BM_DIRTY;
+	uint64		mask = BM_DIRTY;
 	WritebackContext wb_context;
 
 	/*
@@ -3478,7 +3478,7 @@ BufferSync(int flags)
 	for (buf_id = 0; buf_id < NBuffers; buf_id++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
-		uint32		set_bits = 0;
+		uint64		set_bits = 0;
 
 		/*
 		 * Header spinlock is enough to examine BM_DIRTY, see comment in
@@ -3645,7 +3645,7 @@ BufferSync(int flags)
 		 * write the buffer though we didn't need to.  It doesn't seem worth
 		 * guarding against this, though.
 		 */
-		if (pg_atomic_read_u32(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
+		if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
 		{
 			if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
 			{
@@ -4015,7 +4015,7 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 {
 	BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
 	int			result = 0;
-	uint32		buf_state;
+	uint64		buf_state;
 	BufferTag	tag;
 
 	/* Make sure we can handle the pin */
@@ -4264,7 +4264,7 @@ DebugPrintBufferRefcount(Buffer buffer)
 	int32		loccount;
 	char	   *result;
 	ProcNumber	backend;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 	if (BufferIsLocal(buffer))
@@ -4281,9 +4281,9 @@ DebugPrintBufferRefcount(Buffer buffer)
 	}
 
 	/* theoretically we should lock the bufHdr here */
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
-	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%x, refcount=%u %d)",
+	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%" PRIx64 ", refcount=%u %d)",
 					  buffer,
 					  relpathbackend(BufTagGetRelFileLocator(&buf->tag), backend,
 									 BufTagGetForkNum(&buf->tag)).str,
@@ -4383,7 +4383,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	instr_time	io_start;
 	Block		bufBlock;
 	char	   *bufToWrite;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
@@ -4581,7 +4581,7 @@ BufferIsPermanent(Buffer buffer)
 	 * not random garbage.
 	 */
 	bufHdr = GetBufferDescriptor(buffer - 1);
-	return (pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT) != 0;
+	return (pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT) != 0;
 }
 
 /*
@@ -5044,11 +5044,11 @@ FlushRelationBuffers(Relation rel)
 	{
 		for (i = 0; i < NLocBuffer; i++)
 		{
-			uint32		buf_state;
+			uint64		buf_state;
 
 			bufHdr = GetLocalBufferDescriptor(i);
 			if (BufTagMatchesRelFileLocator(&bufHdr->tag, &rel->rd_locator) &&
-				((buf_state = pg_atomic_read_u32(&bufHdr->state)) &
+				((buf_state = pg_atomic_read_u64(&bufHdr->state)) &
 				 (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 			{
 				ErrorContextCallback errcallback;
@@ -5084,7 +5084,7 @@ FlushRelationBuffers(Relation rel)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5156,7 +5156,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 	{
 		SMgrSortArray *srelent = NULL;
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -5405,7 +5405,7 @@ FlushDatabaseBuffers(Oid dbid)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5553,13 +5553,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u32(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
 		(BM_DIRTY | BM_JUST_DIRTIED))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
 		bool		delayChkptFlags = false;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * If we need to protect hint bit updates from torn writes, WAL-log a
@@ -5571,7 +5571,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
 		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT))
+			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5671,8 +5671,8 @@ UnlockBuffers(void)
 
 	if (buf)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		buf_state = LockBufHdr(buf);
 
@@ -5803,8 +5803,8 @@ LockBufferForCleanup(Buffer buffer)
 
 	for (;;)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		/* Try to acquire lock */
 		LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
@@ -5952,7 +5952,7 @@ bool
 ConditionalLockBufferForCleanup(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state,
+	uint64		buf_state,
 				refcount;
 
 	Assert(BufferIsValid(buffer));
@@ -6010,7 +6010,7 @@ bool
 IsBufferCleanupOK(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 
@@ -6066,7 +6066,7 @@ WaitIO(BufferDesc *buf)
 	ConditionVariablePrepareToSleep(cv);
 	for (;;)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 		PgAioWaitRef iow;
 
 		/*
@@ -6140,7 +6140,7 @@ WaitIO(BufferDesc *buf)
 bool
 StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	ResourceOwnerEnlarge(CurrentResourceOwner);
 
@@ -6196,11 +6196,11 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
  * is being released)
  */
 void
-TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 				  bool forget_owner, bool release_aio)
 {
-	uint32		buf_state;
-	uint32		unset_flag_bits = 0;
+	uint64		buf_state;
+	uint64		unset_flag_bits = 0;
 	int			refcount_change = 0;
 
 	buf_state = LockBufHdr(buf);
@@ -6261,7 +6261,7 @@ static void
 AbortBufferIO(Buffer buffer)
 {
 	BufferDesc *buf_hdr = GetBufferDescriptor(buffer - 1);
-	uint32		buf_state;
+	uint64		buf_state;
 
 	buf_state = LockBufHdr(buf_hdr);
 	Assert(buf_state & (BM_IO_IN_PROGRESS | BM_TAG_VALID));
@@ -6355,10 +6355,10 @@ rlocator_comparator(const void *p1, const void *p2)
 /*
  * Lock buffer header - set BM_LOCKED in buffer state.
  */
-uint32
+uint64
 LockBufHdr(BufferDesc *desc)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	Assert(!BufferIsLocal(BufferDescriptorGetBuffer(desc)));
 
@@ -6369,7 +6369,7 @@ LockBufHdr(BufferDesc *desc)
 		 * the spin-delay infrastructure. The work necessary for that shows up
 		 * in profiles and is rarely necessary.
 		 */
-		old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
+		old_buf_state = pg_atomic_fetch_or_u64(&desc->state, BM_LOCKED);
 		if (likely(!(old_buf_state & BM_LOCKED)))
 			break;				/* got lock */
 
@@ -6382,7 +6382,7 @@ LockBufHdr(BufferDesc *desc)
 			while (old_buf_state & BM_LOCKED)
 			{
 				perform_spin_delay(&delayStatus);
-				old_buf_state = pg_atomic_read_u32(&desc->state);
+				old_buf_state = pg_atomic_read_u64(&desc->state);
 			}
 			finish_spin_delay(&delayStatus);
 		}
@@ -6403,20 +6403,20 @@ LockBufHdr(BufferDesc *desc)
  * Obviously the buffer could be locked by the time the value is returned, so
  * this is primarily useful in CAS style loops.
  */
-pg_noinline uint32
+pg_noinline uint64
 WaitBufHdrUnlocked(BufferDesc *buf)
 {
 	SpinDelayStatus delayStatus;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	init_local_spin_delay(&delayStatus);
 
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
 	while (buf_state & BM_LOCKED)
 	{
 		perform_spin_delay(&delayStatus);
-		buf_state = pg_atomic_read_u32(&buf->state);
+		buf_state = pg_atomic_read_u64(&buf->state);
 	}
 
 	finish_spin_delay(&delayStatus);
@@ -6704,12 +6704,12 @@ ResOwnerPrintBufferPin(Datum res)
 static bool
 EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result;
 
 	*buffer_flushed = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -6803,12 +6803,12 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -6855,7 +6855,7 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
@@ -6897,12 +6897,12 @@ static bool
 MarkDirtyUnpinnedBufferInternal(Buffer buf, BufferDesc *desc,
 								bool *buffer_already_dirty)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result = false;
 
 	*buffer_already_dirty = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -7000,7 +7000,7 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_already_dirty;
 
 		CHECK_FOR_INTERRUPTS();
@@ -7054,12 +7054,12 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_already_dirty;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -7110,7 +7110,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		BufferDesc *buf_hdr = is_temp ?
 			GetLocalBufferDescriptor(-buffer - 1)
 			: GetBufferDescriptor(buffer - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * Check that all the buffers are actually ones that could conceivably
@@ -7128,7 +7128,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		}
 
 		if (is_temp)
-			buf_state = pg_atomic_read_u32(&buf_hdr->state);
+			buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		else
 			buf_state = LockBufHdr(buf_hdr);
 
@@ -7166,7 +7166,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		if (is_temp)
 		{
 			buf_state += BUF_REFCOUNT_ONE;
-			pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 		}
 		else
 			UnlockBufHdrExt(buf_hdr, buf_state, 0, 0, 1);
@@ -7352,13 +7352,13 @@ buffer_readv_complete_one(PgAioTargetData *td, uint8 buf_off, Buffer buffer,
 		: GetBufferDescriptor(buffer - 1);
 	BufferTag	tag = buf_hdr->tag;
 	char	   *bufdata = BufferGetBlock(buffer);
-	uint32		set_flag_bits;
+	uint64		set_flag_bits;
 	int			piv_flags;
 
 	/* check that the buffer is in the expected state for a read */
 #ifdef USE_ASSERT_CHECKING
 	{
-		uint32		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(!(buf_state & BM_VALID));
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 9a93fb335fc..b7687836188 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -86,7 +86,7 @@ typedef struct BufferAccessStrategyData
 
 /* Prototypes for internal functions */
 static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
-									 uint32 *buf_state);
+									 uint64 *buf_state);
 static void AddBufferToRing(BufferAccessStrategy strategy,
 							BufferDesc *buf);
 
@@ -171,7 +171,7 @@ ClockSweepTick(void)
  *	before returning.
  */
 BufferDesc *
-StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_ring)
+StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_ring)
 {
 	BufferDesc *buf;
 	int			bgwprocno;
@@ -230,8 +230,8 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 	trycounter = NBuffers;
 	for (;;)
 	{
-		uint32		old_buf_state;
-		uint32		local_buf_state;
+		uint64		old_buf_state;
+		uint64		local_buf_state;
 
 		buf = GetBufferDescriptor(ClockSweepTick());
 
@@ -239,7 +239,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 		 * Check whether the buffer can be used and pin it if so. Do this
 		 * using a CAS loop, to avoid having to lock the buffer header.
 		 */
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			local_buf_state = old_buf_state;
@@ -277,7 +277,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 			{
 				local_buf_state -= BUF_USAGECOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					trycounter = NBuffers;
@@ -289,7 +289,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 				/* pin the buffer if the CAS succeeds */
 				local_buf_state += BUF_REFCOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					/* Found a usable buffer */
@@ -655,12 +655,12 @@ FreeAccessStrategy(BufferAccessStrategy strategy)
  * returning.
  */
 static BufferDesc *
-GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
+GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
 {
 	BufferDesc *buf;
 	Buffer		bufnum;
-	uint32		old_buf_state;
-	uint32		local_buf_state;	/* to avoid repeated (de-)referencing */
+	uint64		old_buf_state;
+	uint64		local_buf_state;	/* to avoid repeated (de-)referencing */
 
 
 	/* Advance to next ring slot */
@@ -682,7 +682,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 	 * Check whether the buffer can be used and pin it if so. Do this using a
 	 * CAS loop, to avoid having to lock the buffer header.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 	for (;;)
 	{
 		local_buf_state = old_buf_state;
@@ -710,7 +710,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 		/* pin the buffer if the CAS succeeds */
 		local_buf_state += BUF_REFCOUNT_ONE;
 
-		if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 										   local_buf_state))
 		{
 			*buf_state = local_buf_state;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index f6e2b1aa288..04a540379a2 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -148,7 +148,7 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 	}
 	else
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		victim_buffer = GetLocalVictimBuffer();
 		bufid = -victim_buffer - 1;
@@ -165,10 +165,10 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 		 */
 		bufHdr->tag = newTag;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
 		buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 		*foundPtr = false;
 	}
@@ -245,12 +245,12 @@ GetLocalVictimBuffer(void)
 
 		if (LocalRefCount[victim_bufid] == 0)
 		{
-			uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 			if (BUF_STATE_GET_USAGECOUNT(buf_state) > 0)
 			{
 				buf_state -= BUF_USAGECOUNT_ONE;
-				pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+				pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 				trycounter = NLocBuffer;
 			}
 			else if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
@@ -286,13 +286,13 @@ GetLocalVictimBuffer(void)
 	 * this buffer is not referenced but it might still be dirty. if that's
 	 * the case, write it out before reusing it!
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY)
 		FlushLocalBuffer(bufHdr, NULL);
 
 	/*
 	 * Remove the victim buffer from the hashtable and mark as invalid.
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID)
 	{
 		InvalidateLocalBuffer(bufHdr, false);
 
@@ -417,7 +417,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 		if (found)
 		{
 			BufferDesc *existing_hdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			UnpinLocalBuffer(BufferDescriptorGetBuffer(victim_buf_hdr));
 
@@ -428,18 +428,18 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 			/*
 			 * Clear the BM_VALID bit, do StartLocalBufferIO() and proceed.
 			 */
-			buf_state = pg_atomic_read_u32(&existing_hdr->state);
+			buf_state = pg_atomic_read_u64(&existing_hdr->state);
 			Assert(buf_state & BM_TAG_VALID);
 			Assert(!(buf_state & BM_DIRTY));
 			buf_state &= ~BM_VALID;
-			pg_atomic_unlocked_write_u32(&existing_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&existing_hdr->state, buf_state);
 
 			/* no need to loop for local buffers */
 			StartLocalBufferIO(existing_hdr, true, false);
 		}
 		else
 		{
-			uint32		buf_state = pg_atomic_read_u32(&victim_buf_hdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&victim_buf_hdr->state);
 
 			Assert(!(buf_state & (BM_VALID | BM_TAG_VALID | BM_DIRTY | BM_JUST_DIRTIED)));
 
@@ -447,7 +447,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 
 			buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 
-			pg_atomic_unlocked_write_u32(&victim_buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&victim_buf_hdr->state, buf_state);
 
 			hresult->id = victim_buf_id;
 
@@ -467,13 +467,13 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 	{
 		Buffer		buf = buffers[i];
 		BufferDesc *buf_hdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		buf_state |= BM_VALID;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 
 	*extended_by = extend_by;
@@ -492,7 +492,7 @@ MarkLocalBufferDirty(Buffer buffer)
 {
 	int			bufid;
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsLocal(buffer));
 
@@ -506,14 +506,14 @@ MarkLocalBufferDirty(Buffer buffer)
 
 	bufHdr = GetLocalBufferDescriptor(bufid);
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	if (!(buf_state & BM_DIRTY))
 		pgBufferUsage.local_blks_dirtied++;
 
 	buf_state |= BM_DIRTY;
 
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -522,7 +522,7 @@ MarkLocalBufferDirty(Buffer buffer)
 bool
 StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * With AIO the buffer could have IO in progress, e.g. when there are two
@@ -542,7 +542,7 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 	/* Once we get here, there is definitely no I/O active on this buffer */
 
 	/* Check if someone else already did the I/O */
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 	if (forInput ? (buf_state & BM_VALID) : !(buf_state & BM_DIRTY))
 	{
 		return false;
@@ -559,11 +559,11 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
  * Like TerminateBufferIO, but for local buffers
  */
 void
-TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bits,
+TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint64 set_flag_bits,
 					   bool release_aio)
 {
 	/* Only need to adjust flags */
-	uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+	uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/* BM_IO_IN_PROGRESS isn't currently used for local buffers */
 
@@ -582,7 +582,7 @@ TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bit
 	}
 
 	buf_state |= set_flag_bits;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 	/* local buffers don't track IO using resowners */
 
@@ -606,7 +606,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(bufHdr);
 	int			bufid = -buffer - 1;
-	uint32		buf_state;
+	uint64		buf_state;
 	LocalBufferLookupEnt *hresult;
 
 	/*
@@ -622,7 +622,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 		Assert(!pgaio_wref_valid(&bufHdr->io_wref));
 	}
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/*
 	 * We need to test not just LocalRefCount[bufid] but also the BufferDesc
@@ -647,7 +647,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 	ClearBufferTag(&bufHdr->tag);
 	buf_state &= ~BUF_FLAG_MASK;
 	buf_state &= ~BUF_USAGECOUNT_MASK;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -671,9 +671,9 @@ DropRelationLocalBuffers(RelFileLocator rlocator, ForkNumber *forkNum,
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (!(buf_state & BM_TAG_VALID) ||
 			!BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -706,9 +706,9 @@ DropRelationAllLocalBuffers(RelFileLocator rlocator)
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if ((buf_state & BM_TAG_VALID) &&
 			BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -804,11 +804,11 @@ InitLocalBuffers(void)
 bool
 PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	Buffer		buffer = BufferDescriptorGetBuffer(buf_hdr);
 	int			bufid = -buffer - 1;
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	if (LocalRefCount[bufid] == 0)
 	{
@@ -819,7 +819,7 @@ PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 		{
 			buf_state += BUF_USAGECOUNT_ONE;
 		}
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/*
 		 * See comment in PinBuffer().
@@ -856,14 +856,14 @@ UnpinLocalBufferNoOwner(Buffer buffer)
 	if (--LocalRefCount[buffid] == 0)
 	{
 		BufferDesc *buf_hdr = GetLocalBufferDescriptor(buffid);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		NLocalPinnedBuffers--;
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state -= BUF_REFCOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/* see comment in UnpinBufferNoOwner */
 		VALGRIND_MAKE_MEM_NOACCESS(LocalBufHdrGetBlock(buf_hdr), BLCKSZ);
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index 0c58e4b265c..529803346ce 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -199,7 +199,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
 		for (i = 0; i < NBuffers; i++)
 		{
 			BufferDesc *bufHdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			CHECK_FOR_INTERRUPTS();
 
@@ -615,7 +615,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -626,7 +626,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 		 * noticeably increase the cost of the function.
 		 */
 		bufHdr = GetBufferDescriptor(i);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (buf_state & BM_VALID)
 		{
@@ -676,7 +676,7 @@ pg_buffercache_usage_counts(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		int			usage_count;
 
 		CHECK_FOR_INTERRUPTS();
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index e046b08f3d5..b1aa8af9ec0 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -308,9 +308,9 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 {
 	Buffer		buf;
 	BufferDesc *buf_hdr;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		was_pinned = false;
-	uint32		unset_bits = 0;
+	uint64		unset_bits = 0;
 
 	/* place buffer in shared buffers without erroring out */
 	buf = ReadBufferExtended(rel, MAIN_FORKNUM, blkno, RBM_ZERO_AND_LOCK, NULL);
@@ -319,7 +319,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	}
 	else
 	{
@@ -340,7 +340,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_state &= ~unset_bits;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 	else
 	{
@@ -489,7 +489,7 @@ invalidate_rel_block(PG_FUNCTION_ARGS)
 
 			LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
-			if (pg_atomic_read_u32(&buf_hdr->state) & BM_DIRTY)
+			if (pg_atomic_read_u64(&buf_hdr->state) & BM_DIRTY)
 			{
 				if (BufferIsLocal(buf))
 					FlushLocalBuffer(buf_hdr, NULL);
@@ -572,7 +572,7 @@ buffer_call_terminate_io(PG_FUNCTION_ARGS)
 	bool		io_error = PG_GETARG_BOOL(3);
 	bool		release_aio = PG_GETARG_BOOL(4);
 	bool		clear_dirty = false;
-	uint32		set_flag_bits = 0;
+	uint64		set_flag_bits = 0;
 
 	if (io_error)
 		set_flag_bits |= BM_IO_ERROR;
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v10-0004-bufmgr-Implement-buffer-content-locks-independen.patch (47.0K, ../../jtg5cu4n6h5lib3kzx66ju4yhh6kmviaud7oq6dtut6c4q4rdi@xwsfoagt3c2b/5-v10-0004-bufmgr-Implement-buffer-content-locks-independen.patch)
  download | inline diff:
From 83cc003fa364dd5e0108c506f8b5c7ddc74f70e7 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 19 Nov 2025 16:37:26 -0500
Subject: [PATCH v10 4/8] bufmgr: Implement buffer content locks independently
 of lwlocks

Until now buffer content locks were implemented using lwlocks. That has the
obvious advantage of not needing a separate efficient implementation of
locks. However, the time for a dedicated buffer content lock implementation
has come:

1) Hint bits are currently set while holding only a share lock. This leads to
   having to copy pages while they are being written out if checksums are
   enabled, which is not cheap. We would like to add AIO writes, however once
   many buffers can be written out at the same time, it gets a lot more
   expensive to copy them, particularly because that copy needs to reside in
   shared buffers (for worker mode to have access to the buffer).

   In addition, modifying buffers while they are being written out can cause
   issues with unbuffered/direct-IO, as some filesystems (like btrfs) do not
   like that, due to filesystem internal checksums getting corrupted.

   The solution to this is to require a new share-exclusive lock-level to set
   hint bits and to write out buffers, making those operations mutually
   exclusive. We could introduce such a lock level into the generic lwlock
   implementation, however it does not look like there would be other users,
   and it does add some overhead into important code paths.

2) For AIO writes we need to be able to race-freely check whether a buffer is
   undergoing IO and whether an exclusive lock on the page can be acquired. That
   is rather hard to do efficiently when the buffer state and the lock state
   are separate atomic variables. This is a major hindrance to allowing writes
   to be done asynchronously.

3) Buffer locks are by far the most frequently taken locks. Optimizing them
   specifically for their use case is worth the effort. E.g. by merging
   content locks into buffer locks we will be able to release a buffer lock
   and pin in one atomic operation.

4) There are more complicated optimizations, like long-lived "super pinned &
   locked" pages, that cannot realistically be implemented with the generic
   lwlock implementation.

Therefore implement content locks inside bufmgr.c. The lockstate is stored as
part of BufferDesc.state. The implementation of buffer content locks is fairly
similar to lwlocks, with a few important differences:

1) An additional lock-level share-exclusive has been added. This lock level
   conflicts with exclusive locks and itself, but not share locks.

2) Error recovery for content locks is implemented as part of the already
   existing private-refcount tracking mechanism in combination with resowners,
   instead of a bespoke mechanism as the case for lwlocks. This means we do
   not need to add dedicated error-recovery code paths to release all content
   locks (like done with LWLockReleaseAll() for lwlocks).

3) The lock state is embedded in BufferDesc.state instead of having its own
   struct.

4) The wakeup logic is a tad more complicated due to needing to support the
   additional lock level

This commit unfortunately introduces some code that is very similar to the
code in lwlock.c, however the code is not equivalent enough to easily merge
it. The future wins that this commit makes possible seem worth the cost.

As of this commit nothing uses the new share-exclusive lock mode. It will be
used in a future commit. It seemed too complicated to introduce the lock-level
in a separate commit.

TODO:
- Decide whether we need to do something about the FIXME, I'm inclined to
  think the reuse of the PGPROC->lw* fields is the lesser evil for now.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Reviewed-by: Greg Burd <greg@burd.me>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  67 +-
 src/include/storage/bufmgr.h                  |  32 +-
 src/include/storage/proc.h                    |   8 +-
 src/backend/storage/buffer/buf_init.c         |   5 +-
 src/backend/storage/buffer/bufmgr.c           | 896 ++++++++++++++++--
 .../utils/activity/wait_event_names.txt       |   3 +
 6 files changed, 919 insertions(+), 92 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index a4d36e9ca01..12086cf6dc7 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -23,6 +23,7 @@
 #include "storage/condition_variable.h"
 #include "storage/lwlock.h"
 #include "storage/procnumber.h"
+#include "storage/proclist_types.h"
 #include "storage/shmem.h"
 #include "storage/smgr.h"
 #include "storage/spin.h"
@@ -35,22 +36,23 @@
  * State of the buffer itself (in order):
  * - 18 bits refcount
  * - 4 bits usage count
- * - 10 bits of flags
+ * - 12 bits of flags
+ * - 18 bits share-lock count
+ * - 1 bit share-exclusive locked
+ * - 1 bit exclusive locked
  *
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
- * NB: A future commit will use a significant portion of the remaining bits to
- * implement buffer locking as part of the state variable.
- *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
 #define BUF_USAGECOUNT_BITS 4
-#define BUF_FLAG_BITS 10
+#define BUF_FLAG_BITS 12
+#define BUF_LOCK_BITS (18+2)
 
-StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
-				 "parts of buffer state space need to equal 32");
+StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS + BUF_LOCK_BITS <= 64,
+				 "parts of buffer state space need to be <= 64");
 
 /* refcount related definitions */
 #define BUF_REFCOUNT_ONE 1
@@ -71,6 +73,19 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 #define BUF_FLAG_MASK \
 	(((UINT64CONST(1) << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
 
+/* lock state related definitions */
+#define BM_LOCK_SHIFT \
+	(BUF_FLAG_SHIFT + BUF_FLAG_BITS)
+#define BM_LOCK_VAL_SHARED \
+	(UINT64CONST(1) << (BM_LOCK_SHIFT))
+#define BM_LOCK_VAL_SHARE_EXCLUSIVE \
+	(UINT64CONST(1) << (BM_LOCK_SHIFT + MAX_BACKENDS_BITS))
+#define BM_LOCK_VAL_EXCLUSIVE \
+	(UINT64CONST(1) << (BM_LOCK_SHIFT + MAX_BACKENDS_BITS + 1))
+#define BM_LOCK_MASK \
+	((((uint64) MAX_BACKENDS) << BM_LOCK_SHIFT) | BM_LOCK_VAL_SHARE_EXCLUSIVE | BM_LOCK_VAL_EXCLUSIVE)
+
+
 /* Get refcount and usagecount from buffer state */
 #define BUF_STATE_GET_REFCOUNT(state) \
 	((uint32)((state) & BUF_REFCOUNT_MASK))
@@ -107,6 +122,17 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 #define BM_CHECKPOINT_NEEDED		BUF_DEFINE_FLAG( 8)
 /* permanent buffer (not unlogged, or init fork) */
 #define BM_PERMANENT				BUF_DEFINE_FLAG( 9)
+/* content lock has waiters */
+#define BM_LOCK_HAS_WAITERS			BUF_DEFINE_FLAG(10)
+/* waiter for content lock has been signalled but not yet run */
+#define BM_LOCK_WAKE_IN_PROGRESS	BUF_DEFINE_FLAG(11)
+
+
+StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
+				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
+StaticAssertDecl(MAX_BACKENDS_BITS <= (BUF_LOCK_BITS - 2),
+				 "MAX_BACKENDS_BITS needs to be <= BUF_LOCK_BITS - 2");
+
 
 /*
  * The maximum allowed value of usage_count represents a tradeoff between
@@ -120,8 +146,6 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 
 StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 				 "BM_MAX_USAGE_COUNT doesn't fit in BUF_USAGECOUNT_BITS bits");
-StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
-				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
 
 /*
  * Buffer tag identifies which disk block the buffer contains.
@@ -265,9 +289,6 @@ BufMappingPartitionLockByIndex(uint32 index)
  * it is held.  However, existing buffer pins may be released while the buffer
  * header spinlock is held, using an atomic subtraction.
  *
- * The LWLock can take care of itself.  The buffer header lock is *not* used
- * to control access to the data in the buffer!
- *
  * If we have the buffer pinned, its tag can't change underneath us, so we can
  * examine the tag without locking the buffer header.  Also, in places we do
  * one-time reads of the flags without bothering to lock the buffer header;
@@ -280,6 +301,15 @@ BufMappingPartitionLockByIndex(uint32 index)
  * wait_backend_pgprocno and setting flag bit BM_PIN_COUNT_WAITER.  At present,
  * there can be only one such waiter per buffer.
  *
+ * The content of buffers is protected via the buffer content lock,
+ * implemented as part of the buffer state. Note that the buffer header lock
+ * is *not* used to control access to the data in the buffer! We used to use
+ * an LWLock to implement the content lock, but having a dedicated
+ * implementation of content locks allows us to implement some otherwise hard
+ * things (e.g.  race-freely checking if AIO is in progress before locking a
+ * buffer exclusively) and enables otherwise impossible optimizations
+ * (e.g. unlocking and unpinning a buffer in one atomic operation).
+ *
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
@@ -321,7 +351,12 @@ typedef struct BufferDesc
 	int			wait_backend_pgprocno;
 
 	PgAioWaitRef io_wref;		/* set iff AIO is in progress */
-	LWLock		content_lock;	/* to lock access to buffer contents */
+
+	/*
+	 * List of PGPROCs waiting for the buffer content lock. Protected by the
+	 * buffer header spinlock.
+	 */
+	proclist_head lock_waiters;
 } BufferDesc;
 
 /*
@@ -408,12 +443,6 @@ BufferDescriptorGetIOCV(const BufferDesc *bdesc)
 	return &(BufferIOCVArray[bdesc->buf_id]).cv;
 }
 
-static inline LWLock *
-BufferDescriptorGetContentLock(const BufferDesc *bdesc)
-{
-	return (LWLock *) (&bdesc->content_lock);
-}
-
 /*
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 715ae96f0f0..a40adf6b2a8 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -203,7 +203,20 @@ extern PGDLLIMPORT int32 *LocalRefCount;
 typedef enum BufferLockMode
 {
 	BUFFER_LOCK_UNLOCK,
+
+	/*
+	 * A share lock conflicts with exclusive locks.
+	 */
 	BUFFER_LOCK_SHARE,
+
+	/*
+	 * A share-exclusive lock conflicts with itself and exclusive locks.
+	 */
+	BUFFER_LOCK_SHARE_EXCLUSIVE,
+
+	/*
+	 * An exclusive lock conflicts with every other lock type.
+	 */
 	BUFFER_LOCK_EXCLUSIVE,
 } BufferLockMode;
 
@@ -302,7 +315,24 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
 extern void UnlockBuffers(void);
-extern void LockBuffer(Buffer buffer, BufferLockMode mode);
+extern void UnlockBuffer(Buffer buffer);
+extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
+
+/*
+ * Handling BUFFER_LOCK_UNLOCK in bufmgr.c leads to sufficiently worse branch
+ * prediction to impact performance. Therefore handle that switch here, where
+ * most of the time `mode` will be a constant and thus can be optimized out by
+ * the compiler.
+ */
+static inline void
+LockBuffer(Buffer buffer, BufferLockMode mode)
+{
+	if (mode == BUFFER_LOCK_UNLOCK)
+		UnlockBuffer(buffer);
+	else
+		LockBufferInternal(buffer, mode);
+}
+
 extern bool ConditionalLockBuffer(Buffer buffer);
 extern void LockBufferForCleanup(Buffer buffer);
 extern bool ConditionalLockBufferForCleanup(Buffer buffer);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index de7b2e0bd2c..039bc8353be 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -242,7 +242,13 @@ struct PGPROC
 	 */
 	bool		recoveryConflictPending;
 
-	/* Info about LWLock the process is currently waiting for, if any. */
+	/*
+	 * Info about LWLock the process is currently waiting for, if any.
+	 *
+	 * This is currently used both for lwlocks and buffer content locks, which
+	 * is acceptable, although not pretty, because a backend can't wait for
+	 * both types of locks at the same time.
+	 */
 	uint8		lwWaiting;		/* see LWLockWaitState */
 	uint8		lwWaitMode;		/* lwlock mode being waited for */
 	proclist_node lwWaitLink;	/* position in LW lock wait list */
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 7d894522526..c0c223b2e32 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
 #include "storage/aio.h"
 #include "storage/buf_internals.h"
 #include "storage/bufmgr.h"
+#include "storage/proclist.h"
 
 BufferDescPadded *BufferDescriptors;
 char	   *BufferBlocks;
@@ -128,9 +129,7 @@ BufferManagerShmemInit(void)
 
 			pgaio_wref_clear(&buf->io_wref);
 
-			LWLockInitialize(BufferDescriptorGetContentLock(buf),
-							 LWTRANCHE_BUFFER_CONTENT);
-
+			proclist_init(&buf->lock_waiters);
 			ConditionVariableInit(BufferDescriptorGetIOCV(buf));
 		}
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index b0de8e45d4d..0d5da094748 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -58,6 +58,7 @@
 #include "storage/ipc.h"
 #include "storage/lmgr.h"
 #include "storage/proc.h"
+#include "storage/proclist.h"
 #include "storage/read_stream.h"
 #include "storage/smgr.h"
 #include "storage/standby.h"
@@ -100,6 +101,12 @@ typedef struct PrivateRefCountData
 	 * How many times has the buffer been pinned by this backend.
 	 */
 	int32		refcount;
+
+	/*
+	 * Is the buffer locked by this backend? BUFFER_LOCK_UNLOCK indicates that
+	 * the buffer is not locked.
+	 */
+	BufferLockMode lockmode;
 } PrivateRefCountData;
 
 typedef struct PrivateRefCountEntry
@@ -210,8 +217,10 @@ static BufferDesc *PinCountWaitBuf = NULL;
  * Each buffer also has a private refcount that keeps track of the number of
  * times the buffer is pinned in the current process.  This is so that the
  * shared refcount needs to be modified only once if a buffer is pinned more
- * than once by an individual backend.  It's also used to check that no buffers
- * are still pinned at the end of transactions and when exiting.
+ * than once by an individual backend.  It's also used to check that no
+ * buffers are still pinned at the end of transactions and when exiting. We
+ * also use this mechanism to track whether this backend has a buffer locked,
+ * and, if so, in what mode.
  *
  *
  * To avoid - as we used to - requiring an array with NBuffers entries to keep
@@ -351,6 +360,7 @@ ReservePrivateRefCountEntry(void)
 		/* clear the whole data member, just for future proofing */
 		memset(&victim_entry->data, 0, sizeof(victim_entry->data));
 		victim_entry->data.refcount = 0;
+		victim_entry->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 		PrivateRefCountOverflowed++;
 	}
@@ -374,6 +384,7 @@ NewPrivateRefCountEntry(Buffer buffer)
 	PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
 	res->buffer = buffer;
 	res->data.refcount = 0;
+	res->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 	/* update cache for the next lookup */
 	PrivateRefCountEntryLast = ReservedRefCountSlot;
@@ -540,6 +551,7 @@ static void
 ForgetPrivateRefCountEntry(PrivateRefCountEntry *ref)
 {
 	Assert(ref->data.refcount == 0);
+	Assert(ref->data.lockmode == BUFFER_LOCK_UNLOCK);
 
 	if (ref >= &PrivateRefCountArray[0] &&
 		ref < &PrivateRefCountArray[REFCOUNT_ARRAY_ENTRIES])
@@ -641,14 +653,27 @@ static void RelationCopyStorageUsingBuffer(RelFileLocator srclocator,
 static void AtProcExit_Buffers(int code, Datum arg);
 static void CheckForBufferLeaks(void);
 #ifdef USE_ASSERT_CHECKING
-static void AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
-									   void *unused_context);
+static void AssertNotCatalogBufferLock(Buffer buffer, BufferLockMode mode);
 #endif
 static int	rlocator_comparator(const void *p1, const void *p2);
 static inline int buffertag_comparator(const BufferTag *ba, const BufferTag *bb);
 static inline int ckpt_buforder_comparator(const CkptSortItem *a, const CkptSortItem *b);
 static int	ts_ckpt_progress_comparator(Datum a, Datum b, void *arg);
 
+static void BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr);
+static bool BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMe(BufferDesc *buf_hdr);
+static inline void BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr);
+static inline int BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr);
+static inline bool BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockDequeueSelf(BufferDesc *buf_hdr);
+static void BufferLockWakeup(BufferDesc *buf_hdr, bool unlocked);
+static void BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate);
+static inline uint64 BufferLockReleaseSub(BufferLockMode mode);
+
 
 /*
  * Implementation of PrefetchBuffer() for shared buffers.
@@ -2306,6 +2331,12 @@ retry:
 		goto retry;
 	}
 
+	/*
+	 * An invalidated buffer should not have any backends waiting to lock the
+	 * buffer, therefore BM_LOCK_WAKE_IN_PROGRESS should not be set.
+	 */
+	Assert(!(buf_state & BM_LOCK_WAKE_IN_PROGRESS));
+
 	/*
 	 * Clear out the buffer's tag and flags.  We must do this to ensure that
 	 * linear scans of the buffer array don't think the buffer is valid.
@@ -2382,6 +2413,12 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 		return false;
 	}
 
+	/*
+	 * An invalidated buffer should not have any backends waiting to lock the
+	 * buffer, therefore BM_LOCK_WAKE_IN_PROGRESS should not be set.
+	 */
+	Assert(!(buf_state & BM_LOCK_WAKE_IN_PROGRESS));
+
 	/*
 	 * Clear out the buffer's tag and flags and usagecount.  This is not
 	 * strictly required, as BM_TAG_VALID/BM_VALID needs to be checked before
@@ -2449,8 +2486,6 @@ again:
 	 */
 	if (buf_state & BM_DIRTY)
 	{
-		LWLock	   *content_lock;
-
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(buf_state & BM_VALID);
 
@@ -2468,8 +2503,7 @@ again:
 		 * one just happens to be trying to split the page the first one got
 		 * from StrategyGetBuffer.)
 		 */
-		content_lock = BufferDescriptorGetContentLock(buf_hdr);
-		if (!LWLockConditionalAcquire(content_lock, LW_SHARED))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -2498,7 +2532,7 @@ again:
 			if (XLogNeedsFlush(lsn)
 				&& StrategyRejectBuffer(strategy, buf_hdr, from_ring))
 			{
-				LWLockRelease(content_lock);
+				LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 				UnpinBuffer(buf_hdr);
 				goto again;
 			}
@@ -2506,7 +2540,7 @@ again:
 
 		/* OK, do the I/O */
 		FlushBuffer(buf_hdr, NULL, IOOBJECT_RELATION, io_context);
-		LWLockRelease(content_lock);
+		LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 		ScheduleBufferTagForWriteback(&BackendWritebackContext, io_context,
 									  &buf_hdr->tag);
@@ -2948,7 +2982,7 @@ BufferIsLockedByMe(Buffer buffer)
 	else
 	{
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr));
+		return BufferLockHeldByMe(bufHdr);
 	}
 }
 
@@ -2973,23 +3007,8 @@ BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode)
 	}
 	else
 	{
-		LWLockMode	lw_mode;
-
-		switch (mode)
-		{
-			case BUFFER_LOCK_EXCLUSIVE:
-				lw_mode = LW_EXCLUSIVE;
-				break;
-			case BUFFER_LOCK_SHARE:
-				lw_mode = LW_SHARED;
-				break;
-			default:
-				pg_unreachable();
-		}
-
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-									lw_mode);
+		return BufferLockHeldByMeInMode(bufHdr, mode);
 	}
 }
 
@@ -3376,7 +3395,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 * I'd better not still hold the buffer content lock. Can't use
 		 * BufferIsLockedByMe(), as that asserts the buffer is pinned.
 		 */
-		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
+		Assert(!BufferLockHeldByMe(buf));
 
 		/* decrement the shared reference count */
 		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
@@ -4198,9 +4217,9 @@ CheckForBufferLeaks(void)
  * Check for exclusive-locked catalog buffers.  This is the core of
  * AssertCouldGetRelation().
  *
- * A backend would self-deadlock on LWLocks if the catalog scan read the
- * exclusive-locked buffer.  The main threat is exclusive-locked buffers of
- * catalogs used in relcache, because a catcache search on any catalog may
+ * A backend would self-deadlock on the content lock if the catalog scan read
+ * the exclusive-locked buffer.  The main threat is exclusive-locked buffers
+ * of catalogs used in relcache, because a catcache search on any catalog may
  * build that catalog's relcache entry.  We don't have an inventory of
  * catalogs relcache uses, so just check buffers of most catalogs.
  *
@@ -4214,26 +4233,45 @@ CheckForBufferLeaks(void)
 void
 AssertBufferLocksPermitCatalogRead(void)
 {
-	ForEachLWLockHeldByMe(AssertNotCatalogBufferLock, NULL);
+	PrivateRefCountEntry *res;
+
+	/* check the array */
+	for (int i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
+	{
+		if (PrivateRefCountArrayKeys[i] != InvalidBuffer)
+		{
+			res = &PrivateRefCountArray[i];
+
+			if (res->buffer == InvalidBuffer)
+				continue;
+
+			AssertNotCatalogBufferLock(res->buffer, res->data.lockmode);
+		}
+	}
+
+	/* if necessary search the hash */
+	if (PrivateRefCountOverflowed)
+	{
+		HASH_SEQ_STATUS hstat;
+
+		hash_seq_init(&hstat, PrivateRefCountHash);
+		while ((res = (PrivateRefCountEntry *) hash_seq_search(&hstat)) != NULL)
+		{
+			AssertNotCatalogBufferLock(res->buffer, res->data.lockmode);
+		}
+	}
 }
 
 static void
-AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
-						   void *unused_context)
+AssertNotCatalogBufferLock(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *bufHdr;
+	BufferDesc *bufHdr = GetBufferDescriptor(buffer - 1);
 	BufferTag	tag;
 	Oid			relid;
 
-	if (mode != LW_EXCLUSIVE)
+	if (mode != BUFFER_LOCK_EXCLUSIVE)
 		return;
 
-	if (!((BufferDescPadded *) lock > BufferDescriptors &&
-		  (BufferDescPadded *) lock < BufferDescriptors + NBuffers))
-		return;					/* not a buffer lock */
-
-	bufHdr = (BufferDesc *)
-		((char *) lock - offsetof(BufferDesc, content_lock));
 	tag = bufHdr->tag;
 
 	/*
@@ -4515,9 +4553,11 @@ static void
 FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 					IOObject io_object, IOContext io_context)
 {
-	LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	Buffer		buffer = BufferDescriptorGetBuffer(buf);
+
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-	LWLockRelease(BufferDescriptorGetContentLock(buf));
+	BufferLockUnlock(buffer, buf);
 }
 
 /*
@@ -5660,9 +5700,10 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
  *
  * Used to clean up after errors.
  *
- * Currently, we can expect that lwlock.c's LWLockReleaseAll() took care
- * of releasing buffer content locks per se; the only thing we need to deal
- * with here is clearing any PIN_COUNT request that was in progress.
+ * Currently, we can expect that resource owner cleanup, via
+ * ResOwnerReleaseBufferPin(), took care of releasing buffer content locks per
+ * se; the only thing we need to deal with here is clearing any PIN_COUNT
+ * request that was in progress.
  */
 void
 UnlockBuffers(void)
@@ -5693,25 +5734,728 @@ UnlockBuffers(void)
 }
 
 /*
- * Acquire or release the content_lock for the buffer.
+ * Acquire the buffer content lock in the specified mode
+ *
+ * If the lock is not available, sleep until it is.
+ *
+ * Side effect: cancel/die interrupts are held off until lock release.
+ *
+ * This uses almost the same locking approach as lwlock.c's
+ * LWLockAcquire(). See documentation at the top of lwlock.c for a more
+ * detailed discussion.
+ *
+ * The reason that this, and most of the other BufferLock* functions, get both
+ * the Buffer and BufferDesc* as parameters, is that looking up one from the
+ * other repeatedly shows up noticeably in profiles.
+ *
+ * Callers should provide a constant for mode, for more efficient code
+ * generation.
+ */
+static inline void
+BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry;
+	int			extraWaits = 0;
+
+	/*
+	 * Get reference to the refcount entry before we hold the lock, it seems
+	 * better to do before holding the lock.
+	 */
+	entry = GetPrivateRefCountEntry(buffer, true);
+
+	/*
+	 * We better not already hold a lock on the buffer.
+	 */
+	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	for (;;)
+	{
+		bool		mustwait;
+		uint32		wait_event;
+
+		/*
+		 * Try to grab the lock the first time, we're not in the waitqueue
+		 * yet/anymore.
+		 */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		if (likely(!mustwait))
+		{
+			break;
+		}
+
+		/*
+		 * Ok, at this point we couldn't grab the lock on the first try. We
+		 * cannot simply queue ourselves to the end of the list and wait to be
+		 * woken up because by now the lock could long have been released.
+		 * Instead add us to the queue and try to grab the lock again. If we
+		 * succeed we need to revert the queuing and be happy, otherwise we
+		 * recheck the lock. If we still couldn't grab it, we know that the
+		 * other locker will see our queue entries when releasing since they
+		 * existed before we checked for the lock.
+		 */
+
+		/* add to the queue */
+		BufferLockQueueSelf(buf_hdr, mode);
+
+		/* we're now guaranteed to be woken up if necessary */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		/* ok, grabbed the lock the second time round, need to undo queueing */
+		if (!mustwait)
+		{
+			BufferLockDequeueSelf(buf_hdr);
+			break;
+		}
+
+		switch (mode)
+		{
+			case BUFFER_LOCK_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_SHARE_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE:
+				wait_event = WAIT_EVENT_BUFFER_SHARED;
+				break;
+			case BUFFER_LOCK_UNLOCK:
+				pg_unreachable();
+
+		}
+		pgstat_report_wait_start(wait_event);
+
+		/*
+		 * Wait until awakened.
+		 *
+		 * It is possible that we get awakened for a reason other than being
+		 * signaled by BufferLockWakeup().  If so, loop back and wait again.
+		 * Once we've gotten the lock, re-increment the sema by the number of
+		 * additional signals received.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		pgstat_report_wait_end();
+
+		/* Retrying, allow BufferLockRelease to release waiters again. */
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_WAKE_IN_PROGRESS);
+	}
+
+	/* Remember that we now hold this lock */
+	entry->data.lockmode = mode;
+
+	/*
+	 * Fix the process wait semaphore's count for any absorbed wakeups.
+	 */
+	while (unlikely(extraWaits-- > 0))
+		PGSemaphoreUnlock(MyProc->sem);
+}
+
+/*
+ * Release a previously acquired buffer content lock.
+ */
+static void
+BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	uint64		oldstate;
+	uint64		sub;
+
+	mode = BufferLockDisownInternal(buffer, buf_hdr);
+
+	/*
+	 * Release my hold on lock, after that it can immediately be acquired by
+	 * others, even if we still have to wakeup other waiters.
+	 */
+	sub = BufferLockReleaseSub(mode);
+
+	oldstate = pg_atomic_sub_fetch_u64(&buf_hdr->state, sub);
+
+	BufferLockProcessRelease(buf_hdr, mode, oldstate);
+
+	/*
+	 * Now okay to allow cancel/die interrupts.
+	 */
+	RESUME_INTERRUPTS();
+}
+
+
+/*
+ * Acquire the content lock for the buffer, but only if we don't have to wait.
+ */
+static bool
+BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry = GetPrivateRefCountEntry(buffer, true);
+	bool		mustwait;
+
+	/*
+	 * We better not already hold a lock on the buffer.
+	 */
+	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	/* Check for the lock */
+	mustwait = BufferLockAttempt(buf_hdr, mode);
+
+	if (mustwait)
+	{
+		/* Failed to get lock, so release interrupt holdoff */
+		RESUME_INTERRUPTS();
+	}
+	else
+	{
+		entry->data.lockmode = mode;
+	}
+
+	return !mustwait;
+}
+
+/*
+ * Internal function that tries to atomically acquire the content lock in the
+ * passed in mode.
+ *
+ * This function will not block waiting for a lock to become free - that's the
+ * caller's job.
+ *
+ * Similar to LWLockAttemptLock().
+ */
+static inline bool
+BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	uint64		old_state;
+
+	/*
+	 * Read once outside the loop, later iterations will get the newer value
+	 * via compare & exchange.
+	 */
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+
+	/* loop until we've determined whether we could acquire the lock or not */
+	while (true)
+	{
+		uint64		desired_state;
+		bool		lock_free;
+
+		desired_state = old_state;
+
+		if (mode == BUFFER_LOCK_EXCLUSIVE)
+		{
+			lock_free = (old_state & BM_LOCK_MASK) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_EXCLUSIVE;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			lock_free = (old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+		}
+		else
+		{
+			lock_free = (old_state & BM_LOCK_VAL_EXCLUSIVE) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARED;
+		}
+
+		/*
+		 * Attempt to swap in the state we are expecting. If we didn't see
+		 * lock to be free, that's just the old value. If we saw it as free,
+		 * we'll attempt to mark it acquired. The reason that we always swap
+		 * in the value is that this doubles as a memory barrier. We could try
+		 * to be smarter and only swap in values if we saw the lock as free,
+		 * but benchmark haven't shown it as beneficial so far.
+		 *
+		 * Retry if the value changed since we last looked at it.
+		 */
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			if (lock_free)
+			{
+				/* Great! Got the lock. */
+				return false;
+			}
+			else
+				return true;	/* somebody else has the lock */
+		}
+	}
+
+	pg_unreachable();
+}
+
+/*
+ * Add ourselves to the end of the content lock's wait queue.
+ */
+static void
+BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	/*
+	 * If we don't have a PGPROC structure, there's no way to wait. This
+	 * should never occur, since MyProc should only be null during shared
+	 * memory initialization.
+	 */
+	if (MyProc == NULL)
+		elog(PANIC, "cannot wait without a PGPROC structure");
+
+	if (MyProc->lwWaiting != LW_WS_NOT_WAITING)
+		elog(PANIC, "queueing for lock while waiting on another one");
+
+	LockBufHdr(buf_hdr);
+
+	/* setting the flag is protected by the spinlock */
+	pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_HAS_WAITERS);
+
+	/*
+	 * FIXME: This is reusing the lwlock fields. That's not a correctness
+	 * issue, a backend can't wait for both an lwlock and a buffer content
+	 * lock at the same time. However, it seems pretty ugly, particularly
+	 * given that the field names have an lw* prefix. But duplicating the
+	 * fields also seems somewhat superfluous.
+	 */
+	MyProc->lwWaiting = LW_WS_WAITING;
+	MyProc->lwWaitMode = mode;
+
+	proclist_push_tail(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	/* Can release the mutex now */
+	UnlockBufHdr(buf_hdr);
+}
+
+/*
+ * Remove ourselves from the waitlist.
+ *
+ * This is used if we queued ourselves because we thought we needed to sleep
+ * but, after further checking, we discovered that we don't actually need to
+ * do so.
+ */
+static void
+BufferLockDequeueSelf(BufferDesc *buf_hdr)
+{
+	bool		on_waitlist;
+
+	LockBufHdr(buf_hdr);
+
+	on_waitlist = MyProc->lwWaiting == LW_WS_WAITING;
+	if (on_waitlist)
+		proclist_delete(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	if (proclist_is_empty(&buf_hdr->lock_waiters) &&
+		(pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS) != 0)
+	{
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_HAS_WAITERS);
+	}
+
+	/* XXX: combine with fetch_and above? */
+	UnlockBufHdr(buf_hdr);
+
+	/* clear waiting state again, nice for debugging */
+	if (on_waitlist)
+		MyProc->lwWaiting = LW_WS_NOT_WAITING;
+	else
+	{
+		int			extraWaits = 0;
+
+
+		/*
+		 * Somebody else dequeued us and has or will wake us up. Deal with the
+		 * superfluous absorption of a wakeup.
+		 */
+
+		/*
+		 * Clear BM_LOCK_WAKE_IN_PROGRESS if somebody woke us before we
+		 * removed ourselves - they'll have set it.
+		 */
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_WAKE_IN_PROGRESS);
+
+		/*
+		 * Now wait for the scheduled wakeup, otherwise our ->lwWaiting would
+		 * get reset at some inconvenient point later. Most of the time this
+		 * will immediately return.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		/*
+		 * Fix the process wait semaphore's count for any absorbed wakeups.
+		 */
+		while (extraWaits-- > 0)
+			PGSemaphoreUnlock(MyProc->sem);
+	}
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * After calling this function it's the callers responsibility to ensure that
+ * the lock gets released, even in case of an error. This only is desirable if
+ * the lock is going to be released in a different process than the process
+ * that acquired it.
+ */
+static inline void
+BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockDisownInternal(buffer, buf_hdr);
+	RESUME_INTERRUPTS();
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * This is the code that can be shared between actually releasing a lock
+ * (BufferLockUnlock()) and just not tracking ownership of the lock anymore
+ * without releasing the lock (BufferLockDisown()).
+ */
+static inline int
+BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	PrivateRefCountEntry *ref;
+
+	ref = GetPrivateRefCountEntry(buffer, false);
+	if (ref == NULL)
+		elog(ERROR, "lock %d is not held", buffer);
+	mode = ref->data.lockmode;
+	ref->data.lockmode = BUFFER_LOCK_UNLOCK;
+
+	return mode;
+}
+
+/*
+ * Wakeup all the lockers that currently have a chance to acquire the lock.
+ *
+ * wake_exclusive indicates whether exclusive lock waiters should be woken up.
+ */
+static void
+BufferLockWakeup(BufferDesc *buf_hdr, bool wake_exclusive)
+{
+	bool		new_wake_in_progress = false;
+	bool		wake_share_exclusive = true;
+	proclist_head wakeup;
+	proclist_mutable_iter iter;
+
+	proclist_init(&wakeup);
+
+	/* lock wait list while collecting backends to wake up */
+	LockBufHdr(buf_hdr);
+
+	proclist_foreach_modify(iter, &buf_hdr->lock_waiters, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		/*
+		 * Already woke up a conflicting lock, so skip over this wait list
+		 * entry.
+		 */
+		if (!wake_exclusive && waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+			continue;
+		if (!wake_share_exclusive && waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+			continue;
+
+		proclist_delete(&buf_hdr->lock_waiters, iter.cur, lwWaitLink);
+		proclist_push_tail(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Prevent additional wakeups until retryer gets to run. Backends that
+		 * are just waiting for the lock to become free don't retry
+		 * automatically.
+		 */
+		new_wake_in_progress = true;
+
+		/*
+		 * Signal that the process isn't on the wait list anymore. This allows
+		 * BufferLockDequeueSelf() to remove itself from the waitlist with a
+		 * proclist_delete(), rather than having to check if it has been
+		 * removed from the list.
+		 */
+		Assert(waiter->lwWaiting == LW_WS_WAITING);
+		waiter->lwWaiting = LW_WS_PENDING_WAKEUP;
+
+		/*
+		 * Don't wakeup further waiters after waking a conflicting waiter.
+		 */
+		if (waiter->lwWaitMode == BUFFER_LOCK_SHARE)
+		{
+			/*
+			 * Share locks conflict with exclusive locks.
+			 */
+			wake_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			/*
+			 * Share-exclusive locks conflict with share-exclusive and
+			 * exclusive locks.
+			 */
+			wake_exclusive = false;
+			wake_share_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+		{
+
+			/*
+			 * Exclusive locks conflict with all other locks, there's no point
+			 * in waking up anybody else.
+			 */
+			break;
+		}
+	}
+
+	Assert(proclist_is_empty(&wakeup) || pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS);
+
+	/* unset required flags, and release lock, in one fell swoop */
+	{
+		uint64		old_state;
+		uint64		desired_state;
+
+		old_state = pg_atomic_read_u64(&buf_hdr->state);
+		while (true)
+		{
+			desired_state = old_state;
+
+			/* compute desired flags */
+
+			if (new_wake_in_progress)
+				desired_state |= BM_LOCK_WAKE_IN_PROGRESS;
+			else
+				desired_state &= ~BM_LOCK_WAKE_IN_PROGRESS;
+
+			if (proclist_is_empty(&buf_hdr->lock_waiters))
+				desired_state &= ~BM_LOCK_HAS_WAITERS;
+
+			desired_state &= ~BM_LOCKED;	/* release lock */
+
+			if (pg_atomic_compare_exchange_u64(&buf_hdr->state, &old_state,
+											   desired_state))
+				break;
+		}
+	}
+
+	/* Awaken any waiters I removed from the queue. */
+	proclist_foreach_modify(iter, &wakeup, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		proclist_delete(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Guarantee that lwWaiting being unset only becomes visible once the
+		 * unlink from the link has completed. Otherwise the target backend
+		 * could be woken up for other reason and enqueue for a new lock - if
+		 * that happens before the list unlink happens, the list would end up
+		 * being corrupted.
+		 *
+		 * The barrier pairs with the LockBufHdr() when enqueuing for another
+		 * lock.
+		 */
+		pg_write_barrier();
+		waiter->lwWaiting = LW_WS_NOT_WAITING;
+		PGSemaphoreUnlock(waiter->sem);
+	}
+}
+
+/*
+ * Compute subtraction from buffer state for a release of a held lock in
+ * `mode`.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places, each needing to compute what to
+ * subtract from the lock state.
+ */
+static inline uint64
+BufferLockReleaseSub(BufferLockMode mode)
+{
+
+	/*
+	 * Turns out that a switch() leads gcc to generate sufficiently worse code
+	 * for this to show up in profiles...
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE)
+		return BM_LOCK_VAL_EXCLUSIVE;
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		return BM_LOCK_VAL_SHARE_EXCLUSIVE;
+	else
+	{
+		Assert(mode == BUFFER_LOCK_SHARE);
+		return BM_LOCK_VAL_SHARED;
+	}
+
+	return 0;					/* keep compiler quiet */
+}
+
+/*
+ * Handle work that needs to be done after releasing a lock that was held in
+ * `mode`, where `lockstate` is the result of the atomic operation modifying
+ * the state variable.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places.
+ */
+static void
+BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate)
+{
+	bool		check_waiters = false;
+	bool		wake_exclusive = false;
+
+	/* nobody else can have that kind of lock */
+	Assert(!(lockstate & BM_LOCK_VAL_EXCLUSIVE));
+
+	/*
+	 * If we're still waiting for backends to get scheduled, don't wake them
+	 * up again. Otherwise check if we need to look through the waitqueue to
+	 * wake other backends.
+	 */
+	if ((lockstate & BM_LOCK_HAS_WAITERS) &&
+		!(lockstate & BM_LOCK_WAKE_IN_PROGRESS))
+	{
+		if ((lockstate & BM_LOCK_MASK) == 0)
+		{
+			/*
+			 * We released a lock and the lock was, in that moment, free. We
+			 * therefore can wake waiters for any kind of lock.
+			 */
+			check_waiters = true;
+			wake_exclusive = true;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			/*
+			 * We released the lock, but another backend still holds a lock.
+			 * We can't have released an exclusive lock, as there couldn't
+			 * have been other lock holders. If we released a share lock, no
+			 * waiters need to be woken up, as there must be other share
+			 * lockers. However, if we held a share-exclusive lock, another
+			 * backend now could acquire a share-exclusive lock.
+			 */
+			check_waiters = true;
+			wake_exclusive = false;
+		}
+	}
+
+	/*
+	 * As waking up waiters requires the spinlock to be acquired, only do so
+	 * if necessary.
+	 */
+	if (check_waiters)
+		BufferLockWakeup(buf_hdr, wake_exclusive);
+}
+
+/*
+ * BufferLockHeldByMeInMode - test whether my process holds the content lock
+ * in the specified mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode == mode;
+
+}
+
+/*
+ * BufferLockHeldByMe - test whether my process holds the content lock in any
+ * mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMe(BufferDesc *buf_hdr)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode != BUFFER_LOCK_UNLOCK;
+}
+
+/*
+ * Release the content lock for the buffer.
+ */
+void
+UnlockBuffer(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+
+	Assert(BufferIsPinned(buffer));
+	if (BufferIsLocal(buffer))
+		return;					/* local buffers need no lock */
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+	BufferLockUnlock(buffer, buf_hdr);
+}
+
+/*
+ * Acquire the content_lock for the buffer.
  */
 void
-LockBuffer(Buffer buffer, BufferLockMode mode)
+LockBufferInternal(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *buf;
+	BufferDesc *buf_hdr;
+
+	/*
+	 * We can't wait if we haven't got a PGPROC.  This should only occur
+	 * during bootstrap or shared memory initialization.  Put an Assert here
+	 * to catch unsafe coding practices.
+	 */
+	Assert(!(MyProc == NULL && IsUnderPostmaster));
+
+	/* handled in LockBuffer() wrapper */
+	Assert(mode != BUFFER_LOCK_UNLOCK);
 
 	Assert(BufferIsPinned(buffer));
 	if (BufferIsLocal(buffer))
 		return;					/* local buffers need no lock */
 
-	buf = GetBufferDescriptor(buffer - 1);
+	buf_hdr = GetBufferDescriptor(buffer - 1);
 
-	if (mode == BUFFER_LOCK_UNLOCK)
-		LWLockRelease(BufferDescriptorGetContentLock(buf));
-	else if (mode == BUFFER_LOCK_SHARE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	/*
+	 * Test the most frequent lock modes first. While a switch (mode) would be
+	 * nice, at least gcc generates considerably worse code for it.
+	 *
+	 * Call BufferLockAcquire() with a constant argument for mode, to generate
+	 * more efficient code for the different lock modes.
+	 */
+	if (mode == BUFFER_LOCK_SHARE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE);
 	else if (mode == BUFFER_LOCK_EXCLUSIVE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_EXCLUSIVE);
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_EXCLUSIVE);
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	else
 		elog(ERROR, "unrecognized buffer lock mode: %d", mode);
 }
@@ -5732,8 +6476,7 @@ ConditionalLockBuffer(Buffer buffer)
 
 	buf = GetBufferDescriptor(buffer - 1);
 
-	return LWLockConditionalAcquire(BufferDescriptorGetContentLock(buf),
-									LW_EXCLUSIVE);
+	return BufferLockConditional(buffer, buf, BUFFER_LOCK_EXCLUSIVE);
 }
 
 /*
@@ -6247,8 +6990,8 @@ TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 /*
  * AbortBufferIO: Clean up active buffer I/O after an error.
  *
- *	All LWLocks we might have held have been released,
- *	but we haven't yet released buffer pins, so the buffer is still pinned.
+ *	All LWLocks & content locks we might have held have been released, but we
+ *	haven't yet released buffer pins, so the buffer is still pinned.
  *
  *	If I/O was in progress, we always set BM_IO_ERROR, even though it's
  *	possible the error condition wasn't related to the I/O.
@@ -6688,7 +7431,28 @@ ResOwnerReleaseBufferPin(Datum res)
 	if (BufferIsLocal(buffer))
 		UnpinLocalBufferNoOwner(buffer);
 	else
+	{
+		PrivateRefCountEntry *ref;
+
+		ref = GetPrivateRefCountEntry(buffer, false);
+
+		/* not having a private refcount would imply resowner corruption */
+		Assert(ref != NULL);
+
+		/*
+		 * If the buffer was locked at the time of the resowner release,
+		 * release the lock now. This should only happen after errors.
+		 */
+		if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
+		{
+			BufferDesc *buf = GetBufferDescriptor(buffer - 1);
+
+			HOLD_INTERRUPTS();	/* match the upcoming RESUME_INTERRUPTS */
+			BufferLockUnlock(buffer, buf);
+		}
+
 		UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
+	}
 }
 
 static char *
@@ -6924,10 +7688,10 @@ MarkDirtyUnpinnedBufferInternal(Buffer buf, BufferDesc *desc,
 	/* If it was not already dirty, mark it as dirty. */
 	if (!(buf_state & BM_DIRTY))
 	{
-		LWLockAcquire(BufferDescriptorGetContentLock(desc), LW_EXCLUSIVE);
+		BufferLockAcquire(buf, desc, BUFFER_LOCK_EXCLUSIVE);
 		MarkBufferDirty(buf);
 		result = true;
-		LWLockRelease(BufferDescriptorGetContentLock(desc));
+		BufferLockUnlock(buf, desc);
 	}
 	else
 		*buffer_already_dirty = true;
@@ -7178,16 +7942,12 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		 */
 		if (is_write && !is_temp)
 		{
-			LWLock	   *content_lock;
-
-			content_lock = BufferDescriptorGetContentLock(buf_hdr);
-
-			Assert(LWLockHeldByMe(content_lock));
+			Assert(BufferLockHeldByMe(buf_hdr));
 
 			/*
 			 * Lock is now owned by AIO subsystem.
 			 */
-			LWLockDisown(content_lock);
+			BufferLockDisown(buffer, buf_hdr);
 		}
 
 		/*
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 3299de23bb3..ced6a510291 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -287,6 +287,9 @@ ABI_compatibility:
 Section: ClassName - WaitEventBuffer
 
 BUFFER_CLEANUP	"Waiting to acquire an exclusive pin on a buffer. Buffer pin waits can be protracted if another process holds an open cursor that last read data from the buffer in question."
+BUFFER_SHARED	"Waiting to acquire a shared lock on a buffer."
+BUFFER_SHARE_EXCLUSIVE	"Waiting to acquire a share exclusive lock on a buffer."
+BUFFER_EXCLUSIVE	"Waiting to acquire a exclusive lock on a buffer."
 
 ABI_compatibility:
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v10-0005-Require-share-exclusive-lock-to-set-hint-bits-an.patch (39.6K, ../../jtg5cu4n6h5lib3kzx66ju4yhh6kmviaud7oq6dtut6c4q4rdi@xwsfoagt3c2b/6-v10-0005-Require-share-exclusive-lock-to-set-hint-bits-an.patch)
  download | inline diff:
From 351fd22b76b09384e868101eeffd05ea9e1f4511 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Fri, 12 Dec 2025 15:31:01 -0500
Subject: [PATCH v10 5/8] Require share-exclusive lock to set hint bits and to
 flush

At the moment hint bits can be set with just a share lock on a page (and,
until 45f658dacb9, in one case even without any lock). Because of this we need
to copy pages while writing them out, as otherwise the checksum could be
corrupted.

The need to copy the page is problematic to implement AIO writes:

1) Instead of just needing a single buffer for a copied page we need one for
   each page that's potentially undergoing I/O
2) To be able to use the "worker" AIO implementation the copied page needs to
   reside in shared memory

It also causes problems for using unbuffered/direct-IO, independent of AIO:
Some filesystems, raid implementations, ... do not tolerate the data being
written out to change during the write. E.g. they may compute internal
checksums that can be invalidated by concurrent modifications, leading e.g. to
filesystem errors (as the case with btrfs).

It also just is plain odd to allow modifications of buffers that are just
share locked.

To address these issue, this commit changes the rules so that modifications to
pages are not allowed anymore while holding a share lock. Instead the new
share-exclusive lock (introduced in FIXME XXXX TODO) allows at most one
backend to modify a buffer while other backends have the same page share
locked. An existing share-lock can be upgraded to a share-exclusive lock, if
there are no conflicting locks. For that
BufferBeginSetHintBits()/BufferFinishSetHintBits() and BufferSetHintBits16()
have been introduced.

To prevent hint bits from being set while the buffer is being written out,
writing out buffers now requires a share-exclusive lock.

The use of share-exclusive to gate setting hint bits means that from now on
only one backend can set hint bits at a time. To allow multiple backends to
set hint bits would require more complicated locking, for setting hint bits
we'd need to store the count of backends currently setting hint bits and we
would need another lock-level for I/O conflicting with the lock-level to set
hint bits. Given that the share-exclusive lock for setting hint bits is only
held for a short time, that backends would often just set the same hint bits
and that the cost of occasionally not setting hint bits in hotly accessed
pages is fairly low, this seems like an acceptable tradeoff.

The biggest change to adapt to this is in heapam. To avoid performance
regressions for sequential scans that need to set a lot of hint bits, we need
to amortize the cost of BufferBeginSetHintBits() for cases where hint bits are
set at a high frequency, HeapTupleSatisfiesMVCCBatch() uses the new
SetHintBitsExt() which defers BufferFinishSetHintBits() until all hint bits on
a page have been set.  Conversely, to avoid regressions in cases where we
can't set hint bits in bulk (because we're looking only at individual tuples),
use BufferSetHintBits16() when setting hint bits without batching.

Several other places also need to be adapted, but those changes are
comparatively simpler.

After this we do not need to copy buffers to write them out anymore. That
change is done separately however.

TODO:
- Update commit reference above
- reflow parts of storage/buffer/README that I didn't reindent to make the
  diff more readable

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf%40gcnactj4z56m
---
 src/include/storage/bufmgr.h                |   4 +
 src/backend/access/gist/gistget.c           |  19 +-
 src/backend/access/hash/hashutil.c          |  10 +-
 src/backend/access/heap/heapam_visibility.c | 130 ++++++--
 src/backend/access/nbtree/nbtinsert.c       |  28 +-
 src/backend/access/nbtree/nbtutils.c        |  16 +-
 src/backend/storage/buffer/README           |  46 ++-
 src/backend/storage/buffer/bufmgr.c         | 329 ++++++++++++++++----
 src/backend/storage/freespace/freespace.c   |  14 +-
 src/backend/storage/freespace/fsmpage.c     |  11 +-
 src/tools/pgindent/typedefs.list            |   1 +
 11 files changed, 474 insertions(+), 134 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index a40adf6b2a8..4017896f951 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -314,6 +314,10 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
+extern bool BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer);
+extern bool BufferBeginSetHintBits(Buffer buffer);
+extern void BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std);
+
 extern void UnlockBuffers(void);
 extern void UnlockBuffer(Buffer buffer);
 extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
diff --git a/src/backend/access/gist/gistget.c b/src/backend/access/gist/gistget.c
index ca0a397b7c3..0bbd365d672 100644
--- a/src/backend/access/gist/gistget.c
+++ b/src/backend/access/gist/gistget.c
@@ -64,11 +64,7 @@ gistkillitems(IndexScanDesc scan)
 	 * safe.
 	 */
 	if (BufferGetLSNAtomic(buffer) != so->curPageLSN)
-	{
-		UnlockReleaseBuffer(buffer);
-		so->numKilled = 0;		/* reset counter */
-		return;
-	}
+		goto unlock;
 
 	Assert(GistPageIsLeaf(page));
 
@@ -78,6 +74,16 @@ gistkillitems(IndexScanDesc scan)
 	 */
 	for (i = 0; i < so->numKilled; i++)
 	{
+		if (!killedsomething)
+		{
+			/*
+			 * Use hint bit infrastructure to be allowed to modify the page
+			 * without holding an exclusive lock.
+			 */
+			if (!BufferBeginSetHintBits(buffer))
+				goto unlock;
+		}
+
 		offnum = so->killedItems[i];
 		iid = PageGetItemId(page, offnum);
 		ItemIdMarkDead(iid);
@@ -87,9 +93,10 @@ gistkillitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		GistMarkPageHasGarbage(page);
-		MarkBufferDirtyHint(buffer, true);
+		BufferFinishSetHintBits(buffer, true, true);
 	}
 
+unlock:
 	UnlockReleaseBuffer(buffer);
 
 	/*
diff --git a/src/backend/access/hash/hashutil.c b/src/backend/access/hash/hashutil.c
index cf7f0b90176..b917c97321a 100644
--- a/src/backend/access/hash/hashutil.c
+++ b/src/backend/access/hash/hashutil.c
@@ -593,6 +593,13 @@ _hash_kill_items(IndexScanDesc scan)
 
 			if (ItemPointerEquals(&ituple->t_tid, &currItem->heapTid))
 			{
+				/*
+				 * Use hint bit infrastructure to be allowed to modify the
+				 * page without holding an exclusive lock.
+				 */
+				if (!BufferBeginSetHintBits(so->currPos.buf))
+					goto unlock_page;
+
 				/* found the item */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -610,9 +617,10 @@ _hash_kill_items(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->hasho_flag |= LH_PAGE_HAS_DEAD_TUPLES;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(so->currPos.buf, true, true);
 	}
 
+unlock_page:
 	if (so->hashso_bucket_buf == so->currPos.buf ||
 		havePin)
 		LockBuffer(so->currPos.buf, BUFFER_LOCK_UNLOCK);
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 75ae268d753..fc64f4343ce 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -80,10 +80,38 @@
 
 
 /*
- * SetHintBits()
+ * To be allowed to set hint bits, SetHintBits() needs to call
+ * BufferBeginSetHintBits(). However, that's not free, and some callsites call
+ * SetHintBits() on many tuples in a row. For those it makes sense to amortize
+ * the cost of BufferBeginSetHintBits(). Additionally it's desirable to defer
+ * the cost of BufferBeginSetHintBits() until a hint bit needs to actually be
+ * set. This enum serves as the necessary state space passed to
+ * SetHintBitsExt().
+ */
+typedef enum SetHintBitsState
+{
+	/* not yet checked if hint bits may be set */
+	SHB_INITIAL,
+	/* failed to get permission to set hint bits, don't check again */
+	SHB_DISABLED,
+	/* allowed to set hint bits */
+	SHB_ENABLED,
+} SetHintBitsState;
+
+/*
+ * SetHintBitsExt()
  *
  * Set commit/abort hint bits on a tuple, if appropriate at this time.
  *
+ * To be allowed to set a hint bit on a tuple, the page must not be undergoing
+ * IO at this time (otherwise we e.g. could corrupt PG's page checksum or even
+ * the filesystem's, as is known to happen with btrfs).
+ *
+ * The right to set a hint bit can be acquired on a page level with
+ * BufferBeginSetHintBits(). Only a single backend gets the right to set hint
+ * bits at a time.  Alternatively, if called with a NULL SetHintBitsState*,
+ * hint bits are set with BufferSetHintBits16().
+ *
  * It is only safe to set a transaction-committed hint bit if we know the
  * transaction's commit record is guaranteed to be flushed to disk before the
  * buffer, or if the table is temporary or unlogged and will be obliterated by
@@ -111,24 +139,67 @@
  * InvalidTransactionId if no check is needed.
  */
 static inline void
-SetHintBits(HeapTupleHeader tuple, Buffer buffer,
-			uint16 infomask, TransactionId xid)
+SetHintBitsExt(HeapTupleHeader tuple, Buffer buffer,
+			   uint16 infomask, TransactionId xid, SetHintBitsState *state)
 {
+	/*
+	 * In batched mode, if we previously did not get permission to set hint
+	 * bits, don't try again - in all likelihood IO is still going on.
+	 */
+	if (state && *state == SHB_DISABLED)
+		return;
+
 	if (TransactionIdIsValid(xid))
 	{
-		/* NB: xid must be known committed here! */
-		XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+		if (BufferIsPermanent(buffer))
+		{
+			/* NB: xid must be known committed here! */
+			XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+
+			if (XLogNeedsFlush(commitLSN) &&
+				BufferGetLSNAtomic(buffer) < commitLSN)
+			{
+				/* not flushed and no LSN interlock, so don't set hint */
+				return;
+			}
+		}
+	}
+
+	/*
+	 * If we're not operating in batch mode, use BufferSetHintBits16() to mark
+	 * the page dirty, that's cheaper than
+	 * BufferBeginSetHintBits()/BufferFinishSetHintBits(). That's important
+	 * for cases where we set a lot of hint bits on a page individually.
+	 */
+	if (!state)
+	{
+		BufferSetHintBits16(&tuple->t_infomask,
+							tuple->t_infomask | infomask, buffer);
+		return;
+	}
 
-		if (BufferIsPermanent(buffer) && XLogNeedsFlush(commitLSN) &&
-			BufferGetLSNAtomic(buffer) < commitLSN)
+	if (*state == SHB_INITIAL)
+	{
+		if (!BufferBeginSetHintBits(buffer))
 		{
-			/* not flushed and no LSN interlock, so don't set hint */
+			*state = SHB_DISABLED;
 			return;
 		}
-	}
 
+		*state = SHB_ENABLED;
+	}
 	tuple->t_infomask |= infomask;
-	MarkBufferDirtyHint(buffer, true);
+}
+
+/*
+ * Simple wrapper around SetHintBitExt(), use when operating on a single
+ * tuple.
+ */
+static inline void
+SetHintBits(HeapTupleHeader tuple, Buffer buffer,
+			uint16 infomask, TransactionId xid)
+{
+	SetHintBitsExt(tuple, buffer, infomask, xid, NULL);
 }
 
 /*
@@ -864,9 +935,9 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
  * inserting/deleting transaction was still running --- which was more cycles
  * and more contention on ProcArrayLock.
  */
-static bool
+static inline bool
 HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
-					   Buffer buffer)
+					   Buffer buffer, SetHintBitsState *state)
 {
 	HeapTupleHeader tuple = htup->t_data;
 
@@ -921,8 +992,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 			if (!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmax(tuple)))
 			{
 				/* deleting subtransaction must have aborted */
-				SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-							InvalidTransactionId);
+				SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+							   InvalidTransactionId, state);
 				return true;
 			}
 
@@ -934,13 +1005,13 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		else if (XidInMVCCSnapshot(HeapTupleHeaderGetRawXmin(tuple), snapshot))
 			return false;
 		else if (TransactionIdDidCommit(HeapTupleHeaderGetRawXmin(tuple)))
-			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-						HeapTupleHeaderGetRawXmin(tuple));
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_COMMITTED,
+						   HeapTupleHeaderGetRawXmin(tuple), state);
 		else
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_INVALID,
+						   InvalidTransactionId, state);
 			return false;
 		}
 	}
@@ -1003,14 +1074,14 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		if (!TransactionIdDidCommit(HeapTupleHeaderGetRawXmax(tuple)))
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+						   InvalidTransactionId, state);
 			return true;
 		}
 
 		/* xmax transaction committed */
-		SetHintBits(tuple, buffer, HEAP_XMAX_COMMITTED,
-					HeapTupleHeaderGetRawXmax(tuple));
+		SetHintBitsExt(tuple, buffer, HEAP_XMAX_COMMITTED,
+					   HeapTupleHeaderGetRawXmax(tuple), state);
 	}
 	else
 	{
@@ -1607,9 +1678,10 @@ HeapTupleSatisfiesHistoricMVCC(HeapTuple htup, Snapshot snapshot,
  * ->vistuples_dense is set to contain the offsets of visible tuples.
  *
  * The reason this is more efficient than HeapTupleSatisfiesMVCC() is that it
- * avoids a cross-translation-unit function call for each tuple and allows the
- * compiler to optimize across calls to HeapTupleSatisfiesMVCC. In the future
- * it will also allow more efficient setting of hint bits.
+ * avoids a cross-translation-unit function call for each tuple, allows the
+ * compiler to optimize across calls to HeapTupleSatisfiesMVCC and allows
+ * setting hint bits more efficiently (see the one BufferFinishSetHintBits()
+ * call below).
  *
  * Returns the number of visible tuples.
  */
@@ -1620,6 +1692,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 							OffsetNumber *vistuples_dense)
 {
 	int			nvis = 0;
+	SetHintBitsState state = SHB_INITIAL;
 
 	Assert(IsMVCCSnapshot(snapshot));
 
@@ -1628,7 +1701,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		bool		valid;
 		HeapTuple	tup = &batchmvcc->tuples[i];
 
-		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer, &state);
 		batchmvcc->visible[i] = valid;
 
 		if (likely(valid))
@@ -1638,6 +1711,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		}
 	}
 
+	if (state == SHB_ENABLED)
+		BufferFinishSetHintBits(buffer, true, true);
+
 	return nvis;
 }
 
@@ -1657,7 +1733,7 @@ HeapTupleSatisfiesVisibility(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 	switch (snapshot->snapshot_type)
 	{
 		case SNAPSHOT_MVCC:
-			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer);
+			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer, NULL);
 		case SNAPSHOT_SELF:
 			return HeapTupleSatisfiesSelf(htup, snapshot, buffer);
 		case SNAPSHOT_ANY:
diff --git a/src/backend/access/nbtree/nbtinsert.c b/src/backend/access/nbtree/nbtinsert.c
index 63eda08f7a2..da43af3ec96 100644
--- a/src/backend/access/nbtree/nbtinsert.c
+++ b/src/backend/access/nbtree/nbtinsert.c
@@ -681,20 +681,28 @@ _bt_check_unique(Relation rel, BTInsertState insertstate, Relation heapRel,
 				{
 					/*
 					 * The conflicting tuple (or all HOT chains pointed to by
-					 * all posting list TIDs) is dead to everyone, so mark the
-					 * index entry killed.
+					 * all posting list TIDs) is dead to everyone, so try to
+					 * mark the index entry killed. It's ok if we're not
+					 * allowed to, this isn't required for correctness.
 					 */
-					ItemIdMarkDead(curitemid);
-					opaque->btpo_flags |= BTP_HAS_GARBAGE;
+					Buffer		buf;
 
-					/*
-					 * Mark buffer with a dirty hint, since state is not
-					 * crucial. Be sure to mark the proper buffer dirty.
-					 */
+					/* Be sure to operate on the proper buffer */
 					if (nbuf != InvalidBuffer)
-						MarkBufferDirtyHint(nbuf, true);
+						buf = nbuf;
 					else
-						MarkBufferDirtyHint(insertstate->buf, true);
+						buf = insertstate->buf;
+
+					/*
+					 * Can't use BufferSetHintBits16() here as we update two
+					 * different locations.
+					 */
+					if (BufferBeginSetHintBits(buf))
+					{
+						ItemIdMarkDead(curitemid);
+						opaque->btpo_flags |= BTP_HAS_GARBAGE;
+						BufferFinishSetHintBits(buf, true, true);
+					}
 				}
 
 				/*
diff --git a/src/backend/access/nbtree/nbtutils.c b/src/backend/access/nbtree/nbtutils.c
index 5c50f0dd1bd..a76d90f2d8e 100644
--- a/src/backend/access/nbtree/nbtutils.c
+++ b/src/backend/access/nbtree/nbtutils.c
@@ -357,10 +357,19 @@ _bt_killitems(IndexScanDesc scan)
 			 * it's possible that multiple processes attempt to do this
 			 * simultaneously, leading to multiple full-page images being sent
 			 * to WAL (if wal_log_hints or data checksums are enabled), which
-			 * is undesirable.
+			 * is undesirable.  We need to use the hint bit infrastructure to
+			 * update the page while just holding a share lock.
 			 */
 			if (killtuple && !ItemIdIsDead(iid))
 			{
+				/*
+				 * If we're not able to set hint bits, there's no point
+				 * continuing.
+				 */
+				if (!killedsomething &&
+					!BufferBeginSetHintBits(buf))
+					goto unlock_page;
+
 				/* found the item/all posting list items */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -371,8 +380,6 @@ _bt_killitems(IndexScanDesc scan)
 	}
 
 	/*
-	 * Since this can be redone later if needed, mark as dirty hint.
-	 *
 	 * Whenever we mark anything LP_DEAD, we also set the page's
 	 * BTP_HAS_GARBAGE flag, which is likewise just a hint.  (Note that we
 	 * only rely on the page-level flag in !heapkeyspace indexes.)
@@ -380,9 +387,10 @@ _bt_killitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->btpo_flags |= BTP_HAS_GARBAGE;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(buf, true, true);
 	}
 
+unlock_page:
 	if (!so->dropPin)
 		_bt_unlockbuf(rel, buf);
 	else
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index 119f31b5d65..8d52533455e 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -25,21 +25,26 @@ that might need to do such a wait is instead handled by waiting to obtain
 the relation-level lock, which is why you'd better hold one first.)  Pins
 may not be held across transaction boundaries, however.
 
-Buffer content locks: there are two kinds of buffer lock, shared and exclusive,
-which act just as you'd expect: multiple backends can hold shared locks on
-the same buffer, but an exclusive lock prevents anyone else from holding
-either shared or exclusive lock.  (These can alternatively be called READ
-and WRITE locks.)  These locks are intended to be short-term: they should not
-be held for long.  Buffer locks are acquired and released by LockBuffer().
-It will *not* work for a single backend to try to acquire multiple locks on
-the same buffer.  One must pin a buffer before trying to lock it.
+Buffer content locks: there are three kinds of buffer lock, shared,
+share-exclusive and exclusive:
+a) multiple backends can hold shared locks on the same buffer
+   (alternatively called a READ lock)
+b) one backend can hold a share-exclusive lock on a buffer while multiple
+   backends can hold a share lock
+c) an exclusive lock prevents anyone else from holding shared, share-exclusive
+   or exclusive lock.
+   (alternatively called a WRITE lock)
+
+These locks are intended to be short-term: they should not be held for long.
+Buffer locks are acquired and released by LockBuffer().  It will *not* work
+for a single backend to try to acquire multiple locks on the same buffer.  One
+must pin a buffer before trying to lock it.
 
 Buffer access rules:
 
-1. To scan a page for tuples, one must hold a pin and either shared or
-exclusive content lock.  To examine the commit status (XIDs and status bits)
-of a tuple in a shared buffer, one must likewise hold a pin and either shared
-or exclusive lock.
+1. To scan a page for tuples, one must hold a pin and at least a share lock.
+To examine the commit status (XIDs and status bits) of a tuple in a shared
+buffer, one must likewise hold a pin and at least a share lock.
 
 2. Once one has determined that a tuple is interesting (visible to the
 current transaction) one may drop the content lock, yet continue to access
@@ -55,9 +60,15 @@ one must hold a pin and an exclusive content lock on the containing buffer.
 This ensures that no one else might see a partially-updated state of the
 tuple while they are doing visibility checks.
 
-4. It is considered OK to update tuple commit status bits (ie, OR the
-values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
-HEAP_XMAX_INVALID into t_infomask) while holding only a shared lock and
+4. Non-critical information on a page ("hint bits") may be modified while
+holding only a share-exclusive lock and pin on the page. To do so in cases
+where only a share lock is already held, use BufferBeginSetHintBits() &
+BufferFinishSetHintBits() (if multiple hint bits are to be set) or
+BufferSetHintBits16() (if a single hint bit is set).
+
+E.g. for heapam, a share-exclusive lock allows to update tuple commit status
+bits (ie, OR the values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
+HEAP_XMAX_INVALID into t_infomask) while holding only a share-exclusive lock and
 pin on a buffer.  This is OK because another backend looking at the tuple
 at about the same time would OR the same bits into the field, so there
 is little or no risk of conflicting update; what's more, if there did
@@ -80,7 +91,6 @@ buffer (increment the refcount) while one is performing the cleanup, but
 it won't be able to actually examine the page until it acquires shared
 or exclusive content lock.
 
-
 Obtaining the lock needed under rule #5 is done by the bufmgr routines
 LockBufferForCleanup() or ConditionalLockBufferForCleanup().  They first get
 an exclusive lock and then check to see if the shared pin count is currently
@@ -96,6 +106,10 @@ VACUUM's use, since we don't allow multiple VACUUMs concurrently on a single
 relation anyway.  Anyone wishing to obtain a cleanup lock outside of recovery
 or a VACUUM must use the conditional variant of the function.
 
+6. To write out a buffer, a share-exclusive lock needs to be held. This
+prevents the buffer from being modified while written out, which could corrupt
+checksums and cause issues on the OS or device level when direct-IO is used.
+
 
 Buffer Manager's Internal Locking
 ---------------------------------
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 0d5da094748..98f473580a4 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2480,9 +2480,8 @@ again:
 	/*
 	 * If the buffer was dirty, try to write it out.  There is a race
 	 * condition here, in that someone might dirty it after we released the
-	 * buffer header lock above, or even while we are writing it out (since
-	 * our share-lock won't prevent hint-bit updates).  We will recheck the
-	 * dirty bit after re-locking the buffer header.
+	 * buffer header lock above.  We will recheck the dirty bit after
+	 * re-locking the buffer header.
 	 */
 	if (buf_state & BM_DIRTY)
 	{
@@ -2490,20 +2489,20 @@ again:
 		Assert(buf_state & BM_VALID);
 
 		/*
-		 * We need a share-lock on the buffer contents to write it out (else
-		 * we might write invalid data, eg because someone else is compacting
-		 * the page contents while we write).  We must use a conditional lock
-		 * acquisition here to avoid deadlock.  Even though the buffer was not
-		 * pinned (and therefore surely not locked) when StrategyGetBuffer
-		 * returned it, someone else could have pinned and exclusive-locked it
-		 * by the time we get here. If we try to get the lock unconditionally,
-		 * we'd block waiting for them; if they later block waiting for us,
-		 * deadlock ensues. (This has been observed to happen when two
-		 * backends are both trying to split btree index pages, and the second
-		 * one just happens to be trying to split the page the first one got
-		 * from StrategyGetBuffer.)
+		 * We need a share-exclusive lock on the buffer contents to write it
+		 * out (else we might write invalid data, eg because someone else is
+		 * compacting the page contents while we write).  We must use a
+		 * conditional lock acquisition here to avoid deadlock.  Even though
+		 * the buffer was not pinned (and therefore surely not locked) when
+		 * StrategyGetBuffer returned it, someone else could have pinned and
+		 * (share-)exclusive-locked it by the time we get here. If we try to
+		 * get the lock unconditionally, we'd block waiting for them; if they
+		 * later block waiting for us, deadlock ensues. (This has been
+		 * observed to happen when two backends are both trying to split btree
+		 * index pages, and the second one just happens to be trying to split
+		 * the page the first one got from StrategyGetBuffer.)
 		 */
-		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -4072,8 +4071,8 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	}
 
 	/*
-	 * Pin it, share-lock it, write it.  (FlushBuffer will do nothing if the
-	 * buffer is clean by the time we've locked it.)
+	 * Pin it, share-exclusive-lock it, write it.  (FlushBuffer will do
+	 * nothing if the buffer is clean by the time we've locked it.)
 	 */
 	PinBuffer_Locked(bufHdr);
 
@@ -4403,11 +4402,8 @@ BufferGetTag(Buffer buffer, RelFileLocator *rlocator, ForkNumber *forknum,
  * However, we will need to force the changes to disk via fsync before
  * we can checkpoint WAL.
  *
- * The caller must hold a pin on the buffer and have share-locked the
- * buffer contents.  (Note: a share-lock does not prevent updates of
- * hint bits in the buffer, so the page could change while the write
- * is in progress, but we assume that that will not invalidate the data
- * written.)
+ * The caller must hold a pin on the buffer and have
+ * (share-)exclusively-locked the buffer contents.
  *
  * If the caller has an smgr reference for the buffer's relation, pass it
  * as the second parameter.  If not, pass NULL.
@@ -4423,6 +4419,9 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	char	   *bufToWrite;
 	uint64		buf_state;
 
+	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(buf, BUFFER_LOCK_SHARE_EXCLUSIVE));
+
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
 	 * someone else flushed the buffer before we could, so we need not do
@@ -4555,7 +4554,7 @@ FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(buf);
 
-	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 	BufferLockUnlock(buffer, buf);
 }
@@ -5474,8 +5473,8 @@ FlushDatabaseBuffers(Oid dbid)
 }
 
 /*
- * Flush a previously, shared or exclusively, locked and pinned buffer to the
- * OS.
+ * Flush a previously, share-exclusively or exclusively, locked and pinned
+ * buffer to the OS.
  */
 void
 FlushOneBuffer(Buffer buffer)
@@ -5548,39 +5547,23 @@ IncrBufferRefCount(Buffer buffer)
 }
 
 /*
- * MarkBufferDirtyHint
+ * Shared-buffer only helper for MarkBufferDirtyHint() and
+ * BufferSetHintBits16().
  *
- *	Mark a buffer dirty for non-critical changes.
- *
- * This is essentially the same as MarkBufferDirty, except:
- *
- * 1. The caller does not write WAL; so if checksums are enabled, we may need
- *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
- * 2. The caller might have only share-lock instead of exclusive-lock on the
- *	  buffer's content lock.
- * 3. This function does not guarantee that the buffer is always marked dirty
- *	  (due to a race condition), so it cannot be used for important changes.
+ * This is separated out because it turns out that the repeated checks for
+ * local buffers, repeated GetBufferDescriptor() and repeated reading of the
+ * buffer's state sufficiently hurts the performance of BufferSetHintBits16().
  */
-void
-MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+static inline void
+MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate, bool buffer_std)
 {
-	BufferDesc *bufHdr;
 	Page		page = BufferGetPage(buffer);
 
-	if (!BufferIsValid(buffer))
-		elog(ERROR, "bad buffer ID: %d", buffer);
-
-	if (BufferIsLocal(buffer))
-	{
-		MarkLocalBufferDirty(buffer);
-		return;
-	}
-
-	bufHdr = GetBufferDescriptor(buffer - 1);
-
 	Assert(GetPrivateRefCount(buffer) > 0);
-	/* here, either share or exclusive lock is OK */
-	Assert(BufferIsLockedByMe(buffer));
+
+	/* here, either share-exclusive or exclusive lock is OK */
+	Assert(BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_SHARE_EXCLUSIVE));
 
 	/*
 	 * This routine might get called many times on the same page, if we are
@@ -5593,8 +5576,8 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
-		(BM_DIRTY | BM_JUST_DIRTIED))
+	if (unlikely((lockstate & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+				 (BM_DIRTY | BM_JUST_DIRTIED)))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
@@ -5610,8 +5593,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * We don't check full_page_writes here because that logic is included
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
-		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
+		if (XLogHintBitIsNeeded() && (lockstate & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5663,17 +5645,19 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 			dirtied = true;		/* Means "will be dirtied by this action" */
 
 			/*
-			 * Set the page LSN if we wrote a backup block. We aren't supposed
-			 * to set this when only holding a share lock but as long as we
-			 * serialise it somehow we're OK. We choose to set LSN while
-			 * holding the buffer header lock, which causes any reader of an
-			 * LSN who holds only a share lock to also obtain a buffer header
-			 * lock before using PageGetLSN(), which is enforced in
-			 * BufferGetLSNAtomic().
+			 * Set the page LSN if we wrote a backup block. To allow backends
+			 * that only hold a share lock on the buffer to read the LSN in a
+			 * tear-free manner, we set the page LSN while holding the buffer
+			 * header lock. This allows any reader of an LSN who holds only a
+			 * share lock to also obtain a buffer header lock before using
+			 * PageGetLSN() to read the LSN in a tear free way. This is done
+			 * in BufferGetLSNAtomic().
 			 *
 			 * If checksums are enabled, you might think we should reset the
 			 * checksum here. That will happen when the page is written
 			 * sometime later in this checkpoint cycle.
+			 *
+			 * FIXME: The start of the comment above needs updating.
 			 */
 			if (XLogRecPtrIsValid(lsn))
 				PageSetLSN(page, lsn);
@@ -5695,6 +5679,41 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	}
 }
 
+/*
+ * MarkBufferDirtyHint
+ *
+ *	Mark a buffer dirty for non-critical changes.
+ *
+ * This is essentially the same as MarkBufferDirty, except:
+ *
+ * 1. The caller does not write WAL; so if checksums are enabled, we may need
+ *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
+ * 2. The caller might have only a share-exclusive-lock instead of an
+ *	  exclusive-lock on the buffer's content lock.
+ * 3. This function does not guarantee that the buffer is always marked dirty
+ *	  (due to a race condition), so it cannot be used for important changes.
+ */
+inline void
+MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+{
+	BufferDesc *bufHdr;
+
+	bufHdr = GetBufferDescriptor(buffer - 1);
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		MarkLocalBufferDirty(buffer);
+		return;
+	}
+
+	MarkSharedBufferDirtyHint(buffer, bufHdr,
+							  pg_atomic_read_u64(&bufHdr->state),
+							  buffer_std);
+}
+
 /*
  * Release buffer content locks for shared buffers.
  *
@@ -6791,6 +6810,188 @@ IsBufferCleanupOK(Buffer buffer)
 	return false;
 }
 
+/*
+ * Helper for BufferBeginSetHintBits() and BufferSetHintBits16().
+ *
+ * This checks if the current lock mode already suffices to allow hint bits
+ * being set and, if not, whether the current lock can be upgraded.
+ */
+static inline bool
+SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
+{
+	uint64		old_state;
+	PrivateRefCountEntry *ref;
+	BufferLockMode mode;
+
+	ref = GetPrivateRefCountEntry(buffer, true);
+
+	if (ref == NULL)
+		elog(ERROR, "lock is not held");
+
+	mode = ref->data.lockmode;
+	if (mode == BUFFER_LOCK_UNLOCK)
+		elog(ERROR, "buffer is not locked");
+
+	/*
+	 * Already am holding a sufficient lock level.
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE || mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+	{
+		*lockstate = pg_atomic_read_u64(&buf_hdr->state);
+		return true;
+	}
+
+	/*
+	 * Only holding a share lock right now, try to upgrade to SHARE_EXCLUSIVE.
+	 */
+	Assert(mode == BUFFER_LOCK_SHARE);
+
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+	while (true)
+	{
+		uint64		desired_state;
+
+		desired_state = old_state;
+
+		/*
+		 * Can't upgrade if somebody else holds the lock in exclusive or
+		 * share-exclusive mode.
+		 */
+		if (unlikely((old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) != 0))
+		{
+			return false;
+		}
+
+		/* currently held lock state */
+		desired_state -= BM_LOCK_VAL_SHARED;
+
+		/* new lock level */
+		desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			ref->data.lockmode = BUFFER_LOCK_SHARE_EXCLUSIVE;
+			*lockstate = desired_state;
+
+			return true;
+		}
+	}
+
+}
+
+/*
+ * Try to acquire the right to set hint bits on the buffer.
+ *
+ * To be allowed to set hint bits, this backend needs to hold either a
+ * share-exclusive or an exclusive lock. In case this backend only holds a
+ * share lock, this function will try to upgrade the lock to
+ * share-exclusive. The caller is only allowed to set hint bits if true is
+ * returned.
+ *
+ * Once BufferBeginSetHintBits() has returned true, hint bits may be set
+ * without further calls to BufferBeginSetHintBits(), until the buffer is
+ * unlocked.
+ *
+ *
+ * Requiring a share-exclusive lock to set hint bits prevents setting hint
+ * bits on buffers that are currently being written out, which could corrupt
+ * the checksum on the page. Flushing buffers also requires a share-exclusive
+ * lock.
+ *
+ * Due to a lock >= share-exclusive being required to set hint bits, only one
+ * backend can set hint bits at a time. Allowing multiple backends to hint
+ * bits would require more complicated locking: For setting hint bits we'd
+ * need to store the count of backends currently setting hint bits, for I/O we
+ * would need another lock-level conflicting with the hint-setting
+ * lock-level. Given that the share-exclusive lock for setting hint bits is
+ * only held for a short time, that backends often would just set the same
+ * hint bits and that the cost of occasionally not setting hint bits in hotly
+ * accessed pages is fairly low, this seems like an acceptable tradeoff.
+ */
+bool
+BufferBeginSetHintBits(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+
+	if (BufferIsLocal(buffer))
+	{
+		/*
+		 * TODO: will need to check for write IO once that's done
+		 * asynchronously.
+		 */
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	return SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate);
+}
+
+/*
+ * End a phase of setting hint bits on this buffer, started with
+ * BufferBeginSetHintBits().
+ *
+ * This would strictly speaking not be required (i.e. the caller could do
+ * MarkBufferDirtyHint() if so desired), but allows us to perform some sanity
+ * checks.
+ */
+void
+BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std)
+{
+	if (!BufferIsLocal(buffer))
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+
+	if (mark_dirty)
+		MarkBufferDirtyHint(buffer, buffer_std);
+}
+
+/*
+ * Try to set a single hint bit in a buffer.
+ *
+ * This is a bit faster than BufferBeginSetHintBits() /
+ * BufferFinishSetHintBits() when setting a single hint bit, but slower than
+ * the former when setting several hint bits.
+ */
+bool
+BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+#ifdef USE_ASSERT_CHECKING
+	char	   *page;
+
+	/* verify that the address is on the page */
+	page = BufferGetPage(buffer);
+	Assert((char *) ptr >= page && (char *) ptr < (page + BLCKSZ));
+#endif
+
+	if (BufferIsLocal(buffer))
+	{
+		*ptr = val;
+
+		MarkLocalBufferDirty(buffer);
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	if (SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate))
+	{
+		*ptr = val;
+
+		MarkSharedBufferDirtyHint(buffer, buf_hdr, lockstate, true);
+
+		return true;
+	}
+
+	return false;
+}
+
 
 /*
  *	Functions for buffer I/O handling
diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index ad337c00871..b9a8f368a63 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -904,13 +904,17 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 	max_avail = fsm_get_max_avail(page);
 
 	/*
-	 * Reset the next slot pointer. This encourages the use of low-numbered
-	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation. We don't bother with marking the page dirty if it wasn't
-	 * already, since this is just a hint.
+	 * Try to reset the next slot pointer. This encourages the use of
+	 * low-numbered pages, increasing the chances that a later vacuum can
+	 * truncate the relation. We don't bother with marking the page dirty if
+	 * it wasn't already, since this is just a hint.
 	 */
 	LockBuffer(buf, BUFFER_LOCK_SHARE);
-	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	if (BufferBeginSetHintBits(buf))
+	{
+		((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+		BufferFinishSetHintBits(buf, false, false);
+	}
 	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 	ReleaseBuffer(buf);
diff --git a/src/backend/storage/freespace/fsmpage.c b/src/backend/storage/freespace/fsmpage.c
index 33ee825529c..e46bf2631fc 100644
--- a/src/backend/storage/freespace/fsmpage.c
+++ b/src/backend/storage/freespace/fsmpage.c
@@ -298,9 +298,18 @@ restart:
 	 * lock and get a garbled next pointer every now and then, than take the
 	 * concurrency hit of an exclusive lock.
 	 *
+	 * Without an exclusive lock, we need to use the hint bit infrastructure
+	 * to be allowed to modify the page.
+	 *
 	 * Wrap-around is handled at the beginning of this function.
 	 */
-	fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+	if (exclusive_lock_held || BufferBeginSetHintBits(buf))
+	{
+		fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+
+		if (!exclusive_lock_held)
+			BufferFinishSetHintBits(buf, false, true);
+	}
 
 	return slot;
 }
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 14dec2d49c1..efea48fcef7 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2750,6 +2750,7 @@ SetConstraintStateData
 SetConstraintTriggerData
 SetExprState
 SetFunctionReturnMode
+SetHintBitsState
 SetOp
 SetOpCmd
 SetOpPath
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v10-0006-WIP-Make-UnlockReleaseBuffer-more-efficient.patch (3.5K, ../../jtg5cu4n6h5lib3kzx66ju4yhh6kmviaud7oq6dtut6c4q4rdi@xwsfoagt3c2b/7-v10-0006-WIP-Make-UnlockReleaseBuffer-more-efficient.patch)
  download | inline diff:
From f8bb42235fde437d49423d26387f16d67c4ed27c Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 19 Nov 2025 15:32:20 -0500
Subject: [PATCH v10 6/8] WIP: Make UnlockReleaseBuffer() more efficient

Now that the buffer content lock is implemented as part of BufferDesc.state,
releasing the lock and unpinning the buffer can be implemented as a single
atomic operation.

Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/nbtree/nbtpage.c | 22 +++++++++++-
 src/backend/storage/buffer/bufmgr.c | 52 ++++++++++++++++++++++++++++-
 2 files changed, 72 insertions(+), 2 deletions(-)

diff --git a/src/backend/access/nbtree/nbtpage.c b/src/backend/access/nbtree/nbtpage.c
index 4125c185e8b..f3e3f67e1fd 100644
--- a/src/backend/access/nbtree/nbtpage.c
+++ b/src/backend/access/nbtree/nbtpage.c
@@ -1007,11 +1007,18 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 
 	Assert(BlockNumberIsValid(blkno));
 	if (BufferIsValid(obuf))
+	{
+		_bt_relbuf(rel, obuf);
+#if 0
+		Assert(BufferGetBlockNumber(obuf) != blkno);
 		_bt_unlockbuf(rel, obuf);
-	buf = ReleaseAndReadBuffer(obuf, rel, blkno);
+#endif
+	}
+	buf = ReadBuffer(rel, blkno);
 	_bt_lockbuf(rel, buf, access);
 
 	_bt_checkpage(rel, buf);
+
 	return buf;
 }
 
@@ -1023,8 +1030,21 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 void
 _bt_relbuf(Relation rel, Buffer buf)
 {
+#if 0
 	_bt_unlockbuf(rel, buf);
 	ReleaseBuffer(buf);
+#else
+	/*
+	 * Buffer is pinned and locked, which means that it is expected to be
+	 * defined and addressable.  Check that proactively.
+	 */
+	VALGRIND_CHECK_MEM_IS_DEFINED(BufferGetPage(buf), BLCKSZ);
+
+	UnlockReleaseBuffer(buf);
+
+	if (!RelationUsesLocalBuffers(rel))
+		VALGRIND_MAKE_MEM_NOACCESS(BufferGetPage(buf), BLCKSZ);
+#endif
 }
 
 /*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 98f473580a4..9574baa36cb 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5511,13 +5511,63 @@ ReleaseBuffer(Buffer buffer)
 /*
  * UnlockReleaseBuffer -- release the content lock and pin on a buffer
  *
- * This is just a shorthand for a common combination.
+ * This is just a, more efficient, shorthand for a common combination.
  */
 void
 UnlockReleaseBuffer(Buffer buffer)
 {
+#if 1
+	int			mode;
+	BufferDesc *buf;
+	PrivateRefCountEntry *ref;
+	uint64		sub;
+	uint64		lockstate;
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		UnpinLocalBuffer(buffer);
+		return;
+	}
+
+	ResourceOwnerForgetBuffer(CurrentResourceOwner, buffer);
+
+	buf = GetBufferDescriptor(buffer - 1);
+
+	mode = BufferLockDisownInternal(buffer, buf);
+
+	/* compute state modification for lock release */
+	sub = BufferLockReleaseSub(mode);
+
+	/* compute state modification for pin release */
+	ref = GetPrivateRefCountEntry(buffer, false);
+	Assert(ref != NULL);
+	Assert(ref->data.refcount > 0);
+	ref->data.refcount--;
+
+	if (ref->data.refcount == 0)
+	{
+		sub |= BUF_REFCOUNT_ONE;
+		ForgetPrivateRefCountEntry(ref);
+	}
+
+	/* perform the lock and pin release in one atomic op */
+	lockstate = pg_atomic_sub_fetch_u64(&buf->state, sub);
+
+	/* wake up waiters etc */
+	BufferLockProcessRelease(buf, mode, lockstate);
+
+	if (lockstate & BM_PIN_COUNT_WAITER)
+		WakePinCountWaiter(buf);
+
+	RESUME_INTERRUPTS();
+
+#else
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 	ReleaseBuffer(buffer);
+#endif
 }
 
 /*
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v10-0007-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch (11.6K, ../../jtg5cu4n6h5lib3kzx66ju4yhh6kmviaud7oq6dtut6c4q4rdi@xwsfoagt3c2b/8-v10-0007-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch)
  download | inline diff:
From 029ae568865d36c6267f5aa963b6f0817c154aba Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 17 Oct 2024 14:14:35 -0400
Subject: [PATCH v10 7/8] WIP: bufmgr: Don't copy pages while writing out

After the series of preceding commits introducing and using
BufferBeginSetHintBits()/BufferSetHintBits16() hint bits are not set
anymore while IO is going on. Therefore we do not need to copy pages while
they are being written out anymore.

TODO: Update comments

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/bufpage.h           |  3 +-
 src/backend/access/hash/hashpage.c      |  2 +-
 src/backend/access/transam/xloginsert.c | 43 ++++++----------------
 src/backend/storage/buffer/bufmgr.c     | 21 +++++------
 src/backend/storage/buffer/localbuf.c   |  2 +-
 src/backend/storage/page/bufpage.c      | 48 ++++---------------------
 src/backend/storage/smgr/bulk_write.c   |  2 +-
 src/test/modules/test_aio/test_aio.c    |  2 +-
 8 files changed, 33 insertions(+), 90 deletions(-)

diff --git a/src/include/storage/bufpage.h b/src/include/storage/bufpage.h
index ae3725b3b81..31ec9a8a047 100644
--- a/src/include/storage/bufpage.h
+++ b/src/include/storage/bufpage.h
@@ -504,7 +504,6 @@ extern void PageIndexMultiDelete(Page page, OffsetNumber *itemnos, int nitems);
 extern void PageIndexTupleDeleteNoCompact(Page page, OffsetNumber offnum);
 extern bool PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 									const void *newtup, Size newsize);
-extern char *PageSetChecksumCopy(Page page, BlockNumber blkno);
-extern void PageSetChecksumInplace(Page page, BlockNumber blkno);
+extern void PageSetChecksum(Page page, BlockNumber blkno);
 
 #endif							/* BUFPAGE_H */
diff --git a/src/backend/access/hash/hashpage.c b/src/backend/access/hash/hashpage.c
index 8e220a3ae16..52c20208c66 100644
--- a/src/backend/access/hash/hashpage.c
+++ b/src/backend/access/hash/hashpage.c
@@ -1029,7 +1029,7 @@ _hash_alloc_buckets(Relation rel, BlockNumber firstblock, uint32 nblocks)
 					zerobuf.data,
 					true);
 
-	PageSetChecksumInplace(page, lastblock);
+	PageSetChecksum(page, lastblock);
 	smgrextend(RelationGetSmgr(rel), MAIN_FORKNUM, lastblock, zerobuf.data,
 			   false);
 
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index 92c48e768c3..4bab484fd5d 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -261,8 +261,11 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	 */
 #ifdef USE_ASSERT_CHECKING
 	if (!(flags & REGBUF_NO_CHANGE))
-		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) &&
-			   BufferIsDirty(buffer));
+	{
+		Assert(BufferIsDirty(buffer));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE));
+	}
 #endif
 
 	if (block_id >= max_registered_block_id)
@@ -1066,7 +1069,7 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * Write a backup block if needed when we are setting a hint. Note that
  * this may be called for a variety of page types, not just heaps.
  *
- * Callable while holding just share lock on the buffer content.
+ * Callable while holding just share-exclusive lock on the buffer content.
  *
  * We can't use the plain backup block mechanism since that relies on the
  * Buffer being exclusively locked. Since some modifications (setting LSN, hint
@@ -1074,6 +1077,8 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * failures. So instead we copy the page and insert the copied data as normal
  * record data.
  *
+ * FIXME: outdated
+ *
  * We only need to do something if page has not yet been full page written in
  * this checkpoint round. The LSN of the inserted wal record is returned if we
  * had to write, InvalidXLogRecPtr otherwise.
@@ -1102,46 +1107,20 @@ XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
 
 	/*
 	 * We assume page LSN is first data on *every* page that can be passed to
-	 * XLogInsert, whether it has the standard page layout or not. Since we're
-	 * only holding a share-lock on the page, we must take the buffer header
-	 * lock when we look at the LSN.
+	 * XLogInsert, whether it has the standard page layout or not.
 	 */
 	lsn = BufferGetLSNAtomic(buffer);
 
 	if (lsn <= RedoRecPtr)
 	{
-		int			flags = 0;
-		PGAlignedBlock copied_buffer;
-		char	   *origdata = (char *) BufferGetBlock(buffer);
-		RelFileLocator rlocator;
-		ForkNumber	forkno;
-		BlockNumber blkno;
-
-		/*
-		 * Copy buffer so we don't have to worry about concurrent hint bit or
-		 * lsn updates. We assume pd_lower/upper cannot be changed without an
-		 * exclusive lock, so the contents bkp are not racy.
-		 */
-		if (buffer_std)
-		{
-			/* Assume we can omit data between pd_lower and pd_upper */
-			Page		page = BufferGetPage(buffer);
-			uint16		lower = ((PageHeader) page)->pd_lower;
-			uint16		upper = ((PageHeader) page)->pd_upper;
-
-			memcpy(copied_buffer.data, origdata, lower);
-			memcpy(copied_buffer.data + upper, origdata + upper, BLCKSZ - upper);
-		}
-		else
-			memcpy(copied_buffer.data, origdata, BLCKSZ);
+		int			flags = REGBUF_NO_CHANGE;
 
 		XLogBeginInsert();
 
 		if (buffer_std)
 			flags |= REGBUF_STANDARD;
 
-		BufferGetTag(buffer, &rlocator, &forkno, &blkno);
-		XLogRegisterBlock(0, &rlocator, forkno, blkno, copied_buffer.data, flags);
+		XLogRegisterBuffer(0, buffer, flags);
 
 		recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI_FOR_HINT);
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 9574baa36cb..e114e64fdd9 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -4416,7 +4416,6 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	ErrorContextCallback errcallback;
 	instr_time	io_start;
 	Block		bufBlock;
-	char	   *bufToWrite;
 	uint64		buf_state;
 
 	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
@@ -4487,12 +4486,8 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	 */
 	bufBlock = BufHdrGetBlock(buf);
 
-	/*
-	 * Update page checksum if desired.  Since we have only shared lock on the
-	 * buffer, other processes might be updating hint bits in it, so we must
-	 * copy the page to private storage if we do checksumming.
-	 */
-	bufToWrite = PageSetChecksumCopy((Page) bufBlock, buf->tag.blockNum);
+	/* Update page checksum if desired. */
+	PageSetChecksum((Page) bufBlock, buf->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
@@ -4502,7 +4497,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	smgrwrite(reln,
 			  BufTagGetForkNum(&buf->tag),
 			  buf->tag.blockNum,
-			  bufToWrite,
+			  bufBlock,
 			  false);
 
 	/*
@@ -4626,8 +4621,8 @@ BufferIsPermanent(Buffer buffer)
 /*
  * BufferGetLSNAtomic
  *		Retrieves the LSN of the buffer atomically using a buffer header lock.
- *		This is necessary for some callers who may not have an exclusive lock
- *		on the buffer.
+ *		This is necessary for some callers who may not have a (share-)exclusive
+ *		lock on the buffer.
  */
 XLogRecPtr
 BufferGetLSNAtomic(Buffer buffer)
@@ -5679,6 +5674,12 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate, b
 			 * It's possible we may enter here without an xid, so it is
 			 * essential that CreateCheckPoint waits for virtual transactions
 			 * rather than full transactionids.
+			 *
+			 * FIXME: I think we now should simply mark the page dirty before
+			 * WAL logging the hint bit - afaikt it then should work just like
+			 * any other buffer write (due to SyncBuffers()/SyncOneBuffer()
+			 * seeing the dirty bit and trying to lock the page
+			 * share-exclusive, and thus having to wait).
 			 */
 			Assert((MyProc->delayChkptFlags & DELAY_CHKPT_START) == 0);
 			MyProc->delayChkptFlags |= DELAY_CHKPT_START;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 04a540379a2..55e17e03acb 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -199,7 +199,7 @@ FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln)
 		reln = smgropen(BufTagGetRelFileLocator(&bufHdr->tag),
 						MyProcNumber);
 
-	PageSetChecksumInplace(localpage, bufHdr->tag.blockNum);
+	PageSetChecksum(localpage, bufHdr->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index de85911e3ac..2072bb1c72c 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -1494,51 +1494,15 @@ PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 /*
  * Set checksum for a page in shared buffers.
  *
- * If checksums are disabled, or if the page is not initialized, just return
- * the input.  Otherwise, we must make a copy of the page before calculating
- * the checksum, to prevent concurrent modifications (e.g. setting hint bits)
- * from making the final checksum invalid.  It doesn't matter if we include or
- * exclude hints during the copy, as long as we write a valid page and
- * associated checksum.
+ * If checksums are disabled, or if the page is not initialized, just
+ * return. Otherwise compute and set the checksum.
  *
- * Returns a pointer to the block-sized data that needs to be written. Uses
- * statically-allocated memory, so the caller must immediately write the
- * returned page and not refer to it again.
- */
-char *
-PageSetChecksumCopy(Page page, BlockNumber blkno)
-{
-	static char *pageCopy = NULL;
-
-	/* If we don't need a checksum, just return the passed-in data */
-	if (PageIsNew(page) || !DataChecksumsEnabled())
-		return page;
-
-	/*
-	 * We allocate the copy space once and use it over on each subsequent
-	 * call.  The point of palloc'ing here, rather than having a static char
-	 * array, is first to ensure adequate alignment for the checksumming code
-	 * and second to avoid wasting space in processes that never call this.
-	 */
-	if (pageCopy == NULL)
-		pageCopy = MemoryContextAllocAligned(TopMemoryContext,
-											 BLCKSZ,
-											 PG_IO_ALIGN_SIZE,
-											 0);
-
-	memcpy(pageCopy, page, BLCKSZ);
-	((PageHeader) pageCopy)->pd_checksum = pg_checksum_page(pageCopy, blkno);
-	return pageCopy;
-}
-
-/*
- * Set checksum for a page in private memory.
- *
- * This must only be used when we know that no other process can be modifying
- * the page buffer.
+ * In the past this needed to be done on a copy of the page, due to the
+ * possibility of e.g. hint bits being set concurrently. However, this is not
+ * necessary anymore as hint bits won't be set while IO is going on.
  */
 void
-PageSetChecksumInplace(Page page, BlockNumber blkno)
+PageSetChecksum(Page page, BlockNumber blkno)
 {
 	/* If we don't need a checksum, just return */
 	if (PageIsNew(page) || !DataChecksumsEnabled())
diff --git a/src/backend/storage/smgr/bulk_write.c b/src/backend/storage/smgr/bulk_write.c
index 36b28824ec8..f3c24082a69 100644
--- a/src/backend/storage/smgr/bulk_write.c
+++ b/src/backend/storage/smgr/bulk_write.c
@@ -279,7 +279,7 @@ smgr_bulk_flush(BulkWriteState *bulkstate)
 		BlockNumber blkno = pending_writes[i].blkno;
 		Page		page = pending_writes[i].buf->data;
 
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 
 		if (blkno >= bulkstate->relsize)
 		{
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index b1aa8af9ec0..2ae4a559fab 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -288,7 +288,7 @@ modify_rel_block(PG_FUNCTION_ARGS)
 	}
 	else
 	{
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 	}
 
 	smgrwrite(RelationGetSmgr(rel),
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v10-0008-WIP-bufmgr-Rename-ResOwnerReleaseBufferPin.patch (3.7K, ../../jtg5cu4n6h5lib3kzx66ju4yhh6kmviaud7oq6dtut6c4q4rdi@xwsfoagt3c2b/9-v10-0008-WIP-bufmgr-Rename-ResOwnerReleaseBufferPin.patch)
  download | inline diff:
From ddc2c9e973090b4989f68a9e2e792088be31a519 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 12 Jan 2026 19:28:51 -0500
Subject: [PATCH v10 8/8] WIP: bufmgr: Rename ResOwnerReleaseBufferPin

This is separate as I'm not yet convinced of the new naming. The comment
probably makes sense regardless.

This is a name suggested a while ago by Melanie.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/buf_internals.h |  6 +++---
 src/backend/storage/buffer/bufmgr.c | 22 ++++++++++++++--------
 2 files changed, 17 insertions(+), 11 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 12086cf6dc7..b6714318154 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -520,18 +520,18 @@ extern PGDLLIMPORT CkptSortItem *CkptBufferIds;
 
 /* ResourceOwner callbacks to hold buffer I/Os and pins */
 extern PGDLLIMPORT const ResourceOwnerDesc buffer_io_resowner_desc;
-extern PGDLLIMPORT const ResourceOwnerDesc buffer_pin_resowner_desc;
+extern PGDLLIMPORT const ResourceOwnerDesc buffer_resowner_desc;
 
 /* Convenience wrappers over ResourceOwnerRemember/Forget */
 static inline void
 ResourceOwnerRememberBuffer(ResourceOwner owner, Buffer buffer)
 {
-	ResourceOwnerRemember(owner, Int32GetDatum(buffer), &buffer_pin_resowner_desc);
+	ResourceOwnerRemember(owner, Int32GetDatum(buffer), &buffer_resowner_desc);
 }
 static inline void
 ResourceOwnerForgetBuffer(ResourceOwner owner, Buffer buffer)
 {
-	ResourceOwnerForget(owner, Int32GetDatum(buffer), &buffer_pin_resowner_desc);
+	ResourceOwnerForget(owner, Int32GetDatum(buffer), &buffer_resowner_desc);
 }
 static inline void
 ResourceOwnerRememberBufferIO(ResourceOwner owner, Buffer buffer)
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index e114e64fdd9..2f39454fd7f 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -263,8 +263,8 @@ static void ForgetPrivateRefCountEntry(PrivateRefCountEntry *ref);
 /* ResourceOwner callbacks to hold in-progress I/Os and buffer pins */
 static void ResOwnerReleaseBufferIO(Datum res);
 static char *ResOwnerPrintBufferIO(Datum res);
-static void ResOwnerReleaseBufferPin(Datum res);
-static char *ResOwnerPrintBufferPin(Datum res);
+static void ResOwnerReleaseBuffer(Datum res);
+static char *ResOwnerPrintBuffer(Datum res);
 
 const ResourceOwnerDesc buffer_io_resowner_desc =
 {
@@ -275,13 +275,13 @@ const ResourceOwnerDesc buffer_io_resowner_desc =
 	.DebugPrint = ResOwnerPrintBufferIO
 };
 
-const ResourceOwnerDesc buffer_pin_resowner_desc =
+const ResourceOwnerDesc buffer_resowner_desc =
 {
-	.name = "buffer pin",
+	.name = "buffer",
 	.release_phase = RESOURCE_RELEASE_BEFORE_LOCKS,
 	.release_priority = RELEASE_PRIO_BUFFER_PINS,
-	.ReleaseResource = ResOwnerReleaseBufferPin,
-	.DebugPrint = ResOwnerPrintBufferPin
+	.ReleaseResource = ResOwnerReleaseBuffer,
+	.DebugPrint = ResOwnerPrintBuffer
 };
 
 /*
@@ -7671,8 +7671,14 @@ ResOwnerPrintBufferIO(Datum res)
 	return psprintf("lost track of buffer IO on buffer %d", buffer);
 }
 
+/*
+ * Release buffer as part of resource owner cleanup. This will only be called
+ * if the buffer is pinned. If this backend held the content lock at the time
+ * of the error we also need to release that (note that it is not possible to
+ * hold a content lock without a pin).
+ */
 static void
-ResOwnerReleaseBufferPin(Datum res)
+ResOwnerReleaseBuffer(Datum res)
 {
 	Buffer		buffer = DatumGetInt32(res);
 
@@ -7708,7 +7714,7 @@ ResOwnerReleaseBufferPin(Datum res)
 }
 
 static char *
-ResOwnerPrintBufferPin(Datum res)
+ResOwnerPrintBuffer(Datum res)
 {
 	return DebugPrintBufferRefcount(DatumGetInt32(res));
 }
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-13 15:05                                           ` Melanie Plageman <melanieplageman@gmail.com>
  2026-01-14 00:49                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  4 siblings, 1 reply; 120+ messages in thread

From: Melanie Plageman @ 2026-01-13 15:05 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Mon, Jan 12, 2026 at 7:33 PM Andres Freund <andres@anarazel.de> wrote:
>
> - added a commit that renames ResOwnerReleaseBufferPin to
>   ResOwnerReleaseBuffer (et al), as it now also releases content locks if held
>
>   I kept this separate as I'm not yet sure about the new name, partially due
>   to there also being a "buffer io" resowner.  I tried "buffer ownership" for
>   the resowner that tracks pins and locks, but that was long and not clearly
>   better.

I didn't look at the patch but I strongly agree that
ResOwnerReleaseBufferPin() should not also release locks, so it should
have a new name. Ironic that ResOwnerReleaseBufferIO() releases pins
and not locks.

What about ResOwnerReleaseBufferClaim() or
ResOwnerReleaseBufferAccess() or ResOwnerReleaseBufferHold()?

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 15:05                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
@ 2026-01-14 00:49                                             ` Andres Freund <andres@anarazel.de>
  2026-01-14 14:17                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-01-14 00:49 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-13 10:05:02 -0500, Melanie Plageman wrote:
> On Mon, Jan 12, 2026 at 7:33 PM Andres Freund <andres@anarazel.de> wrote:
> >
> > - added a commit that renames ResOwnerReleaseBufferPin to
> >   ResOwnerReleaseBuffer (et al), as it now also releases content locks if held
> >
> >   I kept this separate as I'm not yet sure about the new name, partially due
> >   to there also being a "buffer io" resowner.  I tried "buffer ownership" for
> >   the resowner that tracks pins and locks, but that was long and not clearly
> >   better.
> 
> I didn't look at the patch but I strongly agree that
> ResOwnerReleaseBufferPin() should not also release locks, so it should
> have a new name.

OK.

> Ironic that ResOwnerReleaseBufferIO() releases pins and not locks.

Not sure I follow? I don't think it releases pins? And why should it release
locks?


> What about ResOwnerReleaseBufferClaim() or
> ResOwnerReleaseBufferAccess() or ResOwnerReleaseBufferHold()?

I'm inclined to go with just ResOwnerReleaseBuffer() at the moment. Buffer IO
kind of is a subsidiary thing, and it requires holding a pin as well, so it
doesn't feel too wrong.

I also wonder if we could merge BufferIO into the private refcount
infrastructure, similar to how the patches store the lockmode in the private
refcount.  The separate resowner acquisition does show up in profiles when
reading from the kernel page cache, so that'd be a nice (but small)
improvement.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 15:05                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2026-01-14 00:49                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-14 14:17                                               ` Melanie Plageman <melanieplageman@gmail.com>
  2026-01-14 15:20                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Melanie Plageman @ 2026-01-14 14:17 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Tue, Jan 13, 2026 at 7:49 PM Andres Freund <andres@anarazel.de> wrote:
>
> On 2026-01-13 10:05:02 -0500, Melanie Plageman wrote:
>
> > Ironic that ResOwnerReleaseBufferIO() releases pins and not locks.
>
> Not sure I follow? I don't think it releases pins? And why should it release
> locks?

Ah, I must not have actually read it or read the wrong thing.

> I also wonder if we could merge BufferIO into the private refcount
> infrastructure, similar to how the patches store the lockmode in the private
> refcount.  The separate resowner acquisition does show up in profiles when
> reading from the kernel page cache, so that'd be a nice (but small)
> improvement.

When you say "BufferIO", do you mean io_wref in the BufferDesc?

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 15:05                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2026-01-14 00:49                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 14:17                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
@ 2026-01-14 15:20                                                 ` Andres Freund <andres@anarazel.de>
  0 siblings, 0 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-14 15:20 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-14 09:17:22 -0500, Melanie Plageman wrote:
> On Tue, Jan 13, 2026 at 7:49 PM Andres Freund <andres@anarazel.de> wrote:
> > I also wonder if we could merge BufferIO into the private refcount
> > infrastructure, similar to how the patches store the lockmode in the private
> > refcount.  The separate resowner acquisition does show up in profiles when
> > reading from the kernel page cache, so that'd be a nice (but small)
> > improvement.
> 
> When you say "BufferIO", do you mean io_wref in the BufferDesc?

I was trying to refer to ResourceOwnerRememberBufferIO(),
ResourceOwnerForgetBufferIO(), ResOwnerReleaseBufferIO(), etc. That's
basically used to unset BM_IO_IN_PROGRESS when an error occurs while trying to
perform IO.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-14 02:26                                           ` Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:23                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  4 siblings, 1 reply; 120+ messages in thread

From: Chao Li @ 2026-01-14 02:26 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>



> On Jan 13, 2026, at 08:33, Andres Freund <andres@anarazel.de> wrote:
> 
> Hi,
> 
> On 2026-01-12 12:45:03 -0500, Andres Freund wrote:
>> I'm doing another pass through 0003 and will push that if I don't find
>> anything significant.
> 
> Done, after adjust two comments in minor ways.
> 
> 
>> Also working on doing comment polishing of the later patches, found a few
>> things, but not quite enough to be worth reposting yet.
> 
> Here are the remaining commits, with a bit of polish:
> 
> - fixed references to old names in some places (lwlocks, release_ok)
> 
> - Aded an assert that we don't already hold a lock in BufferLockConditional()
> 
> - typo and grammar fixes
> 
> - updated the commit message of the LW_FLAG_RELEASE_OK, as "requested" by
>  Melanie. I hope this explains the situation better.
> 
> - added a commit that renames ResOwnerReleaseBufferPin to
>  ResOwnerReleaseBuffer (et al), as it now also releases content locks if held
> 
>  I kept this separate as I'm not yet sure about the new name, partially due
>  to there also being a "buffer io" resowner.  I tried "buffer ownership" for
>  the resowner that tracks pins and locks, but that was long and not clearly
>  better.
> 
> Greetings,
> 
> Andres Freund
> <v10-0001-lwlock-Invert-meaning-of-LW_FLAG_RELEASE_OK.patch><v10-0002-bufmgr-Make-definitions-related-to-buffer-descri.patch><v10-0003-bufmgr-Change-BufferDesc.state-to-be-a-64-bit-at.patch><v10-0004-bufmgr-Implement-buffer-content-locks-independen.patch><v10-0005-Require-share-exclusive-lock-to-set-hint-bits-an.patch><v10-0006-WIP-Make-UnlockReleaseBuffer-more-efficient.patch><v10-0007-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch><v10-0008-WIP-bufmgr-Rename-ResOwnerReleaseBufferPin.patch>

Hi Andres,

So far I’ve only reviewed 0001 and 0002. I’m not very familiar with this area, so the review has been a bit slow.

Overall, 0001 looks good to me. It renames LW_FLAG_RELEASE_OK to LW_FLAG_WAKE_IN_PROGRESS and inverts the meaning, which makes sense. I only have a small nit on naming: the local variable “new_release_in_progress". I see that it’s inherited from the old name and was updated from “_ok" to “_in_progress", but now that the flag itself is renamed, would it make sense to rename the variable as well? Something like “wake_in_progress" or “new_wake_in_progress" might better reflect the new flag name.

In 0002, a bunch of new macros are introduced. My initial impression wasn’t great, mostly due to the amount of line wrapping. Looking a bit closer, I also noticed some duplication, for example, "BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS" appears more than once; and a small inconsistency between BUF_STATE_GET_REFCOUNT and BUF_STATE_GET_USAGECOUNT (even though the former doesn’t actually need a shift).

I tried a small refactor of the macro definitions in the attached diff to see if things could be made a bit more regular. It introduces a helper macro MASK() and a BUF_REFCOUNT_SHIFT constant, and removes a bit of duplication. If you like it, feel free to take it; otherwise, please just ignore it. Note that, the diff is based on 0002.

(I actually hesitated to attach a diff, because if you’ve already created a CF entry, the attached diff could break the CI tests. If that happens, sorry about that.)

Best regards,
--
Chao Li (Evan)
HighGo Software Co., Ltd.
https://www.highgo.com/

Attachments:

  [application/octet-stream] buf_internals_h.diff (2.2K, ../../AC5E365D-7AD9-47AE-B2C6-25756712B188@gmail.com/2-buf_internals_h.diff)
  download | inline diff:
diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 2f607ea2ac5..34e6c6cd54f 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -49,28 +49,26 @@
 StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 				 "parts of buffer state space need to equal 32");
 
-/* refcount related definitions */
-#define BUF_REFCOUNT_ONE 1
-#define BUF_REFCOUNT_MASK \
-	((1U << BUF_REFCOUNT_BITS) - 1)
+#define BUF_REFCOUNT_SHIFT       0
+#define BUF_USAGECOUNT_SHIFT     (BUF_REFCOUNT_SHIFT + BUF_REFCOUNT_BITS)
+#define BUF_FLAG_SHIFT           (BUF_USAGECOUNT_SHIFT + BUF_USAGECOUNT_BITS)
+
+/* mask generator */
+#define MASK(bits) ((1U << (bits)) - 1)
 
+/* refcount related definitions */
+#define BUF_REFCOUNT_ONE         1U
+#define BUF_REFCOUNT_MASK        (MASK(BUF_REFCOUNT_BITS) << BUF_REFCOUNT_SHIFT)
 /* usage count related definitions */
-#define BUF_USAGECOUNT_SHIFT \
-	BUF_REFCOUNT_BITS
-#define BUF_USAGECOUNT_MASK \
-	(((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_USAGECOUNT_SHIFT))
-#define BUF_USAGECOUNT_ONE \
-	(1U << BUF_REFCOUNT_BITS)
+#define BUF_USAGECOUNT_ONE       (1U << BUF_USAGECOUNT_SHIFT)
+#define BUF_USAGECOUNT_MASK      (MASK(BUF_USAGECOUNT_BITS) << BUF_USAGECOUNT_SHIFT)
 
 /* flags related definitions */
-#define BUF_FLAG_SHIFT \
-	(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS)
-#define BUF_FLAG_MASK \
-	(((1U << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
+#define BUF_FLAG_MASK            (MASK(BUF_FLAG_BITS) << BUF_FLAG_SHIFT)
 
 /* Get refcount and usagecount from buffer state */
 #define BUF_STATE_GET_REFCOUNT(state) \
-	((state) & BUF_REFCOUNT_MASK)
+	(((state) & BUF_REFCOUNT_MASK) >> BUF_REFCOUNT_SHIFT)
 #define BUF_STATE_GET_USAGECOUNT(state) \
 	(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
 
@@ -81,8 +79,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  * entry associated with the buffer's tag.
  */
 
-#define BUF_DEFINE_FLAG(flagno)	\
-	(1U << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + (flagno)))
+#define BUF_DEFINE_FLAG(flagno)	(1U << (BUF_FLAG_SHIFT + (flagno)))
 
 /* buffer header is locked */
 #define BM_LOCKED					BUF_DEFINE_FLAG( 0)

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 02:26                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
@ 2026-01-14 16:23                                             ` Andres Freund <andres@anarazel.de>
  0 siblings, 0 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-14 16:23 UTC (permalink / raw)
  To: Chao Li <li.evan.chao@gmail.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-14 10:26:07 +0800, Chao Li wrote:
> So far I’ve only reviewed 0001 and 0002. I’m not very familiar with this area, so the review has been a bit slow.
> 
> Overall, 0001 looks good to me. It renames LW_FLAG_RELEASE_OK to
> LW_FLAG_WAKE_IN_PROGRESS and inverts the meaning, which makes sense. I only
> have a small nit on naming: the local variable “new_release_in_progress". I
> see that it’s inherited from the old name and was updated from “_ok" to
> “_in_progress", but now that the flag itself is renamed, would it make sense
> to rename the variable as well? Something like “wake_in_progress" or
> “new_wake_in_progress" might better reflect the new flag name.

Agreed that is better. Updated that way.



> In 0002, a bunch of new macros are introduced. My initial impression wasn’t
> great, mostly due to the amount of line wrapping.

I think the previous formatting made it hard to actually write useful comments
and caused line-length problems in the subsequent commits. Lines are cheap.


> Looking a bit closer, I also noticed some duplication, for example,
> "BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS" appears more than once

Yea, that's probably better to avoid. I'll add a fix to that in the commit
changing it to 64bits, I think.


> ; and a small inconsistency between BUF_STATE_GET_REFCOUNT and
> BUF_STATE_GET_USAGECOUNT (even though the former doesn’t actually need a
> shift).

I don't see the point, if we later want to move refcounts elsewhere, we can do
it at that time.


> I tried a small refactor of the macro definitions in the attached diff to
> see if things could be made a bit more regular. It introduces a helper macro
> MASK() and a BUF_REFCOUNT_SHIFT constant, and removes a bit of
> duplication. If you like it, feel free to take it; otherwise, please just
> ignore it. Note that, the diff is based on 0002.

I don't think the MASK thing is an improvement.


> (I actually hesitated to attach a diff, because if you’ve already created a
> CF entry, the attached diff could break the CI tests. If that happens, sorry
> about that.)

FWIW, there's a trick to avoid that: Rename your patch to end in .txt or such.


Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-14 03:41                                           ` Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  4 siblings, 1 reply; 120+ messages in thread

From: Chao Li @ 2026-01-14 03:41 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>



> On Jan 13, 2026, at 08:33, Andres Freund <andres@anarazel.de> wrote:
> 
> Hi,
> 
> On 2026-01-12 12:45:03 -0500, Andres Freund wrote:
>> I'm doing another pass through 0003 and will push that if I don't find
>> anything significant.
> 
> Done, after adjust two comments in minor ways.
> 
> 
>> Also working on doing comment polishing of the later patches, found a few
>> things, but not quite enough to be worth reposting yet.
> 
> Here are the remaining commits, with a bit of polish:
> 
> - fixed references to old names in some places (lwlocks, release_ok)
> 
> - Aded an assert that we don't already hold a lock in BufferLockConditional()
> 
> - typo and grammar fixes
> 
> - updated the commit message of the LW_FLAG_RELEASE_OK, as "requested" by
>  Melanie. I hope this explains the situation better.
> 
> - added a commit that renames ResOwnerReleaseBufferPin to
>  ResOwnerReleaseBuffer (et al), as it now also releases content locks if held
> 
>  I kept this separate as I'm not yet sure about the new name, partially due
>  to there also being a "buffer io" resowner.  I tried "buffer ownership" for
>  the resowner that tracks pins and locks, but that was long and not clearly
>  better.
> 
> Greetings,
> 
> Andres Freund
> <v10-0001-lwlock-Invert-meaning-of-LW_FLAG_RELEASE_OK.patch><v10-0002-bufmgr-Make-definitions-related-to-buffer-descri.patch><v10-0003-bufmgr-Change-BufferDesc.state-to-be-a-64-bit-at.patch><v10-0004-bufmgr-Implement-buffer-content-locks-independen.patch><v10-0005-Require-share-exclusive-lock-to-set-hint-bits-an.patch><v10-0006-WIP-Make-UnlockReleaseBuffer-more-efficient.patch><v10-0007-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch><v10-0008-WIP-bufmgr-Rename-ResOwnerReleaseBufferPin.patch>

A couple of comments on v10-0003, I just noticed 0001 and 0002 have been pushed.

Basically, code changes in 0003 is straightforward, just a couple of small comments:

1
```
- * refcounts in buf_internals.h.  This limitation could be lifted by using a
- * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
- * currently realistic configurations. Even if that limitation were removed,
- * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
- * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
- * 4*MaxBackends without any overflow check.  We check that the configured
- * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
+ * refcounts in buf_internals.h.  This limitation could be lifted, but it's
```

Before this patch, there was room for lifting the limitation. With this patch, state is 64bit already, but the significant 32bit will be used for buffer locking as stated in buf_internals.h, in other words, there is no room for lifting the limitation now. If that’s true, then I think we can remove the statements about lifting limitation.

2. By searching for “LockBufHdr”, I found one place missed to update in contrib/pg_prewarm/autoprewarm.c at line 706:
```
	for (num_blocks = 0, i = 0; i < NBuffers; i++)
	{
		uint32		buf_state; <=== line 706, should change to uint64

		CHECK_FOR_INTERRUPTS();

		bufHdr = GetBufferDescriptor(i);

		/* Lock each buffer header before inspecting. */
		buf_state = LockBufHdr(bufHdr);
```

I will continue reviewing 0004 tomorrow.


Best regards,
--
Chao Li (Evan)
HighGo Software Co., Ltd.
https://www.highgo.com/









^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
@ 2026-01-14 16:30                                             ` Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-01-14 16:30 UTC (permalink / raw)
  To: Chao Li <li.evan.chao@gmail.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-14 11:41:19 +0800, Chao Li wrote:
> Basically, code changes in 0003 is straightforward, just a couple of small comments:
> 
> 1
> ```
> - * refcounts in buf_internals.h.  This limitation could be lifted by using a
> - * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
> - * currently realistic configurations. Even if that limitation were removed,
> - * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
> - * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
> - * 4*MaxBackends without any overflow check.  We check that the configured
> - * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
> + * refcounts in buf_internals.h.  This limitation could be lifted, but it's
> ```
> 
> Before this patch, there was room for lifting the limitation. With this
> patch, state is 64bit already, but the significant 32bit will be used for
> buffer locking as stated in buf_internals.h, in other words, there is no
> room for lifting the limitation now. If that’s true, then I think we can
> remove the statements about lifting limitation.

I'm not following - there's plenty space for more bits if we need that:

 * State of the buffer itself (in order):
 * - 18 bits refcount
 * - 4 bits usage count
 * - 12 bits of flags
 * - 18 bits share-lock count
 * - 1 bit share-exclusive locked
 * - 1 bit exclusive locked

That's 54 bits in total. Which part is in the lower and which in the upper
32bit isn't relevant for anything afaict?


> 2. By searching for “LockBufHdr”, I found one place missed to update in contrib/pg_prewarm/autoprewarm.c at line 706:
> ```
> 	for (num_blocks = 0, i = 0; i < NBuffers; i++)
> 	{
> 		uint32		buf_state; <=== line 706, should change to uint64
> 
> 		CHECK_FOR_INTERRUPTS();
> 
> 		bufHdr = GetBufferDescriptor(i);
> 
> 		/* Lock each buffer header before inspecting. */
> 		buf_state = LockBufHdr(bufHdr);
> ```

Good catch!  I didn't find any other similar omissions...


> I will continue reviewing 0004 tomorrow.

Cool.

I'd like to push

  bufmgr: Change BufferDesc.state to be a 64-bit atomic
  bufmgr: Implement buffer content locks independently of lwlocks

pretty soon, so that we then can concentrate on

  Require share-exclusive lock to set hint bits and to flush

and then subsequently on

  WIP: bufmgr: Don't copy pages while writing out

as there are other patches that have this work as a dependency...

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-14 23:20                                               ` Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Chao Li @ 2026-01-14 23:20 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>



> On Jan 15, 2026, at 00:30, Andres Freund <andres@anarazel.de> wrote:
> 
> Hi,
> 
> On 2026-01-14 11:41:19 +0800, Chao Li wrote:
>> Basically, code changes in 0003 is straightforward, just a couple of small comments:
>> 
>> 1
>> ```
>> - * refcounts in buf_internals.h.  This limitation could be lifted by using a
>> - * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
>> - * currently realistic configurations. Even if that limitation were removed,
>> - * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
>> - * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
>> - * 4*MaxBackends without any overflow check.  We check that the configured
>> - * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
>> + * refcounts in buf_internals.h.  This limitation could be lifted, but it's
>> ```
>> 
>> Before this patch, there was room for lifting the limitation. With this
>> patch, state is 64bit already, but the significant 32bit will be used for
>> buffer locking as stated in buf_internals.h, in other words, there is no
>> room for lifting the limitation now. If that’s true, then I think we can
>> remove the statements about lifting limitation.
> 
> I'm not following - there's plenty space for more bits if we need that:
> 
> * State of the buffer itself (in order):
> * - 18 bits refcount
> * - 4 bits usage count
> * - 12 bits of flags
> * - 18 bits share-lock count
> * - 1 bit share-exclusive locked
> * - 1 bit exclusive locked
> 
> That's 54 bits in total. Which part is in the lower and which in the upper
> 32bit isn't relevant for anything afaict?

Because I saw the comment in buf_internals.h:
```
 * NB: A future commit will use a significant portion of the remaining bits to
* implement buffer locking as part of the state variable.
```
That seems to indicate all the significant 32 bits will be used for buffer locking. Also, there is an assert that concretes the impression:
```
StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
       "parts of buffer state space need to equal 32");
```

So, I thought we can explain 18bit refcount is good enough without mentioning “lifting” that potentially adds confusion to readers. But anyway, this is not a strong opinion. I won’t insist on this comment.

> 
> 
>> 2. By searching for “LockBufHdr”, I found one place missed to update in contrib/pg_prewarm/autoprewarm.c at line 706:
>> ```
>> for (num_blocks = 0, i = 0; i < NBuffers; i++)
>> {
>> uint32 buf_state; <=== line 706, should change to uint64
>> 
>> CHECK_FOR_INTERRUPTS();
>> 
>> bufHdr = GetBufferDescriptor(i);
>> 
>> /* Lock each buffer header before inspecting. */
>> buf_state = LockBufHdr(bufHdr);
>> ```
> 
> Good catch!  I didn't find any other similar omissions...

I saw you have added this occurrence to v11.

> 
> 
>> I will continue reviewing 0004 tomorrow.
> 
> Cool.
> 
> I'd like to push
> 
>  bufmgr: Change BufferDesc.state to be a 64-bit atomic
>  bufmgr: Implement buffer content locks independently of lwlocks
> 
> pretty soon, so that we then can concentrate on

Other than the “lifting” comment, v11 LGTM. But that’s not a strong opinion. I explained more above, if you consider that’s not a problem, I am totally fine.

Best regards,
--
Chao Li (Evan)
HighGo Software Co., Ltd.
https://www.highgo.com/









^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
@ 2026-01-14 23:37                                                 ` Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-01-14 23:37 UTC (permalink / raw)
  To: Chao Li <li.evan.chao@gmail.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-15 07:20:27 +0800, Chao Li wrote:
> > On Jan 15, 2026, at 00:30, Andres Freund <andres@anarazel.de> wrote:
> > On 2026-01-14 11:41:19 +0800, Chao Li wrote:
> >> Basically, code changes in 0003 is straightforward, just a couple of small comments:
> >> 
> >> 1
> >> ```
> >> - * refcounts in buf_internals.h.  This limitation could be lifted by using a
> >> - * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
> >> - * currently realistic configurations. Even if that limitation were removed,
> >> - * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
> >> - * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
> >> - * 4*MaxBackends without any overflow check.  We check that the configured
> >> - * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
> >> + * refcounts in buf_internals.h.  This limitation could be lifted, but it's
> >> ```
> >> 
> >> Before this patch, there was room for lifting the limitation. With this
> >> patch, state is 64bit already, but the significant 32bit will be used for
> >> buffer locking as stated in buf_internals.h, in other words, there is no
> >> room for lifting the limitation now. If that’s true, then I think we can
> >> remove the statements about lifting limitation.
> > 
> > I'm not following - there's plenty space for more bits if we need that:
> > 
> > * State of the buffer itself (in order):
> > * - 18 bits refcount
> > * - 4 bits usage count
> > * - 12 bits of flags
> > * - 18 bits share-lock count
> > * - 1 bit share-exclusive locked
> > * - 1 bit exclusive locked
> > 
> > That's 54 bits in total. Which part is in the lower and which in the upper
> > 32bit isn't relevant for anything afaict?
> 
> Because I saw the comment in buf_internals.h:
> ```
>  * NB: A future commit will use a significant portion of the remaining bits to
> * implement buffer locking as part of the state variable.
> ```
> That seems to indicate all the significant 32 bits will be used for buffer locking.

A significant portion != all. As the above excerpt from the comment shows, the
locking uses 20 bits. We could increase max backends by 5 bits without running
out of bits (we'd need space both in the refcount bitspace as well as the
share-lock bitspace).


> Also, there is an assert that concretes the impression:
> ```
> StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
>        "parts of buffer state space need to equal 32");
> ```

You can see that being relaxed in the subsequent commit, when we start to use
more bits.


Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-15 00:04                                                   ` Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Chao Li @ 2026-01-15 00:04 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>



> On Jan 15, 2026, at 07:37, Andres Freund <andres@anarazel.de> wrote:
> 
> Hi,
> 
> On 2026-01-15 07:20:27 +0800, Chao Li wrote:
>>> On Jan 15, 2026, at 00:30, Andres Freund <andres@anarazel.de> wrote:
>>> On 2026-01-14 11:41:19 +0800, Chao Li wrote:
>>>> Basically, code changes in 0003 is straightforward, just a couple of small comments:
>>>> 
>>>> 1
>>>> ```
>>>> - * refcounts in buf_internals.h.  This limitation could be lifted by using a
>>>> - * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
>>>> - * currently realistic configurations. Even if that limitation were removed,
>>>> - * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
>>>> - * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
>>>> - * 4*MaxBackends without any overflow check.  We check that the configured
>>>> - * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
>>>> + * refcounts in buf_internals.h.  This limitation could be lifted, but it's
>>>> ```
>>>> 
>>>> Before this patch, there was room for lifting the limitation. With this
>>>> patch, state is 64bit already, but the significant 32bit will be used for
>>>> buffer locking as stated in buf_internals.h, in other words, there is no
>>>> room for lifting the limitation now. If that’s true, then I think we can
>>>> remove the statements about lifting limitation.
>>> 
>>> I'm not following - there's plenty space for more bits if we need that:
>>> 
>>> * State of the buffer itself (in order):
>>> * - 18 bits refcount
>>> * - 4 bits usage count
>>> * - 12 bits of flags
>>> * - 18 bits share-lock count
>>> * - 1 bit share-exclusive locked
>>> * - 1 bit exclusive locked
>>> 
>>> That's 54 bits in total. Which part is in the lower and which in the upper
>>> 32bit isn't relevant for anything afaict?
>> 
>> Because I saw the comment in buf_internals.h:
>> ```
>> * NB: A future commit will use a significant portion of the remaining bits to
>> * implement buffer locking as part of the state variable.
>> ```
>> That seems to indicate all the significant 32 bits will be used for buffer locking.
> 
> A significant portion != all. As the above excerpt from the comment shows, the
> locking uses 20 bits. We could increase max backends by 5 bits without running
> out of bits (we'd need space both in the refcount bitspace as well as the
> share-lock bitspace).

Make sense. I think I misread the comment.

> 
> 
>> Also, there is an assert that concretes the impression:
>> ```
>> StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
>>       "parts of buffer state space need to equal 32");
>> ```
> 
> You can see that being relaxed in the subsequent commit, when we start to use
> more bits.
> 

Sure. I plan to review 0003-0005 today. I believe I will get better understanding.

So, 0001 and 0002 LGTM now.

Best regards,
--
Chao Li (Evan)
HighGo Software Co., Ltd.
https://www.highgo.com/









^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
@ 2026-01-15 06:22                                                     ` Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Chao Li @ 2026-01-15 06:22 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>



> On Jan 15, 2026, at 08:04, Chao Li <li.evan.chao@gmail.com> wrote:
> 
> 
> 
>> On Jan 15, 2026, at 07:37, Andres Freund <andres@anarazel.de> wrote:
>> 
>> Hi,
>> 
>> On 2026-01-15 07:20:27 +0800, Chao Li wrote:
>>>> On Jan 15, 2026, at 00:30, Andres Freund <andres@anarazel.de> wrote:
>>>> On 2026-01-14 11:41:19 +0800, Chao Li wrote:
>>>>> Basically, code changes in 0003 is straightforward, just a couple of small comments:
>>>>> 
>>>>> 1
>>>>> ```
>>>>> - * refcounts in buf_internals.h.  This limitation could be lifted by using a
>>>>> - * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
>>>>> - * currently realistic configurations. Even if that limitation were removed,
>>>>> - * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
>>>>> - * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
>>>>> - * 4*MaxBackends without any overflow check.  We check that the configured
>>>>> - * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
>>>>> + * refcounts in buf_internals.h.  This limitation could be lifted, but it's
>>>>> ```
>>>>> 
>>>>> Before this patch, there was room for lifting the limitation. With this
>>>>> patch, state is 64bit already, but the significant 32bit will be used for
>>>>> buffer locking as stated in buf_internals.h, in other words, there is no
>>>>> room for lifting the limitation now. If that’s true, then I think we can
>>>>> remove the statements about lifting limitation.
>>>> 
>>>> I'm not following - there's plenty space for more bits if we need that:
>>>> 
>>>> * State of the buffer itself (in order):
>>>> * - 18 bits refcount
>>>> * - 4 bits usage count
>>>> * - 12 bits of flags
>>>> * - 18 bits share-lock count
>>>> * - 1 bit share-exclusive locked
>>>> * - 1 bit exclusive locked
>>>> 
>>>> That's 54 bits in total. Which part is in the lower and which in the upper
>>>> 32bit isn't relevant for anything afaict?
>>> 
>>> Because I saw the comment in buf_internals.h:
>>> ```
>>> * NB: A future commit will use a significant portion of the remaining bits to
>>> * implement buffer locking as part of the state variable.
>>> ```
>>> That seems to indicate all the significant 32 bits will be used for buffer locking.
>> 
>> A significant portion != all. As the above excerpt from the comment shows, the
>> locking uses 20 bits. We could increase max backends by 5 bits without running
>> out of bits (we'd need space both in the refcount bitspace as well as the
>> share-lock bitspace).
> 
> Make sense. I think I misread the comment.
> 
>> 
>> 
>>> Also, there is an assert that concretes the impression:
>>> ```
>>> StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
>>>      "parts of buffer state space need to equal 32");
>>> ```
>> 
>> You can see that being relaxed in the subsequent commit, when we start to use
>> more bits.
>> 
> 
> Sure. I plan to review 0003-0005 today. I believe I will get better understanding.

I have finished reviewing 0003-0005. Basically, 0003 and 0004 just delete some unused functions. I only searched over the source tree to ensure no usages of them. I spent time on 0005. The code logic are solid already, I didn't find any correctness issue and only got some small comments:

1 - 0004 - commit message
```
subsequent commits fixing a typo an a parameter name.
```

Typo: a typo an a parameter name => a typo in a parameter name

2 - 0005 - bufmgr.c
```
-	 * Pin it, share-lock it, write it.  (FlushBuffer will do nothing if the
-	 * buffer is clean by the time we've locked it.)
+	 * Pin it, share-exclusive-lock it, write it.  (FlushBuffer will do
+	 * nothing if the buffer is clean by the time we've locked it.)
 	 */
 	PinBuffer_Locked(bufHdr);
```

I think we don’t need to mention lock type in this comment, because the logic belongs to FlushUnlockedBuffer(). Also, FlushBuffer is misleading here, because we call FlushUnlockedBuffer() here, and FlushUnlockedBuffer() in turn calls FlushBuffer().

So, I think we can simplify the comment as something like “Pin it and flush the buffer"

3 - 0005 - bufmgr.c
```
+inline void
+MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
```

It’s quite uncommon to extern an inline function. I think usually if we want to make an inline function accessible from external, we define it “static inline” in a header file. So, I guess “inline” is a typo here.

4 - 0005 - bufmgr.c
```
/*
* MarkBufferDirtyHint
*
* Mark a buffer dirty for non-critical changes.
*
* This is essentially the same as MarkBufferDirty, except:
*
* 1. The caller does not write WAL; so if checksums are enabled, we may need
* to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
* 2. The caller might have only a share-exclusive-lock instead of an
* exclusive-lock on the buffer's content lock.
```

For point 2, do we need to mention “For shared buffers”. Because this function also handles local buffer that doesn’t require a lock.

5 - 0005 - bufmgr.c
```
+/*
+ * Helper for BufferBeginSetHintBits() and BufferSetHintBits16().

+static inline bool
+SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
```

In the previous function comment:
```
- * MarkBufferDirtyHint
+ * Shared-buffer only helper for MarkBufferDirtyHint() and
+ * BufferSetHintBits16().
```

It mentions “Shared-buffer only helper”. I think SharedBufferBeginSetHintBits is also only for shared buffer, maybe also add “Shared-buffer only” before “helper” in the comment.

6 - 0005 - bufmgr.c
```
+	}
+
+}
```

Nit: In function SharedBufferBeginSetHintBits, the last empty line is not needed.

7 - 0005 - bufmgr.c
```
+ * This checks if the current lock mode already suffices to allow hint bits
+ * being set and, if not, whether the current lock can be upgraded.
+ */
+static inline bool
+SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
```

Nit: "if not, whether the current lock can be upgraded” might be read as “lock can be upgraded, so the caller still need to take some action to upgrade the lock”, but the function has upgraded the lock when returning true. So, how about explicitly state something like: "if not, it attempts to atomically upgrade it to share-exclusive. Returns true if hint bits may be set (with or without an upgrade), false otherwise."

8 - 0005 - fsmpage.c
```
 * needs to be updated. exclusive_lock_held should be set to true if the
* caller is already holding an exclusive lock, to avoid extra work.
```

The function comment of fsm_search_avail() needs to be updated. exclusive_lock_held should be set to true if the caller is already holding a **share-exclusive or** exclusive lock.

Maybe the parameter name “exclusive_lock_held” can be enhanced as well.

9 - 0005 - fsmpage.c
```
+		if (!exclusive_lock_held)
+			BufferFinishSetHintBits(buf, false, true);
```

Nit: just curious why set the third parameter as “true”? When the second (mark_dirty) is false, the third parameter is not used at all.

10 - 0005 - freespace.c
```
-	 * Reset the next slot pointer. This encourages the use of low-numbered
-	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation. We don't bother with marking the page dirty if it wasn't
-	 * already, since this is just a hint.
+	 * Try to reset the next slot pointer. This encourages the use of
+	 * low-numbered pages, increasing the chances that a later vacuum can
+	 * truncate the relation. We don't bother with marking the page dirty if
+	 * it wasn't already, since this is just a hint.
 	 */
 	LockBuffer(buf, BUFFER_LOCK_SHARE);
-	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	if (BufferBeginSetHintBits(buf))
+	{
+		((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+		BufferFinishSetHintBits(buf, false, false);
+	}
```

Before this patch, we unconditionally set fp_next_slot, now the setting might be skipped. You have add “Try to” in the comment that has explained the possibility of skipping setting fp_next_slot. Would it be better to add a brief statement for what will result in when skipping setting fp_next_slot? Something like “Skipping the update only affects reuse, not correctness”.

Lately, I saw you have done this in nbtinsert.c:
```
* mark the index entry killed. It's ok if we're not
* allowed to, this isn't required for correctness.
```
which, I think, confirmed my comment.

11 - 0005 Maybe not a problem. In nbtree.h:
```
/*
* We need to be able to tell the difference between read and write
* requests for pages, in order to do locking correctly.
*/
#define BT_READ BUFFER_LOCK_SHARE
#define BT_WRITE BUFFER_LOCK_EXCLUSIVE
```

Can the new lock type BUFFER_LOCK_SHARE_EXCLUSIVE be used by nbt? Maybe implicitly upgrading from BT_READ to BUFFER_LOCK_SHARE_EXCLUSIVE is good enough?

12 - 0005 - heapam_visibility.c

After this commit, tuple hint bits may remain unset if we can’t obtain share-exclusive permission. That’s fine because hint bits are advisory and optional, but this is a behavior change. Would it make sense to mention this explicitly and briefly in the SetHintBitsExt() comment?

13 - 0005 - nbtutils.c
```
+				/*
+				 * If we're not able to set hint bits, there's no point
+				 * continuing.
+				 */
+				if (!killedsomething &&
+					!BufferBeginSetHintBits(buf))
+					goto unlock_page;
```
I like this code, because it ensures to only call BufferBeginSetHintBits once. But lately, I saw the same logic in gistget.c:
```
+		if (!killedsomething)
+		{
+			/*
+			 * Use hint bit infrastructure to be allowed to modify the page
+			 * without holding an exclusive lock.
+			 */
+			if (!BufferBeginSetHintBits(buffer))
+				goto unlock;
+		}
```
I just feel the gist version is easier to read, so maybe change the nbt one to be the same as the gist one. But I think the nbt one’s comment is better.

Best regards,
--
Chao Li (Evan)
HighGo Software Co., Ltd.
https://www.highgo.com/









^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
@ 2026-01-15 16:43                                                       ` Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-16 02:36                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  0 siblings, 2 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-15 16:43 UTC (permalink / raw)
  To: Chao Li <li.evan.chao@gmail.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-15 14:22:09 +0800, Chao Li wrote:
> 3 - 0005 - bufmgr.c
> ```
> +inline void
> +MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
> ```
> 
> It’s quite uncommon to extern an inline function. I think usually if we want
> to make an inline function accessible from external, we define it “static
> inline” in a header file. So, I guess “inline” is a typo here.

It works just fine to put an inline into the function definition, even if it's
an external function. That hints the compiler to inline it inside the
translation unit. Which is useful here, because it leads to
MarkBufferDirtyHint() being inlined into BufferFinishSetHintBits().



> 4 - 0005 - bufmgr.c
> ```
> /*
> * MarkBufferDirtyHint
> *
> * Mark a buffer dirty for non-critical changes.
> *
> * This is essentially the same as MarkBufferDirty, except:
> *
> * 1. The caller does not write WAL; so if checksums are enabled, we may need
> * to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
> * 2. The caller might have only a share-exclusive-lock instead of an
> * exclusive-lock on the buffer's content lock.
> ```
> 
> For point 2, do we need to mention “For shared buffers”. Because this function also handles local buffer that doesn’t require a lock.

That's an old comment, just slightly rephrased. And I don't think it matters
that temp buffers don't need to be locked.



> 8 - 0005 - fsmpage.c
> ```
>  * needs to be updated. exclusive_lock_held should be set to true if the
> * caller is already holding an exclusive lock, to avoid extra work.
> ```
> 
> The function comment of fsm_search_avail() needs to be
> updated. exclusive_lock_held should be set to true if the caller is already
> holding a **share-exclusive or** exclusive lock.

Why?


> 9 - 0005 - fsmpage.c
> ```
> +		if (!exclusive_lock_held)
> +			BufferFinishSetHintBits(buf, false, true);
> ```
> 
> Nit: just curious why set the third parameter as “true”? When the second (mark_dirty) is false, the third parameter is not used at all.

That's wrong, indeed. Not because of the mark_dirty, but because they aren't
standard pages.


> 10 - 0005 - freespace.c
> ```
> -	 * Reset the next slot pointer. This encourages the use of low-numbered
> -	 * pages, increasing the chances that a later vacuum can truncate the
> -	 * relation. We don't bother with marking the page dirty if it wasn't
> -	 * already, since this is just a hint.
> +	 * Try to reset the next slot pointer. This encourages the use of
> +	 * low-numbered pages, increasing the chances that a later vacuum can
> +	 * truncate the relation. We don't bother with marking the page dirty if
> +	 * it wasn't already, since this is just a hint.
>  	 */
>  	LockBuffer(buf, BUFFER_LOCK_SHARE);
> -	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
> +	if (BufferBeginSetHintBits(buf))
> +	{
> +		((FSMPage) PageGetContents(page))->fp_next_slot = 0;
> +		BufferFinishSetHintBits(buf, false, false);
> +	}
> ```
> 
> Before this patch, we unconditionally set fp_next_slot, now the setting
> might be skipped. You have add “Try to” in the comment that has explained
> the possibility of skipping setting fp_next_slot. Would it be better to add
> a brief statement for what will result in when skipping setting
> fp_next_slot? Something like “Skipping the update only affects reuse, not
> correctness”.

I don't see the point. The whole paragraph is about how this isn't crucial.


> 11 - 0005 Maybe not a problem. In nbtree.h:
> ```
> /*
> * We need to be able to tell the difference between read and write
> * requests for pages, in order to do locking correctly.
> */
> #define BT_READ BUFFER_LOCK_SHARE
> #define BT_WRITE BUFFER_LOCK_EXCLUSIVE
> ```
> 
> Can the new lock type BUFFER_LOCK_SHARE_EXCLUSIVE be used by nbt?

I don't think at the moment. If we introduce a user, we can add the
define. Orthogonally, I think stuff like BT_READ/BT_WRITE should eventually be
removed, the only thing it does is to make it harder to search the code.


> Maybe implicitly upgrading from BT_READ to BUFFER_LOCK_SHARE_EXCLUSIVE is good enough?

I don't see where that would trivially be safe.


> 12 - 0005 - heapam_visibility.c
> 
> After this commit, tuple hint bits may remain unset if we can’t obtain
> share-exclusive permission. That’s fine because hint bits are advisory and
> optional, but this is a behavior change. Would it make sense to mention this
> explicitly and briefly in the SetHintBitsExt() comment?

That's not new though, we already could skip setting hint bits due to async
commits.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-15 23:02                                                         ` Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-15 23:16                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  1 sibling, 2 replies; 120+ messages in thread

From: Tom Lane @ 2026-01-15 23:02 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Various buildfarm animals are complaining about fcb9c977a,
similarly to this from calliphoridae [1]:

[969/2355] ccache gcc -Isrc/backend/postgres_lib.a.p -Isrc/include -I../pgsql/src/include -I/usr/include/libxml2 -fdiagnostics-color=never -D_FILE_OFFSET_BITS=64 -Wall -Winvalid-pch -O2 -g -fno-strict-aliasing -fwrapv -fexcess-precision=standard -D_GNU_SOURCE -Wmissing-prototypes -Wpointer-arith -Werror=vla -Wendif-labels -Wmissing-format-attribute -Wimplicit-fallthrough=3 -Wcast-function-type -Wshadow=compatible-local -Wformat-security -Wdeclaration-after-statement -Wmissing-variable-declarations -Wno-format-truncation -Wno-stringop-truncation -O1 -ggdb -g3 -fno-omit-frame-pointer -Wall -Wextra -Wno-unused-parameter -Wno-sign-compare -Wno-missing-field-initializers -DCOPY_PARSE_PLAN_TREES -DRAW_EXPRESSION_COVERAGE_TEST -fPIC -isystem /usr/include/mit-krb5 -pthread -DBUILDING_DLL -MD -MQ src/backend/postgres_lib.a.p/storage_buffer_bufmgr.c.o -MF src/backend/postgres_lib.a.p/storage_buffer_bufmgr.c.o.d -o src/backend/postgres_lib.a.p/storage_buffer_bufmgr.c.o -c ../pgsql/src/backend/storage/buffer/bufmgr.c
In file included from ../pgsql/src/include/pgstat.h:24,
                 from ../pgsql/src/backend/storage/buffer/bufmgr.c:52:
In function \342\200\230pgstat_report_wait_start\342\200\231,
    inlined from \342\200\230BufferLockAcquire\342\200\231 at ../pgsql/src/backend/storage/buffer/bufmgr.c:5833:3:
../pgsql/src/include/utils/wait_event.h:75:49: warning: \342\200\230wait_event\342\200\231 may be used uninitialized [-Wmaybe-uninitialized]
   75 |         *(volatile uint32 *) my_wait_event_info = wait_event_info;
      |         ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^~~~~~~~~~~~~~~~~
../pgsql/src/backend/storage/buffer/bufmgr.c: In function \342\200\230BufferLockAcquire\342\200\231:
../pgsql/src/backend/storage/buffer/bufmgr.c:5781:33: note: \342\200\230wait_event\342\200\231 was declared here
 5781 |                 uint32          wait_event;
      |                                 ^~~~~~~~~~

Apparently they do not find the coding in this switch persuasive:

		switch (mode)
		{
			case BUFFER_LOCK_EXCLUSIVE:
				wait_event = WAIT_EVENT_BUFFER_EXCLUSIVE;
				break;
			case BUFFER_LOCK_SHARE_EXCLUSIVE:
				wait_event = WAIT_EVENT_BUFFER_SHARE_EXCLUSIVE;
				break;
			case BUFFER_LOCK_SHARE:
				wait_event = WAIT_EVENT_BUFFER_SHARED;
				break;
			case BUFFER_LOCK_UNLOCK:
				pg_unreachable();

		}

It's not clear to me whether that's more about not believing
pg_unreachable() or more about the lack of a default: case.
I see that this is a modification of code that existed before
fcb9c977a and wasn't being complained of, which makes it even
stranger.

			regards, tom lane

[1] https://buildfarm.postgresql.org/cgi-bin/show_stage_log.pl?nm=calliphoridae&dt=2026-01-15%2019%3...





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
@ 2026-01-15 23:16                                                           ` Andres Freund <andres@anarazel.de>
  2026-01-15 23:19                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  1 sibling, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-01-15 23:16 UTC (permalink / raw)
  To: Tom Lane <tgl@sss.pgh.pa.us>; +Cc: Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-15 18:02:34 -0500, Tom Lane wrote:
> In file included from ../pgsql/src/include/pgstat.h:24,
>                  from ../pgsql/src/backend/storage/buffer/bufmgr.c:52:
> In function \342\200\230pgstat_report_wait_start\342\200\231,
>     inlined from \342\200\230BufferLockAcquire\342\200\231 at ../pgsql/src/backend/storage/buffer/bufmgr.c:5833:3:
> ../pgsql/src/include/utils/wait_event.h:75:49: warning: \342\200\230wait_event\342\200\231 may be used uninitialized [-Wmaybe-uninitialized]
>    75 |         *(volatile uint32 *) my_wait_event_info = wait_event_info;
> [...]
> Apparently they do not find the coding in this switch persuasive:
> 
> 		switch (mode)
> 		{
> 			case BUFFER_LOCK_EXCLUSIVE:
> 				wait_event = WAIT_EVENT_BUFFER_EXCLUSIVE;
> 				break;
> 			case BUFFER_LOCK_SHARE_EXCLUSIVE:
> 				wait_event = WAIT_EVENT_BUFFER_SHARE_EXCLUSIVE;
> 				break;
> 			case BUFFER_LOCK_SHARE:
> 				wait_event = WAIT_EVENT_BUFFER_SHARED;
> 				break;
> 			case BUFFER_LOCK_UNLOCK:
> 				pg_unreachable();
> 
> 		}
> 
> It's not clear to me whether that's more about not believing
> pg_unreachable() or more about the lack of a default: case.

> I see that this is a modification of code that existed before
> fcb9c977a and wasn't being complained of, which makes it even
> stranger.

The code before did have a default:, so I guess that could be the
difference...

I can reproduce it here if I build with -Og or -O1, but not with -O2, so I
guess it's "just" the optimizer falling short somehow.

Huh, interesting, the optimization level dependency made me curious. Turns out
the warning goes away if I actually force the inlining of BufferLockAcquire()
with pg_attribute_always_inline. Presumably because then the compiler realizes
that the mode argument is always a constant that it can evaluate.

Can't quite decide between just using always_inline - after all the buffer
locking really is a rather crucial code path - and just initializing
wait_event to 0. Either seems better than using a default:.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-15 23:16                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-15 23:19                                                             ` Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-15 23:26                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Tom Lane @ 2026-01-15 23:19 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Andres Freund <andres@anarazel.de> writes:
> Can't quite decide between just using always_inline - after all the buffer
> locking really is a rather crucial code path - and just initializing
> wait_event to 0. Either seems better than using a default:.

Yeah, I agree with not using a default: in case we ever extend
the mode enum.  I'd be inclined to fix it like this:

 			case BUFFER_LOCK_UNLOCK:
 				pg_unreachable();
+				/* silence possible compiler warning: */ 
+				wait_event = 0;
  		}

			regards, tom lane





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-15 23:16                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:19                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
@ 2026-01-15 23:26                                                               ` Andres Freund <andres@anarazel.de>
  2026-01-16 12:02                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-01-15 23:26 UTC (permalink / raw)
  To: Tom Lane <tgl@sss.pgh.pa.us>; +Cc: Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On 2026-01-15 18:19:41 -0500, Tom Lane wrote:
> Andres Freund <andres@anarazel.de> writes:
> > Can't quite decide between just using always_inline - after all the buffer
> > locking really is a rather crucial code path - and just initializing
> > wait_event to 0. Either seems better than using a default:.
> 
> Yeah, I agree with not using a default: in case we ever extend
> the mode enum.  I'd be inclined to fix it like this:
> 
>  			case BUFFER_LOCK_UNLOCK:
>  				pg_unreachable();
> +				/* silence possible compiler warning: */ 
> +				wait_event = 0;
>   		}

Unfortunately that doesn't seem to do the trick, at least not locally (gcc
14). I think gcc is considering the case where mode has a value outside of the
known values of the enum.  Hence needing initialize wait_event to zero in the
declaration to avoid the warning.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-15 23:16                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:19                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-15 23:26                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-16 12:02                                                                 ` Andres Freund <andres@anarazel.de>
  0 siblings, 0 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-16 12:02 UTC (permalink / raw)
  To: Tom Lane <tgl@sss.pgh.pa.us>; +Cc: Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-15 18:26:18 -0500, Andres Freund wrote:
> On 2026-01-15 18:19:41 -0500, Tom Lane wrote:
> > Andres Freund <andres@anarazel.de> writes:
> > > Can't quite decide between just using always_inline - after all the buffer
> > > locking really is a rather crucial code path - and just initializing
> > > wait_event to 0. Either seems better than using a default:.
> > 
> > Yeah, I agree with not using a default: in case we ever extend
> > the mode enum.  I'd be inclined to fix it like this:
> > 
> >  			case BUFFER_LOCK_UNLOCK:
> >  				pg_unreachable();
> > +				/* silence possible compiler warning: */ 
> > +				wait_event = 0;
> >   		}
> 
> Unfortunately that doesn't seem to do the trick, at least not locally (gcc
> 14). I think gcc is considering the case where mode has a value outside of the
> known values of the enum.  Hence needing initialize wait_event to zero in the
> declaration to avoid the warning.

I've now pushed the initialization of wait_event, as I'll be only semi-online
until late Monday.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
@ 2026-01-24 19:00                                                           ` Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  1 sibling, 1 reply; 120+ messages in thread

From: Alexander Lakhin @ 2026-01-24 19:00 UTC (permalink / raw)
  To: Tom Lane <tgl@sss.pgh.pa.us>; Andres Freund <andres@anarazel.de>; +Cc: Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hello Andres,

16.01.2026 01:02, Tom Lane wrote:
> Various buildfarm animals are complaining about fcb9c977a,
> similarly to this from calliphoridae [1]:

I've discovered another anomaly introduced with fcb9c977a, this time
run-time:
for i in `seq 300`; do
echo "iteration $i"

echo "
create table t(f1 text);
create index on t using spgist(f1);
insert into t select 'a' from generate_series(1, 9000) g(i);
vacuum analyze t;
insert into t select 'b' from generate_series(1, 1000) g(i);
drop table t;
" | psql >/dev/null -v ON_ERROR_STOP=1 || break;
done

fails for me as below:
...
iteration 39
server closed the connection unexpectedly

Core was generated by `postgres: law regression [local] INSERT                 '.
Program terminated with signal SIGABRT, Aborted.
#0  __pthread_kill_implementation (no_tid=0, signo=6, threadid=<optimized out>) at ./nptl/pthread_kill.c:44

warning: 44     ./nptl/pthread_kill.c: Нет такого файла или каталога
(gdb) bt
#0  __pthread_kill_implementation (no_tid=0, signo=6, threadid=<optimized out>) at ./nptl/pthread_kill.c:44
#1  __pthread_kill_internal (signo=6, threadid=<optimized out>) at ./nptl/pthread_kill.c:78
#2  __GI___pthread_kill (threadid=<optimized out>, signo=signo@entry=6) at ./nptl/pthread_kill.c:89
#3  0x000073625aa4527e in __GI_raise (sig=sig@entry=6) at ../sysdeps/posix/raise.c:26
#4  0x000073625aa288ff in __GI_abort () at ./stdlib/abort.c:79
#5  0x00005a479e2f685f in ExceptionalCondition (conditionName=conditionName@entry=0x5a479e403cb8 "entry->data.lockmode 
== BUFFER_LOCK_UNLOCK", fileName=fileName@entry=0x5a479e37a84f "bufmgr.c", lineNumber=lineNumber@entry=5908) at assert.c:65
#6  0x00005a479e14d87c in BufferLockConditional (buffer=<optimized out>, buf_hdr=0x73624e882b40, 
mode=mode@entry=BUFFER_LOCK_EXCLUSIVE) at bufmgr.c:5908
#7  0x00005a479e14f4d5 in ConditionalLockBuffer (buffer=buffer@entry=13500) at bufmgr.c:6474
#8  0x00005a479de44da9 in SpGistNewBuffer (index=index@entry=0x73625b1f6ce8) at spgutils.c:420
#9  0x00005a479de45397 in allocNewBuffer (index=index@entry=0x73625b1f6ce8, flags=flags@entry=3) at spgutils.c:528
#10 0x00005a479de456a0 in SpGistGetBuffer (index=index@entry=0x73625b1f6ce8, flags=flags@entry=3, needSpace=<optimized 
out>, needSpace@entry=4088, isNew=isNew@entry=0x7ffd8fa1b2a7) at spgutils.c:663
#11 0x00005a479de3c913 in doPickSplit (index=index@entry=0x73625b1f6ce8, state=state@entry=0x7ffd8fa1b740, 
current=current@entry=0x7ffd8fa1b510, parent=parent@entry=0x7ffd8fa1b530, 
newLeafTuple=newLeafTuple@entry=0x5a47cbd7adc8, level=level@entry=4, isNulls=false, isNew=false) at spgdoinsert.c:1046
#12 0x00005a479de3e542 in spgdoinsert (index=index@entry=0x73625b1f6ce8, state=state@entry=0x7ffd8fa1b740, 
heapPtr=heapPtr@entry=0x5a47cbe0b248, datums=datums@entry=0x7ffd8fa1b8d0, isnulls=isnulls@entry=0x7ffd8fa1b8b0) at 
spgdoinsert.c:2134
#13 0x00005a479de40137 in spginsert (index=0x73625b1f6ce8, values=0x7ffd8fa1b8d0, isnull=0x7ffd8fa1b8b0, 
ht_ctid=0x5a47cbe0b248, heapRel=<optimized out>, checkUnique=<optimized out>, indexUnchanged=false, 
indexInfo=0x5a47cbe0b0b8) at spginsert.c:206
#14 0x00005a479df981d8 in ExecInsertIndexTuples (resultRelInfo=resultRelInfo@entry=0x5a47cbe07120, 
slot=slot@entry=0x5a47cbe0b218, estate=estate@entry=0x5a47cbe06c10, update=update@entry=false, 
noDupErr=noDupErr@entry=false, specConflict=specConflict@entry=0x0, arbiterIndexes=0x0, onlySummarizing=false) at 
execIndexing.c:449
#15 0x00005a479dfcc6a9 in ExecInsert (context=context@entry=0x7ffd8fa1bb70, 
resultRelInfo=resultRelInfo@entry=0x5a47cbe07120, slot=slot@entry=0x5a47cbe0b218, canSetTag=<optimized out>, 
inserted_tuple=inserted_tuple@entry=0x0, insert_destrel=insert_destrel@entry=0x0) at nodeModifyTable.c:1240
#16 0x00005a479dfcdf67 in ExecModifyTable (pstate=0x5a47cbe06f10) at nodeModifyTable.c:4485
...

Best regards,
Alexander

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
@ 2026-01-24 20:31                                                             ` Andres Freund <andres@anarazel.de>
  2026-01-24 21:03                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  0 siblings, 2 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-24 20:31 UTC (permalink / raw)
  To: Alexander Lakhin <exclusion@gmail.com>; +Cc: Tom Lane <tgl@sss.pgh.pa.us>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-24 21:00:00 +0200, Alexander Lakhin wrote:
> 16.01.2026 01:02, Tom Lane wrote:
> > Various buildfarm animals are complaining about fcb9c977a,
> > similarly to this from calliphoridae [1]:
>
> I've discovered another anomaly introduced with fcb9c977a, this time
> run-time:

With the other anomaly, you mean the spurious compiler warning or something
else?


> for i in `seq 300`; do
> echo "iteration $i"
>
> echo "
> create table t(f1 text);
> create index on t using spgist(f1);
> insert into t select 'a' from generate_series(1, 9000) g(i);
> vacuum analyze t;
> insert into t select 'b' from generate_series(1, 1000) g(i);
> drop table t;
> " | psql >/dev/null -v ON_ERROR_STOP=1 || break;
> done
>
> fails for me as below:
> ...
> iteration 39
> server closed the connection unexpectedly


Thanks for this report, particularly with the easy reproducer!  How did you
find this?


I think this is more likely to be a spgist bug, not a bug in the patch.  From
what I can tell, spgist tries to conditionally lock a buffer that it itself
already has locked exclusively - that's why the assertion is failing.

I reproduced this locally, and could see in a bt full stack that the buffer
that spgist is trying to lock conditionally, is also referenced as
newInnerBuffer in doPickSplit(). So it's not an issue of bufmgr.c loosing
track of which buffers are locked with what mode.

I haven't yet figured out why spgist ends up with a buffer it already is
using.

We could of course just accept this case and have the conditional lock
acquisition fail, but I think trying to conditionally lock a buffer that you
already lock is indicative of something having gone wrong.  But I'm open to
going there anyway, just to avoid causing problems with previously "working"
code.


Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-24 21:03                                                               ` Andres Freund <andres@anarazel.de>
  1 sibling, 0 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-24 21:03 UTC (permalink / raw)
  To: Alexander Lakhin <exclusion@gmail.com>; +Cc: Tom Lane <tgl@sss.pgh.pa.us>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-24 15:31:14 -0500, Andres Freund wrote:
> I think this is more likely to be a spgist bug, not a bug in the patch.  From
> what I can tell, spgist tries to conditionally lock a buffer that it itself
> already has locked exclusively - that's why the assertion is failing.
> 
> I reproduced this locally, and could see in a bt full stack that the buffer
> that spgist is trying to lock conditionally, is also referenced as
> newInnerBuffer in doPickSplit(). So it's not an issue of bufmgr.c loosing
> track of which buffers are locked with what mode.
> 
> I haven't yet figured out why spgist ends up with a buffer it already is
> using.
> 
> We could of course just accept this case and have the conditional lock
> acquisition fail, but I think trying to conditionally lock a buffer that you
> already lock is indicative of something having gone wrong.  But I'm open to
> going there anyway, just to avoid causing problems with previously "working"
> code.

Looking at the spgist code, and the README, I think we may need to accept the
uglines of silently failing when a backend tries to conditionally lock a
buffer that it itself has already locked.  Even though I still don't
understand how it happens in this this specific case, that doesn't even have
concurrency.

Pretty ... not great ... that spgist does stuff like extending a relation
while holding an exclusive buffer lock.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-24 21:11                                                               ` Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 21:21                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  1 sibling, 2 replies; 120+ messages in thread

From: Tom Lane @ 2026-01-24 21:11 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Andres Freund <andres@anarazel.de> writes:
> I think this is more likely to be a spgist bug, not a bug in the patch.  From
> what I can tell, spgist tries to conditionally lock a buffer that it itself
> already has locked exclusively - that's why the assertion is failing.

I dunno.  It looks to me like the previous LWLock-based implementation
of ConditionalLockBuffer() had no such restriction as

	/*
	 * We better not already hold a lock on the buffer.
	 */
	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);

Maybe I'm missing something, but it looks like the old code would
return false if the buffer was already locked, whether that lock
was held by our process or another one.  SPGist evidently has an
assumption in edge cases that that is the behavior, and I'm not
convinced that it's a good idea to change it.  There may be other
such edge cases we've not tripped over yet.

> We could of course just accept this case and have the conditional lock
> acquisition fail, but I think trying to conditionally lock a buffer that you
> already lock is indicative of something having gone wrong.

I don't really buy this argument.  Yes, within a single function it'd
be silly to lock a buffer and immediately try to lock it again, but
when you consider cases like recursive modifications of index state,
it's *far* from obvious that some lower recursion level might not try
to lock a buffer that some outer level already locked.  In the case at
hand I think it is probably driven by two recursion levels trying to
acquire free space out of the same buffer.  SPGist is expecting the
lower level to fail to get the lock and then go find some free space
elsewhere.  Yeah, we could probably re-code it to get that outcome
in another way, but why?

			regards, tom lane





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
@ 2026-01-24 21:21                                                                 ` Peter Geoghegan <pg@bowt.ie>
  1 sibling, 0 replies; 120+ messages in thread

From: Peter Geoghegan @ 2026-01-24 21:21 UTC (permalink / raw)
  To: Tom Lane <tgl@sss.pgh.pa.us>; +Cc: Andres Freund <andres@anarazel.de>; Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Sat, Jan 24, 2026 at 4:12 PM Tom Lane <tgl@sss.pgh.pa.us> wrote:
> Andres Freund <andres@anarazel.de> writes:
> > We could of course just accept this case and have the conditional lock
> > acquisition fail, but I think trying to conditionally lock a buffer that you
> > already lock is indicative of something having gone wrong.
>
> I don't really buy this argument.  Yes, within a single function it'd
> be silly to lock a buffer and immediately try to lock it again, but
> when you consider cases like recursive modifications of index state,
> it's *far* from obvious that some lower recursion level might not try
> to lock a buffer that some outer level already locked.  In the case at
> hand I think it is probably driven by two recursion levels trying to
> acquire free space out of the same buffer.

Even nbtree has to deal with this. Also in the context of free space
management. See the comments in _bt_allocbuf, particularly the ones
where we explicitly describe a common scenario where we conditionally
lock a buffer that our own caller/backend already has a lock on.

I actually agree with Andres' general sentiment about this kind of
coding pattern; it also seems sloppy to me. But it's hard to see how
we could do better in places such as _bt_allocbuf. At least within the
confines of the current FSM design.

-- 
Peter Geoghegan





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
@ 2026-01-24 23:03                                                                 ` Andres Freund <andres@anarazel.de>
  2026-01-25 00:54                                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  1 sibling, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-01-24 23:03 UTC (permalink / raw)
  To: Tom Lane <tgl@sss.pgh.pa.us>; +Cc: Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-24 16:11:47 -0500, Tom Lane wrote:
> Andres Freund <andres@anarazel.de> writes:
> > I think this is more likely to be a spgist bug, not a bug in the patch.  From
> > what I can tell, spgist tries to conditionally lock a buffer that it itself
> > already has locked exclusively - that's why the assertion is failing.
> 
> I dunno.  It looks to me like the previous LWLock-based implementation
> of ConditionalLockBuffer() had no such restriction as
> 
> 	/*
> 	 * We better not already hold a lock on the buffer.
> 	 */
> 	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
> 
> Maybe I'm missing something, but it looks like the old code would
> return false if the buffer was already locked, whether that lock
> was held by our process or another one.

Yea. I added that assert when (I think) Melanie complained that the new code
wouldn't detect repeated acquisitions or release of the same content lock as
nicely as before.  I guess I went a bit overboard and also added the assertion
to ConditionalLockBuffer().

I'll go and move it into the branch where we actually got the lock, that seems
worth continuing to do.


> > We could of course just accept this case and have the conditional lock
> > acquisition fail, but I think trying to conditionally lock a buffer that you
> > already lock is indicative of something having gone wrong.
> 
> I don't really buy this argument.  Yes, within a single function it'd
> be silly to lock a buffer and immediately try to lock it again, but
> when you consider cases like recursive modifications of index state,
> it's *far* from obvious that some lower recursion level might not try
> to lock a buffer that some outer level already locked.

Yea, there's probably something too that. But I'm not entirely convinced -
consider what happens with an index on a temporary table: A higher level think
it got a buffer with space, but then a lower level uses up that space...


> In the case at hand I think it is probably driven by two recursion levels
> trying to acquire free space out of the same buffer.  SPGist is expecting
> the lower level to fail to get the lock and then go find some free space
> elsewhere.  Yeah, we could probably re-code it to get that outcome in
> another way, but why?

Regardless of the assertion, it still feels like there may be something off
here. Why is the page marked as empty in the FSM, despite actually not being
empty?


Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-25 00:54                                                                   ` Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-29 17:27                                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Tom Lane @ 2026-01-25 00:54 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Andres Freund <andres@anarazel.de> writes:
> On 2026-01-24 16:11:47 -0500, Tom Lane wrote:
>> In the case at hand I think it is probably driven by two recursion levels
>> trying to acquire free space out of the same buffer.  SPGist is expecting
>> the lower level to fail to get the lock and then go find some free space
>> elsewhere.  Yeah, we could probably re-code it to get that outcome in
>> another way, but why?

> Regardless of the assertion, it still feels like there may be something off
> here. Why is the page marked as empty in the FSM, despite actually not being
> empty?

Who said anything about its being empty?  There's X amount of space
used on the page now, the outer level is preparing to add a tuple
of size Y, the inner level is looking for a place to put a tuple
of size Z.  As long as X+Y+Z <= 8K, there's no free-space-related
reason the page wouldn't work for this.  We could actually try to
make it work, but that seems fragile to me: the outer level would
have to be prepared for the possibility that the page changes
under it despite holding exclusive lock.  I think it's reasonable
on complexity grounds to insist that the inner recursion level
go play someplace else.

			regards, tom lane





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-25 00:54                                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
@ 2026-01-29 17:27                                                                     ` Andres Freund <andres@anarazel.de>
  2026-01-29 17:42                                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 18:12                                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
  2026-01-29 20:24                                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  0 siblings, 3 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-29 17:27 UTC (permalink / raw)
  To: Tom Lane <tgl@sss.pgh.pa.us>; Peter Geoghegan <pg@bowt.ie>; +Cc: Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-24 19:54:36 -0500, Tom Lane wrote:
> Andres Freund <andres@anarazel.de> writes:
> > On 2026-01-24 16:11:47 -0500, Tom Lane wrote:
> >> In the case at hand I think it is probably driven by two recursion levels
> >> trying to acquire free space out of the same buffer.  SPGist is expecting
> >> the lower level to fail to get the lock and then go find some free space
> >> elsewhere.  Yeah, we could probably re-code it to get that outcome in
> >> another way, but why?
> 
> > Regardless of the assertion, it still feels like there may be something off
> > here. Why is the page marked as empty in the FSM, despite actually not being
> > empty?
> 
> Who said anything about its being empty?

In the crashes due to Alexander's repro that I have looked at, the page is
returned by SpGistNewBuffer()->GetFreeIndexPage(). Which afaict should only
contain empty pages. It's easy to see how that could happen with concurrency,
but there's none in the test.

I was just trying to repro this again while writing this message, and
interestingly I got the same issue in nbtree this time. Which a) confirms
Peter's statement that the "conditionally locking a buffer we already locked"
issue exists for nbtree b) makes me suspect something odd is happening around
indexfsm.


Anyway, independent of that, the behavior clearly needs to be allowed. Here's
a proposed patch.

At first I was thinking of just removing the assertion without anything else
in place - but I think that's not quite right: We could e.g. be trying to
acquire a share or share-exclusive lock when holding a share lock (or the
reverse), but we can't currently don't keep track of two different lock modes
for the same lock.  Therefore it seems safer to just define it so that
acquiring a conditional lock on a buffer that is already locked by us will
always fail, regardless of what existing lock mode we already hold.  I think
all current callers good with that.

Does that sound reasonable?

We could add support for locking the same buffer multiple times, but I don't
think it'd be worth the complexity and (small) overhead that would bring with
it?  It also seems like allowing that would make it more likely for a backend
to trample over its own state higher up in the call tree.

Greetings,

Andres Freund
From b90f7fbe9958ea785a1302395bb1a201a00f3bc5 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 29 Jan 2026 12:22:02 -0500
Subject: [PATCH v1] bufmgr: Allow conditionally locking of already locked
 buffer

Discussion: https://postgr.es/m/90bd2cbb-49ce-4092-9f61-5ac2ab782c94@gmail.com
---
 src/backend/storage/buffer/bufmgr.c | 18 +++++++++++++++---
 1 file changed, 15 insertions(+), 3 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 6f935648ae9..f5602f4e7e1 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5895,6 +5895,13 @@ BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr)
 
 /*
  * Acquire the content lock for the buffer, but only if we don't have to wait.
+ *
+ * It is allowed to try to conditionally acquire a lock on a buffer that this
+ * backend has already locked, but the lock acquisition will always fail, even
+ * if the new lock acquisition does not conflict with an already held lock
+ * (e.g. two share locks). This is because we don't track per-backend
+ * ownership of multiple lock levels.  That is ok for the current uses of
+ * BufferLockConditional().
  */
 static bool
 BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
@@ -5902,11 +5909,16 @@ BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
 	PrivateRefCountEntry *entry = GetPrivateRefCountEntry(buffer, true);
 	bool		mustwait;
 
-	/*
-	 * We better not already hold a lock on the buffer.
-	 */
 	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
 
+	/*
+	 * As described above, if we're trying to lock a buffer this backend
+	 * already has locked, return false, independent of the existing and
+	 * desired lock level.
+	 */
+	if (entry->data.lockmode != BUFFER_LOCK_UNLOCK)
+		return false;
+
 	/*
 	 * Lock out cancel/die interrupts until we exit the code section protected
 	 * by the content lock.  This ensures that interrupts will not interfere
-- 
2.48.1.76.g4e746b1a31.dirty

Attachments:

  [text/plain] v1-0001-bufmgr-Allow-conditionally-locking-of-already-loc.patch.txt (2.0K, ../../q2o7rfmkcvqhj7ttimno33mb7lklybj6aanliertmfr4g7g7zn@w4cat7zimc5c/2-v1-0001-bufmgr-Allow-conditionally-locking-of-already-loc.patch.txt)
  download | inline diff:
From b90f7fbe9958ea785a1302395bb1a201a00f3bc5 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Thu, 29 Jan 2026 12:22:02 -0500
Subject: [PATCH v1] bufmgr: Allow conditionally locking of already locked
 buffer

Discussion: https://postgr.es/m/90bd2cbb-49ce-4092-9f61-5ac2ab782c94@gmail.com
---
 src/backend/storage/buffer/bufmgr.c | 18 +++++++++++++++---
 1 file changed, 15 insertions(+), 3 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 6f935648ae9..f5602f4e7e1 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5895,6 +5895,13 @@ BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr)
 
 /*
  * Acquire the content lock for the buffer, but only if we don't have to wait.
+ *
+ * It is allowed to try to conditionally acquire a lock on a buffer that this
+ * backend has already locked, but the lock acquisition will always fail, even
+ * if the new lock acquisition does not conflict with an already held lock
+ * (e.g. two share locks). This is because we don't track per-backend
+ * ownership of multiple lock levels.  That is ok for the current uses of
+ * BufferLockConditional().
  */
 static bool
 BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
@@ -5902,11 +5909,16 @@ BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
 	PrivateRefCountEntry *entry = GetPrivateRefCountEntry(buffer, true);
 	bool		mustwait;
 
-	/*
-	 * We better not already hold a lock on the buffer.
-	 */
 	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
 
+	/*
+	 * As described above, if we're trying to lock a buffer this backend
+	 * already has locked, return false, independent of the existing and
+	 * desired lock level.
+	 */
+	if (entry->data.lockmode != BUFFER_LOCK_UNLOCK)
+		return false;
+
 	/*
 	 * Lock out cancel/die interrupts until we exit the code section protected
 	 * by the content lock.  This ensures that interrupts will not interfere
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-25 00:54                                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-29 17:27                                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-29 17:42                                                                       ` Andres Freund <andres@anarazel.de>
  2026-01-29 17:50                                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
  2 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-01-29 17:42 UTC (permalink / raw)
  To: Tom Lane <tgl@sss.pgh.pa.us>; Peter Geoghegan <pg@bowt.ie>; +Cc: Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-29 12:27:10 -0500, Andres Freund wrote:
> In the crashes due to Alexander's repro that I have looked at, the page is
> returned by SpGistNewBuffer()->GetFreeIndexPage(). Which afaict should only
> contain empty pages. It's easy to see how that could happen with concurrency,
> but there's none in the test.

I think I see how that happens - we reenter the page into the FSM during the
VACUUM that's part of the repro, but we previously also "recorded" it with
SpGistSetLastUsedPage(), presumably before it was emptied during vacuum.  The
we reuse the page first via the "last used page" cache, making it not
empty. Then we get the same page via the FSM, which wasn't updated when we
started filling the page via the last-used-page mechanism.


> I was just trying to repro this again while writing this message, and
> interestingly I got the same issue in nbtree this time. Which a) confirms
> Peter's statement that the "conditionally locking a buffer we already locked"
> issue exists for nbtree b) makes me suspect something odd is happening around
> indexfsm.

Not sure how that happens in a single threaded workload yet, though.


> diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
> index 6f935648ae9..f5602f4e7e1 100644
> --- a/src/backend/storage/buffer/bufmgr.c
> +++ b/src/backend/storage/buffer/bufmgr.c
> @@ -5895,6 +5895,13 @@ BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr)
>  
>  /*
>   * Acquire the content lock for the buffer, but only if we don't have to wait.
> + *
> + * It is allowed to try to conditionally acquire a lock on a buffer that this
> + * backend has already locked, but the lock acquisition will always fail, even
> + * if the new lock acquisition does not conflict with an already held lock
> + * (e.g. two share locks). This is because we don't track per-backend
> + * ownership of multiple lock levels.  That is ok for the current uses of
> + * BufferLockConditional().
>   */
>  static bool
>  BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
> @@ -5902,11 +5909,16 @@ BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
>  	PrivateRefCountEntry *entry = GetPrivateRefCountEntry(buffer, true);
>  	bool		mustwait;
>  
> -	/*
> -	 * We better not already hold a lock on the buffer.
> -	 */
>  	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
>  

Err, this should have been removed, I accidentally re-added the hunk while
experimenting.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-25 00:54                                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-29 17:27                                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 17:42                                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-29 17:50                                                                         ` Peter Geoghegan <pg@bowt.ie>
  2026-01-29 18:06                                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Peter Geoghegan @ 2026-01-29 17:50 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Tom Lane <tgl@sss.pgh.pa.us>; Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Thu, Jan 29, 2026 at 12:42 PM Andres Freund <andres@anarazel.de> wrote:
> > -     /*
> > -      * We better not already hold a lock on the buffer.
> > -      */
> >       Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
> >
>
> Err, this should have been removed, I accidentally re-added the hunk while
> experimenting.

I've been running into this assertion failure from time to time while
working on index prefetching. It seems to happen after a hard crash.
I've just been running initdb whenever it happens. It would be nice to
not have to do this again.

-- 
Peter Geoghegan





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-25 00:54                                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-29 17:27                                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 17:42                                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 17:50                                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
@ 2026-01-29 18:06                                                                           ` Andres Freund <andres@anarazel.de>
  2026-01-29 18:33                                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
  2026-01-29 21:49                                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 2 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-29 18:06 UTC (permalink / raw)
  To: Peter Geoghegan <pg@bowt.ie>; +Cc: Tom Lane <tgl@sss.pgh.pa.us>; Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-29 12:50:38 -0500, Peter Geoghegan wrote:
> On Thu, Jan 29, 2026 at 12:42 PM Andres Freund <andres@anarazel.de> wrote:
> > > -     /*
> > > -      * We better not already hold a lock on the buffer.
> > > -      */
> > >       Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
> > >
> >
> > Err, this should have been removed, I accidentally re-added the hunk while
> > experimenting.
> 
> I've been running into this assertion failure from time to time while
> working on index prefetching. It seems to happen after a hard crash.
> I've just been running initdb whenever it happens. It would be nice to
> not have to do this again.

Sure, I'm planning to give folks a bit longer to chime in whether the proposed
behavior is sane and will then push (with an amended commit message and the
fixup discussed here).


After a crash it seems that btree often ends up getting the page from FSM that
it's currently inserting on. Which makes some sense, because presumably that
was the page we filled just before a crash. Wonder if - independent of this
issue - it could make sense to update the FSM during nbtree WAL recovery...

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-25 00:54                                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-29 17:27                                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 17:42                                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 17:50                                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
  2026-01-29 18:06                                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-29 18:33                                                                             ` Peter Geoghegan <pg@bowt.ie>
  2026-01-29 19:29                                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  1 sibling, 1 reply; 120+ messages in thread

From: Peter Geoghegan @ 2026-01-29 18:33 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Tom Lane <tgl@sss.pgh.pa.us>; Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Thu, Jan 29, 2026 at 1:06 PM Andres Freund <andres@anarazel.de> wrote:
> Wonder if - independent of this
> issue - it could make sense to update the FSM during nbtree WAL recovery...

Maybe that would make sense. But I tend to think that we should have a
fully atomic, crash-safe approach to free space management.

Particularly in index AMs, where free space can only ever come in
BLCKSZ units -- the data structure/concurrency rules can be a lot
simpler if it only has to accommodate index AM requirements. Maybe the
WAL-logging could be built into existing index AM record types.

-- 
Peter Geoghegan





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-25 00:54                                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-29 17:27                                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 17:42                                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 17:50                                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
  2026-01-29 18:06                                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 18:33                                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
@ 2026-01-29 19:29                                                                               ` Andres Freund <andres@anarazel.de>
  0 siblings, 0 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-29 19:29 UTC (permalink / raw)
  To: Peter Geoghegan <pg@bowt.ie>; +Cc: Tom Lane <tgl@sss.pgh.pa.us>; Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-29 13:33:02 -0500, Peter Geoghegan wrote:
> On Thu, Jan 29, 2026 at 1:06 PM Andres Freund <andres@anarazel.de> wrote:
> > Wonder if - independent of this
> > issue - it could make sense to update the FSM during nbtree WAL recovery...
> 
> Maybe that would make sense. But I tend to think that we should have a
> fully atomic, crash-safe approach to free space management.

I agree that would be nice, but realistically (as you also say below) that
would have to be embedded into the WAL records that use the page that was
acquired from the FSM.  Maybe we could accept a dedicated WAL record for the
index case, but certainly not in the heap case.

Given that we'd need to embed the record somehow anyway, just adding, for now,
a RecordUsedIndexPage() to the redo of XLOG_BTREE_SPLIT* and
XLOG_BTREE_NEWROOT or such could make sense...

It doesn't seem like it'd be great to have a completely outdated index fsm
after a failover. If the index FSM on the newly promoted node is completely
outdated, due to having been copied at a much earlier time while there were a
lot of free pages, a _bt_allocbuf() could take quite a while...

I'm somewhat surprised it doesn't cause more performance issues to keep btree
pages exclusively locked while extending the relation... If that has to write
out pages and flush the WAL...


> Particularly in index AMs, where free space can only ever come in
> BLCKSZ units -- the data structure/concurrency rules can be a lot
> simpler if it only has to accommodate index AM requirements. Maybe the
> WAL-logging could be built into existing index AM record types.

Yea, I have my doubt that makes sense to share code between the index and heap
use cases. I doubt that having one FSM implementation support variable amount
of "space tracking granularity" really makes sense.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-25 00:54                                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-29 17:27                                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 17:42                                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-29 17:50                                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Peter Geoghegan <pg@bowt.ie>
  2026-01-29 18:06                                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-29 21:49                                                                             ` Andres Freund <andres@anarazel.de>
  1 sibling, 0 replies; 120+ messages in thread

From: Andres Freund @ 2026-01-29 21:49 UTC (permalink / raw)
  To: Peter Geoghegan <pg@bowt.ie>; +Cc: Tom Lane <tgl@sss.pgh.pa.us>; Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-29 13:06:22 -0500, Andres Freund wrote:
> On 2026-01-29 12:50:38 -0500, Peter Geoghegan wrote:
> > On Thu, Jan 29, 2026 at 12:42 PM Andres Freund <andres@anarazel.de> wrote:
> > > > -     /*
> > > > -      * We better not already hold a lock on the buffer.
> > > > -      */
> > > >       Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
> > > >
> > >
> > > Err, this should have been removed, I accidentally re-added the hunk while
> > > experimenting.
> > 
> > I've been running into this assertion failure from time to time while
> > working on index prefetching. It seems to happen after a hard crash.
> > I've just been running initdb whenever it happens. It would be nice to
> > not have to do this again.
> 
> Sure, I'm planning to give folks a bit longer to chime in whether the proposed
> behavior is sane and will then push (with an amended commit message and the
> fixup discussed here).

Pushed the fix now.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-25 00:54                                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-29 17:27                                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-29 18:12                                                                       ` Peter Geoghegan <pg@bowt.ie>
  2 siblings, 0 replies; 120+ messages in thread

From: Peter Geoghegan @ 2026-01-29 18:12 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Tom Lane <tgl@sss.pgh.pa.us>; Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Thu, Jan 29, 2026 at 12:27 PM Andres Freund <andres@anarazel.de> wrote:
> I was just trying to repro this again while writing this message, and
> interestingly I got the same issue in nbtree this time. Which a) confirms
> Peter's statement that the "conditionally locking a buffer we already locked"
> issue exists for nbtree b) makes me suspect something odd is happening around
> indexfsm.

I don't think that there's anything mysterious about it. This is just
how index vacuuming does free space management. It's a consequence of
the fact that VACUUM notices that a page can go in the FSM at one
point, but only actually updates the FSM at some other point.

Nothing stops a backend from finding a recyclable page in the FSM
after a concurrent VACUUM decides that that same page can go in the
FSM, but before actually placing that page in the FSM.  VACUUM doesn't
consider that the FSM might have already known about that page *at
all* -- and so it certainly doesn't try to avoid these kinds of race
conditions. Hence the need for _bt_allocbuf to worry about buffer lock
deadlocks, including even self-deadlock.

-- 
Peter Geoghegan





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 23:02                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 19:00                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-01-24 20:31                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-24 21:11                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-24 23:03                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-25 00:54                                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Tom Lane <tgl@sss.pgh.pa.us>
  2026-01-29 17:27                                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-29 20:24                                                                       ` Tom Lane <tgl@sss.pgh.pa.us>
  2 siblings, 0 replies; 120+ messages in thread

From: Tom Lane @ 2026-01-29 20:24 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Peter Geoghegan <pg@bowt.ie>; Alexander Lakhin <exclusion@gmail.com>; Chao Li <li.evan.chao@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Andres Freund <andres@anarazel.de> writes:
> Anyway, independent of that, the behavior clearly needs to be allowed. Here's
> a proposed patch.

> At first I was thinking of just removing the assertion without anything else
> in place - but I think that's not quite right: We could e.g. be trying to
> acquire a share or share-exclusive lock when holding a share lock (or the
> reverse), but we can't currently don't keep track of two different lock modes
> for the same lock.  Therefore it seems safer to just define it so that
> acquiring a conditional lock on a buffer that is already locked by us will
> always fail, regardless of what existing lock mode we already hold.  I think
> all current callers good with that.

> Does that sound reasonable?

I didn't read the patch, but I agree with this description of what the
behavior should be.

> We could add support for locking the same buffer multiple times, but I don't
> think it'd be worth the complexity and (small) overhead that would bring with
> it?

Also agreed --- I think that's behavior we actively don't want.

> It also seems like allowing that would make it more likely for a backend
> to trample over its own state higher up in the call tree.

Precisely.

			regards, tom lane





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 03:41                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 16:30                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 23:20                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-14 23:37                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-15 00:04                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 06:22                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Chao Li <li.evan.chao@gmail.com>
  2026-01-15 16:43                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-16 02:36                                                         ` Chao Li <li.evan.chao@gmail.com>
  1 sibling, 0 replies; 120+ messages in thread

From: Chao Li @ 2026-01-16 02:36 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>



> On Jan 16, 2026, at 00:43, Andres Freund <andres@anarazel.de> wrote:
> 
> Hi,
> 
> On 2026-01-15 14:22:09 +0800, Chao Li wrote:
>> 3 - 0005 - bufmgr.c
>> ```
>> +inline void
>> +MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
>> ```
>> 
>> It’s quite uncommon to extern an inline function. I think usually if we want
>> to make an inline function accessible from external, we define it “static
>> inline” in a header file. So, I guess “inline” is a typo here.
> 
> It works just fine to put an inline into the function definition, even if it's
> an external function. That hints the compiler to inline it inside the
> translation unit. Which is useful here, because it leads to
> MarkBufferDirtyHint() being inlined into BufferFinishSetHintBits().
> 

Technically, that for sure works. I raised the comment basically because it’s uncommon in the source tree. I just did a search for “inline” across all .c files, only found two non-static inline functions: ReadBufferExtended() and TrackNewBufferPin(), and they are all in bufmgr.c. So, adding a new one should be fine. I should have done the search yesterday while reviewing.

>> 8 - 0005 - fsmpage.c
>> ```
>> * needs to be updated. exclusive_lock_held should be set to true if the
>> * caller is already holding an exclusive lock, to avoid extra work.
>> ```
>> 
>> The function comment of fsm_search_avail() needs to be
>> updated. exclusive_lock_held should be set to true if the caller is already
>> holding a **share-exclusive or** exclusive lock.
> 
> Why?

Ah, I see. We will never use a share-exclusive lock for fsm_search_avail? Currently, there are two callers of fsm_search_avail, one use a share lock, the other uses an exclusive lock. My bad.

>> 10 - 0005 - freespace.c
>> ```
>> - * Reset the next slot pointer. This encourages the use of low-numbered
>> - * pages, increasing the chances that a later vacuum can truncate the
>> - * relation. We don't bother with marking the page dirty if it wasn't
>> - * already, since this is just a hint.
>> + * Try to reset the next slot pointer. This encourages the use of
>> + * low-numbered pages, increasing the chances that a later vacuum can
>> + * truncate the relation. We don't bother with marking the page dirty if
>> + * it wasn't already, since this is just a hint.
>> */
>> LockBuffer(buf, BUFFER_LOCK_SHARE);
>> - ((FSMPage) PageGetContents(page))->fp_next_slot = 0;
>> + if (BufferBeginSetHintBits(buf))
>> + {
>> + ((FSMPage) PageGetContents(page))->fp_next_slot = 0;
>> + BufferFinishSetHintBits(buf, false, false);
>> + }
>> ```
>> 
>> Before this patch, we unconditionally set fp_next_slot, now the setting
>> might be skipped. You have add “Try to” in the comment that has explained
>> the possibility of skipping setting fp_next_slot. Would it be better to add
>> a brief statement for what will result in when skipping setting
>> fp_next_slot? Something like “Skipping the update only affects reuse, not
>> correctness”.
> 
> I don't see the point. The whole paragraph is about how this isn't crucial.
> 

Yeah, I could understand the comment expresses that isn’t crucial, but just felt that’s not “explicit”. I am fine with the current comment. 

Best regards,
--
Chao Li (Evan)
HighGo Software Co., Ltd.
https://www.highgo.com/









^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-01-14 21:20                                           ` Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  4 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-01-14 21:20 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-12 19:33:56 -0500, Andres Freund wrote:
> Here are the remaining commits, with a bit of polish:

I pushed 0001, 0002.

Attached is an updated version of the remaining changes:

- I updated the definition in BUF_DEFINE_FLAG to have a redundant copy of
  BUF_FLAG_SHIFT's "contents", as suggested by Chao

- I realized I had forgotten to remove the BufferContent lwlock tranche

- Updated the FIXME comment about using PGPROC->lw* to not be a fixme anymore,
  it seems nobody is pushing back against that being ugly-but-reasonable for now

- I renamed ResOwnerReleaseBufferPin etc to ResOwnerReleaseBuffer, as I
  suggested nearby.

- Added a commit removing ForEachLWLockHeldByMe, now that it's not used
  anymore. I checked in with Noah, who added it, and he's on-board with that
  plan.

- Added a commit removing LWLockDisown(), LWLockReleaseDisowned(). They were
  added for AIO and AIO doesn't need them anymore, as that's implemented
  purely in bufmgr.c now. I don't see a reason to keep them...

- I reflowed the comments / README in "Require share-exclusive lock to set hint bits and to flush"
  and removed the FIXME about that

- Removed "FIXME: The start of the comment above needs updating." from the
  above commit, I already had rewritten the comment, just hadn't removed the
  FIXME yet


I tried putting the new code in a header, as we had discussed, but that turns
out to not work easily: The locking code needs access to the private-refcount
infrastructure and we can't put the private refcount infrastructure into a
header without making PrivateRef* non-static, which in turn causes slightly
worse code generation.


I'm now working on cleaning up the last two commits. The most crucial bit is
to simplify what happens in MarkSharedBufferDirtyHint(), we afaict can delete
the use of DELAY_CHKPT_START etc and just go to marking the buffer dirty first
and then do the WAL logging, just like normal WAL logging. The previous order
was only required because we were dirtying the page while holding only a
shared lock, which did not conflict with the lock held by SyncBuffers() etc.

There are some comments that arguably should be updated in 0005, but will only
be updated in 0006. I don't really see how to address that without squashing
the two commits though - which I think wouldn't be good, as the necessary
changes are decidedly nontrivial.

Greetings,

Andres Freund

Attachments:

  [text/x-diff] v11-0001-bufmgr-Change-BufferDesc.state-to-be-a-64-bit-at.patch (45.5K, ../../k5j77f3q6ztihnjnx2nqxzyor6fbj2qxcbhzuxhkh2yy63jyfg@p72phigar3n4/2-v11-0001-bufmgr-Change-BufferDesc.state-to-be-a-64-bit-at.patch)
  download | inline diff:
From 65c9d72a531b26f7461392557d354f385b0c404d Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 13 Jan 2026 20:10:32 -0500
Subject: [PATCH v11 1/7] bufmgr: Change BufferDesc.state to be a 64-bit atomic

This is motivated by wanting to merge buffer content locks into
BufferDesc.state in a future commit, rather than having a separate lwlock (see
commit c75ebc657ff for more details). As this change is rather mechanical, it
seems to make sense to split it out into a separate commit, for easier review.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  51 +++---
 src/include/storage/procnumber.h              |  14 +-
 src/backend/storage/buffer/buf_init.c         |   2 +-
 src/backend/storage/buffer/bufmgr.c           | 170 +++++++++---------
 src/backend/storage/buffer/freelist.c         |  24 +--
 src/backend/storage/buffer/localbuf.c         |  72 ++++----
 contrib/pg_buffercache/pg_buffercache_pages.c |   8 +-
 contrib/pg_prewarm/autoprewarm.c              |   2 +-
 src/test/modules/test_aio/test_aio.c          |  12 +-
 9 files changed, 179 insertions(+), 176 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 2f607ea2ac5..e6e788224f5 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -30,7 +30,7 @@
 #include "utils/resowner.h"
 
 /*
- * Buffer state is a single 32-bit variable where following data is combined.
+ * Buffer state is a single 64-bit variable where following data is combined.
  *
  * State of the buffer itself (in order):
  * - 18 bits refcount
@@ -40,6 +40,9 @@
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
+ * NB: A future commit will use a significant portion of the remaining bits to
+ * implement buffer locking as part of the state variable.
+ *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
@@ -52,27 +55,27 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 /* refcount related definitions */
 #define BUF_REFCOUNT_ONE 1
 #define BUF_REFCOUNT_MASK \
-	((1U << BUF_REFCOUNT_BITS) - 1)
+	((UINT64CONST(1) << BUF_REFCOUNT_BITS) - 1)
 
 /* usage count related definitions */
 #define BUF_USAGECOUNT_SHIFT \
 	BUF_REFCOUNT_BITS
 #define BUF_USAGECOUNT_MASK \
-	(((1U << BUF_USAGECOUNT_BITS) - 1) << (BUF_USAGECOUNT_SHIFT))
+	(((UINT64CONST(1) << BUF_USAGECOUNT_BITS) - 1) << (BUF_USAGECOUNT_SHIFT))
 #define BUF_USAGECOUNT_ONE \
-	(1U << BUF_REFCOUNT_BITS)
+	(UINT64CONST(1) << BUF_REFCOUNT_BITS)
 
 /* flags related definitions */
 #define BUF_FLAG_SHIFT \
 	(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS)
 #define BUF_FLAG_MASK \
-	(((1U << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
+	(((UINT64CONST(1) << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
 
 /* Get refcount and usagecount from buffer state */
 #define BUF_STATE_GET_REFCOUNT(state) \
-	((state) & BUF_REFCOUNT_MASK)
+	((uint32)((state) & BUF_REFCOUNT_MASK))
 #define BUF_STATE_GET_USAGECOUNT(state) \
-	(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT)
+	((uint32)(((state) & BUF_USAGECOUNT_MASK) >> BUF_USAGECOUNT_SHIFT))
 
 /*
  * Flags for buffer descriptors
@@ -82,7 +85,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  */
 
 #define BUF_DEFINE_FLAG(flagno)	\
-	(1U << (BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + (flagno)))
+	(UINT64CONST(1) << (BUF_FLAG_SHIFT + (flagno)))
 
 /* buffer header is locked */
 #define BM_LOCKED					BUF_DEFINE_FLAG( 0)
@@ -115,7 +118,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
  */
 #define BM_MAX_USAGE_COUNT	5
 
-StaticAssertDecl(BM_MAX_USAGE_COUNT < (1 << BUF_USAGECOUNT_BITS),
+StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 				 "BM_MAX_USAGE_COUNT doesn't fit in BUF_USAGECOUNT_BITS bits");
 StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
 				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
@@ -280,8 +283,8 @@ BufMappingPartitionLockByIndex(uint32 index)
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
- * atomic operations (i.e. only pg_atomic_read_u32() and
- * pg_atomic_unlocked_write_u32()).
+ * atomic operations (i.e. only pg_atomic_read_u64() and
+ * pg_atomic_unlocked_write_u64()).
  *
  * Be careful to avoid increasing the size of the struct when adding or
  * reordering members.  Keeping it below 64 bytes (the most common CPU
@@ -309,7 +312,7 @@ typedef struct BufferDesc
 	 * State of the buffer, containing flags, refcount and usagecount. See
 	 * BUF_* and BM_* defines at the top of this file.
 	 */
-	pg_atomic_uint32 state;
+	pg_atomic_uint64 state;
 
 	/*
 	 * Backend of pin-count waiter. The buffer header spinlock needs to be
@@ -415,7 +418,7 @@ BufferDescriptorGetContentLock(const BufferDesc *bdesc)
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
  */
-extern uint32 LockBufHdr(BufferDesc *desc);
+extern uint64 LockBufHdr(BufferDesc *desc);
 
 /*
  * Unlock the buffer header.
@@ -426,9 +429,9 @@ extern uint32 LockBufHdr(BufferDesc *desc);
 static inline void
 UnlockBufHdr(BufferDesc *desc)
 {
-	Assert(pg_atomic_read_u32(&desc->state) & BM_LOCKED);
+	Assert(pg_atomic_read_u64(&desc->state) & BM_LOCKED);
 
-	pg_atomic_fetch_sub_u32(&desc->state, BM_LOCKED);
+	pg_atomic_fetch_sub_u64(&desc->state, BM_LOCKED);
 }
 
 /*
@@ -439,14 +442,14 @@ UnlockBufHdr(BufferDesc *desc)
  * Note that this approach would not work for usagecount, since we need to cap
  * the usagecount at BM_MAX_USAGE_COUNT.
  */
-static inline uint32
-UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
-				uint32 set_bits, uint32 unset_bits,
+static inline uint64
+UnlockBufHdrExt(BufferDesc *desc, uint64 old_buf_state,
+				uint64 set_bits, uint64 unset_bits,
 				int refcount_change)
 {
 	for (;;)
 	{
-		uint32		buf_state = old_buf_state;
+		uint64		buf_state = old_buf_state;
 
 		Assert(buf_state & BM_LOCKED);
 
@@ -457,7 +460,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 		if (refcount_change != 0)
 			buf_state += BUF_REFCOUNT_ONE * refcount_change;
 
-		if (pg_atomic_compare_exchange_u32(&desc->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&desc->state, &old_buf_state,
 										   buf_state))
 		{
 			return old_buf_state;
@@ -465,7 +468,7 @@ UnlockBufHdrExt(BufferDesc *desc, uint32 old_buf_state,
 	}
 }
 
-extern uint32 WaitBufHdrUnlocked(BufferDesc *buf);
+extern uint64 WaitBufHdrUnlocked(BufferDesc *buf);
 
 /* in bufmgr.c */
 
@@ -525,14 +528,14 @@ extern void TrackNewBufferPin(Buffer buf);
 
 /* solely to make it easier to write tests */
 extern bool StartBufferIO(BufferDesc *buf, bool forInput, bool nowait);
-extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+extern void TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 							  bool forget_owner, bool release_aio);
 
 
 /* freelist.c */
 extern IOContext IOContextForStrategy(BufferAccessStrategy strategy);
 extern BufferDesc *StrategyGetBuffer(BufferAccessStrategy strategy,
-									 uint32 *buf_state, bool *from_ring);
+									 uint64 *buf_state, bool *from_ring);
 extern bool StrategyRejectBuffer(BufferAccessStrategy strategy,
 								 BufferDesc *buf, bool from_ring);
 
@@ -568,7 +571,7 @@ extern BlockNumber ExtendBufferedRelLocal(BufferManagerRelation bmr,
 										  uint32 *extended_by);
 extern void MarkLocalBufferDirty(Buffer buffer);
 extern void TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty,
-								   uint32 set_flag_bits, bool release_aio);
+								   uint64 set_flag_bits, bool release_aio);
 extern bool StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait);
 extern void FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln);
 extern void InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced);
diff --git a/src/include/storage/procnumber.h b/src/include/storage/procnumber.h
index 30c360ad350..bd9cb3891cc 100644
--- a/src/include/storage/procnumber.h
+++ b/src/include/storage/procnumber.h
@@ -27,13 +27,13 @@ typedef int ProcNumber;
 
 /*
  * Note: MAX_BACKENDS_BITS is 18 as that is the space available for buffer
- * refcounts in buf_internals.h.  This limitation could be lifted by using a
- * 64bit state; but it's unlikely to be worthwhile as 2^18-1 backends exceed
- * currently realistic configurations. Even if that limitation were removed,
- * we still could not a) exceed 2^23-1 because inval.c stores the ProcNumber
- * as a 3-byte signed integer, b) INT_MAX/4 because some places compute
- * 4*MaxBackends without any overflow check.  We check that the configured
- * number of backends does not exceed MAX_BACKENDS in InitializeMaxBackends().
+ * refcounts in buf_internals.h.  This limitation could be lifted, but it's
+ * unlikely to be worthwhile as 2^18-1 backends exceed currently realistic
+ * configurations. Even if that limitation were removed, we still could not a)
+ * exceed 2^23-1 because inval.c stores the ProcNumber as a 3-byte signed
+ * integer, b) INT_MAX/4 because some places compute 4*MaxBackends without any
+ * overflow check.  We check that the configured number of backends does not
+ * exceed MAX_BACKENDS in InitializeMaxBackends().
  */
 #define MAX_BACKENDS_BITS		18
 #define MAX_BACKENDS			((1U << MAX_BACKENDS_BITS)-1)
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 9a312bcc7b3..7d894522526 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -121,7 +121,7 @@ BufferManagerShmemInit(void)
 
 			ClearBufferTag(&buf->tag);
 
-			pg_atomic_init_u32(&buf->state, 0);
+			pg_atomic_init_u64(&buf->state, 0);
 			buf->wait_backend_pgprocno = INVALID_PROC_NUMBER;
 
 			buf->buf_id = i;
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index a036c2aa275..b0de8e45d4d 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -780,7 +780,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 {
 	BufferDesc *bufHdr;
 	BufferTag	tag;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(recent_buffer));
 
@@ -793,7 +793,7 @@ ReadRecentBuffer(RelFileLocator rlocator, ForkNumber forkNum, BlockNumber blockN
 		int			b = -recent_buffer - 1;
 
 		bufHdr = GetLocalBufferDescriptor(b);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		/* Is it still valid and holding the right tag? */
 		if ((buf_state & BM_VALID) && BufferTagsEqual(&tag, &bufHdr->tag))
@@ -1386,8 +1386,8 @@ StartReadBuffersImpl(ReadBuffersOperation *operation,
 				bufHdr = GetLocalBufferDescriptor(-buffers[i] - 1);
 			else
 				bufHdr = GetBufferDescriptor(buffers[i] - 1);
-			Assert(pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID);
-			found = pg_atomic_read_u32(&bufHdr->state) & BM_VALID;
+			Assert(pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID);
+			found = pg_atomic_read_u64(&bufHdr->state) & BM_VALID;
 		}
 		else
 		{
@@ -1613,10 +1613,10 @@ CheckReadBuffersOperation(ReadBuffersOperation *operation, bool is_complete)
 			GetBufferDescriptor(buffer - 1);
 
 		Assert(BufferGetBlockNumber(buffer) == operation->blocknum + i);
-		Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_TAG_VALID);
+		Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_TAG_VALID);
 
 		if (i < operation->nblocks_done)
-			Assert(pg_atomic_read_u32(&buf_hdr->state) & BM_VALID);
+			Assert(pg_atomic_read_u64(&buf_hdr->state) & BM_VALID);
 	}
 #endif
 }
@@ -2083,8 +2083,8 @@ BufferAlloc(SMgrRelation smgr, char relpersistence, ForkNumber forkNum,
 	int			existing_buf_id;
 	Buffer		victim_buffer;
 	BufferDesc *victim_buf_hdr;
-	uint32		victim_buf_state;
-	uint32		set_bits = 0;
+	uint64		victim_buf_state;
+	uint64		set_bits = 0;
 
 	/* Make sure we will have room to remember the buffer pin */
 	ResourceOwnerEnlarge(CurrentResourceOwner);
@@ -2251,7 +2251,7 @@ InvalidateBuffer(BufferDesc *buf)
 	uint32		oldHash;		/* hash value for oldTag */
 	LWLock	   *oldPartitionLock;	/* buffer partition lock for it */
 	uint32		oldFlags;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/* Save the original buffer tag before dropping the spinlock */
 	oldTag = buf->tag;
@@ -2342,7 +2342,7 @@ retry:
 static bool
 InvalidateVictimBuffer(BufferDesc *buf_hdr)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	uint32		hash;
 	LWLock	   *partition_lock;
 	BufferTag	tag;
@@ -2402,10 +2402,10 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 
 	LWLockRelease(partition_lock);
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	Assert(!(buf_state & (BM_DIRTY | BM_VALID | BM_TAG_VALID)));
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u32(&buf_hdr->state)) > 0);
+	Assert(BUF_STATE_GET_REFCOUNT(pg_atomic_read_u64(&buf_hdr->state)) > 0);
 
 	return true;
 }
@@ -2415,7 +2415,7 @@ GetVictimBuffer(BufferAccessStrategy strategy, IOContext io_context)
 {
 	BufferDesc *buf_hdr;
 	Buffer		buf;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		from_ring;
 
 	/*
@@ -2548,7 +2548,7 @@ again:
 
 	/* a final set of sanity checks */
 #ifdef USE_ASSERT_CHECKING
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	Assert(BUF_STATE_GET_REFCOUNT(buf_state) == 1);
 	Assert(!(buf_state & (BM_TAG_VALID | BM_VALID | BM_DIRTY)));
@@ -2839,13 +2839,13 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			 */
 			do
 			{
-				pg_atomic_fetch_and_u32(&existing_hdr->state, ~BM_VALID);
+				pg_atomic_fetch_and_u64(&existing_hdr->state, ~BM_VALID);
 			} while (!StartBufferIO(existing_hdr, true, false));
 		}
 		else
 		{
-			uint32		buf_state;
-			uint32		set_bits = 0;
+			uint64		buf_state;
+			uint64		set_bits = 0;
 
 			buf_state = LockBufHdr(victim_buf_hdr);
 
@@ -3021,7 +3021,7 @@ BufferIsDirty(Buffer buffer)
 		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 	}
 
-	return pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY;
+	return pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY;
 }
 
 /*
@@ -3037,8 +3037,8 @@ void
 MarkBufferDirty(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
-	uint32		old_buf_state;
+	uint64		buf_state;
+	uint64		old_buf_state;
 
 	if (!BufferIsValid(buffer))
 		elog(ERROR, "bad buffer ID: %d", buffer);
@@ -3058,7 +3058,7 @@ MarkBufferDirty(Buffer buffer)
 	 * NB: We have to wait for the buffer header spinlock to be not held, as
 	 * TerminateBufferIO() relies on the spinlock.
 	 */
-	old_buf_state = pg_atomic_read_u32(&bufHdr->state);
+	old_buf_state = pg_atomic_read_u64(&bufHdr->state);
 	for (;;)
 	{
 		if (old_buf_state & BM_LOCKED)
@@ -3069,7 +3069,7 @@ MarkBufferDirty(Buffer buffer)
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
 
-		if (pg_atomic_compare_exchange_u32(&bufHdr->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&bufHdr->state, &old_buf_state,
 										   buf_state))
 			break;
 	}
@@ -3173,10 +3173,10 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 
 	if (ref == NULL)
 	{
-		uint32		buf_state;
-		uint32		old_buf_state;
+		uint64		buf_state;
+		uint64		old_buf_state;
 
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
@@ -3210,7 +3210,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 					buf_state += BUF_USAGECOUNT_ONE;
 			}
 
-			if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+			if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 											   buf_state))
 			{
 				result = (buf_state & BM_VALID) != 0;
@@ -3237,7 +3237,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 		 * that the buffer page is legitimately non-accessible here.  We
 		 * cannot meddle with that.
 		 */
-		result = (pg_atomic_read_u32(&buf->state) & BM_VALID) != 0;
+		result = (pg_atomic_read_u64(&buf->state) & BM_VALID) != 0;
 
 		Assert(ref->data.refcount > 0);
 		ref->data.refcount++;
@@ -3272,7 +3272,7 @@ PinBuffer(BufferDesc *buf, BufferAccessStrategy strategy,
 static void
 PinBuffer_Locked(BufferDesc *buf)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	/*
 	 * As explained, We don't expect any preexisting pins. That allows us to
@@ -3284,7 +3284,7 @@ PinBuffer_Locked(BufferDesc *buf)
 	 * Since we hold the buffer spinlock, we can update the buffer state and
 	 * release the lock in one operation.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 
 	UnlockBufHdrExt(buf, old_buf_state,
 					0, 0, 1);
@@ -3314,7 +3314,7 @@ WakePinCountWaiter(BufferDesc *buf)
 	 * BM_PIN_COUNT_WAITER if it stops waiting for a reason other than this
 	 * backend waking it up.
 	 */
-	uint32		buf_state = LockBufHdr(buf);
+	uint64		buf_state = LockBufHdr(buf);
 
 	if ((buf_state & BM_PIN_COUNT_WAITER) &&
 		BUF_STATE_GET_REFCOUNT(buf_state) == 1)
@@ -3361,7 +3361,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 	ref->data.refcount--;
 	if (ref->data.refcount == 0)
 	{
-		uint32		old_buf_state;
+		uint64		old_buf_state;
 
 		/*
 		 * Mark buffer non-accessible to Valgrind.
@@ -3379,7 +3379,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
 
 		/* decrement the shared reference count */
-		old_buf_state = pg_atomic_fetch_sub_u32(&buf->state, BUF_REFCOUNT_ONE);
+		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
 
 		/* Support LockBufferForCleanup() */
 		if (old_buf_state & BM_PIN_COUNT_WAITER)
@@ -3436,7 +3436,7 @@ TrackNewBufferPin(Buffer buf)
 static void
 BufferSync(int flags)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	int			buf_id;
 	int			num_to_scan;
 	int			num_spaces;
@@ -3446,7 +3446,7 @@ BufferSync(int flags)
 	Oid			last_tsid;
 	binaryheap *ts_heap;
 	int			i;
-	uint32		mask = BM_DIRTY;
+	uint64		mask = BM_DIRTY;
 	WritebackContext wb_context;
 
 	/*
@@ -3478,7 +3478,7 @@ BufferSync(int flags)
 	for (buf_id = 0; buf_id < NBuffers; buf_id++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
-		uint32		set_bits = 0;
+		uint64		set_bits = 0;
 
 		/*
 		 * Header spinlock is enough to examine BM_DIRTY, see comment in
@@ -3645,7 +3645,7 @@ BufferSync(int flags)
 		 * write the buffer though we didn't need to.  It doesn't seem worth
 		 * guarding against this, though.
 		 */
-		if (pg_atomic_read_u32(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
+		if (pg_atomic_read_u64(&bufHdr->state) & BM_CHECKPOINT_NEEDED)
 		{
 			if (SyncOneBuffer(buf_id, false, &wb_context) & BUF_WRITTEN)
 			{
@@ -4015,7 +4015,7 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 {
 	BufferDesc *bufHdr = GetBufferDescriptor(buf_id);
 	int			result = 0;
-	uint32		buf_state;
+	uint64		buf_state;
 	BufferTag	tag;
 
 	/* Make sure we can handle the pin */
@@ -4264,7 +4264,7 @@ DebugPrintBufferRefcount(Buffer buffer)
 	int32		loccount;
 	char	   *result;
 	ProcNumber	backend;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 	if (BufferIsLocal(buffer))
@@ -4281,9 +4281,9 @@ DebugPrintBufferRefcount(Buffer buffer)
 	}
 
 	/* theoretically we should lock the bufHdr here */
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
-	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%x, refcount=%u %d)",
+	result = psprintf("[%03d] (rel=%s, blockNum=%u, flags=0x%" PRIx64 ", refcount=%u %d)",
 					  buffer,
 					  relpathbackend(BufTagGetRelFileLocator(&buf->tag), backend,
 									 BufTagGetForkNum(&buf->tag)).str,
@@ -4383,7 +4383,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	instr_time	io_start;
 	Block		bufBlock;
 	char	   *bufToWrite;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
@@ -4581,7 +4581,7 @@ BufferIsPermanent(Buffer buffer)
 	 * not random garbage.
 	 */
 	bufHdr = GetBufferDescriptor(buffer - 1);
-	return (pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT) != 0;
+	return (pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT) != 0;
 }
 
 /*
@@ -5044,11 +5044,11 @@ FlushRelationBuffers(Relation rel)
 	{
 		for (i = 0; i < NLocBuffer; i++)
 		{
-			uint32		buf_state;
+			uint64		buf_state;
 
 			bufHdr = GetLocalBufferDescriptor(i);
 			if (BufTagMatchesRelFileLocator(&bufHdr->tag, &rel->rd_locator) &&
-				((buf_state = pg_atomic_read_u32(&bufHdr->state)) &
+				((buf_state = pg_atomic_read_u64(&bufHdr->state)) &
 				 (BM_VALID | BM_DIRTY)) == (BM_VALID | BM_DIRTY))
 			{
 				ErrorContextCallback errcallback;
@@ -5084,7 +5084,7 @@ FlushRelationBuffers(Relation rel)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5156,7 +5156,7 @@ FlushRelationsAllBuffers(SMgrRelation *smgrs, int nrels)
 	{
 		SMgrSortArray *srelent = NULL;
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * As in DropRelationBuffers, an unlocked precheck should be safe and
@@ -5405,7 +5405,7 @@ FlushDatabaseBuffers(Oid dbid)
 
 	for (i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		bufHdr = GetBufferDescriptor(i);
 
@@ -5553,13 +5553,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u32(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
 		(BM_DIRTY | BM_JUST_DIRTIED))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
 		bool		delayChkptFlags = false;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * If we need to protect hint bit updates from torn writes, WAL-log a
@@ -5571,7 +5571,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
 		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u32(&bufHdr->state) & BM_PERMANENT))
+			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5671,8 +5671,8 @@ UnlockBuffers(void)
 
 	if (buf)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		buf_state = LockBufHdr(buf);
 
@@ -5803,8 +5803,8 @@ LockBufferForCleanup(Buffer buffer)
 
 	for (;;)
 	{
-		uint32		buf_state;
-		uint32		unset_bits = 0;
+		uint64		buf_state;
+		uint64		unset_bits = 0;
 
 		/* Try to acquire lock */
 		LockBuffer(buffer, BUFFER_LOCK_EXCLUSIVE);
@@ -5952,7 +5952,7 @@ bool
 ConditionalLockBufferForCleanup(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state,
+	uint64		buf_state,
 				refcount;
 
 	Assert(BufferIsValid(buffer));
@@ -6010,7 +6010,7 @@ bool
 IsBufferCleanupOK(Buffer buffer)
 {
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsValid(buffer));
 
@@ -6066,7 +6066,7 @@ WaitIO(BufferDesc *buf)
 	ConditionVariablePrepareToSleep(cv);
 	for (;;)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 		PgAioWaitRef iow;
 
 		/*
@@ -6140,7 +6140,7 @@ WaitIO(BufferDesc *buf)
 bool
 StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	ResourceOwnerEnlarge(CurrentResourceOwner);
 
@@ -6196,11 +6196,11 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
  * is being released)
  */
 void
-TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint32 set_flag_bits,
+TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 				  bool forget_owner, bool release_aio)
 {
-	uint32		buf_state;
-	uint32		unset_flag_bits = 0;
+	uint64		buf_state;
+	uint64		unset_flag_bits = 0;
 	int			refcount_change = 0;
 
 	buf_state = LockBufHdr(buf);
@@ -6261,7 +6261,7 @@ static void
 AbortBufferIO(Buffer buffer)
 {
 	BufferDesc *buf_hdr = GetBufferDescriptor(buffer - 1);
-	uint32		buf_state;
+	uint64		buf_state;
 
 	buf_state = LockBufHdr(buf_hdr);
 	Assert(buf_state & (BM_IO_IN_PROGRESS | BM_TAG_VALID));
@@ -6355,10 +6355,10 @@ rlocator_comparator(const void *p1, const void *p2)
 /*
  * Lock buffer header - set BM_LOCKED in buffer state.
  */
-uint32
+uint64
 LockBufHdr(BufferDesc *desc)
 {
-	uint32		old_buf_state;
+	uint64		old_buf_state;
 
 	Assert(!BufferIsLocal(BufferDescriptorGetBuffer(desc)));
 
@@ -6369,7 +6369,7 @@ LockBufHdr(BufferDesc *desc)
 		 * the spin-delay infrastructure. The work necessary for that shows up
 		 * in profiles and is rarely necessary.
 		 */
-		old_buf_state = pg_atomic_fetch_or_u32(&desc->state, BM_LOCKED);
+		old_buf_state = pg_atomic_fetch_or_u64(&desc->state, BM_LOCKED);
 		if (likely(!(old_buf_state & BM_LOCKED)))
 			break;				/* got lock */
 
@@ -6382,7 +6382,7 @@ LockBufHdr(BufferDesc *desc)
 			while (old_buf_state & BM_LOCKED)
 			{
 				perform_spin_delay(&delayStatus);
-				old_buf_state = pg_atomic_read_u32(&desc->state);
+				old_buf_state = pg_atomic_read_u64(&desc->state);
 			}
 			finish_spin_delay(&delayStatus);
 		}
@@ -6403,20 +6403,20 @@ LockBufHdr(BufferDesc *desc)
  * Obviously the buffer could be locked by the time the value is returned, so
  * this is primarily useful in CAS style loops.
  */
-pg_noinline uint32
+pg_noinline uint64
 WaitBufHdrUnlocked(BufferDesc *buf)
 {
 	SpinDelayStatus delayStatus;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	init_local_spin_delay(&delayStatus);
 
-	buf_state = pg_atomic_read_u32(&buf->state);
+	buf_state = pg_atomic_read_u64(&buf->state);
 
 	while (buf_state & BM_LOCKED)
 	{
 		perform_spin_delay(&delayStatus);
-		buf_state = pg_atomic_read_u32(&buf->state);
+		buf_state = pg_atomic_read_u64(&buf->state);
 	}
 
 	finish_spin_delay(&delayStatus);
@@ -6704,12 +6704,12 @@ ResOwnerPrintBufferPin(Datum res)
 static bool
 EvictUnpinnedBufferInternal(BufferDesc *desc, bool *buffer_flushed)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result;
 
 	*buffer_flushed = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -6803,12 +6803,12 @@ EvictAllUnpinnedBuffers(int32 *buffers_evicted, int32 *buffers_flushed,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -6855,7 +6855,7 @@ EvictRelUnpinnedBuffers(Relation rel, int32 *buffers_evicted,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_flushed;
 
 		CHECK_FOR_INTERRUPTS();
@@ -6897,12 +6897,12 @@ static bool
 MarkDirtyUnpinnedBufferInternal(Buffer buf, BufferDesc *desc,
 								bool *buffer_already_dirty)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		result = false;
 
 	*buffer_already_dirty = false;
 
-	buf_state = pg_atomic_read_u32(&(desc->state));
+	buf_state = pg_atomic_read_u64(&(desc->state));
 	Assert(buf_state & BM_LOCKED);
 
 	if ((buf_state & BM_VALID) == 0)
@@ -7000,7 +7000,7 @@ MarkDirtyRelUnpinnedBuffers(Relation rel,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state = pg_atomic_read_u32(&(desc->state));
+		uint64		buf_state = pg_atomic_read_u64(&(desc->state));
 		bool		buffer_already_dirty;
 
 		CHECK_FOR_INTERRUPTS();
@@ -7054,12 +7054,12 @@ MarkDirtyAllUnpinnedBuffers(int32 *buffers_dirtied,
 	for (int buf = 1; buf <= NBuffers; buf++)
 	{
 		BufferDesc *desc = GetBufferDescriptor(buf - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 		bool		buffer_already_dirty;
 
 		CHECK_FOR_INTERRUPTS();
 
-		buf_state = pg_atomic_read_u32(&desc->state);
+		buf_state = pg_atomic_read_u64(&desc->state);
 		if (!(buf_state & BM_VALID))
 			continue;
 
@@ -7110,7 +7110,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		BufferDesc *buf_hdr = is_temp ?
 			GetLocalBufferDescriptor(-buffer - 1)
 			: GetBufferDescriptor(buffer - 1);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		/*
 		 * Check that all the buffers are actually ones that could conceivably
@@ -7128,7 +7128,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		}
 
 		if (is_temp)
-			buf_state = pg_atomic_read_u32(&buf_hdr->state);
+			buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		else
 			buf_state = LockBufHdr(buf_hdr);
 
@@ -7166,7 +7166,7 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		if (is_temp)
 		{
 			buf_state += BUF_REFCOUNT_ONE;
-			pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 		}
 		else
 			UnlockBufHdrExt(buf_hdr, buf_state, 0, 0, 1);
@@ -7352,13 +7352,13 @@ buffer_readv_complete_one(PgAioTargetData *td, uint8 buf_off, Buffer buffer,
 		: GetBufferDescriptor(buffer - 1);
 	BufferTag	tag = buf_hdr->tag;
 	char	   *bufdata = BufferGetBlock(buffer);
-	uint32		set_flag_bits;
+	uint64		set_flag_bits;
 	int			piv_flags;
 
 	/* check that the buffer is in the expected state for a read */
 #ifdef USE_ASSERT_CHECKING
 	{
-		uint32		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(!(buf_state & BM_VALID));
diff --git a/src/backend/storage/buffer/freelist.c b/src/backend/storage/buffer/freelist.c
index 9a93fb335fc..b7687836188 100644
--- a/src/backend/storage/buffer/freelist.c
+++ b/src/backend/storage/buffer/freelist.c
@@ -86,7 +86,7 @@ typedef struct BufferAccessStrategyData
 
 /* Prototypes for internal functions */
 static BufferDesc *GetBufferFromRing(BufferAccessStrategy strategy,
-									 uint32 *buf_state);
+									 uint64 *buf_state);
 static void AddBufferToRing(BufferAccessStrategy strategy,
 							BufferDesc *buf);
 
@@ -171,7 +171,7 @@ ClockSweepTick(void)
  *	before returning.
  */
 BufferDesc *
-StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_ring)
+StrategyGetBuffer(BufferAccessStrategy strategy, uint64 *buf_state, bool *from_ring)
 {
 	BufferDesc *buf;
 	int			bgwprocno;
@@ -230,8 +230,8 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 	trycounter = NBuffers;
 	for (;;)
 	{
-		uint32		old_buf_state;
-		uint32		local_buf_state;
+		uint64		old_buf_state;
+		uint64		local_buf_state;
 
 		buf = GetBufferDescriptor(ClockSweepTick());
 
@@ -239,7 +239,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 		 * Check whether the buffer can be used and pin it if so. Do this
 		 * using a CAS loop, to avoid having to lock the buffer header.
 		 */
-		old_buf_state = pg_atomic_read_u32(&buf->state);
+		old_buf_state = pg_atomic_read_u64(&buf->state);
 		for (;;)
 		{
 			local_buf_state = old_buf_state;
@@ -277,7 +277,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 			{
 				local_buf_state -= BUF_USAGECOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					trycounter = NBuffers;
@@ -289,7 +289,7 @@ StrategyGetBuffer(BufferAccessStrategy strategy, uint32 *buf_state, bool *from_r
 				/* pin the buffer if the CAS succeeds */
 				local_buf_state += BUF_REFCOUNT_ONE;
 
-				if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+				if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 												   local_buf_state))
 				{
 					/* Found a usable buffer */
@@ -655,12 +655,12 @@ FreeAccessStrategy(BufferAccessStrategy strategy)
  * returning.
  */
 static BufferDesc *
-GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
+GetBufferFromRing(BufferAccessStrategy strategy, uint64 *buf_state)
 {
 	BufferDesc *buf;
 	Buffer		bufnum;
-	uint32		old_buf_state;
-	uint32		local_buf_state;	/* to avoid repeated (de-)referencing */
+	uint64		old_buf_state;
+	uint64		local_buf_state;	/* to avoid repeated (de-)referencing */
 
 
 	/* Advance to next ring slot */
@@ -682,7 +682,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 	 * Check whether the buffer can be used and pin it if so. Do this using a
 	 * CAS loop, to avoid having to lock the buffer header.
 	 */
-	old_buf_state = pg_atomic_read_u32(&buf->state);
+	old_buf_state = pg_atomic_read_u64(&buf->state);
 	for (;;)
 	{
 		local_buf_state = old_buf_state;
@@ -710,7 +710,7 @@ GetBufferFromRing(BufferAccessStrategy strategy, uint32 *buf_state)
 		/* pin the buffer if the CAS succeeds */
 		local_buf_state += BUF_REFCOUNT_ONE;
 
-		if (pg_atomic_compare_exchange_u32(&buf->state, &old_buf_state,
+		if (pg_atomic_compare_exchange_u64(&buf->state, &old_buf_state,
 										   local_buf_state))
 		{
 			*buf_state = local_buf_state;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index f6e2b1aa288..04a540379a2 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -148,7 +148,7 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 	}
 	else
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		victim_buffer = GetLocalVictimBuffer();
 		bufid = -victim_buffer - 1;
@@ -165,10 +165,10 @@ LocalBufferAlloc(SMgrRelation smgr, ForkNumber forkNum, BlockNumber blockNum,
 		 */
 		bufHdr->tag = newTag;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		buf_state &= ~(BUF_FLAG_MASK | BUF_USAGECOUNT_MASK);
 		buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 		*foundPtr = false;
 	}
@@ -245,12 +245,12 @@ GetLocalVictimBuffer(void)
 
 		if (LocalRefCount[victim_bufid] == 0)
 		{
-			uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 			if (BUF_STATE_GET_USAGECOUNT(buf_state) > 0)
 			{
 				buf_state -= BUF_USAGECOUNT_ONE;
-				pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+				pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 				trycounter = NLocBuffer;
 			}
 			else if (BUF_STATE_GET_REFCOUNT(buf_state) > 0)
@@ -286,13 +286,13 @@ GetLocalVictimBuffer(void)
 	 * this buffer is not referenced but it might still be dirty. if that's
 	 * the case, write it out before reusing it!
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_DIRTY)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY)
 		FlushLocalBuffer(bufHdr, NULL);
 
 	/*
 	 * Remove the victim buffer from the hashtable and mark as invalid.
 	 */
-	if (pg_atomic_read_u32(&bufHdr->state) & BM_TAG_VALID)
+	if (pg_atomic_read_u64(&bufHdr->state) & BM_TAG_VALID)
 	{
 		InvalidateLocalBuffer(bufHdr, false);
 
@@ -417,7 +417,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 		if (found)
 		{
 			BufferDesc *existing_hdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			UnpinLocalBuffer(BufferDescriptorGetBuffer(victim_buf_hdr));
 
@@ -428,18 +428,18 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 			/*
 			 * Clear the BM_VALID bit, do StartLocalBufferIO() and proceed.
 			 */
-			buf_state = pg_atomic_read_u32(&existing_hdr->state);
+			buf_state = pg_atomic_read_u64(&existing_hdr->state);
 			Assert(buf_state & BM_TAG_VALID);
 			Assert(!(buf_state & BM_DIRTY));
 			buf_state &= ~BM_VALID;
-			pg_atomic_unlocked_write_u32(&existing_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&existing_hdr->state, buf_state);
 
 			/* no need to loop for local buffers */
 			StartLocalBufferIO(existing_hdr, true, false);
 		}
 		else
 		{
-			uint32		buf_state = pg_atomic_read_u32(&victim_buf_hdr->state);
+			uint64		buf_state = pg_atomic_read_u64(&victim_buf_hdr->state);
 
 			Assert(!(buf_state & (BM_VALID | BM_TAG_VALID | BM_DIRTY | BM_JUST_DIRTIED)));
 
@@ -447,7 +447,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 
 			buf_state |= BM_TAG_VALID | BUF_USAGECOUNT_ONE;
 
-			pg_atomic_unlocked_write_u32(&victim_buf_hdr->state, buf_state);
+			pg_atomic_unlocked_write_u64(&victim_buf_hdr->state, buf_state);
 
 			hresult->id = victim_buf_id;
 
@@ -467,13 +467,13 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 	{
 		Buffer		buf = buffers[i];
 		BufferDesc *buf_hdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		buf_state |= BM_VALID;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 
 	*extended_by = extend_by;
@@ -492,7 +492,7 @@ MarkLocalBufferDirty(Buffer buffer)
 {
 	int			bufid;
 	BufferDesc *bufHdr;
-	uint32		buf_state;
+	uint64		buf_state;
 
 	Assert(BufferIsLocal(buffer));
 
@@ -506,14 +506,14 @@ MarkLocalBufferDirty(Buffer buffer)
 
 	bufHdr = GetLocalBufferDescriptor(bufid);
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	if (!(buf_state & BM_DIRTY))
 		pgBufferUsage.local_blks_dirtied++;
 
 	buf_state |= BM_DIRTY;
 
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -522,7 +522,7 @@ MarkLocalBufferDirty(Buffer buffer)
 bool
 StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 
 	/*
 	 * With AIO the buffer could have IO in progress, e.g. when there are two
@@ -542,7 +542,7 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
 	/* Once we get here, there is definitely no I/O active on this buffer */
 
 	/* Check if someone else already did the I/O */
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 	if (forInput ? (buf_state & BM_VALID) : !(buf_state & BM_DIRTY))
 	{
 		return false;
@@ -559,11 +559,11 @@ StartLocalBufferIO(BufferDesc *bufHdr, bool forInput, bool nowait)
  * Like TerminateBufferIO, but for local buffers
  */
 void
-TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bits,
+TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint64 set_flag_bits,
 					   bool release_aio)
 {
 	/* Only need to adjust flags */
-	uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+	uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/* BM_IO_IN_PROGRESS isn't currently used for local buffers */
 
@@ -582,7 +582,7 @@ TerminateLocalBufferIO(BufferDesc *bufHdr, bool clear_dirty, uint32 set_flag_bit
 	}
 
 	buf_state |= set_flag_bits;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 
 	/* local buffers don't track IO using resowners */
 
@@ -606,7 +606,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(bufHdr);
 	int			bufid = -buffer - 1;
-	uint32		buf_state;
+	uint64		buf_state;
 	LocalBufferLookupEnt *hresult;
 
 	/*
@@ -622,7 +622,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 		Assert(!pgaio_wref_valid(&bufHdr->io_wref));
 	}
 
-	buf_state = pg_atomic_read_u32(&bufHdr->state);
+	buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 	/*
 	 * We need to test not just LocalRefCount[bufid] but also the BufferDesc
@@ -647,7 +647,7 @@ InvalidateLocalBuffer(BufferDesc *bufHdr, bool check_unreferenced)
 	ClearBufferTag(&bufHdr->tag);
 	buf_state &= ~BUF_FLAG_MASK;
 	buf_state &= ~BUF_USAGECOUNT_MASK;
-	pg_atomic_unlocked_write_u32(&bufHdr->state, buf_state);
+	pg_atomic_unlocked_write_u64(&bufHdr->state, buf_state);
 }
 
 /*
@@ -671,9 +671,9 @@ DropRelationLocalBuffers(RelFileLocator rlocator, ForkNumber *forkNum,
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (!(buf_state & BM_TAG_VALID) ||
 			!BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -706,9 +706,9 @@ DropRelationAllLocalBuffers(RelFileLocator rlocator)
 	for (i = 0; i < NLocBuffer; i++)
 	{
 		BufferDesc *bufHdr = GetLocalBufferDescriptor(i);
-		uint32		buf_state;
+		uint64		buf_state;
 
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if ((buf_state & BM_TAG_VALID) &&
 			BufTagMatchesRelFileLocator(&bufHdr->tag, &rlocator))
@@ -804,11 +804,11 @@ InitLocalBuffers(void)
 bool
 PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 {
-	uint32		buf_state;
+	uint64		buf_state;
 	Buffer		buffer = BufferDescriptorGetBuffer(buf_hdr);
 	int			bufid = -buffer - 1;
 
-	buf_state = pg_atomic_read_u32(&buf_hdr->state);
+	buf_state = pg_atomic_read_u64(&buf_hdr->state);
 
 	if (LocalRefCount[bufid] == 0)
 	{
@@ -819,7 +819,7 @@ PinLocalBuffer(BufferDesc *buf_hdr, bool adjust_usagecount)
 		{
 			buf_state += BUF_USAGECOUNT_ONE;
 		}
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/*
 		 * See comment in PinBuffer().
@@ -856,14 +856,14 @@ UnpinLocalBufferNoOwner(Buffer buffer)
 	if (--LocalRefCount[buffid] == 0)
 	{
 		BufferDesc *buf_hdr = GetLocalBufferDescriptor(buffid);
-		uint32		buf_state;
+		uint64		buf_state;
 
 		NLocalPinnedBuffers--;
 
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 		buf_state -= BUF_REFCOUNT_ONE;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 
 		/* see comment in UnpinBufferNoOwner */
 		VALGRIND_MAKE_MEM_NOACCESS(LocalBufHdrGetBlock(buf_hdr), BLCKSZ);
diff --git a/contrib/pg_buffercache/pg_buffercache_pages.c b/contrib/pg_buffercache/pg_buffercache_pages.c
index b682dca658b..dcba3fb5473 100644
--- a/contrib/pg_buffercache/pg_buffercache_pages.c
+++ b/contrib/pg_buffercache/pg_buffercache_pages.c
@@ -199,7 +199,7 @@ pg_buffercache_pages(PG_FUNCTION_ARGS)
 		for (i = 0; i < NBuffers; i++)
 		{
 			BufferDesc *bufHdr;
-			uint32		buf_state;
+			uint64		buf_state;
 
 			CHECK_FOR_INTERRUPTS();
 
@@ -615,7 +615,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr;
-		uint32		buf_state;
+		uint64		buf_state;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -626,7 +626,7 @@ pg_buffercache_summary(PG_FUNCTION_ARGS)
 		 * noticeably increase the cost of the function.
 		 */
 		bufHdr = GetBufferDescriptor(i);
-		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		buf_state = pg_atomic_read_u64(&bufHdr->state);
 
 		if (buf_state & BM_VALID)
 		{
@@ -676,7 +676,7 @@ pg_buffercache_usage_counts(PG_FUNCTION_ARGS)
 	for (int i = 0; i < NBuffers; i++)
 	{
 		BufferDesc *bufHdr = GetBufferDescriptor(i);
-		uint32		buf_state = pg_atomic_read_u32(&bufHdr->state);
+		uint64		buf_state = pg_atomic_read_u64(&bufHdr->state);
 		int			usage_count;
 
 		CHECK_FOR_INTERRUPTS();
diff --git a/contrib/pg_prewarm/autoprewarm.c b/contrib/pg_prewarm/autoprewarm.c
index 3ca7d2ed772..89e187425cc 100644
--- a/contrib/pg_prewarm/autoprewarm.c
+++ b/contrib/pg_prewarm/autoprewarm.c
@@ -703,7 +703,7 @@ apw_dump_now(bool is_bgworker, bool dump_unlogged)
 
 	for (num_blocks = 0, i = 0; i < NBuffers; i++)
 	{
-		uint32		buf_state;
+		uint64		buf_state;
 
 		CHECK_FOR_INTERRUPTS();
 
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index e046b08f3d5..b1aa8af9ec0 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -308,9 +308,9 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 {
 	Buffer		buf;
 	BufferDesc *buf_hdr;
-	uint32		buf_state;
+	uint64		buf_state;
 	bool		was_pinned = false;
-	uint32		unset_bits = 0;
+	uint64		unset_bits = 0;
 
 	/* place buffer in shared buffers without erroring out */
 	buf = ReadBufferExtended(rel, MAIN_FORKNUM, blkno, RBM_ZERO_AND_LOCK, NULL);
@@ -319,7 +319,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_hdr = GetLocalBufferDescriptor(-buf - 1);
-		buf_state = pg_atomic_read_u32(&buf_hdr->state);
+		buf_state = pg_atomic_read_u64(&buf_hdr->state);
 	}
 	else
 	{
@@ -340,7 +340,7 @@ create_toy_buffer(Relation rel, BlockNumber blkno)
 	if (RelationUsesLocalBuffers(rel))
 	{
 		buf_state &= ~unset_bits;
-		pg_atomic_unlocked_write_u32(&buf_hdr->state, buf_state);
+		pg_atomic_unlocked_write_u64(&buf_hdr->state, buf_state);
 	}
 	else
 	{
@@ -489,7 +489,7 @@ invalidate_rel_block(PG_FUNCTION_ARGS)
 
 			LockBuffer(buf, BUFFER_LOCK_EXCLUSIVE);
 
-			if (pg_atomic_read_u32(&buf_hdr->state) & BM_DIRTY)
+			if (pg_atomic_read_u64(&buf_hdr->state) & BM_DIRTY)
 			{
 				if (BufferIsLocal(buf))
 					FlushLocalBuffer(buf_hdr, NULL);
@@ -572,7 +572,7 @@ buffer_call_terminate_io(PG_FUNCTION_ARGS)
 	bool		io_error = PG_GETARG_BOOL(3);
 	bool		release_aio = PG_GETARG_BOOL(4);
 	bool		clear_dirty = false;
-	uint32		set_flag_bits = 0;
+	uint64		set_flag_bits = 0;
 
 	if (io_error)
 		set_flag_bits |= BM_IO_ERROR;
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v11-0002-bufmgr-Implement-buffer-content-locks-independen.patch (51.1K, ../../k5j77f3q6ztihnjnx2nqxzyor6fbj2qxcbhzuxhkh2yy63jyfg@p72phigar3n4/3-v11-0002-bufmgr-Implement-buffer-content-locks-independen.patch)
  download | inline diff:
From 2702c1de6afbc03eed4254482b82dfbb0299c821 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 13 Jan 2026 20:10:32 -0500
Subject: [PATCH v11 2/7] bufmgr: Implement buffer content locks independently
 of lwlocks

Until now buffer content locks were implemented using lwlocks. That has the
obvious advantage of not needing a separate efficient implementation of
locks. However, the time for a dedicated buffer content lock implementation
has come:

1) Hint bits are currently set while holding only a share lock. This leads to
   having to copy pages while they are being written out if checksums are
   enabled, which is not cheap. We would like to add AIO writes, however once
   many buffers can be written out at the same time, it gets a lot more
   expensive to copy them, particularly because that copy needs to reside in
   shared buffers (for worker mode to have access to the buffer).

   In addition, modifying buffers while they are being written out can cause
   issues with unbuffered/direct-IO, as some filesystems (like btrfs) do not
   like that, due to filesystem internal checksums getting corrupted.

   The solution to this is to require a new share-exclusive lock-level to set
   hint bits and to write out buffers, making those operations mutually
   exclusive. We could introduce such a lock-level into the generic lwlock
   implementation, however it does not look like there would be other users,
   and it does add some overhead into important code paths.

2) For AIO writes we need to be able to race-freely check whether a buffer is
   undergoing IO and whether an exclusive lock on the page can be acquired. That
   is rather hard to do efficiently when the buffer state and the lock state
   are separate atomic variables. This is a major hindrance to allowing writes
   to be done asynchronously.

3) Buffer locks are by far the most frequently taken locks. Optimizing them
   specifically for their use case is worth the effort. E.g. by merging
   content locks into buffer locks we will be able to release a buffer lock
   and pin in one atomic operation.

4) There are more complicated optimizations, like long-lived "super pinned &
   locked" pages, that cannot realistically be implemented with the generic
   lwlock implementation.

Therefore implement content locks inside bufmgr.c. The lockstate is stored as
part of BufferDesc.state. The implementation of buffer content locks is fairly
similar to lwlocks, with a few important differences:

1) An additional lock-level share-exclusive has been added. This lock-level
   conflicts with exclusive locks and itself, but not share locks.

2) Error recovery for content locks is implemented as part of the already
   existing private-refcount tracking mechanism in combination with resowners,
   instead of a bespoke mechanism as the case for lwlocks. This means we do
   not need to add dedicated error-recovery code paths to release all content
   locks (like done with LWLockReleaseAll() for lwlocks).

3) The lock state is embedded in BufferDesc.state instead of having its own
   struct.

4) The wakeup logic is a tad more complicated due to needing to support the
   additional lock-level

This commit unfortunately introduces some code that is very similar to the
code in lwlock.c, however the code is not equivalent enough to easily merge
it. The future wins that this commit makes possible seem worth the cost.

As of this commit nothing uses the new share-exclusive lock mode. It will be
used in a future commit. It seemed too complicated to introduce the lock-level
in a separate commit.

It's worth calling out one wart in this commit: Despite content locks not
being lwlocks anymore, they continue to use PGPROC->lw* - that seemed better
than duplicating the relevant infrastructure.

Another thing worth pointing out is that, after this change, content locks are
not reported as LWLock wait events anymore, but as new wait events in the
"Buffer" wait event class (see also 6c5c393b740). The old BufferContent lwlock
tranche has been removed.

Reviewed-by: Melanie Plageman <melanieplageman@gmail.com>
Reviewed-by: Heikki Linnakangas <heikki.linnakangas@iki.fi>
Reviewed-by: Greg Burd <greg@burd.me>
Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
---
 src/include/storage/buf_internals.h           |  73 +-
 src/include/storage/bufmgr.h                  |  32 +-
 src/include/storage/lwlocklist.h              |   1 -
 src/include/storage/proc.h                    |   8 +-
 src/backend/storage/buffer/buf_init.c         |   5 +-
 src/backend/storage/buffer/bufmgr.c           | 916 ++++++++++++++++--
 .../utils/activity/wait_event_names.txt       |   4 +-
 7 files changed, 934 insertions(+), 105 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index e6e788224f5..27f12502d19 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -23,6 +23,7 @@
 #include "storage/condition_variable.h"
 #include "storage/lwlock.h"
 #include "storage/procnumber.h"
+#include "storage/proclist_types.h"
 #include "storage/shmem.h"
 #include "storage/smgr.h"
 #include "storage/spin.h"
@@ -35,22 +36,23 @@
  * State of the buffer itself (in order):
  * - 18 bits refcount
  * - 4 bits usage count
- * - 10 bits of flags
+ * - 12 bits of flags
+ * - 18 bits share-lock count
+ * - 1 bit share-exclusive locked
+ * - 1 bit exclusive locked
  *
  * Combining these values allows to perform some operations without locking
  * the buffer header, by modifying them together with a CAS loop.
  *
- * NB: A future commit will use a significant portion of the remaining bits to
- * implement buffer locking as part of the state variable.
- *
  * The definition of buffer state components is below.
  */
 #define BUF_REFCOUNT_BITS 18
 #define BUF_USAGECOUNT_BITS 4
-#define BUF_FLAG_BITS 10
+#define BUF_FLAG_BITS 12
+#define BUF_LOCK_BITS (18+2)
 
-StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
-				 "parts of buffer state space need to equal 32");
+StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS + BUF_LOCK_BITS <= 64,
+				 "parts of buffer state space need to be <= 64");
 
 /* refcount related definitions */
 #define BUF_REFCOUNT_ONE 1
@@ -71,6 +73,19 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 #define BUF_FLAG_MASK \
 	(((UINT64CONST(1) << BUF_FLAG_BITS) - 1) << BUF_FLAG_SHIFT)
 
+/* lock state related definitions */
+#define BM_LOCK_SHIFT \
+	(BUF_FLAG_SHIFT + BUF_FLAG_BITS)
+#define BM_LOCK_VAL_SHARED \
+	(UINT64CONST(1) << (BM_LOCK_SHIFT))
+#define BM_LOCK_VAL_SHARE_EXCLUSIVE \
+	(UINT64CONST(1) << (BM_LOCK_SHIFT + MAX_BACKENDS_BITS))
+#define BM_LOCK_VAL_EXCLUSIVE \
+	(UINT64CONST(1) << (BM_LOCK_SHIFT + MAX_BACKENDS_BITS + 1))
+#define BM_LOCK_MASK \
+	((((uint64) MAX_BACKENDS) << BM_LOCK_SHIFT) | BM_LOCK_VAL_SHARE_EXCLUSIVE | BM_LOCK_VAL_EXCLUSIVE)
+
+
 /* Get refcount and usagecount from buffer state */
 #define BUF_STATE_GET_REFCOUNT(state) \
 	((uint32)((state) & BUF_REFCOUNT_MASK))
@@ -107,6 +122,17 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 #define BM_CHECKPOINT_NEEDED		BUF_DEFINE_FLAG( 8)
 /* permanent buffer (not unlogged, or init fork) */
 #define BM_PERMANENT				BUF_DEFINE_FLAG( 9)
+/* content lock has waiters */
+#define BM_LOCK_HAS_WAITERS			BUF_DEFINE_FLAG(10)
+/* waiter for content lock has been signalled but not yet run */
+#define BM_LOCK_WAKE_IN_PROGRESS	BUF_DEFINE_FLAG(11)
+
+
+StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
+				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
+StaticAssertDecl(MAX_BACKENDS_BITS <= (BUF_LOCK_BITS - 2),
+				 "MAX_BACKENDS_BITS needs to be <= BUF_LOCK_BITS - 2");
+
 
 /*
  * The maximum allowed value of usage_count represents a tradeoff between
@@ -120,8 +146,6 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS == 32,
 
 StaticAssertDecl(BM_MAX_USAGE_COUNT < (UINT64CONST(1) << BUF_USAGECOUNT_BITS),
 				 "BM_MAX_USAGE_COUNT doesn't fit in BUF_USAGECOUNT_BITS bits");
-StaticAssertDecl(MAX_BACKENDS_BITS <= BUF_REFCOUNT_BITS,
-				 "MAX_BACKENDS_BITS needs to be <= BUF_REFCOUNT_BITS");
 
 /*
  * Buffer tag identifies which disk block the buffer contains.
@@ -265,9 +289,6 @@ BufMappingPartitionLockByIndex(uint32 index)
  * it is held.  However, existing buffer pins may be released while the buffer
  * header spinlock is held, using an atomic subtraction.
  *
- * The LWLock can take care of itself.  The buffer header lock is *not* used
- * to control access to the data in the buffer!
- *
  * If we have the buffer pinned, its tag can't change underneath us, so we can
  * examine the tag without locking the buffer header.  Also, in places we do
  * one-time reads of the flags without bothering to lock the buffer header;
@@ -280,6 +301,15 @@ BufMappingPartitionLockByIndex(uint32 index)
  * wait_backend_pgprocno and setting flag bit BM_PIN_COUNT_WAITER.  At present,
  * there can be only one such waiter per buffer.
  *
+ * The content of buffers is protected via the buffer content lock,
+ * implemented as part of the buffer state. Note that the buffer header lock
+ * is *not* used to control access to the data in the buffer! We used to use
+ * an LWLock to implement the content lock, but having a dedicated
+ * implementation of content locks allows us to implement some otherwise hard
+ * things (e.g. race-freely checking if AIO is in progress before locking a
+ * buffer exclusively) and enables otherwise impossible optimizations
+ * (e.g. unlocking and unpinning a buffer in one atomic operation).
+ *
  * We use this same struct for local buffer headers, but the locks are not
  * used and not all of the flag bits are useful either. To avoid unnecessary
  * overhead, manipulations of the state field should be done without actual
@@ -321,7 +351,12 @@ typedef struct BufferDesc
 	int			wait_backend_pgprocno;
 
 	PgAioWaitRef io_wref;		/* set iff AIO is in progress */
-	LWLock		content_lock;	/* to lock access to buffer contents */
+
+	/*
+	 * List of PGPROCs waiting for the buffer content lock. Protected by the
+	 * buffer header spinlock.
+	 */
+	proclist_head lock_waiters;
 } BufferDesc;
 
 /*
@@ -408,12 +443,6 @@ BufferDescriptorGetIOCV(const BufferDesc *bdesc)
 	return &(BufferIOCVArray[bdesc->buf_id]).cv;
 }
 
-static inline LWLock *
-BufferDescriptorGetContentLock(const BufferDesc *bdesc)
-{
-	return (LWLock *) (&bdesc->content_lock);
-}
-
 /*
  * Functions for acquiring/releasing a shared buffer header's spinlock.  Do
  * not apply these to local buffers!
@@ -491,18 +520,18 @@ extern PGDLLIMPORT CkptSortItem *CkptBufferIds;
 
 /* ResourceOwner callbacks to hold buffer I/Os and pins */
 extern PGDLLIMPORT const ResourceOwnerDesc buffer_io_resowner_desc;
-extern PGDLLIMPORT const ResourceOwnerDesc buffer_pin_resowner_desc;
+extern PGDLLIMPORT const ResourceOwnerDesc buffer_resowner_desc;
 
 /* Convenience wrappers over ResourceOwnerRemember/Forget */
 static inline void
 ResourceOwnerRememberBuffer(ResourceOwner owner, Buffer buffer)
 {
-	ResourceOwnerRemember(owner, Int32GetDatum(buffer), &buffer_pin_resowner_desc);
+	ResourceOwnerRemember(owner, Int32GetDatum(buffer), &buffer_resowner_desc);
 }
 static inline void
 ResourceOwnerForgetBuffer(ResourceOwner owner, Buffer buffer)
 {
-	ResourceOwnerForget(owner, Int32GetDatum(buffer), &buffer_pin_resowner_desc);
+	ResourceOwnerForget(owner, Int32GetDatum(buffer), &buffer_resowner_desc);
 }
 static inline void
 ResourceOwnerRememberBufferIO(ResourceOwner owner, Buffer buffer)
diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index 715ae96f0f0..a40adf6b2a8 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -203,7 +203,20 @@ extern PGDLLIMPORT int32 *LocalRefCount;
 typedef enum BufferLockMode
 {
 	BUFFER_LOCK_UNLOCK,
+
+	/*
+	 * A share lock conflicts with exclusive locks.
+	 */
 	BUFFER_LOCK_SHARE,
+
+	/*
+	 * A share-exclusive lock conflicts with itself and exclusive locks.
+	 */
+	BUFFER_LOCK_SHARE_EXCLUSIVE,
+
+	/*
+	 * An exclusive lock conflicts with every other lock type.
+	 */
 	BUFFER_LOCK_EXCLUSIVE,
 } BufferLockMode;
 
@@ -302,7 +315,24 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
 extern void UnlockBuffers(void);
-extern void LockBuffer(Buffer buffer, BufferLockMode mode);
+extern void UnlockBuffer(Buffer buffer);
+extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
+
+/*
+ * Handling BUFFER_LOCK_UNLOCK in bufmgr.c leads to sufficiently worse branch
+ * prediction to impact performance. Therefore handle that switch here, where
+ * most of the time `mode` will be a constant and thus can be optimized out by
+ * the compiler.
+ */
+static inline void
+LockBuffer(Buffer buffer, BufferLockMode mode)
+{
+	if (mode == BUFFER_LOCK_UNLOCK)
+		UnlockBuffer(buffer);
+	else
+		LockBufferInternal(buffer, mode);
+}
+
 extern bool ConditionalLockBuffer(Buffer buffer);
 extern void LockBufferForCleanup(Buffer buffer);
 extern bool ConditionalLockBufferForCleanup(Buffer buffer);
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index 94f818b9f10..28c8c95c3f4 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -104,7 +104,6 @@ PG_LWLOCKTRANCHE(MULTIXACTMEMBER_BUFFER, MultiXactMemberBuffer)
 PG_LWLOCKTRANCHE(NOTIFY_BUFFER, NotifyBuffer)
 PG_LWLOCKTRANCHE(SERIAL_BUFFER, SerialBuffer)
 PG_LWLOCKTRANCHE(WAL_INSERT, WALInsert)
-PG_LWLOCKTRANCHE(BUFFER_CONTENT, BufferContent)
 PG_LWLOCKTRANCHE(REPLICATION_ORIGIN_STATE, ReplicationOriginState)
 PG_LWLOCKTRANCHE(REPLICATION_SLOT_IO, ReplicationSlotIO)
 PG_LWLOCKTRANCHE(LOCK_FASTPATH, LockFastPath)
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index de7b2e0bd2c..039bc8353be 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -242,7 +242,13 @@ struct PGPROC
 	 */
 	bool		recoveryConflictPending;
 
-	/* Info about LWLock the process is currently waiting for, if any. */
+	/*
+	 * Info about LWLock the process is currently waiting for, if any.
+	 *
+	 * This is currently used both for lwlocks and buffer content locks, which
+	 * is acceptable, although not pretty, because a backend can't wait for
+	 * both types of locks at the same time.
+	 */
 	uint8		lwWaiting;		/* see LWLockWaitState */
 	uint8		lwWaitMode;		/* lwlock mode being waited for */
 	proclist_node lwWaitLink;	/* position in LW lock wait list */
diff --git a/src/backend/storage/buffer/buf_init.c b/src/backend/storage/buffer/buf_init.c
index 7d894522526..c0c223b2e32 100644
--- a/src/backend/storage/buffer/buf_init.c
+++ b/src/backend/storage/buffer/buf_init.c
@@ -17,6 +17,7 @@
 #include "storage/aio.h"
 #include "storage/buf_internals.h"
 #include "storage/bufmgr.h"
+#include "storage/proclist.h"
 
 BufferDescPadded *BufferDescriptors;
 char	   *BufferBlocks;
@@ -128,9 +129,7 @@ BufferManagerShmemInit(void)
 
 			pgaio_wref_clear(&buf->io_wref);
 
-			LWLockInitialize(BufferDescriptorGetContentLock(buf),
-							 LWTRANCHE_BUFFER_CONTENT);
-
+			proclist_init(&buf->lock_waiters);
 			ConditionVariableInit(BufferDescriptorGetIOCV(buf));
 		}
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index b0de8e45d4d..6adf04903cb 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -58,6 +58,7 @@
 #include "storage/ipc.h"
 #include "storage/lmgr.h"
 #include "storage/proc.h"
+#include "storage/proclist.h"
 #include "storage/read_stream.h"
 #include "storage/smgr.h"
 #include "storage/standby.h"
@@ -100,6 +101,12 @@ typedef struct PrivateRefCountData
 	 * How many times has the buffer been pinned by this backend.
 	 */
 	int32		refcount;
+
+	/*
+	 * Is the buffer locked by this backend? BUFFER_LOCK_UNLOCK indicates that
+	 * the buffer is not locked.
+	 */
+	BufferLockMode lockmode;
 } PrivateRefCountData;
 
 typedef struct PrivateRefCountEntry
@@ -210,8 +217,10 @@ static BufferDesc *PinCountWaitBuf = NULL;
  * Each buffer also has a private refcount that keeps track of the number of
  * times the buffer is pinned in the current process.  This is so that the
  * shared refcount needs to be modified only once if a buffer is pinned more
- * than once by an individual backend.  It's also used to check that no buffers
- * are still pinned at the end of transactions and when exiting.
+ * than once by an individual backend.  It's also used to check that no
+ * buffers are still pinned at the end of transactions and when exiting. We
+ * also use this mechanism to track whether this backend has a buffer locked,
+ * and, if so, in what mode.
  *
  *
  * To avoid - as we used to - requiring an array with NBuffers entries to keep
@@ -254,8 +263,8 @@ static void ForgetPrivateRefCountEntry(PrivateRefCountEntry *ref);
 /* ResourceOwner callbacks to hold in-progress I/Os and buffer pins */
 static void ResOwnerReleaseBufferIO(Datum res);
 static char *ResOwnerPrintBufferIO(Datum res);
-static void ResOwnerReleaseBufferPin(Datum res);
-static char *ResOwnerPrintBufferPin(Datum res);
+static void ResOwnerReleaseBuffer(Datum res);
+static char *ResOwnerPrintBuffer(Datum res);
 
 const ResourceOwnerDesc buffer_io_resowner_desc =
 {
@@ -266,13 +275,13 @@ const ResourceOwnerDesc buffer_io_resowner_desc =
 	.DebugPrint = ResOwnerPrintBufferIO
 };
 
-const ResourceOwnerDesc buffer_pin_resowner_desc =
+const ResourceOwnerDesc buffer_resowner_desc =
 {
-	.name = "buffer pin",
+	.name = "buffer",
 	.release_phase = RESOURCE_RELEASE_BEFORE_LOCKS,
 	.release_priority = RELEASE_PRIO_BUFFER_PINS,
-	.ReleaseResource = ResOwnerReleaseBufferPin,
-	.DebugPrint = ResOwnerPrintBufferPin
+	.ReleaseResource = ResOwnerReleaseBuffer,
+	.DebugPrint = ResOwnerPrintBuffer
 };
 
 /*
@@ -351,6 +360,7 @@ ReservePrivateRefCountEntry(void)
 		/* clear the whole data member, just for future proofing */
 		memset(&victim_entry->data, 0, sizeof(victim_entry->data));
 		victim_entry->data.refcount = 0;
+		victim_entry->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 		PrivateRefCountOverflowed++;
 	}
@@ -374,6 +384,7 @@ NewPrivateRefCountEntry(Buffer buffer)
 	PrivateRefCountArrayKeys[ReservedRefCountSlot] = buffer;
 	res->buffer = buffer;
 	res->data.refcount = 0;
+	res->data.lockmode = BUFFER_LOCK_UNLOCK;
 
 	/* update cache for the next lookup */
 	PrivateRefCountEntryLast = ReservedRefCountSlot;
@@ -540,6 +551,7 @@ static void
 ForgetPrivateRefCountEntry(PrivateRefCountEntry *ref)
 {
 	Assert(ref->data.refcount == 0);
+	Assert(ref->data.lockmode == BUFFER_LOCK_UNLOCK);
 
 	if (ref >= &PrivateRefCountArray[0] &&
 		ref < &PrivateRefCountArray[REFCOUNT_ARRAY_ENTRIES])
@@ -641,14 +653,27 @@ static void RelationCopyStorageUsingBuffer(RelFileLocator srclocator,
 static void AtProcExit_Buffers(int code, Datum arg);
 static void CheckForBufferLeaks(void);
 #ifdef USE_ASSERT_CHECKING
-static void AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
-									   void *unused_context);
+static void AssertNotCatalogBufferLock(Buffer buffer, BufferLockMode mode);
 #endif
 static int	rlocator_comparator(const void *p1, const void *p2);
 static inline int buffertag_comparator(const BufferTag *ba, const BufferTag *bb);
 static inline int ckpt_buforder_comparator(const CkptSortItem *a, const CkptSortItem *b);
 static int	ts_ckpt_progress_comparator(Datum a, Datum b, void *arg);
 
+static void BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr);
+static bool BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode);
+static bool BufferLockHeldByMe(BufferDesc *buf_hdr);
+static inline void BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr);
+static inline int BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr);
+static inline bool BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode);
+static void BufferLockDequeueSelf(BufferDesc *buf_hdr);
+static void BufferLockWakeup(BufferDesc *buf_hdr, bool unlocked);
+static void BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate);
+static inline uint64 BufferLockReleaseSub(BufferLockMode mode);
+
 
 /*
  * Implementation of PrefetchBuffer() for shared buffers.
@@ -2306,6 +2331,12 @@ retry:
 		goto retry;
 	}
 
+	/*
+	 * An invalidated buffer should not have any backends waiting to lock the
+	 * buffer, therefore BM_LOCK_WAKE_IN_PROGRESS should not be set.
+	 */
+	Assert(!(buf_state & BM_LOCK_WAKE_IN_PROGRESS));
+
 	/*
 	 * Clear out the buffer's tag and flags.  We must do this to ensure that
 	 * linear scans of the buffer array don't think the buffer is valid.
@@ -2382,6 +2413,12 @@ InvalidateVictimBuffer(BufferDesc *buf_hdr)
 		return false;
 	}
 
+	/*
+	 * An invalidated buffer should not have any backends waiting to lock the
+	 * buffer, therefore BM_LOCK_WAKE_IN_PROGRESS should not be set.
+	 */
+	Assert(!(buf_state & BM_LOCK_WAKE_IN_PROGRESS));
+
 	/*
 	 * Clear out the buffer's tag and flags and usagecount.  This is not
 	 * strictly required, as BM_TAG_VALID/BM_VALID needs to be checked before
@@ -2449,8 +2486,6 @@ again:
 	 */
 	if (buf_state & BM_DIRTY)
 	{
-		LWLock	   *content_lock;
-
 		Assert(buf_state & BM_TAG_VALID);
 		Assert(buf_state & BM_VALID);
 
@@ -2468,8 +2503,7 @@ again:
 		 * one just happens to be trying to split the page the first one got
 		 * from StrategyGetBuffer.)
 		 */
-		content_lock = BufferDescriptorGetContentLock(buf_hdr);
-		if (!LWLockConditionalAcquire(content_lock, LW_SHARED))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -2498,7 +2532,7 @@ again:
 			if (XLogNeedsFlush(lsn)
 				&& StrategyRejectBuffer(strategy, buf_hdr, from_ring))
 			{
-				LWLockRelease(content_lock);
+				LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 				UnpinBuffer(buf_hdr);
 				goto again;
 			}
@@ -2506,7 +2540,7 @@ again:
 
 		/* OK, do the I/O */
 		FlushBuffer(buf_hdr, NULL, IOOBJECT_RELATION, io_context);
-		LWLockRelease(content_lock);
+		LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 		ScheduleBufferTagForWriteback(&BackendWritebackContext, io_context,
 									  &buf_hdr->tag);
@@ -2948,7 +2982,7 @@ BufferIsLockedByMe(Buffer buffer)
 	else
 	{
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMe(BufferDescriptorGetContentLock(bufHdr));
+		return BufferLockHeldByMe(bufHdr);
 	}
 }
 
@@ -2973,23 +3007,8 @@ BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode)
 	}
 	else
 	{
-		LWLockMode	lw_mode;
-
-		switch (mode)
-		{
-			case BUFFER_LOCK_EXCLUSIVE:
-				lw_mode = LW_EXCLUSIVE;
-				break;
-			case BUFFER_LOCK_SHARE:
-				lw_mode = LW_SHARED;
-				break;
-			default:
-				pg_unreachable();
-		}
-
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		return LWLockHeldByMeInMode(BufferDescriptorGetContentLock(bufHdr),
-									lw_mode);
+		return BufferLockHeldByMeInMode(bufHdr, mode);
 	}
 }
 
@@ -3376,7 +3395,7 @@ UnpinBufferNoOwner(BufferDesc *buf)
 		 * I'd better not still hold the buffer content lock. Can't use
 		 * BufferIsLockedByMe(), as that asserts the buffer is pinned.
 		 */
-		Assert(!LWLockHeldByMe(BufferDescriptorGetContentLock(buf)));
+		Assert(!BufferLockHeldByMe(buf));
 
 		/* decrement the shared reference count */
 		old_buf_state = pg_atomic_fetch_sub_u64(&buf->state, BUF_REFCOUNT_ONE);
@@ -4198,9 +4217,9 @@ CheckForBufferLeaks(void)
  * Check for exclusive-locked catalog buffers.  This is the core of
  * AssertCouldGetRelation().
  *
- * A backend would self-deadlock on LWLocks if the catalog scan read the
- * exclusive-locked buffer.  The main threat is exclusive-locked buffers of
- * catalogs used in relcache, because a catcache search on any catalog may
+ * A backend would self-deadlock on the content lock if the catalog scan read
+ * the exclusive-locked buffer.  The main threat is exclusive-locked buffers
+ * of catalogs used in relcache, because a catcache search on any catalog may
  * build that catalog's relcache entry.  We don't have an inventory of
  * catalogs relcache uses, so just check buffers of most catalogs.
  *
@@ -4214,26 +4233,45 @@ CheckForBufferLeaks(void)
 void
 AssertBufferLocksPermitCatalogRead(void)
 {
-	ForEachLWLockHeldByMe(AssertNotCatalogBufferLock, NULL);
+	PrivateRefCountEntry *res;
+
+	/* check the array */
+	for (int i = 0; i < REFCOUNT_ARRAY_ENTRIES; i++)
+	{
+		if (PrivateRefCountArrayKeys[i] != InvalidBuffer)
+		{
+			res = &PrivateRefCountArray[i];
+
+			if (res->buffer == InvalidBuffer)
+				continue;
+
+			AssertNotCatalogBufferLock(res->buffer, res->data.lockmode);
+		}
+	}
+
+	/* if necessary search the hash */
+	if (PrivateRefCountOverflowed)
+	{
+		HASH_SEQ_STATUS hstat;
+
+		hash_seq_init(&hstat, PrivateRefCountHash);
+		while ((res = (PrivateRefCountEntry *) hash_seq_search(&hstat)) != NULL)
+		{
+			AssertNotCatalogBufferLock(res->buffer, res->data.lockmode);
+		}
+	}
 }
 
 static void
-AssertNotCatalogBufferLock(LWLock *lock, LWLockMode mode,
-						   void *unused_context)
+AssertNotCatalogBufferLock(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *bufHdr;
+	BufferDesc *bufHdr = GetBufferDescriptor(buffer - 1);
 	BufferTag	tag;
 	Oid			relid;
 
-	if (mode != LW_EXCLUSIVE)
+	if (mode != BUFFER_LOCK_EXCLUSIVE)
 		return;
 
-	if (!((BufferDescPadded *) lock > BufferDescriptors &&
-		  (BufferDescPadded *) lock < BufferDescriptors + NBuffers))
-		return;					/* not a buffer lock */
-
-	bufHdr = (BufferDesc *)
-		((char *) lock - offsetof(BufferDesc, content_lock));
 	tag = bufHdr->tag;
 
 	/*
@@ -4515,9 +4553,11 @@ static void
 FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 					IOObject io_object, IOContext io_context)
 {
-	LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	Buffer		buffer = BufferDescriptorGetBuffer(buf);
+
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
-	LWLockRelease(BufferDescriptorGetContentLock(buf));
+	BufferLockUnlock(buffer, buf);
 }
 
 /*
@@ -5660,9 +5700,10 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
  *
  * Used to clean up after errors.
  *
- * Currently, we can expect that lwlock.c's LWLockReleaseAll() took care
- * of releasing buffer content locks per se; the only thing we need to deal
- * with here is clearing any PIN_COUNT request that was in progress.
+ * Currently, we can expect that resource owner cleanup, via
+ * ResOwnerReleaseBufferPin(), took care of releasing buffer content locks per
+ * se; the only thing we need to deal with here is clearing any PIN_COUNT
+ * request that was in progress.
  */
 void
 UnlockBuffers(void)
@@ -5693,25 +5734,726 @@ UnlockBuffers(void)
 }
 
 /*
- * Acquire or release the content_lock for the buffer.
+ * Acquire the buffer content lock in the specified mode
+ *
+ * If the lock is not available, sleep until it is.
+ *
+ * Side effect: cancel/die interrupts are held off until lock release.
+ *
+ * This uses almost the same locking approach as lwlock.c's
+ * LWLockAcquire(). See documentation at the top of lwlock.c for a more
+ * detailed discussion.
+ *
+ * The reason that this, and most of the other BufferLock* functions, get both
+ * the Buffer and BufferDesc* as parameters, is that looking up one from the
+ * other repeatedly shows up noticeably in profiles.
+ *
+ * Callers should provide a constant for mode, for more efficient code
+ * generation.
+ */
+static inline void
+BufferLockAcquire(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry;
+	int			extraWaits = 0;
+
+	/*
+	 * Get reference to the refcount entry before we hold the lock, it seems
+	 * better to do before holding the lock.
+	 */
+	entry = GetPrivateRefCountEntry(buffer, true);
+
+	/*
+	 * We better not already hold a lock on the buffer.
+	 */
+	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	for (;;)
+	{
+		bool		mustwait;
+		uint32		wait_event;
+
+		/*
+		 * Try to grab the lock the first time, we're not in the waitqueue
+		 * yet/anymore.
+		 */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		if (likely(!mustwait))
+		{
+			break;
+		}
+
+		/*
+		 * Ok, at this point we couldn't grab the lock on the first try. We
+		 * cannot simply queue ourselves to the end of the list and wait to be
+		 * woken up because by now the lock could long have been released.
+		 * Instead add us to the queue and try to grab the lock again. If we
+		 * succeed we need to revert the queuing and be happy, otherwise we
+		 * recheck the lock. If we still couldn't grab it, we know that the
+		 * other locker will see our queue entries when releasing since they
+		 * existed before we checked for the lock.
+		 */
+
+		/* add to the queue */
+		BufferLockQueueSelf(buf_hdr, mode);
+
+		/* we're now guaranteed to be woken up if necessary */
+		mustwait = BufferLockAttempt(buf_hdr, mode);
+
+		/* ok, grabbed the lock the second time round, need to undo queueing */
+		if (!mustwait)
+		{
+			BufferLockDequeueSelf(buf_hdr);
+			break;
+		}
+
+		switch (mode)
+		{
+			case BUFFER_LOCK_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE_EXCLUSIVE:
+				wait_event = WAIT_EVENT_BUFFER_SHARE_EXCLUSIVE;
+				break;
+			case BUFFER_LOCK_SHARE:
+				wait_event = WAIT_EVENT_BUFFER_SHARED;
+				break;
+			case BUFFER_LOCK_UNLOCK:
+				pg_unreachable();
+
+		}
+		pgstat_report_wait_start(wait_event);
+
+		/*
+		 * Wait until awakened.
+		 *
+		 * It is possible that we get awakened for a reason other than being
+		 * signaled by BufferLockWakeup().  If so, loop back and wait again.
+		 * Once we've gotten the lock, re-increment the sema by the number of
+		 * additional signals received.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		pgstat_report_wait_end();
+
+		/* Retrying, allow BufferLockRelease to release waiters again. */
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_WAKE_IN_PROGRESS);
+	}
+
+	/* Remember that we now hold this lock */
+	entry->data.lockmode = mode;
+
+	/*
+	 * Fix the process wait semaphore's count for any absorbed wakeups.
+	 */
+	while (unlikely(extraWaits-- > 0))
+		PGSemaphoreUnlock(MyProc->sem);
+}
+
+/*
+ * Release a previously acquired buffer content lock.
+ */
+static void
+BufferLockUnlock(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	uint64		oldstate;
+	uint64		sub;
+
+	mode = BufferLockDisownInternal(buffer, buf_hdr);
+
+	/*
+	 * Release my hold on lock, after that it can immediately be acquired by
+	 * others, even if we still have to wakeup other waiters.
+	 */
+	sub = BufferLockReleaseSub(mode);
+
+	oldstate = pg_atomic_sub_fetch_u64(&buf_hdr->state, sub);
+
+	BufferLockProcessRelease(buf_hdr, mode, oldstate);
+
+	/*
+	 * Now okay to allow cancel/die interrupts.
+	 */
+	RESUME_INTERRUPTS();
+}
+
+
+/*
+ * Acquire the content lock for the buffer, but only if we don't have to wait.
+ */
+static bool
+BufferLockConditional(Buffer buffer, BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry = GetPrivateRefCountEntry(buffer, true);
+	bool		mustwait;
+
+	/*
+	 * We better not already hold a lock on the buffer.
+	 */
+	Assert(entry->data.lockmode == BUFFER_LOCK_UNLOCK);
+
+	/*
+	 * Lock out cancel/die interrupts until we exit the code section protected
+	 * by the content lock.  This ensures that interrupts will not interfere
+	 * with manipulations of data structures in shared memory.
+	 */
+	HOLD_INTERRUPTS();
+
+	/* Check for the lock */
+	mustwait = BufferLockAttempt(buf_hdr, mode);
+
+	if (mustwait)
+	{
+		/* Failed to get lock, so release interrupt holdoff */
+		RESUME_INTERRUPTS();
+	}
+	else
+	{
+		entry->data.lockmode = mode;
+	}
+
+	return !mustwait;
+}
+
+/*
+ * Internal function that tries to atomically acquire the content lock in the
+ * passed in mode.
+ *
+ * This function will not block waiting for a lock to become free - that's the
+ * caller's job.
+ *
+ * Similar to LWLockAttemptLock().
+ */
+static inline bool
+BufferLockAttempt(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	uint64		old_state;
+
+	/*
+	 * Read once outside the loop, later iterations will get the newer value
+	 * via compare & exchange.
+	 */
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+
+	/* loop until we've determined whether we could acquire the lock or not */
+	while (true)
+	{
+		uint64		desired_state;
+		bool		lock_free;
+
+		desired_state = old_state;
+
+		if (mode == BUFFER_LOCK_EXCLUSIVE)
+		{
+			lock_free = (old_state & BM_LOCK_MASK) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_EXCLUSIVE;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			lock_free = (old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+		}
+		else
+		{
+			lock_free = (old_state & BM_LOCK_VAL_EXCLUSIVE) == 0;
+			if (lock_free)
+				desired_state += BM_LOCK_VAL_SHARED;
+		}
+
+		/*
+		 * Attempt to swap in the state we are expecting. If we didn't see
+		 * lock to be free, that's just the old value. If we saw it as free,
+		 * we'll attempt to mark it acquired. The reason that we always swap
+		 * in the value is that this doubles as a memory barrier. We could try
+		 * to be smarter and only swap in values if we saw the lock as free,
+		 * but benchmark haven't shown it as beneficial so far.
+		 *
+		 * Retry if the value changed since we last looked at it.
+		 */
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			if (lock_free)
+			{
+				/* Great! Got the lock. */
+				return false;
+			}
+			else
+				return true;	/* somebody else has the lock */
+		}
+	}
+
+	pg_unreachable();
+}
+
+/*
+ * Add ourselves to the end of the content lock's wait queue.
+ */
+static void
+BufferLockQueueSelf(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	/*
+	 * If we don't have a PGPROC structure, there's no way to wait. This
+	 * should never occur, since MyProc should only be null during shared
+	 * memory initialization.
+	 */
+	if (MyProc == NULL)
+		elog(PANIC, "cannot wait without a PGPROC structure");
+
+	if (MyProc->lwWaiting != LW_WS_NOT_WAITING)
+		elog(PANIC, "queueing for lock while waiting on another one");
+
+	LockBufHdr(buf_hdr);
+
+	/* setting the flag is protected by the spinlock */
+	pg_atomic_fetch_or_u64(&buf_hdr->state, BM_LOCK_HAS_WAITERS);
+
+	/*
+	 * These are currently used both for lwlocks and buffer content locks,
+	 * which is acceptable, although not pretty, because a backend can't wait
+	 * for both types of locks at the same time.
+	 */
+	MyProc->lwWaiting = LW_WS_WAITING;
+	MyProc->lwWaitMode = mode;
+
+	proclist_push_tail(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	/* Can release the mutex now */
+	UnlockBufHdr(buf_hdr);
+}
+
+/*
+ * Remove ourselves from the waitlist.
+ *
+ * This is used if we queued ourselves because we thought we needed to sleep
+ * but, after further checking, we discovered that we don't actually need to
+ * do so.
+ */
+static void
+BufferLockDequeueSelf(BufferDesc *buf_hdr)
+{
+	bool		on_waitlist;
+
+	LockBufHdr(buf_hdr);
+
+	on_waitlist = MyProc->lwWaiting == LW_WS_WAITING;
+	if (on_waitlist)
+		proclist_delete(&buf_hdr->lock_waiters, MyProcNumber, lwWaitLink);
+
+	if (proclist_is_empty(&buf_hdr->lock_waiters) &&
+		(pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS) != 0)
+	{
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_HAS_WAITERS);
+	}
+
+	/* XXX: combine with fetch_and above? */
+	UnlockBufHdr(buf_hdr);
+
+	/* clear waiting state again, nice for debugging */
+	if (on_waitlist)
+		MyProc->lwWaiting = LW_WS_NOT_WAITING;
+	else
+	{
+		int			extraWaits = 0;
+
+
+		/*
+		 * Somebody else dequeued us and has or will wake us up. Deal with the
+		 * superfluous absorption of a wakeup.
+		 */
+
+		/*
+		 * Clear BM_LOCK_WAKE_IN_PROGRESS if somebody woke us before we
+		 * removed ourselves - they'll have set it.
+		 */
+		pg_atomic_fetch_and_u64(&buf_hdr->state, ~BM_LOCK_WAKE_IN_PROGRESS);
+
+		/*
+		 * Now wait for the scheduled wakeup, otherwise our ->lwWaiting would
+		 * get reset at some inconvenient point later. Most of the time this
+		 * will immediately return.
+		 */
+		for (;;)
+		{
+			PGSemaphoreLock(MyProc->sem);
+			if (MyProc->lwWaiting == LW_WS_NOT_WAITING)
+				break;
+			extraWaits++;
+		}
+
+		/*
+		 * Fix the process wait semaphore's count for any absorbed wakeups.
+		 */
+		while (extraWaits-- > 0)
+			PGSemaphoreUnlock(MyProc->sem);
+	}
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * After calling this function it's the callers responsibility to ensure that
+ * the lock gets released, even in case of an error. This only is desirable if
+ * the lock is going to be released in a different process than the process
+ * that acquired it.
+ */
+static inline void
+BufferLockDisown(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockDisownInternal(buffer, buf_hdr);
+	RESUME_INTERRUPTS();
+}
+
+/*
+ * Stop treating lock as held by current backend.
+ *
+ * This is the code that can be shared between actually releasing a lock
+ * (BufferLockUnlock()) and just not tracking ownership of the lock anymore
+ * without releasing the lock (BufferLockDisown()).
+ */
+static inline int
+BufferLockDisownInternal(Buffer buffer, BufferDesc *buf_hdr)
+{
+	BufferLockMode mode;
+	PrivateRefCountEntry *ref;
+
+	ref = GetPrivateRefCountEntry(buffer, false);
+	if (ref == NULL)
+		elog(ERROR, "lock %d is not held", buffer);
+	mode = ref->data.lockmode;
+	ref->data.lockmode = BUFFER_LOCK_UNLOCK;
+
+	return mode;
+}
+
+/*
+ * Wakeup all the lockers that currently have a chance to acquire the lock.
+ *
+ * wake_exclusive indicates whether exclusive lock waiters should be woken up.
+ */
+static void
+BufferLockWakeup(BufferDesc *buf_hdr, bool wake_exclusive)
+{
+	bool		new_wake_in_progress = false;
+	bool		wake_share_exclusive = true;
+	proclist_head wakeup;
+	proclist_mutable_iter iter;
+
+	proclist_init(&wakeup);
+
+	/* lock wait list while collecting backends to wake up */
+	LockBufHdr(buf_hdr);
+
+	proclist_foreach_modify(iter, &buf_hdr->lock_waiters, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		/*
+		 * Already woke up a conflicting lock, so skip over this wait list
+		 * entry.
+		 */
+		if (!wake_exclusive && waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+			continue;
+		if (!wake_share_exclusive && waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+			continue;
+
+		proclist_delete(&buf_hdr->lock_waiters, iter.cur, lwWaitLink);
+		proclist_push_tail(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Prevent additional wakeups until retryer gets to run. Backends that
+		 * are just waiting for the lock to become free don't retry
+		 * automatically.
+		 */
+		new_wake_in_progress = true;
+
+		/*
+		 * Signal that the process isn't on the wait list anymore. This allows
+		 * BufferLockDequeueSelf() to remove itself from the waitlist with a
+		 * proclist_delete(), rather than having to check if it has been
+		 * removed from the list.
+		 */
+		Assert(waiter->lwWaiting == LW_WS_WAITING);
+		waiter->lwWaiting = LW_WS_PENDING_WAKEUP;
+
+		/*
+		 * Don't wakeup further waiters after waking a conflicting waiter.
+		 */
+		if (waiter->lwWaitMode == BUFFER_LOCK_SHARE)
+		{
+			/*
+			 * Share locks conflict with exclusive locks.
+			 */
+			wake_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			/*
+			 * Share-exclusive locks conflict with share-exclusive and
+			 * exclusive locks.
+			 */
+			wake_exclusive = false;
+			wake_share_exclusive = false;
+		}
+		else if (waiter->lwWaitMode == BUFFER_LOCK_EXCLUSIVE)
+		{
+
+			/*
+			 * Exclusive locks conflict with all other locks, there's no point
+			 * in waking up anybody else.
+			 */
+			break;
+		}
+	}
+
+	Assert(proclist_is_empty(&wakeup) || pg_atomic_read_u64(&buf_hdr->state) & BM_LOCK_HAS_WAITERS);
+
+	/* unset required flags, and release lock, in one fell swoop */
+	{
+		uint64		old_state;
+		uint64		desired_state;
+
+		old_state = pg_atomic_read_u64(&buf_hdr->state);
+		while (true)
+		{
+			desired_state = old_state;
+
+			/* compute desired flags */
+
+			if (new_wake_in_progress)
+				desired_state |= BM_LOCK_WAKE_IN_PROGRESS;
+			else
+				desired_state &= ~BM_LOCK_WAKE_IN_PROGRESS;
+
+			if (proclist_is_empty(&buf_hdr->lock_waiters))
+				desired_state &= ~BM_LOCK_HAS_WAITERS;
+
+			desired_state &= ~BM_LOCKED;	/* release lock */
+
+			if (pg_atomic_compare_exchange_u64(&buf_hdr->state, &old_state,
+											   desired_state))
+				break;
+		}
+	}
+
+	/* Awaken any waiters I removed from the queue. */
+	proclist_foreach_modify(iter, &wakeup, lwWaitLink)
+	{
+		PGPROC	   *waiter = GetPGProcByNumber(iter.cur);
+
+		proclist_delete(&wakeup, iter.cur, lwWaitLink);
+
+		/*
+		 * Guarantee that lwWaiting being unset only becomes visible once the
+		 * unlink from the link has completed. Otherwise the target backend
+		 * could be woken up for other reason and enqueue for a new lock - if
+		 * that happens before the list unlink happens, the list would end up
+		 * being corrupted.
+		 *
+		 * The barrier pairs with the LockBufHdr() when enqueuing for another
+		 * lock.
+		 */
+		pg_write_barrier();
+		waiter->lwWaiting = LW_WS_NOT_WAITING;
+		PGSemaphoreUnlock(waiter->sem);
+	}
+}
+
+/*
+ * Compute subtraction from buffer state for a release of a held lock in
+ * `mode`.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places, each needing to compute what to
+ * subtract from the lock state.
+ */
+static inline uint64
+BufferLockReleaseSub(BufferLockMode mode)
+{
+
+	/*
+	 * Turns out that a switch() leads gcc to generate sufficiently worse code
+	 * for this to show up in profiles...
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE)
+		return BM_LOCK_VAL_EXCLUSIVE;
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		return BM_LOCK_VAL_SHARE_EXCLUSIVE;
+	else
+	{
+		Assert(mode == BUFFER_LOCK_SHARE);
+		return BM_LOCK_VAL_SHARED;
+	}
+
+	return 0;					/* keep compiler quiet */
+}
+
+/*
+ * Handle work that needs to be done after releasing a lock that was held in
+ * `mode`, where `lockstate` is the result of the atomic operation modifying
+ * the state variable.
+ *
+ * This is separated from BufferLockUnlock() as we want to combine the lock
+ * release with other atomic operations when possible, leading to the lock
+ * release being done in multiple places.
+ */
+static void
+BufferLockProcessRelease(BufferDesc *buf_hdr, BufferLockMode mode, uint64 lockstate)
+{
+	bool		check_waiters = false;
+	bool		wake_exclusive = false;
+
+	/* nobody else can have that kind of lock */
+	Assert(!(lockstate & BM_LOCK_VAL_EXCLUSIVE));
+
+	/*
+	 * If we're still waiting for backends to get scheduled, don't wake them
+	 * up again. Otherwise check if we need to look through the waitqueue to
+	 * wake other backends.
+	 */
+	if ((lockstate & BM_LOCK_HAS_WAITERS) &&
+		!(lockstate & BM_LOCK_WAKE_IN_PROGRESS))
+	{
+		if ((lockstate & BM_LOCK_MASK) == 0)
+		{
+			/*
+			 * We released a lock and the lock was, in that moment, free. We
+			 * therefore can wake waiters for any kind of lock.
+			 */
+			check_waiters = true;
+			wake_exclusive = true;
+		}
+		else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		{
+			/*
+			 * We released the lock, but another backend still holds a lock.
+			 * We can't have released an exclusive lock, as there couldn't
+			 * have been other lock holders. If we released a share lock, no
+			 * waiters need to be woken up, as there must be other share
+			 * lockers. However, if we held a share-exclusive lock, another
+			 * backend now could acquire a share-exclusive lock.
+			 */
+			check_waiters = true;
+			wake_exclusive = false;
+		}
+	}
+
+	/*
+	 * As waking up waiters requires the spinlock to be acquired, only do so
+	 * if necessary.
+	 */
+	if (check_waiters)
+		BufferLockWakeup(buf_hdr, wake_exclusive);
+}
+
+/*
+ * BufferLockHeldByMeInMode - test whether my process holds the content lock
+ * in the specified mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMeInMode(BufferDesc *buf_hdr, BufferLockMode mode)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode == mode;
+
+}
+
+/*
+ * BufferLockHeldByMe - test whether my process holds the content lock in any
+ * mode
+ *
+ * This is meant as debug support only.
+ */
+static bool
+BufferLockHeldByMe(BufferDesc *buf_hdr)
+{
+	PrivateRefCountEntry *entry =
+		GetPrivateRefCountEntry(BufferDescriptorGetBuffer(buf_hdr), false);
+
+	if (!entry)
+		return false;
+	else
+		return entry->data.lockmode != BUFFER_LOCK_UNLOCK;
+}
+
+/*
+ * Release the content lock for the buffer.
+ */
+void
+UnlockBuffer(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+
+	Assert(BufferIsPinned(buffer));
+	if (BufferIsLocal(buffer))
+		return;					/* local buffers need no lock */
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+	BufferLockUnlock(buffer, buf_hdr);
+}
+
+/*
+ * Acquire the content_lock for the buffer.
  */
 void
-LockBuffer(Buffer buffer, BufferLockMode mode)
+LockBufferInternal(Buffer buffer, BufferLockMode mode)
 {
-	BufferDesc *buf;
+	BufferDesc *buf_hdr;
+
+	/*
+	 * We can't wait if we haven't got a PGPROC.  This should only occur
+	 * during bootstrap or shared memory initialization.  Put an Assert here
+	 * to catch unsafe coding practices.
+	 */
+	Assert(!(MyProc == NULL && IsUnderPostmaster));
+
+	/* handled in LockBuffer() wrapper */
+	Assert(mode != BUFFER_LOCK_UNLOCK);
 
 	Assert(BufferIsPinned(buffer));
 	if (BufferIsLocal(buffer))
 		return;					/* local buffers need no lock */
 
-	buf = GetBufferDescriptor(buffer - 1);
+	buf_hdr = GetBufferDescriptor(buffer - 1);
 
-	if (mode == BUFFER_LOCK_UNLOCK)
-		LWLockRelease(BufferDescriptorGetContentLock(buf));
-	else if (mode == BUFFER_LOCK_SHARE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_SHARED);
+	/*
+	 * Test the most frequent lock modes first. While a switch (mode) would be
+	 * nice, at least gcc generates considerably worse code for it.
+	 *
+	 * Call BufferLockAcquire() with a constant argument for mode, to generate
+	 * more efficient code for the different lock modes.
+	 */
+	if (mode == BUFFER_LOCK_SHARE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE);
 	else if (mode == BUFFER_LOCK_EXCLUSIVE)
-		LWLockAcquire(BufferDescriptorGetContentLock(buf), LW_EXCLUSIVE);
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_EXCLUSIVE);
+	else if (mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+		BufferLockAcquire(buffer, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	else
 		elog(ERROR, "unrecognized buffer lock mode: %d", mode);
 }
@@ -5732,8 +6474,7 @@ ConditionalLockBuffer(Buffer buffer)
 
 	buf = GetBufferDescriptor(buffer - 1);
 
-	return LWLockConditionalAcquire(BufferDescriptorGetContentLock(buf),
-									LW_EXCLUSIVE);
+	return BufferLockConditional(buffer, buf, BUFFER_LOCK_EXCLUSIVE);
 }
 
 /*
@@ -6247,8 +6988,8 @@ TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 /*
  * AbortBufferIO: Clean up active buffer I/O after an error.
  *
- *	All LWLocks we might have held have been released,
- *	but we haven't yet released buffer pins, so the buffer is still pinned.
+ *	All LWLocks & content locks we might have held have been released, but we
+ *	haven't yet released buffer pins, so the buffer is still pinned.
  *
  *	If I/O was in progress, we always set BM_IO_ERROR, even though it's
  *	possible the error condition wasn't related to the I/O.
@@ -6676,8 +7417,14 @@ ResOwnerPrintBufferIO(Datum res)
 	return psprintf("lost track of buffer IO on buffer %d", buffer);
 }
 
+/*
+ * Release buffer as part of resource owner cleanup. This will only be called
+ * if the buffer is pinned. If this backend held the content lock at the time
+ * of the error we also need to release that (note that it is not possible to
+ * hold a content lock without a pin).
+ */
 static void
-ResOwnerReleaseBufferPin(Datum res)
+ResOwnerReleaseBuffer(Datum res)
 {
 	Buffer		buffer = DatumGetInt32(res);
 
@@ -6688,11 +7435,32 @@ ResOwnerReleaseBufferPin(Datum res)
 	if (BufferIsLocal(buffer))
 		UnpinLocalBufferNoOwner(buffer);
 	else
+	{
+		PrivateRefCountEntry *ref;
+
+		ref = GetPrivateRefCountEntry(buffer, false);
+
+		/* not having a private refcount would imply resowner corruption */
+		Assert(ref != NULL);
+
+		/*
+		 * If the buffer was locked at the time of the resowner release,
+		 * release the lock now. This should only happen after errors.
+		 */
+		if (ref->data.lockmode != BUFFER_LOCK_UNLOCK)
+		{
+			BufferDesc *buf = GetBufferDescriptor(buffer - 1);
+
+			HOLD_INTERRUPTS();	/* match the upcoming RESUME_INTERRUPTS */
+			BufferLockUnlock(buffer, buf);
+		}
+
 		UnpinBufferNoOwner(GetBufferDescriptor(buffer - 1));
+	}
 }
 
 static char *
-ResOwnerPrintBufferPin(Datum res)
+ResOwnerPrintBuffer(Datum res)
 {
 	return DebugPrintBufferRefcount(DatumGetInt32(res));
 }
@@ -6924,10 +7692,10 @@ MarkDirtyUnpinnedBufferInternal(Buffer buf, BufferDesc *desc,
 	/* If it was not already dirty, mark it as dirty. */
 	if (!(buf_state & BM_DIRTY))
 	{
-		LWLockAcquire(BufferDescriptorGetContentLock(desc), LW_EXCLUSIVE);
+		BufferLockAcquire(buf, desc, BUFFER_LOCK_EXCLUSIVE);
 		MarkBufferDirty(buf);
 		result = true;
-		LWLockRelease(BufferDescriptorGetContentLock(desc));
+		BufferLockUnlock(buf, desc);
 	}
 	else
 		*buffer_already_dirty = true;
@@ -7178,16 +7946,12 @@ buffer_stage_common(PgAioHandle *ioh, bool is_write, bool is_temp)
 		 */
 		if (is_write && !is_temp)
 		{
-			LWLock	   *content_lock;
-
-			content_lock = BufferDescriptorGetContentLock(buf_hdr);
-
-			Assert(LWLockHeldByMe(content_lock));
+			Assert(BufferLockHeldByMe(buf_hdr));
 
 			/*
 			 * Lock is now owned by AIO subsystem.
 			 */
-			LWLockDisown(content_lock);
+			BufferLockDisown(buffer, buf_hdr);
 		}
 
 		/*
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 3299de23bb3..b8936d30d7e 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -287,6 +287,9 @@ ABI_compatibility:
 Section: ClassName - WaitEventBuffer
 
 BUFFER_CLEANUP	"Waiting to acquire an exclusive pin on a buffer. Buffer pin waits can be protracted if another process holds an open cursor that last read data from the buffer in question."
+BUFFER_SHARED	"Waiting to acquire a shared lock on a buffer."
+BUFFER_SHARE_EXCLUSIVE	"Waiting to acquire a share exclusive lock on a buffer."
+BUFFER_EXCLUSIVE	"Waiting to acquire a exclusive lock on a buffer."
 
 ABI_compatibility:
 
@@ -374,7 +377,6 @@ MultiXactMemberBuffer	"Waiting for I/O on a multixact member SLRU buffer."
 NotifyBuffer	"Waiting for I/O on a <command>NOTIFY</command> message SLRU buffer."
 SerialBuffer	"Waiting for I/O on a serializable transaction conflict SLRU buffer."
 WALInsert	"Waiting to insert WAL data into a memory buffer."
-BufferContent	"Waiting to access a data page in memory."
 ReplicationOriginState	"Waiting to read or update the progress of one replication origin."
 ReplicationSlotIO	"Waiting for I/O on a replication slot."
 LockFastPath	"Waiting to read or update a process' fast-path lock information."
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v11-0003-lwlock-Remove-ForEachLWLockHeldByMe.patch (2.2K, ../../k5j77f3q6ztihnjnx2nqxzyor6fbj2qxcbhzuxhkh2yy63jyfg@p72phigar3n4/4-v11-0003-lwlock-Remove-ForEachLWLockHeldByMe.patch)
  download | inline diff:
From d2eabd283e76aeb1da967581d47b0576a104c28e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 14 Jan 2026 13:39:50 -0500
Subject: [PATCH v11 3/7] lwlock: Remove ForEachLWLockHeldByMe

As of commit FIXME-XXX-UPDATEME, ForEachLWLockHeldByMe(), introduced in
f4ece891fc2f, is not used anymore, as content locks are now implemented in
bufmgr.c.  It doesn't seem that likely that a new user of the functionality
will appear all that soon, making removal of the function seem like the most
sensible path. It can easily be added back if necessary.

Discussion: https://postgr.es/m/lneuyxqxamqoayd2ntau3lqjblzdckw6tjgeu4574ezwh4tzlg%40noioxkquezdw
---
 src/include/storage/lwlock.h      |  2 --
 src/backend/storage/lmgr/lwlock.c | 15 ---------------
 2 files changed, 17 deletions(-)

diff --git a/src/include/storage/lwlock.h b/src/include/storage/lwlock.h
index a98d302c602..df589902adc 100644
--- a/src/include/storage/lwlock.h
+++ b/src/include/storage/lwlock.h
@@ -129,8 +129,6 @@ extern void LWLockReleaseClearVar(LWLock *lock, pg_atomic_uint64 *valptr, uint64
 extern void LWLockReleaseAll(void);
 extern void LWLockDisown(LWLock *lock);
 extern void LWLockReleaseDisowned(LWLock *lock, LWLockMode mode);
-extern void ForEachLWLockHeldByMe(void (*callback) (LWLock *, LWLockMode, void *),
-								  void *context);
 extern bool LWLockHeldByMe(LWLock *lock);
 extern bool LWLockAnyHeldByMe(LWLock *lock, int nlocks, size_t stride);
 extern bool LWLockHeldByMeInMode(LWLock *lock, LWLockMode mode);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 415794682a2..2ee0339c52e 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -1955,21 +1955,6 @@ LWLockReleaseAll(void)
 }
 
 
-/*
- * ForEachLWLockHeldByMe - run a callback for each held lock
- *
- * This is meant as debug support only.
- */
-void
-ForEachLWLockHeldByMe(void (*callback) (LWLock *, LWLockMode, void *),
-					  void *context)
-{
-	int			i;
-
-	for (i = 0; i < num_held_lwlocks; i++)
-		callback(held_lwlocks[i].lock, held_lwlocks[i].mode, context);
-}
-
 /*
  * LWLockHeldByMe - test whether my process holds a lock in any mode
  *
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v11-0004-lwlock-Remove-support-for-disowned-lwlwocks.patch (4.6K, ../../k5j77f3q6ztihnjnx2nqxzyor6fbj2qxcbhzuxhkh2yy63jyfg@p72phigar3n4/5-v11-0004-lwlock-Remove-support-for-disowned-lwlwocks.patch)
  download | inline diff:
From 60c879adc2540dad404a05e7f2c20957573b4211 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 14 Jan 2026 13:54:00 -0500
Subject: [PATCH v11 4/7] lwlock: Remove support for disowned lwlwocks

This reverts commit f8d7f29b3e81db59b95e4b5baaa6943178c89fd8, plus parts of
subsequent commits fixing a typo an a parameter name.

Support for disowned lwlocks was added for the benefit of AIO, to be able to
have content locks "owned" by the AIO subsystem. But as of commit
FIXME-XXX-UPDATEME, content locks do not use lwlocks anymore.

It does not seem particularly likely that we need this facility outside of the
AIO use-case, therefore remove the now unused functions.

I did choose to keep the comment added in the aforementioned commit about
lock->owner intentionally being left pointing to the last owner.
---
 src/include/storage/lwlock.h      |  2 -
 src/backend/storage/lmgr/lwlock.c | 71 +++----------------------------
 2 files changed, 6 insertions(+), 67 deletions(-)

diff --git a/src/include/storage/lwlock.h b/src/include/storage/lwlock.h
index df589902adc..9a0290391d0 100644
--- a/src/include/storage/lwlock.h
+++ b/src/include/storage/lwlock.h
@@ -127,8 +127,6 @@ extern bool LWLockAcquireOrWait(LWLock *lock, LWLockMode mode);
 extern void LWLockRelease(LWLock *lock);
 extern void LWLockReleaseClearVar(LWLock *lock, pg_atomic_uint64 *valptr, uint64 val);
 extern void LWLockReleaseAll(void);
-extern void LWLockDisown(LWLock *lock);
-extern void LWLockReleaseDisowned(LWLock *lock, LWLockMode mode);
 extern bool LWLockHeldByMe(LWLock *lock);
 extern bool LWLockAnyHeldByMe(LWLock *lock, int nlocks, size_t stride);
 extern bool LWLockHeldByMeInMode(LWLock *lock, LWLockMode mode);
diff --git a/src/backend/storage/lmgr/lwlock.c b/src/backend/storage/lmgr/lwlock.c
index 2ee0339c52e..a133c97b992 100644
--- a/src/backend/storage/lmgr/lwlock.c
+++ b/src/backend/storage/lmgr/lwlock.c
@@ -1783,25 +1783,18 @@ LWLockUpdateVar(LWLock *lock, pg_atomic_uint64 *valptr, uint64 val)
 
 
 /*
- * Stop treating lock as held by current backend.
- *
- * This is the code that can be shared between actually releasing a lock
- * (LWLockRelease()) and just not tracking ownership of the lock anymore
- * without releasing the lock (LWLockDisown()).
- *
- * Returns the mode in which the lock was held by the current backend.
- *
- * NB: This does not call RESUME_INTERRUPTS(), but leaves that responsibility
- * of the caller.
+ * LWLockRelease - release a previously acquired lock
  *
  * NB: This will leave lock->owner pointing to the current backend (if
  * LOCK_DEBUG is set). This is somewhat intentional, as it makes it easier to
  * debug cases of missing wakeups during lock release.
  */
-static inline LWLockMode
-LWLockDisownInternal(LWLock *lock)
+void
+LWLockRelease(LWLock *lock)
 {
 	LWLockMode	mode;
+	uint32		oldstate;
+	bool		check_waiters;
 	int			i;
 
 	/*
@@ -1821,18 +1814,7 @@ LWLockDisownInternal(LWLock *lock)
 	for (; i < num_held_lwlocks; i++)
 		held_lwlocks[i] = held_lwlocks[i + 1];
 
-	return mode;
-}
-
-/*
- * Helper function to release lock, shared between LWLockRelease() and
- * LWLockReleaseDisowned().
- */
-static void
-LWLockReleaseInternal(LWLock *lock, LWLockMode mode)
-{
-	uint32		oldstate;
-	bool		check_waiters;
+	PRINT_LWDEBUG("LWLockRelease", lock, mode);
 
 	/*
 	 * Release my hold on lock, after that it can immediately be acquired by
@@ -1870,38 +1852,6 @@ LWLockReleaseInternal(LWLock *lock, LWLockMode mode)
 		LOG_LWDEBUG("LWLockRelease", lock, "releasing waiters");
 		LWLockWakeup(lock);
 	}
-}
-
-
-/*
- * Stop treating lock as held by current backend.
- *
- * After calling this function it's the callers responsibility to ensure that
- * the lock gets released (via LWLockReleaseDisowned()), even in case of an
- * error. This only is desirable if the lock is going to be released in a
- * different process than the process that acquired it.
- */
-void
-LWLockDisown(LWLock *lock)
-{
-	LWLockDisownInternal(lock);
-
-	RESUME_INTERRUPTS();
-}
-
-/*
- * LWLockRelease - release a previously acquired lock
- */
-void
-LWLockRelease(LWLock *lock)
-{
-	LWLockMode	mode;
-
-	mode = LWLockDisownInternal(lock);
-
-	PRINT_LWDEBUG("LWLockRelease", lock, mode);
-
-	LWLockReleaseInternal(lock, mode);
 
 	/*
 	 * Now okay to allow cancel/die interrupts.
@@ -1909,15 +1859,6 @@ LWLockRelease(LWLock *lock)
 	RESUME_INTERRUPTS();
 }
 
-/*
- * Release lock previously disowned with LWLockDisown().
- */
-void
-LWLockReleaseDisowned(LWLock *lock, LWLockMode mode)
-{
-	LWLockReleaseInternal(lock, mode);
-}
-
 /*
  * LWLockReleaseClearVar - release a previously acquired lock, reset variable
  */
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v11-0005-Require-share-exclusive-lock-to-set-hint-bits-an.patch (40.8K, ../../k5j77f3q6ztihnjnx2nqxzyor6fbj2qxcbhzuxhkh2yy63jyfg@p72phigar3n4/6-v11-0005-Require-share-exclusive-lock-to-set-hint-bits-an.patch)
  download | inline diff:
From ec49b1a8665a2ce78946b127360f16f92b58b00e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 13 Jan 2026 20:10:32 -0500
Subject: [PATCH v11 5/7] Require share-exclusive lock to set hint bits and to
 flush

At the moment hint bits can be set with just a share lock on a page (and,
until 45f658dacb9, in one case even without any lock). Because of this we need
to copy pages while writing them out, as otherwise the checksum could be
corrupted.

The need to copy the page is problematic to implement AIO writes:

1) Instead of just needing a single buffer for a copied page we need one for
   each page that's potentially undergoing I/O
2) To be able to use the "worker" AIO implementation the copied page needs to
   reside in shared memory

It also causes problems for using unbuffered/direct-IO, independent of AIO:
Some filesystems, raid implementations, ... do not tolerate the data being
written out to change during the write. E.g. they may compute internal
checksums that can be invalidated by concurrent modifications, leading e.g. to
filesystem errors (as the case with btrfs).

It also just is plain odd to allow modifications of buffers that are just
share locked.

To address these issue, this commit changes the rules so that modifications to
pages are not allowed anymore while holding a share lock. Instead the new
share-exclusive lock (introduced in FIXME XXXX TODO) allows at most one
backend to modify a buffer while other backends have the same page share
locked. An existing share-lock can be upgraded to a share-exclusive lock, if
there are no conflicting locks. For that
BufferBeginSetHintBits()/BufferFinishSetHintBits() and BufferSetHintBits16()
have been introduced.

To prevent hint bits from being set while the buffer is being written out,
writing out buffers now requires a share-exclusive lock.

The use of share-exclusive to gate setting hint bits means that from now on
only one backend can set hint bits at a time. To allow multiple backends to
set hint bits would require more complicated locking, for setting hint bits
we'd need to store the count of backends currently setting hint bits and we
would need another lock-level for I/O conflicting with the lock-level to set
hint bits. Given that the share-exclusive lock for setting hint bits is only
held for a short time, that backends would often just set the same hint bits
and that the cost of occasionally not setting hint bits in hotly accessed
pages is fairly low, this seems like an acceptable tradeoff.

The biggest change to adapt to this is in heapam. To avoid performance
regressions for sequential scans that need to set a lot of hint bits, we need
to amortize the cost of BufferBeginSetHintBits() for cases where hint bits are
set at a high frequency, HeapTupleSatisfiesMVCCBatch() uses the new
SetHintBitsExt() which defers BufferFinishSetHintBits() until all hint bits on
a page have been set.  Conversely, to avoid regressions in cases where we
can't set hint bits in bulk (because we're looking only at individual tuples),
use BufferSetHintBits16() when setting hint bits without batching.

Several other places also need to be adapted, but those changes are
comparatively simpler.

After this we do not need to copy buffers to write them out anymore. That
change is done separately however.

TODO:
- Update commit reference above
- Update FIXME comments

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf%40gcnactj4z56m
---
 src/include/storage/bufmgr.h                |   4 +
 src/backend/access/gist/gistget.c           |  19 +-
 src/backend/access/hash/hashutil.c          |  10 +-
 src/backend/access/heap/heapam_visibility.c | 130 ++++++--
 src/backend/access/nbtree/nbtinsert.c       |  28 +-
 src/backend/access/nbtree/nbtutils.c        |  16 +-
 src/backend/storage/buffer/README           |  66 ++--
 src/backend/storage/buffer/bufmgr.c         | 327 ++++++++++++++++----
 src/backend/storage/freespace/freespace.c   |  14 +-
 src/backend/storage/freespace/fsmpage.c     |  11 +-
 src/tools/pgindent/typedefs.list            |   1 +
 11 files changed, 482 insertions(+), 144 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index a40adf6b2a8..4017896f951 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -314,6 +314,10 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
+extern bool BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer);
+extern bool BufferBeginSetHintBits(Buffer buffer);
+extern void BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std);
+
 extern void UnlockBuffers(void);
 extern void UnlockBuffer(Buffer buffer);
 extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
diff --git a/src/backend/access/gist/gistget.c b/src/backend/access/gist/gistget.c
index 11b214eb99b..fc346dc9484 100644
--- a/src/backend/access/gist/gistget.c
+++ b/src/backend/access/gist/gistget.c
@@ -64,11 +64,7 @@ gistkillitems(IndexScanDesc scan)
 	 * safe.
 	 */
 	if (BufferGetLSNAtomic(buffer) != so->curPageLSN)
-	{
-		UnlockReleaseBuffer(buffer);
-		so->numKilled = 0;		/* reset counter */
-		return;
-	}
+		goto unlock;
 
 	Assert(GistPageIsLeaf(page));
 
@@ -78,6 +74,16 @@ gistkillitems(IndexScanDesc scan)
 	 */
 	for (i = 0; i < so->numKilled; i++)
 	{
+		if (!killedsomething)
+		{
+			/*
+			 * Use hint bit infrastructure to be allowed to modify the page
+			 * without holding an exclusive lock.
+			 */
+			if (!BufferBeginSetHintBits(buffer))
+				goto unlock;
+		}
+
 		offnum = so->killedItems[i];
 		iid = PageGetItemId(page, offnum);
 		ItemIdMarkDead(iid);
@@ -87,9 +93,10 @@ gistkillitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		GistMarkPageHasGarbage(page);
-		MarkBufferDirtyHint(buffer, true);
+		BufferFinishSetHintBits(buffer, true, true);
 	}
 
+unlock:
 	UnlockReleaseBuffer(buffer);
 
 	/*
diff --git a/src/backend/access/hash/hashutil.c b/src/backend/access/hash/hashutil.c
index cf7f0b90176..b917c97321a 100644
--- a/src/backend/access/hash/hashutil.c
+++ b/src/backend/access/hash/hashutil.c
@@ -593,6 +593,13 @@ _hash_kill_items(IndexScanDesc scan)
 
 			if (ItemPointerEquals(&ituple->t_tid, &currItem->heapTid))
 			{
+				/*
+				 * Use hint bit infrastructure to be allowed to modify the
+				 * page without holding an exclusive lock.
+				 */
+				if (!BufferBeginSetHintBits(so->currPos.buf))
+					goto unlock_page;
+
 				/* found the item */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -610,9 +617,10 @@ _hash_kill_items(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->hasho_flag |= LH_PAGE_HAS_DEAD_TUPLES;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(so->currPos.buf, true, true);
 	}
 
+unlock_page:
 	if (so->hashso_bucket_buf == so->currPos.buf ||
 		havePin)
 		LockBuffer(so->currPos.buf, BUFFER_LOCK_UNLOCK);
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 75ae268d753..fc64f4343ce 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -80,10 +80,38 @@
 
 
 /*
- * SetHintBits()
+ * To be allowed to set hint bits, SetHintBits() needs to call
+ * BufferBeginSetHintBits(). However, that's not free, and some callsites call
+ * SetHintBits() on many tuples in a row. For those it makes sense to amortize
+ * the cost of BufferBeginSetHintBits(). Additionally it's desirable to defer
+ * the cost of BufferBeginSetHintBits() until a hint bit needs to actually be
+ * set. This enum serves as the necessary state space passed to
+ * SetHintBitsExt().
+ */
+typedef enum SetHintBitsState
+{
+	/* not yet checked if hint bits may be set */
+	SHB_INITIAL,
+	/* failed to get permission to set hint bits, don't check again */
+	SHB_DISABLED,
+	/* allowed to set hint bits */
+	SHB_ENABLED,
+} SetHintBitsState;
+
+/*
+ * SetHintBitsExt()
  *
  * Set commit/abort hint bits on a tuple, if appropriate at this time.
  *
+ * To be allowed to set a hint bit on a tuple, the page must not be undergoing
+ * IO at this time (otherwise we e.g. could corrupt PG's page checksum or even
+ * the filesystem's, as is known to happen with btrfs).
+ *
+ * The right to set a hint bit can be acquired on a page level with
+ * BufferBeginSetHintBits(). Only a single backend gets the right to set hint
+ * bits at a time.  Alternatively, if called with a NULL SetHintBitsState*,
+ * hint bits are set with BufferSetHintBits16().
+ *
  * It is only safe to set a transaction-committed hint bit if we know the
  * transaction's commit record is guaranteed to be flushed to disk before the
  * buffer, or if the table is temporary or unlogged and will be obliterated by
@@ -111,24 +139,67 @@
  * InvalidTransactionId if no check is needed.
  */
 static inline void
-SetHintBits(HeapTupleHeader tuple, Buffer buffer,
-			uint16 infomask, TransactionId xid)
+SetHintBitsExt(HeapTupleHeader tuple, Buffer buffer,
+			   uint16 infomask, TransactionId xid, SetHintBitsState *state)
 {
+	/*
+	 * In batched mode, if we previously did not get permission to set hint
+	 * bits, don't try again - in all likelihood IO is still going on.
+	 */
+	if (state && *state == SHB_DISABLED)
+		return;
+
 	if (TransactionIdIsValid(xid))
 	{
-		/* NB: xid must be known committed here! */
-		XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+		if (BufferIsPermanent(buffer))
+		{
+			/* NB: xid must be known committed here! */
+			XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+
+			if (XLogNeedsFlush(commitLSN) &&
+				BufferGetLSNAtomic(buffer) < commitLSN)
+			{
+				/* not flushed and no LSN interlock, so don't set hint */
+				return;
+			}
+		}
+	}
+
+	/*
+	 * If we're not operating in batch mode, use BufferSetHintBits16() to mark
+	 * the page dirty, that's cheaper than
+	 * BufferBeginSetHintBits()/BufferFinishSetHintBits(). That's important
+	 * for cases where we set a lot of hint bits on a page individually.
+	 */
+	if (!state)
+	{
+		BufferSetHintBits16(&tuple->t_infomask,
+							tuple->t_infomask | infomask, buffer);
+		return;
+	}
 
-		if (BufferIsPermanent(buffer) && XLogNeedsFlush(commitLSN) &&
-			BufferGetLSNAtomic(buffer) < commitLSN)
+	if (*state == SHB_INITIAL)
+	{
+		if (!BufferBeginSetHintBits(buffer))
 		{
-			/* not flushed and no LSN interlock, so don't set hint */
+			*state = SHB_DISABLED;
 			return;
 		}
-	}
 
+		*state = SHB_ENABLED;
+	}
 	tuple->t_infomask |= infomask;
-	MarkBufferDirtyHint(buffer, true);
+}
+
+/*
+ * Simple wrapper around SetHintBitExt(), use when operating on a single
+ * tuple.
+ */
+static inline void
+SetHintBits(HeapTupleHeader tuple, Buffer buffer,
+			uint16 infomask, TransactionId xid)
+{
+	SetHintBitsExt(tuple, buffer, infomask, xid, NULL);
 }
 
 /*
@@ -864,9 +935,9 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
  * inserting/deleting transaction was still running --- which was more cycles
  * and more contention on ProcArrayLock.
  */
-static bool
+static inline bool
 HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
-					   Buffer buffer)
+					   Buffer buffer, SetHintBitsState *state)
 {
 	HeapTupleHeader tuple = htup->t_data;
 
@@ -921,8 +992,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 			if (!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmax(tuple)))
 			{
 				/* deleting subtransaction must have aborted */
-				SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-							InvalidTransactionId);
+				SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+							   InvalidTransactionId, state);
 				return true;
 			}
 
@@ -934,13 +1005,13 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		else if (XidInMVCCSnapshot(HeapTupleHeaderGetRawXmin(tuple), snapshot))
 			return false;
 		else if (TransactionIdDidCommit(HeapTupleHeaderGetRawXmin(tuple)))
-			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-						HeapTupleHeaderGetRawXmin(tuple));
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_COMMITTED,
+						   HeapTupleHeaderGetRawXmin(tuple), state);
 		else
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_INVALID,
+						   InvalidTransactionId, state);
 			return false;
 		}
 	}
@@ -1003,14 +1074,14 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		if (!TransactionIdDidCommit(HeapTupleHeaderGetRawXmax(tuple)))
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+						   InvalidTransactionId, state);
 			return true;
 		}
 
 		/* xmax transaction committed */
-		SetHintBits(tuple, buffer, HEAP_XMAX_COMMITTED,
-					HeapTupleHeaderGetRawXmax(tuple));
+		SetHintBitsExt(tuple, buffer, HEAP_XMAX_COMMITTED,
+					   HeapTupleHeaderGetRawXmax(tuple), state);
 	}
 	else
 	{
@@ -1607,9 +1678,10 @@ HeapTupleSatisfiesHistoricMVCC(HeapTuple htup, Snapshot snapshot,
  * ->vistuples_dense is set to contain the offsets of visible tuples.
  *
  * The reason this is more efficient than HeapTupleSatisfiesMVCC() is that it
- * avoids a cross-translation-unit function call for each tuple and allows the
- * compiler to optimize across calls to HeapTupleSatisfiesMVCC. In the future
- * it will also allow more efficient setting of hint bits.
+ * avoids a cross-translation-unit function call for each tuple, allows the
+ * compiler to optimize across calls to HeapTupleSatisfiesMVCC and allows
+ * setting hint bits more efficiently (see the one BufferFinishSetHintBits()
+ * call below).
  *
  * Returns the number of visible tuples.
  */
@@ -1620,6 +1692,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 							OffsetNumber *vistuples_dense)
 {
 	int			nvis = 0;
+	SetHintBitsState state = SHB_INITIAL;
 
 	Assert(IsMVCCSnapshot(snapshot));
 
@@ -1628,7 +1701,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		bool		valid;
 		HeapTuple	tup = &batchmvcc->tuples[i];
 
-		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer, &state);
 		batchmvcc->visible[i] = valid;
 
 		if (likely(valid))
@@ -1638,6 +1711,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		}
 	}
 
+	if (state == SHB_ENABLED)
+		BufferFinishSetHintBits(buffer, true, true);
+
 	return nvis;
 }
 
@@ -1657,7 +1733,7 @@ HeapTupleSatisfiesVisibility(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 	switch (snapshot->snapshot_type)
 	{
 		case SNAPSHOT_MVCC:
-			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer);
+			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer, NULL);
 		case SNAPSHOT_SELF:
 			return HeapTupleSatisfiesSelf(htup, snapshot, buffer);
 		case SNAPSHOT_ANY:
diff --git a/src/backend/access/nbtree/nbtinsert.c b/src/backend/access/nbtree/nbtinsert.c
index 63eda08f7a2..da43af3ec96 100644
--- a/src/backend/access/nbtree/nbtinsert.c
+++ b/src/backend/access/nbtree/nbtinsert.c
@@ -681,20 +681,28 @@ _bt_check_unique(Relation rel, BTInsertState insertstate, Relation heapRel,
 				{
 					/*
 					 * The conflicting tuple (or all HOT chains pointed to by
-					 * all posting list TIDs) is dead to everyone, so mark the
-					 * index entry killed.
+					 * all posting list TIDs) is dead to everyone, so try to
+					 * mark the index entry killed. It's ok if we're not
+					 * allowed to, this isn't required for correctness.
 					 */
-					ItemIdMarkDead(curitemid);
-					opaque->btpo_flags |= BTP_HAS_GARBAGE;
+					Buffer		buf;
 
-					/*
-					 * Mark buffer with a dirty hint, since state is not
-					 * crucial. Be sure to mark the proper buffer dirty.
-					 */
+					/* Be sure to operate on the proper buffer */
 					if (nbuf != InvalidBuffer)
-						MarkBufferDirtyHint(nbuf, true);
+						buf = nbuf;
 					else
-						MarkBufferDirtyHint(insertstate->buf, true);
+						buf = insertstate->buf;
+
+					/*
+					 * Can't use BufferSetHintBits16() here as we update two
+					 * different locations.
+					 */
+					if (BufferBeginSetHintBits(buf))
+					{
+						ItemIdMarkDead(curitemid);
+						opaque->btpo_flags |= BTP_HAS_GARBAGE;
+						BufferFinishSetHintBits(buf, true, true);
+					}
 				}
 
 				/*
diff --git a/src/backend/access/nbtree/nbtutils.c b/src/backend/access/nbtree/nbtutils.c
index 5c50f0dd1bd..a76d90f2d8e 100644
--- a/src/backend/access/nbtree/nbtutils.c
+++ b/src/backend/access/nbtree/nbtutils.c
@@ -357,10 +357,19 @@ _bt_killitems(IndexScanDesc scan)
 			 * it's possible that multiple processes attempt to do this
 			 * simultaneously, leading to multiple full-page images being sent
 			 * to WAL (if wal_log_hints or data checksums are enabled), which
-			 * is undesirable.
+			 * is undesirable.  We need to use the hint bit infrastructure to
+			 * update the page while just holding a share lock.
 			 */
 			if (killtuple && !ItemIdIsDead(iid))
 			{
+				/*
+				 * If we're not able to set hint bits, there's no point
+				 * continuing.
+				 */
+				if (!killedsomething &&
+					!BufferBeginSetHintBits(buf))
+					goto unlock_page;
+
 				/* found the item/all posting list items */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -371,8 +380,6 @@ _bt_killitems(IndexScanDesc scan)
 	}
 
 	/*
-	 * Since this can be redone later if needed, mark as dirty hint.
-	 *
 	 * Whenever we mark anything LP_DEAD, we also set the page's
 	 * BTP_HAS_GARBAGE flag, which is likewise just a hint.  (Note that we
 	 * only rely on the page-level flag in !heapkeyspace indexes.)
@@ -380,9 +387,10 @@ _bt_killitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->btpo_flags |= BTP_HAS_GARBAGE;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(buf, true, true);
 	}
 
+unlock_page:
 	if (!so->dropPin)
 		_bt_unlockbuf(rel, buf);
 	else
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index 119f31b5d65..b3ff5a0e441 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -25,21 +25,26 @@ that might need to do such a wait is instead handled by waiting to obtain
 the relation-level lock, which is why you'd better hold one first.)  Pins
 may not be held across transaction boundaries, however.
 
-Buffer content locks: there are two kinds of buffer lock, shared and exclusive,
-which act just as you'd expect: multiple backends can hold shared locks on
-the same buffer, but an exclusive lock prevents anyone else from holding
-either shared or exclusive lock.  (These can alternatively be called READ
-and WRITE locks.)  These locks are intended to be short-term: they should not
-be held for long.  Buffer locks are acquired and released by LockBuffer().
-It will *not* work for a single backend to try to acquire multiple locks on
-the same buffer.  One must pin a buffer before trying to lock it.
+Buffer content locks: there are three kinds of buffer lock, shared,
+share-exclusive and exclusive:
+a) multiple backends can hold shared locks on the same buffer
+   (alternatively called a READ lock)
+b) one backend can hold a share-exclusive lock on a buffer while multiple
+   backends can hold a share lock
+c) an exclusive lock prevents anyone else from holding shared, share-exclusive
+   or exclusive lock.
+   (alternatively called a WRITE lock)
+
+These locks are intended to be short-term: they should not be held for long.
+Buffer locks are acquired and released by LockBuffer().  It will *not* work
+for a single backend to try to acquire multiple locks on the same buffer.  One
+must pin a buffer before trying to lock it.
 
 Buffer access rules:
 
-1. To scan a page for tuples, one must hold a pin and either shared or
-exclusive content lock.  To examine the commit status (XIDs and status bits)
-of a tuple in a shared buffer, one must likewise hold a pin and either shared
-or exclusive lock.
+1. To scan a page for tuples, one must hold a pin and at least a share lock.
+To examine the commit status (XIDs and status bits) of a tuple in a shared
+buffer, one must likewise hold a pin and at least a share lock.
 
 2. Once one has determined that a tuple is interesting (visible to the
 current transaction) one may drop the content lock, yet continue to access
@@ -55,19 +60,25 @@ one must hold a pin and an exclusive content lock on the containing buffer.
 This ensures that no one else might see a partially-updated state of the
 tuple while they are doing visibility checks.
 
-4. It is considered OK to update tuple commit status bits (ie, OR the
-values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
-HEAP_XMAX_INVALID into t_infomask) while holding only a shared lock and
-pin on a buffer.  This is OK because another backend looking at the tuple
-at about the same time would OR the same bits into the field, so there
-is little or no risk of conflicting update; what's more, if there did
-manage to be a conflict it would merely mean that one bit-update would
-be lost and need to be done again later.  These four bits are only hints
-(they cache the results of transaction status lookups in pg_xact), so no
-great harm is done if they get reset to zero by conflicting updates.
-Note, however, that a tuple is frozen by setting both HEAP_XMIN_INVALID
-and HEAP_XMIN_COMMITTED; this is a critical update and accordingly requires
-an exclusive buffer lock (and it must also be WAL-logged).
+4. Non-critical information on a page ("hint bits") may be modified while
+holding only a share-exclusive lock and pin on the page. To do so in cases
+where only a share lock is already held, use BufferBeginSetHintBits() &
+BufferFinishSetHintBits() (if multiple hint bits are to be set) or
+BufferSetHintBits16() (if a single hint bit is set).
+
+E.g. for heapam, a share-exclusive lock allows to update tuple commit status
+bits (ie, OR the values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID,
+HEAP_XMAX_COMMITTED, or HEAP_XMAX_INVALID into t_infomask) while holding only
+a share-exclusive lock and pin on a buffer.  This is OK because another
+backend looking at the tuple at about the same time would OR the same bits
+into the field, so there is little or no risk of conflicting update; what's
+more, if there did manage to be a conflict it would merely mean that one
+bit-update would be lost and need to be done again later.  These four bits are
+only hints (they cache the results of transaction status lookups in pg_xact),
+so no great harm is done if they get reset to zero by conflicting updates.
+Note, however, that a tuple is frozen by setting both HEAP_XMIN_INVALID and
+HEAP_XMIN_COMMITTED; this is a critical update and accordingly requires an
+exclusive buffer lock (and it must also be WAL-logged).
 
 5. To physically remove a tuple or compact free space on a page, one
 must hold a pin and an exclusive lock, *and* observe while holding the
@@ -80,7 +91,6 @@ buffer (increment the refcount) while one is performing the cleanup, but
 it won't be able to actually examine the page until it acquires shared
 or exclusive content lock.
 
-
 Obtaining the lock needed under rule #5 is done by the bufmgr routines
 LockBufferForCleanup() or ConditionalLockBufferForCleanup().  They first get
 an exclusive lock and then check to see if the shared pin count is currently
@@ -96,6 +106,10 @@ VACUUM's use, since we don't allow multiple VACUUMs concurrently on a single
 relation anyway.  Anyone wishing to obtain a cleanup lock outside of recovery
 or a VACUUM must use the conditional variant of the function.
 
+6. To write out a buffer, a share-exclusive lock needs to be held. This
+prevents the buffer from being modified while written out, which could corrupt
+checksums and cause issues on the OS or device level when direct-IO is used.
+
 
 Buffer Manager's Internal Locking
 ---------------------------------
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 6adf04903cb..402bf81b269 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2480,9 +2480,8 @@ again:
 	/*
 	 * If the buffer was dirty, try to write it out.  There is a race
 	 * condition here, in that someone might dirty it after we released the
-	 * buffer header lock above, or even while we are writing it out (since
-	 * our share-lock won't prevent hint-bit updates).  We will recheck the
-	 * dirty bit after re-locking the buffer header.
+	 * buffer header lock above.  We will recheck the dirty bit after
+	 * re-locking the buffer header.
 	 */
 	if (buf_state & BM_DIRTY)
 	{
@@ -2490,20 +2489,20 @@ again:
 		Assert(buf_state & BM_VALID);
 
 		/*
-		 * We need a share-lock on the buffer contents to write it out (else
-		 * we might write invalid data, eg because someone else is compacting
-		 * the page contents while we write).  We must use a conditional lock
-		 * acquisition here to avoid deadlock.  Even though the buffer was not
-		 * pinned (and therefore surely not locked) when StrategyGetBuffer
-		 * returned it, someone else could have pinned and exclusive-locked it
-		 * by the time we get here. If we try to get the lock unconditionally,
-		 * we'd block waiting for them; if they later block waiting for us,
-		 * deadlock ensues. (This has been observed to happen when two
-		 * backends are both trying to split btree index pages, and the second
-		 * one just happens to be trying to split the page the first one got
-		 * from StrategyGetBuffer.)
+		 * We need a share-exclusive lock on the buffer contents to write it
+		 * out (else we might write invalid data, eg because someone else is
+		 * compacting the page contents while we write).  We must use a
+		 * conditional lock acquisition here to avoid deadlock.  Even though
+		 * the buffer was not pinned (and therefore surely not locked) when
+		 * StrategyGetBuffer returned it, someone else could have pinned and
+		 * (share-)exclusive-locked it by the time we get here. If we try to
+		 * get the lock unconditionally, we'd block waiting for them; if they
+		 * later block waiting for us, deadlock ensues. (This has been
+		 * observed to happen when two backends are both trying to split btree
+		 * index pages, and the second one just happens to be trying to split
+		 * the page the first one got from StrategyGetBuffer.)
 		 */
-		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -4072,8 +4071,8 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	}
 
 	/*
-	 * Pin it, share-lock it, write it.  (FlushBuffer will do nothing if the
-	 * buffer is clean by the time we've locked it.)
+	 * Pin it, share-exclusive-lock it, write it.  (FlushBuffer will do
+	 * nothing if the buffer is clean by the time we've locked it.)
 	 */
 	PinBuffer_Locked(bufHdr);
 
@@ -4403,11 +4402,8 @@ BufferGetTag(Buffer buffer, RelFileLocator *rlocator, ForkNumber *forknum,
  * However, we will need to force the changes to disk via fsync before
  * we can checkpoint WAL.
  *
- * The caller must hold a pin on the buffer and have share-locked the
- * buffer contents.  (Note: a share-lock does not prevent updates of
- * hint bits in the buffer, so the page could change while the write
- * is in progress, but we assume that that will not invalidate the data
- * written.)
+ * The caller must hold a pin on the buffer and have
+ * (share-)exclusively-locked the buffer contents.
  *
  * If the caller has an smgr reference for the buffer's relation, pass it
  * as the second parameter.  If not, pass NULL.
@@ -4423,6 +4419,9 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	char	   *bufToWrite;
 	uint64		buf_state;
 
+	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(buf, BUFFER_LOCK_SHARE_EXCLUSIVE));
+
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
 	 * someone else flushed the buffer before we could, so we need not do
@@ -4555,7 +4554,7 @@ FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(buf);
 
-	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 	BufferLockUnlock(buffer, buf);
 }
@@ -5474,8 +5473,8 @@ FlushDatabaseBuffers(Oid dbid)
 }
 
 /*
- * Flush a previously, shared or exclusively, locked and pinned buffer to the
- * OS.
+ * Flush a previously, share-exclusively or exclusively, locked and pinned
+ * buffer to the OS.
  */
 void
 FlushOneBuffer(Buffer buffer)
@@ -5548,39 +5547,24 @@ IncrBufferRefCount(Buffer buffer)
 }
 
 /*
- * MarkBufferDirtyHint
+ * Shared-buffer only helper for MarkBufferDirtyHint() and
+ * BufferSetHintBits16().
  *
- *	Mark a buffer dirty for non-critical changes.
- *
- * This is essentially the same as MarkBufferDirty, except:
- *
- * 1. The caller does not write WAL; so if checksums are enabled, we may need
- *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
- * 2. The caller might have only share-lock instead of exclusive-lock on the
- *	  buffer's content lock.
- * 3. This function does not guarantee that the buffer is always marked dirty
- *	  (due to a race condition), so it cannot be used for important changes.
+ * This is separated out because it turns out that the repeated checks for
+ * local buffers, repeated GetBufferDescriptor() and repeated reading of the
+ * buffer's state sufficiently hurts the performance of BufferSetHintBits16().
  */
-void
-MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+static inline void
+MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate,
+						  bool buffer_std)
 {
-	BufferDesc *bufHdr;
 	Page		page = BufferGetPage(buffer);
 
-	if (!BufferIsValid(buffer))
-		elog(ERROR, "bad buffer ID: %d", buffer);
-
-	if (BufferIsLocal(buffer))
-	{
-		MarkLocalBufferDirty(buffer);
-		return;
-	}
-
-	bufHdr = GetBufferDescriptor(buffer - 1);
-
 	Assert(GetPrivateRefCount(buffer) > 0);
-	/* here, either share or exclusive lock is OK */
-	Assert(BufferIsLockedByMe(buffer));
+
+	/* here, either share-exclusive or exclusive lock is OK */
+	Assert(BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_SHARE_EXCLUSIVE));
 
 	/*
 	 * This routine might get called many times on the same page, if we are
@@ -5593,8 +5577,8 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	 * is only intended to be used in cases where failing to write out the
 	 * data would be harmless anyway, it doesn't really matter.
 	 */
-	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
-		(BM_DIRTY | BM_JUST_DIRTIED))
+	if (unlikely((lockstate & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+				 (BM_DIRTY | BM_JUST_DIRTIED)))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		dirtied = false;
@@ -5610,8 +5594,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * We don't check full_page_writes here because that logic is included
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
-		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
+		if (XLogHintBitIsNeeded() && (lockstate & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5663,13 +5646,13 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 			dirtied = true;		/* Means "will be dirtied by this action" */
 
 			/*
-			 * Set the page LSN if we wrote a backup block. We aren't supposed
-			 * to set this when only holding a share lock but as long as we
-			 * serialise it somehow we're OK. We choose to set LSN while
-			 * holding the buffer header lock, which causes any reader of an
-			 * LSN who holds only a share lock to also obtain a buffer header
-			 * lock before using PageGetLSN(), which is enforced in
-			 * BufferGetLSNAtomic().
+			 * Set the page LSN if we wrote a backup block. To allow backends
+			 * that only hold a share lock on the buffer to read the LSN in a
+			 * tear-free manner, we set the page LSN while holding the buffer
+			 * header lock. This allows any reader of an LSN who holds only a
+			 * share lock to also obtain a buffer header lock before using
+			 * PageGetLSN() to read the LSN in a tear free way. This is done
+			 * in BufferGetLSNAtomic().
 			 *
 			 * If checksums are enabled, you might think we should reset the
 			 * checksum here. That will happen when the page is written
@@ -5695,6 +5678,41 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 	}
 }
 
+/*
+ * MarkBufferDirtyHint
+ *
+ *	Mark a buffer dirty for non-critical changes.
+ *
+ * This is essentially the same as MarkBufferDirty, except:
+ *
+ * 1. The caller does not write WAL; so if checksums are enabled, we may need
+ *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
+ * 2. The caller might have only a share-exclusive-lock instead of an
+ *	  exclusive-lock on the buffer's content lock.
+ * 3. This function does not guarantee that the buffer is always marked dirty
+ *	  (due to a race condition), so it cannot be used for important changes.
+ */
+inline void
+MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+{
+	BufferDesc *bufHdr;
+
+	bufHdr = GetBufferDescriptor(buffer - 1);
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		MarkLocalBufferDirty(buffer);
+		return;
+	}
+
+	MarkSharedBufferDirtyHint(buffer, bufHdr,
+							  pg_atomic_read_u64(&bufHdr->state),
+							  buffer_std);
+}
+
 /*
  * Release buffer content locks for shared buffers.
  *
@@ -6789,6 +6807,187 @@ IsBufferCleanupOK(Buffer buffer)
 	return false;
 }
 
+/*
+ * Helper for BufferBeginSetHintBits() and BufferSetHintBits16().
+ *
+ * This checks if the current lock mode already suffices to allow hint bits
+ * being set and, if not, whether the current lock can be upgraded.
+ */
+static inline bool
+SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
+{
+	uint64		old_state;
+	PrivateRefCountEntry *ref;
+	BufferLockMode mode;
+
+	ref = GetPrivateRefCountEntry(buffer, true);
+
+	if (ref == NULL)
+		elog(ERROR, "lock is not held");
+
+	mode = ref->data.lockmode;
+	if (mode == BUFFER_LOCK_UNLOCK)
+		elog(ERROR, "buffer is not locked");
+
+	/*
+	 * Already am holding a sufficient lock level.
+	 */
+	if (mode == BUFFER_LOCK_EXCLUSIVE || mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+	{
+		*lockstate = pg_atomic_read_u64(&buf_hdr->state);
+		return true;
+	}
+
+	/*
+	 * Only holding a share lock right now, try to upgrade to SHARE_EXCLUSIVE.
+	 */
+	Assert(mode == BUFFER_LOCK_SHARE);
+
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+	while (true)
+	{
+		uint64		desired_state;
+
+		desired_state = old_state;
+
+		/*
+		 * Can't upgrade if somebody else holds the lock in exclusive or
+		 * share-exclusive mode.
+		 */
+		if (unlikely((old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) != 0))
+		{
+			return false;
+		}
+
+		/* currently held lock state */
+		desired_state -= BM_LOCK_VAL_SHARED;
+
+		/* new lock level */
+		desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			ref->data.lockmode = BUFFER_LOCK_SHARE_EXCLUSIVE;
+			*lockstate = desired_state;
+
+			return true;
+		}
+	}
+
+}
+
+/*
+ * Try to acquire the right to set hint bits on the buffer.
+ *
+ * To be allowed to set hint bits, this backend needs to hold either a
+ * share-exclusive or an exclusive lock. In case this backend only holds a
+ * share lock, this function will try to upgrade the lock to
+ * share-exclusive. The caller is only allowed to set hint bits if true is
+ * returned.
+ *
+ * Once BufferBeginSetHintBits() has returned true, hint bits may be set
+ * without further calls to BufferBeginSetHintBits(), until the buffer is
+ * unlocked.
+ *
+ *
+ * Requiring a share-exclusive lock to set hint bits prevents setting hint
+ * bits on buffers that are currently being written out, which could corrupt
+ * the checksum on the page. Flushing buffers also requires a share-exclusive
+ * lock.
+ *
+ * Due to a lock >= share-exclusive being required to set hint bits, only one
+ * backend can set hint bits at a time. Allowing multiple backends to hint
+ * bits would require more complicated locking: For setting hint bits we'd
+ * need to store the count of backends currently setting hint bits, for I/O we
+ * would need another lock-level conflicting with the hint-setting
+ * lock-level. Given that the share-exclusive lock for setting hint bits is
+ * only held for a short time, that backends often would just set the same
+ * hint bits and that the cost of occasionally not setting hint bits in hotly
+ * accessed pages is fairly low, this seems like an acceptable tradeoff.
+ */
+bool
+BufferBeginSetHintBits(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+
+	if (BufferIsLocal(buffer))
+	{
+		/*
+		 * NB: Will need to check if there is a write in progress, once it is
+		 * possible for writes to be done asynchronously.
+		 */
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	return SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate);
+}
+
+/*
+ * End a phase of setting hint bits on this buffer, started with
+ * BufferBeginSetHintBits().
+ *
+ * This would strictly speaking not be required (i.e. the caller could do
+ * MarkBufferDirtyHint() if so desired), but allows us to perform some sanity
+ * checks.
+ */
+void
+BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std)
+{
+	if (!BufferIsLocal(buffer))
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+
+	if (mark_dirty)
+		MarkBufferDirtyHint(buffer, buffer_std);
+}
+
+/*
+ * Try to set a single hint bit in a buffer.
+ *
+ * This is a bit faster than BufferBeginSetHintBits() /
+ * BufferFinishSetHintBits() when setting a single hint bit, but slower than
+ * the former when setting several hint bits.
+ */
+bool
+BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+#ifdef USE_ASSERT_CHECKING
+	char	   *page;
+
+	/* verify that the address is on the page */
+	page = BufferGetPage(buffer);
+	Assert((char *) ptr >= page && (char *) ptr < (page + BLCKSZ));
+#endif
+
+	if (BufferIsLocal(buffer))
+	{
+		*ptr = val;
+
+		MarkLocalBufferDirty(buffer);
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	if (SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate))
+	{
+		*ptr = val;
+
+		MarkSharedBufferDirtyHint(buffer, buf_hdr, lockstate, true);
+
+		return true;
+	}
+
+	return false;
+}
+
 
 /*
  *	Functions for buffer I/O handling
diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index ad337c00871..b9a8f368a63 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -904,13 +904,17 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 	max_avail = fsm_get_max_avail(page);
 
 	/*
-	 * Reset the next slot pointer. This encourages the use of low-numbered
-	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation. We don't bother with marking the page dirty if it wasn't
-	 * already, since this is just a hint.
+	 * Try to reset the next slot pointer. This encourages the use of
+	 * low-numbered pages, increasing the chances that a later vacuum can
+	 * truncate the relation. We don't bother with marking the page dirty if
+	 * it wasn't already, since this is just a hint.
 	 */
 	LockBuffer(buf, BUFFER_LOCK_SHARE);
-	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	if (BufferBeginSetHintBits(buf))
+	{
+		((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+		BufferFinishSetHintBits(buf, false, false);
+	}
 	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 	ReleaseBuffer(buf);
diff --git a/src/backend/storage/freespace/fsmpage.c b/src/backend/storage/freespace/fsmpage.c
index 33ee825529c..e46bf2631fc 100644
--- a/src/backend/storage/freespace/fsmpage.c
+++ b/src/backend/storage/freespace/fsmpage.c
@@ -298,9 +298,18 @@ restart:
 	 * lock and get a garbled next pointer every now and then, than take the
 	 * concurrency hit of an exclusive lock.
 	 *
+	 * Without an exclusive lock, we need to use the hint bit infrastructure
+	 * to be allowed to modify the page.
+	 *
 	 * Wrap-around is handled at the beginning of this function.
 	 */
-	fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+	if (exclusive_lock_held || BufferBeginSetHintBits(buf))
+	{
+		fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+
+		if (!exclusive_lock_held)
+			BufferFinishSetHintBits(buf, false, true);
+	}
 
 	return slot;
 }
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 14dec2d49c1..efea48fcef7 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2750,6 +2750,7 @@ SetConstraintStateData
 SetConstraintTriggerData
 SetExprState
 SetFunctionReturnMode
+SetHintBitsState
 SetOp
 SetOpCmd
 SetOpPath
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v11-0006-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch (11.6K, ../../k5j77f3q6ztihnjnx2nqxzyor6fbj2qxcbhzuxhkh2yy63jyfg@p72phigar3n4/7-v11-0006-WIP-bufmgr-Don-t-copy-pages-while-writing-out.patch)
  download | inline diff:
From 3e8707ee794245aee211b7019e512db3ab6da214 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 13 Jan 2026 20:10:32 -0500
Subject: [PATCH v11 6/7] WIP: bufmgr: Don't copy pages while writing out

After the series of preceding commits introducing and using
BufferBeginSetHintBits()/BufferSetHintBits16() hint bits are not set
anymore while IO is going on. Therefore we do not need to copy pages while
they are being written out anymore.

TODO: Update comments

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/bufpage.h           |  3 +-
 src/backend/access/hash/hashpage.c      |  2 +-
 src/backend/access/transam/xloginsert.c | 43 ++++++----------------
 src/backend/storage/buffer/bufmgr.c     | 21 +++++------
 src/backend/storage/buffer/localbuf.c   |  2 +-
 src/backend/storage/page/bufpage.c      | 48 ++++---------------------
 src/backend/storage/smgr/bulk_write.c   |  2 +-
 src/test/modules/test_aio/test_aio.c    |  2 +-
 8 files changed, 33 insertions(+), 90 deletions(-)

diff --git a/src/include/storage/bufpage.h b/src/include/storage/bufpage.h
index ae3725b3b81..31ec9a8a047 100644
--- a/src/include/storage/bufpage.h
+++ b/src/include/storage/bufpage.h
@@ -504,7 +504,6 @@ extern void PageIndexMultiDelete(Page page, OffsetNumber *itemnos, int nitems);
 extern void PageIndexTupleDeleteNoCompact(Page page, OffsetNumber offnum);
 extern bool PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 									const void *newtup, Size newsize);
-extern char *PageSetChecksumCopy(Page page, BlockNumber blkno);
-extern void PageSetChecksumInplace(Page page, BlockNumber blkno);
+extern void PageSetChecksum(Page page, BlockNumber blkno);
 
 #endif							/* BUFPAGE_H */
diff --git a/src/backend/access/hash/hashpage.c b/src/backend/access/hash/hashpage.c
index 8e220a3ae16..52c20208c66 100644
--- a/src/backend/access/hash/hashpage.c
+++ b/src/backend/access/hash/hashpage.c
@@ -1029,7 +1029,7 @@ _hash_alloc_buckets(Relation rel, BlockNumber firstblock, uint32 nblocks)
 					zerobuf.data,
 					true);
 
-	PageSetChecksumInplace(page, lastblock);
+	PageSetChecksum(page, lastblock);
 	smgrextend(RelationGetSmgr(rel), MAIN_FORKNUM, lastblock, zerobuf.data,
 			   false);
 
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index 92c48e768c3..53cfdce8de8 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -261,8 +261,11 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	 */
 #ifdef USE_ASSERT_CHECKING
 	if (!(flags & REGBUF_NO_CHANGE))
-		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) &&
-			   BufferIsDirty(buffer));
+	{
+		Assert(BufferIsDirty(buffer));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE));
+	}
 #endif
 
 	if (block_id >= max_registered_block_id)
@@ -1066,7 +1069,7 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * Write a backup block if needed when we are setting a hint. Note that
  * this may be called for a variety of page types, not just heaps.
  *
- * Callable while holding just share lock on the buffer content.
+ * Callable while holding just a share-exclusive lock on the buffer content.
  *
  * We can't use the plain backup block mechanism since that relies on the
  * Buffer being exclusively locked. Since some modifications (setting LSN, hint
@@ -1074,6 +1077,8 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * failures. So instead we copy the page and insert the copied data as normal
  * record data.
  *
+ * FIXME: outdated
+ *
  * We only need to do something if page has not yet been full page written in
  * this checkpoint round. The LSN of the inserted wal record is returned if we
  * had to write, InvalidXLogRecPtr otherwise.
@@ -1102,46 +1107,20 @@ XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
 
 	/*
 	 * We assume page LSN is first data on *every* page that can be passed to
-	 * XLogInsert, whether it has the standard page layout or not. Since we're
-	 * only holding a share-lock on the page, we must take the buffer header
-	 * lock when we look at the LSN.
+	 * XLogInsert, whether it has the standard page layout or not.
 	 */
 	lsn = BufferGetLSNAtomic(buffer);
 
 	if (lsn <= RedoRecPtr)
 	{
-		int			flags = 0;
-		PGAlignedBlock copied_buffer;
-		char	   *origdata = (char *) BufferGetBlock(buffer);
-		RelFileLocator rlocator;
-		ForkNumber	forkno;
-		BlockNumber blkno;
-
-		/*
-		 * Copy buffer so we don't have to worry about concurrent hint bit or
-		 * lsn updates. We assume pd_lower/upper cannot be changed without an
-		 * exclusive lock, so the contents bkp are not racy.
-		 */
-		if (buffer_std)
-		{
-			/* Assume we can omit data between pd_lower and pd_upper */
-			Page		page = BufferGetPage(buffer);
-			uint16		lower = ((PageHeader) page)->pd_lower;
-			uint16		upper = ((PageHeader) page)->pd_upper;
-
-			memcpy(copied_buffer.data, origdata, lower);
-			memcpy(copied_buffer.data + upper, origdata + upper, BLCKSZ - upper);
-		}
-		else
-			memcpy(copied_buffer.data, origdata, BLCKSZ);
+		int			flags = REGBUF_NO_CHANGE;
 
 		XLogBeginInsert();
 
 		if (buffer_std)
 			flags |= REGBUF_STANDARD;
 
-		BufferGetTag(buffer, &rlocator, &forkno, &blkno);
-		XLogRegisterBlock(0, &rlocator, forkno, blkno, copied_buffer.data, flags);
+		XLogRegisterBuffer(0, buffer, flags);
 
 		recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI_FOR_HINT);
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 402bf81b269..097c8fa67c5 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -4416,7 +4416,6 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	ErrorContextCallback errcallback;
 	instr_time	io_start;
 	Block		bufBlock;
-	char	   *bufToWrite;
 	uint64		buf_state;
 
 	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
@@ -4487,12 +4486,8 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	 */
 	bufBlock = BufHdrGetBlock(buf);
 
-	/*
-	 * Update page checksum if desired.  Since we have only shared lock on the
-	 * buffer, other processes might be updating hint bits in it, so we must
-	 * copy the page to private storage if we do checksumming.
-	 */
-	bufToWrite = PageSetChecksumCopy((Page) bufBlock, buf->tag.blockNum);
+	/* Update page checksum if desired. */
+	PageSetChecksum((Page) bufBlock, buf->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
@@ -4502,7 +4497,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	smgrwrite(reln,
 			  BufTagGetForkNum(&buf->tag),
 			  buf->tag.blockNum,
-			  bufToWrite,
+			  bufBlock,
 			  false);
 
 	/*
@@ -4626,8 +4621,8 @@ BufferIsPermanent(Buffer buffer)
 /*
  * BufferGetLSNAtomic
  *		Retrieves the LSN of the buffer atomically using a buffer header lock.
- *		This is necessary for some callers who may not have an exclusive lock
- *		on the buffer.
+ *		This is necessary for some callers who may not have a (share-)exclusive
+ *		lock on the buffer.
  */
 XLogRecPtr
 BufferGetLSNAtomic(Buffer buffer)
@@ -5630,6 +5625,12 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate,
 			 * It's possible we may enter here without an xid, so it is
 			 * essential that CreateCheckPoint waits for virtual transactions
 			 * rather than full transactionids.
+			 *
+			 * FIXME: I think we now should simply mark the page dirty before
+			 * WAL logging the hint bit - afaict it then should work just like
+			 * any other buffer write (due to SyncBuffers()/SyncOneBuffer()
+			 * seeing the dirty bit and trying to lock the page
+			 * share-exclusive, and thus having to wait).
 			 */
 			Assert((MyProc->delayChkptFlags & DELAY_CHKPT_START) == 0);
 			MyProc->delayChkptFlags |= DELAY_CHKPT_START;
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 04a540379a2..55e17e03acb 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -199,7 +199,7 @@ FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln)
 		reln = smgropen(BufTagGetRelFileLocator(&bufHdr->tag),
 						MyProcNumber);
 
-	PageSetChecksumInplace(localpage, bufHdr->tag.blockNum);
+	PageSetChecksum(localpage, bufHdr->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index de85911e3ac..2072bb1c72c 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -1494,51 +1494,15 @@ PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 /*
  * Set checksum for a page in shared buffers.
  *
- * If checksums are disabled, or if the page is not initialized, just return
- * the input.  Otherwise, we must make a copy of the page before calculating
- * the checksum, to prevent concurrent modifications (e.g. setting hint bits)
- * from making the final checksum invalid.  It doesn't matter if we include or
- * exclude hints during the copy, as long as we write a valid page and
- * associated checksum.
+ * If checksums are disabled, or if the page is not initialized, just
+ * return. Otherwise compute and set the checksum.
  *
- * Returns a pointer to the block-sized data that needs to be written. Uses
- * statically-allocated memory, so the caller must immediately write the
- * returned page and not refer to it again.
- */
-char *
-PageSetChecksumCopy(Page page, BlockNumber blkno)
-{
-	static char *pageCopy = NULL;
-
-	/* If we don't need a checksum, just return the passed-in data */
-	if (PageIsNew(page) || !DataChecksumsEnabled())
-		return page;
-
-	/*
-	 * We allocate the copy space once and use it over on each subsequent
-	 * call.  The point of palloc'ing here, rather than having a static char
-	 * array, is first to ensure adequate alignment for the checksumming code
-	 * and second to avoid wasting space in processes that never call this.
-	 */
-	if (pageCopy == NULL)
-		pageCopy = MemoryContextAllocAligned(TopMemoryContext,
-											 BLCKSZ,
-											 PG_IO_ALIGN_SIZE,
-											 0);
-
-	memcpy(pageCopy, page, BLCKSZ);
-	((PageHeader) pageCopy)->pd_checksum = pg_checksum_page(pageCopy, blkno);
-	return pageCopy;
-}
-
-/*
- * Set checksum for a page in private memory.
- *
- * This must only be used when we know that no other process can be modifying
- * the page buffer.
+ * In the past this needed to be done on a copy of the page, due to the
+ * possibility of e.g. hint bits being set concurrently. However, this is not
+ * necessary anymore as hint bits won't be set while IO is going on.
  */
 void
-PageSetChecksumInplace(Page page, BlockNumber blkno)
+PageSetChecksum(Page page, BlockNumber blkno)
 {
 	/* If we don't need a checksum, just return */
 	if (PageIsNew(page) || !DataChecksumsEnabled())
diff --git a/src/backend/storage/smgr/bulk_write.c b/src/backend/storage/smgr/bulk_write.c
index 36b28824ec8..f3c24082a69 100644
--- a/src/backend/storage/smgr/bulk_write.c
+++ b/src/backend/storage/smgr/bulk_write.c
@@ -279,7 +279,7 @@ smgr_bulk_flush(BulkWriteState *bulkstate)
 		BlockNumber blkno = pending_writes[i].blkno;
 		Page		page = pending_writes[i].buf->data;
 
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 
 		if (blkno >= bulkstate->relsize)
 		{
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index b1aa8af9ec0..2ae4a559fab 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -288,7 +288,7 @@ modify_rel_block(PG_FUNCTION_ARGS)
 	}
 	else
 	{
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 	}
 
 	smgrwrite(RelationGetSmgr(rel),
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v11-0007-WIP-Make-UnlockReleaseBuffer-more-efficient.patch (3.5K, ../../k5j77f3q6ztihnjnx2nqxzyor6fbj2qxcbhzuxhkh2yy63jyfg@p72phigar3n4/8-v11-0007-WIP-Make-UnlockReleaseBuffer-more-efficient.patch)
  download | inline diff:
From 7f263a752faf5017cceb98286c248dbb395b281c Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 13 Jan 2026 20:10:32 -0500
Subject: [PATCH v11 7/7] WIP: Make UnlockReleaseBuffer() more efficient

Now that the buffer content lock is implemented as part of BufferDesc.state,
releasing the lock and unpinning the buffer can be implemented as a single
atomic operation.

Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/nbtree/nbtpage.c | 22 +++++++++++-
 src/backend/storage/buffer/bufmgr.c | 52 ++++++++++++++++++++++++++++-
 2 files changed, 72 insertions(+), 2 deletions(-)

diff --git a/src/backend/access/nbtree/nbtpage.c b/src/backend/access/nbtree/nbtpage.c
index 4125c185e8b..f3e3f67e1fd 100644
--- a/src/backend/access/nbtree/nbtpage.c
+++ b/src/backend/access/nbtree/nbtpage.c
@@ -1007,11 +1007,18 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 
 	Assert(BlockNumberIsValid(blkno));
 	if (BufferIsValid(obuf))
+	{
+		_bt_relbuf(rel, obuf);
+#if 0
+		Assert(BufferGetBlockNumber(obuf) != blkno);
 		_bt_unlockbuf(rel, obuf);
-	buf = ReleaseAndReadBuffer(obuf, rel, blkno);
+#endif
+	}
+	buf = ReadBuffer(rel, blkno);
 	_bt_lockbuf(rel, buf, access);
 
 	_bt_checkpage(rel, buf);
+
 	return buf;
 }
 
@@ -1023,8 +1030,21 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 void
 _bt_relbuf(Relation rel, Buffer buf)
 {
+#if 0
 	_bt_unlockbuf(rel, buf);
 	ReleaseBuffer(buf);
+#else
+	/*
+	 * Buffer is pinned and locked, which means that it is expected to be
+	 * defined and addressable.  Check that proactively.
+	 */
+	VALGRIND_CHECK_MEM_IS_DEFINED(BufferGetPage(buf), BLCKSZ);
+
+	UnlockReleaseBuffer(buf);
+
+	if (!RelationUsesLocalBuffers(rel))
+		VALGRIND_MAKE_MEM_NOACCESS(BufferGetPage(buf), BLCKSZ);
+#endif
 }
 
 /*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 097c8fa67c5..a3a11595b5d 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5506,13 +5506,63 @@ ReleaseBuffer(Buffer buffer)
 /*
  * UnlockReleaseBuffer -- release the content lock and pin on a buffer
  *
- * This is just a shorthand for a common combination.
+ * This is just a, more efficient, shorthand for a common combination.
  */
 void
 UnlockReleaseBuffer(Buffer buffer)
 {
+#if 1
+	int			mode;
+	BufferDesc *buf;
+	PrivateRefCountEntry *ref;
+	uint64		sub;
+	uint64		lockstate;
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		UnpinLocalBuffer(buffer);
+		return;
+	}
+
+	ResourceOwnerForgetBuffer(CurrentResourceOwner, buffer);
+
+	buf = GetBufferDescriptor(buffer - 1);
+
+	mode = BufferLockDisownInternal(buffer, buf);
+
+	/* compute state modification for lock release */
+	sub = BufferLockReleaseSub(mode);
+
+	/* compute state modification for pin release */
+	ref = GetPrivateRefCountEntry(buffer, false);
+	Assert(ref != NULL);
+	Assert(ref->data.refcount > 0);
+	ref->data.refcount--;
+
+	if (ref->data.refcount == 0)
+	{
+		sub |= BUF_REFCOUNT_ONE;
+		ForgetPrivateRefCountEntry(ref);
+	}
+
+	/* perform the lock and pin release in one atomic op */
+	lockstate = pg_atomic_sub_fetch_u64(&buf->state, sub);
+
+	/* wake up waiters etc */
+	BufferLockProcessRelease(buf, mode, lockstate);
+
+	if (lockstate & BM_PIN_COUNT_WAITER)
+		WakePinCountWaiter(buf);
+
+	RESUME_INTERRUPTS();
+
+#else
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 	ReleaseBuffer(buffer);
+#endif
 }
 
 /*
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-02-02 22:33                                             ` Andres Freund <andres@anarazel.de>
  2026-02-06 21:18                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-07 12:38                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-07 12:59                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  0 siblings, 4 replies; 120+ messages in thread

From: Andres Freund @ 2026-02-02 22:33 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-01-14 16:20:58 -0500, Andres Freund wrote:
> I'm now working on cleaning up the last two commits. The most crucial bit is
> to simplify what happens in MarkSharedBufferDirtyHint(), we afaict can delete
> the use of DELAY_CHKPT_START etc and just go to marking the buffer dirty first
> and then do the WAL logging, just like normal WAL logging. The previous order
> was only required because we were dirtying the page while holding only a
> shared lock, which did not conflict with the lock held by SyncBuffers() etc.

I've been working on that.

- A lot of what was special about MarkBufferDirtyHint() isn't needed anymore:

  - The "abnormal" order of WAL logging before marking the buffer dirty was
    only needed because we marked buffers dirty. Which in turn was only needed
    because setting hint bits didn't conflict with flushing the page. With
    share-exclusive they do conflict, and we can switch to the normal order of
    operations, where marking a buffer dirty makes checkpoint wait when the
    buffer is encountered (due to wanting to flush the buffer but not getting
    the lock)


  - Now that we use the normal order of WAL logging, we don't need to delay
    checkpoint starts anymore.

    I think the explanation for why that is ok is correct [1], but it needs to
    be looked at by somebody with experience around this. Maybe Heikki?


  - Thanks to holding share-exclusive lock, nothing can concurrently dirty or
    undirty the buffer. Therefore the comments about spurious failures to mark
    the buffer dirty can be removed.


- I realized that, now that buffers cannot be dirtied while IO is ongoing, we
  don't need BM_JUST_DIRTIED anymore.


- The way MarkBufferDirtyHint() operates was copied into
  heap_inplace_update_and_unlock(). Now that MarkBufferDirtyHint() won't work
  that way anymore, it seems better to go with the alternative approach the
  comments already outlined, namely to only delay updating of the buffer
  contents.

  I've done this in a prequisite commit, as it doesn't actually depend on any
  of the other changes.  Noah, any chance you could take a look at this?


- Lots of minor polish


Greetings,

Andres Freund

[1]
	/*
	 * Update RedoRecPtr so that we can make the right decision. It's possible
	 * that a new checkpoint will start just after GetRedoRecPtr(), but that
	 * is ok, as the buffer is already dirty, ensuring that any BufferSync()
	 * started after the buffer was marked dirty cannot complete without
	 * flushing this buffer.  If a checkpoint started between marking the
	 * buffer dirty and this check, we will emit an unnecessary WAL record (as
	 * the buffer will be written out as part of the checkpoint), but the
	 * window for that is small.
	 */

Attachments:

  [text/x-diff] v12-0001-heapam-Don-t-mimic-MarkBufferDirtyHint-in-inplac.patch (4.0K, ../../5ubipyssiju5twkb7zgqwdr7q2vhpkpmuelxfpanetlk6ofnop@hvxb4g2amb2d/2-v12-0001-heapam-Don-t-mimic-MarkBufferDirtyHint-in-inplac.patch)
  download | inline diff:
From b871916ebc30ca69fbf61aa4f95394c407bcc1cd Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 2 Feb 2026 09:54:01 -0500
Subject: [PATCH v12 1/6] heapam: Don't mimic MarkBufferDirtyHint() in inplace
 updates

Previously heap_inplace_update_and_unlock() used an operation order similar to
MarkBufferDirty(), to reduce the number of different approaches used for
updating buffers.  However, in an upcoming patch, MarkBufferDirtyHint() will
switch to using the update protocol used by most other places (enabled by hint
bits only being set while holding a share-exclusive lock).

Luckily it's pretty easy to adjust heap_inplace_update_and_unlock(), as a
comment already foresaw, we can use the normal order with the slight change of
updating the buffer contents after WAL logging.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/heap/heapam.c | 34 ++++++++++++--------------------
 1 file changed, 13 insertions(+), 21 deletions(-)

diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 3004964ab7f..e387923b9bb 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -6611,11 +6611,11 @@ heap_inplace_update_and_unlock(Relation relation,
 	/*----------
 	 * NO EREPORT(ERROR) from here till changes are complete
 	 *
-	 * Our buffer lock won't stop a reader having already pinned and checked
-	 * visibility for this tuple.  Hence, we write WAL first, then mutate the
-	 * buffer.  Like in MarkBufferDirtyHint() or RecordTransactionCommit(),
-	 * checkpoint delay makes that acceptable.  With the usual order of
-	 * changes, a crash after memcpy() and before XLogInsert() could allow
+	 * Our exclusive buffer lock won't stop a reader having already pinned and
+	 * checked visibility for this tuple. With the usual order of changes
+	 * (i.e. updating the buffer contents before WAL logging), a reader could
+	 * observe our not-yet-persistent update to relfrozenxid and update
+	 * datfrozenxid based on that. A crash in that moment could allow
 	 * datfrozenxid to overtake relfrozenxid:
 	 *
 	 * ["D" is a VACUUM (ONLY_DATABASE_STATS)]
@@ -6627,21 +6627,16 @@ heap_inplace_update_and_unlock(Relation relation,
 	 * [crash]
 	 * [recovery restores datfrozenxid w/o relfrozenxid]
 	 *
-	 * Mimic MarkBufferDirtyHint() subroutine XLogSaveBufferForHint().
-	 * Specifically, use DELAY_CHKPT_START, and copy the buffer to the stack.
-	 * The stack copy facilitates a FPI of the post-mutation block before we
-	 * accept other sessions seeing it.  DELAY_CHKPT_START allows us to
-	 * XLogInsert() before MarkBufferDirty().  Since XLogSaveBufferForHint()
-	 * can operate under BUFFER_LOCK_SHARED, it can't avoid DELAY_CHKPT_START.
-	 * This function, however, likely could avoid it with the following order
-	 * of operations: MarkBufferDirty(), XLogInsert(), memcpy().  Opt to use
-	 * DELAY_CHKPT_START here, too, as a way to have fewer distinct code
-	 * patterns to analyze.  Inplace update isn't so frequent that it should
-	 * pursue the small optimization of skipping DELAY_CHKPT_START.
+	 * As we hold an exclusive lock - preventing the buffer from being written
+	 * out once dirty - we can work around this as follows: MarkBufferDirty(),
+	 * XLogInsert(), memcpy().
+	 *
+	 * That way any action a reader of the in-place-updated value takes will
+	 * be WAL logged after this change.
 	 */
-	Assert((MyProc->delayChkptFlags & DELAY_CHKPT_START) == 0);
 	START_CRIT_SECTION();
-	MyProc->delayChkptFlags |= DELAY_CHKPT_START;
+
+	MarkBufferDirty(buffer);
 
 	/* XLOG stuff */
 	if (RelationNeedsWAL(relation))
@@ -6690,8 +6685,6 @@ heap_inplace_update_and_unlock(Relation relation,
 
 	memcpy(dst, src, newlen);
 
-	MarkBufferDirty(buffer);
-
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 
 	/*
@@ -6700,7 +6693,6 @@ heap_inplace_update_and_unlock(Relation relation,
 	 */
 	AtInplace_Inval();
 
-	MyProc->delayChkptFlags &= ~DELAY_CHKPT_START;
 	END_CRIT_SECTION();
 	UnlockTuple(relation, &tuple->t_self, InplaceUpdateTupleLock);
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v12-0002-Require-share-exclusive-lock-to-set-hint-bits-an.patch (46.4K, ../../5ubipyssiju5twkb7zgqwdr7q2vhpkpmuelxfpanetlk6ofnop@hvxb4g2amb2d/3-v12-0002-Require-share-exclusive-lock-to-set-hint-bits-an.patch)
  download | inline diff:
From 74e4b1a13b0a7e98704ac31c40a3364569e7f260 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 13 Jan 2026 20:10:32 -0500
Subject: [PATCH v12 2/6] Require share-exclusive lock to set hint bits and to
 flush

At the moment hint bits can be set with just a share lock on a page (and,
until 45f658dacb9, in one case even without any lock). Because of this we need
to copy pages while writing them out, as otherwise the checksum could be
corrupted.

The need to copy the page is problematic to implement AIO writes:

1) Instead of just needing a single buffer for a copied page we need one for
   each page that's potentially undergoing I/O
2) To be able to use the "worker" AIO implementation the copied page needs to
   reside in shared memory

It also causes problems for using unbuffered/direct-IO, independent of AIO:
Some filesystems, raid implementations, ... do not tolerate the data being
written out to change during the write. E.g. they may compute internal
checksums that can be invalidated by concurrent modifications, leading e.g. to
filesystem errors (as the case with btrfs).

It also just is plain odd to allow modifications of buffers that are just
share locked.

To address these issues, this commit changes the rules so that modifications
to pages are not allowed anymore while holding a share lock. Instead the new
share-exclusive lock (introduced in fcb9c977aa5) allows at most one backend to
modify a buffer while other backends have the same page share locked. An
existing share-lock can be upgraded to a share-exclusive lock, if there are no
conflicting locks. For that BufferBeginSetHintBits()/BufferFinishSetHintBits()
and BufferSetHintBits16() have been introduced.

To prevent hint bits from being set while the buffer is being written out,
writing out buffers now requires a share-exclusive lock.

The use of share-exclusive to gate setting hint bits means that from now on
only one backend can set hint bits at a time. To allow multiple backends to
set hint bits would require more complicated locking: For setting hint bits
we'd need to store the count of backends currently setting hint bits and we
would need another lock-level for I/O conflicting with the lock-level to set
hint bits. Given that the share-exclusive lock for setting hint bits is only
held for a short time, that backends would often just set the same hint bits
and that the cost of occasionally not setting hint bits in hotly accessed
pages is fairly low, this seems like an acceptable tradeoff.

The biggest change to adapt to this is in heapam. To avoid performance
regressions for sequential scans that need to set a lot of hint bits, we need
to amortize the cost of BufferBeginSetHintBits() for cases where hint bits are
set at a high frequency, HeapTupleSatisfiesMVCCBatch() uses the new
SetHintBitsExt(), which defers BufferFinishSetHintBits() until all hint bits
on a page have been set.  Conversely, to avoid regressions in cases where we
can't set hint bits in bulk (because we're looking only at individual tuples),
use BufferSetHintBits16() when setting hint bits without batching.

Several other places also need to be adapted, but those changes are
comparatively simpler.

After this we do not need to copy buffers to write them out anymore. That
change is done separately however.

Discussion: https://postgr.es/m/fvfmkr5kk4nyex56ejgxj3uzi63isfxovp2biecb4bspbjrze7@az2pljabhnff
Discussion: https://postgr.es/m/stj36ea6yyhoxtqkhpieia2z4krnam7qyetc57rfezgk4zgapf%40gcnactj4z56m
---
 src/include/storage/bufmgr.h                |   4 +
 src/backend/access/gist/gistget.c           |  20 +-
 src/backend/access/hash/hashutil.c          |  14 +-
 src/backend/access/heap/heapam_visibility.c | 130 +++++--
 src/backend/access/nbtree/nbtinsert.c       |  31 +-
 src/backend/access/nbtree/nbtutils.c        |  16 +-
 src/backend/access/transam/xloginsert.c     |  11 +-
 src/backend/storage/buffer/README           |  66 ++--
 src/backend/storage/buffer/bufmgr.c         | 389 +++++++++++++++-----
 src/backend/storage/freespace/freespace.c   |  14 +-
 src/backend/storage/freespace/fsmpage.c     |  11 +-
 src/tools/pgindent/typedefs.list            |   1 +
 12 files changed, 525 insertions(+), 182 deletions(-)

diff --git a/src/include/storage/bufmgr.h b/src/include/storage/bufmgr.h
index a40adf6b2a8..4017896f951 100644
--- a/src/include/storage/bufmgr.h
+++ b/src/include/storage/bufmgr.h
@@ -314,6 +314,10 @@ extern void BufferGetTag(Buffer buffer, RelFileLocator *rlocator,
 
 extern void MarkBufferDirtyHint(Buffer buffer, bool buffer_std);
 
+extern bool BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer);
+extern bool BufferBeginSetHintBits(Buffer buffer);
+extern void BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std);
+
 extern void UnlockBuffers(void);
 extern void UnlockBuffer(Buffer buffer);
 extern void LockBufferInternal(Buffer buffer, BufferLockMode mode);
diff --git a/src/backend/access/gist/gistget.c b/src/backend/access/gist/gistget.c
index 11b214eb99b..606b108a136 100644
--- a/src/backend/access/gist/gistget.c
+++ b/src/backend/access/gist/gistget.c
@@ -64,11 +64,7 @@ gistkillitems(IndexScanDesc scan)
 	 * safe.
 	 */
 	if (BufferGetLSNAtomic(buffer) != so->curPageLSN)
-	{
-		UnlockReleaseBuffer(buffer);
-		so->numKilled = 0;		/* reset counter */
-		return;
-	}
+		goto unlock;
 
 	Assert(GistPageIsLeaf(page));
 
@@ -78,6 +74,17 @@ gistkillitems(IndexScanDesc scan)
 	 */
 	for (i = 0; i < so->numKilled; i++)
 	{
+		if (!killedsomething)
+		{
+			/*
+			 * Use the hint bit infrastructure to check if we can update the
+			 * page while just holding a share lock. If we are not allowed,
+			 * there's no point continuing.
+			 */
+			if (!BufferBeginSetHintBits(buffer))
+				goto unlock;
+		}
+
 		offnum = so->killedItems[i];
 		iid = PageGetItemId(page, offnum);
 		ItemIdMarkDead(iid);
@@ -87,9 +94,10 @@ gistkillitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		GistMarkPageHasGarbage(page);
-		MarkBufferDirtyHint(buffer, true);
+		BufferFinishSetHintBits(buffer, true, true);
 	}
 
+unlock:
 	UnlockReleaseBuffer(buffer);
 
 	/*
diff --git a/src/backend/access/hash/hashutil.c b/src/backend/access/hash/hashutil.c
index cf7f0b90176..3e16119d027 100644
--- a/src/backend/access/hash/hashutil.c
+++ b/src/backend/access/hash/hashutil.c
@@ -593,6 +593,17 @@ _hash_kill_items(IndexScanDesc scan)
 
 			if (ItemPointerEquals(&ituple->t_tid, &currItem->heapTid))
 			{
+				if (!killedsomething)
+				{
+					/*
+					 * Use the hint bit infrastructure to check if we can
+					 * update the page while just holding a share lock. If we
+					 * are not allowed, there's no point continuing.
+					 */
+					if (!BufferBeginSetHintBits(so->currPos.buf))
+						goto unlock_page;
+				}
+
 				/* found the item */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -610,9 +621,10 @@ _hash_kill_items(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->hasho_flag |= LH_PAGE_HAS_DEAD_TUPLES;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(so->currPos.buf, true, true);
 	}
 
+unlock_page:
 	if (so->hashso_bucket_buf == so->currPos.buf ||
 		havePin)
 		LockBuffer(so->currPos.buf, BUFFER_LOCK_UNLOCK);
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 75ae268d753..fc64f4343ce 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -80,10 +80,38 @@
 
 
 /*
- * SetHintBits()
+ * To be allowed to set hint bits, SetHintBits() needs to call
+ * BufferBeginSetHintBits(). However, that's not free, and some callsites call
+ * SetHintBits() on many tuples in a row. For those it makes sense to amortize
+ * the cost of BufferBeginSetHintBits(). Additionally it's desirable to defer
+ * the cost of BufferBeginSetHintBits() until a hint bit needs to actually be
+ * set. This enum serves as the necessary state space passed to
+ * SetHintBitsExt().
+ */
+typedef enum SetHintBitsState
+{
+	/* not yet checked if hint bits may be set */
+	SHB_INITIAL,
+	/* failed to get permission to set hint bits, don't check again */
+	SHB_DISABLED,
+	/* allowed to set hint bits */
+	SHB_ENABLED,
+} SetHintBitsState;
+
+/*
+ * SetHintBitsExt()
  *
  * Set commit/abort hint bits on a tuple, if appropriate at this time.
  *
+ * To be allowed to set a hint bit on a tuple, the page must not be undergoing
+ * IO at this time (otherwise we e.g. could corrupt PG's page checksum or even
+ * the filesystem's, as is known to happen with btrfs).
+ *
+ * The right to set a hint bit can be acquired on a page level with
+ * BufferBeginSetHintBits(). Only a single backend gets the right to set hint
+ * bits at a time.  Alternatively, if called with a NULL SetHintBitsState*,
+ * hint bits are set with BufferSetHintBits16().
+ *
  * It is only safe to set a transaction-committed hint bit if we know the
  * transaction's commit record is guaranteed to be flushed to disk before the
  * buffer, or if the table is temporary or unlogged and will be obliterated by
@@ -111,24 +139,67 @@
  * InvalidTransactionId if no check is needed.
  */
 static inline void
-SetHintBits(HeapTupleHeader tuple, Buffer buffer,
-			uint16 infomask, TransactionId xid)
+SetHintBitsExt(HeapTupleHeader tuple, Buffer buffer,
+			   uint16 infomask, TransactionId xid, SetHintBitsState *state)
 {
+	/*
+	 * In batched mode, if we previously did not get permission to set hint
+	 * bits, don't try again - in all likelihood IO is still going on.
+	 */
+	if (state && *state == SHB_DISABLED)
+		return;
+
 	if (TransactionIdIsValid(xid))
 	{
-		/* NB: xid must be known committed here! */
-		XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+		if (BufferIsPermanent(buffer))
+		{
+			/* NB: xid must be known committed here! */
+			XLogRecPtr	commitLSN = TransactionIdGetCommitLSN(xid);
+
+			if (XLogNeedsFlush(commitLSN) &&
+				BufferGetLSNAtomic(buffer) < commitLSN)
+			{
+				/* not flushed and no LSN interlock, so don't set hint */
+				return;
+			}
+		}
+	}
+
+	/*
+	 * If we're not operating in batch mode, use BufferSetHintBits16() to mark
+	 * the page dirty, that's cheaper than
+	 * BufferBeginSetHintBits()/BufferFinishSetHintBits(). That's important
+	 * for cases where we set a lot of hint bits on a page individually.
+	 */
+	if (!state)
+	{
+		BufferSetHintBits16(&tuple->t_infomask,
+							tuple->t_infomask | infomask, buffer);
+		return;
+	}
 
-		if (BufferIsPermanent(buffer) && XLogNeedsFlush(commitLSN) &&
-			BufferGetLSNAtomic(buffer) < commitLSN)
+	if (*state == SHB_INITIAL)
+	{
+		if (!BufferBeginSetHintBits(buffer))
 		{
-			/* not flushed and no LSN interlock, so don't set hint */
+			*state = SHB_DISABLED;
 			return;
 		}
-	}
 
+		*state = SHB_ENABLED;
+	}
 	tuple->t_infomask |= infomask;
-	MarkBufferDirtyHint(buffer, true);
+}
+
+/*
+ * Simple wrapper around SetHintBitExt(), use when operating on a single
+ * tuple.
+ */
+static inline void
+SetHintBits(HeapTupleHeader tuple, Buffer buffer,
+			uint16 infomask, TransactionId xid)
+{
+	SetHintBitsExt(tuple, buffer, infomask, xid, NULL);
 }
 
 /*
@@ -864,9 +935,9 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
  * inserting/deleting transaction was still running --- which was more cycles
  * and more contention on ProcArrayLock.
  */
-static bool
+static inline bool
 HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
-					   Buffer buffer)
+					   Buffer buffer, SetHintBitsState *state)
 {
 	HeapTupleHeader tuple = htup->t_data;
 
@@ -921,8 +992,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 			if (!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetRawXmax(tuple)))
 			{
 				/* deleting subtransaction must have aborted */
-				SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-							InvalidTransactionId);
+				SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+							   InvalidTransactionId, state);
 				return true;
 			}
 
@@ -934,13 +1005,13 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		else if (XidInMVCCSnapshot(HeapTupleHeaderGetRawXmin(tuple), snapshot))
 			return false;
 		else if (TransactionIdDidCommit(HeapTupleHeaderGetRawXmin(tuple)))
-			SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
-						HeapTupleHeaderGetRawXmin(tuple));
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_COMMITTED,
+						   HeapTupleHeaderGetRawXmin(tuple), state);
 		else
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMIN_INVALID,
+						   InvalidTransactionId, state);
 			return false;
 		}
 	}
@@ -1003,14 +1074,14 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
 		if (!TransactionIdDidCommit(HeapTupleHeaderGetRawXmax(tuple)))
 		{
 			/* it must have aborted or crashed */
-			SetHintBits(tuple, buffer, HEAP_XMAX_INVALID,
-						InvalidTransactionId);
+			SetHintBitsExt(tuple, buffer, HEAP_XMAX_INVALID,
+						   InvalidTransactionId, state);
 			return true;
 		}
 
 		/* xmax transaction committed */
-		SetHintBits(tuple, buffer, HEAP_XMAX_COMMITTED,
-					HeapTupleHeaderGetRawXmax(tuple));
+		SetHintBitsExt(tuple, buffer, HEAP_XMAX_COMMITTED,
+					   HeapTupleHeaderGetRawXmax(tuple), state);
 	}
 	else
 	{
@@ -1607,9 +1678,10 @@ HeapTupleSatisfiesHistoricMVCC(HeapTuple htup, Snapshot snapshot,
  * ->vistuples_dense is set to contain the offsets of visible tuples.
  *
  * The reason this is more efficient than HeapTupleSatisfiesMVCC() is that it
- * avoids a cross-translation-unit function call for each tuple and allows the
- * compiler to optimize across calls to HeapTupleSatisfiesMVCC. In the future
- * it will also allow more efficient setting of hint bits.
+ * avoids a cross-translation-unit function call for each tuple, allows the
+ * compiler to optimize across calls to HeapTupleSatisfiesMVCC and allows
+ * setting hint bits more efficiently (see the one BufferFinishSetHintBits()
+ * call below).
  *
  * Returns the number of visible tuples.
  */
@@ -1620,6 +1692,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 							OffsetNumber *vistuples_dense)
 {
 	int			nvis = 0;
+	SetHintBitsState state = SHB_INITIAL;
 
 	Assert(IsMVCCSnapshot(snapshot));
 
@@ -1628,7 +1701,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		bool		valid;
 		HeapTuple	tup = &batchmvcc->tuples[i];
 
-		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer, &state);
 		batchmvcc->visible[i] = valid;
 
 		if (likely(valid))
@@ -1638,6 +1711,9 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		}
 	}
 
+	if (state == SHB_ENABLED)
+		BufferFinishSetHintBits(buffer, true, true);
+
 	return nvis;
 }
 
@@ -1657,7 +1733,7 @@ HeapTupleSatisfiesVisibility(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 	switch (snapshot->snapshot_type)
 	{
 		case SNAPSHOT_MVCC:
-			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer);
+			return HeapTupleSatisfiesMVCC(htup, snapshot, buffer, NULL);
 		case SNAPSHOT_SELF:
 			return HeapTupleSatisfiesSelf(htup, snapshot, buffer);
 		case SNAPSHOT_ANY:
diff --git a/src/backend/access/nbtree/nbtinsert.c b/src/backend/access/nbtree/nbtinsert.c
index d17aaa5aa0f..796e1513ddf 100644
--- a/src/backend/access/nbtree/nbtinsert.c
+++ b/src/backend/access/nbtree/nbtinsert.c
@@ -681,20 +681,31 @@ _bt_check_unique(Relation rel, BTInsertState insertstate, Relation heapRel,
 				{
 					/*
 					 * The conflicting tuple (or all HOT chains pointed to by
-					 * all posting list TIDs) is dead to everyone, so mark the
-					 * index entry killed.
+					 * all posting list TIDs) is dead to everyone, so try to
+					 * mark the index entry killed. It's ok if we're not
+					 * allowed to, this isn't required for correctness.
 					 */
-					ItemIdMarkDead(curitemid);
-					opaque->btpo_flags |= BTP_HAS_GARBAGE;
+					Buffer		buf;
 
-					/*
-					 * Mark buffer with a dirty hint, since state is not
-					 * crucial. Be sure to mark the proper buffer dirty.
-					 */
+					/* Be sure to operate on the proper buffer */
 					if (nbuf != InvalidBuffer)
-						MarkBufferDirtyHint(nbuf, true);
+						buf = nbuf;
 					else
-						MarkBufferDirtyHint(insertstate->buf, true);
+						buf = insertstate->buf;
+
+					/*
+					 * Use the hint bit infrastructure to check if we can
+					 * update the page while just holding a share lock.
+					 *
+					 * Can't use BufferSetHintBits16() here as we update two
+					 * different locations.
+					 */
+					if (BufferBeginSetHintBits(buf))
+					{
+						ItemIdMarkDead(curitemid);
+						opaque->btpo_flags |= BTP_HAS_GARBAGE;
+						BufferFinishSetHintBits(buf, true, true);
+					}
 				}
 
 				/*
diff --git a/src/backend/access/nbtree/nbtutils.c b/src/backend/access/nbtree/nbtutils.c
index 5c50f0dd1bd..76e6c6fbf88 100644
--- a/src/backend/access/nbtree/nbtutils.c
+++ b/src/backend/access/nbtree/nbtutils.c
@@ -361,6 +361,17 @@ _bt_killitems(IndexScanDesc scan)
 			 */
 			if (killtuple && !ItemIdIsDead(iid))
 			{
+				if (!killedsomething)
+				{
+					/*
+					 * Use the hint bit infrastructure to check if we can
+					 * update the page while just holding a share lock. If we
+					 * are not allowed, there's no point continuing.
+					 */
+					if (!BufferBeginSetHintBits(buf))
+						goto unlock_page;
+				}
+
 				/* found the item/all posting list items */
 				ItemIdMarkDead(iid);
 				killedsomething = true;
@@ -371,8 +382,6 @@ _bt_killitems(IndexScanDesc scan)
 	}
 
 	/*
-	 * Since this can be redone later if needed, mark as dirty hint.
-	 *
 	 * Whenever we mark anything LP_DEAD, we also set the page's
 	 * BTP_HAS_GARBAGE flag, which is likewise just a hint.  (Note that we
 	 * only rely on the page-level flag in !heapkeyspace indexes.)
@@ -380,9 +389,10 @@ _bt_killitems(IndexScanDesc scan)
 	if (killedsomething)
 	{
 		opaque->btpo_flags |= BTP_HAS_GARBAGE;
-		MarkBufferDirtyHint(buf, true);
+		BufferFinishSetHintBits(buf, true, true);
 	}
 
+unlock_page:
 	if (!so->dropPin)
 		_bt_unlockbuf(rel, buf);
 	else
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index d3acaa636c3..bd6e6f06389 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -1077,11 +1077,6 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * We only need to do something if page has not yet been full page written in
  * this checkpoint round. The LSN of the inserted wal record is returned if we
  * had to write, InvalidXLogRecPtr otherwise.
- *
- * It is possible that multiple concurrent backends could attempt to write WAL
- * records. In that case, multiple copies of the same block would be recorded
- * in separate WAL records by different backends, though that is still OK from
- * a correctness perspective.
  */
 XLogRecPtr
 XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
@@ -1102,11 +1097,9 @@ XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
 
 	/*
 	 * We assume page LSN is first data on *every* page that can be passed to
-	 * XLogInsert, whether it has the standard page layout or not. Since we're
-	 * only holding a share-lock on the page, we must take the buffer header
-	 * lock when we look at the LSN.
+	 * XLogInsert, whether it has the standard page layout or not.
 	 */
-	lsn = BufferGetLSNAtomic(buffer);
+	lsn = PageGetLSN(BufferGetPage(buffer));
 
 	if (lsn <= RedoRecPtr)
 	{
diff --git a/src/backend/storage/buffer/README b/src/backend/storage/buffer/README
index 119f31b5d65..b332e002ba1 100644
--- a/src/backend/storage/buffer/README
+++ b/src/backend/storage/buffer/README
@@ -25,21 +25,26 @@ that might need to do such a wait is instead handled by waiting to obtain
 the relation-level lock, which is why you'd better hold one first.)  Pins
 may not be held across transaction boundaries, however.
 
-Buffer content locks: there are two kinds of buffer lock, shared and exclusive,
-which act just as you'd expect: multiple backends can hold shared locks on
-the same buffer, but an exclusive lock prevents anyone else from holding
-either shared or exclusive lock.  (These can alternatively be called READ
-and WRITE locks.)  These locks are intended to be short-term: they should not
-be held for long.  Buffer locks are acquired and released by LockBuffer().
-It will *not* work for a single backend to try to acquire multiple locks on
-the same buffer.  One must pin a buffer before trying to lock it.
+Buffer content locks: there are three kinds of buffer lock, shared,
+share-exclusive and exclusive:
+a) multiple backends can hold shared locks on the same buffer
+   (alternatively called a READ lock)
+b) one backend can hold a share-exclusive lock on a buffer while multiple
+   backends can hold a share lock
+c) an exclusive lock prevents anyone else from holding a shared,
+   share-exclusive or exclusive lock.
+   (alternatively called a WRITE lock)
+
+These locks are intended to be short-term: they should not be held for long.
+Buffer locks are acquired and released by LockBuffer().  It will *not* work
+for a single backend to try to acquire multiple locks on the same buffer.  One
+must pin a buffer before trying to lock it.
 
 Buffer access rules:
 
-1. To scan a page for tuples, one must hold a pin and either shared or
-exclusive content lock.  To examine the commit status (XIDs and status bits)
-of a tuple in a shared buffer, one must likewise hold a pin and either shared
-or exclusive lock.
+1. To scan a page for tuples, one must hold a pin and at least a share lock.
+To examine the commit status (XIDs and status bits) of a tuple in a shared
+buffer, one must likewise hold a pin and at least a share lock.
 
 2. Once one has determined that a tuple is interesting (visible to the
 current transaction) one may drop the content lock, yet continue to access
@@ -55,19 +60,25 @@ one must hold a pin and an exclusive content lock on the containing buffer.
 This ensures that no one else might see a partially-updated state of the
 tuple while they are doing visibility checks.
 
-4. It is considered OK to update tuple commit status bits (ie, OR the
-values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID, HEAP_XMAX_COMMITTED, or
-HEAP_XMAX_INVALID into t_infomask) while holding only a shared lock and
-pin on a buffer.  This is OK because another backend looking at the tuple
-at about the same time would OR the same bits into the field, so there
-is little or no risk of conflicting update; what's more, if there did
-manage to be a conflict it would merely mean that one bit-update would
-be lost and need to be done again later.  These four bits are only hints
-(they cache the results of transaction status lookups in pg_xact), so no
-great harm is done if they get reset to zero by conflicting updates.
-Note, however, that a tuple is frozen by setting both HEAP_XMIN_INVALID
-and HEAP_XMIN_COMMITTED; this is a critical update and accordingly requires
-an exclusive buffer lock (and it must also be WAL-logged).
+4. Non-critical information on a page ("hint bits") may be modified while
+holding only a share-exclusive lock and pin on the page. To do so in cases
+where only a share lock is already held, use BufferBeginSetHintBits() &
+BufferFinishSetHintBits() (if multiple hint bits are to be set) or
+BufferSetHintBits16() (if a single hint bit is set).
+
+E.g. for heapam, a share-exclusive lock allows to update tuple commit status
+bits (ie, OR the values HEAP_XMIN_COMMITTED, HEAP_XMIN_INVALID,
+HEAP_XMAX_COMMITTED, or HEAP_XMAX_INVALID into t_infomask) while holding only
+a share-exclusive lock and pin on a buffer.  This is OK because another
+backend looking at the tuple at about the same time would OR the same bits
+into the field, so there is little or no risk of conflicting update; what's
+more, if there did manage to be a conflict it would merely mean that one
+bit-update would be lost and need to be done again later.  These four bits are
+only hints (they cache the results of transaction status lookups in pg_xact),
+so no great harm is done if they get reset to zero by conflicting updates.
+Note, however, that a tuple is frozen by setting both HEAP_XMIN_INVALID and
+HEAP_XMIN_COMMITTED; this is a critical update and accordingly requires an
+exclusive buffer lock (and it must also be WAL-logged).
 
 5. To physically remove a tuple or compact free space on a page, one
 must hold a pin and an exclusive lock, *and* observe while holding the
@@ -80,7 +91,6 @@ buffer (increment the refcount) while one is performing the cleanup, but
 it won't be able to actually examine the page until it acquires shared
 or exclusive content lock.
 
-
 Obtaining the lock needed under rule #5 is done by the bufmgr routines
 LockBufferForCleanup() or ConditionalLockBufferForCleanup().  They first get
 an exclusive lock and then check to see if the shared pin count is currently
@@ -96,6 +106,10 @@ VACUUM's use, since we don't allow multiple VACUUMs concurrently on a single
 relation anyway.  Anyone wishing to obtain a cleanup lock outside of recovery
 or a VACUUM must use the conditional variant of the function.
 
+6. To write out a buffer, a share-exclusive lock needs to be held. This
+prevents the buffer from being modified while written out, which could corrupt
+checksums and cause issues on the OS or device level when direct-IO is used.
+
 
 Buffer Manager's Internal Locking
 ---------------------------------
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 7241477cac0..3b32b4e0ab1 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2480,9 +2480,8 @@ again:
 	/*
 	 * If the buffer was dirty, try to write it out.  There is a race
 	 * condition here, in that someone might dirty it after we released the
-	 * buffer header lock above, or even while we are writing it out (since
-	 * our share-lock won't prevent hint-bit updates).  We will recheck the
-	 * dirty bit after re-locking the buffer header.
+	 * buffer header lock above.  We will recheck the dirty bit after
+	 * re-locking the buffer header.
 	 */
 	if (buf_state & BM_DIRTY)
 	{
@@ -2490,20 +2489,20 @@ again:
 		Assert(buf_state & BM_VALID);
 
 		/*
-		 * We need a share-lock on the buffer contents to write it out (else
-		 * we might write invalid data, eg because someone else is compacting
-		 * the page contents while we write).  We must use a conditional lock
-		 * acquisition here to avoid deadlock.  Even though the buffer was not
-		 * pinned (and therefore surely not locked) when StrategyGetBuffer
-		 * returned it, someone else could have pinned and exclusive-locked it
-		 * by the time we get here. If we try to get the lock unconditionally,
-		 * we'd block waiting for them; if they later block waiting for us,
-		 * deadlock ensues. (This has been observed to happen when two
-		 * backends are both trying to split btree index pages, and the second
-		 * one just happens to be trying to split the page the first one got
-		 * from StrategyGetBuffer.)
+		 * We need a share-exclusive lock on the buffer contents to write it
+		 * out (else we might write invalid data, eg because someone else is
+		 * compacting the page contents while we write).  We must use a
+		 * conditional lock acquisition here to avoid deadlock.  Even though
+		 * the buffer was not pinned (and therefore surely not locked) when
+		 * StrategyGetBuffer returned it, someone else could have pinned and
+		 * (share-)exclusive-locked it by the time we get here. If we try to
+		 * get the lock unconditionally, we'd block waiting for them; if they
+		 * later block waiting for us, deadlock ensues. (This has been
+		 * observed to happen when two backends are both trying to split btree
+		 * index pages, and the second one just happens to be trying to split
+		 * the page the first one got from StrategyGetBuffer.)
 		 */
-		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE))
+		if (!BufferLockConditional(buf, buf_hdr, BUFFER_LOCK_SHARE_EXCLUSIVE))
 		{
 			/*
 			 * Someone else has locked the buffer, so give it up and loop back
@@ -2516,18 +2515,21 @@ again:
 		/*
 		 * If using a nondefault strategy, and writing the buffer would
 		 * require a WAL flush, let the strategy decide whether to go ahead
-		 * and write/reuse the buffer or to choose another victim.  We need a
-		 * lock to inspect the page LSN, so this can't be done inside
+		 * and write/reuse the buffer or to choose another victim.  We need to
+		 * hold the content lock in at least share-exclusive mode to safely
+		 * inspect the page LSN, so this couldn't have been done inside
 		 * StrategyGetBuffer.
 		 */
 		if (strategy != NULL)
 		{
 			XLogRecPtr	lsn;
 
-			/* Read the LSN while holding buffer header lock */
-			buf_state = LockBufHdr(buf_hdr);
+			/*
+			 * As we now hold at least a share-exclusive lock on the buffer,
+			 * the LSN cannot change during the flush (and thus can't be
+			 * torn).
+			 */
 			lsn = BufferGetLSN(buf_hdr);
-			UnlockBufHdr(buf_hdr);
 
 			if (XLogNeedsFlush(lsn)
 				&& StrategyRejectBuffer(strategy, buf_hdr, from_ring))
@@ -3017,7 +3019,7 @@ BufferIsLockedByMeInMode(Buffer buffer, BufferLockMode mode)
  *
  *		Checks if buffer is already dirty.
  *
- * Buffer must be pinned and exclusive-locked.  (Without an exclusive lock,
+ * Buffer must be pinned and [share-]exclusive-locked.  (Without such a lock,
  * the result may be stale before it's returned.)
  */
 bool
@@ -3037,7 +3039,8 @@ BufferIsDirty(Buffer buffer)
 	else
 	{
 		bufHdr = GetBufferDescriptor(buffer - 1);
-		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
 	}
 
 	return pg_atomic_read_u64(&bufHdr->state) & BM_DIRTY;
@@ -4072,8 +4075,8 @@ SyncOneBuffer(int buf_id, bool skip_recently_used, WritebackContext *wb_context)
 	}
 
 	/*
-	 * Pin it, share-lock it, write it.  (FlushBuffer will do nothing if the
-	 * buffer is clean by the time we've locked it.)
+	 * Pin it, share-exclusive-lock it, write it.  (FlushBuffer will do
+	 * nothing if the buffer is clean by the time we've locked it.)
 	 */
 	PinBuffer_Locked(bufHdr);
 
@@ -4403,11 +4406,8 @@ BufferGetTag(Buffer buffer, RelFileLocator *rlocator, ForkNumber *forknum,
  * However, we will need to force the changes to disk via fsync before
  * we can checkpoint WAL.
  *
- * The caller must hold a pin on the buffer and have share-locked the
- * buffer contents.  (Note: a share-lock does not prevent updates of
- * hint bits in the buffer, so the page could change while the write
- * is in progress, but we assume that that will not invalidate the data
- * written.)
+ * The caller must hold a pin on the buffer and have
+ * (share-)exclusively-locked the buffer contents.
  *
  * If the caller has an smgr reference for the buffer's relation, pass it
  * as the second parameter.  If not, pass NULL.
@@ -4423,6 +4423,9 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	char	   *bufToWrite;
 	uint64		buf_state;
 
+	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(buf, BUFFER_LOCK_SHARE_EXCLUSIVE));
+
 	/*
 	 * Try to start an I/O operation.  If StartBufferIO returns false, then
 	 * someone else flushed the buffer before we could, so we need not do
@@ -4450,8 +4453,8 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	buf_state = LockBufHdr(buf);
 
 	/*
-	 * Run PageGetLSN while holding header lock, since we don't have the
-	 * buffer locked exclusively in all cases.
+	 * As we hold at least a share-exclusive lock on the buffer, the LSN
+	 * cannot change during the flush (and thus can't be torn).
 	 */
 	recptr = BufferGetLSN(buf);
 
@@ -4555,7 +4558,7 @@ FlushUnlockedBuffer(BufferDesc *buf, SMgrRelation reln,
 {
 	Buffer		buffer = BufferDescriptorGetBuffer(buf);
 
-	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE);
+	BufferLockAcquire(buffer, buf, BUFFER_LOCK_SHARE_EXCLUSIVE);
 	FlushBuffer(buf, reln, IOOBJECT_RELATION, IOCONTEXT_NORMAL);
 	BufferLockUnlock(buffer, buf);
 }
@@ -4627,8 +4630,9 @@ BufferIsPermanent(Buffer buffer)
 /*
  * BufferGetLSNAtomic
  *		Retrieves the LSN of the buffer atomically using a buffer header lock.
- *		This is necessary for some callers who may not have an exclusive lock
- *		on the buffer.
+ *		This is necessary for some callers who may only hold a share lock on
+ *		the buffer. A share lock allows a concurrent backend to set hint bits
+ *		on the page, which in turn may require a WAL record to be emitted.
  */
 XLogRecPtr
 BufferGetLSNAtomic(Buffer buffer)
@@ -5474,8 +5478,8 @@ FlushDatabaseBuffers(Oid dbid)
 }
 
 /*
- * Flush a previously, shared or exclusively, locked and pinned buffer to the
- * OS.
+ * Flush a previously, share-exclusively or exclusively, locked and pinned
+ * buffer to the OS.
  */
 void
 FlushOneBuffer(Buffer buffer)
@@ -5548,56 +5552,38 @@ IncrBufferRefCount(Buffer buffer)
 }
 
 /*
- * MarkBufferDirtyHint
+ * Shared-buffer only helper for MarkBufferDirtyHint() and
+ * BufferSetHintBits16().
  *
- *	Mark a buffer dirty for non-critical changes.
- *
- * This is essentially the same as MarkBufferDirty, except:
- *
- * 1. The caller does not write WAL; so if checksums are enabled, we may need
- *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
- * 2. The caller might have only share-lock instead of exclusive-lock on the
- *	  buffer's content lock.
- * 3. This function does not guarantee that the buffer is always marked dirty
- *	  (due to a race condition), so it cannot be used for important changes.
+ * This is separated out because it turns out that the repeated checks for
+ * local buffers, repeated GetBufferDescriptor() and repeated reading of the
+ * buffer's state sufficiently hurts the performance of BufferSetHintBits16().
  */
-void
-MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+static inline void
+MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate,
+						  bool buffer_std)
 {
-	BufferDesc *bufHdr;
 	Page		page = BufferGetPage(buffer);
 
-	if (!BufferIsValid(buffer))
-		elog(ERROR, "bad buffer ID: %d", buffer);
-
-	if (BufferIsLocal(buffer))
-	{
-		MarkLocalBufferDirty(buffer);
-		return;
-	}
-
-	bufHdr = GetBufferDescriptor(buffer - 1);
-
 	Assert(GetPrivateRefCount(buffer) > 0);
-	/* here, either share or exclusive lock is OK */
-	Assert(BufferIsLockedByMe(buffer));
+
+	/* here, either share-exclusive or exclusive lock is OK */
+	Assert(BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_EXCLUSIVE) ||
+		   BufferLockHeldByMeInMode(bufHdr, BUFFER_LOCK_SHARE_EXCLUSIVE));
 
 	/*
 	 * This routine might get called many times on the same page, if we are
 	 * making the first scan after commit of an xact that added/deleted many
-	 * tuples. So, be as quick as we can if the buffer is already dirty.  We
-	 * do this by not acquiring spinlock if it looks like the status bits are
-	 * already set.  Since we make this test unlocked, there's a chance we
-	 * might fail to notice that the flags have just been cleared, and failed
-	 * to reset them, due to memory-ordering issues.  But since this function
-	 * is only intended to be used in cases where failing to write out the
-	 * data would be harmless anyway, it doesn't really matter.
+	 * tuples. So, be as quick as we can if the buffer is already dirty.
+	 *
+	 * As we are holding (at least) a share-exclusive lock, nobody could have
+	 * cleaned or dirtied the page concurrently, so we can just rely on the
+	 * previously fetched value here without any danger of races.
 	 */
-	if ((pg_atomic_read_u64(&bufHdr->state) & (BM_DIRTY | BM_JUST_DIRTIED)) !=
-		(BM_DIRTY | BM_JUST_DIRTIED))
+	if (unlikely((lockstate & (BM_DIRTY | BM_JUST_DIRTIED)) !=
+				 (BM_DIRTY | BM_JUST_DIRTIED)))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
-		bool		dirtied = false;
 		bool		delayChkptFlags = false;
 		uint64		buf_state;
 
@@ -5610,8 +5596,7 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		 * We don't check full_page_writes here because that logic is included
 		 * when we call XLogInsert() since the value changes dynamically.
 		 */
-		if (XLogHintBitIsNeeded() &&
-			(pg_atomic_read_u64(&bufHdr->state) & BM_PERMANENT))
+		if (XLogHintBitIsNeeded() && (lockstate & BM_PERMANENT))
 		{
 			/*
 			 * If we must not write WAL, due to a relfilelocator-specific
@@ -5656,27 +5641,29 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 
 		buf_state = LockBufHdr(bufHdr);
 
+		/*
+		 * It should not be possible for the buffer to already be dirty, see
+		 * comment above.
+		 */
+		Assert(!(buf_state & BM_DIRTY));
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
 
-		if (!(buf_state & BM_DIRTY))
+		if (XLogRecPtrIsValid(lsn))
 		{
-			dirtied = true;		/* Means "will be dirtied by this action" */
-
 			/*
-			 * Set the page LSN if we wrote a backup block. We aren't supposed
-			 * to set this when only holding a share lock but as long as we
-			 * serialise it somehow we're OK. We choose to set LSN while
-			 * holding the buffer header lock, which causes any reader of an
-			 * LSN who holds only a share lock to also obtain a buffer header
-			 * lock before using PageGetLSN(), which is enforced in
-			 * BufferGetLSNAtomic().
+			 * Set the page LSN if we wrote a backup block. To allow backends
+			 * that only hold a share lock on the buffer to read the LSN in a
+			 * tear-free manner, we set the page LSN while holding the buffer
+			 * header lock. This allows any reader of an LSN who holds only a
+			 * share lock to also obtain a buffer header lock before using
+			 * PageGetLSN() to read the LSN in a tear free way. This is done
+			 * in BufferGetLSNAtomic().
 			 *
 			 * If checksums are enabled, you might think we should reset the
 			 * checksum here. That will happen when the page is written
 			 * sometime later in this checkpoint cycle.
 			 */
-			if (XLogRecPtrIsValid(lsn))
-				PageSetLSN(page, lsn);
+			PageSetLSN(page, lsn);
 		}
 
 		UnlockBufHdrExt(bufHdr, buf_state,
@@ -5686,15 +5673,48 @@ MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
 		if (delayChkptFlags)
 			MyProc->delayChkptFlags &= ~DELAY_CHKPT_START;
 
-		if (dirtied)
-		{
-			pgBufferUsage.shared_blks_dirtied++;
-			if (VacuumCostActive)
-				VacuumCostBalance += VacuumCostPageDirty;
-		}
+		pgBufferUsage.shared_blks_dirtied++;
+		if (VacuumCostActive)
+			VacuumCostBalance += VacuumCostPageDirty;
 	}
 }
 
+/*
+ * MarkBufferDirtyHint
+ *
+ *	Mark a buffer dirty for non-critical changes.
+ *
+ * This is essentially the same as MarkBufferDirty, except:
+ *
+ * 1. The caller does not write WAL; so if checksums are enabled, we may need
+ *	  to write an XLOG_FPI_FOR_HINT WAL record to protect against torn pages.
+ * 2. The caller might have only a share-exclusive-lock instead of an
+ *	  exclusive-lock on the buffer's content lock.
+ * 3. This function does not guarantee that the buffer is always marked dirty
+ *	  (it e.g. can't always on a hot standby), so it cannot be used for
+ *	  important changes.
+ */
+inline void
+MarkBufferDirtyHint(Buffer buffer, bool buffer_std)
+{
+	BufferDesc *bufHdr;
+
+	bufHdr = GetBufferDescriptor(buffer - 1);
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		MarkLocalBufferDirty(buffer);
+		return;
+	}
+
+	MarkSharedBufferDirtyHint(buffer, bufHdr,
+							  pg_atomic_read_u64(&bufHdr->state),
+							  buffer_std);
+}
+
 /*
  * Release buffer content locks for shared buffers.
  *
@@ -6796,6 +6816,187 @@ IsBufferCleanupOK(Buffer buffer)
 	return false;
 }
 
+/*
+ * Helper for BufferBeginSetHintBits() and BufferSetHintBits16().
+ *
+ * This checks if the current lock mode already suffices to allow hint bits
+ * being set and, if not, whether the current lock can be upgraded.
+ *
+ * Updates *lockstate when returning true.
+ */
+static inline bool
+SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
+{
+	uint64		old_state;
+	PrivateRefCountEntry *ref;
+	BufferLockMode mode;
+
+	ref = GetPrivateRefCountEntry(buffer, true);
+
+	if (ref == NULL)
+		elog(ERROR, "buffer is not pinned");
+
+	mode = ref->data.lockmode;
+	if (mode == BUFFER_LOCK_UNLOCK)
+		elog(ERROR, "buffer is not locked");
+
+	/* we're done if we are already holding a sufficient lock level */
+	if (mode == BUFFER_LOCK_EXCLUSIVE || mode == BUFFER_LOCK_SHARE_EXCLUSIVE)
+	{
+		*lockstate = pg_atomic_read_u64(&buf_hdr->state);
+		return true;
+	}
+
+	/*
+	 * We are only holding a share lock right now, try to upgrade it to
+	 * SHARE_EXCLUSIVE.
+	 */
+	Assert(mode == BUFFER_LOCK_SHARE);
+
+	old_state = pg_atomic_read_u64(&buf_hdr->state);
+	while (true)
+	{
+		uint64		desired_state;
+
+		desired_state = old_state;
+
+		/*
+		 * Can't upgrade if somebody else holds the lock in exclusive or
+		 * share-exclusive mode.
+		 */
+		if (unlikely((old_state & (BM_LOCK_VAL_EXCLUSIVE | BM_LOCK_VAL_SHARE_EXCLUSIVE)) != 0))
+		{
+			return false;
+		}
+
+		/* currently held lock state */
+		desired_state -= BM_LOCK_VAL_SHARED;
+
+		/* new lock level */
+		desired_state += BM_LOCK_VAL_SHARE_EXCLUSIVE;
+
+		if (likely(pg_atomic_compare_exchange_u64(&buf_hdr->state,
+												  &old_state, desired_state)))
+		{
+			ref->data.lockmode = BUFFER_LOCK_SHARE_EXCLUSIVE;
+			*lockstate = desired_state;
+
+			return true;
+		}
+	}
+}
+
+/*
+ * Try to acquire the right to set hint bits on the buffer.
+ *
+ * To be allowed to set hint bits, this backend needs to hold either a
+ * share-exclusive or an exclusive lock. In case this backend only holds a
+ * share lock, this function will try to upgrade the lock to
+ * share-exclusive. The caller is only allowed to set hint bits if true is
+ * returned.
+ *
+ * Once BufferBeginSetHintBits() has returned true, hint bits may be set
+ * without further calls to BufferBeginSetHintBits(), until the buffer is
+ * unlocked.
+ *
+ *
+ * Requiring a share-exclusive lock to set hint bits prevents setting hint
+ * bits on buffers that are currently being written out, which could corrupt
+ * the checksum on the page. Flushing buffers also requires a share-exclusive
+ * lock.
+ *
+ * Due to a lock >= share-exclusive being required to set hint bits, only one
+ * backend can set hint bits at a time. Allowing multiple backends to set hint
+ * bits would require more complicated locking: For setting hint bits we'd
+ * need to store the count of backends currently setting hint bits, for I/O we
+ * would need another lock-level conflicting with the hint-setting
+ * lock-level. Given that the share-exclusive lock for setting hint bits is
+ * only held for a short time, that backends often would just set the same
+ * hint bits and that the cost of occasionally not setting hint bits in hotly
+ * accessed pages is fairly low, this seems like an acceptable tradeoff.
+ */
+bool
+BufferBeginSetHintBits(Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+
+	if (BufferIsLocal(buffer))
+	{
+		/*
+		 * NB: Will need to check if there is a write in progress, once it is
+		 * possible for writes to be done asynchronously.
+		 */
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	return SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate);
+}
+
+/*
+ * End a phase of setting hint bits on this buffer, started with
+ * BufferBeginSetHintBits().
+ *
+ * This would strictly speaking not be required (i.e. the caller could do
+ * MarkBufferDirtyHint() if so desired), but allows us to perform some sanity
+ * checks.
+ */
+void
+BufferFinishSetHintBits(Buffer buffer, bool mark_dirty, bool buffer_std)
+{
+	if (!BufferIsLocal(buffer))
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE));
+
+	if (mark_dirty)
+		MarkBufferDirtyHint(buffer, buffer_std);
+}
+
+/*
+ * Try to set a single hint bit in a buffer.
+ *
+ * This is a bit faster than BufferBeginSetHintBits() /
+ * BufferFinishSetHintBits() when setting a single hint bit, but slower than
+ * the former when setting several hint bits.
+ */
+bool
+BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)
+{
+	BufferDesc *buf_hdr;
+	uint64		lockstate;
+#ifdef USE_ASSERT_CHECKING
+	char	   *page;
+
+	/* verify that the address is on the page */
+	page = BufferGetPage(buffer);
+	Assert((char *) ptr >= page && (char *) ptr < (page + BLCKSZ));
+#endif
+
+	if (BufferIsLocal(buffer))
+	{
+		*ptr = val;
+
+		MarkLocalBufferDirty(buffer);
+
+		return true;
+	}
+
+	buf_hdr = GetBufferDescriptor(buffer - 1);
+
+	if (SharedBufferBeginSetHintBits(buffer, buf_hdr, &lockstate))
+	{
+		*ptr = val;
+
+		MarkSharedBufferDirtyHint(buffer, buf_hdr, lockstate, true);
+
+		return true;
+	}
+
+	return false;
+}
+
 
 /*
  *	Functions for buffer I/O handling
diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index ad337c00871..b9a8f368a63 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -904,13 +904,17 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 	max_avail = fsm_get_max_avail(page);
 
 	/*
-	 * Reset the next slot pointer. This encourages the use of low-numbered
-	 * pages, increasing the chances that a later vacuum can truncate the
-	 * relation. We don't bother with marking the page dirty if it wasn't
-	 * already, since this is just a hint.
+	 * Try to reset the next slot pointer. This encourages the use of
+	 * low-numbered pages, increasing the chances that a later vacuum can
+	 * truncate the relation. We don't bother with marking the page dirty if
+	 * it wasn't already, since this is just a hint.
 	 */
 	LockBuffer(buf, BUFFER_LOCK_SHARE);
-	((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+	if (BufferBeginSetHintBits(buf))
+	{
+		((FSMPage) PageGetContents(page))->fp_next_slot = 0;
+		BufferFinishSetHintBits(buf, false, false);
+	}
 	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
 	ReleaseBuffer(buf);
diff --git a/src/backend/storage/freespace/fsmpage.c b/src/backend/storage/freespace/fsmpage.c
index 33ee825529c..a2657c4033b 100644
--- a/src/backend/storage/freespace/fsmpage.c
+++ b/src/backend/storage/freespace/fsmpage.c
@@ -298,9 +298,18 @@ restart:
 	 * lock and get a garbled next pointer every now and then, than take the
 	 * concurrency hit of an exclusive lock.
 	 *
+	 * Without an exclusive lock, we need to use the hint bit infrastructure
+	 * to be allowed to modify the page.
+	 *
 	 * Wrap-around is handled at the beginning of this function.
 	 */
-	fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+	if (exclusive_lock_held || BufferBeginSetHintBits(buf))
+	{
+		fsmpage->fp_next_slot = slot + (advancenext ? 1 : 0);
+
+		if (!exclusive_lock_held)
+			BufferFinishSetHintBits(buf, false, false);
+	}
 
 	return slot;
 }
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 9f5ee8fd482..3a42feacdc1 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2757,6 +2757,7 @@ SetConstraintStateData
 SetConstraintTriggerData
 SetExprState
 SetFunctionReturnMode
+SetHintBitsState
 SetOp
 SetOpCmd
 SetOpPath
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v12-0003-bufmgr-Remove-the-now-obsolete-BM_JUST_DIRTIED.patch (6.4K, ../../5ubipyssiju5twkb7zgqwdr7q2vhpkpmuelxfpanetlk6ofnop@hvxb4g2amb2d/4-v12-0003-bufmgr-Remove-the-now-obsolete-BM_JUST_DIRTIED.patch)
  download | inline diff:
From 7fb80c58553812d4485094bbf78e10d8b3dc2c1b Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 2 Feb 2026 13:24:04 -0500
Subject: [PATCH v12 3/6] bufmgr: Remove the, now obsolete, BM_JUST_DIRTIED

Due to the recent changes to use a share-exclusive mode for setting hint bits
and for flushing pages, instead of using share mode as before, a buffer cannot
be dirtied while the flush is ongoing.  The reason we needed JUST_DIRTIED was
to handle the case where the buffer was dirtied while IO was ongoing - which
is not possible anymore.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/buf_internals.h   |  3 +--
 src/backend/storage/buffer/bufmgr.c   | 30 ++++++++-------------------
 src/backend/storage/buffer/localbuf.c |  2 +-
 3 files changed, 11 insertions(+), 24 deletions(-)

diff --git a/src/include/storage/buf_internals.h b/src/include/storage/buf_internals.h
index 27f12502d19..8d1e16b5d51 100644
--- a/src/include/storage/buf_internals.h
+++ b/src/include/storage/buf_internals.h
@@ -114,8 +114,7 @@ StaticAssertDecl(BUF_REFCOUNT_BITS + BUF_USAGECOUNT_BITS + BUF_FLAG_BITS + BUF_L
 #define BM_IO_IN_PROGRESS			BUF_DEFINE_FLAG( 4)
 /* previous I/O failed */
 #define BM_IO_ERROR					BUF_DEFINE_FLAG( 5)
-/* dirtied since write started */
-#define BM_JUST_DIRTIED				BUF_DEFINE_FLAG( 6)
+/* flag bit 6 is not used anymore */
 /* have waiter for sole pin */
 #define BM_PIN_COUNT_WAITER			BUF_DEFINE_FLAG( 7)
 /* must write for checkpoint */
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 3b32b4e0ab1..e462fc799fe 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2886,7 +2886,7 @@ ExtendBufferedRelShared(BufferManagerRelation bmr,
 			buf_state = LockBufHdr(victim_buf_hdr);
 
 			/* some sanity checks while we hold the buffer header lock */
-			Assert(!(buf_state & (BM_VALID | BM_TAG_VALID | BM_DIRTY | BM_JUST_DIRTIED)));
+			Assert(!(buf_state & (BM_VALID | BM_TAG_VALID | BM_DIRTY)));
 			Assert(BUF_STATE_GET_REFCOUNT(buf_state) == 1);
 
 			victim_buf_hdr->tag = tag;
@@ -3089,7 +3089,7 @@ MarkBufferDirty(Buffer buffer)
 		buf_state = old_buf_state;
 
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
-		buf_state |= BM_DIRTY | BM_JUST_DIRTIED;
+		buf_state |= BM_DIRTY;
 
 		if (pg_atomic_compare_exchange_u64(&bufHdr->state, &old_buf_state,
 										   buf_state))
@@ -4421,7 +4421,6 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	instr_time	io_start;
 	Block		bufBlock;
 	char	   *bufToWrite;
-	uint64		buf_state;
 
 	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
 		   BufferLockHeldByMeInMode(buf, BUFFER_LOCK_SHARE_EXCLUSIVE));
@@ -4450,19 +4449,12 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 										reln->smgr_rlocator.locator.dbOid,
 										reln->smgr_rlocator.locator.relNumber);
 
-	buf_state = LockBufHdr(buf);
-
 	/*
 	 * As we hold at least a share-exclusive lock on the buffer, the LSN
 	 * cannot change during the flush (and thus can't be torn).
 	 */
 	recptr = BufferGetLSN(buf);
 
-	/* To check if block content changes while flushing. - vadim 01/17/97 */
-	UnlockBufHdrExt(buf, buf_state,
-					0, BM_JUST_DIRTIED,
-					0);
-
 	/*
 	 * Force XLOG flush up to buffer's LSN.  This implements the basic WAL
 	 * rule that log updates must hit disk before any of the data-file changes
@@ -4480,7 +4472,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	 * disastrous system-wide consequences.  To make sure that can't happen,
 	 * skip the flush if the buffer isn't permanent.
 	 */
-	if (buf_state & BM_PERMANENT)
+	if (pg_atomic_read_u64(&buf->state) & BM_PERMANENT)
 		XLogFlush(recptr);
 
 	/*
@@ -4533,8 +4525,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	pgBufferUsage.shared_blks_written++;
 
 	/*
-	 * Mark the buffer as clean (unless BM_JUST_DIRTIED has become set) and
-	 * end the BM_IO_IN_PROGRESS state.
+	 * Mark the buffer as clean and end the BM_IO_IN_PROGRESS state.
 	 */
 	TerminateBufferIO(buf, true, 0, true, false);
 
@@ -5580,8 +5571,7 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate,
 	 * cleaned or dirtied the page concurrently, so we can just rely on the
 	 * previously fetched value here without any danger of races.
 	 */
-	if (unlikely((lockstate & (BM_DIRTY | BM_JUST_DIRTIED)) !=
-				 (BM_DIRTY | BM_JUST_DIRTIED)))
+	if (unlikely(!(lockstate & BM_DIRTY)))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
 		bool		delayChkptFlags = false;
@@ -5667,7 +5657,7 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate,
 		}
 
 		UnlockBufHdrExt(bufHdr, buf_state,
-						BM_DIRTY | BM_JUST_DIRTIED,
+						BM_DIRTY,
 						0, 0);
 
 		if (delayChkptFlags)
@@ -7131,10 +7121,8 @@ StartBufferIO(BufferDesc *buf, bool forInput, bool nowait)
  *	BM_IO_IN_PROGRESS bit is set for the buffer
  *	The buffer is Pinned
  *
- * If clear_dirty is true and BM_JUST_DIRTIED is not set, we clear the
- * buffer's BM_DIRTY flag.  This is appropriate when terminating a
- * successful write.  The check on BM_JUST_DIRTIED is necessary to avoid
- * marking the buffer clean if it was re-dirtied while we were writing.
+ * If clear_dirty is true, we clear the buffer's BM_DIRTY flag.  This is
+ * appropriate when terminating a successful write.
  *
  * set_flag_bits gets ORed into the buffer's flags.  It must include
  * BM_IO_ERROR in a failure case.  For successful completion it could
@@ -7160,7 +7148,7 @@ TerminateBufferIO(BufferDesc *buf, bool clear_dirty, uint64 set_flag_bits,
 	/* Clear earlier errors, if this IO failed, it'll be marked again */
 	unset_flag_bits |= BM_IO_ERROR;
 
-	if (clear_dirty && !(buf_state & BM_JUST_DIRTIED))
+	if (clear_dirty)
 		unset_flag_bits |= BM_DIRTY | BM_CHECKPOINT_NEEDED;
 
 	if (release_aio)
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 04a540379a2..404c6bccbdd 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -441,7 +441,7 @@ ExtendBufferedRelLocal(BufferManagerRelation bmr,
 		{
 			uint64		buf_state = pg_atomic_read_u64(&victim_buf_hdr->state);
 
-			Assert(!(buf_state & (BM_VALID | BM_TAG_VALID | BM_DIRTY | BM_JUST_DIRTIED)));
+			Assert(!(buf_state & (BM_VALID | BM_TAG_VALID | BM_DIRTY)));
 
 			victim_buf_hdr->tag = tag;
 
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v12-0004-bufmgr-Switch-to-standard-order-in-MarkBufferDir.patch (6.9K, ../../5ubipyssiju5twkb7zgqwdr7q2vhpkpmuelxfpanetlk6ofnop@hvxb4g2amb2d/5-v12-0004-bufmgr-Switch-to-standard-order-in-MarkBufferDir.patch)
  download | inline diff:
From 5d91d2a1359ecabe7b3055ad8593bae50725d54d Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 2 Feb 2026 14:06:45 -0500
Subject: [PATCH v12 4/6] bufmgr: Switch to standard order in
 MarkBufferDirtyHint()

When we were updating hint bits with just a share lock MarkBufferDirtyHint()
had to use a non-standard order of operations, i.e. WAL log the buffer before
marking the buffer dirty. This was required because the lock level used to set
hints did not conflict with the lock level that was used to flush pages, which
would have allowed flushing the page out before the WAL record. The
non-standard order in turn required preventing the checkpoint from starting
between writing the WAL record and flushing out the page.

Now that setting hints and writing out buffers use share-exclusive, we can
revert back to the normal order of operations.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/transam/xloginsert.c | 20 +++++---
 src/backend/storage/buffer/bufmgr.c     | 61 +++++++++++--------------
 2 files changed, 40 insertions(+), 41 deletions(-)

diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index bd6e6f06389..7f27eee5ba1 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -1066,7 +1066,10 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * Write a backup block if needed when we are setting a hint. Note that
  * this may be called for a variety of page types, not just heaps.
  *
- * Callable while holding just share lock on the buffer content.
+ * Callable while holding just a share-exclusive lock on the buffer
+ * content. That suffices to prevent concurrent modifications of the
+ * buffer. The buffer already needs to have been marked dirty by
+ * MarkBufferDirtyHint().
  *
  * We can't use the plain backup block mechanism since that relies on the
  * Buffer being exclusively locked. Since some modifications (setting LSN, hint
@@ -1085,13 +1088,18 @@ XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
 	XLogRecPtr	lsn;
 	XLogRecPtr	RedoRecPtr;
 
-	/*
-	 * Ensure no checkpoint can change our view of RedoRecPtr.
-	 */
-	Assert((MyProc->delayChkptFlags & DELAY_CHKPT_START) != 0);
+	/* this also verifies that we hold an appropriate lock */
+	Assert(BufferIsDirty(buffer));
 
 	/*
-	 * Update RedoRecPtr so that we can make the right decision
+	 * Update RedoRecPtr so that we can make the right decision. It's possible
+	 * that a new checkpoint will start just after GetRedoRecPtr(), but that
+	 * is ok, as the buffer is already dirty, ensuring that any BufferSync()
+	 * started after the buffer was marked dirty cannot complete without
+	 * flushing this buffer.  If a checkpoint started between marking the
+	 * buffer dirty and this check, we will emit an unnecessary WAL record (as
+	 * the buffer will be written out as part of the checkpoint), but the
+	 * window for that is small.
 	 */
 	RedoRecPtr = GetRedoRecPtr();
 
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index e462fc799fe..929466d25fd 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5574,7 +5574,7 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate,
 	if (unlikely(!(lockstate & BM_DIRTY)))
 	{
 		XLogRecPtr	lsn = InvalidXLogRecPtr;
-		bool		delayChkptFlags = false;
+		bool		wal_log = false;
 		uint64		buf_state;
 
 		/*
@@ -5600,35 +5600,18 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate,
 				RelFileLocatorSkippingWAL(BufTagGetRelFileLocator(&bufHdr->tag)))
 				return;
 
-			/*
-			 * If the block is already dirty because we either made a change
-			 * or set a hint already, then we don't need to write a full page
-			 * image.  Note that aggressive cleaning of blocks dirtied by hint
-			 * bit setting would increase the call rate. Bulk setting of hint
-			 * bits would reduce the call rate...
-			 *
-			 * We must issue the WAL record before we mark the buffer dirty.
-			 * Otherwise we might write the page before we write the WAL. That
-			 * causes a race condition, since a checkpoint might occur between
-			 * writing the WAL record and marking the buffer dirty. We solve
-			 * that with a kluge, but one that is already in use during
-			 * transaction commit to prevent race conditions. Basically, we
-			 * simply prevent the checkpoint WAL record from being written
-			 * until we have marked the buffer dirty. We don't start the
-			 * checkpoint flush until we have marked dirty, so our checkpoint
-			 * must flush the change to disk successfully or the checkpoint
-			 * never gets written, so crash recovery will fix.
-			 *
-			 * It's possible we may enter here without an xid, so it is
-			 * essential that CreateCheckPoint waits for virtual transactions
-			 * rather than full transactionids.
-			 */
-			Assert((MyProc->delayChkptFlags & DELAY_CHKPT_START) == 0);
-			MyProc->delayChkptFlags |= DELAY_CHKPT_START;
-			delayChkptFlags = true;
-			lsn = XLogSaveBufferForHint(buffer, buffer_std);
+			wal_log = true;
 		}
 
+		/*
+		 * We must mark the page dirty before we emit the WAL record, as per
+		 * the usual rules, to ensure that BufferSync()/SyncOneBuffer() try to
+		 * flush the buffer, even if we haven't inserted the WAL record yet.
+		 * As we hold at least a share-exclusive lock, checkpoints will wait
+		 * for this backend to be done with the buffer before continuing. If
+		 * we did it the other way round, a checkpoint could start between
+		 * writing the WAL record and marking the buffer dirty.
+		 */
 		buf_state = LockBufHdr(bufHdr);
 
 		/*
@@ -5637,6 +5620,19 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate,
 		 */
 		Assert(!(buf_state & BM_DIRTY));
 		Assert(BUF_STATE_GET_REFCOUNT(buf_state) > 0);
+		UnlockBufHdrExt(bufHdr, buf_state,
+						BM_DIRTY,
+						0, 0);
+
+		/*
+		 * If the block is already dirty because we either made a change or
+		 * set a hint already, then we don't need to write a full page image.
+		 * Note that aggressive cleaning of blocks dirtied by hint bit setting
+		 * would increase the call rate. Bulk setting of hint bits would
+		 * reduce the call rate...
+		 */
+		if (wal_log)
+			lsn = XLogSaveBufferForHint(buffer, buffer_std);
 
 		if (XLogRecPtrIsValid(lsn))
 		{
@@ -5653,16 +5649,11 @@ MarkSharedBufferDirtyHint(Buffer buffer, BufferDesc *bufHdr, uint64 lockstate,
 			 * checksum here. That will happen when the page is written
 			 * sometime later in this checkpoint cycle.
 			 */
+			buf_state = LockBufHdr(bufHdr);
 			PageSetLSN(page, lsn);
+			UnlockBufHdr(bufHdr);
 		}
 
-		UnlockBufHdrExt(bufHdr, buf_state,
-						BM_DIRTY,
-						0, 0);
-
-		if (delayChkptFlags)
-			MyProc->delayChkptFlags &= ~DELAY_CHKPT_START;
-
 		pgBufferUsage.shared_blks_dirtied++;
 		if (VacuumCostActive)
 			VacuumCostBalance += VacuumCostPageDirty;
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v12-0005-bufmgr-Don-t-copy-pages-while-writing-out.patch (10.9K, ../../5ubipyssiju5twkb7zgqwdr7q2vhpkpmuelxfpanetlk6ofnop@hvxb4g2amb2d/6-v12-0005-bufmgr-Don-t-copy-pages-while-writing-out.patch)
  download | inline diff:
From 8464504a923a120381f965cbb3624bdf52b65212 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 13 Jan 2026 20:10:32 -0500
Subject: [PATCH v12 5/6] bufmgr: Don't copy pages while writing out

After the series of preceding commits introducing and using
BufferBeginSetHintBits()/BufferSetHintBits16(), hint bits are not set anymore
while IO is going on. Therefore we do not need to copy pages while they are
being written out anymore.

For the same reason XLogSaveBufferForHint() now does not need to operate on a
copy of the page anymore, but can instead use the normal XLogRegisterBuffer()
mechanism. For that the assertions and comments to XLogRegisterBuffer() had to
be updated to allow share-exclusive locked buffers to be registered.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/include/storage/bufpage.h           |  3 +-
 src/backend/access/hash/hashpage.c      |  2 +-
 src/backend/access/transam/xloginsert.c | 45 +++++------------------
 src/backend/storage/buffer/bufmgr.c     | 11 ++----
 src/backend/storage/buffer/localbuf.c   |  2 +-
 src/backend/storage/page/bufpage.c      | 48 ++++---------------------
 src/backend/storage/smgr/bulk_write.c   |  2 +-
 src/test/modules/test_aio/test_aio.c    |  2 +-
 8 files changed, 23 insertions(+), 92 deletions(-)

diff --git a/src/include/storage/bufpage.h b/src/include/storage/bufpage.h
index ae3725b3b81..31ec9a8a047 100644
--- a/src/include/storage/bufpage.h
+++ b/src/include/storage/bufpage.h
@@ -504,7 +504,6 @@ extern void PageIndexMultiDelete(Page page, OffsetNumber *itemnos, int nitems);
 extern void PageIndexTupleDeleteNoCompact(Page page, OffsetNumber offnum);
 extern bool PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 									const void *newtup, Size newsize);
-extern char *PageSetChecksumCopy(Page page, BlockNumber blkno);
-extern void PageSetChecksumInplace(Page page, BlockNumber blkno);
+extern void PageSetChecksum(Page page, BlockNumber blkno);
 
 #endif							/* BUFPAGE_H */
diff --git a/src/backend/access/hash/hashpage.c b/src/backend/access/hash/hashpage.c
index 8e220a3ae16..52c20208c66 100644
--- a/src/backend/access/hash/hashpage.c
+++ b/src/backend/access/hash/hashpage.c
@@ -1029,7 +1029,7 @@ _hash_alloc_buckets(Relation rel, BlockNumber firstblock, uint32 nblocks)
 					zerobuf.data,
 					true);
 
-	PageSetChecksumInplace(page, lastblock);
+	PageSetChecksum(page, lastblock);
 	smgrextend(RelationGetSmgr(rel), MAIN_FORKNUM, lastblock, zerobuf.data,
 			   false);
 
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index 7f27eee5ba1..c15c6aa7161 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -251,9 +251,9 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	Assert(begininsert_called);
 
 	/*
-	 * Ordinarily, buffer should be exclusive-locked and marked dirty before
-	 * we get here, otherwise we could end up violating one of the rules in
-	 * access/transam/README.
+	 * Ordinarily, the buffer should be exclusive-locked (or share-exclusive
+	 * in case of hint bits) and marked dirty before we get here, otherwise we
+	 * could end up violating one of the rules in access/transam/README.
 	 *
 	 * Some callers intentionally register a clean page and never update that
 	 * page's LSN; in that case they can pass the flag REGBUF_NO_CHANGE to
@@ -261,8 +261,11 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	 */
 #ifdef USE_ASSERT_CHECKING
 	if (!(flags & REGBUF_NO_CHANGE))
-		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) &&
-			   BufferIsDirty(buffer));
+	{
+		Assert(BufferIsDirty(buffer));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE));
+	}
 #endif
 
 	if (block_id >= max_registered_block_id)
@@ -1071,12 +1074,6 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * buffer. The buffer already needs to have been marked dirty by
  * MarkBufferDirtyHint().
  *
- * We can't use the plain backup block mechanism since that relies on the
- * Buffer being exclusively locked. Since some modifications (setting LSN, hint
- * bits) are allowed in a sharelocked buffer that can lead to wal checksum
- * failures. So instead we copy the page and insert the copied data as normal
- * record data.
- *
  * We only need to do something if page has not yet been full page written in
  * this checkpoint round. The LSN of the inserted wal record is returned if we
  * had to write, InvalidXLogRecPtr otherwise.
@@ -1112,37 +1109,13 @@ XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
 	if (lsn <= RedoRecPtr)
 	{
 		int			flags = 0;
-		PGAlignedBlock copied_buffer;
-		char	   *origdata = (char *) BufferGetBlock(buffer);
-		RelFileLocator rlocator;
-		ForkNumber	forkno;
-		BlockNumber blkno;
-
-		/*
-		 * Copy buffer so we don't have to worry about concurrent hint bit or
-		 * lsn updates. We assume pd_lower/upper cannot be changed without an
-		 * exclusive lock, so the contents bkp are not racy.
-		 */
-		if (buffer_std)
-		{
-			/* Assume we can omit data between pd_lower and pd_upper */
-			Page		page = BufferGetPage(buffer);
-			uint16		lower = ((PageHeader) page)->pd_lower;
-			uint16		upper = ((PageHeader) page)->pd_upper;
-
-			memcpy(copied_buffer.data, origdata, lower);
-			memcpy(copied_buffer.data + upper, origdata + upper, BLCKSZ - upper);
-		}
-		else
-			memcpy(copied_buffer.data, origdata, BLCKSZ);
 
 		XLogBeginInsert();
 
 		if (buffer_std)
 			flags |= REGBUF_STANDARD;
 
-		BufferGetTag(buffer, &rlocator, &forkno, &blkno);
-		XLogRegisterBlock(0, &rlocator, forkno, blkno, copied_buffer.data, flags);
+		XLogRegisterBuffer(0, buffer, flags);
 
 		recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI_FOR_HINT);
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 929466d25fd..4a9107cb47a 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -4420,7 +4420,6 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	ErrorContextCallback errcallback;
 	instr_time	io_start;
 	Block		bufBlock;
-	char	   *bufToWrite;
 
 	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
 		   BufferLockHeldByMeInMode(buf, BUFFER_LOCK_SHARE_EXCLUSIVE));
@@ -4483,12 +4482,8 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	 */
 	bufBlock = BufHdrGetBlock(buf);
 
-	/*
-	 * Update page checksum if desired.  Since we have only shared lock on the
-	 * buffer, other processes might be updating hint bits in it, so we must
-	 * copy the page to private storage if we do checksumming.
-	 */
-	bufToWrite = PageSetChecksumCopy((Page) bufBlock, buf->tag.blockNum);
+	/* Update page checksum if desired. */
+	PageSetChecksum((Page) bufBlock, buf->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
@@ -4498,7 +4493,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	smgrwrite(reln,
 			  BufTagGetForkNum(&buf->tag),
 			  buf->tag.blockNum,
-			  bufToWrite,
+			  bufBlock,
 			  false);
 
 	/*
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 404c6bccbdd..b69398c6375 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -199,7 +199,7 @@ FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln)
 		reln = smgropen(BufTagGetRelFileLocator(&bufHdr->tag),
 						MyProcNumber);
 
-	PageSetChecksumInplace(localpage, bufHdr->tag.blockNum);
+	PageSetChecksum(localpage, bufHdr->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index de85911e3ac..5cc92e68079 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -1494,51 +1494,15 @@ PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 /*
  * Set checksum for a page in shared buffers.
  *
- * If checksums are disabled, or if the page is not initialized, just return
- * the input.  Otherwise, we must make a copy of the page before calculating
- * the checksum, to prevent concurrent modifications (e.g. setting hint bits)
- * from making the final checksum invalid.  It doesn't matter if we include or
- * exclude hints during the copy, as long as we write a valid page and
- * associated checksum.
+ * If checksums are disabled, or if the page is not initialized, just
+ * return. Otherwise compute and set the checksum.
  *
- * Returns a pointer to the block-sized data that needs to be written. Uses
- * statically-allocated memory, so the caller must immediately write the
- * returned page and not refer to it again.
- */
-char *
-PageSetChecksumCopy(Page page, BlockNumber blkno)
-{
-	static char *pageCopy = NULL;
-
-	/* If we don't need a checksum, just return the passed-in data */
-	if (PageIsNew(page) || !DataChecksumsEnabled())
-		return page;
-
-	/*
-	 * We allocate the copy space once and use it over on each subsequent
-	 * call.  The point of palloc'ing here, rather than having a static char
-	 * array, is first to ensure adequate alignment for the checksumming code
-	 * and second to avoid wasting space in processes that never call this.
-	 */
-	if (pageCopy == NULL)
-		pageCopy = MemoryContextAllocAligned(TopMemoryContext,
-											 BLCKSZ,
-											 PG_IO_ALIGN_SIZE,
-											 0);
-
-	memcpy(pageCopy, page, BLCKSZ);
-	((PageHeader) pageCopy)->pd_checksum = pg_checksum_page(pageCopy, blkno);
-	return pageCopy;
-}
-
-/*
- * Set checksum for a page in private memory.
- *
- * This must only be used when we know that no other process can be modifying
- * the page buffer.
+ * In the past this needed to be done on a copy of the page, due to the
+ * possibility of e.g., hint bits being set concurrently. However, this is not
+ * necessary anymore as hint bits won't be set while IO is going on.
  */
 void
-PageSetChecksumInplace(Page page, BlockNumber blkno)
+PageSetChecksum(Page page, BlockNumber blkno)
 {
 	/* If we don't need a checksum, just return */
 	if (PageIsNew(page) || !DataChecksumsEnabled())
diff --git a/src/backend/storage/smgr/bulk_write.c b/src/backend/storage/smgr/bulk_write.c
index 36b28824ec8..f3c24082a69 100644
--- a/src/backend/storage/smgr/bulk_write.c
+++ b/src/backend/storage/smgr/bulk_write.c
@@ -279,7 +279,7 @@ smgr_bulk_flush(BulkWriteState *bulkstate)
 		BlockNumber blkno = pending_writes[i].blkno;
 		Page		page = pending_writes[i].buf->data;
 
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 
 		if (blkno >= bulkstate->relsize)
 		{
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index b1aa8af9ec0..2ae4a559fab 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -288,7 +288,7 @@ modify_rel_block(PG_FUNCTION_ARGS)
 	}
 	else
 	{
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 	}
 
 	smgrwrite(RelationGetSmgr(rel),
-- 
2.48.1.76.g4e746b1a31.dirty

  [text/x-diff] v12-0006-WIP-Make-UnlockReleaseBuffer-more-efficient.patch (3.5K, ../../5ubipyssiju5twkb7zgqwdr7q2vhpkpmuelxfpanetlk6ofnop@hvxb4g2amb2d/7-v12-0006-WIP-Make-UnlockReleaseBuffer-more-efficient.patch)
  download | inline diff:
From 1ff37e5e56b848a6f4f5fe1869112a93e8a7bf04 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 13 Jan 2026 20:10:32 -0500
Subject: [PATCH v12 6/6] WIP: Make UnlockReleaseBuffer() more efficient

Now that the buffer content lock is implemented as part of BufferDesc.state,
releasing the lock and unpinning the buffer can be implemented as a single
atomic operation.

Author:
Reviewed-By:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/access/nbtree/nbtpage.c | 22 +++++++++++-
 src/backend/storage/buffer/bufmgr.c | 52 ++++++++++++++++++++++++++++-
 2 files changed, 72 insertions(+), 2 deletions(-)

diff --git a/src/backend/access/nbtree/nbtpage.c b/src/backend/access/nbtree/nbtpage.c
index 4125c185e8b..f3e3f67e1fd 100644
--- a/src/backend/access/nbtree/nbtpage.c
+++ b/src/backend/access/nbtree/nbtpage.c
@@ -1007,11 +1007,18 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 
 	Assert(BlockNumberIsValid(blkno));
 	if (BufferIsValid(obuf))
+	{
+		_bt_relbuf(rel, obuf);
+#if 0
+		Assert(BufferGetBlockNumber(obuf) != blkno);
 		_bt_unlockbuf(rel, obuf);
-	buf = ReleaseAndReadBuffer(obuf, rel, blkno);
+#endif
+	}
+	buf = ReadBuffer(rel, blkno);
 	_bt_lockbuf(rel, buf, access);
 
 	_bt_checkpage(rel, buf);
+
 	return buf;
 }
 
@@ -1023,8 +1030,21 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 void
 _bt_relbuf(Relation rel, Buffer buf)
 {
+#if 0
 	_bt_unlockbuf(rel, buf);
 	ReleaseBuffer(buf);
+#else
+	/*
+	 * Buffer is pinned and locked, which means that it is expected to be
+	 * defined and addressable.  Check that proactively.
+	 */
+	VALGRIND_CHECK_MEM_IS_DEFINED(BufferGetPage(buf), BLCKSZ);
+
+	UnlockReleaseBuffer(buf);
+
+	if (!RelationUsesLocalBuffers(rel))
+		VALGRIND_MAKE_MEM_NOACCESS(BufferGetPage(buf), BLCKSZ);
+#endif
 }
 
 /*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 4a9107cb47a..8a4fb7c30d7 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5502,13 +5502,63 @@ ReleaseBuffer(Buffer buffer)
 /*
  * UnlockReleaseBuffer -- release the content lock and pin on a buffer
  *
- * This is just a shorthand for a common combination.
+ * This is just a, more efficient, shorthand for a common combination.
  */
 void
 UnlockReleaseBuffer(Buffer buffer)
 {
+#if 1
+	int			mode;
+	BufferDesc *buf;
+	PrivateRefCountEntry *ref;
+	uint64		sub;
+	uint64		lockstate;
+
+	if (!BufferIsValid(buffer))
+		elog(ERROR, "bad buffer ID: %d", buffer);
+
+	if (BufferIsLocal(buffer))
+	{
+		UnpinLocalBuffer(buffer);
+		return;
+	}
+
+	ResourceOwnerForgetBuffer(CurrentResourceOwner, buffer);
+
+	buf = GetBufferDescriptor(buffer - 1);
+
+	mode = BufferLockDisownInternal(buffer, buf);
+
+	/* compute state modification for lock release */
+	sub = BufferLockReleaseSub(mode);
+
+	/* compute state modification for pin release */
+	ref = GetPrivateRefCountEntry(buffer, false);
+	Assert(ref != NULL);
+	Assert(ref->data.refcount > 0);
+	ref->data.refcount--;
+
+	if (ref->data.refcount == 0)
+	{
+		sub |= BUF_REFCOUNT_ONE;
+		ForgetPrivateRefCountEntry(ref);
+	}
+
+	/* perform the lock and pin release in one atomic op */
+	lockstate = pg_atomic_sub_fetch_u64(&buf->state, sub);
+
+	/* wake up waiters etc */
+	BufferLockProcessRelease(buf, mode, lockstate);
+
+	if (lockstate & BM_PIN_COUNT_WAITER)
+		WakePinCountWaiter(buf);
+
+	RESUME_INTERRUPTS();
+
+#else
 	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 	ReleaseBuffer(buffer);
+#endif
 }
 
 /*
-- 
2.48.1.76.g4e746b1a31.dirty

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-02-06 21:18                                               ` Kirill Reshke <reshkekirill@gmail.com>
  3 siblings, 0 replies; 120+ messages in thread

From: Kirill Reshke @ 2026-02-06 21:18 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Noah Misch <noah@leadboat.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Tue, 3 Feb 2026 at 03:33, Andres Freund <andres@anarazel.de> wrote:
>
> Hi,
>
> On 2026-01-14 16:20:58 -0500, Andres Freund wrote:
> > I'm now working on cleaning up the last two commits. The most crucial bit is
> > to simplify what happens in MarkSharedBufferDirtyHint(), we afaict can delete
> > the use of DELAY_CHKPT_START etc and just go to marking the buffer dirty first
> > and then do the WAL logging, just like normal WAL logging. The previous order
> > was only required because we were dirtying the page while holding only a
> > shared lock, which did not conflict with the lock held by SyncBuffers() etc.
>
> I've been working on that.
>
> - A lot of what was special about MarkBufferDirtyHint() isn't needed anymore:
>
>   - The "abnormal" order of WAL logging before marking the buffer dirty was
>     only needed because we marked buffers dirty. Which in turn was only needed
>     because setting hint bits didn't conflict with flushing the page. With
>     share-exclusive they do conflict, and we can switch to the normal order of
>     operations, where marking a buffer dirty makes checkpoint wait when the
>     buffer is encountered (due to wanting to flush the buffer but not getting
>     the lock)
>
>
>   - Now that we use the normal order of WAL logging, we don't need to delay
>     checkpoint starts anymore.
>
>     I think the explanation for why that is ok is correct [1], but it needs to
>     be looked at by somebody with experience around this. Maybe Heikki?
>
>
>   - Thanks to holding share-exclusive lock, nothing can concurrently dirty or
>     undirty the buffer. Therefore the comments about spurious failures to mark
>     the buffer dirty can be removed.
>
>
> - I realized that, now that buffers cannot be dirtied while IO is ongoing, we
>   don't need BM_JUST_DIRTIED anymore.
>
>
> - The way MarkBufferDirtyHint() operates was copied into
>   heap_inplace_update_and_unlock(). Now that MarkBufferDirtyHint() won't work
>   that way anymore, it seems better to go with the alternative approach the
>   comments already outlined, namely to only delay updating of the buffer
>   contents.
>
>   I've done this in a prequisite commit, as it doesn't actually depend on any
>   of the other changes.  Noah, any chance you could take a look at this?
>
>
> - Lots of minor polish
>
>
> Greetings,
>
> Andres Freund
>
> [1]
>         /*
>          * Update RedoRecPtr so that we can make the right decision. It's possible
>          * that a new checkpoint will start just after GetRedoRecPtr(), but that
>          * is ok, as the buffer is already dirty, ensuring that any BufferSync()
>          * started after the buffer was marked dirty cannot complete without
>          * flushing this buffer.  If a checkpoint started between marking the
>          * buffer dirty and this check, we will emit an unnecessary WAL record (as
>          * the buffer will be written out as part of the checkpoint), but the
>          * window for that is small.
>          */

Hi!
I was reviewing new patches in this thread, and noticed your changes
in v12-0002, in gistkillitems. This makes me think you can be
interested in reviewing [0].

Anyway, v12-0003, on other patches I don't have an opinion (yet).

[0] https://commitfest.postgresql.org/patch/6399/
-- 
Best regards,
Kirill Reshke





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-02-07 10:44                                               ` Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  3 siblings, 1 reply; 120+ messages in thread

From: Heikki Linnakangas @ 2026-02-07 10:44 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; Melanie Plageman <melanieplageman@gmail.com>; Noah Misch <noah@leadboat.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On 03/02/2026 00:33, Andres Freund wrote:
> - The way MarkBufferDirtyHint() operates was copied into
>    heap_inplace_update_and_unlock(). Now that MarkBufferDirtyHint() won't work
>    that way anymore, it seems better to go with the alternative approach the
>    comments already outlined, namely to only delay updating of the buffer
>    contents.
> 
>    I've done this in a prequisite commit, as it doesn't actually depend on any
>    of the other changes.  Noah, any chance you could take a look at this?

Patch 0001 Looks correct to me. However:

> 	 * ["D" is a VACUUM (ONLY_DATABASE_STATS)]
> 	 * ["R" is a VACUUM tbl]
> 	 * D: vac_update_datfrozenxid() -> systable_beginscan(pg_class)
> 	 * D: systable_getnext() returns pg_class tuple of tbl
> 	 * R: memcpy() into pg_class tuple of tbl
> 	 * D: raise pg_database.datfrozenxid, XLogInsert(), finish
> 	 * [crash]
> 	 * [recovery restores datfrozenxid w/o relfrozenxid]
> 	 *
> 	 * As we hold an exclusive lock - preventing the buffer from being written
> 	 * out once dirty - we can work around this as follows: MarkBufferDirty(),
> 	 * XLogInsert(), memcpy().

That last reference to 'memcpy' is a little orphaned now. The comment 
used to talk about the stack copy of the page, but now there's no 
mention of that except for this reference to memcpy(). To make things 
worse, the steps have "memcpy() into pg_class tuple of tbl", so one 
could think that the "memcpy" refers to that.

How about this:

	 * We avoid that by using a temporary copy of the buffer to hide our
	 * change from other backends until it's been WAL-logged. We apply our
	 * change to the temporary copy and WAL-log it before modifying the real
	 * page. That way any action a reader of the in-place-updated value takes
	 * will be WAL logged after this change.

- Heikki





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
@ 2026-02-15 19:52                                                 ` Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Noah Misch @ 2026-02-15 19:52 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Sat, Feb 07, 2026 at 12:44:25PM +0200, Heikki Linnakangas wrote:
> On 03/02/2026 00:33, Andres Freund wrote:
> > - The way MarkBufferDirtyHint() operates was copied into
> >    heap_inplace_update_and_unlock(). Now that MarkBufferDirtyHint() won't work
> >    that way anymore, it seems better to go with the alternative approach the
> >    comments already outlined, namely to only delay updating of the buffer
> >    contents.
> > 
> >    I've done this in a prequisite commit, as it doesn't actually depend on any
> >    of the other changes.  Noah, any chance you could take a look at this?

v12-0001-heapam-Don-t-mimic-MarkBufferDirtyHint-in-inplac.patch looks good.

> Patch 0001 Looks correct to me. However:
> 
> > 	 * ["D" is a VACUUM (ONLY_DATABASE_STATS)]
> > 	 * ["R" is a VACUUM tbl]
> > 	 * D: vac_update_datfrozenxid() -> systable_beginscan(pg_class)
> > 	 * D: systable_getnext() returns pg_class tuple of tbl
> > 	 * R: memcpy() into pg_class tuple of tbl
> > 	 * D: raise pg_database.datfrozenxid, XLogInsert(), finish
> > 	 * [crash]
> > 	 * [recovery restores datfrozenxid w/o relfrozenxid]
> > 	 *
> > 	 * As we hold an exclusive lock - preventing the buffer from being written
> > 	 * out once dirty - we can work around this as follows: MarkBufferDirty(),
> > 	 * XLogInsert(), memcpy().
> 
> That last reference to 'memcpy' is a little orphaned now. The comment used
> to talk about the stack copy of the page, but now there's no mention of that
> except for this reference to memcpy(). To make things worse, the steps have
> "memcpy() into pg_class tuple of tbl", so one could think that the "memcpy"
> refers to that.

"memcpy" does refer to "memcpy() into pg_class tuple of tbl", so I don't see
that as orphaned.  Nonetheless:

> How about this:
> 
> 	 * We avoid that by using a temporary copy of the buffer to hide our
> 	 * change from other backends until it's been WAL-logged. We apply our
> 	 * change to the temporary copy and WAL-log it before modifying the real
> 	 * page. That way any action a reader of the in-place-updated value takes
> 	 * will be WAL logged after this change.

Either v12 or v12 w/ this edit is fine with me.  I find this proposed text
redundant with nearby comment "register block matching what buffer will look
like after changes", so I mildly prefer v12.





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
@ 2026-03-11 22:40                                                   ` Andres Freund <andres@anarazel.de>
  2026-03-13 08:00                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-03-25 21:34                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  0 siblings, 2 replies; 120+ messages in thread

From: Andres Freund @ 2026-03-11 22:40 UTC (permalink / raw)
  To: Noah Misch <noah@leadboat.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-02-15 11:52:39 -0800, Noah Misch wrote:
> On Sat, Feb 07, 2026 at 12:44:25PM +0200, Heikki Linnakangas wrote:
> > On 03/02/2026 00:33, Andres Freund wrote:
> > > - The way MarkBufferDirtyHint() operates was copied into
> > >    heap_inplace_update_and_unlock(). Now that MarkBufferDirtyHint() won't work
> > >    that way anymore, it seems better to go with the alternative approach the
> > >    comments already outlined, namely to only delay updating of the buffer
> > >    contents.
> > > 
> > >    I've done this in a prequisite commit, as it doesn't actually depend on any
> > >    of the other changes.  Noah, any chance you could take a look at this?
> 
> v12-0001-heapam-Don-t-mimic-MarkBufferDirtyHint-in-inplac.patch looks good.

> > How about this:
> > 
> > 	 * We avoid that by using a temporary copy of the buffer to hide our
> > 	 * change from other backends until it's been WAL-logged. We apply our
> > 	 * change to the temporary copy and WAL-log it before modifying the real
> > 	 * page. That way any action a reader of the in-place-updated value takes
> > 	 * will be WAL logged after this change.
> 
> Either v12 or v12 w/ this edit is fine with me.  I find this proposed text
> redundant with nearby comment "register block matching what buffer will look
> like after changes", so I mildly prefer v12.

Thanks for the review!


I pushed this and many of the later patches in the series.  Here are updated
versions of the remaining changes.  The last two previously were one commit
with "WIP" in the title. The first one has, I think, not had a lot of review -
but it's also not a complicated change.


I see decent performance improvements with a fully s_b resident pipelined
pgbench -S with 0002+0003, ~7-8% on an older small two socket machine.

The improvement is just from reducing the number of atomic operations on
contended cachelines (i.e. inner btree pages).

Without pipelining the difference is smaller (1-2%), because of the context
switches are the bigger bottleneck.


More extreme worloads involving an index nested loop join benefit
more. E.g. the setup and query from
https://anarazel.de/talks/2024-05-29-pgconf-dev-c2c/postgres-perf-c2c.pdf
slide 23, show a 25% improvement on the same 2 socket machine.


We could probably do something similar for the also very common combination of
PinBuffer() + LockBuffer(), but I think it'd be a fair bit more complicated,
and would require new APIs, rather than just using existing APIs more widely.

Greetings,

Andres Freund

Attachments:

  [text/x-diff] v13-0001-bufmgr-Don-t-copy-pages-while-writing-out.patch (10.9K, ../../mheeefrtikvgjnjsenocvo3afj7vlpr5rljkzevssxs77n2zdt@sugqtjhryydo/2-v13-0001-bufmgr-Don-t-copy-pages-while-writing-out.patch)
  download | inline diff:
From 66ef53cdd3d5a99539c984488e11dd330ecbdf50 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 13 Jan 2026 20:10:32 -0500
Subject: [PATCH v13 1/3] bufmgr: Don't copy pages while writing out

After the series of preceding commits introducing and using
BufferBeginSetHintBits()/BufferSetHintBits16(), hint bits are not set anymore
while IO is going on. Therefore we do not need to copy pages while they are
being written out anymore.

For the same reason XLogSaveBufferForHint() now does not need to operate on a
copy of the page anymore, but can instead use the normal XLogRegisterBuffer()
mechanism. For that the assertions and comments to XLogRegisterBuffer() had to
be updated to allow share-exclusive locked buffers to be registered.

Discussion: https://postgr.es/m/5ubipyssiju5twkb7zgqwdr7q2vhpkpmuelxfpanetlk6ofnop@hvxb4g2amb2d
---
 src/include/storage/bufpage.h           |  3 +-
 src/backend/access/hash/hashpage.c      |  2 +-
 src/backend/access/transam/xloginsert.c | 45 +++++------------------
 src/backend/storage/buffer/bufmgr.c     | 11 ++----
 src/backend/storage/buffer/localbuf.c   |  2 +-
 src/backend/storage/page/bufpage.c      | 48 ++++---------------------
 src/backend/storage/smgr/bulk_write.c   |  2 +-
 src/test/modules/test_aio/test_aio.c    |  2 +-
 8 files changed, 23 insertions(+), 92 deletions(-)

diff --git a/src/include/storage/bufpage.h b/src/include/storage/bufpage.h
index 3de58ba4312..e5267b93fe6 100644
--- a/src/include/storage/bufpage.h
+++ b/src/include/storage/bufpage.h
@@ -537,7 +537,6 @@ extern void PageIndexMultiDelete(Page page, OffsetNumber *itemnos, int nitems);
 extern void PageIndexTupleDeleteNoCompact(Page page, OffsetNumber offnum);
 extern bool PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 									const void *newtup, Size newsize);
-extern char *PageSetChecksumCopy(Page page, BlockNumber blkno);
-extern void PageSetChecksumInplace(Page page, BlockNumber blkno);
+extern void PageSetChecksum(Page page, BlockNumber blkno);
 
 #endif							/* BUFPAGE_H */
diff --git a/src/backend/access/hash/hashpage.c b/src/backend/access/hash/hashpage.c
index 8e220a3ae16..52c20208c66 100644
--- a/src/backend/access/hash/hashpage.c
+++ b/src/backend/access/hash/hashpage.c
@@ -1029,7 +1029,7 @@ _hash_alloc_buckets(Relation rel, BlockNumber firstblock, uint32 nblocks)
 					zerobuf.data,
 					true);
 
-	PageSetChecksumInplace(page, lastblock);
+	PageSetChecksum(page, lastblock);
 	smgrextend(RelationGetSmgr(rel), MAIN_FORKNUM, lastblock, zerobuf.data,
 			   false);
 
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index ac3c1a78396..baa12cfb3ef 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -251,9 +251,9 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	Assert(begininsert_called);
 
 	/*
-	 * Ordinarily, buffer should be exclusive-locked and marked dirty before
-	 * we get here, otherwise we could end up violating one of the rules in
-	 * access/transam/README.
+	 * Ordinarily, the buffer should be exclusive-locked (or share-exclusive
+	 * in case of hint bits) and marked dirty before we get here, otherwise we
+	 * could end up violating one of the rules in access/transam/README.
 	 *
 	 * Some callers intentionally register a clean page and never update that
 	 * page's LSN; in that case they can pass the flag REGBUF_NO_CHANGE to
@@ -261,8 +261,11 @@ XLogRegisterBuffer(uint8 block_id, Buffer buffer, uint8 flags)
 	 */
 #ifdef USE_ASSERT_CHECKING
 	if (!(flags & REGBUF_NO_CHANGE))
-		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) &&
-			   BufferIsDirty(buffer));
+	{
+		Assert(BufferIsDirty(buffer));
+		Assert(BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_EXCLUSIVE) ||
+			   BufferIsLockedByMeInMode(buffer, BUFFER_LOCK_SHARE_EXCLUSIVE));
+	}
 #endif
 
 	if (block_id >= max_registered_block_id)
@@ -1071,12 +1074,6 @@ XLogCheckBufferNeedsBackup(Buffer buffer)
  * buffer. The buffer already needs to have been marked dirty by
  * MarkBufferDirtyHint().
  *
- * We can't use the plain backup block mechanism since that relies on the
- * Buffer being exclusively locked. Since some modifications (setting LSN, hint
- * bits) are allowed in a sharelocked buffer that can lead to wal checksum
- * failures. So instead we copy the page and insert the copied data as normal
- * record data.
- *
  * We only need to do something if page has not yet been full page written in
  * this checkpoint round. The LSN of the inserted wal record is returned if we
  * had to write, InvalidXLogRecPtr otherwise.
@@ -1112,37 +1109,13 @@ XLogSaveBufferForHint(Buffer buffer, bool buffer_std)
 	if (lsn <= RedoRecPtr)
 	{
 		int			flags = 0;
-		PGAlignedBlock copied_buffer;
-		char	   *origdata = (char *) BufferGetBlock(buffer);
-		RelFileLocator rlocator;
-		ForkNumber	forkno;
-		BlockNumber blkno;
-
-		/*
-		 * Copy buffer so we don't have to worry about concurrent hint bit or
-		 * lsn updates. We assume pd_lower/upper cannot be changed without an
-		 * exclusive lock, so the contents bkp are not racy.
-		 */
-		if (buffer_std)
-		{
-			/* Assume we can omit data between pd_lower and pd_upper */
-			Page		page = BufferGetPage(buffer);
-			uint16		lower = ((PageHeader) page)->pd_lower;
-			uint16		upper = ((PageHeader) page)->pd_upper;
-
-			memcpy(copied_buffer.data, origdata, lower);
-			memcpy(copied_buffer.data + upper, origdata + upper, BLCKSZ - upper);
-		}
-		else
-			memcpy(copied_buffer.data, origdata, BLCKSZ);
 
 		XLogBeginInsert();
 
 		if (buffer_std)
 			flags |= REGBUF_STANDARD;
 
-		BufferGetTag(buffer, &rlocator, &forkno, &blkno);
-		XLogRegisterBlock(0, &rlocator, forkno, blkno, copied_buffer.data, flags);
+		XLogRegisterBuffer(0, buffer, flags);
 
 		recptr = XLogInsert(RM_XLOG_ID, XLOG_FPI_FOR_HINT);
 	}
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 6ded968e163..e5c06882395 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -4416,7 +4416,6 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	ErrorContextCallback errcallback;
 	instr_time	io_start;
 	Block		bufBlock;
-	char	   *bufToWrite;
 
 	Assert(BufferLockHeldByMeInMode(buf, BUFFER_LOCK_EXCLUSIVE) ||
 		   BufferLockHeldByMeInMode(buf, BUFFER_LOCK_SHARE_EXCLUSIVE));
@@ -4479,12 +4478,8 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	 */
 	bufBlock = BufHdrGetBlock(buf);
 
-	/*
-	 * Update page checksum if desired.  Since we have only shared lock on the
-	 * buffer, other processes might be updating hint bits in it, so we must
-	 * copy the page to private storage if we do checksumming.
-	 */
-	bufToWrite = PageSetChecksumCopy((Page) bufBlock, buf->tag.blockNum);
+	/* Update page checksum if desired. */
+	PageSetChecksum((Page) bufBlock, buf->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
@@ -4494,7 +4489,7 @@ FlushBuffer(BufferDesc *buf, SMgrRelation reln, IOObject io_object,
 	smgrwrite(reln,
 			  BufTagGetForkNum(&buf->tag),
 			  buf->tag.blockNum,
-			  bufToWrite,
+			  bufBlock,
 			  false);
 
 	/*
diff --git a/src/backend/storage/buffer/localbuf.c b/src/backend/storage/buffer/localbuf.c
index 404c6bccbdd..b69398c6375 100644
--- a/src/backend/storage/buffer/localbuf.c
+++ b/src/backend/storage/buffer/localbuf.c
@@ -199,7 +199,7 @@ FlushLocalBuffer(BufferDesc *bufHdr, SMgrRelation reln)
 		reln = smgropen(BufTagGetRelFileLocator(&bufHdr->tag),
 						MyProcNumber);
 
-	PageSetChecksumInplace(localpage, bufHdr->tag.blockNum);
+	PageSetChecksum(localpage, bufHdr->tag.blockNum);
 
 	io_start = pgstat_prepare_io_time(track_io_timing);
 
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index de85911e3ac..5cc92e68079 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -1494,51 +1494,15 @@ PageIndexTupleOverwrite(Page page, OffsetNumber offnum,
 /*
  * Set checksum for a page in shared buffers.
  *
- * If checksums are disabled, or if the page is not initialized, just return
- * the input.  Otherwise, we must make a copy of the page before calculating
- * the checksum, to prevent concurrent modifications (e.g. setting hint bits)
- * from making the final checksum invalid.  It doesn't matter if we include or
- * exclude hints during the copy, as long as we write a valid page and
- * associated checksum.
+ * If checksums are disabled, or if the page is not initialized, just
+ * return. Otherwise compute and set the checksum.
  *
- * Returns a pointer to the block-sized data that needs to be written. Uses
- * statically-allocated memory, so the caller must immediately write the
- * returned page and not refer to it again.
- */
-char *
-PageSetChecksumCopy(Page page, BlockNumber blkno)
-{
-	static char *pageCopy = NULL;
-
-	/* If we don't need a checksum, just return the passed-in data */
-	if (PageIsNew(page) || !DataChecksumsEnabled())
-		return page;
-
-	/*
-	 * We allocate the copy space once and use it over on each subsequent
-	 * call.  The point of palloc'ing here, rather than having a static char
-	 * array, is first to ensure adequate alignment for the checksumming code
-	 * and second to avoid wasting space in processes that never call this.
-	 */
-	if (pageCopy == NULL)
-		pageCopy = MemoryContextAllocAligned(TopMemoryContext,
-											 BLCKSZ,
-											 PG_IO_ALIGN_SIZE,
-											 0);
-
-	memcpy(pageCopy, page, BLCKSZ);
-	((PageHeader) pageCopy)->pd_checksum = pg_checksum_page(pageCopy, blkno);
-	return pageCopy;
-}
-
-/*
- * Set checksum for a page in private memory.
- *
- * This must only be used when we know that no other process can be modifying
- * the page buffer.
+ * In the past this needed to be done on a copy of the page, due to the
+ * possibility of e.g., hint bits being set concurrently. However, this is not
+ * necessary anymore as hint bits won't be set while IO is going on.
  */
 void
-PageSetChecksumInplace(Page page, BlockNumber blkno)
+PageSetChecksum(Page page, BlockNumber blkno)
 {
 	/* If we don't need a checksum, just return */
 	if (PageIsNew(page) || !DataChecksumsEnabled())
diff --git a/src/backend/storage/smgr/bulk_write.c b/src/backend/storage/smgr/bulk_write.c
index 36b28824ec8..f3c24082a69 100644
--- a/src/backend/storage/smgr/bulk_write.c
+++ b/src/backend/storage/smgr/bulk_write.c
@@ -279,7 +279,7 @@ smgr_bulk_flush(BulkWriteState *bulkstate)
 		BlockNumber blkno = pending_writes[i].blkno;
 		Page		page = pending_writes[i].buf->data;
 
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 
 		if (blkno >= bulkstate->relsize)
 		{
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index b1aa8af9ec0..2ae4a559fab 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -288,7 +288,7 @@ modify_rel_block(PG_FUNCTION_ARGS)
 	}
 	else
 	{
-		PageSetChecksumInplace(page, blkno);
+		PageSetChecksum(page, blkno);
 	}
 
 	smgrwrite(RelationGetSmgr(rel),
-- 
2.53.0.1.gb2826b52eb

  [text/x-diff] v13-0002-Use-UnlockReleaseBuffer-in-more-places.patch (7.8K, ../../mheeefrtikvgjnjsenocvo3afj7vlpr5rljkzevssxs77n2zdt@sugqtjhryydo/3-v13-0002-Use-UnlockReleaseBuffer-in-more-places.patch)
  download | inline diff:
From abf51df70c27ad0fd255ffe5877a25b556709baa Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Wed, 11 Mar 2026 15:12:53 -0400
Subject: [PATCH v13 2/3] Use UnlockReleaseBuffer() in more places

An upcoming commit will make UnlockReleaseBuffer() considerably faster and
more scalable than doing LockBuffer(BUFFER_LOCK_UNLOCK); ReleaseBuffer();. But
it's a small performance benefit even as-is.

Most of the callsites changed in this patch are not performance sensitive,
however some, like the nbtree ones, are in critical paths.

This patch changes all the easily convertible places over to
UnlockReleaseBuffer() mainly because I needed to check all of them anyway, and
reducing cases where the operations are done separately makes the checking
easier.

Discussion: https://postgr.es/m/
---
 src/backend/access/heap/heapam.c          |  6 ++--
 src/backend/access/heap/hio.c             |  7 +++--
 src/backend/access/nbtree/nbtpage.c       | 36 +++++++++++++++++++----
 src/backend/storage/buffer/bufmgr.c       |  3 +-
 src/backend/storage/freespace/freespace.c |  3 +-
 contrib/amcheck/verify_gin.c              |  6 ++--
 contrib/pageinspect/rawpage.c             |  3 +-
 contrib/pgstattuple/pgstatindex.c         |  3 +-
 src/test/modules/test_aio/test_aio.c      |  7 ++---
 9 files changed, 44 insertions(+), 30 deletions(-)

diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 8f1c11a9350..7b988ff36df 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -1687,8 +1687,7 @@ heap_fetch(Relation relation,
 	offnum = ItemPointerGetOffsetNumber(tid);
 	if (offnum < FirstOffsetNumber || offnum > PageGetMaxOffsetNumber(page))
 	{
-		LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
-		ReleaseBuffer(buffer);
+		UnlockReleaseBuffer(buffer);
 		*userbuf = InvalidBuffer;
 		tuple->t_data = NULL;
 		return false;
@@ -1704,8 +1703,7 @@ heap_fetch(Relation relation,
 	 */
 	if (!ItemIdIsNormal(lp))
 	{
-		LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
-		ReleaseBuffer(buffer);
+		UnlockReleaseBuffer(buffer);
 		*userbuf = InvalidBuffer;
 		tuple->t_data = NULL;
 		return false;
diff --git a/src/backend/access/heap/hio.c b/src/backend/access/heap/hio.c
index d26ceacd38c..1097f44a74e 100644
--- a/src/backend/access/heap/hio.c
+++ b/src/backend/access/heap/hio.c
@@ -711,14 +711,15 @@ loop:
 		 * unlock the two buffers in, so this can be slightly simpler than the
 		 * code above.
 		 */
-		LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 		if (otherBuffer == InvalidBuffer)
-			ReleaseBuffer(buffer);
+			UnlockReleaseBuffer(buffer);
 		else if (otherBlock != targetBlock)
 		{
+			UnlockReleaseBuffer(buffer);
 			LockBuffer(otherBuffer, BUFFER_LOCK_UNLOCK);
-			ReleaseBuffer(buffer);
 		}
+		else
+			LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
 
 		/* Is there an ongoing bulk extension? */
 		if (bistate && bistate->next_free != InvalidBlockNumber)
diff --git a/src/backend/access/nbtree/nbtpage.c b/src/backend/access/nbtree/nbtpage.c
index 4125c185e8b..acefe68a382 100644
--- a/src/backend/access/nbtree/nbtpage.c
+++ b/src/backend/access/nbtree/nbtpage.c
@@ -1007,24 +1007,48 @@ _bt_relandgetbuf(Relation rel, Buffer obuf, BlockNumber blkno, int access)
 
 	Assert(BlockNumberIsValid(blkno));
 	if (BufferIsValid(obuf))
-		_bt_unlockbuf(rel, obuf);
-	buf = ReleaseAndReadBuffer(obuf, rel, blkno);
+	{
+		if (BufferGetBlockNumber(obuf) == blkno)
+		{
+			/* trade in old lock mode for new lock */
+			_bt_unlockbuf(rel, obuf);
+			buf = obuf;
+		}
+		else
+		{
+			/* release lock and pin at once, that's a bit more efficient */
+			_bt_relbuf(rel, obuf);
+			buf = ReadBuffer(rel, blkno);
+		}
+	}
+	else
+		buf = ReadBuffer(rel, blkno);
+
 	_bt_lockbuf(rel, buf, access);
-
 	_bt_checkpage(rel, buf);
+
 	return buf;
 }
 
 /*
  *	_bt_relbuf() -- release a locked buffer.
  *
- * Lock and pin (refcount) are both dropped.
+ * Lock and pin (refcount) are both dropped. This is a bit more efficient than
+ * doing the two operations separately.
  */
 void
 _bt_relbuf(Relation rel, Buffer buf)
 {
-	_bt_unlockbuf(rel, buf);
-	ReleaseBuffer(buf);
+	/*
+	 * Buffer is pinned and locked, which means that it is expected to be
+	 * defined and addressable.  Check that proactively.
+	 */
+	VALGRIND_CHECK_MEM_IS_DEFINED(BufferGetPage(buf), BLCKSZ);
+
+	UnlockReleaseBuffer(buf);
+
+	if (!RelationUsesLocalBuffers(rel))
+		VALGRIND_MAKE_MEM_NOACCESS(BufferGetPage(buf), BLCKSZ);
 }
 
 /*
diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index e5c06882395..53b327200d7 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -2531,8 +2531,7 @@ again:
 			XLogNeedsFlush(BufferGetLSN(buf_hdr)) &&
 			StrategyRejectBuffer(strategy, buf_hdr, from_ring))
 		{
-			LockBuffer(buf, BUFFER_LOCK_UNLOCK);
-			UnpinBuffer(buf_hdr);
+			UnlockReleaseBuffer(buf);
 			goto again;
 		}
 
diff --git a/src/backend/storage/freespace/freespace.c b/src/backend/storage/freespace/freespace.c
index b9a8f368a63..40d67a96178 100644
--- a/src/backend/storage/freespace/freespace.c
+++ b/src/backend/storage/freespace/freespace.c
@@ -915,9 +915,8 @@ fsm_vacuum_page(Relation rel, FSMAddress addr,
 		((FSMPage) PageGetContents(page))->fp_next_slot = 0;
 		BufferFinishSetHintBits(buf, false, false);
 	}
-	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
 
-	ReleaseBuffer(buf);
+	UnlockReleaseBuffer(buf);
 
 	return max_avail;
 }
diff --git a/contrib/amcheck/verify_gin.c b/contrib/amcheck/verify_gin.c
index f7a15a21467..abfad07d5e4 100644
--- a/contrib/amcheck/verify_gin.c
+++ b/contrib/amcheck/verify_gin.c
@@ -368,8 +368,7 @@ gin_check_posting_tree_parent_keys_consistency(Relation rel, BlockNumber posting
 				stack->next = ptr;
 			}
 		}
-		LockBuffer(buffer, GIN_UNLOCK);
-		ReleaseBuffer(buffer);
+		UnlockReleaseBuffer(buffer);
 
 		/* Step to next item in the queue */
 		stack_next = stack->next;
@@ -642,8 +641,7 @@ gin_check_parent_keys_consistency(Relation rel,
 			prev_attnum = current_attnum;
 		}
 
-		LockBuffer(buffer, GIN_UNLOCK);
-		ReleaseBuffer(buffer);
+		UnlockReleaseBuffer(buffer);
 
 		/* Step to next item in the queue */
 		stack_next = stack->next;
diff --git a/contrib/pageinspect/rawpage.c b/contrib/pageinspect/rawpage.c
index 86fe245cac5..f3996dc39fc 100644
--- a/contrib/pageinspect/rawpage.c
+++ b/contrib/pageinspect/rawpage.c
@@ -193,8 +193,7 @@ get_raw_page_internal(text *relname, ForkNumber forknum, BlockNumber blkno)
 
 	memcpy(raw_page_data, BufferGetPage(buf), BLCKSZ);
 
-	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
-	ReleaseBuffer(buf);
+	UnlockReleaseBuffer(buf);
 
 	relation_close(rel, AccessShareLock);
 
diff --git a/contrib/pgstattuple/pgstatindex.c b/contrib/pgstattuple/pgstatindex.c
index ef723af1f19..a4716bd1b36 100644
--- a/contrib/pgstattuple/pgstatindex.c
+++ b/contrib/pgstattuple/pgstatindex.c
@@ -323,8 +323,7 @@ pgstatindex_impl(Relation rel, FunctionCallInfo fcinfo)
 			indexStat.internal_pages++;
 
 		/* Unlock and release buffer */
-		LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
-		ReleaseBuffer(buffer);
+		UnlockReleaseBuffer(buffer);
 	}
 
 	relation_close(rel, AccessShareLock);
diff --git a/src/test/modules/test_aio/test_aio.c b/src/test/modules/test_aio/test_aio.c
index 2ae4a559fab..138e1259dfd 100644
--- a/src/test/modules/test_aio/test_aio.c
+++ b/src/test/modules/test_aio/test_aio.c
@@ -221,9 +221,7 @@ modify_rel_block(PG_FUNCTION_ARGS)
 	 */
 	memcpy(page, BufferGetPage(buf), BLCKSZ);
 
-	LockBuffer(buf, BUFFER_LOCK_UNLOCK);
-
-	ReleaseBuffer(buf);
+	UnlockReleaseBuffer(buf);
 
 	/*
 	 * Don't want to have a buffer in-memory that's marked valid where the
@@ -496,8 +494,7 @@ invalidate_rel_block(PG_FUNCTION_ARGS)
 				else
 					FlushOneBuffer(buf);
 			}
-			LockBuffer(buf, BUFFER_LOCK_UNLOCK);
-			ReleaseBuffer(buf);
+			UnlockReleaseBuffer(buf);
 
 			if (BufferIsLocal(buf))
 				InvalidateLocalBuffer(GetLocalBufferDescriptor(-buf - 1), true);
-- 
2.53.0.1.gb2826b52eb

  [text/x-diff] v13-0003-bufmgr-Make-UnlockReleaseBuffer-more-efficient.patch (2.6K, ../../mheeefrtikvgjnjsenocvo3afj7vlpr5rljkzevssxs77n2zdt@sugqtjhryydo/4-v13-0003-bufmgr-Make-UnlockReleaseBuffer-more-efficient.patch)
  download | inline diff:
From 80eee071308c2548cb0dfc3f2dbdf04bd4b151bf Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 3 Feb 2026 11:01:58 -0500
Subject: [PATCH v13 3/3] bufmgr: Make UnlockReleaseBuffer() more efficient

Now that the buffer content lock is implemented as part of BufferDesc.state,
releasing the lock and unpinning the buffer can be implemented as a single
atomic operation.

This improves workloads that have heavy contention on a small number of
buffers substantially, I e.g., see a ~20% improvement for pipelined readonly
pgbench on an older two socket machine.

Author:
Reviewed-by:
Discussion: https://postgr.es/m/
Backpatch:
---
 src/backend/storage/buffer/bufmgr.c | 55 +++++++++++++++++++++++++++--
 1 file changed, 52 insertions(+), 3 deletions(-)

diff --git a/src/backend/storage/buffer/bufmgr.c b/src/backend/storage/buffer/bufmgr.c
index 53b327200d7..9147e4c127b 100644
--- a/src/backend/storage/buffer/bufmgr.c
+++ b/src/backend/storage/buffer/bufmgr.c
@@ -5508,13 +5508,62 @@ ReleaseBuffer(Buffer buffer)
 /*
  * UnlockReleaseBuffer -- release the content lock and pin on a buffer
  *
- * This is just a shorthand for a common combination.
+ * This is just a, more efficient, shorthand for a common combination.
  */
 void
 UnlockReleaseBuffer(Buffer buffer)
 {
-	LockBuffer(buffer, BUFFER_LOCK_UNLOCK);
-	ReleaseBuffer(buffer);
+	int			mode;
+	BufferDesc *buf;
+	PrivateRefCountEntry *ref;
+	uint64		sub;
+	uint64		lockstate;
+
+	Assert(BufferIsPinned(buffer));
+
+	if (BufferIsLocal(buffer))
+	{
+		UnpinLocalBuffer(buffer);
+		return;
+	}
+
+	ResourceOwnerForgetBuffer(CurrentResourceOwner, buffer);
+
+	buf = GetBufferDescriptor(buffer - 1);
+
+	mode = BufferLockDisownInternal(buffer, buf);
+
+	/* compute state modification for lock release */
+	sub = BufferLockReleaseSub(mode);
+
+	/* compute state modification for pin release */
+	ref = GetPrivateRefCountEntry(buffer, false);
+	Assert(ref != NULL);
+	Assert(ref->data.refcount > 0);
+	ref->data.refcount--;
+
+	/* no more backend local pins, reduce shared pin count */
+	if (likely(ref->data.refcount == 0))
+	{
+		sub |= BUF_REFCOUNT_ONE;
+		ForgetPrivateRefCountEntry(ref);
+	}
+
+	/* perform the lock and pin release in one atomic op */
+	lockstate = pg_atomic_sub_fetch_u64(&buf->state, sub);
+
+	/* wake up waiters for the lock */
+	BufferLockProcessRelease(buf, mode, lockstate);
+
+	/* wake up waiter for the pin release */
+	if (lockstate & BM_PIN_COUNT_WAITER)
+		WakePinCountWaiter(buf);
+
+	/*
+	 * Now okay to allow cancel/die interrupts again, were held when the lock
+	 * was acquired.
+	 */
+	RESUME_INTERRUPTS();
 }
 
 /*
-- 
2.53.0.1.gb2826b52eb

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-03-13 08:00                                                     ` Alexander Lakhin <exclusion@gmail.com>
  2026-03-13 15:55                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  1 sibling, 1 reply; 120+ messages in thread

From: Alexander Lakhin @ 2026-03-13 08:00 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; Noah Misch <noah@leadboat.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hello Andres,

12.03.2026 00:40, Andres Freund wrote:
> I pushed this and many of the later patches in the series.  Here are updated
> versions of the remaining changes.  The last two previously were one commit
> with "WIP" in the title. The first one has, I think, not had a lot of review -
> but it's also not a complicated change.

I've discovered that starting from 82467f627, the following query:
SET cpu_operator_cost = 1000;
CREATE TABLE t (i INT);
INSERT INTO T SELECT 1 FROM generate_series(1, 1000) a;
CREATE INDEX hi on t USING HASH (i);
DELETE FROM t WHERE i = 1;
DELETE FROM t WHERE i = 1;

triggers
TRAP: failed Assert("BufferIsValid(buffer)"), File: "bufmgr.c", Line: 497, PID: 3942058

#4  0x000079a60ae288ff in __GI_abort () at ./stdlib/abort.c:79
#5  0x00005a68d9343eef in ExceptionalCondition (conditionName=conditionName@entry=0x5a68d93ac27d "BufferIsValid(buffer)",
     fileName=fileName@entry=0x5a68d93c99ef "bufmgr.c", lineNumber=lineNumber@entry=497) at assert.c:65
#6  0x00005a68d91a18eb in GetPrivateRefCountEntry (do_move=true, buffer=<optimized out>) at bufmgr.c:497
#7  SharedBufferBeginSetHintBits (lockstate=<synthetic pointer>, buf_hdr=0x79e5febbbc40, buffer=<optimized out>)
     at bufmgr.c:6830
#8  BufferBeginSetHintBits (buffer=<optimized out>) at bufmgr.c:6931
#9  0x00005a68d8e3c862 in _hash_kill_items (scan=<optimized out>) at hashutil.c:603
#10 0x00005a68d8e3b7c3 in _hash_next (scan=0x5a68e735f938, dir=<optimized out>) at hashsearch.c:69
#11 0x00005a68d8e616ce in index_getnext_tid (scan=scan@entry=0x5a68e735f938, direction=direction@entry=ForwardScanDirection)
     at indexam.c:647
...
#25 0x00005a68d91eb4ad in exec_simple_query (query_string=0x5a68e7270120 "DELETE FROM t WHERE i = 1;") at postgres.c:1277
...

Could you please look at this?

Best regards,
Alexander

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-13 08:00                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
@ 2026-03-13 15:55                                                       ` Andres Freund <andres@anarazel.de>
  2026-03-17 20:50                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-03-13 15:55 UTC (permalink / raw)
  To: Alexander Lakhin <exclusion@gmail.com>; Alexander Kuzmenkov <akuzmenkov@tigerdata.com>; +Cc: Noah Misch <noah@leadboat.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-03-13 10:00:00 +0200, Alexander Lakhin wrote:
> Hello Andres,
> 
> 12.03.2026 00:40, Andres Freund wrote:
> > I pushed this and many of the later patches in the series.  Here are updated
> > versions of the remaining changes.  The last two previously were one commit
> > with "WIP" in the title. The first one has, I think, not had a lot of review -
> > but it's also not a complicated change.
> 
> I've discovered that starting from 82467f627, the following query:
> SET cpu_operator_cost = 1000;
> CREATE TABLE t (i INT);
> INSERT INTO T SELECT 1 FROM generate_series(1, 1000) a;
> CREATE INDEX hi on t USING HASH (i);
> DELETE FROM t WHERE i = 1;
> DELETE FROM t WHERE i = 1;
> 
> triggers
> TRAP: failed Assert("BufferIsValid(buffer)"), File: "bufmgr.c", Line: 497, PID: 3942058
> 
> #4  0x000079a60ae288ff in __GI_abort () at ./stdlib/abort.c:79
> #5  0x00005a68d9343eef in ExceptionalCondition (conditionName=conditionName@entry=0x5a68d93ac27d "BufferIsValid(buffer)",
>     fileName=fileName@entry=0x5a68d93c99ef "bufmgr.c", lineNumber=lineNumber@entry=497) at assert.c:65
> #6  0x00005a68d91a18eb in GetPrivateRefCountEntry (do_move=true, buffer=<optimized out>) at bufmgr.c:497
> #7  SharedBufferBeginSetHintBits (lockstate=<synthetic pointer>, buf_hdr=0x79e5febbbc40, buffer=<optimized out>)
>     at bufmgr.c:6830
> #8  BufferBeginSetHintBits (buffer=<optimized out>) at bufmgr.c:6931
> #9  0x00005a68d8e3c862 in _hash_kill_items (scan=<optimized out>) at hashutil.c:603
> #10 0x00005a68d8e3b7c3 in _hash_next (scan=0x5a68e735f938, dir=<optimized out>) at hashsearch.c:69
> #11 0x00005a68d8e616ce in index_getnext_tid (scan=scan@entry=0x5a68e735f938, direction=direction@entry=ForwardScanDirection)
>     at indexam.c:647
> ...
> #25 0x00005a68d91eb4ad in exec_simple_query (query_string=0x5a68e7270120 "DELETE FROM t WHERE i = 1;") at postgres.c:1277
> ...
> 
> Could you please look at this?

Yea, it's a stupid small mistake. Alexander Kuzmenkov reported it late
afternoon yesterday, privately as I just noticed, and I was too tired to make
sure an added test wouldn't have stability issues.

Will fix in the next few hours.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-13 08:00                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Alexander Lakhin <exclusion@gmail.com>
  2026-03-13 15:55                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-03-17 20:50                                                         ` Andres Freund <andres@anarazel.de>
  0 siblings, 0 replies; 120+ messages in thread

From: Andres Freund @ 2026-03-17 20:50 UTC (permalink / raw)
  To: Alexander Lakhin <exclusion@gmail.com>; Alexander Kuzmenkov <akuzmenkov@tigerdata.com>; +Cc: Noah Misch <noah@leadboat.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-03-13 11:55:53 -0400, Andres Freund wrote:
> On 2026-03-13 10:00:00 +0200, Alexander Lakhin wrote:
> > Hello Andres,
> > 
> > 12.03.2026 00:40, Andres Freund wrote:
> > > I pushed this and many of the later patches in the series.  Here are updated
> > > versions of the remaining changes.  The last two previously were one commit
> > > with "WIP" in the title. The first one has, I think, not had a lot of review -
> > > but it's also not a complicated change.
> > 
> > I've discovered that starting from 82467f627, the following query:
> > SET cpu_operator_cost = 1000;
> > CREATE TABLE t (i INT);
> > INSERT INTO T SELECT 1 FROM generate_series(1, 1000) a;
> > CREATE INDEX hi on t USING HASH (i);
> > DELETE FROM t WHERE i = 1;
> > DELETE FROM t WHERE i = 1;
> > 
> > triggers
> > TRAP: failed Assert("BufferIsValid(buffer)"), File: "bufmgr.c", Line: 497, PID: 3942058
> > 
> > #4  0x000079a60ae288ff in __GI_abort () at ./stdlib/abort.c:79
> > #5  0x00005a68d9343eef in ExceptionalCondition (conditionName=conditionName@entry=0x5a68d93ac27d "BufferIsValid(buffer)",
> >     fileName=fileName@entry=0x5a68d93c99ef "bufmgr.c", lineNumber=lineNumber@entry=497) at assert.c:65
> > #6  0x00005a68d91a18eb in GetPrivateRefCountEntry (do_move=true, buffer=<optimized out>) at bufmgr.c:497
> > #7  SharedBufferBeginSetHintBits (lockstate=<synthetic pointer>, buf_hdr=0x79e5febbbc40, buffer=<optimized out>)
> >     at bufmgr.c:6830
> > #8  BufferBeginSetHintBits (buffer=<optimized out>) at bufmgr.c:6931
> > #9  0x00005a68d8e3c862 in _hash_kill_items (scan=<optimized out>) at hashutil.c:603
> > #10 0x00005a68d8e3b7c3 in _hash_next (scan=0x5a68e735f938, dir=<optimized out>) at hashsearch.c:69
> > #11 0x00005a68d8e616ce in index_getnext_tid (scan=scan@entry=0x5a68e735f938, direction=direction@entry=ForwardScanDirection)
> >     at indexam.c:647
> > ...
> > #25 0x00005a68d91eb4ad in exec_simple_query (query_string=0x5a68e7270120 "DELETE FROM t WHERE i = 1;") at postgres.c:1277
> > ...
> > 
> > Could you please look at this?
> 
> Yea, it's a stupid small mistake. Alexander Kuzmenkov reported it late
> afternoon yesterday, privately as I just noticed, and I was too tired to make
> sure an added test wouldn't have stability issues.
> 
> Will fix in the next few hours.

Took longer, sorry.  But it's pushed now.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-03-25 21:34                                                     ` Melanie Plageman <melanieplageman@gmail.com>
  2026-03-25 22:35                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  1 sibling, 1 reply; 120+ messages in thread

From: Melanie Plageman @ 2026-03-25 21:34 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Noah Misch <noah@leadboat.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Wed, Mar 11, 2026 at 6:40 PM Andres Freund <andres@anarazel.de> wrote:
>
> I pushed this and many of the later patches in the series.  Here are updated
> versions of the remaining changes.  The last two previously were one commit
> with "WIP" in the title. The first one has, I think, not had a lot of review -
> but it's also not a complicated change.

0001 looks good except for the comment above PageSetChecksum() that
says it is only for shared buffers and a stray reference to the
no-longer-present bufToWrite variable in a comment around line 4490 in
bufmgr.c

0002
diff --git a/src/backend/access/nbtree/nbtpage.c
b/src/backend/access/nbtree/nbtpage.c
index cc9c45dc40c..ad700e590e8 100644
--- a/src/backend/access/nbtree/nbtpage.c
+++ b/src/backend/access/nbtree/nbtpage.c
@@ -1011,24 +1011,48 @@ _bt_relandgetbuf(Relation rel, Buffer obuf,
BlockNumber blkno, int access)
    Assert(BlockNumberIsValid(blkno));
    if (BufferIsValid(obuf))
-       _bt_unlockbuf(rel, obuf);
-   buf = ReleaseAndReadBuffer(obuf, rel, blkno);
-   _bt_lockbuf(rel, buf, access);
+   {
+       if (BufferGetBlockNumber(obuf) == blkno)
+       {
+           /* trade in old lock mode for new lock */
+           _bt_unlockbuf(rel, obuf);
+           buf = obuf;
+       }
+       else
+       {
+           /* release lock and pin at once, that's a bit more efficient */
+           _bt_relbuf(rel, obuf);
+           buf = ReadBuffer(rel, blkno);
+       }
+   }
+   else
+       buf = ReadBuffer(rel, blkno);

Not related to this patch, but why do we unlock and relock it when
obuf has the block we need? Couldn't we pass lock mode and then just
do nothing if it is the right lockmode?

Setting that aside, I presume we don't need to check the fork and
relfilelocator (as ReleaseAndReadBuffer() did) because this code knows
it will be the same?

Anyway, LGTM.

0003
AFAICT, this does what you claim. I don't really know what else to
look when reviewing it, if I'm being honest. As such, I diligently fed
it through AI which suggested you may have lost a
        VALGRIND_MAKE_MEM_NOACCESS(BufHdrGetBlock(buf), BLCKSZ);
which sounds right to me and like something you should fix.

Also, I'd say this comment
+   /*
+    * Now okay to allow cancel/die interrupts again, were held when the lock
+    * was acquired.
+    */

needs a "which" after the comma to read smoothly.

- Melanie





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-25 21:34                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
@ 2026-03-25 22:35                                                       ` Andres Freund <andres@anarazel.de>
  2026-03-27 20:00                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-03-25 22:35 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Noah Misch <noah@leadboat.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-03-25 17:34:33 -0400, Melanie Plageman wrote:
> On Wed, Mar 11, 2026 at 6:40 PM Andres Freund <andres@anarazel.de> wrote:
> >
> > I pushed this and many of the later patches in the series.  Here are updated
> > versions of the remaining changes.  The last two previously were one commit
> > with "WIP" in the title. The first one has, I think, not had a lot of review -
> > but it's also not a complicated change.
> 
> 0001 looks good except for the comment above PageSetChecksum() that
> says it is only for shared buffers and a stray reference to the
> no-longer-present bufToWrite variable in a comment around line 4490 in
> bufmgr.c

Thanks for catching these.

Updated the PageSetChecksum() comment to

 * Set checksum on a page.
 *
 * If the page is in shared buffers, it needs to be locked in at least
 * share-exclusive mode.
...


> 0002
> diff --git a/src/backend/access/nbtree/nbtpage.c
> b/src/backend/access/nbtree/nbtpage.c
> index cc9c45dc40c..ad700e590e8 100644
> --- a/src/backend/access/nbtree/nbtpage.c
> +++ b/src/backend/access/nbtree/nbtpage.c
> @@ -1011,24 +1011,48 @@ _bt_relandgetbuf(Relation rel, Buffer obuf,
> BlockNumber blkno, int access)
>     Assert(BlockNumberIsValid(blkno));
>     if (BufferIsValid(obuf))
> -       _bt_unlockbuf(rel, obuf);
> -   buf = ReleaseAndReadBuffer(obuf, rel, blkno);
> -   _bt_lockbuf(rel, buf, access);
> +   {
> +       if (BufferGetBlockNumber(obuf) == blkno)
> +       {
> +           /* trade in old lock mode for new lock */
> +           _bt_unlockbuf(rel, obuf);
> +           buf = obuf;
> +       }
> +       else
> +       {
> +           /* release lock and pin at once, that's a bit more efficient */
> +           _bt_relbuf(rel, obuf);
> +           buf = ReadBuffer(rel, blkno);
> +       }
> +   }
> +   else
> +       buf = ReadBuffer(rel, blkno);
> 
> Not related to this patch, but why do we unlock and relock it when
> obuf has the block we need? Couldn't we pass lock mode and then just
> do nothing if it is the right lockmode?

I think it's very unlikely that it's called at any frequency with the same
buffer and lockmode. What would be the point of calling _bt_relandgetbuf() if
that's the case.


> Setting that aside, I presume we don't need to check the fork and
> relfilelocator (as ReleaseAndReadBuffer() did) because this code knows
> it will be the same?

Yea, it's a single index, so there can't be a different relfilenode.


> 0003
> AFAICT, this does what you claim. I don't really know what else to
> look when reviewing it, if I'm being honest. As such, I diligently fed
> it through AI which suggested you may have lost a
>         VALGRIND_MAKE_MEM_NOACCESS(BufHdrGetBlock(buf), BLCKSZ);
> which sounds right to me and like something you should fix.

Good catch Melai.


> Also, I'd say this comment
> +   /*
> +    * Now okay to allow cancel/die interrupts again, were held when the lock
> +    * was acquired.
> +    */
> 
> needs a "which" after the comma to read smoothly.

Fixed.


Running it through valgrind and then will work on reading through one more
time and pushing them.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-25 21:34                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2026-03-25 22:35                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-03-27 20:00                                                         ` Andres Freund <andres@anarazel.de>
  2026-03-31 16:02                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Yura Sokolov <y.sokolov@postgrespro.ru>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-03-27 20:00 UTC (permalink / raw)
  To: Melanie Plageman <melanieplageman@gmail.com>; +Cc: Noah Misch <noah@leadboat.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-03-25 18:35:55 -0400, Andres Freund wrote:
> Running it through valgrind and then will work on reading through one more
> time and pushing them.

And done.

Phew, this project took way longer than I'd though it'd take.

Greetings,

Andres





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-25 21:34                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2026-03-25 22:35                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-27 20:00                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-03-31 16:02                                                           ` Yura Sokolov <y.sokolov@postgrespro.ru>
  2026-03-31 22:05                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Yura Sokolov @ 2026-03-31 16:02 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; Melanie Plageman <melanieplageman@gmail.com>; +Cc: Noah Misch <noah@leadboat.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

27.03.2026 23:00, Andres Freund wrote:
> Hi,
> 
> On 2026-03-25 18:35:55 -0400, Andres Freund wrote:
>> Running it through valgrind and then will work on reading through one more
>> time and pushing them.
> 
> And done.
> 
> Phew, this project took way longer than I'd though it'd take.

In addition to bug with BM_IO_ERROR [1] , I found race condition in
PinBuffer in this lines of code:

	if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
		return false;

	/*
	 * We're not allowed to increase the refcount while the buffer
	 * header spinlock is held. Wait for the lock to be released.
	 */
	if (old_buf_state & BM_LOCKED)
		old_buf_state = WaitBufHdrUnlocked(buf);

While we waited for buffer header for being unlocked, it may become
invalid, isn't it?
Therefore, check related to skip_if_not_valid have to happen after waiting.

....

Another question: previously we had to wait for buffer for being unlocked
because UnlockBufHdr wrote to buf->state unconditionally, therefore our pin
increment could be lost.
Now UnlockBufHdr and UnlockBufHdrExt does proper atomic operations and
preserves concurrent changes. Are we still need to wait?

Most of time PinBuffer is called protected by BufTable's partition LWLock,
therefore buffer may not be changed in dramatic way.

But call in ReadRecentBuffer is the exception. It is not protected by
partition lock and have to make additional checks. That is why you
introduced skip_if_not_valid.

Does optimization of ReadRecentBuffer pays for WaitBufHdrUnlocked?

[1]
https://www.postgresql.org/message-id/57a8ea70-32a8-4596-bb68-5d2990b83380%40postgrespro.ru

-- 
regards
Yura Sokolov aka funny-falcon





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-25 21:34                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2026-03-25 22:35                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-27 20:00                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-31 16:02                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Yura Sokolov <y.sokolov@postgrespro.ru>
@ 2026-03-31 22:05                                                             ` Andres Freund <andres@anarazel.de>
  2026-04-01 00:29                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-03-31 22:05 UTC (permalink / raw)
  To: Yura Sokolov <y.sokolov@postgrespro.ru>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Noah Misch <noah@leadboat.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-03-31 19:02:33 +0300, Yura Sokolov wrote:
> 27.03.2026 23:00, Andres Freund wrote:
> > On 2026-03-25 18:35:55 -0400, Andres Freund wrote:
> >> Running it through valgrind and then will work on reading through one more
> >> time and pushing them.
> > 
> > And done.
> > 
> > Phew, this project took way longer than I'd though it'd take.
> 
> In addition to bug with BM_IO_ERROR [1] , I found race condition in
> PinBuffer in this lines of code:
> 
> 	if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
> 		return false;
> 
> 	/*
> 	 * We're not allowed to increase the refcount while the buffer
> 	 * header spinlock is held. Wait for the lock to be released.
> 	 */
> 	if (old_buf_state & BM_LOCKED)
> 		old_buf_state = WaitBufHdrUnlocked(buf);
> 
> While we waited for buffer header for being unlocked, it may become
> invalid, isn't it?
> Therefore, check related to skip_if_not_valid have to happen after waiting.

Yea, that does seem wrong.  Not sure how it ended up that way.

I think it may be better to add a continue after the WaitBufHdrUnlocked(), so
that we restart the loop, rather than moving the skip_if_not_valid check.


> ....
> 
> Another question: previously we had to wait for buffer for being unlocked
> because UnlockBufHdr wrote to buf->state unconditionally, therefore our pin
> increment could be lost.
> Now UnlockBufHdr and UnlockBufHdrExt does proper atomic operations and
> preserves concurrent changes. Are we still need to wait?

Yes.


> Most of time PinBuffer is called protected by BufTable's partition LWLock,
> therefore buffer may not be changed in dramatic way.

I don't think the partition locks are sufficient protection for everything. We
have a few places in the code that want to be able to modify the buffer state
depending on whether the buffer is already pinned, and I don't think all of
them currently hold the relevant buffer mapping partition's lock.  If pinning
were not to wait for an existing header lock, such checks would not easily be
doable.

Perhaps we could fix all the relevant places by acquiring the partition lock
in a few more places. But I think that'd be going in exactly the opposite
direction we should to. The partition locks are quite contended locks and we
should work on getting rid of them eventually. Building them into the
protection model seems quite unwise.


I think many of the places that currently do rely on the buffer header
spinlock can be converted to CAS loops.

I'm not really sure how much that's worth though - the WaitBufHdrLocked() in
PinBuffer() is pretty hard to hit in realistic workloads. What would be really
nice, is to be able to replace the CAS() with an atomic add (since those are
considerably faster), but that's not really possible regardless of the need
for WaitBufHdrLocked(), because we can't just add BUF_USAGECOUNT_ONE, as that
would allow increasing the usage count too far.

I would like to eventually narrow the definition of the buffer header spinlock
to just be about the "identity" of the buffer, which then would only be needed
by things like DropRelationBuffers() and when changing the buffer's identity.


> But call in ReadRecentBuffer is the exception. It is not protected by
> partition lock and have to make additional checks. That is why you
> introduced skip_if_not_valid.
> 
> Does optimization of ReadRecentBuffer pays for WaitBufHdrUnlocked?

As mentioned above, I don't think it's just ReadRecentBuffer that relies on
the buffer header spinlock preventing new pins.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-25 21:34                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2026-03-25 22:35                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-27 20:00                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-31 16:02                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Yura Sokolov <y.sokolov@postgrespro.ru>
  2026-03-31 22:05                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-04-01 00:29                                                               ` Andres Freund <andres@anarazel.de>
  2026-04-03 10:06                                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) =?utf-8?B?Y2NhNTUwNw==?= <cca5507@qq.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-04-01 00:29 UTC (permalink / raw)
  To: Yura Sokolov <y.sokolov@postgrespro.ru>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Noah Misch <noah@leadboat.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-03-31 18:05:46 -0400, Andres Freund wrote:
> On 2026-03-31 19:02:33 +0300, Yura Sokolov wrote:
> > 27.03.2026 23:00, Andres Freund wrote:
> > > On 2026-03-25 18:35:55 -0400, Andres Freund wrote:
> > >> Running it through valgrind and then will work on reading through one more
> > >> time and pushing them.
> > >
> > > And done.
> > >
> > > Phew, this project took way longer than I'd though it'd take.
> >
> > In addition to bug with BM_IO_ERROR [1] , I found race condition in
> > PinBuffer in this lines of code:
> >
> > 	if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID)))
> > 		return false;
> >
> > 	/*
> > 	 * We're not allowed to increase the refcount while the buffer
> > 	 * header spinlock is held. Wait for the lock to be released.
> > 	 */
> > 	if (old_buf_state & BM_LOCKED)
> > 		old_buf_state = WaitBufHdrUnlocked(buf);
> >
> > While we waited for buffer header for being unlocked, it may become
> > invalid, isn't it?
> > Therefore, check related to skip_if_not_valid have to happen after waiting.
>
> Yea, that does seem wrong.  Not sure how it ended up that way.
>
> I think it may be better to add a continue after the WaitBufHdrUnlocked(), so
> that we restart the loop, rather than moving the skip_if_not_valid check.

Done that way. Thanks for finding & reporting this, well spotted!

Greetings,

Andres





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 10:44                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-15 19:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Noah Misch <noah@leadboat.com>
  2026-03-11 22:40                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-25 21:34                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2026-03-25 22:35                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-27 20:00                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-31 16:02                                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Yura Sokolov <y.sokolov@postgrespro.ru>
  2026-03-31 22:05                                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-04-01 00:29                                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-04-03 10:06                                                                 ` =?utf-8?B?Y2NhNTUwNw==?= <cca5507@qq.com>
  0 siblings, 0 replies; 120+ messages in thread

From: cca5507 @ 2026-04-03 10:06 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; pgsql-hackers <pgsql-hackers@lists.postgresql.org>

Hi,

I find some outdated comments in src/backend/storage/buffer/README:

```
Note that a buffer header's spinlock does not control access to the data
held within the buffer.  Each buffer header also contains an LWLock, the
"buffer content lock", that *does* represent the right to access the data
in the buffer.  It is used per the rules above.
```

"Each buffer header also contains an LWLock" is outdated.

```
The background writer takes shared content lock on a buffer while writing it
out (and anyone else who flushes buffer contents to disk must do so too).
This ensures that the page image transferred to disk is reasonably consistent.
We might miss a hint-bit update or two but that isn't a problem, for the same
reasons mentioned under buffer access rules.
```

"The background writer takes shared content lock ...", should be "share-exclusive content lock".

"We might miss a hint-bit update or two ...", maybe already fixed by share-exclusive content lock?

--
Regards,
ChangAo Chen


^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-02-07 12:38                                               ` Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-08 18:38                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  3 siblings, 1 reply; 120+ messages in thread

From: Heikki Linnakangas @ 2026-02-07 12:38 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; Melanie Plageman <melanieplageman@gmail.com>; Noah Misch <noah@leadboat.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On 03/02/2026 00:33, Andres Freund wrote:
>    - Now that we use the normal order of WAL logging, we don't need to delay
>      checkpoint starts anymore.
> 
>      I think the explanation for why that is ok is correct [1], but it needs to
>      be looked at by somebody with experience around this. Maybe Heikki?

So that's patch 0004 "bufmgr: Switch to standard order in 
MarkBufferDirtyHint()". Yes, looks correct to me.

> 	/*
> 	 * Update RedoRecPtr so that we can make the right decision. It's possible
> 	 * that a new checkpoint will start just after GetRedoRecPtr(), but that
> 	 * is ok, as the buffer is already dirty, ensuring that any BufferSync()
> 	 * started after the buffer was marked dirty cannot complete without
> 	 * flushing this buffer.  If a checkpoint started between marking the
> 	 * buffer dirty and this check, we will emit an unnecessary WAL record (as
> 	 * the buffer will be written out as part of the checkpoint), but the
> 	 * window for that is small.
> 	 */
> 	RedoRecPtr = GetRedoRecPtr();

That "small window" is actually pretty big if you think of it a little 
more loosely. Our rule is that we write the full page image if a 
checkpoint has started since the page LSN, but that's very conservative 
already. It would be sufficient to write the full page image only if the 
checkpoint has already flushed the page. This small window is just a 
special case of that conservatism.

I've been thinking of trying track that more accurately for a long time, 
because it would smoothen the WAL spike when a checkpoint begins.

That gets off-topic, but my point is that it feels a little silly to 
mention that small window when there's the other giant panoramic window 
next to it.

- Heikki






^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 12:38                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
@ 2026-02-08 18:38                                                 ` Andres Freund <andres@anarazel.de>
  2026-02-09 19:54                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-02-08 18:38 UTC (permalink / raw)
  To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Noah Misch <noah@leadboat.com>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-02-07 14:38:53 +0200, Heikki Linnakangas wrote:
> On 03/02/2026 00:33, Andres Freund wrote:
> >    - Now that we use the normal order of WAL logging, we don't need to delay
> >      checkpoint starts anymore.
> >
> >      I think the explanation for why that is ok is correct [1], but it needs to
> >      be looked at by somebody with experience around this. Maybe Heikki?
>
> So that's patch 0004 "bufmgr: Switch to standard order in
> MarkBufferDirtyHint()". Yes, looks correct to me.

Thanks for checking!  Somehow I went back and forth about it being right
multiple times...


> > 	/*
> > 	 * Update RedoRecPtr so that we can make the right decision. It's possible
> > 	 * that a new checkpoint will start just after GetRedoRecPtr(), but that
> > 	 * is ok, as the buffer is already dirty, ensuring that any BufferSync()
> > 	 * started after the buffer was marked dirty cannot complete without
> > 	 * flushing this buffer.  If a checkpoint started between marking the
> > 	 * buffer dirty and this check, we will emit an unnecessary WAL record (as
> > 	 * the buffer will be written out as part of the checkpoint), but the
> > 	 * window for that is small.
> > 	 */
> > 	RedoRecPtr = GetRedoRecPtr();
>
> That "small window" is actually pretty big if you think of it a little more
> loosely. Our rule is that we write the full page image if a checkpoint has
> started since the page LSN, but that's very conservative already. It would
> be sufficient to write the full page image only if the checkpoint has
> already flushed the page. This small window is just a special case of that
> conservatism.

I mainly want to mention that window because I have to think about it when
analyzing the correctness of the approach. If the window is not mentioned, at
least I have to think about whether the window is dangerous in some form.


> It would be sufficient to write the full page image only if the checkpoint
> has already flushed the page.

Today that would probably not quite be sufficient, due to issues around
re-dirtying the page during checkpointer's flush (and thus needing to be
written out again, with the chance of a torn write that has no FPI to repair
it). But that will soon be impossible.


I think the actual rule would need to be more complicated, I think we would
need to generate an FPI for the first modification after the checkpoint flush,
even though the LSN is newer than the redo LSN, because we didn't generate one
earlier?  Otherwise we could get into a situation where there is no non-torn
on-disk page version after a later crash, I think?

Consider:

1) modify page w/ FPI
2) redo pointer determined at X
3) modify page w/o FPI, as the page hasn't yet been flushed at X+1
4) checkpointer flushes page
5) checkpoint completes, at X+2
6) page is dirtied, w/o FPI X+3, as X+1 > X
7) in the middle of writing out the page, we crash, the page is torn

For recovery we will replay starting from position X. Then will replay the
record from 3), which will be skipped due to the LSN. Then we will replay X+3,
which either will be skipped due to the LSN condition (if the page header
survived the torn page), leading to the changes to the "old portion" of the
torn page not being replayed, or we will replay the WAL record, applying it to
a torn page (or failing to read in the page due to checksum errors).

If we only needed to think about buffers that stay in memory, we could "just"
tackle this by remember that the page will need to be FPId during the next
modification in the BufferDesc, but that doesn't help us if the page is
evicted and reread...



> I've been thinking of trying track that more accurately for a long time,
> because it would smoothen the WAL spike when a checkpoint begins.

It'd indeed be nice to improve that. Another thing it'd be helpful is widening
when we can write out hint bits on standbys.

If the rule were just that we can skip an FPI if the page still needs to be
written out by the checkpoint, it'd be fairly simple - we could utilize
BM_CHECKPOINT_NEEDED. But as hinted at above, I think it's a it more
complicated.


Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 12:38                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-08 18:38                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-02-09 19:54                                                   ` Kirill Reshke <reshkekirill@gmail.com>
  0 siblings, 0 replies; 120+ messages in thread

From: Kirill Reshke @ 2026-02-09 19:54 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Noah Misch <noah@leadboat.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On Sun, 8 Feb 2026 at 23:38, Andres Freund <andres@anarazel.de> wrote:
>

> Consider:
>
> 1) modify page w/ FPI
> 2) redo pointer determined at X
> 3) modify page w/o FPI, as the page hasn't yet been flushed at X+1
> 4) checkpointer flushes page
> 5) checkpoint completes, at X+2
> 6) page is dirtied, w/o FPI X+3, as X+1 > X
> 7) in the middle of writing out the page, we crash, the page is torn
>
> For recovery we will replay starting from position X. Then will replay the
> record from 3), which will be skipped due to the LSN. Then we will replay X+3,
> which either will be skipped due to the LSN condition (if the page header
> survived the torn page), leading to the changes to the "old portion" of the
> torn page not being replayed, or we will replay the WAL record, applying it to
> a torn page (or failing to read in the page due to checksum errors).
>
> If we only needed to think about buffers that stay in memory, we could "just"
> tackle this by remember that the page will need to be FPId during the next
> modification in the BufferDesc, but that doesn't help us if the page is
> evicted and reread...
>
>

Hmm, after thinking about this, I wonder if we can actually have a TAP
test for this sequence of events?
Maybe it would be desirable to execute some rare recovery code path.
But I'm unsure if there is any reliable way to have an OS to have a
buffer in page cache, but not on disk when evicted.


-- 
Best regards,
Kirill Reshke





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-02-07 12:59                                               ` Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-09 01:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  3 siblings, 1 reply; 120+ messages in thread

From: Heikki Linnakangas @ 2026-02-07 12:59 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; Melanie Plageman <melanieplageman@gmail.com>; Noah Misch <noah@leadboat.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

A few minor nitpicks on v12 below. Other than these and the comments I 
wrote in separate emails, looks good to me.

> @@ -371,8 +382,6 @@ _bt_killitems(IndexScanDesc scan)
>         }
>  
>         /*
> -        * Since this can be redone later if needed, mark as dirty hint.
> -        *
>          * Whenever we mark anything LP_DEAD, we also set the page's
>          * BTP_HAS_GARBAGE flag, which is likewise just a hint.  (Note that we
>          * only rely on the page-level flag in !heapkeyspace indexes.)

Seems a bit random to remove that.

> +/*
> + * Try to set a single hint bit in a buffer.
> + *
> + * This is a bit faster than BufferBeginSetHintBits() /
> + * BufferFinishSetHintBits() when setting a single hint bit, but slower than
> + * the former when setting several hint bits.
> + */
> +bool
> +BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)

This could use some more explanation. The point is that this does "*ptr 
= val", if it's allowed to set hint bits. That's not obvious. And 
"single hint bit" isn't really accurate, as you could update multiple 
bits in *ptr with one call.

> 	/*
> 	 * If the buffer was dirty, try to write it out.  There is a race
> 	 * condition here, in that someone might dirty it after we released the
> 	 * buffer header lock above.  We will recheck the dirty bit after
> 	 * re-locking the buffer header.
> 	 */

It's not clear what "above" means in that paragraph. Where do we release 
the buffer header lock? In StrategyGetBuffer?

(This is not actually new with this patch; it goes back to commit 
5e89985928. Before that, there was a call to PinBuffer_Locked() which 
released the spinlock.)

> @@ -2516,18 +2515,21 @@ again:
>                 /*
>                  * If using a nondefault strategy, and writing the buffer would
>                  * require a WAL flush, let the strategy decide whether to go ahead
> -                * and write/reuse the buffer or to choose another victim.  We need a
> -                * lock to inspect the page LSN, so this can't be done inside
> +                * and write/reuse the buffer or to choose another victim.  We need to
> +                * hold the content lock in at least share-exclusive mode to safely
> +                * inspect the page LSN, so this couldn't have been done inside
>                  * StrategyGetBuffer.
>                  */
>                 if (strategy != NULL)
>                 {
>                         XLogRecPtr      lsn;
>  
> -                       /* Read the LSN while holding buffer header lock */
> -                       buf_state = LockBufHdr(buf_hdr);
> +                       /*
> +                        * As we now hold at least a share-exclusive lock on the buffer,
> +                        * the LSN cannot change during the flush (and thus can't be
> +                        * torn).
> +                        */
>                         lsn = BufferGetLSN(buf_hdr);
> -                       UnlockBufHdr(buf_hdr);
>  
>                         if (XLogNeedsFlush(lsn)
>                                 && StrategyRejectBuffer(strategy, buf_hdr, from_ring))

I think the second comment is redundant with the first one. Let's just 
remove it.

> +/*
> + * Helper for BufferBeginSetHintBits() and BufferSetHintBits16().
> + *
> + * This checks if the current lock mode already suffices to allow hint bits
> + * being set and, if not, whether the current lock can be upgraded.
> + *
> + * Updates *lockstate when returning true.
> + */
> +static inline bool
> +SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)

Would be good to be more explicit what returning true/false here means.

- Heikki






^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 12:59                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
@ 2026-02-09 01:52                                                 ` Andres Freund <andres@anarazel.de>
  2026-02-09 10:14                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-02-09 01:52 UTC (permalink / raw)
  To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Noah Misch <noah@leadboat.com>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-02-07 14:59:34 +0200, Heikki Linnakangas wrote:
> A few minor nitpicks on v12 below. Other than these and the comments I wrote
> in separate emails, looks good to me.
> 
> > @@ -371,8 +382,6 @@ _bt_killitems(IndexScanDesc scan)
> >         }
> >         /*
> > -        * Since this can be redone later if needed, mark as dirty hint.
> > -        *
> >          * Whenever we mark anything LP_DEAD, we also set the page's
> >          * BTP_HAS_GARBAGE flag, which is likewise just a hint.  (Note that we
> >          * only rely on the page-level flag in !heapkeyspace indexes.)
> 
> Seems a bit random to remove that.

Fair.


> > +/*
> > + * Try to set a single hint bit in a buffer.
> > + *
> > + * This is a bit faster than BufferBeginSetHintBits() /
> > + * BufferFinishSetHintBits() when setting a single hint bit, but slower than
> > + * the former when setting several hint bits.
> > + */
> > +bool
> > +BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)
> 
> This could use some more explanation. The point is that this does "*ptr =
> val", if it's allowed to set hint bits. That's not obvious. And "single hint
> bit" isn't really accurate, as you could update multiple bits in *ptr with
> one call.

Agreed.  I updated it to

 * Try to set hint bits on a single 16bit value in a buffer.
 *
 * If hint bits are allowed to be set, set *ptr = val, try mark the buffer
 * dirty and return true. Otherwise false is returned.
 *
 * *ptr needs to be a pointer to memory within the buffer.
 *
 * This is a bit faster than BufferBeginSetHintBits() /
 * BufferFinishSetHintBits() when setting hints once in a buffer, but slower
 * than the former when setting hint bits multiple times in the same buffer.



> > 	/*
> > 	 * If the buffer was dirty, try to write it out.  There is a race
> > 	 * condition here, in that someone might dirty it after we released the
> > 	 * buffer header lock above.  We will recheck the dirty bit after
> > 	 * re-locking the buffer header.
> > 	 */
> 
> It's not clear what "above" means in that paragraph. Where do we release the
> buffer header lock? In StrategyGetBuffer?
>
> (This is not actually new with this patch; it goes back to commit
> 5e89985928. Before that, there was a call to PinBuffer_Locked() which
> released the spinlock.)

Yea, looks like I should have edited the comment in that commit. Updated to:

	 * If the buffer was dirty, try to write it out.  There is a race
	 * condition here, another backend could dirty the buffer between
	 * StrategyGetBuffer() checking that it is not in use and invalidating the
	 * buffer below. That's addressed by InvalidateVictimBuffer() verifying
	 * that the buffer is not dirty.


> > @@ -2516,18 +2515,21 @@ again:
> >                 /*
> >                  * If using a nondefault strategy, and writing the buffer would
> >                  * require a WAL flush, let the strategy decide whether to go ahead
> > -                * and write/reuse the buffer or to choose another victim.  We need a
> > -                * lock to inspect the page LSN, so this can't be done inside
> > +                * and write/reuse the buffer or to choose another victim.  We need to
> > +                * hold the content lock in at least share-exclusive mode to safely
> > +                * inspect the page LSN, so this couldn't have been done inside
> >                  * StrategyGetBuffer.
> >                  */
> >                 if (strategy != NULL)
> >                 {
> >                         XLogRecPtr      lsn;
> > -                       /* Read the LSN while holding buffer header lock */
> > -                       buf_state = LockBufHdr(buf_hdr);
> > +                       /*
> > +                        * As we now hold at least a share-exclusive lock on the buffer,
> > +                        * the LSN cannot change during the flush (and thus can't be
> > +                        * torn).
> > +                        */
> >                         lsn = BufferGetLSN(buf_hdr);
> > -                       UnlockBufHdr(buf_hdr);
> >                         if (XLogNeedsFlush(lsn)
> >                                 && StrategyRejectBuffer(strategy, buf_hdr, from_ring))
> 
> I think the second comment is redundant with the first one. Let's just
> remove it.

Agreed.


> > +/*
> > + * Helper for BufferBeginSetHintBits() and BufferSetHintBits16().
> > + *
> > + * This checks if the current lock mode already suffices to allow hint bits
> > + * being set and, if not, whether the current lock can be upgraded.
> > + *
> > + * Updates *lockstate when returning true.
> > + */
> > +static inline bool
> > +SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
> 
> Would be good to be more explicit what returning true/false here means.

Hm, ISTM that'd just end up restating the comments for
BufferBeginSetHintBits(). Given SharedBufferBeginSetHintBits() is just an
implementation detail for BufferBeginSetHintBits/BufferSetHintBits16, I don't
think it's worth restating the details here.

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 12:59                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-09 01:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-02-09 10:14                                                   ` Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-09 22:19                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Heikki Linnakangas @ 2026-02-09 10:14 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Noah Misch <noah@leadboat.com>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

On 09/02/2026 03:52, Andres Freund wrote:
> On 2026-02-07 14:59:34 +0200, Heikki Linnakangas wrote:
>>> +/*
>>> + * Try to set a single hint bit in a buffer.
>>> + *
>>> + * This is a bit faster than BufferBeginSetHintBits() /
>>> + * BufferFinishSetHintBits() when setting a single hint bit, but slower than
>>> + * the former when setting several hint bits.
>>> + */
>>> +bool
>>> +BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)
>>
>> This could use some more explanation. The point is that this does "*ptr =
>> val", if it's allowed to set hint bits. That's not obvious. And "single hint
>> bit" isn't really accurate, as you could update multiple bits in *ptr with
>> one call.
> 
> Agreed.  I updated it to
> 
>   * Try to set hint bits on a single 16bit value in a buffer.
>   *
>   * If hint bits are allowed to be set, set *ptr = val, try mark the buffer
>   * dirty and return true. Otherwise false is returned.
>   *
>   * *ptr needs to be a pointer to memory within the buffer.
>   *
>   * This is a bit faster than BufferBeginSetHintBits() /
>   * BufferFinishSetHintBits() when setting hints once in a buffer, but slower
>   * than the former when setting hint bits multiple times in the same buffer.

+1. Instead of "try mark the buffer dirty", I'd say just "mark the 
buffer dirty". The only reason it might not to mark the buffer dirty is 
that it was already marked dirty, right? I wouldn't call that a failure.

>>> 	/*
>>> 	 * If the buffer was dirty, try to write it out.  There is a race
>>> 	 * condition here, in that someone might dirty it after we released the
>>> 	 * buffer header lock above.  We will recheck the dirty bit after
>>> 	 * re-locking the buffer header.
>>> 	 */
>>
>> It's not clear what "above" means in that paragraph. Where do we release the
>> buffer header lock? In StrategyGetBuffer?
>>
>> (This is not actually new with this patch; it goes back to commit
>> 5e89985928. Before that, there was a call to PinBuffer_Locked() which
>> released the spinlock.)
> 
> Yea, looks like I should have edited the comment in that commit. Updated to:
> 
> 	 * If the buffer was dirty, try to write it out.  There is a race
> 	 * condition here, another backend could dirty the buffer between
> 	 * StrategyGetBuffer() checking that it is not in use and invalidating the
> 	 * buffer below. That's addressed by InvalidateVictimBuffer() verifying
> 	 * that the buffer is not dirty.

+1

>>> +/*
>>> + * Helper for BufferBeginSetHintBits() and BufferSetHintBits16().
>>> + *
>>> + * This checks if the current lock mode already suffices to allow hint bits
>>> + * being set and, if not, whether the current lock can be upgraded.
>>> + *
>>> + * Updates *lockstate when returning true.
>>> + */
>>> +static inline bool
>>> +SharedBufferBeginSetHintBits(Buffer buffer, BufferDesc *buf_hdr, uint64 *lockstate)
>>
>> Would be good to be more explicit what returning true/false here means.
> 
> Hm, ISTM that'd just end up restating the comments for
> BufferBeginSetHintBits(). Given SharedBufferBeginSetHintBits() is just an
> implementation detail for BufferBeginSetHintBits/BufferSetHintBits16, I don't
> think it's worth restating the details here.

Ok fair.

- Heikki






^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-14 21:20                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-02 22:33                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-07 12:59                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2026-02-09 01:52                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-09 10:14                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
@ 2026-02-09 22:19                                                     ` Andres Freund <andres@anarazel.de>
  0 siblings, 0 replies; 120+ messages in thread

From: Andres Freund @ 2026-02-09 22:19 UTC (permalink / raw)
  To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Melanie Plageman <melanieplageman@gmail.com>; Noah Misch <noah@leadboat.com>; Kirill Reshke <reshkekirill@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-02-09 12:14:47 +0200, Heikki Linnakangas wrote:
> On 09/02/2026 03:52, Andres Freund wrote:
> > On 2026-02-07 14:59:34 +0200, Heikki Linnakangas wrote:
> > > > +/*
> > > > + * Try to set a single hint bit in a buffer.
> > > > + *
> > > > + * This is a bit faster than BufferBeginSetHintBits() /
> > > > + * BufferFinishSetHintBits() when setting a single hint bit, but slower than
> > > > + * the former when setting several hint bits.
> > > > + */
> > > > +bool
> > > > +BufferSetHintBits16(uint16 *ptr, uint16 val, Buffer buffer)
> > > 
> > > This could use some more explanation. The point is that this does "*ptr =
> > > val", if it's allowed to set hint bits. That's not obvious. And "single hint
> > > bit" isn't really accurate, as you could update multiple bits in *ptr with
> > > one call.
> > 
> > Agreed.  I updated it to
> > 
> >   * Try to set hint bits on a single 16bit value in a buffer.
> >   *
> >   * If hint bits are allowed to be set, set *ptr = val, try mark the buffer
> >   * dirty and return true. Otherwise false is returned.
> >   *
> >   * *ptr needs to be a pointer to memory within the buffer.
> >   *
> >   * This is a bit faster than BufferBeginSetHintBits() /
> >   * BufferFinishSetHintBits() when setting hints once in a buffer, but slower
> >   * than the former when setting hint bits multiple times in the same buffer.
> 
> +1. Instead of "try mark the buffer dirty", I'd say just "mark the buffer
> dirty". The only reason it might not to mark the buffer dirty is that it was
> already marked dirty, right? I wouldn't call that a failure.

It's not quite the only reason:

		/*
		 * If we need to protect hint bit updates from torn writes, WAL-log a
		 * full page image of the page. This full page image is only necessary
		 * if the hint bit update is the first change to the page since the
		 * last checkpoint.
		 *
		 * We don't check full_page_writes here because that logic is included
		 * when we call XLogInsert() since the value changes dynamically.
		 */
		if (XLogHintBitIsNeeded() && (lockstate & BM_PERMANENT))
		{
			/*
			 * If we must not write WAL, due to a relfilelocator-specific
			 * condition or being in recovery, don't dirty the page.  We can
			 * set the hint, just not dirty the page as a result so the hint
			 * is lost when we evict the page or shutdown.
			 *
			 * See src/backend/storage/page/README for longer discussion.
			 */
			if (RecoveryInProgress() ||
				RelFileLocatorSkippingWAL(BufTagGetRelFileLocator(&bufHdr->tag)))
				return;

			wal_log = true;
		}

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-02-09 11:42                                           ` Antonin Houska <ah@cybertec.at>
  2026-02-09 22:16                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  4 siblings, 1 reply; 120+ messages in thread

From: Antonin Houska @ 2026-02-09 11:42 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Andres Freund <andres@anarazel.de> wrote:

> On 2026-01-12 12:45:03 -0500, Andres Freund wrote:
> > I'm doing another pass through 0003 and will push that if I don't find
> > anything significant.
> 
> Done, after adjust two comments in minor ways.

I suppose this is commit 0b96e734c590.

While troubleshooting REPACK issue [1], I realized that
HeapTupleSatisfiesMVCCBatch() can also be called during logical decoding - in
that case we need to use a historic MVCC snapshot. My proposal to fix the
problem is attached.

[1] https://www.postgresql.org/message-id/CADzfLwWNv5QDn6qmxCRV-p_ijSTGwNcEZFCOXt09+RmpSG2=+w@mail.gmail...

-- 
Antonin Houska
Web: https://www.cybertec-postgresql.com

Attachments:

  [text/x-diff] fix_batch_visibility_checks.diff (564B, ../../61812.1770637345@localhost/2-fix_batch_visibility_checks.diff)
  download | inline diff:
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 75ae268d753..685a938bd68 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -1628,7 +1649,7 @@ HeapTupleSatisfiesMVCCBatch(Snapshot snapshot, Buffer buffer,
 		bool		valid;
 		HeapTuple	tup = &batchmvcc->tuples[i];
 
-		valid = HeapTupleSatisfiesMVCC(tup, snapshot, buffer);
+		valid = HeapTupleSatisfiesVisibility(tup, snapshot, buffer);
 		batchmvcc->visible[i] = valid;
 
 		if (likely(valid))

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-09 11:42                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
@ 2026-02-09 22:16                                             ` Andres Freund <andres@anarazel.de>
  2026-02-10 07:46                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-02-09 22:16 UTC (permalink / raw)
  To: Antonin Houska <ah@cybertec.at>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-02-09 12:42:25 +0100, Antonin Houska wrote:
> Andres Freund <andres@anarazel.de> wrote:
> 
> > On 2026-01-12 12:45:03 -0500, Andres Freund wrote:
> > > I'm doing another pass through 0003 and will push that if I don't find
> > > anything significant.
> > 
> > Done, after adjust two comments in minor ways.
> 
> I suppose this is commit 0b96e734c590.
> 
> While troubleshooting REPACK issue [1], I realized that
> HeapTupleSatisfiesMVCCBatch() can also be called during logical decoding - in
> that case we need to use a historic MVCC snapshot.

Huh. Indeed. That's unintentional - the path should never have been reached,
we are checking that an MVCC snapshot is used. Unfortunately, somebody
(i.e. probably me) at some point defined the relevant macro as

/* This macro encodes the knowledge of which snapshots are MVCC-safe */
#define IsMVCCSnapshot(snapshot)  \
	((snapshot)->snapshot_type == SNAPSHOT_MVCC || \
	 (snapshot)->snapshot_type == SNAPSHOT_HISTORIC_MVCC)

Which makes sense for some places, but not for plenty others.

The reason this didn't cause more widespread issues is that during logical
decoding we mostly don't use sequential scans etc that are affected by the
these paths.


> My proposal to fix the problem is attached.

That's imo not at all the right fix - it'd make visibility during seqscans
checking noticeably slower.


I think we ought to instead restrict the page-at-a-time scans to only happen
with "real" mvcc snapshots. I.e. this:

	/*
	 * Disable page-at-a-time mode if it's not a MVCC-safe snapshot.
	 */
	if (!(snapshot && IsMVCCSnapshot(snapshot)))
		scan->rs_base.rs_flags &= ~SO_ALLOW_PAGEMODE;

should trigger for historic snapshots as well.


Does that fix the issue for you?


What's your reproducer?


Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-09 11:42                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-02-09 22:16                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-02-10 07:46                                               ` Antonin Houska <ah@cybertec.at>
  2026-02-10 16:49                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Antonin Houska @ 2026-02-10 07:46 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Andres Freund <andres@anarazel.de> wrote:

> > While troubleshooting REPACK issue [1], I realized that
> > HeapTupleSatisfiesMVCCBatch() can also be called during logical decoding - in
> > that case we need to use a historic MVCC snapshot.
> 
> Huh. Indeed. That's unintentional - the path should never have been reached,
> we are checking that an MVCC snapshot is used. Unfortunately, somebody
> (i.e. probably me) at some point defined the relevant macro as
> 
> /* This macro encodes the knowledge of which snapshots are MVCC-safe */
> #define IsMVCCSnapshot(snapshot)  \
> 	((snapshot)->snapshot_type == SNAPSHOT_MVCC || \
> 	 (snapshot)->snapshot_type == SNAPSHOT_HISTORIC_MVCC)
> 
> Which makes sense for some places, but not for plenty others.
> 
> The reason this didn't cause more widespread issues is that during logical
> decoding we mostly don't use sequential scans etc that are affected by the
> these paths.

> > My proposal to fix the problem is attached.
> 
> That's imo not at all the right fix - it'd make visibility during seqscans
> checking noticeably slower.

ok

> I think we ought to instead restrict the page-at-a-time scans to only happen
> with "real" mvcc snapshots. I.e. this:
> 
> 	/*
> 	 * Disable page-at-a-time mode if it's not a MVCC-safe snapshot.
> 	 */
> 	if (!(snapshot && IsMVCCSnapshot(snapshot)))
> 		scan->rs_base.rs_flags &= ~SO_ALLOW_PAGEMODE;
>
> should trigger for historic snapshots as well.

I suppose you mean changing it to

	if (!(snapshot && IsMVCCSnapshot(snapshot) &&
		  !IsHistoricMVCCSnapshot(snapshot)))
		scan->rs_base.rs_flags &= ~SO_ALLOW_PAGEMODE;

> Does that fix the issue for you?

Yes, with this change, I don't hit the problem anymore.

> What's your reproducer?

Check out this branch

https://github.com/michail-nikolaev/postgres/tree/repack_concurrently_repro_22

and run t/008_repack_concurrently.pl in contrib/amcheck. The error we saw in
most cases was "ERROR: cache lookup failed for relation". I noticed that the
related pg_class entries had hint bits set incorrectly, so I added the
following to see when exactly it happens (actually I used lower elevel first,
to find out that the decoding worker is responsible for the problem):

diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 75ae268d753..ebf38460873 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -73,6 +73,7 @@
 #include "access/transam.h"
 #include "access/xact.h"
 #include "access/xlog.h"
+#include "commands/cluster.h"
 #include "storage/bufmgr.h"
 #include "storage/procarray.h"
 #include "utils/builtins.h"
@@ -938,6 +939,8 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
                                                HeapTupleHeaderGetRawXmin(tuple));
                else
                {
+                       if (am_decoding_for_repack())
+                               elog(PANIC, "HEAP_XMIN_INVALID set");
                        /* it must have aborted or crashed */
                        SetHintBits(tuple, buffer, HEAP_XMIN_INVALID,
                                                InvalidTransactionId);

The backtrace looked like:

    #3  0x0000000000c367ad errfinish (postgres + 0x8367ad)
    #4  0x0000000000517ee1 HeapTupleSatisfiesMVCC (postgres + 0x117ee1)
    #5  0x0000000000518e14 HeapTupleSatisfiesMVCCBatch (postgres + 0x118e14)
    #6  0x000000000050483e page_collect_tuples (postgres + 0x10483e)
    #7  0x0000000000504a8e heap_prepare_pagescan (postgres + 0x104a8e)
    #8  0x0000000000505344 heapgettup_pagemode (postgres + 0x105344)
    #9  0x0000000000505cff heap_getnextslot (postgres + 0x105cff)
    #10 0x000000000052ca08 table_scan_getnextslot (postgres + 0x12ca08)
    #11 0x000000000052d4ed systable_getnext (postgres + 0x12d4ed)
    #12 0x0000000000c2098e ScanPgRelation (postgres + 0x82098e)
    #13 0x0000000000c23aa2 RelationReloadIndexInfo (postgres + 0x823aa2)
    #14 0x0000000000c24376 RelationRebuildRelation (postgres + 0x824376)
    #15 0x0000000000c23784 RelationIdGetRelation (postgres + 0x823784)
    #16 0x00000000004b558f relation_open (postgres + 0xb558f)
    #17 0x000000000052e073 index_open (postgres + 0x12e073)
    #18 0x000000000052d0c5 systable_beginscan (postgres + 0x12d0c5)
    #19 0x0000000000c2097e ScanPgRelation (postgres + 0x82097e)
    #20 0x0000000000c23dde RelationReloadNailed (postgres + 0x823dde)
    #21 0x0000000000c24399 RelationRebuildRelation (postgres + 0x824399)
    #22 0x0000000000c23784 RelationIdGetRelation (postgres + 0x823784)
    #23 0x00000000004b558f relation_open (postgres + 0xb558f)
    #24 0x0000000000573ab1 table_open (postgres + 0x173ab1)
    #25 0x0000000000c20919 ScanPgRelation (postgres + 0x820919)
    #26 0x0000000000c21c98 RelationBuildDesc (postgres + 0x821c98)
    #27 0x0000000000c237d6 RelationIdGetRelation (postgres + 0x8237d6)
    #28 0x00000000009a028a ReorderBufferProcessTXN (postgres + 0x5a028a)
    #29 0x00000000009a10ed ReorderBufferReplay (postgres + 0x5a10ed)
    #30 0x00000000009a116b ReorderBufferCommit (postgres + 0x5a116b)
    #31 0x000000000098c7fe DecodeCommit (postgres + 0x58c7fe)
    #32 0x000000000098ba67 xact_decode (postgres + 0x58ba67)
    #33 0x000000000098b685 LogicalDecodingProcessRecord (postgres + 0x58b685)
    #34 0x00000000006a3e4b decode_concurrent_changes (postgres + 0x2a3e4b)
    #35 0x00000000006a5ea3 repack_worker_internal (postgres + 0x2a5ea3)
    #36 0x00000000006a5d5e RepackWorkerMain (postgres + 0x2a5d5e)
    #37 0x0000000000961b2b BackgroundWorkerMain (postgres + 0x561b2b)
    #38 0x0000000000964a61 postmaster_child_launch (postgres + 0x564a61)
    #39 0x000000000096b58d StartBackgroundWorker (postgres + 0x56b58d)
    #40 0x000000000096b80a maybe_start_bgworkers (postgres + 0x56b80a)
    #41 0x000000000096a5d9 LaunchMissingBackgroundProcesses (postgres + 0x56a5d9)
    #42 0x00000000009683fc ServerLoop (postgres + 0x5683fc)
    #43 0x0000000000967d37 PostmasterMain (postgres + 0x567d37)
    #44 0x0000000000813421 main (postgres + 0x413421)

Thanks.

-- 
Antonin Houska
Web: https://www.cybertec-postgresql.com





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-09 11:42                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-02-09 22:16                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-10 07:46                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
@ 2026-02-10 16:49                                                 ` Andres Freund <andres@anarazel.de>
  2026-02-12 10:36                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-02-10 16:49 UTC (permalink / raw)
  To: Antonin Houska <ah@cybertec.at>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-02-10 08:46:27 +0100, Antonin Houska wrote:
> Andres Freund <andres@anarazel.de> wrote:
> > I think we ought to instead restrict the page-at-a-time scans to only happen
> > with "real" mvcc snapshots. I.e. this:
> > 
> > 	/*
> > 	 * Disable page-at-a-time mode if it's not a MVCC-safe snapshot.
> > 	 */
> > 	if (!(snapshot && IsMVCCSnapshot(snapshot)))
> > 		scan->rs_base.rs_flags &= ~SO_ALLOW_PAGEMODE;
> >
> > should trigger for historic snapshots as well.
> 
> I suppose you mean changing it to
> 
> 	if (!(snapshot && IsMVCCSnapshot(snapshot) &&
> 		  !IsHistoricMVCCSnapshot(snapshot)))
> 		scan->rs_base.rs_flags &= ~SO_ALLOW_PAGEMODE;

Yes.

For something committable, I think we should probably split IsMVCCSnapshot
into IsMVCCSnapshot(), just accepting SNAPSHOT_MVCC, and IsMVCCLikeSnapshot()
accepting both SNAPSHOT_MVCC and SNAPSHOT_HISTORIC_MVCC. And then go through
all the existing callers of IsMVCCSnapshot() - only about half should stay
as-is, I think.


> > Does that fix the issue for you?
> 
> Yes, with this change, I don't hit the problem anymore.

Great!

Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-09 11:42                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-02-09 22:16                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-10 07:46                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-02-10 16:49                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-02-12 10:36                                                   ` Antonin Houska <ah@cybertec.at>
  2026-03-11 23:09                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Antonin Houska @ 2026-02-12 10:36 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Andres Freund <andres@anarazel.de> wrote:

> For something committable, I think we should probably split IsMVCCSnapshot
> into IsMVCCSnapshot(), just accepting SNAPSHOT_MVCC, and IsMVCCLikeSnapshot()
> accepting both SNAPSHOT_MVCC and SNAPSHOT_HISTORIC_MVCC. And then go through
> all the existing callers of IsMVCCSnapshot() - only about half should stay
> as-is, I think.

The attached patch tries to do that.

-- 
Antonin Houska
Web: https://www.cybertec-postgresql.com

Attachments:

  [text/x-diff] 0001-Refine-checking-of-snapshot-type.patch (3.3K, ../../196082.1770892568@localhost/2-0001-Refine-checking-of-snapshot-type.patch)
  download | inline diff:
From dcdbaf3095e632a1f7f65f3abc43eccff0249d4c Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Thu, 12 Feb 2026 11:14:00 +0100
Subject: [PATCH] Refine checking of snapshot type.

It appears to be confusing if IsMVCCSnapshot() evaluates to true for both
"regular" and "historic" MVCC snapshot. This patch restricts the meaning of
the macro to the "regular" MVCC snapshot, and introduces a new macro
IsMVCCLikeSnapshot() to recognize both types.

IsMVCCLikeSnapshot() is only used in functions that can (supposedly) be called
during logical decoding.
---
 src/backend/access/heap/heapam_handler.c | 2 +-
 src/backend/access/index/indexam.c       | 2 +-
 src/backend/access/nbtree/nbtree.c       | 2 +-
 src/include/utils/snapmgr.h              | 6 ++++--
 4 files changed, 7 insertions(+), 5 deletions(-)

diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index cbef73e5d4b..332c788bab2 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -159,7 +159,7 @@ heapam_index_fetch_tuple(struct IndexFetchTableData *scan,
 		 * Only in a non-MVCC snapshot can more than one member of the HOT
 		 * chain be visible.
 		 */
-		*call_again = !IsMVCCSnapshot(snapshot);
+		*call_again = !IsMVCCLikeSnapshot(snapshot);
 
 		slot->tts_tableOid = RelationGetRelid(scan->rel);
 		ExecStoreBufferHeapTuple(&bslot->base.tupdata, slot, hscan->xs_cbuf);
diff --git a/src/backend/access/index/indexam.c b/src/backend/access/index/indexam.c
index 4ed0508c605..80623350d6f 100644
--- a/src/backend/access/index/indexam.c
+++ b/src/backend/access/index/indexam.c
@@ -445,7 +445,7 @@ index_markpos(IndexScanDesc scan)
 void
 index_restrpos(IndexScanDesc scan)
 {
-	Assert(IsMVCCSnapshot(scan->xs_snapshot));
+	Assert(IsMVCCLikeSnapshot(scan->xs_snapshot));
 
 	SCAN_CHECKS;
 	CHECK_SCAN_PROCEDURE(amrestrpos);
diff --git a/src/backend/access/nbtree/nbtree.c b/src/backend/access/nbtree/nbtree.c
index 3dec1ee657d..07ba5997fbc 100644
--- a/src/backend/access/nbtree/nbtree.c
+++ b/src/backend/access/nbtree/nbtree.c
@@ -422,7 +422,7 @@ btrescan(IndexScanDesc scan, ScanKey scankey, int nscankeys,
 	 * Note: so->dropPin should never change across rescans.
 	 */
 	so->dropPin = (!scan->xs_want_itup &&
-				   IsMVCCSnapshot(scan->xs_snapshot) &&
+				   IsMVCCLikeSnapshot(scan->xs_snapshot) &&
 				   RelationNeedsWAL(scan->indexRelation) &&
 				   scan->heapRelation != NULL);
 
diff --git a/src/include/utils/snapmgr.h b/src/include/utils/snapmgr.h
index b8c01a291a1..dd5aaae6953 100644
--- a/src/include/utils/snapmgr.h
+++ b/src/include/utils/snapmgr.h
@@ -53,12 +53,14 @@ extern PGDLLIMPORT SnapshotData SnapshotToastData;
 
 /* This macro encodes the knowledge of which snapshots are MVCC-safe */
 #define IsMVCCSnapshot(snapshot)  \
-	((snapshot)->snapshot_type == SNAPSHOT_MVCC || \
-	 (snapshot)->snapshot_type == SNAPSHOT_HISTORIC_MVCC)
+	((snapshot)->snapshot_type == SNAPSHOT_MVCC)
 
 #define IsHistoricMVCCSnapshot(snapshot)  \
 	((snapshot)->snapshot_type == SNAPSHOT_HISTORIC_MVCC)
 
+#define IsMVCCLikeSnapshot(snapshot)  \
+	(IsMVCCSnapshot(snapshot) || IsHistoricMVCCSnapshot(snapshot))
+
 extern Snapshot GetTransactionSnapshot(void);
 extern Snapshot GetLatestSnapshot(void);
 extern void SnapshotSetCommandId(CommandId curcid);
-- 
2.47.3

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-09 11:42                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-02-09 22:16                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-10 07:46                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-02-10 16:49                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-12 10:36                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
@ 2026-03-11 23:09                                                     ` Andres Freund <andres@anarazel.de>
  2026-03-13 15:01                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  0 siblings, 1 reply; 120+ messages in thread

From: Andres Freund @ 2026-03-11 23:09 UTC (permalink / raw)
  To: Antonin Houska <ah@cybertec.at>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-02-12 11:36:08 +0100, Antonin Houska wrote:
> Andres Freund <andres@anarazel.de> wrote:
>
> > For something committable, I think we should probably split IsMVCCSnapshot
> > into IsMVCCSnapshot(), just accepting SNAPSHOT_MVCC, and IsMVCCLikeSnapshot()
> > accepting both SNAPSHOT_MVCC and SNAPSHOT_HISTORIC_MVCC. And then go through
> > all the existing callers of IsMVCCSnapshot() - only about half should stay
> > as-is, I think.
>
> The attached patch tries to do that.

Thanks!


> From dcdbaf3095e632a1f7f65f3abc43eccff0249d4c Mon Sep 17 00:00:00 2001
> From: Antonin Houska <ah@cybertec.at>
> Date: Thu, 12 Feb 2026 11:14:00 +0100
> Subject: [PATCH] Refine checking of snapshot type.
>
> It appears to be confusing if IsMVCCSnapshot() evaluates to true for both
> "regular" and "historic" MVCC snapshot. This patch restricts the meaning of
> the macro to the "regular" MVCC snapshot, and introduces a new macro
> IsMVCCLikeSnapshot() to recognize both types.
>
> IsMVCCLikeSnapshot() is only used in functions that can (supposedly) be called
> during logical decoding.

I think I agree with where you selected IsMVCCSnapshot() and where you
selected IsMVCCLikeSnapshot().


> diff --git a/src/include/utils/snapmgr.h b/src/include/utils/snapmgr.h
> index b8c01a291a1..dd5aaae6953 100644
> --- a/src/include/utils/snapmgr.h
> +++ b/src/include/utils/snapmgr.h
> @@ -53,12 +53,14 @@ extern PGDLLIMPORT SnapshotData SnapshotToastData;
>
>  /* This macro encodes the knowledge of which snapshots are MVCC-safe */
>  #define IsMVCCSnapshot(snapshot)  \
> -	((snapshot)->snapshot_type == SNAPSHOT_MVCC || \
> -	 (snapshot)->snapshot_type == SNAPSHOT_HISTORIC_MVCC)
> +	((snapshot)->snapshot_type == SNAPSHOT_MVCC)
>
>  #define IsHistoricMVCCSnapshot(snapshot)  \
>  	((snapshot)->snapshot_type == SNAPSHOT_HISTORIC_MVCC)
>
> +#define IsMVCCLikeSnapshot(snapshot)  \
> +	(IsMVCCSnapshot(snapshot) || IsHistoricMVCCSnapshot(snapshot))
> +
>  extern Snapshot GetTransactionSnapshot(void);
>  extern Snapshot GetLatestSnapshot(void);
>  extern void SnapshotSetCommandId(CommandId curcid);

Probably need to update the comments a bit.  What about something like


/*
 * Is the snapshot implemented as an MVCC snapshot (i.e. it uses
 * SNAPSHOT_MVCC).  If so, there will be at most be one visible row in a chain
 * of updated tuples, and each visible tuple will be seen exactly once.
 */
#define IsMVCCSnapshot(snapshot)  \
...

/*
 * Is the snapshot either an MVCC snapshot or has equivalent visibility
 * semantics (see IsMVCCSnapshot()).
 */
#define IsMVCCLikeSnapshot(snapshot)  \


Greetings,

Andres Freund





^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-09 11:42                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-02-09 22:16                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-10 07:46                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-02-10 16:49                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-12 10:36                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-03-11 23:09                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
@ 2026-03-13 15:01                                                       ` Antonin Houska <ah@cybertec.at>
  2026-03-13 17:55                                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  0 siblings, 1 reply; 120+ messages in thread

From: Antonin Houska @ 2026-03-13 15:01 UTC (permalink / raw)
  To: Andres Freund <andres@anarazel.de>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Andres Freund <andres@anarazel.de> wrote:

> Probably need to update the comments a bit.  What about something like
> 
> 
> /*
>  * Is the snapshot implemented as an MVCC snapshot (i.e. it uses
>  * SNAPSHOT_MVCC).  If so, there will be at most be one visible row in a chain
>  * of updated tuples, and each visible tuple will be seen exactly once.
>  */
> #define IsMVCCSnapshot(snapshot)  \

The ", and each visible tuple ..." part seemed to me redundant, so I omitted
it. If you think I'm wrong, please add it yourself when committing the patch.

I also added a comment to the IsHistoricMVCCSnapshot(), trying to explain what
"historic" means.

-- 
Antonin Houska
Web: https://www.cybertec-postgresql.com

Attachments:

  [text/x-diff] v2-Refine-checking-of-snapshot-type.patch (3.9K, ../../38845.1773414100@localhost/2-v2-Refine-checking-of-snapshot-type.patch)
  download | inline diff:
From 6c391d437ec61ff16b2a459a799d2030142e656f Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Thu, 12 Feb 2026 11:14:00 +0100
Subject: [PATCH] Refine checking of snapshot type.

It appears to be confusing if IsMVCCSnapshot() evaluates to true for both
"regular" and "historic" MVCC snapshot. This patch restricts the meaning of
the macro to the "regular" MVCC snapshot, and introduces a new macro
IsMVCCLikeSnapshot() to recognize both types.

IsMVCCLikeSnapshot() is only used in functions that can (supposedly) be called
during logical decoding.
---
 src/backend/access/heap/heapam_handler.c |  2 +-
 src/backend/access/index/indexam.c       |  2 +-
 src/backend/access/nbtree/nbtree.c       |  2 +-
 src/include/utils/snapmgr.h              | 21 ++++++++++++++++++---
 4 files changed, 21 insertions(+), 6 deletions(-)

diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index 5137d2510ea..2d1ee9ac95d 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -159,7 +159,7 @@ heapam_index_fetch_tuple(struct IndexFetchTableData *scan,
 		 * Only in a non-MVCC snapshot can more than one member of the HOT
 		 * chain be visible.
 		 */
-		*call_again = !IsMVCCSnapshot(snapshot);
+		*call_again = !IsMVCCLikeSnapshot(snapshot);
 
 		slot->tts_tableOid = RelationGetRelid(scan->rel);
 		ExecStoreBufferHeapTuple(&bslot->base.tupdata, slot, hscan->xs_cbuf);
diff --git a/src/backend/access/index/indexam.c b/src/backend/access/index/indexam.c
index 43f64a0e721..5eb7e99ad3e 100644
--- a/src/backend/access/index/indexam.c
+++ b/src/backend/access/index/indexam.c
@@ -445,7 +445,7 @@ index_markpos(IndexScanDesc scan)
 void
 index_restrpos(IndexScanDesc scan)
 {
-	Assert(IsMVCCSnapshot(scan->xs_snapshot));
+	Assert(IsMVCCLikeSnapshot(scan->xs_snapshot));
 
 	SCAN_CHECKS;
 	CHECK_SCAN_PROCEDURE(amrestrpos);
diff --git a/src/backend/access/nbtree/nbtree.c b/src/backend/access/nbtree/nbtree.c
index 6d0a6f27f3f..cdd81d147cc 100644
--- a/src/backend/access/nbtree/nbtree.c
+++ b/src/backend/access/nbtree/nbtree.c
@@ -423,7 +423,7 @@ btrescan(IndexScanDesc scan, ScanKey scankey, int nscankeys,
 	 * Note: so->dropPin should never change across rescans.
 	 */
 	so->dropPin = (!scan->xs_want_itup &&
-				   IsMVCCSnapshot(scan->xs_snapshot) &&
+				   IsMVCCLikeSnapshot(scan->xs_snapshot) &&
 				   RelationNeedsWAL(scan->indexRelation) &&
 				   scan->heapRelation != NULL);
 
diff --git a/src/include/utils/snapmgr.h b/src/include/utils/snapmgr.h
index b8c01a291a1..4f1e910bc75 100644
--- a/src/include/utils/snapmgr.h
+++ b/src/include/utils/snapmgr.h
@@ -51,14 +51,29 @@ extern PGDLLIMPORT SnapshotData SnapshotToastData;
 	((snapshotdata).snapshot_type = SNAPSHOT_NON_VACUUMABLE, \
 	 (snapshotdata).vistest = (vistestp))
 
-/* This macro encodes the knowledge of which snapshots are MVCC-safe */
+/*
+ * Is the snapshot implemented as an MVCC snapshot (i.e. it uses
+ * SNAPSHOT_MVCC)? If so, there will be at most one visible tuple in a chain
+ * of updated tuples.
+ */
 #define IsMVCCSnapshot(snapshot)  \
-	((snapshot)->snapshot_type == SNAPSHOT_MVCC || \
-	 (snapshot)->snapshot_type == SNAPSHOT_HISTORIC_MVCC)
+	((snapshot)->snapshot_type == SNAPSHOT_MVCC)
 
+/*
+ * Special kind of MVCC snapshot, to be used during logical decoding. The
+ * visibility is checked from the perspective of an already committed
+ * transaction, which we're trying to decode.
+ */
 #define IsHistoricMVCCSnapshot(snapshot)  \
 	((snapshot)->snapshot_type == SNAPSHOT_HISTORIC_MVCC)
 
+/*
+ * Is the snapshot either an MVCC snapshot or has equivalent visibility
+ * semantics (see IsMVCCSnapshot())?
+ */
+#define IsMVCCLikeSnapshot(snapshot)  \
+	(IsMVCCSnapshot(snapshot) || IsHistoricMVCCSnapshot(snapshot))
+
 extern Snapshot GetTransactionSnapshot(void);
 extern Snapshot GetLatestSnapshot(void);
 extern void SnapshotSetCommandId(CommandId curcid);
-- 
2.47.3

^ permalink  raw  reply  [nested|flat] 120+ messages in thread

* Re: Buffer locking is special (hints, checksums, AIO writes)
  2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-15 23:05 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-09-22 22:14   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-04 07:05     ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-06 22:55       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-07 16:40         ` Re: Buffer locking is special (hints, checksums, AIO writes) Matthias van de Meent <boekewurm+postgres@gmail.com>
  2025-10-09 20:35           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-10-09 21:16             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-20 02:47               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-11-25 15:44                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Melanie Plageman <melanieplageman@gmail.com>
  2025-11-25 16:54                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-03 00:47                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-17 09:25                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-17 14:54                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:03                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 17:20                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Heikki Linnakangas <hlinnaka@iki.fi>
  2025-12-18 22:06                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2025-12-18 23:39                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 00:29                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-09 08:08                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Kirill Reshke <reshkekirill@gmail.com>
  2026-01-12 17:45                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-01-13 00:33                                         ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-09 11:42                                           ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-02-09 22:16                                             ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-10 07:46                                               ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-02-10 16:49                                                 ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-02-12 10:36                                                   ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
  2026-03-11 23:09                                                     ` Re: Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
  2026-03-13 15:01                                                       ` Re: Buffer locking is special (hints, checksums, AIO writes) Antonin Houska <ah@cybertec.at>
@ 2026-03-13 17:55                                                         ` Andres Freund <andres@anarazel.de>
  0 siblings, 0 replies; 120+ messages in thread

From: Andres Freund @ 2026-03-13 17:55 UTC (permalink / raw)
  To: Antonin Houska <ah@cybertec.at>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Melanie Plageman <melanieplageman@gmail.com>; Matthias van de Meent <boekewurm+postgres@gmail.com>; pgsql-hackers@postgresql.org, Thomas Munro <thomas.munro@gmail.com>; Noah Misch <noah@leadboat.com>; Robert Haas <robertmhaas@gmail.com>; Michael Paquier <michael.paquier@gmail.com>

Hi,

On 2026-03-13 16:01:40 +0100, Antonin Houska wrote:
> Andres Freund <andres@anarazel.de> wrote:
> 
> > Probably need to update the comments a bit.  What about something like
> > 
> > 
> > /*
> >  * Is the snapshot implemented as an MVCC snapshot (i.e. it uses
> >  * SNAPSHOT_MVCC).  If so, there will be at most be one visible row in a chain
> >  * of updated tuples, and each visible tuple will be seen exactly once.
> >  */
> > #define IsMVCCSnapshot(snapshot)  \
> 
> The ", and each visible tuple ..." part seemed to me redundant, so I omitted
> it. If you think I'm wrong, please add it yourself when committing the patch.

It's relevant in that many non-mvcc scan types do *not* guarantee that (i.e. a
tuple may never be seen, e.g. because the new version of the tuple is placed
before the current scan position of a scan and the old version of the tuple is
not considered visible anymore).


> I also added a comment to the IsHistoricMVCCSnapshot(), trying to explain what
> "historic" means.

Good idea.


Pushed with slightly revised comments and a different commit message (I
thought it was important to explain that this fixes breakage during logical
decoding, even if currently hard to reach).


Thanks for the report and patch!

- Andres





^ permalink  raw  reply  [nested|flat] 120+ messages in thread


end of thread, other threads:[~2026-04-03 10:06 UTC | newest]

Thread overview: 120+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2025-08-22 19:44 Buffer locking is special (hints, checksums, AIO writes) Andres Freund <andres@anarazel.de>
2025-08-23 10:31 ` Mihail Nikalayeu <mihailnikalayeu@gmail.com>
2025-08-26 20:21 ` Robert Haas <robertmhaas@gmail.com>
2025-08-26 21:00   ` Andres Freund <andres@anarazel.de>
2025-08-27 00:14     ` Noah Misch <noah@leadboat.com>
2025-08-27 14:03       ` Robert Haas <robertmhaas@gmail.com>
2025-09-01 02:03         ` Michael Paquier <michael@paquier.xyz>
2025-08-27 16:18       ` Andres Freund <andres@anarazel.de>
2025-08-27 19:14         ` Noah Misch <noah@leadboat.com>
2025-08-27 19:29           ` Andres Freund <andres@anarazel.de>
2025-08-27 23:04             ` Noah Misch <noah@leadboat.com>
2025-08-27 19:22         ` Robert Haas <robertmhaas@gmail.com>
2025-09-15 23:05 ` Andres Freund <andres@anarazel.de>
2025-09-22 22:14   ` Andres Freund <andres@anarazel.de>
2025-10-04 07:05     ` Matthias van de Meent <boekewurm+postgres@gmail.com>
2025-10-06 22:55       ` Andres Freund <andres@anarazel.de>
2025-10-07 16:40         ` Matthias van de Meent <boekewurm+postgres@gmail.com>
2025-10-09 20:35           ` Andres Freund <andres@anarazel.de>
2025-10-09 21:16             ` Andres Freund <andres@anarazel.de>
2025-11-20 02:47               ` Andres Freund <andres@anarazel.de>
2025-11-20 19:08                 ` Greg Burd <greg@burd.me>
2025-11-20 20:51                   ` Andres Freund <andres@anarazel.de>
2025-11-25 00:09                     ` Thomas Munro <thomas.munro@gmail.com>
2025-11-21 17:52                 ` Melanie Plageman <melanieplageman@gmail.com>
2025-12-01 20:28                   ` Andres Freund <andres@anarazel.de>
2025-12-01 20:41                     ` Melanie Plageman <melanieplageman@gmail.com>
2025-11-24 20:57                 ` Andres Freund <andres@anarazel.de>
2025-11-24 21:04                   ` Melanie Plageman <melanieplageman@gmail.com>
2025-11-25 00:17                     ` Andres Freund <andres@anarazel.de>
2025-11-25 15:44                 ` Melanie Plageman <melanieplageman@gmail.com>
2025-11-25 16:54                   ` Andres Freund <andres@anarazel.de>
2025-11-25 20:02                     ` Melanie Plageman <melanieplageman@gmail.com>
2025-11-25 20:46                       ` Andres Freund <andres@anarazel.de>
2025-11-25 21:23                         ` Melanie Plageman <melanieplageman@gmail.com>
2025-12-02 08:01                         ` Heikki Linnakangas <hlinnaka@iki.fi>
2025-12-02 13:20                           ` Andres Freund <andres@anarazel.de>
2025-12-02 13:38                             ` Heikki Linnakangas <hlinnaka@iki.fi>
2025-12-03 00:47                     ` Andres Freund <andres@anarazel.de>
2025-12-03 01:12                       ` Peter Geoghegan <pg@bowt.ie>
2025-12-03 01:18                         ` Andres Freund <andres@anarazel.de>
2025-12-03 16:03                       ` Andres Freund <andres@anarazel.de>
2025-12-17 09:25                       ` Heikki Linnakangas <hlinnaka@iki.fi>
2025-12-17 14:54                         ` Andres Freund <andres@anarazel.de>
2025-12-18 17:03                           ` Andres Freund <andres@anarazel.de>
2025-12-18 17:20                             ` Heikki Linnakangas <hlinnaka@iki.fi>
2025-12-18 22:06                               ` Andres Freund <andres@anarazel.de>
2025-12-18 23:39                                 ` Andres Freund <andres@anarazel.de>
2026-01-09 00:29                                   ` Andres Freund <andres@anarazel.de>
2026-01-09 08:08                                     ` Kirill Reshke <reshkekirill@gmail.com>
2026-01-12 17:45                                       ` Andres Freund <andres@anarazel.de>
2026-01-12 22:27                                         ` Melanie Plageman <melanieplageman@gmail.com>
2026-01-12 23:22                                           ` Andres Freund <andres@anarazel.de>
2026-01-13 14:59                                             ` Melanie Plageman <melanieplageman@gmail.com>
2026-01-13 00:33                                         ` Andres Freund <andres@anarazel.de>
2026-01-13 15:05                                           ` Melanie Plageman <melanieplageman@gmail.com>
2026-01-14 00:49                                             ` Andres Freund <andres@anarazel.de>
2026-01-14 14:17                                               ` Melanie Plageman <melanieplageman@gmail.com>
2026-01-14 15:20                                                 ` Andres Freund <andres@anarazel.de>
2026-01-14 02:26                                           ` Chao Li <li.evan.chao@gmail.com>
2026-01-14 16:23                                             ` Andres Freund <andres@anarazel.de>
2026-01-14 03:41                                           ` Chao Li <li.evan.chao@gmail.com>
2026-01-14 16:30                                             ` Andres Freund <andres@anarazel.de>
2026-01-14 23:20                                               ` Chao Li <li.evan.chao@gmail.com>
2026-01-14 23:37                                                 ` Andres Freund <andres@anarazel.de>
2026-01-15 00:04                                                   ` Chao Li <li.evan.chao@gmail.com>
2026-01-15 06:22                                                     ` Chao Li <li.evan.chao@gmail.com>
2026-01-15 16:43                                                       ` Andres Freund <andres@anarazel.de>
2026-01-15 23:02                                                         ` Tom Lane <tgl@sss.pgh.pa.us>
2026-01-15 23:16                                                           ` Andres Freund <andres@anarazel.de>
2026-01-15 23:19                                                             ` Tom Lane <tgl@sss.pgh.pa.us>
2026-01-15 23:26                                                               ` Andres Freund <andres@anarazel.de>
2026-01-16 12:02                                                                 ` Andres Freund <andres@anarazel.de>
2026-01-24 19:00                                                           ` Alexander Lakhin <exclusion@gmail.com>
2026-01-24 20:31                                                             ` Andres Freund <andres@anarazel.de>
2026-01-24 21:03                                                               ` Andres Freund <andres@anarazel.de>
2026-01-24 21:11                                                               ` Tom Lane <tgl@sss.pgh.pa.us>
2026-01-24 21:21                                                                 ` Peter Geoghegan <pg@bowt.ie>
2026-01-24 23:03                                                                 ` Andres Freund <andres@anarazel.de>
2026-01-25 00:54                                                                   ` Tom Lane <tgl@sss.pgh.pa.us>
2026-01-29 17:27                                                                     ` Andres Freund <andres@anarazel.de>
2026-01-29 17:42                                                                       ` Andres Freund <andres@anarazel.de>
2026-01-29 17:50                                                                         ` Peter Geoghegan <pg@bowt.ie>
2026-01-29 18:06                                                                           ` Andres Freund <andres@anarazel.de>
2026-01-29 18:33                                                                             ` Peter Geoghegan <pg@bowt.ie>
2026-01-29 19:29                                                                               ` Andres Freund <andres@anarazel.de>
2026-01-29 21:49                                                                             ` Andres Freund <andres@anarazel.de>
2026-01-29 18:12                                                                       ` Peter Geoghegan <pg@bowt.ie>
2026-01-29 20:24                                                                       ` Tom Lane <tgl@sss.pgh.pa.us>
2026-01-16 02:36                                                         ` Chao Li <li.evan.chao@gmail.com>
2026-01-14 21:20                                           ` Andres Freund <andres@anarazel.de>
2026-02-02 22:33                                             ` Andres Freund <andres@anarazel.de>
2026-02-06 21:18                                               ` Kirill Reshke <reshkekirill@gmail.com>
2026-02-07 10:44                                               ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-15 19:52                                                 ` Noah Misch <noah@leadboat.com>
2026-03-11 22:40                                                   ` Andres Freund <andres@anarazel.de>
2026-03-13 08:00                                                     ` Alexander Lakhin <exclusion@gmail.com>
2026-03-13 15:55                                                       ` Andres Freund <andres@anarazel.de>
2026-03-17 20:50                                                         ` Andres Freund <andres@anarazel.de>
2026-03-25 21:34                                                     ` Melanie Plageman <melanieplageman@gmail.com>
2026-03-25 22:35                                                       ` Andres Freund <andres@anarazel.de>
2026-03-27 20:00                                                         ` Andres Freund <andres@anarazel.de>
2026-03-31 16:02                                                           ` Yura Sokolov <y.sokolov@postgrespro.ru>
2026-03-31 22:05                                                             ` Andres Freund <andres@anarazel.de>
2026-04-01 00:29                                                               ` Andres Freund <andres@anarazel.de>
2026-04-03 10:06                                                                 ` =?utf-8?B?Y2NhNTUwNw==?= <cca5507@qq.com>
2026-02-07 12:38                                               ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-08 18:38                                                 ` Andres Freund <andres@anarazel.de>
2026-02-09 19:54                                                   ` Kirill Reshke <reshkekirill@gmail.com>
2026-02-07 12:59                                               ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 01:52                                                 ` Andres Freund <andres@anarazel.de>
2026-02-09 10:14                                                   ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-02-09 22:19                                                     ` Andres Freund <andres@anarazel.de>
2026-02-09 11:42                                           ` Antonin Houska <ah@cybertec.at>
2026-02-09 22:16                                             ` Andres Freund <andres@anarazel.de>
2026-02-10 07:46                                               ` Antonin Houska <ah@cybertec.at>
2026-02-10 16:49                                                 ` Andres Freund <andres@anarazel.de>
2026-02-12 10:36                                                   ` Antonin Houska <ah@cybertec.at>
2026-03-11 23:09                                                     ` Andres Freund <andres@anarazel.de>
2026-03-13 15:01                                                       ` Antonin Houska <ah@cybertec.at>
2026-03-13 17:55                                                         ` Andres Freund <andres@anarazel.de>

This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox