agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedslab allocator performance issues
29+ messages / 6 participants
[nested] [flat]
* slab allocator performance issues
@ 2021-07-17 19:43 Andres Freund <andres@anarazel.de>
2021-07-17 19:53 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
0 siblings, 2 replies; 29+ messages in thread
From: Andres Freund @ 2021-07-17 19:43 UTC (permalink / raw)
To: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Hi,
I just tried to use the slab allocator for a case where aset.c was
bloating memory usage substantially. First: It worked wonders for memory
usage, nearly eliminating overhead.
But it turned out to cause a *substantial* slowdown. With aset the
allocator is barely in the profile. With slab the profile is dominated
by allocator performance.
slab:
NOTICE: 00000: 100000000 ordered insertions in 5.216287 seconds, 19170724/sec
LOCATION: bfm_test_insert_bulk, radix.c:2880
Overhead Command Shared Object Symbol
+ 28.27% postgres postgres [.] SlabAlloc
+ 9.64% postgres bdbench.so [.] bfm_delete
+ 9.03% postgres bdbench.so [.] bfm_set
+ 8.39% postgres bdbench.so [.] bfm_lookup
+ 8.36% postgres bdbench.so [.] bfm_set_leaf.constprop.0
+ 8.16% postgres libc-2.31.so [.] __memmove_avx_unaligned_erms
+ 6.88% postgres bdbench.so [.] bfm_delete_leaf
+ 3.24% postgres libc-2.31.so [.] _int_malloc
+ 2.58% postgres bdbench.so [.] bfm_tests
+ 2.33% postgres postgres [.] SlabFree
+ 1.29% postgres libc-2.31.so [.] _int_free
+ 1.09% postgres libc-2.31.so [.] unlink_chunk.constprop.0
aset:
NOTICE: 00000: 100000000 ordered insertions in 2.082602 seconds, 48016848/sec
LOCATION: bfm_test_insert_bulk, radix.c:2880
+ 16.43% postgres bdbench.so [.] bfm_lookup
+ 15.38% postgres bdbench.so [.] bfm_delete
+ 12.82% postgres libc-2.31.so [.] __memmove_avx_unaligned_erms
+ 12.65% postgres bdbench.so [.] bfm_set
+ 12.15% postgres bdbench.so [.] bfm_set_leaf.constprop.0
+ 10.57% postgres bdbench.so [.] bfm_delete_leaf
+ 4.05% postgres bdbench.so [.] bfm_tests
+ 2.93% postgres [kernel.vmlinux] [k] clear_page_erms
+ 1.59% postgres postgres [.] AllocSetAlloc
+ 1.15% postgres bdbench.so [.] memmove@plt
+ 1.06% postgres bdbench.so [.] bfm_grow_leaf_16
OS:
NOTICE: 00000: 100000000 ordered insertions in 2.089790 seconds, 47851690/sec
LOCATION: bfm_test_insert_bulk, radix.c:2880
That is somewhat surprising - part of the promise of a slab allocator is
that it's fast...
This is caused by multiple issues, I think. Some of which seems fairly easy to
fix.
1) If allocations are short-lived slab.c, can end up constantly
freeing/initializing blocks. Which requires fairly expensively iterating over
all potential chunks in the block and initializing it. Just to then free that
memory again after a small number of allocations. The extreme case of this is
when there are phases of alloc/free of a single allocation.
I "fixed" this by adding a few && slab->nblocks > 1 in SlabFree() and the
problem vastly reduced. Instead of a 0.4x slowdown it's 0.88x. Of course that
only works if the problem is with the only, so it's not a great
approach. Perhaps just keeping the last allocated block around would work?
2) SlabChunkIndex() in SlabFree() is slow. It requires a 64bit division, taking
up ~50% of the cycles in SlabFree(). A 64bit div, according to [1] , has a
latency of 35-88 cycles on skylake-x (and a reverse throughput of 21-83,
i.e. no parallelism). While it's getting a bit faster on icelake / zen 3, it's
still slow enough there to be very worrisome.
I don't see a way to get around the division while keeping the freelist
structure as is. But:
ISTM that we only need the index because of the free-chunk list, right? Why
don't we make the chunk list use actual pointers? Is it concern that that'd
increase the minimum allocation size? If so, I see two ways around that:
First, we could make the index just the offset from the start of the block,
that's much cheaper to calculate. Second, we could store the next pointer in
SlabChunk->slab/block instead (making it a union) - while on the freelist we
don't need to dereference those, right?
I suspect both would also make the block initialization a bit cheaper.
That should also accelerate SlabBlockGetChunk(), which currently shows up as
an imul, which isn't exactly fast either (and uses a lot of execution ports).
3) Every free/alloc needing to unlink from slab->freelist[i] and then relink
shows up prominently in profiles. That's ~tripling the number of cachelines
touched in the happy path, with unpredictable accesses to boot.
Perhaps we could reduce the precision of slab->freelist indexing to amortize
that cost? I.e. not have one slab->freelist entry for each nfree, but instead
have an upper limit on the number of freelists?
4) Less of a performance, and more of a usability issue: The constant
block size strikes me as problematic. Most users of an allocator can
sometimes be used with a small amount of data, and sometimes with a
large amount.
Greetings,
Andres Freund
[1] https://www.agner.org/optimize/instruction_tables.pdf
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
@ 2021-07-17 19:53 ` Andres Freund <andres@anarazel.de>
1 sibling, 0 replies; 29+ messages in thread
From: Andres Freund @ 2021-07-17 19:53 UTC (permalink / raw)
To: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Hi,
On 2021-07-17 12:43:33 -0700, Andres Freund wrote:
> 2) SlabChunkIndex() in SlabFree() is slow. It requires a 64bit division, taking
> up ~50% of the cycles in SlabFree(). A 64bit div, according to [1] , has a
> latency of 35-88 cycles on skylake-x (and a reverse throughput of 21-83,
> i.e. no parallelism). While it's getting a bit faster on icelake / zen 3, it's
> still slow enough there to be very worrisome.
>
> I don't see a way to get around the division while keeping the freelist
> structure as is. But:
>
> ISTM that we only need the index because of the free-chunk list, right? Why
> don't we make the chunk list use actual pointers? Is it concern that that'd
> increase the minimum allocation size? If so, I see two ways around that:
> First, we could make the index just the offset from the start of the block,
> that's much cheaper to calculate. Second, we could store the next pointer in
> SlabChunk->slab/block instead (making it a union) - while on the freelist we
> don't need to dereference those, right?
>
> I suspect both would also make the block initialization a bit cheaper.
>
> That should also accelerate SlabBlockGetChunk(), which currently shows up as
> an imul, which isn't exactly fast either (and uses a lot of execution ports).
Oh - I just saw that effectively the allocation size already is a
uintptr_t at minimum. I had only seen
/* Make sure the linked list node fits inside a freed chunk */
if (chunkSize < sizeof(int))
chunkSize = sizeof(int);
but it's followed by
/* chunk, including SLAB header (both addresses nicely aligned) */
fullChunkSize = sizeof(SlabChunk) + MAXALIGN(chunkSize);
which means we are reserving enough space for a pointer on just about
any platform already? Seems we can just make that official and reserve
space for a pointer as part of the chunk size rounding up, instead of
fullChunkSize?
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
@ 2021-07-17 20:35 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
1 sibling, 1 reply; 29+ messages in thread
From: Tomas Vondra @ 2021-07-17 20:35 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Hi,
On 7/17/21 9:43 PM, Andres Freund wrote:
> Hi,
>
> I just tried to use the slab allocator for a case where aset.c was
> bloating memory usage substantially. First: It worked wonders for memory
> usage, nearly eliminating overhead.
>
> But it turned out to cause a *substantial* slowdown. With aset the
> allocator is barely in the profile. With slab the profile is dominated
> by allocator performance.
>
> slab:
> NOTICE: 00000: 100000000 ordered insertions in 5.216287 seconds, 19170724/sec
> LOCATION: bfm_test_insert_bulk, radix.c:2880
> Overhead Command Shared Object Symbol
>
> + 28.27% postgres postgres [.] SlabAlloc
> + 9.64% postgres bdbench.so [.] bfm_delete
> + 9.03% postgres bdbench.so [.] bfm_set
> + 8.39% postgres bdbench.so [.] bfm_lookup
> + 8.36% postgres bdbench.so [.] bfm_set_leaf.constprop.0
> + 8.16% postgres libc-2.31.so [.] __memmove_avx_unaligned_erms
> + 6.88% postgres bdbench.so [.] bfm_delete_leaf
> + 3.24% postgres libc-2.31.so [.] _int_malloc
> + 2.58% postgres bdbench.so [.] bfm_tests
> + 2.33% postgres postgres [.] SlabFree
> + 1.29% postgres libc-2.31.so [.] _int_free
> + 1.09% postgres libc-2.31.so [.] unlink_chunk.constprop.0
>
> aset:
>
> NOTICE: 00000: 100000000 ordered insertions in 2.082602 seconds, 48016848/sec
> LOCATION: bfm_test_insert_bulk, radix.c:2880
>
> + 16.43% postgres bdbench.so [.] bfm_lookup
> + 15.38% postgres bdbench.so [.] bfm_delete
> + 12.82% postgres libc-2.31.so [.] __memmove_avx_unaligned_erms
> + 12.65% postgres bdbench.so [.] bfm_set
> + 12.15% postgres bdbench.so [.] bfm_set_leaf.constprop.0
> + 10.57% postgres bdbench.so [.] bfm_delete_leaf
> + 4.05% postgres bdbench.so [.] bfm_tests
> + 2.93% postgres [kernel.vmlinux] [k] clear_page_erms
> + 1.59% postgres postgres [.] AllocSetAlloc
> + 1.15% postgres bdbench.so [.] memmove@plt
> + 1.06% postgres bdbench.so [.] bfm_grow_leaf_16
>
> OS:
> NOTICE: 00000: 100000000 ordered insertions in 2.089790 seconds, 47851690/sec
> LOCATION: bfm_test_insert_bulk, radix.c:2880
>
>
> That is somewhat surprising - part of the promise of a slab allocator is
> that it's fast...
>
>
> This is caused by multiple issues, I think. Some of which seems fairly easy to
> fix.
>
> 1) If allocations are short-lived slab.c, can end up constantly
> freeing/initializing blocks. Which requires fairly expensively iterating over
> all potential chunks in the block and initializing it. Just to then free that
> memory again after a small number of allocations. The extreme case of this is
> when there are phases of alloc/free of a single allocation.
>
> I "fixed" this by adding a few && slab->nblocks > 1 in SlabFree() and the
> problem vastly reduced. Instead of a 0.4x slowdown it's 0.88x. Of course that
> only works if the problem is with the only, so it's not a great
> approach. Perhaps just keeping the last allocated block around would work?
>
+1
I think it makes perfect sense to not free the blocks immediately, and
keep one (or a small number) as a cache. I'm not sure why we decided not
to have a "keeper" block, but I suspect memory consumption was my main
concern at that point. But I never expected the cost to be this high.
>
> 2) SlabChunkIndex() in SlabFree() is slow. It requires a 64bit division, taking
> up ~50% of the cycles in SlabFree(). A 64bit div, according to [1] , has a
> latency of 35-88 cycles on skylake-x (and a reverse throughput of 21-83,
> i.e. no parallelism). While it's getting a bit faster on icelake / zen 3, it's
> still slow enough there to be very worrisome.
>
> I don't see a way to get around the division while keeping the freelist
> structure as is. But:
>
> ISTM that we only need the index because of the free-chunk list, right? Why
> don't we make the chunk list use actual pointers? Is it concern that that'd
> increase the minimum allocation size? If so, I see two ways around that:
> First, we could make the index just the offset from the start of the block,
> that's much cheaper to calculate. Second, we could store the next pointer in
> SlabChunk->slab/block instead (making it a union) - while on the freelist we
> don't need to dereference those, right?
>
> I suspect both would also make the block initialization a bit cheaper.
>
> That should also accelerate SlabBlockGetChunk(), which currently shows up as
> an imul, which isn't exactly fast either (and uses a lot of execution ports).
>
Hmm, I think you're right we could simply use the pointers, but I have
not tried that.
>
> 3) Every free/alloc needing to unlink from slab->freelist[i] and then relink
> shows up prominently in profiles. That's ~tripling the number of cachelines
> touched in the happy path, with unpredictable accesses to boot.
>
> Perhaps we could reduce the precision of slab->freelist indexing to amortize
> that cost? I.e. not have one slab->freelist entry for each nfree, but instead
> have an upper limit on the number of freelists?
>
Yeah. The purpose of organizing the freelists like this is to prioritize
the "more full" blocks" when allocating new chunks, in the hope that the
"less full" blocks will end up empty and freed faster.
But this is naturally imprecise, and strongly depends on the workload,
of course, and I bet for most cases a less precise approach would work
just as fine.
I'm not sure how exactly would the upper limit you propose work, but
perhaps we could group the blocks for nfree ranges, say [0-15], [16-31]
and so on. So after the alloc/free we'd calculate the new freelist index
as (nfree/16) and only moved the block if the index changed. This would
reduce the overhead to 1/16 and probably even more in practice.
Of course, we could also say we have e.g. 8 freelists and work the
ranges backwards from that, I guess that's what you propose.
>
> 4) Less of a performance, and more of a usability issue: The constant
> block size strikes me as problematic. Most users of an allocator can
> sometimes be used with a small amount of data, and sometimes with a
> large amount.
>
I doubt this is worth the effort, really. The constant block size makes
various places much simpler (both to code and reason about), so this
should not make a huge difference in performance. And IMHO the block
size is mostly an implementation detail, so I don't see that as a
usability issue.
regards
--
Tomas Vondra
EnterpriseDB: http://www.enterprisedb.com
The Enterprise PostgreSQL Company
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
@ 2021-07-17 21:14 ` Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
0 siblings, 1 reply; 29+ messages in thread
From: Andres Freund @ 2021-07-17 21:14 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@enterprisedb.com>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Hi,
On 2021-07-17 22:35:07 +0200, Tomas Vondra wrote:
> On 7/17/21 9:43 PM, Andres Freund wrote:
> > 1) If allocations are short-lived slab.c, can end up constantly
> > freeing/initializing blocks. Which requires fairly expensively iterating over
> > all potential chunks in the block and initializing it. Just to then free that
> > memory again after a small number of allocations. The extreme case of this is
> > when there are phases of alloc/free of a single allocation.
> >
> > I "fixed" this by adding a few && slab->nblocks > 1 in SlabFree() and the
> > problem vastly reduced. Instead of a 0.4x slowdown it's 0.88x. Of course that
> > only works if the problem is with the only, so it's not a great
> > approach. Perhaps just keeping the last allocated block around would work?
> >
>
> +1
>
> I think it makes perfect sense to not free the blocks immediately, and keep
> one (or a small number) as a cache. I'm not sure why we decided not to have
> a "keeper" block, but I suspect memory consumption was my main concern at
> that point. But I never expected the cost to be this high.
I think one free block might be too low in some cases. It's pretty
common to have workloads where the number of allocations is "bursty",
and it's imo one case where one might justifiably want to use a slab
allocator... Perhaps a portion of a high watermark? Or a portion of the
in use blocks?
Hm. I wonder if we should just not populate the freelist eagerly, to
drive down the initialization cost. I.e. have a separate allocation path
for chunks that have never been allocated, by having a
SlabBlock->free_offset or such.
Sure, it adds a branch to the allocation happy path, but it also makes the
first allocation for a chunk cheaper, because there's no need to get the next
element from the freelist (adding a likely cache miss). And it should make the
allocation of a new block faster by a lot.
> > 2) SlabChunkIndex() in SlabFree() is slow. It requires a 64bit division, taking
> > up ~50% of the cycles in SlabFree(). A 64bit div, according to [1] , has a
> > latency of 35-88 cycles on skylake-x (and a reverse throughput of 21-83,
> > i.e. no parallelism). While it's getting a bit faster on icelake / zen 3, it's
> > still slow enough there to be very worrisome.
> >
> > I don't see a way to get around the division while keeping the freelist
> > structure as is. But:
> >
> > ISTM that we only need the index because of the free-chunk list, right? Why
> > don't we make the chunk list use actual pointers? Is it concern that that'd
> > increase the minimum allocation size? If so, I see two ways around that:
> > First, we could make the index just the offset from the start of the block,
> > that's much cheaper to calculate. Second, we could store the next pointer in
> > SlabChunk->slab/block instead (making it a union) - while on the freelist we
> > don't need to dereference those, right?
> >
> > I suspect both would also make the block initialization a bit cheaper.
> >
> > That should also accelerate SlabBlockGetChunk(), which currently shows up as
> > an imul, which isn't exactly fast either (and uses a lot of execution ports).
> >
>
> Hmm, I think you're right we could simply use the pointers, but I have not
> tried that.
I quickly tried that, and it does seem to improve matters considerably. The
block initialization still shows up as expensive, but not as bad. The div and
imul are gone (exept in an assertion build right now). The list manipulation
still is visible.
> > 3) Every free/alloc needing to unlink from slab->freelist[i] and then relink
> > shows up prominently in profiles. That's ~tripling the number of cachelines
> > touched in the happy path, with unpredictable accesses to boot.
> >
> > Perhaps we could reduce the precision of slab->freelist indexing to amortize
> > that cost? I.e. not have one slab->freelist entry for each nfree, but instead
> > have an upper limit on the number of freelists?
> >
>
> Yeah. The purpose of organizing the freelists like this is to prioritize the
> "more full" blocks" when allocating new chunks, in the hope that the "less
> full" blocks will end up empty and freed faster.
>
> But this is naturally imprecise, and strongly depends on the workload, of
> course, and I bet for most cases a less precise approach would work just as
> fine.
>
> I'm not sure how exactly would the upper limit you propose work, but perhaps
> we could group the blocks for nfree ranges, say [0-15], [16-31] and so on.
> So after the alloc/free we'd calculate the new freelist index as (nfree/16)
> and only moved the block if the index changed. This would reduce the
> overhead to 1/16 and probably even more in practice.
> Of course, we could also say we have e.g. 8 freelists and work the ranges
> backwards from that, I guess that's what you propose.
Yea, I was thinking something along those lines. As you say, ow about there's
always at most 8 freelists or such. During initialization we compute a shift
that distributes chunksPerBlock from 0 to 8. Then we only need to perform list
manipulation if block->nfree >> slab->freelist_shift != --block->nfree >> slab->freelist_shift.
That seems nice from a memory usage POV as well - for small and frequent
allocations needing chunksPerBlock freelists isn't great. The freelists are on
a fair number of cachelines right now.
> > 4) Less of a performance, and more of a usability issue: The constant
> > block size strikes me as problematic. Most users of an allocator can
> > sometimes be used with a small amount of data, and sometimes with a
> > large amount.
> >
>
> I doubt this is worth the effort, really. The constant block size makes
> various places much simpler (both to code and reason about), so this should
> not make a huge difference in performance. And IMHO the block size is mostly
> an implementation detail, so I don't see that as a usability issue.
Hm? It's something the user has to specify, so I it's not really an
implementation detail. It needs to be specified without sufficient
information, as well, since externally one doesn't know how much memory the
block header and chunk headers + rounding up will use, so computing a good
block size isn't easy. I've wondered whether it should just be a count...
Why do you not think it's relevant for performance? Either one causes too much
memory usage by using a too large block size, wasting memory, or one ends up
loosing perf through frequent allocations?
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
@ 2021-07-17 22:46 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
0 siblings, 1 reply; 29+ messages in thread
From: Tomas Vondra @ 2021-07-17 22:46 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On 7/17/21 11:14 PM, Andres Freund wrote:
> Hi,
>
> On 2021-07-17 22:35:07 +0200, Tomas Vondra wrote:
>> On 7/17/21 9:43 PM, Andres Freund wrote:
>>> 1) If allocations are short-lived slab.c, can end up constantly
>>> freeing/initializing blocks. Which requires fairly expensively iterating over
>>> all potential chunks in the block and initializing it. Just to then free that
>>> memory again after a small number of allocations. The extreme case of this is
>>> when there are phases of alloc/free of a single allocation.
>>>
>>> I "fixed" this by adding a few && slab->nblocks > 1 in SlabFree() and the
>>> problem vastly reduced. Instead of a 0.4x slowdown it's 0.88x. Of course that
>>> only works if the problem is with the only, so it's not a great
>>> approach. Perhaps just keeping the last allocated block around would work?
>>>
>>
>> +1
>>
>> I think it makes perfect sense to not free the blocks immediately, and keep
>> one (or a small number) as a cache. I'm not sure why we decided not to have
>> a "keeper" block, but I suspect memory consumption was my main concern at
>> that point. But I never expected the cost to be this high.
>
> I think one free block might be too low in some cases. It's pretty
> common to have workloads where the number of allocations is "bursty",
> and it's imo one case where one might justifiably want to use a slab
> allocator... Perhaps a portion of a high watermark? Or a portion of the
> in use blocks?
>
I think the portion of watermark would be problematic for cases with one
huge transaction - that'll set a high watermark, and we'll keep way too
many free blocks. But the portion of in use blocks might work, I think.
> Hm. I wonder if we should just not populate the freelist eagerly, to
> drive down the initialization cost. I.e. have a separate allocation path
> for chunks that have never been allocated, by having a
> SlabBlock->free_offset or such.
>
> Sure, it adds a branch to the allocation happy path, but it also makes the
> first allocation for a chunk cheaper, because there's no need to get the next
> element from the freelist (adding a likely cache miss). And it should make the
> allocation of a new block faster by a lot.
>
Not sure what you mean by 'not populate eagerly' so can't comment :-(
>
>>> 2) SlabChunkIndex() in SlabFree() is slow. It requires a 64bit division, taking
>>> up ~50% of the cycles in SlabFree(). A 64bit div, according to [1] , has a
>>> latency of 35-88 cycles on skylake-x (and a reverse throughput of 21-83,
>>> i.e. no parallelism). While it's getting a bit faster on icelake / zen 3, it's
>>> still slow enough there to be very worrisome.
>>>
>>> I don't see a way to get around the division while keeping the freelist
>>> structure as is. But:
>>>
>>> ISTM that we only need the index because of the free-chunk list, right? Why
>>> don't we make the chunk list use actual pointers? Is it concern that that'd
>>> increase the minimum allocation size? If so, I see two ways around that:
>>> First, we could make the index just the offset from the start of the block,
>>> that's much cheaper to calculate. Second, we could store the next pointer in
>>> SlabChunk->slab/block instead (making it a union) - while on the freelist we
>>> don't need to dereference those, right?
>>>
>>> I suspect both would also make the block initialization a bit cheaper.
>>>
>>> That should also accelerate SlabBlockGetChunk(), which currently shows up as
>>> an imul, which isn't exactly fast either (and uses a lot of execution ports).
>>>
>>
>> Hmm, I think you're right we could simply use the pointers, but I have not
>> tried that.
>
> I quickly tried that, and it does seem to improve matters considerably. The
> block initialization still shows up as expensive, but not as bad. The div and
> imul are gone (exept in an assertion build right now). The list manipulation
> still is visible.
>
Understood. I didn't expect this to be a full solution.
>
>>> 3) Every free/alloc needing to unlink from slab->freelist[i] and then relink
>>> shows up prominently in profiles. That's ~tripling the number of cachelines
>>> touched in the happy path, with unpredictable accesses to boot.
>>>
>>> Perhaps we could reduce the precision of slab->freelist indexing to amortize
>>> that cost? I.e. not have one slab->freelist entry for each nfree, but instead
>>> have an upper limit on the number of freelists?
>>>
>>
>> Yeah. The purpose of organizing the freelists like this is to prioritize the
>> "more full" blocks" when allocating new chunks, in the hope that the "less
>> full" blocks will end up empty and freed faster.
>>
>> But this is naturally imprecise, and strongly depends on the workload, of
>> course, and I bet for most cases a less precise approach would work just as
>> fine.
>>
>> I'm not sure how exactly would the upper limit you propose work, but perhaps
>> we could group the blocks for nfree ranges, say [0-15], [16-31] and so on.
>> So after the alloc/free we'd calculate the new freelist index as (nfree/16)
>> and only moved the block if the index changed. This would reduce the
>> overhead to 1/16 and probably even more in practice.
>
>> Of course, we could also say we have e.g. 8 freelists and work the ranges
>> backwards from that, I guess that's what you propose.
>
> Yea, I was thinking something along those lines. As you say, ow about there's
> always at most 8 freelists or such. During initialization we compute a shift
> that distributes chunksPerBlock from 0 to 8. Then we only need to perform list
> manipulation if block->nfree >> slab->freelist_shift != --block->nfree >> slab->freelist_shift.
>
> That seems nice from a memory usage POV as well - for small and frequent
> allocations needing chunksPerBlock freelists isn't great. The freelists are on
> a fair number of cachelines right now.
>
Agreed.
>
>>> 4) Less of a performance, and more of a usability issue: The constant
>>> block size strikes me as problematic. Most users of an allocator can
>>> sometimes be used with a small amount of data, and sometimes with a
>>> large amount.
>>>
>>
>> I doubt this is worth the effort, really. The constant block size makes
>> various places much simpler (both to code and reason about), so this should
>> not make a huge difference in performance. And IMHO the block size is mostly
>> an implementation detail, so I don't see that as a usability issue.
>
> Hm? It's something the user has to specify, so I it's not really an
> implementation detail. It needs to be specified without sufficient
> information, as well, since externally one doesn't know how much memory the
> block header and chunk headers + rounding up will use, so computing a good
> block size isn't easy. I've wondered whether it should just be a count...
>
I think this is mixing two problems - how to specify the block size, and
whether the block size is constant (as in slab) or grows over time (as
in allocset).
As for specifying the block size, I agree maybe setting chunk count and
deriving the bytes from that might be easier to use, for the reasons you
mentioned.
But growing the block size seems problematic for long-lived contexts
with workloads that change a lot over time - imagine e.g. decoding many
small transactions, with one huge transaction mixed in. The one huge
transaction will grow the block size, and we'll keep using it forever.
But in that case we might have just as well allocate the large blocks
from the start, I guess.
> Why do you not think it's relevant for performance? Either one causes too much
> memory usage by using a too large block size, wasting memory, or one ends up
> loosing perf through frequent allocations?
>
True. I simply would not expect this to make a huge difference - I may
be wrong, and I'm sure there are workloads where it matters. But I still
think it's easier to just use larger blocks than to make the slab code
more complex.
regards
--
Tomas Vondra
EnterpriseDB: http://www.enterprisedb.com
The Enterprise PostgreSQL Company
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
@ 2021-07-17 23:10 ` Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
0 siblings, 1 reply; 29+ messages in thread
From: Andres Freund @ 2021-07-17 23:10 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@enterprisedb.com>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Hi,
On 2021-07-18 00:46:09 +0200, Tomas Vondra wrote:
> On 7/17/21 11:14 PM, Andres Freund wrote:
> > Hm. I wonder if we should just not populate the freelist eagerly, to
> > drive down the initialization cost. I.e. have a separate allocation path
> > for chunks that have never been allocated, by having a
> > SlabBlock->free_offset or such.
> >
> > Sure, it adds a branch to the allocation happy path, but it also makes the
> > first allocation for a chunk cheaper, because there's no need to get the next
> > element from the freelist (adding a likely cache miss). And it should make the
> > allocation of a new block faster by a lot.
> >
>
> Not sure what you mean by 'not populate eagerly' so can't comment :-(
Instead of populating a linked list with all chunks upon creation of a block -
which requires touching a fair bit of memory - keep a per-block pointer (or an
offset) into "unused" area of the block. When allocating from the block and
theres still "unused" memory left, use that, instead of bothering with the
freelist.
I tried that, and it nearly got slab up to the allocation/freeing performance
of aset.c (while winning after allocation, due to the higher memory density).
> > > > 4) Less of a performance, and more of a usability issue: The constant
> > > > block size strikes me as problematic. Most users of an allocator can
> > > > sometimes be used with a small amount of data, and sometimes with a
> > > > large amount.
> > > >
> > >
> > > I doubt this is worth the effort, really. The constant block size makes
> > > various places much simpler (both to code and reason about), so this should
> > > not make a huge difference in performance. And IMHO the block size is mostly
> > > an implementation detail, so I don't see that as a usability issue.
> >
> > Hm? It's something the user has to specify, so I it's not really an
> > implementation detail. It needs to be specified without sufficient
> > information, as well, since externally one doesn't know how much memory the
> > block header and chunk headers + rounding up will use, so computing a good
> > block size isn't easy. I've wondered whether it should just be a count...
> >
>
> I think this is mixing two problems - how to specify the block size, and
> whether the block size is constant (as in slab) or grows over time (as in
> allocset).
That was in response to the "implementation detail" bit solely.
> But growing the block size seems problematic for long-lived contexts with
> workloads that change a lot over time - imagine e.g. decoding many small
> transactions, with one huge transaction mixed in. The one huge transaction
> will grow the block size, and we'll keep using it forever. But in that case
> we might have just as well allocate the large blocks from the start, I
> guess.
I was thinking of capping the growth fairly low. I don't think after a 16x
growth or so you're likely to still see allocation performance gains with
slab. And I don't think that'd be too bad for decoding - we'd start with a
small initial block size, and in many workloads that will be enough, and just
workloads where that doesn't suffice will adapt performance wise. And: Medium
term I wouldn't expect reorderbuffer.c to stay the only slab.c user...
> > Why do you not think it's relevant for performance? Either one causes too much
> > memory usage by using a too large block size, wasting memory, or one ends up
> > loosing perf through frequent allocations?
>
> True. I simply would not expect this to make a huge difference - I may be
> wrong, and I'm sure there are workloads where it matters. But I still think
> it's easier to just use larger blocks than to make the slab code more
> complex.
IDK. I'm looking at using slab as part of a radix tree implementation right
now. Which I'd hope to be used in various different situations. So it's hard
to choose the right block size - and it does seem to matter for performance.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
@ 2021-07-18 01:06 ` Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
0 siblings, 1 reply; 29+ messages in thread
From: Andres Freund @ 2021-07-18 01:06 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@enterprisedb.com>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Hi,
On 2021-07-17 16:10:19 -0700, Andres Freund wrote:
> Instead of populating a linked list with all chunks upon creation of a block -
> which requires touching a fair bit of memory - keep a per-block pointer (or an
> offset) into "unused" area of the block. When allocating from the block and
> theres still "unused" memory left, use that, instead of bothering with the
> freelist.
>
> I tried that, and it nearly got slab up to the allocation/freeing performance
> of aset.c (while winning after allocation, due to the higher memory density).
Combining that with limiting the number of freelists, and some
microoptimizations, allocation performance is now on par.
Freeing still seems to be a tad slower, mostly because SlabFree()
practically is immediately stalled on fetching the block, whereas
AllocSetFree() can happily speculate ahead and do work like computing
the freelist index. And then aset only needs to access memory inside the
context - which is much more likely to be in cache than a freelist
inside a block (there are many more).
But that's ok, I think. It's close and it's only a small share of the
overall runtime of my workload...
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
@ 2021-07-18 17:23 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
0 siblings, 1 reply; 29+ messages in thread
From: Tomas Vondra @ 2021-07-18 17:23 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On 7/18/21 3:06 AM, Andres Freund wrote:
> Hi,
>
> On 2021-07-17 16:10:19 -0700, Andres Freund wrote:
>> Instead of populating a linked list with all chunks upon creation of a block -
>> which requires touching a fair bit of memory - keep a per-block pointer (or an
>> offset) into "unused" area of the block. When allocating from the block and
>> theres still "unused" memory left, use that, instead of bothering with the
>> freelist.
>>
>> I tried that, and it nearly got slab up to the allocation/freeing performance
>> of aset.c (while winning after allocation, due to the higher memory density).
>
> Combining that with limiting the number of freelists, and some
> microoptimizations, allocation performance is now on par.
>
> Freeing still seems to be a tad slower, mostly because SlabFree()
> practically is immediately stalled on fetching the block, whereas
> AllocSetFree() can happily speculate ahead and do work like computing
> the freelist index. And then aset only needs to access memory inside the
> context - which is much more likely to be in cache than a freelist
> inside a block (there are many more).
>
> But that's ok, I think. It's close and it's only a small share of the
> overall runtime of my workload...
>
Sounds great! Thanks for investigating this and for the improvements.
It might be good to do some experiments to see how the changes affect
memory consumption for practical workloads. I'm willing to spend soem
time on that, if needed.
regards
--
Tomas Vondra
EnterpriseDB: http://www.enterprisedb.com
The Enterprise PostgreSQL Company
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
@ 2021-07-19 20:56 ` Andres Freund <andres@anarazel.de>
2021-07-19 22:03 ` Re: slab allocator performance issues Ranier Vilela <ranier.vf@gmail.com>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
0 siblings, 2 replies; 29+ messages in thread
From: Andres Freund @ 2021-07-19 20:56 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@enterprisedb.com>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Hi,
On 2021-07-18 19:23:31 +0200, Tomas Vondra wrote:
> Sounds great! Thanks for investigating this and for the improvements.
>
> It might be good to do some experiments to see how the changes affect memory
> consumption for practical workloads. I'm willing to spend soem time on that,
> if needed.
I've attached my changes. They're in a rough shape right now, but I
think good enough for an initial look.
Greetings,
Andres Freund
Attachments:
[text/x-diff] 0001-WIP-optimize-allocations-by-separating-hot-from-cold.patch (33.2K, ../../20210719205603.lzh5rkuv7lsoq4mp@alap3.anarazel.de/2-0001-WIP-optimize-allocations-by-separating-hot-from-cold.patch)
download | inline diff:
From af4cd1f0b64cd52d7eab342493e3dfd6b0d8388e Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 19 Jul 2021 12:55:51 -0700
Subject: [PATCH 1/2] WIP: optimize allocations by separating hot from cold
paths.
---
src/include/nodes/memnodes.h | 4 +-
src/include/utils/memutils.h | 13 +
src/backend/utils/mmgr/aset.c | 537 ++++++++++++++--------------
src/backend/utils/mmgr/generation.c | 22 +-
src/backend/utils/mmgr/mcxt.c | 179 +++-------
src/backend/utils/mmgr/slab.c | 14 +-
6 files changed, 354 insertions(+), 415 deletions(-)
diff --git a/src/include/nodes/memnodes.h b/src/include/nodes/memnodes.h
index e6a757d6a07..8a42d2ff999 100644
--- a/src/include/nodes/memnodes.h
+++ b/src/include/nodes/memnodes.h
@@ -57,10 +57,10 @@ typedef void (*MemoryStatsPrintFunc) (MemoryContext context, void *passthru,
typedef struct MemoryContextMethods
{
- void *(*alloc) (MemoryContext context, Size size);
+ void *(*alloc) (MemoryContext context, Size size, int flags);
/* call this free_p in case someone #define's free() */
void (*free_p) (MemoryContext context, void *pointer);
- void *(*realloc) (MemoryContext context, void *pointer, Size size);
+ void *(*realloc) (MemoryContext context, void *pointer, Size size, int flags);
void (*reset) (MemoryContext context);
void (*delete_context) (MemoryContext context);
Size (*get_chunk_space) (MemoryContext context, void *pointer);
diff --git a/src/include/utils/memutils.h b/src/include/utils/memutils.h
index ff872274d44..2f75b4cca46 100644
--- a/src/include/utils/memutils.h
+++ b/src/include/utils/memutils.h
@@ -147,6 +147,19 @@ extern void MemoryContextCreate(MemoryContext node,
extern void HandleLogMemoryContextInterrupt(void);
extern void ProcessLogMemoryContextInterrupt(void);
+extern void *MemoryContextAllocationFailure(MemoryContext context, Size size, int flags);
+
+extern void MemoryContextSizeFailure(MemoryContext context, Size size, int flags) pg_attribute_noreturn();
+
+static inline void
+MemoryContextCheckSize(MemoryContext context, Size size, int flags)
+{
+ if (unlikely(!AllocSizeIsValid(size)))
+ {
+ if (!(flags & MCXT_ALLOC_HUGE) || !AllocHugeSizeIsValid(size))
+ MemoryContextSizeFailure(context, size, flags);
+ }
+}
/*
* Memory-context-type-specific functions
diff --git a/src/backend/utils/mmgr/aset.c b/src/backend/utils/mmgr/aset.c
index 77872e77bcd..00878354392 100644
--- a/src/backend/utils/mmgr/aset.c
+++ b/src/backend/utils/mmgr/aset.c
@@ -263,9 +263,9 @@ static AllocSetFreeList context_freelists[2] =
/*
* These functions implement the MemoryContext API for AllocSet contexts.
*/
-static void *AllocSetAlloc(MemoryContext context, Size size);
+static void *AllocSetAlloc(MemoryContext context, Size size, int flags);
static void AllocSetFree(MemoryContext context, void *pointer);
-static void *AllocSetRealloc(MemoryContext context, void *pointer, Size size);
+static void *AllocSetRealloc(MemoryContext context, void *pointer, Size size, int flags);
static void AllocSetReset(MemoryContext context);
static void AllocSetDelete(MemoryContext context);
static Size AllocSetGetChunkSpace(MemoryContext context, void *pointer);
@@ -704,266 +704,10 @@ AllocSetDelete(MemoryContext context)
free(set);
}
-/*
- * AllocSetAlloc
- * Returns pointer to allocated memory of given size or NULL if
- * request could not be completed; memory is added to the set.
- *
- * No request may exceed:
- * MAXALIGN_DOWN(SIZE_MAX) - ALLOC_BLOCKHDRSZ - ALLOC_CHUNKHDRSZ
- * All callers use a much-lower limit.
- *
- * Note: when using valgrind, it doesn't matter how the returned allocation
- * is marked, as mcxt.c will set it to UNDEFINED. In some paths we will
- * return space that is marked NOACCESS - AllocSetRealloc has to beware!
- */
-static void *
-AllocSetAlloc(MemoryContext context, Size size)
+static inline void *
+AllocSetAllocReturnChunk(AllocSet set, Size size, AllocChunk chunk, Size chunk_size)
{
- AllocSet set = (AllocSet) context;
- AllocBlock block;
- AllocChunk chunk;
- int fidx;
- Size chunk_size;
- Size blksize;
-
- AssertArg(AllocSetIsValid(set));
-
- /*
- * If requested size exceeds maximum for chunks, allocate an entire block
- * for this request.
- */
- if (size > set->allocChunkLimit)
- {
- chunk_size = MAXALIGN(size);
- blksize = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
- block = (AllocBlock) malloc(blksize);
- if (block == NULL)
- return NULL;
-
- context->mem_allocated += blksize;
-
- block->aset = set;
- block->freeptr = block->endptr = ((char *) block) + blksize;
-
- chunk = (AllocChunk) (((char *) block) + ALLOC_BLOCKHDRSZ);
- chunk->aset = set;
- chunk->size = chunk_size;
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = size;
- /* set mark to catch clobber of "unused" space */
- if (size < chunk_size)
- set_sentinel(AllocChunkGetPointer(chunk), size);
-#endif
-#ifdef RANDOMIZE_ALLOCATED_MEMORY
- /* fill the allocated space with junk */
- randomize_mem((char *) AllocChunkGetPointer(chunk), size);
-#endif
-
- /*
- * Stick the new block underneath the active allocation block, if any,
- * so that we don't lose the use of the space remaining therein.
- */
- if (set->blocks != NULL)
- {
- block->prev = set->blocks;
- block->next = set->blocks->next;
- if (block->next)
- block->next->prev = block;
- set->blocks->next = block;
- }
- else
- {
- block->prev = NULL;
- block->next = NULL;
- set->blocks = block;
- }
-
- /* Ensure any padding bytes are marked NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS((char *) AllocChunkGetPointer(chunk) + size,
- chunk_size - size);
-
- /* Disallow external access to private part of chunk header. */
- VALGRIND_MAKE_MEM_NOACCESS(chunk, ALLOCCHUNK_PRIVATE_LEN);
-
- return AllocChunkGetPointer(chunk);
- }
-
- /*
- * Request is small enough to be treated as a chunk. Look in the
- * corresponding free list to see if there is a free chunk we could reuse.
- * If one is found, remove it from the free list, make it again a member
- * of the alloc set and return its data address.
- */
- fidx = AllocSetFreeIndex(size);
- chunk = set->freelist[fidx];
- if (chunk != NULL)
- {
- Assert(chunk->size >= size);
-
- set->freelist[fidx] = (AllocChunk) chunk->aset;
-
- chunk->aset = (void *) set;
-
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = size;
- /* set mark to catch clobber of "unused" space */
- if (size < chunk->size)
- set_sentinel(AllocChunkGetPointer(chunk), size);
-#endif
-#ifdef RANDOMIZE_ALLOCATED_MEMORY
- /* fill the allocated space with junk */
- randomize_mem((char *) AllocChunkGetPointer(chunk), size);
-#endif
-
- /* Ensure any padding bytes are marked NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS((char *) AllocChunkGetPointer(chunk) + size,
- chunk->size - size);
-
- /* Disallow external access to private part of chunk header. */
- VALGRIND_MAKE_MEM_NOACCESS(chunk, ALLOCCHUNK_PRIVATE_LEN);
-
- return AllocChunkGetPointer(chunk);
- }
-
- /*
- * Choose the actual chunk size to allocate.
- */
- chunk_size = (1 << ALLOC_MINBITS) << fidx;
- Assert(chunk_size >= size);
-
- /*
- * If there is enough room in the active allocation block, we will put the
- * chunk into that block. Else must start a new one.
- */
- if ((block = set->blocks) != NULL)
- {
- Size availspace = block->endptr - block->freeptr;
-
- if (availspace < (chunk_size + ALLOC_CHUNKHDRSZ))
- {
- /*
- * The existing active (top) block does not have enough room for
- * the requested allocation, but it might still have a useful
- * amount of space in it. Once we push it down in the block list,
- * we'll never try to allocate more space from it. So, before we
- * do that, carve up its free space into chunks that we can put on
- * the set's freelists.
- *
- * Because we can only get here when there's less than
- * ALLOC_CHUNK_LIMIT left in the block, this loop cannot iterate
- * more than ALLOCSET_NUM_FREELISTS-1 times.
- */
- while (availspace >= ((1 << ALLOC_MINBITS) + ALLOC_CHUNKHDRSZ))
- {
- Size availchunk = availspace - ALLOC_CHUNKHDRSZ;
- int a_fidx = AllocSetFreeIndex(availchunk);
-
- /*
- * In most cases, we'll get back the index of the next larger
- * freelist than the one we need to put this chunk on. The
- * exception is when availchunk is exactly a power of 2.
- */
- if (availchunk != ((Size) 1 << (a_fidx + ALLOC_MINBITS)))
- {
- a_fidx--;
- Assert(a_fidx >= 0);
- availchunk = ((Size) 1 << (a_fidx + ALLOC_MINBITS));
- }
-
- chunk = (AllocChunk) (block->freeptr);
-
- /* Prepare to initialize the chunk header. */
- VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
-
- block->freeptr += (availchunk + ALLOC_CHUNKHDRSZ);
- availspace -= (availchunk + ALLOC_CHUNKHDRSZ);
-
- chunk->size = availchunk;
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = 0; /* mark it free */
-#endif
- chunk->aset = (void *) set->freelist[a_fidx];
- set->freelist[a_fidx] = chunk;
- }
-
- /* Mark that we need to create a new block */
- block = NULL;
- }
- }
-
- /*
- * Time to create a new regular (multi-chunk) block?
- */
- if (block == NULL)
- {
- Size required_size;
-
- /*
- * The first such block has size initBlockSize, and we double the
- * space in each succeeding block, but not more than maxBlockSize.
- */
- blksize = set->nextBlockSize;
- set->nextBlockSize <<= 1;
- if (set->nextBlockSize > set->maxBlockSize)
- set->nextBlockSize = set->maxBlockSize;
-
- /*
- * If initBlockSize is less than ALLOC_CHUNK_LIMIT, we could need more
- * space... but try to keep it a power of 2.
- */
- required_size = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
- while (blksize < required_size)
- blksize <<= 1;
-
- /* Try to allocate it */
- block = (AllocBlock) malloc(blksize);
-
- /*
- * We could be asking for pretty big blocks here, so cope if malloc
- * fails. But give up if there's less than 1 MB or so available...
- */
- while (block == NULL && blksize > 1024 * 1024)
- {
- blksize >>= 1;
- if (blksize < required_size)
- break;
- block = (AllocBlock) malloc(blksize);
- }
-
- if (block == NULL)
- return NULL;
-
- context->mem_allocated += blksize;
-
- block->aset = set;
- block->freeptr = ((char *) block) + ALLOC_BLOCKHDRSZ;
- block->endptr = ((char *) block) + blksize;
-
- /* Mark unallocated space NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS(block->freeptr,
- blksize - ALLOC_BLOCKHDRSZ);
-
- block->prev = NULL;
- block->next = set->blocks;
- if (block->next)
- block->next->prev = block;
- set->blocks = block;
- }
-
- /*
- * OK, do the allocation
- */
- chunk = (AllocChunk) (block->freeptr);
-
- /* Prepare to initialize the chunk header. */
- VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
-
- block->freeptr += (chunk_size + ALLOC_CHUNKHDRSZ);
- Assert(block->freeptr <= block->endptr);
-
chunk->aset = (void *) set;
- chunk->size = chunk_size;
#ifdef MEMORY_CONTEXT_CHECKING
chunk->requested_size = size;
/* set mark to catch clobber of "unused" space */
@@ -985,6 +729,273 @@ AllocSetAlloc(MemoryContext context, Size size)
return AllocChunkGetPointer(chunk);
}
+static void * pg_noinline
+AllocSetAllocLarge(AllocSet set, Size size, int flags)
+{
+ AllocBlock block;
+ AllocChunk chunk;
+ Size chunk_size;
+ Size blksize;
+
+ /* check size, only allocation path where the limits could be hit */
+ MemoryContextCheckSize(&set->header, size, flags);
+
+ AssertArg(AllocSetIsValid(set));
+
+ chunk_size = MAXALIGN(size);
+ blksize = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
+ block = (AllocBlock) malloc(blksize);
+ if (block == NULL)
+ return NULL;
+
+ set->header.mem_allocated += blksize;
+
+ block->aset = set;
+ block->freeptr = block->endptr = ((char *) block) + blksize;
+
+ /*
+ * Stick the new block underneath the active allocation block, if any,
+ * so that we don't lose the use of the space remaining therein.
+ */
+ if (set->blocks != NULL)
+ {
+ block->prev = set->blocks;
+ block->next = set->blocks->next;
+ if (block->next)
+ block->next->prev = block;
+ set->blocks->next = block;
+ }
+ else
+ {
+ block->prev = NULL;
+ block->next = NULL;
+ set->blocks = block;
+ }
+
+ chunk = (AllocChunk) (((char *) block) + ALLOC_BLOCKHDRSZ);
+ chunk->size = chunk_size;
+
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk_size);
+}
+
+static void * pg_noinline
+AllocSetAllocFromNewBlock(AllocSet set, Size size, Size chunk_size)
+{
+ AllocBlock block;
+ Size blksize;
+ Size required_size;
+ AllocChunk chunk;
+
+ /*
+ * The first such block has size initBlockSize, and we double the
+ * space in each succeeding block, but not more than maxBlockSize.
+ */
+ blksize = set->nextBlockSize;
+ set->nextBlockSize <<= 1;
+ if (set->nextBlockSize > set->maxBlockSize)
+ set->nextBlockSize = set->maxBlockSize;
+
+ /*
+ * If initBlockSize is less than ALLOC_CHUNK_LIMIT, we could need more
+ * space... but try to keep it a power of 2.
+ */
+ required_size = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
+ while (blksize < required_size)
+ blksize <<= 1;
+
+ /* Try to allocate it */
+ block = (AllocBlock) malloc(blksize);
+
+ /*
+ * We could be asking for pretty big blocks here, so cope if malloc
+ * fails. But give up if there's less than 1 MB or so available...
+ */
+ while (block == NULL && blksize > 1024 * 1024)
+ {
+ blksize >>= 1;
+ if (blksize < required_size)
+ break;
+ block = (AllocBlock) malloc(blksize);
+ }
+
+ if (block == NULL)
+ return NULL;
+
+ set->header.mem_allocated += blksize;
+
+ block->aset = set;
+ block->freeptr = ((char *) block) + ALLOC_BLOCKHDRSZ;
+ block->endptr = ((char *) block) + blksize;
+
+ /* Mark unallocated space NOACCESS. */
+ VALGRIND_MAKE_MEM_NOACCESS(block->freeptr,
+ blksize - ALLOC_BLOCKHDRSZ);
+
+ block->prev = NULL;
+ block->next = set->blocks;
+ if (block->next)
+ block->next->prev = block;
+ set->blocks = block;
+
+ /*
+ * OK, do the allocation
+ */
+ chunk = (AllocChunk) (block->freeptr);
+
+ /* Prepare to initialize the chunk header. */
+ VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
+
+ block->freeptr += (chunk_size + ALLOC_CHUNKHDRSZ);
+ Assert(block->freeptr <= block->endptr);
+
+ chunk->size = chunk_size;
+
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk_size);
+}
+
+static void * pg_noinline
+AllocSetAllocCarveOldAndAlloc(AllocSet set, Size size, Size chunk_size, AllocBlock block, Size availspace)
+{
+ AllocChunk chunk;
+
+ /*
+ * The existing active (top) block does not have enough room for
+ * the requested allocation, but it might still have a useful
+ * amount of space in it. Once we push it down in the block list,
+ * we'll never try to allocate more space from it. So, before we
+ * do that, carve up its free space into chunks that we can put on
+ * the set's freelists.
+ *
+ * Because we can only get here when there's less than
+ * ALLOC_CHUNK_LIMIT left in the block, this loop cannot iterate
+ * more than ALLOCSET_NUM_FREELISTS-1 times.
+ */
+ while (availspace >= ((1 << ALLOC_MINBITS) + ALLOC_CHUNKHDRSZ))
+ {
+ Size availchunk = availspace - ALLOC_CHUNKHDRSZ;
+ int a_fidx = AllocSetFreeIndex(availchunk);
+
+ /*
+ * In most cases, we'll get back the index of the next larger
+ * freelist than the one we need to put this chunk on. The
+ * exception is when availchunk is exactly a power of 2.
+ */
+ if (availchunk != ((Size) 1 << (a_fidx + ALLOC_MINBITS)))
+ {
+ a_fidx--;
+ Assert(a_fidx >= 0);
+ availchunk = ((Size) 1 << (a_fidx + ALLOC_MINBITS));
+ }
+
+ chunk = (AllocChunk) (block->freeptr);
+
+ /* Prepare to initialize the chunk header. */
+ VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
+
+ block->freeptr += (availchunk + ALLOC_CHUNKHDRSZ);
+ availspace -= (availchunk + ALLOC_CHUNKHDRSZ);
+
+ chunk->size = availchunk;
+#ifdef MEMORY_CONTEXT_CHECKING
+ chunk->requested_size = 0; /* mark it free */
+#endif
+ chunk->aset = (void *) set->freelist[a_fidx];
+ set->freelist[a_fidx] = chunk;
+ }
+
+ return AllocSetAllocFromNewBlock(set, size, chunk_size);
+}
+
+/*
+ * AllocSetAlloc
+ * Returns pointer to allocated memory of given size or NULL if
+ * request could not be completed; memory is added to the set.
+ *
+ * No request may exceed:
+ * MAXALIGN_DOWN(SIZE_MAX) - ALLOC_BLOCKHDRSZ - ALLOC_CHUNKHDRSZ
+ * All callers use a much-lower limit.
+ *
+ * Note: when using valgrind, it doesn't matter how the returned allocation
+ * is marked, as mcxt.c will set it to UNDEFINED. In some paths we will
+ * return space that is marked NOACCESS - AllocSetRealloc has to beware!
+ */
+static void *
+AllocSetAlloc(MemoryContext context, Size size, int flags)
+{
+ AllocSet set = (AllocSet) context;
+ AllocBlock block;
+ AllocChunk chunk;
+ int fidx;
+ Size chunk_size;
+
+ AssertArg(AllocSetIsValid(set));
+
+ /*
+ * If requested size exceeds maximum for chunks, allocate an entire block
+ * for this request.
+ */
+ if (unlikely(size > set->allocChunkLimit))
+ return AllocSetAllocLarge(set, size, flags);
+
+ /*
+ * Request is small enough to be treated as a chunk. Look in the
+ * corresponding free list to see if there is a free chunk we could reuse.
+ * If one is found, remove it from the free list, make it again a member
+ * of the alloc set and return its data address.
+ */
+ fidx = AllocSetFreeIndex(size);
+ chunk = set->freelist[fidx];
+ if (chunk != NULL)
+ {
+ Assert(chunk->size >= size);
+
+ set->freelist[fidx] = (AllocChunk) chunk->aset;
+
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk->size);
+ }
+
+ /*
+ * Choose the actual chunk size to allocate.
+ */
+ chunk_size = (1 << ALLOC_MINBITS) << fidx;
+ Assert(chunk_size >= size);
+
+ /*
+ * If there is enough room in the active allocation block, we will put the
+ * chunk into that block. Else must start a new one.
+ */
+ if ((block = set->blocks) != NULL)
+ {
+ Size availspace = block->endptr - block->freeptr;
+
+ if (unlikely(availspace < (chunk_size + ALLOC_CHUNKHDRSZ)))
+ return AllocSetAllocCarveOldAndAlloc(set, size, chunk_size,
+ block, availspace);
+ }
+ else if (unlikely(block == NULL))
+ {
+ /*
+ * Time to create a new regular (multi-chunk) block.
+ */
+ return AllocSetAllocFromNewBlock(set, size, chunk_size);
+ }
+
+ /*
+ * OK, do the allocation
+ */
+ chunk = (AllocChunk) (block->freeptr);
+
+ /* Prepare to initialize the chunk header. */
+ VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
+
+ chunk->size = chunk_size;
+
+ block->freeptr += (chunk_size + ALLOC_CHUNKHDRSZ);
+ Assert(block->freeptr <= block->endptr);
+
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk_size);
+}
+
/*
* AllocSetFree
* Frees allocated memory; memory is removed from the set.
@@ -1072,7 +1083,7 @@ AllocSetFree(MemoryContext context, void *pointer)
* request size.)
*/
static void *
-AllocSetRealloc(MemoryContext context, void *pointer, Size size)
+AllocSetRealloc(MemoryContext context, void *pointer, Size size, int flags)
{
AllocSet set = (AllocSet) context;
AllocChunk chunk = AllocPointerGetChunk(pointer);
@@ -1081,6 +1092,8 @@ AllocSetRealloc(MemoryContext context, void *pointer, Size size)
/* Allow access to private part of chunk header. */
VALGRIND_MAKE_MEM_DEFINED(chunk, ALLOCCHUNK_PRIVATE_LEN);
+ MemoryContextCheckSize(context, size, flags);
+
oldsize = chunk->size;
#ifdef MEMORY_CONTEXT_CHECKING
@@ -1260,7 +1273,7 @@ AllocSetRealloc(MemoryContext context, void *pointer, Size size)
AllocPointer newPointer;
/* allocate new chunk */
- newPointer = AllocSetAlloc((MemoryContext) set, size);
+ newPointer = AllocSetAlloc((MemoryContext) set, size, flags);
/* leave immediately if request was not completed */
if (newPointer == NULL)
diff --git a/src/backend/utils/mmgr/generation.c b/src/backend/utils/mmgr/generation.c
index 584cd614da0..5dde15654d7 100644
--- a/src/backend/utils/mmgr/generation.c
+++ b/src/backend/utils/mmgr/generation.c
@@ -146,9 +146,9 @@ struct GenerationChunk
/*
* These functions implement the MemoryContext API for Generation contexts.
*/
-static void *GenerationAlloc(MemoryContext context, Size size);
+static void *GenerationAlloc(MemoryContext context, Size size, int flags);
static void GenerationFree(MemoryContext context, void *pointer);
-static void *GenerationRealloc(MemoryContext context, void *pointer, Size size);
+static void *GenerationRealloc(MemoryContext context, void *pointer, Size size, int flags);
static void GenerationReset(MemoryContext context);
static void GenerationDelete(MemoryContext context);
static Size GenerationGetChunkSpace(MemoryContext context, void *pointer);
@@ -323,7 +323,7 @@ GenerationDelete(MemoryContext context)
* return space that is marked NOACCESS - GenerationRealloc has to beware!
*/
static void *
-GenerationAlloc(MemoryContext context, Size size)
+GenerationAlloc(MemoryContext context, Size size, int flags)
{
GenerationContext *set = (GenerationContext *) context;
GenerationBlock *block;
@@ -335,9 +335,12 @@ GenerationAlloc(MemoryContext context, Size size)
{
Size blksize = chunk_size + Generation_BLOCKHDRSZ + Generation_CHUNKHDRSZ;
+ /* check size, only allocation path where the limits could be hit */
+ MemoryContextCheckSize(context, size, flags);
+
block = (GenerationBlock *) malloc(blksize);
if (block == NULL)
- return NULL;
+ return MemoryContextAllocationFailure(context, size, flags);
context->mem_allocated += blksize;
@@ -392,7 +395,7 @@ GenerationAlloc(MemoryContext context, Size size)
block = (GenerationBlock *) malloc(blksize);
if (block == NULL)
- return NULL;
+ return MemoryContextAllocationFailure(context, size, flags);
context->mem_allocated += blksize;
@@ -520,13 +523,15 @@ GenerationFree(MemoryContext context, void *pointer)
* into the old chunk - in that case we just update chunk header.
*/
static void *
-GenerationRealloc(MemoryContext context, void *pointer, Size size)
+GenerationRealloc(MemoryContext context, void *pointer, Size size, int flags)
{
GenerationContext *set = (GenerationContext *) context;
GenerationChunk *chunk = GenerationPointerGetChunk(pointer);
GenerationPointer newPointer;
Size oldsize;
+ MemoryContextCheckSize(context, size, flags);
+
/* Allow access to private part of chunk header. */
VALGRIND_MAKE_MEM_DEFINED(chunk, GENERATIONCHUNK_PRIVATE_LEN);
@@ -596,14 +601,15 @@ GenerationRealloc(MemoryContext context, void *pointer, Size size)
}
/* allocate new chunk */
- newPointer = GenerationAlloc((MemoryContext) set, size);
+ newPointer = GenerationAlloc((MemoryContext) set, size, flags);
/* leave immediately if request was not completed */
if (newPointer == NULL)
{
/* Disallow external access to private part of chunk header. */
VALGRIND_MAKE_MEM_NOACCESS(chunk, GENERATIONCHUNK_PRIVATE_LEN);
- return NULL;
+ /* again? */
+ return MemoryContextAllocationFailure(context, size, flags);
}
/*
diff --git a/src/backend/utils/mmgr/mcxt.c b/src/backend/utils/mmgr/mcxt.c
index 6919a732804..13125c4a562 100644
--- a/src/backend/utils/mmgr/mcxt.c
+++ b/src/backend/utils/mmgr/mcxt.c
@@ -867,28 +867,9 @@ MemoryContextAlloc(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
-
- /*
- * Here, and elsewhere in this module, we show the target context's
- * "name" but not its "ident" (if any) in user-visible error messages.
- * The "ident" string might contain security-sensitive data, such as
- * values in SQL commands.
- */
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -910,21 +891,10 @@ MemoryContextAllocZero(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -948,21 +918,10 @@ MemoryContextAllocZeroAligned(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -983,26 +942,11 @@ MemoryContextAllocExtended(MemoryContext context, Size size, int flags)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (((flags & MCXT_ALLOC_HUGE) != 0 && !AllocHugeSizeIsValid(size)) ||
- ((flags & MCXT_ALLOC_HUGE) == 0 && !AllocSizeIsValid(size)))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
+ ret = context->methods->alloc(context, size, flags);
if (unlikely(ret == NULL))
- {
- if ((flags & MCXT_ALLOC_NO_OOM) == 0)
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
return NULL;
- }
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1067,22 +1011,10 @@ palloc(Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
-
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1099,21 +1031,10 @@ palloc0(Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1132,26 +1053,11 @@ palloc_extended(Size size, int flags)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (((flags & MCXT_ALLOC_HUGE) != 0 && !AllocHugeSizeIsValid(size)) ||
- ((flags & MCXT_ALLOC_HUGE) == 0 && !AllocSizeIsValid(size)))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
+ ret = context->methods->alloc(context, size, 0);
if (unlikely(ret == NULL))
- {
- if ((flags & MCXT_ALLOC_NO_OOM) == 0)
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
return NULL;
- }
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1184,24 +1090,13 @@ repalloc(void *pointer, Size size)
MemoryContext context = GetMemoryChunkContext(pointer);
void *ret;
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
AssertNotInCriticalSection(context);
/* isReset must be false already */
Assert(!context->isReset);
- ret = context->methods->realloc(context, pointer, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->realloc(context, pointer, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_CHANGE(context, pointer, ret, size);
@@ -1222,21 +1117,9 @@ MemoryContextAllocHuge(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocHugeSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
-
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, MCXT_ALLOC_HUGE);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1254,16 +1137,23 @@ repalloc_huge(void *pointer, Size size)
MemoryContext context = GetMemoryChunkContext(pointer);
void *ret;
- if (!AllocHugeSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
AssertNotInCriticalSection(context);
/* isReset must be false already */
Assert(!context->isReset);
- ret = context->methods->realloc(context, pointer, size);
- if (unlikely(ret == NULL))
+ ret = context->methods->realloc(context, pointer, size, MCXT_ALLOC_HUGE);
+ Assert(ret != NULL);
+
+ VALGRIND_MEMPOOL_CHANGE(context, pointer, ret, size);
+
+ return ret;
+}
+
+void *
+MemoryContextAllocationFailure(MemoryContext context, Size size, int flags)
+{
+ if ((flags & MCXT_ALLOC_NO_OOM) == 0)
{
MemoryContextStats(TopMemoryContext);
ereport(ERROR,
@@ -1273,9 +1163,13 @@ repalloc_huge(void *pointer, Size size)
size, context->name)));
}
- VALGRIND_MEMPOOL_CHANGE(context, pointer, ret, size);
+ return NULL;
+}
- return ret;
+void
+MemoryContextSizeFailure(MemoryContext context, Size size, int flags)
+{
+ elog(ERROR, "invalid memory alloc request size %zu", size);
}
/*
@@ -1298,6 +1192,13 @@ MemoryContextStrdup(MemoryContext context, const char *string)
char *
pstrdup(const char *in)
{
+ /*
+ * Here, and elsewhere in this module, we show the target context's
+ * "name" but not its "ident" (if any) in user-visible error messages.
+ * The "ident" string might contain security-sensitive data, such as
+ * values in SQL commands.
+ */
+
return MemoryContextStrdup(CurrentMemoryContext, in);
}
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index 553dd7f6674..e4b8275045f 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -126,9 +126,9 @@ typedef struct SlabChunk
/*
* These functions implement the MemoryContext API for Slab contexts.
*/
-static void *SlabAlloc(MemoryContext context, Size size);
+static void *SlabAlloc(MemoryContext context, Size size, int flags);
static void SlabFree(MemoryContext context, void *pointer);
-static void *SlabRealloc(MemoryContext context, void *pointer, Size size);
+static void *SlabRealloc(MemoryContext context, void *pointer, Size size, int flags);
static void SlabReset(MemoryContext context);
static void SlabDelete(MemoryContext context);
static Size SlabGetChunkSpace(MemoryContext context, void *pointer);
@@ -337,7 +337,7 @@ SlabDelete(MemoryContext context)
* request could not be completed; memory is added to the slab.
*/
static void *
-SlabAlloc(MemoryContext context, Size size)
+SlabAlloc(MemoryContext context, Size size, int flags)
{
SlabContext *slab = castNode(SlabContext, context);
SlabBlock *block;
@@ -346,6 +346,12 @@ SlabAlloc(MemoryContext context, Size size)
Assert(slab);
+ /*
+ * XXX: Probably no need to check for huge allocations, we only support
+ * one size? Which could theoretically be huge, but that'd not make
+ * sense...
+ */
+
Assert((slab->minFreeChunks >= 0) &&
(slab->minFreeChunks < slab->chunksPerBlock));
@@ -583,7 +589,7 @@ SlabFree(MemoryContext context, void *pointer)
* realloc is usually used to enlarge the chunk.
*/
static void *
-SlabRealloc(MemoryContext context, void *pointer, Size size)
+SlabRealloc(MemoryContext context, void *pointer, Size size, int flags)
{
SlabContext *slab = castNode(SlabContext, context);
--
2.32.0.rc2
[text/x-diff] 0002-WIP-slab-performance.patch (23.6K, ../../20210719205603.lzh5rkuv7lsoq4mp@alap3.anarazel.de/3-0002-WIP-slab-performance.patch)
download | inline diff:
From 8331119ab405eb8e29cb7a76ea8bfbf7a8914d68 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 19 Jul 2021 13:48:14 -0700
Subject: [PATCH 2/2] WIP: slab performance.
Discussion: https://postgr.es/m/20210717194333.mr5io3zup3kxahfm@alap3.anarazel.de
---
src/backend/utils/mmgr/slab.c | 454 ++++++++++++++++++++++------------
1 file changed, 295 insertions(+), 159 deletions(-)
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index e4b8275045f..c979553dcdc 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -23,15 +23,18 @@
* global (context) level. This is possible as the chunk size (and thus also
* the number of chunks per block) is fixed.
*
- * On each block, free chunks are tracked in a simple linked list. Contents
- * of free chunks is replaced with an index of the next free chunk, forming
- * a very simple linked list. Each block also contains a counter of free
- * chunks. Combined with the local block-level freelist, it makes it trivial
- * to eventually free the whole block.
+ * On each block, never allocated chunks are tracked by a simple offset, and
+ * free chunks are tracked in a simple linked list. The offset approach
+ * avoids needing to iterate over all chunks when allocating a new block,
+ * which would cause page faults and cache pollution. Contents of free chunks
+ * is replaced with a pointer to the next free chunk, forming a very simple
+ * linked list. Each block also contains a counter of free chunks. Combined
+ * with the local block-level freelist, it makes it trivial to eventually
+ * free the whole block.
*
* At the context level, we use 'freelist' to track blocks ordered by number
* of free chunks, starting with blocks having a single allocated chunk, and
- * with completely full blocks on the tail.
+ * with completely full blocks on the tail. XXX
*
* This also allows various optimizations - for example when searching for
* free chunk, the allocator reuses space from the fullest blocks first, in
@@ -44,7 +47,7 @@
* case this performs as if the pointer was not maintained.
*
* We cache the freelist index for the blocks with the fewest free chunks
- * (minFreeChunks), so that we don't have to search the freelist on every
+ * (minFreeChunkIndex), so that we don't have to search the freelist on every
* SlabAlloc() call, which is quite expensive.
*
*-------------------------------------------------------------------------
@@ -56,6 +59,17 @@
#include "utils/memdebug.h"
#include "utils/memutils.h"
+struct SlabBlock;
+struct SlabChunk;
+
+/*
+ * Number of actual freelists + 1, for full blocks. Full blocks are always at
+ * offset 0.
+ */
+#define SLAB_FREELIST_COUNT 9
+
+#define SLAB_RETAIN_EMPTY_BLOCK_COUNT 10
+
/*
* SlabContext is a specialized implementation of MemoryContext.
*/
@@ -68,13 +82,16 @@ typedef struct SlabContext
Size blockSize; /* block size */
Size headerSize; /* allocated size of context header */
int chunksPerBlock; /* number of chunks per block */
- int minFreeChunks; /* min number of free chunks in any block */
+ int minFreeChunksIndex; /* min number of free chunks in any block XXX */
int nblocks; /* number of blocks allocated */
#ifdef MEMORY_CONTEXT_CHECKING
bool *freechunks; /* bitmap of free chunks in a block */
#endif
+ int nfreeblocks;
+ dlist_head freeblocks;
+ int freelist_shift;
/* blocks with free space, grouped by number of free chunks: */
- dlist_head freelist[FLEXIBLE_ARRAY_MEMBER];
+ dlist_head freelist[SLAB_FREELIST_COUNT];
} SlabContext;
/*
@@ -83,13 +100,15 @@ typedef struct SlabContext
*
* node: doubly-linked list of blocks in global freelist
* nfree: number of free chunks in this block
- * firstFreeChunk: index of the first free chunk
+ * firstFreeChunk: first free chunk
*/
typedef struct SlabBlock
{
dlist_node node; /* doubly-linked list */
- int nfree; /* number of free chunks */
- int firstFreeChunk; /* index of the first free chunk in the block */
+ int nfree; /* number of chunks on freelist + unused */
+ int nunused; /* number of unused chunks */
+ struct SlabChunk *unused; /* */
+ struct SlabChunk *firstFreeChunk; /* first free chunk in the block */
} SlabBlock;
/*
@@ -123,6 +142,35 @@ typedef struct SlabChunk
#define SlabChunkIndex(slab, block, chunk) \
(((char *) chunk - SlabBlockStart(block)) / slab->fullChunkSize)
+static inline uint8
+SlabFreelistIndex(SlabContext *slab, int freecount)
+{
+ uint8 index;
+
+ Assert(freecount <= slab->chunksPerBlock);
+
+ /*
+ * Ensure, without a branch, that index 0 is only used for blocks entirely
+ * without free chunks.
+ * XXX: There probably is a cheaper way to do this. Needing to shift twice
+ * by slab->freelist_shift isn't great.
+ */
+ index = (freecount + (1 << slab->freelist_shift) - 1) >> slab->freelist_shift;
+
+ if (freecount == 0)
+ Assert(index == 0);
+ else
+ Assert(index > 0 && index < (SLAB_FREELIST_COUNT));
+
+ return index;
+}
+
+static inline dlist_head*
+SlabFreelist(SlabContext *slab, int freecount)
+{
+ return &slab->freelist[SlabFreelistIndex(slab, freecount)];
+}
+
/*
* These functions implement the MemoryContext API for Slab contexts.
*/
@@ -179,7 +227,6 @@ SlabContextCreate(MemoryContext parent,
{
int chunksPerBlock;
Size fullChunkSize;
- Size freelistSize;
Size headerSize;
SlabContext *slab;
int i;
@@ -192,11 +239,11 @@ SlabContextCreate(MemoryContext parent,
"padding calculation in SlabChunk is wrong");
/* Make sure the linked list node fits inside a freed chunk */
- if (chunkSize < sizeof(int))
- chunkSize = sizeof(int);
+ if (chunkSize < MAXALIGN(sizeof(void *)))
+ chunkSize = MAXALIGN(sizeof(void *));
/* chunk, including SLAB header (both addresses nicely aligned) */
- fullChunkSize = sizeof(SlabChunk) + MAXALIGN(chunkSize);
+ fullChunkSize = sizeof(SlabChunk) + chunkSize;
/* Make sure the block can store at least one chunk. */
if (blockSize < fullChunkSize + sizeof(SlabBlock))
@@ -206,16 +253,14 @@ SlabContextCreate(MemoryContext parent,
/* Compute maximum number of chunks per block */
chunksPerBlock = (blockSize - sizeof(SlabBlock)) / fullChunkSize;
- /* The freelist starts with 0, ends with chunksPerBlock. */
- freelistSize = sizeof(dlist_head) * (chunksPerBlock + 1);
-
/*
* Allocate the context header. Unlike aset.c, we never try to combine
* this with the first regular block; not worth the extra complication.
+ * XXX: What's the evidence for that?
*/
/* Size of the memory context header */
- headerSize = offsetof(SlabContext, freelist) + freelistSize;
+ headerSize = sizeof(SlabContext);
#ifdef MEMORY_CONTEXT_CHECKING
@@ -249,17 +294,34 @@ SlabContextCreate(MemoryContext parent,
slab->blockSize = blockSize;
slab->headerSize = headerSize;
slab->chunksPerBlock = chunksPerBlock;
- slab->minFreeChunks = 0;
+ slab->minFreeChunksIndex = 0;
slab->nblocks = 0;
+ slab->nfreeblocks = 0;
+
+ /*
+ * Compute a shift that guarantees that shifting chunksPerBlock with it
+ * yields is smaller than SLAB_FREELIST_COUNT - 1 (one freelist is used for full blocks).
+ */
+ slab->freelist_shift = 0;
+ while ((slab->chunksPerBlock >> slab->freelist_shift) >= (SLAB_FREELIST_COUNT - 1))
+ slab->freelist_shift++;
+
+#if 0
+ elog(LOG, "freelist shift for %d chunks of size %zu is %d, block size %zu",
+ slab->chunksPerBlock, slab->fullChunkSize, slab->freelist_shift,
+ slab->blockSize);
+#endif
+
+ dlist_init(&slab->freeblocks);
/* initialize the freelist slots */
- for (i = 0; i < (slab->chunksPerBlock + 1); i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
dlist_init(&slab->freelist[i]);
#ifdef MEMORY_CONTEXT_CHECKING
/* set the freechunks pointer right after the freelists array */
slab->freechunks
- = (bool *) slab + offsetof(SlabContext, freelist) + freelistSize;
+ = (bool *) slab + sizeof(SlabContext);
#endif
/* Finally, do the type-independent part of context creation */
@@ -292,8 +354,27 @@ SlabReset(MemoryContext context)
SlabCheck(context);
#endif
+ /* release retained empty blocks */
+ {
+ dlist_mutable_iter miter;
+
+ dlist_foreach_modify(miter, &slab->freeblocks)
+ {
+ SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
+
+ dlist_delete(miter.cur);
+
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ slab->nblocks--;
+ context->mem_allocated -= slab->blockSize;
+ }
+ }
+
/* walk over freelists and free the blocks */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
dlist_mutable_iter miter;
@@ -312,7 +393,7 @@ SlabReset(MemoryContext context)
}
}
- slab->minFreeChunks = 0;
+ slab->minFreeChunksIndex = 0;
Assert(slab->nblocks == 0);
Assert(context->mem_allocated == 0);
@@ -342,7 +423,6 @@ SlabAlloc(MemoryContext context, Size size, int flags)
SlabContext *slab = castNode(SlabContext, context);
SlabBlock *block;
SlabChunk *chunk;
- int idx;
Assert(slab);
@@ -352,8 +432,8 @@ SlabAlloc(MemoryContext context, Size size, int flags)
* sense...
*/
- Assert((slab->minFreeChunks >= 0) &&
- (slab->minFreeChunks < slab->chunksPerBlock));
+ Assert((slab->minFreeChunksIndex >= 0) &&
+ (slab->minFreeChunksIndex < SLAB_FREELIST_COUNT));
/* make sure we only allow correct request size */
if (size != slab->chunkSize)
@@ -367,25 +447,33 @@ SlabAlloc(MemoryContext context, Size size, int flags)
* slab->minFreeChunks == 0 means there are no blocks with free chunks,
* thanks to how minFreeChunks is updated at the end of SlabAlloc().
*/
- if (slab->minFreeChunks == 0)
+ if (unlikely(slab->minFreeChunksIndex == 0))
{
- block = (SlabBlock *) malloc(slab->blockSize);
+ if (slab->nfreeblocks > 0)
+ {
+ dlist_node *node;
- if (block == NULL)
- return NULL;
+ node = dlist_pop_head_node(&slab->freeblocks);
+ block = dlist_container(SlabBlock, node, node);
+ slab->nfreeblocks--;
+ }
+ else
+ {
+ block = (SlabBlock *) malloc(slab->blockSize);
+
+ if (unlikely(block == NULL))
+ return NULL;
+
+ slab->nblocks += 1;
+ context->mem_allocated += slab->blockSize;
+ }
block->nfree = slab->chunksPerBlock;
- block->firstFreeChunk = 0;
+ block->firstFreeChunk = NULL;
+ block->nunused = slab->chunksPerBlock;
+ block->unused = (SlabChunk *) SlabBlockStart(block);
- /*
- * Put all the chunks on a freelist. Walk the chunks and point each
- * one to the next one.
- */
- for (idx = 0; idx < slab->chunksPerBlock; idx++)
- {
- chunk = SlabBlockGetChunk(slab, block, idx);
- *(int32 *) SlabChunkGetPointer(chunk) = (idx + 1);
- }
+ slab->minFreeChunksIndex = SlabFreelistIndex(slab, slab->chunksPerBlock);
/*
* And add it to the last freelist with all chunks empty.
@@ -393,32 +481,54 @@ SlabAlloc(MemoryContext context, Size size, int flags)
* We know there are no blocks in the freelist, otherwise we wouldn't
* need a new block.
*/
- Assert(dlist_is_empty(&slab->freelist[slab->chunksPerBlock]));
+ Assert(dlist_is_empty(SlabFreelist(slab, slab->chunksPerBlock)));
- dlist_push_head(&slab->freelist[slab->chunksPerBlock], &block->node);
+ dlist_push_head(SlabFreelist(slab, slab->chunksPerBlock),
+ &block->node);
- slab->minFreeChunks = slab->chunksPerBlock;
- slab->nblocks += 1;
- context->mem_allocated += slab->blockSize;
+ chunk = block->unused;
+ block->unused = (SlabChunk *)(((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
+ else
+ {
+ /* grab the block from the freelist */
+ block = dlist_head_element(SlabBlock, node,
+ &slab->freelist[slab->minFreeChunksIndex]);
+
+ /* make sure we actually got a valid block, with matching nfree */
+ Assert(block != NULL);
+ Assert(slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree));
+ Assert(block->nfree > 0);
+
+ /* we know index of the first free chunk in the block */
+ if (block->nunused > 0)
+ {
+ chunk = block->unused;
+ block->unused = (SlabChunk *)(((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
+ else
+ {
+ chunk = block->firstFreeChunk;
+
+ /*
+ * Remove the chunk from the freelist head. The index of the next free
+ * chunk is stored in the chunk itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(void*));
+ block->firstFreeChunk = *(SlabChunk **) SlabChunkGetPointer(chunk);
+ }
}
- /* grab the block from the freelist (even the new block is there) */
- block = dlist_head_element(SlabBlock, node,
- &slab->freelist[slab->minFreeChunks]);
+ /* make sure the chunk is in the block and that it's marked as empty (XXX?) */
+ Assert((char *) chunk >= SlabBlockStart(block));
+ Assert((char *) chunk < (((char *)block) + slab->blockSize));
- /* make sure we actually got a valid block, with matching nfree */
- Assert(block != NULL);
- Assert(slab->minFreeChunks == block->nfree);
- Assert(block->nfree > 0);
-
- /* we know index of the first free chunk in the block */
- idx = block->firstFreeChunk;
-
- /* make sure the chunk index is valid, and that it's marked as empty */
- Assert((idx >= 0) && (idx < slab->chunksPerBlock));
-
- /* compute the chunk location block start (after the block header) */
- chunk = SlabBlockGetChunk(slab, block, idx);
+ Assert(block->firstFreeChunk == NULL || (
+ (block->firstFreeChunk >= (SlabChunk *) SlabBlockStart(block)) &&
+ block->firstFreeChunk <= (SlabChunk *) (((char *)block) + slab->blockSize))
+ );
/*
* Update the block nfree count, and also the minFreeChunks as we've
@@ -426,48 +536,6 @@ SlabAlloc(MemoryContext context, Size size, int flags)
* (because that's how we chose the block).
*/
block->nfree--;
- slab->minFreeChunks = block->nfree;
-
- /*
- * Remove the chunk from the freelist head. The index of the next free
- * chunk is stored in the chunk itself.
- */
- VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(int32));
- block->firstFreeChunk = *(int32 *) SlabChunkGetPointer(chunk);
-
- Assert(block->firstFreeChunk >= 0);
- Assert(block->firstFreeChunk <= slab->chunksPerBlock);
-
- Assert((block->nfree != 0 &&
- block->firstFreeChunk < slab->chunksPerBlock) ||
- (block->nfree == 0 &&
- block->firstFreeChunk == slab->chunksPerBlock));
-
- /* move the whole block to the right place in the freelist */
- dlist_delete(&block->node);
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
-
- /*
- * And finally update minFreeChunks, i.e. the index to the block with the
- * lowest number of free chunks. We only need to do that when the block
- * got full (otherwise we know the current block is the right one). We'll
- * simply walk the freelist until we find a non-empty entry.
- */
- if (slab->minFreeChunks == 0)
- {
- for (idx = 1; idx <= slab->chunksPerBlock; idx++)
- {
- if (dlist_is_empty(&slab->freelist[idx]))
- continue;
-
- /* found a non-empty freelist */
- slab->minFreeChunks = idx;
- break;
- }
- }
-
- if (slab->minFreeChunks == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
/* Prepare to initialize the chunk header. */
VALGRIND_MAKE_MEM_UNDEFINED(chunk, sizeof(SlabChunk));
@@ -475,6 +543,45 @@ SlabAlloc(MemoryContext context, Size size, int flags)
chunk->block = block;
chunk->slab = slab;
+ Assert(slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree + 1));
+
+ /* move the whole block to the right place in the freelist */
+ if (unlikely(SlabFreelistIndex(slab, block->nfree) != slab->minFreeChunksIndex))
+ {
+ slab->minFreeChunksIndex = SlabFreelistIndex(slab, block->nfree);
+ dlist_delete(&block->node);
+ dlist_push_head(SlabFreelist(slab, block->nfree), &block->node);
+
+
+ /*
+ * And finally update minFreeChunks, i.e. the index to the block with the
+ * lowest number of free chunks. We only need to do that when the block
+ * got full (otherwise we know the current block is the right one). We'll
+ * simply walk the freelist until we find a non-empty entry.
+ */
+ if (slab->minFreeChunksIndex == 0)
+ {
+ for (int idx = 1; idx < SLAB_FREELIST_COUNT; idx++)
+ {
+ if (dlist_is_empty(&slab->freelist[idx]))
+ continue;
+
+ /* found a non-empty freelist */
+ slab->minFreeChunksIndex = idx;
+ break;
+ }
+ }
+ }
+
+#if 0
+ /*
+ * FIXME: I don't understand what this ever did? It should be unreachable
+ * I think?
+ */
+ if (slab->minFreeChunks == slab->chunksPerBlock)
+ slab->minFreeChunks = 0;
+#endif
+
#ifdef MEMORY_CONTEXT_CHECKING
/* slab mark to catch clobber of "unused" space */
if (slab->chunkSize < (slab->fullChunkSize - sizeof(SlabChunk)))
@@ -496,6 +603,61 @@ SlabAlloc(MemoryContext context, Size size, int flags)
return SlabChunkGetPointer(chunk);
}
+static void pg_noinline
+SlabFreeSlow(SlabContext *slab, SlabBlock *block)
+{
+ dlist_delete(&block->node);
+
+ /*
+ * See if we need to update the minFreeChunks field for the slab - we only
+ * need to do that if the block had that number of free chunks
+ * before we freed one. In that case, we check if there still are blocks
+ * in the original freelist and we either keep the current value (if there
+ * still are blocks) or increment it by one (the new block is still the
+ * one with minimum free chunks).
+ *
+ * The one exception is when the block will get completely free - in that
+ * case we will free it, se we can't use it for minFreeChunks. It however
+ * means there are no more blocks with free chunks.
+ */
+ if (slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree - 1))
+ {
+ /* Have we removed the last chunk from the freelist? */
+ if (dlist_is_empty(&slab->freelist[slab->minFreeChunksIndex]))
+ {
+ /* but if we made the block entirely free, we'll free it */
+ if (block->nfree == slab->chunksPerBlock)
+ slab->minFreeChunksIndex = 0;
+ else
+ slab->minFreeChunksIndex++;
+ }
+ }
+
+ /* If the block is now completely empty, free it (in a way). */
+ if (block->nfree == slab->chunksPerBlock)
+ {
+ /*
+ * To avoid constantly freeing/allocating blocks in bursty patterns
+ * (on most crucially cases of repeatedly allocating and freeing a
+ * single chunk), retain a small number of blocks.
+ */
+ if (slab->nfreeblocks < SLAB_RETAIN_EMPTY_BLOCK_COUNT)
+ {
+ dlist_push_head(&slab->freeblocks, &block->node);
+ slab->nfreeblocks++;
+ }
+ else
+ {
+ slab->nblocks--;
+ slab->header.mem_allocated -= slab->blockSize;
+ free(block);
+ }
+ }
+ else
+ dlist_push_head(SlabFreelist(slab, block->nfree), &block->node);
+
+}
+
/*
* SlabFree
* Frees allocated memory; memory is removed from the slab.
@@ -503,7 +665,6 @@ SlabAlloc(MemoryContext context, Size size, int flags)
static void
SlabFree(MemoryContext context, void *pointer)
{
- int idx;
SlabContext *slab = castNode(SlabContext, context);
SlabChunk *chunk = SlabPointerGetChunk(pointer);
SlabBlock *block = chunk->block;
@@ -516,61 +677,26 @@ SlabFree(MemoryContext context, void *pointer)
slab->header.name, chunk);
#endif
- /* compute index of the chunk with respect to block start */
- idx = SlabChunkIndex(slab, block, chunk);
-
/* add chunk to freelist, and update block nfree count */
- *(int32 *) pointer = block->firstFreeChunk;
- block->firstFreeChunk = idx;
+ *(SlabChunk **) pointer = block->firstFreeChunk;
+ block->firstFreeChunk = chunk;
block->nfree++;
Assert(block->nfree > 0);
Assert(block->nfree <= slab->chunksPerBlock);
#ifdef CLOBBER_FREED_MEMORY
- /* XXX don't wipe the int32 index, used for block-level freelist */
- wipe_mem((char *) pointer + sizeof(int32),
- slab->chunkSize - sizeof(int32));
+ /* XXX don't wipe the SlabChunk* index, used for block-level freelist */
+ wipe_mem((char *) pointer + sizeof(SlabChunk*),
+ slab->chunkSize - sizeof(SlabChunk*));
#endif
/* remove the block from a freelist */
- dlist_delete(&block->node);
-
- /*
- * See if we need to update the minFreeChunks field for the slab - we only
- * need to do that if there the block had that number of free chunks
- * before we freed one. In that case, we check if there still are blocks
- * in the original freelist and we either keep the current value (if there
- * still are blocks) or increment it by one (the new block is still the
- * one with minimum free chunks).
- *
- * The one exception is when the block will get completely free - in that
- * case we will free it, se we can't use it for minFreeChunks. It however
- * means there are no more blocks with free chunks.
- */
- if (slab->minFreeChunks == (block->nfree - 1))
+ if (SlabFreelistIndex(slab, block->nfree) != SlabFreelistIndex(slab, block->nfree - 1))
{
- /* Have we removed the last chunk from the freelist? */
- if (dlist_is_empty(&slab->freelist[slab->minFreeChunks]))
- {
- /* but if we made the block entirely free, we'll free it */
- if (block->nfree == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
- else
- slab->minFreeChunks++;
- }
+ SlabFreeSlow(slab, block);
}
- /* If the block is now completely empty, free it. */
- if (block->nfree == slab->chunksPerBlock)
- {
- free(block);
- slab->nblocks--;
- context->mem_allocated -= slab->blockSize;
- }
- else
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
-
Assert(slab->nblocks >= 0);
Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
}
@@ -657,7 +783,7 @@ SlabStats(MemoryContext context,
/* Include context header in totalspace */
totalspace = slab->headerSize;
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
dlist_iter iter;
@@ -714,7 +840,7 @@ SlabCheck(MemoryContext context)
Assert(slab->chunksPerBlock > 0);
/* walk all the freelists */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
int j,
nfree;
@@ -723,20 +849,19 @@ SlabCheck(MemoryContext context)
/* walk all blocks on this freelist */
dlist_foreach(iter, &slab->freelist[i])
{
- int idx;
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+ SlabChunk *cur_chunk;
/*
* Make sure the number of free chunks (in the block header)
* matches position in the freelist.
*/
- if (block->nfree != i)
+ if (SlabFreelistIndex(slab, block->nfree) != i)
elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match freelist %d",
name, block->nfree, block, i);
/* reset the bitmap of free chunks for this block */
memset(slab->freechunks, 0, (slab->chunksPerBlock * sizeof(bool)));
- idx = block->firstFreeChunk;
/*
* Now walk through the chunks, count the free ones and also
@@ -744,20 +869,31 @@ SlabCheck(MemoryContext context)
* freelist is stored within the chunks themselves, we have to
* walk through the chunks and construct our own bitmap.
*/
-
+ cur_chunk = block->firstFreeChunk;
nfree = 0;
- while (idx < slab->chunksPerBlock)
+ while (cur_chunk != NULL)
{
- SlabChunk *chunk;
+ int idx = SlabChunkIndex(slab, block, cur_chunk);
/* count the chunk as free, add it to the bitmap */
nfree++;
slab->freechunks[idx] = true;
/* read index of the next free chunk */
- chunk = SlabBlockGetChunk(slab, block, idx);
- VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(int32));
- idx = *(int32 *) SlabChunkGetPointer(chunk);
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(cur_chunk), sizeof(SlabChunk **));
+ cur_chunk = *(SlabChunk **) SlabChunkGetPointer(cur_chunk);
+ }
+
+ cur_chunk = block->unused;
+ for (j = 0; j < block->nunused; j++)
+ {
+ int idx = SlabChunkIndex(slab, block, cur_chunk);
+
+ /* count the chunk as free, add it to the bitmap */
+ nfree++;
+ slab->freechunks[idx] = true;
+
+ cur_chunk = (SlabChunk *)(((char *) cur_chunk) + slab->fullChunkSize);
}
for (j = 0; j < slab->chunksPerBlock; j++)
--
2.32.0.rc2
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
@ 2021-07-19 22:03 ` Ranier Vilela <ranier.vf@gmail.com>
2021-07-20 14:15 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
1 sibling, 1 reply; 29+ messages in thread
From: Ranier Vilela @ 2021-07-19 22:03 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; pgsql-hackers; Tomas Vondra <tv@fuzzy.cz>
Em seg., 19 de jul. de 2021 Ã s 17:56, Andres Freund <andres@anarazel.de>
escreveu:
> Hi,
>
> On 2021-07-18 19:23:31 +0200, Tomas Vondra wrote:
> > Sounds great! Thanks for investigating this and for the improvements.
> >
> > It might be good to do some experiments to see how the changes affect
> memory
> > consumption for practical workloads. I'm willing to spend soem time on
> that,
> > if needed.
>
> I've attached my changes. They're in a rough shape right now, but I
> think good enough for an initial look.
>
Hi Andres, I take a look.
Perhaps you would agree with me that in the most absolute of times, malloc
will not fail.
So it makes more sense to test:
if (ret != NULL)
than
if (ret == NULL)
What might help branch prediction.
With this change wins too, the possibility
to reduce the scope of some variable.
Example:
+static void * pg_noinline
+AllocSetAllocLarge(AllocSet set, Size size, int flags)
+{
+ AllocBlock block;
+ Size chunk_size;
+ Size blksize;
+
+ /* check size, only allocation path where the limits could be hit */
+ MemoryContextCheckSize(&set->header, size, flags);
+
+ AssertArg(AllocSetIsValid(set));
+
+ chunk_size = MAXALIGN(size);
+ blksize = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
+ block = (AllocBlock) malloc(blksize);
+ if (block != NULL)
+ {
+ AllocChunk chunk;
+
+ set->header.mem_allocated += blksize;
+
+ block->aset = set;
+ block->freeptr = block->endptr = ((char *) block) + blksize;
+
+ /*
+ * Stick the new block underneath the active allocation block, if
any,
+ * so that we don't lose the use of the space remaining therein.
+ */
+ if (set->blocks != NULL)
+ {
+ block->prev = set->blocks;
+ block->next = set->blocks->next;
+ if (block->next)
+ block->next->prev = block;
+ set->blocks->next = block;
+ }
+ else
+ {
+ block->prev = NULL;
+ block->next = NULL;
+ set->blocks = block;
+ }
+
+ chunk = (AllocChunk) (((char *) block) + ALLOC_BLOCKHDRSZ);
+ chunk->size = chunk_size;
+
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk_size);
+ }
+
+ return NULL;
+}
regards,
Ranier Vilela
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-19 22:03 ` Re: slab allocator performance issues Ranier Vilela <ranier.vf@gmail.com>
@ 2021-07-20 14:15 ` David Rowley <dgrowleyml@gmail.com>
2021-07-20 14:24 ` Re: slab allocator performance issues Ranier Vilela <ranier.vf@gmail.com>
0 siblings, 1 reply; 29+ messages in thread
From: David Rowley @ 2021-07-20 14:15 UTC (permalink / raw)
To: Ranier Vilela <ranier.vf@gmail.com>; +Cc: Andres Freund <andres@anarazel.de>; Tomas Vondra <tomas.vondra@enterprisedb.com>; pgsql-hackers; Tomas Vondra <tv@fuzzy.cz>
On Tue, 20 Jul 2021 at 10:04, Ranier Vilela <ranier.vf@gmail.com> wrote:
> Perhaps you would agree with me that in the most absolute of times, malloc will not fail.
> So it makes more sense to test:
> if (ret != NULL)
> than
> if (ret == NULL)
I think it'd be better to use unlikely() for that.
David
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-19 22:03 ` Re: slab allocator performance issues Ranier Vilela <ranier.vf@gmail.com>
2021-07-20 14:15 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
@ 2021-07-20 14:24 ` Ranier Vilela <ranier.vf@gmail.com>
0 siblings, 0 replies; 29+ messages in thread
From: Ranier Vilela @ 2021-07-20 14:24 UTC (permalink / raw)
To: David Rowley <dgrowleyml@gmail.com>; +Cc: Andres Freund <andres@anarazel.de>; Tomas Vondra <tomas.vondra@enterprisedb.com>; pgsql-hackers; Tomas Vondra <tv@fuzzy.cz>
Em ter., 20 de jul. de 2021 Ã s 11:15, David Rowley <dgrowleyml@gmail.com>
escreveu:
> On Tue, 20 Jul 2021 at 10:04, Ranier Vilela <ranier.vf@gmail.com> wrote:
> > Perhaps you would agree with me that in the most absolute of times,
> malloc will not fail.
> > So it makes more sense to test:
> > if (ret != NULL)
> > than
> > if (ret == NULL)
>
> I think it'd be better to use unlikely() for that.
>
Sure, it can be, but in this case, there is no way to reduce the scope.
regards,
Ranier Vilela
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
@ 2021-08-01 17:59 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
1 sibling, 1 reply; 29+ messages in thread
From: Tomas Vondra @ 2021-08-01 17:59 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Hi,
I spent a bit of time benchmarking this - the first patch adds an
extension with three functions, each executing a slightly different
allocation pattern:
1) FIFO (allocates and frees in the same order)
2) LIFO (frees in reverse order)
3) random
Each function can also do a custom number of iterations, each allocating
and freeing certain number of chunks. The bench.sql script executes
three combinations (for each pattern)
1) no loops
2) increase: 100 loops, each freeing 10k chunks and allocating 15k
3) decrease: 100 loops, each freeing 10k chunks and allocating 5k
The idea is to test simple one-time allocation, and workloads that mix
allocations and frees (to see how the changes affect reuse etc.).
The script tests this with a range of block sizes (1k-32k) and chunk
sizes (32B-512B).
In the attached .ods file with results, the "comparison" sheets are the
interesting ones - the last couple columns compare the main metrics for
the two patches (labeled patch-1 and patch-2) to master.
Overall, the results look quite good - patch-1 is mostly on par with
master, with maybe 5% variability in both directions. That's expected,
considering the patch does not aim to improve performance.
The second patch brings some nice improvements - 30%-50% in most cases
(for both allocation and free) seems pretty nice. But for the "increase"
FIFO pattern (incrementally allocating/freeing more memory) there's a
significant regression - particularly for the allocation time. In some
cases (larger chunks, block size does not matter too much) it jumps from
25ms to almost 200ms.
This seems unfortunate - the allocation pattern (FIFO, allocating more
memory over time) seems pretty common, and the slowdown is significant.
regards
--
Tomas Vondra
EnterpriseDB: http://www.enterprisedb.com
The Enterprise PostgreSQL Company
Attachments:
[text/x-patch] 0001-slab-bench.patch (16.3K, ../../6717ca90-a2ea-b17d-f544-37f5181b1175@enterprisedb.com/2-0001-slab-bench.patch)
download | inline diff:
From d458dd2996632cce68bc659ecdeedb989c5681b3 Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas.vondra@postgresql.org>
Date: Sat, 31 Jul 2021 22:55:36 +0200
Subject: [PATCH 1/3] slab bench
---
contrib/slab_bench/.gitignore | 4 +
contrib/slab_bench/Makefile | 21 ++
contrib/slab_bench/bench.sql | 30 ++
contrib/slab_bench/slab_bench--1.0.sql | 16 +
contrib/slab_bench/slab_bench.c | 405 +++++++++++++++++++++++++
contrib/slab_bench/slab_bench.control | 4 +
6 files changed, 480 insertions(+)
create mode 100644 contrib/slab_bench/.gitignore
create mode 100644 contrib/slab_bench/Makefile
create mode 100644 contrib/slab_bench/bench.sql
create mode 100644 contrib/slab_bench/slab_bench--1.0.sql
create mode 100644 contrib/slab_bench/slab_bench.c
create mode 100644 contrib/slab_bench/slab_bench.control
diff --git a/contrib/slab_bench/.gitignore b/contrib/slab_bench/.gitignore
new file mode 100644
index 0000000000..5dcb3ff972
--- /dev/null
+++ b/contrib/slab_bench/.gitignore
@@ -0,0 +1,4 @@
+# Generated subdirectories
+/log/
+/results/
+/tmp_check/
diff --git a/contrib/slab_bench/Makefile b/contrib/slab_bench/Makefile
new file mode 100644
index 0000000000..c423bbd4ca
--- /dev/null
+++ b/contrib/slab_bench/Makefile
@@ -0,0 +1,21 @@
+# contrib/slab_bench/Makefile
+
+MODULE_big = slab_bench
+OBJS = slab_bench.o
+
+EXTENSION = slab_bench
+DATA = slab_bench--1.0.sql
+PGFILEDESC = "slab_bench - slab context benchmarking functions"
+
+REGRESS = slab_bench
+
+ifdef USE_PGXS
+PG_CONFIG = pg_config
+PGXS := $(shell $(PG_CONFIG) --pgxs)
+include $(PGXS)
+else
+subdir = contrib/slab_bench
+top_builddir = ../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+endif
diff --git a/contrib/slab_bench/bench.sql b/contrib/slab_bench/bench.sql
new file mode 100644
index 0000000000..bcf6e05de0
--- /dev/null
+++ b/contrib/slab_bench/bench.sql
@@ -0,0 +1,30 @@
+CREATE EXTENSION slab_bench;
+
+\o fifo-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+\o lifo-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+\o random-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+
+\o fifo-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+\o lifo-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+\o random-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+
+\o fifo-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+\o lifo-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+\o random-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 100, 10000, 5000) x;
diff --git a/contrib/slab_bench/slab_bench--1.0.sql b/contrib/slab_bench/slab_bench--1.0.sql
new file mode 100644
index 0000000000..0a23134d74
--- /dev/null
+++ b/contrib/slab_bench/slab_bench--1.0.sql
@@ -0,0 +1,16 @@
+/* slab_bench--1.0.sql */
+
+-- complain if script is sourced in psql, rather than via CREATE EXTENSION
+\echo Use "CREATE EXTENSION slab_bench" to load this file. \quit
+
+CREATE FUNCTION slab_bench_random(nallocs bigint, block_size bigint, alloc_size bigint, loops int, free_cnt int, alloc_cnt int, out mem_allocated bigint, out alloc_ms bigint, out free_ms bigint)
+AS 'MODULE_PATHNAME', 'slab_bench_random'
+LANGUAGE C VOLATILE STRICT;
+
+CREATE FUNCTION slab_bench_fifo(nallocs bigint, block_size bigint, alloc_size bigint, loops int, free_cnt int, alloc_cnt int, out mem_allocated bigint, out alloc_ms bigint, out free_ms bigint)
+AS 'MODULE_PATHNAME', 'slab_bench_fifo'
+LANGUAGE C VOLATILE STRICT;
+
+CREATE FUNCTION slab_bench_lifo(nallocs bigint, block_size bigint, alloc_size bigint, loops int, free_cnt int, alloc_cnt int, out mem_allocated bigint, out alloc_ms bigint, out free_ms bigint)
+AS 'MODULE_PATHNAME', 'slab_bench_lifo'
+LANGUAGE C VOLATILE STRICT;
diff --git a/contrib/slab_bench/slab_bench.c b/contrib/slab_bench/slab_bench.c
new file mode 100644
index 0000000000..d2249ce4a2
--- /dev/null
+++ b/contrib/slab_bench/slab_bench.c
@@ -0,0 +1,405 @@
+/*-------------------------------------------------------------------------
+ *
+ * slab_bench.c
+ *
+ * helper functions to benchmark slab context with different workloads
+ *-------------------------------------------------------------------------
+ */
+#include "postgres.h"
+
+#include <sys/time.h>
+
+#include "funcapi.h"
+#include "miscadmin.h"
+
+PG_MODULE_MAGIC;
+
+PG_FUNCTION_INFO_V1(slab_bench_random);
+PG_FUNCTION_INFO_V1(slab_bench_fifo);
+PG_FUNCTION_INFO_V1(slab_bench_lifo);
+
+typedef struct Chunk {
+ int random;
+ void *ptr;
+} Chunk;
+
+static int
+chunk_index_cmp(const void *a, const void *b)
+{
+ Chunk *ca = (Chunk *) a;
+ Chunk *cb = (Chunk *) b;
+
+ if (ca->random < cb->random)
+ return -1;
+ else if (ca->random > cb->random)
+ return 1;
+
+ return 0;
+}
+
+Datum
+slab_bench_random(PG_FUNCTION_ARGS)
+{
+ MemoryContext cxt,
+ oldcxt;
+ Chunk *chunks;
+ int64 i, j;
+ int64 nallocs = PG_GETARG_INT64(0);
+ int64 blockSize = PG_GETARG_INT64(1);
+ int64 chunkSize = PG_GETARG_INT64(2);
+
+ int nloops = PG_GETARG_INT32(3);
+ int free_cnt = PG_GETARG_INT32(4);
+ int alloc_cnt = PG_GETARG_INT32(5);
+
+ struct timeval start_time,
+ end_time;
+ int64 alloc_time = 0,
+ free_time = 0;
+ int64 mem_allocated;
+
+ TupleDesc tupdesc;
+ Datum result;
+ HeapTuple tuple;
+ Datum values[9];
+ bool nulls[9];
+
+ int maxchunks;
+
+ maxchunks = nallocs + nloops * Max(0, alloc_cnt - free_cnt);
+
+ cxt = SlabContextCreate(CurrentMemoryContext, "slab_bench", blockSize, chunkSize);
+
+ chunks = (Chunk *) palloc(maxchunks * sizeof(Chunk));
+
+ /* allocate the chunks in random order */
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ gettimeofday(&start_time, NULL);
+
+ for (i = 0; i < nallocs; i++)
+ chunks[i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = MemoryContextMemAllocated(cxt, true);
+
+ /* do the requested number of free/alloc loops */
+ for (j = 0; j < nloops; j++)
+ {
+ /* randomize the indexes */
+ for (i = 0; i < nallocs; i++)
+ chunks[i].random = random();
+
+ qsort(chunks, nallocs, sizeof(Chunk), chunk_index_cmp);
+
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 0; i < free_cnt; i++)
+ pfree(chunks[i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs -= free_cnt;
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ memmove(chunks, &chunks[free_cnt], nallocs * sizeof(Chunk));
+
+
+ /* allocate alloc_cnt chunks at the end */
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 0; i < alloc_cnt; i++)
+ chunks[nallocs + i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs += alloc_cnt;
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = Max(mem_allocated, MemoryContextMemAllocated(cxt, true));
+ }
+
+ /* release the chunks in random order */
+ for (i = 0; i < nallocs; i++)
+ chunks[i].random = random();
+
+ qsort(chunks, nallocs, sizeof(Chunk), chunk_index_cmp);
+
+ gettimeofday(&start_time, NULL);
+
+ for (i = 0; i < nallocs; i++)
+ pfree(chunks[i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+
+ values[0] = Int64GetDatum(mem_allocated);
+ values[1] = Int64GetDatum(alloc_time);
+ values[2] = Int64GetDatum(free_time);
+
+ memset(nulls, 0, sizeof(nulls));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ result = HeapTupleGetDatum(tuple);
+
+ PG_RETURN_DATUM(result);
+}
+
+Datum
+slab_bench_fifo(PG_FUNCTION_ARGS)
+{
+ MemoryContext cxt,
+ oldcxt;
+ Chunk *chunks;
+ int64 i, j;
+ int64 nallocs = PG_GETARG_INT64(0);
+ int64 blockSize = PG_GETARG_INT64(1);
+ int64 chunkSize = PG_GETARG_INT64(2);
+
+ int nloops = PG_GETARG_INT32(3);
+ int free_cnt = PG_GETARG_INT32(4);
+ int alloc_cnt = PG_GETARG_INT32(5);
+
+ struct timeval start_time,
+ end_time;
+ int64 alloc_time = 0,
+ free_time = 0;
+ int64 mem_allocated;
+
+ TupleDesc tupdesc;
+ Datum result;
+ HeapTuple tuple;
+ Datum values[9];
+ bool nulls[9];
+
+ int maxchunks;
+
+ maxchunks = nallocs + nloops * Max(0, alloc_cnt - free_cnt);
+
+ cxt = SlabContextCreate(CurrentMemoryContext, "slab_bench", blockSize, chunkSize);
+
+ chunks = (Chunk *) palloc(maxchunks * sizeof(Chunk));
+
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ gettimeofday(&start_time, NULL);
+
+ for (i = 0; i < nallocs; i++)
+ chunks[i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = MemoryContextMemAllocated(cxt, true);
+
+
+ /* do the requested number of free/alloc loops */
+ for (j = 0; j < nloops; j++)
+ {
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 0; i < free_cnt; i++)
+ pfree(chunks[i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs -= free_cnt;
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ memmove(chunks, &chunks[free_cnt], nallocs * sizeof(Chunk));
+
+ /* allocate alloc_cnt chunks at the end */
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 0; i < alloc_cnt; i++)
+ chunks[nallocs + i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs += alloc_cnt;
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = Max(mem_allocated, MemoryContextMemAllocated(cxt, true));
+ }
+
+
+ gettimeofday(&start_time, NULL);
+
+ for (i = 0; i < nallocs; i++)
+ pfree(chunks[i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+
+ values[0] = Int64GetDatum(mem_allocated);
+ values[1] = Int64GetDatum(alloc_time);
+ values[2] = Int64GetDatum(free_time);
+
+ memset(nulls, 0, sizeof(nulls));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ result = HeapTupleGetDatum(tuple);
+
+ PG_RETURN_DATUM(result);
+}
+
+Datum
+slab_bench_lifo(PG_FUNCTION_ARGS)
+{
+ MemoryContext cxt,
+ oldcxt;
+ Chunk *chunks;
+ int64 i, j;
+ int64 nallocs = PG_GETARG_INT64(0);
+ int64 blockSize = PG_GETARG_INT64(1);
+ int64 chunkSize = PG_GETARG_INT64(2);
+
+ int nloops = PG_GETARG_INT32(3);
+ int free_cnt = PG_GETARG_INT32(4);
+ int alloc_cnt = PG_GETARG_INT32(5);
+
+ struct timeval start_time,
+ end_time;
+ int64 alloc_time = 0,
+ free_time = 0;
+ int64 mem_allocated;
+
+ TupleDesc tupdesc;
+ Datum result;
+ HeapTuple tuple;
+ Datum values[9];
+ bool nulls[9];
+
+ int maxchunks;
+
+ maxchunks = nallocs + nloops * Max(0, alloc_cnt - free_cnt);
+
+ cxt = SlabContextCreate(CurrentMemoryContext, "slab_bench", blockSize, chunkSize);
+
+ chunks = (Chunk *) palloc(maxchunks * sizeof(Chunk));
+
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ /* palloc benchmark */
+ gettimeofday(&start_time, NULL);
+
+ for (i = 0; i < nallocs; i++)
+ chunks[i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = MemoryContextMemAllocated(cxt, true);
+
+
+ /* do the requested number of free/alloc loops */
+ for (j = 0; j < nloops; j++)
+ {
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 1; i <= free_cnt; i++)
+ pfree(chunks[nallocs - i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs -= free_cnt;
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ /* allocate alloc_cnt chunks at the end */
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 0; i < alloc_cnt; i++)
+ chunks[nallocs + i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs += alloc_cnt;
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = Max(mem_allocated, MemoryContextMemAllocated(cxt, true));
+ }
+
+
+ gettimeofday(&start_time, NULL);
+
+ for (i = (nallocs - 1); i >= 0; i--)
+ pfree(chunks[i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000L +
+ (end_time.tv_usec - start_time.tv_usec) / 1000L;
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+
+ values[0] = Int64GetDatum(mem_allocated);
+ values[1] = Int64GetDatum(alloc_time);
+ values[2] = Int64GetDatum(free_time);
+
+ memset(nulls, 0, sizeof(nulls));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ result = HeapTupleGetDatum(tuple);
+
+ MemoryContextDelete(cxt);
+
+ PG_RETURN_DATUM(result);
+}
diff --git a/contrib/slab_bench/slab_bench.control b/contrib/slab_bench/slab_bench.control
new file mode 100644
index 0000000000..290809c6f7
--- /dev/null
+++ b/contrib/slab_bench/slab_bench.control
@@ -0,0 +1,4 @@
+# slab_bench extension
+comment = 'functions for benchmarking slab context'
+default_version = '1.0'
+module_pathname = '$libdir/slab_bench'
--
2.31.1
[text/x-patch] 0002-WIP-optimize-allocations-by-separating-hot-from-cold.patch (31.7K, ../../6717ca90-a2ea-b17d-f544-37f5181b1175@enterprisedb.com/3-0002-WIP-optimize-allocations-by-separating-hot-from-cold.patch)
download | inline diff:
From ae0e274cee0585fd91fbbe27a20cb575aefaef48 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 19 Jul 2021 12:55:51 -0700
Subject: [PATCH 2/3] WIP: optimize allocations by separating hot from cold
paths.
---
src/backend/utils/mmgr/aset.c | 429 ++++++++++++++--------------
src/backend/utils/mmgr/generation.c | 22 +-
src/backend/utils/mmgr/mcxt.c | 179 +++---------
src/backend/utils/mmgr/slab.c | 14 +-
src/include/nodes/memnodes.h | 4 +-
src/include/utils/memutils.h | 13 +
6 files changed, 300 insertions(+), 361 deletions(-)
diff --git a/src/backend/utils/mmgr/aset.c b/src/backend/utils/mmgr/aset.c
index 77872e77bc..0087835439 100644
--- a/src/backend/utils/mmgr/aset.c
+++ b/src/backend/utils/mmgr/aset.c
@@ -263,9 +263,9 @@ static AllocSetFreeList context_freelists[2] =
/*
* These functions implement the MemoryContext API for AllocSet contexts.
*/
-static void *AllocSetAlloc(MemoryContext context, Size size);
+static void *AllocSetAlloc(MemoryContext context, Size size, int flags);
static void AllocSetFree(MemoryContext context, void *pointer);
-static void *AllocSetRealloc(MemoryContext context, void *pointer, Size size);
+static void *AllocSetRealloc(MemoryContext context, void *pointer, Size size, int flags);
static void AllocSetReset(MemoryContext context);
static void AllocSetDelete(MemoryContext context);
static Size AllocSetGetChunkSpace(MemoryContext context, void *pointer);
@@ -704,6 +704,208 @@ AllocSetDelete(MemoryContext context)
free(set);
}
+static inline void *
+AllocSetAllocReturnChunk(AllocSet set, Size size, AllocChunk chunk, Size chunk_size)
+{
+ chunk->aset = (void *) set;
+#ifdef MEMORY_CONTEXT_CHECKING
+ chunk->requested_size = size;
+ /* set mark to catch clobber of "unused" space */
+ if (size < chunk->size)
+ set_sentinel(AllocChunkGetPointer(chunk), size);
+#endif
+#ifdef RANDOMIZE_ALLOCATED_MEMORY
+ /* fill the allocated space with junk */
+ randomize_mem((char *) AllocChunkGetPointer(chunk), size);
+#endif
+
+ /* Ensure any padding bytes are marked NOACCESS. */
+ VALGRIND_MAKE_MEM_NOACCESS((char *) AllocChunkGetPointer(chunk) + size,
+ chunk_size - size);
+
+ /* Disallow external access to private part of chunk header. */
+ VALGRIND_MAKE_MEM_NOACCESS(chunk, ALLOCCHUNK_PRIVATE_LEN);
+
+ return AllocChunkGetPointer(chunk);
+}
+
+static void * pg_noinline
+AllocSetAllocLarge(AllocSet set, Size size, int flags)
+{
+ AllocBlock block;
+ AllocChunk chunk;
+ Size chunk_size;
+ Size blksize;
+
+ /* check size, only allocation path where the limits could be hit */
+ MemoryContextCheckSize(&set->header, size, flags);
+
+ AssertArg(AllocSetIsValid(set));
+
+ chunk_size = MAXALIGN(size);
+ blksize = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
+ block = (AllocBlock) malloc(blksize);
+ if (block == NULL)
+ return NULL;
+
+ set->header.mem_allocated += blksize;
+
+ block->aset = set;
+ block->freeptr = block->endptr = ((char *) block) + blksize;
+
+ /*
+ * Stick the new block underneath the active allocation block, if any,
+ * so that we don't lose the use of the space remaining therein.
+ */
+ if (set->blocks != NULL)
+ {
+ block->prev = set->blocks;
+ block->next = set->blocks->next;
+ if (block->next)
+ block->next->prev = block;
+ set->blocks->next = block;
+ }
+ else
+ {
+ block->prev = NULL;
+ block->next = NULL;
+ set->blocks = block;
+ }
+
+ chunk = (AllocChunk) (((char *) block) + ALLOC_BLOCKHDRSZ);
+ chunk->size = chunk_size;
+
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk_size);
+}
+
+static void * pg_noinline
+AllocSetAllocFromNewBlock(AllocSet set, Size size, Size chunk_size)
+{
+ AllocBlock block;
+ Size blksize;
+ Size required_size;
+ AllocChunk chunk;
+
+ /*
+ * The first such block has size initBlockSize, and we double the
+ * space in each succeeding block, but not more than maxBlockSize.
+ */
+ blksize = set->nextBlockSize;
+ set->nextBlockSize <<= 1;
+ if (set->nextBlockSize > set->maxBlockSize)
+ set->nextBlockSize = set->maxBlockSize;
+
+ /*
+ * If initBlockSize is less than ALLOC_CHUNK_LIMIT, we could need more
+ * space... but try to keep it a power of 2.
+ */
+ required_size = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
+ while (blksize < required_size)
+ blksize <<= 1;
+
+ /* Try to allocate it */
+ block = (AllocBlock) malloc(blksize);
+
+ /*
+ * We could be asking for pretty big blocks here, so cope if malloc
+ * fails. But give up if there's less than 1 MB or so available...
+ */
+ while (block == NULL && blksize > 1024 * 1024)
+ {
+ blksize >>= 1;
+ if (blksize < required_size)
+ break;
+ block = (AllocBlock) malloc(blksize);
+ }
+
+ if (block == NULL)
+ return NULL;
+
+ set->header.mem_allocated += blksize;
+
+ block->aset = set;
+ block->freeptr = ((char *) block) + ALLOC_BLOCKHDRSZ;
+ block->endptr = ((char *) block) + blksize;
+
+ /* Mark unallocated space NOACCESS. */
+ VALGRIND_MAKE_MEM_NOACCESS(block->freeptr,
+ blksize - ALLOC_BLOCKHDRSZ);
+
+ block->prev = NULL;
+ block->next = set->blocks;
+ if (block->next)
+ block->next->prev = block;
+ set->blocks = block;
+
+ /*
+ * OK, do the allocation
+ */
+ chunk = (AllocChunk) (block->freeptr);
+
+ /* Prepare to initialize the chunk header. */
+ VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
+
+ block->freeptr += (chunk_size + ALLOC_CHUNKHDRSZ);
+ Assert(block->freeptr <= block->endptr);
+
+ chunk->size = chunk_size;
+
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk_size);
+}
+
+static void * pg_noinline
+AllocSetAllocCarveOldAndAlloc(AllocSet set, Size size, Size chunk_size, AllocBlock block, Size availspace)
+{
+ AllocChunk chunk;
+
+ /*
+ * The existing active (top) block does not have enough room for
+ * the requested allocation, but it might still have a useful
+ * amount of space in it. Once we push it down in the block list,
+ * we'll never try to allocate more space from it. So, before we
+ * do that, carve up its free space into chunks that we can put on
+ * the set's freelists.
+ *
+ * Because we can only get here when there's less than
+ * ALLOC_CHUNK_LIMIT left in the block, this loop cannot iterate
+ * more than ALLOCSET_NUM_FREELISTS-1 times.
+ */
+ while (availspace >= ((1 << ALLOC_MINBITS) + ALLOC_CHUNKHDRSZ))
+ {
+ Size availchunk = availspace - ALLOC_CHUNKHDRSZ;
+ int a_fidx = AllocSetFreeIndex(availchunk);
+
+ /*
+ * In most cases, we'll get back the index of the next larger
+ * freelist than the one we need to put this chunk on. The
+ * exception is when availchunk is exactly a power of 2.
+ */
+ if (availchunk != ((Size) 1 << (a_fidx + ALLOC_MINBITS)))
+ {
+ a_fidx--;
+ Assert(a_fidx >= 0);
+ availchunk = ((Size) 1 << (a_fidx + ALLOC_MINBITS));
+ }
+
+ chunk = (AllocChunk) (block->freeptr);
+
+ /* Prepare to initialize the chunk header. */
+ VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
+
+ block->freeptr += (availchunk + ALLOC_CHUNKHDRSZ);
+ availspace -= (availchunk + ALLOC_CHUNKHDRSZ);
+
+ chunk->size = availchunk;
+#ifdef MEMORY_CONTEXT_CHECKING
+ chunk->requested_size = 0; /* mark it free */
+#endif
+ chunk->aset = (void *) set->freelist[a_fidx];
+ set->freelist[a_fidx] = chunk;
+ }
+
+ return AllocSetAllocFromNewBlock(set, size, chunk_size);
+}
+
/*
* AllocSetAlloc
* Returns pointer to allocated memory of given size or NULL if
@@ -718,14 +920,13 @@ AllocSetDelete(MemoryContext context)
* return space that is marked NOACCESS - AllocSetRealloc has to beware!
*/
static void *
-AllocSetAlloc(MemoryContext context, Size size)
+AllocSetAlloc(MemoryContext context, Size size, int flags)
{
AllocSet set = (AllocSet) context;
AllocBlock block;
AllocChunk chunk;
int fidx;
Size chunk_size;
- Size blksize;
AssertArg(AllocSetIsValid(set));
@@ -733,61 +934,8 @@ AllocSetAlloc(MemoryContext context, Size size)
* If requested size exceeds maximum for chunks, allocate an entire block
* for this request.
*/
- if (size > set->allocChunkLimit)
- {
- chunk_size = MAXALIGN(size);
- blksize = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
- block = (AllocBlock) malloc(blksize);
- if (block == NULL)
- return NULL;
-
- context->mem_allocated += blksize;
-
- block->aset = set;
- block->freeptr = block->endptr = ((char *) block) + blksize;
-
- chunk = (AllocChunk) (((char *) block) + ALLOC_BLOCKHDRSZ);
- chunk->aset = set;
- chunk->size = chunk_size;
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = size;
- /* set mark to catch clobber of "unused" space */
- if (size < chunk_size)
- set_sentinel(AllocChunkGetPointer(chunk), size);
-#endif
-#ifdef RANDOMIZE_ALLOCATED_MEMORY
- /* fill the allocated space with junk */
- randomize_mem((char *) AllocChunkGetPointer(chunk), size);
-#endif
-
- /*
- * Stick the new block underneath the active allocation block, if any,
- * so that we don't lose the use of the space remaining therein.
- */
- if (set->blocks != NULL)
- {
- block->prev = set->blocks;
- block->next = set->blocks->next;
- if (block->next)
- block->next->prev = block;
- set->blocks->next = block;
- }
- else
- {
- block->prev = NULL;
- block->next = NULL;
- set->blocks = block;
- }
-
- /* Ensure any padding bytes are marked NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS((char *) AllocChunkGetPointer(chunk) + size,
- chunk_size - size);
-
- /* Disallow external access to private part of chunk header. */
- VALGRIND_MAKE_MEM_NOACCESS(chunk, ALLOCCHUNK_PRIVATE_LEN);
-
- return AllocChunkGetPointer(chunk);
- }
+ if (unlikely(size > set->allocChunkLimit))
+ return AllocSetAllocLarge(set, size, flags);
/*
* Request is small enough to be treated as a chunk. Look in the
@@ -803,27 +951,7 @@ AllocSetAlloc(MemoryContext context, Size size)
set->freelist[fidx] = (AllocChunk) chunk->aset;
- chunk->aset = (void *) set;
-
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = size;
- /* set mark to catch clobber of "unused" space */
- if (size < chunk->size)
- set_sentinel(AllocChunkGetPointer(chunk), size);
-#endif
-#ifdef RANDOMIZE_ALLOCATED_MEMORY
- /* fill the allocated space with junk */
- randomize_mem((char *) AllocChunkGetPointer(chunk), size);
-#endif
-
- /* Ensure any padding bytes are marked NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS((char *) AllocChunkGetPointer(chunk) + size,
- chunk->size - size);
-
- /* Disallow external access to private part of chunk header. */
- VALGRIND_MAKE_MEM_NOACCESS(chunk, ALLOCCHUNK_PRIVATE_LEN);
-
- return AllocChunkGetPointer(chunk);
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk->size);
}
/*
@@ -840,115 +968,16 @@ AllocSetAlloc(MemoryContext context, Size size)
{
Size availspace = block->endptr - block->freeptr;
- if (availspace < (chunk_size + ALLOC_CHUNKHDRSZ))
- {
- /*
- * The existing active (top) block does not have enough room for
- * the requested allocation, but it might still have a useful
- * amount of space in it. Once we push it down in the block list,
- * we'll never try to allocate more space from it. So, before we
- * do that, carve up its free space into chunks that we can put on
- * the set's freelists.
- *
- * Because we can only get here when there's less than
- * ALLOC_CHUNK_LIMIT left in the block, this loop cannot iterate
- * more than ALLOCSET_NUM_FREELISTS-1 times.
- */
- while (availspace >= ((1 << ALLOC_MINBITS) + ALLOC_CHUNKHDRSZ))
- {
- Size availchunk = availspace - ALLOC_CHUNKHDRSZ;
- int a_fidx = AllocSetFreeIndex(availchunk);
-
- /*
- * In most cases, we'll get back the index of the next larger
- * freelist than the one we need to put this chunk on. The
- * exception is when availchunk is exactly a power of 2.
- */
- if (availchunk != ((Size) 1 << (a_fidx + ALLOC_MINBITS)))
- {
- a_fidx--;
- Assert(a_fidx >= 0);
- availchunk = ((Size) 1 << (a_fidx + ALLOC_MINBITS));
- }
-
- chunk = (AllocChunk) (block->freeptr);
-
- /* Prepare to initialize the chunk header. */
- VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
-
- block->freeptr += (availchunk + ALLOC_CHUNKHDRSZ);
- availspace -= (availchunk + ALLOC_CHUNKHDRSZ);
-
- chunk->size = availchunk;
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = 0; /* mark it free */
-#endif
- chunk->aset = (void *) set->freelist[a_fidx];
- set->freelist[a_fidx] = chunk;
- }
-
- /* Mark that we need to create a new block */
- block = NULL;
- }
+ if (unlikely(availspace < (chunk_size + ALLOC_CHUNKHDRSZ)))
+ return AllocSetAllocCarveOldAndAlloc(set, size, chunk_size,
+ block, availspace);
}
-
- /*
- * Time to create a new regular (multi-chunk) block?
- */
- if (block == NULL)
+ else if (unlikely(block == NULL))
{
- Size required_size;
-
- /*
- * The first such block has size initBlockSize, and we double the
- * space in each succeeding block, but not more than maxBlockSize.
- */
- blksize = set->nextBlockSize;
- set->nextBlockSize <<= 1;
- if (set->nextBlockSize > set->maxBlockSize)
- set->nextBlockSize = set->maxBlockSize;
-
- /*
- * If initBlockSize is less than ALLOC_CHUNK_LIMIT, we could need more
- * space... but try to keep it a power of 2.
- */
- required_size = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
- while (blksize < required_size)
- blksize <<= 1;
-
- /* Try to allocate it */
- block = (AllocBlock) malloc(blksize);
-
/*
- * We could be asking for pretty big blocks here, so cope if malloc
- * fails. But give up if there's less than 1 MB or so available...
+ * Time to create a new regular (multi-chunk) block.
*/
- while (block == NULL && blksize > 1024 * 1024)
- {
- blksize >>= 1;
- if (blksize < required_size)
- break;
- block = (AllocBlock) malloc(blksize);
- }
-
- if (block == NULL)
- return NULL;
-
- context->mem_allocated += blksize;
-
- block->aset = set;
- block->freeptr = ((char *) block) + ALLOC_BLOCKHDRSZ;
- block->endptr = ((char *) block) + blksize;
-
- /* Mark unallocated space NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS(block->freeptr,
- blksize - ALLOC_BLOCKHDRSZ);
-
- block->prev = NULL;
- block->next = set->blocks;
- if (block->next)
- block->next->prev = block;
- set->blocks = block;
+ return AllocSetAllocFromNewBlock(set, size, chunk_size);
}
/*
@@ -959,30 +988,12 @@ AllocSetAlloc(MemoryContext context, Size size)
/* Prepare to initialize the chunk header. */
VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
- block->freeptr += (chunk_size + ALLOC_CHUNKHDRSZ);
- Assert(block->freeptr <= block->endptr);
-
- chunk->aset = (void *) set;
chunk->size = chunk_size;
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = size;
- /* set mark to catch clobber of "unused" space */
- if (size < chunk->size)
- set_sentinel(AllocChunkGetPointer(chunk), size);
-#endif
-#ifdef RANDOMIZE_ALLOCATED_MEMORY
- /* fill the allocated space with junk */
- randomize_mem((char *) AllocChunkGetPointer(chunk), size);
-#endif
-
- /* Ensure any padding bytes are marked NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS((char *) AllocChunkGetPointer(chunk) + size,
- chunk_size - size);
- /* Disallow external access to private part of chunk header. */
- VALGRIND_MAKE_MEM_NOACCESS(chunk, ALLOCCHUNK_PRIVATE_LEN);
+ block->freeptr += (chunk_size + ALLOC_CHUNKHDRSZ);
+ Assert(block->freeptr <= block->endptr);
- return AllocChunkGetPointer(chunk);
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk_size);
}
/*
@@ -1072,7 +1083,7 @@ AllocSetFree(MemoryContext context, void *pointer)
* request size.)
*/
static void *
-AllocSetRealloc(MemoryContext context, void *pointer, Size size)
+AllocSetRealloc(MemoryContext context, void *pointer, Size size, int flags)
{
AllocSet set = (AllocSet) context;
AllocChunk chunk = AllocPointerGetChunk(pointer);
@@ -1081,6 +1092,8 @@ AllocSetRealloc(MemoryContext context, void *pointer, Size size)
/* Allow access to private part of chunk header. */
VALGRIND_MAKE_MEM_DEFINED(chunk, ALLOCCHUNK_PRIVATE_LEN);
+ MemoryContextCheckSize(context, size, flags);
+
oldsize = chunk->size;
#ifdef MEMORY_CONTEXT_CHECKING
@@ -1260,7 +1273,7 @@ AllocSetRealloc(MemoryContext context, void *pointer, Size size)
AllocPointer newPointer;
/* allocate new chunk */
- newPointer = AllocSetAlloc((MemoryContext) set, size);
+ newPointer = AllocSetAlloc((MemoryContext) set, size, flags);
/* leave immediately if request was not completed */
if (newPointer == NULL)
diff --git a/src/backend/utils/mmgr/generation.c b/src/backend/utils/mmgr/generation.c
index 584cd614da..5dde15654d 100644
--- a/src/backend/utils/mmgr/generation.c
+++ b/src/backend/utils/mmgr/generation.c
@@ -146,9 +146,9 @@ struct GenerationChunk
/*
* These functions implement the MemoryContext API for Generation contexts.
*/
-static void *GenerationAlloc(MemoryContext context, Size size);
+static void *GenerationAlloc(MemoryContext context, Size size, int flags);
static void GenerationFree(MemoryContext context, void *pointer);
-static void *GenerationRealloc(MemoryContext context, void *pointer, Size size);
+static void *GenerationRealloc(MemoryContext context, void *pointer, Size size, int flags);
static void GenerationReset(MemoryContext context);
static void GenerationDelete(MemoryContext context);
static Size GenerationGetChunkSpace(MemoryContext context, void *pointer);
@@ -323,7 +323,7 @@ GenerationDelete(MemoryContext context)
* return space that is marked NOACCESS - GenerationRealloc has to beware!
*/
static void *
-GenerationAlloc(MemoryContext context, Size size)
+GenerationAlloc(MemoryContext context, Size size, int flags)
{
GenerationContext *set = (GenerationContext *) context;
GenerationBlock *block;
@@ -335,9 +335,12 @@ GenerationAlloc(MemoryContext context, Size size)
{
Size blksize = chunk_size + Generation_BLOCKHDRSZ + Generation_CHUNKHDRSZ;
+ /* check size, only allocation path where the limits could be hit */
+ MemoryContextCheckSize(context, size, flags);
+
block = (GenerationBlock *) malloc(blksize);
if (block == NULL)
- return NULL;
+ return MemoryContextAllocationFailure(context, size, flags);
context->mem_allocated += blksize;
@@ -392,7 +395,7 @@ GenerationAlloc(MemoryContext context, Size size)
block = (GenerationBlock *) malloc(blksize);
if (block == NULL)
- return NULL;
+ return MemoryContextAllocationFailure(context, size, flags);
context->mem_allocated += blksize;
@@ -520,13 +523,15 @@ GenerationFree(MemoryContext context, void *pointer)
* into the old chunk - in that case we just update chunk header.
*/
static void *
-GenerationRealloc(MemoryContext context, void *pointer, Size size)
+GenerationRealloc(MemoryContext context, void *pointer, Size size, int flags)
{
GenerationContext *set = (GenerationContext *) context;
GenerationChunk *chunk = GenerationPointerGetChunk(pointer);
GenerationPointer newPointer;
Size oldsize;
+ MemoryContextCheckSize(context, size, flags);
+
/* Allow access to private part of chunk header. */
VALGRIND_MAKE_MEM_DEFINED(chunk, GENERATIONCHUNK_PRIVATE_LEN);
@@ -596,14 +601,15 @@ GenerationRealloc(MemoryContext context, void *pointer, Size size)
}
/* allocate new chunk */
- newPointer = GenerationAlloc((MemoryContext) set, size);
+ newPointer = GenerationAlloc((MemoryContext) set, size, flags);
/* leave immediately if request was not completed */
if (newPointer == NULL)
{
/* Disallow external access to private part of chunk header. */
VALGRIND_MAKE_MEM_NOACCESS(chunk, GENERATIONCHUNK_PRIVATE_LEN);
- return NULL;
+ /* again? */
+ return MemoryContextAllocationFailure(context, size, flags);
}
/*
diff --git a/src/backend/utils/mmgr/mcxt.c b/src/backend/utils/mmgr/mcxt.c
index 6919a73280..13125c4a56 100644
--- a/src/backend/utils/mmgr/mcxt.c
+++ b/src/backend/utils/mmgr/mcxt.c
@@ -867,28 +867,9 @@ MemoryContextAlloc(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
-
- /*
- * Here, and elsewhere in this module, we show the target context's
- * "name" but not its "ident" (if any) in user-visible error messages.
- * The "ident" string might contain security-sensitive data, such as
- * values in SQL commands.
- */
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -910,21 +891,10 @@ MemoryContextAllocZero(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -948,21 +918,10 @@ MemoryContextAllocZeroAligned(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -983,26 +942,11 @@ MemoryContextAllocExtended(MemoryContext context, Size size, int flags)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (((flags & MCXT_ALLOC_HUGE) != 0 && !AllocHugeSizeIsValid(size)) ||
- ((flags & MCXT_ALLOC_HUGE) == 0 && !AllocSizeIsValid(size)))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
+ ret = context->methods->alloc(context, size, flags);
if (unlikely(ret == NULL))
- {
- if ((flags & MCXT_ALLOC_NO_OOM) == 0)
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
return NULL;
- }
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1067,22 +1011,10 @@ palloc(Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
-
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1099,21 +1031,10 @@ palloc0(Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1132,26 +1053,11 @@ palloc_extended(Size size, int flags)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (((flags & MCXT_ALLOC_HUGE) != 0 && !AllocHugeSizeIsValid(size)) ||
- ((flags & MCXT_ALLOC_HUGE) == 0 && !AllocSizeIsValid(size)))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
+ ret = context->methods->alloc(context, size, 0);
if (unlikely(ret == NULL))
- {
- if ((flags & MCXT_ALLOC_NO_OOM) == 0)
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
return NULL;
- }
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1184,24 +1090,13 @@ repalloc(void *pointer, Size size)
MemoryContext context = GetMemoryChunkContext(pointer);
void *ret;
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
AssertNotInCriticalSection(context);
/* isReset must be false already */
Assert(!context->isReset);
- ret = context->methods->realloc(context, pointer, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->realloc(context, pointer, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_CHANGE(context, pointer, ret, size);
@@ -1222,21 +1117,9 @@ MemoryContextAllocHuge(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocHugeSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
-
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, MCXT_ALLOC_HUGE);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1254,16 +1137,23 @@ repalloc_huge(void *pointer, Size size)
MemoryContext context = GetMemoryChunkContext(pointer);
void *ret;
- if (!AllocHugeSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
AssertNotInCriticalSection(context);
/* isReset must be false already */
Assert(!context->isReset);
- ret = context->methods->realloc(context, pointer, size);
- if (unlikely(ret == NULL))
+ ret = context->methods->realloc(context, pointer, size, MCXT_ALLOC_HUGE);
+ Assert(ret != NULL);
+
+ VALGRIND_MEMPOOL_CHANGE(context, pointer, ret, size);
+
+ return ret;
+}
+
+void *
+MemoryContextAllocationFailure(MemoryContext context, Size size, int flags)
+{
+ if ((flags & MCXT_ALLOC_NO_OOM) == 0)
{
MemoryContextStats(TopMemoryContext);
ereport(ERROR,
@@ -1273,9 +1163,13 @@ repalloc_huge(void *pointer, Size size)
size, context->name)));
}
- VALGRIND_MEMPOOL_CHANGE(context, pointer, ret, size);
+ return NULL;
+}
- return ret;
+void
+MemoryContextSizeFailure(MemoryContext context, Size size, int flags)
+{
+ elog(ERROR, "invalid memory alloc request size %zu", size);
}
/*
@@ -1298,6 +1192,13 @@ MemoryContextStrdup(MemoryContext context, const char *string)
char *
pstrdup(const char *in)
{
+ /*
+ * Here, and elsewhere in this module, we show the target context's
+ * "name" but not its "ident" (if any) in user-visible error messages.
+ * The "ident" string might contain security-sensitive data, such as
+ * values in SQL commands.
+ */
+
return MemoryContextStrdup(CurrentMemoryContext, in);
}
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index 553dd7f667..e4b8275045 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -126,9 +126,9 @@ typedef struct SlabChunk
/*
* These functions implement the MemoryContext API for Slab contexts.
*/
-static void *SlabAlloc(MemoryContext context, Size size);
+static void *SlabAlloc(MemoryContext context, Size size, int flags);
static void SlabFree(MemoryContext context, void *pointer);
-static void *SlabRealloc(MemoryContext context, void *pointer, Size size);
+static void *SlabRealloc(MemoryContext context, void *pointer, Size size, int flags);
static void SlabReset(MemoryContext context);
static void SlabDelete(MemoryContext context);
static Size SlabGetChunkSpace(MemoryContext context, void *pointer);
@@ -337,7 +337,7 @@ SlabDelete(MemoryContext context)
* request could not be completed; memory is added to the slab.
*/
static void *
-SlabAlloc(MemoryContext context, Size size)
+SlabAlloc(MemoryContext context, Size size, int flags)
{
SlabContext *slab = castNode(SlabContext, context);
SlabBlock *block;
@@ -346,6 +346,12 @@ SlabAlloc(MemoryContext context, Size size)
Assert(slab);
+ /*
+ * XXX: Probably no need to check for huge allocations, we only support
+ * one size? Which could theoretically be huge, but that'd not make
+ * sense...
+ */
+
Assert((slab->minFreeChunks >= 0) &&
(slab->minFreeChunks < slab->chunksPerBlock));
@@ -583,7 +589,7 @@ SlabFree(MemoryContext context, void *pointer)
* realloc is usually used to enlarge the chunk.
*/
static void *
-SlabRealloc(MemoryContext context, void *pointer, Size size)
+SlabRealloc(MemoryContext context, void *pointer, Size size, int flags)
{
SlabContext *slab = castNode(SlabContext, context);
diff --git a/src/include/nodes/memnodes.h b/src/include/nodes/memnodes.h
index e6a757d6a0..8a42d2ff99 100644
--- a/src/include/nodes/memnodes.h
+++ b/src/include/nodes/memnodes.h
@@ -57,10 +57,10 @@ typedef void (*MemoryStatsPrintFunc) (MemoryContext context, void *passthru,
typedef struct MemoryContextMethods
{
- void *(*alloc) (MemoryContext context, Size size);
+ void *(*alloc) (MemoryContext context, Size size, int flags);
/* call this free_p in case someone #define's free() */
void (*free_p) (MemoryContext context, void *pointer);
- void *(*realloc) (MemoryContext context, void *pointer, Size size);
+ void *(*realloc) (MemoryContext context, void *pointer, Size size, int flags);
void (*reset) (MemoryContext context);
void (*delete_context) (MemoryContext context);
Size (*get_chunk_space) (MemoryContext context, void *pointer);
diff --git a/src/include/utils/memutils.h b/src/include/utils/memutils.h
index ff872274d4..2f75b4cca4 100644
--- a/src/include/utils/memutils.h
+++ b/src/include/utils/memutils.h
@@ -147,6 +147,19 @@ extern void MemoryContextCreate(MemoryContext node,
extern void HandleLogMemoryContextInterrupt(void);
extern void ProcessLogMemoryContextInterrupt(void);
+extern void *MemoryContextAllocationFailure(MemoryContext context, Size size, int flags);
+
+extern void MemoryContextSizeFailure(MemoryContext context, Size size, int flags) pg_attribute_noreturn();
+
+static inline void
+MemoryContextCheckSize(MemoryContext context, Size size, int flags)
+{
+ if (unlikely(!AllocSizeIsValid(size)))
+ {
+ if (!(flags & MCXT_ALLOC_HUGE) || !AllocHugeSizeIsValid(size))
+ MemoryContextSizeFailure(context, size, flags);
+ }
+}
/*
* Memory-context-type-specific functions
--
2.31.1
[text/x-patch] 0003-WIP-slab-performance.patch (23.4K, ../../6717ca90-a2ea-b17d-f544-37f5181b1175@enterprisedb.com/4-0003-WIP-slab-performance.patch)
download | inline diff:
From adeeaac000dc16f5c0fc0c447be438a0c37f6cd3 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 19 Jul 2021 13:48:14 -0700
Subject: [PATCH 3/3] WIP: slab performance.
Discussion: https://postgr.es/m/20210717194333.mr5io3zup3kxahfm@alap3.anarazel.de
---
src/backend/utils/mmgr/slab.c | 434 ++++++++++++++++++++++------------
1 file changed, 285 insertions(+), 149 deletions(-)
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index e4b8275045..c979553dcd 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -23,15 +23,18 @@
* global (context) level. This is possible as the chunk size (and thus also
* the number of chunks per block) is fixed.
*
- * On each block, free chunks are tracked in a simple linked list. Contents
- * of free chunks is replaced with an index of the next free chunk, forming
- * a very simple linked list. Each block also contains a counter of free
- * chunks. Combined with the local block-level freelist, it makes it trivial
- * to eventually free the whole block.
+ * On each block, never allocated chunks are tracked by a simple offset, and
+ * free chunks are tracked in a simple linked list. The offset approach
+ * avoids needing to iterate over all chunks when allocating a new block,
+ * which would cause page faults and cache pollution. Contents of free chunks
+ * is replaced with a pointer to the next free chunk, forming a very simple
+ * linked list. Each block also contains a counter of free chunks. Combined
+ * with the local block-level freelist, it makes it trivial to eventually
+ * free the whole block.
*
* At the context level, we use 'freelist' to track blocks ordered by number
* of free chunks, starting with blocks having a single allocated chunk, and
- * with completely full blocks on the tail.
+ * with completely full blocks on the tail. XXX
*
* This also allows various optimizations - for example when searching for
* free chunk, the allocator reuses space from the fullest blocks first, in
@@ -44,7 +47,7 @@
* case this performs as if the pointer was not maintained.
*
* We cache the freelist index for the blocks with the fewest free chunks
- * (minFreeChunks), so that we don't have to search the freelist on every
+ * (minFreeChunkIndex), so that we don't have to search the freelist on every
* SlabAlloc() call, which is quite expensive.
*
*-------------------------------------------------------------------------
@@ -56,6 +59,17 @@
#include "utils/memdebug.h"
#include "utils/memutils.h"
+struct SlabBlock;
+struct SlabChunk;
+
+/*
+ * Number of actual freelists + 1, for full blocks. Full blocks are always at
+ * offset 0.
+ */
+#define SLAB_FREELIST_COUNT 9
+
+#define SLAB_RETAIN_EMPTY_BLOCK_COUNT 10
+
/*
* SlabContext is a specialized implementation of MemoryContext.
*/
@@ -68,13 +82,16 @@ typedef struct SlabContext
Size blockSize; /* block size */
Size headerSize; /* allocated size of context header */
int chunksPerBlock; /* number of chunks per block */
- int minFreeChunks; /* min number of free chunks in any block */
+ int minFreeChunksIndex; /* min number of free chunks in any block XXX */
int nblocks; /* number of blocks allocated */
#ifdef MEMORY_CONTEXT_CHECKING
bool *freechunks; /* bitmap of free chunks in a block */
#endif
+ int nfreeblocks;
+ dlist_head freeblocks;
+ int freelist_shift;
/* blocks with free space, grouped by number of free chunks: */
- dlist_head freelist[FLEXIBLE_ARRAY_MEMBER];
+ dlist_head freelist[SLAB_FREELIST_COUNT];
} SlabContext;
/*
@@ -83,13 +100,15 @@ typedef struct SlabContext
*
* node: doubly-linked list of blocks in global freelist
* nfree: number of free chunks in this block
- * firstFreeChunk: index of the first free chunk
+ * firstFreeChunk: first free chunk
*/
typedef struct SlabBlock
{
dlist_node node; /* doubly-linked list */
- int nfree; /* number of free chunks */
- int firstFreeChunk; /* index of the first free chunk in the block */
+ int nfree; /* number of chunks on freelist + unused */
+ int nunused; /* number of unused chunks */
+ struct SlabChunk *unused; /* */
+ struct SlabChunk *firstFreeChunk; /* first free chunk in the block */
} SlabBlock;
/*
@@ -123,6 +142,35 @@ typedef struct SlabChunk
#define SlabChunkIndex(slab, block, chunk) \
(((char *) chunk - SlabBlockStart(block)) / slab->fullChunkSize)
+static inline uint8
+SlabFreelistIndex(SlabContext *slab, int freecount)
+{
+ uint8 index;
+
+ Assert(freecount <= slab->chunksPerBlock);
+
+ /*
+ * Ensure, without a branch, that index 0 is only used for blocks entirely
+ * without free chunks.
+ * XXX: There probably is a cheaper way to do this. Needing to shift twice
+ * by slab->freelist_shift isn't great.
+ */
+ index = (freecount + (1 << slab->freelist_shift) - 1) >> slab->freelist_shift;
+
+ if (freecount == 0)
+ Assert(index == 0);
+ else
+ Assert(index > 0 && index < (SLAB_FREELIST_COUNT));
+
+ return index;
+}
+
+static inline dlist_head*
+SlabFreelist(SlabContext *slab, int freecount)
+{
+ return &slab->freelist[SlabFreelistIndex(slab, freecount)];
+}
+
/*
* These functions implement the MemoryContext API for Slab contexts.
*/
@@ -179,7 +227,6 @@ SlabContextCreate(MemoryContext parent,
{
int chunksPerBlock;
Size fullChunkSize;
- Size freelistSize;
Size headerSize;
SlabContext *slab;
int i;
@@ -192,11 +239,11 @@ SlabContextCreate(MemoryContext parent,
"padding calculation in SlabChunk is wrong");
/* Make sure the linked list node fits inside a freed chunk */
- if (chunkSize < sizeof(int))
- chunkSize = sizeof(int);
+ if (chunkSize < MAXALIGN(sizeof(void *)))
+ chunkSize = MAXALIGN(sizeof(void *));
/* chunk, including SLAB header (both addresses nicely aligned) */
- fullChunkSize = sizeof(SlabChunk) + MAXALIGN(chunkSize);
+ fullChunkSize = sizeof(SlabChunk) + chunkSize;
/* Make sure the block can store at least one chunk. */
if (blockSize < fullChunkSize + sizeof(SlabBlock))
@@ -206,16 +253,14 @@ SlabContextCreate(MemoryContext parent,
/* Compute maximum number of chunks per block */
chunksPerBlock = (blockSize - sizeof(SlabBlock)) / fullChunkSize;
- /* The freelist starts with 0, ends with chunksPerBlock. */
- freelistSize = sizeof(dlist_head) * (chunksPerBlock + 1);
-
/*
* Allocate the context header. Unlike aset.c, we never try to combine
* this with the first regular block; not worth the extra complication.
+ * XXX: What's the evidence for that?
*/
/* Size of the memory context header */
- headerSize = offsetof(SlabContext, freelist) + freelistSize;
+ headerSize = sizeof(SlabContext);
#ifdef MEMORY_CONTEXT_CHECKING
@@ -249,17 +294,34 @@ SlabContextCreate(MemoryContext parent,
slab->blockSize = blockSize;
slab->headerSize = headerSize;
slab->chunksPerBlock = chunksPerBlock;
- slab->minFreeChunks = 0;
+ slab->minFreeChunksIndex = 0;
slab->nblocks = 0;
+ slab->nfreeblocks = 0;
+
+ /*
+ * Compute a shift that guarantees that shifting chunksPerBlock with it
+ * yields is smaller than SLAB_FREELIST_COUNT - 1 (one freelist is used for full blocks).
+ */
+ slab->freelist_shift = 0;
+ while ((slab->chunksPerBlock >> slab->freelist_shift) >= (SLAB_FREELIST_COUNT - 1))
+ slab->freelist_shift++;
+
+#if 0
+ elog(LOG, "freelist shift for %d chunks of size %zu is %d, block size %zu",
+ slab->chunksPerBlock, slab->fullChunkSize, slab->freelist_shift,
+ slab->blockSize);
+#endif
+
+ dlist_init(&slab->freeblocks);
/* initialize the freelist slots */
- for (i = 0; i < (slab->chunksPerBlock + 1); i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
dlist_init(&slab->freelist[i]);
#ifdef MEMORY_CONTEXT_CHECKING
/* set the freechunks pointer right after the freelists array */
slab->freechunks
- = (bool *) slab + offsetof(SlabContext, freelist) + freelistSize;
+ = (bool *) slab + sizeof(SlabContext);
#endif
/* Finally, do the type-independent part of context creation */
@@ -292,8 +354,27 @@ SlabReset(MemoryContext context)
SlabCheck(context);
#endif
+ /* release retained empty blocks */
+ {
+ dlist_mutable_iter miter;
+
+ dlist_foreach_modify(miter, &slab->freeblocks)
+ {
+ SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
+
+ dlist_delete(miter.cur);
+
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ slab->nblocks--;
+ context->mem_allocated -= slab->blockSize;
+ }
+ }
+
/* walk over freelists and free the blocks */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
dlist_mutable_iter miter;
@@ -312,7 +393,7 @@ SlabReset(MemoryContext context)
}
}
- slab->minFreeChunks = 0;
+ slab->minFreeChunksIndex = 0;
Assert(slab->nblocks == 0);
Assert(context->mem_allocated == 0);
@@ -342,7 +423,6 @@ SlabAlloc(MemoryContext context, Size size, int flags)
SlabContext *slab = castNode(SlabContext, context);
SlabBlock *block;
SlabChunk *chunk;
- int idx;
Assert(slab);
@@ -352,8 +432,8 @@ SlabAlloc(MemoryContext context, Size size, int flags)
* sense...
*/
- Assert((slab->minFreeChunks >= 0) &&
- (slab->minFreeChunks < slab->chunksPerBlock));
+ Assert((slab->minFreeChunksIndex >= 0) &&
+ (slab->minFreeChunksIndex < SLAB_FREELIST_COUNT));
/* make sure we only allow correct request size */
if (size != slab->chunkSize)
@@ -367,58 +447,88 @@ SlabAlloc(MemoryContext context, Size size, int flags)
* slab->minFreeChunks == 0 means there are no blocks with free chunks,
* thanks to how minFreeChunks is updated at the end of SlabAlloc().
*/
- if (slab->minFreeChunks == 0)
+ if (unlikely(slab->minFreeChunksIndex == 0))
{
- block = (SlabBlock *) malloc(slab->blockSize);
+ if (slab->nfreeblocks > 0)
+ {
+ dlist_node *node;
- if (block == NULL)
- return NULL;
+ node = dlist_pop_head_node(&slab->freeblocks);
+ block = dlist_container(SlabBlock, node, node);
+ slab->nfreeblocks--;
+ }
+ else
+ {
+ block = (SlabBlock *) malloc(slab->blockSize);
- block->nfree = slab->chunksPerBlock;
- block->firstFreeChunk = 0;
+ if (unlikely(block == NULL))
+ return NULL;
- /*
- * Put all the chunks on a freelist. Walk the chunks and point each
- * one to the next one.
- */
- for (idx = 0; idx < slab->chunksPerBlock; idx++)
- {
- chunk = SlabBlockGetChunk(slab, block, idx);
- *(int32 *) SlabChunkGetPointer(chunk) = (idx + 1);
+ slab->nblocks += 1;
+ context->mem_allocated += slab->blockSize;
}
+ block->nfree = slab->chunksPerBlock;
+ block->firstFreeChunk = NULL;
+ block->nunused = slab->chunksPerBlock;
+ block->unused = (SlabChunk *) SlabBlockStart(block);
+
+ slab->minFreeChunksIndex = SlabFreelistIndex(slab, slab->chunksPerBlock);
+
/*
* And add it to the last freelist with all chunks empty.
*
* We know there are no blocks in the freelist, otherwise we wouldn't
* need a new block.
*/
- Assert(dlist_is_empty(&slab->freelist[slab->chunksPerBlock]));
+ Assert(dlist_is_empty(SlabFreelist(slab, slab->chunksPerBlock)));
- dlist_push_head(&slab->freelist[slab->chunksPerBlock], &block->node);
+ dlist_push_head(SlabFreelist(slab, slab->chunksPerBlock),
+ &block->node);
- slab->minFreeChunks = slab->chunksPerBlock;
- slab->nblocks += 1;
- context->mem_allocated += slab->blockSize;
+ chunk = block->unused;
+ block->unused = (SlabChunk *)(((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
}
+ else
+ {
+ /* grab the block from the freelist */
+ block = dlist_head_element(SlabBlock, node,
+ &slab->freelist[slab->minFreeChunksIndex]);
- /* grab the block from the freelist (even the new block is there) */
- block = dlist_head_element(SlabBlock, node,
- &slab->freelist[slab->minFreeChunks]);
+ /* make sure we actually got a valid block, with matching nfree */
+ Assert(block != NULL);
+ Assert(slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree));
+ Assert(block->nfree > 0);
- /* make sure we actually got a valid block, with matching nfree */
- Assert(block != NULL);
- Assert(slab->minFreeChunks == block->nfree);
- Assert(block->nfree > 0);
+ /* we know index of the first free chunk in the block */
+ if (block->nunused > 0)
+ {
+ chunk = block->unused;
+ block->unused = (SlabChunk *)(((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
+ else
+ {
+ chunk = block->firstFreeChunk;
- /* we know index of the first free chunk in the block */
- idx = block->firstFreeChunk;
+ /*
+ * Remove the chunk from the freelist head. The index of the next free
+ * chunk is stored in the chunk itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(void*));
+ block->firstFreeChunk = *(SlabChunk **) SlabChunkGetPointer(chunk);
+ }
+ }
- /* make sure the chunk index is valid, and that it's marked as empty */
- Assert((idx >= 0) && (idx < slab->chunksPerBlock));
+ /* make sure the chunk is in the block and that it's marked as empty (XXX?) */
+ Assert((char *) chunk >= SlabBlockStart(block));
+ Assert((char *) chunk < (((char *)block) + slab->blockSize));
- /* compute the chunk location block start (after the block header) */
- chunk = SlabBlockGetChunk(slab, block, idx);
+ Assert(block->firstFreeChunk == NULL || (
+ (block->firstFreeChunk >= (SlabChunk *) SlabBlockStart(block)) &&
+ block->firstFreeChunk <= (SlabChunk *) (((char *)block) + slab->blockSize))
+ );
/*
* Update the block nfree count, and also the minFreeChunks as we've
@@ -426,54 +536,51 @@ SlabAlloc(MemoryContext context, Size size, int flags)
* (because that's how we chose the block).
*/
block->nfree--;
- slab->minFreeChunks = block->nfree;
- /*
- * Remove the chunk from the freelist head. The index of the next free
- * chunk is stored in the chunk itself.
- */
- VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(int32));
- block->firstFreeChunk = *(int32 *) SlabChunkGetPointer(chunk);
+ /* Prepare to initialize the chunk header. */
+ VALGRIND_MAKE_MEM_UNDEFINED(chunk, sizeof(SlabChunk));
- Assert(block->firstFreeChunk >= 0);
- Assert(block->firstFreeChunk <= slab->chunksPerBlock);
+ chunk->block = block;
+ chunk->slab = slab;
- Assert((block->nfree != 0 &&
- block->firstFreeChunk < slab->chunksPerBlock) ||
- (block->nfree == 0 &&
- block->firstFreeChunk == slab->chunksPerBlock));
+ Assert(slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree + 1));
/* move the whole block to the right place in the freelist */
- dlist_delete(&block->node);
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
-
- /*
- * And finally update minFreeChunks, i.e. the index to the block with the
- * lowest number of free chunks. We only need to do that when the block
- * got full (otherwise we know the current block is the right one). We'll
- * simply walk the freelist until we find a non-empty entry.
- */
- if (slab->minFreeChunks == 0)
+ if (unlikely(SlabFreelistIndex(slab, block->nfree) != slab->minFreeChunksIndex))
{
- for (idx = 1; idx <= slab->chunksPerBlock; idx++)
+ slab->minFreeChunksIndex = SlabFreelistIndex(slab, block->nfree);
+ dlist_delete(&block->node);
+ dlist_push_head(SlabFreelist(slab, block->nfree), &block->node);
+
+
+ /*
+ * And finally update minFreeChunks, i.e. the index to the block with the
+ * lowest number of free chunks. We only need to do that when the block
+ * got full (otherwise we know the current block is the right one). We'll
+ * simply walk the freelist until we find a non-empty entry.
+ */
+ if (slab->minFreeChunksIndex == 0)
{
- if (dlist_is_empty(&slab->freelist[idx]))
- continue;
+ for (int idx = 1; idx < SLAB_FREELIST_COUNT; idx++)
+ {
+ if (dlist_is_empty(&slab->freelist[idx]))
+ continue;
- /* found a non-empty freelist */
- slab->minFreeChunks = idx;
- break;
+ /* found a non-empty freelist */
+ slab->minFreeChunksIndex = idx;
+ break;
+ }
}
}
+#if 0
+ /*
+ * FIXME: I don't understand what this ever did? It should be unreachable
+ * I think?
+ */
if (slab->minFreeChunks == slab->chunksPerBlock)
slab->minFreeChunks = 0;
-
- /* Prepare to initialize the chunk header. */
- VALGRIND_MAKE_MEM_UNDEFINED(chunk, sizeof(SlabChunk));
-
- chunk->block = block;
- chunk->slab = slab;
+#endif
#ifdef MEMORY_CONTEXT_CHECKING
/* slab mark to catch clobber of "unused" space */
@@ -496,6 +603,61 @@ SlabAlloc(MemoryContext context, Size size, int flags)
return SlabChunkGetPointer(chunk);
}
+static void pg_noinline
+SlabFreeSlow(SlabContext *slab, SlabBlock *block)
+{
+ dlist_delete(&block->node);
+
+ /*
+ * See if we need to update the minFreeChunks field for the slab - we only
+ * need to do that if the block had that number of free chunks
+ * before we freed one. In that case, we check if there still are blocks
+ * in the original freelist and we either keep the current value (if there
+ * still are blocks) or increment it by one (the new block is still the
+ * one with minimum free chunks).
+ *
+ * The one exception is when the block will get completely free - in that
+ * case we will free it, se we can't use it for minFreeChunks. It however
+ * means there are no more blocks with free chunks.
+ */
+ if (slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree - 1))
+ {
+ /* Have we removed the last chunk from the freelist? */
+ if (dlist_is_empty(&slab->freelist[slab->minFreeChunksIndex]))
+ {
+ /* but if we made the block entirely free, we'll free it */
+ if (block->nfree == slab->chunksPerBlock)
+ slab->minFreeChunksIndex = 0;
+ else
+ slab->minFreeChunksIndex++;
+ }
+ }
+
+ /* If the block is now completely empty, free it (in a way). */
+ if (block->nfree == slab->chunksPerBlock)
+ {
+ /*
+ * To avoid constantly freeing/allocating blocks in bursty patterns
+ * (on most crucially cases of repeatedly allocating and freeing a
+ * single chunk), retain a small number of blocks.
+ */
+ if (slab->nfreeblocks < SLAB_RETAIN_EMPTY_BLOCK_COUNT)
+ {
+ dlist_push_head(&slab->freeblocks, &block->node);
+ slab->nfreeblocks++;
+ }
+ else
+ {
+ slab->nblocks--;
+ slab->header.mem_allocated -= slab->blockSize;
+ free(block);
+ }
+ }
+ else
+ dlist_push_head(SlabFreelist(slab, block->nfree), &block->node);
+
+}
+
/*
* SlabFree
* Frees allocated memory; memory is removed from the slab.
@@ -503,7 +665,6 @@ SlabAlloc(MemoryContext context, Size size, int flags)
static void
SlabFree(MemoryContext context, void *pointer)
{
- int idx;
SlabContext *slab = castNode(SlabContext, context);
SlabChunk *chunk = SlabPointerGetChunk(pointer);
SlabBlock *block = chunk->block;
@@ -516,61 +677,26 @@ SlabFree(MemoryContext context, void *pointer)
slab->header.name, chunk);
#endif
- /* compute index of the chunk with respect to block start */
- idx = SlabChunkIndex(slab, block, chunk);
-
/* add chunk to freelist, and update block nfree count */
- *(int32 *) pointer = block->firstFreeChunk;
- block->firstFreeChunk = idx;
+ *(SlabChunk **) pointer = block->firstFreeChunk;
+ block->firstFreeChunk = chunk;
block->nfree++;
Assert(block->nfree > 0);
Assert(block->nfree <= slab->chunksPerBlock);
#ifdef CLOBBER_FREED_MEMORY
- /* XXX don't wipe the int32 index, used for block-level freelist */
- wipe_mem((char *) pointer + sizeof(int32),
- slab->chunkSize - sizeof(int32));
+ /* XXX don't wipe the SlabChunk* index, used for block-level freelist */
+ wipe_mem((char *) pointer + sizeof(SlabChunk*),
+ slab->chunkSize - sizeof(SlabChunk*));
#endif
/* remove the block from a freelist */
- dlist_delete(&block->node);
-
- /*
- * See if we need to update the minFreeChunks field for the slab - we only
- * need to do that if there the block had that number of free chunks
- * before we freed one. In that case, we check if there still are blocks
- * in the original freelist and we either keep the current value (if there
- * still are blocks) or increment it by one (the new block is still the
- * one with minimum free chunks).
- *
- * The one exception is when the block will get completely free - in that
- * case we will free it, se we can't use it for minFreeChunks. It however
- * means there are no more blocks with free chunks.
- */
- if (slab->minFreeChunks == (block->nfree - 1))
+ if (SlabFreelistIndex(slab, block->nfree) != SlabFreelistIndex(slab, block->nfree - 1))
{
- /* Have we removed the last chunk from the freelist? */
- if (dlist_is_empty(&slab->freelist[slab->minFreeChunks]))
- {
- /* but if we made the block entirely free, we'll free it */
- if (block->nfree == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
- else
- slab->minFreeChunks++;
- }
+ SlabFreeSlow(slab, block);
}
- /* If the block is now completely empty, free it. */
- if (block->nfree == slab->chunksPerBlock)
- {
- free(block);
- slab->nblocks--;
- context->mem_allocated -= slab->blockSize;
- }
- else
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
-
Assert(slab->nblocks >= 0);
Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
}
@@ -657,7 +783,7 @@ SlabStats(MemoryContext context,
/* Include context header in totalspace */
totalspace = slab->headerSize;
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
dlist_iter iter;
@@ -714,7 +840,7 @@ SlabCheck(MemoryContext context)
Assert(slab->chunksPerBlock > 0);
/* walk all the freelists */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
int j,
nfree;
@@ -723,20 +849,19 @@ SlabCheck(MemoryContext context)
/* walk all blocks on this freelist */
dlist_foreach(iter, &slab->freelist[i])
{
- int idx;
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+ SlabChunk *cur_chunk;
/*
* Make sure the number of free chunks (in the block header)
* matches position in the freelist.
*/
- if (block->nfree != i)
+ if (SlabFreelistIndex(slab, block->nfree) != i)
elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match freelist %d",
name, block->nfree, block, i);
/* reset the bitmap of free chunks for this block */
memset(slab->freechunks, 0, (slab->chunksPerBlock * sizeof(bool)));
- idx = block->firstFreeChunk;
/*
* Now walk through the chunks, count the free ones and also
@@ -744,20 +869,31 @@ SlabCheck(MemoryContext context)
* freelist is stored within the chunks themselves, we have to
* walk through the chunks and construct our own bitmap.
*/
-
+ cur_chunk = block->firstFreeChunk;
nfree = 0;
- while (idx < slab->chunksPerBlock)
+ while (cur_chunk != NULL)
{
- SlabChunk *chunk;
+ int idx = SlabChunkIndex(slab, block, cur_chunk);
/* count the chunk as free, add it to the bitmap */
nfree++;
slab->freechunks[idx] = true;
/* read index of the next free chunk */
- chunk = SlabBlockGetChunk(slab, block, idx);
- VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(int32));
- idx = *(int32 *) SlabChunkGetPointer(chunk);
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(cur_chunk), sizeof(SlabChunk **));
+ cur_chunk = *(SlabChunk **) SlabChunkGetPointer(cur_chunk);
+ }
+
+ cur_chunk = block->unused;
+ for (j = 0; j < block->nunused; j++)
+ {
+ int idx = SlabChunkIndex(slab, block, cur_chunk);
+
+ /* count the chunk as free, add it to the bitmap */
+ nfree++;
+ slab->freechunks[idx] = true;
+
+ cur_chunk = (SlabChunk *)(((char *) cur_chunk) + slab->fullChunkSize);
}
for (j = 0; j < slab->chunksPerBlock; j++)
--
2.31.1
[application/vnd.oasis.opendocument.spreadsheet] slab-results.ods (498.5K, ../../6717ca90-a2ea-b17d-f544-37f5181b1175@enterprisedb.com/5-slab-results.ods)
download
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
@ 2021-08-01 21:07 ` Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
0 siblings, 1 reply; 29+ messages in thread
From: Andres Freund @ 2021-08-01 21:07 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@enterprisedb.com>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Hi,
On 2021-08-01 19:59:18 +0200, Tomas Vondra wrote:
> In the attached .ods file with results, the "comparison" sheets are the
> interesting ones - the last couple columns compare the main metrics for the
> two patches (labeled patch-1 and patch-2) to master.
I assume with patch-1/2 you mean the ones after the benchmark patch
itself?
> Overall, the results look quite good - patch-1 is mostly on par with master,
> with maybe 5% variability in both directions. That's expected, considering
> the patch does not aim to improve performance.
Not for slab anyway...
> The second patch brings some nice improvements - 30%-50% in most cases (for
> both allocation and free) seems pretty nice. But for the "increase" FIFO
> pattern (incrementally allocating/freeing more memory) there's a significant
> regression - particularly for the allocation time. In some cases (larger
> chunks, block size does not matter too much) it jumps from 25ms to almost
> 200ms.
I'm not surprised to see some memory usage increase some, but that
degree of time overhead does surprise me. ISTM there's something wrong.
It'd probably worth benchmarking the different improvements inside the
WIP: slab performance. patch. There's some that I'd expect to be all
around improvements, whereas others likely aren't quite that clear
cut. I assume you'd prefer that I split the patch up?
> This seems unfortunate - the allocation pattern (FIFO, allocating more
> memory over time) seems pretty common, and the slowdown is significant.
Did you analyze what causes the regressions?
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
@ 2021-08-01 22:01 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
0 siblings, 1 reply; 29+ messages in thread
From: Tomas Vondra @ 2021-08-01 22:01 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On 8/1/21 11:07 PM, Andres Freund wrote:
> Hi,
>
> On 2021-08-01 19:59:18 +0200, Tomas Vondra wrote:
>> In the attached .ods file with results, the "comparison" sheets are the
>> interesting ones - the last couple columns compare the main metrics for the
>> two patches (labeled patch-1 and patch-2) to master.
>
> I assume with patch-1/2 you mean the ones after the benchmark patch
> itself?
>
Yes, those are the two WIP patches you shared on 19/7.
>
>> Overall, the results look quite good - patch-1 is mostly on par with master,
>> with maybe 5% variability in both directions. That's expected, considering
>> the patch does not aim to improve performance.
>
> Not for slab anyway...
>
Maybe the hot/cold separation could have some effect, but probably not
for the workloads I've tested.
>
>> The second patch brings some nice improvements - 30%-50% in most cases (for
>> both allocation and free) seems pretty nice. But for the "increase" FIFO
>> pattern (incrementally allocating/freeing more memory) there's a significant
>> regression - particularly for the allocation time. In some cases (larger
>> chunks, block size does not matter too much) it jumps from 25ms to almost
>> 200ms.
>
> I'm not surprised to see some memory usage increase some, but that
> degree of time overhead does surprise me. ISTM there's something wrong.
>
Yeah, the higher amount of allocated memory is due to the couple fields
added to the SlabBlock struct, but even that only affects a single case
with 480B chunks and 1kB blocks. Seems fine to me, especially if we end
up growing the slab blocks.
Not sure about the allocation time, though.
> It'd probably worth benchmarking the different improvements inside the
> WIP: slab performance. patch. There's some that I'd expect to be all
> around improvements, whereas others likely aren't quite that clear
> cut. I assume you'd prefer that I split the patch up?
>
Yeah, if you split that patch into sensible parts, I'll benchmark those.
Also, we can add more interesting workloads if you have some ideas.
>
>> This seems unfortunate - the allocation pattern (FIFO, allocating more
>> memory over time) seems pretty common, and the slowdown is significant.
>
> Did you analyze what causes the regressions?
>
No, not yet. I'll run the same set of benchmarks for the Generation,
discussed in the other thread, and then I'll investigate this. But if
you split the patch, that'd probably help.
regards
--
Tomas Vondra
EnterpriseDB: http://www.enterprisedb.com
The Enterprise PostgreSQL Company
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
@ 2021-08-03 13:33 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
0 siblings, 1 reply; 29+ messages in thread
From: Tomas Vondra @ 2021-08-03 13:33 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
FWIW I tried running the benchmarks again, with some minor changes in
the extension code - most importantly, the time is counted in microsecs
(instead of milisecs).
I suspected the rounding might have been causing some rounding errors
(essentially not counting anything below 1ms, because it rounds to 0),
and the results are a bit different.
On the i5-2500k machine it's an improvement across the board, while on
the bigger Xeon e5-2620v3 machine it shows roughly the same regression
for the "decreasing" allocation pattern.
There's another issue in the benchmarking script - the queries are meant
to do multiple runs for each combination of parameters, but it's written
in a way that simply runs it once and then does cross product with the
generate_sequence(1,5). I'll look into fixing that, but judging by the
stability of results for similar chunk sizes it won't change much.
regards
--
Tomas Vondra
EnterpriseDB: http://www.enterprisedb.com
The Enterprise PostgreSQL Company
Attachments:
[text/x-patch] 0001-slab-bench-v2.patch (17.3K, ../../5d83ac71-9054-208d-a48a-24d9d6b61f55@enterprisedb.com/2-0001-slab-bench-v2.patch)
download | inline diff:
From d71310e248bb5224a35d90a691ba8d2b58a96b02 Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas.vondra@postgresql.org>
Date: Sat, 31 Jul 2021 22:55:36 +0200
Subject: [PATCH 1/3] slab bench
---
contrib/slab_bench/.gitignore | 4 +
contrib/slab_bench/Makefile | 21 ++
contrib/slab_bench/bench.sql | 40 +++
contrib/slab_bench/slab_bench--1.0.sql | 16 +
contrib/slab_bench/slab_bench.c | 411 +++++++++++++++++++++++++
contrib/slab_bench/slab_bench.control | 4 +
6 files changed, 496 insertions(+)
create mode 100644 contrib/slab_bench/.gitignore
create mode 100644 contrib/slab_bench/Makefile
create mode 100644 contrib/slab_bench/bench.sql
create mode 100644 contrib/slab_bench/slab_bench--1.0.sql
create mode 100644 contrib/slab_bench/slab_bench.c
create mode 100644 contrib/slab_bench/slab_bench.control
diff --git a/contrib/slab_bench/.gitignore b/contrib/slab_bench/.gitignore
new file mode 100644
index 0000000000..5dcb3ff972
--- /dev/null
+++ b/contrib/slab_bench/.gitignore
@@ -0,0 +1,4 @@
+# Generated subdirectories
+/log/
+/results/
+/tmp_check/
diff --git a/contrib/slab_bench/Makefile b/contrib/slab_bench/Makefile
new file mode 100644
index 0000000000..c423bbd4ca
--- /dev/null
+++ b/contrib/slab_bench/Makefile
@@ -0,0 +1,21 @@
+# contrib/slab_bench/Makefile
+
+MODULE_big = slab_bench
+OBJS = slab_bench.o
+
+EXTENSION = slab_bench
+DATA = slab_bench--1.0.sql
+PGFILEDESC = "slab_bench - slab context benchmarking functions"
+
+REGRESS = slab_bench
+
+ifdef USE_PGXS
+PG_CONFIG = pg_config
+PGXS := $(shell $(PG_CONFIG) --pgxs)
+include $(PGXS)
+else
+subdir = contrib/slab_bench
+top_builddir = ../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+endif
diff --git a/contrib/slab_bench/bench.sql b/contrib/slab_bench/bench.sql
new file mode 100644
index 0000000000..d476a386c2
--- /dev/null
+++ b/contrib/slab_bench/bench.sql
@@ -0,0 +1,40 @@
+CREATE EXTENSION slab_bench;
+
+\o fifo-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+\o lifo-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+\o random-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+
+\o fifo-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+\o lifo-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+\o random-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+
+\o fifo-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+\o lifo-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+\o random-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+
+\o fifo-cycle.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000, block_size, chunk_size, 10000, 1000, 1000) x;
+
+\o lifo-cycle.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000, block_size, chunk_size, 10000, 1000, 1000) x;
+
+\o random-cycle.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000, block_size, chunk_size, 10000, 1000, 1000) x;
diff --git a/contrib/slab_bench/slab_bench--1.0.sql b/contrib/slab_bench/slab_bench--1.0.sql
new file mode 100644
index 0000000000..0a23134d74
--- /dev/null
+++ b/contrib/slab_bench/slab_bench--1.0.sql
@@ -0,0 +1,16 @@
+/* slab_bench--1.0.sql */
+
+-- complain if script is sourced in psql, rather than via CREATE EXTENSION
+\echo Use "CREATE EXTENSION slab_bench" to load this file. \quit
+
+CREATE FUNCTION slab_bench_random(nallocs bigint, block_size bigint, alloc_size bigint, loops int, free_cnt int, alloc_cnt int, out mem_allocated bigint, out alloc_ms bigint, out free_ms bigint)
+AS 'MODULE_PATHNAME', 'slab_bench_random'
+LANGUAGE C VOLATILE STRICT;
+
+CREATE FUNCTION slab_bench_fifo(nallocs bigint, block_size bigint, alloc_size bigint, loops int, free_cnt int, alloc_cnt int, out mem_allocated bigint, out alloc_ms bigint, out free_ms bigint)
+AS 'MODULE_PATHNAME', 'slab_bench_fifo'
+LANGUAGE C VOLATILE STRICT;
+
+CREATE FUNCTION slab_bench_lifo(nallocs bigint, block_size bigint, alloc_size bigint, loops int, free_cnt int, alloc_cnt int, out mem_allocated bigint, out alloc_ms bigint, out free_ms bigint)
+AS 'MODULE_PATHNAME', 'slab_bench_lifo'
+LANGUAGE C VOLATILE STRICT;
diff --git a/contrib/slab_bench/slab_bench.c b/contrib/slab_bench/slab_bench.c
new file mode 100644
index 0000000000..97ac10d7c3
--- /dev/null
+++ b/contrib/slab_bench/slab_bench.c
@@ -0,0 +1,411 @@
+/*-------------------------------------------------------------------------
+ *
+ * slab_bench.c
+ *
+ * helper functions to benchmark slab context with different workloads
+ *-------------------------------------------------------------------------
+ */
+#include "postgres.h"
+
+#include <sys/time.h>
+
+#include "funcapi.h"
+#include "miscadmin.h"
+
+PG_MODULE_MAGIC;
+
+PG_FUNCTION_INFO_V1(slab_bench_random);
+PG_FUNCTION_INFO_V1(slab_bench_fifo);
+PG_FUNCTION_INFO_V1(slab_bench_lifo);
+
+typedef struct Chunk {
+ int random;
+ void *ptr;
+} Chunk;
+
+static int
+chunk_index_cmp(const void *a, const void *b)
+{
+ Chunk *ca = (Chunk *) a;
+ Chunk *cb = (Chunk *) b;
+
+ if (ca->random < cb->random)
+ return -1;
+ else if (ca->random > cb->random)
+ return 1;
+
+ return 0;
+}
+
+Datum
+slab_bench_random(PG_FUNCTION_ARGS)
+{
+ MemoryContext cxt,
+ oldcxt;
+ Chunk *chunks;
+ int64 i, j;
+ int64 nallocs = PG_GETARG_INT64(0);
+ int64 blockSize = PG_GETARG_INT64(1);
+ int64 chunkSize = PG_GETARG_INT64(2);
+
+ int nloops = PG_GETARG_INT32(3);
+ int free_cnt = PG_GETARG_INT32(4);
+ int alloc_cnt = PG_GETARG_INT32(5);
+
+ struct timeval start_time,
+ end_time;
+ int64 alloc_time = 0,
+ free_time = 0;
+ int64 mem_allocated;
+
+ TupleDesc tupdesc;
+ Datum result;
+ HeapTuple tuple;
+ Datum values[9];
+ bool nulls[9];
+
+ int maxchunks;
+
+ maxchunks = nallocs + nloops * Max(0, alloc_cnt - free_cnt);
+
+ cxt = SlabContextCreate(CurrentMemoryContext, "slab_bench", blockSize, chunkSize);
+
+ chunks = (Chunk *) palloc(maxchunks * sizeof(Chunk));
+
+ /* allocate the chunks in random order */
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ gettimeofday(&start_time, NULL);
+
+ for (i = 0; i < nallocs; i++)
+ chunks[i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = MemoryContextMemAllocated(cxt, true);
+
+ /* do the requested number of free/alloc loops */
+ for (j = 0; j < nloops; j++)
+ {
+ CHECK_FOR_INTERRUPTS();
+
+ /* randomize the indexes */
+ for (i = 0; i < nallocs; i++)
+ chunks[i].random = random();
+
+ qsort(chunks, nallocs, sizeof(Chunk), chunk_index_cmp);
+
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 0; i < Min(nallocs, free_cnt); i++)
+ pfree(chunks[i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs -= Min(nallocs, free_cnt);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ memmove(chunks, &chunks[free_cnt], nallocs * sizeof(Chunk));
+
+
+ /* allocate alloc_cnt chunks at the end */
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 0; i < alloc_cnt; i++)
+ chunks[nallocs + i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs += alloc_cnt;
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = Max(mem_allocated, MemoryContextMemAllocated(cxt, true));
+ }
+
+ /* release the chunks in random order */
+ for (i = 0; i < nallocs; i++)
+ chunks[i].random = random();
+
+ qsort(chunks, nallocs, sizeof(Chunk), chunk_index_cmp);
+
+ gettimeofday(&start_time, NULL);
+
+ for (i = 0; i < nallocs; i++)
+ pfree(chunks[i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+
+ values[0] = Int64GetDatum(mem_allocated);
+ values[1] = Int64GetDatum(alloc_time);
+ values[2] = Int64GetDatum(free_time);
+
+ memset(nulls, 0, sizeof(nulls));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ result = HeapTupleGetDatum(tuple);
+
+ PG_RETURN_DATUM(result);
+}
+
+Datum
+slab_bench_fifo(PG_FUNCTION_ARGS)
+{
+ MemoryContext cxt,
+ oldcxt;
+ Chunk *chunks;
+ int64 i, j;
+ int64 nallocs = PG_GETARG_INT64(0);
+ int64 blockSize = PG_GETARG_INT64(1);
+ int64 chunkSize = PG_GETARG_INT64(2);
+
+ int nloops = PG_GETARG_INT32(3);
+ int free_cnt = PG_GETARG_INT32(4);
+ int alloc_cnt = PG_GETARG_INT32(5);
+
+ struct timeval start_time,
+ end_time;
+ int64 alloc_time = 0,
+ free_time = 0;
+ int64 mem_allocated;
+
+ TupleDesc tupdesc;
+ Datum result;
+ HeapTuple tuple;
+ Datum values[9];
+ bool nulls[9];
+
+ int maxchunks;
+
+ maxchunks = nallocs + nloops * Max(0, alloc_cnt - free_cnt);
+
+ cxt = SlabContextCreate(CurrentMemoryContext, "slab_bench", blockSize, chunkSize);
+
+ chunks = (Chunk *) palloc(maxchunks * sizeof(Chunk));
+
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ gettimeofday(&start_time, NULL);
+
+ for (i = 0; i < nallocs; i++)
+ chunks[i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = MemoryContextMemAllocated(cxt, true);
+
+
+ /* do the requested number of free/alloc loops */
+ for (j = 0; j < nloops; j++)
+ {
+ CHECK_FOR_INTERRUPTS();
+
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 0; i < Min(nallocs, free_cnt); i++)
+ pfree(chunks[i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs -= Min(nallocs, free_cnt);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ memmove(chunks, &chunks[free_cnt], nallocs * sizeof(Chunk));
+
+ /* allocate alloc_cnt chunks at the end */
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 0; i < alloc_cnt; i++)
+ chunks[nallocs + i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs += alloc_cnt;
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = Max(mem_allocated, MemoryContextMemAllocated(cxt, true));
+ }
+
+
+ gettimeofday(&start_time, NULL);
+
+ for (i = 0; i < nallocs; i++)
+ pfree(chunks[i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+
+ values[0] = Int64GetDatum(mem_allocated);
+ values[1] = Int64GetDatum(alloc_time);
+ values[2] = Int64GetDatum(free_time);
+
+ memset(nulls, 0, sizeof(nulls));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ result = HeapTupleGetDatum(tuple);
+
+ PG_RETURN_DATUM(result);
+}
+
+Datum
+slab_bench_lifo(PG_FUNCTION_ARGS)
+{
+ MemoryContext cxt,
+ oldcxt;
+ Chunk *chunks;
+ int64 i, j;
+ int64 nallocs = PG_GETARG_INT64(0);
+ int64 blockSize = PG_GETARG_INT64(1);
+ int64 chunkSize = PG_GETARG_INT64(2);
+
+ int nloops = PG_GETARG_INT32(3);
+ int free_cnt = PG_GETARG_INT32(4);
+ int alloc_cnt = PG_GETARG_INT32(5);
+
+ struct timeval start_time,
+ end_time;
+ int64 alloc_time = 0,
+ free_time = 0;
+ int64 mem_allocated;
+
+ TupleDesc tupdesc;
+ Datum result;
+ HeapTuple tuple;
+ Datum values[9];
+ bool nulls[9];
+
+ int maxchunks;
+
+ maxchunks = nallocs + nloops * Max(0, alloc_cnt - free_cnt);
+
+ cxt = SlabContextCreate(CurrentMemoryContext, "slab_bench", blockSize, chunkSize);
+
+ chunks = (Chunk *) palloc(maxchunks * sizeof(Chunk));
+
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ /* palloc benchmark */
+ gettimeofday(&start_time, NULL);
+
+ for (i = 0; i < nallocs; i++)
+ chunks[i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = MemoryContextMemAllocated(cxt, true);
+
+
+ /* do the requested number of free/alloc loops */
+ for (j = 0; j < nloops; j++)
+ {
+ CHECK_FOR_INTERRUPTS();
+
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 1; i <= Min(nallocs, free_cnt); i++)
+ pfree(chunks[nallocs - i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs -= Min(nallocs, free_cnt);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ /* allocate alloc_cnt chunks at the end */
+ gettimeofday(&start_time, NULL);
+
+ /* free the first free_cnt chunks */
+ for (i = 0; i < alloc_cnt; i++)
+ chunks[nallocs + i].ptr = palloc(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ nallocs += alloc_cnt;
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ MemoryContextSwitchTo(oldcxt);
+
+ mem_allocated = Max(mem_allocated, MemoryContextMemAllocated(cxt, true));
+ }
+
+
+ gettimeofday(&start_time, NULL);
+
+ for (i = (nallocs - 1); i >= 0; i--)
+ pfree(chunks[i].ptr);
+
+ gettimeofday(&end_time, NULL);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+
+ values[0] = Int64GetDatum(mem_allocated);
+ values[1] = Int64GetDatum(alloc_time);
+ values[2] = Int64GetDatum(free_time);
+
+ memset(nulls, 0, sizeof(nulls));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ result = HeapTupleGetDatum(tuple);
+
+ MemoryContextDelete(cxt);
+
+ PG_RETURN_DATUM(result);
+}
diff --git a/contrib/slab_bench/slab_bench.control b/contrib/slab_bench/slab_bench.control
new file mode 100644
index 0000000000..290809c6f7
--- /dev/null
+++ b/contrib/slab_bench/slab_bench.control
@@ -0,0 +1,4 @@
+# slab_bench extension
+comment = 'functions for benchmarking slab context'
+default_version = '1.0'
+module_pathname = '$libdir/slab_bench'
--
2.31.1
[text/x-patch] 0002-WIP-optimize-allocations-by-separating-hot-from-c-v2.patch (31.7K, ../../5d83ac71-9054-208d-a48a-24d9d6b61f55@enterprisedb.com/3-0002-WIP-optimize-allocations-by-separating-hot-from-c-v2.patch)
download | inline diff:
From b46dbb73869067a08746d0aa3bfbdb3a2835a992 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 19 Jul 2021 12:55:51 -0700
Subject: [PATCH 2/3] WIP: optimize allocations by separating hot from cold
paths.
---
src/backend/utils/mmgr/aset.c | 429 ++++++++++++++--------------
src/backend/utils/mmgr/generation.c | 22 +-
src/backend/utils/mmgr/mcxt.c | 179 +++---------
src/backend/utils/mmgr/slab.c | 14 +-
src/include/nodes/memnodes.h | 4 +-
src/include/utils/memutils.h | 13 +
6 files changed, 300 insertions(+), 361 deletions(-)
diff --git a/src/backend/utils/mmgr/aset.c b/src/backend/utils/mmgr/aset.c
index 77872e77bc..0087835439 100644
--- a/src/backend/utils/mmgr/aset.c
+++ b/src/backend/utils/mmgr/aset.c
@@ -263,9 +263,9 @@ static AllocSetFreeList context_freelists[2] =
/*
* These functions implement the MemoryContext API for AllocSet contexts.
*/
-static void *AllocSetAlloc(MemoryContext context, Size size);
+static void *AllocSetAlloc(MemoryContext context, Size size, int flags);
static void AllocSetFree(MemoryContext context, void *pointer);
-static void *AllocSetRealloc(MemoryContext context, void *pointer, Size size);
+static void *AllocSetRealloc(MemoryContext context, void *pointer, Size size, int flags);
static void AllocSetReset(MemoryContext context);
static void AllocSetDelete(MemoryContext context);
static Size AllocSetGetChunkSpace(MemoryContext context, void *pointer);
@@ -704,6 +704,208 @@ AllocSetDelete(MemoryContext context)
free(set);
}
+static inline void *
+AllocSetAllocReturnChunk(AllocSet set, Size size, AllocChunk chunk, Size chunk_size)
+{
+ chunk->aset = (void *) set;
+#ifdef MEMORY_CONTEXT_CHECKING
+ chunk->requested_size = size;
+ /* set mark to catch clobber of "unused" space */
+ if (size < chunk->size)
+ set_sentinel(AllocChunkGetPointer(chunk), size);
+#endif
+#ifdef RANDOMIZE_ALLOCATED_MEMORY
+ /* fill the allocated space with junk */
+ randomize_mem((char *) AllocChunkGetPointer(chunk), size);
+#endif
+
+ /* Ensure any padding bytes are marked NOACCESS. */
+ VALGRIND_MAKE_MEM_NOACCESS((char *) AllocChunkGetPointer(chunk) + size,
+ chunk_size - size);
+
+ /* Disallow external access to private part of chunk header. */
+ VALGRIND_MAKE_MEM_NOACCESS(chunk, ALLOCCHUNK_PRIVATE_LEN);
+
+ return AllocChunkGetPointer(chunk);
+}
+
+static void * pg_noinline
+AllocSetAllocLarge(AllocSet set, Size size, int flags)
+{
+ AllocBlock block;
+ AllocChunk chunk;
+ Size chunk_size;
+ Size blksize;
+
+ /* check size, only allocation path where the limits could be hit */
+ MemoryContextCheckSize(&set->header, size, flags);
+
+ AssertArg(AllocSetIsValid(set));
+
+ chunk_size = MAXALIGN(size);
+ blksize = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
+ block = (AllocBlock) malloc(blksize);
+ if (block == NULL)
+ return NULL;
+
+ set->header.mem_allocated += blksize;
+
+ block->aset = set;
+ block->freeptr = block->endptr = ((char *) block) + blksize;
+
+ /*
+ * Stick the new block underneath the active allocation block, if any,
+ * so that we don't lose the use of the space remaining therein.
+ */
+ if (set->blocks != NULL)
+ {
+ block->prev = set->blocks;
+ block->next = set->blocks->next;
+ if (block->next)
+ block->next->prev = block;
+ set->blocks->next = block;
+ }
+ else
+ {
+ block->prev = NULL;
+ block->next = NULL;
+ set->blocks = block;
+ }
+
+ chunk = (AllocChunk) (((char *) block) + ALLOC_BLOCKHDRSZ);
+ chunk->size = chunk_size;
+
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk_size);
+}
+
+static void * pg_noinline
+AllocSetAllocFromNewBlock(AllocSet set, Size size, Size chunk_size)
+{
+ AllocBlock block;
+ Size blksize;
+ Size required_size;
+ AllocChunk chunk;
+
+ /*
+ * The first such block has size initBlockSize, and we double the
+ * space in each succeeding block, but not more than maxBlockSize.
+ */
+ blksize = set->nextBlockSize;
+ set->nextBlockSize <<= 1;
+ if (set->nextBlockSize > set->maxBlockSize)
+ set->nextBlockSize = set->maxBlockSize;
+
+ /*
+ * If initBlockSize is less than ALLOC_CHUNK_LIMIT, we could need more
+ * space... but try to keep it a power of 2.
+ */
+ required_size = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
+ while (blksize < required_size)
+ blksize <<= 1;
+
+ /* Try to allocate it */
+ block = (AllocBlock) malloc(blksize);
+
+ /*
+ * We could be asking for pretty big blocks here, so cope if malloc
+ * fails. But give up if there's less than 1 MB or so available...
+ */
+ while (block == NULL && blksize > 1024 * 1024)
+ {
+ blksize >>= 1;
+ if (blksize < required_size)
+ break;
+ block = (AllocBlock) malloc(blksize);
+ }
+
+ if (block == NULL)
+ return NULL;
+
+ set->header.mem_allocated += blksize;
+
+ block->aset = set;
+ block->freeptr = ((char *) block) + ALLOC_BLOCKHDRSZ;
+ block->endptr = ((char *) block) + blksize;
+
+ /* Mark unallocated space NOACCESS. */
+ VALGRIND_MAKE_MEM_NOACCESS(block->freeptr,
+ blksize - ALLOC_BLOCKHDRSZ);
+
+ block->prev = NULL;
+ block->next = set->blocks;
+ if (block->next)
+ block->next->prev = block;
+ set->blocks = block;
+
+ /*
+ * OK, do the allocation
+ */
+ chunk = (AllocChunk) (block->freeptr);
+
+ /* Prepare to initialize the chunk header. */
+ VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
+
+ block->freeptr += (chunk_size + ALLOC_CHUNKHDRSZ);
+ Assert(block->freeptr <= block->endptr);
+
+ chunk->size = chunk_size;
+
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk_size);
+}
+
+static void * pg_noinline
+AllocSetAllocCarveOldAndAlloc(AllocSet set, Size size, Size chunk_size, AllocBlock block, Size availspace)
+{
+ AllocChunk chunk;
+
+ /*
+ * The existing active (top) block does not have enough room for
+ * the requested allocation, but it might still have a useful
+ * amount of space in it. Once we push it down in the block list,
+ * we'll never try to allocate more space from it. So, before we
+ * do that, carve up its free space into chunks that we can put on
+ * the set's freelists.
+ *
+ * Because we can only get here when there's less than
+ * ALLOC_CHUNK_LIMIT left in the block, this loop cannot iterate
+ * more than ALLOCSET_NUM_FREELISTS-1 times.
+ */
+ while (availspace >= ((1 << ALLOC_MINBITS) + ALLOC_CHUNKHDRSZ))
+ {
+ Size availchunk = availspace - ALLOC_CHUNKHDRSZ;
+ int a_fidx = AllocSetFreeIndex(availchunk);
+
+ /*
+ * In most cases, we'll get back the index of the next larger
+ * freelist than the one we need to put this chunk on. The
+ * exception is when availchunk is exactly a power of 2.
+ */
+ if (availchunk != ((Size) 1 << (a_fidx + ALLOC_MINBITS)))
+ {
+ a_fidx--;
+ Assert(a_fidx >= 0);
+ availchunk = ((Size) 1 << (a_fidx + ALLOC_MINBITS));
+ }
+
+ chunk = (AllocChunk) (block->freeptr);
+
+ /* Prepare to initialize the chunk header. */
+ VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
+
+ block->freeptr += (availchunk + ALLOC_CHUNKHDRSZ);
+ availspace -= (availchunk + ALLOC_CHUNKHDRSZ);
+
+ chunk->size = availchunk;
+#ifdef MEMORY_CONTEXT_CHECKING
+ chunk->requested_size = 0; /* mark it free */
+#endif
+ chunk->aset = (void *) set->freelist[a_fidx];
+ set->freelist[a_fidx] = chunk;
+ }
+
+ return AllocSetAllocFromNewBlock(set, size, chunk_size);
+}
+
/*
* AllocSetAlloc
* Returns pointer to allocated memory of given size or NULL if
@@ -718,14 +920,13 @@ AllocSetDelete(MemoryContext context)
* return space that is marked NOACCESS - AllocSetRealloc has to beware!
*/
static void *
-AllocSetAlloc(MemoryContext context, Size size)
+AllocSetAlloc(MemoryContext context, Size size, int flags)
{
AllocSet set = (AllocSet) context;
AllocBlock block;
AllocChunk chunk;
int fidx;
Size chunk_size;
- Size blksize;
AssertArg(AllocSetIsValid(set));
@@ -733,61 +934,8 @@ AllocSetAlloc(MemoryContext context, Size size)
* If requested size exceeds maximum for chunks, allocate an entire block
* for this request.
*/
- if (size > set->allocChunkLimit)
- {
- chunk_size = MAXALIGN(size);
- blksize = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
- block = (AllocBlock) malloc(blksize);
- if (block == NULL)
- return NULL;
-
- context->mem_allocated += blksize;
-
- block->aset = set;
- block->freeptr = block->endptr = ((char *) block) + blksize;
-
- chunk = (AllocChunk) (((char *) block) + ALLOC_BLOCKHDRSZ);
- chunk->aset = set;
- chunk->size = chunk_size;
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = size;
- /* set mark to catch clobber of "unused" space */
- if (size < chunk_size)
- set_sentinel(AllocChunkGetPointer(chunk), size);
-#endif
-#ifdef RANDOMIZE_ALLOCATED_MEMORY
- /* fill the allocated space with junk */
- randomize_mem((char *) AllocChunkGetPointer(chunk), size);
-#endif
-
- /*
- * Stick the new block underneath the active allocation block, if any,
- * so that we don't lose the use of the space remaining therein.
- */
- if (set->blocks != NULL)
- {
- block->prev = set->blocks;
- block->next = set->blocks->next;
- if (block->next)
- block->next->prev = block;
- set->blocks->next = block;
- }
- else
- {
- block->prev = NULL;
- block->next = NULL;
- set->blocks = block;
- }
-
- /* Ensure any padding bytes are marked NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS((char *) AllocChunkGetPointer(chunk) + size,
- chunk_size - size);
-
- /* Disallow external access to private part of chunk header. */
- VALGRIND_MAKE_MEM_NOACCESS(chunk, ALLOCCHUNK_PRIVATE_LEN);
-
- return AllocChunkGetPointer(chunk);
- }
+ if (unlikely(size > set->allocChunkLimit))
+ return AllocSetAllocLarge(set, size, flags);
/*
* Request is small enough to be treated as a chunk. Look in the
@@ -803,27 +951,7 @@ AllocSetAlloc(MemoryContext context, Size size)
set->freelist[fidx] = (AllocChunk) chunk->aset;
- chunk->aset = (void *) set;
-
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = size;
- /* set mark to catch clobber of "unused" space */
- if (size < chunk->size)
- set_sentinel(AllocChunkGetPointer(chunk), size);
-#endif
-#ifdef RANDOMIZE_ALLOCATED_MEMORY
- /* fill the allocated space with junk */
- randomize_mem((char *) AllocChunkGetPointer(chunk), size);
-#endif
-
- /* Ensure any padding bytes are marked NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS((char *) AllocChunkGetPointer(chunk) + size,
- chunk->size - size);
-
- /* Disallow external access to private part of chunk header. */
- VALGRIND_MAKE_MEM_NOACCESS(chunk, ALLOCCHUNK_PRIVATE_LEN);
-
- return AllocChunkGetPointer(chunk);
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk->size);
}
/*
@@ -840,115 +968,16 @@ AllocSetAlloc(MemoryContext context, Size size)
{
Size availspace = block->endptr - block->freeptr;
- if (availspace < (chunk_size + ALLOC_CHUNKHDRSZ))
- {
- /*
- * The existing active (top) block does not have enough room for
- * the requested allocation, but it might still have a useful
- * amount of space in it. Once we push it down in the block list,
- * we'll never try to allocate more space from it. So, before we
- * do that, carve up its free space into chunks that we can put on
- * the set's freelists.
- *
- * Because we can only get here when there's less than
- * ALLOC_CHUNK_LIMIT left in the block, this loop cannot iterate
- * more than ALLOCSET_NUM_FREELISTS-1 times.
- */
- while (availspace >= ((1 << ALLOC_MINBITS) + ALLOC_CHUNKHDRSZ))
- {
- Size availchunk = availspace - ALLOC_CHUNKHDRSZ;
- int a_fidx = AllocSetFreeIndex(availchunk);
-
- /*
- * In most cases, we'll get back the index of the next larger
- * freelist than the one we need to put this chunk on. The
- * exception is when availchunk is exactly a power of 2.
- */
- if (availchunk != ((Size) 1 << (a_fidx + ALLOC_MINBITS)))
- {
- a_fidx--;
- Assert(a_fidx >= 0);
- availchunk = ((Size) 1 << (a_fidx + ALLOC_MINBITS));
- }
-
- chunk = (AllocChunk) (block->freeptr);
-
- /* Prepare to initialize the chunk header. */
- VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
-
- block->freeptr += (availchunk + ALLOC_CHUNKHDRSZ);
- availspace -= (availchunk + ALLOC_CHUNKHDRSZ);
-
- chunk->size = availchunk;
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = 0; /* mark it free */
-#endif
- chunk->aset = (void *) set->freelist[a_fidx];
- set->freelist[a_fidx] = chunk;
- }
-
- /* Mark that we need to create a new block */
- block = NULL;
- }
+ if (unlikely(availspace < (chunk_size + ALLOC_CHUNKHDRSZ)))
+ return AllocSetAllocCarveOldAndAlloc(set, size, chunk_size,
+ block, availspace);
}
-
- /*
- * Time to create a new regular (multi-chunk) block?
- */
- if (block == NULL)
+ else if (unlikely(block == NULL))
{
- Size required_size;
-
- /*
- * The first such block has size initBlockSize, and we double the
- * space in each succeeding block, but not more than maxBlockSize.
- */
- blksize = set->nextBlockSize;
- set->nextBlockSize <<= 1;
- if (set->nextBlockSize > set->maxBlockSize)
- set->nextBlockSize = set->maxBlockSize;
-
- /*
- * If initBlockSize is less than ALLOC_CHUNK_LIMIT, we could need more
- * space... but try to keep it a power of 2.
- */
- required_size = chunk_size + ALLOC_BLOCKHDRSZ + ALLOC_CHUNKHDRSZ;
- while (blksize < required_size)
- blksize <<= 1;
-
- /* Try to allocate it */
- block = (AllocBlock) malloc(blksize);
-
/*
- * We could be asking for pretty big blocks here, so cope if malloc
- * fails. But give up if there's less than 1 MB or so available...
+ * Time to create a new regular (multi-chunk) block.
*/
- while (block == NULL && blksize > 1024 * 1024)
- {
- blksize >>= 1;
- if (blksize < required_size)
- break;
- block = (AllocBlock) malloc(blksize);
- }
-
- if (block == NULL)
- return NULL;
-
- context->mem_allocated += blksize;
-
- block->aset = set;
- block->freeptr = ((char *) block) + ALLOC_BLOCKHDRSZ;
- block->endptr = ((char *) block) + blksize;
-
- /* Mark unallocated space NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS(block->freeptr,
- blksize - ALLOC_BLOCKHDRSZ);
-
- block->prev = NULL;
- block->next = set->blocks;
- if (block->next)
- block->next->prev = block;
- set->blocks = block;
+ return AllocSetAllocFromNewBlock(set, size, chunk_size);
}
/*
@@ -959,30 +988,12 @@ AllocSetAlloc(MemoryContext context, Size size)
/* Prepare to initialize the chunk header. */
VALGRIND_MAKE_MEM_UNDEFINED(chunk, ALLOC_CHUNKHDRSZ);
- block->freeptr += (chunk_size + ALLOC_CHUNKHDRSZ);
- Assert(block->freeptr <= block->endptr);
-
- chunk->aset = (void *) set;
chunk->size = chunk_size;
-#ifdef MEMORY_CONTEXT_CHECKING
- chunk->requested_size = size;
- /* set mark to catch clobber of "unused" space */
- if (size < chunk->size)
- set_sentinel(AllocChunkGetPointer(chunk), size);
-#endif
-#ifdef RANDOMIZE_ALLOCATED_MEMORY
- /* fill the allocated space with junk */
- randomize_mem((char *) AllocChunkGetPointer(chunk), size);
-#endif
-
- /* Ensure any padding bytes are marked NOACCESS. */
- VALGRIND_MAKE_MEM_NOACCESS((char *) AllocChunkGetPointer(chunk) + size,
- chunk_size - size);
- /* Disallow external access to private part of chunk header. */
- VALGRIND_MAKE_MEM_NOACCESS(chunk, ALLOCCHUNK_PRIVATE_LEN);
+ block->freeptr += (chunk_size + ALLOC_CHUNKHDRSZ);
+ Assert(block->freeptr <= block->endptr);
- return AllocChunkGetPointer(chunk);
+ return AllocSetAllocReturnChunk(set, size, chunk, chunk_size);
}
/*
@@ -1072,7 +1083,7 @@ AllocSetFree(MemoryContext context, void *pointer)
* request size.)
*/
static void *
-AllocSetRealloc(MemoryContext context, void *pointer, Size size)
+AllocSetRealloc(MemoryContext context, void *pointer, Size size, int flags)
{
AllocSet set = (AllocSet) context;
AllocChunk chunk = AllocPointerGetChunk(pointer);
@@ -1081,6 +1092,8 @@ AllocSetRealloc(MemoryContext context, void *pointer, Size size)
/* Allow access to private part of chunk header. */
VALGRIND_MAKE_MEM_DEFINED(chunk, ALLOCCHUNK_PRIVATE_LEN);
+ MemoryContextCheckSize(context, size, flags);
+
oldsize = chunk->size;
#ifdef MEMORY_CONTEXT_CHECKING
@@ -1260,7 +1273,7 @@ AllocSetRealloc(MemoryContext context, void *pointer, Size size)
AllocPointer newPointer;
/* allocate new chunk */
- newPointer = AllocSetAlloc((MemoryContext) set, size);
+ newPointer = AllocSetAlloc((MemoryContext) set, size, flags);
/* leave immediately if request was not completed */
if (newPointer == NULL)
diff --git a/src/backend/utils/mmgr/generation.c b/src/backend/utils/mmgr/generation.c
index 584cd614da..5dde15654d 100644
--- a/src/backend/utils/mmgr/generation.c
+++ b/src/backend/utils/mmgr/generation.c
@@ -146,9 +146,9 @@ struct GenerationChunk
/*
* These functions implement the MemoryContext API for Generation contexts.
*/
-static void *GenerationAlloc(MemoryContext context, Size size);
+static void *GenerationAlloc(MemoryContext context, Size size, int flags);
static void GenerationFree(MemoryContext context, void *pointer);
-static void *GenerationRealloc(MemoryContext context, void *pointer, Size size);
+static void *GenerationRealloc(MemoryContext context, void *pointer, Size size, int flags);
static void GenerationReset(MemoryContext context);
static void GenerationDelete(MemoryContext context);
static Size GenerationGetChunkSpace(MemoryContext context, void *pointer);
@@ -323,7 +323,7 @@ GenerationDelete(MemoryContext context)
* return space that is marked NOACCESS - GenerationRealloc has to beware!
*/
static void *
-GenerationAlloc(MemoryContext context, Size size)
+GenerationAlloc(MemoryContext context, Size size, int flags)
{
GenerationContext *set = (GenerationContext *) context;
GenerationBlock *block;
@@ -335,9 +335,12 @@ GenerationAlloc(MemoryContext context, Size size)
{
Size blksize = chunk_size + Generation_BLOCKHDRSZ + Generation_CHUNKHDRSZ;
+ /* check size, only allocation path where the limits could be hit */
+ MemoryContextCheckSize(context, size, flags);
+
block = (GenerationBlock *) malloc(blksize);
if (block == NULL)
- return NULL;
+ return MemoryContextAllocationFailure(context, size, flags);
context->mem_allocated += blksize;
@@ -392,7 +395,7 @@ GenerationAlloc(MemoryContext context, Size size)
block = (GenerationBlock *) malloc(blksize);
if (block == NULL)
- return NULL;
+ return MemoryContextAllocationFailure(context, size, flags);
context->mem_allocated += blksize;
@@ -520,13 +523,15 @@ GenerationFree(MemoryContext context, void *pointer)
* into the old chunk - in that case we just update chunk header.
*/
static void *
-GenerationRealloc(MemoryContext context, void *pointer, Size size)
+GenerationRealloc(MemoryContext context, void *pointer, Size size, int flags)
{
GenerationContext *set = (GenerationContext *) context;
GenerationChunk *chunk = GenerationPointerGetChunk(pointer);
GenerationPointer newPointer;
Size oldsize;
+ MemoryContextCheckSize(context, size, flags);
+
/* Allow access to private part of chunk header. */
VALGRIND_MAKE_MEM_DEFINED(chunk, GENERATIONCHUNK_PRIVATE_LEN);
@@ -596,14 +601,15 @@ GenerationRealloc(MemoryContext context, void *pointer, Size size)
}
/* allocate new chunk */
- newPointer = GenerationAlloc((MemoryContext) set, size);
+ newPointer = GenerationAlloc((MemoryContext) set, size, flags);
/* leave immediately if request was not completed */
if (newPointer == NULL)
{
/* Disallow external access to private part of chunk header. */
VALGRIND_MAKE_MEM_NOACCESS(chunk, GENERATIONCHUNK_PRIVATE_LEN);
- return NULL;
+ /* again? */
+ return MemoryContextAllocationFailure(context, size, flags);
}
/*
diff --git a/src/backend/utils/mmgr/mcxt.c b/src/backend/utils/mmgr/mcxt.c
index 6919a73280..13125c4a56 100644
--- a/src/backend/utils/mmgr/mcxt.c
+++ b/src/backend/utils/mmgr/mcxt.c
@@ -867,28 +867,9 @@ MemoryContextAlloc(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
-
- /*
- * Here, and elsewhere in this module, we show the target context's
- * "name" but not its "ident" (if any) in user-visible error messages.
- * The "ident" string might contain security-sensitive data, such as
- * values in SQL commands.
- */
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -910,21 +891,10 @@ MemoryContextAllocZero(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -948,21 +918,10 @@ MemoryContextAllocZeroAligned(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -983,26 +942,11 @@ MemoryContextAllocExtended(MemoryContext context, Size size, int flags)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (((flags & MCXT_ALLOC_HUGE) != 0 && !AllocHugeSizeIsValid(size)) ||
- ((flags & MCXT_ALLOC_HUGE) == 0 && !AllocSizeIsValid(size)))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
+ ret = context->methods->alloc(context, size, flags);
if (unlikely(ret == NULL))
- {
- if ((flags & MCXT_ALLOC_NO_OOM) == 0)
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
return NULL;
- }
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1067,22 +1011,10 @@ palloc(Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
-
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1099,21 +1031,10 @@ palloc0(Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1132,26 +1053,11 @@ palloc_extended(Size size, int flags)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (((flags & MCXT_ALLOC_HUGE) != 0 && !AllocHugeSizeIsValid(size)) ||
- ((flags & MCXT_ALLOC_HUGE) == 0 && !AllocSizeIsValid(size)))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
- ret = context->methods->alloc(context, size);
+ ret = context->methods->alloc(context, size, 0);
if (unlikely(ret == NULL))
- {
- if ((flags & MCXT_ALLOC_NO_OOM) == 0)
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
return NULL;
- }
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1184,24 +1090,13 @@ repalloc(void *pointer, Size size)
MemoryContext context = GetMemoryChunkContext(pointer);
void *ret;
- if (!AllocSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
AssertNotInCriticalSection(context);
/* isReset must be false already */
Assert(!context->isReset);
- ret = context->methods->realloc(context, pointer, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->realloc(context, pointer, size, 0);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_CHANGE(context, pointer, ret, size);
@@ -1222,21 +1117,9 @@ MemoryContextAllocHuge(MemoryContext context, Size size)
AssertArg(MemoryContextIsValid(context));
AssertNotInCriticalSection(context);
- if (!AllocHugeSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
context->isReset = false;
-
- ret = context->methods->alloc(context, size);
- if (unlikely(ret == NULL))
- {
- MemoryContextStats(TopMemoryContext);
- ereport(ERROR,
- (errcode(ERRCODE_OUT_OF_MEMORY),
- errmsg("out of memory"),
- errdetail("Failed on request of size %zu in memory context \"%s\".",
- size, context->name)));
- }
+ ret = context->methods->alloc(context, size, MCXT_ALLOC_HUGE);
+ Assert(ret != NULL);
VALGRIND_MEMPOOL_ALLOC(context, ret, size);
@@ -1254,16 +1137,23 @@ repalloc_huge(void *pointer, Size size)
MemoryContext context = GetMemoryChunkContext(pointer);
void *ret;
- if (!AllocHugeSizeIsValid(size))
- elog(ERROR, "invalid memory alloc request size %zu", size);
-
AssertNotInCriticalSection(context);
/* isReset must be false already */
Assert(!context->isReset);
- ret = context->methods->realloc(context, pointer, size);
- if (unlikely(ret == NULL))
+ ret = context->methods->realloc(context, pointer, size, MCXT_ALLOC_HUGE);
+ Assert(ret != NULL);
+
+ VALGRIND_MEMPOOL_CHANGE(context, pointer, ret, size);
+
+ return ret;
+}
+
+void *
+MemoryContextAllocationFailure(MemoryContext context, Size size, int flags)
+{
+ if ((flags & MCXT_ALLOC_NO_OOM) == 0)
{
MemoryContextStats(TopMemoryContext);
ereport(ERROR,
@@ -1273,9 +1163,13 @@ repalloc_huge(void *pointer, Size size)
size, context->name)));
}
- VALGRIND_MEMPOOL_CHANGE(context, pointer, ret, size);
+ return NULL;
+}
- return ret;
+void
+MemoryContextSizeFailure(MemoryContext context, Size size, int flags)
+{
+ elog(ERROR, "invalid memory alloc request size %zu", size);
}
/*
@@ -1298,6 +1192,13 @@ MemoryContextStrdup(MemoryContext context, const char *string)
char *
pstrdup(const char *in)
{
+ /*
+ * Here, and elsewhere in this module, we show the target context's
+ * "name" but not its "ident" (if any) in user-visible error messages.
+ * The "ident" string might contain security-sensitive data, such as
+ * values in SQL commands.
+ */
+
return MemoryContextStrdup(CurrentMemoryContext, in);
}
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index 553dd7f667..e4b8275045 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -126,9 +126,9 @@ typedef struct SlabChunk
/*
* These functions implement the MemoryContext API for Slab contexts.
*/
-static void *SlabAlloc(MemoryContext context, Size size);
+static void *SlabAlloc(MemoryContext context, Size size, int flags);
static void SlabFree(MemoryContext context, void *pointer);
-static void *SlabRealloc(MemoryContext context, void *pointer, Size size);
+static void *SlabRealloc(MemoryContext context, void *pointer, Size size, int flags);
static void SlabReset(MemoryContext context);
static void SlabDelete(MemoryContext context);
static Size SlabGetChunkSpace(MemoryContext context, void *pointer);
@@ -337,7 +337,7 @@ SlabDelete(MemoryContext context)
* request could not be completed; memory is added to the slab.
*/
static void *
-SlabAlloc(MemoryContext context, Size size)
+SlabAlloc(MemoryContext context, Size size, int flags)
{
SlabContext *slab = castNode(SlabContext, context);
SlabBlock *block;
@@ -346,6 +346,12 @@ SlabAlloc(MemoryContext context, Size size)
Assert(slab);
+ /*
+ * XXX: Probably no need to check for huge allocations, we only support
+ * one size? Which could theoretically be huge, but that'd not make
+ * sense...
+ */
+
Assert((slab->minFreeChunks >= 0) &&
(slab->minFreeChunks < slab->chunksPerBlock));
@@ -583,7 +589,7 @@ SlabFree(MemoryContext context, void *pointer)
* realloc is usually used to enlarge the chunk.
*/
static void *
-SlabRealloc(MemoryContext context, void *pointer, Size size)
+SlabRealloc(MemoryContext context, void *pointer, Size size, int flags)
{
SlabContext *slab = castNode(SlabContext, context);
diff --git a/src/include/nodes/memnodes.h b/src/include/nodes/memnodes.h
index e6a757d6a0..8a42d2ff99 100644
--- a/src/include/nodes/memnodes.h
+++ b/src/include/nodes/memnodes.h
@@ -57,10 +57,10 @@ typedef void (*MemoryStatsPrintFunc) (MemoryContext context, void *passthru,
typedef struct MemoryContextMethods
{
- void *(*alloc) (MemoryContext context, Size size);
+ void *(*alloc) (MemoryContext context, Size size, int flags);
/* call this free_p in case someone #define's free() */
void (*free_p) (MemoryContext context, void *pointer);
- void *(*realloc) (MemoryContext context, void *pointer, Size size);
+ void *(*realloc) (MemoryContext context, void *pointer, Size size, int flags);
void (*reset) (MemoryContext context);
void (*delete_context) (MemoryContext context);
Size (*get_chunk_space) (MemoryContext context, void *pointer);
diff --git a/src/include/utils/memutils.h b/src/include/utils/memutils.h
index ff872274d4..2f75b4cca4 100644
--- a/src/include/utils/memutils.h
+++ b/src/include/utils/memutils.h
@@ -147,6 +147,19 @@ extern void MemoryContextCreate(MemoryContext node,
extern void HandleLogMemoryContextInterrupt(void);
extern void ProcessLogMemoryContextInterrupt(void);
+extern void *MemoryContextAllocationFailure(MemoryContext context, Size size, int flags);
+
+extern void MemoryContextSizeFailure(MemoryContext context, Size size, int flags) pg_attribute_noreturn();
+
+static inline void
+MemoryContextCheckSize(MemoryContext context, Size size, int flags)
+{
+ if (unlikely(!AllocSizeIsValid(size)))
+ {
+ if (!(flags & MCXT_ALLOC_HUGE) || !AllocHugeSizeIsValid(size))
+ MemoryContextSizeFailure(context, size, flags);
+ }
+}
/*
* Memory-context-type-specific functions
--
2.31.1
[text/x-patch] 0003-WIP-slab-performance-v2.patch (23.4K, ../../5d83ac71-9054-208d-a48a-24d9d6b61f55@enterprisedb.com/4-0003-WIP-slab-performance-v2.patch)
download | inline diff:
From 2da723e692164fe136fc7d43130b111fd98a756a Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Mon, 19 Jul 2021 13:48:14 -0700
Subject: [PATCH 3/3] WIP: slab performance.
Discussion: https://postgr.es/m/20210717194333.mr5io3zup3kxahfm@alap3.anarazel.de
---
src/backend/utils/mmgr/slab.c | 434 ++++++++++++++++++++++------------
1 file changed, 285 insertions(+), 149 deletions(-)
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index e4b8275045..c979553dcd 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -23,15 +23,18 @@
* global (context) level. This is possible as the chunk size (and thus also
* the number of chunks per block) is fixed.
*
- * On each block, free chunks are tracked in a simple linked list. Contents
- * of free chunks is replaced with an index of the next free chunk, forming
- * a very simple linked list. Each block also contains a counter of free
- * chunks. Combined with the local block-level freelist, it makes it trivial
- * to eventually free the whole block.
+ * On each block, never allocated chunks are tracked by a simple offset, and
+ * free chunks are tracked in a simple linked list. The offset approach
+ * avoids needing to iterate over all chunks when allocating a new block,
+ * which would cause page faults and cache pollution. Contents of free chunks
+ * is replaced with a pointer to the next free chunk, forming a very simple
+ * linked list. Each block also contains a counter of free chunks. Combined
+ * with the local block-level freelist, it makes it trivial to eventually
+ * free the whole block.
*
* At the context level, we use 'freelist' to track blocks ordered by number
* of free chunks, starting with blocks having a single allocated chunk, and
- * with completely full blocks on the tail.
+ * with completely full blocks on the tail. XXX
*
* This also allows various optimizations - for example when searching for
* free chunk, the allocator reuses space from the fullest blocks first, in
@@ -44,7 +47,7 @@
* case this performs as if the pointer was not maintained.
*
* We cache the freelist index for the blocks with the fewest free chunks
- * (minFreeChunks), so that we don't have to search the freelist on every
+ * (minFreeChunkIndex), so that we don't have to search the freelist on every
* SlabAlloc() call, which is quite expensive.
*
*-------------------------------------------------------------------------
@@ -56,6 +59,17 @@
#include "utils/memdebug.h"
#include "utils/memutils.h"
+struct SlabBlock;
+struct SlabChunk;
+
+/*
+ * Number of actual freelists + 1, for full blocks. Full blocks are always at
+ * offset 0.
+ */
+#define SLAB_FREELIST_COUNT 9
+
+#define SLAB_RETAIN_EMPTY_BLOCK_COUNT 10
+
/*
* SlabContext is a specialized implementation of MemoryContext.
*/
@@ -68,13 +82,16 @@ typedef struct SlabContext
Size blockSize; /* block size */
Size headerSize; /* allocated size of context header */
int chunksPerBlock; /* number of chunks per block */
- int minFreeChunks; /* min number of free chunks in any block */
+ int minFreeChunksIndex; /* min number of free chunks in any block XXX */
int nblocks; /* number of blocks allocated */
#ifdef MEMORY_CONTEXT_CHECKING
bool *freechunks; /* bitmap of free chunks in a block */
#endif
+ int nfreeblocks;
+ dlist_head freeblocks;
+ int freelist_shift;
/* blocks with free space, grouped by number of free chunks: */
- dlist_head freelist[FLEXIBLE_ARRAY_MEMBER];
+ dlist_head freelist[SLAB_FREELIST_COUNT];
} SlabContext;
/*
@@ -83,13 +100,15 @@ typedef struct SlabContext
*
* node: doubly-linked list of blocks in global freelist
* nfree: number of free chunks in this block
- * firstFreeChunk: index of the first free chunk
+ * firstFreeChunk: first free chunk
*/
typedef struct SlabBlock
{
dlist_node node; /* doubly-linked list */
- int nfree; /* number of free chunks */
- int firstFreeChunk; /* index of the first free chunk in the block */
+ int nfree; /* number of chunks on freelist + unused */
+ int nunused; /* number of unused chunks */
+ struct SlabChunk *unused; /* */
+ struct SlabChunk *firstFreeChunk; /* first free chunk in the block */
} SlabBlock;
/*
@@ -123,6 +142,35 @@ typedef struct SlabChunk
#define SlabChunkIndex(slab, block, chunk) \
(((char *) chunk - SlabBlockStart(block)) / slab->fullChunkSize)
+static inline uint8
+SlabFreelistIndex(SlabContext *slab, int freecount)
+{
+ uint8 index;
+
+ Assert(freecount <= slab->chunksPerBlock);
+
+ /*
+ * Ensure, without a branch, that index 0 is only used for blocks entirely
+ * without free chunks.
+ * XXX: There probably is a cheaper way to do this. Needing to shift twice
+ * by slab->freelist_shift isn't great.
+ */
+ index = (freecount + (1 << slab->freelist_shift) - 1) >> slab->freelist_shift;
+
+ if (freecount == 0)
+ Assert(index == 0);
+ else
+ Assert(index > 0 && index < (SLAB_FREELIST_COUNT));
+
+ return index;
+}
+
+static inline dlist_head*
+SlabFreelist(SlabContext *slab, int freecount)
+{
+ return &slab->freelist[SlabFreelistIndex(slab, freecount)];
+}
+
/*
* These functions implement the MemoryContext API for Slab contexts.
*/
@@ -179,7 +227,6 @@ SlabContextCreate(MemoryContext parent,
{
int chunksPerBlock;
Size fullChunkSize;
- Size freelistSize;
Size headerSize;
SlabContext *slab;
int i;
@@ -192,11 +239,11 @@ SlabContextCreate(MemoryContext parent,
"padding calculation in SlabChunk is wrong");
/* Make sure the linked list node fits inside a freed chunk */
- if (chunkSize < sizeof(int))
- chunkSize = sizeof(int);
+ if (chunkSize < MAXALIGN(sizeof(void *)))
+ chunkSize = MAXALIGN(sizeof(void *));
/* chunk, including SLAB header (both addresses nicely aligned) */
- fullChunkSize = sizeof(SlabChunk) + MAXALIGN(chunkSize);
+ fullChunkSize = sizeof(SlabChunk) + chunkSize;
/* Make sure the block can store at least one chunk. */
if (blockSize < fullChunkSize + sizeof(SlabBlock))
@@ -206,16 +253,14 @@ SlabContextCreate(MemoryContext parent,
/* Compute maximum number of chunks per block */
chunksPerBlock = (blockSize - sizeof(SlabBlock)) / fullChunkSize;
- /* The freelist starts with 0, ends with chunksPerBlock. */
- freelistSize = sizeof(dlist_head) * (chunksPerBlock + 1);
-
/*
* Allocate the context header. Unlike aset.c, we never try to combine
* this with the first regular block; not worth the extra complication.
+ * XXX: What's the evidence for that?
*/
/* Size of the memory context header */
- headerSize = offsetof(SlabContext, freelist) + freelistSize;
+ headerSize = sizeof(SlabContext);
#ifdef MEMORY_CONTEXT_CHECKING
@@ -249,17 +294,34 @@ SlabContextCreate(MemoryContext parent,
slab->blockSize = blockSize;
slab->headerSize = headerSize;
slab->chunksPerBlock = chunksPerBlock;
- slab->minFreeChunks = 0;
+ slab->minFreeChunksIndex = 0;
slab->nblocks = 0;
+ slab->nfreeblocks = 0;
+
+ /*
+ * Compute a shift that guarantees that shifting chunksPerBlock with it
+ * yields is smaller than SLAB_FREELIST_COUNT - 1 (one freelist is used for full blocks).
+ */
+ slab->freelist_shift = 0;
+ while ((slab->chunksPerBlock >> slab->freelist_shift) >= (SLAB_FREELIST_COUNT - 1))
+ slab->freelist_shift++;
+
+#if 0
+ elog(LOG, "freelist shift for %d chunks of size %zu is %d, block size %zu",
+ slab->chunksPerBlock, slab->fullChunkSize, slab->freelist_shift,
+ slab->blockSize);
+#endif
+
+ dlist_init(&slab->freeblocks);
/* initialize the freelist slots */
- for (i = 0; i < (slab->chunksPerBlock + 1); i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
dlist_init(&slab->freelist[i]);
#ifdef MEMORY_CONTEXT_CHECKING
/* set the freechunks pointer right after the freelists array */
slab->freechunks
- = (bool *) slab + offsetof(SlabContext, freelist) + freelistSize;
+ = (bool *) slab + sizeof(SlabContext);
#endif
/* Finally, do the type-independent part of context creation */
@@ -292,8 +354,27 @@ SlabReset(MemoryContext context)
SlabCheck(context);
#endif
+ /* release retained empty blocks */
+ {
+ dlist_mutable_iter miter;
+
+ dlist_foreach_modify(miter, &slab->freeblocks)
+ {
+ SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
+
+ dlist_delete(miter.cur);
+
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ slab->nblocks--;
+ context->mem_allocated -= slab->blockSize;
+ }
+ }
+
/* walk over freelists and free the blocks */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
dlist_mutable_iter miter;
@@ -312,7 +393,7 @@ SlabReset(MemoryContext context)
}
}
- slab->minFreeChunks = 0;
+ slab->minFreeChunksIndex = 0;
Assert(slab->nblocks == 0);
Assert(context->mem_allocated == 0);
@@ -342,7 +423,6 @@ SlabAlloc(MemoryContext context, Size size, int flags)
SlabContext *slab = castNode(SlabContext, context);
SlabBlock *block;
SlabChunk *chunk;
- int idx;
Assert(slab);
@@ -352,8 +432,8 @@ SlabAlloc(MemoryContext context, Size size, int flags)
* sense...
*/
- Assert((slab->minFreeChunks >= 0) &&
- (slab->minFreeChunks < slab->chunksPerBlock));
+ Assert((slab->minFreeChunksIndex >= 0) &&
+ (slab->minFreeChunksIndex < SLAB_FREELIST_COUNT));
/* make sure we only allow correct request size */
if (size != slab->chunkSize)
@@ -367,58 +447,88 @@ SlabAlloc(MemoryContext context, Size size, int flags)
* slab->minFreeChunks == 0 means there are no blocks with free chunks,
* thanks to how minFreeChunks is updated at the end of SlabAlloc().
*/
- if (slab->minFreeChunks == 0)
+ if (unlikely(slab->minFreeChunksIndex == 0))
{
- block = (SlabBlock *) malloc(slab->blockSize);
+ if (slab->nfreeblocks > 0)
+ {
+ dlist_node *node;
- if (block == NULL)
- return NULL;
+ node = dlist_pop_head_node(&slab->freeblocks);
+ block = dlist_container(SlabBlock, node, node);
+ slab->nfreeblocks--;
+ }
+ else
+ {
+ block = (SlabBlock *) malloc(slab->blockSize);
- block->nfree = slab->chunksPerBlock;
- block->firstFreeChunk = 0;
+ if (unlikely(block == NULL))
+ return NULL;
- /*
- * Put all the chunks on a freelist. Walk the chunks and point each
- * one to the next one.
- */
- for (idx = 0; idx < slab->chunksPerBlock; idx++)
- {
- chunk = SlabBlockGetChunk(slab, block, idx);
- *(int32 *) SlabChunkGetPointer(chunk) = (idx + 1);
+ slab->nblocks += 1;
+ context->mem_allocated += slab->blockSize;
}
+ block->nfree = slab->chunksPerBlock;
+ block->firstFreeChunk = NULL;
+ block->nunused = slab->chunksPerBlock;
+ block->unused = (SlabChunk *) SlabBlockStart(block);
+
+ slab->minFreeChunksIndex = SlabFreelistIndex(slab, slab->chunksPerBlock);
+
/*
* And add it to the last freelist with all chunks empty.
*
* We know there are no blocks in the freelist, otherwise we wouldn't
* need a new block.
*/
- Assert(dlist_is_empty(&slab->freelist[slab->chunksPerBlock]));
+ Assert(dlist_is_empty(SlabFreelist(slab, slab->chunksPerBlock)));
- dlist_push_head(&slab->freelist[slab->chunksPerBlock], &block->node);
+ dlist_push_head(SlabFreelist(slab, slab->chunksPerBlock),
+ &block->node);
- slab->minFreeChunks = slab->chunksPerBlock;
- slab->nblocks += 1;
- context->mem_allocated += slab->blockSize;
+ chunk = block->unused;
+ block->unused = (SlabChunk *)(((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
}
+ else
+ {
+ /* grab the block from the freelist */
+ block = dlist_head_element(SlabBlock, node,
+ &slab->freelist[slab->minFreeChunksIndex]);
- /* grab the block from the freelist (even the new block is there) */
- block = dlist_head_element(SlabBlock, node,
- &slab->freelist[slab->minFreeChunks]);
+ /* make sure we actually got a valid block, with matching nfree */
+ Assert(block != NULL);
+ Assert(slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree));
+ Assert(block->nfree > 0);
- /* make sure we actually got a valid block, with matching nfree */
- Assert(block != NULL);
- Assert(slab->minFreeChunks == block->nfree);
- Assert(block->nfree > 0);
+ /* we know index of the first free chunk in the block */
+ if (block->nunused > 0)
+ {
+ chunk = block->unused;
+ block->unused = (SlabChunk *)(((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
+ else
+ {
+ chunk = block->firstFreeChunk;
- /* we know index of the first free chunk in the block */
- idx = block->firstFreeChunk;
+ /*
+ * Remove the chunk from the freelist head. The index of the next free
+ * chunk is stored in the chunk itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(void*));
+ block->firstFreeChunk = *(SlabChunk **) SlabChunkGetPointer(chunk);
+ }
+ }
- /* make sure the chunk index is valid, and that it's marked as empty */
- Assert((idx >= 0) && (idx < slab->chunksPerBlock));
+ /* make sure the chunk is in the block and that it's marked as empty (XXX?) */
+ Assert((char *) chunk >= SlabBlockStart(block));
+ Assert((char *) chunk < (((char *)block) + slab->blockSize));
- /* compute the chunk location block start (after the block header) */
- chunk = SlabBlockGetChunk(slab, block, idx);
+ Assert(block->firstFreeChunk == NULL || (
+ (block->firstFreeChunk >= (SlabChunk *) SlabBlockStart(block)) &&
+ block->firstFreeChunk <= (SlabChunk *) (((char *)block) + slab->blockSize))
+ );
/*
* Update the block nfree count, and also the minFreeChunks as we've
@@ -426,54 +536,51 @@ SlabAlloc(MemoryContext context, Size size, int flags)
* (because that's how we chose the block).
*/
block->nfree--;
- slab->minFreeChunks = block->nfree;
- /*
- * Remove the chunk from the freelist head. The index of the next free
- * chunk is stored in the chunk itself.
- */
- VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(int32));
- block->firstFreeChunk = *(int32 *) SlabChunkGetPointer(chunk);
+ /* Prepare to initialize the chunk header. */
+ VALGRIND_MAKE_MEM_UNDEFINED(chunk, sizeof(SlabChunk));
- Assert(block->firstFreeChunk >= 0);
- Assert(block->firstFreeChunk <= slab->chunksPerBlock);
+ chunk->block = block;
+ chunk->slab = slab;
- Assert((block->nfree != 0 &&
- block->firstFreeChunk < slab->chunksPerBlock) ||
- (block->nfree == 0 &&
- block->firstFreeChunk == slab->chunksPerBlock));
+ Assert(slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree + 1));
/* move the whole block to the right place in the freelist */
- dlist_delete(&block->node);
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
-
- /*
- * And finally update minFreeChunks, i.e. the index to the block with the
- * lowest number of free chunks. We only need to do that when the block
- * got full (otherwise we know the current block is the right one). We'll
- * simply walk the freelist until we find a non-empty entry.
- */
- if (slab->minFreeChunks == 0)
+ if (unlikely(SlabFreelistIndex(slab, block->nfree) != slab->minFreeChunksIndex))
{
- for (idx = 1; idx <= slab->chunksPerBlock; idx++)
+ slab->minFreeChunksIndex = SlabFreelistIndex(slab, block->nfree);
+ dlist_delete(&block->node);
+ dlist_push_head(SlabFreelist(slab, block->nfree), &block->node);
+
+
+ /*
+ * And finally update minFreeChunks, i.e. the index to the block with the
+ * lowest number of free chunks. We only need to do that when the block
+ * got full (otherwise we know the current block is the right one). We'll
+ * simply walk the freelist until we find a non-empty entry.
+ */
+ if (slab->minFreeChunksIndex == 0)
{
- if (dlist_is_empty(&slab->freelist[idx]))
- continue;
+ for (int idx = 1; idx < SLAB_FREELIST_COUNT; idx++)
+ {
+ if (dlist_is_empty(&slab->freelist[idx]))
+ continue;
- /* found a non-empty freelist */
- slab->minFreeChunks = idx;
- break;
+ /* found a non-empty freelist */
+ slab->minFreeChunksIndex = idx;
+ break;
+ }
}
}
+#if 0
+ /*
+ * FIXME: I don't understand what this ever did? It should be unreachable
+ * I think?
+ */
if (slab->minFreeChunks == slab->chunksPerBlock)
slab->minFreeChunks = 0;
-
- /* Prepare to initialize the chunk header. */
- VALGRIND_MAKE_MEM_UNDEFINED(chunk, sizeof(SlabChunk));
-
- chunk->block = block;
- chunk->slab = slab;
+#endif
#ifdef MEMORY_CONTEXT_CHECKING
/* slab mark to catch clobber of "unused" space */
@@ -496,6 +603,61 @@ SlabAlloc(MemoryContext context, Size size, int flags)
return SlabChunkGetPointer(chunk);
}
+static void pg_noinline
+SlabFreeSlow(SlabContext *slab, SlabBlock *block)
+{
+ dlist_delete(&block->node);
+
+ /*
+ * See if we need to update the minFreeChunks field for the slab - we only
+ * need to do that if the block had that number of free chunks
+ * before we freed one. In that case, we check if there still are blocks
+ * in the original freelist and we either keep the current value (if there
+ * still are blocks) or increment it by one (the new block is still the
+ * one with minimum free chunks).
+ *
+ * The one exception is when the block will get completely free - in that
+ * case we will free it, se we can't use it for minFreeChunks. It however
+ * means there are no more blocks with free chunks.
+ */
+ if (slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree - 1))
+ {
+ /* Have we removed the last chunk from the freelist? */
+ if (dlist_is_empty(&slab->freelist[slab->minFreeChunksIndex]))
+ {
+ /* but if we made the block entirely free, we'll free it */
+ if (block->nfree == slab->chunksPerBlock)
+ slab->minFreeChunksIndex = 0;
+ else
+ slab->minFreeChunksIndex++;
+ }
+ }
+
+ /* If the block is now completely empty, free it (in a way). */
+ if (block->nfree == slab->chunksPerBlock)
+ {
+ /*
+ * To avoid constantly freeing/allocating blocks in bursty patterns
+ * (on most crucially cases of repeatedly allocating and freeing a
+ * single chunk), retain a small number of blocks.
+ */
+ if (slab->nfreeblocks < SLAB_RETAIN_EMPTY_BLOCK_COUNT)
+ {
+ dlist_push_head(&slab->freeblocks, &block->node);
+ slab->nfreeblocks++;
+ }
+ else
+ {
+ slab->nblocks--;
+ slab->header.mem_allocated -= slab->blockSize;
+ free(block);
+ }
+ }
+ else
+ dlist_push_head(SlabFreelist(slab, block->nfree), &block->node);
+
+}
+
/*
* SlabFree
* Frees allocated memory; memory is removed from the slab.
@@ -503,7 +665,6 @@ SlabAlloc(MemoryContext context, Size size, int flags)
static void
SlabFree(MemoryContext context, void *pointer)
{
- int idx;
SlabContext *slab = castNode(SlabContext, context);
SlabChunk *chunk = SlabPointerGetChunk(pointer);
SlabBlock *block = chunk->block;
@@ -516,61 +677,26 @@ SlabFree(MemoryContext context, void *pointer)
slab->header.name, chunk);
#endif
- /* compute index of the chunk with respect to block start */
- idx = SlabChunkIndex(slab, block, chunk);
-
/* add chunk to freelist, and update block nfree count */
- *(int32 *) pointer = block->firstFreeChunk;
- block->firstFreeChunk = idx;
+ *(SlabChunk **) pointer = block->firstFreeChunk;
+ block->firstFreeChunk = chunk;
block->nfree++;
Assert(block->nfree > 0);
Assert(block->nfree <= slab->chunksPerBlock);
#ifdef CLOBBER_FREED_MEMORY
- /* XXX don't wipe the int32 index, used for block-level freelist */
- wipe_mem((char *) pointer + sizeof(int32),
- slab->chunkSize - sizeof(int32));
+ /* XXX don't wipe the SlabChunk* index, used for block-level freelist */
+ wipe_mem((char *) pointer + sizeof(SlabChunk*),
+ slab->chunkSize - sizeof(SlabChunk*));
#endif
/* remove the block from a freelist */
- dlist_delete(&block->node);
-
- /*
- * See if we need to update the minFreeChunks field for the slab - we only
- * need to do that if there the block had that number of free chunks
- * before we freed one. In that case, we check if there still are blocks
- * in the original freelist and we either keep the current value (if there
- * still are blocks) or increment it by one (the new block is still the
- * one with minimum free chunks).
- *
- * The one exception is when the block will get completely free - in that
- * case we will free it, se we can't use it for minFreeChunks. It however
- * means there are no more blocks with free chunks.
- */
- if (slab->minFreeChunks == (block->nfree - 1))
+ if (SlabFreelistIndex(slab, block->nfree) != SlabFreelistIndex(slab, block->nfree - 1))
{
- /* Have we removed the last chunk from the freelist? */
- if (dlist_is_empty(&slab->freelist[slab->minFreeChunks]))
- {
- /* but if we made the block entirely free, we'll free it */
- if (block->nfree == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
- else
- slab->minFreeChunks++;
- }
+ SlabFreeSlow(slab, block);
}
- /* If the block is now completely empty, free it. */
- if (block->nfree == slab->chunksPerBlock)
- {
- free(block);
- slab->nblocks--;
- context->mem_allocated -= slab->blockSize;
- }
- else
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
-
Assert(slab->nblocks >= 0);
Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
}
@@ -657,7 +783,7 @@ SlabStats(MemoryContext context,
/* Include context header in totalspace */
totalspace = slab->headerSize;
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
dlist_iter iter;
@@ -714,7 +840,7 @@ SlabCheck(MemoryContext context)
Assert(slab->chunksPerBlock > 0);
/* walk all the freelists */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
int j,
nfree;
@@ -723,20 +849,19 @@ SlabCheck(MemoryContext context)
/* walk all blocks on this freelist */
dlist_foreach(iter, &slab->freelist[i])
{
- int idx;
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+ SlabChunk *cur_chunk;
/*
* Make sure the number of free chunks (in the block header)
* matches position in the freelist.
*/
- if (block->nfree != i)
+ if (SlabFreelistIndex(slab, block->nfree) != i)
elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match freelist %d",
name, block->nfree, block, i);
/* reset the bitmap of free chunks for this block */
memset(slab->freechunks, 0, (slab->chunksPerBlock * sizeof(bool)));
- idx = block->firstFreeChunk;
/*
* Now walk through the chunks, count the free ones and also
@@ -744,20 +869,31 @@ SlabCheck(MemoryContext context)
* freelist is stored within the chunks themselves, we have to
* walk through the chunks and construct our own bitmap.
*/
-
+ cur_chunk = block->firstFreeChunk;
nfree = 0;
- while (idx < slab->chunksPerBlock)
+ while (cur_chunk != NULL)
{
- SlabChunk *chunk;
+ int idx = SlabChunkIndex(slab, block, cur_chunk);
/* count the chunk as free, add it to the bitmap */
nfree++;
slab->freechunks[idx] = true;
/* read index of the next free chunk */
- chunk = SlabBlockGetChunk(slab, block, idx);
- VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(int32));
- idx = *(int32 *) SlabChunkGetPointer(chunk);
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(cur_chunk), sizeof(SlabChunk **));
+ cur_chunk = *(SlabChunk **) SlabChunkGetPointer(cur_chunk);
+ }
+
+ cur_chunk = block->unused;
+ for (j = 0; j < block->nunused; j++)
+ {
+ int idx = SlabChunkIndex(slab, block, cur_chunk);
+
+ /* count the chunk as free, add it to the bitmap */
+ nfree++;
+ slab->freechunks[idx] = true;
+
+ cur_chunk = (SlabChunk *)(((char *) cur_chunk) + slab->fullChunkSize);
}
for (j = 0; j < slab->chunksPerBlock; j++)
--
2.31.1
[application/vnd.oasis.opendocument.spreadsheet] slab i5.ods (587.8K, ../../5d83ac71-9054-208d-a48a-24d9d6b61f55@enterprisedb.com/5-slab%20i5.ods)
download
[application/vnd.oasis.opendocument.spreadsheet] slab xeon.ods (593.6K, ../../5d83ac71-9054-208d-a48a-24d9d6b61f55@enterprisedb.com/6-slab%20xeon.ods)
download
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
@ 2021-09-10 21:06 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-05 15:31 ` Re: slab allocator performance issues Robert Haas <robertmhaas@gmail.com>
0 siblings, 2 replies; 29+ messages in thread
From: Tomas Vondra @ 2021-09-10 21:06 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Hi,
I've been investigating the regressions in some of the benchmark
results, together with the generation context benchmarks [1].
Turns out it's pretty difficult to benchmark this, because the results
strongly depend on what the backend did before. For example if I run
slab_bench_fifo with the "decreasing" test for 32kB blocks and 512B
chunks, I get this:
select * from slab_bench_fifo(1000000, 32768, 512, 100, 10000, 5000);
mem_allocated | alloc_ms | free_ms
---------------+----------+---------
528547840 | 155394 | 87440
i.e. palloc() takes ~155ms and pfree() ~87ms (and these result are
stable, the numbers don't change much with more runs).
But if I run a set of "lifo" tests in the backend first, the results
look like this:
mem_allocated | alloc_ms | free_ms
---------------+----------+---------
528547840 | 41728 | 71524
(1 row)
so the pallocs are suddenly about ~4x faster. Clearly, what the backend
did before may have pretty dramatic impact on results, even for simple
benchmarks like this.
Note: The benchmark was a single SQL script, running all the different
workloads in the same backend.
I did a fair amount of perf profiling, and the main difference between
the slow and fast runs seems to be this:
0 page-faults:u
0 minor-faults:u
0 major-faults:u
vs
20,634,153 page-faults:u
20,634,153 minor-faults:u
0 major-faults:u
Attached is a more complete perf stat output, but the page faults seem
to be the main issue. My theory is that in the "fast" case, the past
backend activity puts the glibc memory management into a state that
prevents page faults in the benchmark.
But of course, this theory may be incomplete - for example it's not
clear why running the benchmark repeatedly would not "condition" the
backend the same way. But it doesn't - it's ~150ms even for repeated runs.
Secondly, I'm not sure this explains why some of the timings actually
got much slower with the 0003 patch, when the sequence of the steps is
still the same. Of course, it's possible 0003 changes the allocation
pattern a bit, interfering with glibc memory management.
This leads to a couple of interesting questions, I think:
1) I've only tested this on Linux, with glibc. I wonder how it'd behave
on other platforms, or with other allocators.
2) Which cases are more important? When the backend was warmed up, or
when each benchmark runs in a new backend? It seems the "new backend" is
something like a "worst case" leading to more page faults, so maybe
that's the thing to watch. OTOH it's unlikely to have a completely new
backend, so maybe not.
3) Can this teach us something about how to allocate stuff, to better
"prepare" the backend for future allocations? For example, it's a bit
strange that repeated runs of the same benchmark don't do the trick, for
some reason.
regards
[1]
https://www.postgresql.org/message-id/bcdd4e3e-c12d-cd2b-7ead-a91ad416100a%40enterprisedb.com
--
Tomas Vondra
EnterpriseDB: http://www.enterprisedb.com
The Enterprise PostgreSQL Company
Performance counter stats for process id '11829':
219,869,739,737 cycles:u
119,722,280,436 instructions:u # 0.54 insn per cycle
# 1.56 stalled cycles per insn
7,784,566,858 cache-references:u
2,487,257,287 cache-misses:u # 31.951 % of all cache refs
5,942,054,520 bus-cycles:u
0 page-faults:u
0 minor-faults:u
0 major-faults:u
187,181,661,719 stalled-cycles-frontend:u # 85.13% frontend cycles idle
144,274,017,071 stalled-cycles-backend:u # 65.62% backend cycles idle
60.000876248 seconds time elapsed
Performance counter stats for process id '11886':
145,093,090,692 cycles:u
74,986,543,212 instructions:u # 0.52 insn per cycle
# 2.63 stalled cycles per insn
4,753,764,781 cache-references:u
1,342,653,549 cache-misses:u # 28.244 % of all cache refs
3,925,175,515 bus-cycles:u
20,634,153 page-faults:u
20,634,153 minor-faults:u
0 major-faults:u
197,130,461,632 stalled-cycles-frontend:u # 135.86% frontend cycles idle
168,434,343,213 stalled-cycles-backend:u # 116.09% backend cycles idle
60.000891867 seconds time elapsed
Attachments:
[text/plain] perf.stat.txt (2.0K, ../../a5ccda91-d9fc-49c5-b3c7-c81528b938c5@enterprisedb.com/2-perf.stat.txt)
download | inline:
Performance counter stats for process id '11829':
219,869,739,737 cycles:u
119,722,280,436 instructions:u # 0.54 insn per cycle
# 1.56 stalled cycles per insn
7,784,566,858 cache-references:u
2,487,257,287 cache-misses:u # 31.951 % of all cache refs
5,942,054,520 bus-cycles:u
0 page-faults:u
0 minor-faults:u
0 major-faults:u
187,181,661,719 stalled-cycles-frontend:u # 85.13% frontend cycles idle
144,274,017,071 stalled-cycles-backend:u # 65.62% backend cycles idle
60.000876248 seconds time elapsed
Performance counter stats for process id '11886':
145,093,090,692 cycles:u
74,986,543,212 instructions:u # 0.52 insn per cycle
# 2.63 stalled cycles per insn
4,753,764,781 cache-references:u
1,342,653,549 cache-misses:u # 28.244 % of all cache refs
3,925,175,515 bus-cycles:u
20,634,153 page-faults:u
20,634,153 minor-faults:u
0 major-faults:u
197,130,461,632 stalled-cycles-frontend:u # 135.86% frontend cycles idle
168,434,343,213 stalled-cycles-backend:u # 116.09% backend cycles idle
60.000891867 seconds time elapsed
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
@ 2022-10-12 09:37 ` David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
1 sibling, 1 reply; 29+ messages in thread
From: David Rowley @ 2022-10-12 09:37 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@enterprisedb.com>; +Cc: Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On Sat, 11 Sept 2021 at 09:07, Tomas Vondra
<tomas.vondra@enterprisedb.com> wrote:
> I've been investigating the regressions in some of the benchmark
> results, together with the generation context benchmarks [1].
I've not looked into the regression you found with this yet, but I did
rebase the patch. slab.c has seen quite a number of changes recently.
I didn't spend a lot of time checking over the patch. I mainly wanted
to see what the performance was like before reviewing in too much
detail.
To test the performance, I used [1] and ran:
select pg_allocate_memory_test(<nbytes>, 1024*1024,
10::bigint*1024*1024*1024, 'slab');
that basically allocates chunks of <nbytes> and keeps around 1MB of
them at a time and allocates a total of 10GBs of them.
I saw:
Master:
16 byte chunk = 8754.678 ms
32 byte chunk = 4511.725 ms
64 byte chunk = 2244.885 ms
128 byte chunk = 1135.349 ms
256 byte chunk = 548.030 ms
512 byte chunk = 272.017 ms
1024 byte chunk = 144.618 ms
Master + attached patch:
16 byte chunk = 5255.974 ms
32 byte chunk = 2640.807 ms
64 byte chunk = 1328.949 ms
128 byte chunk = 668.078 ms
256 byte chunk = 330.564 ms
512 byte chunk = 166.844 ms
1024 byte chunk = 85.399 ms
So patched runs in about 60% of the time that master runs in.
I plan to look at the patch in a bit more detail and see if I can
recreate and figure out the regression that Tomas reported. For now, I
just want to share the rebased patch.
The only thing I really adjusted from Andres' version is to instead of
using pointers for the linked list block freelist, I made it store the
number of bytes into the block that the chunk is. This means we can
use 4 bytes instead of 8 bytes for these pointers. The block size is
limited to 1GB now anyway, so 32-bit is large enough for these
offsets.
David
[1] https://www.postgresql.org/message-id/attachment/137056/allocate_performance_functions.patch.txt
From 2b89d993b0294d5c0fe0a8333bccc555337ac979 Mon Sep 17 00:00:00 2001
From: David Rowley <dgrowley@gmail.com>
Date: Wed, 12 Oct 2022 09:30:24 +1300
Subject: [PATCH v3] WIP: slab performance.
Author: Andres Freund
Discussion: https://postgr.es/m/20210717194333.mr5io3zup3kxahfm@alap3.anarazel.de
---
src/backend/utils/mmgr/slab.c | 449 ++++++++++++++++++++++------------
1 file changed, 294 insertions(+), 155 deletions(-)
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index 1a0b28f9ea..58b8e2d67c 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -23,15 +23,18 @@
* global (context) level. This is possible as the chunk size (and thus also
* the number of chunks per block) is fixed.
*
- * On each block, free chunks are tracked in a simple linked list. Contents
- * of free chunks is replaced with an index of the next free chunk, forming
- * a very simple linked list. Each block also contains a counter of free
- * chunks. Combined with the local block-level freelist, it makes it trivial
- * to eventually free the whole block.
+ * On each block, never allocated chunks are tracked by a simple offset, and
+ * free chunks are tracked in a simple linked list. The offset approach
+ * avoids needing to iterate over all chunks when allocating a new block,
+ * which would cause page faults and cache pollution. Contents of free chunks
+ * is replaced with a the offset (in bytes) from the block pointer to the
+ * next free chunk, forming a very simple linked list. Each block also
+ * contains a counter of free chunks. Combined with the local block-level
+ * freelist, it makes it trivial to eventually free the whole block.
*
* At the context level, we use 'freelist' to track blocks ordered by number
* of free chunks, starting with blocks having a single allocated chunk, and
- * with completely full blocks on the tail.
+ * with completely full blocks on the tail. XXX
*
* This also allows various optimizations - for example when searching for
* free chunk, the allocator reuses space from the fullest blocks first, in
@@ -44,7 +47,7 @@
* case this performs as if the pointer was not maintained.
*
* We cache the freelist index for the blocks with the fewest free chunks
- * (minFreeChunks), so that we don't have to search the freelist on every
+ * (minFreeChunkIndex), so that we don't have to search the freelist on every
* SlabAlloc() call, which is quite expensive.
*
*-------------------------------------------------------------------------
@@ -60,6 +63,17 @@
#define Slab_BLOCKHDRSZ MAXALIGN(sizeof(SlabBlock))
+struct SlabBlock;
+struct SlabChunk;
+
+/*
+ * Number of actual freelists + 1, for full blocks. Full blocks are always at
+ * offset 0.
+ */
+#define SLAB_FREELIST_COUNT 9
+
+#define SLAB_RETAIN_EMPTY_BLOCK_COUNT 10
+
/*
* SlabContext is a specialized implementation of MemoryContext.
*/
@@ -72,13 +86,16 @@ typedef struct SlabContext
Size blockSize; /* block size */
Size headerSize; /* allocated size of context header */
int chunksPerBlock; /* number of chunks per block */
- int minFreeChunks; /* min number of free chunks in any block */
+ int minFreeChunksIndex; /* min number of free chunks in any block XXX */
int nblocks; /* number of blocks allocated */
#ifdef MEMORY_CONTEXT_CHECKING
bool *freechunks; /* bitmap of free chunks in a block */
#endif
+ int nfreeblocks;
+ dlist_head freeblocks;
+ int freelist_shift;
/* blocks with free space, grouped by number of free chunks: */
- dlist_head freelist[FLEXIBLE_ARRAY_MEMBER];
+ dlist_head freelist[SLAB_FREELIST_COUNT];
} SlabContext;
/*
@@ -92,8 +109,11 @@ typedef struct SlabContext
typedef struct SlabBlock
{
dlist_node node; /* doubly-linked list */
- int nfree; /* number of free chunks */
- int firstFreeChunk; /* index of the first free chunk in the block */
+ int nfree; /* number of chunks on freelist + unused */
+ int nunused; /* number of unused chunks */
+ int32 firstFreeOffset; /* offset of bytes from block to the first
+ * free chunk in the block. */
+ MemoryChunk *unused; /* */
SlabContext *slab; /* owning context */
} SlabBlock;
@@ -125,6 +145,34 @@ typedef struct SlabBlock
#define SlabBlockIsValid(block) \
(PointerIsValid(block) && SlabIsValid((block)->slab))
+static inline uint8
+SlabFreelistIndex(SlabContext *slab, int freecount)
+{
+ uint8 index;
+
+ Assert(freecount <= slab->chunksPerBlock);
+
+ /*
+ * Ensure, without a branch, that index 0 is only used for blocks entirely
+ * without free chunks.
+ * XXX: There probably is a cheaper way to do this. Needing to shift twice
+ * by slab->freelist_shift isn't great.
+ */
+ index = (freecount + (1 << slab->freelist_shift) - 1) >> slab->freelist_shift;
+
+ if (freecount == 0)
+ Assert(index == 0);
+ else
+ Assert(index > 0 && index < (SLAB_FREELIST_COUNT));
+
+ return index;
+}
+
+static inline dlist_head*
+SlabFreelist(SlabContext *slab, int freecount)
+{
+ return &slab->freelist[SlabFreelistIndex(slab, freecount)];
+}
/*
* SlabContextCreate
@@ -145,7 +193,6 @@ SlabContextCreate(MemoryContext parent,
{
int chunksPerBlock;
Size fullChunkSize;
- Size freelistSize;
Size headerSize;
SlabContext *slab;
int i;
@@ -155,9 +202,12 @@ SlabContextCreate(MemoryContext parent,
"sizeof(MemoryChunk) is not maxaligned");
Assert(MAXALIGN(chunkSize) <= MEMORYCHUNK_MAX_VALUE);
- /* Make sure the linked list node fits inside a freed chunk */
- if (chunkSize < sizeof(int))
- chunkSize = sizeof(int);
+ /*
+ * Ensure there's enough space to store the offset to the next free chunk
+ * in the memory of free'd chunks.
+ */
+ if (chunkSize < sizeof(uint32))
+ chunkSize = sizeof(uint32);
/* chunk, including SLAB header (both addresses nicely aligned) */
#ifdef MEMORY_CONTEXT_CHECKING
@@ -175,16 +225,14 @@ SlabContextCreate(MemoryContext parent,
/* Compute maximum number of chunks per block */
chunksPerBlock = (blockSize - Slab_BLOCKHDRSZ) / fullChunkSize;
- /* The freelist starts with 0, ends with chunksPerBlock. */
- freelistSize = sizeof(dlist_head) * (chunksPerBlock + 1);
-
/*
* Allocate the context header. Unlike aset.c, we never try to combine
* this with the first regular block; not worth the extra complication.
+ * XXX: What's the evidence for that?
*/
/* Size of the memory context header */
- headerSize = offsetof(SlabContext, freelist) + freelistSize;
+ headerSize = sizeof(SlabContext);
#ifdef MEMORY_CONTEXT_CHECKING
@@ -218,17 +266,34 @@ SlabContextCreate(MemoryContext parent,
slab->blockSize = blockSize;
slab->headerSize = headerSize;
slab->chunksPerBlock = chunksPerBlock;
- slab->minFreeChunks = 0;
+ slab->minFreeChunksIndex = 0;
slab->nblocks = 0;
+ slab->nfreeblocks = 0;
+
+ /*
+ * Compute a shift that guarantees that shifting chunksPerBlock with it
+ * yields is smaller than SLAB_FREELIST_COUNT - 1 (one freelist is used for full blocks).
+ */
+ slab->freelist_shift = 0;
+ while ((slab->chunksPerBlock >> slab->freelist_shift) >= (SLAB_FREELIST_COUNT - 1))
+ slab->freelist_shift++;
+
+#if 0
+ elog(LOG, "freelist shift for %d chunks of size %zu is %d, block size %zu",
+ slab->chunksPerBlock, slab->fullChunkSize, slab->freelist_shift,
+ slab->blockSize);
+#endif
+
+ dlist_init(&slab->freeblocks);
/* initialize the freelist slots */
- for (i = 0; i < (slab->chunksPerBlock + 1); i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
dlist_init(&slab->freelist[i]);
#ifdef MEMORY_CONTEXT_CHECKING
/* set the freechunks pointer right after the freelists array */
slab->freechunks
- = (bool *) slab + offsetof(SlabContext, freelist) + freelistSize;
+ = (bool *) slab + sizeof(SlabContext);
#endif
/* Finally, do the type-independent part of context creation */
@@ -261,8 +326,27 @@ SlabReset(MemoryContext context)
SlabCheck(context);
#endif
+ /* release retained empty blocks */
+ {
+ dlist_mutable_iter miter;
+
+ dlist_foreach_modify(miter, &slab->freeblocks)
+ {
+ SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
+
+ dlist_delete(miter.cur);
+
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ slab->nblocks--;
+ context->mem_allocated -= slab->blockSize;
+ }
+ }
+
/* walk over freelists and free the blocks */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
dlist_mutable_iter miter;
@@ -281,7 +365,7 @@ SlabReset(MemoryContext context)
}
}
- slab->minFreeChunks = 0;
+ slab->minFreeChunksIndex = 0;
Assert(slab->nblocks == 0);
Assert(context->mem_allocated == 0);
@@ -311,12 +395,11 @@ SlabAlloc(MemoryContext context, Size size)
SlabContext *slab = (SlabContext *) context;
SlabBlock *block;
MemoryChunk *chunk;
- int idx;
AssertArg(SlabIsValid(slab));
- Assert((slab->minFreeChunks >= 0) &&
- (slab->minFreeChunks < slab->chunksPerBlock));
+ Assert((slab->minFreeChunksIndex >= 0) &&
+ (slab->minFreeChunksIndex < SLAB_FREELIST_COUNT));
/* make sure we only allow correct request size */
if (size != slab->chunkSize)
@@ -327,62 +410,93 @@ SlabAlloc(MemoryContext context, Size size)
* If there are no free chunks in any existing block, create a new block
* and put it to the last freelist bucket.
*
- * slab->minFreeChunks == 0 means there are no blocks with free chunks,
- * thanks to how minFreeChunks is updated at the end of SlabAlloc().
+ * slab->minFreeChunksIndex == 0 means there are no blocks with free
+ * chunks, thanks to how minFreeChunks is updated at the end of
+ * SlabAlloc().
*/
- if (slab->minFreeChunks == 0)
+ if (unlikely(slab->minFreeChunksIndex == 0))
{
- block = (SlabBlock *) malloc(slab->blockSize);
+ if (slab->nfreeblocks > 0)
+ {
+ dlist_node *node;
- if (block == NULL)
- return NULL;
+ node = dlist_pop_head_node(&slab->freeblocks);
+ block = dlist_container(SlabBlock, node, node);
+ slab->nfreeblocks--;
+ }
+ else
+ {
+ block = (SlabBlock *) malloc(slab->blockSize);
- block->nfree = slab->chunksPerBlock;
- block->firstFreeChunk = 0;
- block->slab = slab;
+ if (unlikely(block == NULL))
+ return NULL;
- /*
- * Put all the chunks on a freelist. Walk the chunks and point each
- * one to the next one.
- */
- for (idx = 0; idx < slab->chunksPerBlock; idx++)
- {
- chunk = SlabBlockGetChunk(slab, block, idx);
- *(int32 *) MemoryChunkGetPointer(chunk) = (idx + 1);
+ slab->nblocks += 1;
+ block->slab = slab;
+ context->mem_allocated += slab->blockSize;
}
+ block->nfree = slab->chunksPerBlock;
+ block->firstFreeOffset = -1;
+ block->nunused = slab->chunksPerBlock;
+ block->unused = (MemoryChunk *) SlabBlockStart(block);
+
+ slab->minFreeChunksIndex = SlabFreelistIndex(slab, slab->chunksPerBlock);
+
/*
* And add it to the last freelist with all chunks empty.
*
* We know there are no blocks in the freelist, otherwise we wouldn't
* need a new block.
*/
- Assert(dlist_is_empty(&slab->freelist[slab->chunksPerBlock]));
+ Assert(dlist_is_empty(SlabFreelist(slab, slab->chunksPerBlock)));
- dlist_push_head(&slab->freelist[slab->chunksPerBlock], &block->node);
+ dlist_push_head(SlabFreelist(slab, slab->chunksPerBlock),
+ &block->node);
- slab->minFreeChunks = slab->chunksPerBlock;
- slab->nblocks += 1;
- context->mem_allocated += slab->blockSize;
+ chunk = block->unused;
+ block->unused = (MemoryChunk *) (((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
}
+ else
+ {
+ /* grab the block from the freelist */
+ block = dlist_head_element(SlabBlock, node,
+ &slab->freelist[slab->minFreeChunksIndex]);
- /* grab the block from the freelist (even the new block is there) */
- block = dlist_head_element(SlabBlock, node,
- &slab->freelist[slab->minFreeChunks]);
+ /* make sure we actually got a valid block, with matching nfree */
+ Assert(block != NULL);
+ Assert(slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree));
+ Assert(block->nfree > 0);
- /* make sure we actually got a valid block, with matching nfree */
- Assert(block != NULL);
- Assert(slab->minFreeChunks == block->nfree);
- Assert(block->nfree > 0);
+ /* we know index of the first free chunk in the block */
+ if (block->nunused > 0)
+ {
+ chunk = block->unused;
+ block->unused = (MemoryChunk *) (((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
+ else
+ {
+ chunk = (MemoryChunk *) ((char *) block + block->firstFreeOffset);
- /* we know index of the first free chunk in the block */
- idx = block->firstFreeChunk;
+ /*
+ * Remove the chunk from the freelist head. The index of the next free
+ * chunk is stored in the chunk itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(void*));
+ block->firstFreeOffset = *(int32 *) SlabChunkGetPointer(chunk);
+ }
+ }
- /* make sure the chunk index is valid, and that it's marked as empty */
- Assert((idx >= 0) && (idx < slab->chunksPerBlock));
+ /* make sure the chunk is in the block and that it's marked as empty (XXX?) */
+ Assert((char *) chunk >= SlabBlockStart(block));
+ Assert((char *) chunk < (((char *)block) + slab->blockSize));
- /* compute the chunk location block start (after the block header) */
- chunk = SlabBlockGetChunk(slab, block, idx);
+ Assert(block->firstFreeChunk == NULL || (
+ (block->firstFreeChunk >= (SlabChunk *) SlabBlockStart(block)) &&
+ block->firstFreeChunk <= (SlabChunk *) (((char *)block) + slab->blockSize))
+ );
/*
* Update the block nfree count, and also the minFreeChunks as we've
@@ -390,48 +504,6 @@ SlabAlloc(MemoryContext context, Size size)
* (because that's how we chose the block).
*/
block->nfree--;
- slab->minFreeChunks = block->nfree;
-
- /*
- * Remove the chunk from the freelist head. The index of the next free
- * chunk is stored in the chunk itself.
- */
- VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- block->firstFreeChunk = *(int32 *) MemoryChunkGetPointer(chunk);
-
- Assert(block->firstFreeChunk >= 0);
- Assert(block->firstFreeChunk <= slab->chunksPerBlock);
-
- Assert((block->nfree != 0 &&
- block->firstFreeChunk < slab->chunksPerBlock) ||
- (block->nfree == 0 &&
- block->firstFreeChunk == slab->chunksPerBlock));
-
- /* move the whole block to the right place in the freelist */
- dlist_delete(&block->node);
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
-
- /*
- * And finally update minFreeChunks, i.e. the index to the block with the
- * lowest number of free chunks. We only need to do that when the block
- * got full (otherwise we know the current block is the right one). We'll
- * simply walk the freelist until we find a non-empty entry.
- */
- if (slab->minFreeChunks == 0)
- {
- for (idx = 1; idx <= slab->chunksPerBlock; idx++)
- {
- if (dlist_is_empty(&slab->freelist[idx]))
- continue;
-
- /* found a non-empty freelist */
- slab->minFreeChunks = idx;
- break;
- }
- }
-
- if (slab->minFreeChunks == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
/* Prepare to initialize the chunk header. */
VALGRIND_MAKE_MEM_UNDEFINED(chunk, Slab_CHUNKHDRSZ);
@@ -458,6 +530,60 @@ SlabAlloc(MemoryContext context, Size size)
return MemoryChunkGetPointer(chunk);
}
+static void pg_noinline
+SlabFreeSlow(SlabContext *slab, SlabBlock *block)
+{
+ dlist_delete(&block->node);
+
+ /*
+ * See if we need to update the minFreeChunks field for the slab - we only
+ * need to do that if the block had that number of free chunks
+ * before we freed one. In that case, we check if there still are blocks
+ * in the original freelist and we either keep the current value (if there
+ * still are blocks) or increment it by one (the new block is still the
+ * one with minimum free chunks).
+ *
+ * The one exception is when the block will get completely free - in that
+ * case we will free it, se we can't use it for minFreeChunks. It however
+ * means there are no more blocks with free chunks.
+ */
+ if (slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree - 1))
+ {
+ /* Have we removed the last chunk from the freelist? */
+ if (dlist_is_empty(&slab->freelist[slab->minFreeChunksIndex]))
+ {
+ /* but if we made the block entirely free, we'll free it */
+ if (block->nfree == slab->chunksPerBlock)
+ slab->minFreeChunksIndex = 0;
+ else
+ slab->minFreeChunksIndex++;
+ }
+ }
+
+ /* If the block is now completely empty, free it (in a way). */
+ if (block->nfree == slab->chunksPerBlock)
+ {
+ /*
+ * To avoid constantly freeing/allocating blocks in bursty patterns
+ * (on most crucially cases of repeatedly allocating and freeing a
+ * single chunk), retain a small number of blocks.
+ */
+ if (slab->nfreeblocks < SLAB_RETAIN_EMPTY_BLOCK_COUNT)
+ {
+ dlist_push_head(&slab->freeblocks, &block->node);
+ slab->nfreeblocks++;
+ }
+ else
+ {
+ slab->nblocks--;
+ slab->header.mem_allocated -= slab->blockSize;
+ free(block);
+ }
+ }
+ else
+ dlist_push_head(SlabFreelist(slab, block->nfree), &block->node);
+}
+
/*
* SlabFree
* Frees allocated memory; memory is removed from the slab.
@@ -468,7 +594,6 @@ SlabFree(void *pointer)
MemoryChunk *chunk = PointerGetMemoryChunk(pointer);
SlabBlock *block = MemoryChunkGetBlock(chunk);
SlabContext *slab;
- int idx;
/*
* For speed reasons we just Assert that the referenced block is good.
@@ -478,6 +603,45 @@ SlabFree(void *pointer)
AssertArg(SlabBlockIsValid(block));
slab = block->slab;
+ Assert(slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree + 1));
+
+ /* move the whole block to the right place in the freelist */
+ if (unlikely(SlabFreelistIndex(slab, block->nfree) != slab->minFreeChunksIndex))
+ {
+ slab->minFreeChunksIndex = SlabFreelistIndex(slab, block->nfree);
+ dlist_delete(&block->node);
+ dlist_push_head(SlabFreelist(slab, block->nfree), &block->node);
+
+
+ /*
+ * And finally update minFreeChunks, i.e. the index to the block with the
+ * lowest number of free chunks. We only need to do that when the block
+ * got full (otherwise we know the current block is the right one). We'll
+ * simply walk the freelist until we find a non-empty entry.
+ */
+ if (slab->minFreeChunksIndex == 0)
+ {
+ for (int idx = 1; idx < SLAB_FREELIST_COUNT; idx++)
+ {
+ if (dlist_is_empty(&slab->freelist[idx]))
+ continue;
+
+ /* found a non-empty freelist */
+ slab->minFreeChunksIndex = idx;
+ break;
+ }
+ }
+ }
+
+#if 0
+ /*
+ * FIXME: I don't understand what this ever did? It should be unreachable
+ * I think?
+ */
+ if (slab->minFreeChunks == slab->chunksPerBlock)
+ slab->minFreeChunks = 0;
+#endif
+
#ifdef MEMORY_CONTEXT_CHECKING
/* Test for someone scribbling on unused space in chunk */
Assert(slab->chunkSize < (slab->fullChunkSize - Slab_CHUNKHDRSZ));
@@ -486,60 +650,22 @@ SlabFree(void *pointer)
slab->header.name, chunk);
#endif
- /* compute index of the chunk with respect to block start */
- idx = SlabChunkIndex(slab, block, chunk);
-
/* add chunk to freelist, and update block nfree count */
- *(int32 *) pointer = block->firstFreeChunk;
- block->firstFreeChunk = idx;
+ *(MemoryChunk **) pointer = (MemoryChunk *) ((char *) block + block->firstFreeOffset);
+ block->firstFreeOffset = (char *) chunk - (char *) block;
block->nfree++;
Assert(block->nfree > 0);
Assert(block->nfree <= slab->chunksPerBlock);
#ifdef CLOBBER_FREED_MEMORY
- /* XXX don't wipe the int32 index, used for block-level freelist */
+ /* XXX don't wipe the int32 offset, used for block-level freelist */
wipe_mem((char *) pointer + sizeof(int32),
slab->chunkSize - sizeof(int32));
#endif
- /* remove the block from a freelist */
- dlist_delete(&block->node);
-
- /*
- * See if we need to update the minFreeChunks field for the slab - we only
- * need to do that if there the block had that number of free chunks
- * before we freed one. In that case, we check if there still are blocks
- * in the original freelist and we either keep the current value (if there
- * still are blocks) or increment it by one (the new block is still the
- * one with minimum free chunks).
- *
- * The one exception is when the block will get completely free - in that
- * case we will free it, se we can't use it for minFreeChunks. It however
- * means there are no more blocks with free chunks.
- */
- if (slab->minFreeChunks == (block->nfree - 1))
- {
- /* Have we removed the last chunk from the freelist? */
- if (dlist_is_empty(&slab->freelist[slab->minFreeChunks]))
- {
- /* but if we made the block entirely free, we'll free it */
- if (block->nfree == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
- else
- slab->minFreeChunks++;
- }
- }
-
- /* If the block is now completely empty, free it. */
- if (block->nfree == slab->chunksPerBlock)
- {
- free(block);
- slab->nblocks--;
- slab->header.mem_allocated -= slab->blockSize;
- }
- else
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
+ if (SlabFreelistIndex(slab, block->nfree) != SlabFreelistIndex(slab, block->nfree - 1))
+ SlabFreeSlow(slab, block);
Assert(slab->nblocks >= 0);
Assert(slab->nblocks * slab->blockSize == slab->header.mem_allocated);
@@ -656,7 +782,7 @@ SlabStats(MemoryContext context,
/* Include context header in totalspace */
totalspace = slab->headerSize;
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
dlist_iter iter;
@@ -713,7 +839,7 @@ SlabCheck(MemoryContext context)
Assert(slab->chunksPerBlock > 0);
/* walk all the freelists */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
int j,
nfree;
@@ -722,14 +848,14 @@ SlabCheck(MemoryContext context)
/* walk all blocks on this freelist */
dlist_foreach(iter, &slab->freelist[i])
{
- int idx;
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+ MemoryChunk *cur_chunk;
/*
* Make sure the number of free chunks (in the block header)
* matches position in the freelist.
*/
- if (block->nfree != i)
+ if (SlabFreelistIndex(slab, block->nfree) != i)
elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match freelist %d",
name, block->nfree, block, i);
@@ -740,7 +866,6 @@ SlabCheck(MemoryContext context)
/* reset the bitmap of free chunks for this block */
memset(slab->freechunks, 0, (slab->chunksPerBlock * sizeof(bool)));
- idx = block->firstFreeChunk;
/*
* Now walk through the chunks, count the free ones and also
@@ -749,10 +874,12 @@ SlabCheck(MemoryContext context)
* walk through the chunks and construct our own bitmap.
*/
+ cur_chunk = (MemoryChunk *) ((char *) block) + block->firstChunkOffset);
nfree = 0;
- while (idx < slab->chunksPerBlock)
+ while (cur_chunk != NULL)
{
MemoryChunk *chunk;
+ int idx = SlabChunkIndex(slab, block, cur_chunk);
/* count the chunk as free, add it to the bitmap */
nfree++;
@@ -761,9 +888,21 @@ SlabCheck(MemoryContext context)
/* read index of the next free chunk */
chunk = SlabBlockGetChunk(slab, block, idx);
VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- idx = *(int32 *) MemoryChunkGetPointer(chunk);
+ cur_chunk = (char *) block + (int32) SlabChunkGetPointer(cur_chunk);
}
+ cur_chunk = block->unused;
+ for (j = 0; j < block->nunused; j++)
+ {
+ int idx = SlabChunkIndex(slab, block, cur_chunk);
+
+ /* count the chunk as free, add it to the bitmap */
+ nfree++;
+ slab->freechunks[idx] = true;
+
+ cur_chunk = (MemoryChunk *) (((char *) cur_chunk) + slab->fullChunkSize);
+ }
+
for (j = 0; j < slab->chunksPerBlock; j++)
{
/* non-zero bit in the bitmap means chunk the chunk is used */
--
2.35.1.windows.2
Attachments:
[text/plain] v3-0001-WIP-slab-performance.patch (23.6K, ../../CAApHDvoxVxFN0DXYyn6tDdg6s7wx2sVrVJ_JSCZxrfd-s86j8Q@mail.gmail.com/2-v3-0001-WIP-slab-performance.patch)
download | inline diff:
From 2b89d993b0294d5c0fe0a8333bccc555337ac979 Mon Sep 17 00:00:00 2001
From: David Rowley <dgrowley@gmail.com>
Date: Wed, 12 Oct 2022 09:30:24 +1300
Subject: [PATCH v3] WIP: slab performance.
Author: Andres Freund
Discussion: https://postgr.es/m/20210717194333.mr5io3zup3kxahfm@alap3.anarazel.de
---
src/backend/utils/mmgr/slab.c | 449 ++++++++++++++++++++++------------
1 file changed, 294 insertions(+), 155 deletions(-)
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index 1a0b28f9ea..58b8e2d67c 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -23,15 +23,18 @@
* global (context) level. This is possible as the chunk size (and thus also
* the number of chunks per block) is fixed.
*
- * On each block, free chunks are tracked in a simple linked list. Contents
- * of free chunks is replaced with an index of the next free chunk, forming
- * a very simple linked list. Each block also contains a counter of free
- * chunks. Combined with the local block-level freelist, it makes it trivial
- * to eventually free the whole block.
+ * On each block, never allocated chunks are tracked by a simple offset, and
+ * free chunks are tracked in a simple linked list. The offset approach
+ * avoids needing to iterate over all chunks when allocating a new block,
+ * which would cause page faults and cache pollution. Contents of free chunks
+ * is replaced with a the offset (in bytes) from the block pointer to the
+ * next free chunk, forming a very simple linked list. Each block also
+ * contains a counter of free chunks. Combined with the local block-level
+ * freelist, it makes it trivial to eventually free the whole block.
*
* At the context level, we use 'freelist' to track blocks ordered by number
* of free chunks, starting with blocks having a single allocated chunk, and
- * with completely full blocks on the tail.
+ * with completely full blocks on the tail. XXX
*
* This also allows various optimizations - for example when searching for
* free chunk, the allocator reuses space from the fullest blocks first, in
@@ -44,7 +47,7 @@
* case this performs as if the pointer was not maintained.
*
* We cache the freelist index for the blocks with the fewest free chunks
- * (minFreeChunks), so that we don't have to search the freelist on every
+ * (minFreeChunkIndex), so that we don't have to search the freelist on every
* SlabAlloc() call, which is quite expensive.
*
*-------------------------------------------------------------------------
@@ -60,6 +63,17 @@
#define Slab_BLOCKHDRSZ MAXALIGN(sizeof(SlabBlock))
+struct SlabBlock;
+struct SlabChunk;
+
+/*
+ * Number of actual freelists + 1, for full blocks. Full blocks are always at
+ * offset 0.
+ */
+#define SLAB_FREELIST_COUNT 9
+
+#define SLAB_RETAIN_EMPTY_BLOCK_COUNT 10
+
/*
* SlabContext is a specialized implementation of MemoryContext.
*/
@@ -72,13 +86,16 @@ typedef struct SlabContext
Size blockSize; /* block size */
Size headerSize; /* allocated size of context header */
int chunksPerBlock; /* number of chunks per block */
- int minFreeChunks; /* min number of free chunks in any block */
+ int minFreeChunksIndex; /* min number of free chunks in any block XXX */
int nblocks; /* number of blocks allocated */
#ifdef MEMORY_CONTEXT_CHECKING
bool *freechunks; /* bitmap of free chunks in a block */
#endif
+ int nfreeblocks;
+ dlist_head freeblocks;
+ int freelist_shift;
/* blocks with free space, grouped by number of free chunks: */
- dlist_head freelist[FLEXIBLE_ARRAY_MEMBER];
+ dlist_head freelist[SLAB_FREELIST_COUNT];
} SlabContext;
/*
@@ -92,8 +109,11 @@ typedef struct SlabContext
typedef struct SlabBlock
{
dlist_node node; /* doubly-linked list */
- int nfree; /* number of free chunks */
- int firstFreeChunk; /* index of the first free chunk in the block */
+ int nfree; /* number of chunks on freelist + unused */
+ int nunused; /* number of unused chunks */
+ int32 firstFreeOffset; /* offset of bytes from block to the first
+ * free chunk in the block. */
+ MemoryChunk *unused; /* */
SlabContext *slab; /* owning context */
} SlabBlock;
@@ -125,6 +145,34 @@ typedef struct SlabBlock
#define SlabBlockIsValid(block) \
(PointerIsValid(block) && SlabIsValid((block)->slab))
+static inline uint8
+SlabFreelistIndex(SlabContext *slab, int freecount)
+{
+ uint8 index;
+
+ Assert(freecount <= slab->chunksPerBlock);
+
+ /*
+ * Ensure, without a branch, that index 0 is only used for blocks entirely
+ * without free chunks.
+ * XXX: There probably is a cheaper way to do this. Needing to shift twice
+ * by slab->freelist_shift isn't great.
+ */
+ index = (freecount + (1 << slab->freelist_shift) - 1) >> slab->freelist_shift;
+
+ if (freecount == 0)
+ Assert(index == 0);
+ else
+ Assert(index > 0 && index < (SLAB_FREELIST_COUNT));
+
+ return index;
+}
+
+static inline dlist_head*
+SlabFreelist(SlabContext *slab, int freecount)
+{
+ return &slab->freelist[SlabFreelistIndex(slab, freecount)];
+}
/*
* SlabContextCreate
@@ -145,7 +193,6 @@ SlabContextCreate(MemoryContext parent,
{
int chunksPerBlock;
Size fullChunkSize;
- Size freelistSize;
Size headerSize;
SlabContext *slab;
int i;
@@ -155,9 +202,12 @@ SlabContextCreate(MemoryContext parent,
"sizeof(MemoryChunk) is not maxaligned");
Assert(MAXALIGN(chunkSize) <= MEMORYCHUNK_MAX_VALUE);
- /* Make sure the linked list node fits inside a freed chunk */
- if (chunkSize < sizeof(int))
- chunkSize = sizeof(int);
+ /*
+ * Ensure there's enough space to store the offset to the next free chunk
+ * in the memory of free'd chunks.
+ */
+ if (chunkSize < sizeof(uint32))
+ chunkSize = sizeof(uint32);
/* chunk, including SLAB header (both addresses nicely aligned) */
#ifdef MEMORY_CONTEXT_CHECKING
@@ -175,16 +225,14 @@ SlabContextCreate(MemoryContext parent,
/* Compute maximum number of chunks per block */
chunksPerBlock = (blockSize - Slab_BLOCKHDRSZ) / fullChunkSize;
- /* The freelist starts with 0, ends with chunksPerBlock. */
- freelistSize = sizeof(dlist_head) * (chunksPerBlock + 1);
-
/*
* Allocate the context header. Unlike aset.c, we never try to combine
* this with the first regular block; not worth the extra complication.
+ * XXX: What's the evidence for that?
*/
/* Size of the memory context header */
- headerSize = offsetof(SlabContext, freelist) + freelistSize;
+ headerSize = sizeof(SlabContext);
#ifdef MEMORY_CONTEXT_CHECKING
@@ -218,17 +266,34 @@ SlabContextCreate(MemoryContext parent,
slab->blockSize = blockSize;
slab->headerSize = headerSize;
slab->chunksPerBlock = chunksPerBlock;
- slab->minFreeChunks = 0;
+ slab->minFreeChunksIndex = 0;
slab->nblocks = 0;
+ slab->nfreeblocks = 0;
+
+ /*
+ * Compute a shift that guarantees that shifting chunksPerBlock with it
+ * yields is smaller than SLAB_FREELIST_COUNT - 1 (one freelist is used for full blocks).
+ */
+ slab->freelist_shift = 0;
+ while ((slab->chunksPerBlock >> slab->freelist_shift) >= (SLAB_FREELIST_COUNT - 1))
+ slab->freelist_shift++;
+
+#if 0
+ elog(LOG, "freelist shift for %d chunks of size %zu is %d, block size %zu",
+ slab->chunksPerBlock, slab->fullChunkSize, slab->freelist_shift,
+ slab->blockSize);
+#endif
+
+ dlist_init(&slab->freeblocks);
/* initialize the freelist slots */
- for (i = 0; i < (slab->chunksPerBlock + 1); i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
dlist_init(&slab->freelist[i]);
#ifdef MEMORY_CONTEXT_CHECKING
/* set the freechunks pointer right after the freelists array */
slab->freechunks
- = (bool *) slab + offsetof(SlabContext, freelist) + freelistSize;
+ = (bool *) slab + sizeof(SlabContext);
#endif
/* Finally, do the type-independent part of context creation */
@@ -261,8 +326,27 @@ SlabReset(MemoryContext context)
SlabCheck(context);
#endif
+ /* release retained empty blocks */
+ {
+ dlist_mutable_iter miter;
+
+ dlist_foreach_modify(miter, &slab->freeblocks)
+ {
+ SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
+
+ dlist_delete(miter.cur);
+
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ slab->nblocks--;
+ context->mem_allocated -= slab->blockSize;
+ }
+ }
+
/* walk over freelists and free the blocks */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
dlist_mutable_iter miter;
@@ -281,7 +365,7 @@ SlabReset(MemoryContext context)
}
}
- slab->minFreeChunks = 0;
+ slab->minFreeChunksIndex = 0;
Assert(slab->nblocks == 0);
Assert(context->mem_allocated == 0);
@@ -311,12 +395,11 @@ SlabAlloc(MemoryContext context, Size size)
SlabContext *slab = (SlabContext *) context;
SlabBlock *block;
MemoryChunk *chunk;
- int idx;
AssertArg(SlabIsValid(slab));
- Assert((slab->minFreeChunks >= 0) &&
- (slab->minFreeChunks < slab->chunksPerBlock));
+ Assert((slab->minFreeChunksIndex >= 0) &&
+ (slab->minFreeChunksIndex < SLAB_FREELIST_COUNT));
/* make sure we only allow correct request size */
if (size != slab->chunkSize)
@@ -327,62 +410,93 @@ SlabAlloc(MemoryContext context, Size size)
* If there are no free chunks in any existing block, create a new block
* and put it to the last freelist bucket.
*
- * slab->minFreeChunks == 0 means there are no blocks with free chunks,
- * thanks to how minFreeChunks is updated at the end of SlabAlloc().
+ * slab->minFreeChunksIndex == 0 means there are no blocks with free
+ * chunks, thanks to how minFreeChunks is updated at the end of
+ * SlabAlloc().
*/
- if (slab->minFreeChunks == 0)
+ if (unlikely(slab->minFreeChunksIndex == 0))
{
- block = (SlabBlock *) malloc(slab->blockSize);
+ if (slab->nfreeblocks > 0)
+ {
+ dlist_node *node;
- if (block == NULL)
- return NULL;
+ node = dlist_pop_head_node(&slab->freeblocks);
+ block = dlist_container(SlabBlock, node, node);
+ slab->nfreeblocks--;
+ }
+ else
+ {
+ block = (SlabBlock *) malloc(slab->blockSize);
- block->nfree = slab->chunksPerBlock;
- block->firstFreeChunk = 0;
- block->slab = slab;
+ if (unlikely(block == NULL))
+ return NULL;
- /*
- * Put all the chunks on a freelist. Walk the chunks and point each
- * one to the next one.
- */
- for (idx = 0; idx < slab->chunksPerBlock; idx++)
- {
- chunk = SlabBlockGetChunk(slab, block, idx);
- *(int32 *) MemoryChunkGetPointer(chunk) = (idx + 1);
+ slab->nblocks += 1;
+ block->slab = slab;
+ context->mem_allocated += slab->blockSize;
}
+ block->nfree = slab->chunksPerBlock;
+ block->firstFreeOffset = -1;
+ block->nunused = slab->chunksPerBlock;
+ block->unused = (MemoryChunk *) SlabBlockStart(block);
+
+ slab->minFreeChunksIndex = SlabFreelistIndex(slab, slab->chunksPerBlock);
+
/*
* And add it to the last freelist with all chunks empty.
*
* We know there are no blocks in the freelist, otherwise we wouldn't
* need a new block.
*/
- Assert(dlist_is_empty(&slab->freelist[slab->chunksPerBlock]));
+ Assert(dlist_is_empty(SlabFreelist(slab, slab->chunksPerBlock)));
- dlist_push_head(&slab->freelist[slab->chunksPerBlock], &block->node);
+ dlist_push_head(SlabFreelist(slab, slab->chunksPerBlock),
+ &block->node);
- slab->minFreeChunks = slab->chunksPerBlock;
- slab->nblocks += 1;
- context->mem_allocated += slab->blockSize;
+ chunk = block->unused;
+ block->unused = (MemoryChunk *) (((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
}
+ else
+ {
+ /* grab the block from the freelist */
+ block = dlist_head_element(SlabBlock, node,
+ &slab->freelist[slab->minFreeChunksIndex]);
- /* grab the block from the freelist (even the new block is there) */
- block = dlist_head_element(SlabBlock, node,
- &slab->freelist[slab->minFreeChunks]);
+ /* make sure we actually got a valid block, with matching nfree */
+ Assert(block != NULL);
+ Assert(slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree));
+ Assert(block->nfree > 0);
- /* make sure we actually got a valid block, with matching nfree */
- Assert(block != NULL);
- Assert(slab->minFreeChunks == block->nfree);
- Assert(block->nfree > 0);
+ /* we know index of the first free chunk in the block */
+ if (block->nunused > 0)
+ {
+ chunk = block->unused;
+ block->unused = (MemoryChunk *) (((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
+ else
+ {
+ chunk = (MemoryChunk *) ((char *) block + block->firstFreeOffset);
- /* we know index of the first free chunk in the block */
- idx = block->firstFreeChunk;
+ /*
+ * Remove the chunk from the freelist head. The index of the next free
+ * chunk is stored in the chunk itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(void*));
+ block->firstFreeOffset = *(int32 *) SlabChunkGetPointer(chunk);
+ }
+ }
- /* make sure the chunk index is valid, and that it's marked as empty */
- Assert((idx >= 0) && (idx < slab->chunksPerBlock));
+ /* make sure the chunk is in the block and that it's marked as empty (XXX?) */
+ Assert((char *) chunk >= SlabBlockStart(block));
+ Assert((char *) chunk < (((char *)block) + slab->blockSize));
- /* compute the chunk location block start (after the block header) */
- chunk = SlabBlockGetChunk(slab, block, idx);
+ Assert(block->firstFreeChunk == NULL || (
+ (block->firstFreeChunk >= (SlabChunk *) SlabBlockStart(block)) &&
+ block->firstFreeChunk <= (SlabChunk *) (((char *)block) + slab->blockSize))
+ );
/*
* Update the block nfree count, and also the minFreeChunks as we've
@@ -390,48 +504,6 @@ SlabAlloc(MemoryContext context, Size size)
* (because that's how we chose the block).
*/
block->nfree--;
- slab->minFreeChunks = block->nfree;
-
- /*
- * Remove the chunk from the freelist head. The index of the next free
- * chunk is stored in the chunk itself.
- */
- VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- block->firstFreeChunk = *(int32 *) MemoryChunkGetPointer(chunk);
-
- Assert(block->firstFreeChunk >= 0);
- Assert(block->firstFreeChunk <= slab->chunksPerBlock);
-
- Assert((block->nfree != 0 &&
- block->firstFreeChunk < slab->chunksPerBlock) ||
- (block->nfree == 0 &&
- block->firstFreeChunk == slab->chunksPerBlock));
-
- /* move the whole block to the right place in the freelist */
- dlist_delete(&block->node);
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
-
- /*
- * And finally update minFreeChunks, i.e. the index to the block with the
- * lowest number of free chunks. We only need to do that when the block
- * got full (otherwise we know the current block is the right one). We'll
- * simply walk the freelist until we find a non-empty entry.
- */
- if (slab->minFreeChunks == 0)
- {
- for (idx = 1; idx <= slab->chunksPerBlock; idx++)
- {
- if (dlist_is_empty(&slab->freelist[idx]))
- continue;
-
- /* found a non-empty freelist */
- slab->minFreeChunks = idx;
- break;
- }
- }
-
- if (slab->minFreeChunks == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
/* Prepare to initialize the chunk header. */
VALGRIND_MAKE_MEM_UNDEFINED(chunk, Slab_CHUNKHDRSZ);
@@ -458,6 +530,60 @@ SlabAlloc(MemoryContext context, Size size)
return MemoryChunkGetPointer(chunk);
}
+static void pg_noinline
+SlabFreeSlow(SlabContext *slab, SlabBlock *block)
+{
+ dlist_delete(&block->node);
+
+ /*
+ * See if we need to update the minFreeChunks field for the slab - we only
+ * need to do that if the block had that number of free chunks
+ * before we freed one. In that case, we check if there still are blocks
+ * in the original freelist and we either keep the current value (if there
+ * still are blocks) or increment it by one (the new block is still the
+ * one with minimum free chunks).
+ *
+ * The one exception is when the block will get completely free - in that
+ * case we will free it, se we can't use it for minFreeChunks. It however
+ * means there are no more blocks with free chunks.
+ */
+ if (slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree - 1))
+ {
+ /* Have we removed the last chunk from the freelist? */
+ if (dlist_is_empty(&slab->freelist[slab->minFreeChunksIndex]))
+ {
+ /* but if we made the block entirely free, we'll free it */
+ if (block->nfree == slab->chunksPerBlock)
+ slab->minFreeChunksIndex = 0;
+ else
+ slab->minFreeChunksIndex++;
+ }
+ }
+
+ /* If the block is now completely empty, free it (in a way). */
+ if (block->nfree == slab->chunksPerBlock)
+ {
+ /*
+ * To avoid constantly freeing/allocating blocks in bursty patterns
+ * (on most crucially cases of repeatedly allocating and freeing a
+ * single chunk), retain a small number of blocks.
+ */
+ if (slab->nfreeblocks < SLAB_RETAIN_EMPTY_BLOCK_COUNT)
+ {
+ dlist_push_head(&slab->freeblocks, &block->node);
+ slab->nfreeblocks++;
+ }
+ else
+ {
+ slab->nblocks--;
+ slab->header.mem_allocated -= slab->blockSize;
+ free(block);
+ }
+ }
+ else
+ dlist_push_head(SlabFreelist(slab, block->nfree), &block->node);
+}
+
/*
* SlabFree
* Frees allocated memory; memory is removed from the slab.
@@ -468,7 +594,6 @@ SlabFree(void *pointer)
MemoryChunk *chunk = PointerGetMemoryChunk(pointer);
SlabBlock *block = MemoryChunkGetBlock(chunk);
SlabContext *slab;
- int idx;
/*
* For speed reasons we just Assert that the referenced block is good.
@@ -478,6 +603,45 @@ SlabFree(void *pointer)
AssertArg(SlabBlockIsValid(block));
slab = block->slab;
+ Assert(slab->minFreeChunksIndex == SlabFreelistIndex(slab, block->nfree + 1));
+
+ /* move the whole block to the right place in the freelist */
+ if (unlikely(SlabFreelistIndex(slab, block->nfree) != slab->minFreeChunksIndex))
+ {
+ slab->minFreeChunksIndex = SlabFreelistIndex(slab, block->nfree);
+ dlist_delete(&block->node);
+ dlist_push_head(SlabFreelist(slab, block->nfree), &block->node);
+
+
+ /*
+ * And finally update minFreeChunks, i.e. the index to the block with the
+ * lowest number of free chunks. We only need to do that when the block
+ * got full (otherwise we know the current block is the right one). We'll
+ * simply walk the freelist until we find a non-empty entry.
+ */
+ if (slab->minFreeChunksIndex == 0)
+ {
+ for (int idx = 1; idx < SLAB_FREELIST_COUNT; idx++)
+ {
+ if (dlist_is_empty(&slab->freelist[idx]))
+ continue;
+
+ /* found a non-empty freelist */
+ slab->minFreeChunksIndex = idx;
+ break;
+ }
+ }
+ }
+
+#if 0
+ /*
+ * FIXME: I don't understand what this ever did? It should be unreachable
+ * I think?
+ */
+ if (slab->minFreeChunks == slab->chunksPerBlock)
+ slab->minFreeChunks = 0;
+#endif
+
#ifdef MEMORY_CONTEXT_CHECKING
/* Test for someone scribbling on unused space in chunk */
Assert(slab->chunkSize < (slab->fullChunkSize - Slab_CHUNKHDRSZ));
@@ -486,60 +650,22 @@ SlabFree(void *pointer)
slab->header.name, chunk);
#endif
- /* compute index of the chunk with respect to block start */
- idx = SlabChunkIndex(slab, block, chunk);
-
/* add chunk to freelist, and update block nfree count */
- *(int32 *) pointer = block->firstFreeChunk;
- block->firstFreeChunk = idx;
+ *(MemoryChunk **) pointer = (MemoryChunk *) ((char *) block + block->firstFreeOffset);
+ block->firstFreeOffset = (char *) chunk - (char *) block;
block->nfree++;
Assert(block->nfree > 0);
Assert(block->nfree <= slab->chunksPerBlock);
#ifdef CLOBBER_FREED_MEMORY
- /* XXX don't wipe the int32 index, used for block-level freelist */
+ /* XXX don't wipe the int32 offset, used for block-level freelist */
wipe_mem((char *) pointer + sizeof(int32),
slab->chunkSize - sizeof(int32));
#endif
- /* remove the block from a freelist */
- dlist_delete(&block->node);
-
- /*
- * See if we need to update the minFreeChunks field for the slab - we only
- * need to do that if there the block had that number of free chunks
- * before we freed one. In that case, we check if there still are blocks
- * in the original freelist and we either keep the current value (if there
- * still are blocks) or increment it by one (the new block is still the
- * one with minimum free chunks).
- *
- * The one exception is when the block will get completely free - in that
- * case we will free it, se we can't use it for minFreeChunks. It however
- * means there are no more blocks with free chunks.
- */
- if (slab->minFreeChunks == (block->nfree - 1))
- {
- /* Have we removed the last chunk from the freelist? */
- if (dlist_is_empty(&slab->freelist[slab->minFreeChunks]))
- {
- /* but if we made the block entirely free, we'll free it */
- if (block->nfree == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
- else
- slab->minFreeChunks++;
- }
- }
-
- /* If the block is now completely empty, free it. */
- if (block->nfree == slab->chunksPerBlock)
- {
- free(block);
- slab->nblocks--;
- slab->header.mem_allocated -= slab->blockSize;
- }
- else
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
+ if (SlabFreelistIndex(slab, block->nfree) != SlabFreelistIndex(slab, block->nfree - 1))
+ SlabFreeSlow(slab, block);
Assert(slab->nblocks >= 0);
Assert(slab->nblocks * slab->blockSize == slab->header.mem_allocated);
@@ -656,7 +782,7 @@ SlabStats(MemoryContext context,
/* Include context header in totalspace */
totalspace = slab->headerSize;
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
dlist_iter iter;
@@ -713,7 +839,7 @@ SlabCheck(MemoryContext context)
Assert(slab->chunksPerBlock > 0);
/* walk all the freelists */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ for (i = 0; i < SLAB_FREELIST_COUNT; i++)
{
int j,
nfree;
@@ -722,14 +848,14 @@ SlabCheck(MemoryContext context)
/* walk all blocks on this freelist */
dlist_foreach(iter, &slab->freelist[i])
{
- int idx;
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+ MemoryChunk *cur_chunk;
/*
* Make sure the number of free chunks (in the block header)
* matches position in the freelist.
*/
- if (block->nfree != i)
+ if (SlabFreelistIndex(slab, block->nfree) != i)
elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match freelist %d",
name, block->nfree, block, i);
@@ -740,7 +866,6 @@ SlabCheck(MemoryContext context)
/* reset the bitmap of free chunks for this block */
memset(slab->freechunks, 0, (slab->chunksPerBlock * sizeof(bool)));
- idx = block->firstFreeChunk;
/*
* Now walk through the chunks, count the free ones and also
@@ -749,10 +874,12 @@ SlabCheck(MemoryContext context)
* walk through the chunks and construct our own bitmap.
*/
+ cur_chunk = (MemoryChunk *) ((char *) block) + block->firstChunkOffset);
nfree = 0;
- while (idx < slab->chunksPerBlock)
+ while (cur_chunk != NULL)
{
MemoryChunk *chunk;
+ int idx = SlabChunkIndex(slab, block, cur_chunk);
/* count the chunk as free, add it to the bitmap */
nfree++;
@@ -761,9 +888,21 @@ SlabCheck(MemoryContext context)
/* read index of the next free chunk */
chunk = SlabBlockGetChunk(slab, block, idx);
VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- idx = *(int32 *) MemoryChunkGetPointer(chunk);
+ cur_chunk = (char *) block + (int32) SlabChunkGetPointer(cur_chunk);
}
+ cur_chunk = block->unused;
+ for (j = 0; j < block->nunused; j++)
+ {
+ int idx = SlabChunkIndex(slab, block, cur_chunk);
+
+ /* count the chunk as free, add it to the bitmap */
+ nfree++;
+ slab->freechunks[idx] = true;
+
+ cur_chunk = (MemoryChunk *) (((char *) cur_chunk) + slab->fullChunkSize);
+ }
+
for (j = 0; j < slab->chunksPerBlock; j++)
{
/* non-zero bit in the bitmap means chunk the chunk is used */
--
2.35.1.windows.2
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
@ 2022-11-11 09:20 ` John Naylor <john.naylor@enterprisedb.com>
2022-12-05 08:02 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
0 siblings, 1 reply; 29+ messages in thread
From: John Naylor @ 2022-11-11 09:20 UTC (permalink / raw)
To: David Rowley <dgrowleyml@gmail.com>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On Wed, Oct 12, 2022 at 4:37 PM David Rowley <dgrowleyml@gmail.com> wrote:
> [v3]
+ /*
+ * Compute a shift that guarantees that shifting chunksPerBlock with it
+ * yields is smaller than SLAB_FREELIST_COUNT - 1 (one freelist is used
for full blocks).
+ */
+ slab->freelist_shift = 0;
+ while ((slab->chunksPerBlock >> slab->freelist_shift) >=
(SLAB_FREELIST_COUNT - 1))
+ slab->freelist_shift++;
+ /*
+ * Ensure, without a branch, that index 0 is only used for blocks entirely
+ * without free chunks.
+ * XXX: There probably is a cheaper way to do this. Needing to shift twice
+ * by slab->freelist_shift isn't great.
+ */
+ index = (freecount + (1 << slab->freelist_shift) - 1) >>
slab->freelist_shift;
How about something like
#define SLAB_FREELIST_COUNT ((1<<3) + 1)
index = (freecount & (SLAB_FREELIST_COUNT - 2)) + (freecount != 0);
and dispense with both freelist_shift and the loop that computes it?
--
John Naylor
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
@ 2022-12-05 08:02 ` David Rowley <dgrowleyml@gmail.com>
2022-12-05 10:18 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
0 siblings, 1 reply; 29+ messages in thread
From: David Rowley @ 2022-12-05 08:02 UTC (permalink / raw)
To: John Naylor <john.naylor@enterprisedb.com>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On Fri, 11 Nov 2022 at 22:20, John Naylor <john.naylor@enterprisedb.com> wrote:
>
>
> On Wed, Oct 12, 2022 at 4:37 PM David Rowley <dgrowleyml@gmail.com> wrote:
> > [v3]
>
> + /*
> + * Compute a shift that guarantees that shifting chunksPerBlock with it
> + * yields is smaller than SLAB_FREELIST_COUNT - 1 (one freelist is used for full blocks).
> + */
> + slab->freelist_shift = 0;
> + while ((slab->chunksPerBlock >> slab->freelist_shift) >= (SLAB_FREELIST_COUNT - 1))
> + slab->freelist_shift++;
>
> + /*
> + * Ensure, without a branch, that index 0 is only used for blocks entirely
> + * without free chunks.
> + * XXX: There probably is a cheaper way to do this. Needing to shift twice
> + * by slab->freelist_shift isn't great.
> + */
> + index = (freecount + (1 << slab->freelist_shift) - 1) >> slab->freelist_shift;
>
> How about something like
>
> #define SLAB_FREELIST_COUNT ((1<<3) + 1)
> index = (freecount & (SLAB_FREELIST_COUNT - 2)) + (freecount != 0);
Doesn't this create a sort of round-robin use of the free list? What
we want is a sort of "histogram" bucket set of free lists so we can
group together blocks that have a close-enough free number of chunks.
Unless I'm mistaken, I think what you have doesn't do that.
I wondered if simply:
index = -(-freecount >> slab->freelist_shift);
would be faster than Andres' version. I tried it out and on my AMD
machine, it's about the same speed. Same on a Raspberry Pi 4.
Going by [2], the instructions are very different with each method, so
other machines with different latencies on those instructions might
show something different. I attached what I used to test if anyone
else wants a go.
AMD Zen2
$ ./freecount 2000000000
Test 'a' in 0.922766 seconds
Test 'd' in 0.922762 seconds (0.000433% faster)
RPI4
$ ./freecount 2000000000
Test 'a' in 3.341350 seconds
Test 'd' in 3.338690 seconds (0.079672% faster)
That was gcc. Trying it with clang, it went in a little heavy-handed
and optimized out my loop, so some more trickery might be needed for a
useful test on that compiler.
David
[2] https://godbolt.org/z/dh95TohEG
#include <stdio.h>
#include <time.h>
#include <stdlib.h>
#define SLAB_FREELIST_COUNT 9
int main(int argc, char **argv)
{
clock_t start, end;
double v1_time, v2_time;
int freecount;
int freelist_shift = 0;
int chunksPerBlock;
int index;
int zerocount = 0;
if (argc < 2)
{
printf("Syntax: %s <chunksPerBlock>\n", argv[0]);
return -1;
}
chunksPerBlock = atoi(argv[1]);
printf("chunksPerBlock = %d\n", chunksPerBlock);
while ((chunksPerBlock >> freelist_shift) >= (SLAB_FREELIST_COUNT - 1))
freelist_shift++;
printf("freelist_shift = %d\n", freelist_shift);
start = clock();
for (freecount = 0; freecount <= chunksPerBlock; freecount++)
{
index = (freecount + (1 << freelist_shift) - 1) >> freelist_shift;
/* try to prevent optimizing the above out */
if (index == 0)
zerocount++;
}
end = clock();
v1_time = (double) (end - start) / CLOCKS_PER_SEC;
printf("Test 'a' in %f seconds\n", v1_time);
printf("zerocount = %d\n", zerocount);
zerocount = 0;
start = clock();
for (freecount = 0; freecount <= chunksPerBlock; freecount++)
{
index = -(-freecount >> freelist_shift);
/* try to prevent optimizing the above out */
if (index == 0)
zerocount++;
}
end = clock();
v2_time = (double) (end - start) / CLOCKS_PER_SEC;
printf("Test 'd' in %f seconds (%f%% faster)\n", v2_time, v1_time / v2_time * 100.0 - 100.0);
printf("zerocount = %d\n", zerocount);
return 0;
}
Attachments:
[text/plain] freecount.c (1.4K, ../../CAApHDvoUN6Kh3KBrybprL0ApLuGe_oBfjqYGW_Dff21VzZDw4Q@mail.gmail.com/2-freecount.c)
download | inline:
#include <stdio.h>
#include <time.h>
#include <stdlib.h>
#define SLAB_FREELIST_COUNT 9
int main(int argc, char **argv)
{
clock_t start, end;
double v1_time, v2_time;
int freecount;
int freelist_shift = 0;
int chunksPerBlock;
int index;
int zerocount = 0;
if (argc < 2)
{
printf("Syntax: %s <chunksPerBlock>\n", argv[0]);
return -1;
}
chunksPerBlock = atoi(argv[1]);
printf("chunksPerBlock = %d\n", chunksPerBlock);
while ((chunksPerBlock >> freelist_shift) >= (SLAB_FREELIST_COUNT - 1))
freelist_shift++;
printf("freelist_shift = %d\n", freelist_shift);
start = clock();
for (freecount = 0; freecount <= chunksPerBlock; freecount++)
{
index = (freecount + (1 << freelist_shift) - 1) >> freelist_shift;
/* try to prevent optimizing the above out */
if (index == 0)
zerocount++;
}
end = clock();
v1_time = (double) (end - start) / CLOCKS_PER_SEC;
printf("Test 'a' in %f seconds\n", v1_time);
printf("zerocount = %d\n", zerocount);
zerocount = 0;
start = clock();
for (freecount = 0; freecount <= chunksPerBlock; freecount++)
{
index = -(-freecount >> freelist_shift);
/* try to prevent optimizing the above out */
if (index == 0)
zerocount++;
}
end = clock();
v2_time = (double) (end - start) / CLOCKS_PER_SEC;
printf("Test 'd' in %f seconds (%f%% faster)\n", v2_time, v1_time / v2_time * 100.0 - 100.0);
printf("zerocount = %d\n", zerocount);
return 0;
}
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-05 08:02 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
@ 2022-12-05 10:18 ` John Naylor <john.naylor@enterprisedb.com>
2022-12-10 04:01 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
0 siblings, 1 reply; 29+ messages in thread
From: John Naylor @ 2022-12-05 10:18 UTC (permalink / raw)
To: David Rowley <dgrowleyml@gmail.com>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On Mon, Dec 5, 2022 at 3:02 PM David Rowley <dgrowleyml@gmail.com> wrote:
>
> On Fri, 11 Nov 2022 at 22:20, John Naylor <john.naylor@enterprisedb.com>
wrote:
> > #define SLAB_FREELIST_COUNT ((1<<3) + 1)
> > index = (freecount & (SLAB_FREELIST_COUNT - 2)) + (freecount != 0);
>
> Doesn't this create a sort of round-robin use of the free list? What
> we want is a sort of "histogram" bucket set of free lists so we can
> group together blocks that have a close-enough free number of chunks.
> Unless I'm mistaken, I think what you have doesn't do that.
The intent must have slipped my mind along the way.
> I wondered if simply:
>
> index = -(-freecount >> slab->freelist_shift);
>
> would be faster than Andres' version. I tried it out and on my AMD
> machine, it's about the same speed. Same on a Raspberry Pi 4.
>
> Going by [2], the instructions are very different with each method, so
> other machines with different latencies on those instructions might
> show something different. I attached what I used to test if anyone
> else wants a go.
I get about 0.1% difference on my machine. Both ways boil down to (on gcc)
3 instructions with low latency. The later ones need the prior results to
execute, which I think is what the XXX comment "isn't great" was referring
to. The new coding is more mysterious (does it do the right thing on all
platforms?), so I guess the original is still the way to go unless we get a
better idea.
--
John Naylor
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-05 08:02 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-05 10:18 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
@ 2022-12-10 04:01 ` David Rowley <dgrowleyml@gmail.com>
2022-12-12 07:13 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
0 siblings, 1 reply; 29+ messages in thread
From: David Rowley @ 2022-12-10 04:01 UTC (permalink / raw)
To: John Naylor <john.naylor@enterprisedb.com>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
.On Mon, 5 Dec 2022 at 23:18, John Naylor <john.naylor@enterprisedb.com> wrote:
>
>
> On Mon, Dec 5, 2022 at 3:02 PM David Rowley <dgrowleyml@gmail.com> wrote:
> > Going by [2], the instructions are very different with each method, so
> > other machines with different latencies on those instructions might
> > show something different. I attached what I used to test if anyone
> > else wants a go.
>
> I get about 0.1% difference on my machine. Both ways boil down to (on gcc) 3 instructions with low latency. The later ones need the prior results to execute, which I think is what the XXX comment "isn't great" was referring to. The new coding is more mysterious (does it do the right thing on all platforms?), so I guess the original is still the way to go unless we get a better idea.
I don't think it would work well on a one's complement machine, but I
don't think we support those going by the comments above RIGHTMOST_ONE
in bitmapset.c. In anycase, I found that it wasn't any faster than
what Andres wrote. In fact, even changing the code there to "index =
(freecount > 0);" seems to do very little to increase performance. I
do see that having 3 freelist items performs a decent amount better
than having 9. However, which workload is run may alter the result of
that, assuming that keeping new allocations on fuller blocks is a
winning strategy for the CPU's caches.
I've now done quite a bit more work on Andres' patch to try and get it
into (hopefully) somewhere close to a committable shape.
I'm fairly happy with what's there now. It does seem to perform much
better than current master, especially so when the workload would have
caused master to continually malloc and free an entire block when
freeing and allocating a single chunk.
I've done some basic benchmarking, mostly using the attached alloc_bench patch.
If I run:
select *, round(slab_result / aset_result * 100 - 100,1)::text || '%'
as slab_slower_by
from (
select
chunk_size,
keep_chunks,
sum(pg_allocate_memory_test(chunk_size, chunk_size * keep_chunks,
1024*1024*1024, 'slab'))::numeric(1000,3) as slab_result,
sum(pg_allocate_memory_test(chunk_size, chunk_size * keep_chunks,
1024*1024*1024, 'aset'))::numeric(1000,3) as aset_result
from
(values(1),(10),(50),(100),(200),(300),(400),(500),(1000),(2000),(3000),(4000),(5000),(10000))
v1(keep_chunks),
(values(64),(128),(256),(512),(1024)) v2(chunk_size)
group by rollup(1,2)
);
The results for the 64-byte chunk are shown in the attached chart.
It's not as fast as aset, but much faster than master's slab.c
The first blue bar of the chart is well above the vertical axis. It
took master 1118958 milliseconds for that test. The attached patch
took 194 ms. The rest of the tests seem to put the patched code around
somewhere in the middle between the unpatched code and aset.c's
performance.
The benchmark I did was entirely a FIFO workload that keeps around
"keep_chunks" at once before starting to free the oldest chunks. I
know Tomas has some LIFO and random benchmarks. I edited that code [1]
a little to add support for other context types so that a comparison
could be done more easily, however, I'm getting very weird performance
results where sometimes it runs about twice as fast (Tomas mentioned
he got this too). I'm not seeing that with my own benchmarking
function, so I'm wondering if there's something weird going on with
the benchmark itself rather than the slab.c code.
I've likely made much more changes than I can list here, but here are
a few of the more major ones:
1. Make use of dclist for empty blocks
2. In SlabAlloc() allocate chunks from the freelist before the unused list.
3. Added output for showing information about empty blocks in the
SlabStats output.
4. Renamed the context freelists to blocklist. I found this was
likely just confusing things with the block-level freelist. In any
case, it seemed weird to have freelist[0] store full blocks. Not much
free there! I renamed to blocklist[].
5. I did a round of changing the order of the fields in SlabBlock.
This seems to affect performance quite a bit. Having the context first
seems to improve performance. Having the blocklist[] node last also
helps.
6. Removed nblocks and headerSize from SlabContext. headerSize is no
longer needed. nblocks was only really used for Asserts and
SlabIsEmpty. I changed the Asserts to use a local count of blocks and
changed SlabIsEmpty to look at the context's mem_allocated.
7. There's now no integer division in any of the alloc and free code.
The only time we divide by fullChunkSize is in the check function.
8. When using a block from the emptyblock list, I changed the code to
not re-init the block. It's now used as it was left previously. This
means no longer having to build the freelist again.
9. Updated all comments to reflect the current state of the code.
Some things I thought about but didn't do:
a. Change the size of SlabContext's chunkSize, fullChunkSize and
blockSize to be uint32 instead of Size. It might be possible to get
SlabContext below 128 bytes with a bit more work.
b. I could have done a bit more experimentation with unlikely() and
likely() to move less frequently accessed code off into a cold area.
For #2 above, I didn't really see much change in performance when I
swapped the order of what we allocate from first. I expected free
chunks would be better as they've been used and are seemingly more
likely to be in some CPU cache than one of the unused chunks. I might
need a different allocation pattern than the one I used to highlight
that fact though.
David
[1] https://github.com/david-rowley/postgres/tree/alloc_bench_contrib
From b970576654bbce5c57690b8ab49d0b4376d1b5d3 Mon Sep 17 00:00:00 2001
From: David Rowley <dgrowley@gmail.com>
Date: Wed, 12 Oct 2022 09:30:24 +1300
Subject: [PATCH v4] Improve the performance of the slab memory allocator
Author: Andres Freund, David Rowley
Discussion: https://postgr.es/m/20210717194333.mr5io3zup3kxahfm@alap3.anarazel.de
---
src/backend/utils/mmgr/slab.c | 750 ++++++++++++++++++++++------------
1 file changed, 490 insertions(+), 260 deletions(-)
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index 6df0839b6a..2b95a6c061 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -3,8 +3,8 @@
* slab.c
* SLAB allocator definitions.
*
- * SLAB is a MemoryContext implementation designed for cases where large
- * numbers of equally-sized objects are allocated (and freed).
+ * SLAB is a MemoryContext implementation designed for cases where a large
+ * numbers of equally-sized objects can be allocated and freed efficiently.
*
*
* Portions Copyright (c) 2017-2022, PostgreSQL Global Development Group
@@ -19,33 +19,41 @@
* into chunks of exactly the right size (plus alignment), not wasting any
* memory.
*
- * The information about free chunks is maintained both at the block level and
- * global (context) level. This is possible as the chunk size (and thus also
- * the number of chunks per block) is fixed.
+ * Slab can also help reduce memory fragmentation where chunks remain stored
+ * on blocks with no other or only a few other allocated chunks. If this
+ * happens then the entire block cannot be freed until the last remaining
+ * chunk is freed. Slab helps work around this problem by prioritizing
+ * storing newly allocated chunks starting with the fullest blocks first.
+ * This makes it more likely that blocks with only a small number of
+ * remaining chunks eventually get freed when the final remaining chunk is
+ * freed.
*
- * On each block, free chunks are tracked in a simple linked list. Contents
- * of free chunks is replaced with an index of the next free chunk, forming
- * a very simple linked list. Each block also contains a counter of free
- * chunks. Combined with the local block-level freelist, it makes it trivial
- * to eventually free the whole block.
+ * We are easily able to find the fullest block to store a new chunk on as we
+ * maintain a list of all blocks and partition that list by the block's
+ * number of free chunks. This does mean having to possibly move the block
+ * onto another list whenever we allocate or free a chunk from it. However,
+ * this problem is significantly reduced by only having a small fixed number
+ * of lists for blocks, each of which allows storage of blocks with ranges of
+ * free chunks. The block only needs to be moved to another list when the
+ * number of free chunks crosses the range boundary with another block list.
*
- * At the context level, we use 'freelist' to track blocks ordered by number
- * of free chunks, starting with blocks having a single allocated chunk, and
- * with completely full blocks on the tail.
+ * Within each block, we maintain a list of chunks which are free to be used
+ * by new allocations. This is done in the form of a linked list where we
+ * store a pointer to the MemoryChunk that's next in the free list within the
+ * previous chunk's memory. The first of such chunks is pointed to from
+ * the block's "freehead" pointer.
*
- * This also allows various optimizations - for example when searching for
- * free chunk, the allocator reuses space from the fullest blocks first, in
- * the hope that some of the less full blocks will get completely empty (and
- * returned back to the OS).
- *
- * For each block, we maintain pointer to the first free chunk - this is quite
- * cheap and allows us to skip all the preceding used chunks, eliminating
- * a significant number of lookups in many common usage patterns. In the worst
- * case this performs as if the pointer was not maintained.
- *
- * We cache the freelist index for the blocks with the fewest free chunks
- * (minFreeChunks), so that we don't have to search the freelist on every
- * SlabAlloc() call, which is quite expensive.
+ * When we allocate a new block, technically all chunks are free, however, to
+ * avoid having to write out the entire block to set the linked list for the
+ * free chunks for every chunk in the block, we instead store a pointer to
+ * the next "unused" chunk on the block and keep track of how many of these
+ * unused chunks there are. In a newly allocated block, instead of
+ * populating the block's freelist, we simply consume the first unused chunk.
+ * It's only when these chunks are later freed that they go onto the block's
+ * freelist. When a block has both unused chunks and chunks on the freelist,
+ * we give priority to using the freelist chunks as the memory of these
+ * chunks is more likely to be in the CPUs caches as they've previously been
+ * used, whereas the unused ones have not yet been used.
*
*-------------------------------------------------------------------------
*/
@@ -60,6 +68,26 @@
#define Slab_BLOCKHDRSZ MAXALIGN(sizeof(SlabBlock))
+#ifdef MEMORY_CONTEXT_CHECKING
+/*
+ * Size of the memory required to store the SlabContext.
+ * MEMORY_CONTEXT_CHECKING builds need some extra member for the freechunks
+ * array.
+ */
+#define Slab_CONTEXT_HDRSZ(cpb) (sizeof(SlabContext) + ((cpb) * sizeof(bool)))
+#else
+#define Slab_CONTEXT_HDRSZ(cpb) sizeof(SlabContext)
+#endif
+
+/*
+ * The number of partitions to divide the used blocks list into based their
+ * number of free chunks. There must be at least 2.
+ */
+#define SLAB_BLOCKLIST_COUNT 3
+
+/* The maximum number of completely empty blocks to keep around to recycle */
+#define SLAB_MAXIMUM_EMPTY_BLOCKS 10
+
/*
* SlabContext is a specialized implementation of MemoryContext.
*/
@@ -67,56 +95,94 @@ typedef struct SlabContext
{
MemoryContextData header; /* Standard memory-context fields */
/* Allocation parameters for this context: */
- Size chunkSize; /* chunk size */
- Size fullChunkSize; /* chunk size including header and alignment */
- Size blockSize; /* block size */
- Size headerSize; /* allocated size of context header */
- int chunksPerBlock; /* number of chunks per block */
- int minFreeChunks; /* min number of free chunks in any block */
- int nblocks; /* number of blocks allocated */
+ Size chunkSize; /* the requested (non-aligned) chunk size */
+ Size fullChunkSize; /* chunk size with chunk header and alignment */
+ Size blockSize; /* the size to make each block of chunks */
+ int32 chunksPerBlock; /* number of chunks per block */
+ int32 curBlocklistIndex; /* index into the blocklist[] element
+ * containing the fullest, blocks */
#ifdef MEMORY_CONTEXT_CHECKING
bool *freechunks; /* bitmap of free chunks in a block */
#endif
- /* blocks with free space, grouped by number of free chunks: */
- dlist_head freelist[FLEXIBLE_ARRAY_MEMBER];
+
+ int32 blocklist_shift; /* number of bits to shift the nfree count
+ * by to get the index into blocklist[] */
+ dclist_head emptyblocks; /* Up to SLAB_MAXIMUM_EMPTY_BLOCKS blocks
+ * which are empty and ready to be reused */
+
+ /*
+ * Blocks with free space, grouped by the number of free chunks they
+ * contain. Completely full blocks are stored in the 0th element.
+ * Completely empty blocks are freed or stored in emptyblocks.
+ */
+ dlist_head blocklist[SLAB_BLOCKLIST_COUNT];
} SlabContext;
/*
* SlabBlock
- * Structure of a single block in SLAB allocator.
+ * Structure of a single slab block.
*
- * node: doubly-linked list of blocks in global freelist
- * nfree: number of free chunks in this block
- * firstFreeChunk: index of the first free chunk
+ * slab: pointer back to the owning MemoryContext
+ * nfree: number of chunks on the block which are unallocated
+ * nunused: number of chunks on the block unallocated and not on the block's
+ * freelist.
+ * freehead: linked-list header storing a pointer to the first free chunk on
+ * the block. Subsequent pointers are stored in the chunk's memory. NULL
+ * indicates the end of the list.
+ * unused: pointer to the next chunk which has yet to be used.
+ * node: doubly-linked list node for the context's blocklist
*/
typedef struct SlabBlock
{
- dlist_node node; /* doubly-linked list */
- int nfree; /* number of free chunks */
- int firstFreeChunk; /* index of the first free chunk in the block */
SlabContext *slab; /* owning context */
+ int32 nfree; /* number of chunks on freelist + unused */
+ int32 nunused; /* number of unused chunks */
+ MemoryChunk *freehead; /* pointer to the first free chunk */
+ MemoryChunk *unused; /* pointer to the next unused chunk */
+ dlist_node node; /* doubly-linked list for blocklist[] */
} SlabBlock;
#define Slab_CHUNKHDRSZ sizeof(MemoryChunk)
-#define SlabPointerGetChunk(ptr) \
- ((MemoryChunk *)(((char *)(ptr)) - sizeof(MemoryChunk)))
#define SlabChunkGetPointer(chk) \
- ((void *)(((char *)(chk)) + sizeof(MemoryChunk)))
-#define SlabBlockGetChunk(slab, block, idx) \
+ ((void *) (((char *) (chk)) + sizeof(MemoryChunk)))
+
+/*
+ * SlabBlockGetChunk
+ * Obtain a pointer to the Nth chunk in the block
+ */
+#define SlabBlockGetChunk(slab, block, N) \
((MemoryChunk *) ((char *) (block) + Slab_BLOCKHDRSZ \
- + (idx * slab->fullChunkSize)))
-#define SlabBlockStart(block) \
- ((char *) block + Slab_BLOCKHDRSZ)
+ + ((N) * (slab)->fullChunkSize)))
+
+#ifdef MEMORY_CONTEXT_CHECKING
+
+/*
+ * SlabChunkIndex
+ * Get the 0-based index of how many chunks into the block the given
+ * chunk is.
+*/
#define SlabChunkIndex(slab, block, chunk) \
- (((char *) chunk - SlabBlockStart(block)) / slab->fullChunkSize)
+ (((char *) (chunk) - (char *) SlabBlockGetChunk(slab, block, 0)) / \
+ (slab)->fullChunkSize)
+
+/*
+ * SlabChunkMod
+ * A MemoryChunk should always be at an address which is a multiple of
+ * fullChunkSize starting from the 0th chunk position. This will return
+ * non-zero if it's not.
+ */
+#define SlabChunkMod(slab, block, chunk) \
+ (((char *) (chunk) - (char *) SlabBlockGetChunk(slab, block, 0)) % \
+ (slab)->fullChunkSize)
+
+#endif
/*
* SlabIsValid
* True iff set is valid slab allocation set.
*/
-#define SlabIsValid(set) \
- (PointerIsValid(set) && IsA(set, SlabContext))
+#define SlabIsValid(set) (PointerIsValid(set) && IsA(set, SlabContext))
/*
* SlabBlockIsValid
@@ -125,6 +191,59 @@ typedef struct SlabBlock
#define SlabBlockIsValid(block) \
(PointerIsValid(block) && SlabIsValid((block)->slab))
+/*
+ * SlabBlocklistIndex
+ * Determine the blocklist index that a block should be in for the given
+ * number of free chunks.
+ */
+static inline int32
+SlabBlocklistIndex(SlabContext *slab, int nfree)
+{
+ int32 index;
+ int32 blocklist_shift = slab->blocklist_shift;
+
+ Assert(nfree <= slab->chunksPerBlock);
+
+ /*
+ * Determine the blocklist index based on the number of free chunks. We
+ * must enure that 0 free chunks is dedicated to index 0. Everything else
+ * must be >= 1 and < SLAB_BLOCKLIST_COUNT.
+ */
+ index = (nfree + (1 << blocklist_shift) - 1) >> blocklist_shift;
+
+ if (nfree == 0)
+ Assert(index == 0);
+ else
+ Assert(index >= 1 && index < SLAB_BLOCKLIST_COUNT);
+
+ return index;
+}
+
+/*
+ * SlabFindNextBlockListIndex
+ * Search blocklist for blocks which have free chunks and return the
+ * index of the blocklist found containing at least 1 block with free
+ * chunks. If no block can be found we return 0.
+ *
+ * Note: We give priority to full blocks so that these are filled before more
+ * empty blocks. This is done to increase the chances that mostly-empty
+ * blocks will eventually become completely empty so they can be freed.
+ */
+static int32
+SlabFindNextBlockListIndex(SlabContext *slab)
+{
+ /* start at 1. blocklist[0] is for full blocks. */
+ for (int i = 1; i < SLAB_BLOCKLIST_COUNT; i++)
+ {
+ if (dlist_is_empty(&slab->blocklist[i]))
+ continue;
+
+ return i;
+ }
+
+ /* no blocks with free space */
+ return 0;
+}
/*
* SlabContextCreate
@@ -145,8 +264,6 @@ SlabContextCreate(MemoryContext parent,
{
int chunksPerBlock;
Size fullChunkSize;
- Size freelistSize;
- Size headerSize;
SlabContext *slab;
int i;
@@ -155,11 +272,14 @@ SlabContextCreate(MemoryContext parent,
"sizeof(MemoryChunk) is not maxaligned");
Assert(MAXALIGN(chunkSize) <= MEMORYCHUNK_MAX_VALUE);
- /* Make sure the linked list node fits inside a freed chunk */
- if (chunkSize < sizeof(int))
- chunkSize = sizeof(int);
+ /*
+ * Ensure there's enough space to store the pointer to the next free chunk
+ * in the memory of the (otherwise) unused allocation.
+ */
+ if (chunkSize < sizeof(MemoryChunk *))
+ chunkSize = sizeof(MemoryChunk *);
- /* chunk, including SLAB header (both addresses nicely aligned) */
+ /* length of the maxaligned chunk including the chunk header */
#ifdef MEMORY_CONTEXT_CHECKING
/* ensure there's always space for the sentinel byte */
fullChunkSize = Slab_CHUNKHDRSZ + MAXALIGN(chunkSize + 1);
@@ -167,36 +287,17 @@ SlabContextCreate(MemoryContext parent,
fullChunkSize = Slab_CHUNKHDRSZ + MAXALIGN(chunkSize);
#endif
- /* Make sure the block can store at least one chunk. */
- if (blockSize < fullChunkSize + Slab_BLOCKHDRSZ)
- elog(ERROR, "block size %zu for slab is too small for %zu chunks",
- blockSize, chunkSize);
-
- /* Compute maximum number of chunks per block */
+ /* compute the number of chunks that will fit on each block */
chunksPerBlock = (blockSize - Slab_BLOCKHDRSZ) / fullChunkSize;
- /* The freelist starts with 0, ends with chunksPerBlock. */
- freelistSize = sizeof(dlist_head) * (chunksPerBlock + 1);
-
- /*
- * Allocate the context header. Unlike aset.c, we never try to combine
- * this with the first regular block; not worth the extra complication.
- */
+ /* Make sure the block can store at least one chunk. */
+ if (chunksPerBlock == 0)
+ elog(ERROR, "block size %zu for slab is too small for %zu-byte chunks",
+ blockSize, chunkSize);
- /* Size of the memory context header */
- headerSize = offsetof(SlabContext, freelist) + freelistSize;
-#ifdef MEMORY_CONTEXT_CHECKING
- /*
- * With memory checking, we need to allocate extra space for the bitmap of
- * free chunks. The bitmap is an array of bools, so we don't need to worry
- * about alignment.
- */
- headerSize += chunksPerBlock * sizeof(bool);
-#endif
-
- slab = (SlabContext *) malloc(headerSize);
+ slab = (SlabContext *) malloc(Slab_CONTEXT_HDRSZ(chunksPerBlock));
if (slab == NULL)
{
MemoryContextStats(TopMemoryContext);
@@ -216,19 +317,33 @@ SlabContextCreate(MemoryContext parent,
slab->chunkSize = chunkSize;
slab->fullChunkSize = fullChunkSize;
slab->blockSize = blockSize;
- slab->headerSize = headerSize;
slab->chunksPerBlock = chunksPerBlock;
- slab->minFreeChunks = 0;
- slab->nblocks = 0;
+ slab->curBlocklistIndex = 0;
- /* initialize the freelist slots */
- for (i = 0; i < (slab->chunksPerBlock + 1); i++)
- dlist_init(&slab->freelist[i]);
+ /*
+ * Compute a shift that guarantees that shifting chunksPerBlock with it is
+ * < SLAB_BLOCKLIST_COUNT - 1. The reason that we subtract 1 from
+ * SLAB_BLOCKLIST_COUNT in this calculation is that we reserve the 0th
+ * blocklist element for blocks which have no free chunks.
+ *
+ * We calculate the number of bits to shift by rather than a divisor to
+ * divide by as performing division each time we need to find the
+ * blocklist index would be much slower.
+ */
+ slab->blocklist_shift = 0;
+ while ((slab->chunksPerBlock >> slab->blocklist_shift) >= (SLAB_BLOCKLIST_COUNT - 1))
+ slab->blocklist_shift++;
+
+ /* initialize the list to store empty blocks to be recycled */
+ dclist_init(&slab->emptyblocks);
+
+ /* initialize the blocklist slots */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
+ dlist_init(&slab->blocklist[i]);
#ifdef MEMORY_CONTEXT_CHECKING
- /* set the freechunks pointer right after the freelists array */
- slab->freechunks
- = (bool *) slab + offsetof(SlabContext, freelist) + freelistSize;
+ /* set the freechunks pointer right after the end of the context */
+ slab->freechunks = (bool *) ((char *) slab + sizeof(SlabContext));
#endif
/* Finally, do the type-independent part of context creation */
@@ -252,6 +367,7 @@ void
SlabReset(MemoryContext context)
{
SlabContext *slab = (SlabContext *) context;
+ dlist_mutable_iter miter;
int i;
Assert(SlabIsValid(slab));
@@ -261,12 +377,24 @@ SlabReset(MemoryContext context)
SlabCheck(context);
#endif
- /* walk over freelists and free the blocks */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /* release any retained empty blocks */
+ dclist_foreach_modify(miter, &slab->emptyblocks)
{
- dlist_mutable_iter miter;
+ SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
+
+ dclist_delete_from(&slab->emptyblocks, miter.cur);
- dlist_foreach_modify(miter, &slab->freelist[i])
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ context->mem_allocated -= slab->blockSize;
+ }
+
+ /* walk over blocklist and free the blocks */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
+ {
+ dlist_foreach_modify(miter, &slab->blocklist[i])
{
SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
@@ -276,14 +404,12 @@ SlabReset(MemoryContext context)
wipe_mem(block, slab->blockSize);
#endif
free(block);
- slab->nblocks--;
context->mem_allocated -= slab->blockSize;
}
}
- slab->minFreeChunks = 0;
+ slab->curBlocklistIndex = 0;
- Assert(slab->nblocks == 0);
Assert(context->mem_allocated == 0);
}
@@ -311,12 +437,12 @@ SlabAlloc(MemoryContext context, Size size)
SlabContext *slab = (SlabContext *) context;
SlabBlock *block;
MemoryChunk *chunk;
- int idx;
Assert(SlabIsValid(slab));
- Assert((slab->minFreeChunks >= 0) &&
- (slab->minFreeChunks < slab->chunksPerBlock));
+ /* sanity check that this is pointing to a valid blocklist */
+ Assert(slab->curBlocklistIndex >= 0);
+ Assert(slab->curBlocklistIndex <= SlabBlocklistIndex(slab, slab->chunksPerBlock));
/* make sure we only allow correct request size */
if (size != slab->chunkSize)
@@ -324,114 +450,153 @@ SlabAlloc(MemoryContext context, Size size)
size, slab->chunkSize);
/*
- * If there are no free chunks in any existing block, create a new block
- * and put it to the last freelist bucket.
- *
- * slab->minFreeChunks == 0 means there are no blocks with free chunks,
- * thanks to how minFreeChunks is updated at the end of SlabAlloc().
+ * Handle the case when there are no partially filled blocks available.
+ * SlabFree() will have updated the curBlocklistIndex setting it to zero
+ * to indicate that it has freed the final block. Also later in
+ * SlabAlloc() we will set the curBlocklistIndex to zero if we end up
+ * filling the final block.
*/
- if (slab->minFreeChunks == 0)
+ if (unlikely(slab->curBlocklistIndex == 0))
{
- block = (SlabBlock *) malloc(slab->blockSize);
+ dlist_head *blocklist;
+ int blocklist_idx;
- if (block == NULL)
- return NULL;
+ /* to save allocating a new one, first check the empty blocks list */
+ if (dclist_count(&slab->emptyblocks) > 0)
+ {
+ dlist_node *node = dclist_pop_head_node(&slab->emptyblocks);
- block->nfree = slab->chunksPerBlock;
- block->firstFreeChunk = 0;
- block->slab = slab;
+ block = dlist_container(SlabBlock, node, node);
- /*
- * Put all the chunks on a freelist. Walk the chunks and point each
- * one to the next one.
- */
- for (idx = 0; idx < slab->chunksPerBlock; idx++)
- {
- chunk = SlabBlockGetChunk(slab, block, idx);
- *(int32 *) MemoryChunkGetPointer(chunk) = (idx + 1);
+ /*
+ * SlabFree() should have left this block in a valid state with
+ * all chunks free. Ensure that's the case.
+ */
+ Assert(block->nfree == slab->chunksPerBlock);
+
+ if (block->freehead != NULL)
+ {
+ chunk = block->freehead;
+
+ /*
+ * Pop the chunk from the linked list of free chunks. The
+ * pointer to the next free chunk is stored in the chunk
+ * itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(MemoryChunk *));
+ block->freehead = *(MemoryChunk **) SlabChunkGetPointer(chunk);
+
+ Assert(block->freehead == NULL ||
+ ((char *) block->freehead >= (char *) block &&
+ (char *) block->freehead < (char *) block + slab->blockSize &&
+ SlabChunkMod(slab, block, block->freehead) == 0));
+ }
+ else
+ {
+ Assert(block->nunused > 0);
+
+ chunk = block->unused;
+ block->unused = (MemoryChunk *) (((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
}
+ else
+ {
+ block = (SlabBlock *) malloc(slab->blockSize);
- /*
- * And add it to the last freelist with all chunks empty.
- *
- * We know there are no blocks in the freelist, otherwise we wouldn't
- * need a new block.
- */
- Assert(dlist_is_empty(&slab->freelist[slab->chunksPerBlock]));
+ if (unlikely(block == NULL))
+ return NULL;
- dlist_push_head(&slab->freelist[slab->chunksPerBlock], &block->node);
+ block->slab = slab;
+ context->mem_allocated += slab->blockSize;
- slab->minFreeChunks = slab->chunksPerBlock;
- slab->nblocks += 1;
- context->mem_allocated += slab->blockSize;
- }
+ /* use the first chunk in the new block */
+ chunk = SlabBlockGetChunk(slab, block, 0);
- /* grab the block from the freelist (even the new block is there) */
- block = dlist_head_element(SlabBlock, node,
- &slab->freelist[slab->minFreeChunks]);
+ block->nfree = slab->chunksPerBlock;
+ block->unused = SlabBlockGetChunk(slab, block, 1);
+ block->freehead = NULL;
+ block->nunused = slab->chunksPerBlock - 1;
+ }
- /* make sure we actually got a valid block, with matching nfree */
- Assert(block != NULL);
- Assert(slab->minFreeChunks == block->nfree);
- Assert(block->nfree > 0);
+ /* find the blocklist element for storing blocks with 1 used chunk */
+ blocklist_idx = SlabBlocklistIndex(slab, slab->chunksPerBlock - 1);
+ blocklist = &slab->blocklist[blocklist_idx];
- /* we know index of the first free chunk in the block */
- idx = block->firstFreeChunk;
+ /* this better be empty. We just added a block thinking it was */
+ Assert(dlist_is_empty(blocklist));
- /* make sure the chunk index is valid, and that it's marked as empty */
- Assert((idx >= 0) && (idx < slab->chunksPerBlock));
+ dlist_push_head(blocklist, &block->node);
- /* compute the chunk location block start (after the block header) */
- chunk = SlabBlockGetChunk(slab, block, idx);
+ slab->curBlocklistIndex = blocklist_idx;
+ }
+ else
+ {
+ dlist_head *blocklist = &slab->blocklist[slab->curBlocklistIndex];
+ int new_blocklist_idx;
- /*
- * Update the block nfree count, and also the minFreeChunks as we've
- * decreased nfree for a block with the minimum number of free chunks
- * (because that's how we chose the block).
- */
- block->nfree--;
- slab->minFreeChunks = block->nfree;
+ Assert(!dlist_is_empty(blocklist));
- /*
- * Remove the chunk from the freelist head. The index of the next free
- * chunk is stored in the chunk itself.
- */
- VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- block->firstFreeChunk = *(int32 *) MemoryChunkGetPointer(chunk);
+ /* grab the block from the blocklist */
+ block = dlist_head_element(SlabBlock, node, blocklist);
- Assert(block->firstFreeChunk >= 0);
- Assert(block->firstFreeChunk <= slab->chunksPerBlock);
+ /* make sure we actually got a valid block, with matching nfree */
+ Assert(block != NULL);
+ Assert(slab->curBlocklistIndex == SlabBlocklistIndex(slab, block->nfree));
+ Assert(block->nfree > 0);
- Assert((block->nfree != 0 &&
- block->firstFreeChunk < slab->chunksPerBlock) ||
- (block->nfree == 0 &&
- block->firstFreeChunk == slab->chunksPerBlock));
+ /*
+ * Use previously free'd chunks first. Unused chunks are less likely
+ * to be cached by the CPU.
+ */
+ if (block->freehead != NULL)
+ {
+ chunk = block->freehead;
- /* move the whole block to the right place in the freelist */
- dlist_delete(&block->node);
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
+ /*
+ * Pop the chunk from the linked list of free chunks. The pointer
+ * to the next free chunk is stored in the chunk itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(MemoryChunk *));
+ block->freehead = *(MemoryChunk **) SlabChunkGetPointer(chunk);
- /*
- * And finally update minFreeChunks, i.e. the index to the block with the
- * lowest number of free chunks. We only need to do that when the block
- * got full (otherwise we know the current block is the right one). We'll
- * simply walk the freelist until we find a non-empty entry.
- */
- if (slab->minFreeChunks == 0)
- {
- for (idx = 1; idx <= slab->chunksPerBlock; idx++)
+ Assert(block->freehead == NULL ||
+ ((char *) block->freehead >= (char *) block &&
+ (char *) block->freehead < (char *) block + slab->blockSize &&
+ SlabChunkMod(slab, block, block->freehead) == 0));
+ }
+ else
{
- if (dlist_is_empty(&slab->freelist[idx]))
- continue;
+ Assert(block->nunused > 0);
- /* found a non-empty freelist */
- slab->minFreeChunks = idx;
- break;
+ chunk = block->unused;
+ block->unused = (MemoryChunk *) (((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
+
+ /* get the new blocklist index based on the new free chunk count */
+ new_blocklist_idx = SlabBlocklistIndex(slab, block->nfree - 1);
+
+ /*
+ * Handle the case where the blocklist index changes. This also deals
+ * with blocks becoming full as only full blocks go at index 0.
+ */
+ if (unlikely(slab->curBlocklistIndex != new_blocklist_idx))
+ {
+ dlist_delete_from(blocklist, &block->node);
+ dlist_push_head(&slab->blocklist[new_blocklist_idx], &block->node);
+
+ if (dlist_is_empty(blocklist))
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
}
}
- if (slab->minFreeChunks == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
+ /* check that the chunk pointer is actually somewhere on the block */
+ Assert(chunk >= SlabBlockGetChunk(slab, block, 0));
+ Assert(chunk <= SlabBlockGetChunk(slab, block, slab->chunksPerBlock));
+
+ /* update the block nfree count */
+ block->nfree--;
/* Prepare to initialize the chunk header. */
VALGRIND_MAKE_MEM_UNDEFINED(chunk, Slab_CHUNKHDRSZ);
@@ -453,8 +618,6 @@ SlabAlloc(MemoryContext context, Size size)
randomize_mem((char *) MemoryChunkGetPointer(chunk), size);
#endif
- Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
-
return MemoryChunkGetPointer(chunk);
}
@@ -468,7 +631,8 @@ SlabFree(void *pointer)
MemoryChunk *chunk = PointerGetMemoryChunk(pointer);
SlabBlock *block = MemoryChunkGetBlock(chunk);
SlabContext *slab;
- int idx;
+ int curBlocklistIdx;
+ int newBlocklistIdx;
/*
* For speed reasons we just Assert that the referenced block is good.
@@ -486,63 +650,82 @@ SlabFree(void *pointer)
slab->header.name, chunk);
#endif
- /* compute index of the chunk with respect to block start */
- idx = SlabChunkIndex(slab, block, chunk);
+ /* push this chunk onto the head of the free list */
+ *(MemoryChunk **) pointer = block->freehead;
+ block->freehead = chunk;
- /* add chunk to freelist, and update block nfree count */
- *(int32 *) pointer = block->firstFreeChunk;
- block->firstFreeChunk = idx;
block->nfree++;
Assert(block->nfree > 0);
Assert(block->nfree <= slab->chunksPerBlock);
#ifdef CLOBBER_FREED_MEMORY
- /* XXX don't wipe the int32 index, used for block-level freelist */
- wipe_mem((char *) pointer + sizeof(int32),
- slab->chunkSize - sizeof(int32));
+ /* don't wipe the free list MemoryChunk pointer stored in the chunk */
+ wipe_mem((char *) pointer + sizeof(MemoryChunk *),
+ slab->chunkSize - sizeof(MemoryChunk *));
#endif
- /* remove the block from a freelist */
- dlist_delete(&block->node);
+ curBlocklistIdx = SlabBlocklistIndex(slab, block->nfree - 1);
+ newBlocklistIdx = SlabBlocklistIndex(slab, block->nfree);
/*
- * See if we need to update the minFreeChunks field for the slab - we only
- * need to do that if there the block had that number of free chunks
- * before we freed one. In that case, we check if there still are blocks
- * in the original freelist and we either keep the current value (if there
- * still are blocks) or increment it by one (the new block is still the
- * one with minimum free chunks).
- *
- * The one exception is when the block will get completely free - in that
- * case we will free it, se we can't use it for minFreeChunks. It however
- * means there are no more blocks with free chunks.
+ * Check if the block needs to be moved to another element on the
+ * blocklist based on it now having 1 more free chunk.
*/
- if (slab->minFreeChunks == (block->nfree - 1))
+ if (unlikely(curBlocklistIdx != newBlocklistIdx))
{
- /* Have we removed the last chunk from the freelist? */
- if (dlist_is_empty(&slab->freelist[slab->minFreeChunks]))
+ /* do the move */
+ dlist_delete_from(&slab->blocklist[curBlocklistIdx], &block->node);
+ dlist_push_head(&slab->blocklist[newBlocklistIdx], &block->node);
+
+ /*
+ * It's possible that we've no blocks in the blocklist at the
+ * curBlocklistIndex position. When this happens we must find the
+ * next blocklist index which contains blocks. We can be certain
+ * we'll find a block as at least one must exist for the chunk we're
+ * currently freeing.
+ */
+ if (slab->curBlocklistIndex == curBlocklistIdx &&
+ dlist_is_empty(&slab->blocklist[curBlocklistIdx]))
{
- /* but if we made the block entirely free, we'll free it */
- if (block->nfree == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
- else
- slab->minFreeChunks++;
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
+ Assert(slab->curBlocklistIndex > 0);
}
}
- /* If the block is now completely empty, free it. */
- if (block->nfree == slab->chunksPerBlock)
+ /* Handle when a block becomes completely empty */
+ if (unlikely(block->nfree == slab->chunksPerBlock))
{
- free(block);
- slab->nblocks--;
- slab->header.mem_allocated -= slab->blockSize;
- }
- else
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
+ /* remove the block */
+ dlist_delete_from(&slab->blocklist[newBlocklistIdx], &block->node);
- Assert(slab->nblocks >= 0);
- Assert(slab->nblocks * slab->blockSize == slab->header.mem_allocated);
+ /*
+ * To avoid thrashing malloc/free, we keep a list of empty blocks that
+ * we can reuse again insteading having to malloc a new one.
+ */
+ if (dclist_count(&slab->emptyblocks) < SLAB_MAXIMUM_EMPTY_BLOCKS)
+ dclist_push_head(&slab->emptyblocks, &block->node);
+ else
+ {
+ /*
+ * When we have enough empty blocks stored already, we actually
+ * free the block.
+ */
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ slab->header.mem_allocated -= slab->blockSize;
+ }
+
+ /*
+ * Check if we need to reset the blocklist index. This is required
+ * when the blocklist we're using has become completely empty.
+ */
+ if (slab->curBlocklistIndex == newBlocklistIdx &&
+ dlist_is_empty(&slab->blocklist[newBlocklistIdx]))
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
+ }
}
/*
@@ -622,11 +805,9 @@ SlabGetChunkSpace(void *pointer)
bool
SlabIsEmpty(MemoryContext context)
{
- SlabContext *slab = (SlabContext *) context;
-
- Assert(SlabIsValid(slab));
+ Assert(SlabIsValid((SlabContext *) context));
- return (slab->nblocks == 0);
+ return (context->mem_allocated > 0);
}
/*
@@ -654,13 +835,16 @@ SlabStats(MemoryContext context,
Assert(SlabIsValid(slab));
/* Include context header in totalspace */
- totalspace = slab->headerSize;
+ totalspace = Slab_CONTEXT_HDRSZ(slab->chunksPerBlock);
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /* Add the space consumed by blocks in the emptyblocks list */
+ totalspace += dclist_count(&slab->emptyblocks) * slab->blockSize;
+
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
{
dlist_iter iter;
- dlist_foreach(iter, &slab->freelist[i])
+ dlist_foreach(iter, &slab->blocklist[i])
{
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
@@ -676,9 +860,9 @@ SlabStats(MemoryContext context,
char stats_string[200];
snprintf(stats_string, sizeof(stats_string),
- "%zu total in %zu blocks; %zu free (%zu chunks); %zu used",
- totalspace, nblocks, freespace, freechunks,
- totalspace - freespace);
+ "%zu total in %zu blocks; %u empty blocks; %zu free (%zu chunks); %zu used",
+ totalspace, nblocks, dclist_count(&slab->emptyblocks),
+ freespace, freechunks, totalspace - freespace);
printfunc(context, passthru, stats_string, print_to_stderr);
}
@@ -707,31 +891,45 @@ SlabCheck(MemoryContext context)
{
SlabContext *slab = (SlabContext *) context;
int i;
+ int nblocks = 0;
const char *name = slab->header.name;
+ dlist_iter iter;
Assert(SlabIsValid(slab));
Assert(slab->chunksPerBlock > 0);
- /* walk all the freelists */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /*
+ * Have a look at the empty blocks. These should have all their chunks
+ * marked as free. Ensure that's the case.
+ */
+ dclist_foreach(iter, &slab->emptyblocks)
+ {
+ SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+
+ if (block->nfree != slab->chunksPerBlock)
+ elog(WARNING, "problem in slab %s: empty block %p should have %d free chunks but only has %d",
+ name, block, slab->chunksPerBlock, block->nfree);
+ }
+
+ /* walk all the block lists */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
{
int j,
nfree;
- dlist_iter iter;
- /* walk all blocks on this freelist */
- dlist_foreach(iter, &slab->freelist[i])
+ /* walk all blocks on this blocklist */
+ dlist_foreach(iter, &slab->blocklist[i])
{
- int idx;
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+ MemoryChunk *cur_chunk;
/*
* Make sure the number of free chunks (in the block header)
- * matches position in the freelist.
+ * matches position in the blocklist.
*/
- if (block->nfree != i)
- elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match freelist %d",
- name, block->nfree, block, i);
+ if (SlabBlocklistIndex(slab, block->nfree) != i)
+ elog(WARNING, "problem in slab %s: block %p is on blocklist %d but should be on blocklist %d",
+ name, block, i, SlabBlocklistIndex(slab, block->nfree));
/* make sure the slab pointer correctly points to this context */
if (block->slab != slab)
@@ -740,28 +938,55 @@ SlabCheck(MemoryContext context)
/* reset the bitmap of free chunks for this block */
memset(slab->freechunks, 0, (slab->chunksPerBlock * sizeof(bool)));
- idx = block->firstFreeChunk;
+ nfree = 0;
/*
- * Now walk through the chunks, count the free ones and also
- * perform some additional checks for the used ones. As the chunk
- * freelist is stored within the chunks themselves, we have to
- * walk through the chunks and construct our own bitmap.
+ * Walk through the free list chunks and count the number of free
+ * chunks.
*/
+ cur_chunk = block->freehead;
+ while (cur_chunk != NULL)
+ {
+ int chunkidx = SlabChunkIndex(slab, block, cur_chunk);
+ intptr_t chunkmod = SlabChunkMod(slab, block, cur_chunk);
+
+ /*
+ * Make sure the free list link points to something on the
+ * block
+ */
+ if ((char *) cur_chunk <= (char *) block ||
+ (char *) cur_chunk >= (char *) block + slab->blockSize ||
+ chunkmod != 0)
+ elog(WARNING, "problem in slab %s: bogus free list link %p in block %p",
+ name, cur_chunk, block);
- nfree = 0;
- while (idx < slab->chunksPerBlock)
+ /* count the chunk as free, add it to the bitmap */
+ nfree++;
+ slab->freechunks[chunkidx] = true;
+
+ /* read pointer of the next free chunk */
+ VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(cur_chunk), sizeof(MemoryChunk *));
+ cur_chunk = *(MemoryChunk **) SlabChunkGetPointer(cur_chunk);
+ }
+
+ /*
+ * count the remaining free chunks that have yet to make it onto
+ * the block's free list.
+ */
+ cur_chunk = block->unused;
+ for (j = 0; j < block->nunused; j++)
{
- MemoryChunk *chunk;
+ int chunkidx = SlabChunkIndex(slab, block, cur_chunk);
+
+ /* make sure we've not stepped off the end of the block */
+ Assert((char *) cur_chunk < (char *) block + slab->blockSize);
/* count the chunk as free, add it to the bitmap */
nfree++;
- slab->freechunks[idx] = true;
+ slab->freechunks[chunkidx] = true;
- /* read index of the next free chunk */
- chunk = SlabBlockGetChunk(slab, block, idx);
- VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- idx = *(int32 *) MemoryChunkGetPointer(chunk);
+ /* move forward 1 chunk */
+ cur_chunk = (MemoryChunk *) (((char *) cur_chunk) + slab->fullChunkSize);
}
for (j = 0; j < slab->chunksPerBlock; j++)
@@ -795,10 +1020,15 @@ SlabCheck(MemoryContext context)
if (nfree != block->nfree)
elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match bitmap %d",
name, block->nfree, block, nfree);
+
+ nblocks++;
}
}
- Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
+ /* the stored empty blocks are tracked in mem_allocated too */
+ nblocks += dclist_count(&slab->emptyblocks);
+
+ Assert(nblocks * slab->blockSize == context->mem_allocated);
}
#endif /* MEMORY_CONTEXT_CHECKING */
--
2.35.1.windows.2
diff --git a/src/backend/utils/adt/mcxtfuncs.c b/src/backend/utils/adt/mcxtfuncs.c
index 4add219553..8d04e8aad3 100644
--- a/src/backend/utils/adt/mcxtfuncs.c
+++ b/src/backend/utils/adt/mcxtfuncs.c
@@ -15,6 +15,8 @@
#include "postgres.h"
+#include <time.h>
+
#include "funcapi.h"
#include "miscadmin.h"
#include "mb/pg_wchar.h"
@@ -193,3 +195,214 @@ pg_log_backend_memory_contexts(PG_FUNCTION_ARGS)
PG_RETURN_BOOL(true);
}
+
+typedef struct AllocateTestNext
+{
+ struct AllocateTestNext *next; /* ptr to the next allocation */
+} AllocateTestNext;
+
+/* #define ALLOCATE_TEST_DEBUG */
+/*
+ * pg_allocate_memory_test
+ * Used to test the performance of a memory context types
+ */
+Datum
+pg_allocate_memory_test(PG_FUNCTION_ARGS)
+{
+ int32 chunk_size = PG_GETARG_INT32(0);
+ int64 keep_memory = PG_GETARG_INT64(1);
+ int64 total_alloc = PG_GETARG_INT64(2);
+ text *context_type_text = PG_GETARG_TEXT_PP(3);
+ char *context_type;
+ int64 curr_memory_use = 0;
+ int64 remaining_alloc_bytes = total_alloc;
+ MemoryContext context;
+ MemoryContext oldContext;
+ AllocateTestNext *next_free_ptr = NULL;
+ AllocateTestNext *last_alloc = NULL;
+ clock_t start, end;
+
+ if (chunk_size < sizeof(AllocateTestNext))
+ elog(ERROR, "chunk_size (%d) must be at least %ld bytes", chunk_size,
+ sizeof(AllocateTestNext));
+ if (keep_memory > total_alloc)
+ elog(ERROR, "keep_memory (" INT64_FORMAT ") must be less than total_alloc (" INT64_FORMAT ")",
+ keep_memory, total_alloc);
+
+ context_type = text_to_cstring(context_type_text);
+
+ start = clock();
+
+ if (strcmp(context_type, "generation") == 0)
+ context = GenerationContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "aset") == 0)
+ context = AllocSetContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "slab") == 0)
+ context = SlabContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_MAXSIZE,
+ chunk_size);
+ else
+ elog(ERROR, "context_type must be \"generation\", \"aset\" or \"slab\"");
+
+ oldContext = MemoryContextSwitchTo(context);
+
+ while (remaining_alloc_bytes > 0)
+ {
+ AllocateTestNext *curr_alloc;
+
+ CHECK_FOR_INTERRUPTS();
+
+ /* Allocate the memory and update the counters */
+ curr_alloc = (AllocateTestNext *) palloc(chunk_size);
+ remaining_alloc_bytes -= chunk_size;
+ curr_memory_use += chunk_size;
+
+#ifdef ALLOCATE_TEST_DEBUG
+ elog(NOTICE, "alloc %p (curr_memory_use " INT64_FORMAT " bytes, remaining_alloc_bytes " INT64_FORMAT ")", curr_alloc, curr_memory_use, remaining_alloc_bytes);
+#endif
+
+ /*
+ * Point the last allocate to this one so that we can free allocations
+ * starting with the oldest first.
+ */
+ curr_alloc->next = NULL;
+ if (last_alloc != NULL)
+ last_alloc->next = curr_alloc;
+
+ if (next_free_ptr == NULL)
+ {
+ /*
+ * Remember the first chunk to free. We will follow the ->next
+ * pointers to find the next chunk to free when freeing memory
+ */
+ next_free_ptr = curr_alloc;
+ }
+
+ /*
+ * If the currently allocated memory has reached or exceeded the amount
+ * of memory we want to keep allocated at once then we'd better free
+ * some. Since all allocations are the same size we only need to free
+ * one allocation per loop.
+ */
+ if (curr_memory_use >= keep_memory)
+ {
+ AllocateTestNext *next = next_free_ptr->next;
+
+ /* free the memory and update the current memory usage */
+ pfree(next_free_ptr);
+ curr_memory_use -= chunk_size;
+
+#ifdef ALLOCATE_TEST_DEBUG
+ elog(NOTICE, "free %p (curr_memory_use " INT64_FORMAT " bytes, remaining_alloc_bytes " INT64_FORMAT ")", next_free_ptr, curr_memory_use, remaining_alloc_bytes);
+#endif
+ /* get the next chunk to free */
+ next_free_ptr = next;
+ }
+
+ if (curr_memory_use > 0)
+ last_alloc = curr_alloc;
+ else
+ last_alloc = NULL;
+ }
+
+ /* cleanup loop -- pfree remaining memory */
+ while (next_free_ptr != NULL)
+ {
+ AllocateTestNext *next = next_free_ptr->next;
+
+ /* free the memory and update the current memory usage */
+ pfree(next_free_ptr);
+ curr_memory_use -= chunk_size;
+
+#ifdef ALLOCATE_TEST_DEBUG
+ elog(NOTICE, "free %p (curr_memory_use " INT64_FORMAT " bytes, remaining_alloc_bytes " INT64_FORMAT ")", next_free_ptr, curr_memory_use, remaining_alloc_bytes);
+#endif
+
+ next_free_ptr = next;
+ }
+
+ MemoryContextSwitchTo(oldContext);
+
+ end = clock();
+
+ PG_RETURN_FLOAT8((double) (end - start) / CLOCKS_PER_SEC);
+}
+
+Datum
+pg_allocate_memory_test_reset(PG_FUNCTION_ARGS)
+{
+ int32 chunk_size = PG_GETARG_INT32(0);
+ int64 keep_memory = PG_GETARG_INT64(1);
+ int64 total_alloc = PG_GETARG_INT64(2);
+ text *context_type_text = PG_GETARG_TEXT_PP(3);
+ char *context_type;
+ int64 curr_memory_use = 0;
+ int64 remaining_alloc_bytes = total_alloc;
+ MemoryContext context;
+ MemoryContext oldContext;
+ clock_t start, end;
+
+ if (chunk_size < 1)
+ elog(ERROR, "size of chunk must be above 0");
+ if (keep_memory > total_alloc)
+ elog(ERROR, "keep_memory (" INT64_FORMAT ") must be less than total_alloc (" INT64_FORMAT ")",
+ keep_memory, total_alloc);
+
+ context_type = text_to_cstring(context_type_text);
+
+ start = clock();
+
+ if (strcmp(context_type, "generation") == 0)
+ context = GenerationContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "aset") == 0)
+ context = AllocSetContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "slab") == 0)
+ context = SlabContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_MAXSIZE,
+ chunk_size);
+ else
+ elog(ERROR, "context_type must be \"generation\", \"aset\" or \"slab\"");
+
+ oldContext = MemoryContextSwitchTo(context);
+
+ while (remaining_alloc_bytes > 0)
+ {
+ CHECK_FOR_INTERRUPTS();
+
+ /* Allocate the memory and update the counters */
+ (void) palloc(chunk_size);
+ remaining_alloc_bytes -= chunk_size;
+ curr_memory_use += chunk_size;
+
+#ifdef ALLOCATE_TEST_DEBUG
+ elog(NOTICE, "alloc %p (curr_memory_use " INT64_FORMAT " bytes, remaining_alloc_bytes " INT64_FORMAT ")", curr_alloc, curr_memory_use, remaining_alloc_bytes);
+#endif
+
+ /*
+ * If the currently allocated memory has reached or exceeded the amount
+ * of memory we want to keep allocated at once then reset the context.
+ */
+ if (curr_memory_use >= keep_memory)
+ {
+ curr_memory_use = 0;
+ MemoryContextReset(context);
+ }
+ }
+
+ MemoryContextSwitchTo(oldContext);
+ MemoryContextDelete(context);
+
+ end = clock();
+
+ PG_RETURN_FLOAT8((double) (end - start) / CLOCKS_PER_SEC);
+}
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 719599649a..6bb53529a0 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8143,6 +8143,18 @@
prorettype => 'bool', proargtypes => 'int4',
prosrc => 'pg_log_backend_memory_contexts' },
+# just for testing memory context allocation speed
+{ oid => '9319', descr => 'for testing performance of allocation and freeing',
+ proname => 'pg_allocate_memory_test', provolatile => 'v',
+ prorettype => 'float8', proargtypes => 'int4 int8 int8 text',
+ prosrc => 'pg_allocate_memory_test' },
+
+# just for testing memory context allocation speed
+{ oid => '9320', descr => 'for testing performance of allocation and resetting',
+ proname => 'pg_allocate_memory_test_reset', provolatile => 'v',
+ prorettype => 'float8', proargtypes => 'int4 int8 int8 text',
+ prosrc => 'pg_allocate_memory_test_reset' },
+
# non-persistent series generator
{ oid => '1066', descr => 'non-persistent series generator',
proname => 'generate_series', prorows => '1000',
Attachments:
[text/plain] v4-0001-Improve-the-performance-of-the-slab-memory-alloca.patch (37.3K, ../../CAApHDvr0pKjrGVKFg2VqSUjbB1QxsKobu0PhD11qkSO_caL3dg@mail.gmail.com/2-v4-0001-Improve-the-performance-of-the-slab-memory-alloca.patch)
download | inline diff:
From b970576654bbce5c57690b8ab49d0b4376d1b5d3 Mon Sep 17 00:00:00 2001
From: David Rowley <dgrowley@gmail.com>
Date: Wed, 12 Oct 2022 09:30:24 +1300
Subject: [PATCH v4] Improve the performance of the slab memory allocator
Author: Andres Freund, David Rowley
Discussion: https://postgr.es/m/20210717194333.mr5io3zup3kxahfm@alap3.anarazel.de
---
src/backend/utils/mmgr/slab.c | 750 ++++++++++++++++++++++------------
1 file changed, 490 insertions(+), 260 deletions(-)
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index 6df0839b6a..2b95a6c061 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -3,8 +3,8 @@
* slab.c
* SLAB allocator definitions.
*
- * SLAB is a MemoryContext implementation designed for cases where large
- * numbers of equally-sized objects are allocated (and freed).
+ * SLAB is a MemoryContext implementation designed for cases where a large
+ * numbers of equally-sized objects can be allocated and freed efficiently.
*
*
* Portions Copyright (c) 2017-2022, PostgreSQL Global Development Group
@@ -19,33 +19,41 @@
* into chunks of exactly the right size (plus alignment), not wasting any
* memory.
*
- * The information about free chunks is maintained both at the block level and
- * global (context) level. This is possible as the chunk size (and thus also
- * the number of chunks per block) is fixed.
+ * Slab can also help reduce memory fragmentation where chunks remain stored
+ * on blocks with no other or only a few other allocated chunks. If this
+ * happens then the entire block cannot be freed until the last remaining
+ * chunk is freed. Slab helps work around this problem by prioritizing
+ * storing newly allocated chunks starting with the fullest blocks first.
+ * This makes it more likely that blocks with only a small number of
+ * remaining chunks eventually get freed when the final remaining chunk is
+ * freed.
*
- * On each block, free chunks are tracked in a simple linked list. Contents
- * of free chunks is replaced with an index of the next free chunk, forming
- * a very simple linked list. Each block also contains a counter of free
- * chunks. Combined with the local block-level freelist, it makes it trivial
- * to eventually free the whole block.
+ * We are easily able to find the fullest block to store a new chunk on as we
+ * maintain a list of all blocks and partition that list by the block's
+ * number of free chunks. This does mean having to possibly move the block
+ * onto another list whenever we allocate or free a chunk from it. However,
+ * this problem is significantly reduced by only having a small fixed number
+ * of lists for blocks, each of which allows storage of blocks with ranges of
+ * free chunks. The block only needs to be moved to another list when the
+ * number of free chunks crosses the range boundary with another block list.
*
- * At the context level, we use 'freelist' to track blocks ordered by number
- * of free chunks, starting with blocks having a single allocated chunk, and
- * with completely full blocks on the tail.
+ * Within each block, we maintain a list of chunks which are free to be used
+ * by new allocations. This is done in the form of a linked list where we
+ * store a pointer to the MemoryChunk that's next in the free list within the
+ * previous chunk's memory. The first of such chunks is pointed to from
+ * the block's "freehead" pointer.
*
- * This also allows various optimizations - for example when searching for
- * free chunk, the allocator reuses space from the fullest blocks first, in
- * the hope that some of the less full blocks will get completely empty (and
- * returned back to the OS).
- *
- * For each block, we maintain pointer to the first free chunk - this is quite
- * cheap and allows us to skip all the preceding used chunks, eliminating
- * a significant number of lookups in many common usage patterns. In the worst
- * case this performs as if the pointer was not maintained.
- *
- * We cache the freelist index for the blocks with the fewest free chunks
- * (minFreeChunks), so that we don't have to search the freelist on every
- * SlabAlloc() call, which is quite expensive.
+ * When we allocate a new block, technically all chunks are free, however, to
+ * avoid having to write out the entire block to set the linked list for the
+ * free chunks for every chunk in the block, we instead store a pointer to
+ * the next "unused" chunk on the block and keep track of how many of these
+ * unused chunks there are. In a newly allocated block, instead of
+ * populating the block's freelist, we simply consume the first unused chunk.
+ * It's only when these chunks are later freed that they go onto the block's
+ * freelist. When a block has both unused chunks and chunks on the freelist,
+ * we give priority to using the freelist chunks as the memory of these
+ * chunks is more likely to be in the CPUs caches as they've previously been
+ * used, whereas the unused ones have not yet been used.
*
*-------------------------------------------------------------------------
*/
@@ -60,6 +68,26 @@
#define Slab_BLOCKHDRSZ MAXALIGN(sizeof(SlabBlock))
+#ifdef MEMORY_CONTEXT_CHECKING
+/*
+ * Size of the memory required to store the SlabContext.
+ * MEMORY_CONTEXT_CHECKING builds need some extra member for the freechunks
+ * array.
+ */
+#define Slab_CONTEXT_HDRSZ(cpb) (sizeof(SlabContext) + ((cpb) * sizeof(bool)))
+#else
+#define Slab_CONTEXT_HDRSZ(cpb) sizeof(SlabContext)
+#endif
+
+/*
+ * The number of partitions to divide the used blocks list into based their
+ * number of free chunks. There must be at least 2.
+ */
+#define SLAB_BLOCKLIST_COUNT 3
+
+/* The maximum number of completely empty blocks to keep around to recycle */
+#define SLAB_MAXIMUM_EMPTY_BLOCKS 10
+
/*
* SlabContext is a specialized implementation of MemoryContext.
*/
@@ -67,56 +95,94 @@ typedef struct SlabContext
{
MemoryContextData header; /* Standard memory-context fields */
/* Allocation parameters for this context: */
- Size chunkSize; /* chunk size */
- Size fullChunkSize; /* chunk size including header and alignment */
- Size blockSize; /* block size */
- Size headerSize; /* allocated size of context header */
- int chunksPerBlock; /* number of chunks per block */
- int minFreeChunks; /* min number of free chunks in any block */
- int nblocks; /* number of blocks allocated */
+ Size chunkSize; /* the requested (non-aligned) chunk size */
+ Size fullChunkSize; /* chunk size with chunk header and alignment */
+ Size blockSize; /* the size to make each block of chunks */
+ int32 chunksPerBlock; /* number of chunks per block */
+ int32 curBlocklistIndex; /* index into the blocklist[] element
+ * containing the fullest, blocks */
#ifdef MEMORY_CONTEXT_CHECKING
bool *freechunks; /* bitmap of free chunks in a block */
#endif
- /* blocks with free space, grouped by number of free chunks: */
- dlist_head freelist[FLEXIBLE_ARRAY_MEMBER];
+
+ int32 blocklist_shift; /* number of bits to shift the nfree count
+ * by to get the index into blocklist[] */
+ dclist_head emptyblocks; /* Up to SLAB_MAXIMUM_EMPTY_BLOCKS blocks
+ * which are empty and ready to be reused */
+
+ /*
+ * Blocks with free space, grouped by the number of free chunks they
+ * contain. Completely full blocks are stored in the 0th element.
+ * Completely empty blocks are freed or stored in emptyblocks.
+ */
+ dlist_head blocklist[SLAB_BLOCKLIST_COUNT];
} SlabContext;
/*
* SlabBlock
- * Structure of a single block in SLAB allocator.
+ * Structure of a single slab block.
*
- * node: doubly-linked list of blocks in global freelist
- * nfree: number of free chunks in this block
- * firstFreeChunk: index of the first free chunk
+ * slab: pointer back to the owning MemoryContext
+ * nfree: number of chunks on the block which are unallocated
+ * nunused: number of chunks on the block unallocated and not on the block's
+ * freelist.
+ * freehead: linked-list header storing a pointer to the first free chunk on
+ * the block. Subsequent pointers are stored in the chunk's memory. NULL
+ * indicates the end of the list.
+ * unused: pointer to the next chunk which has yet to be used.
+ * node: doubly-linked list node for the context's blocklist
*/
typedef struct SlabBlock
{
- dlist_node node; /* doubly-linked list */
- int nfree; /* number of free chunks */
- int firstFreeChunk; /* index of the first free chunk in the block */
SlabContext *slab; /* owning context */
+ int32 nfree; /* number of chunks on freelist + unused */
+ int32 nunused; /* number of unused chunks */
+ MemoryChunk *freehead; /* pointer to the first free chunk */
+ MemoryChunk *unused; /* pointer to the next unused chunk */
+ dlist_node node; /* doubly-linked list for blocklist[] */
} SlabBlock;
#define Slab_CHUNKHDRSZ sizeof(MemoryChunk)
-#define SlabPointerGetChunk(ptr) \
- ((MemoryChunk *)(((char *)(ptr)) - sizeof(MemoryChunk)))
#define SlabChunkGetPointer(chk) \
- ((void *)(((char *)(chk)) + sizeof(MemoryChunk)))
-#define SlabBlockGetChunk(slab, block, idx) \
+ ((void *) (((char *) (chk)) + sizeof(MemoryChunk)))
+
+/*
+ * SlabBlockGetChunk
+ * Obtain a pointer to the Nth chunk in the block
+ */
+#define SlabBlockGetChunk(slab, block, N) \
((MemoryChunk *) ((char *) (block) + Slab_BLOCKHDRSZ \
- + (idx * slab->fullChunkSize)))
-#define SlabBlockStart(block) \
- ((char *) block + Slab_BLOCKHDRSZ)
+ + ((N) * (slab)->fullChunkSize)))
+
+#ifdef MEMORY_CONTEXT_CHECKING
+
+/*
+ * SlabChunkIndex
+ * Get the 0-based index of how many chunks into the block the given
+ * chunk is.
+*/
#define SlabChunkIndex(slab, block, chunk) \
- (((char *) chunk - SlabBlockStart(block)) / slab->fullChunkSize)
+ (((char *) (chunk) - (char *) SlabBlockGetChunk(slab, block, 0)) / \
+ (slab)->fullChunkSize)
+
+/*
+ * SlabChunkMod
+ * A MemoryChunk should always be at an address which is a multiple of
+ * fullChunkSize starting from the 0th chunk position. This will return
+ * non-zero if it's not.
+ */
+#define SlabChunkMod(slab, block, chunk) \
+ (((char *) (chunk) - (char *) SlabBlockGetChunk(slab, block, 0)) % \
+ (slab)->fullChunkSize)
+
+#endif
/*
* SlabIsValid
* True iff set is valid slab allocation set.
*/
-#define SlabIsValid(set) \
- (PointerIsValid(set) && IsA(set, SlabContext))
+#define SlabIsValid(set) (PointerIsValid(set) && IsA(set, SlabContext))
/*
* SlabBlockIsValid
@@ -125,6 +191,59 @@ typedef struct SlabBlock
#define SlabBlockIsValid(block) \
(PointerIsValid(block) && SlabIsValid((block)->slab))
+/*
+ * SlabBlocklistIndex
+ * Determine the blocklist index that a block should be in for the given
+ * number of free chunks.
+ */
+static inline int32
+SlabBlocklistIndex(SlabContext *slab, int nfree)
+{
+ int32 index;
+ int32 blocklist_shift = slab->blocklist_shift;
+
+ Assert(nfree <= slab->chunksPerBlock);
+
+ /*
+ * Determine the blocklist index based on the number of free chunks. We
+ * must enure that 0 free chunks is dedicated to index 0. Everything else
+ * must be >= 1 and < SLAB_BLOCKLIST_COUNT.
+ */
+ index = (nfree + (1 << blocklist_shift) - 1) >> blocklist_shift;
+
+ if (nfree == 0)
+ Assert(index == 0);
+ else
+ Assert(index >= 1 && index < SLAB_BLOCKLIST_COUNT);
+
+ return index;
+}
+
+/*
+ * SlabFindNextBlockListIndex
+ * Search blocklist for blocks which have free chunks and return the
+ * index of the blocklist found containing at least 1 block with free
+ * chunks. If no block can be found we return 0.
+ *
+ * Note: We give priority to full blocks so that these are filled before more
+ * empty blocks. This is done to increase the chances that mostly-empty
+ * blocks will eventually become completely empty so they can be freed.
+ */
+static int32
+SlabFindNextBlockListIndex(SlabContext *slab)
+{
+ /* start at 1. blocklist[0] is for full blocks. */
+ for (int i = 1; i < SLAB_BLOCKLIST_COUNT; i++)
+ {
+ if (dlist_is_empty(&slab->blocklist[i]))
+ continue;
+
+ return i;
+ }
+
+ /* no blocks with free space */
+ return 0;
+}
/*
* SlabContextCreate
@@ -145,8 +264,6 @@ SlabContextCreate(MemoryContext parent,
{
int chunksPerBlock;
Size fullChunkSize;
- Size freelistSize;
- Size headerSize;
SlabContext *slab;
int i;
@@ -155,11 +272,14 @@ SlabContextCreate(MemoryContext parent,
"sizeof(MemoryChunk) is not maxaligned");
Assert(MAXALIGN(chunkSize) <= MEMORYCHUNK_MAX_VALUE);
- /* Make sure the linked list node fits inside a freed chunk */
- if (chunkSize < sizeof(int))
- chunkSize = sizeof(int);
+ /*
+ * Ensure there's enough space to store the pointer to the next free chunk
+ * in the memory of the (otherwise) unused allocation.
+ */
+ if (chunkSize < sizeof(MemoryChunk *))
+ chunkSize = sizeof(MemoryChunk *);
- /* chunk, including SLAB header (both addresses nicely aligned) */
+ /* length of the maxaligned chunk including the chunk header */
#ifdef MEMORY_CONTEXT_CHECKING
/* ensure there's always space for the sentinel byte */
fullChunkSize = Slab_CHUNKHDRSZ + MAXALIGN(chunkSize + 1);
@@ -167,36 +287,17 @@ SlabContextCreate(MemoryContext parent,
fullChunkSize = Slab_CHUNKHDRSZ + MAXALIGN(chunkSize);
#endif
- /* Make sure the block can store at least one chunk. */
- if (blockSize < fullChunkSize + Slab_BLOCKHDRSZ)
- elog(ERROR, "block size %zu for slab is too small for %zu chunks",
- blockSize, chunkSize);
-
- /* Compute maximum number of chunks per block */
+ /* compute the number of chunks that will fit on each block */
chunksPerBlock = (blockSize - Slab_BLOCKHDRSZ) / fullChunkSize;
- /* The freelist starts with 0, ends with chunksPerBlock. */
- freelistSize = sizeof(dlist_head) * (chunksPerBlock + 1);
-
- /*
- * Allocate the context header. Unlike aset.c, we never try to combine
- * this with the first regular block; not worth the extra complication.
- */
+ /* Make sure the block can store at least one chunk. */
+ if (chunksPerBlock == 0)
+ elog(ERROR, "block size %zu for slab is too small for %zu-byte chunks",
+ blockSize, chunkSize);
- /* Size of the memory context header */
- headerSize = offsetof(SlabContext, freelist) + freelistSize;
-#ifdef MEMORY_CONTEXT_CHECKING
- /*
- * With memory checking, we need to allocate extra space for the bitmap of
- * free chunks. The bitmap is an array of bools, so we don't need to worry
- * about alignment.
- */
- headerSize += chunksPerBlock * sizeof(bool);
-#endif
-
- slab = (SlabContext *) malloc(headerSize);
+ slab = (SlabContext *) malloc(Slab_CONTEXT_HDRSZ(chunksPerBlock));
if (slab == NULL)
{
MemoryContextStats(TopMemoryContext);
@@ -216,19 +317,33 @@ SlabContextCreate(MemoryContext parent,
slab->chunkSize = chunkSize;
slab->fullChunkSize = fullChunkSize;
slab->blockSize = blockSize;
- slab->headerSize = headerSize;
slab->chunksPerBlock = chunksPerBlock;
- slab->minFreeChunks = 0;
- slab->nblocks = 0;
+ slab->curBlocklistIndex = 0;
- /* initialize the freelist slots */
- for (i = 0; i < (slab->chunksPerBlock + 1); i++)
- dlist_init(&slab->freelist[i]);
+ /*
+ * Compute a shift that guarantees that shifting chunksPerBlock with it is
+ * < SLAB_BLOCKLIST_COUNT - 1. The reason that we subtract 1 from
+ * SLAB_BLOCKLIST_COUNT in this calculation is that we reserve the 0th
+ * blocklist element for blocks which have no free chunks.
+ *
+ * We calculate the number of bits to shift by rather than a divisor to
+ * divide by as performing division each time we need to find the
+ * blocklist index would be much slower.
+ */
+ slab->blocklist_shift = 0;
+ while ((slab->chunksPerBlock >> slab->blocklist_shift) >= (SLAB_BLOCKLIST_COUNT - 1))
+ slab->blocklist_shift++;
+
+ /* initialize the list to store empty blocks to be recycled */
+ dclist_init(&slab->emptyblocks);
+
+ /* initialize the blocklist slots */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
+ dlist_init(&slab->blocklist[i]);
#ifdef MEMORY_CONTEXT_CHECKING
- /* set the freechunks pointer right after the freelists array */
- slab->freechunks
- = (bool *) slab + offsetof(SlabContext, freelist) + freelistSize;
+ /* set the freechunks pointer right after the end of the context */
+ slab->freechunks = (bool *) ((char *) slab + sizeof(SlabContext));
#endif
/* Finally, do the type-independent part of context creation */
@@ -252,6 +367,7 @@ void
SlabReset(MemoryContext context)
{
SlabContext *slab = (SlabContext *) context;
+ dlist_mutable_iter miter;
int i;
Assert(SlabIsValid(slab));
@@ -261,12 +377,24 @@ SlabReset(MemoryContext context)
SlabCheck(context);
#endif
- /* walk over freelists and free the blocks */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /* release any retained empty blocks */
+ dclist_foreach_modify(miter, &slab->emptyblocks)
{
- dlist_mutable_iter miter;
+ SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
+
+ dclist_delete_from(&slab->emptyblocks, miter.cur);
- dlist_foreach_modify(miter, &slab->freelist[i])
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ context->mem_allocated -= slab->blockSize;
+ }
+
+ /* walk over blocklist and free the blocks */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
+ {
+ dlist_foreach_modify(miter, &slab->blocklist[i])
{
SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
@@ -276,14 +404,12 @@ SlabReset(MemoryContext context)
wipe_mem(block, slab->blockSize);
#endif
free(block);
- slab->nblocks--;
context->mem_allocated -= slab->blockSize;
}
}
- slab->minFreeChunks = 0;
+ slab->curBlocklistIndex = 0;
- Assert(slab->nblocks == 0);
Assert(context->mem_allocated == 0);
}
@@ -311,12 +437,12 @@ SlabAlloc(MemoryContext context, Size size)
SlabContext *slab = (SlabContext *) context;
SlabBlock *block;
MemoryChunk *chunk;
- int idx;
Assert(SlabIsValid(slab));
- Assert((slab->minFreeChunks >= 0) &&
- (slab->minFreeChunks < slab->chunksPerBlock));
+ /* sanity check that this is pointing to a valid blocklist */
+ Assert(slab->curBlocklistIndex >= 0);
+ Assert(slab->curBlocklistIndex <= SlabBlocklistIndex(slab, slab->chunksPerBlock));
/* make sure we only allow correct request size */
if (size != slab->chunkSize)
@@ -324,114 +450,153 @@ SlabAlloc(MemoryContext context, Size size)
size, slab->chunkSize);
/*
- * If there are no free chunks in any existing block, create a new block
- * and put it to the last freelist bucket.
- *
- * slab->minFreeChunks == 0 means there are no blocks with free chunks,
- * thanks to how minFreeChunks is updated at the end of SlabAlloc().
+ * Handle the case when there are no partially filled blocks available.
+ * SlabFree() will have updated the curBlocklistIndex setting it to zero
+ * to indicate that it has freed the final block. Also later in
+ * SlabAlloc() we will set the curBlocklistIndex to zero if we end up
+ * filling the final block.
*/
- if (slab->minFreeChunks == 0)
+ if (unlikely(slab->curBlocklistIndex == 0))
{
- block = (SlabBlock *) malloc(slab->blockSize);
+ dlist_head *blocklist;
+ int blocklist_idx;
- if (block == NULL)
- return NULL;
+ /* to save allocating a new one, first check the empty blocks list */
+ if (dclist_count(&slab->emptyblocks) > 0)
+ {
+ dlist_node *node = dclist_pop_head_node(&slab->emptyblocks);
- block->nfree = slab->chunksPerBlock;
- block->firstFreeChunk = 0;
- block->slab = slab;
+ block = dlist_container(SlabBlock, node, node);
- /*
- * Put all the chunks on a freelist. Walk the chunks and point each
- * one to the next one.
- */
- for (idx = 0; idx < slab->chunksPerBlock; idx++)
- {
- chunk = SlabBlockGetChunk(slab, block, idx);
- *(int32 *) MemoryChunkGetPointer(chunk) = (idx + 1);
+ /*
+ * SlabFree() should have left this block in a valid state with
+ * all chunks free. Ensure that's the case.
+ */
+ Assert(block->nfree == slab->chunksPerBlock);
+
+ if (block->freehead != NULL)
+ {
+ chunk = block->freehead;
+
+ /*
+ * Pop the chunk from the linked list of free chunks. The
+ * pointer to the next free chunk is stored in the chunk
+ * itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(MemoryChunk *));
+ block->freehead = *(MemoryChunk **) SlabChunkGetPointer(chunk);
+
+ Assert(block->freehead == NULL ||
+ ((char *) block->freehead >= (char *) block &&
+ (char *) block->freehead < (char *) block + slab->blockSize &&
+ SlabChunkMod(slab, block, block->freehead) == 0));
+ }
+ else
+ {
+ Assert(block->nunused > 0);
+
+ chunk = block->unused;
+ block->unused = (MemoryChunk *) (((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
}
+ else
+ {
+ block = (SlabBlock *) malloc(slab->blockSize);
- /*
- * And add it to the last freelist with all chunks empty.
- *
- * We know there are no blocks in the freelist, otherwise we wouldn't
- * need a new block.
- */
- Assert(dlist_is_empty(&slab->freelist[slab->chunksPerBlock]));
+ if (unlikely(block == NULL))
+ return NULL;
- dlist_push_head(&slab->freelist[slab->chunksPerBlock], &block->node);
+ block->slab = slab;
+ context->mem_allocated += slab->blockSize;
- slab->minFreeChunks = slab->chunksPerBlock;
- slab->nblocks += 1;
- context->mem_allocated += slab->blockSize;
- }
+ /* use the first chunk in the new block */
+ chunk = SlabBlockGetChunk(slab, block, 0);
- /* grab the block from the freelist (even the new block is there) */
- block = dlist_head_element(SlabBlock, node,
- &slab->freelist[slab->minFreeChunks]);
+ block->nfree = slab->chunksPerBlock;
+ block->unused = SlabBlockGetChunk(slab, block, 1);
+ block->freehead = NULL;
+ block->nunused = slab->chunksPerBlock - 1;
+ }
- /* make sure we actually got a valid block, with matching nfree */
- Assert(block != NULL);
- Assert(slab->minFreeChunks == block->nfree);
- Assert(block->nfree > 0);
+ /* find the blocklist element for storing blocks with 1 used chunk */
+ blocklist_idx = SlabBlocklistIndex(slab, slab->chunksPerBlock - 1);
+ blocklist = &slab->blocklist[blocklist_idx];
- /* we know index of the first free chunk in the block */
- idx = block->firstFreeChunk;
+ /* this better be empty. We just added a block thinking it was */
+ Assert(dlist_is_empty(blocklist));
- /* make sure the chunk index is valid, and that it's marked as empty */
- Assert((idx >= 0) && (idx < slab->chunksPerBlock));
+ dlist_push_head(blocklist, &block->node);
- /* compute the chunk location block start (after the block header) */
- chunk = SlabBlockGetChunk(slab, block, idx);
+ slab->curBlocklistIndex = blocklist_idx;
+ }
+ else
+ {
+ dlist_head *blocklist = &slab->blocklist[slab->curBlocklistIndex];
+ int new_blocklist_idx;
- /*
- * Update the block nfree count, and also the minFreeChunks as we've
- * decreased nfree for a block with the minimum number of free chunks
- * (because that's how we chose the block).
- */
- block->nfree--;
- slab->minFreeChunks = block->nfree;
+ Assert(!dlist_is_empty(blocklist));
- /*
- * Remove the chunk from the freelist head. The index of the next free
- * chunk is stored in the chunk itself.
- */
- VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- block->firstFreeChunk = *(int32 *) MemoryChunkGetPointer(chunk);
+ /* grab the block from the blocklist */
+ block = dlist_head_element(SlabBlock, node, blocklist);
- Assert(block->firstFreeChunk >= 0);
- Assert(block->firstFreeChunk <= slab->chunksPerBlock);
+ /* make sure we actually got a valid block, with matching nfree */
+ Assert(block != NULL);
+ Assert(slab->curBlocklistIndex == SlabBlocklistIndex(slab, block->nfree));
+ Assert(block->nfree > 0);
- Assert((block->nfree != 0 &&
- block->firstFreeChunk < slab->chunksPerBlock) ||
- (block->nfree == 0 &&
- block->firstFreeChunk == slab->chunksPerBlock));
+ /*
+ * Use previously free'd chunks first. Unused chunks are less likely
+ * to be cached by the CPU.
+ */
+ if (block->freehead != NULL)
+ {
+ chunk = block->freehead;
- /* move the whole block to the right place in the freelist */
- dlist_delete(&block->node);
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
+ /*
+ * Pop the chunk from the linked list of free chunks. The pointer
+ * to the next free chunk is stored in the chunk itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(MemoryChunk *));
+ block->freehead = *(MemoryChunk **) SlabChunkGetPointer(chunk);
- /*
- * And finally update minFreeChunks, i.e. the index to the block with the
- * lowest number of free chunks. We only need to do that when the block
- * got full (otherwise we know the current block is the right one). We'll
- * simply walk the freelist until we find a non-empty entry.
- */
- if (slab->minFreeChunks == 0)
- {
- for (idx = 1; idx <= slab->chunksPerBlock; idx++)
+ Assert(block->freehead == NULL ||
+ ((char *) block->freehead >= (char *) block &&
+ (char *) block->freehead < (char *) block + slab->blockSize &&
+ SlabChunkMod(slab, block, block->freehead) == 0));
+ }
+ else
{
- if (dlist_is_empty(&slab->freelist[idx]))
- continue;
+ Assert(block->nunused > 0);
- /* found a non-empty freelist */
- slab->minFreeChunks = idx;
- break;
+ chunk = block->unused;
+ block->unused = (MemoryChunk *) (((char *) block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
+
+ /* get the new blocklist index based on the new free chunk count */
+ new_blocklist_idx = SlabBlocklistIndex(slab, block->nfree - 1);
+
+ /*
+ * Handle the case where the blocklist index changes. This also deals
+ * with blocks becoming full as only full blocks go at index 0.
+ */
+ if (unlikely(slab->curBlocklistIndex != new_blocklist_idx))
+ {
+ dlist_delete_from(blocklist, &block->node);
+ dlist_push_head(&slab->blocklist[new_blocklist_idx], &block->node);
+
+ if (dlist_is_empty(blocklist))
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
}
}
- if (slab->minFreeChunks == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
+ /* check that the chunk pointer is actually somewhere on the block */
+ Assert(chunk >= SlabBlockGetChunk(slab, block, 0));
+ Assert(chunk <= SlabBlockGetChunk(slab, block, slab->chunksPerBlock));
+
+ /* update the block nfree count */
+ block->nfree--;
/* Prepare to initialize the chunk header. */
VALGRIND_MAKE_MEM_UNDEFINED(chunk, Slab_CHUNKHDRSZ);
@@ -453,8 +618,6 @@ SlabAlloc(MemoryContext context, Size size)
randomize_mem((char *) MemoryChunkGetPointer(chunk), size);
#endif
- Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
-
return MemoryChunkGetPointer(chunk);
}
@@ -468,7 +631,8 @@ SlabFree(void *pointer)
MemoryChunk *chunk = PointerGetMemoryChunk(pointer);
SlabBlock *block = MemoryChunkGetBlock(chunk);
SlabContext *slab;
- int idx;
+ int curBlocklistIdx;
+ int newBlocklistIdx;
/*
* For speed reasons we just Assert that the referenced block is good.
@@ -486,63 +650,82 @@ SlabFree(void *pointer)
slab->header.name, chunk);
#endif
- /* compute index of the chunk with respect to block start */
- idx = SlabChunkIndex(slab, block, chunk);
+ /* push this chunk onto the head of the free list */
+ *(MemoryChunk **) pointer = block->freehead;
+ block->freehead = chunk;
- /* add chunk to freelist, and update block nfree count */
- *(int32 *) pointer = block->firstFreeChunk;
- block->firstFreeChunk = idx;
block->nfree++;
Assert(block->nfree > 0);
Assert(block->nfree <= slab->chunksPerBlock);
#ifdef CLOBBER_FREED_MEMORY
- /* XXX don't wipe the int32 index, used for block-level freelist */
- wipe_mem((char *) pointer + sizeof(int32),
- slab->chunkSize - sizeof(int32));
+ /* don't wipe the free list MemoryChunk pointer stored in the chunk */
+ wipe_mem((char *) pointer + sizeof(MemoryChunk *),
+ slab->chunkSize - sizeof(MemoryChunk *));
#endif
- /* remove the block from a freelist */
- dlist_delete(&block->node);
+ curBlocklistIdx = SlabBlocklistIndex(slab, block->nfree - 1);
+ newBlocklistIdx = SlabBlocklistIndex(slab, block->nfree);
/*
- * See if we need to update the minFreeChunks field for the slab - we only
- * need to do that if there the block had that number of free chunks
- * before we freed one. In that case, we check if there still are blocks
- * in the original freelist and we either keep the current value (if there
- * still are blocks) or increment it by one (the new block is still the
- * one with minimum free chunks).
- *
- * The one exception is when the block will get completely free - in that
- * case we will free it, se we can't use it for minFreeChunks. It however
- * means there are no more blocks with free chunks.
+ * Check if the block needs to be moved to another element on the
+ * blocklist based on it now having 1 more free chunk.
*/
- if (slab->minFreeChunks == (block->nfree - 1))
+ if (unlikely(curBlocklistIdx != newBlocklistIdx))
{
- /* Have we removed the last chunk from the freelist? */
- if (dlist_is_empty(&slab->freelist[slab->minFreeChunks]))
+ /* do the move */
+ dlist_delete_from(&slab->blocklist[curBlocklistIdx], &block->node);
+ dlist_push_head(&slab->blocklist[newBlocklistIdx], &block->node);
+
+ /*
+ * It's possible that we've no blocks in the blocklist at the
+ * curBlocklistIndex position. When this happens we must find the
+ * next blocklist index which contains blocks. We can be certain
+ * we'll find a block as at least one must exist for the chunk we're
+ * currently freeing.
+ */
+ if (slab->curBlocklistIndex == curBlocklistIdx &&
+ dlist_is_empty(&slab->blocklist[curBlocklistIdx]))
{
- /* but if we made the block entirely free, we'll free it */
- if (block->nfree == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
- else
- slab->minFreeChunks++;
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
+ Assert(slab->curBlocklistIndex > 0);
}
}
- /* If the block is now completely empty, free it. */
- if (block->nfree == slab->chunksPerBlock)
+ /* Handle when a block becomes completely empty */
+ if (unlikely(block->nfree == slab->chunksPerBlock))
{
- free(block);
- slab->nblocks--;
- slab->header.mem_allocated -= slab->blockSize;
- }
- else
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
+ /* remove the block */
+ dlist_delete_from(&slab->blocklist[newBlocklistIdx], &block->node);
- Assert(slab->nblocks >= 0);
- Assert(slab->nblocks * slab->blockSize == slab->header.mem_allocated);
+ /*
+ * To avoid thrashing malloc/free, we keep a list of empty blocks that
+ * we can reuse again insteading having to malloc a new one.
+ */
+ if (dclist_count(&slab->emptyblocks) < SLAB_MAXIMUM_EMPTY_BLOCKS)
+ dclist_push_head(&slab->emptyblocks, &block->node);
+ else
+ {
+ /*
+ * When we have enough empty blocks stored already, we actually
+ * free the block.
+ */
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ slab->header.mem_allocated -= slab->blockSize;
+ }
+
+ /*
+ * Check if we need to reset the blocklist index. This is required
+ * when the blocklist we're using has become completely empty.
+ */
+ if (slab->curBlocklistIndex == newBlocklistIdx &&
+ dlist_is_empty(&slab->blocklist[newBlocklistIdx]))
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
+ }
}
/*
@@ -622,11 +805,9 @@ SlabGetChunkSpace(void *pointer)
bool
SlabIsEmpty(MemoryContext context)
{
- SlabContext *slab = (SlabContext *) context;
-
- Assert(SlabIsValid(slab));
+ Assert(SlabIsValid((SlabContext *) context));
- return (slab->nblocks == 0);
+ return (context->mem_allocated > 0);
}
/*
@@ -654,13 +835,16 @@ SlabStats(MemoryContext context,
Assert(SlabIsValid(slab));
/* Include context header in totalspace */
- totalspace = slab->headerSize;
+ totalspace = Slab_CONTEXT_HDRSZ(slab->chunksPerBlock);
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /* Add the space consumed by blocks in the emptyblocks list */
+ totalspace += dclist_count(&slab->emptyblocks) * slab->blockSize;
+
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
{
dlist_iter iter;
- dlist_foreach(iter, &slab->freelist[i])
+ dlist_foreach(iter, &slab->blocklist[i])
{
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
@@ -676,9 +860,9 @@ SlabStats(MemoryContext context,
char stats_string[200];
snprintf(stats_string, sizeof(stats_string),
- "%zu total in %zu blocks; %zu free (%zu chunks); %zu used",
- totalspace, nblocks, freespace, freechunks,
- totalspace - freespace);
+ "%zu total in %zu blocks; %u empty blocks; %zu free (%zu chunks); %zu used",
+ totalspace, nblocks, dclist_count(&slab->emptyblocks),
+ freespace, freechunks, totalspace - freespace);
printfunc(context, passthru, stats_string, print_to_stderr);
}
@@ -707,31 +891,45 @@ SlabCheck(MemoryContext context)
{
SlabContext *slab = (SlabContext *) context;
int i;
+ int nblocks = 0;
const char *name = slab->header.name;
+ dlist_iter iter;
Assert(SlabIsValid(slab));
Assert(slab->chunksPerBlock > 0);
- /* walk all the freelists */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /*
+ * Have a look at the empty blocks. These should have all their chunks
+ * marked as free. Ensure that's the case.
+ */
+ dclist_foreach(iter, &slab->emptyblocks)
+ {
+ SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+
+ if (block->nfree != slab->chunksPerBlock)
+ elog(WARNING, "problem in slab %s: empty block %p should have %d free chunks but only has %d",
+ name, block, slab->chunksPerBlock, block->nfree);
+ }
+
+ /* walk all the block lists */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
{
int j,
nfree;
- dlist_iter iter;
- /* walk all blocks on this freelist */
- dlist_foreach(iter, &slab->freelist[i])
+ /* walk all blocks on this blocklist */
+ dlist_foreach(iter, &slab->blocklist[i])
{
- int idx;
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+ MemoryChunk *cur_chunk;
/*
* Make sure the number of free chunks (in the block header)
- * matches position in the freelist.
+ * matches position in the blocklist.
*/
- if (block->nfree != i)
- elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match freelist %d",
- name, block->nfree, block, i);
+ if (SlabBlocklistIndex(slab, block->nfree) != i)
+ elog(WARNING, "problem in slab %s: block %p is on blocklist %d but should be on blocklist %d",
+ name, block, i, SlabBlocklistIndex(slab, block->nfree));
/* make sure the slab pointer correctly points to this context */
if (block->slab != slab)
@@ -740,28 +938,55 @@ SlabCheck(MemoryContext context)
/* reset the bitmap of free chunks for this block */
memset(slab->freechunks, 0, (slab->chunksPerBlock * sizeof(bool)));
- idx = block->firstFreeChunk;
+ nfree = 0;
/*
- * Now walk through the chunks, count the free ones and also
- * perform some additional checks for the used ones. As the chunk
- * freelist is stored within the chunks themselves, we have to
- * walk through the chunks and construct our own bitmap.
+ * Walk through the free list chunks and count the number of free
+ * chunks.
*/
+ cur_chunk = block->freehead;
+ while (cur_chunk != NULL)
+ {
+ int chunkidx = SlabChunkIndex(slab, block, cur_chunk);
+ intptr_t chunkmod = SlabChunkMod(slab, block, cur_chunk);
+
+ /*
+ * Make sure the free list link points to something on the
+ * block
+ */
+ if ((char *) cur_chunk <= (char *) block ||
+ (char *) cur_chunk >= (char *) block + slab->blockSize ||
+ chunkmod != 0)
+ elog(WARNING, "problem in slab %s: bogus free list link %p in block %p",
+ name, cur_chunk, block);
- nfree = 0;
- while (idx < slab->chunksPerBlock)
+ /* count the chunk as free, add it to the bitmap */
+ nfree++;
+ slab->freechunks[chunkidx] = true;
+
+ /* read pointer of the next free chunk */
+ VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(cur_chunk), sizeof(MemoryChunk *));
+ cur_chunk = *(MemoryChunk **) SlabChunkGetPointer(cur_chunk);
+ }
+
+ /*
+ * count the remaining free chunks that have yet to make it onto
+ * the block's free list.
+ */
+ cur_chunk = block->unused;
+ for (j = 0; j < block->nunused; j++)
{
- MemoryChunk *chunk;
+ int chunkidx = SlabChunkIndex(slab, block, cur_chunk);
+
+ /* make sure we've not stepped off the end of the block */
+ Assert((char *) cur_chunk < (char *) block + slab->blockSize);
/* count the chunk as free, add it to the bitmap */
nfree++;
- slab->freechunks[idx] = true;
+ slab->freechunks[chunkidx] = true;
- /* read index of the next free chunk */
- chunk = SlabBlockGetChunk(slab, block, idx);
- VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- idx = *(int32 *) MemoryChunkGetPointer(chunk);
+ /* move forward 1 chunk */
+ cur_chunk = (MemoryChunk *) (((char *) cur_chunk) + slab->fullChunkSize);
}
for (j = 0; j < slab->chunksPerBlock; j++)
@@ -795,10 +1020,15 @@ SlabCheck(MemoryContext context)
if (nfree != block->nfree)
elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match bitmap %d",
name, block->nfree, block, nfree);
+
+ nblocks++;
}
}
- Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
+ /* the stored empty blocks are tracked in mem_allocated too */
+ nblocks += dclist_count(&slab->emptyblocks);
+
+ Assert(nblocks * slab->blockSize == context->mem_allocated);
}
#endif /* MEMORY_CONTEXT_CHECKING */
--
2.35.1.windows.2
[text/plain] alloc_bench.patch (7.8K, ../../CAApHDvr0pKjrGVKFg2VqSUjbB1QxsKobu0PhD11qkSO_caL3dg@mail.gmail.com/3-alloc_bench.patch)
download | inline diff:
diff --git a/src/backend/utils/adt/mcxtfuncs.c b/src/backend/utils/adt/mcxtfuncs.c
index 4add219553..8d04e8aad3 100644
--- a/src/backend/utils/adt/mcxtfuncs.c
+++ b/src/backend/utils/adt/mcxtfuncs.c
@@ -15,6 +15,8 @@
#include "postgres.h"
+#include <time.h>
+
#include "funcapi.h"
#include "miscadmin.h"
#include "mb/pg_wchar.h"
@@ -193,3 +195,214 @@ pg_log_backend_memory_contexts(PG_FUNCTION_ARGS)
PG_RETURN_BOOL(true);
}
+
+typedef struct AllocateTestNext
+{
+ struct AllocateTestNext *next; /* ptr to the next allocation */
+} AllocateTestNext;
+
+/* #define ALLOCATE_TEST_DEBUG */
+/*
+ * pg_allocate_memory_test
+ * Used to test the performance of a memory context types
+ */
+Datum
+pg_allocate_memory_test(PG_FUNCTION_ARGS)
+{
+ int32 chunk_size = PG_GETARG_INT32(0);
+ int64 keep_memory = PG_GETARG_INT64(1);
+ int64 total_alloc = PG_GETARG_INT64(2);
+ text *context_type_text = PG_GETARG_TEXT_PP(3);
+ char *context_type;
+ int64 curr_memory_use = 0;
+ int64 remaining_alloc_bytes = total_alloc;
+ MemoryContext context;
+ MemoryContext oldContext;
+ AllocateTestNext *next_free_ptr = NULL;
+ AllocateTestNext *last_alloc = NULL;
+ clock_t start, end;
+
+ if (chunk_size < sizeof(AllocateTestNext))
+ elog(ERROR, "chunk_size (%d) must be at least %ld bytes", chunk_size,
+ sizeof(AllocateTestNext));
+ if (keep_memory > total_alloc)
+ elog(ERROR, "keep_memory (" INT64_FORMAT ") must be less than total_alloc (" INT64_FORMAT ")",
+ keep_memory, total_alloc);
+
+ context_type = text_to_cstring(context_type_text);
+
+ start = clock();
+
+ if (strcmp(context_type, "generation") == 0)
+ context = GenerationContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "aset") == 0)
+ context = AllocSetContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "slab") == 0)
+ context = SlabContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_MAXSIZE,
+ chunk_size);
+ else
+ elog(ERROR, "context_type must be \"generation\", \"aset\" or \"slab\"");
+
+ oldContext = MemoryContextSwitchTo(context);
+
+ while (remaining_alloc_bytes > 0)
+ {
+ AllocateTestNext *curr_alloc;
+
+ CHECK_FOR_INTERRUPTS();
+
+ /* Allocate the memory and update the counters */
+ curr_alloc = (AllocateTestNext *) palloc(chunk_size);
+ remaining_alloc_bytes -= chunk_size;
+ curr_memory_use += chunk_size;
+
+#ifdef ALLOCATE_TEST_DEBUG
+ elog(NOTICE, "alloc %p (curr_memory_use " INT64_FORMAT " bytes, remaining_alloc_bytes " INT64_FORMAT ")", curr_alloc, curr_memory_use, remaining_alloc_bytes);
+#endif
+
+ /*
+ * Point the last allocate to this one so that we can free allocations
+ * starting with the oldest first.
+ */
+ curr_alloc->next = NULL;
+ if (last_alloc != NULL)
+ last_alloc->next = curr_alloc;
+
+ if (next_free_ptr == NULL)
+ {
+ /*
+ * Remember the first chunk to free. We will follow the ->next
+ * pointers to find the next chunk to free when freeing memory
+ */
+ next_free_ptr = curr_alloc;
+ }
+
+ /*
+ * If the currently allocated memory has reached or exceeded the amount
+ * of memory we want to keep allocated at once then we'd better free
+ * some. Since all allocations are the same size we only need to free
+ * one allocation per loop.
+ */
+ if (curr_memory_use >= keep_memory)
+ {
+ AllocateTestNext *next = next_free_ptr->next;
+
+ /* free the memory and update the current memory usage */
+ pfree(next_free_ptr);
+ curr_memory_use -= chunk_size;
+
+#ifdef ALLOCATE_TEST_DEBUG
+ elog(NOTICE, "free %p (curr_memory_use " INT64_FORMAT " bytes, remaining_alloc_bytes " INT64_FORMAT ")", next_free_ptr, curr_memory_use, remaining_alloc_bytes);
+#endif
+ /* get the next chunk to free */
+ next_free_ptr = next;
+ }
+
+ if (curr_memory_use > 0)
+ last_alloc = curr_alloc;
+ else
+ last_alloc = NULL;
+ }
+
+ /* cleanup loop -- pfree remaining memory */
+ while (next_free_ptr != NULL)
+ {
+ AllocateTestNext *next = next_free_ptr->next;
+
+ /* free the memory and update the current memory usage */
+ pfree(next_free_ptr);
+ curr_memory_use -= chunk_size;
+
+#ifdef ALLOCATE_TEST_DEBUG
+ elog(NOTICE, "free %p (curr_memory_use " INT64_FORMAT " bytes, remaining_alloc_bytes " INT64_FORMAT ")", next_free_ptr, curr_memory_use, remaining_alloc_bytes);
+#endif
+
+ next_free_ptr = next;
+ }
+
+ MemoryContextSwitchTo(oldContext);
+
+ end = clock();
+
+ PG_RETURN_FLOAT8((double) (end - start) / CLOCKS_PER_SEC);
+}
+
+Datum
+pg_allocate_memory_test_reset(PG_FUNCTION_ARGS)
+{
+ int32 chunk_size = PG_GETARG_INT32(0);
+ int64 keep_memory = PG_GETARG_INT64(1);
+ int64 total_alloc = PG_GETARG_INT64(2);
+ text *context_type_text = PG_GETARG_TEXT_PP(3);
+ char *context_type;
+ int64 curr_memory_use = 0;
+ int64 remaining_alloc_bytes = total_alloc;
+ MemoryContext context;
+ MemoryContext oldContext;
+ clock_t start, end;
+
+ if (chunk_size < 1)
+ elog(ERROR, "size of chunk must be above 0");
+ if (keep_memory > total_alloc)
+ elog(ERROR, "keep_memory (" INT64_FORMAT ") must be less than total_alloc (" INT64_FORMAT ")",
+ keep_memory, total_alloc);
+
+ context_type = text_to_cstring(context_type_text);
+
+ start = clock();
+
+ if (strcmp(context_type, "generation") == 0)
+ context = GenerationContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "aset") == 0)
+ context = AllocSetContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "slab") == 0)
+ context = SlabContextCreate(CurrentMemoryContext,
+ "pg_allocate_memory_test",
+ ALLOCSET_DEFAULT_MAXSIZE,
+ chunk_size);
+ else
+ elog(ERROR, "context_type must be \"generation\", \"aset\" or \"slab\"");
+
+ oldContext = MemoryContextSwitchTo(context);
+
+ while (remaining_alloc_bytes > 0)
+ {
+ CHECK_FOR_INTERRUPTS();
+
+ /* Allocate the memory and update the counters */
+ (void) palloc(chunk_size);
+ remaining_alloc_bytes -= chunk_size;
+ curr_memory_use += chunk_size;
+
+#ifdef ALLOCATE_TEST_DEBUG
+ elog(NOTICE, "alloc %p (curr_memory_use " INT64_FORMAT " bytes, remaining_alloc_bytes " INT64_FORMAT ")", curr_alloc, curr_memory_use, remaining_alloc_bytes);
+#endif
+
+ /*
+ * If the currently allocated memory has reached or exceeded the amount
+ * of memory we want to keep allocated at once then reset the context.
+ */
+ if (curr_memory_use >= keep_memory)
+ {
+ curr_memory_use = 0;
+ MemoryContextReset(context);
+ }
+ }
+
+ MemoryContextSwitchTo(oldContext);
+ MemoryContextDelete(context);
+
+ end = clock();
+
+ PG_RETURN_FLOAT8((double) (end - start) / CLOCKS_PER_SEC);
+}
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index 719599649a..6bb53529a0 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -8143,6 +8143,18 @@
prorettype => 'bool', proargtypes => 'int4',
prosrc => 'pg_log_backend_memory_contexts' },
+# just for testing memory context allocation speed
+{ oid => '9319', descr => 'for testing performance of allocation and freeing',
+ proname => 'pg_allocate_memory_test', provolatile => 'v',
+ prorettype => 'float8', proargtypes => 'int4 int8 int8 text',
+ prosrc => 'pg_allocate_memory_test' },
+
+# just for testing memory context allocation speed
+{ oid => '9320', descr => 'for testing performance of allocation and resetting',
+ proname => 'pg_allocate_memory_test_reset', provolatile => 'v',
+ prorettype => 'float8', proargtypes => 'int4 int8 int8 text',
+ prosrc => 'pg_allocate_memory_test_reset' },
+
# non-persistent series generator
{ oid => '1066', descr => 'non-persistent series generator',
proname => 'generate_series', prorows => '1000',
[image/png] slab_64byte_chunks_chart.png (52.8K, ../../CAApHDvr0pKjrGVKFg2VqSUjbB1QxsKobu0PhD11qkSO_caL3dg@mail.gmail.com/4-slab_64byte_chunks_chart.png)
download | view image
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-05 08:02 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-05 10:18 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-10 04:01 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
@ 2022-12-12 07:13 ` John Naylor <john.naylor@enterprisedb.com>
2022-12-13 00:49 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
0 siblings, 1 reply; 29+ messages in thread
From: John Naylor @ 2022-12-12 07:13 UTC (permalink / raw)
To: David Rowley <dgrowleyml@gmail.com>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On Sat, Dec 10, 2022 at 11:02 AM David Rowley <dgrowleyml@gmail.com> wrote:
> [v4]
Thanks for working on this!
I ran an in-situ benchmark using the v13 radix tree patchset ([1] WIP but
should be useful enough for testing allocation speed), only applying the
first five, which are local-memory only. The benchmark is not meant to
represent a realistic workload, and primarily stresses traversal and
allocation of the smallest node type. Minimum of five, with turbo-boost
off, on recent Intel laptop hardware:
v13-0001 to 0005:
# select * from bench_load_random_int(500 * 1000);
mem_allocated | load_ms
---------------+---------
151123432 | 222
47.06% postgres postgres [.] rt_set
22.89% postgres postgres [.] SlabAlloc
9.65% postgres postgres [.] rt_node_insert_inner.isra.0
5.94% postgres [unknown] [k] 0xffffffffb5e011b7
3.62% postgres postgres [.] MemoryContextAlloc
2.70% postgres libc.so.6 [.] __memmove_avx_unaligned_erms
2.60% postgres postgres [.] SlabFree
+ v4 slab:
# select * from bench_load_random_int(500 * 1000);
mem_allocated | load_ms
---------------+---------
152463112 | 213
52.42% postgres postgres [.] rt_set
12.80% postgres postgres [.] SlabAlloc
9.38% postgres postgres [.] rt_node_insert_inner.isra.0
7.87% postgres [unknown] [k] 0xffffffffb5e011b7
4.98% postgres postgres [.] SlabFree
While allocation is markedly improved, freeing looks worse here. The
proportion is surprising because only about 2% of nodes are freed during
the load, but doing that takes up 10-40% of the time compared to allocating.
num_keys = 500000, height = 7
n4 = 2501016, n15 = 56932, n32 = 270, n125 = 0, n256 = 257
Sidenote: I don't recall ever seeing vsyscall (I think that's what the
0xffffffffb5e011b7 address is referring to) in a profile, so not sure what
is happening there.
[1]
https://www.postgresql.org/message-id/CAFBsxsHNE621mGuPhd7kxaGc22vMkoSu7R4JW9Zan1jjorGy3g%40mail.gma...
--
John Naylor
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-05 08:02 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-05 10:18 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-10 04:01 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-12 07:13 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
@ 2022-12-13 00:49 ` David Rowley <dgrowleyml@gmail.com>
2022-12-14 10:37 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
0 siblings, 1 reply; 29+ messages in thread
From: David Rowley @ 2022-12-13 00:49 UTC (permalink / raw)
To: John Naylor <john.naylor@enterprisedb.com>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
Thanks for testing the patch.
On Mon, 12 Dec 2022 at 20:14, John Naylor <john.naylor@enterprisedb.com> wrote:
> v13-0001 to 0005:
> 2.60% postgres postgres [.] SlabFree
> + v4 slab:
> 4.98% postgres postgres [.] SlabFree
>
> While allocation is markedly improved, freeing looks worse here. The proportion is surprising because only about 2% of nodes are freed during the load, but doing that takes up 10-40% of the time compared to allocating.
I've tried to reproduce this with the v13 patches applied and I'm not
really getting the same as you are. To run the function 100 times I
used:
select x, a.* from generate_series(1,100) x(x), lateral (select * from
bench_load_random_int(500 * 1000 * (1+x-x))) a;
(I had to add the * (1+x-x) to add a lateral dependency to stop the
function just being executed once)
v13-0001 - 0005 gives me:
37.71% postgres [.] rt_set
19.24% postgres [.] SlabAlloc
8.73% [kernel] [k] clear_page_rep
5.21% postgres [.] rt_node_insert_inner.isra.0
2.63% [kernel] [k] asm_exc_page_fault
2.24% postgres [.] SlabFree
and fairly consistently 122 ms runtime per call.
Applying v4 slab patch I get:
41.06% postgres [.] rt_set
10.84% postgres [.] SlabAlloc
9.01% [kernel] [k] clear_page_rep
6.49% postgres [.] rt_node_insert_inner.isra.0
2.76% postgres [.] SlabFree
and fairly consistently 112 ms per call.
I wonder if you can consistently get the same result on another
compiler or after patching something like master~50 or master~100.
Maybe it's just a code alignment thing.
Looking at the annotation of perf report for SlabFree with the patched
version I see:
│
│ /* push this chunk onto the head of the free list */
│ *(MemoryChunk **) pointer = block->freehead;
0.09 │ mov 0x10(%r8),%rax
│ slab = block->slab;
59.15 │ mov (%r8),%rbp
│ *(MemoryChunk **) pointer = block->freehead;
9.43 │ mov %rax,(%rdi)
│ block->freehead = chunk;
│
│ block->nfree++;
I think what that's telling me is that dereferencing the block's
memory is slow, likely due to that particular cache line not being
cached any longer. I tried running the test with 10,000 ints instead
of 500,000 so that there would be less CPU cache pressure. I see:
29.76 │ mov (%r8),%rbp
│ *(MemoryChunk **) pointer = block->freehead;
12.72 │ mov %rax,(%rdi)
│ block->freehead = chunk;
│
│ block->nfree++;
│ mov 0x8(%r8),%eax
│ block->freehead = chunk;
4.27 │ mov %rdx,0x10(%r8)
│ SlabBlocklistIndex():
│ index = (nfree + (1 << blocklist_shift) - 1) >> blocklist_shift;
│ mov $0x1,%edx
│ SlabFree():
│ block->nfree++;
│ lea 0x1(%rax),%edi
│ mov %edi,0x8(%r8)
│ SlabBlocklistIndex():
│ int32 blocklist_shift = slab->blocklist_shift;
│ mov 0x70(%rbp),%ecx
│ index = (nfree + (1 << blocklist_shift) - 1) >> blocklist_shift;
8.46 │ shl %cl,%edx
various other instructions in SlabFree are proportionally taking
longer now. For example the bitshift at the end was insignificant
previously. That indicates to me that this is due to caching effects.
We must fetch the block in SlabFree() in both versions. It's possible
that something is going on in SlabAlloc() that is causing more useful
cachelines to be evicted, but (I think) one of primary design goals
Andres was going for was to reduce that. For example not having to
write out the freelist for an entire block when the block is first
allocated means not having to load possibly all cache lines for the
entire block anymore.
I tried looking at perf stat during the run.
Without slab changes:
drowley@amd3990x:~$ sudo perf stat --pid=74922 sleep 2
Performance counter stats for process id '74922':
2,000.74 msec task-clock # 1.000 CPUs utilized
4 context-switches # 1.999 /sec
0 cpu-migrations # 0.000 /sec
578,139 page-faults # 288.963 K/sec
8,614,687,392 cycles # 4.306 GHz
(83.21%)
682,574,688 stalled-cycles-frontend # 7.92% frontend
cycles idle (83.33%)
4,822,904,271 stalled-cycles-backend # 55.98% backend
cycles idle (83.41%)
11,447,124,105 instructions # 1.33 insn per cycle
# 0.42 stalled
cycles per insn (83.41%)
1,947,647,575 branches # 973.464 M/sec
(83.41%)
13,914,897 branch-misses # 0.71% of all
branches (83.24%)
2.000924020 seconds time elapsed
With slab changes:
drowley@amd3990x:~$ sudo perf stat --pid=75967 sleep 2
Performance counter stats for process id '75967':
2,000.89 msec task-clock # 1.000 CPUs utilized
1 context-switches # 0.500 /sec
0 cpu-migrations # 0.000 /sec
607,423 page-faults # 303.576 K/sec
8,566,091,176 cycles # 4.281 GHz
(83.21%)
737,839,390 stalled-cycles-frontend # 8.61% frontend
cycles idle (83.32%)
4,454,357,725 stalled-cycles-backend # 52.00% backend
cycles idle (83.41%)
10,760,559,837 instructions # 1.26 insn per cycle
# 0.41 stalled
cycles per insn (83.41%)
1,872,047,962 branches # 935.606 M/sec
(83.41%)
14,928,953 branch-misses # 0.80% of all
branches (83.25%)
2.000960610 seconds time elapsed
It would be interesting to see if your perf stat output is showing
something significantly different with and without the slab changes.
It does not seem impossible that due to the slab changes having to
look at less memory in SlabAlloc() that that's moving some additional
requirements for SlabFree() to fetch cache lines that in the unpatched
version would have already been available. If that is the case, then
I think we shouldn't worry about it unless we can find some workload
that demonstrates an overall performance regression with the patch. I
just don't quite have enough perf experience to know how I might go
about proving that.
David
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-05 08:02 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-05 10:18 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-10 04:01 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-12 07:13 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-13 00:49 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
@ 2022-12-14 10:37 ` John Naylor <john.naylor@enterprisedb.com>
2022-12-20 03:35 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
0 siblings, 1 reply; 29+ messages in thread
From: John Naylor @ 2022-12-14 10:37 UTC (permalink / raw)
To: David Rowley <dgrowleyml@gmail.com>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On Tue, Dec 13, 2022 at 7:50 AM David Rowley <dgrowleyml@gmail.com> wrote:
>
> Thanks for testing the patch.
>
> On Mon, 12 Dec 2022 at 20:14, John Naylor <john.naylor@enterprisedb.com>
wrote:
> > While allocation is markedly improved, freeing looks worse here. The
proportion is surprising because only about 2% of nodes are freed during
the load, but doing that takes up 10-40% of the time compared to allocating.
>
> I've tried to reproduce this with the v13 patches applied and I'm not
> really getting the same as you are. To run the function 100 times I
> used:
>
> select x, a.* from generate_series(1,100) x(x), lateral (select * from
> bench_load_random_int(500 * 1000 * (1+x-x))) a;
Simply running over a longer period of time like this makes the SlabFree
difference much closer to your results, so it doesn't seem out of line
anymore. Here SlabAlloc seems to take maybe 2/3 of the time of current
slab, with a 5% reduction in total time:
500k ints:
v13-0001-0005
average of 30: 217ms
47.61% postgres postgres [.] rt_set
20.99% postgres postgres [.] SlabAlloc
10.00% postgres postgres [.] rt_node_insert_inner.isra.0
6.87% postgres [unknown] [k] 0xffffffffbce011b7
3.53% postgres postgres [.] MemoryContextAlloc
2.82% postgres postgres [.] SlabFree
+slab v4
average of 30: 206ms
51.13% postgres postgres [.] rt_set
14.08% postgres postgres [.] SlabAlloc
11.41% postgres postgres [.] rt_node_insert_inner.isra.0
7.44% postgres [unknown] [k] 0xffffffffbce011b7
3.89% postgres postgres [.] MemoryContextAlloc
3.39% postgres postgres [.] SlabFree
It doesn't look mysterious anymore, but I went ahead and took some more
perf measurements, including for cache misses. My naive impression is that
we're spending a bit more time waiting for data, but having to do less work
with it once we get it, which is consistent with your earlier comments:
perf stat -p $pid sleep 2
v13:
2,001.55 msec task-clock:u # 1.000 CPUs
utilized
0 context-switches:u # 0.000 /sec
0 cpu-migrations:u # 0.000 /sec
311,690 page-faults:u # 155.724 K/sec
3,128,740,701 cycles:u # 1.563 GHz
4,739,333,861 instructions:u # 1.51 insn
per cycle
820,014,588 branches:u # 409.690 M/sec
7,385,923 branch-misses:u # 0.90% of all
branches
+slab v4:
2,001.09 msec task-clock:u # 1.000 CPUs
utilized
0 context-switches:u # 0.000 /sec
0 cpu-migrations:u # 0.000 /sec
326,017 page-faults:u # 162.920 K/sec
3,016,668,818 cycles:u # 1.508 GHz
4,324,863,908 instructions:u # 1.43 insn
per cycle
761,839,927 branches:u # 380.712 M/sec
7,718,366 branch-misses:u # 1.01% of all
branches
perf stat -e LLC-loads,LLC-loads-misses -p $pid sleep 2
min/max of 3 runs:
v13: LL cache misses: 25.08% - 25.41%
+slab v4: LL cache misses: 25.74% - 26.01%
--
John Naylor
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-05 08:02 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-05 10:18 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-10 04:01 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-12 07:13 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-13 00:49 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-14 10:37 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
@ 2022-12-20 03:35 ` David Rowley <dgrowleyml@gmail.com>
2022-12-20 08:19 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
0 siblings, 1 reply; 29+ messages in thread
From: David Rowley @ 2022-12-20 03:35 UTC (permalink / raw)
To: John Naylor <john.naylor@enterprisedb.com>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
I've spent quite a bit more time on the slab changes now and I've
attached a v3 patch.
One of the major things that I've spent time on is benchmarking this.
I'm aware that Tomas wrote some functions to benchmark. I've taken
those and made some modifications to allow the memory context type to
be specified as a function parameter. This allows me to easily
compare the performance of slab with both aset and generation.
Another change that I made to Tomas' module was how the random
ordering part works. What I wanted was the ability to specify how
randomly to pfree the chunks and test various "degrees-of-randomness"
to see how that affects the performance. What I ended up coming up
with was the ability to specify the number of "random segments". This
controls how many groups we split all allocated chunks into to
randomise. If there is 1 random segment, then that's just randomising
over all chunks. If there are 10 random segments, then we split the
array of allocated chunks into 10 portions based on either FIFO or
LIFO order, then randomise the order of the chunks only within each of
those segments. This allows us to test FIFO/LIFO allocation patterns
with and without random and any degrees of that in between. If the
random segments is set to 0, then no randomisation is done.
Another change I made to Tomas' code was, I'm now using palloc0()
instead of palloc() and I'm also checking the first byte of the
allocated chunk is '\0' before pfreeing it. What I was finding was
that pfree was showing as highly dominant in perf output due to it
having to deference the MemoryChunk to find the context-type bits.
pfree had to do this as none of the calling code had touched any of
the memory in the chunk. I felt it was unrealistic to be pallocing
memory and not doing anything with it and then pfreeing it without
having done anything with it. Mostly this just moves the
responsibilities around of which function is penalised in having to
load the cache line. I mostly did this as I was struggling to make any
sense of perf's output.
I've attached alloc_bench_contrib.patch which I used for testing.
I've also attached a spreadsheet with the benchmark results. The
general summary from having done those is that slab is now generally
now on-par with aset in terms of palloc performance. Previously slab
was performing at about half the speed of aset unless CPU cache
pressure became more significant, in which case the performance is
dominated by fetching cache lines from RAM. However, the new code
still makes meaningful improvements even under heavy CPU cache
pressure. When it comes to pfree performance, the updated slab code
is much faster than it was previously, but not quite on-par with aset
or generation.
The attached spreadsheet is broken down into 3 tabs. Each tab is
testing a chunk size and a fixed total number of chunks allocated at
once. Within each tab, I'm testing FIFO and then LIFO allocation
patterns each with a different degree of randomness introduced, as I
described above. In none of the tests was the patched version slower
than the unpatched version.
One pending question I had was about SlabStats where we list free
chunks. Since we now have a list of emptyblocks, I wasn't too sure if
the chunks from those should be included in that total. I currently
am not including them, but I have added some additional information to
list the number of completely empty blocks that we've got in the
emptyblocks list.
Some follow-up work that I'm thinking is a good idea:
1. Reduce the SlabContext's chunkSize, fullChunkSize and blockSize
fields from Size down to uint32. These have no need to be 64 bits. We
don't allow slab blocks over 1GB since c6e0fe1f2. I thought of doing
this separately as we might need to rationalise the equivalent fields
in aset.c and generation.c. Those can have external chunks, so I'm
not 100% sure if we should do that there or not yet. I just didn't
want to touch those files in this effort.
2. Slab should probably gain the ability to grow the block size as
aset and generation both do. Since the performance of the slab context
is good now, we might want to use it for hash join's 32kb chunks, but
I doubt we can without the block size growth.
I'm planning on pushing the attached v3 patch shortly. I've spent
several days reading over this and testing it in detail along with
adding additional features to the SlabCheck code to find more
inconsistencies.
David
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index c2f9bb6ad3..b6f8ed28a7 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -4,7 +4,8 @@
* SLAB allocator definitions.
*
* SLAB is a MemoryContext implementation designed for cases where large
- * numbers of equally-sized objects are allocated (and freed).
+ * numbers of equally-sized objects can be allocated and freed efficiently
+ * with minimal memory wastage and fragmentation.
*
*
* Portions Copyright (c) 2017-2022, PostgreSQL Global Development Group
@@ -16,36 +17,51 @@
* NOTE:
* The constant allocation size allows significant simplification and various
* optimizations over more general purpose allocators. The blocks are carved
- * into chunks of exactly the right size (plus alignment), not wasting any
- * memory.
+ * into chunks of exactly the right size, wasting only the space required to
+ * MAXALIGN the allocated chunks.
*
- * The information about free chunks is maintained both at the block level and
- * global (context) level. This is possible as the chunk size (and thus also
- * the number of chunks per block) is fixed.
+ * Slab can also help reduce memory fragmentation in cases where longer-lived
+ * chunks remain stored on blocks while most of the other chunks have already
+ * been pfree'd. We give priority to putting new allocations into the
+ * "fullest" block. This help avoid having too many sparsely used blocks
+ * around and allows blocks to more easily become completely unused which
+ * allows them to be eventually free'd.
*
- * On each block, free chunks are tracked in a simple linked list. Contents
- * of free chunks is replaced with an index of the next free chunk, forming
- * a very simple linked list. Each block also contains a counter of free
- * chunks. Combined with the local block-level freelist, it makes it trivial
- * to eventually free the whole block.
+ * We identify the "fullest" block to put new allocations on by using a block
+ * from the lowest populated element of the context's "blocklist" array.
+ * This is an array of dlists containing blocks which we partition by the
+ * number of free chunks which block has. Blocks with fewer free chunks are
+ * stored in a lower indexed dlist array slot. Full blocks go on the 0th
+ * element of the blocklist array. So that we don't have to have too many
+ * elements in the array, each dlist in the array is responsible for a range
+ * of free chunks. When a chunk is palloc'd or pfree'd we may need to move
+ * the block onto another dlist if the number of free chunks crosses the
+ * range boundary that the current list is responsible for. Having just a
+ * few blocklist elements reduces the number of times we must move the block
+ * onto another dlist element.
*
- * At the context level, we use 'freelist' to track blocks ordered by number
- * of free chunks, starting with blocks having a single allocated chunk, and
- * with completely full blocks on the tail.
+ * We keep track of free chunks within each block by using a block-level free
+ * list. We consult this list when we allocate a new chunk in the block.
+ * The free list is a linked list, the head of which is pointed to with
+ * SlabBlock's freehead field. Each subsequent list item is stored in the
+ * free chunk's memory. We ensure chunks are large enough to store this
+ * address.
*
- * This also allows various optimizations - for example when searching for
- * free chunk, the allocator reuses space from the fullest blocks first, in
- * the hope that some of the less full blocks will get completely empty (and
- * returned back to the OS).
- *
- * For each block, we maintain pointer to the first free chunk - this is quite
- * cheap and allows us to skip all the preceding used chunks, eliminating
- * a significant number of lookups in many common usage patterns. In the worst
- * case this performs as if the pointer was not maintained.
- *
- * We cache the freelist index for the blocks with the fewest free chunks
- * (minFreeChunks), so that we don't have to search the freelist on every
- * SlabAlloc() call, which is quite expensive.
+ * When we allocate a new block, technically all chunks are free, however, to
+ * avoid having to write out the entire block to set the linked list for the
+ * free chunks for every chunk in the block, we instead store a pointer to
+ * the next "unused" chunk on the block and keep track of how many of these
+ * unused chunks there are. When a new block is malloc'd, all chunks are
+ * unused. The unused pointer starts with the first chunk on the block and
+ * as chunks are allocated, the unused pointer is incremented. As chunks are
+ * pfree'd, the unused pointer never goes backwards. The unused pointer can
+ * be thought of as a high watermark for the maximum number of chunks in the
+ * block which have been in use concurrently. When a chunk is pfree'd the
+ * chunk is put onto the head of the free list and the unused pointer is not
+ * changed. We only consume more unused chunks if we run out of free chunks
+ * on the free list. This method effectively gives priority to using
+ * previously used chunks over previously unused chunks, which should perform
+ * better due to CPU caching effects.
*
*-------------------------------------------------------------------------
*/
@@ -60,6 +76,27 @@
#define Slab_BLOCKHDRSZ MAXALIGN(sizeof(SlabBlock))
+#ifdef MEMORY_CONTEXT_CHECKING
+/*
+ * Size of the memory required to store the SlabContext.
+ * MEMORY_CONTEXT_CHECKING builds need some extra memory for the isChunkFree
+ * array.
+ */
+#define Slab_CONTEXT_HDRSZ(chnksperblk) \
+ (sizeof(SlabContext) + ((chnksperblk) * sizeof(bool)))
+#else
+#define Slab_CONTEXT_HDRSZ(chnksperblk) sizeof(SlabContext)
+#endif
+
+/*
+ * The number of partitions to divide the blocklist into based their number of
+ * free chunks. There must be at least 2.
+ */
+#define SLAB_BLOCKLIST_COUNT 3
+
+/* The maximum number of completely empty blocks to keep around for reuse. */
+#define SLAB_MAXIMUM_EMPTY_BLOCKS 10
+
/*
* SlabContext is a specialized implementation of MemoryContext.
*/
@@ -67,64 +104,206 @@ typedef struct SlabContext
{
MemoryContextData header; /* Standard memory-context fields */
/* Allocation parameters for this context: */
- Size chunkSize; /* chunk size */
- Size fullChunkSize; /* chunk size including header and alignment */
- Size blockSize; /* block size */
- Size headerSize; /* allocated size of context header */
- int chunksPerBlock; /* number of chunks per block */
- int minFreeChunks; /* min number of free chunks in any block */
- int nblocks; /* number of blocks allocated */
+ Size chunkSize; /* the requested (non-aligned) chunk size */
+ Size fullChunkSize; /* chunk size with chunk header and alignment */
+ Size blockSize; /* the size to make each block of chunks */
+ int32 chunksPerBlock; /* number of chunks that fit in 1 block */
+ int32 curBlocklistIndex; /* index into the blocklist[] element
+ * containing the fullest, blocks */
#ifdef MEMORY_CONTEXT_CHECKING
- bool *freechunks; /* bitmap of free chunks in a block */
+ bool *isChunkFree; /* array to mark free chunks in a block during
+ * SlabCheck */
#endif
- /* blocks with free space, grouped by number of free chunks: */
- dlist_head freelist[FLEXIBLE_ARRAY_MEMBER];
+
+ int32 blocklist_shift; /* number of bits to shift the nfree count
+ * by to get the index into blocklist[] */
+ dclist_head emptyblocks; /* empty blocks to use up first instead of
+ * mallocing new blocks */
+
+ /*
+ * Blocks with free space, grouped by the number of free chunks they
+ * contain. Completely full blocks are stored in the 0th element.
+ * Completely empty blocks are stored in emptyblocks or free'd if we have
+ * enough empty blocks already.
+ */
+ dlist_head blocklist[SLAB_BLOCKLIST_COUNT];
} SlabContext;
/*
* SlabBlock
- * Structure of a single block in SLAB allocator.
+ * Structure of a single slab block.
*
- * node: doubly-linked list of blocks in global freelist
- * nfree: number of free chunks in this block
- * firstFreeChunk: index of the first free chunk
+ * slab: pointer back to the owning MemoryContext
+ * nfree: number of chunks on the block which are unallocated
+ * nunused: number of chunks on the block unallocated and not on the block's
+ * freelist.
+ * freehead: linked-list header storing a pointer to the first free chunk on
+ * the block. Subsequent pointers are stored in the chunk's memory. NULL
+ * indicates the end of the list.
+ * unused: pointer to the next chunk which has yet to be used.
+ * node: doubly-linked list node for the context's blocklist
*/
typedef struct SlabBlock
{
- dlist_node node; /* doubly-linked list */
- int nfree; /* number of free chunks */
- int firstFreeChunk; /* index of the first free chunk in the block */
SlabContext *slab; /* owning context */
+ int32 nfree; /* number of chunks on free + unused chunks */
+ int32 nunused; /* number of unused chunks */
+ MemoryChunk *freehead; /* pointer to the first free chunk */
+ MemoryChunk *unused; /* pointer to the next unused chunk */
+ dlist_node node; /* doubly-linked list for blocklist[] */
} SlabBlock;
#define Slab_CHUNKHDRSZ sizeof(MemoryChunk)
-#define SlabPointerGetChunk(ptr) \
- ((MemoryChunk *)(((char *)(ptr)) - sizeof(MemoryChunk)))
#define SlabChunkGetPointer(chk) \
- ((void *)(((char *)(chk)) + sizeof(MemoryChunk)))
-#define SlabBlockGetChunk(slab, block, idx) \
+ ((void *) (((char *) (chk)) + sizeof(MemoryChunk)))
+
+/*
+ * SlabBlockGetChunk
+ * Obtain a pointer to the nth (0-based) chunk in the block
+ */
+#define SlabBlockGetChunk(slab, block, n) \
((MemoryChunk *) ((char *) (block) + Slab_BLOCKHDRSZ \
- + (idx * slab->fullChunkSize)))
-#define SlabBlockStart(block) \
- ((char *) block + Slab_BLOCKHDRSZ)
+ + ((n) * (slab)->fullChunkSize)))
+
+#if defined(MEMORY_CONTEXT_CHECKING) || defined(USE_ASSERT_CHECKING)
+
+/*
+ * SlabChunkIndex
+ * Get the 0-based index of how many chunks into the block the given
+ * chunk is.
+*/
#define SlabChunkIndex(slab, block, chunk) \
- (((char *) chunk - SlabBlockStart(block)) / slab->fullChunkSize)
+ (((char *) (chunk) - (char *) SlabBlockGetChunk(slab, block, 0)) / \
+ (slab)->fullChunkSize)
+
+/*
+ * SlabChunkMod
+ * A MemoryChunk should always be at an address which is a multiple of
+ * fullChunkSize starting from the 0th chunk position. This will return
+ * non-zero if it's not.
+ */
+#define SlabChunkMod(slab, block, chunk) \
+ (((char *) (chunk) - (char *) SlabBlockGetChunk(slab, block, 0)) % \
+ (slab)->fullChunkSize)
+
+#endif
/*
* SlabIsValid
- * True iff set is valid slab allocation set.
+ * True iff set is a valid slab allocation set.
*/
-#define SlabIsValid(set) \
- (PointerIsValid(set) && IsA(set, SlabContext))
+#define SlabIsValid(set) (PointerIsValid(set) && IsA(set, SlabContext))
/*
* SlabBlockIsValid
- * True iff block is valid block of slab allocation set.
+ * True iff block is a valid block of slab allocation set.
*/
#define SlabBlockIsValid(block) \
(PointerIsValid(block) && SlabIsValid((block)->slab))
+/*
+ * SlabBlocklistIndex
+ * Determine the blocklist index that a block should be in for the given
+ * number of free chunks.
+ */
+static inline int32
+SlabBlocklistIndex(SlabContext *slab, int nfree)
+{
+ int32 index;
+ int32 blocklist_shift = slab->blocklist_shift;
+
+ Assert(nfree >= 0 && nfree <= slab->chunksPerBlock);
+
+ /*
+ * Determine the blocklist index based on the number of free chunks. We
+ * must ensure that 0 free chunks is dedicated to index 0. Everything else
+ * must be >= 1 and < SLAB_BLOCKLIST_COUNT.
+ *
+ * To make this as efficient as possible, we exploit some two's complement
+ * arithmetic where we reverse the sign before bit shifting. This results
+ * in an nfree of 0 using index 0 and anything non-zero staying non-zero.
+ * This is exploiting 0 and -0 being the same in two's complement. When
+ * we're done, we just need to flip the sign back over again for a
+ * positive index.
+ */
+ index = -((-nfree) >> blocklist_shift);
+
+ if (nfree == 0)
+ Assert(index == 0);
+ else
+ Assert(index >= 1 && index < SLAB_BLOCKLIST_COUNT);
+
+ return index;
+}
+
+/*
+ * SlabFindNextBlockListIndex
+ * Search blocklist for blocks which have free chunks and return the
+ * index of the blocklist found containing at least 1 block with free
+ * chunks. If no block can be found we return 0.
+ *
+ * Note: We give priority to fuller blocks so that these are filled before
+ * emptier blocks. This is done to increase the chances that mostly-empty
+ * blocks will eventually become completely empty so they can be free'd.
+ */
+static int32
+SlabFindNextBlockListIndex(SlabContext *slab)
+{
+ /* start at 1 as blocklist[0] is for full blocks. */
+ for (int i = 1; i < SLAB_BLOCKLIST_COUNT; i++)
+ {
+ /* return the first found non-empty index */
+ if (!dlist_is_empty(&slab->blocklist[i]))
+ return i;
+ }
+
+ /* no blocks with free space */
+ return 0;
+}
+
+/*
+ * SlabGetNextFreeChunk
+ * Return the next free chunk in block and update the block to account
+ * for the returned chunk now being used.
+ */
+static inline MemoryChunk *
+SlabGetNextFreeChunk(SlabContext *slab, SlabBlock *block)
+{
+ MemoryChunk *chunk;
+
+ Assert(block->nfree > 0);
+
+ if (block->freehead != NULL)
+ {
+ chunk = block->freehead;
+
+ /*
+ * Pop the chunk from the linked list of free chunks. The pointer to
+ * the next free chunk is stored in the chunk itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(MemoryChunk *));
+ block->freehead = *(MemoryChunk **) SlabChunkGetPointer(chunk);
+
+ /* check nothing stomped on the free chunk's memory */
+ Assert(block->freehead == NULL ||
+ (block->freehead >= SlabBlockGetChunk(slab, block, 0) &&
+ block->freehead <= SlabBlockGetChunk(slab, block, slab->chunksPerBlock - 1) &&
+ SlabChunkMod(slab, block, block->freehead) == 0));
+ }
+ else
+ {
+ Assert(block->nunused > 0);
+
+ chunk = block->unused;
+ block->unused = (MemoryChunk *) (((char *)block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
+
+ block->nfree--;
+
+ return chunk;
+}
/*
* SlabContextCreate
@@ -145,8 +324,6 @@ SlabContextCreate(MemoryContext parent,
{
int chunksPerBlock;
Size fullChunkSize;
- Size freelistSize;
- Size headerSize;
SlabContext *slab;
int i;
@@ -155,11 +332,14 @@ SlabContextCreate(MemoryContext parent,
"sizeof(MemoryChunk) is not maxaligned");
Assert(MAXALIGN(chunkSize) <= MEMORYCHUNK_MAX_VALUE);
- /* Make sure the linked list node fits inside a freed chunk */
- if (chunkSize < sizeof(int))
- chunkSize = sizeof(int);
+ /*
+ * Ensure there's enough space to store the pointer to the next free chunk
+ * in the memory of the (otherwise) unused allocation.
+ */
+ if (chunkSize < sizeof(MemoryChunk *))
+ chunkSize = sizeof(MemoryChunk *);
- /* chunk, including SLAB header (both addresses nicely aligned) */
+ /* length of the maxaligned chunk including the chunk header */
#ifdef MEMORY_CONTEXT_CHECKING
/* ensure there's always space for the sentinel byte */
fullChunkSize = Slab_CHUNKHDRSZ + MAXALIGN(chunkSize + 1);
@@ -167,36 +347,17 @@ SlabContextCreate(MemoryContext parent,
fullChunkSize = Slab_CHUNKHDRSZ + MAXALIGN(chunkSize);
#endif
- /* Make sure the block can store at least one chunk. */
- if (blockSize < fullChunkSize + Slab_BLOCKHDRSZ)
- elog(ERROR, "block size %zu for slab is too small for %zu chunks",
- blockSize, chunkSize);
-
- /* Compute maximum number of chunks per block */
+ /* compute the number of chunks that will fit on each block */
chunksPerBlock = (blockSize - Slab_BLOCKHDRSZ) / fullChunkSize;
- /* The freelist starts with 0, ends with chunksPerBlock. */
- freelistSize = sizeof(dlist_head) * (chunksPerBlock + 1);
-
- /*
- * Allocate the context header. Unlike aset.c, we never try to combine
- * this with the first regular block; not worth the extra complication.
- */
+ /* Make sure the block can store at least one chunk. */
+ if (chunksPerBlock == 0)
+ elog(ERROR, "block size %zu for slab is too small for %zu-byte chunks",
+ blockSize, chunkSize);
- /* Size of the memory context header */
- headerSize = offsetof(SlabContext, freelist) + freelistSize;
-#ifdef MEMORY_CONTEXT_CHECKING
-
- /*
- * With memory checking, we need to allocate extra space for the bitmap of
- * free chunks. The bitmap is an array of bools, so we don't need to worry
- * about alignment.
- */
- headerSize += chunksPerBlock * sizeof(bool);
-#endif
- slab = (SlabContext *) malloc(headerSize);
+ slab = (SlabContext *) malloc(Slab_CONTEXT_HDRSZ(chunksPerBlock));
if (slab == NULL)
{
MemoryContextStats(TopMemoryContext);
@@ -216,19 +377,33 @@ SlabContextCreate(MemoryContext parent,
slab->chunkSize = chunkSize;
slab->fullChunkSize = fullChunkSize;
slab->blockSize = blockSize;
- slab->headerSize = headerSize;
slab->chunksPerBlock = chunksPerBlock;
- slab->minFreeChunks = 0;
- slab->nblocks = 0;
+ slab->curBlocklistIndex = 0;
- /* initialize the freelist slots */
- for (i = 0; i < (slab->chunksPerBlock + 1); i++)
- dlist_init(&slab->freelist[i]);
+ /*
+ * Compute a shift that guarantees that shifting chunksPerBlock with it is
+ * < SLAB_BLOCKLIST_COUNT - 1. The reason that we subtract 1 from
+ * SLAB_BLOCKLIST_COUNT in this calculation is that we reserve the 0th
+ * blocklist element for blocks which have no free chunks.
+ *
+ * We calculate the number of bits to shift by rather than a divisor to
+ * divide by as performing division each time we need to find the
+ * blocklist index would be much slower.
+ */
+ slab->blocklist_shift = 0;
+ while ((slab->chunksPerBlock >> slab->blocklist_shift) >= (SLAB_BLOCKLIST_COUNT - 1))
+ slab->blocklist_shift++;
+
+ /* initialize the list to store empty blocks to be reused */
+ dclist_init(&slab->emptyblocks);
+
+ /* initialize each blocklist slot */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
+ dlist_init(&slab->blocklist[i]);
#ifdef MEMORY_CONTEXT_CHECKING
- /* set the freechunks pointer right after the freelists array */
- slab->freechunks
- = (bool *) slab + offsetof(SlabContext, freelist) + freelistSize;
+ /* set the isChunkFree pointer right after the end of the context */
+ slab->isChunkFree = (bool *) ((char *) slab + sizeof(SlabContext));
#endif
/* Finally, do the type-independent part of context creation */
@@ -252,6 +427,7 @@ void
SlabReset(MemoryContext context)
{
SlabContext *slab = (SlabContext *) context;
+ dlist_mutable_iter miter;
int i;
Assert(SlabIsValid(slab));
@@ -261,12 +437,24 @@ SlabReset(MemoryContext context)
SlabCheck(context);
#endif
- /* walk over freelists and free the blocks */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /* release any retained empty blocks */
+ dclist_foreach_modify(miter, &slab->emptyblocks)
{
- dlist_mutable_iter miter;
+ SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
+
+ dclist_delete_from(&slab->emptyblocks, miter.cur);
- dlist_foreach_modify(miter, &slab->freelist[i])
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ context->mem_allocated -= slab->blockSize;
+ }
+
+ /* walk over blocklist and free the blocks */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
+ {
+ dlist_foreach_modify(miter, &slab->blocklist[i])
{
SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
@@ -276,14 +464,12 @@ SlabReset(MemoryContext context)
wipe_mem(block, slab->blockSize);
#endif
free(block);
- slab->nblocks--;
context->mem_allocated -= slab->blockSize;
}
}
- slab->minFreeChunks = 0;
+ slab->curBlocklistIndex = 0;
- Assert(slab->nblocks == 0);
Assert(context->mem_allocated == 0);
}
@@ -302,7 +488,7 @@ SlabDelete(MemoryContext context)
/*
* SlabAlloc
- * Returns pointer to allocated memory of given size or NULL if
+ * Returns a pointer to allocated memory of given size or NULL if
* request could not be completed; memory is added to the slab.
*/
void *
@@ -311,127 +497,118 @@ SlabAlloc(MemoryContext context, Size size)
SlabContext *slab = (SlabContext *) context;
SlabBlock *block;
MemoryChunk *chunk;
- int idx;
Assert(SlabIsValid(slab));
- Assert((slab->minFreeChunks >= 0) &&
- (slab->minFreeChunks < slab->chunksPerBlock));
+ /* sanity check that this is pointing to a valid blocklist */
+ Assert(slab->curBlocklistIndex >= 0);
+ Assert(slab->curBlocklistIndex <= SlabBlocklistIndex(slab, slab->chunksPerBlock));
/* make sure we only allow correct request size */
- if (size != slab->chunkSize)
+ if (unlikely(size != slab->chunkSize))
elog(ERROR, "unexpected alloc chunk size %zu (expected %zu)",
size, slab->chunkSize);
/*
- * If there are no free chunks in any existing block, create a new block
- * and put it to the last freelist bucket.
- *
- * slab->minFreeChunks == 0 means there are no blocks with free chunks,
- * thanks to how minFreeChunks is updated at the end of SlabAlloc().
+ * Handle the case when there are no partially filled blocks available.
+ * SlabFree() will have updated the curBlocklistIndex setting it to zero
+ * to indicate that it has freed the final block. Also later in
+ * SlabAlloc() we will set the curBlocklistIndex to zero if we end up
+ * filling the final block.
*/
- if (slab->minFreeChunks == 0)
+ if (unlikely(slab->curBlocklistIndex == 0))
{
- block = (SlabBlock *) malloc(slab->blockSize);
+ dlist_head *blocklist;
+ int blocklist_idx;
+
+ /* to save allocating a new one, first check the empty blocks list */
+ if (dclist_count(&slab->emptyblocks) > 0)
+ {
+ dlist_node *node = dclist_pop_head_node(&slab->emptyblocks);
- if (block == NULL)
- return NULL;
+ block = dlist_container(SlabBlock, node, node);
- block->nfree = slab->chunksPerBlock;
- block->firstFreeChunk = 0;
- block->slab = slab;
+ /*
+ * SlabFree() should have left this block in a valid state with
+ * all chunks free. Ensure that's the case.
+ */
+ Assert(block->nfree == slab->chunksPerBlock);
- /*
- * Put all the chunks on a freelist. Walk the chunks and point each
- * one to the next one.
- */
- for (idx = 0; idx < slab->chunksPerBlock; idx++)
- {
- chunk = SlabBlockGetChunk(slab, block, idx);
- *(int32 *) MemoryChunkGetPointer(chunk) = (idx + 1);
+ /* fetch the next chunk from this block */
+ chunk = SlabGetNextFreeChunk(slab, block);
}
+ else
+ {
+ block = (SlabBlock *) malloc(slab->blockSize);
- /*
- * And add it to the last freelist with all chunks empty.
- *
- * We know there are no blocks in the freelist, otherwise we wouldn't
- * need a new block.
- */
- Assert(dlist_is_empty(&slab->freelist[slab->chunksPerBlock]));
+ if (unlikely(block == NULL))
+ return NULL;
- dlist_push_head(&slab->freelist[slab->chunksPerBlock], &block->node);
+ block->slab = slab;
+ context->mem_allocated += slab->blockSize;
- slab->minFreeChunks = slab->chunksPerBlock;
- slab->nblocks += 1;
- context->mem_allocated += slab->blockSize;
- }
+ /* use the first chunk in the new block */
+ chunk = SlabBlockGetChunk(slab, block, 0);
- /* grab the block from the freelist (even the new block is there) */
- block = dlist_head_element(SlabBlock, node,
- &slab->freelist[slab->minFreeChunks]);
+ block->nfree = slab->chunksPerBlock - 1;
+ block->unused = SlabBlockGetChunk(slab, block, 1);
+ block->freehead = NULL;
+ block->nunused = slab->chunksPerBlock - 1;
+ }
- /* make sure we actually got a valid block, with matching nfree */
- Assert(block != NULL);
- Assert(slab->minFreeChunks == block->nfree);
- Assert(block->nfree > 0);
+ /* find the blocklist element for storing blocks with 1 used chunk */
+ blocklist_idx = SlabBlocklistIndex(slab, slab->chunksPerBlock - 1);
+ blocklist = &slab->blocklist[blocklist_idx];
- /* we know index of the first free chunk in the block */
- idx = block->firstFreeChunk;
+ /* this better be empty. We just added a block thinking it was */
+ Assert(dlist_is_empty(blocklist));
- /* make sure the chunk index is valid, and that it's marked as empty */
- Assert((idx >= 0) && (idx < slab->chunksPerBlock));
+ dlist_push_head(blocklist, &block->node);
- /* compute the chunk location block start (after the block header) */
- chunk = SlabBlockGetChunk(slab, block, idx);
+ slab->curBlocklistIndex = blocklist_idx;
+ }
+ else
+ {
+ dlist_head *blocklist = &slab->blocklist[slab->curBlocklistIndex];
+ int new_blocklist_idx;
- /*
- * Update the block nfree count, and also the minFreeChunks as we've
- * decreased nfree for a block with the minimum number of free chunks
- * (because that's how we chose the block).
- */
- block->nfree--;
- slab->minFreeChunks = block->nfree;
+ Assert(!dlist_is_empty(blocklist));
- /*
- * Remove the chunk from the freelist head. The index of the next free
- * chunk is stored in the chunk itself.
- */
- VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- block->firstFreeChunk = *(int32 *) MemoryChunkGetPointer(chunk);
+ /* grab the block from the blocklist */
+ block = dlist_head_element(SlabBlock, node, blocklist);
- Assert(block->firstFreeChunk >= 0);
- Assert(block->firstFreeChunk <= slab->chunksPerBlock);
+ /* make sure we actually got a valid block, with matching nfree */
+ Assert(block != NULL);
+ Assert(slab->curBlocklistIndex == SlabBlocklistIndex(slab, block->nfree));
+ Assert(block->nfree > 0);
- Assert((block->nfree != 0 &&
- block->firstFreeChunk < slab->chunksPerBlock) ||
- (block->nfree == 0 &&
- block->firstFreeChunk == slab->chunksPerBlock));
+ /* fetch the next chunk from this block */
+ chunk = SlabGetNextFreeChunk(slab, block);
- /* move the whole block to the right place in the freelist */
- dlist_delete(&block->node);
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
+ /* get the new blocklist index based on the new free chunk count */
+ new_blocklist_idx = SlabBlocklistIndex(slab, block->nfree);
- /*
- * And finally update minFreeChunks, i.e. the index to the block with the
- * lowest number of free chunks. We only need to do that when the block
- * got full (otherwise we know the current block is the right one). We'll
- * simply walk the freelist until we find a non-empty entry.
- */
- if (slab->minFreeChunks == 0)
- {
- for (idx = 1; idx <= slab->chunksPerBlock; idx++)
+ /*
+ * Handle the case where the blocklist index changes. This also deals
+ * with blocks becoming full as only full blocks go at index 0.
+ */
+ if (unlikely(slab->curBlocklistIndex != new_blocklist_idx))
{
- if (dlist_is_empty(&slab->freelist[idx]))
- continue;
+ dlist_delete_from(blocklist, &block->node);
+ dlist_push_head(&slab->blocklist[new_blocklist_idx], &block->node);
- /* found a non-empty freelist */
- slab->minFreeChunks = idx;
- break;
+ if (dlist_is_empty(blocklist))
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
}
}
- if (slab->minFreeChunks == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
+ /*
+ * Check that the chunk pointer is actually somewhere on the block and is
+ * aligned as expected.
+ */
+ Assert(chunk >= SlabBlockGetChunk(slab, block, 0));
+ Assert(chunk <= SlabBlockGetChunk(slab, block, slab->chunksPerBlock - 1));
+ Assert(SlabChunkMod(slab, block, chunk) == 0);
/* Prepare to initialize the chunk header. */
VALGRIND_MAKE_MEM_UNDEFINED(chunk, Slab_CHUNKHDRSZ);
@@ -453,8 +630,6 @@ SlabAlloc(MemoryContext context, Size size)
randomize_mem((char *) MemoryChunkGetPointer(chunk), size);
#endif
- Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
-
return MemoryChunkGetPointer(chunk);
}
@@ -468,7 +643,8 @@ SlabFree(void *pointer)
MemoryChunk *chunk = PointerGetMemoryChunk(pointer);
SlabBlock *block = MemoryChunkGetBlock(chunk);
SlabContext *slab;
- int idx;
+ int curBlocklistIdx;
+ int newBlocklistIdx;
/*
* For speed reasons we just Assert that the referenced block is good.
@@ -486,63 +662,82 @@ SlabFree(void *pointer)
slab->header.name, chunk);
#endif
- /* compute index of the chunk with respect to block start */
- idx = SlabChunkIndex(slab, block, chunk);
+ /* push this chunk onto the head of the block's free list */
+ *(MemoryChunk **) pointer = block->freehead;
+ block->freehead = chunk;
- /* add chunk to freelist, and update block nfree count */
- *(int32 *) pointer = block->firstFreeChunk;
- block->firstFreeChunk = idx;
block->nfree++;
Assert(block->nfree > 0);
Assert(block->nfree <= slab->chunksPerBlock);
#ifdef CLOBBER_FREED_MEMORY
- /* XXX don't wipe the int32 index, used for block-level freelist */
- wipe_mem((char *) pointer + sizeof(int32),
- slab->chunkSize - sizeof(int32));
+ /* don't wipe the free list MemoryChunk pointer stored in the chunk */
+ wipe_mem((char *) pointer + sizeof(MemoryChunk *),
+ slab->chunkSize - sizeof(MemoryChunk *));
#endif
- /* remove the block from a freelist */
- dlist_delete(&block->node);
+ curBlocklistIdx = SlabBlocklistIndex(slab, block->nfree - 1);
+ newBlocklistIdx = SlabBlocklistIndex(slab, block->nfree);
/*
- * See if we need to update the minFreeChunks field for the slab - we only
- * need to do that if there the block had that number of free chunks
- * before we freed one. In that case, we check if there still are blocks
- * in the original freelist and we either keep the current value (if there
- * still are blocks) or increment it by one (the new block is still the
- * one with minimum free chunks).
- *
- * The one exception is when the block will get completely free - in that
- * case we will free it, se we can't use it for minFreeChunks. It however
- * means there are no more blocks with free chunks.
+ * Check if the block needs to be moved to another element on the
+ * blocklist based on it now having 1 more free chunk.
*/
- if (slab->minFreeChunks == (block->nfree - 1))
+ if (unlikely(curBlocklistIdx != newBlocklistIdx))
{
- /* Have we removed the last chunk from the freelist? */
- if (dlist_is_empty(&slab->freelist[slab->minFreeChunks]))
+ /* do the move */
+ dlist_delete_from(&slab->blocklist[curBlocklistIdx], &block->node);
+ dlist_push_head(&slab->blocklist[newBlocklistIdx], &block->node);
+
+ /*
+ * It's possible that we've no blocks in the blocklist at the
+ * curBlocklistIndex position. When this happens we must find the
+ * next blocklist index which contains blocks. We can be certain
+ * we'll find a block as at least one must exist for the chunk we're
+ * currently freeing.
+ */
+ if (slab->curBlocklistIndex == curBlocklistIdx &&
+ dlist_is_empty(&slab->blocklist[curBlocklistIdx]))
{
- /* but if we made the block entirely free, we'll free it */
- if (block->nfree == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
- else
- slab->minFreeChunks++;
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
+ Assert(slab->curBlocklistIndex > 0);
}
}
- /* If the block is now completely empty, free it. */
- if (block->nfree == slab->chunksPerBlock)
+ /* Handle when a block becomes completely empty */
+ if (unlikely(block->nfree == slab->chunksPerBlock))
{
- free(block);
- slab->nblocks--;
- slab->header.mem_allocated -= slab->blockSize;
- }
- else
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
+ /* remove the block */
+ dlist_delete_from(&slab->blocklist[newBlocklistIdx], &block->node);
+
+ /*
+ * To avoid thrashing malloc/free, we keep a list of empty blocks that
+ * we can reuse again instead of having to malloc a new one.
+ */
+ if (dclist_count(&slab->emptyblocks) < SLAB_MAXIMUM_EMPTY_BLOCKS)
+ dclist_push_head(&slab->emptyblocks, &block->node);
+ else
+ {
+ /*
+ * When we have enough empty blocks stored already, we actually
+ * free the block.
+ */
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ slab->header.mem_allocated -= slab->blockSize;
+ }
- Assert(slab->nblocks >= 0);
- Assert(slab->nblocks * slab->blockSize == slab->header.mem_allocated);
+ /*
+ * Check if we need to reset the blocklist index. This is required
+ * when the blocklist this block is on has become completely empty.
+ */
+ if (slab->curBlocklistIndex == newBlocklistIdx &&
+ dlist_is_empty(&slab->blocklist[newBlocklistIdx]))
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
+ }
}
/*
@@ -617,16 +812,14 @@ SlabGetChunkSpace(void *pointer)
/*
* SlabIsEmpty
- * Is an Slab empty of any allocated space?
+ * Is the slab empty of any allocated space?
*/
bool
SlabIsEmpty(MemoryContext context)
{
- SlabContext *slab = (SlabContext *) context;
-
- Assert(SlabIsValid(slab));
+ Assert(SlabIsValid((SlabContext *) context));
- return (slab->nblocks == 0);
+ return (context->mem_allocated == 0);
}
/*
@@ -654,13 +847,16 @@ SlabStats(MemoryContext context,
Assert(SlabIsValid(slab));
/* Include context header in totalspace */
- totalspace = slab->headerSize;
+ totalspace = Slab_CONTEXT_HDRSZ(slab->chunksPerBlock);
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /* Add the space consumed by blocks in the emptyblocks list */
+ totalspace += dclist_count(&slab->emptyblocks) * slab->blockSize;
+
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
{
dlist_iter iter;
- dlist_foreach(iter, &slab->freelist[i])
+ dlist_foreach(iter, &slab->blocklist[i])
{
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
@@ -675,10 +871,11 @@ SlabStats(MemoryContext context,
{
char stats_string[200];
+ /* XXX should we include free chunks on empty blocks? */
snprintf(stats_string, sizeof(stats_string),
- "%zu total in %zu blocks; %zu free (%zu chunks); %zu used",
- totalspace, nblocks, freespace, freechunks,
- totalspace - freespace);
+ "%zu total in %zu blocks; %u empty blocks; %zu free (%zu chunks); %zu used",
+ totalspace, nblocks, dclist_count(&slab->emptyblocks),
+ freespace, freechunks, totalspace - freespace);
printfunc(context, passthru, stats_string, print_to_stderr);
}
@@ -696,7 +893,7 @@ SlabStats(MemoryContext context,
/*
* SlabCheck
- * Walk through chunks and check consistency of memory.
+ * Walk through all blocks looking for inconsistencies.
*
* NOTE: report errors as WARNING, *not* ERROR or FATAL. Otherwise you'll
* find yourself in an infinite loop when trouble occurs, because this
@@ -707,67 +904,112 @@ SlabCheck(MemoryContext context)
{
SlabContext *slab = (SlabContext *) context;
int i;
+ int nblocks = 0;
const char *name = slab->header.name;
+ dlist_iter iter;
Assert(SlabIsValid(slab));
Assert(slab->chunksPerBlock > 0);
- /* walk all the freelists */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /*
+ * Have a look at the empty blocks. These should have all their chunks
+ * marked as free. Ensure that's the case.
+ */
+ dclist_foreach(iter, &slab->emptyblocks)
+ {
+ SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+
+ if (block->nfree != slab->chunksPerBlock)
+ elog(WARNING, "problem in slab %s: empty block %p should have %d free chunks but has %d chunks free",
+ name, block, slab->chunksPerBlock, block->nfree);
+ }
+
+ /* walk the non-empty block lists */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
{
int j,
nfree;
- dlist_iter iter;
- /* walk all blocks on this freelist */
- dlist_foreach(iter, &slab->freelist[i])
+ /* walk all blocks on this blocklist */
+ dlist_foreach(iter, &slab->blocklist[i])
{
- int idx;
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+ MemoryChunk *cur_chunk;
/*
* Make sure the number of free chunks (in the block header)
- * matches position in the freelist.
+ * matches the position in the blocklist.
*/
- if (block->nfree != i)
- elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match freelist %d",
- name, block->nfree, block, i);
+ if (SlabBlocklistIndex(slab, block->nfree) != i)
+ elog(WARNING, "problem in slab %s: block %p is on blocklist %d but should be on blocklist %d",
+ name, block, i, SlabBlocklistIndex(slab, block->nfree));
+
+ /* make sure the block is not empty */
+ if (block->nfree <= 0)
+ elog(WARNING, "problem in slab %s: empty block %p incorrectly stored on blocklist element %d",
+ name, block, i);
/* make sure the slab pointer correctly points to this context */
if (block->slab != slab)
elog(WARNING, "problem in slab %s: bogus slab link in block %p",
name, block);
- /* reset the bitmap of free chunks for this block */
- memset(slab->freechunks, 0, (slab->chunksPerBlock * sizeof(bool)));
- idx = block->firstFreeChunk;
+ /* reset the array of free chunks for this block */
+ memset(slab->isChunkFree, 0, (slab->chunksPerBlock * sizeof(bool)));
+ nfree = 0;
+ /* walk through the block's free list chunks */
+ cur_chunk = block->freehead;
+ while (cur_chunk != NULL)
+ {
+ int chunkidx = SlabChunkIndex(slab, block, cur_chunk);
+
+ /*
+ * Ensure the free list link points to something on the block
+ * at an address aligned according to the full chunk size.
+ */
+ if (cur_chunk < SlabBlockGetChunk(slab, block, 0) ||
+ cur_chunk > SlabBlockGetChunk(slab, block, slab->chunksPerBlock - 1) ||
+ SlabChunkMod(slab, block, cur_chunk) != 0)
+ elog(WARNING, "problem in slab %s: bogus free list link %p in block %p",
+ name, cur_chunk, block);
+
+ /* count the chunk and mark it free on the free chunk array */
+ nfree++;
+ slab->isChunkFree[chunkidx] = true;
+
+ /* read pointer of the next free chunk */
+ VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(cur_chunk), sizeof(MemoryChunk *));
+ cur_chunk = *(MemoryChunk **) SlabChunkGetPointer(cur_chunk);
+ }
+
+ /* check that the unused pointer matches what nunused claims */
+ if (SlabBlockGetChunk(slab, block, slab->chunksPerBlock - block->nunused) !=
+ block->unused)
+ elog(WARNING, "problem in slab %s: mismatch detected between nunused chunks and unused pointer in block %p",
+ name, block);
/*
- * Now walk through the chunks, count the free ones and also
- * perform some additional checks for the used ones. As the chunk
- * freelist is stored within the chunks themselves, we have to
- * walk through the chunks and construct our own bitmap.
+ * count the remaining free chunks that have yet to make it onto
+ * the block's free list.
*/
-
- nfree = 0;
- while (idx < slab->chunksPerBlock)
+ cur_chunk = block->unused;
+ for (j = 0; j < block->nunused; j++)
{
- MemoryChunk *chunk;
+ int chunkidx = SlabChunkIndex(slab, block, cur_chunk);
+
- /* count the chunk as free, add it to the bitmap */
+ /* count the chunk as free and mark it as so in the array */
nfree++;
- slab->freechunks[idx] = true;
+ if (chunkidx < slab->chunksPerBlock)
+ slab->isChunkFree[chunkidx] = true;
- /* read index of the next free chunk */
- chunk = SlabBlockGetChunk(slab, block, idx);
- VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- idx = *(int32 *) MemoryChunkGetPointer(chunk);
+ /* move forward 1 chunk */
+ cur_chunk = (MemoryChunk *) (((char *) cur_chunk) + slab->fullChunkSize);
}
for (j = 0; j < slab->chunksPerBlock; j++)
{
- /* non-zero bit in the bitmap means chunk the chunk is used */
- if (!slab->freechunks[j])
+ if (!slab->isChunkFree[j])
{
MemoryChunk *chunk = SlabBlockGetChunk(slab, block, j);
SlabBlock *chunkblock = (SlabBlock *) MemoryChunkGetBlock(chunk);
@@ -793,12 +1035,17 @@ SlabCheck(MemoryContext context)
* in the block header).
*/
if (nfree != block->nfree)
- elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match bitmap %d",
- name, block->nfree, block, nfree);
+ elog(WARNING, "problem in slab %s: nfree in block %p is %d but %d chunk were found as free",
+ name, block, block->nfree, nfree);
+
+ nblocks++;
}
}
- Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
+ /* the stored empty blocks are tracked in mem_allocated too */
+ nblocks += dclist_count(&slab->emptyblocks);
+
+ Assert(nblocks * slab->blockSize == context->mem_allocated);
}
#endif /* MEMORY_CONTEXT_CHECKING */
diff --git a/contrib/alloc_bench/.gitignore b/contrib/alloc_bench/.gitignore
new file mode 100644
index 0000000000..5dcb3ff972
--- /dev/null
+++ b/contrib/alloc_bench/.gitignore
@@ -0,0 +1,4 @@
+# Generated subdirectories
+/log/
+/results/
+/tmp_check/
diff --git a/contrib/alloc_bench/Makefile b/contrib/alloc_bench/Makefile
new file mode 100644
index 0000000000..c672b29bdb
--- /dev/null
+++ b/contrib/alloc_bench/Makefile
@@ -0,0 +1,21 @@
+# contrib/alloc_bench/Makefile
+
+MODULE_big = alloc_bench
+OBJS = alloc_bench.o
+
+EXTENSION = alloc_bench
+DATA = alloc_bench--1.0.sql
+PGFILEDESC = "alloc_bench - memory context benchmarking functions"
+
+REGRESS = alloc_bench
+
+ifdef USE_PGXS
+PG_CONFIG = pg_config
+PGXS := $(shell $(PG_CONFIG) --pgxs)
+include $(PGXS)
+else
+subdir = contrib/alloc_bench
+top_builddir = ../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+endif
diff --git a/contrib/alloc_bench/alloc_bench--1.0.sql b/contrib/alloc_bench/alloc_bench--1.0.sql
new file mode 100644
index 0000000000..2ca604a682
--- /dev/null
+++ b/contrib/alloc_bench/alloc_bench--1.0.sql
@@ -0,0 +1,9 @@
+/* alloc_bench--1.0.sql */
+
+-- complain if script is sourced in psql, rather than via CREATE EXTENSION
+\echo Use "CREATE EXTENSION alloc_bench" to load this file. \quit
+
+CREATE FUNCTION alloc_bench(context_type text, pattern text, nchunks bigint, block_size bigint, chunk_size bigint, nloops int, random_segments int, out mem_allocated bigint, out alloc_us bigint, out free_us bigint)
+AS 'MODULE_PATHNAME', 'alloc_bench'
+LANGUAGE C VOLATILE STRICT;
+
diff --git a/contrib/alloc_bench/alloc_bench.c b/contrib/alloc_bench/alloc_bench.c
new file mode 100644
index 0000000000..f76501c053
--- /dev/null
+++ b/contrib/alloc_bench/alloc_bench.c
@@ -0,0 +1,199 @@
+/*-------------------------------------------------------------------------
+ *
+ * alloc_bench.c
+ *
+ * helper functions to benchmark memory contexts with different workloads
+ *-------------------------------------------------------------------------
+ */
+#include "postgres.h"
+
+#include <sys/time.h>
+
+#include "funcapi.h"
+#include "miscadmin.h"
+#include "utils/builtins.h"
+
+PG_MODULE_MAGIC;
+
+PG_FUNCTION_INFO_V1(alloc_bench);
+
+typedef enum AllocationPattern {
+ FIFO,
+ LIFO
+} AllocationPattern;
+
+static int random_segment_size;
+static int random_fifo = 0;
+
+static int
+int64_random_compare(const void *a, const void *b)
+{
+ const int64 ia = *(const int64 *) a;
+ const int64 ib = *(const int64 *) b;
+ int rss = random_segment_size;
+
+ if ((ia / rss * rss) + (random() % rss) > ib)
+ return random_fifo == 0 ? -1 : 1;
+ else
+ return random_fifo == 0 ? 1 : -1;
+}
+
+Datum
+alloc_bench(PG_FUNCTION_ARGS)
+{
+ MemoryContext cxt,
+ oldcxt;
+ void **chunks;
+ int64 *indexes;
+ int64 j;
+ char *context_type;
+ text *context_type_text = PG_GETARG_TEXT_PP(0);
+ char *pattern_str;
+ text *pattern_text = PG_GETARG_TEXT_PP(1);
+ int64 nchunks = PG_GETARG_INT64(2);
+ int64 blockSize = PG_GETARG_INT64(3);
+ int64 chunkSize = PG_GETARG_INT64(4);
+
+ int nloops = PG_GETARG_INT32(5);
+ int random_segments = PG_GETARG_INT32(6);
+
+ struct timeval start_time,
+ end_time;
+ int64 alloc_time = 0,
+ free_time = 0;
+ int64 mem_allocated;
+
+ TupleDesc tupdesc;
+ Datum result;
+ HeapTuple tuple;
+ Datum values[9];
+ bool nulls[9];
+ AllocationPattern pattern;
+
+ context_type = text_to_cstring(context_type_text);
+
+ if (strcmp(context_type, "generation") == 0)
+ cxt = GenerationContextCreate(CurrentMemoryContext,
+ "alloc_bench",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "aset") == 0)
+ cxt = AllocSetContextCreate(CurrentMemoryContext,
+ "alloc_bench",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "slab") == 0)
+ cxt = SlabContextCreate(CurrentMemoryContext,
+ "alloc_bench",
+ blockSize,
+ chunkSize);
+ else
+ elog(ERROR, "%s is not a valid context type. context_type must be \"generation\", \"aset\" or \"slab\"",
+ context_type);
+
+ pattern_str = text_to_cstring(pattern_text);
+
+ if (strcmp(pattern_str, "fifo") == 0)
+ pattern = FIFO;
+ else if (strcmp(pattern_str, "lifo") == 0)
+ pattern = LIFO;
+ else
+ elog(ERROR, "%s is not a valid allocation pattern. Must be \"fifo\" or \"lifo\"",
+ pattern_str);
+
+ chunks = (void **) palloc(nchunks * sizeof(void *));
+ indexes = (int64 *) palloc(nchunks * sizeof(int64));
+
+ /* set the indexes so according to the allocation pattern */
+ switch (pattern)
+ {
+ case FIFO:
+ for (int64 i = 0; i < nchunks; i++)
+ indexes[i] = i;
+ break;
+ case LIFO:
+ for (int64 i = 0; i < nchunks; i++)
+ indexes[i] = nchunks - i - 1;
+ break;
+ }
+
+ /*
+ * When a non-zero random_segments is specified, we randomize the
+ * order of the pfree's so they're random within their own segment.
+ * This means we don't entirely randomize the order, but if there are
+ * say, 10 segments and 60 chunks, we only randomize the first 6
+ * chunks then the next 6. Specifiying a lower number of
+ * random_segments means the FIFO or LIFO pattern becomes more random.
+ */
+ if (random_segments > 0)
+ {
+ random_segment_size = nchunks / random_segments;
+ random_fifo = (pattern == FIFO);
+ qsort(indexes, nchunks, sizeof(int64), int64_random_compare);
+ }
+
+ mem_allocated = 0;
+
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ /* do the requested number of pfree/palloc loops */
+ for (j = 0; j < nloops; j++)
+ {
+ CHECK_FOR_INTERRUPTS();
+
+ gettimeofday(&start_time, NULL);
+
+ /*
+ * We do palloc0 instead of palloc to simulate touching the
+ * memory. This will load cachelines, which is likely more
+ * realistic than just allocating memory and not doing
+ * anything with it.
+ */
+ for (int64 i = 0; i < nchunks; i++)
+ chunks[i] = palloc0(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ gettimeofday(&start_time, NULL);
+
+ for (int64 i = 0; i < nchunks; i++)
+ {
+ char *ptr = (char *) chunks[indexes[i]];
+
+ /*
+ * Free the chunk, but first touch the first cacheline
+ * of the chunk to simulate that we've just done
+ * something real with this memory.
+ */
+ if (ptr[0] == '\0')
+ pfree(ptr);
+ }
+
+ gettimeofday(&end_time, NULL);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ mem_allocated = Max(mem_allocated, MemoryContextMemAllocated(cxt, true));
+ }
+
+ MemoryContextSwitchTo(oldcxt);
+ MemoryContextDelete(cxt);
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+
+ values[0] = Int64GetDatum(mem_allocated);
+ values[1] = Int64GetDatum(alloc_time);
+ values[2] = Int64GetDatum(free_time);
+
+ memset(nulls, 0, sizeof(nulls));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ result = HeapTupleGetDatum(tuple);
+
+ PG_RETURN_DATUM(result);
+}
+
diff --git a/contrib/alloc_bench/alloc_bench.control b/contrib/alloc_bench/alloc_bench.control
new file mode 100644
index 0000000000..597480ed5e
--- /dev/null
+++ b/contrib/alloc_bench/alloc_bench.control
@@ -0,0 +1,4 @@
+# alloc_bench extension
+comment = 'functions for benchmarking memory context'
+default_version = '1.0'
+module_pathname = '$libdir/alloc_bench'
diff --git a/contrib/alloc_bench/bench.sql b/contrib/alloc_bench/bench.sql
new file mode 100644
index 0000000000..d476a386c2
--- /dev/null
+++ b/contrib/alloc_bench/bench.sql
@@ -0,0 +1,40 @@
+CREATE EXTENSION slab_bench;
+
+\o fifo-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+\o lifo-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+\o random-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+
+\o fifo-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+\o lifo-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+\o random-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+
+\o fifo-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+\o lifo-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+\o random-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+
+\o fifo-cycle.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000, block_size, chunk_size, 10000, 1000, 1000) x;
+
+\o lifo-cycle.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000, block_size, chunk_size, 10000, 1000, 1000) x;
+
+\o random-cycle.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000, block_size, chunk_size, 10000, 1000, 1000) x;
Attachments:
[text/plain] slab_changes_v3.patch (39.9K, ../../CAApHDvq6eUdLJxAUdSmukGiiTQNT79cNtntL=3FE52T_AP3XDQ@mail.gmail.com/2-slab_changes_v3.patch)
download | inline diff:
diff --git a/src/backend/utils/mmgr/slab.c b/src/backend/utils/mmgr/slab.c
index c2f9bb6ad3..b6f8ed28a7 100644
--- a/src/backend/utils/mmgr/slab.c
+++ b/src/backend/utils/mmgr/slab.c
@@ -4,7 +4,8 @@
* SLAB allocator definitions.
*
* SLAB is a MemoryContext implementation designed for cases where large
- * numbers of equally-sized objects are allocated (and freed).
+ * numbers of equally-sized objects can be allocated and freed efficiently
+ * with minimal memory wastage and fragmentation.
*
*
* Portions Copyright (c) 2017-2022, PostgreSQL Global Development Group
@@ -16,36 +17,51 @@
* NOTE:
* The constant allocation size allows significant simplification and various
* optimizations over more general purpose allocators. The blocks are carved
- * into chunks of exactly the right size (plus alignment), not wasting any
- * memory.
+ * into chunks of exactly the right size, wasting only the space required to
+ * MAXALIGN the allocated chunks.
*
- * The information about free chunks is maintained both at the block level and
- * global (context) level. This is possible as the chunk size (and thus also
- * the number of chunks per block) is fixed.
+ * Slab can also help reduce memory fragmentation in cases where longer-lived
+ * chunks remain stored on blocks while most of the other chunks have already
+ * been pfree'd. We give priority to putting new allocations into the
+ * "fullest" block. This help avoid having too many sparsely used blocks
+ * around and allows blocks to more easily become completely unused which
+ * allows them to be eventually free'd.
*
- * On each block, free chunks are tracked in a simple linked list. Contents
- * of free chunks is replaced with an index of the next free chunk, forming
- * a very simple linked list. Each block also contains a counter of free
- * chunks. Combined with the local block-level freelist, it makes it trivial
- * to eventually free the whole block.
+ * We identify the "fullest" block to put new allocations on by using a block
+ * from the lowest populated element of the context's "blocklist" array.
+ * This is an array of dlists containing blocks which we partition by the
+ * number of free chunks which block has. Blocks with fewer free chunks are
+ * stored in a lower indexed dlist array slot. Full blocks go on the 0th
+ * element of the blocklist array. So that we don't have to have too many
+ * elements in the array, each dlist in the array is responsible for a range
+ * of free chunks. When a chunk is palloc'd or pfree'd we may need to move
+ * the block onto another dlist if the number of free chunks crosses the
+ * range boundary that the current list is responsible for. Having just a
+ * few blocklist elements reduces the number of times we must move the block
+ * onto another dlist element.
*
- * At the context level, we use 'freelist' to track blocks ordered by number
- * of free chunks, starting with blocks having a single allocated chunk, and
- * with completely full blocks on the tail.
+ * We keep track of free chunks within each block by using a block-level free
+ * list. We consult this list when we allocate a new chunk in the block.
+ * The free list is a linked list, the head of which is pointed to with
+ * SlabBlock's freehead field. Each subsequent list item is stored in the
+ * free chunk's memory. We ensure chunks are large enough to store this
+ * address.
*
- * This also allows various optimizations - for example when searching for
- * free chunk, the allocator reuses space from the fullest blocks first, in
- * the hope that some of the less full blocks will get completely empty (and
- * returned back to the OS).
- *
- * For each block, we maintain pointer to the first free chunk - this is quite
- * cheap and allows us to skip all the preceding used chunks, eliminating
- * a significant number of lookups in many common usage patterns. In the worst
- * case this performs as if the pointer was not maintained.
- *
- * We cache the freelist index for the blocks with the fewest free chunks
- * (minFreeChunks), so that we don't have to search the freelist on every
- * SlabAlloc() call, which is quite expensive.
+ * When we allocate a new block, technically all chunks are free, however, to
+ * avoid having to write out the entire block to set the linked list for the
+ * free chunks for every chunk in the block, we instead store a pointer to
+ * the next "unused" chunk on the block and keep track of how many of these
+ * unused chunks there are. When a new block is malloc'd, all chunks are
+ * unused. The unused pointer starts with the first chunk on the block and
+ * as chunks are allocated, the unused pointer is incremented. As chunks are
+ * pfree'd, the unused pointer never goes backwards. The unused pointer can
+ * be thought of as a high watermark for the maximum number of chunks in the
+ * block which have been in use concurrently. When a chunk is pfree'd the
+ * chunk is put onto the head of the free list and the unused pointer is not
+ * changed. We only consume more unused chunks if we run out of free chunks
+ * on the free list. This method effectively gives priority to using
+ * previously used chunks over previously unused chunks, which should perform
+ * better due to CPU caching effects.
*
*-------------------------------------------------------------------------
*/
@@ -60,6 +76,27 @@
#define Slab_BLOCKHDRSZ MAXALIGN(sizeof(SlabBlock))
+#ifdef MEMORY_CONTEXT_CHECKING
+/*
+ * Size of the memory required to store the SlabContext.
+ * MEMORY_CONTEXT_CHECKING builds need some extra memory for the isChunkFree
+ * array.
+ */
+#define Slab_CONTEXT_HDRSZ(chnksperblk) \
+ (sizeof(SlabContext) + ((chnksperblk) * sizeof(bool)))
+#else
+#define Slab_CONTEXT_HDRSZ(chnksperblk) sizeof(SlabContext)
+#endif
+
+/*
+ * The number of partitions to divide the blocklist into based their number of
+ * free chunks. There must be at least 2.
+ */
+#define SLAB_BLOCKLIST_COUNT 3
+
+/* The maximum number of completely empty blocks to keep around for reuse. */
+#define SLAB_MAXIMUM_EMPTY_BLOCKS 10
+
/*
* SlabContext is a specialized implementation of MemoryContext.
*/
@@ -67,64 +104,206 @@ typedef struct SlabContext
{
MemoryContextData header; /* Standard memory-context fields */
/* Allocation parameters for this context: */
- Size chunkSize; /* chunk size */
- Size fullChunkSize; /* chunk size including header and alignment */
- Size blockSize; /* block size */
- Size headerSize; /* allocated size of context header */
- int chunksPerBlock; /* number of chunks per block */
- int minFreeChunks; /* min number of free chunks in any block */
- int nblocks; /* number of blocks allocated */
+ Size chunkSize; /* the requested (non-aligned) chunk size */
+ Size fullChunkSize; /* chunk size with chunk header and alignment */
+ Size blockSize; /* the size to make each block of chunks */
+ int32 chunksPerBlock; /* number of chunks that fit in 1 block */
+ int32 curBlocklistIndex; /* index into the blocklist[] element
+ * containing the fullest, blocks */
#ifdef MEMORY_CONTEXT_CHECKING
- bool *freechunks; /* bitmap of free chunks in a block */
+ bool *isChunkFree; /* array to mark free chunks in a block during
+ * SlabCheck */
#endif
- /* blocks with free space, grouped by number of free chunks: */
- dlist_head freelist[FLEXIBLE_ARRAY_MEMBER];
+
+ int32 blocklist_shift; /* number of bits to shift the nfree count
+ * by to get the index into blocklist[] */
+ dclist_head emptyblocks; /* empty blocks to use up first instead of
+ * mallocing new blocks */
+
+ /*
+ * Blocks with free space, grouped by the number of free chunks they
+ * contain. Completely full blocks are stored in the 0th element.
+ * Completely empty blocks are stored in emptyblocks or free'd if we have
+ * enough empty blocks already.
+ */
+ dlist_head blocklist[SLAB_BLOCKLIST_COUNT];
} SlabContext;
/*
* SlabBlock
- * Structure of a single block in SLAB allocator.
+ * Structure of a single slab block.
*
- * node: doubly-linked list of blocks in global freelist
- * nfree: number of free chunks in this block
- * firstFreeChunk: index of the first free chunk
+ * slab: pointer back to the owning MemoryContext
+ * nfree: number of chunks on the block which are unallocated
+ * nunused: number of chunks on the block unallocated and not on the block's
+ * freelist.
+ * freehead: linked-list header storing a pointer to the first free chunk on
+ * the block. Subsequent pointers are stored in the chunk's memory. NULL
+ * indicates the end of the list.
+ * unused: pointer to the next chunk which has yet to be used.
+ * node: doubly-linked list node for the context's blocklist
*/
typedef struct SlabBlock
{
- dlist_node node; /* doubly-linked list */
- int nfree; /* number of free chunks */
- int firstFreeChunk; /* index of the first free chunk in the block */
SlabContext *slab; /* owning context */
+ int32 nfree; /* number of chunks on free + unused chunks */
+ int32 nunused; /* number of unused chunks */
+ MemoryChunk *freehead; /* pointer to the first free chunk */
+ MemoryChunk *unused; /* pointer to the next unused chunk */
+ dlist_node node; /* doubly-linked list for blocklist[] */
} SlabBlock;
#define Slab_CHUNKHDRSZ sizeof(MemoryChunk)
-#define SlabPointerGetChunk(ptr) \
- ((MemoryChunk *)(((char *)(ptr)) - sizeof(MemoryChunk)))
#define SlabChunkGetPointer(chk) \
- ((void *)(((char *)(chk)) + sizeof(MemoryChunk)))
-#define SlabBlockGetChunk(slab, block, idx) \
+ ((void *) (((char *) (chk)) + sizeof(MemoryChunk)))
+
+/*
+ * SlabBlockGetChunk
+ * Obtain a pointer to the nth (0-based) chunk in the block
+ */
+#define SlabBlockGetChunk(slab, block, n) \
((MemoryChunk *) ((char *) (block) + Slab_BLOCKHDRSZ \
- + (idx * slab->fullChunkSize)))
-#define SlabBlockStart(block) \
- ((char *) block + Slab_BLOCKHDRSZ)
+ + ((n) * (slab)->fullChunkSize)))
+
+#if defined(MEMORY_CONTEXT_CHECKING) || defined(USE_ASSERT_CHECKING)
+
+/*
+ * SlabChunkIndex
+ * Get the 0-based index of how many chunks into the block the given
+ * chunk is.
+*/
#define SlabChunkIndex(slab, block, chunk) \
- (((char *) chunk - SlabBlockStart(block)) / slab->fullChunkSize)
+ (((char *) (chunk) - (char *) SlabBlockGetChunk(slab, block, 0)) / \
+ (slab)->fullChunkSize)
+
+/*
+ * SlabChunkMod
+ * A MemoryChunk should always be at an address which is a multiple of
+ * fullChunkSize starting from the 0th chunk position. This will return
+ * non-zero if it's not.
+ */
+#define SlabChunkMod(slab, block, chunk) \
+ (((char *) (chunk) - (char *) SlabBlockGetChunk(slab, block, 0)) % \
+ (slab)->fullChunkSize)
+
+#endif
/*
* SlabIsValid
- * True iff set is valid slab allocation set.
+ * True iff set is a valid slab allocation set.
*/
-#define SlabIsValid(set) \
- (PointerIsValid(set) && IsA(set, SlabContext))
+#define SlabIsValid(set) (PointerIsValid(set) && IsA(set, SlabContext))
/*
* SlabBlockIsValid
- * True iff block is valid block of slab allocation set.
+ * True iff block is a valid block of slab allocation set.
*/
#define SlabBlockIsValid(block) \
(PointerIsValid(block) && SlabIsValid((block)->slab))
+/*
+ * SlabBlocklistIndex
+ * Determine the blocklist index that a block should be in for the given
+ * number of free chunks.
+ */
+static inline int32
+SlabBlocklistIndex(SlabContext *slab, int nfree)
+{
+ int32 index;
+ int32 blocklist_shift = slab->blocklist_shift;
+
+ Assert(nfree >= 0 && nfree <= slab->chunksPerBlock);
+
+ /*
+ * Determine the blocklist index based on the number of free chunks. We
+ * must ensure that 0 free chunks is dedicated to index 0. Everything else
+ * must be >= 1 and < SLAB_BLOCKLIST_COUNT.
+ *
+ * To make this as efficient as possible, we exploit some two's complement
+ * arithmetic where we reverse the sign before bit shifting. This results
+ * in an nfree of 0 using index 0 and anything non-zero staying non-zero.
+ * This is exploiting 0 and -0 being the same in two's complement. When
+ * we're done, we just need to flip the sign back over again for a
+ * positive index.
+ */
+ index = -((-nfree) >> blocklist_shift);
+
+ if (nfree == 0)
+ Assert(index == 0);
+ else
+ Assert(index >= 1 && index < SLAB_BLOCKLIST_COUNT);
+
+ return index;
+}
+
+/*
+ * SlabFindNextBlockListIndex
+ * Search blocklist for blocks which have free chunks and return the
+ * index of the blocklist found containing at least 1 block with free
+ * chunks. If no block can be found we return 0.
+ *
+ * Note: We give priority to fuller blocks so that these are filled before
+ * emptier blocks. This is done to increase the chances that mostly-empty
+ * blocks will eventually become completely empty so they can be free'd.
+ */
+static int32
+SlabFindNextBlockListIndex(SlabContext *slab)
+{
+ /* start at 1 as blocklist[0] is for full blocks. */
+ for (int i = 1; i < SLAB_BLOCKLIST_COUNT; i++)
+ {
+ /* return the first found non-empty index */
+ if (!dlist_is_empty(&slab->blocklist[i]))
+ return i;
+ }
+
+ /* no blocks with free space */
+ return 0;
+}
+
+/*
+ * SlabGetNextFreeChunk
+ * Return the next free chunk in block and update the block to account
+ * for the returned chunk now being used.
+ */
+static inline MemoryChunk *
+SlabGetNextFreeChunk(SlabContext *slab, SlabBlock *block)
+{
+ MemoryChunk *chunk;
+
+ Assert(block->nfree > 0);
+
+ if (block->freehead != NULL)
+ {
+ chunk = block->freehead;
+
+ /*
+ * Pop the chunk from the linked list of free chunks. The pointer to
+ * the next free chunk is stored in the chunk itself.
+ */
+ VALGRIND_MAKE_MEM_DEFINED(SlabChunkGetPointer(chunk), sizeof(MemoryChunk *));
+ block->freehead = *(MemoryChunk **) SlabChunkGetPointer(chunk);
+
+ /* check nothing stomped on the free chunk's memory */
+ Assert(block->freehead == NULL ||
+ (block->freehead >= SlabBlockGetChunk(slab, block, 0) &&
+ block->freehead <= SlabBlockGetChunk(slab, block, slab->chunksPerBlock - 1) &&
+ SlabChunkMod(slab, block, block->freehead) == 0));
+ }
+ else
+ {
+ Assert(block->nunused > 0);
+
+ chunk = block->unused;
+ block->unused = (MemoryChunk *) (((char *)block->unused) + slab->fullChunkSize);
+ block->nunused--;
+ }
+
+ block->nfree--;
+
+ return chunk;
+}
/*
* SlabContextCreate
@@ -145,8 +324,6 @@ SlabContextCreate(MemoryContext parent,
{
int chunksPerBlock;
Size fullChunkSize;
- Size freelistSize;
- Size headerSize;
SlabContext *slab;
int i;
@@ -155,11 +332,14 @@ SlabContextCreate(MemoryContext parent,
"sizeof(MemoryChunk) is not maxaligned");
Assert(MAXALIGN(chunkSize) <= MEMORYCHUNK_MAX_VALUE);
- /* Make sure the linked list node fits inside a freed chunk */
- if (chunkSize < sizeof(int))
- chunkSize = sizeof(int);
+ /*
+ * Ensure there's enough space to store the pointer to the next free chunk
+ * in the memory of the (otherwise) unused allocation.
+ */
+ if (chunkSize < sizeof(MemoryChunk *))
+ chunkSize = sizeof(MemoryChunk *);
- /* chunk, including SLAB header (both addresses nicely aligned) */
+ /* length of the maxaligned chunk including the chunk header */
#ifdef MEMORY_CONTEXT_CHECKING
/* ensure there's always space for the sentinel byte */
fullChunkSize = Slab_CHUNKHDRSZ + MAXALIGN(chunkSize + 1);
@@ -167,36 +347,17 @@ SlabContextCreate(MemoryContext parent,
fullChunkSize = Slab_CHUNKHDRSZ + MAXALIGN(chunkSize);
#endif
- /* Make sure the block can store at least one chunk. */
- if (blockSize < fullChunkSize + Slab_BLOCKHDRSZ)
- elog(ERROR, "block size %zu for slab is too small for %zu chunks",
- blockSize, chunkSize);
-
- /* Compute maximum number of chunks per block */
+ /* compute the number of chunks that will fit on each block */
chunksPerBlock = (blockSize - Slab_BLOCKHDRSZ) / fullChunkSize;
- /* The freelist starts with 0, ends with chunksPerBlock. */
- freelistSize = sizeof(dlist_head) * (chunksPerBlock + 1);
-
- /*
- * Allocate the context header. Unlike aset.c, we never try to combine
- * this with the first regular block; not worth the extra complication.
- */
+ /* Make sure the block can store at least one chunk. */
+ if (chunksPerBlock == 0)
+ elog(ERROR, "block size %zu for slab is too small for %zu-byte chunks",
+ blockSize, chunkSize);
- /* Size of the memory context header */
- headerSize = offsetof(SlabContext, freelist) + freelistSize;
-#ifdef MEMORY_CONTEXT_CHECKING
-
- /*
- * With memory checking, we need to allocate extra space for the bitmap of
- * free chunks. The bitmap is an array of bools, so we don't need to worry
- * about alignment.
- */
- headerSize += chunksPerBlock * sizeof(bool);
-#endif
- slab = (SlabContext *) malloc(headerSize);
+ slab = (SlabContext *) malloc(Slab_CONTEXT_HDRSZ(chunksPerBlock));
if (slab == NULL)
{
MemoryContextStats(TopMemoryContext);
@@ -216,19 +377,33 @@ SlabContextCreate(MemoryContext parent,
slab->chunkSize = chunkSize;
slab->fullChunkSize = fullChunkSize;
slab->blockSize = blockSize;
- slab->headerSize = headerSize;
slab->chunksPerBlock = chunksPerBlock;
- slab->minFreeChunks = 0;
- slab->nblocks = 0;
+ slab->curBlocklistIndex = 0;
- /* initialize the freelist slots */
- for (i = 0; i < (slab->chunksPerBlock + 1); i++)
- dlist_init(&slab->freelist[i]);
+ /*
+ * Compute a shift that guarantees that shifting chunksPerBlock with it is
+ * < SLAB_BLOCKLIST_COUNT - 1. The reason that we subtract 1 from
+ * SLAB_BLOCKLIST_COUNT in this calculation is that we reserve the 0th
+ * blocklist element for blocks which have no free chunks.
+ *
+ * We calculate the number of bits to shift by rather than a divisor to
+ * divide by as performing division each time we need to find the
+ * blocklist index would be much slower.
+ */
+ slab->blocklist_shift = 0;
+ while ((slab->chunksPerBlock >> slab->blocklist_shift) >= (SLAB_BLOCKLIST_COUNT - 1))
+ slab->blocklist_shift++;
+
+ /* initialize the list to store empty blocks to be reused */
+ dclist_init(&slab->emptyblocks);
+
+ /* initialize each blocklist slot */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
+ dlist_init(&slab->blocklist[i]);
#ifdef MEMORY_CONTEXT_CHECKING
- /* set the freechunks pointer right after the freelists array */
- slab->freechunks
- = (bool *) slab + offsetof(SlabContext, freelist) + freelistSize;
+ /* set the isChunkFree pointer right after the end of the context */
+ slab->isChunkFree = (bool *) ((char *) slab + sizeof(SlabContext));
#endif
/* Finally, do the type-independent part of context creation */
@@ -252,6 +427,7 @@ void
SlabReset(MemoryContext context)
{
SlabContext *slab = (SlabContext *) context;
+ dlist_mutable_iter miter;
int i;
Assert(SlabIsValid(slab));
@@ -261,12 +437,24 @@ SlabReset(MemoryContext context)
SlabCheck(context);
#endif
- /* walk over freelists and free the blocks */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /* release any retained empty blocks */
+ dclist_foreach_modify(miter, &slab->emptyblocks)
{
- dlist_mutable_iter miter;
+ SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
+
+ dclist_delete_from(&slab->emptyblocks, miter.cur);
- dlist_foreach_modify(miter, &slab->freelist[i])
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ context->mem_allocated -= slab->blockSize;
+ }
+
+ /* walk over blocklist and free the blocks */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
+ {
+ dlist_foreach_modify(miter, &slab->blocklist[i])
{
SlabBlock *block = dlist_container(SlabBlock, node, miter.cur);
@@ -276,14 +464,12 @@ SlabReset(MemoryContext context)
wipe_mem(block, slab->blockSize);
#endif
free(block);
- slab->nblocks--;
context->mem_allocated -= slab->blockSize;
}
}
- slab->minFreeChunks = 0;
+ slab->curBlocklistIndex = 0;
- Assert(slab->nblocks == 0);
Assert(context->mem_allocated == 0);
}
@@ -302,7 +488,7 @@ SlabDelete(MemoryContext context)
/*
* SlabAlloc
- * Returns pointer to allocated memory of given size or NULL if
+ * Returns a pointer to allocated memory of given size or NULL if
* request could not be completed; memory is added to the slab.
*/
void *
@@ -311,127 +497,118 @@ SlabAlloc(MemoryContext context, Size size)
SlabContext *slab = (SlabContext *) context;
SlabBlock *block;
MemoryChunk *chunk;
- int idx;
Assert(SlabIsValid(slab));
- Assert((slab->minFreeChunks >= 0) &&
- (slab->minFreeChunks < slab->chunksPerBlock));
+ /* sanity check that this is pointing to a valid blocklist */
+ Assert(slab->curBlocklistIndex >= 0);
+ Assert(slab->curBlocklistIndex <= SlabBlocklistIndex(slab, slab->chunksPerBlock));
/* make sure we only allow correct request size */
- if (size != slab->chunkSize)
+ if (unlikely(size != slab->chunkSize))
elog(ERROR, "unexpected alloc chunk size %zu (expected %zu)",
size, slab->chunkSize);
/*
- * If there are no free chunks in any existing block, create a new block
- * and put it to the last freelist bucket.
- *
- * slab->minFreeChunks == 0 means there are no blocks with free chunks,
- * thanks to how minFreeChunks is updated at the end of SlabAlloc().
+ * Handle the case when there are no partially filled blocks available.
+ * SlabFree() will have updated the curBlocklistIndex setting it to zero
+ * to indicate that it has freed the final block. Also later in
+ * SlabAlloc() we will set the curBlocklistIndex to zero if we end up
+ * filling the final block.
*/
- if (slab->minFreeChunks == 0)
+ if (unlikely(slab->curBlocklistIndex == 0))
{
- block = (SlabBlock *) malloc(slab->blockSize);
+ dlist_head *blocklist;
+ int blocklist_idx;
+
+ /* to save allocating a new one, first check the empty blocks list */
+ if (dclist_count(&slab->emptyblocks) > 0)
+ {
+ dlist_node *node = dclist_pop_head_node(&slab->emptyblocks);
- if (block == NULL)
- return NULL;
+ block = dlist_container(SlabBlock, node, node);
- block->nfree = slab->chunksPerBlock;
- block->firstFreeChunk = 0;
- block->slab = slab;
+ /*
+ * SlabFree() should have left this block in a valid state with
+ * all chunks free. Ensure that's the case.
+ */
+ Assert(block->nfree == slab->chunksPerBlock);
- /*
- * Put all the chunks on a freelist. Walk the chunks and point each
- * one to the next one.
- */
- for (idx = 0; idx < slab->chunksPerBlock; idx++)
- {
- chunk = SlabBlockGetChunk(slab, block, idx);
- *(int32 *) MemoryChunkGetPointer(chunk) = (idx + 1);
+ /* fetch the next chunk from this block */
+ chunk = SlabGetNextFreeChunk(slab, block);
}
+ else
+ {
+ block = (SlabBlock *) malloc(slab->blockSize);
- /*
- * And add it to the last freelist with all chunks empty.
- *
- * We know there are no blocks in the freelist, otherwise we wouldn't
- * need a new block.
- */
- Assert(dlist_is_empty(&slab->freelist[slab->chunksPerBlock]));
+ if (unlikely(block == NULL))
+ return NULL;
- dlist_push_head(&slab->freelist[slab->chunksPerBlock], &block->node);
+ block->slab = slab;
+ context->mem_allocated += slab->blockSize;
- slab->minFreeChunks = slab->chunksPerBlock;
- slab->nblocks += 1;
- context->mem_allocated += slab->blockSize;
- }
+ /* use the first chunk in the new block */
+ chunk = SlabBlockGetChunk(slab, block, 0);
- /* grab the block from the freelist (even the new block is there) */
- block = dlist_head_element(SlabBlock, node,
- &slab->freelist[slab->minFreeChunks]);
+ block->nfree = slab->chunksPerBlock - 1;
+ block->unused = SlabBlockGetChunk(slab, block, 1);
+ block->freehead = NULL;
+ block->nunused = slab->chunksPerBlock - 1;
+ }
- /* make sure we actually got a valid block, with matching nfree */
- Assert(block != NULL);
- Assert(slab->minFreeChunks == block->nfree);
- Assert(block->nfree > 0);
+ /* find the blocklist element for storing blocks with 1 used chunk */
+ blocklist_idx = SlabBlocklistIndex(slab, slab->chunksPerBlock - 1);
+ blocklist = &slab->blocklist[blocklist_idx];
- /* we know index of the first free chunk in the block */
- idx = block->firstFreeChunk;
+ /* this better be empty. We just added a block thinking it was */
+ Assert(dlist_is_empty(blocklist));
- /* make sure the chunk index is valid, and that it's marked as empty */
- Assert((idx >= 0) && (idx < slab->chunksPerBlock));
+ dlist_push_head(blocklist, &block->node);
- /* compute the chunk location block start (after the block header) */
- chunk = SlabBlockGetChunk(slab, block, idx);
+ slab->curBlocklistIndex = blocklist_idx;
+ }
+ else
+ {
+ dlist_head *blocklist = &slab->blocklist[slab->curBlocklistIndex];
+ int new_blocklist_idx;
- /*
- * Update the block nfree count, and also the minFreeChunks as we've
- * decreased nfree for a block with the minimum number of free chunks
- * (because that's how we chose the block).
- */
- block->nfree--;
- slab->minFreeChunks = block->nfree;
+ Assert(!dlist_is_empty(blocklist));
- /*
- * Remove the chunk from the freelist head. The index of the next free
- * chunk is stored in the chunk itself.
- */
- VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- block->firstFreeChunk = *(int32 *) MemoryChunkGetPointer(chunk);
+ /* grab the block from the blocklist */
+ block = dlist_head_element(SlabBlock, node, blocklist);
- Assert(block->firstFreeChunk >= 0);
- Assert(block->firstFreeChunk <= slab->chunksPerBlock);
+ /* make sure we actually got a valid block, with matching nfree */
+ Assert(block != NULL);
+ Assert(slab->curBlocklistIndex == SlabBlocklistIndex(slab, block->nfree));
+ Assert(block->nfree > 0);
- Assert((block->nfree != 0 &&
- block->firstFreeChunk < slab->chunksPerBlock) ||
- (block->nfree == 0 &&
- block->firstFreeChunk == slab->chunksPerBlock));
+ /* fetch the next chunk from this block */
+ chunk = SlabGetNextFreeChunk(slab, block);
- /* move the whole block to the right place in the freelist */
- dlist_delete(&block->node);
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
+ /* get the new blocklist index based on the new free chunk count */
+ new_blocklist_idx = SlabBlocklistIndex(slab, block->nfree);
- /*
- * And finally update minFreeChunks, i.e. the index to the block with the
- * lowest number of free chunks. We only need to do that when the block
- * got full (otherwise we know the current block is the right one). We'll
- * simply walk the freelist until we find a non-empty entry.
- */
- if (slab->minFreeChunks == 0)
- {
- for (idx = 1; idx <= slab->chunksPerBlock; idx++)
+ /*
+ * Handle the case where the blocklist index changes. This also deals
+ * with blocks becoming full as only full blocks go at index 0.
+ */
+ if (unlikely(slab->curBlocklistIndex != new_blocklist_idx))
{
- if (dlist_is_empty(&slab->freelist[idx]))
- continue;
+ dlist_delete_from(blocklist, &block->node);
+ dlist_push_head(&slab->blocklist[new_blocklist_idx], &block->node);
- /* found a non-empty freelist */
- slab->minFreeChunks = idx;
- break;
+ if (dlist_is_empty(blocklist))
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
}
}
- if (slab->minFreeChunks == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
+ /*
+ * Check that the chunk pointer is actually somewhere on the block and is
+ * aligned as expected.
+ */
+ Assert(chunk >= SlabBlockGetChunk(slab, block, 0));
+ Assert(chunk <= SlabBlockGetChunk(slab, block, slab->chunksPerBlock - 1));
+ Assert(SlabChunkMod(slab, block, chunk) == 0);
/* Prepare to initialize the chunk header. */
VALGRIND_MAKE_MEM_UNDEFINED(chunk, Slab_CHUNKHDRSZ);
@@ -453,8 +630,6 @@ SlabAlloc(MemoryContext context, Size size)
randomize_mem((char *) MemoryChunkGetPointer(chunk), size);
#endif
- Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
-
return MemoryChunkGetPointer(chunk);
}
@@ -468,7 +643,8 @@ SlabFree(void *pointer)
MemoryChunk *chunk = PointerGetMemoryChunk(pointer);
SlabBlock *block = MemoryChunkGetBlock(chunk);
SlabContext *slab;
- int idx;
+ int curBlocklistIdx;
+ int newBlocklistIdx;
/*
* For speed reasons we just Assert that the referenced block is good.
@@ -486,63 +662,82 @@ SlabFree(void *pointer)
slab->header.name, chunk);
#endif
- /* compute index of the chunk with respect to block start */
- idx = SlabChunkIndex(slab, block, chunk);
+ /* push this chunk onto the head of the block's free list */
+ *(MemoryChunk **) pointer = block->freehead;
+ block->freehead = chunk;
- /* add chunk to freelist, and update block nfree count */
- *(int32 *) pointer = block->firstFreeChunk;
- block->firstFreeChunk = idx;
block->nfree++;
Assert(block->nfree > 0);
Assert(block->nfree <= slab->chunksPerBlock);
#ifdef CLOBBER_FREED_MEMORY
- /* XXX don't wipe the int32 index, used for block-level freelist */
- wipe_mem((char *) pointer + sizeof(int32),
- slab->chunkSize - sizeof(int32));
+ /* don't wipe the free list MemoryChunk pointer stored in the chunk */
+ wipe_mem((char *) pointer + sizeof(MemoryChunk *),
+ slab->chunkSize - sizeof(MemoryChunk *));
#endif
- /* remove the block from a freelist */
- dlist_delete(&block->node);
+ curBlocklistIdx = SlabBlocklistIndex(slab, block->nfree - 1);
+ newBlocklistIdx = SlabBlocklistIndex(slab, block->nfree);
/*
- * See if we need to update the minFreeChunks field for the slab - we only
- * need to do that if there the block had that number of free chunks
- * before we freed one. In that case, we check if there still are blocks
- * in the original freelist and we either keep the current value (if there
- * still are blocks) or increment it by one (the new block is still the
- * one with minimum free chunks).
- *
- * The one exception is when the block will get completely free - in that
- * case we will free it, se we can't use it for minFreeChunks. It however
- * means there are no more blocks with free chunks.
+ * Check if the block needs to be moved to another element on the
+ * blocklist based on it now having 1 more free chunk.
*/
- if (slab->minFreeChunks == (block->nfree - 1))
+ if (unlikely(curBlocklistIdx != newBlocklistIdx))
{
- /* Have we removed the last chunk from the freelist? */
- if (dlist_is_empty(&slab->freelist[slab->minFreeChunks]))
+ /* do the move */
+ dlist_delete_from(&slab->blocklist[curBlocklistIdx], &block->node);
+ dlist_push_head(&slab->blocklist[newBlocklistIdx], &block->node);
+
+ /*
+ * It's possible that we've no blocks in the blocklist at the
+ * curBlocklistIndex position. When this happens we must find the
+ * next blocklist index which contains blocks. We can be certain
+ * we'll find a block as at least one must exist for the chunk we're
+ * currently freeing.
+ */
+ if (slab->curBlocklistIndex == curBlocklistIdx &&
+ dlist_is_empty(&slab->blocklist[curBlocklistIdx]))
{
- /* but if we made the block entirely free, we'll free it */
- if (block->nfree == slab->chunksPerBlock)
- slab->minFreeChunks = 0;
- else
- slab->minFreeChunks++;
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
+ Assert(slab->curBlocklistIndex > 0);
}
}
- /* If the block is now completely empty, free it. */
- if (block->nfree == slab->chunksPerBlock)
+ /* Handle when a block becomes completely empty */
+ if (unlikely(block->nfree == slab->chunksPerBlock))
{
- free(block);
- slab->nblocks--;
- slab->header.mem_allocated -= slab->blockSize;
- }
- else
- dlist_push_head(&slab->freelist[block->nfree], &block->node);
+ /* remove the block */
+ dlist_delete_from(&slab->blocklist[newBlocklistIdx], &block->node);
+
+ /*
+ * To avoid thrashing malloc/free, we keep a list of empty blocks that
+ * we can reuse again instead of having to malloc a new one.
+ */
+ if (dclist_count(&slab->emptyblocks) < SLAB_MAXIMUM_EMPTY_BLOCKS)
+ dclist_push_head(&slab->emptyblocks, &block->node);
+ else
+ {
+ /*
+ * When we have enough empty blocks stored already, we actually
+ * free the block.
+ */
+#ifdef CLOBBER_FREED_MEMORY
+ wipe_mem(block, slab->blockSize);
+#endif
+ free(block);
+ slab->header.mem_allocated -= slab->blockSize;
+ }
- Assert(slab->nblocks >= 0);
- Assert(slab->nblocks * slab->blockSize == slab->header.mem_allocated);
+ /*
+ * Check if we need to reset the blocklist index. This is required
+ * when the blocklist this block is on has become completely empty.
+ */
+ if (slab->curBlocklistIndex == newBlocklistIdx &&
+ dlist_is_empty(&slab->blocklist[newBlocklistIdx]))
+ slab->curBlocklistIndex = SlabFindNextBlockListIndex(slab);
+ }
}
/*
@@ -617,16 +812,14 @@ SlabGetChunkSpace(void *pointer)
/*
* SlabIsEmpty
- * Is an Slab empty of any allocated space?
+ * Is the slab empty of any allocated space?
*/
bool
SlabIsEmpty(MemoryContext context)
{
- SlabContext *slab = (SlabContext *) context;
-
- Assert(SlabIsValid(slab));
+ Assert(SlabIsValid((SlabContext *) context));
- return (slab->nblocks == 0);
+ return (context->mem_allocated == 0);
}
/*
@@ -654,13 +847,16 @@ SlabStats(MemoryContext context,
Assert(SlabIsValid(slab));
/* Include context header in totalspace */
- totalspace = slab->headerSize;
+ totalspace = Slab_CONTEXT_HDRSZ(slab->chunksPerBlock);
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /* Add the space consumed by blocks in the emptyblocks list */
+ totalspace += dclist_count(&slab->emptyblocks) * slab->blockSize;
+
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
{
dlist_iter iter;
- dlist_foreach(iter, &slab->freelist[i])
+ dlist_foreach(iter, &slab->blocklist[i])
{
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
@@ -675,10 +871,11 @@ SlabStats(MemoryContext context,
{
char stats_string[200];
+ /* XXX should we include free chunks on empty blocks? */
snprintf(stats_string, sizeof(stats_string),
- "%zu total in %zu blocks; %zu free (%zu chunks); %zu used",
- totalspace, nblocks, freespace, freechunks,
- totalspace - freespace);
+ "%zu total in %zu blocks; %u empty blocks; %zu free (%zu chunks); %zu used",
+ totalspace, nblocks, dclist_count(&slab->emptyblocks),
+ freespace, freechunks, totalspace - freespace);
printfunc(context, passthru, stats_string, print_to_stderr);
}
@@ -696,7 +893,7 @@ SlabStats(MemoryContext context,
/*
* SlabCheck
- * Walk through chunks and check consistency of memory.
+ * Walk through all blocks looking for inconsistencies.
*
* NOTE: report errors as WARNING, *not* ERROR or FATAL. Otherwise you'll
* find yourself in an infinite loop when trouble occurs, because this
@@ -707,67 +904,112 @@ SlabCheck(MemoryContext context)
{
SlabContext *slab = (SlabContext *) context;
int i;
+ int nblocks = 0;
const char *name = slab->header.name;
+ dlist_iter iter;
Assert(SlabIsValid(slab));
Assert(slab->chunksPerBlock > 0);
- /* walk all the freelists */
- for (i = 0; i <= slab->chunksPerBlock; i++)
+ /*
+ * Have a look at the empty blocks. These should have all their chunks
+ * marked as free. Ensure that's the case.
+ */
+ dclist_foreach(iter, &slab->emptyblocks)
+ {
+ SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+
+ if (block->nfree != slab->chunksPerBlock)
+ elog(WARNING, "problem in slab %s: empty block %p should have %d free chunks but has %d chunks free",
+ name, block, slab->chunksPerBlock, block->nfree);
+ }
+
+ /* walk the non-empty block lists */
+ for (i = 0; i < SLAB_BLOCKLIST_COUNT; i++)
{
int j,
nfree;
- dlist_iter iter;
- /* walk all blocks on this freelist */
- dlist_foreach(iter, &slab->freelist[i])
+ /* walk all blocks on this blocklist */
+ dlist_foreach(iter, &slab->blocklist[i])
{
- int idx;
SlabBlock *block = dlist_container(SlabBlock, node, iter.cur);
+ MemoryChunk *cur_chunk;
/*
* Make sure the number of free chunks (in the block header)
- * matches position in the freelist.
+ * matches the position in the blocklist.
*/
- if (block->nfree != i)
- elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match freelist %d",
- name, block->nfree, block, i);
+ if (SlabBlocklistIndex(slab, block->nfree) != i)
+ elog(WARNING, "problem in slab %s: block %p is on blocklist %d but should be on blocklist %d",
+ name, block, i, SlabBlocklistIndex(slab, block->nfree));
+
+ /* make sure the block is not empty */
+ if (block->nfree <= 0)
+ elog(WARNING, "problem in slab %s: empty block %p incorrectly stored on blocklist element %d",
+ name, block, i);
/* make sure the slab pointer correctly points to this context */
if (block->slab != slab)
elog(WARNING, "problem in slab %s: bogus slab link in block %p",
name, block);
- /* reset the bitmap of free chunks for this block */
- memset(slab->freechunks, 0, (slab->chunksPerBlock * sizeof(bool)));
- idx = block->firstFreeChunk;
+ /* reset the array of free chunks for this block */
+ memset(slab->isChunkFree, 0, (slab->chunksPerBlock * sizeof(bool)));
+ nfree = 0;
+ /* walk through the block's free list chunks */
+ cur_chunk = block->freehead;
+ while (cur_chunk != NULL)
+ {
+ int chunkidx = SlabChunkIndex(slab, block, cur_chunk);
+
+ /*
+ * Ensure the free list link points to something on the block
+ * at an address aligned according to the full chunk size.
+ */
+ if (cur_chunk < SlabBlockGetChunk(slab, block, 0) ||
+ cur_chunk > SlabBlockGetChunk(slab, block, slab->chunksPerBlock - 1) ||
+ SlabChunkMod(slab, block, cur_chunk) != 0)
+ elog(WARNING, "problem in slab %s: bogus free list link %p in block %p",
+ name, cur_chunk, block);
+
+ /* count the chunk and mark it free on the free chunk array */
+ nfree++;
+ slab->isChunkFree[chunkidx] = true;
+
+ /* read pointer of the next free chunk */
+ VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(cur_chunk), sizeof(MemoryChunk *));
+ cur_chunk = *(MemoryChunk **) SlabChunkGetPointer(cur_chunk);
+ }
+
+ /* check that the unused pointer matches what nunused claims */
+ if (SlabBlockGetChunk(slab, block, slab->chunksPerBlock - block->nunused) !=
+ block->unused)
+ elog(WARNING, "problem in slab %s: mismatch detected between nunused chunks and unused pointer in block %p",
+ name, block);
/*
- * Now walk through the chunks, count the free ones and also
- * perform some additional checks for the used ones. As the chunk
- * freelist is stored within the chunks themselves, we have to
- * walk through the chunks and construct our own bitmap.
+ * count the remaining free chunks that have yet to make it onto
+ * the block's free list.
*/
-
- nfree = 0;
- while (idx < slab->chunksPerBlock)
+ cur_chunk = block->unused;
+ for (j = 0; j < block->nunused; j++)
{
- MemoryChunk *chunk;
+ int chunkidx = SlabChunkIndex(slab, block, cur_chunk);
+
- /* count the chunk as free, add it to the bitmap */
+ /* count the chunk as free and mark it as so in the array */
nfree++;
- slab->freechunks[idx] = true;
+ if (chunkidx < slab->chunksPerBlock)
+ slab->isChunkFree[chunkidx] = true;
- /* read index of the next free chunk */
- chunk = SlabBlockGetChunk(slab, block, idx);
- VALGRIND_MAKE_MEM_DEFINED(MemoryChunkGetPointer(chunk), sizeof(int32));
- idx = *(int32 *) MemoryChunkGetPointer(chunk);
+ /* move forward 1 chunk */
+ cur_chunk = (MemoryChunk *) (((char *) cur_chunk) + slab->fullChunkSize);
}
for (j = 0; j < slab->chunksPerBlock; j++)
{
- /* non-zero bit in the bitmap means chunk the chunk is used */
- if (!slab->freechunks[j])
+ if (!slab->isChunkFree[j])
{
MemoryChunk *chunk = SlabBlockGetChunk(slab, block, j);
SlabBlock *chunkblock = (SlabBlock *) MemoryChunkGetBlock(chunk);
@@ -793,12 +1035,17 @@ SlabCheck(MemoryContext context)
* in the block header).
*/
if (nfree != block->nfree)
- elog(WARNING, "problem in slab %s: number of free chunks %d in block %p does not match bitmap %d",
- name, block->nfree, block, nfree);
+ elog(WARNING, "problem in slab %s: nfree in block %p is %d but %d chunk were found as free",
+ name, block, block->nfree, nfree);
+
+ nblocks++;
}
}
- Assert(slab->nblocks * slab->blockSize == context->mem_allocated);
+ /* the stored empty blocks are tracked in mem_allocated too */
+ nblocks += dclist_count(&slab->emptyblocks);
+
+ Assert(nblocks * slab->blockSize == context->mem_allocated);
}
#endif /* MEMORY_CONTEXT_CHECKING */
[text/plain] alloc_bench_contrib.patch (11.3K, ../../CAApHDvq6eUdLJxAUdSmukGiiTQNT79cNtntL=3FE52T_AP3XDQ@mail.gmail.com/3-alloc_bench_contrib.patch)
download | inline diff:
diff --git a/contrib/alloc_bench/.gitignore b/contrib/alloc_bench/.gitignore
new file mode 100644
index 0000000000..5dcb3ff972
--- /dev/null
+++ b/contrib/alloc_bench/.gitignore
@@ -0,0 +1,4 @@
+# Generated subdirectories
+/log/
+/results/
+/tmp_check/
diff --git a/contrib/alloc_bench/Makefile b/contrib/alloc_bench/Makefile
new file mode 100644
index 0000000000..c672b29bdb
--- /dev/null
+++ b/contrib/alloc_bench/Makefile
@@ -0,0 +1,21 @@
+# contrib/alloc_bench/Makefile
+
+MODULE_big = alloc_bench
+OBJS = alloc_bench.o
+
+EXTENSION = alloc_bench
+DATA = alloc_bench--1.0.sql
+PGFILEDESC = "alloc_bench - memory context benchmarking functions"
+
+REGRESS = alloc_bench
+
+ifdef USE_PGXS
+PG_CONFIG = pg_config
+PGXS := $(shell $(PG_CONFIG) --pgxs)
+include $(PGXS)
+else
+subdir = contrib/alloc_bench
+top_builddir = ../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+endif
diff --git a/contrib/alloc_bench/alloc_bench--1.0.sql b/contrib/alloc_bench/alloc_bench--1.0.sql
new file mode 100644
index 0000000000..2ca604a682
--- /dev/null
+++ b/contrib/alloc_bench/alloc_bench--1.0.sql
@@ -0,0 +1,9 @@
+/* alloc_bench--1.0.sql */
+
+-- complain if script is sourced in psql, rather than via CREATE EXTENSION
+\echo Use "CREATE EXTENSION alloc_bench" to load this file. \quit
+
+CREATE FUNCTION alloc_bench(context_type text, pattern text, nchunks bigint, block_size bigint, chunk_size bigint, nloops int, random_segments int, out mem_allocated bigint, out alloc_us bigint, out free_us bigint)
+AS 'MODULE_PATHNAME', 'alloc_bench'
+LANGUAGE C VOLATILE STRICT;
+
diff --git a/contrib/alloc_bench/alloc_bench.c b/contrib/alloc_bench/alloc_bench.c
new file mode 100644
index 0000000000..f76501c053
--- /dev/null
+++ b/contrib/alloc_bench/alloc_bench.c
@@ -0,0 +1,199 @@
+/*-------------------------------------------------------------------------
+ *
+ * alloc_bench.c
+ *
+ * helper functions to benchmark memory contexts with different workloads
+ *-------------------------------------------------------------------------
+ */
+#include "postgres.h"
+
+#include <sys/time.h>
+
+#include "funcapi.h"
+#include "miscadmin.h"
+#include "utils/builtins.h"
+
+PG_MODULE_MAGIC;
+
+PG_FUNCTION_INFO_V1(alloc_bench);
+
+typedef enum AllocationPattern {
+ FIFO,
+ LIFO
+} AllocationPattern;
+
+static int random_segment_size;
+static int random_fifo = 0;
+
+static int
+int64_random_compare(const void *a, const void *b)
+{
+ const int64 ia = *(const int64 *) a;
+ const int64 ib = *(const int64 *) b;
+ int rss = random_segment_size;
+
+ if ((ia / rss * rss) + (random() % rss) > ib)
+ return random_fifo == 0 ? -1 : 1;
+ else
+ return random_fifo == 0 ? 1 : -1;
+}
+
+Datum
+alloc_bench(PG_FUNCTION_ARGS)
+{
+ MemoryContext cxt,
+ oldcxt;
+ void **chunks;
+ int64 *indexes;
+ int64 j;
+ char *context_type;
+ text *context_type_text = PG_GETARG_TEXT_PP(0);
+ char *pattern_str;
+ text *pattern_text = PG_GETARG_TEXT_PP(1);
+ int64 nchunks = PG_GETARG_INT64(2);
+ int64 blockSize = PG_GETARG_INT64(3);
+ int64 chunkSize = PG_GETARG_INT64(4);
+
+ int nloops = PG_GETARG_INT32(5);
+ int random_segments = PG_GETARG_INT32(6);
+
+ struct timeval start_time,
+ end_time;
+ int64 alloc_time = 0,
+ free_time = 0;
+ int64 mem_allocated;
+
+ TupleDesc tupdesc;
+ Datum result;
+ HeapTuple tuple;
+ Datum values[9];
+ bool nulls[9];
+ AllocationPattern pattern;
+
+ context_type = text_to_cstring(context_type_text);
+
+ if (strcmp(context_type, "generation") == 0)
+ cxt = GenerationContextCreate(CurrentMemoryContext,
+ "alloc_bench",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "aset") == 0)
+ cxt = AllocSetContextCreate(CurrentMemoryContext,
+ "alloc_bench",
+ ALLOCSET_DEFAULT_SIZES);
+ else if (strcmp(context_type, "slab") == 0)
+ cxt = SlabContextCreate(CurrentMemoryContext,
+ "alloc_bench",
+ blockSize,
+ chunkSize);
+ else
+ elog(ERROR, "%s is not a valid context type. context_type must be \"generation\", \"aset\" or \"slab\"",
+ context_type);
+
+ pattern_str = text_to_cstring(pattern_text);
+
+ if (strcmp(pattern_str, "fifo") == 0)
+ pattern = FIFO;
+ else if (strcmp(pattern_str, "lifo") == 0)
+ pattern = LIFO;
+ else
+ elog(ERROR, "%s is not a valid allocation pattern. Must be \"fifo\" or \"lifo\"",
+ pattern_str);
+
+ chunks = (void **) palloc(nchunks * sizeof(void *));
+ indexes = (int64 *) palloc(nchunks * sizeof(int64));
+
+ /* set the indexes so according to the allocation pattern */
+ switch (pattern)
+ {
+ case FIFO:
+ for (int64 i = 0; i < nchunks; i++)
+ indexes[i] = i;
+ break;
+ case LIFO:
+ for (int64 i = 0; i < nchunks; i++)
+ indexes[i] = nchunks - i - 1;
+ break;
+ }
+
+ /*
+ * When a non-zero random_segments is specified, we randomize the
+ * order of the pfree's so they're random within their own segment.
+ * This means we don't entirely randomize the order, but if there are
+ * say, 10 segments and 60 chunks, we only randomize the first 6
+ * chunks then the next 6. Specifiying a lower number of
+ * random_segments means the FIFO or LIFO pattern becomes more random.
+ */
+ if (random_segments > 0)
+ {
+ random_segment_size = nchunks / random_segments;
+ random_fifo = (pattern == FIFO);
+ qsort(indexes, nchunks, sizeof(int64), int64_random_compare);
+ }
+
+ mem_allocated = 0;
+
+ oldcxt = MemoryContextSwitchTo(cxt);
+
+ /* do the requested number of pfree/palloc loops */
+ for (j = 0; j < nloops; j++)
+ {
+ CHECK_FOR_INTERRUPTS();
+
+ gettimeofday(&start_time, NULL);
+
+ /*
+ * We do palloc0 instead of palloc to simulate touching the
+ * memory. This will load cachelines, which is likely more
+ * realistic than just allocating memory and not doing
+ * anything with it.
+ */
+ for (int64 i = 0; i < nchunks; i++)
+ chunks[i] = palloc0(chunkSize);
+
+ gettimeofday(&end_time, NULL);
+
+ alloc_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ gettimeofday(&start_time, NULL);
+
+ for (int64 i = 0; i < nchunks; i++)
+ {
+ char *ptr = (char *) chunks[indexes[i]];
+
+ /*
+ * Free the chunk, but first touch the first cacheline
+ * of the chunk to simulate that we've just done
+ * something real with this memory.
+ */
+ if (ptr[0] == '\0')
+ pfree(ptr);
+ }
+
+ gettimeofday(&end_time, NULL);
+
+ free_time += (end_time.tv_sec - start_time.tv_sec) * 1000000L +
+ (end_time.tv_usec - start_time.tv_usec);
+
+ mem_allocated = Max(mem_allocated, MemoryContextMemAllocated(cxt, true));
+ }
+
+ MemoryContextSwitchTo(oldcxt);
+ MemoryContextDelete(cxt);
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ elog(ERROR, "return type must be a row type");
+
+ values[0] = Int64GetDatum(mem_allocated);
+ values[1] = Int64GetDatum(alloc_time);
+ values[2] = Int64GetDatum(free_time);
+
+ memset(nulls, 0, sizeof(nulls));
+
+ tuple = heap_form_tuple(tupdesc, values, nulls);
+ result = HeapTupleGetDatum(tuple);
+
+ PG_RETURN_DATUM(result);
+}
+
diff --git a/contrib/alloc_bench/alloc_bench.control b/contrib/alloc_bench/alloc_bench.control
new file mode 100644
index 0000000000..597480ed5e
--- /dev/null
+++ b/contrib/alloc_bench/alloc_bench.control
@@ -0,0 +1,4 @@
+# alloc_bench extension
+comment = 'functions for benchmarking memory context'
+default_version = '1.0'
+module_pathname = '$libdir/alloc_bench'
diff --git a/contrib/alloc_bench/bench.sql b/contrib/alloc_bench/bench.sql
new file mode 100644
index 0000000000..d476a386c2
--- /dev/null
+++ b/contrib/alloc_bench/bench.sql
@@ -0,0 +1,40 @@
+CREATE EXTENSION slab_bench;
+
+\o fifo-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+\o lifo-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+\o random-no-loops.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 0, 0, 0) x;
+
+
+\o fifo-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+\o lifo-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+\o random-increase.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 100, 10000, 15000) x;
+
+
+\o fifo-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+\o lifo-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+\o random-decrease.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000000, block_size, chunk_size, 100, 10000, 5000) x;
+
+
+\o fifo-cycle.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_fifo(1000, block_size, chunk_size, 10000, 1000, 1000) x;
+
+\o lifo-cycle.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_lifo(1000, block_size, chunk_size, 10000, 1000, 1000) x;
+
+\o random-cycle.data
+select run, block_size, chunk_size, 1000000 * chunk_size, x.* from generate_series(1,5) r(run), generate_series(32,512,32) a(chunk_size), (values (1024), (2048), (4096), (8192), (16384), (32768)) AS b(block_size), lateral slab_bench_random(1000, block_size, chunk_size, 10000, 1000, 1000) x;
[application/vnd.oasis.opendocument.spreadsheet] Slab Allocator performance improvements.ods (97.2K, ../../CAApHDvq6eUdLJxAUdSmukGiiTQNT79cNtntL=3FE52T_AP3XDQ@mail.gmail.com/4-Slab%20Allocator%20performance%20improvements.ods)
download
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-05 08:02 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-05 10:18 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-10 04:01 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-12 07:13 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-13 00:49 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-14 10:37 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-20 03:35 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
@ 2022-12-20 08:19 ` John Naylor <john.naylor@enterprisedb.com>
2022-12-20 08:51 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
0 siblings, 1 reply; 29+ messages in thread
From: John Naylor @ 2022-12-20 08:19 UTC (permalink / raw)
To: David Rowley <dgrowleyml@gmail.com>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On Tue, Dec 20, 2022 at 10:36 AM David Rowley <dgrowleyml@gmail.com> wrote:
>
> I'm planning on pushing the attached v3 patch shortly. I've spent
> several days reading over this and testing it in detail along with
> adding additional features to the SlabCheck code to find more
> inconsistencies.
FWIW, I reran my test from last week and got similar results.
--
John Naylor
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-05 08:02 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-05 10:18 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-10 04:01 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-12 07:13 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-13 00:49 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-14 10:37 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
2022-12-20 03:35 ` Re: slab allocator performance issues David Rowley <dgrowleyml@gmail.com>
2022-12-20 08:19 ` Re: slab allocator performance issues John Naylor <john.naylor@enterprisedb.com>
@ 2022-12-20 08:51 ` David Rowley <dgrowleyml@gmail.com>
0 siblings, 0 replies; 29+ messages in thread
From: David Rowley @ 2022-12-20 08:51 UTC (permalink / raw)
To: John Naylor <john.naylor@enterprisedb.com>; +Cc: Tomas Vondra <tomas.vondra@enterprisedb.com>; Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On Tue, 20 Dec 2022 at 21:19, John Naylor <john.naylor@enterprisedb.com> wrote:
>
>
> On Tue, Dec 20, 2022 at 10:36 AM David Rowley <dgrowleyml@gmail.com> wrote:
> >
> > I'm planning on pushing the attached v3 patch shortly. I've spent
> > several days reading over this and testing it in detail along with
> > adding additional features to the SlabCheck code to find more
> > inconsistencies.
>
> FWIW, I reran my test from last week and got similar results.
Thanks a lot for testing that stuff last week. I got a bit engrossed
in the perf weirdness and forgot to reply. I found they made much
more sense after using palloc0 and touching the allocated memory just
before freeing. I think this is a more realistic test.
I've now pushed the patch after making a small adjustment to the
version I sent earlier.
David
^ permalink raw reply [nested|flat] 29+ messages in thread
* Re: slab allocator performance issues
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 17:59 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Re: slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Re: slab allocator performance issues Tomas Vondra <tomas.vondra@enterprisedb.com>
@ 2022-12-05 15:31 ` Robert Haas <robertmhaas@gmail.com>
1 sibling, 0 replies; 29+ messages in thread
From: Robert Haas @ 2022-12-05 15:31 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@enterprisedb.com>; +Cc: Andres Freund <andres@anarazel.de>; pgsql-hackers@postgresql.org, Tomas Vondra <tv@fuzzy.cz>
On Fri, Sep 10, 2021 at 5:07 PM Tomas Vondra
<tomas.vondra@enterprisedb.com> wrote:
> Turns out it's pretty difficult to benchmark this, because the results
> strongly depend on what the backend did before.
What you report here seems to be mostly cold-cache effects, with which
I don't think we need to be overly concerned. We don't want big
regressions in the cold-cache case, but there is always going to be
some overhead when a new backend starts up, because you've got to
fault some pages into the heap/malloc arena/whatever before they can
be efficiently accessed. What would be more concerning is if we found
out that the performance depended heavily on the internal state of the
allocator. For example, suppose you have two warmup routines W1 and
W2, each of which touches the same amount of total memory, but with
different details. Then you have a benchmark B. If you do W1-B and
W2-B and the time for B varies dramatically between them, then you've
maybe got an issue. For instance, it could indicate that the allocator
has issue when the old and new allocations are very different sizes,
or something like that.
--
Robert Haas
EDB: http://www.enterprisedb.com
^ permalink raw reply [nested|flat] 29+ messages in thread
end of thread, other threads:[~2022-12-20 08:51 UTC | newest]
Thread overview: 29+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2021-07-17 19:43 slab allocator performance issues Andres Freund <andres@anarazel.de>
2021-07-17 19:53 ` Andres Freund <andres@anarazel.de>
2021-07-17 20:35 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 21:14 ` Andres Freund <andres@anarazel.de>
2021-07-17 22:46 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-17 23:10 ` Andres Freund <andres@anarazel.de>
2021-07-18 01:06 ` Andres Freund <andres@anarazel.de>
2021-07-18 17:23 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-07-19 20:56 ` Andres Freund <andres@anarazel.de>
2021-07-19 22:03 ` Ranier Vilela <ranier.vf@gmail.com>
2021-07-20 14:15 ` David Rowley <dgrowleyml@gmail.com>
2021-07-20 14:24 ` Ranier Vilela <ranier.vf@gmail.com>
2021-08-01 17:59 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-01 21:07 ` Andres Freund <andres@anarazel.de>
2021-08-01 22:01 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-08-03 13:33 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2021-09-10 21:06 ` Tomas Vondra <tomas.vondra@enterprisedb.com>
2022-10-12 09:37 ` David Rowley <dgrowleyml@gmail.com>
2022-11-11 09:20 ` John Naylor <john.naylor@enterprisedb.com>
2022-12-05 08:02 ` David Rowley <dgrowleyml@gmail.com>
2022-12-05 10:18 ` John Naylor <john.naylor@enterprisedb.com>
2022-12-10 04:01 ` David Rowley <dgrowleyml@gmail.com>
2022-12-12 07:13 ` John Naylor <john.naylor@enterprisedb.com>
2022-12-13 00:49 ` David Rowley <dgrowleyml@gmail.com>
2022-12-14 10:37 ` John Naylor <john.naylor@enterprisedb.com>
2022-12-20 03:35 ` David Rowley <dgrowleyml@gmail.com>
2022-12-20 08:19 ` John Naylor <john.naylor@enterprisedb.com>
2022-12-20 08:51 ` David Rowley <dgrowleyml@gmail.com>
2022-12-05 15:31 ` Robert Haas <robertmhaas@gmail.com>
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox