agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedauto-vectorize varbit bitwise operators
4+ messages / 3 participants
[nested] [flat]
* auto-vectorize varbit bitwise operators
@ 2026-09-02 18:57 Nathan Bossart <nathandbossart@gmail.com>
2026-09-02 19:42 ` Re: auto-vectorize varbit bitwise operators Heikki Linnakangas <hlinnaka@iki.fi>
2026-09-03 02:01 ` Re: auto-vectorize varbit bitwise operators Chao Li <li.evan.chao@gmail.com>
0 siblings, 2 replies; 4+ messages in thread
From: Nathan Bossart @ 2026-09-02 18:57 UTC (permalink / raw)
To: pgsql-hackers
After adding -ftree-vectorize and moving the loop boundary computations to
outside the loops, popular compilers will auto-vectorize bit_and(),
bit_or(), bitxor(), and bitnot(). My testing indicates this produces some
nice speedups, but I haven't yet done anything scientific enough to share.
--
nathan
From 027ed15aa336948445918d85d01fed6eef0df6c0 Mon Sep 17 00:00:00 2001
From: Nathan Bossart <nathan@postgresql.org>
Date: Wed, 2 Sep 2026 13:52:14 -0500
Subject: [PATCH v1 1/1] Auto-vectorize the varbit bitwise operators.
Compile varbit.c with -ftree-vectorize where available, and adjust
the loops in bit_and(), bit_or(), bitxor(), and bitnot() so that
they are amenable to being auto-vectorized. (Mainly, that involves
reading the byte count into a local variable before the loop. As
written, the bound is recomputed from the varlena header on every
iteration, because the compiler must assume that the loop might
modify the header.)
---
src/backend/utils/adt/Makefile | 3 ++-
src/backend/utils/adt/meson.build | 9 ++++-----
src/backend/utils/adt/varbit.c | 25 +++++++++++++++++--------
3 files changed, 23 insertions(+), 14 deletions(-)
diff --git a/src/backend/utils/adt/Makefile b/src/backend/utils/adt/Makefile
index 0c7621957c1..94ac5169940 100644
--- a/src/backend/utils/adt/Makefile
+++ b/src/backend/utils/adt/Makefile
@@ -150,8 +150,9 @@ clean:
like.o: like.c like_match.c
-# Some code in numeric.c benefits from auto-vectorization
+# Some code benefits from auto-vectorization
numeric.o: CFLAGS += ${CFLAGS_VECTORIZE}
+varbit.o: CFLAGS += ${CFLAGS_VECTORIZE}
varlena.o: varlena.c levenshtein.c
diff --git a/src/backend/utils/adt/meson.build b/src/backend/utils/adt/meson.build
index d793f8145f6..04d6f8aa168 100644
--- a/src/backend/utils/adt/meson.build
+++ b/src/backend/utils/adt/meson.build
@@ -1,14 +1,14 @@
# Copyright (c) 2022-2026, PostgreSQL Global Development Group
-# Some code in numeric.c benefits from auto-vectorization
-numeric_backend_lib = static_library('numeric_backend_lib',
- 'numeric.c',
+# Some code benefits from auto-vectorization
+vectorize_backend_lib = static_library('vectorize_backend_lib',
+ ['numeric.c', 'varbit.c'],
dependencies: backend_build_deps,
kwargs: internal_lib_args,
c_args: vectorize_cflags,
)
-backend_link_with += numeric_backend_lib
+backend_link_with += vectorize_backend_lib
backend_sources += files(
'acl.c',
@@ -118,7 +118,6 @@ backend_sources += files(
'tsvector_op.c',
'tsvector_parser.c',
'uuid.c',
- 'varbit.c',
'varchar.c',
'varlena.c',
'version.c',
diff --git a/src/backend/utils/adt/varbit.c b/src/backend/utils/adt/varbit.c
index fe49ba851ac..cbb60e68335 100644
--- a/src/backend/utils/adt/varbit.c
+++ b/src/backend/utils/adt/varbit.c
@@ -1251,6 +1251,7 @@ bit_and(PG_FUNCTION_ARGS)
uint8 *p1,
*p2,
*r;
+ size_t nbytes;
bitlen1 = VARBITLEN(arg1);
bitlen2 = VARBITLEN(arg2);
@@ -1267,8 +1268,9 @@ bit_and(PG_FUNCTION_ARGS)
p1 = VARBITS(arg1);
p2 = VARBITS(arg2);
r = VARBITS(result);
- for (size_t i = 0; i < VARBITBYTES(arg1); i++)
- *r++ = *p1++ & *p2++;
+ nbytes = VARBITBYTES(arg1);
+ for (size_t i = 0; i < nbytes; i++)
+ r[i] = p1[i] & p2[i];
/* Padding is not needed as & of 0 pads is 0 */
@@ -1291,6 +1293,7 @@ bit_or(PG_FUNCTION_ARGS)
uint8 *p1,
*p2,
*r;
+ size_t nbytes;
bitlen1 = VARBITLEN(arg1);
bitlen2 = VARBITLEN(arg2);
@@ -1306,8 +1309,9 @@ bit_or(PG_FUNCTION_ARGS)
p1 = VARBITS(arg1);
p2 = VARBITS(arg2);
r = VARBITS(result);
- for (size_t i = 0; i < VARBITBYTES(arg1); i++)
- *r++ = *p1++ | *p2++;
+ nbytes = VARBITBYTES(arg1);
+ for (size_t i = 0; i < nbytes; i++)
+ r[i] = p1[i] | p2[i];
/* Padding is not needed as | of 0 pads is 0 */
@@ -1330,6 +1334,7 @@ bitxor(PG_FUNCTION_ARGS)
uint8 *p1,
*p2,
*r;
+ size_t nbytes;
bitlen1 = VARBITLEN(arg1);
bitlen2 = VARBITLEN(arg2);
@@ -1346,8 +1351,9 @@ bitxor(PG_FUNCTION_ARGS)
p1 = VARBITS(arg1);
p2 = VARBITS(arg2);
r = VARBITS(result);
- for (size_t i = 0; i < VARBITBYTES(arg1); i++)
- *r++ = *p1++ ^ *p2++;
+ nbytes = VARBITBYTES(arg1);
+ for (size_t i = 0; i < nbytes; i++)
+ r[i] = p1[i] ^ p2[i];
/* Padding is not needed as ^ of 0 pads is 0 */
@@ -1365,6 +1371,7 @@ bitnot(PG_FUNCTION_ARGS)
VarBit *result;
uint8 *p,
*r;
+ size_t nbytes;
result = (VarBit *) palloc(VARSIZE(arg));
SET_VARSIZE(result, VARSIZE(arg));
@@ -1372,8 +1379,10 @@ bitnot(PG_FUNCTION_ARGS)
p = VARBITS(arg);
r = VARBITS(result);
- for (; p < VARBITEND(arg); p++)
- *r++ = ~*p;
+ nbytes = VARBITBYTES(arg);
+ for (size_t i = 0; i < nbytes; i++)
+ r[i] = ~p[i];
+ r += nbytes;
/* Must zero-pad the result, because extra bits are surely 1's here */
VARBIT_PAD_LAST(result, r);
--
2.55.0
Attachments:
[text/plain] v1-0001-Auto-vectorize-the-varbit-bitwise-operators.patch (4.4K, ../../aphxrw6NkYkLMtZN@nathan/2-v1-0001-Auto-vectorize-the-varbit-bitwise-operators.patch)
download | inline diff:
From 027ed15aa336948445918d85d01fed6eef0df6c0 Mon Sep 17 00:00:00 2001
From: Nathan Bossart <nathan@postgresql.org>
Date: Wed, 2 Sep 2026 13:52:14 -0500
Subject: [PATCH v1 1/1] Auto-vectorize the varbit bitwise operators.
Compile varbit.c with -ftree-vectorize where available, and adjust
the loops in bit_and(), bit_or(), bitxor(), and bitnot() so that
they are amenable to being auto-vectorized. (Mainly, that involves
reading the byte count into a local variable before the loop. As
written, the bound is recomputed from the varlena header on every
iteration, because the compiler must assume that the loop might
modify the header.)
---
src/backend/utils/adt/Makefile | 3 ++-
src/backend/utils/adt/meson.build | 9 ++++-----
src/backend/utils/adt/varbit.c | 25 +++++++++++++++++--------
3 files changed, 23 insertions(+), 14 deletions(-)
diff --git a/src/backend/utils/adt/Makefile b/src/backend/utils/adt/Makefile
index 0c7621957c1..94ac5169940 100644
--- a/src/backend/utils/adt/Makefile
+++ b/src/backend/utils/adt/Makefile
@@ -150,8 +150,9 @@ clean:
like.o: like.c like_match.c
-# Some code in numeric.c benefits from auto-vectorization
+# Some code benefits from auto-vectorization
numeric.o: CFLAGS += ${CFLAGS_VECTORIZE}
+varbit.o: CFLAGS += ${CFLAGS_VECTORIZE}
varlena.o: varlena.c levenshtein.c
diff --git a/src/backend/utils/adt/meson.build b/src/backend/utils/adt/meson.build
index d793f8145f6..04d6f8aa168 100644
--- a/src/backend/utils/adt/meson.build
+++ b/src/backend/utils/adt/meson.build
@@ -1,14 +1,14 @@
# Copyright (c) 2022-2026, PostgreSQL Global Development Group
-# Some code in numeric.c benefits from auto-vectorization
-numeric_backend_lib = static_library('numeric_backend_lib',
- 'numeric.c',
+# Some code benefits from auto-vectorization
+vectorize_backend_lib = static_library('vectorize_backend_lib',
+ ['numeric.c', 'varbit.c'],
dependencies: backend_build_deps,
kwargs: internal_lib_args,
c_args: vectorize_cflags,
)
-backend_link_with += numeric_backend_lib
+backend_link_with += vectorize_backend_lib
backend_sources += files(
'acl.c',
@@ -118,7 +118,6 @@ backend_sources += files(
'tsvector_op.c',
'tsvector_parser.c',
'uuid.c',
- 'varbit.c',
'varchar.c',
'varlena.c',
'version.c',
diff --git a/src/backend/utils/adt/varbit.c b/src/backend/utils/adt/varbit.c
index fe49ba851ac..cbb60e68335 100644
--- a/src/backend/utils/adt/varbit.c
+++ b/src/backend/utils/adt/varbit.c
@@ -1251,6 +1251,7 @@ bit_and(PG_FUNCTION_ARGS)
uint8 *p1,
*p2,
*r;
+ size_t nbytes;
bitlen1 = VARBITLEN(arg1);
bitlen2 = VARBITLEN(arg2);
@@ -1267,8 +1268,9 @@ bit_and(PG_FUNCTION_ARGS)
p1 = VARBITS(arg1);
p2 = VARBITS(arg2);
r = VARBITS(result);
- for (size_t i = 0; i < VARBITBYTES(arg1); i++)
- *r++ = *p1++ & *p2++;
+ nbytes = VARBITBYTES(arg1);
+ for (size_t i = 0; i < nbytes; i++)
+ r[i] = p1[i] & p2[i];
/* Padding is not needed as & of 0 pads is 0 */
@@ -1291,6 +1293,7 @@ bit_or(PG_FUNCTION_ARGS)
uint8 *p1,
*p2,
*r;
+ size_t nbytes;
bitlen1 = VARBITLEN(arg1);
bitlen2 = VARBITLEN(arg2);
@@ -1306,8 +1309,9 @@ bit_or(PG_FUNCTION_ARGS)
p1 = VARBITS(arg1);
p2 = VARBITS(arg2);
r = VARBITS(result);
- for (size_t i = 0; i < VARBITBYTES(arg1); i++)
- *r++ = *p1++ | *p2++;
+ nbytes = VARBITBYTES(arg1);
+ for (size_t i = 0; i < nbytes; i++)
+ r[i] = p1[i] | p2[i];
/* Padding is not needed as | of 0 pads is 0 */
@@ -1330,6 +1334,7 @@ bitxor(PG_FUNCTION_ARGS)
uint8 *p1,
*p2,
*r;
+ size_t nbytes;
bitlen1 = VARBITLEN(arg1);
bitlen2 = VARBITLEN(arg2);
@@ -1346,8 +1351,9 @@ bitxor(PG_FUNCTION_ARGS)
p1 = VARBITS(arg1);
p2 = VARBITS(arg2);
r = VARBITS(result);
- for (size_t i = 0; i < VARBITBYTES(arg1); i++)
- *r++ = *p1++ ^ *p2++;
+ nbytes = VARBITBYTES(arg1);
+ for (size_t i = 0; i < nbytes; i++)
+ r[i] = p1[i] ^ p2[i];
/* Padding is not needed as ^ of 0 pads is 0 */
@@ -1365,6 +1371,7 @@ bitnot(PG_FUNCTION_ARGS)
VarBit *result;
uint8 *p,
*r;
+ size_t nbytes;
result = (VarBit *) palloc(VARSIZE(arg));
SET_VARSIZE(result, VARSIZE(arg));
@@ -1372,8 +1379,10 @@ bitnot(PG_FUNCTION_ARGS)
p = VARBITS(arg);
r = VARBITS(result);
- for (; p < VARBITEND(arg); p++)
- *r++ = ~*p;
+ nbytes = VARBITBYTES(arg);
+ for (size_t i = 0; i < nbytes; i++)
+ r[i] = ~p[i];
+ r += nbytes;
/* Must zero-pad the result, because extra bits are surely 1's here */
VARBIT_PAD_LAST(result, r);
--
2.55.0
^ permalink raw reply [nested|flat] 4+ messages in thread
* Re: auto-vectorize varbit bitwise operators
2026-09-02 18:57 auto-vectorize varbit bitwise operators Nathan Bossart <nathandbossart@gmail.com>
@ 2026-09-02 19:42 ` Heikki Linnakangas <hlinnaka@iki.fi>
1 sibling, 0 replies; 4+ messages in thread
From: Heikki Linnakangas @ 2026-09-02 19:42 UTC (permalink / raw)
To: Nathan Bossart <nathandbossart@gmail.com>; pgsql-hackers
On 02/09/2026 21:57, Nathan Bossart wrote:
> After adding -ftree-vectorize and moving the loop boundary computations to
> outside the loops, popular compilers will auto-vectorize bit_and(),
> bit_or(), bitxor(), and bitnot(). My testing indicates this produces some
> nice speedups, but I haven't yet done anything scientific enough to share.
Dunno if these are particularly worth optimizing, but regardless of
performance, I think the loop code looks better with these changes anyway.
- Heikki
^ permalink raw reply [nested|flat] 4+ messages in thread
* Re: auto-vectorize varbit bitwise operators
2026-09-02 18:57 auto-vectorize varbit bitwise operators Nathan Bossart <nathandbossart@gmail.com>
@ 2026-09-03 02:01 ` Chao Li <li.evan.chao@gmail.com>
2026-09-03 21:11 ` Re: auto-vectorize varbit bitwise operators Nathan Bossart <nathandbossart@gmail.com>
1 sibling, 1 reply; 4+ messages in thread
From: Chao Li @ 2026-09-03 02:01 UTC (permalink / raw)
To: Nathan Bossart <nathandbossart@gmail.com>; +Cc: pgsql-hackers
> On Sep 3, 2026, at 02:57, Nathan Bossart <nathandbossart@gmail.com> wrote:
>
> After adding -ftree-vectorize and moving the loop boundary computations to
> outside the loops, popular compilers will auto-vectorize bit_and(),
> bit_or(), bitxor(), and bitnot(). My testing indicates this produces some
> nice speedups, but I haven't yet done anything scientific enough to share.
>
> --
> nathan
> <v1-0001-Auto-vectorize-the-varbit-bitwise-operators.patch>
I think this should improve performance on CPUs with SIMD support. It also improves code readability.
Best regards,
--
Chao Li (Evan)
HighGo Software Co., Ltd.
https://www.highgo.com/
^ permalink raw reply [nested|flat] 4+ messages in thread
* Re: auto-vectorize varbit bitwise operators
2026-09-02 18:57 auto-vectorize varbit bitwise operators Nathan Bossart <nathandbossart@gmail.com>
2026-09-03 02:01 ` Re: auto-vectorize varbit bitwise operators Chao Li <li.evan.chao@gmail.com>
@ 2026-09-03 21:11 ` Nathan Bossart <nathandbossart@gmail.com>
0 siblings, 0 replies; 4+ messages in thread
From: Nathan Bossart @ 2026-09-03 21:11 UTC (permalink / raw)
To: Chao Li <li.evan.chao@gmail.com>; +Cc: pgsql-hackers
Committed, thanks for looking!
--
nathan
^ permalink raw reply [nested|flat] 4+ messages in thread
end of thread, other threads:[~2026-09-03 21:11 UTC | newest]
Thread overview: 4+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2026-09-02 18:57 auto-vectorize varbit bitwise operators Nathan Bossart <nathandbossart@gmail.com>
2026-09-02 19:42 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-09-03 02:01 ` Chao Li <li.evan.chao@gmail.com>
2026-09-03 21:11 ` Nathan Bossart <nathandbossart@gmail.com>
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox