Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1x1q9s-005Rdm-0R for pgsql-hackers@arkaria.postgresql.org; Wed, 02 Sep 2026 18:58:04 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.96) (envelope-from ) id 1x1q9r-00DnKh-0W for pgsql-hackers@arkaria.postgresql.org; Wed, 02 Sep 2026 18:58:03 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1x1q9q-00DnKZ-2h for pgsql-hackers@lists.postgresql.org; Wed, 02 Sep 2026 18:58:02 +0000 Received: from mail-oa1-x2e.google.com ([2001:4860:4864:20::2e]) by magus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 (Exim 4.98.2) (envelope-from ) id 1x1q9o-00000002amS-2SBp for pgsql-hackers@postgresql.org; Wed, 02 Sep 2026 18:58:02 +0000 Received: by mail-oa1-x2e.google.com with SMTP id 586e51a60fabf-45e2fd33133so913131fac.1 for ; Wed, 02 Sep 2026 11:58:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1788375478; x=1788980278; darn=postgresql.org; h=content-disposition:content-type:mime-version:message-id:subject:to :from:date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=8rHtM7/PLuyz98nPvirNWkRmf9NeuiNBuLmekY+p/iQ=; b=nU1KNlL2bpwNnjazYczPzoGlJSRQ7taJ7lQcV/pyRkz++7pFvHWDLLNQ2IIw4/gRvI rP2a4izpcczmOL2xSke/8F7WIMPnF1c4X/RPVr2m55B6Gz39iaakI7AH5F+p61Zb1Lb0 Tf71xOIxRdolhOns0mTpzWP5Gi+Ei0mw9GrvdDmTJwLg7Zr4F6C8CnAWfEtFwW15Eo/6 GifeVR9+iFbn+Dkw3SypbIjKHJZtnbkNe08XFnqTd31YuWefWdu69u81iuXwJedA95x3 RA3CrpouU+jJ8lWKC3XtT5ak+bq4M26cr2jd9itDNMKPTZsfWdSLX/nkIISF73r0zlWf wsSA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1788375478; x=1788980278; h=content-disposition:content-type:mime-version:message-id:subject:to :from:date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=8rHtM7/PLuyz98nPvirNWkRmf9NeuiNBuLmekY+p/iQ=; b=RmMPrg5I0FDsdTCJALOtyCk4Md6z4LZ9AXkAbz3bG8sOWqnMc84/3enS8RAwwGtNVE yEt3mif1B5CjzRsGAMlwbKsNpCuIw5CqyE+iZYXVZMQAZEfE1lJeCEJGixJFkL4vwZ6y 7JNDILNkOVZ7v+/BKhL2Hfx+qsnkQQzKMIFibGuLd+BvSlei8bKjwALMTHY1vIgkf3ez 4HpuwxclmEjJpUbpOgd4J8/I3+j6bY9VQ2CfB33k0gVUD30Q4QXIIr0oJis5I2qqLfZn CsvMFjFvuxgnAIGivuKfaM2r6YGxEAbDewfBxt7AiaTDHmrm/UcuUAjZcsZajrTyi/Me i50Q== X-Gm-Message-State: AFuF++mgKNietk7+8r2TUFq9LGjr4sh/QwNBqoE1uWqF3hkJOiiE1Czx DGoinqadRFSvxDrvJKlQwbjf/SgYcWIITpAXF7dk8s2ou7DniPiRbOT+P3RYNA== X-Gm-Gg: AYBFou2BUw1CSYyHso4NzdZHjmRsTKMKT8Y935wwc81isvFaSigplK9Ekh73jGFSuOm O1QMTaz576L3pn5U9Y62LR7BNh7pe8xdZ9dalk1ZKY2aWdTALxgYBvCkGjyG5dua85OwcStn0TX YHYxZa++rvK53EO0d+5ABTcDR/vGh+ihdUkDaXYx+HC3OTKFUitEIPJ6aOnNwH/S2oXwQjVcSZ/ P3ypb5U9pMZhVomoSF3mriC5+LED5ixIHjHsf5gtjKfS/NdS4Jwas3n2noycjFhsZWz/xwa7h3m 5BJjMCIObK/qsy9EZL1tBK2CYJjyYr1KsW5Z5rB0scQGjbfjPbGpxZl+9nBwPAzzOwH9B/zFe++ 80jacQfJXKd003eFUTkWl8ZTiq0eR1vi/TidBE4aWpDf5VL9HcfdorBWENYvHjXtIUzNx3hXHo/ SJlTLmzzw4lF7VG+EWz8THHHLEZxBFetmcYkICPHcJXaHDT6eTYGxaNQ5o444alQU+oIc7IEmrH 6Fni6s1jFOqOURDjEsOx2RXY5pp6jyE4ZE6LO3dT3Fx1PIoRcRCjHPh5cB50y1wSCkbL0oZxGq1 X9KZAzoSI5WSbMY= X-Received: by 2002:a05:6870:1b83:b0:465:127b:2620 with SMTP id 586e51a60fabf-46f878783a2mr6795226fac.17.1788375478565; Wed, 02 Sep 2026 11:57:58 -0700 (PDT) Received: from nathan (162-195-168-172.lightspeed.stlsmo.sbcglobal.net. [162.195.168.172]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-46f30baa0d4sm4834787fac.6.2026.09.02.11.57.52 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 02 Sep 2026 11:57:57 -0700 (PDT) Date: Wed, 2 Sep 2026 13:57:51 -0500 From: Nathan Bossart To: pgsql-hackers@postgresql.org Subject: auto-vectorize varbit bitwise operators Message-ID: MIME-Version: 1.0 Content-Type: multipart/mixed; boundary="/PKgyr8biqo/qZ1W" Content-Disposition: inline List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk --/PKgyr8biqo/qZ1W Content-Type: text/plain; charset=us-ascii Content-Disposition: inline After adding -ftree-vectorize and moving the loop boundary computations to outside the loops, popular compilers will auto-vectorize bit_and(), bit_or(), bitxor(), and bitnot(). My testing indicates this produces some nice speedups, but I haven't yet done anything scientific enough to share. -- nathan --/PKgyr8biqo/qZ1W Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename=v1-0001-Auto-vectorize-the-varbit-bitwise-operators.patch From 027ed15aa336948445918d85d01fed6eef0df6c0 Mon Sep 17 00:00:00 2001 From: Nathan Bossart Date: Wed, 2 Sep 2026 13:52:14 -0500 Subject: [PATCH v1 1/1] Auto-vectorize the varbit bitwise operators. Compile varbit.c with -ftree-vectorize where available, and adjust the loops in bit_and(), bit_or(), bitxor(), and bitnot() so that they are amenable to being auto-vectorized. (Mainly, that involves reading the byte count into a local variable before the loop. As written, the bound is recomputed from the varlena header on every iteration, because the compiler must assume that the loop might modify the header.) --- src/backend/utils/adt/Makefile | 3 ++- src/backend/utils/adt/meson.build | 9 ++++----- src/backend/utils/adt/varbit.c | 25 +++++++++++++++++-------- 3 files changed, 23 insertions(+), 14 deletions(-) diff --git a/src/backend/utils/adt/Makefile b/src/backend/utils/adt/Makefile index 0c7621957c1..94ac5169940 100644 --- a/src/backend/utils/adt/Makefile +++ b/src/backend/utils/adt/Makefile @@ -150,8 +150,9 @@ clean: like.o: like.c like_match.c -# Some code in numeric.c benefits from auto-vectorization +# Some code benefits from auto-vectorization numeric.o: CFLAGS += ${CFLAGS_VECTORIZE} +varbit.o: CFLAGS += ${CFLAGS_VECTORIZE} varlena.o: varlena.c levenshtein.c diff --git a/src/backend/utils/adt/meson.build b/src/backend/utils/adt/meson.build index d793f8145f6..04d6f8aa168 100644 --- a/src/backend/utils/adt/meson.build +++ b/src/backend/utils/adt/meson.build @@ -1,14 +1,14 @@ # Copyright (c) 2022-2026, PostgreSQL Global Development Group -# Some code in numeric.c benefits from auto-vectorization -numeric_backend_lib = static_library('numeric_backend_lib', - 'numeric.c', +# Some code benefits from auto-vectorization +vectorize_backend_lib = static_library('vectorize_backend_lib', + ['numeric.c', 'varbit.c'], dependencies: backend_build_deps, kwargs: internal_lib_args, c_args: vectorize_cflags, ) -backend_link_with += numeric_backend_lib +backend_link_with += vectorize_backend_lib backend_sources += files( 'acl.c', @@ -118,7 +118,6 @@ backend_sources += files( 'tsvector_op.c', 'tsvector_parser.c', 'uuid.c', - 'varbit.c', 'varchar.c', 'varlena.c', 'version.c', diff --git a/src/backend/utils/adt/varbit.c b/src/backend/utils/adt/varbit.c index fe49ba851ac..cbb60e68335 100644 --- a/src/backend/utils/adt/varbit.c +++ b/src/backend/utils/adt/varbit.c @@ -1251,6 +1251,7 @@ bit_and(PG_FUNCTION_ARGS) uint8 *p1, *p2, *r; + size_t nbytes; bitlen1 = VARBITLEN(arg1); bitlen2 = VARBITLEN(arg2); @@ -1267,8 +1268,9 @@ bit_and(PG_FUNCTION_ARGS) p1 = VARBITS(arg1); p2 = VARBITS(arg2); r = VARBITS(result); - for (size_t i = 0; i < VARBITBYTES(arg1); i++) - *r++ = *p1++ & *p2++; + nbytes = VARBITBYTES(arg1); + for (size_t i = 0; i < nbytes; i++) + r[i] = p1[i] & p2[i]; /* Padding is not needed as & of 0 pads is 0 */ @@ -1291,6 +1293,7 @@ bit_or(PG_FUNCTION_ARGS) uint8 *p1, *p2, *r; + size_t nbytes; bitlen1 = VARBITLEN(arg1); bitlen2 = VARBITLEN(arg2); @@ -1306,8 +1309,9 @@ bit_or(PG_FUNCTION_ARGS) p1 = VARBITS(arg1); p2 = VARBITS(arg2); r = VARBITS(result); - for (size_t i = 0; i < VARBITBYTES(arg1); i++) - *r++ = *p1++ | *p2++; + nbytes = VARBITBYTES(arg1); + for (size_t i = 0; i < nbytes; i++) + r[i] = p1[i] | p2[i]; /* Padding is not needed as | of 0 pads is 0 */ @@ -1330,6 +1334,7 @@ bitxor(PG_FUNCTION_ARGS) uint8 *p1, *p2, *r; + size_t nbytes; bitlen1 = VARBITLEN(arg1); bitlen2 = VARBITLEN(arg2); @@ -1346,8 +1351,9 @@ bitxor(PG_FUNCTION_ARGS) p1 = VARBITS(arg1); p2 = VARBITS(arg2); r = VARBITS(result); - for (size_t i = 0; i < VARBITBYTES(arg1); i++) - *r++ = *p1++ ^ *p2++; + nbytes = VARBITBYTES(arg1); + for (size_t i = 0; i < nbytes; i++) + r[i] = p1[i] ^ p2[i]; /* Padding is not needed as ^ of 0 pads is 0 */ @@ -1365,6 +1371,7 @@ bitnot(PG_FUNCTION_ARGS) VarBit *result; uint8 *p, *r; + size_t nbytes; result = (VarBit *) palloc(VARSIZE(arg)); SET_VARSIZE(result, VARSIZE(arg)); @@ -1372,8 +1379,10 @@ bitnot(PG_FUNCTION_ARGS) p = VARBITS(arg); r = VARBITS(result); - for (; p < VARBITEND(arg); p++) - *r++ = ~*p; + nbytes = VARBITBYTES(arg); + for (size_t i = 0; i < nbytes; i++) + r[i] = ~p[i]; + r += nbytes; /* Must zero-pad the result, because extra bits are surely 1's here */ VARBIT_PAD_LAST(result, r); -- 2.55.0 --/PKgyr8biqo/qZ1W--