agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedCrash issue in PG18.5 regression
16+ messages / 6 participants
[nested] [flat]
* Crash issue in PG18.5 regression
@ 2026-08-11 05:16 Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>
0 siblings, 1 reply; 16+ messages in thread
From: Masashi Kamura (Fujitsu) @ 2026-08-11 05:16 UTC (permalink / raw)
To: 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>
Hi,
We found that the program crashes when following the steps below.
1) Create the instance
initdb -D data --encoding=UTF8 --no-locale
2) Execute following SQL
SELECT to_date('01 ŞUB 2010', 'DD TMMON YYYY');
We are analyzing the cause and the following commit seems the cause.
https://github.com/postgres/postgres/commit/011384ba45f
Could you please check this?
FYI)
- This did not occur in PG17. We suspect that this is because PG17 is using `str_toupper()`.
- Build option is as below :
--enable-nls --with-libedit-preferred --with-openssl --with-krb-srvnam=postgres --with-gssapi --with-ldap --with-libcurl --with-libnuma --with-ossp-uuid --with-libxml --with-libxslt --with-perl --with-python --with-tcl --with-tclconfig=/usr/lib64 --with-pam --with-lz4 --with-zstd --enable-tap-tests --with-selinux TCLSH=/usr/bin/tclsh 'CC=gcc ' CFLAGS=-O2 'CPPFLAGS=-DLINUX_OOM_SCORE_ADJ=0 -DLINUX_OOM_ADJ' 'LDFLAGS= -Wl,-rpath,'\''$$ORIGIN/../lib'\'',--enable-new-dtags' --with-llvm LLVM_CONFIG=/usr/bin/llvm-config CLANG=/usr/bin/clang
Regards,
Masashi Kamura
Fujitsu Limited
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-11 06:32 Álvaro Herrera <alvherre@kurilemu.de>
parent: Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>
0 siblings, 2 replies; 16+ messages in thread
From: Álvaro Herrera @ 2026-08-11 06:32 UTC (permalink / raw)
To: Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; +Cc: 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>; Jeff Davis <pgsql@j-davis.com>
On 2026-Aug-11, Masashi Kamura (Fujitsu) wrote:
> Hi,
>
> We found that the program crashes when following the steps below.
>
> 1) Create the instance
> initdb -D data --encoding=UTF8 --no-locale
>
> 2) Execute following SQL
> SELECT to_date('01 ŞUB 2010', 'DD TMMON YYYY');
>
> We are analyzing the cause and the following commit seems the cause.
> https://github.com/postgres/postgres/commit/011384ba45f
>
> Could you please check this?
I confirm that this crashes with my regular build options also, as long
as initdb --no-locale is used. The backtrace from the crash point is
#0 __GI___towupper_l (wc=74, locale=locale@entry=0x0) at ./wctype/wcfuncs_l.c:69
#1 0x0000563291e315b8 in strupper_libc_mb (dest=0x7ffe3664f100 "\002", destsize=80, src=0x5632b0c066d8 "Jan",
srclen=3, locale=0x5632b0c02898) at ../../source/REL_18_STABLE/src/backend/utils/adt/pg_locale_libc.c:398
#2 strupper_libc (dst=dst@entry=0x7ffe3664f100 "\002", dstsize=dstsize@entry=80, src=src@entry=0x5632b0c066d8 "Jan",
srclen=<optimized out>, locale=locale@entry=0x5632b0c02898)
at ../../source/REL_18_STABLE/src/backend/utils/adt/pg_locale_libc.c:147
#3 0x0000563291e2f029 in pg_strupper (dst=dst@entry=0x7ffe3664f100 "\002", dstsize=dstsize@entry=80,
src=src@entry=0x5632b0c066d8 "Jan", srclen=<optimized out>, locale=locale@entry=0x5632b0c02898)
at ../../source/REL_18_STABLE/src/backend/utils/adt/pg_locale.c:1325
The relevant code in src/backend/utils/adt/pg_locale_libc.c's
strupper_libc_mb() from frame 1 is
397 │ for (curr_char = 0; workspace[curr_char] != 0; curr_char++)
398 │ workspace[curr_char] = towupper_l(workspace[curr_char], loc);
where the important detail is that 'loc' is 0, which is not a valid
locale handle.
The locale code is quite the maze, but I'll see if I can find why is the
locale object not initialized.
--
Álvaro Herrera PostgreSQL Developer — https://www.EnterpriseDB.com/
Maybe there's lots of data loss but the records of data loss are also lost.
(Lincoln Yeoh)
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-11 09:41 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Álvaro Herrera <alvherre@kurilemu.de>
1 sibling, 3 replies; 16+ messages in thread
From: Heikki Linnakangas @ 2026-08-11 09:41 UTC (permalink / raw)
To: Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; +Cc: 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>; Jeff Davis <pgsql@j-davis.com>
On 11/08/2026 12:19, Heikki Linnakangas wrote:
> I think the best fix is to make pg_strupper() in REL_18_STABLE also work
> with the C locale. It's an accident waiting to happen if it doesn't.
> (And same for all the other pg_str*() functions, of course)
Like the attached.
There's some code duplication: the strupper_c() function is essentially
the same as asc_toupper(), and str_toupper() wouldn't really need to
have the special case for C locale anymore, it could just rely on
pg_strupper() now. But that's so in 'master' too, so I think cleaning
that up should be left for a separate patch.
- Heikki
Attachments:
[text/x-patch] 0001-Fix-pg_strupper-lower-fold-functions-work-with-C-loc.patch (3.5K, ../../f72a8fbe-7ef1-4cd3-8a4b-fd9106480510@iki.fi/2-0001-Fix-pg_strupper-lower-fold-functions-work-with-C-loc.patch)
download | inline diff:
From bb623fc827b367393e2378846fb12c52a947edfe Mon Sep 17 00:00:00 2001
From: Heikki Linnakangas <heikki.linnakangas@iki.fi>
Date: Tue, 11 Aug 2026 12:32:49 +0300
Subject: [PATCH 1/1] Fix pg_strupper/lower/fold() functions work with C locale
---
src/backend/utils/adt/pg_locale.c | 68 +++++++++++++++++++++++++++++--
1 file changed, 64 insertions(+), 4 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 78243b2c795..2f8d5fea8f2 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1273,11 +1273,64 @@ get_collation_actual_version(char collprovider, const char *collcollate)
return collversion;
}
+/* lowercasing/casefolding in C locale */
+static size_t
+strlower_c(char *dst, size_t dstsize, const char *src, size_t srclen)
+{
+ int i;
+
+ for (i = 0; i < srclen && i < dstsize; i++)
+ dst[i] = pg_ascii_tolower(src[i]);
+ if (i < dstsize)
+ dst[i] = '\0';
+ return srclen;
+}
+
+/* titlecasing in C locale */
+static size_t
+strtitle_c(char *dst, size_t dstsize, const char *src, size_t srclen)
+{
+ bool wasalnum = false;
+ int i;
+
+ for (i = 0; i < srclen && i < dstsize; i++)
+ {
+ char c = src[i];
+
+ if (wasalnum)
+ dst[i] = pg_ascii_tolower(c);
+ else
+ dst[i] = pg_ascii_toupper(c);
+
+ wasalnum = ((c >= '0' && c <= '9') ||
+ (c >= 'A' && c <= 'Z') ||
+ (c >= 'a' && c <= 'z'));
+ }
+ if (i < dstsize)
+ dst[i] = '\0';
+ return srclen;
+}
+
+/* uppercasing in C locale */
+static size_t
+strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
+{
+ int i;
+
+ for (i = 0; i < srclen && i < dstsize; i++)
+ dst[i] = pg_ascii_toupper(src[i]);
+ if (i < dstsize)
+ dst[i] = '\0';
+ return srclen;
+}
+
size_t
pg_strlower(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
- if (locale->provider == COLLPROVIDER_BUILTIN)
+ if (locale->ctype_is_c)
+ return strlower_c(dst, dstsize, src, srclen);
+ else if (locale->provider == COLLPROVIDER_BUILTIN)
return strlower_builtin(dst, dstsize, src, srclen, locale);
#ifdef USE_ICU
else if (locale->provider == COLLPROVIDER_ICU)
@@ -1296,7 +1349,9 @@ size_t
pg_strtitle(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
- if (locale->provider == COLLPROVIDER_BUILTIN)
+ if (locale->ctype_is_c)
+ return strtitle_c(dst, dstsize, src, srclen);
+ else if (locale->provider == COLLPROVIDER_BUILTIN)
return strtitle_builtin(dst, dstsize, src, srclen, locale);
#ifdef USE_ICU
else if (locale->provider == COLLPROVIDER_ICU)
@@ -1315,7 +1370,9 @@ size_t
pg_strupper(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
- if (locale->provider == COLLPROVIDER_BUILTIN)
+ if (locale->ctype_is_c)
+ return strupper_c(dst, dstsize, src, srclen);
+ else if (locale->provider == COLLPROVIDER_BUILTIN)
return strupper_builtin(dst, dstsize, src, srclen, locale);
#ifdef USE_ICU
else if (locale->provider == COLLPROVIDER_ICU)
@@ -1334,7 +1391,10 @@ size_t
pg_strfold(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
- if (locale->provider == COLLPROVIDER_BUILTIN)
+ /* in the C locale, casefolding is the same as lowercasing */
+ if (locale->ctype_is_c)
+ return strlower_c(dst, dstsize, src, srclen);
+ else if (locale->provider == COLLPROVIDER_BUILTIN)
return strfold_builtin(dst, dstsize, src, srclen, locale);
#ifdef USE_ICU
else if (locale->provider == COLLPROVIDER_ICU)
--
2.47.3
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-11 10:23 Álvaro Herrera <alvherre@kurilemu.de>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
2 siblings, 1 reply; 16+ messages in thread
From: Álvaro Herrera @ 2026-08-11 10:23 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>; Jeff Davis <pgsql@j-davis.com>
On 2026-Aug-11, Heikki Linnakangas wrote:
> On 11/08/2026 12:19, Heikki Linnakangas wrote:
> > I think the best fix is to make pg_strupper() in REL_18_STABLE also work
> > with the C locale. It's an accident waiting to happen if it doesn't.
> > (And same for all the other pg_str*() functions, of course)
>
> Like the attached.
Hm, this looks really similar to what I wrote (attached here for the
curious), except you chose a different layer (directly in pg_locale.c,
whereas I put mine under the libc implementation). Probably yours is
the better choice since it also covers the builtin provider.
> There's some code duplication: the strupper_c() function is essentially the
> same as asc_toupper(), and str_toupper() wouldn't really need to have the
> special case for C locale anymore, it could just rely on pg_strupper() now.
> But that's so in 'master' too, so I think cleaning that up should be left
> for a separate patch.
Agreed.
It's unclear to me how to get the strtitle() thing called, since the
only caller seems to be str_initcap() which will use asc_initcap anyway.
Maybe such a cleanup should remove some part of this code.
--
Álvaro Herrera 48°01'N 7°57'E — https://www.EnterpriseDB.com/
Attachments:
[text/x-diff] 0001-Hardcode-str-lower-upper-title-for-the-C-locale.patch (4.5K, ../../anr2M1Wj_31k6aPD@alvherre.pgsql/2-0001-Hardcode-str-lower-upper-title-for-the-C-locale.patch)
download | inline diff:
From e6df99c6ba63216297630a3b7057fb5ac1735db6 Mon Sep 17 00:00:00 2001
From: =?UTF-8?q?=C3=81lvaro=20Herrera?= <alvherre@kurilemu.de>
Date: Tue, 11 Aug 2026 12:11:35 +0200
Subject: [PATCH] Hardcode str{lower,upper,title}() for the C locale
This avoids a crash when those functions are called directly rather than
via str_to{lower,upper,title} directly.
---
src/backend/utils/adt/pg_locale_libc.c | 64 +++++++++++++++++++++++++-
1 file changed, 62 insertions(+), 2 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale_libc.c b/src/backend/utils/adt/pg_locale_libc.c
index 31274f069d2..afacb727efe 100644
--- a/src/backend/utils/adt/pg_locale_libc.c
+++ b/src/backend/utils/adt/pg_locale_libc.c
@@ -66,18 +66,24 @@ static int strncoll_libc_win32_utf8(const char *arg1, ssize_t len1,
pg_locale_t locale);
#endif
+static size_t strlower_libc_c(char *dest, size_t dstsize,
+ const char *src, ssize_t srclen);
static size_t strlower_libc_sb(char *dest, size_t destsize,
const char *src, ssize_t srclen,
pg_locale_t locale);
static size_t strlower_libc_mb(char *dest, size_t destsize,
const char *src, ssize_t srclen,
pg_locale_t locale);
+static size_t strtitle_libc_c(char *dest, size_t dstsize,
+ const char *src, ssize_t srclen);
static size_t strtitle_libc_sb(char *dest, size_t destsize,
const char *src, ssize_t srclen,
pg_locale_t locale);
static size_t strtitle_libc_mb(char *dest, size_t destsize,
const char *src, ssize_t srclen,
pg_locale_t locale);
+static size_t strupper_libc_c(char *dest, size_t dstsize,
+ const char *src, ssize_t srclen);
static size_t strupper_libc_sb(char *dest, size_t destsize,
const char *src, ssize_t srclen,
pg_locale_t locale);
@@ -123,7 +129,9 @@ size_t
strlower_libc(char *dst, size_t dstsize, const char *src,
ssize_t srclen, pg_locale_t locale)
{
- if (pg_database_encoding_max_length() > 1)
+ if (locale->ctype_is_c)
+ return strlower_libc_c(dst, dstsize, src, srclen);
+ else if (pg_database_encoding_max_length() > 1)
return strlower_libc_mb(dst, dstsize, src, srclen, locale);
else
return strlower_libc_sb(dst, dstsize, src, srclen, locale);
@@ -133,6 +141,8 @@ size_t
strtitle_libc(char *dst, size_t dstsize, const char *src,
ssize_t srclen, pg_locale_t locale)
{
+ if (locale->ctype_is_c)
+ return strtitle_libc_c(dst, dstsize, src, srclen);
if (pg_database_encoding_max_length() > 1)
return strtitle_libc_mb(dst, dstsize, src, srclen, locale);
else
@@ -143,12 +153,26 @@ size_t
strupper_libc(char *dst, size_t dstsize, const char *src,
ssize_t srclen, pg_locale_t locale)
{
- if (pg_database_encoding_max_length() > 1)
+ if (locale->ctype_is_c)
+ return strupper_libc_c(dst, dstsize, src, srclen);
+ else if (pg_database_encoding_max_length() > 1)
return strupper_libc_mb(dst, dstsize, src, srclen, locale);
else
return strupper_libc_sb(dst, dstsize, src, srclen, locale);
}
+static size_t
+strlower_libc_c(char *dest, size_t dstsize, const char *src, ssize_t srclen)
+{
+ int i;
+
+ for (i = 0; i < srclen && i < dstsize; i++)
+ dest[i] = pg_ascii_tolower(src[i]);
+ if (i < dstsize)
+ dest[i] = '\0';
+ return srclen;
+}
+
static size_t
strlower_libc_sb(char *dest, size_t destsize, const char *src, ssize_t srclen,
pg_locale_t locale)
@@ -280,6 +304,30 @@ strtitle_libc_sb(char *dest, size_t destsize, const char *src, ssize_t srclen,
return srclen;
}
+static size_t
+strtitle_libc_c(char *dest, size_t dstsize, const char *src, ssize_t srclen)
+{
+ bool wasalnum = false;
+ int i;
+
+ for (i = 0; i < srclen && i < dstsize; i++)
+ {
+ char c = src[i];
+
+ if (wasalnum)
+ dest[i] = pg_ascii_tolower(c);
+ else
+ dest[i] = pg_ascii_toupper(c);
+
+ wasalnum = ((c >= '0' && c <= '9') ||
+ (c >= 'A' && c <= 'Z') ||
+ (c >= 'a' && c <= 'z'));
+ }
+ if (i < dstsize)
+ dest[i] = '\0';
+ return srclen;
+}
+
static size_t
strtitle_libc_mb(char *dest, size_t destsize, const char *src, ssize_t srclen,
pg_locale_t locale)
@@ -335,6 +383,18 @@ strtitle_libc_mb(char *dest, size_t destsize, const char *src, ssize_t srclen,
return result_size;
}
+static size_t
+strupper_libc_c(char *dest, size_t dstsize, const char *src, ssize_t srclen)
+{
+ int i;
+
+ for (i = 0; i < srclen && i < dstsize; i++)
+ dest[i] = pg_ascii_toupper(src[i]);
+ if (i < dstsize)
+ dest[i] = '\0';
+ return srclen;
+}
+
static size_t
strupper_libc_sb(char *dest, size_t destsize, const char *src, ssize_t srclen,
pg_locale_t locale)
--
2.47.3
^ permalink raw reply [nested|flat] 16+ messages in thread
* RE: Crash issue in PG18.5 regression
@ 2026-08-11 11:23 Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>
parent: Álvaro Herrera <alvherre@kurilemu.de>
0 siblings, 0 replies; 16+ messages in thread
From: Masashi Kamura (Fujitsu) @ 2026-08-11 11:23 UTC (permalink / raw)
To: 'Álvaro Herrera' <alvherre@kurilemu.de>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>; Jeff Davis <pgsql@j-davis.com>
Hi Heikki san, Álvaro san,
Thank you for sharing the patch.
I have confirmed both patches passed the test without any issues.
Regards,
Masashi Kamura
Fujitsu Limited
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-11 16:55 Andres Freund <andres@anarazel.de>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
2 siblings, 1 reply; 16+ messages in thread
From: Andres Freund @ 2026-08-11 16:55 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>; Jeff Davis <pgsql@j-davis.com>
Hi,
On 2026-08-11 12:41:49 +0300, Heikki Linnakangas wrote:
> On 11/08/2026 12:19, Heikki Linnakangas wrote:
> > I think the best fix is to make pg_strupper() in REL_18_STABLE also work
> > with the C locale. It's an accident waiting to happen if it doesn't.
> > (And same for all the other pg_str*() functions, of course)
>
> Like the attached.
>
> There's some code duplication: the strupper_c() function is essentially the
> same as asc_toupper(), and str_toupper() wouldn't really need to have the
> special case for C locale anymore, it could just rely on pg_strupper() now.
> But that's so in 'master' too, so I think cleaning that up should be left
> for a separate patch.
I wonder if we also ought to do something about the other uses of
pg_locale_t->info.lt in pg_locale_libc.c? That's at least strncoll_libc(),
strnxfrm_libc(). I think all the in-core callers guard them, but the
protection seems mighty far away in some cases.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-11 17:02 Tom Lane <tgl@sss.pgh.pa.us>
parent: Andres Freund <andres@anarazel.de>
0 siblings, 1 reply; 16+ messages in thread
From: Tom Lane @ 2026-08-11 17:02 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>; Jeff Davis <pgsql@j-davis.com>
Andres Freund <andres@anarazel.de> writes:
> On 2026-08-11 12:41:49 +0300, Heikki Linnakangas wrote:
>> Like the attached.
> I wonder if we also ought to do something about the other uses of
> pg_locale_t->info.lt in pg_locale_libc.c? That's at least strncoll_libc(),
> strnxfrm_libc(). I think all the in-core callers guard them, but the
> protection seems mighty far away in some cases.
We don't have a lot of time to think about this, and AFAICS the
security patch only added calls to pg_strupper and pg_strlower.
So as long as those are protected I'm content to ship.
regards, tom lane
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-11 18:26 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Tom Lane <tgl@sss.pgh.pa.us>
0 siblings, 0 replies; 16+ messages in thread
From: Heikki Linnakangas @ 2026-08-11 18:26 UTC (permalink / raw)
To: Tom Lane <tgl@sss.pgh.pa.us>; Andres Freund <andres@anarazel.de>; +Cc: Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>; Jeff Davis <pgsql@j-davis.com>
On 11/08/2026 20:02, Tom Lane wrote:
> Andres Freund <andres@anarazel.de> writes:
>> On 2026-08-11 12:41:49 +0300, Heikki Linnakangas wrote:
>>> Like the attached.
>
>> I wonder if we also ought to do something about the other uses of
>> pg_locale_t->info.lt in pg_locale_libc.c? That's at least strncoll_libc(),
>> strnxfrm_libc(). I think all the in-core callers guard them, but the
>> protection seems mighty far away in some cases.
>
> We don't have a lot of time to think about this, and AFAICS the
> security patch only added calls to pg_strupper and pg_strlower.
> So as long as those are protected I'm content to ship.
+1. It'd be good to look at those, but not right now.
I have pushed the fix.
- Heikki
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-11 18:31 Andres Freund <andres@anarazel.de>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
2 siblings, 2 replies; 16+ messages in thread
From: Andres Freund @ 2026-08-11 18:31 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>; Jeff Davis <pgsql@j-davis.com>
Hi,
On 2026-08-11 12:41:49 +0300, Heikki Linnakangas wrote:
> +/* lowercasing/casefolding in C locale */
> +static size_t
> +strlower_c(char *dst, size_t dstsize, const char *src, size_t srclen)
> +{
> + int i;
> +
> + for (i = 0; i < srclen && i < dstsize; i++)
> + dst[i] = pg_ascii_tolower(src[i]);
> + if (i < dstsize)
> + dst[i] = '\0';
> + return srclen;
> +}
Hm. If I infer the pg_strlower() API correctly - it's utterly underdocumented
- it seems to be inteded to support a few things in 18:
1) srclen = -1 works
Inferred from unicode_strlower()'s comment:
* String src must be encoded in UTF-8. If srclen < 0, src must be
* NUL-terminated.
Also note that srclen is ssize_t. This changed in 6d22c67c3bf5 (recently).
2) The required length for the conversion is returned, even if the destination
is too short (including when dstlen = 0)
* Result string is stored in dst, truncating if larger than dstsize. If
* dstsize is greater than the result length, dst will be NUL-terminated;
* otherwise not.
*
* If dstsize is zero, dst may be NULL. This is useful for calculating the
* required buffer size before allocating.
It's really a guessing game though, due to the religious under documentation
of the generic functions. Why does unicode_strlower() have docs, but
pg_strlower() does not?
Both don't seem quite right given this implementation.
I don't quite know whether we need to fix these, given the lack of problematic
uses in tree, the time pressure, but it also seems like a recipe for future
disaster to leave it like this.
I'd also make i size_t, given that the input is size_t. Perhaps practically
no problem, but I see no reason to not use size_t here.
It also seems like we really ought to have an actually reachable, currently
crashing, to_date() call in the tests? It seems concerning that
seq_search_localized(), casefold_str_cmp() are completely uncovered today, and
quite obviously we can't be relied upon to get this right.
https://coverage.postgresql.org/src/backend/utils/adt/formatting.c.gcov.html#L2379
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-11 19:20 Jeff Davis <pgsql@j-davis.com>
parent: Álvaro Herrera <alvherre@kurilemu.de>
1 sibling, 0 replies; 16+ messages in thread
From: Jeff Davis @ 2026-08-11 19:20 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; +Cc: 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>
On Tue, 2026-08-11 at 12:19 +0300, Heikki Linnakangas wrote:
> Interestingly this only fails on REL_18_STABLE. On REL_19_STABLE,
> pg_strupper() checks if locale->ctype is NULL, and does the
> equivalent
> of asc_toupper() internally.
Ugh. I should have backported 1476028225, sorry :-(
Thank you for catching it quickly, Masashi!
Regards,
Jeff Davis
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-11 20:04 Tom Lane <tgl@sss.pgh.pa.us>
parent: Andres Freund <andres@anarazel.de>
1 sibling, 1 reply; 16+ messages in thread
From: Tom Lane @ 2026-08-11 20:04 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>; Jeff Davis <pgsql@j-davis.com>
Andres Freund <andres@anarazel.de> writes:
> It also seems like we really ought to have an actually reachable, currently
> crashing, to_date() call in the tests? It seems concerning that
> seq_search_localized(), casefold_str_cmp() are completely uncovered today, and
> quite obviously we can't be relied upon to get this right.
All that code is reached when I run the core regression tests under
LANG=C.utf8 or LANG=en_US.utf8, except for the "As last resort"
stanza at the bottom of seq_search_localized()'s loop. I suppose the
coverage.postgresql.org animal is either not Linux or doesn't test
any UTF8 encoding, but that's not the fault of our test cases, and
it doesn't reflect what I think actually happens in the buildfarm.
Yeah, it'd be good if we could devise a test case that reaches the
"As last resort" bit, but that's irrelevant to the current problem.
The reason we failed to notice this sooner is that the crash is only
reached with (a) locale = "C" and (b) either a multi-byte encoding,
so that we reach strupper_libc_mb, or a single-byte encoding with
some high-bit-set characters, so that strupper_libc_sb invokes libc.
The regression test cases that might have noticed this are in
collate.linux.utf8.sql, so we need locale = "C" + encoding = UTF8 +
a Linux test machine that has a reasonable set of locales installed.
That would have been enough to find it, except that the buildfarm
client doesn't have any easy way to test locale = "C" with
encoding = UTF8. It will test locale = "C" with encoding SQL_ASCII,
which doesn't run collate.linux.utf8.sql, and it will test other
cases as set up by the machine owner, but there's no way to tell it
to use that specific locale+encoding combination. I've tried
"LANG=C.utf8", but that doesn't reach the crash, probably because
it doesn't cause us to take the locale_is_c optimization paths.
(Should it? I'm unsure.)
So I'm not seeing a huge failure to test here. We missed a very
narrow combination of cases.
regards, tom lane
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-11 23:05 Andres Freund <andres@anarazel.de>
parent: Tom Lane <tgl@sss.pgh.pa.us>
0 siblings, 0 replies; 16+ messages in thread
From: Andres Freund @ 2026-08-11 23:05 UTC (permalink / raw)
To: Tom Lane <tgl@sss.pgh.pa.us>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>; Jeff Davis <pgsql@j-davis.com>
Hi,
On 2026-08-11 16:04:15 -0400, Tom Lane wrote:
> Andres Freund <andres@anarazel.de> writes:
> > It also seems like we really ought to have an actually reachable, currently
> > crashing, to_date() call in the tests? It seems concerning that
> > seq_search_localized(), casefold_str_cmp() are completely uncovered today, and
> > quite obviously we can't be relied upon to get this right.
>
> All that code is reached when I run the core regression tests under
> LANG=C.utf8 or LANG=en_US.utf8, except for the "As last resort"
> stanza at the bottom of seq_search_localized()'s loop. I suppose the
> coverage.postgresql.org animal is either not Linux or doesn't test
> any UTF8 encoding, but that's not the fault of our test cases, and
> it doesn't reflect what I think actually happens in the buildfarm.
I tested locally on a database that triggered the problem, the skip condition
query was true, which lead me to hastily misread what the query is trying to
do. Turns out the reason it doesn't run here - and likely the reason that it
doesn't run for coverage.pg.o - is that the test afaict *never* matches for a
meson build :(
no_enc2[3902736][1]=# SELECT version();
┌────────────────────────────────────────────────────────────────────┐
│ version │
├────────────────────────────────────────────────────────────────────┤
│ PostgreSQL 20devel on x86_64-linux, compiled by gcc-16.1.0, 64-bit │
└────────────────────────────────────────────────────────────────────┘
(1 row)
That will obviously never match "linux-gnu".
I guess I should start a separate thread about that.
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-14 05:17 Jeff Davis <pgsql@j-davis.com>
parent: Andres Freund <andres@anarazel.de>
1 sibling, 1 reply; 16+ messages in thread
From: Jeff Davis @ 2026-08-14 05:17 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>
On Tue, 2026-08-11 at 14:31 -0400, Andres Freund wrote:
> Hm. If I infer the pg_strlower() API correctly - it's utterly
> underdocumented
Agreed. Patch attached.
> I'd also make i size_t, given that the input is size_t. Perhaps
> practically
> no problem, but I see no reason to not use size_t here.
Patch attached for that, too.
I also attached patches to make all the functions work with
collate_is_c, and fixed up the -1 API in 18.
Regards,
Jeff Davis
Attachments:
[text/x-patch] vPG18-0001-Fixup-5f003855e7-for-srclen-0.patch (1.9K, ../../3d2bf8bebb21ded729431fa25c92d2425fa615f5.camel@j-davis.com/2-vPG18-0001-Fixup-5f003855e7-for-srclen-0.patch)
download | inline diff:
From 3dabc8c4ccc9b4b0d6dbe8f87353f511d6862bb4 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Thu, 13 Aug 2026 20:57:20 -0700
Subject: [PATCH vPG18 1/4] Fixup 5f003855e7 for srclen < 0.
No actual problem because no callers used that aspect of the API.
Only commit to 18, because that part of the API was removed in commit
6d22c67c3b.
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
---
src/backend/utils/adt/pg_locale.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 2f8d5fea8f2..9c721efb52f 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1328,6 +1328,8 @@ size_t
pg_strlower(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+
if (locale->ctype_is_c)
return strlower_c(dst, dstsize, src, srclen);
else if (locale->provider == COLLPROVIDER_BUILTIN)
@@ -1349,6 +1351,8 @@ size_t
pg_strtitle(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+
if (locale->ctype_is_c)
return strtitle_c(dst, dstsize, src, srclen);
else if (locale->provider == COLLPROVIDER_BUILTIN)
@@ -1370,6 +1374,8 @@ size_t
pg_strupper(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+
if (locale->ctype_is_c)
return strupper_c(dst, dstsize, src, srclen);
else if (locale->provider == COLLPROVIDER_BUILTIN)
@@ -1391,6 +1397,8 @@ size_t
pg_strfold(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+
/* in the C locale, casefolding is the same as lowercasing */
if (locale->ctype_is_c)
return strlower_c(dst, dstsize, src, srclen);
--
2.43.0
[text/x-patch] vPG18-0002-pg_locale.c-unicode_case.c-use-size_t-for-iter.patch (2.3K, ../../3d2bf8bebb21ded729431fa25c92d2425fa615f5.camel@j-davis.com/3-vPG18-0002-pg_locale.c-unicode_case.c-use-size_t-for-iter.patch)
download | inline diff:
From 6f11163f52b320f1b4ef1cb66a06a73f210de46f Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Thu, 13 Aug 2026 21:12:34 -0700
Subject: [PATCH vPG18 2/4] pg_locale.c, unicode_case.c: use size_t for
iteration.
No actual problem, just cleanup. Only relevant to 18 and 19.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
---
src/backend/utils/adt/pg_locale.c | 6 +++---
src/common/unicode_case.c | 4 ++--
2 files changed, 5 insertions(+), 5 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 9c721efb52f..8e79ccbd9f6 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1277,7 +1277,7 @@ get_collation_actual_version(char collprovider, const char *collcollate)
static size_t
strlower_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
dst[i] = pg_ascii_tolower(src[i]);
@@ -1291,7 +1291,7 @@ static size_t
strtitle_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
bool wasalnum = false;
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
{
@@ -1315,7 +1315,7 @@ strtitle_c(char *dst, size_t dstsize, const char *src, size_t srclen)
static size_t
strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
dst[i] = pg_ascii_toupper(src[i]);
diff --git a/src/common/unicode_case.c b/src/common/unicode_case.c
index 8639b203e0c..86e0b57d7d3 100644
--- a/src/common/unicode_case.c
+++ b/src/common/unicode_case.c
@@ -341,7 +341,7 @@ check_final_sigma(const unsigned char *str, size_t len, size_t offset)
int ulen;
/* iterate backwards looking for preceding character */
- for (int i = offset; i > 0;)
+ for (size_t i = offset; i > 0;)
{
/* skip backwards through continuation bytes */
i--;
@@ -369,7 +369,7 @@ check_final_sigma(const unsigned char *str, size_t len, size_t offset)
ulen = utf8_mblen((const unsigned char *) str + offset);
/* iterate forward looking for following character */
- for (int i = offset + ulen; i < len;)
+ for (size_t i = offset + ulen; i < len;)
{
ulen = utf8_mblen((const unsigned char *) str + i);
--
2.43.0
[text/x-patch] vPG18-0003-Add-missing-comments-in-pg_locale.c.patch (6.4K, ../../3d2bf8bebb21ded729431fa25c92d2425fa615f5.camel@j-davis.com/4-vPG18-0003-Add-missing-comments-in-pg_locale.c.patch)
download | inline diff:
From 70ebdd061f9ae158b1674453e4a6e03db1b74815 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:32:19 -0700
Subject: [PATCH vPG18 3/4] Add missing comments in pg_locale.c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v3nniwcrxejmcfvz56xbd22hphprqleuornd6hqkmw2bl7kgmz@cnytz2ee5ltk
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 76 +++++++++++++++++++++++++------
1 file changed, 63 insertions(+), 13 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 8e79ccbd9f6..c21619f85cd 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1324,6 +1324,20 @@ strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
return srclen;
}
+/*
+ * pg_strlower()
+ *
+ * Convert src to lowercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If srclen is
+ * -1, src must be NUL-terminated. If dstsize is zero, dst may be NULL, which
+ * is useful for calculating the required buffer size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strlower(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
@@ -1347,6 +1361,20 @@ pg_strlower(char *dst, size_t dstsize, const char *src, ssize_t srclen,
return 0; /* keep compiler quiet */
}
+/*
+ * pg_strtitle()
+ *
+ * Convert src to titlecase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If srclen is
+ * -1, src must be NUL-terminated. If dstsize is zero, dst may be NULL, which
+ * is useful for calculating the required buffer size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strtitle(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
@@ -1370,6 +1398,20 @@ pg_strtitle(char *dst, size_t dstsize, const char *src, ssize_t srclen,
return 0; /* keep compiler quiet */
}
+/*
+ * pg_strupper()
+ *
+ * Convert src to uppercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If srclen is
+ * -1, src must be NUL-terminated. If dstsize is zero, dst may be NULL, which
+ * is useful for calculating the required buffer size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strupper(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
@@ -1393,6 +1435,19 @@ pg_strupper(char *dst, size_t dstsize, const char *src, ssize_t srclen,
return 0; /* keep compiler quiet */
}
+/*
+ * pg_strfold()
+ *
+ * Casefold src, and return the result length (not including terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If srclen is
+ * -1, src must be NUL-terminated. If dstsize is zero, dst may be NULL, which
+ * is useful for calculating the required buffer size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strfold(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
@@ -1432,12 +1487,10 @@ pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
/*
* pg_strncoll
*
- * Call ucol_strcollUTF8(), ucol_strcoll(), strcoll_l() or wcscoll_l() as
- * appropriate for the given locale, platform, and database encoding. If the
- * locale is not specified, use the database collation.
+ * Compare strings according to the given locale.
*
- * The input strings must be encoded in the database encoding. If an input
- * string is NUL-terminated, its length may be specified as -1.
+ * Strings must be encoded in the database encoding with no embedded NULs. If
+ * an input string is NUL-terminated, its length may be specified as -1.
*
* The caller is responsible for breaking ties if the collation is
* deterministic; this maintains consistency with pg_strnxfrm(), which cannot
@@ -1453,9 +1506,6 @@ pg_strncoll(const char *arg1, ssize_t len1, const char *arg2, ssize_t len2,
/*
* Return true if the collation provider supports pg_strxfrm() and
* pg_strnxfrm(); otherwise false.
- *
- *
- * No similar problem is known for the ICU provider.
*/
bool
pg_strxfrm_enabled(pg_locale_t locale)
@@ -1486,9 +1536,9 @@ pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
* ordinary strcmp() on transformed strings is equivalent to pg_strcoll() on
* untransformed strings.
*
- * The input string must be encoded in the database encoding. If the input
- * string is NUL-terminated, its length may be specified as -1. If 'destsize'
- * is zero, 'dest' may be NULL.
+ * String must be encoded in the database encoding with no embedded NULs. If
+ * srclen is -1, src must be NUL-terminated. If 'destsize' is zero, 'dest'
+ * may be NULL.
*
* Not all providers support pg_strnxfrm() safely. The caller should check
* pg_strxfrm_enabled() first, otherwise this function may return wrong
@@ -1534,8 +1584,8 @@ pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
* memcmp() on the byte sequence is equivalent to pg_strncoll() on
* untransformed strings. The result is not nul-terminated.
*
- * The input string must be encoded in the database encoding. If the input
- * string is NUL-terminated, its length may be specified as -1.
+ * String must be encoded in the database encoding with no embedded NULs. If
+ * the input string is NUL-terminated, its length may be specified as -1.
*
* Not all providers support pg_strnxfrm_prefix() safely. The caller should
* check pg_strxfrm_prefix_enabled() first, otherwise this function may return
--
2.43.0
[text/x-patch] vPG18-0004-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch (4.0K, ../../3d2bf8bebb21ded729431fa25c92d2425fa615f5.camel@j-davis.com/5-vPG18-0004-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch)
download | inline diff:
From 27e336a2add2660fa28ba410dc8888aac7be9fb3 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:46:48 -0700
Subject: [PATCH vPG18 4/4] Ensure all pg_locale.h APIs work with collate_is_c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 64 ++++++++++++++++++++++++++++---
1 file changed, 58 insertions(+), 6 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index c21619f85cd..ad9f416ec13 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1481,7 +1481,10 @@ pg_strfold(char *dst, size_t dstsize, const char *src, ssize_t srclen,
int
pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
{
- return locale->collate->strncoll(arg1, -1, arg2, -1, locale);
+ if (locale->collate == NULL)
+ return strcmp(arg1, arg2);
+ else
+ return locale->collate->strncoll(arg1, -1, arg2, -1, locale);
}
/*
@@ -1500,7 +1503,20 @@ int
pg_strncoll(const char *arg1, ssize_t len1, const char *arg2, ssize_t len2,
pg_locale_t locale)
{
- return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
+ if (locale->collate == NULL)
+ {
+ int result;
+
+ len1 = (len1 < 0) ? strlen(arg1) : len1;
+ len2 = (len2 < 0) ? strlen(arg2) : len2;
+ result = memcmp(arg1, arg2, Min(len1, len2));
+
+ if ((result == 0) && (len1 != len2))
+ result = (len1 < len2) ? -1 : 1;
+ return result;
+ }
+ else
+ return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
}
/*
@@ -1510,6 +1526,9 @@ pg_strncoll(const char *arg1, ssize_t len1, const char *arg2, ssize_t len2,
bool
pg_strxfrm_enabled(pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ return true;
+
/*
* locale->collate->strnxfrm is still a required method, even if it may
* have the wrong behavior, because the planner uses it for estimates in
@@ -1526,7 +1545,10 @@ pg_strxfrm_enabled(pg_locale_t locale)
size_t
pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
{
- return locale->collate->strnxfrm(dest, destsize, src, -1, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strnxfrm(dest, destsize, src, -1, locale);
}
/*
@@ -1552,6 +1574,18 @@ size_t
pg_strnxfrm(char *dest, size_t destsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ {
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+
+ if (destsize > srclen)
+ {
+ memcpy(dest, src, srclen);
+ dest[srclen] = '\0';
+ }
+
+ return srclen;
+ }
return locale->collate->strnxfrm(dest, destsize, src, srclen, locale);
}
@@ -1562,7 +1596,10 @@ pg_strnxfrm(char *dest, size_t destsize, const char *src, ssize_t srclen,
bool
pg_strxfrm_prefix_enabled(pg_locale_t locale)
{
- return (locale->collate->strnxfrm_prefix != NULL);
+ if (locale->collate == NULL)
+ return true;
+ else
+ return (locale->collate->strnxfrm_prefix != NULL);
}
/*
@@ -1574,7 +1611,10 @@ size_t
pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
pg_locale_t locale)
{
- return locale->collate->strnxfrm_prefix(dest, destsize, src, -1, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm_prefix(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strnxfrm_prefix(dest, destsize, src, -1, locale);
}
/*
@@ -1599,7 +1639,19 @@ size_t
pg_strnxfrm_prefix(char *dest, size_t destsize, const char *src,
ssize_t srclen, pg_locale_t locale)
{
- return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
+ if (locale->collate == NULL)
+ {
+ size_t len;
+
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+ len = Min(srclen, destsize);
+
+ if (destsize > 0)
+ memcpy(dest, src, len);
+ return len;
+ }
+ else
+ return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
}
/*
--
2.43.0
[text/x-patch] vPG19-0001-pg_locale.c-unicode_case.c-use-size_t-for-iter.patch (2.3K, ../../3d2bf8bebb21ded729431fa25c92d2425fa615f5.camel@j-davis.com/6-vPG19-0001-pg_locale.c-unicode_case.c-use-size_t-for-iter.patch)
download | inline diff:
From c1a532d38ae63f27ff705a6bbdec78ca5aed5b6f Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Thu, 13 Aug 2026 21:12:34 -0700
Subject: [PATCH vPG19 1/3] pg_locale.c, unicode_case.c: use size_t for
iteration.
No actual problem, just cleanup. Only relevant to 18 and 19.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
---
src/backend/utils/adt/pg_locale.c | 6 +++---
src/common/unicode_case.c | 4 ++--
2 files changed, 5 insertions(+), 5 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 11d48a3916e..9eb99487e57 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1270,7 +1270,7 @@ get_collation_actual_version(char collprovider, const char *collcollate)
static size_t
strlower_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
dst[i] = pg_ascii_tolower(src[i]);
@@ -1284,7 +1284,7 @@ static size_t
strtitle_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
bool wasalnum = false;
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
{
@@ -1308,7 +1308,7 @@ strtitle_c(char *dst, size_t dstsize, const char *src, size_t srclen)
static size_t
strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
dst[i] = pg_ascii_toupper(src[i]);
diff --git a/src/common/unicode_case.c b/src/common/unicode_case.c
index dd5b3ba86d0..744b9116b12 100644
--- a/src/common/unicode_case.c
+++ b/src/common/unicode_case.c
@@ -336,7 +336,7 @@ check_final_sigma(const unsigned char *str, size_t len, size_t offset)
int ulen;
/* iterate backwards looking for preceding character */
- for (int i = offset; i > 0;)
+ for (size_t i = offset; i > 0;)
{
/* skip backwards through continuation bytes */
i--;
@@ -364,7 +364,7 @@ check_final_sigma(const unsigned char *str, size_t len, size_t offset)
ulen = utf8_mblen((const unsigned char *) str + offset);
/* iterate forward looking for following character */
- for (int i = offset + ulen; i < len;)
+ for (size_t i = offset + ulen; i < len;)
{
ulen = utf8_mblen((const unsigned char *) str + i);
--
2.43.0
[text/x-patch] vPG19-0002-Add-missing-comments-in-pg_locale.c.patch (5.9K, ../../3d2bf8bebb21ded729431fa25c92d2425fa615f5.camel@j-davis.com/7-vPG19-0002-Add-missing-comments-in-pg_locale.c.patch)
download | inline diff:
From 81ad1161ed864e1572e5319eecec037670e940b4 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:32:19 -0700
Subject: [PATCH vPG19 2/3] Add missing comments in pg_locale.c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v3nniwcrxejmcfvz56xbd22hphprqleuornd6hqkmw2bl7kgmz@cnytz2ee5ltk
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 70 ++++++++++++++++++++++++++-----
1 file changed, 60 insertions(+), 10 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 9eb99487e57..da32590396f 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1317,6 +1317,20 @@ strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
return srclen;
}
+/*
+ * pg_strlower()
+ *
+ * Convert src to lowercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strlower(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1327,6 +1341,20 @@ pg_strlower(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strlower(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strtitle()
+ *
+ * Convert src to titlecase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strtitle(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1337,6 +1365,20 @@ pg_strtitle(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strtitle(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strupper()
+ *
+ * Convert src to uppercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strupper(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1347,6 +1389,19 @@ pg_strupper(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strupper(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strfold()
+ *
+ * Casefold src, and return the result length (not including terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strfold(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1392,11 +1447,9 @@ pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
/*
* pg_strncoll
*
- * Call ucol_strcollUTF8(), ucol_strcoll(), strcoll_l() or wcscoll_l() as
- * appropriate for the given locale, platform, and database encoding. If the
- * locale is not specified, use the database collation.
+ * Compare strings according to the given locale.
*
- * The input strings must be encoded in the database encoding.
+ * Strings must be encoded in the database encoding with no embedded NULs.
*
* The caller is responsible for breaking ties if the collation is
* deterministic; this maintains consistency with pg_strnxfrm(), which cannot
@@ -1412,9 +1465,6 @@ pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
/*
* Return true if the collation provider supports pg_strxfrm() and
* pg_strnxfrm(); otherwise false.
- *
- *
- * No similar problem is known for the ICU provider.
*/
bool
pg_strxfrm_enabled(pg_locale_t locale)
@@ -1445,8 +1495,8 @@ pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
* ordinary strcmp() on transformed strings is equivalent to pg_strcoll() on
* untransformed strings.
*
- * The input string must be encoded in the database encoding. If 'destsize' is
- * zero, 'dest' may be NULL.
+ * String must be encoded in the database encoding with no embedded NULs. If
+ * 'destsize' is zero, 'dest' may be NULL.
*
* Not all providers support pg_strnxfrm() safely. The caller should check
* pg_strxfrm_enabled() first, otherwise this function may return wrong
@@ -1492,7 +1542,7 @@ pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
* memcmp() on the byte sequence is equivalent to pg_strncoll() on
* untransformed strings. The result is not nul-terminated.
*
- * The input string must be encoded in the database encoding.
+ * String must be encoded in the database encoding with no embedded NULs.
*
* Not all providers support pg_strnxfrm_prefix() safely. The caller should
* check pg_strxfrm_prefix_enabled() first, otherwise this function may return
--
2.43.0
[text/x-patch] vPG19-0003-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch (3.8K, ../../3d2bf8bebb21ded729431fa25c92d2425fa615f5.camel@j-davis.com/8-vPG19-0003-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch)
download | inline diff:
From acf021658ccc38e01bda2fb9900c1d0ecacf1535 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:46:48 -0700
Subject: [PATCH vPG19 3/3] Ensure all pg_locale.h APIs work with collate_is_c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 55 +++++++++++++++++++++++++++----
1 file changed, 49 insertions(+), 6 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index da32590396f..2b501a7b6b7 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1441,7 +1441,10 @@ pg_downcase_ident(char *dst, size_t dstsize, const char *src, size_t srclen)
int
pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
{
- return locale->collate->strcoll(arg1, arg2, locale);
+ if (locale->collate == NULL)
+ return strcmp(arg1, arg2);
+ else
+ return locale->collate->strcoll(arg1, arg2, locale);
}
/*
@@ -1459,7 +1462,16 @@ int
pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
pg_locale_t locale)
{
- return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
+ if (locale->collate == NULL)
+ {
+ int result = memcmp(arg1, arg2, Min(len1, len2));
+
+ if ((result == 0) && (len1 != len2))
+ result = (len1 < len2) ? -1 : 1;
+ return result;
+ }
+ else
+ return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
}
/*
@@ -1469,6 +1481,9 @@ pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
bool
pg_strxfrm_enabled(pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ return true;
+
/*
* locale->collate->strnxfrm is still a required method, even if it may
* have the wrong behavior, because the planner uses it for estimates in
@@ -1485,7 +1500,10 @@ pg_strxfrm_enabled(pg_locale_t locale)
size_t
pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
{
- return locale->collate->strxfrm(dest, destsize, src, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strxfrm(dest, destsize, src, locale);
}
/*
@@ -1510,6 +1528,16 @@ size_t
pg_strnxfrm(char *dest, size_t destsize, const char *src, size_t srclen,
pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ {
+ if (destsize > srclen)
+ {
+ memcpy(dest, src, srclen);
+ dest[srclen] = '\0';
+ }
+
+ return srclen;
+ }
return locale->collate->strnxfrm(dest, destsize, src, srclen, locale);
}
@@ -1520,7 +1548,10 @@ pg_strnxfrm(char *dest, size_t destsize, const char *src, size_t srclen,
bool
pg_strxfrm_prefix_enabled(pg_locale_t locale)
{
- return (locale->collate->strnxfrm_prefix != NULL);
+ if (locale->collate == NULL)
+ return true;
+ else
+ return (locale->collate->strnxfrm_prefix != NULL);
}
/*
@@ -1532,7 +1563,10 @@ size_t
pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
pg_locale_t locale)
{
- return locale->collate->strxfrm_prefix(dest, destsize, src, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm_prefix(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strxfrm_prefix(dest, destsize, src, locale);
}
/*
@@ -1556,7 +1590,16 @@ size_t
pg_strnxfrm_prefix(char *dest, size_t destsize, const char *src,
size_t srclen, pg_locale_t locale)
{
- return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
+ if (locale->collate == NULL)
+ {
+ size_t len = Min(srclen, destsize);
+
+ if (destsize > 0)
+ memcpy(dest, src, len);
+ return len;
+ }
+ else
+ return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
}
bool
--
2.43.0
[text/x-patch] vPG20-0001-Add-missing-comments-in-pg_locale.c.patch (5.9K, ../../3d2bf8bebb21ded729431fa25c92d2425fa615f5.camel@j-davis.com/9-vPG20-0001-Add-missing-comments-in-pg_locale.c.patch)
download | inline diff:
From daaea4c338afb1c1389ffbfa6ac33d0bc38a5b48 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:32:19 -0700
Subject: [PATCH vPG20 1/2] Add missing comments in pg_locale.c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v3nniwcrxejmcfvz56xbd22hphprqleuornd6hqkmw2bl7kgmz@cnytz2ee5ltk
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 70 ++++++++++++++++++++++++++-----
1 file changed, 60 insertions(+), 10 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 9eb99487e57..da32590396f 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1317,6 +1317,20 @@ strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
return srclen;
}
+/*
+ * pg_strlower()
+ *
+ * Convert src to lowercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strlower(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1327,6 +1341,20 @@ pg_strlower(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strlower(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strtitle()
+ *
+ * Convert src to titlecase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strtitle(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1337,6 +1365,20 @@ pg_strtitle(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strtitle(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strupper()
+ *
+ * Convert src to uppercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strupper(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1347,6 +1389,19 @@ pg_strupper(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strupper(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strfold()
+ *
+ * Casefold src, and return the result length (not including terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strfold(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1392,11 +1447,9 @@ pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
/*
* pg_strncoll
*
- * Call ucol_strcollUTF8(), ucol_strcoll(), strcoll_l() or wcscoll_l() as
- * appropriate for the given locale, platform, and database encoding. If the
- * locale is not specified, use the database collation.
+ * Compare strings according to the given locale.
*
- * The input strings must be encoded in the database encoding.
+ * Strings must be encoded in the database encoding with no embedded NULs.
*
* The caller is responsible for breaking ties if the collation is
* deterministic; this maintains consistency with pg_strnxfrm(), which cannot
@@ -1412,9 +1465,6 @@ pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
/*
* Return true if the collation provider supports pg_strxfrm() and
* pg_strnxfrm(); otherwise false.
- *
- *
- * No similar problem is known for the ICU provider.
*/
bool
pg_strxfrm_enabled(pg_locale_t locale)
@@ -1445,8 +1495,8 @@ pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
* ordinary strcmp() on transformed strings is equivalent to pg_strcoll() on
* untransformed strings.
*
- * The input string must be encoded in the database encoding. If 'destsize' is
- * zero, 'dest' may be NULL.
+ * String must be encoded in the database encoding with no embedded NULs. If
+ * 'destsize' is zero, 'dest' may be NULL.
*
* Not all providers support pg_strnxfrm() safely. The caller should check
* pg_strxfrm_enabled() first, otherwise this function may return wrong
@@ -1492,7 +1542,7 @@ pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
* memcmp() on the byte sequence is equivalent to pg_strncoll() on
* untransformed strings. The result is not nul-terminated.
*
- * The input string must be encoded in the database encoding.
+ * String must be encoded in the database encoding with no embedded NULs.
*
* Not all providers support pg_strnxfrm_prefix() safely. The caller should
* check pg_strxfrm_prefix_enabled() first, otherwise this function may return
--
2.43.0
[text/x-patch] vPG20-0002-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch (3.8K, ../../3d2bf8bebb21ded729431fa25c92d2425fa615f5.camel@j-davis.com/10-vPG20-0002-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch)
download | inline diff:
From 7fbcfeeeb04ecefe917c29e53237d79d4447795e Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:46:48 -0700
Subject: [PATCH vPG20 2/2] Ensure all pg_locale.h APIs work with collate_is_c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 55 +++++++++++++++++++++++++++----
1 file changed, 49 insertions(+), 6 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index da32590396f..2b501a7b6b7 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1441,7 +1441,10 @@ pg_downcase_ident(char *dst, size_t dstsize, const char *src, size_t srclen)
int
pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
{
- return locale->collate->strcoll(arg1, arg2, locale);
+ if (locale->collate == NULL)
+ return strcmp(arg1, arg2);
+ else
+ return locale->collate->strcoll(arg1, arg2, locale);
}
/*
@@ -1459,7 +1462,16 @@ int
pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
pg_locale_t locale)
{
- return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
+ if (locale->collate == NULL)
+ {
+ int result = memcmp(arg1, arg2, Min(len1, len2));
+
+ if ((result == 0) && (len1 != len2))
+ result = (len1 < len2) ? -1 : 1;
+ return result;
+ }
+ else
+ return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
}
/*
@@ -1469,6 +1481,9 @@ pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
bool
pg_strxfrm_enabled(pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ return true;
+
/*
* locale->collate->strnxfrm is still a required method, even if it may
* have the wrong behavior, because the planner uses it for estimates in
@@ -1485,7 +1500,10 @@ pg_strxfrm_enabled(pg_locale_t locale)
size_t
pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
{
- return locale->collate->strxfrm(dest, destsize, src, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strxfrm(dest, destsize, src, locale);
}
/*
@@ -1510,6 +1528,16 @@ size_t
pg_strnxfrm(char *dest, size_t destsize, const char *src, size_t srclen,
pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ {
+ if (destsize > srclen)
+ {
+ memcpy(dest, src, srclen);
+ dest[srclen] = '\0';
+ }
+
+ return srclen;
+ }
return locale->collate->strnxfrm(dest, destsize, src, srclen, locale);
}
@@ -1520,7 +1548,10 @@ pg_strnxfrm(char *dest, size_t destsize, const char *src, size_t srclen,
bool
pg_strxfrm_prefix_enabled(pg_locale_t locale)
{
- return (locale->collate->strnxfrm_prefix != NULL);
+ if (locale->collate == NULL)
+ return true;
+ else
+ return (locale->collate->strnxfrm_prefix != NULL);
}
/*
@@ -1532,7 +1563,10 @@ size_t
pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
pg_locale_t locale)
{
- return locale->collate->strxfrm_prefix(dest, destsize, src, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm_prefix(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strxfrm_prefix(dest, destsize, src, locale);
}
/*
@@ -1556,7 +1590,16 @@ size_t
pg_strnxfrm_prefix(char *dest, size_t destsize, const char *src,
size_t srclen, pg_locale_t locale)
{
- return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
+ if (locale->collate == NULL)
+ {
+ size_t len = Min(srclen, destsize);
+
+ if (destsize > 0)
+ memcpy(dest, src, len);
+ return len;
+ }
+ else
+ return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
}
bool
--
2.43.0
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-15 21:53 Jeff Davis <pgsql@j-davis.com>
parent: Jeff Davis <pgsql@j-davis.com>
0 siblings, 1 reply; 16+ messages in thread
From: Jeff Davis @ 2026-08-15 21:53 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>
On Thu, 2026-08-13 at 22:17 -0700, Jeff Davis wrote:
> On Tue, 2026-08-11 at 14:31 -0400, Andres Freund wrote:
> > Hm. If I infer the pg_strlower() API correctly - it's utterly
> > underdocumented
>
> Agreed. Patch attached.
>
> > I'd also make i size_t, given that the input is size_t. Perhaps
> > practically
> > no problem, but I see no reason to not use size_t here.
>
> Patch attached for that, too.
>
> I also attached patches to make all the functions work with
> collate_is_c, and fixed up the -1 API in 18.
Now with a C test module (made with AI assistance).
I plan to start committing these fairly soon. I'm not sure whether to
backport the C test module, but I included the patches to do so.
Regards,
Jeff Davis
Attachments:
[text/x-patch] vPG18-0001-Fixup-5f003855e7-for-srclen-0.patch (1.9K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/2-vPG18-0001-Fixup-5f003855e7-for-srclen-0.patch)
download | inline diff:
From 77e5ec7124d4530fead31794cfee62905d7254d2 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Thu, 13 Aug 2026 20:57:20 -0700
Subject: [PATCH vPG18 1/5] Fixup 5f003855e7 for srclen < 0.
No actual problem because no callers used that aspect of the API.
Only commit to 18, because that part of the API was removed in commit
6d22c67c3b.
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
---
src/backend/utils/adt/pg_locale.c | 8 ++++++++
1 file changed, 8 insertions(+)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 2f8d5fea8f2..9c721efb52f 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1328,6 +1328,8 @@ size_t
pg_strlower(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+
if (locale->ctype_is_c)
return strlower_c(dst, dstsize, src, srclen);
else if (locale->provider == COLLPROVIDER_BUILTIN)
@@ -1349,6 +1351,8 @@ size_t
pg_strtitle(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+
if (locale->ctype_is_c)
return strtitle_c(dst, dstsize, src, srclen);
else if (locale->provider == COLLPROVIDER_BUILTIN)
@@ -1370,6 +1374,8 @@ size_t
pg_strupper(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+
if (locale->ctype_is_c)
return strupper_c(dst, dstsize, src, srclen);
else if (locale->provider == COLLPROVIDER_BUILTIN)
@@ -1391,6 +1397,8 @@ size_t
pg_strfold(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+
/* in the C locale, casefolding is the same as lowercasing */
if (locale->ctype_is_c)
return strlower_c(dst, dstsize, src, srclen);
--
2.43.0
[text/x-patch] vPG18-0002-pg_locale.c-unicode_case.c-use-size_t-for-iter.patch (2.3K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/3-vPG18-0002-pg_locale.c-unicode_case.c-use-size_t-for-iter.patch)
download | inline diff:
From f669e9fafb9bf435c946dcc2c2dd98efb4aa46cb Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Thu, 13 Aug 2026 21:12:34 -0700
Subject: [PATCH vPG18 2/5] pg_locale.c, unicode_case.c: use size_t for
iteration.
No actual problem, just cleanup. Only relevant to 18 and 19.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
---
src/backend/utils/adt/pg_locale.c | 6 +++---
src/common/unicode_case.c | 4 ++--
2 files changed, 5 insertions(+), 5 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 9c721efb52f..8e79ccbd9f6 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1277,7 +1277,7 @@ get_collation_actual_version(char collprovider, const char *collcollate)
static size_t
strlower_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
dst[i] = pg_ascii_tolower(src[i]);
@@ -1291,7 +1291,7 @@ static size_t
strtitle_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
bool wasalnum = false;
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
{
@@ -1315,7 +1315,7 @@ strtitle_c(char *dst, size_t dstsize, const char *src, size_t srclen)
static size_t
strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
dst[i] = pg_ascii_toupper(src[i]);
diff --git a/src/common/unicode_case.c b/src/common/unicode_case.c
index 8639b203e0c..86e0b57d7d3 100644
--- a/src/common/unicode_case.c
+++ b/src/common/unicode_case.c
@@ -341,7 +341,7 @@ check_final_sigma(const unsigned char *str, size_t len, size_t offset)
int ulen;
/* iterate backwards looking for preceding character */
- for (int i = offset; i > 0;)
+ for (size_t i = offset; i > 0;)
{
/* skip backwards through continuation bytes */
i--;
@@ -369,7 +369,7 @@ check_final_sigma(const unsigned char *str, size_t len, size_t offset)
ulen = utf8_mblen((const unsigned char *) str + offset);
/* iterate forward looking for following character */
- for (int i = offset + ulen; i < len;)
+ for (size_t i = offset + ulen; i < len;)
{
ulen = utf8_mblen((const unsigned char *) str + i);
--
2.43.0
[text/x-patch] vPG18-0003-Add-missing-comments-in-pg_locale.c.patch (6.4K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/4-vPG18-0003-Add-missing-comments-in-pg_locale.c.patch)
download | inline diff:
From a8e72bb34c979ac43ee8480fd10ed4e73db47b10 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:32:19 -0700
Subject: [PATCH vPG18 3/5] Add missing comments in pg_locale.c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v3nniwcrxejmcfvz56xbd22hphprqleuornd6hqkmw2bl7kgmz@cnytz2ee5ltk
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 76 +++++++++++++++++++++++++------
1 file changed, 63 insertions(+), 13 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 8e79ccbd9f6..c21619f85cd 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1324,6 +1324,20 @@ strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
return srclen;
}
+/*
+ * pg_strlower()
+ *
+ * Convert src to lowercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If srclen is
+ * -1, src must be NUL-terminated. If dstsize is zero, dst may be NULL, which
+ * is useful for calculating the required buffer size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strlower(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
@@ -1347,6 +1361,20 @@ pg_strlower(char *dst, size_t dstsize, const char *src, ssize_t srclen,
return 0; /* keep compiler quiet */
}
+/*
+ * pg_strtitle()
+ *
+ * Convert src to titlecase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If srclen is
+ * -1, src must be NUL-terminated. If dstsize is zero, dst may be NULL, which
+ * is useful for calculating the required buffer size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strtitle(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
@@ -1370,6 +1398,20 @@ pg_strtitle(char *dst, size_t dstsize, const char *src, ssize_t srclen,
return 0; /* keep compiler quiet */
}
+/*
+ * pg_strupper()
+ *
+ * Convert src to uppercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If srclen is
+ * -1, src must be NUL-terminated. If dstsize is zero, dst may be NULL, which
+ * is useful for calculating the required buffer size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strupper(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
@@ -1393,6 +1435,19 @@ pg_strupper(char *dst, size_t dstsize, const char *src, ssize_t srclen,
return 0; /* keep compiler quiet */
}
+/*
+ * pg_strfold()
+ *
+ * Casefold src, and return the result length (not including terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If srclen is
+ * -1, src must be NUL-terminated. If dstsize is zero, dst may be NULL, which
+ * is useful for calculating the required buffer size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strfold(char *dst, size_t dstsize, const char *src, ssize_t srclen,
pg_locale_t locale)
@@ -1432,12 +1487,10 @@ pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
/*
* pg_strncoll
*
- * Call ucol_strcollUTF8(), ucol_strcoll(), strcoll_l() or wcscoll_l() as
- * appropriate for the given locale, platform, and database encoding. If the
- * locale is not specified, use the database collation.
+ * Compare strings according to the given locale.
*
- * The input strings must be encoded in the database encoding. If an input
- * string is NUL-terminated, its length may be specified as -1.
+ * Strings must be encoded in the database encoding with no embedded NULs. If
+ * an input string is NUL-terminated, its length may be specified as -1.
*
* The caller is responsible for breaking ties if the collation is
* deterministic; this maintains consistency with pg_strnxfrm(), which cannot
@@ -1453,9 +1506,6 @@ pg_strncoll(const char *arg1, ssize_t len1, const char *arg2, ssize_t len2,
/*
* Return true if the collation provider supports pg_strxfrm() and
* pg_strnxfrm(); otherwise false.
- *
- *
- * No similar problem is known for the ICU provider.
*/
bool
pg_strxfrm_enabled(pg_locale_t locale)
@@ -1486,9 +1536,9 @@ pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
* ordinary strcmp() on transformed strings is equivalent to pg_strcoll() on
* untransformed strings.
*
- * The input string must be encoded in the database encoding. If the input
- * string is NUL-terminated, its length may be specified as -1. If 'destsize'
- * is zero, 'dest' may be NULL.
+ * String must be encoded in the database encoding with no embedded NULs. If
+ * srclen is -1, src must be NUL-terminated. If 'destsize' is zero, 'dest'
+ * may be NULL.
*
* Not all providers support pg_strnxfrm() safely. The caller should check
* pg_strxfrm_enabled() first, otherwise this function may return wrong
@@ -1534,8 +1584,8 @@ pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
* memcmp() on the byte sequence is equivalent to pg_strncoll() on
* untransformed strings. The result is not nul-terminated.
*
- * The input string must be encoded in the database encoding. If the input
- * string is NUL-terminated, its length may be specified as -1.
+ * String must be encoded in the database encoding with no embedded NULs. If
+ * the input string is NUL-terminated, its length may be specified as -1.
*
* Not all providers support pg_strnxfrm_prefix() safely. The caller should
* check pg_strxfrm_prefix_enabled() first, otherwise this function may return
--
2.43.0
[text/x-patch] vPG18-0004-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch (4.0K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/5-vPG18-0004-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch)
download | inline diff:
From a578ea2245f071f3ba7897ce3063c8926ba765b5 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:46:48 -0700
Subject: [PATCH vPG18 4/5] Ensure all pg_locale.h APIs work with collate_is_c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 64 ++++++++++++++++++++++++++++---
1 file changed, 58 insertions(+), 6 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index c21619f85cd..ad9f416ec13 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1481,7 +1481,10 @@ pg_strfold(char *dst, size_t dstsize, const char *src, ssize_t srclen,
int
pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
{
- return locale->collate->strncoll(arg1, -1, arg2, -1, locale);
+ if (locale->collate == NULL)
+ return strcmp(arg1, arg2);
+ else
+ return locale->collate->strncoll(arg1, -1, arg2, -1, locale);
}
/*
@@ -1500,7 +1503,20 @@ int
pg_strncoll(const char *arg1, ssize_t len1, const char *arg2, ssize_t len2,
pg_locale_t locale)
{
- return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
+ if (locale->collate == NULL)
+ {
+ int result;
+
+ len1 = (len1 < 0) ? strlen(arg1) : len1;
+ len2 = (len2 < 0) ? strlen(arg2) : len2;
+ result = memcmp(arg1, arg2, Min(len1, len2));
+
+ if ((result == 0) && (len1 != len2))
+ result = (len1 < len2) ? -1 : 1;
+ return result;
+ }
+ else
+ return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
}
/*
@@ -1510,6 +1526,9 @@ pg_strncoll(const char *arg1, ssize_t len1, const char *arg2, ssize_t len2,
bool
pg_strxfrm_enabled(pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ return true;
+
/*
* locale->collate->strnxfrm is still a required method, even if it may
* have the wrong behavior, because the planner uses it for estimates in
@@ -1526,7 +1545,10 @@ pg_strxfrm_enabled(pg_locale_t locale)
size_t
pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
{
- return locale->collate->strnxfrm(dest, destsize, src, -1, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strnxfrm(dest, destsize, src, -1, locale);
}
/*
@@ -1552,6 +1574,18 @@ size_t
pg_strnxfrm(char *dest, size_t destsize, const char *src, ssize_t srclen,
pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ {
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+
+ if (destsize > srclen)
+ {
+ memcpy(dest, src, srclen);
+ dest[srclen] = '\0';
+ }
+
+ return srclen;
+ }
return locale->collate->strnxfrm(dest, destsize, src, srclen, locale);
}
@@ -1562,7 +1596,10 @@ pg_strnxfrm(char *dest, size_t destsize, const char *src, ssize_t srclen,
bool
pg_strxfrm_prefix_enabled(pg_locale_t locale)
{
- return (locale->collate->strnxfrm_prefix != NULL);
+ if (locale->collate == NULL)
+ return true;
+ else
+ return (locale->collate->strnxfrm_prefix != NULL);
}
/*
@@ -1574,7 +1611,10 @@ size_t
pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
pg_locale_t locale)
{
- return locale->collate->strnxfrm_prefix(dest, destsize, src, -1, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm_prefix(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strnxfrm_prefix(dest, destsize, src, -1, locale);
}
/*
@@ -1599,7 +1639,19 @@ size_t
pg_strnxfrm_prefix(char *dest, size_t destsize, const char *src,
ssize_t srclen, pg_locale_t locale)
{
- return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
+ if (locale->collate == NULL)
+ {
+ size_t len;
+
+ srclen = (srclen < 0) ? strlen(src) : srclen;
+ len = Min(srclen, destsize);
+
+ if (destsize > 0)
+ memcpy(dest, src, len);
+ return len;
+ }
+ else
+ return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
}
/*
--
2.43.0
[text/x-patch] vPG18-0005-Add-C-test-module-for-pg_locale.h-APIs.patch (14.3K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/6-vPG18-0005-Add-C-test-module-for-pg_locale.h-APIs.patch)
download | inline diff:
From adeef54179a11f5b7817054bc53f8b7b584d2b6a Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Sat, 15 Aug 2026 12:26:09 -0700
Subject: [PATCH vPG18 5/5] Add C test module for pg_locale.h APIs.
Test the API independently to account for fallback paths that aren't
adequately tested from SQL.
The backport to 18 also tests the previously-supported behavior where
a size of -1 meant that the string was NUL-terminated. That behavior
was later removed in 19.
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/test/modules/Makefile | 1 +
src/test/modules/meson.build | 1 +
src/test/modules/test_pg_locale/.gitignore | 4 +
src/test/modules/test_pg_locale/Makefile | 23 +++
src/test/modules/test_pg_locale/README | 2 +
.../expected/test_pg_locale.out | 38 ++++
src/test/modules/test_pg_locale/meson.build | 33 ++++
.../test_pg_locale/sql/test_pg_locale.sql | 23 +++
.../test_pg_locale/test_pg_locale--1.0.sql | 8 +
.../modules/test_pg_locale/test_pg_locale.c | 169 ++++++++++++++++++
.../test_pg_locale/test_pg_locale.control | 4 +
11 files changed, 306 insertions(+)
create mode 100644 src/test/modules/test_pg_locale/.gitignore
create mode 100644 src/test/modules/test_pg_locale/Makefile
create mode 100644 src/test/modules/test_pg_locale/README
create mode 100644 src/test/modules/test_pg_locale/expected/test_pg_locale.out
create mode 100644 src/test/modules/test_pg_locale/meson.build
create mode 100644 src/test/modules/test_pg_locale/sql/test_pg_locale.sql
create mode 100644 src/test/modules/test_pg_locale/test_pg_locale--1.0.sql
create mode 100644 src/test/modules/test_pg_locale/test_pg_locale.c
create mode 100644 src/test/modules/test_pg_locale/test_pg_locale.control
diff --git a/src/test/modules/Makefile b/src/test/modules/Makefile
index 4e82d6f1517..9ee74b66797 100644
--- a/src/test/modules/Makefile
+++ b/src/test/modules/Makefile
@@ -33,6 +33,7 @@ SUBDIRS = \
test_oat_hooks \
test_parser \
test_pg_dump \
+ test_pg_locale \
test_predtest \
test_radixtree \
test_rbtree \
diff --git a/src/test/modules/meson.build b/src/test/modules/meson.build
index 9a957351ab6..85ef25deadf 100644
--- a/src/test/modules/meson.build
+++ b/src/test/modules/meson.build
@@ -32,6 +32,7 @@ subdir('test_misc')
subdir('test_oat_hooks')
subdir('test_parser')
subdir('test_pg_dump')
+subdir('test_pg_locale')
subdir('test_predtest')
subdir('test_radixtree')
subdir('test_rbtree')
diff --git a/src/test/modules/test_pg_locale/.gitignore b/src/test/modules/test_pg_locale/.gitignore
new file mode 100644
index 00000000000..5dcb3ff9723
--- /dev/null
+++ b/src/test/modules/test_pg_locale/.gitignore
@@ -0,0 +1,4 @@
+# Generated subdirectories
+/log/
+/results/
+/tmp_check/
diff --git a/src/test/modules/test_pg_locale/Makefile b/src/test/modules/test_pg_locale/Makefile
new file mode 100644
index 00000000000..9b051f8a697
--- /dev/null
+++ b/src/test/modules/test_pg_locale/Makefile
@@ -0,0 +1,23 @@
+# src/test/modules/test_pg_locale/Makefile
+
+MODULE_big = test_pg_locale
+OBJS = \
+ $(WIN32RES) \
+ test_pg_locale.o
+PGFILEDESC = "test_pg_locale - test code for pg_locale.h APIs"
+
+EXTENSION = test_pg_locale
+DATA = test_pg_locale--1.0.sql
+
+REGRESS = test_pg_locale
+
+ifdef USE_PGXS
+PG_CONFIG = pg_config
+PGXS := $(shell $(PG_CONFIG) --pgxs)
+include $(PGXS)
+else
+subdir = src/test/modules/test_pg_locale
+top_builddir = ../../../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+endif
diff --git a/src/test/modules/test_pg_locale/README b/src/test/modules/test_pg_locale/README
new file mode 100644
index 00000000000..d95af97c005
--- /dev/null
+++ b/src/test/modules/test_pg_locale/README
@@ -0,0 +1,2 @@
+Calls pg_locale.h wrappers directly. Ordinary SQL tests do not reach
+the C-locale fallbacks because in-tree callers special-case collate_is_c.
diff --git a/src/test/modules/test_pg_locale/expected/test_pg_locale.out b/src/test/modules/test_pg_locale/expected/test_pg_locale.out
new file mode 100644
index 00000000000..edbba284552
--- /dev/null
+++ b/src/test/modules/test_pg_locale/expected/test_pg_locale.out
@@ -0,0 +1,38 @@
+CREATE EXTENSION test_pg_locale;
+--
+-- These tests don't produce any interesting output. We're checking that
+-- the operations complete without crashing and that none of their internal
+-- sanity tests fail.
+--
+-- to_regcollation() returns NULL when the collation is absent or is not
+-- usable in the current database encoding. The function is STRICT, so
+-- those cases are skipped.
+--
+-- Libc C. Available in every database.
+SELECT test_pg_locale_apis(to_regcollation('"C"'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- Builtin C (collate and ctype). Usable only in UTF8 databases.
+SELECT test_pg_locale_apis(to_regcollation('ucs_basic'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- Builtin C.UTF-8 (C collate, Unicode ctype). Same encoding restriction.
+SELECT test_pg_locale_apis(to_regcollation('pg_c_utf8'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- en-x-icu is present when ICU collations were imported at initdb.
+SELECT test_pg_locale_apis(to_regcollation('en-x-icu'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
diff --git a/src/test/modules/test_pg_locale/meson.build b/src/test/modules/test_pg_locale/meson.build
new file mode 100644
index 00000000000..f5097464ad2
--- /dev/null
+++ b/src/test/modules/test_pg_locale/meson.build
@@ -0,0 +1,33 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+test_pg_locale_sources = files(
+ 'test_pg_locale.c',
+)
+
+if host_system == 'windows'
+ test_pg_locale_sources += rc_lib_gen.process(win32ver_rc, extra_args: [
+ '--NAME', 'test_pg_locale',
+ '--FILEDESC', 'test_pg_locale - test code for pg_locale.h APIs',])
+endif
+
+test_pg_locale = shared_module('test_pg_locale',
+ test_pg_locale_sources,
+ kwargs: pg_test_mod_args,
+)
+test_install_libs += test_pg_locale
+
+test_install_data += files(
+ 'test_pg_locale.control',
+ 'test_pg_locale--1.0.sql',
+)
+
+tests += {
+ 'name': 'test_pg_locale',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'test_pg_locale',
+ ],
+ },
+}
diff --git a/src/test/modules/test_pg_locale/sql/test_pg_locale.sql b/src/test/modules/test_pg_locale/sql/test_pg_locale.sql
new file mode 100644
index 00000000000..212012feb9c
--- /dev/null
+++ b/src/test/modules/test_pg_locale/sql/test_pg_locale.sql
@@ -0,0 +1,23 @@
+CREATE EXTENSION test_pg_locale;
+
+--
+-- These tests don't produce any interesting output. We're checking that
+-- the operations complete without crashing and that none of their internal
+-- sanity tests fail.
+--
+-- to_regcollation() returns NULL when the collation is absent or is not
+-- usable in the current database encoding. The function is STRICT, so
+-- those cases are skipped.
+--
+
+-- Libc C. Available in every database.
+SELECT test_pg_locale_apis(to_regcollation('"C"'));
+
+-- Builtin C (collate and ctype). Usable only in UTF8 databases.
+SELECT test_pg_locale_apis(to_regcollation('ucs_basic'));
+
+-- Builtin C.UTF-8 (C collate, Unicode ctype). Same encoding restriction.
+SELECT test_pg_locale_apis(to_regcollation('pg_c_utf8'));
+
+-- en-x-icu is present when ICU collations were imported at initdb.
+SELECT test_pg_locale_apis(to_regcollation('en-x-icu'));
diff --git a/src/test/modules/test_pg_locale/test_pg_locale--1.0.sql b/src/test/modules/test_pg_locale/test_pg_locale--1.0.sql
new file mode 100644
index 00000000000..134c4befa06
--- /dev/null
+++ b/src/test/modules/test_pg_locale/test_pg_locale--1.0.sql
@@ -0,0 +1,8 @@
+/* src/test/modules/test_pg_locale/test_pg_locale--1.0.sql */
+
+-- complain if script is sourced in psql, rather than via CREATE EXTENSION
+\echo Use "CREATE EXTENSION test_pg_locale" to load this file. \quit
+
+CREATE FUNCTION test_pg_locale_apis(oid)
+ RETURNS pg_catalog.void
+ AS 'MODULE_PATHNAME' LANGUAGE C STRICT;
diff --git a/src/test/modules/test_pg_locale/test_pg_locale.c b/src/test/modules/test_pg_locale/test_pg_locale.c
new file mode 100644
index 00000000000..79041fb3f69
--- /dev/null
+++ b/src/test/modules/test_pg_locale/test_pg_locale.c
@@ -0,0 +1,169 @@
+/*-------------------------------------------------------------------------
+ *
+ * test_pg_locale.c
+ * Call pg_locale.h wrappers directly.
+ *
+ * SQL callers special-case collate_is_c, so the C-locale fallbacks are
+ * not reached by ordinary regression tests.
+ *
+ * Copyright (c) 2026, PostgreSQL Global Development Group
+ *
+ * IDENTIFICATION
+ * src/test/modules/test_pg_locale/test_pg_locale.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "fmgr.h"
+#include "utils/pg_locale.h"
+
+PG_MODULE_MAGIC;
+
+static void
+test_case_mapping(pg_locale_t locale)
+{
+ char buf[32];
+ size_t n;
+
+ n = pg_strlower(NULL, 0, "AbC", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strlower() size probe returned %zu, expected 3", n);
+ n = pg_strlower(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strlower() produced \"%s\"", buf);
+
+ n = pg_strupper(NULL, 0, "AbC", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strupper() size probe returned %zu, expected 3", n);
+ n = pg_strupper(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "ABC") != 0)
+ elog(ERROR, "pg_strupper() produced \"%s\"", buf);
+
+ n = pg_strfold(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strfold() produced \"%s\"", buf);
+
+ buf[0] = '\0';
+ n = pg_strtitle(buf, sizeof(buf), "hello-world", 11, locale);
+ if (n != 11)
+ elog(ERROR, "pg_strtitle() returned %zu, expected 11", n);
+ if (locale->ctype_is_c && strcmp(buf, "Hello-World") != 0)
+ elog(ERROR, "pg_strtitle() produced \"%s\"", buf);
+
+ n = pg_strlower(buf, sizeof(buf), "AbC", -1, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strlower() with srclen -1 produced \"%s\"", buf);
+}
+
+static void
+test_collate(pg_locale_t locale)
+{
+ char buf[32];
+ char pfx[8];
+ char x1[8];
+ char x2[8];
+ size_t n;
+
+ if (pg_strcoll("abc", "abc", locale) != 0 ||
+ pg_strncoll("abc", 3, "abc", 3, locale) != 0 ||
+ pg_strcoll("", "", locale) != 0)
+ elog(ERROR, "equal strings did not compare equal");
+
+ if (locale->collate_is_c)
+ {
+ if (locale->collate != NULL)
+ elog(ERROR, "collate_is_c but collate methods are set");
+ if (pg_strcoll("abc", "abd", locale) >= 0 ||
+ pg_strcoll("abd", "abc", locale) <= 0 ||
+ pg_strncoll("ab", 2, "abc", 3, locale) >= 0 ||
+ pg_strncoll("abc", 3, "ab", 2, locale) <= 0 ||
+ pg_strncoll("xyz", 3, "abc", 2, locale) <= 0)
+ elog(ERROR, "C-locale comparison result is wrong");
+
+ if (!pg_strxfrm_enabled(locale))
+ elog(ERROR, "pg_strxfrm_enabled() is false for C locale");
+ n = pg_strnxfrm(NULL, 0, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() size probe returned %zu, expected 3", n);
+ n = pg_strnxfrm(buf, 4, "abc", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strnxfrm() produced \"%s\"", buf);
+ n = pg_strxfrm(buf, "abc", 4, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strxfrm() produced \"%s\"", buf);
+ n = pg_strnxfrm(buf, 3, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() destsize==srclen returned %zu", n);
+ n = pg_strnxfrm(buf, 2, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() short dest returned %zu", n);
+
+ if (!pg_strxfrm_prefix_enabled(locale))
+ elog(ERROR, "pg_strxfrm_prefix_enabled() is false for C locale");
+ n = pg_strnxfrm_prefix(NULL, 0, "abcdef", 6, locale);
+ if (n != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() destsize 0 returned %zu", n);
+ n = pg_strnxfrm_prefix(pfx, 2, "abcdef", 6, locale);
+ if (n != 2 || memcmp(pfx, "ab", 2) != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() produced a wrong prefix");
+ n = pg_strnxfrm_prefix(pfx, sizeof(pfx), "abc", 3, locale);
+ if (n != 3 || memcmp(pfx, "abc", 3) != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() destsize>=srclen produced a wrong result");
+ n = pg_strxfrm_prefix(pfx, "abcdef", 2, locale);
+ if (n != 2 || memcmp(pfx, "ab", 2) != 0)
+ elog(ERROR, "pg_strxfrm_prefix() produced a wrong prefix");
+
+ if (pg_strxfrm(x1, "abc", sizeof(x1), locale) >= sizeof(x1) ||
+ pg_strxfrm(x2, "abd", sizeof(x2), locale) >= sizeof(x2) ||
+ (strcmp(x1, x2) < 0) != (pg_strcoll("abc", "abd", locale) < 0))
+ elog(ERROR, "pg_strxfrm() disagrees with pg_strcoll()");
+ }
+ else
+ {
+ char *tmp;
+
+ if (locale->collate == NULL)
+ elog(ERROR, "collate methods missing for non-C locale");
+
+ n = pg_strnxfrm(NULL, 0, "abc", 3, locale);
+ tmp = palloc(n + 1);
+ if (pg_strnxfrm(tmp, n + 1, "abc", 3, locale) > n)
+ elog(ERROR, "pg_strnxfrm() grew on the second call");
+ pfree(tmp);
+
+ if (pg_strxfrm_prefix_enabled(locale) &&
+ pg_strnxfrm_prefix(pfx, sizeof(pfx), "abc", 3, locale) > sizeof(pfx))
+ elog(ERROR, "pg_strnxfrm_prefix() exceeded destsize");
+ }
+
+ if (pg_strncoll("abc", -1, "abc", -1, locale) != 0)
+ elog(ERROR, "pg_strncoll() with length -1 failed");
+ if (locale->collate_is_c)
+ {
+ n = pg_strnxfrm(buf, sizeof(buf), "abc", -1, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strnxfrm() with srclen -1 produced \"%s\"", buf);
+ n = pg_strnxfrm_prefix(pfx, sizeof(pfx), "abc", -1, locale);
+ if (n != 3 || memcmp(pfx, "abc", 3) != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() with srclen -1 produced a wrong result");
+ }
+}
+
+PG_FUNCTION_INFO_V1(test_pg_locale_apis);
+
+Datum
+test_pg_locale_apis(PG_FUNCTION_ARGS)
+{
+ pg_locale_t locale;
+
+ locale = pg_newlocale_from_collation(PG_GETARG_OID(0));
+ if (locale == NULL)
+ elog(ERROR, "pg_newlocale_from_collation() returned NULL");
+
+ test_collate(locale);
+ test_case_mapping(locale);
+
+ PG_RETURN_VOID();
+}
diff --git a/src/test/modules/test_pg_locale/test_pg_locale.control b/src/test/modules/test_pg_locale/test_pg_locale.control
new file mode 100644
index 00000000000..6b224d04a1b
--- /dev/null
+++ b/src/test/modules/test_pg_locale/test_pg_locale.control
@@ -0,0 +1,4 @@
+comment = 'Test code for pg_locale.h APIs'
+default_version = '1.0'
+module_pathname = '$libdir/test_pg_locale'
+relocatable = true
--
2.43.0
[text/x-patch] vPG19-0001-pg_locale.c-unicode_case.c-use-size_t-for-iter.patch (2.3K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/7-vPG19-0001-pg_locale.c-unicode_case.c-use-size_t-for-iter.patch)
download | inline diff:
From 3d63edf5d0132614d8007845220517cd6dfd2f37 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Thu, 13 Aug 2026 21:12:34 -0700
Subject: [PATCH vPG19 1/4] pg_locale.c, unicode_case.c: use size_t for
iteration.
No actual problem, just cleanup. Only relevant to 18 and 19.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
---
src/backend/utils/adt/pg_locale.c | 6 +++---
src/common/unicode_case.c | 4 ++--
2 files changed, 5 insertions(+), 5 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 11d48a3916e..9eb99487e57 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1270,7 +1270,7 @@ get_collation_actual_version(char collprovider, const char *collcollate)
static size_t
strlower_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
dst[i] = pg_ascii_tolower(src[i]);
@@ -1284,7 +1284,7 @@ static size_t
strtitle_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
bool wasalnum = false;
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
{
@@ -1308,7 +1308,7 @@ strtitle_c(char *dst, size_t dstsize, const char *src, size_t srclen)
static size_t
strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
{
- int i;
+ size_t i;
for (i = 0; i < srclen && i < dstsize; i++)
dst[i] = pg_ascii_toupper(src[i]);
diff --git a/src/common/unicode_case.c b/src/common/unicode_case.c
index dd5b3ba86d0..744b9116b12 100644
--- a/src/common/unicode_case.c
+++ b/src/common/unicode_case.c
@@ -336,7 +336,7 @@ check_final_sigma(const unsigned char *str, size_t len, size_t offset)
int ulen;
/* iterate backwards looking for preceding character */
- for (int i = offset; i > 0;)
+ for (size_t i = offset; i > 0;)
{
/* skip backwards through continuation bytes */
i--;
@@ -364,7 +364,7 @@ check_final_sigma(const unsigned char *str, size_t len, size_t offset)
ulen = utf8_mblen((const unsigned char *) str + offset);
/* iterate forward looking for following character */
- for (int i = offset + ulen; i < len;)
+ for (size_t i = offset + ulen; i < len;)
{
ulen = utf8_mblen((const unsigned char *) str + i);
--
2.43.0
[text/x-patch] vPG19-0002-Add-missing-comments-in-pg_locale.c.patch (5.9K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/8-vPG19-0002-Add-missing-comments-in-pg_locale.c.patch)
download | inline diff:
From 2d63e62bc454828630f68aa5bfff486a47540bff Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:32:19 -0700
Subject: [PATCH vPG19 2/4] Add missing comments in pg_locale.c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v3nniwcrxejmcfvz56xbd22hphprqleuornd6hqkmw2bl7kgmz@cnytz2ee5ltk
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 70 ++++++++++++++++++++++++++-----
1 file changed, 60 insertions(+), 10 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 9eb99487e57..da32590396f 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1317,6 +1317,20 @@ strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
return srclen;
}
+/*
+ * pg_strlower()
+ *
+ * Convert src to lowercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strlower(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1327,6 +1341,20 @@ pg_strlower(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strlower(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strtitle()
+ *
+ * Convert src to titlecase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strtitle(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1337,6 +1365,20 @@ pg_strtitle(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strtitle(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strupper()
+ *
+ * Convert src to uppercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strupper(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1347,6 +1389,19 @@ pg_strupper(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strupper(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strfold()
+ *
+ * Casefold src, and return the result length (not including terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strfold(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1392,11 +1447,9 @@ pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
/*
* pg_strncoll
*
- * Call ucol_strcollUTF8(), ucol_strcoll(), strcoll_l() or wcscoll_l() as
- * appropriate for the given locale, platform, and database encoding. If the
- * locale is not specified, use the database collation.
+ * Compare strings according to the given locale.
*
- * The input strings must be encoded in the database encoding.
+ * Strings must be encoded in the database encoding with no embedded NULs.
*
* The caller is responsible for breaking ties if the collation is
* deterministic; this maintains consistency with pg_strnxfrm(), which cannot
@@ -1412,9 +1465,6 @@ pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
/*
* Return true if the collation provider supports pg_strxfrm() and
* pg_strnxfrm(); otherwise false.
- *
- *
- * No similar problem is known for the ICU provider.
*/
bool
pg_strxfrm_enabled(pg_locale_t locale)
@@ -1445,8 +1495,8 @@ pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
* ordinary strcmp() on transformed strings is equivalent to pg_strcoll() on
* untransformed strings.
*
- * The input string must be encoded in the database encoding. If 'destsize' is
- * zero, 'dest' may be NULL.
+ * String must be encoded in the database encoding with no embedded NULs. If
+ * 'destsize' is zero, 'dest' may be NULL.
*
* Not all providers support pg_strnxfrm() safely. The caller should check
* pg_strxfrm_enabled() first, otherwise this function may return wrong
@@ -1492,7 +1542,7 @@ pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
* memcmp() on the byte sequence is equivalent to pg_strncoll() on
* untransformed strings. The result is not nul-terminated.
*
- * The input string must be encoded in the database encoding.
+ * String must be encoded in the database encoding with no embedded NULs.
*
* Not all providers support pg_strnxfrm_prefix() safely. The caller should
* check pg_strxfrm_prefix_enabled() first, otherwise this function may return
--
2.43.0
[text/x-patch] vPG19-0003-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch (3.8K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/9-vPG19-0003-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch)
download | inline diff:
From a93385298c5350b85d2b9033b5ea88d0775e383d Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:46:48 -0700
Subject: [PATCH vPG19 3/4] Ensure all pg_locale.h APIs work with collate_is_c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 55 +++++++++++++++++++++++++++----
1 file changed, 49 insertions(+), 6 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index da32590396f..2b501a7b6b7 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1441,7 +1441,10 @@ pg_downcase_ident(char *dst, size_t dstsize, const char *src, size_t srclen)
int
pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
{
- return locale->collate->strcoll(arg1, arg2, locale);
+ if (locale->collate == NULL)
+ return strcmp(arg1, arg2);
+ else
+ return locale->collate->strcoll(arg1, arg2, locale);
}
/*
@@ -1459,7 +1462,16 @@ int
pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
pg_locale_t locale)
{
- return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
+ if (locale->collate == NULL)
+ {
+ int result = memcmp(arg1, arg2, Min(len1, len2));
+
+ if ((result == 0) && (len1 != len2))
+ result = (len1 < len2) ? -1 : 1;
+ return result;
+ }
+ else
+ return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
}
/*
@@ -1469,6 +1481,9 @@ pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
bool
pg_strxfrm_enabled(pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ return true;
+
/*
* locale->collate->strnxfrm is still a required method, even if it may
* have the wrong behavior, because the planner uses it for estimates in
@@ -1485,7 +1500,10 @@ pg_strxfrm_enabled(pg_locale_t locale)
size_t
pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
{
- return locale->collate->strxfrm(dest, destsize, src, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strxfrm(dest, destsize, src, locale);
}
/*
@@ -1510,6 +1528,16 @@ size_t
pg_strnxfrm(char *dest, size_t destsize, const char *src, size_t srclen,
pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ {
+ if (destsize > srclen)
+ {
+ memcpy(dest, src, srclen);
+ dest[srclen] = '\0';
+ }
+
+ return srclen;
+ }
return locale->collate->strnxfrm(dest, destsize, src, srclen, locale);
}
@@ -1520,7 +1548,10 @@ pg_strnxfrm(char *dest, size_t destsize, const char *src, size_t srclen,
bool
pg_strxfrm_prefix_enabled(pg_locale_t locale)
{
- return (locale->collate->strnxfrm_prefix != NULL);
+ if (locale->collate == NULL)
+ return true;
+ else
+ return (locale->collate->strnxfrm_prefix != NULL);
}
/*
@@ -1532,7 +1563,10 @@ size_t
pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
pg_locale_t locale)
{
- return locale->collate->strxfrm_prefix(dest, destsize, src, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm_prefix(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strxfrm_prefix(dest, destsize, src, locale);
}
/*
@@ -1556,7 +1590,16 @@ size_t
pg_strnxfrm_prefix(char *dest, size_t destsize, const char *src,
size_t srclen, pg_locale_t locale)
{
- return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
+ if (locale->collate == NULL)
+ {
+ size_t len = Min(srclen, destsize);
+
+ if (destsize > 0)
+ memcpy(dest, src, len);
+ return len;
+ }
+ else
+ return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
}
bool
--
2.43.0
[text/x-patch] vPG19-0004-Add-C-test-module-for-pg_locale.h-APIs.patch (13.6K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/10-vPG19-0004-Add-C-test-module-for-pg_locale.h-APIs.patch)
download | inline diff:
From d59ba32bd3a72442a49aee71a740784a1087b94b Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Sat, 15 Aug 2026 12:26:09 -0700
Subject: [PATCH vPG19 4/4] Add C test module for pg_locale.h APIs.
Test the API independently to account for fallback paths that aren't
adequately tested from SQL.
The backport to 18 also tests the previously-supported behavior where
a size of -1 meant that the string was NUL-terminated. That behavior
was later removed in 19.
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/test/modules/Makefile | 1 +
src/test/modules/meson.build | 1 +
src/test/modules/test_pg_locale/.gitignore | 4 +
src/test/modules/test_pg_locale/Makefile | 23 +++
src/test/modules/test_pg_locale/README | 2 +
.../expected/test_pg_locale.out | 38 +++++
src/test/modules/test_pg_locale/meson.build | 33 ++++
.../test_pg_locale/sql/test_pg_locale.sql | 23 +++
.../test_pg_locale/test_pg_locale--1.0.sql | 8 +
.../modules/test_pg_locale/test_pg_locale.c | 153 ++++++++++++++++++
.../test_pg_locale/test_pg_locale.control | 4 +
11 files changed, 290 insertions(+)
create mode 100644 src/test/modules/test_pg_locale/.gitignore
create mode 100644 src/test/modules/test_pg_locale/Makefile
create mode 100644 src/test/modules/test_pg_locale/README
create mode 100644 src/test/modules/test_pg_locale/expected/test_pg_locale.out
create mode 100644 src/test/modules/test_pg_locale/meson.build
create mode 100644 src/test/modules/test_pg_locale/sql/test_pg_locale.sql
create mode 100644 src/test/modules/test_pg_locale/test_pg_locale--1.0.sql
create mode 100644 src/test/modules/test_pg_locale/test_pg_locale.c
create mode 100644 src/test/modules/test_pg_locale/test_pg_locale.control
diff --git a/src/test/modules/Makefile b/src/test/modules/Makefile
index 098bb8142ae..8a2b09bd11e 100644
--- a/src/test/modules/Makefile
+++ b/src/test/modules/Makefile
@@ -41,6 +41,7 @@ SUBDIRS = \
test_oat_hooks \
test_parser \
test_pg_dump \
+ test_pg_locale \
test_plan_advice \
test_predtest \
test_radixtree \
diff --git a/src/test/modules/meson.build b/src/test/modules/meson.build
index 4bca42bb370..71c4035b1b3 100644
--- a/src/test/modules/meson.build
+++ b/src/test/modules/meson.build
@@ -42,6 +42,7 @@ subdir('test_misc')
subdir('test_oat_hooks')
subdir('test_parser')
subdir('test_pg_dump')
+subdir('test_pg_locale')
subdir('test_plan_advice')
subdir('test_predtest')
subdir('test_radixtree')
diff --git a/src/test/modules/test_pg_locale/.gitignore b/src/test/modules/test_pg_locale/.gitignore
new file mode 100644
index 00000000000..5dcb3ff9723
--- /dev/null
+++ b/src/test/modules/test_pg_locale/.gitignore
@@ -0,0 +1,4 @@
+# Generated subdirectories
+/log/
+/results/
+/tmp_check/
diff --git a/src/test/modules/test_pg_locale/Makefile b/src/test/modules/test_pg_locale/Makefile
new file mode 100644
index 00000000000..9b051f8a697
--- /dev/null
+++ b/src/test/modules/test_pg_locale/Makefile
@@ -0,0 +1,23 @@
+# src/test/modules/test_pg_locale/Makefile
+
+MODULE_big = test_pg_locale
+OBJS = \
+ $(WIN32RES) \
+ test_pg_locale.o
+PGFILEDESC = "test_pg_locale - test code for pg_locale.h APIs"
+
+EXTENSION = test_pg_locale
+DATA = test_pg_locale--1.0.sql
+
+REGRESS = test_pg_locale
+
+ifdef USE_PGXS
+PG_CONFIG = pg_config
+PGXS := $(shell $(PG_CONFIG) --pgxs)
+include $(PGXS)
+else
+subdir = src/test/modules/test_pg_locale
+top_builddir = ../../../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+endif
diff --git a/src/test/modules/test_pg_locale/README b/src/test/modules/test_pg_locale/README
new file mode 100644
index 00000000000..d95af97c005
--- /dev/null
+++ b/src/test/modules/test_pg_locale/README
@@ -0,0 +1,2 @@
+Calls pg_locale.h wrappers directly. Ordinary SQL tests do not reach
+the C-locale fallbacks because in-tree callers special-case collate_is_c.
diff --git a/src/test/modules/test_pg_locale/expected/test_pg_locale.out b/src/test/modules/test_pg_locale/expected/test_pg_locale.out
new file mode 100644
index 00000000000..edbba284552
--- /dev/null
+++ b/src/test/modules/test_pg_locale/expected/test_pg_locale.out
@@ -0,0 +1,38 @@
+CREATE EXTENSION test_pg_locale;
+--
+-- These tests don't produce any interesting output. We're checking that
+-- the operations complete without crashing and that none of their internal
+-- sanity tests fail.
+--
+-- to_regcollation() returns NULL when the collation is absent or is not
+-- usable in the current database encoding. The function is STRICT, so
+-- those cases are skipped.
+--
+-- Libc C. Available in every database.
+SELECT test_pg_locale_apis(to_regcollation('"C"'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- Builtin C (collate and ctype). Usable only in UTF8 databases.
+SELECT test_pg_locale_apis(to_regcollation('ucs_basic'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- Builtin C.UTF-8 (C collate, Unicode ctype). Same encoding restriction.
+SELECT test_pg_locale_apis(to_regcollation('pg_c_utf8'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- en-x-icu is present when ICU collations were imported at initdb.
+SELECT test_pg_locale_apis(to_regcollation('en-x-icu'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
diff --git a/src/test/modules/test_pg_locale/meson.build b/src/test/modules/test_pg_locale/meson.build
new file mode 100644
index 00000000000..f5097464ad2
--- /dev/null
+++ b/src/test/modules/test_pg_locale/meson.build
@@ -0,0 +1,33 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+test_pg_locale_sources = files(
+ 'test_pg_locale.c',
+)
+
+if host_system == 'windows'
+ test_pg_locale_sources += rc_lib_gen.process(win32ver_rc, extra_args: [
+ '--NAME', 'test_pg_locale',
+ '--FILEDESC', 'test_pg_locale - test code for pg_locale.h APIs',])
+endif
+
+test_pg_locale = shared_module('test_pg_locale',
+ test_pg_locale_sources,
+ kwargs: pg_test_mod_args,
+)
+test_install_libs += test_pg_locale
+
+test_install_data += files(
+ 'test_pg_locale.control',
+ 'test_pg_locale--1.0.sql',
+)
+
+tests += {
+ 'name': 'test_pg_locale',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'test_pg_locale',
+ ],
+ },
+}
diff --git a/src/test/modules/test_pg_locale/sql/test_pg_locale.sql b/src/test/modules/test_pg_locale/sql/test_pg_locale.sql
new file mode 100644
index 00000000000..212012feb9c
--- /dev/null
+++ b/src/test/modules/test_pg_locale/sql/test_pg_locale.sql
@@ -0,0 +1,23 @@
+CREATE EXTENSION test_pg_locale;
+
+--
+-- These tests don't produce any interesting output. We're checking that
+-- the operations complete without crashing and that none of their internal
+-- sanity tests fail.
+--
+-- to_regcollation() returns NULL when the collation is absent or is not
+-- usable in the current database encoding. The function is STRICT, so
+-- those cases are skipped.
+--
+
+-- Libc C. Available in every database.
+SELECT test_pg_locale_apis(to_regcollation('"C"'));
+
+-- Builtin C (collate and ctype). Usable only in UTF8 databases.
+SELECT test_pg_locale_apis(to_regcollation('ucs_basic'));
+
+-- Builtin C.UTF-8 (C collate, Unicode ctype). Same encoding restriction.
+SELECT test_pg_locale_apis(to_regcollation('pg_c_utf8'));
+
+-- en-x-icu is present when ICU collations were imported at initdb.
+SELECT test_pg_locale_apis(to_regcollation('en-x-icu'));
diff --git a/src/test/modules/test_pg_locale/test_pg_locale--1.0.sql b/src/test/modules/test_pg_locale/test_pg_locale--1.0.sql
new file mode 100644
index 00000000000..134c4befa06
--- /dev/null
+++ b/src/test/modules/test_pg_locale/test_pg_locale--1.0.sql
@@ -0,0 +1,8 @@
+/* src/test/modules/test_pg_locale/test_pg_locale--1.0.sql */
+
+-- complain if script is sourced in psql, rather than via CREATE EXTENSION
+\echo Use "CREATE EXTENSION test_pg_locale" to load this file. \quit
+
+CREATE FUNCTION test_pg_locale_apis(oid)
+ RETURNS pg_catalog.void
+ AS 'MODULE_PATHNAME' LANGUAGE C STRICT;
diff --git a/src/test/modules/test_pg_locale/test_pg_locale.c b/src/test/modules/test_pg_locale/test_pg_locale.c
new file mode 100644
index 00000000000..122152f5fcb
--- /dev/null
+++ b/src/test/modules/test_pg_locale/test_pg_locale.c
@@ -0,0 +1,153 @@
+/*-------------------------------------------------------------------------
+ *
+ * test_pg_locale.c
+ * Call pg_locale.h wrappers directly.
+ *
+ * SQL callers special-case collate_is_c, so the C-locale fallbacks are
+ * not reached by ordinary regression tests.
+ *
+ * Copyright (c) 2026, PostgreSQL Global Development Group
+ *
+ * IDENTIFICATION
+ * src/test/modules/test_pg_locale/test_pg_locale.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "fmgr.h"
+#include "utils/pg_locale.h"
+
+PG_MODULE_MAGIC;
+
+static void
+test_case_mapping(pg_locale_t locale)
+{
+ char buf[32];
+ size_t n;
+
+ n = pg_strlower(NULL, 0, "AbC", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strlower() size probe returned %zu, expected 3", n);
+ n = pg_strlower(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strlower() produced \"%s\"", buf);
+
+ n = pg_strupper(NULL, 0, "AbC", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strupper() size probe returned %zu, expected 3", n);
+ n = pg_strupper(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "ABC") != 0)
+ elog(ERROR, "pg_strupper() produced \"%s\"", buf);
+
+ n = pg_strfold(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strfold() produced \"%s\"", buf);
+
+ buf[0] = '\0';
+ n = pg_strtitle(buf, sizeof(buf), "hello-world", 11, locale);
+ if (n != 11)
+ elog(ERROR, "pg_strtitle() returned %zu, expected 11", n);
+ if (locale->ctype_is_c && strcmp(buf, "Hello-World") != 0)
+ elog(ERROR, "pg_strtitle() produced \"%s\"", buf);
+}
+
+static void
+test_collate(pg_locale_t locale)
+{
+ char buf[32];
+ char pfx[8];
+ char x1[8];
+ char x2[8];
+ size_t n;
+
+ if (pg_strcoll("abc", "abc", locale) != 0 ||
+ pg_strncoll("abc", 3, "abc", 3, locale) != 0 ||
+ pg_strcoll("", "", locale) != 0)
+ elog(ERROR, "equal strings did not compare equal");
+
+ if (locale->collate_is_c)
+ {
+ if (locale->collate != NULL)
+ elog(ERROR, "collate_is_c but collate methods are set");
+ if (pg_strcoll("abc", "abd", locale) >= 0 ||
+ pg_strcoll("abd", "abc", locale) <= 0 ||
+ pg_strncoll("ab", 2, "abc", 3, locale) >= 0 ||
+ pg_strncoll("abc", 3, "ab", 2, locale) <= 0 ||
+ pg_strncoll("xyz", 3, "abc", 2, locale) <= 0)
+ elog(ERROR, "C-locale comparison result is wrong");
+
+ if (!pg_strxfrm_enabled(locale))
+ elog(ERROR, "pg_strxfrm_enabled() is false for C locale");
+ n = pg_strnxfrm(NULL, 0, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() size probe returned %zu, expected 3", n);
+ n = pg_strnxfrm(buf, 4, "abc", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strnxfrm() produced \"%s\"", buf);
+ n = pg_strxfrm(buf, "abc", 4, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strxfrm() produced \"%s\"", buf);
+ n = pg_strnxfrm(buf, 3, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() destsize==srclen returned %zu", n);
+ n = pg_strnxfrm(buf, 2, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() short dest returned %zu", n);
+
+ if (!pg_strxfrm_prefix_enabled(locale))
+ elog(ERROR, "pg_strxfrm_prefix_enabled() is false for C locale");
+ n = pg_strnxfrm_prefix(NULL, 0, "abcdef", 6, locale);
+ if (n != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() destsize 0 returned %zu", n);
+ n = pg_strnxfrm_prefix(pfx, 2, "abcdef", 6, locale);
+ if (n != 2 || memcmp(pfx, "ab", 2) != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() produced a wrong prefix");
+ n = pg_strnxfrm_prefix(pfx, sizeof(pfx), "abc", 3, locale);
+ if (n != 3 || memcmp(pfx, "abc", 3) != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() destsize>=srclen produced a wrong result");
+ n = pg_strxfrm_prefix(pfx, "abcdef", 2, locale);
+ if (n != 2 || memcmp(pfx, "ab", 2) != 0)
+ elog(ERROR, "pg_strxfrm_prefix() produced a wrong prefix");
+
+ if (pg_strxfrm(x1, "abc", sizeof(x1), locale) >= sizeof(x1) ||
+ pg_strxfrm(x2, "abd", sizeof(x2), locale) >= sizeof(x2) ||
+ (strcmp(x1, x2) < 0) != (pg_strcoll("abc", "abd", locale) < 0))
+ elog(ERROR, "pg_strxfrm() disagrees with pg_strcoll()");
+ }
+ else
+ {
+ char *tmp;
+
+ if (locale->collate == NULL)
+ elog(ERROR, "collate methods missing for non-C locale");
+
+ n = pg_strnxfrm(NULL, 0, "abc", 3, locale);
+ tmp = palloc(n + 1);
+ if (pg_strnxfrm(tmp, n + 1, "abc", 3, locale) > n)
+ elog(ERROR, "pg_strnxfrm() grew on the second call");
+ pfree(tmp);
+
+ if (pg_strxfrm_prefix_enabled(locale) &&
+ pg_strnxfrm_prefix(pfx, sizeof(pfx), "abc", 3, locale) > sizeof(pfx))
+ elog(ERROR, "pg_strnxfrm_prefix() exceeded destsize");
+ }
+}
+
+PG_FUNCTION_INFO_V1(test_pg_locale_apis);
+
+Datum
+test_pg_locale_apis(PG_FUNCTION_ARGS)
+{
+ pg_locale_t locale;
+
+ locale = pg_newlocale_from_collation(PG_GETARG_OID(0));
+ if (locale == NULL)
+ elog(ERROR, "pg_newlocale_from_collation() returned NULL");
+
+ test_collate(locale);
+ test_case_mapping(locale);
+
+ PG_RETURN_VOID();
+}
diff --git a/src/test/modules/test_pg_locale/test_pg_locale.control b/src/test/modules/test_pg_locale/test_pg_locale.control
new file mode 100644
index 00000000000..6b224d04a1b
--- /dev/null
+++ b/src/test/modules/test_pg_locale/test_pg_locale.control
@@ -0,0 +1,4 @@
+comment = 'Test code for pg_locale.h APIs'
+default_version = '1.0'
+module_pathname = '$libdir/test_pg_locale'
+relocatable = true
--
2.43.0
[text/x-patch] vPG20-0001-Add-missing-comments-in-pg_locale.c.patch (5.9K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/11-vPG20-0001-Add-missing-comments-in-pg_locale.c.patch)
download | inline diff:
From 497a4685ddbcf2880a076604721a3e2b53453f5e Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:32:19 -0700
Subject: [PATCH vPG20 1/3] Add missing comments in pg_locale.c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v3nniwcrxejmcfvz56xbd22hphprqleuornd6hqkmw2bl7kgmz@cnytz2ee5ltk
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 70 ++++++++++++++++++++++++++-----
1 file changed, 60 insertions(+), 10 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index 9eb99487e57..da32590396f 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1317,6 +1317,20 @@ strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
return srclen;
}
+/*
+ * pg_strlower()
+ *
+ * Convert src to lowercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strlower(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1327,6 +1341,20 @@ pg_strlower(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strlower(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strtitle()
+ *
+ * Convert src to titlecase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strtitle(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1337,6 +1365,20 @@ pg_strtitle(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strtitle(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strupper()
+ *
+ * Convert src to uppercase, and return the result length (not including
+ * terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strupper(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1347,6 +1389,19 @@ pg_strupper(char *dst, size_t dstsize, const char *src, size_t srclen,
return locale->ctype->strupper(dst, dstsize, src, srclen, locale);
}
+/*
+ * pg_strfold()
+ *
+ * Casefold src, and return the result length (not including terminating NUL).
+ *
+ * src must be in the database encoding with no embedded NULs. If dstsize is
+ * zero, dst may be NULL, which is useful for calculating the required buffer
+ * size before allocating.
+ *
+ * If the result length is less than dstsize, the NUL-terminated result is
+ * stored in dst. Otherwise, the contents of dst are undefined, and the
+ * caller should use the return value to resize the buffer and retry.
+ */
size_t
pg_strfold(char *dst, size_t dstsize, const char *src, size_t srclen,
pg_locale_t locale)
@@ -1392,11 +1447,9 @@ pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
/*
* pg_strncoll
*
- * Call ucol_strcollUTF8(), ucol_strcoll(), strcoll_l() or wcscoll_l() as
- * appropriate for the given locale, platform, and database encoding. If the
- * locale is not specified, use the database collation.
+ * Compare strings according to the given locale.
*
- * The input strings must be encoded in the database encoding.
+ * Strings must be encoded in the database encoding with no embedded NULs.
*
* The caller is responsible for breaking ties if the collation is
* deterministic; this maintains consistency with pg_strnxfrm(), which cannot
@@ -1412,9 +1465,6 @@ pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
/*
* Return true if the collation provider supports pg_strxfrm() and
* pg_strnxfrm(); otherwise false.
- *
- *
- * No similar problem is known for the ICU provider.
*/
bool
pg_strxfrm_enabled(pg_locale_t locale)
@@ -1445,8 +1495,8 @@ pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
* ordinary strcmp() on transformed strings is equivalent to pg_strcoll() on
* untransformed strings.
*
- * The input string must be encoded in the database encoding. If 'destsize' is
- * zero, 'dest' may be NULL.
+ * String must be encoded in the database encoding with no embedded NULs. If
+ * 'destsize' is zero, 'dest' may be NULL.
*
* Not all providers support pg_strnxfrm() safely. The caller should check
* pg_strxfrm_enabled() first, otherwise this function may return wrong
@@ -1492,7 +1542,7 @@ pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
* memcmp() on the byte sequence is equivalent to pg_strncoll() on
* untransformed strings. The result is not nul-terminated.
*
- * The input string must be encoded in the database encoding.
+ * String must be encoded in the database encoding with no embedded NULs.
*
* Not all providers support pg_strnxfrm_prefix() safely. The caller should
* check pg_strxfrm_prefix_enabled() first, otherwise this function may return
--
2.43.0
[text/x-patch] vPG20-0002-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch (3.8K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/12-vPG20-0002-Ensure-all-pg_locale.h-APIs-work-with-collate_.patch)
download | inline diff:
From ba30a89ef3c424b4436c91373001668daef068bf Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Wed, 12 Aug 2026 07:46:48 -0700
Subject: [PATCH vPG20 2/3] Ensure all pg_locale.h APIs work with collate_is_c.
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 55 +++++++++++++++++++++++++++----
1 file changed, 49 insertions(+), 6 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index da32590396f..2b501a7b6b7 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1441,7 +1441,10 @@ pg_downcase_ident(char *dst, size_t dstsize, const char *src, size_t srclen)
int
pg_strcoll(const char *arg1, const char *arg2, pg_locale_t locale)
{
- return locale->collate->strcoll(arg1, arg2, locale);
+ if (locale->collate == NULL)
+ return strcmp(arg1, arg2);
+ else
+ return locale->collate->strcoll(arg1, arg2, locale);
}
/*
@@ -1459,7 +1462,16 @@ int
pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
pg_locale_t locale)
{
- return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
+ if (locale->collate == NULL)
+ {
+ int result = memcmp(arg1, arg2, Min(len1, len2));
+
+ if ((result == 0) && (len1 != len2))
+ result = (len1 < len2) ? -1 : 1;
+ return result;
+ }
+ else
+ return locale->collate->strncoll(arg1, len1, arg2, len2, locale);
}
/*
@@ -1469,6 +1481,9 @@ pg_strncoll(const char *arg1, size_t len1, const char *arg2, size_t len2,
bool
pg_strxfrm_enabled(pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ return true;
+
/*
* locale->collate->strnxfrm is still a required method, even if it may
* have the wrong behavior, because the planner uses it for estimates in
@@ -1485,7 +1500,10 @@ pg_strxfrm_enabled(pg_locale_t locale)
size_t
pg_strxfrm(char *dest, const char *src, size_t destsize, pg_locale_t locale)
{
- return locale->collate->strxfrm(dest, destsize, src, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strxfrm(dest, destsize, src, locale);
}
/*
@@ -1510,6 +1528,16 @@ size_t
pg_strnxfrm(char *dest, size_t destsize, const char *src, size_t srclen,
pg_locale_t locale)
{
+ if (locale->collate == NULL)
+ {
+ if (destsize > srclen)
+ {
+ memcpy(dest, src, srclen);
+ dest[srclen] = '\0';
+ }
+
+ return srclen;
+ }
return locale->collate->strnxfrm(dest, destsize, src, srclen, locale);
}
@@ -1520,7 +1548,10 @@ pg_strnxfrm(char *dest, size_t destsize, const char *src, size_t srclen,
bool
pg_strxfrm_prefix_enabled(pg_locale_t locale)
{
- return (locale->collate->strnxfrm_prefix != NULL);
+ if (locale->collate == NULL)
+ return true;
+ else
+ return (locale->collate->strnxfrm_prefix != NULL);
}
/*
@@ -1532,7 +1563,10 @@ size_t
pg_strxfrm_prefix(char *dest, const char *src, size_t destsize,
pg_locale_t locale)
{
- return locale->collate->strxfrm_prefix(dest, destsize, src, locale);
+ if (locale->collate == NULL)
+ return pg_strnxfrm_prefix(dest, destsize, src, strlen(src), locale);
+ else
+ return locale->collate->strxfrm_prefix(dest, destsize, src, locale);
}
/*
@@ -1556,7 +1590,16 @@ size_t
pg_strnxfrm_prefix(char *dest, size_t destsize, const char *src,
size_t srclen, pg_locale_t locale)
{
- return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
+ if (locale->collate == NULL)
+ {
+ size_t len = Min(srclen, destsize);
+
+ if (destsize > 0)
+ memcpy(dest, src, len);
+ return len;
+ }
+ else
+ return locale->collate->strnxfrm_prefix(dest, destsize, src, srclen, locale);
}
bool
--
2.43.0
[text/x-patch] vPG20-0003-Add-C-test-module-for-pg_locale.h-APIs.patch (13.6K, ../../6c479549b51b738da3065fd979af8bd1ea80b654.camel@j-davis.com/13-vPG20-0003-Add-C-test-module-for-pg_locale.h-APIs.patch)
download | inline diff:
From 4b201e6dd10d59ea5b8d636f03b73e4d5ff4fdc2 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Sat, 15 Aug 2026 12:26:09 -0700
Subject: [PATCH vPG20 3/3] Add C test module for pg_locale.h APIs.
Test the API independently to account for fallback paths that aren't
adequately tested from SQL.
The backport to 18 also tests the previously-supported behavior where
a size of -1 meant that the string was NUL-terminated. That behavior
was later removed in 19.
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/test/modules/Makefile | 1 +
src/test/modules/meson.build | 1 +
src/test/modules/test_pg_locale/.gitignore | 4 +
src/test/modules/test_pg_locale/Makefile | 23 +++
src/test/modules/test_pg_locale/README | 2 +
.../expected/test_pg_locale.out | 38 +++++
src/test/modules/test_pg_locale/meson.build | 33 ++++
.../test_pg_locale/sql/test_pg_locale.sql | 23 +++
.../test_pg_locale/test_pg_locale--1.0.sql | 8 +
.../modules/test_pg_locale/test_pg_locale.c | 153 ++++++++++++++++++
.../test_pg_locale/test_pg_locale.control | 4 +
11 files changed, 290 insertions(+)
create mode 100644 src/test/modules/test_pg_locale/.gitignore
create mode 100644 src/test/modules/test_pg_locale/Makefile
create mode 100644 src/test/modules/test_pg_locale/README
create mode 100644 src/test/modules/test_pg_locale/expected/test_pg_locale.out
create mode 100644 src/test/modules/test_pg_locale/meson.build
create mode 100644 src/test/modules/test_pg_locale/sql/test_pg_locale.sql
create mode 100644 src/test/modules/test_pg_locale/test_pg_locale--1.0.sql
create mode 100644 src/test/modules/test_pg_locale/test_pg_locale.c
create mode 100644 src/test/modules/test_pg_locale/test_pg_locale.control
diff --git a/src/test/modules/Makefile b/src/test/modules/Makefile
index 098bb8142ae..8a2b09bd11e 100644
--- a/src/test/modules/Makefile
+++ b/src/test/modules/Makefile
@@ -41,6 +41,7 @@ SUBDIRS = \
test_oat_hooks \
test_parser \
test_pg_dump \
+ test_pg_locale \
test_plan_advice \
test_predtest \
test_radixtree \
diff --git a/src/test/modules/meson.build b/src/test/modules/meson.build
index 4bca42bb370..71c4035b1b3 100644
--- a/src/test/modules/meson.build
+++ b/src/test/modules/meson.build
@@ -42,6 +42,7 @@ subdir('test_misc')
subdir('test_oat_hooks')
subdir('test_parser')
subdir('test_pg_dump')
+subdir('test_pg_locale')
subdir('test_plan_advice')
subdir('test_predtest')
subdir('test_radixtree')
diff --git a/src/test/modules/test_pg_locale/.gitignore b/src/test/modules/test_pg_locale/.gitignore
new file mode 100644
index 00000000000..5dcb3ff9723
--- /dev/null
+++ b/src/test/modules/test_pg_locale/.gitignore
@@ -0,0 +1,4 @@
+# Generated subdirectories
+/log/
+/results/
+/tmp_check/
diff --git a/src/test/modules/test_pg_locale/Makefile b/src/test/modules/test_pg_locale/Makefile
new file mode 100644
index 00000000000..9b051f8a697
--- /dev/null
+++ b/src/test/modules/test_pg_locale/Makefile
@@ -0,0 +1,23 @@
+# src/test/modules/test_pg_locale/Makefile
+
+MODULE_big = test_pg_locale
+OBJS = \
+ $(WIN32RES) \
+ test_pg_locale.o
+PGFILEDESC = "test_pg_locale - test code for pg_locale.h APIs"
+
+EXTENSION = test_pg_locale
+DATA = test_pg_locale--1.0.sql
+
+REGRESS = test_pg_locale
+
+ifdef USE_PGXS
+PG_CONFIG = pg_config
+PGXS := $(shell $(PG_CONFIG) --pgxs)
+include $(PGXS)
+else
+subdir = src/test/modules/test_pg_locale
+top_builddir = ../../../..
+include $(top_builddir)/src/Makefile.global
+include $(top_srcdir)/contrib/contrib-global.mk
+endif
diff --git a/src/test/modules/test_pg_locale/README b/src/test/modules/test_pg_locale/README
new file mode 100644
index 00000000000..d95af97c005
--- /dev/null
+++ b/src/test/modules/test_pg_locale/README
@@ -0,0 +1,2 @@
+Calls pg_locale.h wrappers directly. Ordinary SQL tests do not reach
+the C-locale fallbacks because in-tree callers special-case collate_is_c.
diff --git a/src/test/modules/test_pg_locale/expected/test_pg_locale.out b/src/test/modules/test_pg_locale/expected/test_pg_locale.out
new file mode 100644
index 00000000000..edbba284552
--- /dev/null
+++ b/src/test/modules/test_pg_locale/expected/test_pg_locale.out
@@ -0,0 +1,38 @@
+CREATE EXTENSION test_pg_locale;
+--
+-- These tests don't produce any interesting output. We're checking that
+-- the operations complete without crashing and that none of their internal
+-- sanity tests fail.
+--
+-- to_regcollation() returns NULL when the collation is absent or is not
+-- usable in the current database encoding. The function is STRICT, so
+-- those cases are skipped.
+--
+-- Libc C. Available in every database.
+SELECT test_pg_locale_apis(to_regcollation('"C"'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- Builtin C (collate and ctype). Usable only in UTF8 databases.
+SELECT test_pg_locale_apis(to_regcollation('ucs_basic'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- Builtin C.UTF-8 (C collate, Unicode ctype). Same encoding restriction.
+SELECT test_pg_locale_apis(to_regcollation('pg_c_utf8'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- en-x-icu is present when ICU collations were imported at initdb.
+SELECT test_pg_locale_apis(to_regcollation('en-x-icu'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
diff --git a/src/test/modules/test_pg_locale/meson.build b/src/test/modules/test_pg_locale/meson.build
new file mode 100644
index 00000000000..f5097464ad2
--- /dev/null
+++ b/src/test/modules/test_pg_locale/meson.build
@@ -0,0 +1,33 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+test_pg_locale_sources = files(
+ 'test_pg_locale.c',
+)
+
+if host_system == 'windows'
+ test_pg_locale_sources += rc_lib_gen.process(win32ver_rc, extra_args: [
+ '--NAME', 'test_pg_locale',
+ '--FILEDESC', 'test_pg_locale - test code for pg_locale.h APIs',])
+endif
+
+test_pg_locale = shared_module('test_pg_locale',
+ test_pg_locale_sources,
+ kwargs: pg_test_mod_args,
+)
+test_install_libs += test_pg_locale
+
+test_install_data += files(
+ 'test_pg_locale.control',
+ 'test_pg_locale--1.0.sql',
+)
+
+tests += {
+ 'name': 'test_pg_locale',
+ 'sd': meson.current_source_dir(),
+ 'bd': meson.current_build_dir(),
+ 'regress': {
+ 'sql': [
+ 'test_pg_locale',
+ ],
+ },
+}
diff --git a/src/test/modules/test_pg_locale/sql/test_pg_locale.sql b/src/test/modules/test_pg_locale/sql/test_pg_locale.sql
new file mode 100644
index 00000000000..212012feb9c
--- /dev/null
+++ b/src/test/modules/test_pg_locale/sql/test_pg_locale.sql
@@ -0,0 +1,23 @@
+CREATE EXTENSION test_pg_locale;
+
+--
+-- These tests don't produce any interesting output. We're checking that
+-- the operations complete without crashing and that none of their internal
+-- sanity tests fail.
+--
+-- to_regcollation() returns NULL when the collation is absent or is not
+-- usable in the current database encoding. The function is STRICT, so
+-- those cases are skipped.
+--
+
+-- Libc C. Available in every database.
+SELECT test_pg_locale_apis(to_regcollation('"C"'));
+
+-- Builtin C (collate and ctype). Usable only in UTF8 databases.
+SELECT test_pg_locale_apis(to_regcollation('ucs_basic'));
+
+-- Builtin C.UTF-8 (C collate, Unicode ctype). Same encoding restriction.
+SELECT test_pg_locale_apis(to_regcollation('pg_c_utf8'));
+
+-- en-x-icu is present when ICU collations were imported at initdb.
+SELECT test_pg_locale_apis(to_regcollation('en-x-icu'));
diff --git a/src/test/modules/test_pg_locale/test_pg_locale--1.0.sql b/src/test/modules/test_pg_locale/test_pg_locale--1.0.sql
new file mode 100644
index 00000000000..134c4befa06
--- /dev/null
+++ b/src/test/modules/test_pg_locale/test_pg_locale--1.0.sql
@@ -0,0 +1,8 @@
+/* src/test/modules/test_pg_locale/test_pg_locale--1.0.sql */
+
+-- complain if script is sourced in psql, rather than via CREATE EXTENSION
+\echo Use "CREATE EXTENSION test_pg_locale" to load this file. \quit
+
+CREATE FUNCTION test_pg_locale_apis(oid)
+ RETURNS pg_catalog.void
+ AS 'MODULE_PATHNAME' LANGUAGE C STRICT;
diff --git a/src/test/modules/test_pg_locale/test_pg_locale.c b/src/test/modules/test_pg_locale/test_pg_locale.c
new file mode 100644
index 00000000000..122152f5fcb
--- /dev/null
+++ b/src/test/modules/test_pg_locale/test_pg_locale.c
@@ -0,0 +1,153 @@
+/*-------------------------------------------------------------------------
+ *
+ * test_pg_locale.c
+ * Call pg_locale.h wrappers directly.
+ *
+ * SQL callers special-case collate_is_c, so the C-locale fallbacks are
+ * not reached by ordinary regression tests.
+ *
+ * Copyright (c) 2026, PostgreSQL Global Development Group
+ *
+ * IDENTIFICATION
+ * src/test/modules/test_pg_locale/test_pg_locale.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "fmgr.h"
+#include "utils/pg_locale.h"
+
+PG_MODULE_MAGIC;
+
+static void
+test_case_mapping(pg_locale_t locale)
+{
+ char buf[32];
+ size_t n;
+
+ n = pg_strlower(NULL, 0, "AbC", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strlower() size probe returned %zu, expected 3", n);
+ n = pg_strlower(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strlower() produced \"%s\"", buf);
+
+ n = pg_strupper(NULL, 0, "AbC", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strupper() size probe returned %zu, expected 3", n);
+ n = pg_strupper(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "ABC") != 0)
+ elog(ERROR, "pg_strupper() produced \"%s\"", buf);
+
+ n = pg_strfold(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strfold() produced \"%s\"", buf);
+
+ buf[0] = '\0';
+ n = pg_strtitle(buf, sizeof(buf), "hello-world", 11, locale);
+ if (n != 11)
+ elog(ERROR, "pg_strtitle() returned %zu, expected 11", n);
+ if (locale->ctype_is_c && strcmp(buf, "Hello-World") != 0)
+ elog(ERROR, "pg_strtitle() produced \"%s\"", buf);
+}
+
+static void
+test_collate(pg_locale_t locale)
+{
+ char buf[32];
+ char pfx[8];
+ char x1[8];
+ char x2[8];
+ size_t n;
+
+ if (pg_strcoll("abc", "abc", locale) != 0 ||
+ pg_strncoll("abc", 3, "abc", 3, locale) != 0 ||
+ pg_strcoll("", "", locale) != 0)
+ elog(ERROR, "equal strings did not compare equal");
+
+ if (locale->collate_is_c)
+ {
+ if (locale->collate != NULL)
+ elog(ERROR, "collate_is_c but collate methods are set");
+ if (pg_strcoll("abc", "abd", locale) >= 0 ||
+ pg_strcoll("abd", "abc", locale) <= 0 ||
+ pg_strncoll("ab", 2, "abc", 3, locale) >= 0 ||
+ pg_strncoll("abc", 3, "ab", 2, locale) <= 0 ||
+ pg_strncoll("xyz", 3, "abc", 2, locale) <= 0)
+ elog(ERROR, "C-locale comparison result is wrong");
+
+ if (!pg_strxfrm_enabled(locale))
+ elog(ERROR, "pg_strxfrm_enabled() is false for C locale");
+ n = pg_strnxfrm(NULL, 0, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() size probe returned %zu, expected 3", n);
+ n = pg_strnxfrm(buf, 4, "abc", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strnxfrm() produced \"%s\"", buf);
+ n = pg_strxfrm(buf, "abc", 4, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strxfrm() produced \"%s\"", buf);
+ n = pg_strnxfrm(buf, 3, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() destsize==srclen returned %zu", n);
+ n = pg_strnxfrm(buf, 2, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() short dest returned %zu", n);
+
+ if (!pg_strxfrm_prefix_enabled(locale))
+ elog(ERROR, "pg_strxfrm_prefix_enabled() is false for C locale");
+ n = pg_strnxfrm_prefix(NULL, 0, "abcdef", 6, locale);
+ if (n != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() destsize 0 returned %zu", n);
+ n = pg_strnxfrm_prefix(pfx, 2, "abcdef", 6, locale);
+ if (n != 2 || memcmp(pfx, "ab", 2) != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() produced a wrong prefix");
+ n = pg_strnxfrm_prefix(pfx, sizeof(pfx), "abc", 3, locale);
+ if (n != 3 || memcmp(pfx, "abc", 3) != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() destsize>=srclen produced a wrong result");
+ n = pg_strxfrm_prefix(pfx, "abcdef", 2, locale);
+ if (n != 2 || memcmp(pfx, "ab", 2) != 0)
+ elog(ERROR, "pg_strxfrm_prefix() produced a wrong prefix");
+
+ if (pg_strxfrm(x1, "abc", sizeof(x1), locale) >= sizeof(x1) ||
+ pg_strxfrm(x2, "abd", sizeof(x2), locale) >= sizeof(x2) ||
+ (strcmp(x1, x2) < 0) != (pg_strcoll("abc", "abd", locale) < 0))
+ elog(ERROR, "pg_strxfrm() disagrees with pg_strcoll()");
+ }
+ else
+ {
+ char *tmp;
+
+ if (locale->collate == NULL)
+ elog(ERROR, "collate methods missing for non-C locale");
+
+ n = pg_strnxfrm(NULL, 0, "abc", 3, locale);
+ tmp = palloc(n + 1);
+ if (pg_strnxfrm(tmp, n + 1, "abc", 3, locale) > n)
+ elog(ERROR, "pg_strnxfrm() grew on the second call");
+ pfree(tmp);
+
+ if (pg_strxfrm_prefix_enabled(locale) &&
+ pg_strnxfrm_prefix(pfx, sizeof(pfx), "abc", 3, locale) > sizeof(pfx))
+ elog(ERROR, "pg_strnxfrm_prefix() exceeded destsize");
+ }
+}
+
+PG_FUNCTION_INFO_V1(test_pg_locale_apis);
+
+Datum
+test_pg_locale_apis(PG_FUNCTION_ARGS)
+{
+ pg_locale_t locale;
+
+ locale = pg_newlocale_from_collation(PG_GETARG_OID(0));
+ if (locale == NULL)
+ elog(ERROR, "pg_newlocale_from_collation() returned NULL");
+
+ test_collate(locale);
+ test_case_mapping(locale);
+
+ PG_RETURN_VOID();
+}
diff --git a/src/test/modules/test_pg_locale/test_pg_locale.control b/src/test/modules/test_pg_locale/test_pg_locale.control
new file mode 100644
index 00000000000..6b224d04a1b
--- /dev/null
+++ b/src/test/modules/test_pg_locale/test_pg_locale.control
@@ -0,0 +1,4 @@
+comment = 'Test code for pg_locale.h APIs'
+default_version = '1.0'
+module_pathname = '$libdir/test_pg_locale'
+relocatable = true
--
2.43.0
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-18 18:36 Andres Freund <andres@anarazel.de>
parent: Jeff Davis <pgsql@j-davis.com>
0 siblings, 1 reply; 16+ messages in thread
From: Andres Freund @ 2026-08-18 18:36 UTC (permalink / raw)
To: Jeff Davis <pgsql@j-davis.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>
Hi,
On 2026-08-15 14:53:39 -0700, Jeff Davis wrote:
> On Thu, 2026-08-13 at 22:17 -0700, Jeff Davis wrote:
> > On Tue, 2026-08-11 at 14:31 -0400, Andres Freund wrote:
> > > Hm. If I infer the pg_strlower() API correctly - it's utterly
> > > underdocumented
> >
> > Agreed. Patch attached.
> >
> > > I'd also make i size_t, given that the input is size_t. Perhaps
> > > practically
> > > no problem, but I see no reason to not use size_t here.
> >
> > Patch attached for that, too.
> >
> > I also attached patches to make all the functions work with
> > collate_is_c, and fixed up the -1 API in 18.
>
> Now with a C test module (made with AI assistance).
>
> I plan to start committing these fairly soon. I'm not sure whether to
> backport the C test module, but I included the patches to do so.
Thanks for working on these.
> From f669e9fafb9bf435c946dcc2c2dd98efb4aa46cb Mon Sep 17 00:00:00 2001
> From: Jeff Davis <jeff@j-davis.com>
> Date: Thu, 13 Aug 2026 21:12:34 -0700
> Subject: [PATCH vPG18 2/5] pg_locale.c, unicode_case.c: use size_t for
> iteration.
>
> No actual problem, just cleanup. Only relevant to 18 and 19.
FWIW, I think this is actually a bug, it just turns out that we don't know of
any callers that hit it.
> From a8e72bb34c979ac43ee8480fd10ed4e73db47b10 Mon Sep 17 00:00:00 2001
> From: Jeff Davis <jeff@j-davis.com>
> Date: Wed, 12 Aug 2026 07:32:19 -0700
> Subject: [PATCH vPG18 3/5] Add missing comments in pg_locale.c.
>
> Suggested-by: Andres Freund <andres@anarazel.de>
> Discussion: https://postgr.es/m/v3nniwcrxejmcfvz56xbd22hphprqleuornd6hqkmw2bl7kgmz@cnytz2ee5ltk
> Backpatch-through: 18
Nice, that's much better than before.
> +/*
> + * pg_strtitle()
> + *
> + * Convert src to titlecase, and return the result length (not including
> + * terminating NUL).
Might not hurt to actually say what titlecase and folded strings are.
> From adeef54179a11f5b7817054bc53f8b7b584d2b6a Mon Sep 17 00:00:00 2001
> From: Jeff Davis <jeff@j-davis.com>
> Date: Sat, 15 Aug 2026 12:26:09 -0700
> Subject: [PATCH vPG18 5/5] Add C test module for pg_locale.h APIs.
>
> Test the API independently to account for fallback paths that aren't
> adequately tested from SQL.
>
> The backport to 18 also tests the previously-supported behavior where
> a size of -1 meant that the string was NUL-terminated. That behavior
> was later removed in 19.
>
> Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
> Backpatch-through: 18
> ---
> src/test/modules/Makefile | 1 +
> src/test/modules/meson.build | 1 +
> src/test/modules/test_pg_locale/.gitignore | 4 +
> src/test/modules/test_pg_locale/Makefile | 23 +++
> src/test/modules/test_pg_locale/README | 2 +
> .../expected/test_pg_locale.out | 38 ++++
> src/test/modules/test_pg_locale/meson.build | 33 ++++
> .../test_pg_locale/sql/test_pg_locale.sql | 23 +++
> .../test_pg_locale/test_pg_locale--1.0.sql | 8 +
> .../modules/test_pg_locale/test_pg_locale.c | 169 ++++++++++++++++++
> .../test_pg_locale/test_pg_locale.control | 4 +
> 11 files changed, 306 insertions(+)
> create mode 100644 src/test/modules/test_pg_locale/.gitignore
> create mode 100644 src/test/modules/test_pg_locale/Makefile
> create mode 100644 src/test/modules/test_pg_locale/README
> create mode 100644 src/test/modules/test_pg_locale/expected/test_pg_locale.out
> create mode 100644 src/test/modules/test_pg_locale/meson.build
> create mode 100644 src/test/modules/test_pg_locale/sql/test_pg_locale.sql
> create mode 100644 src/test/modules/test_pg_locale/test_pg_locale--1.0.sql
> create mode 100644 src/test/modules/test_pg_locale/test_pg_locale.c
> create mode 100644 src/test/modules/test_pg_locale/test_pg_locale.control
I very much like that this is tested now, but this really need its own
initdb'd cluster? Every full testrun writes ginormous amounts of data (~73GB
for one master run on macos!), due to the number of clusters we create, and
the amount is growing from release to release at an alarming clip.
Sometimes that's unavoidable, because you need a server configured in a
specific way, the tests take a good while and should therefore run
concurrently, or such. But that shouldn't be the case her. Can't you stuff
this into regress.c or such?
Other than that complaint, I'd probably backpatch this. Seems unlikely to be
flappy or such?
Greetings,
Andres Freund
^ permalink raw reply [nested|flat] 16+ messages in thread
* Re: Crash issue in PG18.5 regression
@ 2026-08-20 18:49 Jeff Davis <pgsql@j-davis.com>
parent: Andres Freund <andres@anarazel.de>
0 siblings, 0 replies; 16+ messages in thread
From: Jeff Davis @ 2026-08-20 18:49 UTC (permalink / raw)
To: Andres Freund <andres@anarazel.de>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Álvaro Herrera <alvherre@kurilemu.de>; Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>; 'pgsql-hackers@lists.postgresql.org' <pgsql-hackers@lists.postgresql.org>
On Tue, 2026-08-18 at 14:36 -0400, Andres Freund wrote:
> FWIW, I think this is actually a bug, it just turns out that we don't
> know of
> any callers that hit it.
Committed.
Also committed the change to support collate_is_c always, for
consistency with ctype_is_c.
>
> > +/*
> > + * pg_strtitle()
> > + *
> > + * Convert src to titlecase, and return the result length (not
> > including
> > + * terminating NUL).
>
> Might not hurt to actually say what titlecase and folded strings are.
Attached follow-up patch that includes explanatory comments, and adds a
couple tiny SQL tests to cover the new examples.
>
> Sometimes that's unavoidable, because you need a server configured in
> a
> specific way, the tests take a good while and should therefore run
> concurrently, or such. But that shouldn't be the case her. Can't you
> stuff
> this into regress.c or such?
Attached new test patch that just adds it to regress.c and calls it
from misc_functions.sql.
> Other than that complaint, I'd probably backpatch this. Seems
> unlikely to be
> flappy or such?
Agreed. It would have caught 27e2afb492.
Thank you for looking at it.
Regards,
Jeff Davis
Attachments:
[text/x-patch] v3.pg20-0001-Add-C-test-function-for-pg_locale.h-APIs.patch (8.4K, ../../fe35594a54923c94142390dabb365b737c581f2a.camel@j-davis.com/2-v3.pg20-0001-Add-C-test-function-for-pg_locale.h-APIs.patch)
download | inline diff:
From 0b631dd835e54d5d2d7596a42cbd2bc38e9d87d8 Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Sat, 15 Aug 2026 12:26:09 -0700
Subject: [PATCH v3.pg20 1/2] Add C test function for pg_locale.h APIs.
Test the API independently to account for fallback paths that aren't
adequately tested from SQL.
The backport to 18 also tests the previously-supported behavior where
a size of -1 meant that the string was NUL-terminated. That behavior
was later removed in 19.
Reviewed-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/test/regress/expected/misc_functions.out | 40 ++++++
src/test/regress/regress.c | 134 +++++++++++++++++++
src/test/regress/sql/misc_functions.sql | 25 ++++
3 files changed, 199 insertions(+)
diff --git a/src/test/regress/expected/misc_functions.out b/src/test/regress/expected/misc_functions.out
index c3261bff209..2990e0c4f28 100644
--- a/src/test/regress/expected/misc_functions.out
+++ b/src/test/regress/expected/misc_functions.out
@@ -861,3 +861,43 @@ SELECT test_instr_time();
t
(1 row)
+--
+-- C tests for pg_locale.h APIs. No interesting output; tests will
+-- ERROR upon failure.
+--
+-- The test function is STRICT, so tests will be skipped if the
+-- collation is unavailable in the current database encoding
+-- (to_regcollation() will return NULL).
+--
+CREATE FUNCTION test_pg_locale_apis(oid)
+ RETURNS void
+ AS :'regresslib'
+ LANGUAGE C STRICT;
+-- Libc C. Available in every database.
+SELECT test_pg_locale_apis(to_regcollation('"C"'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- Builtin C (collate and ctype). Usable only in UTF8 databases.
+SELECT test_pg_locale_apis(to_regcollation('ucs_basic'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- Builtin C.UTF-8 (C collate, Unicode ctype). Same encoding restriction.
+SELECT test_pg_locale_apis(to_regcollation('pg_c_utf8'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
+-- en-x-icu is present when ICU collations were imported at initdb.
+SELECT test_pg_locale_apis(to_regcollation('en-x-icu'));
+ test_pg_locale_apis
+---------------------
+
+(1 row)
+
diff --git a/src/test/regress/regress.c b/src/test/regress/regress.c
index c2975ffc1b0..c72ee31cdce 100644
--- a/src/test/regress/regress.c
+++ b/src/test/regress/regress.c
@@ -48,6 +48,7 @@
#include "utils/builtins.h"
#include "utils/geo_decls.h"
#include "utils/memutils.h"
+#include "utils/pg_locale.h"
#include "utils/rel.h"
#include "utils/typcache.h"
@@ -1498,3 +1499,136 @@ test_pglz_decompress(PG_FUNCTION_ARGS)
SET_VARSIZE(result, dlen + VARHDRSZ);
PG_RETURN_BYTEA_P(result);
}
+
+static void
+test_case_mapping(pg_locale_t locale)
+{
+ char buf[32];
+ size_t n;
+
+ n = pg_strlower(NULL, 0, "AbC", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strlower() size probe returned %zu, expected 3", n);
+ n = pg_strlower(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strlower() produced \"%s\"", buf);
+
+ n = pg_strupper(NULL, 0, "AbC", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strupper() size probe returned %zu, expected 3", n);
+ n = pg_strupper(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "ABC") != 0)
+ elog(ERROR, "pg_strupper() produced \"%s\"", buf);
+
+ n = pg_strfold(buf, 4, "AbC", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strfold() produced \"%s\"", buf);
+
+ buf[0] = '\0';
+ n = pg_strtitle(buf, sizeof(buf), "hello-world", 11, locale);
+ if (n != 11)
+ elog(ERROR, "pg_strtitle() returned %zu, expected 11", n);
+ if (locale->ctype_is_c && strcmp(buf, "Hello-World") != 0)
+ elog(ERROR, "pg_strtitle() produced \"%s\"", buf);
+}
+
+static void
+test_collate(pg_locale_t locale)
+{
+ char buf[32];
+ char pfx[8];
+ char x1[8];
+ char x2[8];
+ size_t n;
+
+ if (pg_strcoll("abc", "abc", locale) != 0 ||
+ pg_strncoll("abc", 3, "abc", 3, locale) != 0 ||
+ pg_strcoll("", "", locale) != 0)
+ elog(ERROR, "equal strings did not compare equal");
+
+ if (locale->collate_is_c)
+ {
+ if (locale->collate != NULL)
+ elog(ERROR, "collate_is_c but collate methods are set");
+ if (pg_strcoll("abc", "abd", locale) >= 0 ||
+ pg_strcoll("abd", "abc", locale) <= 0 ||
+ pg_strncoll("ab", 2, "abc", 3, locale) >= 0 ||
+ pg_strncoll("abc", 3, "ab", 2, locale) <= 0 ||
+ pg_strncoll("xyz", 3, "abc", 2, locale) <= 0)
+ elog(ERROR, "C-locale comparison result is wrong");
+
+ if (!pg_strxfrm_enabled(locale))
+ elog(ERROR, "pg_strxfrm_enabled() is false for C locale");
+ n = pg_strnxfrm(NULL, 0, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() size probe returned %zu, expected 3", n);
+ n = pg_strnxfrm(buf, 4, "abc", 3, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strnxfrm() produced \"%s\"", buf);
+ n = pg_strxfrm(buf, "abc", 4, locale);
+ if (n != 3 || strcmp(buf, "abc") != 0)
+ elog(ERROR, "pg_strxfrm() produced \"%s\"", buf);
+ n = pg_strnxfrm(buf, 3, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() destsize==srclen returned %zu", n);
+ n = pg_strnxfrm(buf, 2, "abc", 3, locale);
+ if (n != 3)
+ elog(ERROR, "pg_strnxfrm() short dest returned %zu", n);
+
+ if (!pg_strxfrm_prefix_enabled(locale))
+ elog(ERROR, "pg_strxfrm_prefix_enabled() is false for C locale");
+ n = pg_strnxfrm_prefix(NULL, 0, "abcdef", 6, locale);
+ if (n != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() destsize 0 returned %zu", n);
+ n = pg_strnxfrm_prefix(pfx, 2, "abcdef", 6, locale);
+ if (n != 2 || memcmp(pfx, "ab", 2) != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() produced a wrong prefix");
+ n = pg_strnxfrm_prefix(pfx, sizeof(pfx), "abc", 3, locale);
+ if (n != 3 || memcmp(pfx, "abc", 3) != 0)
+ elog(ERROR, "pg_strnxfrm_prefix() destsize>=srclen produced a wrong result");
+ n = pg_strxfrm_prefix(pfx, "abcdef", 2, locale);
+ if (n != 2 || memcmp(pfx, "ab", 2) != 0)
+ elog(ERROR, "pg_strxfrm_prefix() produced a wrong prefix");
+
+ if (pg_strxfrm(x1, "abc", sizeof(x1), locale) >= sizeof(x1) ||
+ pg_strxfrm(x2, "abd", sizeof(x2), locale) >= sizeof(x2) ||
+ (strcmp(x1, x2) < 0) != (pg_strcoll("abc", "abd", locale) < 0))
+ elog(ERROR, "pg_strxfrm() disagrees with pg_strcoll()");
+ }
+ else
+ {
+ char *tmp;
+
+ if (locale->collate == NULL)
+ elog(ERROR, "collate methods missing for non-C locale");
+
+ n = pg_strnxfrm(NULL, 0, "abc", 3, locale);
+ tmp = palloc(n + 1);
+ if (pg_strnxfrm(tmp, n + 1, "abc", 3, locale) > n)
+ elog(ERROR, "pg_strnxfrm() grew on the second call");
+ pfree(tmp);
+
+ if (pg_strxfrm_prefix_enabled(locale) &&
+ pg_strnxfrm_prefix(pfx, sizeof(pfx), "abc", 3, locale) > sizeof(pfx))
+ elog(ERROR, "pg_strnxfrm_prefix() exceeded destsize");
+ }
+}
+
+/*
+ * Test pg_locale.h APIs directly, to cover cases not easily reachable by SQL.
+ */
+PG_FUNCTION_INFO_V1(test_pg_locale_apis);
+Datum
+test_pg_locale_apis(PG_FUNCTION_ARGS)
+{
+ pg_locale_t locale;
+
+ locale = pg_newlocale_from_collation(PG_GETARG_OID(0));
+ if (locale == NULL)
+ elog(ERROR, "pg_newlocale_from_collation() returned NULL");
+
+ test_collate(locale);
+ test_case_mapping(locale);
+
+ PG_RETURN_VOID();
+}
diff --git a/src/test/regress/sql/misc_functions.sql b/src/test/regress/sql/misc_functions.sql
index 946ee5726cd..950d9ab1a4a 100644
--- a/src/test/regress/sql/misc_functions.sql
+++ b/src/test/regress/sql/misc_functions.sql
@@ -356,3 +356,28 @@ CREATE FUNCTION test_instr_time()
AS :'regresslib'
LANGUAGE C;
SELECT test_instr_time();
+
+--
+-- C tests for pg_locale.h APIs. No interesting output; tests will
+-- ERROR upon failure.
+--
+-- The test function is STRICT, so tests will be skipped if the
+-- collation is unavailable in the current database encoding
+-- (to_regcollation() will return NULL).
+--
+CREATE FUNCTION test_pg_locale_apis(oid)
+ RETURNS void
+ AS :'regresslib'
+ LANGUAGE C STRICT;
+
+-- Libc C. Available in every database.
+SELECT test_pg_locale_apis(to_regcollation('"C"'));
+
+-- Builtin C (collate and ctype). Usable only in UTF8 databases.
+SELECT test_pg_locale_apis(to_regcollation('ucs_basic'));
+
+-- Builtin C.UTF-8 (C collate, Unicode ctype). Same encoding restriction.
+SELECT test_pg_locale_apis(to_regcollation('pg_c_utf8'));
+
+-- en-x-icu is present when ICU collations were imported at initdb.
+SELECT test_pg_locale_apis(to_regcollation('en-x-icu'));
--
2.43.0
[text/x-patch] v3.pg20-0002-pg_locale.c-add-explanatory-comments-and-tes.patch (7.9K, ../../fe35594a54923c94142390dabb365b737c581f2a.camel@j-davis.com/3-v3.pg20-0002-pg_locale.c-add-explanatory-comments-and-tes.patch)
download | inline diff:
From 83d7b5f26799bb27038036c01dbefaf72d6aee4b Mon Sep 17 00:00:00 2001
From: Jeff Davis <jeff@j-davis.com>
Date: Tue, 18 Aug 2026 18:56:24 -0700
Subject: [PATCH v3.pg20 2/2] pg_locale.c: add explanatory comments and test
the examples.
Explain the purpose and caveats of case conversion functions, rather
than just the API. Also add tests to cover the examples.
Reviewed-by: Andres Freund <andres@anarazel.de>
Suggested-by: Andres Freund <andres@anarazel.de>
Discussion: https://postgr.es/m/v36ssaygf7grb3qzfsjhtdzi7kqd45ds56nyuf7gi5qjml4qbb@ezmfqzmhlrs2
Backpatch-through: 18
---
src/backend/utils/adt/pg_locale.c | 72 +++++++++++++++++++++-
src/test/regress/expected/collate.utf8.out | 8 ++-
src/test/regress/sql/collate.utf8.sql | 4 +-
3 files changed, 80 insertions(+), 4 deletions(-)
diff --git a/src/backend/utils/adt/pg_locale.c b/src/backend/utils/adt/pg_locale.c
index b6e6ee8928b..af9fdd87bf4 100644
--- a/src/backend/utils/adt/pg_locale.c
+++ b/src/backend/utils/adt/pg_locale.c
@@ -1317,12 +1317,58 @@ strupper_c(char *dst, size_t dstsize, const char *src, size_t srclen)
return srclen;
}
+/*
+ * Case Mapping Complexities
+ *
+ * pg_strlower(), pg_strtitle(), pg_strupper(), and pg_strfold() are based on
+ * case mapping. Below are general notes on the complexities of Unicode case
+ * mapping, which are relevant to most locales other than "C", but which vary
+ * significantly among different providers and locales. See the Unicode
+ * Standard and the provider implementation for details.
+ *
+ * Some characters map to more than one other character, so the result may
+ * have more characters than the original. Unicode defines the maximum string
+ * expansion to be 3x the code points (not necessarily bytes); see Unicode
+ * 17.0 section 5.18.2. Examples: U+0130 LATIN CAPITAL LETTER I WITH DOT
+ * ABOVE lowercases to "i" followed by U+0307 COMBINING DOT ABOVE; U+FB01
+ * LATIN SMALL LIGATURE FI titlecases to "Fi"; U+0390 GREEK SMALL LETTER IOTA
+ * WITH DIALYTIKA AND TONOS uppercases to <0399 0308 0301>; U+00DF LATIN SMALL
+ * LETTER SHARP S casefolds to "ss".
+ *
+ * Some characters have more than two forms, e.g. U+03A3 GREEK CAPITAL LETTER
+ * SIGMA, U+03C3 GREEK SMALL LETTER SIGMA, and U+03C2 GREEK SMALL LETTER FINAL
+ * SIGMA (the first is uppercase and the latter two are both lowercase). The
+ * form used depends on context within the string.
+ *
+ * Some characters have special titlecase forms to use for the initial letter
+ * of a word, if available; otherwise uppercase is used. Example: U+01F2
+ * LATIN CAPITAL LETTER D WITH SMALL LETTER Z.
+ *
+ * Titlecasing requires finding a word boundary, which is dependent on the
+ * provider and locale. The semantics of identifying word boundaries may
+ * differ from the semantics used to choose a particular form (e.g. U+03C3
+ * GREEK SMALL LETTER SIGMA vs. U+03C2 GREEK SMALL LETTER FINAL SIGMA).
+ *
+ * Mappings may depend on the provider and the version of Unicode on which it
+ * is based. If mapping only assigned code points, the results of casefolding
+ * are guaranteed to be stable across Unicode versions. Unassigned code
+ * points map to themselves, so are subject to change if the provider updates
+ * Unicode. Therefore, casefolding strings of assigned code points is the
+ * safest mapping for callers that will store the result, e.g. an expression
+ * index.
+ */
+
/*
* pg_strlower()
*
* Convert src to lowercase, and return the result length (not including
* terminating NUL).
*
+ * Lowercasing is intended for human-readable display. If the goal is to
+ * convert to a canonical caseless form, see pg_strfold().
+ *
+ * See Case Mapping Complexities comment above.
+ *
* src must be in the database encoding with no embedded NULs. If dstsize is
* zero, dst may be NULL, which is useful for calculating the required buffer
* size before allocating.
@@ -1347,6 +1393,13 @@ pg_strlower(char *dst, size_t dstsize, const char *src, size_t srclen,
* Convert src to titlecase, and return the result length (not including
* terminating NUL).
*
+ * Titlecasing is intended for human-readable display. A titlecase string has
+ * the initial letter of each word uppercased (or changed to a special
+ * titlecase form, if available), and all other characters lowercased. Used
+ * to implement the SQL INITCAP() function.
+ *
+ * See Case Mapping Complexities comment above.
+ *
* src must be in the database encoding with no embedded NULs. If dstsize is
* zero, dst may be NULL, which is useful for calculating the required buffer
* size before allocating.
@@ -1371,6 +1424,11 @@ pg_strtitle(char *dst, size_t dstsize, const char *src, size_t srclen,
* Convert src to uppercase, and return the result length (not including
* terminating NUL).
*
+ * Uppercasing is intended for human-readable display. If the goal is to
+ * convert to a canonical caseless form, see pg_strfold().
+ *
+ * See Case Mapping Complexities comment above.
+ *
* src must be in the database encoding with no embedded NULs. If dstsize is
* zero, dst may be NULL, which is useful for calculating the required buffer
* size before allocating.
@@ -1392,7 +1450,19 @@ pg_strupper(char *dst, size_t dstsize, const char *src, size_t srclen,
/*
* pg_strfold()
*
- * Casefold src, and return the result length (not including terminating NUL).
+ * Casefold src, and return the result length (not including terminating
+ * NUL).
+ *
+ * Casefolding produces a canonical string such that, iff the casefolded
+ * strings are equal, the original strings are a case-insensitive match (the
+ * strength of this guarantee depends on normalization, provider and locale).
+ * In practice the result is similar to lowercasing, but the purpose is
+ * different: lowercasing is for human-readable display; whereas casefolding
+ * is meant to canonicalize complex mappings reliably without regard for
+ * display. Unicode guarantees that casefolding is stable across versions if
+ * the original string consists only of assigned code points.
+ *
+ * See Case Mapping Complexities comment above.
*
* src must be in the database encoding with no embedded NULs. If dstsize is
* zero, dst may be NULL, which is useful for calculating the required buffer
diff --git a/src/test/regress/expected/collate.utf8.out b/src/test/regress/expected/collate.utf8.out
index cdd1a37ba18..3b8d6392732 100644
--- a/src/test/regress/expected/collate.utf8.out
+++ b/src/test/regress/expected/collate.utf8.out
@@ -194,7 +194,9 @@ INSERT INTO test_pg_unicode_fast VALUES
(U&'Λλ 1a \FF11a'),
('ȺȺȺ'),
('ⱥⱥⱥ'),
- ('ⱥȺ');
+ ('ⱥȺ'),
+ (U&'\FB01'),
+ (U&'\0390');
SELECT
t, lower(t), initcap(t), upper(t),
length(convert_to(t, 'UTF8')) AS t_bytes,
@@ -211,7 +213,9 @@ SELECT
ȺȺȺ | ⱥⱥⱥ | Ⱥⱥⱥ | ȺȺȺ | 6 | 9 | 8 | 6
ⱥⱥⱥ | ⱥⱥⱥ | Ⱥⱥⱥ | ȺȺȺ | 9 | 9 | 8 | 6
ⱥȺ | ⱥⱥ | Ⱥⱥ | ȺȺ | 5 | 6 | 5 | 4
-(7 rows)
+ fi | fi | Fi | FI | 3 | 3 | 2 | 2
+ ΐ | ΐ | Ϊ́ | Ϊ́ | 2 | 2 | 6 | 6
+(9 rows)
DROP TABLE test_pg_unicode_fast;
-- test Final_Sigma
diff --git a/src/test/regress/sql/collate.utf8.sql b/src/test/regress/sql/collate.utf8.sql
index 52cf068dd0c..6875f6f9e35 100644
--- a/src/test/regress/sql/collate.utf8.sql
+++ b/src/test/regress/sql/collate.utf8.sql
@@ -107,7 +107,9 @@ INSERT INTO test_pg_unicode_fast VALUES
(U&'Λλ 1a \FF11a'),
('ȺȺȺ'),
('ⱥⱥⱥ'),
- ('ⱥȺ');
+ ('ⱥȺ'),
+ (U&'\FB01'),
+ (U&'\0390');
SELECT
t, lower(t), initcap(t), upper(t),
--
2.43.0
^ permalink raw reply [nested|flat] 16+ messages in thread
end of thread, other threads:[~2026-08-20 18:49 UTC | newest]
Thread overview: 16+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2026-08-11 05:16 Crash issue in PG18.5 regression Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>
2026-08-11 06:32 ` Álvaro Herrera <alvherre@kurilemu.de>
2026-08-11 09:41 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-08-11 10:23 ` Álvaro Herrera <alvherre@kurilemu.de>
2026-08-11 11:23 ` Masashi Kamura (Fujitsu) <kamura.masashi@fujitsu.com>
2026-08-11 16:55 ` Andres Freund <andres@anarazel.de>
2026-08-11 17:02 ` Tom Lane <tgl@sss.pgh.pa.us>
2026-08-11 18:26 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-08-11 18:31 ` Andres Freund <andres@anarazel.de>
2026-08-11 20:04 ` Tom Lane <tgl@sss.pgh.pa.us>
2026-08-11 23:05 ` Andres Freund <andres@anarazel.de>
2026-08-14 05:17 ` Jeff Davis <pgsql@j-davis.com>
2026-08-15 21:53 ` Jeff Davis <pgsql@j-davis.com>
2026-08-18 18:36 ` Andres Freund <andres@anarazel.de>
2026-08-20 18:49 ` Jeff Davis <pgsql@j-davis.com>
2026-08-11 19:20 ` Jeff Davis <pgsql@j-davis.com>
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox