pg.ddx.io pgsql-committers@postgresql.org mailing list archivehelp / color / mirror / Atom feed
pgsql: Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 4+ messages / 1 participants [nested] [flat]
* pgsql: Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 @ 2026-08-26 20:25 Andrew Dunstan <andrew@dunslane.net> 0 siblings, 0 replies; 4+ messages in thread From: Andrew Dunstan @ 2026-08-26 20:25 UTC (permalink / raw) To: pgsql-committers@lists.postgresql.org Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 test_regex_utf8 decides whether to run by looking at the database encoding alone, but two of its cases, [[:graph:]] and [[:print:]] over E'xᔀሷ', depend on the ctype as well. In a database with encoding UTF8 and locale C they match just the x, because isgraph() and isprint() are false for anything outside ASCII, and the file fails. No buildfarm animal builds such a cluster, which is why this went unnoticed, and why the to_date() crash in 18.5 went undetected for want of exactly this coverage. A pending buildfarm client change will let an animal be configured that way. Fix by giving the two cases an explicit collation, so that they exercise a fixed Unicode ctype instead of whatever the database happened to be initialized with. test_regex() already passes its input collation down to the regex compiler. The expected results are unchanged; only the echoed queries differ. Backpatch-through: 17 (15 and 16 get a different fix) Reviewed-by: Jonathan Gonzalez V. <jonathan@abdiel.eu> Branch ------ master Details ------- https://git.postgresql.org/pg/commitdiff/b941cace8b2547c9b597fd43c880ef23c26184ea Modified Files -------------- src/test/modules/test_regex/expected/test_regex_utf8.out | 6 ++++-- src/test/modules/test_regex/sql/test_regex_utf8.sql | 6 ++++-- 2 files changed, 8 insertions(+), 4 deletions(-) ^ permalink raw reply [nested|flat] 4+ messages in thread
* pgsql: Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 @ 2026-08-26 20:25 Andrew Dunstan <andrew@dunslane.net> 0 siblings, 0 replies; 4+ messages in thread From: Andrew Dunstan @ 2026-08-26 20:25 UTC (permalink / raw) To: pgsql-committers@lists.postgresql.org Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 test_regex_utf8 decides whether to run by looking at the database encoding alone, but two of its cases, [[:graph:]] and [[:print:]] over E'xᔀሷ', depend on the ctype as well. In a database with encoding UTF8 and locale C they match just the x, because isgraph() and isprint() are false for anything outside ASCII, and the file fails. No buildfarm animal builds such a cluster, which is why this went unnoticed, and why the to_date() crash in 18.5 went undetected for want of exactly this coverage. A pending buildfarm client change will let an animal be configured that way. Fix by giving the two cases an explicit collation, so that they exercise a fixed Unicode ctype instead of whatever the database happened to be initialized with. test_regex() already passes its input collation down to the regex compiler. The expected results are unchanged; only the echoed queries differ. Backpatch-through: 17 (15 and 16 get a different fix) Reviewed-by: Jonathan Gonzalez V. <jonathan@abdiel.eu> Branch ------ REL_19_STABLE Details ------- https://git.postgresql.org/pg/commitdiff/073bd832772fa9ee2460d8420d2275687fbe88bc Modified Files -------------- src/test/modules/test_regex/expected/test_regex_utf8.out | 6 ++++-- src/test/modules/test_regex/sql/test_regex_utf8.sql | 6 ++++-- 2 files changed, 8 insertions(+), 4 deletions(-) ^ permalink raw reply [nested|flat] 4+ messages in thread
* pgsql: Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 @ 2026-08-26 20:25 Andrew Dunstan <andrew@dunslane.net> 0 siblings, 0 replies; 4+ messages in thread From: Andrew Dunstan @ 2026-08-26 20:25 UTC (permalink / raw) To: pgsql-committers@lists.postgresql.org Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 test_regex_utf8 decides whether to run by looking at the database encoding alone, but two of its cases, [[:graph:]] and [[:print:]] over E'xᔀሷ', depend on the ctype as well. In a database with encoding UTF8 and locale C they match just the x, because isgraph() and isprint() are false for anything outside ASCII, and the file fails. No buildfarm animal builds such a cluster, which is why this went unnoticed, and why the to_date() crash in 18.5 went undetected for want of exactly this coverage. A pending buildfarm client change will let an animal be configured that way. Fix by giving the two cases an explicit collation, so that they exercise a fixed Unicode ctype instead of whatever the database happened to be initialized with. test_regex() already passes its input collation down to the regex compiler. The expected results are unchanged; only the echoed queries differ. Backpatch-through: 17 (15 and 16 get a different fix) Reviewed-by: Jonathan Gonzalez V. <jonathan@abdiel.eu> Branch ------ REL_18_STABLE Details ------- https://git.postgresql.org/pg/commitdiff/8554d746964dd6ceaea827e7edbc45a1bc5948c9 Modified Files -------------- src/test/modules/test_regex/expected/test_regex_utf8.out | 6 ++++-- src/test/modules/test_regex/sql/test_regex_utf8.sql | 6 ++++-- 2 files changed, 8 insertions(+), 4 deletions(-) ^ permalink raw reply [nested|flat] 4+ messages in thread
* pgsql: Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 @ 2026-08-26 20:25 Andrew Dunstan <andrew@dunslane.net> 0 siblings, 0 replies; 4+ messages in thread From: Andrew Dunstan @ 2026-08-26 20:25 UTC (permalink / raw) To: pgsql-committers@lists.postgresql.org Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 test_regex_utf8 decides whether to run by looking at the database encoding alone, but two of its cases, [[:graph:]] and [[:print:]] over E'xᔀሷ', depend on the ctype as well. In a database with encoding UTF8 and locale C they match just the x, because isgraph() and isprint() are false for anything outside ASCII, and the file fails. No buildfarm animal builds such a cluster, which is why this went unnoticed, and why the to_date() crash in 18.5 went undetected for want of exactly this coverage. A pending buildfarm client change will let an animal be configured that way. Fix by giving the two cases an explicit collation, so that they exercise a fixed Unicode ctype instead of whatever the database happened to be initialized with. test_regex() already passes its input collation down to the regex compiler. The expected results are unchanged; only the echoed queries differ. Backpatch-through: 17 (15 and 16 get a different fix) Reviewed-by: Jonathan Gonzalez V. <jonathan@abdiel.eu> Branch ------ REL_17_STABLE Details ------- https://git.postgresql.org/pg/commitdiff/33b8a42dfba090cf140f8e5d01c4eddf2c08100c Modified Files -------------- src/test/modules/test_regex/expected/test_regex_utf8.out | 6 ++++-- src/test/modules/test_regex/sql/test_regex_utf8.sql | 6 ++++-- 2 files changed, 8 insertions(+), 4 deletions(-) ^ permalink raw reply [nested|flat] 4+ messages in thread
end of thread, other threads:[~2026-08-26 20:25 UTC | newest] Thread overview: 4+ messages (download: mbox mbox.gz follow: Atom feed) -- links below jump to the message on this page -- 2026-08-26 20:25 pgsql: Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 Andrew Dunstan <andrew@dunslane.net> 2026-08-26 20:25 pgsql: Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 Andrew Dunstan <andrew@dunslane.net> 2026-08-26 20:25 pgsql: Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 Andrew Dunstan <andrew@dunslane.net> 2026-08-26 20:25 pgsql: Pin two ctype-dependent test_regex_utf8 cases to pg_c_utf8 Andrew Dunstan <andrew@dunslane.net>
This inbox is served by DDX for PostgreSQL; see mirroring instructions for how to clone and mirror all data and code used for this inbox