pg.ddx.io  pgsql-committers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S
7+ messages / 1 participants
[nested] [flat]

* pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S
@ 2026-08-25 01:24 Richard Guo <rguo@postgresql.org>
  0 siblings, 0 replies; 7+ messages in thread

From: Richard Guo @ 2026-08-25 01:24 UTC (permalink / raw)
  To: pgsql-committers@lists.postgresql.org

Don't assume DISTINCT ON implies uniqueness when the tlist has SRFs

query_is_distinct_for() treated a subquery's DISTINCT ON clause as
proof that its output is unique over the DISTINCT ON columns, even if
the targetlist contains set-returning functions.  That's not true:
when the query has an ORDER BY, the planner postpones evaluation of
SRFs that are not DISTINCT ON or ORDER BY columns until after the
Unique step, so the subquery can produce duplicates of the DISTINCT ON
columns.  Relying on this bogus uniqueness proof allowed join removal
and unique-inner joins to produce wrong results.

Plain DISTINCT is not affected, since all tlist columns are DISTINCT
columns there, and so any SRFs get expanded before the Unique step.

To fix, make query_supports_distinctness() and query_is_distinct_for()
refuse to prove distinctness via DISTINCT ON if the targetlist
contains any SRFs.  This is more conservative than necessary, since
the SRFs are only postponed when there is an ORDER BY and none of them
appear in a sort/group column, but it doesn't seem worth the trouble
to check that precisely.

Author: Richard Guo <guofenglinux@gmail.com>
Reviewed-by: Tom Lane <tgl@sss.pgh.pa.us>
Discussion: https://postgr.es/m/CAMbWs4-hfd1Pyy_zBejsVUSy-3dx16rz2hgUakkKnAg3qg2q=Q@mail.gmail.com
Backpatch-through: 14

Branch
------
master

Details
-------
https://git.postgresql.org/pg/commitdiff/1bf74c806c41fcd382d5a0fab226e027f454a576

Modified Files
--------------
src/backend/optimizer/plan/analyzejoins.c | 16 ++++++++++-----
src/test/regress/expected/join.out        | 33 +++++++++++++++++++++++++++++++
src/test/regress/sql/join.sql             | 12 +++++++++++
3 files changed, 56 insertions(+), 5 deletions(-)



^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S
@ 2026-08-25 01:24 Richard Guo <rguo@postgresql.org>
  0 siblings, 0 replies; 7+ messages in thread

From: Richard Guo @ 2026-08-25 01:24 UTC (permalink / raw)
  To: pgsql-committers@lists.postgresql.org

Don't assume DISTINCT ON implies uniqueness when the tlist has SRFs

query_is_distinct_for() treated a subquery's DISTINCT ON clause as
proof that its output is unique over the DISTINCT ON columns, even if
the targetlist contains set-returning functions.  That's not true:
when the query has an ORDER BY, the planner postpones evaluation of
SRFs that are not DISTINCT ON or ORDER BY columns until after the
Unique step, so the subquery can produce duplicates of the DISTINCT ON
columns.  Relying on this bogus uniqueness proof allowed join removal
and unique-inner joins to produce wrong results.

Plain DISTINCT is not affected, since all tlist columns are DISTINCT
columns there, and so any SRFs get expanded before the Unique step.

To fix, make query_supports_distinctness() and query_is_distinct_for()
refuse to prove distinctness via DISTINCT ON if the targetlist
contains any SRFs.  This is more conservative than necessary, since
the SRFs are only postponed when there is an ORDER BY and none of them
appear in a sort/group column, but it doesn't seem worth the trouble
to check that precisely.

Author: Richard Guo <guofenglinux@gmail.com>
Reviewed-by: Tom Lane <tgl@sss.pgh.pa.us>
Discussion: https://postgr.es/m/CAMbWs4-hfd1Pyy_zBejsVUSy-3dx16rz2hgUakkKnAg3qg2q=Q@mail.gmail.com
Backpatch-through: 14

Branch
------
REL_19_STABLE

Details
-------
https://git.postgresql.org/pg/commitdiff/d446ca2c459c5541c257fbff05ec5a0bdbeb6a0c

Modified Files
--------------
src/backend/optimizer/plan/analyzejoins.c | 16 ++++++++++-----
src/test/regress/expected/join.out        | 33 +++++++++++++++++++++++++++++++
src/test/regress/sql/join.sql             | 12 +++++++++++
3 files changed, 56 insertions(+), 5 deletions(-)



^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S
@ 2026-08-25 01:24 Richard Guo <rguo@postgresql.org>
  0 siblings, 0 replies; 7+ messages in thread

From: Richard Guo @ 2026-08-25 01:24 UTC (permalink / raw)
  To: pgsql-committers@lists.postgresql.org

Don't assume DISTINCT ON implies uniqueness when the tlist has SRFs

query_is_distinct_for() treated a subquery's DISTINCT ON clause as
proof that its output is unique over the DISTINCT ON columns, even if
the targetlist contains set-returning functions.  That's not true:
when the query has an ORDER BY, the planner postpones evaluation of
SRFs that are not DISTINCT ON or ORDER BY columns until after the
Unique step, so the subquery can produce duplicates of the DISTINCT ON
columns.  Relying on this bogus uniqueness proof allowed join removal
and unique-inner joins to produce wrong results.

Plain DISTINCT is not affected, since all tlist columns are DISTINCT
columns there, and so any SRFs get expanded before the Unique step.

To fix, make query_supports_distinctness() and query_is_distinct_for()
refuse to prove distinctness via DISTINCT ON if the targetlist
contains any SRFs.  This is more conservative than necessary, since
the SRFs are only postponed when there is an ORDER BY and none of them
appear in a sort/group column, but it doesn't seem worth the trouble
to check that precisely.

Author: Richard Guo <guofenglinux@gmail.com>
Reviewed-by: Tom Lane <tgl@sss.pgh.pa.us>
Discussion: https://postgr.es/m/CAMbWs4-hfd1Pyy_zBejsVUSy-3dx16rz2hgUakkKnAg3qg2q=Q@mail.gmail.com
Backpatch-through: 14

Branch
------
REL_18_STABLE

Details
-------
https://git.postgresql.org/pg/commitdiff/4c66f172a09296b08d53526f802ddd2b461bd7e8

Modified Files
--------------
src/backend/optimizer/plan/analyzejoins.c | 16 ++++++++++-----
src/test/regress/expected/join.out        | 33 +++++++++++++++++++++++++++++++
src/test/regress/sql/join.sql             | 12 +++++++++++
3 files changed, 56 insertions(+), 5 deletions(-)



^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S
@ 2026-08-25 01:24 Richard Guo <rguo@postgresql.org>
  0 siblings, 0 replies; 7+ messages in thread

From: Richard Guo @ 2026-08-25 01:24 UTC (permalink / raw)
  To: pgsql-committers@lists.postgresql.org

Don't assume DISTINCT ON implies uniqueness when the tlist has SRFs

query_is_distinct_for() treated a subquery's DISTINCT ON clause as
proof that its output is unique over the DISTINCT ON columns, even if
the targetlist contains set-returning functions.  That's not true:
when the query has an ORDER BY, the planner postpones evaluation of
SRFs that are not DISTINCT ON or ORDER BY columns until after the
Unique step, so the subquery can produce duplicates of the DISTINCT ON
columns.  Relying on this bogus uniqueness proof allowed join removal
and unique-inner joins to produce wrong results.

Plain DISTINCT is not affected, since all tlist columns are DISTINCT
columns there, and so any SRFs get expanded before the Unique step.

To fix, make query_supports_distinctness() and query_is_distinct_for()
refuse to prove distinctness via DISTINCT ON if the targetlist
contains any SRFs.  This is more conservative than necessary, since
the SRFs are only postponed when there is an ORDER BY and none of them
appear in a sort/group column, but it doesn't seem worth the trouble
to check that precisely.

Author: Richard Guo <guofenglinux@gmail.com>
Reviewed-by: Tom Lane <tgl@sss.pgh.pa.us>
Discussion: https://postgr.es/m/CAMbWs4-hfd1Pyy_zBejsVUSy-3dx16rz2hgUakkKnAg3qg2q=Q@mail.gmail.com
Backpatch-through: 14

Branch
------
REL_17_STABLE

Details
-------
https://git.postgresql.org/pg/commitdiff/a4eb938b33557193d90c1b396afd2c274e28b07e

Modified Files
--------------
src/backend/optimizer/plan/analyzejoins.c | 16 ++++++++++-----
src/test/regress/expected/join.out        | 33 +++++++++++++++++++++++++++++++
src/test/regress/sql/join.sql             | 12 +++++++++++
3 files changed, 56 insertions(+), 5 deletions(-)



^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S
@ 2026-08-25 01:24 Richard Guo <rguo@postgresql.org>
  0 siblings, 0 replies; 7+ messages in thread

From: Richard Guo @ 2026-08-25 01:24 UTC (permalink / raw)
  To: pgsql-committers@lists.postgresql.org

Don't assume DISTINCT ON implies uniqueness when the tlist has SRFs

query_is_distinct_for() treated a subquery's DISTINCT ON clause as
proof that its output is unique over the DISTINCT ON columns, even if
the targetlist contains set-returning functions.  That's not true:
when the query has an ORDER BY, the planner postpones evaluation of
SRFs that are not DISTINCT ON or ORDER BY columns until after the
Unique step, so the subquery can produce duplicates of the DISTINCT ON
columns.  Relying on this bogus uniqueness proof allowed join removal
and unique-inner joins to produce wrong results.

Plain DISTINCT is not affected, since all tlist columns are DISTINCT
columns there, and so any SRFs get expanded before the Unique step.

To fix, make query_supports_distinctness() and query_is_distinct_for()
refuse to prove distinctness via DISTINCT ON if the targetlist
contains any SRFs.  This is more conservative than necessary, since
the SRFs are only postponed when there is an ORDER BY and none of them
appear in a sort/group column, but it doesn't seem worth the trouble
to check that precisely.

Author: Richard Guo <guofenglinux@gmail.com>
Reviewed-by: Tom Lane <tgl@sss.pgh.pa.us>
Discussion: https://postgr.es/m/CAMbWs4-hfd1Pyy_zBejsVUSy-3dx16rz2hgUakkKnAg3qg2q=Q@mail.gmail.com
Backpatch-through: 14

Branch
------
REL_16_STABLE

Details
-------
https://git.postgresql.org/pg/commitdiff/8fc5b57ae644a6ef1c8dcc28ce09dae583949ffc

Modified Files
--------------
src/backend/optimizer/plan/analyzejoins.c | 16 ++++++++++-----
src/test/regress/expected/join.out        | 33 +++++++++++++++++++++++++++++++
src/test/regress/sql/join.sql             | 12 +++++++++++
3 files changed, 56 insertions(+), 5 deletions(-)



^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S
@ 2026-08-25 01:24 Richard Guo <rguo@postgresql.org>
  0 siblings, 0 replies; 7+ messages in thread

From: Richard Guo @ 2026-08-25 01:24 UTC (permalink / raw)
  To: pgsql-committers@lists.postgresql.org

Don't assume DISTINCT ON implies uniqueness when the tlist has SRFs

query_is_distinct_for() treated a subquery's DISTINCT ON clause as
proof that its output is unique over the DISTINCT ON columns, even if
the targetlist contains set-returning functions.  That's not true:
when the query has an ORDER BY, the planner postpones evaluation of
SRFs that are not DISTINCT ON or ORDER BY columns until after the
Unique step, so the subquery can produce duplicates of the DISTINCT ON
columns.  Relying on this bogus uniqueness proof allowed join removal
and unique-inner joins to produce wrong results.

Plain DISTINCT is not affected, since all tlist columns are DISTINCT
columns there, and so any SRFs get expanded before the Unique step.

To fix, make query_supports_distinctness() and query_is_distinct_for()
refuse to prove distinctness via DISTINCT ON if the targetlist
contains any SRFs.  This is more conservative than necessary, since
the SRFs are only postponed when there is an ORDER BY and none of them
appear in a sort/group column, but it doesn't seem worth the trouble
to check that precisely.

Author: Richard Guo <guofenglinux@gmail.com>
Reviewed-by: Tom Lane <tgl@sss.pgh.pa.us>
Discussion: https://postgr.es/m/CAMbWs4-hfd1Pyy_zBejsVUSy-3dx16rz2hgUakkKnAg3qg2q=Q@mail.gmail.com
Backpatch-through: 14

Branch
------
REL_15_STABLE

Details
-------
https://git.postgresql.org/pg/commitdiff/201f7192e07763a57081a889dd2ba8c06d17c105

Modified Files
--------------
src/backend/optimizer/plan/analyzejoins.c | 16 ++++++++++-----
src/test/regress/expected/join.out        | 33 +++++++++++++++++++++++++++++++
src/test/regress/sql/join.sql             | 12 +++++++++++
3 files changed, 56 insertions(+), 5 deletions(-)



^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S
@ 2026-08-25 01:24 Richard Guo <rguo@postgresql.org>
  0 siblings, 0 replies; 7+ messages in thread

From: Richard Guo @ 2026-08-25 01:24 UTC (permalink / raw)
  To: pgsql-committers@lists.postgresql.org

Don't assume DISTINCT ON implies uniqueness when the tlist has SRFs

query_is_distinct_for() treated a subquery's DISTINCT ON clause as
proof that its output is unique over the DISTINCT ON columns, even if
the targetlist contains set-returning functions.  That's not true:
when the query has an ORDER BY, the planner postpones evaluation of
SRFs that are not DISTINCT ON or ORDER BY columns until after the
Unique step, so the subquery can produce duplicates of the DISTINCT ON
columns.  Relying on this bogus uniqueness proof allowed join removal
and unique-inner joins to produce wrong results.

Plain DISTINCT is not affected, since all tlist columns are DISTINCT
columns there, and so any SRFs get expanded before the Unique step.

To fix, make query_supports_distinctness() and query_is_distinct_for()
refuse to prove distinctness via DISTINCT ON if the targetlist
contains any SRFs.  This is more conservative than necessary, since
the SRFs are only postponed when there is an ORDER BY and none of them
appear in a sort/group column, but it doesn't seem worth the trouble
to check that precisely.

Author: Richard Guo <guofenglinux@gmail.com>
Reviewed-by: Tom Lane <tgl@sss.pgh.pa.us>
Discussion: https://postgr.es/m/CAMbWs4-hfd1Pyy_zBejsVUSy-3dx16rz2hgUakkKnAg3qg2q=Q@mail.gmail.com
Backpatch-through: 14

Branch
------
REL_14_STABLE

Details
-------
https://git.postgresql.org/pg/commitdiff/7e29fcd1ae0083a7fd45f14ab7b9ad3ea0c97e9c

Modified Files
--------------
src/backend/optimizer/plan/analyzejoins.c | 16 ++++++++++-----
src/test/regress/expected/join.out        | 33 +++++++++++++++++++++++++++++++
src/test/regress/sql/join.sql             | 12 +++++++++++
3 files changed, 56 insertions(+), 5 deletions(-)



^ permalink  raw  reply  [nested|flat] 7+ messages in thread


end of thread, other threads:[~2026-08-25 01:24 UTC | newest]

Thread overview: 7+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2026-08-25 01:24 pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S Richard Guo <rguo@postgresql.org>
2026-08-25 01:24 pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S Richard Guo <rguo@postgresql.org>
2026-08-25 01:24 pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S Richard Guo <rguo@postgresql.org>
2026-08-25 01:24 pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S Richard Guo <rguo@postgresql.org>
2026-08-25 01:24 pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S Richard Guo <rguo@postgresql.org>
2026-08-25 01:24 pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S Richard Guo <rguo@postgresql.org>
2026-08-25 01:24 pgsql: Don't assume DISTINCT ON implies uniqueness when the tlist has S Richard Guo <rguo@postgresql.org>

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox