pg.ddx.io pgsql-hackers@postgresql.org mailing list archive
help / color / mirror / Atom feedRe: multivariate statistics (v19)
70+ messages / 15 participants
[nested] [flat]
* Re: multivariate statistics (v19)
@ 2016-08-03 01:58 Tomas Vondra <tomas.vondra@2ndquadrant.com>
1 sibling, 3 replies; 70+ messages in thread
From: Tomas Vondra @ 2016-08-03 01:58 UTC (permalink / raw)
To: Tatsuo Ishii <ishii@postgresql.org>; +Cc: robertmhaas@gmail.com; david@pgmasters.net; tgl@sss.pgh.pa.us; alvherre@2ndquadrant.com; petr@2ndquadrant.com; jeff.janes@gmail.com; pgsql-hackers
Hi,
Attached is v19 of the "multivariate stats" patch series - essentially
v18 rebased on top of current master. Aside from a few bug fixes, the
main improvement is addition of SGML docs demonstrating the statistics
in a way similar to the current "Row Estimation Examples" (and the docs
are actually in the same section). I've tried to keep the right amount
of technical detail (and pointing to the right README for additional
details), but this may need improvements. I have not written docs
explaining how statistics may be combined yet (more about this later).
There are two general design questions that I'd like to get feedback on:
1) enriching the query tree with multivariate statistics info
Right now all the stuff related to multivariate statistics estimation
happens in clausesel.c - matching condition to statistics, selection of
statistics to use (if there are multiple usable stats), etc. So pretty
much all this info is internal to clausesel.c and does not get outside.
I'm starting to think that some of the steps (matching quals to stats,
selection of stats) should happen in a "preprocess" step before the
actual estimation, storing the information (which stats to use, etc.) in
a new type of node in the query tree - something like RestrictInfo.
I believe this needs to happen sometime after deconstruct_jointree() as
that builds RestrictInfos nodes, and looking at planmain.c, right after
extract_restriction_or_clauses seems about right. Haven't tried, though.
This would move all the "statistics selection" logic from clausesel.c,
separating it from the "actual estimation" and simplifying the code.
But more importantly, I think we'll need to show some of the data in
EXPLAIN output. With per-column statistics it's fairly straightforward
to determine which statistics are used and how. But with multivariate
stats things are often more complicated - there may be multiple
candidate statistics (e.g. histograms covering different subsets of the
conditions), it's possible to apply them in different orders, etc.
But EXPLAIN can't show the info if it's ephemeral and available only
within clausesel.c (and thrown away after the estimation).
2) combining multiple statistics
I think the ability to combine multivariate statistics (covering
different subsets of conditions) is important and useful, but I'm
starting to think that the current implementation may not be the correct
one (which is why I haven't written the SGML docs about this part of the
patch series yet).
Assume there's a table "t" with 3 columns (a, b, c), and that we're
estimating query:
SELECT * FROM t WHERE a = 1 AND b = 2 AND c = 3
but that we only have two statistics (a,b) and (b,c). The current patch
does about this:
P(a=1,b=2,c=3) = P(a=1,b=2) * P(c=3|b=2)
i.e. it estimates the first two conditions using (a,b), and then
estimates (c=3) using (b,c) with "b=2" as a condition. Now, this is very
efficient, but it only works as long as the query contains conditions
"connecting" the two statistics. So if we remove the "b=2" condition
from the query, this stops working.
But it's possible to do this differently, e.g. by doing this:
P(a=1) * P(c=3|a=1)
where P(c=3|a=1) is using (b,c), but uses (a,b) to restrict the set of
buckets (if the statistics is a histogram) to consider. In pseudo-code,
it might look like this:
buckets = {}
foreach bucket x in (b,c):
foreach bucket y in (a,b):
if y matches (a=1) and overlap(x,y):
buckets := buckets + x
which is the part of (b,c) matching (a=1), allowing us to compute the
conditional probability.
It may get more complicated, of course. In particular, there may be
different types of statistics, and we need to be able to "match" them
against each other. With just MCV lists and histograms that's probably
easy enough, but if we add other types of statistics, it may get way
more complicated.
I still think this is a useful capability, but perhaps there are better
ideas how to do that. In any case, it only affects the last part of the
patch (0006).
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
Attachments:
[application/x-compressed-tar] multivariate-stats-v19.tgz (150.8K, ../../0924a7a9-a219-1366-d3bc-2ddd98bc9269@2ndquadrant.com/2-multivariate-stats-v19.tgz)
download
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-05 04:24 Michael Paquier <michael.paquier@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
2 siblings, 1 reply; 70+ messages in thread
From: Michael Paquier @ 2016-08-05 04:24 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; pgsql-hackers
On Wed, Aug 3, 2016 at 10:58 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> Attached is v19 of the "multivariate stats" patch series - essentially v18
> rebased on top of current master. Aside from a few bug fixes, the main
> improvement is addition of SGML docs demonstrating the statistics in a way
> similar to the current "Row Estimation Examples" (and the docs are actually
> in the same section). I've tried to keep the right amount of technical
> detail (and pointing to the right README for additional details), but this
> may need improvements. I have not written docs explaining how statistics may
> be combined yet (more about this later).
What we have here is quite something:
$ git diff master --stat | tail -n1
77 files changed, 12809 insertions(+), 65 deletions(-)
I will try to get familiar on the topic and added myself as a reviewer
of this patch. Hopefully I'll get feedback soon.
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-05 17:38 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Michael Paquier <michael.paquier@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Tomas Vondra @ 2016-08-05 17:38 UTC (permalink / raw)
To: Michael Paquier <michael.paquier@gmail.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; pgsql-hackers
On 08/05/2016 06:24 AM, Michael Paquier wrote:
> On Wed, Aug 3, 2016 at 10:58 AM, Tomas Vondra
> <tomas.vondra@2ndquadrant.com> wrote:
>> Attached is v19 of the "multivariate stats" patch series - essentially v18
>> rebased on top of current master. Aside from a few bug fixes, the main
>> improvement is addition of SGML docs demonstrating the statistics in a way
>> similar to the current "Row Estimation Examples" (and the docs are actually
>> in the same section). I've tried to keep the right amount of technical
>> detail (and pointing to the right README for additional details), but this
>> may need improvements. I have not written docs explaining how statistics may
>> be combined yet (more about this later).
>
> What we have here is quite something:
> $ git diff master --stat | tail -n1
> 77 files changed, 12809 insertions(+), 65 deletions(-)
> I will try to get familiar on the topic and added myself as a reviewer
> of this patch. Hopefully I'll get feedback soon.
Yes, it's a large patch. Although 25% of the insertions are SGML docs,
regression tests and READMEs, and large part of the remaining ~9k
insertions are comments. But it may still be overwhelming, no doubt
about that.
FWIW, if someone is interested in the patch but is unsure where to
start, I'm ready to help with that as much as possible. For example if
you happen to go to PostgresOpen, feel free to drag me to a corner and
ask me as many questions as you want ...
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-05 22:21 Michael Paquier <michael.paquier@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
0 siblings, 0 replies; 70+ messages in thread
From: Michael Paquier @ 2016-08-05 22:21 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; pgsql-hackers
On Sat, Aug 6, 2016 at 2:38 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> On 08/05/2016 06:24 AM, Michael Paquier wrote:
>>
>> On Wed, Aug 3, 2016 at 10:58 AM, Tomas Vondra
>> <tomas.vondra@2ndquadrant.com> wrote:
>>>
>>> Attached is v19 of the "multivariate stats" patch series - essentially
>>> v18
>>> rebased on top of current master. Aside from a few bug fixes, the main
>>> improvement is addition of SGML docs demonstrating the statistics in a
>>> way
>>> similar to the current "Row Estimation Examples" (and the docs are
>>> actually
>>> in the same section). I've tried to keep the right amount of technical
>>> detail (and pointing to the right README for additional details), but
>>> this
>>> may need improvements. I have not written docs explaining how statistics
>>> may
>>> be combined yet (more about this later).
>>
>>
>> What we have here is quite something:
>> $ git diff master --stat | tail -n1
>> 77 files changed, 12809 insertions(+), 65 deletions(-)
>> I will try to get familiar on the topic and added myself as a reviewer
>> of this patch. Hopefully I'll get feedback soon.
>
>
> Yes, it's a large patch. Although 25% of the insertions are SGML docs,
> regression tests and READMEs, and large part of the remaining ~9k insertions
> are comments. But it may still be overwhelming, no doubt about that.
>
> FWIW, if someone is interested in the patch but is unsure where to start,
> I'm ready to help with that as much as possible. For example if you happen
> to go to PostgresOpen, feel free to drag me to a corner and ask me as many
> questions as you want ...
Sure. Only PGconf SV is on my track this year.
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-10 04:41 Michael Paquier <michael.paquier@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
2 siblings, 2 replies; 70+ messages in thread
From: Michael Paquier @ 2016-08-10 04:41 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; pgsql-hackers
On Wed, Aug 3, 2016 at 10:58 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> 1) enriching the query tree with multivariate statistics info
>
> Right now all the stuff related to multivariate statistics estimation
> happens in clausesel.c - matching condition to statistics, selection of
> statistics to use (if there are multiple usable stats), etc. So pretty much
> all this info is internal to clausesel.c and does not get outside.
This does not seem bad to me as first sight but...
> I'm starting to think that some of the steps (matching quals to stats,
> selection of stats) should happen in a "preprocess" step before the actual
> estimation, storing the information (which stats to use, etc.) in a new type
> of node in the query tree - something like RestrictInfo.
>
> I believe this needs to happen sometime after deconstruct_jointree() as that
> builds RestrictInfos nodes, and looking at planmain.c, right after
> extract_restriction_or_clauses seems about right. Haven't tried, though.
>
> This would move all the "statistics selection" logic from clausesel.c,
> separating it from the "actual estimation" and simplifying the code.
>
> But more importantly, I think we'll need to show some of the data in EXPLAIN
> output. With per-column statistics it's fairly straightforward to determine
> which statistics are used and how. But with multivariate stats things are
> often more complicated - there may be multiple candidate statistics (e.g.
> histograms covering different subsets of the conditions), it's possible to
> apply them in different orders, etc.
>
> But EXPLAIN can't show the info if it's ephemeral and available only within
> clausesel.c (and thrown away after the estimation).
This gives a good reason to not do that in clauserel.c, it would be
really cool to be able to get some information regarding the stats
used with a simple EXPLAIN.
> 2) combining multiple statistics
>
> I think the ability to combine multivariate statistics (covering different
> subsets of conditions) is important and useful, but I'm starting to think
> that the current implementation may not be the correct one (which is why I
> haven't written the SGML docs about this part of the patch series yet).
>
> Assume there's a table "t" with 3 columns (a, b, c), and that we're
> estimating query:
>
> SELECT * FROM t WHERE a = 1 AND b = 2 AND c = 3
>
> but that we only have two statistics (a,b) and (b,c). The current patch does
> about this:
>
> P(a=1,b=2,c=3) = P(a=1,b=2) * P(c=3|b=2)
>
> i.e. it estimates the first two conditions using (a,b), and then estimates
> (c=3) using (b,c) with "b=2" as a condition. Now, this is very efficient,
> but it only works as long as the query contains conditions "connecting" the
> two statistics. So if we remove the "b=2" condition from the query, this
> stops working.
This is trying to make the algorithm smarter than the user, which is
something I'd think we could live without. In this case statistics on
(a,c) or (a,b,c) are missing. And what if the user does not want to
make use of stats for (a,c) because he only defined (a,b) and (b,c)?
Patch 0001: there have been comments about that before, and you have
put the checks on RestrictInfo in a couple of variables of
pull_varnos_walker, so nothing to say from here.
Patch 0002:
+ <para>
+ <command>CREATE STATISTICS</command> will create a new multivariate
+ statistics on the table. The statistics will be created in the in the
+ current database. The statistics will be owned by the user issuing
+ the command.
+ </para>
s/in the/in the/.
+ <para>
+ Create table <structname>t1</> with two functionally dependent columns, i.e.
+ knowledge of a value in the first column is sufficient for detemining the
+ value in the other column. Then functional dependencies are built on those
+ columns:
s/detemining/determining/
+ <para>
+ If a schema name is given (for example, <literal>CREATE STATISTICS
+ myschema.mystat ...</>) then the statistics is created in the specified
+ schema. Otherwise it is created in the current schema. The name of
+ the table must be distinct from the name of any other statistics in the
+ same schema.
+ </para>
I would just assume that a statistics is located on the schema of the
relation it depends on. So the thing that may be better to do is just:
- Register the OID of the table a statistics depends on but not the schema.
- Give up on those query extensions related to the schema.
- Allow the same statistics name to be used for multiple tables.
- Just fail if a statistics name is being reused on the table again.
It may be better to complain about that even if the column list is
different.
- Register the dependency between the statistics and the table.
+ALTER STATISTICS <replaceable class="parameter">name</replaceable>
OWNER TO { <replaceable class="PARAMETER">new_owner</replaceable> |
CURRENT_USER | SESSION_USER }
On the same line, is OWNER TO really necessary? I could have assumed
that if a user is able to query the set of columns related to a
statistics, he should have access to it.
=# create statistics aa_a_b3 on aam (a, b) with (dependencies);
ERROR: 23505: duplicate key value violates unique constraint
"pg_mv_statistic_name_index"
DETAIL: Key (staname, stanamespace)=(aa_a_b3, 2200) already exists.
SCHEMA NAME: pg_catalog
TABLE NAME: pg_mv_statistic
CONSTRAINT NAME: pg_mv_statistic_name_index
LOCATION: _bt_check_unique, nbtinsert.c:433
When creating a multivariate function with a name that already exists,
this error message should be more friendly.
=# create table aa (a int, b int);
CREATE TABLE
=# create view aav as select * from aa;
CREATE VIEW
=# create statistics aab_v on aav (a, b) with (dependencies);
CREATE STATISTICS
Why do views and foreign tables support this command? This code also
mentions that this case is not actually supported:
+ /* multivariate stats are supported on tables and matviews */
+ if (rel->rd_rel->relkind == RELKIND_RELATION ||
+ rel->rd_rel->relkind == RELKIND_MATVIEW)
+ tupdesc = RelationGetDescr(rel);
};
+
/*
Spurious noise in the patch.
+ /* check that at least some statistics were requested */
+ if (!build_dependencies)
+ ereport(ERROR,
+ (errcode(ERRCODE_SYNTAX_ERROR),
+ errmsg("no statistics type (dependencies) was requested")));
So, WITH (dependencies) is mandatory in any case. Why not just
dropping it from the first cut then?
pg_mv_stats shows only the attribute numbers of the columns it has
stats on, I think that those should be the column names. [...after a
while...], as it is mentioned here:
+ * TODO Would be nice if this printed column names (instead of just attnums).
Does this work properly with DDL deparsing? If yes, could it be
possible to add tests in test_ddl_deparse? This is a new object type,
so those look necessary I think.
Statistics definition reorder the columns by itself depending on their
order. For example:
create table aa (a int, b int);
create statistics aas on aa(b, a) with (dependencies);
\d aa
"public.aas" (dependencies) ON (a, b)
As this defines a correlation between multiple columns, isn't it wrong
to assume that (b, a) and (a, b) are always the same correlation? I
don't recall such properties as being always commutative (old
memories, I suck at stats in general). [...reading README...] So this
is caused by the implementation limitations that only limit the
analysis between interactions of two columns. Still it seems incorrect
to reorder the user-visible portion.
The comment on top of get_relation_info needs to be updated to mention
that mvstatlist gets fetched as well.
+ while (HeapTupleIsValid(htup = systable_getnext(indscan)))
+ /* TODO maybe include only already built statistics? */
+ result = insert_ordered_oid(result, HeapTupleGetOid(htup));
I haven't looked at the rest yet of the series yet, but I'd think that
including the ones not built may be a good idea to let caller do
itself more filtering. Of course this depends on the next series...
+typedef struct MVDependencyData
+{
+ int nattributes; /* number of attributes */
+ int16 attributes[1]; /* attribute numbers */
+} MVDependencyData;
You need to look for FLEXIBLE_ARRAY_MEMBER here. Same for MVDependenciesData.
+++ b/src/test/regress/serial_schedule
@@ -167,3 +167,4 @@ test: with
test: xml
test: event_trigger
test: stats
+test: mv_dependencies
This test is not listed in parallel_schedule.
s/Apllying/Applying/
There is a lot of mumbo-jumbo regarding the way dependencies are
stored with mainly serialize_mv_dependencies and
deserialize_mv_dependencies that operates them from bytea/dep trees.
That's not cool and not portable because pg_mv_statistic represents
that as pure bytea. I would suggest creating a generic data type that
does those operations, named like pg_dependency_tree and then use that
in those new catalogs. pg_node_tree is a precedent of such a thing.
New features could as well make use of this new data type of we are
able to design that in a way generic enough, so that would be a base
patch that the current 0002 applies on top of.
Regarding psql:
- The new commands lack psql completion, that would ease the use of
the new commands.
- Would it make sense to have a backslash command to show the list of
statistics?
Congratulations. I just looked at 25% of the overall patch and my mind
is already blown away, but I am catching up with the rest...
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-10 11:33 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Michael Paquier <michael.paquier@gmail.com>
1 sibling, 2 replies; 70+ messages in thread
From: Tomas Vondra @ 2016-08-10 11:33 UTC (permalink / raw)
To: Michael Paquier <michael.paquier@gmail.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; pgsql-hackers
On 08/10/2016 06:41 AM, Michael Paquier wrote:
> On Wed, Aug 3, 2016 at 10:58 AM, Tomas Vondra
> <tomas.vondra@2ndquadrant.com> wrote:
...
>> But more importantly, I think we'll need to show some of the data in EXPLAIN
>> output. With per-column statistics it's fairly straightforward to determine
>> which statistics are used and how. But with multivariate stats things are
>> often more complicated - there may be multiple candidate statistics (e.g.
>> histograms covering different subsets of the conditions), it's possible to
>> apply them in different orders, etc.
>>
>> But EXPLAIN can't show the info if it's ephemeral and available only within
>> clausesel.c (and thrown away after the estimation).
>
> This gives a good reason to not do that in clauserel.c, it would be
> really cool to be able to get some information regarding the stats
> used with a simple EXPLAIN.
>
I think there are two separate questions:
(a) Whether the query plan is "enriched" with information about
statistics, or whether this information is ephemeral and available only
in clausesel.c.
(b) Where exactly this enrichment happens.
Theoretically we might enrich the query plan (add nodes with info about
the statistics), so that EXPLAIN gets the info, and it might still
happen in clausesel.c.
>> 2) combining multiple statistics
>>
>> I think the ability to combine multivariate statistics (covering different
>> subsets of conditions) is important and useful, but I'm starting to think
>> that the current implementation may not be the correct one (which is why I
>> haven't written the SGML docs about this part of the patch series yet).
>>
>> Assume there's a table "t" with 3 columns (a, b, c), and that we're
>> estimating query:
>>
>> SELECT * FROM t WHERE a = 1 AND b = 2 AND c = 3
>>
>> but that we only have two statistics (a,b) and (b,c). The current patch does
>> about this:
>>
>> P(a=1,b=2,c=3) = P(a=1,b=2) * P(c=3|b=2)
>>
>> i.e. it estimates the first two conditions using (a,b), and then estimates
>> (c=3) using (b,c) with "b=2" as a condition. Now, this is very efficient,
>> but it only works as long as the query contains conditions "connecting" the
>> two statistics. So if we remove the "b=2" condition from the query, this
>> stops working.
>
> This is trying to make the algorithm smarter than the user, which is
> something I'd think we could live without. In this case statistics on
> (a,c) or (a,b,c) are missing. And what if the user does not want to
> make use of stats for (a,c) because he only defined (a,b) and (b,c)?
>
I don't think so. Obviously, if you have statistics covering all the
conditions - great, we can't really do better than that.
But there's a crucial relation between the number of dimensions of the
statistics and accuracy of the statistics. Let's say you have statistics
on 8 columns, and you split each dimension twice to build a histogram -
that's 256 buckets right there, and we only get ~50% selectivity in each
dimension (the actual histogram building algorithm is more complex, but
you get the idea).
I see this as probably the most interesting part of the patch, and quite
useful. But we'll definitely get the single-statistics estimate first,
no doubt about that.
> Patch 0001: there have been comments about that before, and you have
> put the checks on RestrictInfo in a couple of variables of
> pull_varnos_walker, so nothing to say from here.
>
I don't follow. Are you suggesting 0001 is a reasonable fix, or that
there's a proposed solution?
> Patch 0002:
> + <para>
> + <command>CREATE STATISTICS</command> will create a new multivariate
> + statistics on the table. The statistics will be created in the in the
> + current database. The statistics will be owned by the user issuing
> + the command.
> + </para>
> s/in the/in the/.
>
> + <para>
> + Create table <structname>t1</> with two functionally dependent columns, i.e.
> + knowledge of a value in the first column is sufficient for detemining the
> + value in the other column. Then functional dependencies are built on those
> + columns:
> s/detemining/determining/
>
> + <para>
> + If a schema name is given (for example, <literal>CREATE STATISTICS
> + myschema.mystat ...</>) then the statistics is created in the specified
> + schema. Otherwise it is created in the current schema. The name of
> + the table must be distinct from the name of any other statistics in the
> + same schema.
> + </para>
> I would just assume that a statistics is located on the schema of the
> relation it depends on. So the thing that may be better to do is just:
> - Register the OID of the table a statistics depends on but not the schema.
> - Give up on those query extensions related to the schema.
> - Allow the same statistics name to be used for multiple tables.
> - Just fail if a statistics name is being reused on the table again.
> It may be better to complain about that even if the column list is
> different.
> - Register the dependency between the statistics and the table.
The idea is that the syntax should work even for statistics built on
multiple tables, e.g. to provide better statistics for joins. That's why
the schema may be specified (as each table might be in different
schema), and so on.
>
> +ALTER STATISTICS <replaceable class="parameter">name</replaceable>
> OWNER TO { <replaceable class="PARAMETER">new_owner</replaceable> |
> CURRENT_USER | SESSION_USER }
> On the same line, is OWNER TO really necessary? I could have assumed
> that if a user is able to query the set of columns related to a
> statistics, he should have access to it.
>
Not sure, TBH. I think I've reused ALTER INDEX syntax, but now I see
it's actually ignored with a warning.
> =# create statistics aa_a_b3 on aam (a, b) with (dependencies);
> ERROR: 23505: duplicate key value violates unique constraint
> "pg_mv_statistic_name_index"
> DETAIL: Key (staname, stanamespace)=(aa_a_b3, 2200) already exists.
> SCHEMA NAME: pg_catalog
> TABLE NAME: pg_mv_statistic
> CONSTRAINT NAME: pg_mv_statistic_name_index
> LOCATION: _bt_check_unique, nbtinsert.c:433
> When creating a multivariate function with a name that already exists,
> this error message should be more friendly.
Yes, agreed.
>
> =# create table aa (a int, b int);
> CREATE TABLE
> =# create view aav as select * from aa;
> CREATE VIEW
> =# create statistics aab_v on aav (a, b) with (dependencies);
> CREATE STATISTICS
> Why do views and foreign tables support this command? This code also
> mentions that this case is not actually supported:
> + /* multivariate stats are supported on tables and matviews */
> + if (rel->rd_rel->relkind == RELKIND_RELATION ||
> + rel->rd_rel->relkind == RELKIND_MATVIEW)
> + tupdesc = RelationGetDescr(rel);
>
> };
Yes, seems like a bug.
>
> +
> /*
> Spurious noise in the patch.
>
> + /* check that at least some statistics were requested */
> + if (!build_dependencies)
> + ereport(ERROR,
> + (errcode(ERRCODE_SYNTAX_ERROR),
> + errmsg("no statistics type (dependencies) was requested")));
> So, WITH (dependencies) is mandatory in any case. Why not just
> dropping it from the first cut then?
Because the follow-up patches extend this to require at least one
statistics type. So in 0004 it becomes
if (!(build_dependencies || build_mcv))
and in 0005 it's
if (!(build_dependencies || build_mcv || build_histogram))
We might drop it from 0002 (and assume build_dependencies=true), and
then add the check in 0004. But it seems a bit pointless.
>
> pg_mv_stats shows only the attribute numbers of the columns it has
> stats on, I think that those should be the column names. [...after a
> while...], as it is mentioned here:
> + * TODO Would be nice if this printed column names (instead of just attnums).
Yeah.
>
> Does this work properly with DDL deparsing? If yes, could it be
> possible to add tests in test_ddl_deparse? This is a new object type,
> so those look necessary I think.
>
I haven't done anything with DDL deparsing, so I think the answer is
"no" and needs to be added to a TODO.
> Statistics definition reorder the columns by itself depending on their
> order. For example:
> create table aa (a int, b int);
> create statistics aas on aa(b, a) with (dependencies);
> \d aa
> "public.aas" (dependencies) ON (a, b)
> As this defines a correlation between multiple columns, isn't it wrong
> to assume that (b, a) and (a, b) are always the same correlation? I
> don't recall such properties as being always commutative (old
> memories, I suck at stats in general). [...reading README...] So this
> is caused by the implementation limitations that only limit the
> analysis between interactions of two columns. Still it seems incorrect
> to reorder the user-visible portion.
I don't follow. If you talk about Pearson's correlation, that clearly
does not depend on the order of columns - it's perfectly independent of
that. If you talk about about correlation in the wider sense (i.e.
arbitrary dependence between columns), that might depend - but I don't
remember a single piece of the patch where this might be a problem.
Also, which README states that we can only analyze interactions between
two columns? That's pretty clearly not the case - the patch should
handle dependencies between more columns without any problems.
>
> The comment on top of get_relation_info needs to be updated to mention
> that mvstatlist gets fetched as well.
>
> + while (HeapTupleIsValid(htup = systable_getnext(indscan)))
> + /* TODO maybe include only already built statistics? */
> + result = insert_ordered_oid(result, HeapTupleGetOid(htup));
> I haven't looked at the rest yet of the series yet, but I'd think that
> including the ones not built may be a good idea to let caller do
> itself more filtering. Of course this depends on the next series...
>
Probably, although the more I'm thinking about this the more I think
I'll rework this along the lines of the foreign-key-estimation patch,
i.e. preprocessing called from planmain.c (adding info to the query
plan), estimation in clausesel.c etc. Which also affects this bit,
because the foreign keys are also loaded elsewhere, IIRC.
> +typedef struct MVDependencyData
> +{
> + int nattributes; /* number of attributes */
> + int16 attributes[1]; /* attribute numbers */
> +} MVDependencyData;
> You need to look for FLEXIBLE_ARRAY_MEMBER here. Same for MVDependenciesData.
>
> +++ b/src/test/regress/serial_schedule
> @@ -167,3 +167,4 @@ test: with
> test: xml
> test: event_trigger
> test: stats
> +test: mv_dependencies
> This test is not listed in parallel_schedule.
>
> s/Apllying/Applying/
>
> There is a lot of mumbo-jumbo regarding the way dependencies are
> stored with mainly serialize_mv_dependencies and
> deserialize_mv_dependencies that operates them from bytea/dep trees.
> That's not cool and not portable because pg_mv_statistic represents
> that as pure bytea. I would suggest creating a generic data type that
> does those operations, named like pg_dependency_tree and then use that
> in those new catalogs. pg_node_tree is a precedent of such a thing.
> New features could as well make use of this new data type of we are
> able to design that in a way generic enough, so that would be a base
> patch that the current 0002 applies on top of.
Interesting idea, haven't thought about that. So are you suggesting to
add a data type for each statistics type (dependencies, MCV, histogram,
...)?
>
> Regarding psql:
> - The new commands lack psql completion, that would ease the use of
> the new commands.
> - Would it make sense to have a backslash command to show the list of
> statistics?
>
Yeah, that's on the TODO.
> Congratulations. I just looked at 25% of the overall patch and my mind
> is already blown away, but I am catching up with the rest...
>
Thanks for looking.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-10 11:50 Petr Jelinek <petr@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
1 sibling, 1 reply; 70+ messages in thread
From: Petr Jelinek @ 2016-08-10 11:50 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Michael Paquier <michael.paquier@gmail.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; pgsql-hackers
On 10/08/16 13:33, Tomas Vondra wrote:
> On 08/10/2016 06:41 AM, Michael Paquier wrote:
>> On Wed, Aug 3, 2016 at 10:58 AM, Tomas Vondra
>>> 2) combining multiple statistics
>>>
>>> I think the ability to combine multivariate statistics (covering
>>> different
>>> subsets of conditions) is important and useful, but I'm starting to
>>> think
>>> that the current implementation may not be the correct one (which is
>>> why I
>>> haven't written the SGML docs about this part of the patch series yet).
>>>
>>> Assume there's a table "t" with 3 columns (a, b, c), and that we're
>>> estimating query:
>>>
>>> SELECT * FROM t WHERE a = 1 AND b = 2 AND c = 3
>>>
>>> but that we only have two statistics (a,b) and (b,c). The current
>>> patch does
>>> about this:
>>>
>>> P(a=1,b=2,c=3) = P(a=1,b=2) * P(c=3|b=2)
>>>
>>> i.e. it estimates the first two conditions using (a,b), and then
>>> estimates
>>> (c=3) using (b,c) with "b=2" as a condition. Now, this is very
>>> efficient,
>>> but it only works as long as the query contains conditions
>>> "connecting" the
>>> two statistics. So if we remove the "b=2" condition from the query, this
>>> stops working.
>>
>> This is trying to make the algorithm smarter than the user, which is
>> something I'd think we could live without. In this case statistics on
>> (a,c) or (a,b,c) are missing. And what if the user does not want to
>> make use of stats for (a,c) because he only defined (a,b) and (b,c)?
>>
>
> I don't think so. Obviously, if you have statistics covering all the
> conditions - great, we can't really do better than that.
>
> But there's a crucial relation between the number of dimensions of the
> statistics and accuracy of the statistics. Let's say you have statistics
> on 8 columns, and you split each dimension twice to build a histogram -
> that's 256 buckets right there, and we only get ~50% selectivity in each
> dimension (the actual histogram building algorithm is more complex, but
> you get the idea).
>
I think it makes sense to pursue this, but I also think we can easily
live with not having it in the first version that gets committed and
doing it as follow-up patch.
--
Petr Jelinek http://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-10 12:23 Michael Paquier <michael.paquier@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
1 sibling, 1 reply; 70+ messages in thread
From: Michael Paquier @ 2016-08-10 12:23 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; pgsql-hackers
On Wed, Aug 10, 2016 at 8:33 PM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> On 08/10/2016 06:41 AM, Michael Paquier wrote:
>> Patch 0001: there have been comments about that before, and you have
>> put the checks on RestrictInfo in a couple of variables of
>> pull_varnos_walker, so nothing to say from here.
>>
>
> I don't follow. Are you suggesting 0001 is a reasonable fix, or that there's
> a proposed solution?
I think that's reasonable.
>> Patch 0002:
>> + <para>
>> + <command>CREATE STATISTICS</command> will create a new multivariate
>> + statistics on the table. The statistics will be created in the in the
>> + current database. The statistics will be owned by the user issuing
>> + the command.
>> + </para>
>> s/in the/in the/.
>>
>> + <para>
>> + Create table <structname>t1</> with two functionally dependent
>> columns, i.e.
>> + knowledge of a value in the first column is sufficient for detemining
>> the
>> + value in the other column. Then functional dependencies are built on
>> those
>> + columns:
>> s/detemining/determining/
>>
>> + <para>
>> + If a schema name is given (for example, <literal>CREATE STATISTICS
>> + myschema.mystat ...</>) then the statistics is created in the
>> specified
>> + schema. Otherwise it is created in the current schema. The name of
>> + the table must be distinct from the name of any other statistics in
>> the
>> + same schema.
>> + </para>
>> I would just assume that a statistics is located on the schema of the
>> relation it depends on. So the thing that may be better to do is just:
>> - Register the OID of the table a statistics depends on but not the
>> schema.
>> - Give up on those query extensions related to the schema.
>> - Allow the same statistics name to be used for multiple tables.
>> - Just fail if a statistics name is being reused on the table again.
>> It may be better to complain about that even if the column list is
>> different.
>> - Register the dependency between the statistics and the table.
>
> The idea is that the syntax should work even for statistics built on
> multiple tables, e.g. to provide better statistics for joins. That's why the
> schema may be specified (as each table might be in different schema), and so
> on.
So you mean that the same statistics could be shared between tables?
But as this is visibly not a concept introduced yet in this set of
patches, why not just cut it off for now to simplify the whole? If
there is no schema-related field in pg_mv_statistics we could still
add it later if it proves to be useful.
>> +
>> /*
>> Spurious noise in the patch.
>>
>> + /* check that at least some statistics were requested */
>> + if (!build_dependencies)
>> + ereport(ERROR,
>> + (errcode(ERRCODE_SYNTAX_ERROR),
>> + errmsg("no statistics type (dependencies) was
>> requested")));
>> So, WITH (dependencies) is mandatory in any case. Why not just
>> dropping it from the first cut then?
>
>
> Because the follow-up patches extend this to require at least one statistics
> type. So in 0004 it becomes
>
> if (!(build_dependencies || build_mcv))
>
> and in 0005 it's
>
> if (!(build_dependencies || build_mcv || build_histogram))
>
> We might drop it from 0002 (and assume build_dependencies=true), and then
> add the check in 0004. But it seems a bit pointless.
This is a complicated set of patches. I'd think that we should try to
simplify things as much as possible first, and the WITH clause is not
mandatory to have as of 0002.
>> Statistics definition reorder the columns by itself depending on their
>> order. For example:
>> create table aa (a int, b int);
>> create statistics aas on aa(b, a) with (dependencies);
>> \d aa
>> "public.aas" (dependencies) ON (a, b)
>> As this defines a correlation between multiple columns, isn't it wrong
>> to assume that (b, a) and (a, b) are always the same correlation? I
>> don't recall such properties as being always commutative (old
>> memories, I suck at stats in general). [...reading README...] So this
>> is caused by the implementation limitations that only limit the
>> analysis between interactions of two columns. Still it seems incorrect
>> to reorder the user-visible portion.
>
> I don't follow. If you talk about Pearson's correlation, that clearly does
> not depend on the order of columns - it's perfectly independent of that. If
> you talk about about correlation in the wider sense (i.e. arbitrary
> dependence between columns), that might depend - but I don't remember a
> single piece of the patch where this might be a problem.
Yes, based on what is done in the patch that may not be a problem, but
I am wondering if this is not restricting things too much.
> Also, which README states that we can only analyze interactions between two
> columns? That's pretty clearly not the case - the patch should handle
> dependencies between more columns without any problems.
I have noticed that the patch evaluates all the set of permutations
possible using a column list, it seems to me though that say if we
have three columns (a,b,c) listed in a statistics, (a,b) => c and
(b,a) => c are two different things.
>> There is a lot of mumbo-jumbo regarding the way dependencies are
>> stored with mainly serialize_mv_dependencies and
>> deserialize_mv_dependencies that operates them from bytea/dep trees.
>> That's not cool and not portable because pg_mv_statistic represents
>> that as pure bytea. I would suggest creating a generic data type that
>> does those operations, named like pg_dependency_tree and then use that
>> in those new catalogs. pg_node_tree is a precedent of such a thing.
>> New features could as well make use of this new data type of we are
>> able to design that in a way generic enough, so that would be a base
>> patch that the current 0002 applies on top of.
>
>
> Interesting idea, haven't thought about that. So are you suggesting to add a
> data type for each statistics type (dependencies, MCV, histogram, ...)?
Yes that would be something like that, it would be actually perhaps
better to have one single data type, and be able to switch between
each model easily instead of putting byteas in the catalog.
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-10 12:24 Michael Paquier <michael.paquier@gmail.com>
parent: Petr Jelinek <petr@2ndquadrant.com>
0 siblings, 1 reply; 70+ messages in thread
From: Michael Paquier @ 2016-08-10 12:24 UTC (permalink / raw)
To: Petr Jelinek <petr@2ndquadrant.com>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; pgsql-hackers
On Wed, Aug 10, 2016 at 8:50 PM, Petr Jelinek <petr@2ndquadrant.com> wrote:
> On 10/08/16 13:33, Tomas Vondra wrote:
>>
>> On 08/10/2016 06:41 AM, Michael Paquier wrote:
>>>
>>> On Wed, Aug 3, 2016 at 10:58 AM, Tomas Vondra
>>>>
>>>> 2) combining multiple statistics
>>>>
>>>>
>>>> I think the ability to combine multivariate statistics (covering
>>>> different
>>>> subsets of conditions) is important and useful, but I'm starting to
>>>> think
>>>> that the current implementation may not be the correct one (which is
>>>> why I
>>>> haven't written the SGML docs about this part of the patch series yet).
>>>>
>>>> Assume there's a table "t" with 3 columns (a, b, c), and that we're
>>>> estimating query:
>>>>
>>>> SELECT * FROM t WHERE a = 1 AND b = 2 AND c = 3
>>>>
>>>> but that we only have two statistics (a,b) and (b,c). The current
>>>> patch does
>>>> about this:
>>>>
>>>> P(a=1,b=2,c=3) = P(a=1,b=2) * P(c=3|b=2)
>>>>
>>>> i.e. it estimates the first two conditions using (a,b), and then
>>>> estimates
>>>> (c=3) using (b,c) with "b=2" as a condition. Now, this is very
>>>> efficient,
>>>> but it only works as long as the query contains conditions
>>>> "connecting" the
>>>> two statistics. So if we remove the "b=2" condition from the query, this
>>>> stops working.
>>>
>>>
>>> This is trying to make the algorithm smarter than the user, which is
>>> something I'd think we could live without. In this case statistics on
>>> (a,c) or (a,b,c) are missing. And what if the user does not want to
>>> make use of stats for (a,c) because he only defined (a,b) and (b,c)?
>>>
>>
>> I don't think so. Obviously, if you have statistics covering all the
>> conditions - great, we can't really do better than that.
>>
>> But there's a crucial relation between the number of dimensions of the
>> statistics and accuracy of the statistics. Let's say you have statistics
>> on 8 columns, and you split each dimension twice to build a histogram -
>> that's 256 buckets right there, and we only get ~50% selectivity in each
>> dimension (the actual histogram building algorithm is more complex, but
>> you get the idea).
>
> I think it makes sense to pursue this, but I also think we can easily live
> with not having it in the first version that gets committed and doing it as
> follow-up patch.
This patch is large and complicated enough. As this is not a mandatory
piece to get a basic support, I'd suggest just to drop that for later.
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-10 13:29 Ants Aasma <ants.aasma@eesti.ee>
1 sibling, 1 reply; 70+ messages in thread
From: Ants Aasma @ 2016-08-10 13:29 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; petr@2ndquadrant.com, Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Wed, Aug 3, 2016 at 4:58 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> 2) combining multiple statistics
>
> I think the ability to combine multivariate statistics (covering different
> subsets of conditions) is important and useful, but I'm starting to think
> that the current implementation may not be the correct one (which is why I
> haven't written the SGML docs about this part of the patch series yet).
While researching this topic a few years ago I came across a paper on
this exact topic called "Consistently Estimating the Selectivity of
Conjuncts of Predicates" [1]. While effective it seems to be quite
heavy-weight, so would probably need support for tiered optimization.
[1] https://courses.cs.washington.edu/courses/cse544/11wi/papers/markl-vldb-2005.pdf
Regards,
Ants Aasma
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-10 18:07 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Ants Aasma <ants.aasma@eesti.ee>
0 siblings, 0 replies; 70+ messages in thread
From: Tomas Vondra @ 2016-08-10 18:07 UTC (permalink / raw)
To: Ants Aasma <ants.aasma@eesti.ee>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; petr@2ndquadrant.com, Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 08/10/2016 03:29 PM, Ants Aasma wrote:
> On Wed, Aug 3, 2016 at 4:58 AM, Tomas Vondra
> <tomas.vondra@2ndquadrant.com> wrote:
>> 2) combining multiple statistics
>>
>> I think the ability to combine multivariate statistics (covering different
>> subsets of conditions) is important and useful, but I'm starting to think
>> that the current implementation may not be the correct one (which is why I
>> haven't written the SGML docs about this part of the patch series yet).
>
> While researching this topic a few years ago I came across a paper on
> this exact topic called "Consistently Estimating the Selectivity of
> Conjuncts of Predicates" [1]. While effective it seems to be quite
> heavy-weight, so would probably need support for tiered optimization.
>
> [1] https://courses.cs.washington.edu/courses/cse544/11wi/papers/markl-vldb-2005.pdf
>
I think I've read that paper some time ago, and IIRC it's solving the
same problem but in a very different way - instead of combining the
statistics directly, it relies on the "partial" selectivities and then
estimates the total selectivity using the maximum-entropy principle.
I think it's a nice idea and it probably works fine in many cases, but
it kinda throws away part of the information (that we could get by
matching the statistics against each other directly). But I'll keep that
paper in mind, and we can revisit this solution later.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-10 18:09 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Michael Paquier <michael.paquier@gmail.com>
0 siblings, 0 replies; 70+ messages in thread
From: Tomas Vondra @ 2016-08-10 18:09 UTC (permalink / raw)
To: Michael Paquier <michael.paquier@gmail.com>; Petr Jelinek <petr@2ndquadrant.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; pgsql-hackers
On 08/10/2016 02:24 PM, Michael Paquier wrote:
> On Wed, Aug 10, 2016 at 8:50 PM, Petr Jelinek <petr@2ndquadrant.com> wrote:
>> On 10/08/16 13:33, Tomas Vondra wrote:
>>>
>>> On 08/10/2016 06:41 AM, Michael Paquier wrote:
>>>>
>>>> On Wed, Aug 3, 2016 at 10:58 AM, Tomas Vondra
>>>>>
>>>>> 2) combining multiple statistics
>>>>>
>>>>>
>>>>> I think the ability to combine multivariate statistics (covering
>>>>> different
>>>>> subsets of conditions) is important and useful, but I'm starting to
>>>>> think
>>>>> that the current implementation may not be the correct one (which is
>>>>> why I
>>>>> haven't written the SGML docs about this part of the patch series yet).
>>>>>
>>>>> Assume there's a table "t" with 3 columns (a, b, c), and that we're
>>>>> estimating query:
>>>>>
>>>>> SELECT * FROM t WHERE a = 1 AND b = 2 AND c = 3
>>>>>
>>>>> but that we only have two statistics (a,b) and (b,c). The current
>>>>> patch does
>>>>> about this:
>>>>>
>>>>> P(a=1,b=2,c=3) = P(a=1,b=2) * P(c=3|b=2)
>>>>>
>>>>> i.e. it estimates the first two conditions using (a,b), and then
>>>>> estimates
>>>>> (c=3) using (b,c) with "b=2" as a condition. Now, this is very
>>>>> efficient,
>>>>> but it only works as long as the query contains conditions
>>>>> "connecting" the
>>>>> two statistics. So if we remove the "b=2" condition from the query, this
>>>>> stops working.
>>>>
>>>>
>>>> This is trying to make the algorithm smarter than the user, which is
>>>> something I'd think we could live without. In this case statistics on
>>>> (a,c) or (a,b,c) are missing. And what if the user does not want to
>>>> make use of stats for (a,c) because he only defined (a,b) and (b,c)?
>>>>
>>>
>>> I don't think so. Obviously, if you have statistics covering all the
>>> conditions - great, we can't really do better than that.
>>>
>>> But there's a crucial relation between the number of dimensions of the
>>> statistics and accuracy of the statistics. Let's say you have statistics
>>> on 8 columns, and you split each dimension twice to build a histogram -
>>> that's 256 buckets right there, and we only get ~50% selectivity in each
>>> dimension (the actual histogram building algorithm is more complex, but
>>> you get the idea).
>>
>> I think it makes sense to pursue this, but I also think we can easily live
>> with not having it in the first version that gets committed and doing it as
>> follow-up patch.
>
> This patch is large and complicated enough. As this is not a mandatory
> piece to get a basic support, I'd suggest just to drop that for later.
Which is why combining multiple statistics is in part 0006 and all the
previous parts simply choose the single "best" statistics ;-)
I'm perfectly fine with committing just the first few parts, and leaving
0006 for the next major version.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-10 18:34 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Michael Paquier <michael.paquier@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Tomas Vondra @ 2016-08-10 18:34 UTC (permalink / raw)
To: Michael Paquier <michael.paquier@gmail.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; pgsql-hackers
On 08/10/2016 02:23 PM, Michael Paquier wrote:
> On Wed, Aug 10, 2016 at 8:33 PM, Tomas Vondra
> <tomas.vondra@2ndquadrant.com> wrote:
>> On 08/10/2016 06:41 AM, Michael Paquier wrote:
>>> Patch 0001: there have been comments about that before, and you have
>>> put the checks on RestrictInfo in a couple of variables of
>>> pull_varnos_walker, so nothing to say from here.
>>>
>>
>> I don't follow. Are you suggesting 0001 is a reasonable fix, or that there's
>> a proposed solution?
>
> I think that's reasonable.
>
Well, to me the 0001 feels more like a temporary workaround rather than
a proper solution. I just don't know how to deal with it so I've kept it
for now. Pretty sure there will be complaints that adding RestrictInfo
to the expression walkers is not a nice idea.
>> ...
>>
>> The idea is that the syntax should work even for statistics built on
>> multiple tables, e.g. to provide better statistics for joins. That's why the
>> schema may be specified (as each table might be in different schema), and so
>> on.
>
> So you mean that the same statistics could be shared between tables?
> But as this is visibly not a concept introduced yet in this set of
> patches, why not just cut it off for now to simplify the whole? If
> there is no schema-related field in pg_mv_statistics we could still
> add it later if it proves to be useful.
>
Yes, I think creating statistics on multiple tables is one of the
possible future directions. One of the previous patch versions
introduced ALTER TABLE ... ADD STATISTICS syntax, but that ran into
issues in gram.y, and given the multi-table possibilities the CREATE
STATISTICS seems like a much better idea anyway.
But I guess you're right we may make this a bit more strict now, and
relax it in the future if needed. For example as we only support
single-table statistics at this point, we may remove the schema and
always create the statistics in the schema of the table.
But I don't think we should make the statistics names unique only within
a table (instead of within the schema).
The difference between those two cases is that if we allow multi-table
statistics in the future, we can simply allow specifying the schema and
everything will work just fine. But it'd break the second case, as it
might result in conflicts in existing schemas.
I do realize this might be seen as a case of "future proofing" based on
dubious predictions of how something might work, but OTOH this (schema
inherited from table, unique within a schema) is pretty consistent with
how this work for indexes.
>>> +
>>> /*
>>> Spurious noise in the patch.
>>>
>>> + /* check that at least some statistics were requested */
>>> + if (!build_dependencies)
>>> + ereport(ERROR,
>>> + (errcode(ERRCODE_SYNTAX_ERROR),
>>> + errmsg("no statistics type (dependencies) was
>>> requested")));
>>> So, WITH (dependencies) is mandatory in any case. Why not just
>>> dropping it from the first cut then?
>>
>>
>> Because the follow-up patches extend this to require at least one statistics
>> type. So in 0004 it becomes
>>
>> if (!(build_dependencies || build_mcv))
>>
>> and in 0005 it's
>>
>> if (!(build_dependencies || build_mcv || build_histogram))
>>
>> We might drop it from 0002 (and assume build_dependencies=true), and then
>> add the check in 0004. But it seems a bit pointless.
>
> This is a complicated set of patches. I'd think that we should try to
> simplify things as much as possible first, and the WITH clause is not
> mandatory to have as of 0002.
>
OK, I can remove the WITH from the 0002 part. Not a big deal.
>>> Statistics definition reorder the columns by itself depending on their
>>> order. For example:
>>> create table aa (a int, b int);
>>> create statistics aas on aa(b, a) with (dependencies);
>>> \d aa
>>> "public.aas" (dependencies) ON (a, b)
>>> As this defines a correlation between multiple columns, isn't it wrong
>>> to assume that (b, a) and (a, b) are always the same correlation? I
>>> don't recall such properties as being always commutative (old
>>> memories, I suck at stats in general). [...reading README...] So this
>>> is caused by the implementation limitations that only limit the
>>> analysis between interactions of two columns. Still it seems incorrect
>>> to reorder the user-visible portion.
>>
>> I don't follow. If you talk about Pearson's correlation, that clearly does
>> not depend on the order of columns - it's perfectly independent of that. If
>> you talk about about correlation in the wider sense (i.e. arbitrary
>> dependence between columns), that might depend - but I don't remember a
>> single piece of the patch where this might be a problem.
>
> Yes, based on what is done in the patch that may not be a problem, but
> I am wondering if this is not restricting things too much.
>
Let's keep the code as it is. If we run into this issue in the future,
we can easily relax this - there's nothing depending on the ordering of
attnums, IIRC.
>> Also, which README states that we can only analyze interactions between two
>> columns? That's pretty clearly not the case - the patch should handle
>> dependencies between more columns without any problems.
>
> I have noticed that the patch evaluates all the set of permutations
> possible using a column list, it seems to me though that say if we
> have three columns (a,b,c) listed in a statistics, (a,b) => c and
> (b,a) => c are two different things.
>
Yes, those are two different functional dependencies, of course. But the
algorithm (during ANALYZE) should discover all of them, and even the
examples are using three columns, so I'm not sure what you mean by
"analyze interactions between two columns"?
>>> There is a lot of mumbo-jumbo regarding the way dependencies are
>>> stored with mainly serialize_mv_dependencies and
>>> deserialize_mv_dependencies that operates them from bytea/dep trees.
>>> That's not cool and not portable because pg_mv_statistic represents
>>> that as pure bytea. I would suggest creating a generic data type that
>>> does those operations, named like pg_dependency_tree and then use that
>>> in those new catalogs. pg_node_tree is a precedent of such a thing.
>>> New features could as well make use of this new data type of we are
>>> able to design that in a way generic enough, so that would be a base
>>> patch that the current 0002 applies on top of.
>>
>>
>> Interesting idea, haven't thought about that. So are you suggesting to add a
>> data type for each statistics type (dependencies, MCV, histogram, ...)?
>
> Yes that would be something like that, it would be actually perhaps
> better to have one single data type, and be able to switch between
> each model easily instead of putting byteas in the catalog.
Hmmm, not sure about that. For example what about combinations of
statistics - e.g. when we have MCV list on the most common values and a
histogram on the rest? Should we store both as a single value, or would
that be in two separate values, or what?
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-11 05:55 Michael Paquier <michael.paquier@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
0 siblings, 0 replies; 70+ messages in thread
From: Michael Paquier @ 2016-08-11 05:55 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; pgsql-hackers
On Thu, Aug 11, 2016 at 3:34 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> On 08/10/2016 02:23 PM, Michael Paquier wrote:
>>
>> On Wed, Aug 10, 2016 at 8:33 PM, Tomas Vondra
>> <tomas.vondra@2ndquadrant.com> wrote:
>>> The idea is that the syntax should work even for statistics built on
>>> multiple tables, e.g. to provide better statistics for joins. That's why
>>> the
>>> schema may be specified (as each table might be in different schema), and
>>> so
>>> on.
>>
>>
>> So you mean that the same statistics could be shared between tables?
>> But as this is visibly not a concept introduced yet in this set of
>> patches, why not just cut it off for now to simplify the whole? If
>> there is no schema-related field in pg_mv_statistics we could still
>> add it later if it proves to be useful.
>>
>
> Yes, I think creating statistics on multiple tables is one of the possible
> future directions. One of the previous patch versions introduced ALTER TABLE
> ... ADD STATISTICS syntax, but that ran into issues in gram.y, and given the
> multi-table possibilities the CREATE STATISTICS seems like a much better
> idea anyway.
>
> But I guess you're right we may make this a bit more strict now, and relax
> it in the future if needed. For example as we only support single-table
> statistics at this point, we may remove the schema and always create the
> statistics in the schema of the table.
This would simplify the code the code a bit so I'd suggest removing
that from the first shot. If there is demand for it, keeping the
infrastructure open for this extension is what we had better do.
> But I don't think we should make the statistics names unique only within a
> table (instead of within the schema).
They could be made unique using (name, table_oid, column_list).
>>>> There is a lot of mumbo-jumbo regarding the way dependencies are
>>>> stored with mainly serialize_mv_dependencies and
>>>> deserialize_mv_dependencies that operates them from bytea/dep trees.
>>>> That's not cool and not portable because pg_mv_statistic represents
>>>> that as pure bytea. I would suggest creating a generic data type that
>>>> does those operations, named like pg_dependency_tree and then use that
>>>> in those new catalogs. pg_node_tree is a precedent of such a thing.
>>>> New features could as well make use of this new data type of we are
>>>> able to design that in a way generic enough, so that would be a base
>>>> patch that the current 0002 applies on top of.
>>>
>>>
>>>
>>> Interesting idea, haven't thought about that. So are you suggesting to
>>> add a
>>> data type for each statistics type (dependencies, MCV, histogram, ...)?
>>
>>
>> Yes that would be something like that, it would be actually perhaps
>> better to have one single data type, and be able to switch between
>> each model easily instead of putting byteas in the catalog.
>
> Hmmm, not sure about that. For example what about combinations of statistics
> - e.g. when we have MCV list on the most common values and a histogram on
> the rest? Should we store both as a single value, or would that be in two
> separate values, or what?
The same statistics can combine two different things, using different
columns may depend on how readable things get...
Btw, for the format we could get inspired from pg_node_tree, with pg_stat_tree:
{HISTOGRAM :arg {BUCKET :index 0 :minvals ... }}
{DEPENDENCY :arg {:elt "a => c" ...} ... }
{MVC :arg {:index 0 :values {0,0} ... } ... }
Please consider that as a tentative idea to make things more friendly.
Others may have a different opinion on the matter.
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-15 20:50 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Michael Paquier <michael.paquier@gmail.com>
1 sibling, 0 replies; 70+ messages in thread
From: Tomas Vondra @ 2016-08-15 20:50 UTC (permalink / raw)
To: Michael Paquier <michael.paquier@gmail.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; Robert Haas <robertmhaas@gmail.com>; david@pgmasters.net, Tom Lane <tgl@sss.pgh.pa.us>; Alvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; pgsql-hackers
On 08/10/2016 06:41 AM, Michael Paquier wrote:
> On Wed, Aug 3, 2016 at 10:58 AM, Tomas Vondra
> <tomas.vondra@2ndquadrant.com> wrote:
>> 1) enriching the query tree with multivariate statistics info
>>
>> Right now all the stuff related to multivariate statistics estimation
>> happens in clausesel.c - matching condition to statistics, selection of
>> statistics to use (if there are multiple usable stats), etc. So pretty much
>> all this info is internal to clausesel.c and does not get outside.
>
> This does not seem bad to me as first sight but...
>
>> I'm starting to think that some of the steps (matching quals to stats,
>> selection of stats) should happen in a "preprocess" step before the actual
>> estimation, storing the information (which stats to use, etc.) in a new type
>> of node in the query tree - something like RestrictInfo.
>>
>> I believe this needs to happen sometime after deconstruct_jointree() as that
>> builds RestrictInfos nodes, and looking at planmain.c, right after
>> extract_restriction_or_clauses seems about right. Haven't tried, though.
>>
>> This would move all the "statistics selection" logic from clausesel.c,
>> separating it from the "actual estimation" and simplifying the code.
>>
>> But more importantly, I think we'll need to show some of the data in EXPLAIN
>> output. With per-column statistics it's fairly straightforward to determine
>> which statistics are used and how. But with multivariate stats things are
>> often more complicated - there may be multiple candidate statistics (e.g.
>> histograms covering different subsets of the conditions), it's possible to
>> apply them in different orders, etc.
>>
>> But EXPLAIN can't show the info if it's ephemeral and available only within
>> clausesel.c (and thrown away after the estimation).
>
> This gives a good reason to not do that in clauserel.c, it would be
> really cool to be able to get some information regarding the stats
> used with a simple EXPLAIN.
I've been thinking about this, and I'm afraid it's way more complicated
in practice. It essentially means doing something like
rel->baserestrictinfo = enrichWithStatistics(rel->baserestrictinfo);
for each table (and in the future maybe also for joins etc.) But as the
name suggests the list should only include RestrictInfo nodes, which
seems to contradict the transformation.
For example with conditions
WHERE (a=1) AND (b=2) AND (c=3)
the list will contain 3 RestrictInfos. But if there's a statistics on
(a,b,c), we need to note that somehow - my plan was to inject a node
storing this information, something like (a bit simplified):
StatisticsInfo {
Oid statisticsoid; /* OID of the statistics */
List *mvconditions; /* estimate using the statistics */
List *otherconditions; /* estimate the old way */
}
But that'd clearly violate the assumption that baserestrictinfo only
contains RestrictInfo. I don't think it's feasible (or desirable) to
rework all the places to expect both RestrictInfo and the new node.
I can think of two alternatives:
1) keep the transformed list as separate list, next to baserestrictinfo
This obviously fixes the issue, as each caller can decide which node it
wants. But it also means we need to maintain two lists instead of one,
and keep them synchronized.
2) embed the information into the existing tree
It might be possible to store the information in existing nodes, i.e.
each node would track whether it's estimated the "old way" or using
multivariate statistics (and which one). But it would require changing
many of the existing nodes (at least those compatible with multivariate
statistics: currently OpExpr, NullTest, ...).
And it also seems fairly difficult to reconstruct the information during
the estimation, as it'd be necessary to look for other nodes to be
estimated by the same statistics. Which seems to defeat the idea of
preprocessing to some degree.
So I'm not sure what's the best solution. I'm leaning to (1), i.e.
keeping a separate list, but I'd welcome other ideas.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-23 17:03 Robert Haas <robertmhaas@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
2 siblings, 1 reply; 70+ messages in thread
From: Robert Haas @ 2016-08-23 17:03 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Tue, Aug 2, 2016 at 9:58 PM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> Attached is v19 of the "multivariate stats" patch series - essentially v18
> rebased on top of current master.
Tom:
ISTR that you were going to try to look at this patch set. It seems
from the discussion that it's not really ready for serious
consideration for commit yet, but also that some high-level design
comments from you at this stage could go a long way toward making sure
that the final form of the patch is something that will be acceptable.
I'd really like to see us get some kind of capability along these
lines, but I'm sure it will go a lot better if you or Dean handle it
than if I try to do it ... not to mention that there are only so many
hours in the day.
--
Robert Haas
EnterpriseDB: http://www.enterprisedb.com
The Enterprise PostgreSQL Company
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-08-30 06:54 Michael Paquier <michael.paquier@gmail.com>
parent: Robert Haas <robertmhaas@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Michael Paquier @ 2016-08-30 06:54 UTC (permalink / raw)
To: Robert Haas <robertmhaas@gmail.com>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Wed, Aug 24, 2016 at 2:03 AM, Robert Haas <robertmhaas@gmail.com> wrote:
> ISTR that you were going to try to look at this patch set. It seems
> from the discussion that it's not really ready for serious
> consideration for commit yet, but also that some high-level design
> comments from you at this stage could go a long way toward making sure
> that the final form of the patch is something that will be acceptable.
>
> I'd really like to see us get some kind of capability along these
> lines, but I'm sure it will go a lot better if you or Dean handle it
> than if I try to do it ... not to mention that there are only so many
> hours in the day.
Agreed. What I have been able to look until now was the high-level
structure of the patch, and I think that we should really shave 0002
and simplify it to get a core infrastructure in place, but the core
patch is at another level, and it would be good to get some feedback
regarding the structure of the patch and if it is moving in the good
direction is good or not.
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-09-12 14:08 Dean Rasheed <dean.a.rasheed@gmail.com>
parent: Michael Paquier <michael.paquier@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Dean Rasheed @ 2016-09-12 14:08 UTC (permalink / raw)
To: Michael Paquier <michael.paquier@gmail.com>; +Cc: Robert Haas <robertmhaas@gmail.com>; Tomas Vondra <tomas.vondra@2ndquadrant.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 3 August 2016 at 02:58, Tomas Vondra <tomas.vondra@2ndquadrant.com> wrote:
> Attached is v19 of the "multivariate stats" patch series
Hi,
I started looking at this - just at a very high level - I've not read
much of the detail yet, but here are some initial review comments.
I think the overall infrastructure approach for CREATE STATISTICS
makes sense, and I agree with other suggestions upthread that it would
be useful to be able to build statistics on arbitrary expressions,
although that doesn't need to be part of this patch, it's useful to
keep that in mind as a possible future extension of this initial
design.
I can imagine it being useful to be able to create user-defined
statistics on an arbitrary list of expressions, and I think that would
include univariate as well as multivariate statistics. Perhaps that's
something to take into account in the naming of things, e.g., as David
Rowley suggested, something like pg_statistic_ext, rather than
pg_mv_statistic.
I also like the idea that this might one day be extended to support
statistics across multiple tables, although I think that might be
challenging to achieve -- you'd need a method of taking a random
sample of rows from a join between 2 or more tables. However, if the
intention is to be able to support that one day, I think that needs to
be accounted for in the syntax now -- specifically, I think it will be
too limiting to only support things extending the current syntax of
the form table1(col1, col2, ...), table2(col1, col2, ...), because
that precludes building statistics on an expression referring to
columns from more than one table. So I think we should plan further
ahead and use a syntax giving greater flexibility in the future, for
example something structured more like a query (like CREATE VIEW):
CREATE STATISTICS name
[ WITH (options) ]
ON expression [, ...]
FROM table [, ...]
WHERE condition
where the first version of the patch would only support expressions
that are simple column references, and would require at least 2 such
columns from a single table with no WHERE clause, i.e.:
CREATE STATISTICS name
[ WITH (options) ]
ON column1, column2 [, ...]
FROM table
For multi-table statistics, a WHERE clause would typically be needed
to specify how the tables are expected to be joined, but potentially
such a clause might also be useful in single-table statistics, to
build partial statistics on a commonly queried subset of the table,
just like a partial index.
Of course, I'm not suggesting that the current patch do any of that --
it's big enough as it is. I'm just throwing out possible future
directions this might go in, so that we don't get painted into a
corner when designing the syntax for the current patch.
Regarding the statistics themselves, I read the description of soft
functional dependencies, and I'm somewhat skeptical about that
algorithm. I don't like the arbitrary thresholds or the sudden jump
from independence to dependence and clause reduction. As others have
said, I think this should account for a continuous spectrum of
dependence from fully independent to fully dependent, and combine
clause selectivities in a way based on the degree of dependence. For
example, if you computed an estimate for the fraction 'f' of the
table's rows for which a -> b, then it might be reasonable to combine
the selectivities using
P(a,b) = P(a) * (f + (1-f) * P(b))
Of course, having just a single number that tells you the columns are
correlated, tells you nothing about whether the clauses on those
columns are consistent with that correlation. For example, in the
following table
CREATE TABLE t(a int, b int);
INSERT INTO t SELECT x/10, ((x/10)*789)%100 FROM generate_series(0,999) g(x);
'b' is functionally dependent on 'a' (and vice versa), but if you
query the rows with a<50 and with b<50, those clauses behave
essentially independently, because they're not consistent with the
functional dependence between 'a' and 'b', so the best way to combine
their selectivities is just to multiply them, as we currently do.
So whilst it may be interesting to determine that 'b' is functionally
dependent on 'a', it's not obvious whether that fact by itself should
be used in the selectivity estimates. Perhaps it should, on the
grounds that it's best to attempt to use all the available
information, but only if there are no more detailed statistics
available. In any case, knowing that there is a correlation can be
used as an indicator that it may be worthwhile to build more detailed
multivariate statistics like a MCV list or a histogram on those
columns.
Looking at the ndistinct coefficient 'q', I think it would be better
if the recorded statistic were just the estimate for
ndistinct(a,b,...) rather than a ratio of ndistinct values. That's a
more fundamental statistic, and it's easier to document and easier to
interpret. Also, I don't believe that the coefficient 'q' is the right
number to use for clause estimation:
Looking at README.ndistinct, I'm skeptical about the selectivity
estimation argument. In the case where a -> b, you'd have q =
ndistinct(b), so then P(a=1 & b=2) would become 1/ndistinct(a), which
is fine for a uniform distribution. But typically, there would be
univariate statistics on a and b, so if for example a=1 were 100x more
likely than average, you'd probably know that and the existing code
computing P(a=1) would reflect that, whereas simply using P(a=1 & b=2)
= 1/ndistinct(a) would be a significant underestimate, since it would
be ignoring known information about the distribution of a.
But likewise if, as is later argued, you were to use 'q' as a
correction factor applied to the individual clause selectivities, you
could end up with significant overestimates: if you said P(a=1 & b=2)
= q * P(a=1) * P(b=2), and a=1 were 100x more likely than average, and
a -> b, then b=2 would also be 100x more likely than average (assuming
that b=2 was the value implied by the functional dependency), and that
would also be reflected in the univariate statics on b, so then you'd
end up with an overall selectivity of around 10000/ndistinct(a), which
would be 100x too big. In fact, since a -> b means that q =
ndistinct(b), there's a good chance of hitting data for which q * P(b)
is greater than 1, so this formula would lead to a combined
selectivity greater than P(a), which is obviously nonsense.
Having a better estimate for ndistinct(a,b,...) looks very useful by
itself for GROUP BY estimation, and there may be other places that
would benefit from it, but I don't think it's the best statistic for
determining functional dependence or combining clause selectivities.
That's as much as I've looked at so far. It's such a big patch that
it's difficult to consider all at once. I think perhaps the smallest
committable self-contained unit providing a tangible benefit would be
something containing the core infrastructure plus the ndistinct
estimate and the improved GROUP BY estimation.
Regards,
Dean
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-09-13 22:01 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Dean Rasheed <dean.a.rasheed@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Tomas Vondra @ 2016-09-13 22:01 UTC (permalink / raw)
To: Dean Rasheed <dean.a.rasheed@gmail.com>; Michael Paquier <michael.paquier@gmail.com>; +Cc: Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Hi,
Thanks for looking into this!
On 09/12/2016 04:08 PM, Dean Rasheed wrote:
> On 3 August 2016 at 02:58, Tomas Vondra <tomas.vondra@2ndquadrant.com> wrote:
>> Attached is v19 of the "multivariate stats" patch series
>
> Hi,
>
> I started looking at this - just at a very high level - I've not read
> much of the detail yet, but here are some initial review comments.
>
> I think the overall infrastructure approach for CREATE STATISTICS
> makes sense, and I agree with other suggestions upthread that it would
> be useful to be able to build statistics on arbitrary expressions,
> although that doesn't need to be part of this patch, it's useful to
> keep that in mind as a possible future extension of this initial
> design.
>
> I can imagine it being useful to be able to create user-defined
> statistics on an arbitrary list of expressions, and I think that would
> include univariate as well as multivariate statistics. Perhaps that's
> something to take into account in the naming of things, e.g., as David
> Rowley suggested, something like pg_statistic_ext, rather than
> pg_mv_statistic.
>
> I also like the idea that this might one day be extended to support
> statistics across multiple tables, although I think that might be
> challenging to achieve -- you'd need a method of taking a random
> sample of rows from a join between 2 or more tables. However, if the
> intention is to be able to support that one day, I think that needs to
> be accounted for in the syntax now -- specifically, I think it will be
> too limiting to only support things extending the current syntax of
> the form table1(col1, col2, ...), table2(col1, col2, ...), because
> that precludes building statistics on an expression referring to
> columns from more than one table. So I think we should plan further
> ahead and use a syntax giving greater flexibility in the future, for
> example something structured more like a query (like CREATE VIEW):
>
> CREATE STATISTICS name
> [ WITH (options) ]
> ON expression [, ...]
> FROM table [, ...]
> WHERE condition
>
> where the first version of the patch would only support expressions
> that are simple column references, and would require at least 2 such
> columns from a single table with no WHERE clause, i.e.:
>
> CREATE STATISTICS name
> [ WITH (options) ]
> ON column1, column2 [, ...]
> FROM table
>
> For multi-table statistics, a WHERE clause would typically be needed
> to specify how the tables are expected to be joined, but potentially
> such a clause might also be useful in single-table statistics, to
> build partial statistics on a commonly queried subset of the table,
> just like a partial index.
Hmm, the "partial statistics" idea seems interesting, It would allow us
to provide additional / more detailed statistics only for a subset of a
table.
I'm however not sure about the join case - how would the syntax work
with outer joins? But as you said, we only need
CREATE STATISTICS name
[ WITH (options) ]
ON (column1, column2 [, ...])
FROM table
WHERE condition
until we add support for join statistics.
>
> Regarding the statistics themselves, I read the description of soft
> functional dependencies, and I'm somewhat skeptical about that
> algorithm. I don't like the arbitrary thresholds or the sudden jump
> from independence to dependence and clause reduction. As others have
> said, I think this should account for a continuous spectrum of
> dependence from fully independent to fully dependent, and combine
> clause selectivities in a way based on the degree of dependence. For
> example, if you computed an estimate for the fraction 'f' of the
> table's rows for which a -> b, then it might be reasonable to combine
> the selectivities using
>
> P(a,b) = P(a) * (f + (1-f) * P(b))
>
Yeah, I agree that the thresholds resulting in sudden changes between
"dependent" and "not dependent" are annoying. The question is whether it
makes sense to fix that, though - the functional dependencies were meant
as the simplest form of statistics, allowing us to get the rest of the
infrastructure in.
I'm OK with replacing the true/false dependencies with a degree of
dependency between 0 and 1, but I'm a bit afraid it'll result in
complaints that the first patch got too large / complicated.
It also contradicts the idea of using functional dependencies as a
low-overhead type of statistics, filtering the list of clauses that need
to be estimated using more expensive types of statistics (MCV lists,
histograms, ...). Switching to a degree of dependency would prevent
removal of "unnecessary" clauses.
> Of course, having just a single number that tells you the columns are
> correlated, tells you nothing about whether the clauses on those
> columns are consistent with that correlation. For example, in the
> following table
>
> CREATE TABLE t(a int, b int);
> INSERT INTO t SELECT x/10, ((x/10)*789)%100 FROM generate_series(0,999) g(x);
>
> 'b' is functionally dependent on 'a' (and vice versa), but if you
> query the rows with a<50 and with b<50, those clauses behave
> essentially independently, because they're not consistent with the
> functional dependence between 'a' and 'b', so the best way to combine
> their selectivities is just to multiply them, as we currently do.
>
> So whilst it may be interesting to determine that 'b' is functionally
> dependent on 'a', it's not obvious whether that fact by itself should
> be used in the selectivity estimates. Perhaps it should, on the
> grounds that it's best to attempt to use all the available
> information, but only if there are no more detailed statistics
> available. In any case, knowing that there is a correlation can be
> used as an indicator that it may be worthwhile to build more detailed
> multivariate statistics like a MCV list or a histogram on those
> columns.
>
Right. IIRC this is actually described in the README as "incompatible
conditions". While implementing it, I concluded that this is OK and it's
up to the developer to decide whether the queries are compatible with
the "assumption of compatibility". But maybe this is reasoning is bogus
and makes (the current implementation of) functional dependencies
unusable in practice.
But I like the idea of reverting the order from
(a) look for functional dependencies
(b) reduce the clauses using functional dependencies
(c) estimate the rest using multivariate MCV/histograms
to
(a) estimate the rest using multivariate MCV/histograms
(b) try to apply functional dependencies on the remaining clauses
It contradicts the idea of functional dependencies as "low-overhead
statistics" but maybe it's worth it.
>
> Looking at the ndistinct coefficient 'q', I think it would be better
> if the recorded statistic were just the estimate for
> ndistinct(a,b,...) rather than a ratio of ndistinct values. That's a
> more fundamental statistic, and it's easier to document and easier to
> interpret. Also, I don't believe that the coefficient 'q' is the right
> number to use for clause estimation:
>
IIRC the reason why I stored the coefficient instead of the ndistinct()
values is that the coefficients are not directly related to number of
rows in the original relation, so you can apply it directly to whatever
cardinality estimate you have.
Otherwise it's mostly the same information - it's trivial to compute one
from the other.
>
> Looking at README.ndistinct, I'm skeptical about the selectivity
> estimation argument. In the case where a -> b, you'd have q =
> ndistinct(b), so then P(a=1 & b=2) would become 1/ndistinct(a), which
> is fine for a uniform distribution. But typically, there would be
> univariate statistics on a and b, so if for example a=1 were 100x more
> likely than average, you'd probably know that and the existing code
> computing P(a=1) would reflect that, whereas simply using P(a=1 & b=2)
> = 1/ndistinct(a) would be a significant underestimate, since it would
> be ignoring known information about the distribution of a.
>
> But likewise if, as is later argued, you were to use 'q' as a
> correction factor applied to the individual clause selectivities, you
> could end up with significant overestimates: if you said P(a=1 & b=2)
> = q * P(a=1) * P(b=2), and a=1 were 100x more likely than average, and
> a -> b, then b=2 would also be 100x more likely than average (assuming
> that b=2 was the value implied by the functional dependency), and that
> would also be reflected in the univariate statics on b, so then you'd
> end up with an overall selectivity of around 10000/ndistinct(a), which
> would be 100x too big. In fact, since a -> b means that q =
> ndistinct(b), there's a good chance of hitting data for which q * P(b)
> is greater than 1, so this formula would lead to a combined
> selectivity greater than P(a), which is obviously nonsense.
Well, yeah. The
P(a=1) = 1/ndistinct(a)
was really just a simplification for the uniform distribution, and
looking at "q" as a correction factor is much more practical - no doubt
about that.
As for the overestimated and underestimates - I don't think we can
entirely prevent that. We're essentially replacing one assumption (AVIA)
with other assumptions (homogenity for ndistinct, compatibility for
functional dependencies), hoping that those assumptions are weaker in
some sense. But there'll always be cases that break those assumptions
and I don't think we can prevent that.
Unlike the functional dependencies, this "homogenity" assumption is not
dependent on the queries at all, so it should be possible to verify it
during ANALYZE.
Also, maybe we could/should use the same approach as for functional
dependencies, i.e. try using more detailed statistics first and then
apply ndistinct coefficients only on the remaining clauses?
>
> Having a better estimate for ndistinct(a,b,...) looks very useful by
> itself for GROUP BY estimation, and there may be other places that
> would benefit from it, but I don't think it's the best statistic for
> determining functional dependence or combining clause selectivities.
>
Not sure. I think it may be very useful type of statistics, but I'm not
going to fight for this very hard. I'm fine with ignoring this
statistics type for now, getting the other "detailed" statistics types
(MCV, histograms) in and then revisiting this.
> That's as much as I've looked at so far. It's such a big patch that
> it's difficult to consider all at once. I think perhaps the smallest
> committable self-contained unit providing a tangible benefit would be
> something containing the core infrastructure plus the ndistinct
> estimate and the improved GROUP BY estimation.
>
FWIW I find the ndistinct statistics as rather uninteresting (at least
compared to the other types of statistics), which is why it's the last
patch in the patch series. Perhaps I shouldn't have include it at all,
as it's just a distraction.
regards
Dean
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-09-30 11:10 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
0 siblings, 2 replies; 70+ messages in thread
From: Heikki Linnakangas @ 2016-09-30 11:10 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Dean Rasheed <dean.a.rasheed@gmail.com>; Michael Paquier <michael.paquier@gmail.com>; +Cc: Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
This patch set is in pretty good shape, the only problem is that it's so
big that no-one seems to have the time or courage to do the final
touches and commit it. If we just focus on the functional dependencies
part for now, I think we might get somewhere. I peeked at the MCV and
histogram patches too, and I think they make total sense as well, and
are a natural extension of the functional dependencies patch. So if we
just focus on that for now, I don't think we will paint ourselves in the
corner.
(more below)
On 09/14/2016 01:01 AM, Tomas Vondra wrote:
> On 09/12/2016 04:08 PM, Dean Rasheed wrote:
>> Regarding the statistics themselves, I read the description of soft
>> functional dependencies, and I'm somewhat skeptical about that
>> algorithm. I don't like the arbitrary thresholds or the sudden jump
>> from independence to dependence and clause reduction. As others have
>> said, I think this should account for a continuous spectrum of
>> dependence from fully independent to fully dependent, and combine
>> clause selectivities in a way based on the degree of dependence. For
>> example, if you computed an estimate for the fraction 'f' of the
>> table's rows for which a -> b, then it might be reasonable to combine
>> the selectivities using
>>
>> P(a,b) = P(a) * (f + (1-f) * P(b))
>>
>
> Yeah, I agree that the thresholds resulting in sudden changes between
> "dependent" and "not dependent" are annoying. The question is whether it
> makes sense to fix that, though - the functional dependencies were meant
> as the simplest form of statistics, allowing us to get the rest of the
> infrastructure in.
>
> I'm OK with replacing the true/false dependencies with a degree of
> dependency between 0 and 1, but I'm a bit afraid it'll result in
> complaints that the first patch got too large / complicated.
+1 for using a floating degree between 0 and 1, rather than a boolean.
> It also contradicts the idea of using functional dependencies as a
> low-overhead type of statistics, filtering the list of clauses that need
> to be estimated using more expensive types of statistics (MCV lists,
> histograms, ...). Switching to a degree of dependency would prevent
> removal of "unnecessary" clauses.
That sounds OK to me, although I'm not deeply familiar with this patch yet.
>> Of course, having just a single number that tells you the columns are
>> correlated, tells you nothing about whether the clauses on those
>> columns are consistent with that correlation. For example, in the
>> following table
>>
>> CREATE TABLE t(a int, b int);
>> INSERT INTO t SELECT x/10, ((x/10)*789)%100 FROM generate_series(0,999) g(x);
>>
>> 'b' is functionally dependent on 'a' (and vice versa), but if you
>> query the rows with a<50 and with b<50, those clauses behave
>> essentially independently, because they're not consistent with the
>> functional dependence between 'a' and 'b', so the best way to combine
>> their selectivities is just to multiply them, as we currently do.
>>
>> So whilst it may be interesting to determine that 'b' is functionally
>> dependent on 'a', it's not obvious whether that fact by itself should
>> be used in the selectivity estimates. Perhaps it should, on the
>> grounds that it's best to attempt to use all the available
>> information, but only if there are no more detailed statistics
>> available. In any case, knowing that there is a correlation can be
>> used as an indicator that it may be worthwhile to build more detailed
>> multivariate statistics like a MCV list or a histogram on those
>> columns.
>
> Right. IIRC this is actually described in the README as "incompatible
> conditions". While implementing it, I concluded that this is OK and it's
> up to the developer to decide whether the queries are compatible with
> the "assumption of compatibility". But maybe this is reasoning is bogus
> and makes (the current implementation of) functional dependencies
> unusable in practice.
I think that's OK. It seems like a good assumption that the conditions
are "compatible" with the functional dependency. For two reasons:
1) A query with compatible clauses is much more likely to occur in real
life. Why would you run a query with an incompatible ZIP and city clauses?
2) If the conditions were in fact incompatible, the query is likely to
return 0 rows, and will bail out very quickly, even if the estimates are
way off and you choose a non-optimal plan. There are exceptions, of
course: an index scan might be able to conclude that there are no rows
much quicker than a seqscan, but as a general rule of thumb, a query
that returns 0 rows isn't very sensitive to the chosen plan.
And of course, as long as we're not collecting these statistics
automatically, if it doesn't work for your application, just don't
collect them.
I fear that using "statistics" as the name of the new object might get a
bit awkward. "statistics" is a plural, but we use it as the name of a
single object, like "pants" or "scissors". Not sure I have any better
ideas though. "estimator"? "statistics collection"? Or perhaps it should
be singular, "statistic". I note that you actually called the system
table "pg_mv_statistic", in singular.
I'm not a big fan of storing the stats as just a bytea blob, and having
to use special functions to interpret it. By looking at the patch, it's
not clear to me what we actually store for functional dependencies. A
list of attribute numbers? Could we store them simply as an int[]? (I'm
not a big fan of the hack in pg_statistic, that allows storing arrays of
any data type in the same column, though. But for functional
dependencies, I don't think we need that.)
Overall, this is going to be a great feature!
- Heikki
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-10-03 01:46 Michael Paquier <michael.paquier@gmail.com>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
1 sibling, 1 reply; 70+ messages in thread
From: Michael Paquier @ 2016-10-03 01:46 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Dean Rasheed <dean.a.rasheed@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Fri, Sep 30, 2016 at 8:10 PM, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
> This patch set is in pretty good shape, the only problem is that it's so big
> that no-one seems to have the time or courage to do the final touches and
> commit it.
Did you see my suggestions about simplifying its SQL structure? You
could shave some code without impacting the base set of features.
> I fear that using "statistics" as the name of the new object might get a bit
> awkward. "statistics" is a plural, but we use it as the name of a single
> object, like "pants" or "scissors". Not sure I have any better ideas though.
> "estimator"? "statistics collection"? Or perhaps it should be singular,
> "statistic". I note that you actually called the system table
> "pg_mv_statistic", in singular.
>
> I'm not a big fan of storing the stats as just a bytea blob, and having to
> use special functions to interpret it. By looking at the patch, it's not
> clear to me what we actually store for functional dependencies. A list of
> attribute numbers? Could we store them simply as an int[]? (I'm not a big
> fan of the hack in pg_statistic, that allows storing arrays of any data type
> in the same column, though. But for functional dependencies, I don't think
> we need that.)
I am marking this patch as returned with feedback for now.
> Overall, this is going to be a great feature!
+1.
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-10-03 11:25 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Michael Paquier <michael.paquier@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Heikki Linnakangas @ 2016-10-03 11:25 UTC (permalink / raw)
To: Michael Paquier <michael.paquier@gmail.com>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Dean Rasheed <dean.a.rasheed@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 10/03/2016 04:46 AM, Michael Paquier wrote:
> On Fri, Sep 30, 2016 at 8:10 PM, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>> This patch set is in pretty good shape, the only problem is that it's so big
>> that no-one seems to have the time or courage to do the final touches and
>> commit it.
>
> Did you see my suggestions about simplifying its SQL structure? You
> could shave some code without impacting the base set of features.
Yeah. The idea was to use something like pg_node_tree to store all the
different kinds of statistics, the histogram, the MCV, and the
functional dependencies, in one datum. Or JSON, maybe. It sounds better
than an opaque bytea blob, although I'd prefer something more
relational. For the functional dependencies, I think we could get away
with a simple float array, so let's do that in the first cut, and
revisit this for the MCV and histogram later. Separate columns for the
functional dependencies, the MCVs, and the histogram, probably makes
sense anyway.
- Heikki
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-10-04 03:25 Michael Paquier <michael.paquier@gmail.com>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
0 siblings, 1 reply; 70+ messages in thread
From: Michael Paquier @ 2016-10-04 03:25 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Dean Rasheed <dean.a.rasheed@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Mon, Oct 3, 2016 at 8:25 PM, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
> Yeah. The idea was to use something like pg_node_tree to store all the
> different kinds of statistics, the histogram, the MCV, and the functional
> dependencies, in one datum. Or JSON, maybe. It sounds better than an opaque
> bytea blob, although I'd prefer something more relational. For the
> functional dependencies, I think we could get away with a simple float
> array, so let's do that in the first cut, and revisit this for the MCV and
> histogram later.
OK. A second thing was related to the use of schemas in the new system
catalogs. As mentioned in [1], those could be removed.
[1]: https://www.postgresql.org/message-id/CAB7nPqTU40Q5_NSgHVoMJfbyH1HDtqMbFDJ+kwFJSpam35b3Qg@mail.gmail....
> Separate columns for the functional dependencies, the MCVs,
> and the histogram, probably makes sense anyway.
Probably..
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-10-04 07:37 Dean Rasheed <dean.a.rasheed@gmail.com>
parent: Michael Paquier <michael.paquier@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Dean Rasheed @ 2016-10-04 07:37 UTC (permalink / raw)
To: Michael Paquier <michael.paquier@gmail.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Tomas Vondra <tomas.vondra@2ndquadrant.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 4 October 2016 at 04:25, Michael Paquier <michael.paquier@gmail.com> wrote:
> OK. A second thing was related to the use of schemas in the new system
> catalogs. As mentioned in [1], those could be removed.
> [1]: https://www.postgresql.org/message-id/CAB7nPqTU40Q5_NSgHVoMJfbyH1HDtqMbFDJ+kwFJSpam35b3Qg@mail.gmail....
>
That doesn't work, because if the intention is to be able to one day
support statistics across multiple tables, you can't assume that the
statistics are in the same schema as the table.
In fact, if multi-table statistics are to be allowed in the future, I
think you want to move away from thinking of statistics as depending
on and referring to a single table, and handle them more like views --
i.e, store a pg_node_tree representing the from_clause and add
multiple dependencies at statistics creation time. That was what I was
getting at upthread when I suggested the alternate syntax, and also
answers Tomas' question about how JOIN might one day be supported.
Of course, if we don't think that we will ever support multi-table
statistics, that all goes away, and you may as well make the
statistics name local to the table, but I think that's a bit limiting.
One way or the other, I think this is a question that needs to be
answered now. My vote is to leave expansion room to support
multi-table statistics in the future.
Regards,
Dean
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-10-04 07:49 Dean Rasheed <dean.a.rasheed@gmail.com>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
1 sibling, 1 reply; 70+ messages in thread
From: Dean Rasheed @ 2016-10-04 07:49 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 30 September 2016 at 12:10, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
> I fear that using "statistics" as the name of the new object might get a bit
> awkward. "statistics" is a plural, but we use it as the name of a single
> object, like "pants" or "scissors". Not sure I have any better ideas though.
> "estimator"? "statistics collection"? Or perhaps it should be singular,
> "statistic". I note that you actually called the system table
> "pg_mv_statistic", in singular.
>
I think it's OK. The functional dependency is a single statistic, but
MCV lists and histograms are multiple statistics (multiple facts about
the data sampled), so in general when you create one of these new
objects, you are creating multiple statistics about the data. Also I
find "CREATE STATISTIC" just sounds a bit clumsy compared to "CREATE
STATISTICS".
The convention for naming system catalogs seems to be to use the
singular for tables and plural for views, so I guess we should stick
with that. It doesn't seem like the end of the world that it doesn't
match the user-facing syntax. A bigger concern is the use of "mv" in
the name, because as has already been pointed out, this table may also
in the future be used to store univariate expression and partial
statistics, so I think we should drop the "mv" and go with something
like pg_statistic_ext, or some other more general name.
Regards,
Dean
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-10-04 08:15 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Dean Rasheed <dean.a.rasheed@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Heikki Linnakangas @ 2016-10-04 08:15 UTC (permalink / raw)
To: Dean Rasheed <dean.a.rasheed@gmail.com>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 10/04/2016 10:49 AM, Dean Rasheed wrote:
> On 30 September 2016 at 12:10, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>> I fear that using "statistics" as the name of the new object might get a bit
>> awkward. "statistics" is a plural, but we use it as the name of a single
>> object, like "pants" or "scissors". Not sure I have any better ideas though.
>> "estimator"? "statistics collection"? Or perhaps it should be singular,
>> "statistic". I note that you actually called the system table
>> "pg_mv_statistic", in singular.
>
> I think it's OK. The functional dependency is a single statistic, but
> MCV lists and histograms are multiple statistics (multiple facts about
> the data sampled), so in general when you create one of these new
> objects, you are creating multiple statistics about the data.
Ok. I don't really have any better ideas, was just hoping that someone
else would.
> Also I find "CREATE STATISTIC" just sounds a bit clumsy compared to
> "CREATE STATISTICS".
Agreed.
> The convention for naming system catalogs seems to be to use the
> singular for tables and plural for views, so I guess we should stick
> with that.
However, for tables and views, each object you store in those views is a
"table" or "view", but with this thing, the object you store is
"statistics". Would you have a catalog table called "pg_scissor"?
We call the current system table "pg_statistic", though. I agree we
should call it pg_mv_statistic, in singular, to follow the example of
pg_statistic.
Of course, the user-friendly system view on top of that is called
"pg_stats", just to confuse things more :-).
> It doesn't seem like the end of the world that it doesn't
> match the user-facing syntax. A bigger concern is the use of "mv" in
> the name, because as has already been pointed out, this table may also
> in the future be used to store univariate expression and partial
> statistics, so I think we should drop the "mv" and go with something
> like pg_statistic_ext, or some other more general name.
Also, "mv" makes me think of materialized views, which is completely
unrelated to this.
- Heikki
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-10-04 08:51 Gavin Flower <GavinFlower@archidevsys.co.nz>
parent: Dean Rasheed <dean.a.rasheed@gmail.com>
0 siblings, 0 replies; 70+ messages in thread
From: Gavin Flower @ 2016-10-04 08:51 UTC (permalink / raw)
To: Dean Rasheed <dean.a.rasheed@gmail.com>; Michael Paquier <michael.paquier@gmail.com>; +Cc: Heikki Linnakangas <hlinnaka@iki.fi>; Tomas Vondra <tomas.vondra@2ndquadrant.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 04/10/16 20:37, Dean Rasheed wrote:
> On 4 October 2016 at 04:25, Michael Paquier <michael.paquier@gmail.com> wrote:
>> OK. A second thing was related to the use of schemas in the new system
>> catalogs. As mentioned in [1], those could be removed.
>> [1]: https://www.postgresql.org/message-id/CAB7nPqTU40Q5_NSgHVoMJfbyH1HDtqMbFDJ+kwFJSpam35b3Qg@mail.gmail....
>>
> That doesn't work, because if the intention is to be able to one day
> support statistics across multiple tables, you can't assume that the
> statistics are in the same schema as the table.
>
> In fact, if multi-table statistics are to be allowed in the future, I
> think you want to move away from thinking of statistics as depending
> on and referring to a single table, and handle them more like views --
> i.e, store a pg_node_tree representing the from_clause and add
> multiple dependencies at statistics creation time. That was what I was
> getting at upthread when I suggested the alternate syntax, and also
> answers Tomas' question about how JOIN might one day be supported.
>
> Of course, if we don't think that we will ever support multi-table
> statistics, that all goes away, and you may as well make the
> statistics name local to the table, but I think that's a bit limiting.
> One way or the other, I think this is a question that needs to be
> answered now. My vote is to leave expansion room to support
> multi-table statistics in the future.
>
> Regards,
> Dean
>
>
I can see multi-table statistics being useful if one is trying to
optimise indexes for multiple joins.
Am assuming that the statistics can be accessed by the user as well as
the planner? (I've only lightly followed this thread, so I might have
missed, significant relevant details!)
Cheers,
Gavin
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-10-04 09:21 Dean Rasheed <dean.a.rasheed@gmail.com>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
0 siblings, 1 reply; 70+ messages in thread
From: Dean Rasheed @ 2016-10-04 09:21 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 4 October 2016 at 09:15, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
> However, for tables and views, each object you store in those views is a
> "table" or "view", but with this thing, the object you store is
> "statistics". Would you have a catalog table called "pg_scissor"?
>
No, probably not (unless it was storing individual scissor blades).
However, in this case, we have related pre-existing catalog tables, so...
> We call the current system table "pg_statistic", though. I agree we should
> call it pg_mv_statistic, in singular, to follow the example of pg_statistic.
>
> Of course, the user-friendly system view on top of that is called
> "pg_stats", just to confuse things more :-).
>
I agree. Given where we are, with a pg_statistic table and a pg_stats
view, I think the least worst solution is to have a pg_statistic_ext
table, and then maybe a pg_stats_ext view.
>> It doesn't seem like the end of the world that it doesn't
>> match the user-facing syntax. A bigger concern is the use of "mv" in
>> the name, because as has already been pointed out, this table may also
>> in the future be used to store univariate expression and partial
>> statistics, so I think we should drop the "mv" and go with something
>> like pg_statistic_ext, or some other more general name.
>
>
> Also, "mv" makes me think of materialized views, which is completely
> unrelated to this.
>
Yeah, I hadn't thought of that.
Regards,
Dean
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-10-11 03:39 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Dean Rasheed <dean.a.rasheed@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Tomas Vondra @ 2016-10-11 03:39 UTC (permalink / raw)
To: Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Hi everyone,
thanks for the reviews. Let me sum the feedback so far, and outline my
plans for the next patch version that I'd like to submit for CF 2016-11.
1) syntax changes
I agree with the changes proposed by Dean, although only a subset of the
syntax is going to be supported until we add support for either join or
partial statistics. So something like this:
CREATE STATISTICS name
[ WITH (options) ]
ON (column1, column2 [, ...])
FROM table
That should be a difficult change.
2) catalog names
I'm not sure what are the best names, so I'm fine with using whatever is
the consensus.
That being said, I'm not sure I like extending the catalog to also
support non-multivariate statistics (like for example statistics on
expressions). While that would be a clearly useful feature, it seems
like a slightly different use case and perhaps a separate catalog would
be better. So maybe pg_statistic_ext is not the best name.
3) special data type(s) to store statistics
I agree using an opaque bytea value is not very nice. I see Heikki
proposed using something like pg_node_tree, and maybe storing all the
statistics in a single value.
I assume the pg_node_tree was meant only as an inspiration how to build
pseudo-type on top of a varlena value. I agree that's a good idea, and I
plan to do something like that - say adding pg_mcv, pg_histogram,
pg_ndistinct and pg_dependencies data types.
Heikki also mentioned that maybe JSONB would be a good way to store the
statistics. I don't think so - firstly, it only supports a subset of
data types, so we'd be unable to store statistics for some data types
(or we'd have to store them as text, which sucks). Also, there's a fair
amount of smartness in how the statistics are stored (e.g. how the
histogram bucket boundaries are deduplicated, or how the estimation uses
the serialized representation directly). We'd lose all of that when
using JSONB.
Similarly for storing all the statistics in a single value - I see no
reason why keeping the statistics in separate columns would be a bad
idea (after all, that's kinda the point of relational databases). Also,
there are perfectly valid cases when the caller only needs a particular
type of statistic - e.g. when estimating GROUP BY we'll only need the
ndistinct coefficients. Why should we force the caller to fetch and
detoast everything, and throw away probably 99% of that?
So my plan here is to define pseudo types similar to how pg_node_tree is
defined. That does not seem like a tremendous amount of work.
4) functional dependencies
Several people mentioned they don't like how functional dependencies are
detected at ANALYZE time, particularly that there's a sudden jump
between 0 and 1. Instead, a continuous "dependency degree" between 0 and
1 was proposed.
I'm fine with that, although that makes "clause reduction" (deciding
that we don't need to estimate one of the clauses at all, as it's
implied by some other clause) impossible. But that's fine, the
functional dependencies will still be much less expensive than the other
statistics.
I'm wondering how will this interact with transitivity, though. IIRC the
current implementation is able to detect transitive dependencies and use
that to reduce storage space etc.
In any case, this significantly complicates the functional dependencies,
which were meant as a trivial type of statistics, mostly to establish
the shared infrastructure. Which brings me to ndistinct.
5) ndistinct
So far, the ndistinct coefficients were lumped at the very end of the
patch, and the statistic was only built but not used for any sort of
estimation. I agree with Dean that perhaps it'd be better to move this
to the very beginning, and use it as the simplest statistic to build the
infrastructure instead of functional dependencies (which only gets truer
due to the changes in functional dependencies, discussed in the
preceding section).
I think it's probably a good idea and I plan to do that, so the patch
series will probably look like this:
* 001 - CREATE STATISTICS infrastucture with ndistinct coefficients
* 002 - use ndistinct coefficients to improve GROUP BY estimates
* 003 - use ndistinct coefficients in clausesel.c (not sure)
* 004 - add functional dependencies (build + clausesel.c)
* 005 - add multivariate MCV (build + clausesel.c)
* 006 - add multivariate histograms (build + clausesel.c)
I'm not sure about using the ndistinct coefficients in clausesel.c to
estimate regular conditions - it's the place for which ndistinct
coefficients were originally proposed by Kyotaro-san, but I seem to
remember it was non-trivial to choose the best statistics when there
were other types of stats available. But I'll look into that.
6) combining statistics
I've decided not to re-submit this part of the patch until the basic
functionality gets in. I do think it's a very useful feature (despite
having my doubts about the existing implementation), but it clearly
distracts people.
Instead, the patch will use some simple selection strategy (e.g. using a
single statistics covering most conditions) or perhaps something more
advanced (e.g. non-overlapping statistics). But nothing complicated.
7) enriching the query plan
Sadly, none of the reviews provides any sort of feedback on how to
enrich the query plan with information about statistics (instead of
doing that in clausesel.c in ad-hoc ephemeral manner).
So I'm still a bit stuck on this :-(
8) join statistics
Not directly related to the current patch, but I recommend reading this
paper quantifying impact of each part of query optimizer (estimates,
cost model, plan enumeration):
http://www.vldb.org/pvldb/vol9/p204-leis.pdf
The one conclusion that I take from it is we really need to think about
improving the join estimates, somehow. Because it's by far the most
significant source of issues (and the hardest one to fix).
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-10-29 19:23 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
0 siblings, 1 reply; 70+ messages in thread
From: Tomas Vondra @ 2016-10-29 19:23 UTC (permalink / raw)
To: Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Hi,
Attached is v20 of the multivariate statistics patch series, doing
mostly the changes outlined in the preceding e-mail from October 11.
The patch series currently has these parts:
* 0001 : (FIX) teach pull_varno about RestrictInfo
* 0002 : (PATCH) shared infrastructure and ndistinct coefficients
* 0003 : (PATCH) functional dependencies (only the ANALYZE part)
* 0004 : (PATCH) selectivity estimation using functional dependencies
* 0005 : (PATCH) multivariate MCV lists
* 0006 : (PATCH) multivariate histograms
* 0007 : (WIP) selectivity estimation using ndistinct coefficients
* 0008 : (WIP) use multiple statistics for estimation
* 0009 : (WIP) psql tab completion basics
Let me elaborate about the main changes in this version:
1) rework CREATE STATISTICS to what Dean Rasheed proposed in [1]:
-----------------------------------------------------------------------
CREATE STATISTICS name WITH (options) ON (columns) FROM table
This allows adding support for statistics on joins, expressions
referencing multiple tables, and partial statistics (with WHERE
predicates, similar to indexes). Although those things are not
implemented (and I don't know if/when that happens), it's good the
syntax supports it.
I've been thinking about using "CREATE STATISTIC" instead, but I decided
to stick with "STATISTICS" for two reasons. Firstly it's possible to
create multiple statistics in a single command, for example by using
WITH (mcv,histogram). And secondly, we already hava "ALTER TABLE ... SET
STATISTICS n" (although that tweaks the statistics target for a column,
not the statistics on the column).
2) no changes to catalog names
-----------------------------------------------------------------------
Clearly, naming things is one of the hardest things in computer science.
I don't have a good idea what names would be better than the current
ones. In any case, this is fairly trivial to do.
3) special data types for statistics
-----------------------------------------------------------------------
Heikki proposed to invent a new data type, similar to pg_node_tree. I do
agree that storing the stats in plain bytea (i.e. catalog having bytea
columns) was not particularly convenient, but I'm not sure how much of
pg_node_tree Heikki wanted to copy.
In particular, I'm not sure whether Heikki's idea was store all the
statistics together in a single Datum, serialized into a text string
(similar to pg_node_tree).
I don't think that would be a good idea, as the statistics may be quite
large and complex, and deserializing them from text format would be
quite expensive. For pg_node_tree that's not a major issue because the
values are usually fairly small. Similarly, packing everything into a
single datum would force the planner to parse/unpack everything, even if
it needs just a small piece (e.g. the ndistinct coefficients, but not
histograms).
So I've decided to invent new data types, one for each statistic type:
* pg_ndistinct
* pg_dependencies
* pg_mcv_list
* pg_histogram
Similarly to pg_node_tree those data types only support output, i.e.
both 'recv' and 'in' functions do elog(ERROR). But while pg_node_tree is
stored as text, those new data types are still bytea.
I do believe this is a good solution, and it allows casting the data
types to text easily, as it simply calls the out function.
The statistics however do not store attnums in the bytea, just indexes
into pg_mv_statistic.stakeys. That means the out functions can't print
column names in the output, or values (because without the attnum we
don't know the type, and thus can't lookup the proper out function).
I don't think there's a good solution for that (I was thinking about
storing the attnums/typeoid in the statistics itself, but that seems
fairly ugly). And I'm quite happy with those new data types.
4) replace functional dependencies with ndistinct (in the first patch)
-----------------------------------------------------------------------
As the ndistinct coeffients are simpler than functional dependencies,
I've decided to use them in the fist patch in the series, which
implements the shared infrastructure. This does not mean throwing away
functional dependencies entirely, just moving them to a later patch.
5) rework of ndistinct coefficients
-----------------------------------------------------------------------
The ndistinct coefficients were also significantly reworked. Instead of
computing and storing the value for the exact combination of attributes,
the new version computes ndistinct for all combinations of attributes.
So for example with CREATE STATISTICS x ON (a,b,c) the old patch only
computed ndistinct on (a,b,c), while the new patch computes ndistinct on
{(a,b,c), (a,b), (a,c), (b,c)}. This makes it way more powerful.
The first patch (0002) only uses this in estimate_num_groups to improve
GROUP BY estimates. A later patch (0007) shows how it might be used for
selectivity estimation, but it's a very early WIP at this point.
Also, I'm not sure we should use ndistinct coefficients this way,
because of the "homogenity" assumption, similarly to functional
dependencies. Functional dependencies are used only for selectivity
estimation, so it's quite easy not to use them if they don't work for
that purpose. But ndistinct coefficients are also used for GROUP BY
estimation, where the homogenity assumption is not such a big deal. So I
expect people to add ndistinct, get better GROUP BY estimates but
sometimes worse selectivity estimates - not great, I guess.
But the selectivity estimation using ndistinct coefficients is very
simple right now - in particular it does not use the per-clause
selectivities at all, it simply assumes the whole selectivity is
1/ndistinct for the combination of columns.
Functional dependencies use this formula to combine the selectivities:
P(a,b) = P(a) * [f + (1-f)*P(b)]
so maybe there's something similar for ndistinct coefficients? I mean,
let's say we know ndistinc(a), ndistinct(b), ndistinct(a,b) and P(a)
and P(b). How do we compute P(a,b)?
5) rework functional dependencies
-----------------------------------------------------------------------
Based on Dean's feedback, I've reworked functional dependencies to use
continuous "degree" of validity (instead of true/false behavior,
resulting in sudden changes in behavior).
This significantly reduced the amount of code, because the old patch
tried to identify transitive dependencies (to minimize time and storage
requirements). Switching to continuous degree makes this impossible (or
at least far more complicated), so I've simply ripped all of this out.
This means the statistics will be larger and ANALYZE will take more
time, the differences are fairly small in practice, and the estimation
actually seems to work better.
6) MCV and histogram changes
-----------------------------------------------------------------------
Those statistics types are mostly unchanged, except for a few minor bug
fixes and removal of remove max_mcv_items and max_buckets options.
Those options were meant to allow users to limit the size of the
statistics, but the implementation was ignoring them so far. So I've
ripped them out, and if needed we may reintroduce them later.
7) no more (elaborate) combinations of statistics
-----------------------------------------------------------------------
I've ripped out the patch that combined multiple statistics in very
elaborate way - it was overly complex, possibly wrong, but most
importantly it distracted people from the preceding patches. So I've
ripped this out, and instead replaced that with a very simple approach
that allows using multiple statistics on different subsets if the clause
list. So for example
WHERE (a=1) AND (b=1) AND (c=1) AND (d=1)
may benefit from two statistics, one on (a,b) and second on (c,d). It's
very simple approach, but it does the trick for many cases and is better
than "single statistics" limitation.
The 0008 patch is actually very simple, essentially adding just a loop
into the code blocks, so I think it's quite likely this will get merged
into the preceding patches.
8) reduce table sizes used in regression tests
-----------------------------------------------------------------------
Some of the regression tests used quite large tables (with up to 1M
rows), which had two issues - long runtimes and unstability (because the
ANALYZE sample is only 30k rows, so there were sometimes small changes
due to picking a different sample). I've limited the table sizes to 30k
rows.
8) open / unsolved questions
-----------------------------------------------------------------------
The main open question is still whether clausesel.c is the best place to
do all the heavy lifting (particularly matching clauses and statistics,
and deciding which statistics to use). I suspect some of that should be
done elsewhere (earlier in the planning), enriching the query tree
somehow. Then clausesel.c would "only" compute the estimates, and it
would also allow showing the info in EXPLAIN.
I'm not particularly happy with the changes in claselist_selectivity
look right now - there are three almost identical blocks, so this would
deserve some refactoring. But I'd like to get some feedback first.
regards
[1]
https://www.postgresql.org/message-id/CAEZATCUtGR+U5+QTwjHhe9rLG2nguEysHQ5NaqcK=VbJ78VQFA@mail.gmail...
[2]
https://www.postgresql.org/message-id/1c7e4e63-769b-f8ce-f245-85ef4f59fcba%40iki.fi
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
Attachments:
[application/x-compressed-tar] multivariate-stats-v20.tgz (140.6K, ../../61e71067-9461-d785-b4a6-6e8a08996d5f@2ndquadrant.com/2-multivariate-stats-v20.tgz)
download
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-12-12 11:26 Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
0 siblings, 1 reply; 70+ messages in thread
From: Amit Langote @ 2016-12-12 11:26 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Hi Tomas,
On 2016/10/30 4:23, Tomas Vondra wrote:
> Hi,
>
> Attached is v20 of the multivariate statistics patch series, doing mostly
> the changes outlined in the preceding e-mail from October 11.
>
> The patch series currently has these parts:
>
> * 0001 : (FIX) teach pull_varno about RestrictInfo
> * 0002 : (PATCH) shared infrastructure and ndistinct coefficients
> * 0003 : (PATCH) functional dependencies (only the ANALYZE part)
> * 0004 : (PATCH) selectivity estimation using functional dependencies
> * 0005 : (PATCH) multivariate MCV lists
> * 0006 : (PATCH) multivariate histograms
> * 0007 : (WIP) selectivity estimation using ndistinct coefficients
> * 0008 : (WIP) use multiple statistics for estimation
> * 0009 : (WIP) psql tab completion basics
Unfortunately, this failed to compile because of the duplicate_oids error.
Partitioning patch consumed same OIDs as used in this patch.
I will try to read the patches in some more detail, but in the meantime,
here are some comments/nitpicks on the documentation:
No updates to doc/src/sgml/catalogs.sgml?
+ <para>
+ The examples presented in <xref linkend="row-estimation-examples"> used
+ statistics about individual columns to compute selectivity estimates.
+ When estimating conditions on multiple columns, the planner assumes
+ independence and multiplies the selectivities. When the columns are
+ correlated, the independence assumption is violated, and the estimates
+ may be seriously off, resulting in poor plan choices.
+ </para>
The term independence is used in isolation - independence of what?
Independence of the distributions of values in separate columns? Also,
the phrase "seriously off" could perhaps be replaced by more rigorous
terminology; it might be unclear to some readers. Perhaps: wildly
inaccurate, :)
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1;
+ QUERY PLAN
+-------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..170.00 rows=100 width=8) (actual
time=0.031..2.870 rows=100 loops=1)
+ Filter: (a = 1)
+ Rows Removed by Filter: 9900
+ Planning time: 0.092 ms
+ Execution time: 3.103 ms
Is there a reason why examples in "67.2. Multivariate Statistics" (like
the one above) use EXPLAIN ANALYZE, whereas those in "67.1. Row Estimation
Examples" (also, other relevant chapters) uses just EXPLAIN.
+ the final 0.01% estimate. The plan however shows that this results in
+ a significant under-estimate, as the actual number of rows matching the
s/under-estimate/underestimate/g
+ <para>
+ For additional details about multivariate statistics, see
+ <filename>src/backend/utils/mvstats/README.statsc</>. There are additional
+ <literal>README</> for each type of statistics, mentioned in the following
+ sections.
+ </para>
Referring to source tree READMEs seems novel around this portion of the
documentation, but I think not too far away, there are some references.
This is under the VII. Internals chapter anyway, so that might be OK.
In any case, s/README.statsc/README.stats/g
Also, s/additional README/additional READMEs/g (tags omitted for brevity)
+ used in definitions of database normal forms. When simplified, saying
that
+ <literal>b</> is functionally dependent on <literal>a</> means that
Maybe, s/When simplified/In simple terms/g
+ In normalized databases, only functional dependencies on primary keys
+ and super keys are allowed. In practice however many data sets are not
+ fully normalized, for example thanks to intentional denormalization for
+ performance reasons. The table <literal>t</> is an example of a data
+ with functional dependencies. As <literal>a=b</> for all rows in the
+ table, <literal>a</> is functionally dependent on <literal>b</> and
+ <literal>b</> is functionally dependent on <literal>a</literal>.
"super keys" sounds like a new term.
s/for example thanks to/for example, thanks to/g (or due to instead of
thanks to)
How about: s/an example of a data with/an example of a schema with/g
Perhaps, s/a=b/a = b/g (additional white space)
+ Similarly to per-column statistics, multivariate statistics are stored in
I notice that "similar to" is used more often than "similarly to". But
that might be OK.
+ This shows that the statistics is defined on table <structname>t</>,
Perhaps: the statistics is -> the statistics are or the statistic is
+ lists <structfield>attnums</structfield> of the columns (references
+ <structname>pg_attribute</structname>).
While this text may be OK on the catalog description page, it might be
better to expand attnums here as "attribute numbers" dropping the
parenthesized phrase altogether.
+<programlisting>
+SELECT pg_mv_stats_dependencies_show(stadeps)
+ FROM pg_mv_statistic WHERE staname = 's1';
+
+ pg_mv_stats_dependencies_show
+-------------------------------
+ (1) => 2, (2) => 1
+(1 row)
+</programlisting>
Couldn't this somehow show actual column names, instead of attribute numbers?
Will read more later.
Thanks,
Amit
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-12-12 21:50 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>
0 siblings, 3 replies; 70+ messages in thread
From: Tomas Vondra @ 2016-12-12 21:50 UTC (permalink / raw)
To: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Hi Amit,
attached is v21 of the patch series, rebased to current master
(resolving the duplicate OID and a few trivial merge conflicts), and
also fixing some of the issues you reported.
On 12/12/2016 12:26 PM, Amit Langote wrote:
>
> Hi Tomas,
>
> On 2016/10/30 4:23, Tomas Vondra wrote:
>> Hi,
>>
>> Attached is v20 of the multivariate statistics patch series, doing mostly
>> the changes outlined in the preceding e-mail from October 11.
>>
>> The patch series currently has these parts:
>>
>> * 0001 : (FIX) teach pull_varno about RestrictInfo
>> * 0002 : (PATCH) shared infrastructure and ndistinct coefficients
>> * 0003 : (PATCH) functional dependencies (only the ANALYZE part)
>> * 0004 : (PATCH) selectivity estimation using functional dependencies
>> * 0005 : (PATCH) multivariate MCV lists
>> * 0006 : (PATCH) multivariate histograms
>> * 0007 : (WIP) selectivity estimation using ndistinct coefficients
>> * 0008 : (WIP) use multiple statistics for estimation
>> * 0009 : (WIP) psql tab completion basics
>
> Unfortunately, this failed to compile because of the duplicate_oids error.
> Partitioning patch consumed same OIDs as used in this patch.
>
Fixed, should compile fine now (even each patch in the series).
> I will try to read the patches in some more detail, but in the meantime,
> here are some comments/nitpicks on the documentation:
>
> No updates to doc/src/sgml/catalogs.sgml?
>
Good point. I've added a section for the pg_mv_statistic catalog.
> + <para>
> + The examples presented in <xref linkend="row-estimation-examples"> used
> + statistics about individual columns to compute selectivity estimates.
> + When estimating conditions on multiple columns, the planner assumes
> + independence and multiplies the selectivities. When the columns are
> + correlated, the independence assumption is violated, and the estimates
> + may be seriously off, resulting in poor plan choices.
> + </para>
>
> The term independence is used in isolation - independence of what?
> Independence of the distributions of values in separate columns? Also,
> the phrase "seriously off" could perhaps be replaced by more rigorous
> terminology; it might be unclear to some readers. Perhaps: wildly
> inaccurate, :)
>
I've reworded this to "independence of the conditions" and "off by
several orders of magnitude". Hope that's better.
> +<programlisting>
> +EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1;
> + QUERY PLAN
> +-------------------------------------------------------------------------------------------------
> + Seq Scan on t (cost=0.00..170.00 rows=100 width=8) (actual
> time=0.031..2.870 rows=100 loops=1)
> + Filter: (a = 1)
> + Rows Removed by Filter: 9900
> + Planning time: 0.092 ms
> + Execution time: 3.103 ms
>
> Is there a reason why examples in "67.2. Multivariate Statistics" (like
> the one above) use EXPLAIN ANALYZE, whereas those in "67.1. Row Estimation
> Examples" (also, other relevant chapters) uses just EXPLAIN.
>
Yes, the reason is that while 67.1 shows how the optimizer estimates row
counts and constructs the plan (so EXPLAIN is sufficient), 67.2
demonstrates how the estimates are inaccurate with respect to the actual
row counts. Thus the EXPLAIN ANALYZE.
> + the final 0.01% estimate. The plan however shows that this results in
> + a significant under-estimate, as the actual number of rows matching the
>
> s/under-estimate/underestimate/g
>
> + <para>
> + For additional details about multivariate statistics, see
> + <filename>src/backend/utils/mvstats/README.statsc</>. There are additional
> + <literal>README</> for each type of statistics, mentioned in the following
> + sections.
> + </para>
>
> Referring to source tree READMEs seems novel around this portion of the
> documentation, but I think not too far away, there are some references.
> This is under the VII. Internals chapter anyway, so that might be OK.
>
I think the there's a threshold when the detail becomes too detailed for
the sgml docs - say, when it discusses some implementation details, at
which point a README is more appropriate. I don't know if I got it
entirely right with the docs, though, so perhaps some bits may move in
either direction.
> In any case, s/README.statsc/README.stats/g
>
> Also, s/additional README/additional READMEs/g (tags omitted for brevity)
>
> + used in definitions of database normal forms. When simplified, saying
> that
> + <literal>b</> is functionally dependent on <literal>a</> means that
>
Fixed.
> Maybe, s/When simplified/In simple terms/g
>
> + In normalized databases, only functional dependencies on primary keys
> + and super keys are allowed. In practice however many data sets are not
> + fully normalized, for example thanks to intentional denormalization for
> + performance reasons. The table <literal>t</> is an example of a data
> + with functional dependencies. As <literal>a=b</> for all rows in the
> + table, <literal>a</> is functionally dependent on <literal>b</> and
> + <literal>b</> is functionally dependent on <literal>a</literal>.
>
> "super keys" sounds like a new term.
>
Actually no, "super key" is a term defined in normal forms.
> s/for example thanks to/for example, thanks to/g (or due to instead of
> thanks to)
>
> How about: s/an example of a data with/an example of a schema with/g
>
I think "example of data set" is better. Reworded.
> Perhaps, s/a=b/a = b/g (additional white space)
>
> + Similarly to per-column statistics, multivariate statistics are stored in
>
> I notice that "similar to" is used more often than "similarly to". But
> that might be OK.
>
Not sure.
> + This shows that the statistics is defined on table <structname>t</>,
>
> Perhaps: the statistics is -> the statistics are or the statistic is
>
As that paragraph is only about functional dependencies, I think
'statistic is' is more appropriate.
> + lists <structfield>attnums</structfield> of the columns (references
> + <structname>pg_attribute</structname>).
>
> While this text may be OK on the catalog description page, it might be
> better to expand attnums here as "attribute numbers" dropping the
> parenthesized phrase altogether.
>
Not sure. I've reworded it like this:
This shows that the statistic is defined on table <structname>t</>,
<structfield>attnums</structfield> lists attribute numbers of columns
(references <structname>pg_attribute</structname>). It also shows
Does that sound better?
> +<programlisting>
> +SELECT pg_mv_stats_dependencies_show(stadeps)
> + FROM pg_mv_statistic WHERE staname = 's1';
> +
> + pg_mv_stats_dependencies_show
> +-------------------------------
> + (1) => 2, (2) => 1
> +(1 row)
> +</programlisting>
>
> Couldn't this somehow show actual column names, instead of attribute numbers?
>
Yeah, I was thinking about that too. The trouble is that's table-level
metadata, so we don't have that kind of info serialized within the data
type (e.g. because it would not handle column renames etc.).
It might be possible to explicitly pass the table OID as a parameter of
the function, but it seemed a bit ugly to me.
FWIW, as I wrote in this thread, the place where this patch series needs
feedback most desperately is integration into the optimizer. Currently
all the magic happens in clausesel.c and does not leave it.I think it
would be good to move some of that (particularly the choice of
statistics to apply) to an earlier stage, and store the information
within the plan tree itself, so that it's available outside clausesel.c
(e.g. for EXPLAIN - showing which stats were picked seems useful).
I was thinking it might work similarly to the foreign key estimation
patch (100340e2). It might even be more efficient, as the current code
may end repeating the selection of statistics multiple times. But
enriching the plan tree turned out to be way more invasive than I'm
comfortable with (but maybe that'd be OK).
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
Attachments:
[application/x-compressed-tar] multivariate-stats-v21.tgz (142.5K, ../../72eeb3d5-c406-93b0-8ff8-11b31789f683@2ndquadrant.com/2-multivariate-stats-v21.tgz)
download
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-12-30 13:05 Petr Jelinek <petr.jelinek@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
2 siblings, 0 replies; 70+ messages in thread
From: Petr Jelinek @ 2016-12-30 13:05 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 12/12/16 22:50, Tomas Vondra wrote:
>> +<programlisting>
>> +SELECT pg_mv_stats_dependencies_show(stadeps)
>> + FROM pg_mv_statistic WHERE staname = 's1';
>> +
>> + pg_mv_stats_dependencies_show
>> +-------------------------------
>> + (1) => 2, (2) => 1
>> +(1 row)
>> +</programlisting>
>>
>> Couldn't this somehow show actual column names, instead of attribute
>> numbers?
>>
>
> Yeah, I was thinking about that too. The trouble is that's table-level
> metadata, so we don't have that kind of info serialized within the data
> type (e.g. because it would not handle column renames etc.).
>
> It might be possible to explicitly pass the table OID as a parameter of
> the function, but it seemed a bit ugly to me.
I think it makes sense to have such function, this is not out function
so I think it's ok for it to have the oid as input, especially since in
the use-case shown above you can use starelid easily.
>
> FWIW, as I wrote in this thread, the place where this patch series needs
> feedback most desperately is integration into the optimizer. Currently
> all the magic happens in clausesel.c and does not leave it.I think it
> would be good to move some of that (particularly the choice of
> statistics to apply) to an earlier stage, and store the information
> within the plan tree itself, so that it's available outside clausesel.c
> (e.g. for EXPLAIN - showing which stats were picked seems useful).
>
> I was thinking it might work similarly to the foreign key estimation
> patch (100340e2). It might even be more efficient, as the current code
> may end repeating the selection of statistics multiple times. But
> enriching the plan tree turned out to be way more invasive than I'm
> comfortable with (but maybe that'd be OK).
>
In theory it seems like possibly reasonable approach to me, mainly
because mv statistics are user defined objects. I guess we'd have to see
at least some PoC to see how invasive it is. But I ultimately think that
feedback from a committer who is more familiar with planner is needed here.
--
Petr Jelinek http://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2016-12-30 13:12 Petr Jelinek <petr.jelinek@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
2 siblings, 1 reply; 70+ messages in thread
From: Petr Jelinek @ 2016-12-30 13:12 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 12/12/16 22:50, Tomas Vondra wrote:
> On 12/12/2016 12:26 PM, Amit Langote wrote:
>>
>> Hi Tomas,
>>
>> On 2016/10/30 4:23, Tomas Vondra wrote:
>>> Hi,
>>>
>>> Attached is v20 of the multivariate statistics patch series, doing
>>> mostly
>>> the changes outlined in the preceding e-mail from October 11.
>>>
>>> The patch series currently has these parts:
>>>
>>> * 0001 : (FIX) teach pull_varno about RestrictInfo
>>> * 0002 : (PATCH) shared infrastructure and ndistinct coefficients
Hi,
I went over these two (IMHO those could easily be considered as minimal
committable set even if the user visible functionality they provide is
rather limited).
> dropping statistics
> -------------------
>
> The statistics may be dropped automatically using DROP STATISTICS.
>
> After ALTER TABLE ... DROP COLUMN, statistics referencing are:
>
> (a) dropped, if the statistics would reference only one column
>
> (b) retained, but modified on the next ANALYZE
This should be documented in user visible form if you plan to keep it
(it does make sense to me).
> + therefore perfectly correlated. Providing additional information about
> + correlation between columns is the purpose of multivariate statistics,
> + and the rest of this section thoroughly explains how the planner
> + leverages them to improve estimates.
> + </para>
> +
> + <para>
> + For additional details about multivariate statistics, see
> + <filename>src/backend/utils/mvstats/README.stats</>. There are additional
> + <literal>READMEs</> for each type of statistics, mentioned in the following
> + sections.
> + </para>
> +
> + </sect1>
I don't think this qualifies as "thoroughly explains" ;)
> +
> +Oid
> +get_statistics_oid(List *names, bool missing_ok)
No comment?
> + case OBJECT_STATISTICS:
> + msg = gettext_noop("statistics \"%s\" does not exist, skipping");
> + name = NameListToString(objname);
> + break;
This sounds somewhat weird (plural vs singular).
> + * XXX Maybe this should check for duplicate stats. Although it's not clear
> + * what "duplicate" would mean here (wheter to compare only keys or also
> + * options). Moreover, we don't do such checks for indexes, although those
> + * store tuples and recreating a new index may be a way to fix bloat (which
> + * is a problem statistics don't have).
> + */
> +ObjectAddress
> +CreateStatistics(CreateStatsStmt *stmt)
I don't think we should check duplicates TBH so I would remove the XXX
(also "wheter" is typo but if you remove that paragraph it does not matter).
> + if (true)
> + {
Huh?
> +
> +List *
> +RelationGetMVStatList(Relation relation)
> +{
...
> +
> +void
> +update_mv_stats(Oid mvoid, MVNDistinct ndistinct,
> + int2vector *attrs, VacAttrStats **stats)
...
> +static double
> +ndistinct_for_combination(double totalrows, int numrows, HeapTuple *rows,
> + int2vector *attrs, VacAttrStats **stats,
> + int k, int *combination)
> +{
Again, these deserve comment.
I'll try to look at other patches in the series as time permits.
--
Petr Jelinek http://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-03 13:42 Dilip Kumar <dilipbalaut@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
2 siblings, 1 reply; 70+ messages in thread
From: Dilip Kumar @ 2017-01-03 13:42 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Tue, Dec 13, 2016 at 3:20 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> attached is v21 of the patch series, rebased to current master (resolving
> the duplicate OID and a few trivial merge conflicts), and also fixing some
> of the issues you reported.
I wanted to test the grouping estimation behaviour with TPCH, While
testing I found some crash so I thought of reporting it.
My setup detail:
TPCH scale factor : 5
Applied all the patch for 21 series, and ran below queries.
postgres=# analyze part;
ANALYZE
postgres=# CREATE STATISTICS s2 WITH (ndistinct) on (p_brand, p_type,
p_size) from part;
CREATE STATISTICS
postgres=# analyze part;
server closed the connection unexpectedly
This probably means the server terminated abnormally
before or while processing the request.
The connection to the server was lost. Attempting reset: Failed.
I think it should be easily reproducible, in case it's not I can send
call stack or core dump.
--
Regards,
Dilip Kumar
EnterpriseDB: http://www.enterprisedb.com
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-03 16:22 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Dilip Kumar <dilipbalaut@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Tomas Vondra @ 2017-01-03 16:22 UTC (permalink / raw)
To: Dilip Kumar <dilipbalaut@gmail.com>; +Cc: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 01/03/2017 02:42 PM, Dilip Kumar wrote:
> On Tue, Dec 13, 2016 at 3:20 AM, Tomas Vondra
> <tomas.vondra@2ndquadrant.com> wrote:
>> attached is v21 of the patch series, rebased to current master (resolving
>> the duplicate OID and a few trivial merge conflicts), and also fixing some
>> of the issues you reported.
>
> I wanted to test the grouping estimation behaviour with TPCH, While
> testing I found some crash so I thought of reporting it.
>
> My setup detail:
> TPCH scale factor : 5
> Applied all the patch for 21 series, and ran below queries.
>
> postgres=# analyze part;
> ANALYZE
> postgres=# CREATE STATISTICS s2 WITH (ndistinct) on (p_brand, p_type,
> p_size) from part;
> CREATE STATISTICS
> postgres=# analyze part;
> server closed the connection unexpectedly
> This probably means the server terminated abnormally
> before or while processing the request.
> The connection to the server was lost. Attempting reset: Failed.
>
> I think it should be easily reproducible, in case it's not I can send
> call stack or core dump.
>
Thanks for the report. It was trivial to reproduce and it turned out to
be a fairly simple bug. Will send a new version of the patch soon.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-03 21:55 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Petr Jelinek <petr.jelinek@2ndquadrant.com>
0 siblings, 0 replies; 70+ messages in thread
From: Tomas Vondra @ 2017-01-03 21:55 UTC (permalink / raw)
To: Petr Jelinek <petr.jelinek@2ndquadrant.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 12/30/2016 02:12 PM, Petr Jelinek wrote:
> On 12/12/16 22:50, Tomas Vondra wrote:
>> On 12/12/2016 12:26 PM, Amit Langote wrote:
>>>
>>> Hi Tomas,
>>>
>>> On 2016/10/30 4:23, Tomas Vondra wrote:
>>>> Hi,
>>>>
>>>> Attached is v20 of the multivariate statistics patch series, doing
>>>> mostly
>>>> the changes outlined in the preceding e-mail from October 11.
>>>>
>>>> The patch series currently has these parts:
>>>>
>>>> * 0001 : (FIX) teach pull_varno about RestrictInfo
>>>> * 0002 : (PATCH) shared infrastructure and ndistinct coefficients
>
> Hi,
>
> I went over these two (IMHO those could easily be considered as minimal
> committable set even if the user visible functionality they provide is
> rather limited).
>
Yes, although I still have my doubts 0001 is the right way to make
pull_varnos work. It's probably related to the bigger design question,
because moving the statistics selection to an earlier phase could make
it unnecessary I guess.
>> dropping statistics
>> -------------------
>>
>> The statistics may be dropped automatically using DROP STATISTICS.
>>
>> After ALTER TABLE ... DROP COLUMN, statistics referencing are:
>>
>> (a) dropped, if the statistics would reference only one column
>>
>> (b) retained, but modified on the next ANALYZE
>
> This should be documented in user visible form if you plan to keep it
> (it does make sense to me).
>
Yes, I plan to keep it. I agree it should be documented, probably on the
ALTER TABLE page (and linked from CREATE/DROP statistics pages).
>> + therefore perfectly correlated. Providing additional information about
>> + correlation between columns is the purpose of multivariate statistics,
>> + and the rest of this section thoroughly explains how the planner
>> + leverages them to improve estimates.
>> + </para>
>> +
>> + <para>
>> + For additional details about multivariate statistics, see
>> + <filename>src/backend/utils/mvstats/README.stats</>. There are additional
>> + <literal>READMEs</> for each type of statistics, mentioned in the following
>> + sections.
>> + </para>
>> +
>> + </sect1>
>
> I don't think this qualifies as "thoroughly explains" ;)
>
OK, I'll drop the "thoroughly" ;-)
>> +
>> +Oid
>> +get_statistics_oid(List *names, bool missing_ok)
>
> No comment?
>
>> + case OBJECT_STATISTICS:
>> + msg = gettext_noop("statistics \"%s\" does not exist, skipping");
>> + name = NameListToString(objname);
>> + break;
>
> This sounds somewhat weird (plural vs singular).
>
Ah, right - it should be either "statistic ... does not" or "statistics
... do not". I think "statistics" is the right choice here, because (a)
we have CREATE STATISTICS and (b) it may be a combination of statistics,
e.g. histogram + MCV.
>> + * XXX Maybe this should check for duplicate stats. Although it's not clear
>> + * what "duplicate" would mean here (wheter to compare only keys or also
>> + * options). Moreover, we don't do such checks for indexes, although those
>> + * store tuples and recreating a new index may be a way to fix bloat (which
>> + * is a problem statistics don't have).
>> + */
>> +ObjectAddress
>> +CreateStatistics(CreateStatsStmt *stmt)
>
> I don't think we should check duplicates TBH so I would remove the XXX
> (also "wheter" is typo but if you remove that paragraph it does not matter).
>
Yes, I came to the same conclusion - we can only really check for exact
matches (same set of columns, same choice of statistic types), but
that's fairly useless. I'll remove the XXX.
>> + if (true)
>> + {
>
> Huh?
>
Yeah, that's a bit weird pattern. It's a remainder of copy-pasting the
preceding block, which looks like this
if (hasindex)
{
...
}
But we've decided to not add similar flag for the statistics. I'll move
the block to a separate function (instead of merging it directly into
the function, which is already a bit largeish).
>> +
>> +List *
>> +RelationGetMVStatList(Relation relation)
>> +{
> ...
>> +
>> +void
>> +update_mv_stats(Oid mvoid, MVNDistinct ndistinct,
>> + int2vector *attrs, VacAttrStats **stats)
> ...
>> +static double
>> +ndistinct_for_combination(double totalrows, int numrows, HeapTuple *rows,
>> + int2vector *attrs, VacAttrStats **stats,
>> + int k, int *combination)
>> +{
>
>
> Again, these deserve comment.
>
OK, will add.
> I'll try to look at other patches in the series as time permits.
thanks
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-04 02:35 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
0 siblings, 4 replies; 70+ messages in thread
From: Tomas Vondra @ 2017-01-04 02:35 UTC (permalink / raw)
To: Dilip Kumar <dilipbalaut@gmail.com>; +Cc: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 01/03/2017 05:22 PM, Tomas Vondra wrote:
> On 01/03/2017 02:42 PM, Dilip Kumar wrote:
...
>> I think it should be easily reproducible, in case it's not I can send
>> call stack or core dump.
>>
>
> Thanks for the report. It was trivial to reproduce and it turned out to
> be a fairly simple bug. Will send a new version of the patch soon.
>
Attached is v22 of the patch series, rebased to current master and
fixing the reported bug. I haven't made any other changes - the issues
reported by Petr are mostly minor, so I've decided to wait a bit more
for (hopefully) other reviews.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
Attachments:
[binary/octet-stream] 0001-teach-pull_-varno-varattno-_walker-about-Restric-v22.patch (1.4K, ../../696aa95c-2411-9b2b-f36e-65b66bf47c88@2ndquadrant.com/2-0001-teach-pull_-varno-varattno-_walker-about-Restric-v22.patch)
download | inline diff:
From d242a85fd3d21a48c2e3b01dc7cfec92d2a40268 Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:35:05 +0200
Subject: [PATCH 1/9] teach pull_(varno|varattno)_walker about RestrictInfo
otherwise pull_varnos fails when processing OR clauses
---
src/backend/optimizer/util/var.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
diff --git a/src/backend/optimizer/util/var.c b/src/backend/optimizer/util/var.c
index 292e1f4..9228a46 100644
--- a/src/backend/optimizer/util/var.c
+++ b/src/backend/optimizer/util/var.c
@@ -196,6 +196,13 @@ pull_varnos_walker(Node *node, pull_varnos_context *context)
context->sublevels_up--;
return result;
}
+ if (IsA(node, RestrictInfo))
+ {
+ RestrictInfo *rinfo = (RestrictInfo*)node;
+ context->varnos = bms_add_members(context->varnos,
+ rinfo->clause_relids);
+ return false;
+ }
return expression_tree_walker(node, pull_varnos_walker,
(void *) context);
}
@@ -244,6 +251,15 @@ pull_varattnos_walker(Node *node, pull_varattnos_context *context)
return false;
}
+ if (IsA(node, RestrictInfo))
+ {
+ RestrictInfo *rinfo = (RestrictInfo *)node;
+
+ return expression_tree_walker((Node*)rinfo->clause,
+ pull_varattnos_walker,
+ (void*) context);
+ }
+
/* Should not find an unplanned subquery */
Assert(!IsA(node, Query));
--
2.5.5
[binary/octet-stream] 0002-PATCH-shared-infrastructure-and-ndistinct-coeffi-v22.patch (138.6K, ../../696aa95c-2411-9b2b-f36e-65b66bf47c88@2ndquadrant.com/3-0002-PATCH-shared-infrastructure-and-ndistinct-coeffi-v22.patch)
download | inline diff:
From 1d633082cbc2d62a0fa412c4417331aaf0331772 Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:35:47 +0200
Subject: [PATCH 2/9] PATCH: shared infrastructure and ndistinct coefficients
Basic infrastructure shared by all kinds of multivariate stats, most
importantly:
- adds a new system catalog (pg_mv_statistic)
- CREATE STATISTICS name ON (columns) FROM table
- DROP STATISTICS name
- ALTER STATISTICS ... OWNER TO / SET SCHEMA / RENAME
- implementation of ndistinct coefficients (the simplest type of
multivariate statistics)
- computing ndistinct coefficients during ANALYZE
- updates existing regression tests (new catalog etc.)
- modifies estimate_num_groups() to use ndistinct if available
The current implementation requires a valid 'ltopr' for the columns, so
that we can sort the sample rows in various ways, both in this patch
and other kinds of statistics. Maybe this restriction could be relaxed
in the future, requiring just 'eqopr' in case of stats not sorting the
data (e.g. functional dependencies and MCV lists).
Some of the stats implemented in follow-up patches (e.g. functional
dependencies and MCV list with limited functionality) might be made
to work with hashes of the values. That would save a lot of space
for storing the statistics, and it would be sufficient for estimating
equality conditions.
creating statistics
-------------------
Statistics are created by CREATE STATISTICS command, with this syntax:
CREATE STATISTICS statistics_name ON (columns) FROM table
where 'statistics_name' may be a fully-qualified name (i.e. specifying
a schema). It's expected that we'll eventually add support for join
statistics, referencing tables that may be located in different schemas,
so we can't make the name unique per-table (like constraints), and we
can't just pick one of the table schemas.
dropping statistics
-------------------
The statistics may be dropped automatically using DROP STATISTICS.
After ALTER TABLE ... DROP COLUMN, statistics referencing are:
(a) dropped, if the statistics would reference only one column
(b) retained, but modified on the next ANALYZE
The goal of the lazy cleanup is not to disrupt the optimizer, but
arguably this is over-engineering and it might also made work just
like for indexes by simply dropping all dependent statistics on
ALTER TABLE ... DROP COLUMN. If the user wants to minimize impact,
the smaller statistics needs to be created explicitly in advance.
This also adds a simple list of statistics to \d in psql.
ndistinct coefficients
----------------------
The patch only implements a very simple type of statistics, tracking
the number of groups for different combinations of columns. For
example given columns (a,b,c) the statistics will estimate number
of distinct combinations of values in (a,b), (a,c), (b,c) and (a,b,c).
This is then used in estimate_num_groups() for estimating cardinality
of GROUP BY and similar clauses.
pg_ndistinct data type
----------------------
The patch introduces pg_ndistinct, a new varlena data type used for
serialized version of ndistinct coefficients. Internally it's just
a bytea value, but it allows us to control casting, input/output
and so on. It's somewhat inspired by pg_node_tree.
---
doc/src/sgml/catalogs.sgml | 108 +++++
doc/src/sgml/planstats.sgml | 141 +++++++
doc/src/sgml/ref/allfiles.sgml | 3 +
doc/src/sgml/ref/alter_statistics.sgml | 115 ++++++
doc/src/sgml/ref/create_statistics.sgml | 153 +++++++
doc/src/sgml/ref/drop_statistics.sgml | 91 +++++
doc/src/sgml/reference.sgml | 3 +
src/backend/catalog/Makefile | 1 +
src/backend/catalog/aclchk.c | 27 ++
src/backend/catalog/dependency.c | 11 +-
src/backend/catalog/heap.c | 101 +++++
src/backend/catalog/namespace.c | 51 +++
src/backend/catalog/objectaddress.c | 54 +++
src/backend/catalog/system_views.sql | 10 +
src/backend/commands/Makefile | 6 +-
src/backend/commands/alter.c | 3 +
src/backend/commands/analyze.c | 8 +
src/backend/commands/dropcmds.c | 4 +
src/backend/commands/event_trigger.c | 3 +
src/backend/commands/statscmds.c | 259 ++++++++++++
src/backend/nodes/copyfuncs.c | 16 +
src/backend/nodes/outfuncs.c | 18 +
src/backend/optimizer/util/plancat.c | 59 +++
src/backend/parser/gram.y | 58 ++-
src/backend/tcop/utility.c | 14 +
src/backend/utils/Makefile | 2 +-
src/backend/utils/adt/selfuncs.c | 168 +++++++-
src/backend/utils/cache/relcache.c | 59 +++
src/backend/utils/cache/syscache.c | 23 ++
src/backend/utils/mvstats/Makefile | 17 +
src/backend/utils/mvstats/README.ndistinct | 22 +
src/backend/utils/mvstats/README.stats | 98 +++++
src/backend/utils/mvstats/common.c | 384 ++++++++++++++++++
src/backend/utils/mvstats/common.h | 80 ++++
src/backend/utils/mvstats/mvdist.c | 585 +++++++++++++++++++++++++++
src/bin/psql/describe.c | 44 ++
src/include/catalog/dependency.h | 5 +-
src/include/catalog/heap.h | 1 +
src/include/catalog/indexing.h | 7 +
src/include/catalog/namespace.h | 2 +
src/include/catalog/pg_cast.h | 4 +
src/include/catalog/pg_mv_statistic.h | 78 ++++
src/include/catalog/pg_proc.h | 9 +
src/include/catalog/pg_type.h | 4 +
src/include/catalog/toasting.h | 1 +
src/include/commands/defrem.h | 4 +
src/include/nodes/nodes.h | 2 +
src/include/nodes/parsenodes.h | 11 +
src/include/nodes/relation.h | 27 ++
src/include/utils/acl.h | 1 +
src/include/utils/builtins.h | 4 +
src/include/utils/mvstats.h | 60 +++
src/include/utils/rel.h | 4 +
src/include/utils/relcache.h | 1 +
src/include/utils/syscache.h | 2 +
src/test/regress/expected/mv_ndistinct.out | 117 ++++++
src/test/regress/expected/object_address.out | 7 +-
src/test/regress/expected/opr_sanity.out | 3 +-
src/test/regress/expected/rules.out | 8 +
src/test/regress/expected/sanity_check.out | 1 +
src/test/regress/expected/type_sanity.out | 13 +-
src/test/regress/parallel_schedule | 3 +
src/test/regress/serial_schedule | 1 +
src/test/regress/sql/mv_ndistinct.sql | 68 ++++
src/test/regress/sql/object_address.sql | 4 +-
65 files changed, 3232 insertions(+), 19 deletions(-)
create mode 100644 doc/src/sgml/ref/alter_statistics.sgml
create mode 100644 doc/src/sgml/ref/create_statistics.sgml
create mode 100644 doc/src/sgml/ref/drop_statistics.sgml
create mode 100644 src/backend/commands/statscmds.c
create mode 100644 src/backend/utils/mvstats/Makefile
create mode 100644 src/backend/utils/mvstats/README.ndistinct
create mode 100644 src/backend/utils/mvstats/README.stats
create mode 100644 src/backend/utils/mvstats/common.c
create mode 100644 src/backend/utils/mvstats/common.h
create mode 100644 src/backend/utils/mvstats/mvdist.c
create mode 100644 src/include/catalog/pg_mv_statistic.h
create mode 100644 src/include/utils/mvstats.h
create mode 100644 src/test/regress/expected/mv_ndistinct.out
create mode 100644 src/test/regress/sql/mv_ndistinct.sql
diff --git a/doc/src/sgml/catalogs.sgml b/doc/src/sgml/catalogs.sgml
index 4930506..b39ca69 100644
--- a/doc/src/sgml/catalogs.sgml
+++ b/doc/src/sgml/catalogs.sgml
@@ -201,6 +201,11 @@
</row>
<row>
+ <entry><link linkend="catalog-pg-mv-statistic"><structname>pg_mv_statistic</structname></link></entry>
+ <entry>multivariate statistics</entry>
+ </row>
+
+ <row>
<entry><link linkend="catalog-pg-namespace"><structname>pg_namespace</structname></link></entry>
<entry>schemas</entry>
</row>
@@ -4196,6 +4201,109 @@
</table>
</sect1>
+ <sect1 id="catalog-pg-mv-statistic">
+ <title><structname>pg_mv_statistic</structname></title>
+
+ <indexterm zone="catalog-pg-mv-statistic">
+ <primary>pg_mv_statistic</primary>
+ </indexterm>
+
+ <para>
+ The catalog <structname>pg_mv_statistic</structname>
+ holds multivariate statistics about combinations of columns.
+ </para>
+
+ <table>
+ <title><structname>pg_mv_statistic</> Columns</title>
+
+ <tgroup cols="4">
+ <thead>
+ <row>
+ <entry>Name</entry>
+ <entry>Type</entry>
+ <entry>References</entry>
+ <entry>Description</entry>
+ </row>
+ </thead>
+
+ <tbody>
+
+ <row>
+ <entry><structfield>starelid</structfield></entry>
+ <entry><type>oid</type></entry>
+ <entry><literal><link linkend="catalog-pg-class"><structname>pg_class</structname></link>.oid</literal></entry>
+ <entry>The table that the described columns belongs to</entry>
+ </row>
+
+ <row>
+ <entry><structfield>staname</structfield></entry>
+ <entry><type>name</type></entry>
+ <entry></entry>
+ <entry>Name of the statistic.</entry>
+ </row>
+
+ <row>
+ <entry><structfield>stanamespace</structfield></entry>
+ <entry><type>oid</type></entry>
+ <entry><literal><link linkend="catalog-pg-namespace"><structname>pg_namespace</structname></link>.oid</literal></entry>
+ <entry>
+ The OID of the namespace that contains this statistic
+ </entry>
+ </row>
+
+ <row>
+ <entry><structfield>staowner</structfield></entry>
+ <entry><type>oid</type></entry>
+ <entry><literal><link linkend="catalog-pg-authid"><structname>pg_authid</structname></link>.oid</literal></entry>
+ <entry>Owner of the statistic</entry>
+ </row>
+
+ <row>
+ <entry><structfield>ndist_enabled</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, ndistinct coefficients will be computed for the combination of
+ columns, covered by the statistics. This does not mean the coefficients
+ are already computed, though.
+ </entry>
+ </row>
+
+ <row>
+ <entry><structfield>ndist_built</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, ndistinct coefficients are already computed and available for
+ use during query estimation.
+ </entry>
+ </row>
+
+ <row>
+ <entry><structfield>stakeys</structfield></entry>
+ <entry><type>int2vector</type></entry>
+ <entry><literal><link linkend="catalog-pg-attribute"><structname>pg_attribute</structname></link>.attnum</literal></entry>
+ <entry>
+ This is an array of values that indicate which table columns this
+ statistic covers. For example a value of <literal>1 3</literal> would
+ mean that the first and the third table columns make up the statistic key.
+ </entry>
+ </row>
+
+ <row>
+ <entry><structfield>standist</structfield></entry>
+ <entry><type>pg_ndistinct</type></entry>
+ <entry></entry>
+ <entry>
+ Ndistict coefficients, serialized as <structname>pg_ndistinct</> type.
+ </entry>
+ </row>
+
+ </tbody>
+ </tgroup>
+ </table>
+ </sect1>
+
<sect1 id="catalog-pg-namespace">
<title><structname>pg_namespace</structname></title>
diff --git a/doc/src/sgml/planstats.sgml b/doc/src/sgml/planstats.sgml
index 1a482d3..e9248b4 100644
--- a/doc/src/sgml/planstats.sgml
+++ b/doc/src/sgml/planstats.sgml
@@ -448,4 +448,145 @@ rows = (outer_cardinality * inner_cardinality) * selectivity
</sect1>
+ <sect1 id="multivariate-statistics">
+ <title>Multivariate Statistics</title>
+
+ <indexterm zone="multivariate-statistics">
+ <primary>multivariate statistics</primary>
+ <secondary>planner</secondary>
+ </indexterm>
+
+ <para>
+ The examples presented in <xref linkend="row-estimation-examples"> used
+ statistics about individual columns to compute selectivity estimates.
+ When estimating conditions on multiple columns, the planner assumes
+ independence of the conditions and multiplies the selectivities. When the
+ columns are correlated, the independence assumption is violated, and the
+ estimates may be off by several orders of magnitude, resulting in poor
+ plan choices.
+ </para>
+
+ <para>
+ The examples presented below demonstrate such estimation errors on simple
+ data sets, and also how to resolve them by creating multivariate statistics
+ using <command>CREATE STATISTICS</> command.
+ </para>
+
+ <para>
+ Let's start with a very simple data set - a table with two columns,
+ containing exactly the same values:
+
+<programlisting>
+CREATE TABLE t (a INT, b INT);
+INSERT INTO t SELECT i % 100, i % 100 FROM generate_series(1, 10000) s(i);
+ANALYZE t;
+</programlisting>
+
+ As explained in <xref linkend="planner-stats">, the planner can determine
+ cardinality of <structname>t</structname> using the number of pages and
+ rows is looked up in <structname>pg_class</structname>:
+
+<programlisting>
+SELECT relpages, reltuples FROM pg_class WHERE relname = 't';
+
+ relpages | reltuples
+----------+-----------
+ 45 | 10000
+</programlisting>
+
+ The data distribution is very simple - there are only 100 distinct values
+ in each column, uniformly distributed.
+ </para>
+
+ <para>
+ The following example shows the result of estimating a <literal>WHERE</>
+ condition on the <structfield>a</> column:
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1;
+ QUERY PLAN
+-------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..170.00 rows=100 width=8) (actual time=0.031..2.870 rows=100 loops=1)
+ Filter: (a = 1)
+ Rows Removed by Filter: 9900
+ Planning time: 0.092 ms
+ Execution time: 3.103 ms
+(5 rows)
+</programlisting>
+
+ The planner examines the condition and computes the estimate using
+ <function>eqsel</>, the selectivity function for <literal>=</>, and
+ statistics stored in the <structname>pg_stats</> table. In this case
+ the planner estimates the condition matches 1% rows, and by comparing
+ the estimated and actual number of rows, we see that the estimate is
+ very accurate (in fact exact, as the table is very small).
+ </para>
+
+ <para>
+ Adding a condition on the second column results in the following plan:
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
+ QUERY PLAN
+-----------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=1 width=8) (actual time=0.033..3.006 rows=100 loops=1)
+ Filter: ((a = 1) AND (b = 1))
+ Rows Removed by Filter: 9900
+ Planning time: 0.121 ms
+ Execution time: 3.220 ms
+(5 rows)
+</programlisting>
+
+ The planner estimates the selectivity for each condition individually,
+ arriving to the 1% estimates as above, and then multiplies them, getting
+ the final 0.01% estimate. The plan however shows that this results in
+ a significant underestimate, as the actual number of rows matching the
+ conditions is two orders of magnitude higher than estimated.
+ </para>
+
+ <para>
+ Overestimates, i.e. errors in the opposite direction, are also possible.
+ Consider for example the following combination of range conditions, each
+ matching
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a <= 49 AND b > 49;
+ QUERY PLAN
+------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=2500 width=8) (actual time=1.607..1.607 rows=0 loops=1)
+ Filter: ((a <= 49) AND (b > 49))
+ Rows Removed by Filter: 10000
+ Planning time: 0.050 ms
+ Execution time: 1.623 ms
+(5 rows)
+</programlisting>
+
+ The planner examines both <literal>WHERE</> clauses and estimates them
+ using the <function>scalarltsel</> and <function>scalargtsel</> functions,
+ specified as the selectivity functions matching the <literal><=</> and
+ <literal>></literal> operators. Both conditions match 50% of the
+ table, and assuming independence the planner multiplies them to compute
+ the total estimate of 25%. However as the explain output shows, the actual
+ number of rows is 0, because the columns are correlated and the conditions
+ contradict each other.
+ </para>
+
+ <para>
+ Both estimation errors are caused by violation of the independence
+ assumption, as the two columns contain exactly the same values, and are
+ therefore perfectly correlated. Providing additional information about
+ correlation between columns is the purpose of multivariate statistics,
+ and the rest of this section thoroughly explains how the planner
+ leverages them to improve estimates.
+ </para>
+
+ <para>
+ For additional details about multivariate statistics, see
+ <filename>src/backend/utils/mvstats/README.stats</>. There are additional
+ <literal>READMEs</> for each type of statistics, mentioned in the following
+ sections.
+ </para>
+
+ </sect1>
+
</chapter>
diff --git a/doc/src/sgml/ref/allfiles.sgml b/doc/src/sgml/ref/allfiles.sgml
index 77667bd..75977af 100644
--- a/doc/src/sgml/ref/allfiles.sgml
+++ b/doc/src/sgml/ref/allfiles.sgml
@@ -32,6 +32,7 @@ Complete list of usable sgml source files in this directory.
<!ENTITY alterServer SYSTEM "alter_server.sgml">
<!ENTITY alterSequence SYSTEM "alter_sequence.sgml">
<!ENTITY alterSystem SYSTEM "alter_system.sgml">
+<!ENTITY alterStatistics SYSTEM "alter_statistics.sgml">
<!ENTITY alterTable SYSTEM "alter_table.sgml">
<!ENTITY alterTableSpace SYSTEM "alter_tablespace.sgml">
<!ENTITY alterTSConfig SYSTEM "alter_tsconfig.sgml">
@@ -77,6 +78,7 @@ Complete list of usable sgml source files in this directory.
<!ENTITY createSchema SYSTEM "create_schema.sgml">
<!ENTITY createSequence SYSTEM "create_sequence.sgml">
<!ENTITY createServer SYSTEM "create_server.sgml">
+<!ENTITY createStatistics SYSTEM "create_statistics.sgml">
<!ENTITY createTable SYSTEM "create_table.sgml">
<!ENTITY createTableAs SYSTEM "create_table_as.sgml">
<!ENTITY createTableSpace SYSTEM "create_tablespace.sgml">
@@ -121,6 +123,7 @@ Complete list of usable sgml source files in this directory.
<!ENTITY dropSchema SYSTEM "drop_schema.sgml">
<!ENTITY dropSequence SYSTEM "drop_sequence.sgml">
<!ENTITY dropServer SYSTEM "drop_server.sgml">
+<!ENTITY dropStatistics SYSTEM "drop_statistics.sgml">
<!ENTITY dropTable SYSTEM "drop_table.sgml">
<!ENTITY dropTableSpace SYSTEM "drop_tablespace.sgml">
<!ENTITY dropTransform SYSTEM "drop_transform.sgml">
diff --git a/doc/src/sgml/ref/alter_statistics.sgml b/doc/src/sgml/ref/alter_statistics.sgml
new file mode 100644
index 0000000..3f477cb
--- /dev/null
+++ b/doc/src/sgml/ref/alter_statistics.sgml
@@ -0,0 +1,115 @@
+<!--
+doc/src/sgml/ref/alter_statistics.sgml
+PostgreSQL documentation
+-->
+
+<refentry id="SQL-ALTERSTATISTICS">
+ <indexterm zone="sql-alterstatistics">
+ <primary>ALTER STATISTICS</primary>
+ </indexterm>
+
+ <refmeta>
+ <refentrytitle>ALTER STATISTICS</refentrytitle>
+ <manvolnum>7</manvolnum>
+ <refmiscinfo>SQL - Language Statements</refmiscinfo>
+ </refmeta>
+
+ <refnamediv>
+ <refname>ALTER STATISTICS</refname>
+ <refpurpose>
+ change the definition of a multivariate statistics
+ </refpurpose>
+ </refnamediv>
+
+ <refsynopsisdiv>
+<synopsis>
+ALTER STATISTICS <replaceable class="parameter">name</replaceable> OWNER TO { <replaceable class="PARAMETER">new_owner</replaceable> | CURRENT_USER | SESSION_USER }
+ALTER STATISTICS <replaceable class="parameter">name</replaceable> RENAME TO <replaceable class="parameter">new_name</replaceable>
+ALTER STATISTICS <replaceable class="parameter">name</replaceable> SET SCHEMA <replaceable class="parameter">new_schema</replaceable>
+</synopsis>
+ </refsynopsisdiv>
+
+ <refsect1>
+ <title>Description</title>
+
+ <para>
+ <command>ALTER STATISTICS</command> changes the parameters of an existing
+ multivariate statistics. Any parameters not specifically set in the
+ <command>ALTER STATISTICS</command> command retain their prior settings.
+ </para>
+
+ <para>
+ You must own the statistics to use <command>ALTER STATISTICS</>.
+ To change a statistics' schema, you must also have <literal>CREATE</>
+ privilege on the new schema.
+ To alter the owner, you must also be a direct or indirect member of the new
+ owning role, and that role must have <literal>CREATE</literal> privilege on
+ the statistics' schema. (These restrictions enforce that altering the owner
+ doesn't do anything you couldn't do by dropping and recreating the statistics.
+ However, a superuser can alter ownership of any statistics anyway.)
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>Parameters</title>
+
+ <para>
+ <variablelist>
+ <varlistentry>
+ <term><replaceable class="parameter">name</replaceable></term>
+ <listitem>
+ <para>
+ The name (optionally schema-qualified) of a statistics to be altered.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="PARAMETER">new_owner</replaceable></term>
+ <listitem>
+ <para>
+ The user name of the new owner of the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="parameter">new_name</replaceable></term>
+ <listitem>
+ <para>
+ The new name for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="parameter">new_schema</replaceable></term>
+ <listitem>
+ <para>
+ The new schema for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ </variablelist>
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>Compatibility</title>
+
+ <para>
+ There's no <command>ALTER STATISTICS</command> command in the SQL standard.
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>See Also</title>
+
+ <simplelist type="inline">
+ <member><xref linkend="sql-createstatistics"></member>
+ <member><xref linkend="sql-dropstatistics"></member>
+ </simplelist>
+ </refsect1>
+
+</refentry>
diff --git a/doc/src/sgml/ref/create_statistics.sgml b/doc/src/sgml/ref/create_statistics.sgml
new file mode 100644
index 0000000..7fa118c
--- /dev/null
+++ b/doc/src/sgml/ref/create_statistics.sgml
@@ -0,0 +1,153 @@
+<!--
+doc/src/sgml/ref/create_statistics.sgml
+PostgreSQL documentation
+-->
+
+<refentry id="SQL-CREATESTATISTICS">
+ <indexterm zone="sql-createstatistics">
+ <primary>CREATE STATISTICS</primary>
+ </indexterm>
+
+ <refmeta>
+ <refentrytitle>CREATE STATISTICS</refentrytitle>
+ <manvolnum>7</manvolnum>
+ <refmiscinfo>SQL - Language Statements</refmiscinfo>
+ </refmeta>
+
+ <refnamediv>
+ <refname>CREATE STATISTICS</refname>
+ <refpurpose>define a new statistics</refpurpose>
+ </refnamediv>
+
+ <refsynopsisdiv>
+<synopsis>
+CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_name</replaceable> ON (
+ <replaceable class="PARAMETER">column_name</replaceable>, <replaceable class="PARAMETER">column_name</replaceable> [, ...])
+ FROM <replaceable class="PARAMETER">table_name</replaceable>
+</synopsis>
+
+ </refsynopsisdiv>
+
+ <refsect1 id="SQL-CREATESTATISTICS-description">
+ <title>Description</title>
+
+ <para>
+ <command>CREATE STATISTICS</command> will create a new multivariate
+ statistics on the table. The statistics will be created in the in the
+ current database. The statistics will be owned by the user issuing
+ the command.
+ </para>
+
+ <para>
+ If a schema name is given (for example, <literal>CREATE STATISTICS
+ myschema.mystat ...</>) then the statistics is created in the specified
+ schema. Otherwise it is created in the current schema. The name of
+ the table must be distinct from the name of any other statistics in the
+ same schema.
+ </para>
+
+ <para>
+ To be able to create a table, you must have <literal>USAGE</literal>
+ privilege on all column types or the type in the <literal>OF</literal>
+ clause, respectively.
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>Parameters</title>
+
+ <variablelist>
+
+ <varlistentry>
+ <term><literal>IF NOT EXISTS</></term>
+ <listitem>
+ <para>
+ Do not throw an error if a statistics with the same name already exists.
+ A notice is issued in this case. Note that there is no guarantee that
+ the existing statistics is anything like the one that would have been
+ created.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="PARAMETER">statistics_name</replaceable></term>
+ <listitem>
+ <para>
+ The name (optionally schema-qualified) of the statistics to be created.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="PARAMETER">table_name</replaceable></term>
+ <listitem>
+ <para>
+ The name (optionally schema-qualified) of the table the statistics should
+ be created on.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="PARAMETER">column_name</replaceable></term>
+ <listitem>
+ <para>
+ The name of a column to be included in the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ </variablelist>
+
+ </refsect1>
+
+ <refsect1 id="SQL-CREATESTATISTICS-examples">
+ <title>Examples</title>
+
+ <para>
+ Create table <structname>t1</> with two functionally dependent columns, i.e.
+ knowledge of a value in the first column is sufficient for detemining the
+ value in the other column. Then functional dependencies are built on those
+ columns:
+
+<programlisting>
+CREATE TABLE t1 (
+ a int,
+ b int
+);
+
+INSERT INTO t1 SELECT i/100, i/500
+ FROM generate_series(1,1000000) s(i);
+
+CREATE STATISTICS s1 ON (a, b) FROM t1;
+
+ANALYZE t1;
+
+-- valid combination of values
+EXPLAIN ANALYZE SELECT * FROM t1 WHERE (a = 1) AND (b = 0);
+
+-- invalid combination of values
+EXPLAIN ANALYZE SELECT * FROM t1 WHERE (a = 1) AND (b = 1);
+</programlisting>
+ </para>
+
+ </refsect1>
+
+ <refsect1>
+ <title>Compatibility</title>
+
+ <para>
+ There's no <command>CREATE STATISTICS</command> command in the SQL standard.
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>See Also</title>
+
+ <simplelist type="inline">
+ <member><xref linkend="sql-alterstatistics"></member>
+ <member><xref linkend="sql-dropstatistics"></member>
+ </simplelist>
+ </refsect1>
+</refentry>
diff --git a/doc/src/sgml/ref/drop_statistics.sgml b/doc/src/sgml/ref/drop_statistics.sgml
new file mode 100644
index 0000000..dd9047a
--- /dev/null
+++ b/doc/src/sgml/ref/drop_statistics.sgml
@@ -0,0 +1,91 @@
+<!--
+doc/src/sgml/ref/drop_statistics.sgml
+PostgreSQL documentation
+-->
+
+<refentry id="SQL-DROPSTATISTICS">
+ <indexterm zone="sql-dropstatistics">
+ <primary>DROP STATISTICS</primary>
+ </indexterm>
+
+ <refmeta>
+ <refentrytitle>DROP STATISTICS</refentrytitle>
+ <manvolnum>7</manvolnum>
+ <refmiscinfo>SQL - Language Statements</refmiscinfo>
+ </refmeta>
+
+ <refnamediv>
+ <refname>DROP STATISTICS</refname>
+ <refpurpose>remove a statistics</refpurpose>
+ </refnamediv>
+
+ <refsynopsisdiv>
+<synopsis>
+DROP STATISTICS [ IF EXISTS ] <replaceable class="PARAMETER">name</replaceable> [, ...]
+</synopsis>
+ </refsynopsisdiv>
+
+ <refsect1>
+ <title>Description</title>
+
+ <para>
+ <command>DROP STATISTICS</command> removes statistics from the database.
+ Only the statistics owner, the schema owner, and superuser can drop a
+ statistics.
+ </para>
+
+ </refsect1>
+
+ <refsect1>
+ <title>Parameters</title>
+
+ <variablelist>
+ <varlistentry>
+ <term><literal>IF EXISTS</literal></term>
+ <listitem>
+ <para>
+ Do not throw an error if the statistics does not exist. A notice is
+ issued in this case.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="PARAMETER">name</replaceable></term>
+ <listitem>
+ <para>
+ The name (optionally schema-qualified) of the statistics to drop.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ </variablelist>
+ </refsect1>
+
+ <refsect1>
+ <title>Examples</title>
+
+ <para>
+ ...
+ </para>
+
+ </refsect1>
+
+ <refsect1>
+ <title>Compatibility</title>
+
+ <para>
+ There's no <command>DROP STATISTICS</command> command in the SQL standard.
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>See Also</title>
+
+ <simplelist type="inline">
+ <member><xref linkend="sql-alterstatistics"></member>
+ <member><xref linkend="sql-createstatistics"></member>
+ </simplelist>
+ </refsect1>
+
+</refentry>
diff --git a/doc/src/sgml/reference.sgml b/doc/src/sgml/reference.sgml
index 8acdff1..ba7e17b 100644
--- a/doc/src/sgml/reference.sgml
+++ b/doc/src/sgml/reference.sgml
@@ -59,6 +59,7 @@
&alterSchema;
&alterSequence;
&alterServer;
+ &alterStatistics;
&alterSystem;
&alterTable;
&alterTableSpace;
@@ -105,6 +106,7 @@
&createSchema;
&createSequence;
&createServer;
+ &createStatistics;
&createTable;
&createTableAs;
&createTableSpace;
@@ -149,6 +151,7 @@
&dropSchema;
&dropSequence;
&dropServer;
+ &dropStatistics;
&dropTable;
&dropTableSpace;
&dropTSConfig;
diff --git a/src/backend/catalog/Makefile b/src/backend/catalog/Makefile
index cd38c8a..9fc46b4 100644
--- a/src/backend/catalog/Makefile
+++ b/src/backend/catalog/Makefile
@@ -32,6 +32,7 @@ POSTGRES_BKI_SRCS = $(addprefix $(top_srcdir)/src/include/catalog/,\
pg_attrdef.h pg_constraint.h pg_inherits.h pg_index.h pg_operator.h \
pg_opfamily.h pg_opclass.h pg_am.h pg_amop.h pg_amproc.h \
pg_language.h pg_largeobject_metadata.h pg_largeobject.h pg_aggregate.h \
+ pg_mv_statistic.h \
pg_statistic.h pg_rewrite.h pg_trigger.h pg_event_trigger.h pg_description.h \
pg_cast.h pg_enum.h pg_namespace.h pg_conversion.h pg_depend.h \
pg_database.h pg_db_role_setting.h pg_tablespace.h pg_pltemplate.h \
diff --git a/src/backend/catalog/aclchk.c b/src/backend/catalog/aclchk.c
index fb6c276..7db7de9 100644
--- a/src/backend/catalog/aclchk.c
+++ b/src/backend/catalog/aclchk.c
@@ -40,6 +40,7 @@
#include "catalog/pg_language.h"
#include "catalog/pg_largeobject.h"
#include "catalog/pg_largeobject_metadata.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_operator.h"
@@ -5071,6 +5072,32 @@ pg_extension_ownercheck(Oid ext_oid, Oid roleid)
}
/*
+ * Ownership check for a multivariate statistics (specified by OID).
+ */
+bool
+pg_statistics_ownercheck(Oid stat_oid, Oid roleid)
+{
+ HeapTuple tuple;
+ Oid ownerId;
+
+ /* Superusers bypass all permission checking. */
+ if (superuser_arg(roleid))
+ return true;
+
+ tuple = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(stat_oid));
+ if (!HeapTupleIsValid(tuple))
+ ereport(ERROR,
+ (errcode(ERRCODE_UNDEFINED_OBJECT),
+ errmsg("statistics with OID %u does not exist", stat_oid)));
+
+ ownerId = ((Form_pg_mv_statistic) GETSTRUCT(tuple))->staowner;
+
+ ReleaseSysCache(tuple);
+
+ return has_privs_of_role(roleid, ownerId);
+}
+
+/*
* Check whether specified role has CREATEROLE privilege (or is a superuser)
*
* Note: roles do not have owners per se; instead we use this test in
diff --git a/src/backend/catalog/dependency.c b/src/backend/catalog/dependency.c
index 18a14bf..f2afa1c 100644
--- a/src/backend/catalog/dependency.c
+++ b/src/backend/catalog/dependency.c
@@ -42,6 +42,7 @@
#include "catalog/pg_init_privs.h"
#include "catalog/pg_language.h"
#include "catalog/pg_largeobject.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_operator.h"
@@ -164,7 +165,8 @@ static const Oid object_classes[] = {
ExtensionRelationId, /* OCLASS_EXTENSION */
EventTriggerRelationId, /* OCLASS_EVENT_TRIGGER */
PolicyRelationId, /* OCLASS_POLICY */
- TransformRelationId /* OCLASS_TRANSFORM */
+ TransformRelationId, /* OCLASS_TRANSFORM */
+ MvStatisticRelationId /* OCLASS_STATISTICS */
};
@@ -1248,6 +1250,10 @@ doDeletion(const ObjectAddress *object, int flags)
DropTransformById(object->objectId);
break;
+ case OCLASS_STATISTICS:
+ RemoveStatisticsById(object->objectId);
+ break;
+
default:
elog(ERROR, "unrecognized object class: %u",
object->classId);
@@ -2405,6 +2411,9 @@ getObjectClass(const ObjectAddress *object)
case TransformRelationId:
return OCLASS_TRANSFORM;
+
+ case MvStatisticRelationId:
+ return OCLASS_STATISTICS;
}
/* shouldn't get here */
diff --git a/src/backend/catalog/heap.c b/src/backend/catalog/heap.c
index e5d6aec..a0fbb7a 100644
--- a/src/backend/catalog/heap.c
+++ b/src/backend/catalog/heap.c
@@ -48,6 +48,7 @@
#include "catalog/pg_constraint_fn.h"
#include "catalog/pg_foreign_table.h"
#include "catalog/pg_inherits.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_partitioned_table.h"
@@ -1625,7 +1626,10 @@ RemoveAttributeById(Oid relid, AttrNumber attnum)
heap_close(attr_rel, RowExclusiveLock);
if (attnum > 0)
+ {
RemoveStatistics(relid, attnum);
+ RemoveMVStatistics(relid, attnum);
+ }
relation_close(rel, NoLock);
}
@@ -1876,6 +1880,11 @@ heap_drop_with_catalog(Oid relid)
RemoveStatistics(relid, 0);
/*
+ * delete multi-variate statistics
+ */
+ RemoveMVStatistics(relid, 0);
+
+ /*
* delete attribute tuples
*/
DeleteAttributeTuples(relid);
@@ -2794,6 +2803,98 @@ RemoveStatistics(Oid relid, AttrNumber attnum)
/*
+ * RemoveMVStatistics --- remove entries in pg_mv_statistic for a rel
+ *
+ * If attnum is zero, remove all entries for rel; else remove only the one(s)
+ * for that column.
+ */
+void
+RemoveMVStatistics(Oid relid, AttrNumber attnum)
+{
+ Relation pgmvstatistic;
+ TupleDesc tupdesc = NULL;
+ SysScanDesc scan;
+ ScanKeyData key;
+ HeapTuple tuple;
+
+ /*
+ * When dropping a column, we'll drop statistics with a single remaining
+ * (undropped column). To do that, we need the tuple descriptor.
+ *
+ * We already have the relation locked (as we're running ALTER TABLE ...
+ * DROP COLUMN), so we'll just get the descriptor here.
+ */
+ if (attnum != 0)
+ {
+ Relation rel = relation_open(relid, NoLock);
+
+ /* multivariate stats are supported on tables and matviews */
+ if (rel->rd_rel->relkind == RELKIND_RELATION ||
+ rel->rd_rel->relkind == RELKIND_MATVIEW)
+ tupdesc = RelationGetDescr(rel);
+
+ relation_close(rel, NoLock);
+ }
+
+ if (tupdesc == NULL)
+ return;
+
+ pgmvstatistic = heap_open(MvStatisticRelationId, RowExclusiveLock);
+
+ ScanKeyInit(&key,
+ Anum_pg_mv_statistic_starelid,
+ BTEqualStrategyNumber, F_OIDEQ,
+ ObjectIdGetDatum(relid));
+
+ scan = systable_beginscan(pgmvstatistic,
+ MvStatisticRelidIndexId,
+ true, NULL, 1, &key);
+
+ /* we must loop even when attnum != 0, in case of inherited stats */
+ while (HeapTupleIsValid(tuple = systable_getnext(scan)))
+ {
+ bool delete = true;
+
+ if (attnum != 0)
+ {
+ Datum adatum;
+ bool isnull;
+ int i;
+ int ncolumns = 0;
+ ArrayType *arr;
+ int16 *attnums;
+
+ /* get the columns */
+ adatum = SysCacheGetAttr(MVSTATOID, tuple,
+ Anum_pg_mv_statistic_stakeys, &isnull);
+ Assert(!isnull);
+
+ arr = DatumGetArrayTypeP(adatum);
+ attnums = (int16 *) ARR_DATA_PTR(arr);
+
+ for (i = 0; i < ARR_DIMS(arr)[0]; i++)
+ {
+ /* count the column unless it's has been / is being dropped */
+ if ((!tupdesc->attrs[attnums[i] - 1]->attisdropped) &&
+ (attnums[i] != attnum))
+ ncolumns += 1;
+ }
+
+ /* delete if there are less than two attributes */
+ delete = (ncolumns < 2);
+ }
+
+ if (delete)
+ simple_heap_delete(pgmvstatistic, &tuple->t_self);
+ }
+
+ systable_endscan(scan);
+
+ heap_close(pgmvstatistic, RowExclusiveLock);
+}
+
+
+/*
* RelationTruncateIndexes - truncate all indexes associated
* with the heap relation to zero tuples.
*
diff --git a/src/backend/catalog/namespace.c b/src/backend/catalog/namespace.c
index e3cfe22..2b6bbf1 100644
--- a/src/backend/catalog/namespace.c
+++ b/src/backend/catalog/namespace.c
@@ -4251,3 +4251,54 @@ pg_is_other_temp_schema(PG_FUNCTION_ARGS)
PG_RETURN_BOOL(isOtherTempNamespace(oid));
}
+
+Oid
+get_statistics_oid(List *names, bool missing_ok)
+{
+ char *schemaname;
+ char *stats_name;
+ Oid namespaceId;
+ Oid stats_oid = InvalidOid;
+ ListCell *l;
+
+ /* deconstruct the name list */
+ DeconstructQualifiedName(names, &schemaname, &stats_name);
+
+ if (schemaname)
+ {
+ /* use exact schema given */
+ namespaceId = LookupExplicitNamespace(schemaname, missing_ok);
+ if (missing_ok && !OidIsValid(namespaceId))
+ stats_oid = InvalidOid;
+ else
+ stats_oid = GetSysCacheOid2(MVSTATNAMENSP,
+ PointerGetDatum(stats_name),
+ ObjectIdGetDatum(namespaceId));
+ }
+ else
+ {
+ /* search for it in search path */
+ recomputeNamespacePath();
+
+ foreach(l, activeSearchPath)
+ {
+ namespaceId = lfirst_oid(l);
+
+ if (namespaceId == myTempNamespace)
+ continue; /* do not look in temp namespace */
+ stats_oid = GetSysCacheOid2(MVSTATNAMENSP,
+ PointerGetDatum(stats_name),
+ ObjectIdGetDatum(namespaceId));
+ if (OidIsValid(stats_oid))
+ break;
+ }
+ }
+
+ if (!OidIsValid(stats_oid) && !missing_ok)
+ ereport(ERROR,
+ (errcode(ERRCODE_UNDEFINED_OBJECT),
+ errmsg("statistics \"%s\" does not exist",
+ NameListToString(names))));
+
+ return stats_oid;
+}
diff --git a/src/backend/catalog/objectaddress.c b/src/backend/catalog/objectaddress.c
index bb4b080..7b2c079 100644
--- a/src/backend/catalog/objectaddress.c
+++ b/src/backend/catalog/objectaddress.c
@@ -39,6 +39,7 @@
#include "catalog/pg_language.h"
#include "catalog/pg_largeobject.h"
#include "catalog/pg_largeobject_metadata.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_opfamily.h"
@@ -450,9 +451,22 @@ static const ObjectPropertyType ObjectProperty[] =
Anum_pg_type_typacl,
ACL_KIND_TYPE,
true
+ },
+ {
+ MvStatisticRelationId,
+ MvStatisticOidIndexId,
+ MVSTATOID,
+ MVSTATNAMENSP,
+ Anum_pg_mv_statistic_staname,
+ Anum_pg_mv_statistic_stanamespace,
+ Anum_pg_mv_statistic_staowner,
+ InvalidAttrNumber, /* no ACL (same as relation) */
+ -1, /* no ACL */
+ true
}
};
+
/*
* This struct maps the string object types as returned by
* getObjectTypeDescription into ObjType enum values. Note that some enum
@@ -656,6 +670,10 @@ static const struct object_type_map
/* OCLASS_TRANSFORM */
{
"transform", OBJECT_TRANSFORM
+ },
+ /* OBJECT_STATISTICS */
+ {
+ "statistics", OBJECT_STATISTICS
}
};
@@ -930,6 +948,11 @@ get_object_address(ObjectType objtype, List *objname, List *objargs,
address = get_object_address_defacl(objname, objargs,
missing_ok);
break;
+ case OBJECT_STATISTICS:
+ address.classId = MvStatisticRelationId;
+ address.objectId = get_statistics_oid(objname, missing_ok);
+ address.objectSubId = 0;
+ break;
default:
elog(ERROR, "unrecognized objtype: %d", (int) objtype);
/* placate compiler, in case it thinks elog might return */
@@ -2237,6 +2260,10 @@ check_object_ownership(Oid roleid, ObjectType objtype, ObjectAddress address,
(errcode(ERRCODE_INSUFFICIENT_PRIVILEGE),
errmsg("must be superuser")));
break;
+ case OBJECT_STATISTICS:
+ if (!pg_statistics_ownercheck(address.objectId, roleid))
+ aclcheck_error_type(ACLCHECK_NOT_OWNER, address.objectId);
+ break;
default:
elog(ERROR, "unrecognized object type: %d",
(int) objtype);
@@ -3677,6 +3704,10 @@ getObjectTypeDescription(const ObjectAddress *object)
appendStringInfoString(&buffer, "access method");
break;
+ case OCLASS_STATISTICS:
+ appendStringInfoString(&buffer, "statistics");
+ break;
+
default:
appendStringInfo(&buffer, "unrecognized %u", object->classId);
break;
@@ -4648,6 +4679,29 @@ getObjectIdentityParts(const ObjectAddress *object,
}
break;
+ case OCLASS_STATISTICS:
+ {
+ HeapTuple tup;
+ Form_pg_mv_statistic formStatistic;
+ char *schema;
+
+ tup = SearchSysCache1(MVSTATOID,
+ ObjectIdGetDatum(object->objectId));
+ if (!HeapTupleIsValid(tup))
+ elog(ERROR, "cache lookup failed for statistics %u",
+ object->objectId);
+ formStatistic = (Form_pg_mv_statistic) GETSTRUCT(tup);
+ schema = get_namespace_name_or_temp(formStatistic->stanamespace);
+ appendStringInfoString(&buffer,
+ quote_qualified_identifier(schema,
+ NameStr(formStatistic->staname)));
+ if (objname)
+ *objname = list_make2(schema,
+ pstrdup(NameStr(formStatistic->staname)));
+ ReleaseSysCache(tup);
+ }
+ break;
+
default:
appendStringInfo(&buffer, "unrecognized object %u %u %d",
object->classId,
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 649cef8..7eb356e 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -181,6 +181,16 @@ CREATE OR REPLACE VIEW pg_sequences AS
WHERE NOT pg_is_other_temp_schema(N.oid)
AND relkind = 'S';
+CREATE VIEW pg_mv_stats AS
+ SELECT
+ N.nspname AS schemaname,
+ C.relname AS tablename,
+ S.staname AS staname,
+ S.stakeys AS attnums,
+ length(s.standist) AS ndistbytes
+ FROM (pg_mv_statistic S JOIN pg_class C ON (C.oid = S.starelid))
+ LEFT JOIN pg_namespace N ON (N.oid = C.relnamespace);
+
CREATE VIEW pg_stats WITH (security_barrier) AS
SELECT
nspname AS schemaname,
diff --git a/src/backend/commands/Makefile b/src/backend/commands/Makefile
index 6b3742c..3203c01 100644
--- a/src/backend/commands/Makefile
+++ b/src/backend/commands/Makefile
@@ -18,8 +18,8 @@ OBJS = amcmds.o aggregatecmds.o alter.o analyze.o async.o cluster.o comment.o \
event_trigger.o explain.o extension.o foreigncmds.o functioncmds.o \
indexcmds.o lockcmds.o matview.o operatorcmds.o opclasscmds.o \
policy.o portalcmds.o prepare.o proclang.o \
- schemacmds.o seclabel.o sequence.o tablecmds.o tablespace.o trigger.o \
- tsearchcmds.o typecmds.o user.o vacuum.o vacuumlazy.o \
- variable.o view.o
+ schemacmds.o seclabel.o sequence.o statscmds.o \
+ tablecmds.o tablespace.o trigger.o tsearchcmds.o typecmds.o \
+ user.o vacuum.o vacuumlazy.o variable.o view.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/commands/alter.c b/src/backend/commands/alter.c
index 03c0433..0716635 100644
--- a/src/backend/commands/alter.c
+++ b/src/backend/commands/alter.c
@@ -359,6 +359,7 @@ ExecRenameStmt(RenameStmt *stmt)
case OBJECT_OPCLASS:
case OBJECT_OPFAMILY:
case OBJECT_LANGUAGE:
+ case OBJECT_STATISTICS:
case OBJECT_TSCONFIGURATION:
case OBJECT_TSDICTIONARY:
case OBJECT_TSPARSER:
@@ -473,6 +474,7 @@ ExecAlterObjectSchemaStmt(AlterObjectSchemaStmt *stmt,
case OBJECT_OPERATOR:
case OBJECT_OPCLASS:
case OBJECT_OPFAMILY:
+ case OBJECT_STATISTICS:
case OBJECT_TSCONFIGURATION:
case OBJECT_TSDICTIONARY:
case OBJECT_TSPARSER:
@@ -780,6 +782,7 @@ ExecAlterOwnerStmt(AlterOwnerStmt *stmt)
case OBJECT_OPERATOR:
case OBJECT_OPCLASS:
case OBJECT_OPFAMILY:
+ case OBJECT_STATISTICS:
case OBJECT_TABLESPACE:
case OBJECT_TSDICTIONARY:
case OBJECT_TSCONFIGURATION:
diff --git a/src/backend/commands/analyze.c b/src/backend/commands/analyze.c
index f4afcd9..7bb3cfe 100644
--- a/src/backend/commands/analyze.c
+++ b/src/backend/commands/analyze.c
@@ -17,6 +17,7 @@
#include <math.h>
#include "access/multixact.h"
+#include "access/sysattr.h"
#include "access/transam.h"
#include "access/tupconvert.h"
#include "access/tuptoaster.h"
@@ -27,6 +28,7 @@
#include "catalog/indexing.h"
#include "catalog/pg_collation.h"
#include "catalog/pg_inherits_fn.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "commands/dbcommands.h"
#include "commands/tablecmds.h"
@@ -45,10 +47,13 @@
#include "storage/procarray.h"
#include "utils/acl.h"
#include "utils/attoptcache.h"
+#include "utils/builtins.h"
#include "utils/datum.h"
+#include "utils/fmgroids.h"
#include "utils/guc.h"
#include "utils/lsyscache.h"
#include "utils/memutils.h"
+#include "utils/mvstats.h"
#include "utils/pg_rusage.h"
#include "utils/sampling.h"
#include "utils/sortsupport.h"
@@ -559,6 +564,9 @@ do_analyze_rel(Relation onerel, int options, VacuumParams *params,
update_attstats(RelationGetRelid(Irel[ind]), false,
thisdata->attr_cnt, thisdata->vacattrstats);
}
+
+ /* Build multivariate stats (if there are any). */
+ build_mv_stats(onerel, totalrows, numrows, rows, attr_cnt, vacattrstats);
}
/*
diff --git a/src/backend/commands/dropcmds.c b/src/backend/commands/dropcmds.c
index 61ff8f2..070348b 100644
--- a/src/backend/commands/dropcmds.c
+++ b/src/backend/commands/dropcmds.c
@@ -296,6 +296,10 @@ does_not_exist_skipping(ObjectType objtype, List *objname, List *objargs)
msg = gettext_noop("schema \"%s\" does not exist, skipping");
name = NameListToString(objname);
break;
+ case OBJECT_STATISTICS:
+ msg = gettext_noop("statistics \"%s\" does not exist, skipping");
+ name = NameListToString(objname);
+ break;
case OBJECT_TSPARSER:
if (!schema_does_not_exist_skipping(objname, &msg, &name))
{
diff --git a/src/backend/commands/event_trigger.c b/src/backend/commands/event_trigger.c
index e87fce7..f9ce2a5 100644
--- a/src/backend/commands/event_trigger.c
+++ b/src/backend/commands/event_trigger.c
@@ -111,6 +111,7 @@ static event_trigger_support_data event_trigger_support[] = {
{"SCHEMA", true},
{"SEQUENCE", true},
{"SERVER", true},
+ {"STATISTICS", true},
{"TABLE", true},
{"TABLESPACE", false},
{"TRANSFORM", true},
@@ -1106,6 +1107,7 @@ EventTriggerSupportsObjectType(ObjectType obtype)
case OBJECT_RULE:
case OBJECT_SCHEMA:
case OBJECT_SEQUENCE:
+ case OBJECT_STATISTICS:
case OBJECT_TABCONSTRAINT:
case OBJECT_TABLE:
case OBJECT_TRANSFORM:
@@ -1168,6 +1170,7 @@ EventTriggerSupportsObjectClass(ObjectClass objclass)
case OCLASS_EXTENSION:
case OCLASS_POLICY:
case OCLASS_AM:
+ case OCLASS_STATISTICS:
return true;
}
diff --git a/src/backend/commands/statscmds.c b/src/backend/commands/statscmds.c
new file mode 100644
index 0000000..8453dc4
--- /dev/null
+++ b/src/backend/commands/statscmds.c
@@ -0,0 +1,259 @@
+/*-------------------------------------------------------------------------
+ *
+ * statscmds.c
+ * Commands for creating and altering multivariate statistics
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/commands/statscmds.c
+ *
+ *-------------------------------------------------------------------------
+ */
+#include "postgres.h"
+
+#include "access/relscan.h"
+#include "catalog/dependency.h"
+#include "catalog/indexing.h"
+#include "catalog/namespace.h"
+#include "catalog/pg_mv_statistic.h"
+#include "catalog/pg_namespace.h"
+#include "commands/defrem.h"
+#include "miscadmin.h"
+#include "utils/builtins.h"
+#include "utils/inval.h"
+#include "utils/memutils.h"
+#include "utils/mvstats.h"
+#include "utils/rel.h"
+#include "utils/syscache.h"
+
+
+/* used for sorting the attnums in ExecCreateStatistics */
+static int
+compare_int16(const void *a, const void *b)
+{
+ return memcmp(a, b, sizeof(int16));
+}
+
+/*
+ * Implements the CREATE STATISTICS name ON (columns) FROM table
+ *
+ * We do require that the types support sorting (ltopr), although some
+ * statistics might work with equality only.
+ *
+ * XXX Maybe this should check for duplicate stats. Although it's not clear
+ * what "duplicate" would mean here (wheter to compare only keys or also
+ * options). Moreover, we don't do such checks for indexes, although those
+ * store tuples and recreating a new index may be a way to fix bloat (which
+ * is a problem statistics don't have).
+ */
+ObjectAddress
+CreateStatistics(CreateStatsStmt *stmt)
+{
+ int i;
+ ListCell *l;
+ int16 attnums[MVSTATS_MAX_DIMENSIONS];
+ int numcols = 0;
+ ObjectAddress address = InvalidObjectAddress;
+ char *namestr;
+ NameData staname;
+ Oid statoid;
+ Oid namespaceId;
+
+ HeapTuple htup;
+ Datum values[Natts_pg_mv_statistic];
+ bool nulls[Natts_pg_mv_statistic];
+ int2vector *stakeys;
+ Relation mvstatrel;
+ Relation rel;
+ Oid relid;
+ ObjectAddress parentobject,
+ childobject;
+
+ Assert(IsA(stmt, CreateStatsStmt));
+
+ /* resolve the pieces of the name (namespace etc.) */
+ namespaceId = QualifiedNameGetCreationNamespace(stmt->defnames, &namestr);
+ namestrcpy(&staname, namestr);
+
+ /*
+ * If if_not_exists was given and the statistics already exists, bail out.
+ */
+ if (stmt->if_not_exists &&
+ SearchSysCacheExists2(MVSTATNAMENSP,
+ PointerGetDatum(&staname),
+ ObjectIdGetDatum(namespaceId)))
+ {
+ ereport(NOTICE,
+ (errcode(ERRCODE_DUPLICATE_OBJECT),
+ errmsg("statistics \"%s\" already exists, skipping",
+ namestr)));
+ return InvalidObjectAddress;
+ }
+
+ rel = heap_openrv(stmt->relation, AccessExclusiveLock);
+ relid = RelationGetRelid(rel);
+
+ /*
+ * Transform column names to array of attnums. While doing that, we
+ * also enforce the maximum number of keys.
+ */
+ foreach(l, stmt->keys)
+ {
+ char *attname = strVal(lfirst(l));
+ HeapTuple atttuple;
+
+ atttuple = SearchSysCacheAttName(relid, attname);
+
+ if (!HeapTupleIsValid(atttuple))
+ ereport(ERROR,
+ (errcode(ERRCODE_UNDEFINED_COLUMN),
+ errmsg("column \"%s\" referenced in statistics does not exist",
+ attname)));
+
+ /* more than MVSTATS_MAX_DIMENSIONS columns not allowed */
+ if (numcols >= MVSTATS_MAX_DIMENSIONS)
+ ereport(ERROR,
+ (errcode(ERRCODE_TOO_MANY_COLUMNS),
+ errmsg("cannot have more than %d keys in a statistics",
+ MVSTATS_MAX_DIMENSIONS)));
+
+ attnums[numcols] = ((Form_pg_attribute) GETSTRUCT(atttuple))->attnum;
+ ReleaseSysCache(atttuple);
+ numcols++;
+ }
+
+ /*
+ * Check that at least two columns were specified in the statement.
+ * The upper bound was already checked in the loop above.
+ */
+ if (numcols < 2)
+ ereport(ERROR,
+ (errcode(ERRCODE_TOO_MANY_COLUMNS),
+ errmsg("statistics require at least 2 columns")));
+
+ /*
+ * Sort the attnums, which makes detecting duplicies somewhat
+ * easier, and it does not hurt (it does not affect the efficiency,
+ * onlike for indexes, for example).
+ */
+ qsort(attnums, numcols, sizeof(int16), compare_int16);
+
+ /*
+ * Look for duplicities in the list of columns. The attnums are sorted
+ * so just check consecutive elements.
+ */
+ for (i = 1; i < numcols; i++)
+ if (attnums[i] == attnums[i-1])
+ ereport(ERROR,
+ (errcode(ERRCODE_UNDEFINED_COLUMN),
+ errmsg("duplicate column name in statistics definition")));
+
+ stakeys = buildint2vector(attnums, numcols);
+
+ /*
+ * Everything seems fine, so let's build the pg_mv_statistic entry.
+ * At this point we obviously only have the keys and options.
+ */
+
+ memset(values, 0, sizeof(values));
+ memset(nulls, false, sizeof(nulls));
+
+ /* metadata */
+ values[Anum_pg_mv_statistic_starelid - 1] = ObjectIdGetDatum(relid);
+ values[Anum_pg_mv_statistic_staname - 1] = NameGetDatum(&staname);
+ values[Anum_pg_mv_statistic_stanamespace - 1] = ObjectIdGetDatum(namespaceId);
+ values[Anum_pg_mv_statistic_staowner - 1] = ObjectIdGetDatum(GetUserId());
+
+ values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(stakeys);
+
+ /* enabled statistics */
+ values[Anum_pg_mv_statistic_ndist_enabled - 1] = BoolGetDatum(true);
+
+ nulls[Anum_pg_mv_statistic_standist - 1] = true;
+
+ /* insert the tuple into pg_mv_statistic */
+ mvstatrel = heap_open(MvStatisticRelationId, RowExclusiveLock);
+
+ htup = heap_form_tuple(mvstatrel->rd_att, values, nulls);
+
+ simple_heap_insert(mvstatrel, htup);
+
+ CatalogUpdateIndexes(mvstatrel, htup);
+
+ statoid = HeapTupleGetOid(htup);
+
+ heap_freetuple(htup);
+
+ /*
+ * Add a dependency on a table, so that stats get dropped on DROP TABLE.
+ */
+ ObjectAddressSet(parentobject, RelationRelationId, relid);
+ ObjectAddressSet(childobject, MvStatisticRelationId, statoid);
+
+ recordDependencyOn(&childobject, &parentobject, DEPENDENCY_AUTO);
+
+ /*
+ * Also add dependency on the schema (to drop statistics on DROP SCHEMA).
+ * This is not handled automatically by DROP TABLE because statistics have
+ * their own schema.
+ */
+ ObjectAddressSet(parentobject, NamespaceRelationId, namespaceId);
+
+ recordDependencyOn(&childobject, &parentobject, DEPENDENCY_AUTO);
+
+ heap_close(mvstatrel, RowExclusiveLock);
+
+ relation_close(rel, NoLock);
+
+ /*
+ * Invalidate relcache so that others see the new statistics.
+ */
+ CacheInvalidateRelcache(rel);
+
+ ObjectAddressSet(address, MvStatisticRelationId, statoid);
+
+ return address;
+}
+
+
+/*
+ * Implements the DROP STATISTICS
+ *
+ * DROP STATISTICS stats_name
+ */
+void
+RemoveStatisticsById(Oid statsOid)
+{
+ Relation relation;
+ Oid relid;
+ Relation rel;
+ HeapTuple tup;
+ Form_pg_mv_statistic mvstat;
+
+ /*
+ * Delete the pg_proc tuple.
+ */
+ relation = heap_open(MvStatisticRelationId, RowExclusiveLock);
+
+ tup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(statsOid));
+
+ if (!HeapTupleIsValid(tup)) /* should not happen */
+ elog(ERROR, "cache lookup failed for statistics %u", statsOid);
+
+ mvstat = (Form_pg_mv_statistic) GETSTRUCT(tup);
+ relid = mvstat->starelid;
+
+ rel = heap_open(relid, AccessExclusiveLock);
+
+ simple_heap_delete(relation, &tup->t_self);
+
+ CacheInvalidateRelcache(rel);
+
+ ReleaseSysCache(tup);
+
+ heap_close(relation, RowExclusiveLock);
+ heap_close(rel, NoLock);
+}
diff --git a/src/backend/nodes/copyfuncs.c b/src/backend/nodes/copyfuncs.c
index 6955298..256f8c6 100644
--- a/src/backend/nodes/copyfuncs.c
+++ b/src/backend/nodes/copyfuncs.c
@@ -4252,6 +4252,19 @@ _copyPartitionCmd(const PartitionCmd *from)
return newnode;
}
+static CreateStatsStmt *
+_copyCreateStatsStmt(const CreateStatsStmt *from)
+{
+ CreateStatsStmt *newnode = makeNode(CreateStatsStmt);
+
+ COPY_NODE_FIELD(defnames);
+ COPY_NODE_FIELD(relation);
+ COPY_NODE_FIELD(keys);
+ COPY_SCALAR_FIELD(if_not_exists);
+
+ return newnode;
+}
+
/* ****************************************************************
* pg_list.h copy functions
* ****************************************************************
@@ -5154,6 +5167,9 @@ copyObject(const void *from)
case T_CommonTableExpr:
retval = _copyCommonTableExpr(from);
break;
+ case T_CreateStatsStmt:
+ retval = _copyCreateStatsStmt(from);
+ break;
case T_FuncWithArgs:
retval = _copyFuncWithArgs(from);
break;
diff --git a/src/backend/nodes/outfuncs.c b/src/backend/nodes/outfuncs.c
index 9fe9873..1c2c200 100644
--- a/src/backend/nodes/outfuncs.c
+++ b/src/backend/nodes/outfuncs.c
@@ -2171,6 +2171,21 @@ _outForeignKeyOptInfo(StringInfo str, const ForeignKeyOptInfo *node)
}
static void
+_outMVStatisticInfo(StringInfo str, const MVStatisticInfo *node)
+{
+ WRITE_NODE_TYPE("MVSTATISTICINFO");
+
+ /* NB: this isn't a complete set of fields */
+ WRITE_OID_FIELD(mvoid);
+
+ /* enabled statistics */
+ WRITE_BOOL_FIELD(ndist_enabled);
+
+ /* built/available statistics */
+ WRITE_BOOL_FIELD(ndist_built);
+}
+
+static void
_outEquivalenceClass(StringInfo str, const EquivalenceClass *node)
{
/*
@@ -3764,6 +3779,9 @@ outNode(StringInfo str, const void *obj)
case T_PlannerParamItem:
_outPlannerParamItem(str, obj);
break;
+ case T_MVStatisticInfo:
+ _outMVStatisticInfo(str, obj);
+ break;
case T_ExtensibleNode:
_outExtensibleNode(str, obj);
diff --git a/src/backend/optimizer/util/plancat.c b/src/backend/optimizer/util/plancat.c
index 72272d9..16a90ea 100644
--- a/src/backend/optimizer/util/plancat.c
+++ b/src/backend/optimizer/util/plancat.c
@@ -29,6 +29,7 @@
#include "catalog/heap.h"
#include "catalog/partition.h"
#include "catalog/pg_am.h"
+#include "catalog/pg_mv_statistic.h"
#include "foreign/fdwapi.h"
#include "miscadmin.h"
#include "nodes/makefuncs.h"
@@ -41,7 +42,9 @@
#include "parser/parsetree.h"
#include "rewrite/rewriteManip.h"
#include "storage/bufmgr.h"
+#include "utils/builtins.h"
#include "utils/lsyscache.h"
+#include "utils/syscache.h"
#include "utils/rel.h"
#include "utils/snapmgr.h"
@@ -100,6 +103,7 @@ get_relation_info(PlannerInfo *root, Oid relationObjectId, bool inhparent,
Relation relation;
bool hasindex;
List *indexinfos = NIL;
+ List *stainfos = NIL;
/*
* We need not lock the relation since it was already locked, either by
@@ -397,6 +401,61 @@ get_relation_info(PlannerInfo *root, Oid relationObjectId, bool inhparent,
rel->indexlist = indexinfos;
+ if (true)
+ {
+ List *mvstatoidlist;
+ ListCell *l;
+
+ mvstatoidlist = RelationGetMVStatList(relation);
+
+ foreach(l, mvstatoidlist)
+ {
+ ArrayType *arr;
+ Datum adatum;
+ bool isnull;
+ Oid mvoid = lfirst_oid(l);
+ Form_pg_mv_statistic mvstat;
+ MVStatisticInfo *info;
+
+ HeapTuple htup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(mvoid));
+
+ mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
+
+ /* unavailable stats are not interesting for the planner */
+ if (mvstat->ndist_built)
+ {
+ info = makeNode(MVStatisticInfo);
+
+ info->mvoid = mvoid;
+ info->rel = rel;
+
+ /* enabled statistics */
+ info->ndist_enabled = mvstat->ndist_enabled;
+
+ /* built/available statistics */
+ info->ndist_built = mvstat->ndist_built;
+
+ /* stakeys */
+ adatum = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_stakeys, &isnull);
+ Assert(!isnull);
+
+ arr = DatumGetArrayTypeP(adatum);
+
+ info->stakeys = buildint2vector((int16 *) ARR_DATA_PTR(arr),
+ ARR_DIMS(arr)[0]);
+
+ stainfos = lcons(info, stainfos);
+ }
+
+ ReleaseSysCache(htup);
+ }
+
+ list_free(mvstatoidlist);
+ }
+
+ rel->mvstatlist = stainfos;
+
/* Grab foreign-table info using the relcache, while we have it */
if (relation->rd_rel->relkind == RELKIND_FOREIGN_TABLE)
{
diff --git a/src/backend/parser/gram.y b/src/backend/parser/gram.y
index 08cf5b7..3fd1fac 100644
--- a/src/backend/parser/gram.y
+++ b/src/backend/parser/gram.y
@@ -248,7 +248,7 @@ static Node *makeRecursiveViewSelect(char *relname, List *aliases, Node *query);
ConstraintsSetStmt CopyStmt CreateAsStmt CreateCastStmt
CreateDomainStmt CreateExtensionStmt CreateGroupStmt CreateOpClassStmt
CreateOpFamilyStmt AlterOpFamilyStmt CreatePLangStmt
- CreateSchemaStmt CreateSeqStmt CreateStmt CreateTableSpaceStmt
+ CreateSchemaStmt CreateSeqStmt CreateStmt CreateStatsStmt CreateTableSpaceStmt
CreateFdwStmt CreateForeignServerStmt CreateForeignTableStmt
CreateAssertStmt CreateTransformStmt CreateTrigStmt CreateEventTrigStmt
CreateUserStmt CreateUserMappingStmt CreateRoleStmt CreatePolicyStmt
@@ -834,6 +834,7 @@ stmt :
| CreateSchemaStmt
| CreateSeqStmt
| CreateStmt
+ | CreateStatsStmt
| CreateTableSpaceStmt
| CreateTransformStmt
| CreateTrigStmt
@@ -3713,6 +3714,34 @@ OptConsTableSpace: USING INDEX TABLESPACE name { $$ = $4; }
ExistingIndex: USING INDEX index_name { $$ = $3; }
;
+/*****************************************************************************
+ *
+ * QUERY :
+ * CREATE STATISTICS stats_name ON relname (columns) WITH (options)
+ *
+ *****************************************************************************/
+
+
+CreateStatsStmt: CREATE STATISTICS any_name ON '(' columnList ')' FROM qualified_name
+ {
+ CreateStatsStmt *n = makeNode(CreateStatsStmt);
+ n->defnames = $3;
+ n->relation = $9;
+ n->keys = $6;
+ n->if_not_exists = false;
+ $$ = (Node *)n;
+ }
+ | CREATE STATISTICS IF_P NOT EXISTS any_name ON '(' columnList ')' FROM qualified_name
+ {
+ CreateStatsStmt *n = makeNode(CreateStatsStmt);
+ n->defnames = $6;
+ n->relation = $12;
+ n->keys = $9;
+ n->if_not_exists = true;
+ $$ = (Node *)n;
+ }
+ ;
+
/*****************************************************************************
*
@@ -6050,6 +6079,7 @@ drop_type: TABLE { $$ = OBJECT_TABLE; }
| TEXT_P SEARCH DICTIONARY { $$ = OBJECT_TSDICTIONARY; }
| TEXT_P SEARCH TEMPLATE { $$ = OBJECT_TSTEMPLATE; }
| TEXT_P SEARCH CONFIGURATION { $$ = OBJECT_TSCONFIGURATION; }
+ | STATISTICS { $$ = OBJECT_STATISTICS; }
;
any_name_list:
@@ -8435,6 +8465,15 @@ RenameStmt: ALTER AGGREGATE aggregate_with_argtypes RENAME TO name
n->missing_ok = false;
$$ = (Node *)n;
}
+ | ALTER STATISTICS any_name RENAME TO name
+ {
+ RenameStmt *n = makeNode(RenameStmt);
+ n->renameType = OBJECT_STATISTICS;
+ n->object = $3;
+ n->newname = $6;
+ n->missing_ok = false;
+ $$ = (Node *)n;
+ }
;
opt_column: COLUMN { $$ = COLUMN; }
@@ -8720,6 +8759,15 @@ AlterObjectSchemaStmt:
n->missing_ok = false;
$$ = (Node *)n;
}
+ | ALTER STATISTICS any_name SET SCHEMA name
+ {
+ AlterObjectSchemaStmt *n = makeNode(AlterObjectSchemaStmt);
+ n->objectType = OBJECT_STATISTICS;
+ n->object = $3;
+ n->newschema = $6;
+ n->missing_ok = false;
+ $$ = (Node *)n;
+ }
;
/*****************************************************************************
@@ -8910,6 +8958,14 @@ AlterOwnerStmt: ALTER AGGREGATE aggregate_with_argtypes OWNER TO RoleSpec
n->newowner = $7;
$$ = (Node *)n;
}
+ | ALTER STATISTICS name OWNER TO RoleSpec
+ {
+ AlterOwnerStmt *n = makeNode(AlterOwnerStmt);
+ n->objectType = OBJECT_STATISTICS;
+ n->object = list_make1(makeString($3));
+ n->newowner = $6;
+ $$ = (Node *)n;
+ }
;
diff --git a/src/backend/tcop/utility.c b/src/backend/tcop/utility.c
index 2e89ad7..b19632f 100644
--- a/src/backend/tcop/utility.c
+++ b/src/backend/tcop/utility.c
@@ -1562,6 +1562,10 @@ ProcessUtilitySlow(ParseState *pstate,
address = CreateAccessMethod((CreateAmStmt *) parsetree);
break;
+ case T_CreateStatsStmt: /* CREATE STATISTICS */
+ address = CreateStatistics((CreateStatsStmt *) parsetree);
+ break;
+
default:
elog(ERROR, "unrecognized node type: %d",
(int) nodeTag(parsetree));
@@ -1920,6 +1924,9 @@ AlterObjectTypeCommandTag(ObjectType objtype)
case OBJECT_MATVIEW:
tag = "ALTER MATERIALIZED VIEW";
break;
+ case OBJECT_STATISTICS:
+ tag = "ALTER STATISTICS";
+ break;
default:
tag = "???";
break;
@@ -2205,6 +2212,9 @@ CreateCommandTag(Node *parsetree)
case OBJECT_ACCESS_METHOD:
tag = "DROP ACCESS METHOD";
break;
+ case OBJECT_STATISTICS:
+ tag = "DROP STATISTICS";
+ break;
default:
tag = "???";
}
@@ -2583,6 +2593,10 @@ CreateCommandTag(Node *parsetree)
tag = "EXECUTE";
break;
+ case T_CreateStatsStmt:
+ tag = "CREATE STATISTICS";
+ break;
+
case T_DeallocateStmt:
{
DeallocateStmt *stmt = (DeallocateStmt *) parsetree;
diff --git a/src/backend/utils/Makefile b/src/backend/utils/Makefile
index 8374533..eba0352 100644
--- a/src/backend/utils/Makefile
+++ b/src/backend/utils/Makefile
@@ -9,7 +9,7 @@ top_builddir = ../../..
include $(top_builddir)/src/Makefile.global
OBJS = fmgrtab.o
-SUBDIRS = adt cache error fmgr hash init mb misc mmgr resowner sort time
+SUBDIRS = adt cache error fmgr hash init mb misc mmgr mvstats resowner sort time
# location of Catalog.pm
catalogdir = $(top_srcdir)/src/backend/catalog
diff --git a/src/backend/utils/adt/selfuncs.c b/src/backend/utils/adt/selfuncs.c
index 4973396..a33549c 100644
--- a/src/backend/utils/adt/selfuncs.c
+++ b/src/backend/utils/adt/selfuncs.c
@@ -132,6 +132,7 @@
#include "utils/fmgroids.h"
#include "utils/index_selfuncs.h"
#include "utils/lsyscache.h"
+#include "utils/mvstats.h"
#include "utils/nabstime.h"
#include "utils/pg_locale.h"
#include "utils/rel.h"
@@ -206,6 +207,8 @@ static Const *string_to_const(const char *str, Oid datatype);
static Const *string_to_bytea_const(const char *str, size_t str_len);
static List *add_predicate_to_quals(IndexOptInfo *index, List *indexQuals);
+static double find_ndistinct(PlannerInfo *root, RelOptInfo *rel, List *varinfos,
+ bool *found);
/*
* eqsel - Selectivity of "=" for any data types.
@@ -3435,12 +3438,26 @@ estimate_num_groups(PlannerInfo *root, List *groupExprs, double input_rows,
* don't know by how much. We should never clamp to less than the
* largest ndistinct value for any of the Vars, though, since
* there will surely be at least that many groups.
+ *
+ * However we don't need to do this if we have ndistinct stats on
+ * the columns - in that case we can simply use the coefficient
+ * to get the (probably way more accurate) estimate.
+ *
+ * XXX Might benefit from some refactoring, mixing the ndistinct
+ * coefficients and clamp seems a bit unfortunate.
*/
double clamp = rel->tuples;
if (relvarcount > 1)
{
- clamp *= 0.1;
+ bool found;
+ double ndist = find_ndistinct(root, rel, varinfos, &found);
+
+ if (found)
+ reldistinct = ndist;
+ else
+ clamp *= 0.1;
+
if (clamp < relmaxndistinct)
{
clamp = relmaxndistinct;
@@ -3449,6 +3466,7 @@ estimate_num_groups(PlannerInfo *root, List *groupExprs, double input_rows,
clamp = rel->tuples;
}
}
+
if (reldistinct > clamp)
reldistinct = clamp;
@@ -7599,3 +7617,151 @@ brincostestimate(PlannerInfo *root, IndexPath *path, double loop_count,
/* XXX what about pages_per_range? */
}
+
+/*
+ * Find applicable ndistinct statistics and compute the coefficient to
+ * correct the estimate (simply a product of per-column ndistincts).
+ *
+ * XXX Currently we only look for a perfect match, i.e. a single ndistinct
+ * estimate exactly matching all the columns of the statistics. This may be
+ * a bit problematic as adding a column (not covered by the ndistinct stats)
+ * will prevent us from using the stats entirely. So instead this needs to
+ * estimate the covered attributes, and then combine that with the extra
+ * attributes somehow (probably the old way).
+ */
+static double
+find_ndistinct(PlannerInfo *root, RelOptInfo *rel, List *varinfos, bool *found)
+{
+ ListCell *lc;
+ Bitmapset *attnums = NULL;
+ VariableStatData vardata;
+
+ /* assume we haven't found any suitable ndistinct statistics */
+ *found = false;
+
+ /* bail out immediately if the table has no multivariate statistics */
+ if (!rel->mvstatlist)
+ return 0.0;
+
+ foreach(lc, varinfos)
+ {
+ GroupVarInfo *varinfo = (GroupVarInfo *) lfirst(lc);
+
+ if (varinfo->rel != rel)
+ continue;
+
+ /* FIXME handle expressions in general only */
+
+ /*
+ * examine the variable (or expression) so that we know which
+ * attribute we're dealing with - we need this for matching the
+ * ndistinct coefficient
+ *
+ * FIXME probably might remember this from estimate_num_groups
+ */
+ examine_variable(root, varinfo->var, 0, &vardata);
+
+ if (HeapTupleIsValid(vardata.statsTuple))
+ {
+ Form_pg_statistic stats
+ = (Form_pg_statistic) GETSTRUCT(vardata.statsTuple);
+
+ attnums = bms_add_member(attnums, stats->staattnum);
+
+ ReleaseVariableStats(vardata);
+ }
+ }
+
+ /* look for a matching ndistinct statistics */
+ foreach (lc, rel->mvstatlist)
+ {
+ int i, k;
+ bool matches;
+ MVStatisticInfo *info = (MVStatisticInfo *)lfirst(lc);
+
+ /* skip statistics without ndistinct coefficient built */
+ if (!info->ndist_built)
+ continue;
+
+ /*
+ * Only ndistinct stats covering all Vars are acceptable, which can't
+ * happen if the statistics has fewer attributes than we have Vars.
+ */
+ if (bms_num_members(attnums) > info->stakeys->dim1)
+ continue;
+
+ /* check that all Vars are covered by the statistic */
+ matches = true; /* assume match until we find unmatched attribute */
+ k = -1;
+ while ((k = bms_next_member(attnums, k)) >= 0)
+ {
+ bool attr_found = false;
+ for (i = 0; i < info->stakeys->dim1; i++)
+ {
+ if (info->stakeys->values[i] == k)
+ {
+ attr_found = true;
+ break;
+ }
+ }
+
+ /* found attribute not covered by this ndistinct stats, skip */
+ if (!attr_found)
+ {
+ matches = false;
+ break;
+ }
+ }
+
+ if (! matches)
+ continue;
+
+ /* hey, this statistics matches! great, let's extract the value */
+ *found = true;
+
+ {
+ int j;
+ MVNDistinct stat = load_mv_ndistinct(info->mvoid);
+
+ for (j = 0; j < stat->nitems; j++)
+ {
+ bool item_matches = true;
+ MVNDistinctItem * item = &stat->items[j];
+
+ /* not the right item (different number of attributes) */
+ if (item->nattrs != bms_num_members(attnums))
+ continue;
+
+ /* check the attribute numbers */
+ k = -1;
+ while ((k = bms_next_member(attnums, k)) >= 0)
+ {
+ bool attr_found = false;
+ for (i = 0; i < item->nattrs; i++)
+ {
+ if (info->stakeys->values[item->attrs[i]] == k)
+ {
+ attr_found = true;
+ break;
+ }
+ }
+
+ if (! attr_found)
+ {
+ item_matches = false;
+ break;
+ }
+ }
+
+ if (! item_matches)
+ continue;
+
+ return item->ndistinct;
+ }
+ }
+ }
+
+ Assert(!(*found));
+
+ return 0.0;
+}
diff --git a/src/backend/utils/cache/relcache.c b/src/backend/utils/cache/relcache.c
index 2a68359..83688eb 100644
--- a/src/backend/utils/cache/relcache.c
+++ b/src/backend/utils/cache/relcache.c
@@ -49,6 +49,7 @@
#include "catalog/pg_auth_members.h"
#include "catalog/pg_constraint.h"
#include "catalog/pg_database.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_partitioned_table.h"
@@ -4440,6 +4441,62 @@ RelationGetIndexList(Relation relation)
return result;
}
+
+List *
+RelationGetMVStatList(Relation relation)
+{
+ Relation indrel;
+ SysScanDesc indscan;
+ ScanKeyData skey;
+ HeapTuple htup;
+ List *result;
+ List *oldlist;
+ MemoryContext oldcxt;
+
+ /* Quick exit if we already computed the list. */
+ if (relation->rd_mvstatvalid != 0)
+ return list_copy(relation->rd_mvstatlist);
+
+ /*
+ * We build the list we intend to return (in the caller's context) while
+ * doing the scan. After successfully completing the scan, we copy that
+ * list into the relcache entry. This avoids cache-context memory leakage
+ * if we get some sort of error partway through.
+ */
+ result = NIL;
+
+ /* Prepare to scan pg_index for entries having indrelid = this rel. */
+ ScanKeyInit(&skey,
+ Anum_pg_mv_statistic_starelid,
+ BTEqualStrategyNumber, F_OIDEQ,
+ ObjectIdGetDatum(RelationGetRelid(relation)));
+
+ indrel = heap_open(MvStatisticRelationId, AccessShareLock);
+ indscan = systable_beginscan(indrel, MvStatisticRelidIndexId, true,
+ NULL, 1, &skey);
+
+ while (HeapTupleIsValid(htup = systable_getnext(indscan)))
+ /* TODO maybe include only already built statistics? */
+ result = insert_ordered_oid(result, HeapTupleGetOid(htup));
+
+ systable_endscan(indscan);
+
+ heap_close(indrel, AccessShareLock);
+
+ /* Now save a copy of the completed list in the relcache entry. */
+ oldcxt = MemoryContextSwitchTo(CacheMemoryContext);
+ oldlist = relation->rd_mvstatlist;
+ relation->rd_mvstatlist = list_copy(result);
+
+ relation->rd_mvstatvalid = true;
+ MemoryContextSwitchTo(oldcxt);
+
+ /* Don't leak the old list, if there is one */
+ list_free(oldlist);
+
+ return result;
+}
+
/*
* insert_ordered_oid
* Insert a new Oid into a sorted list of Oids, preserving ordering
@@ -5411,6 +5468,8 @@ load_relcache_init_file(bool shared)
rel->rd_indexattr = NULL;
rel->rd_keyattr = NULL;
rel->rd_idattr = NULL;
+ rel->rd_mvstatvalid = false;
+ rel->rd_mvstatlist = NIL;
rel->rd_createSubid = InvalidSubTransactionId;
rel->rd_newRelfilenodeSubid = InvalidSubTransactionId;
rel->rd_amcache = NULL;
diff --git a/src/backend/utils/cache/syscache.c b/src/backend/utils/cache/syscache.c
index e87fe0e..4d3e8cd 100644
--- a/src/backend/utils/cache/syscache.c
+++ b/src/backend/utils/cache/syscache.c
@@ -44,6 +44,7 @@
#include "catalog/pg_foreign_server.h"
#include "catalog/pg_foreign_table.h"
#include "catalog/pg_language.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_operator.h"
@@ -504,6 +505,28 @@ static const struct cachedesc cacheinfo[] = {
},
4
},
+ {MvStatisticRelationId, /* MVSTATNAMENSP */
+ MvStatisticNameIndexId,
+ 2,
+ {
+ Anum_pg_mv_statistic_staname,
+ Anum_pg_mv_statistic_stanamespace,
+ 0,
+ 0
+ },
+ 4
+ },
+ {MvStatisticRelationId, /* MVSTATOID */
+ MvStatisticOidIndexId,
+ 1,
+ {
+ ObjectIdAttributeNumber,
+ 0,
+ 0,
+ 0
+ },
+ 4
+ },
{NamespaceRelationId, /* NAMESPACENAME */
NamespaceNameIndexId,
1,
diff --git a/src/backend/utils/mvstats/Makefile b/src/backend/utils/mvstats/Makefile
new file mode 100644
index 0000000..7295d46
--- /dev/null
+++ b/src/backend/utils/mvstats/Makefile
@@ -0,0 +1,17 @@
+#-------------------------------------------------------------------------
+#
+# Makefile--
+# Makefile for utils/mvstats
+#
+# IDENTIFICATION
+# src/backend/utils/mvstats/Makefile
+#
+#-------------------------------------------------------------------------
+
+subdir = src/backend/utils/mvstats
+top_builddir = ../../../..
+include $(top_builddir)/src/Makefile.global
+
+OBJS = common.o mvdist.o
+
+include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/utils/mvstats/README.ndistinct b/src/backend/utils/mvstats/README.ndistinct
new file mode 100644
index 0000000..9365b17
--- /dev/null
+++ b/src/backend/utils/mvstats/README.ndistinct
@@ -0,0 +1,22 @@
+ndistinct coefficients
+======================
+
+Estimating number of groups in a combination of columns (e.g. for GROUP BY)
+is tricky, and the estimation error is often significant.
+
+The ndistinct coefficients address this by storing ndistinct estimates not
+only for individual columns, but also for (all) combinations of columns.
+So for example given three columns (a,b,c) the statistics will estimate
+ndistinct for (a,b), (a,c), (b,c) and (a,b,c). The per-column estimates
+are already available in pg_statistic.
+
+
+GROUP BY estimation (estimate_num_groups)
+-----------------------------------------
+
+Although ndistinct coefficient might be used for selectivity estimation
+(of equality conditions in WHERE clause), that is not implemented at this
+point.
+
+Instead, ndistinct coefficients are only used in estimate_num_groups() to
+estimate grouped queries.
diff --git a/src/backend/utils/mvstats/README.stats b/src/backend/utils/mvstats/README.stats
new file mode 100644
index 0000000..865a75d
--- /dev/null
+++ b/src/backend/utils/mvstats/README.stats
@@ -0,0 +1,98 @@
+Multivariate statististics
+==========================
+
+When estimating various quantities (e.g. condition selectivities) the default
+approach relies on the assumption of independence. In practice that's often
+not true, resulting in estimation errors.
+
+Multivariate stats track different types of dependencies between the columns,
+hopefully improving the estimates.
+
+
+Types of statistics
+-------------------
+
+Currently we only have two kinds of multivariate statistics
+
+ (a) soft functional dependencies (README.dependencies)
+
+ (b) ndistinct coefficients
+
+
+Compatible clause types
+-----------------------
+
+Each type of statistics may be used to estimate some subset of clause types.
+
+ (a) functional dependencies - equality clauses (AND), possibly IS NULL
+
+Currently only simple operator clauses (Var op Const) are supported, but it's
+possible to support more complex clause types, e.g. (Var op Var).
+
+
+Complex clauses
+---------------
+
+We also support estimating more complex clauses - essentially AND/OR clauses
+with (Var op Const) as leaves, as long as all the referenced attributes are
+covered by a single statistics.
+
+For example this condition
+
+ (a=1) AND ((b=2) OR ((c=3) AND (d=4)))
+
+may be estimated using statistics on (a,b,c,d). If we only have statistics on
+(b,c,d) we may estimate the second part, and estimate (a=1) using simple stats.
+
+If we only have statistics on (a,b,c) we can't apply it at all at this point,
+but it's worth pointing out clauselist_selectivity() works recursively and when
+handling the second part (the OR-clause), we'll be able to apply the statistics.
+
+Note: The multi-statistics estimation patch also makes it possible to pass some
+clauses as 'conditions' into the deeper parts of the expression tree.
+
+
+Selectivity estimation
+----------------------
+
+When estimating selectivity, we aim to achieve several things:
+
+ (a) maximize the estimate accuracy
+
+ (b) minimize the overhead, especially when no suitable multivariate stats
+ exist (so if you are not using multivariate stats, there's no overhead)
+
+This clauselist_selectivity() performs several inexpensive checks first, before
+even attempting to do the more expensive estimation.
+
+ (1) check if there are multivariate stats on the relation
+
+ (2) check there are at least two attributes referenced by clauses compatible
+ with multivariate statistics (equality clauses for func. dependencies)
+
+ (3) perform reduction of equality clauses using func. dependencies
+
+ (4) estimate the reduced list of clauses using regular statistics
+
+Whenever we find there are no suitable stats, we skip the expensive steps.
+
+
+Size of sample in ANALYZE
+-------------------------
+When performing ANALYZE, the number of rows to sample is determined as
+
+ (300 * statistics_target)
+
+That works reasonably well for statistics on individual columns, but perhaps
+it's not enough for multivariate statistics. Papers analyzing estimation errors
+all use samples proportional to the table (usually finding that 1-3% of the
+table is enough to build accurate stats).
+
+The requested accuracy (number of MCV items or histogram bins) should also
+be considered when determining the sample size, and in multivariate statistics
+those are not necessarily limited by statistics_target.
+
+This however merits further discussion, because collecting the sample is quite
+expensive and increasing it further would make ANALYZE even more painful.
+Judging by the experiments with the current implementation, the fixed size
+seems to work reasonably well for now, so we leave this as a future work.
diff --git a/src/backend/utils/mvstats/common.c b/src/backend/utils/mvstats/common.c
new file mode 100644
index 0000000..c4a2644
--- /dev/null
+++ b/src/backend/utils/mvstats/common.c
@@ -0,0 +1,384 @@
+/*-------------------------------------------------------------------------
+ *
+ * common.c
+ * POSTGRES multivariate statistics
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/common.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "common.h"
+
+static VacAttrStats **lookup_var_attr_stats(int2vector *attrs,
+ int natts, VacAttrStats **vacattrstats);
+
+static List *list_mv_stats(Oid relid);
+
+
+/*
+ * Compute requested multivariate stats, using the rows sampled for the
+ * plain (single-column) stats.
+ *
+ * This fetches a list of stats from pg_mv_statistic, computes the stats
+ * and serializes them back into the catalog (as bytea values).
+ */
+void
+build_mv_stats(Relation onerel, double totalrows,
+ int numrows, HeapTuple *rows,
+ int natts, VacAttrStats **vacattrstats)
+{
+ ListCell *lc;
+ List *mvstats;
+
+ TupleDesc tupdesc = RelationGetDescr(onerel);
+
+ /*
+ * Fetch defined MV groups from pg_mv_statistic, and then compute the MV
+ * statistics (histograms for now).
+ */
+ mvstats = list_mv_stats(RelationGetRelid(onerel));
+
+ foreach(lc, mvstats)
+ {
+ int j;
+ MVStatisticInfo *stat = (MVStatisticInfo *) lfirst(lc);
+ MVNDistinct ndistinct = NULL;
+
+ VacAttrStats **stats = NULL;
+ int numatts = 0;
+
+ /* int2 vector of attnums the stats should be computed on */
+ int2vector *attrs = stat->stakeys;
+
+ /* see how many of the columns are not dropped */
+ for (j = 0; j < attrs->dim1; j++)
+ if (!tupdesc->attrs[attrs->values[j] - 1]->attisdropped)
+ numatts += 1;
+
+ /* if there are dropped attributes, build a filtered int2vector */
+ if (numatts != attrs->dim1)
+ {
+ int16 *tmp = palloc0(numatts * sizeof(int16));
+ int attnum = 0;
+
+ for (j = 0; j < attrs->dim1; j++)
+ if (!tupdesc->attrs[attrs->values[j] - 1]->attisdropped)
+ tmp[attnum++] = attrs->values[j];
+
+ pfree(attrs);
+ attrs = buildint2vector(tmp, numatts);
+ }
+
+ /* filter only the interesting vacattrstats records */
+ stats = lookup_var_attr_stats(attrs, natts, vacattrstats);
+
+ /* check allowed number of dimensions */
+ Assert((attrs->dim1 >= 2) && (attrs->dim1 <= MVSTATS_MAX_DIMENSIONS));
+
+ /* compute ndistinct coefficients */
+ if (stat->ndist_enabled)
+ ndistinct = build_mv_ndistinct(totalrows, numrows, rows, attrs, stats);
+
+ /* store the statistics in the catalog */
+ update_mv_stats(stat->mvoid, ndistinct, attrs, stats);
+ }
+}
+
+/*
+ * Lookup the VacAttrStats info for the selected columns, with indexes
+ * matching the attrs vector (to make it easy to work with when
+ * computing multivariate stats).
+ */
+static VacAttrStats **
+lookup_var_attr_stats(int2vector *attrs, int natts, VacAttrStats **vacattrstats)
+{
+ int i,
+ j;
+ int numattrs = attrs->dim1;
+ VacAttrStats **stats = (VacAttrStats **) palloc0(numattrs * sizeof(VacAttrStats *));
+
+ /* lookup VacAttrStats info for the requested columns (same attnum) */
+ for (i = 0; i < numattrs; i++)
+ {
+ stats[i] = NULL;
+ for (j = 0; j < natts; j++)
+ {
+ if (attrs->values[i] == vacattrstats[j]->tupattnum)
+ {
+ stats[i] = vacattrstats[j];
+ break;
+ }
+ }
+
+ /*
+ * Check that we found the info, that the attnum matches and that
+ * there's the requested 'lt' operator and that the type is
+ * 'passed-by-value'.
+ */
+ Assert(stats[i] != NULL);
+ Assert(stats[i]->tupattnum == attrs->values[i]);
+
+ /*
+ * FIXME This is rather ugly way to check for 'ltopr' (which is
+ * defined for 'scalar' attributes).
+ */
+ Assert(((StdAnalyzeData *) stats[i]->extra_data)->ltopr != InvalidOid);
+ }
+
+ return stats;
+}
+
+/*
+ * Fetch list of MV stats defined on a table, without the actual data
+ * for histograms, MCV lists etc.
+ */
+static List *
+list_mv_stats(Oid relid)
+{
+ Relation indrel;
+ SysScanDesc indscan;
+ ScanKeyData skey;
+ HeapTuple htup;
+ List *result = NIL;
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ ScanKeyInit(&skey,
+ Anum_pg_mv_statistic_starelid,
+ BTEqualStrategyNumber, F_OIDEQ,
+ ObjectIdGetDatum(relid));
+
+ indrel = heap_open(MvStatisticRelationId, AccessShareLock);
+ indscan = systable_beginscan(indrel, MvStatisticRelidIndexId, true,
+ NULL, 1, &skey);
+
+ while (HeapTupleIsValid(htup = systable_getnext(indscan)))
+ {
+ MVStatisticInfo *info = makeNode(MVStatisticInfo);
+ Form_pg_mv_statistic stats = (Form_pg_mv_statistic) GETSTRUCT(htup);
+
+ info->mvoid = HeapTupleGetOid(htup);
+ info->stakeys = buildint2vector(stats->stakeys.values, stats->stakeys.dim1);
+ info->ndist_enabled = stats->ndist_enabled;
+ info->ndist_built = stats->ndist_built;
+
+ result = lappend(result, info);
+ }
+
+ systable_endscan(indscan);
+
+ heap_close(indrel, AccessShareLock);
+
+ /*
+ * TODO maybe save the list into relcache, as in RelationGetIndexList
+ * (which was used as an inspiration of this one)?.
+ */
+
+ return result;
+}
+
+void
+update_mv_stats(Oid mvoid, MVNDistinct ndistinct,
+ int2vector *attrs, VacAttrStats **stats)
+{
+ HeapTuple stup,
+ oldtup;
+ Datum values[Natts_pg_mv_statistic];
+ bool nulls[Natts_pg_mv_statistic];
+ bool replaces[Natts_pg_mv_statistic];
+
+ Relation sd = heap_open(MvStatisticRelationId, RowExclusiveLock);
+
+ memset(nulls, 1, Natts_pg_mv_statistic * sizeof(bool));
+ memset(replaces, 0, Natts_pg_mv_statistic * sizeof(bool));
+ memset(values, 0, Natts_pg_mv_statistic * sizeof(Datum));
+
+ /*
+ * Construct a new pg_mv_statistic tuple - replace only the histogram and
+ * MCV list, depending whether it actually was computed.
+ */
+ if (ndistinct != NULL)
+ {
+ bytea *data = serialize_mv_ndistinct(ndistinct);
+
+ nulls[Anum_pg_mv_statistic_standist -1] = (data == NULL);
+ values[Anum_pg_mv_statistic_standist-1] = PointerGetDatum(data);
+ }
+
+ /* always replace the value (either by bytea or NULL) */
+ replaces[Anum_pg_mv_statistic_standist - 1] = true;
+
+ /* always change the availability flags */
+ nulls[Anum_pg_mv_statistic_ndist_built - 1] = false;
+ nulls[Anum_pg_mv_statistic_stakeys - 1] = false;
+
+ /* use the new attnums, in case we removed some dropped ones */
+ replaces[Anum_pg_mv_statistic_ndist_built - 1] = true;
+ replaces[Anum_pg_mv_statistic_stakeys - 1] = true;
+
+ values[Anum_pg_mv_statistic_ndist_built - 1] = BoolGetDatum(ndistinct != NULL);
+
+ values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(attrs);
+
+ /* Is there already a pg_mv_statistic tuple for this attribute? */
+ oldtup = SearchSysCache1(MVSTATOID,
+ ObjectIdGetDatum(mvoid));
+
+ if (HeapTupleIsValid(oldtup))
+ {
+ /* Yes, replace it */
+ stup = heap_modify_tuple(oldtup,
+ RelationGetDescr(sd),
+ values,
+ nulls,
+ replaces);
+ ReleaseSysCache(oldtup);
+ simple_heap_update(sd, &stup->t_self, stup);
+ }
+ else
+ elog(ERROR, "invalid pg_mv_statistic record (oid=%d)", mvoid);
+
+ /* update indexes too */
+ CatalogUpdateIndexes(sd, stup);
+
+ heap_freetuple(stup);
+
+ heap_close(sd, RowExclusiveLock);
+}
+
+/* multi-variate stats comparator */
+
+/*
+ * qsort_arg comparator for sorting Datums (MV stats)
+ *
+ * This does not maintain the tupnoLink array.
+ */
+int
+compare_scalars_simple(const void *a, const void *b, void *arg)
+{
+ Datum da = *(Datum *) a;
+ Datum db = *(Datum *) b;
+ SortSupport ssup = (SortSupport) arg;
+
+ return ApplySortComparator(da, false, db, false, ssup);
+}
+
+/*
+ * qsort_arg comparator for sorting data when partitioning a MV bucket
+ */
+int
+compare_scalars_partition(const void *a, const void *b, void *arg)
+{
+ Datum da = ((ScalarItem *) a)->value;
+ Datum db = ((ScalarItem *) b)->value;
+ SortSupport ssup = (SortSupport) arg;
+
+ return ApplySortComparator(da, false, db, false, ssup);
+}
+
+/* initialize multi-dimensional sort */
+MultiSortSupport
+multi_sort_init(int ndims)
+{
+ MultiSortSupport mss;
+
+ Assert(ndims >= 2);
+
+ mss = (MultiSortSupport) palloc0(offsetof(MultiSortSupportData, ssup)
+ +sizeof(SortSupportData) * ndims);
+
+ mss->ndims = ndims;
+
+ return mss;
+}
+
+/*
+ * add sort into for dimension 'dim' (index into vacattrstats) to mss,
+ * at the position 'sortattr'
+ */
+void
+multi_sort_add_dimension(MultiSortSupport mss, int sortdim,
+ int dim, VacAttrStats **vacattrstats)
+{
+ /* first, lookup StdAnalyzeData for the dimension (attribute) */
+ SortSupportData ssup;
+ StdAnalyzeData *tmp = (StdAnalyzeData *) vacattrstats[dim]->extra_data;
+
+ Assert(mss != NULL);
+ Assert(sortdim < mss->ndims);
+
+ /* initialize sort support, etc. */
+ memset(&ssup, 0, sizeof(ssup));
+ ssup.ssup_cxt = CurrentMemoryContext;
+
+ /* We always use the default collation for statistics */
+ ssup.ssup_collation = DEFAULT_COLLATION_OID;
+ ssup.ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(tmp->ltopr, &ssup);
+
+ mss->ssup[sortdim] = ssup;
+}
+
+/* compare all the dimensions in the selected order */
+int
+multi_sort_compare(const void *a, const void *b, void *arg)
+{
+ int i;
+ SortItem *ia = (SortItem *) a;
+ SortItem *ib = (SortItem *) b;
+
+ MultiSortSupport mss = (MultiSortSupport) arg;
+
+ for (i = 0; i < mss->ndims; i++)
+ {
+ int compare;
+
+ compare = ApplySortComparator(ia->values[i], ia->isnull[i],
+ ib->values[i], ib->isnull[i],
+ &mss->ssup[i]);
+
+ if (compare != 0)
+ return compare;
+
+ }
+
+ /* equal by default */
+ return 0;
+}
+
+/* compare selected dimension */
+int
+multi_sort_compare_dim(int dim, const SortItem *a, const SortItem *b,
+ MultiSortSupport mss)
+{
+ return ApplySortComparator(a->values[dim], a->isnull[dim],
+ b->values[dim], b->isnull[dim],
+ &mss->ssup[dim]);
+}
+
+int
+multi_sort_compare_dims(int start, int end,
+ const SortItem *a, const SortItem *b,
+ MultiSortSupport mss)
+{
+ int dim;
+
+ for (dim = start; dim <= end; dim++)
+ {
+ int r = ApplySortComparator(a->values[dim], a->isnull[dim],
+ b->values[dim], b->isnull[dim],
+ &mss->ssup[dim]);
+
+ if (r != 0)
+ return r;
+ }
+
+ return 0;
+}
diff --git a/src/backend/utils/mvstats/common.h b/src/backend/utils/mvstats/common.h
new file mode 100644
index 0000000..e471c88
--- /dev/null
+++ b/src/backend/utils/mvstats/common.h
@@ -0,0 +1,80 @@
+/*-------------------------------------------------------------------------
+ *
+ * common.h
+ * POSTGRES multivariate statistics
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/common.h
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "access/sysattr.h"
+#include "access/tuptoaster.h"
+#include "catalog/indexing.h"
+#include "catalog/pg_collation.h"
+#include "catalog/pg_mv_statistic.h"
+#include "foreign/fdwapi.h"
+#include "postmaster/autovacuum.h"
+#include "storage/lmgr.h"
+#include "utils/builtins.h"
+#include "utils/datum.h"
+#include "utils/fmgroids.h"
+#include "utils/mvstats.h"
+#include "utils/sortsupport.h"
+#include "utils/syscache.h"
+
+
+/* FIXME private structure copied from analyze.c */
+
+typedef struct
+{
+ Oid eqopr; /* '=' operator for datatype, if any */
+ Oid eqfunc; /* and associated function */
+ Oid ltopr; /* '<' operator for datatype, if any */
+} StdAnalyzeData;
+
+typedef struct
+{
+ Datum value; /* a data value */
+ int tupno; /* position index for tuple it came from */
+} ScalarItem;
+
+/* multi-sort */
+typedef struct MultiSortSupportData
+{
+ int ndims; /* number of dimensions supported by the */
+ SortSupportData ssup[1]; /* sort support data for each dimension */
+} MultiSortSupportData;
+
+typedef MultiSortSupportData *MultiSortSupport;
+
+typedef struct SortItem
+{
+ Datum *values;
+ bool *isnull;
+} SortItem;
+
+MultiSortSupport multi_sort_init(int ndims);
+
+void multi_sort_add_dimension(MultiSortSupport mss, int sortdim,
+ int dim, VacAttrStats **vacattrstats);
+
+int multi_sort_compare(const void *a, const void *b, void *arg);
+
+int multi_sort_compare_dim(int dim, const SortItem *a,
+ const SortItem *b, MultiSortSupport mss);
+
+int multi_sort_compare_dims(int start, int end, const SortItem *a,
+ const SortItem *b, MultiSortSupport mss);
+
+/* comparators, used when constructing multivariate stats */
+int compare_scalars_simple(const void *a, const void *b, void *arg);
+int compare_scalars_partition(const void *a, const void *b, void *arg);
diff --git a/src/backend/utils/mvstats/mvdist.c b/src/backend/utils/mvstats/mvdist.c
new file mode 100644
index 0000000..9cb0a00
--- /dev/null
+++ b/src/backend/utils/mvstats/mvdist.c
@@ -0,0 +1,585 @@
+/*-------------------------------------------------------------------------
+ *
+ * mvdist.c
+ * POSTGRES multivariate distinct coefficients
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/mvdist.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include <math.h>
+
+#include "common.h"
+#include "utils/bytea.h"
+#include "utils/lsyscache.h"
+
+static double estimate_ndistinct(double totalrows, int numrows, int d, int f1);
+
+/* internal state for generator of k-combinations of n elements */
+typedef struct CombinationGeneratorData
+{
+
+ int k; /* size of the combination */
+ int current; /* index of the next combination to return */
+
+ int ncombinations; /* number of combinations (size of array) */
+ int *combinations; /* array of pre-built combinations */
+
+} CombinationGeneratorData;
+
+typedef CombinationGeneratorData *CombinationGenerator;
+
+/* generator API */
+static CombinationGenerator generator_init(int2vector *attrs, int k);
+static void generator_free(CombinationGenerator state);
+static int *generator_next(CombinationGenerator state, int2vector *attrs);
+
+static int n_choose_k(int n, int k);
+static int num_combinations(int n);
+static double ndistinct_for_combination(double totalrows, int numrows,
+ HeapTuple *rows, int2vector *attrs, VacAttrStats **stats,
+ int k, int *combination);
+
+/*
+ * Compute ndistinct coefficient for the combination of attributes. This
+ * computes the ndistinct estimate using the same estimator used in analyze.c
+ * and then computes the coefficient.
+ */
+MVNDistinct
+build_mv_ndistinct(double totalrows, int numrows, HeapTuple *rows,
+ int2vector *attrs, VacAttrStats **stats)
+{
+ int i, k;
+ int numattrs = attrs->dim1;
+ int numcombs = num_combinations(numattrs);
+
+ MVNDistinct result;
+
+ result = palloc0(offsetof(MVNDistinctData, items) +
+ numcombs * sizeof(MVNDistinctItem));
+
+ result->nitems = numcombs;
+
+ i = 0;
+ for (k = 2; k <= numattrs; k++)
+ {
+ int * combination;
+ CombinationGenerator generator;
+
+ generator = generator_init(attrs, k);
+
+ while ((combination = generator_next(generator, attrs)))
+ {
+ MVNDistinctItem *item = &result->items[i++];
+
+ item->nattrs = k;
+ item->ndistinct = ndistinct_for_combination(totalrows, numrows, rows,
+ attrs, stats, k, combination);
+
+ item->attrs = palloc(k * sizeof(int));
+ memcpy(item->attrs, combination, k * sizeof(int));
+
+ /* must not overflow the output array */
+ Assert(i <= result->nitems);
+ }
+
+ generator_free(generator);
+ }
+
+ /* must consume exactly the whole output array */
+ Assert(i == result->nitems);
+
+ return result;
+}
+
+static double
+ndistinct_for_combination(double totalrows, int numrows, HeapTuple *rows,
+ int2vector *attrs, VacAttrStats **stats,
+ int k, int *combination)
+{
+ int i, j;
+ int f1, cnt, d;
+ int nmultiple, summultiple;
+ MultiSortSupport mss = multi_sort_init(k);
+
+ /*
+ * It's possible to sort the sample rows directly, but this seemed
+ * somehow simpler / less error prone. Another option would be to
+ * allocate the arrays for each SortItem separately, but that'd be
+ * significant overhead (not just CPU, but especially memory bloat).
+ */
+ SortItem * items = (SortItem*)palloc0(numrows * sizeof(SortItem));
+
+ Datum *values = (Datum*)palloc0(sizeof(Datum) * numrows * k);
+ bool *isnull = (bool*)palloc0(sizeof(bool) * numrows * k);
+
+ Assert((k >= 2) && (k <= attrs->dim1));
+
+ for (i = 0; i < numrows; i++)
+ {
+ items[i].values = &values[i * k];
+ items[i].isnull = &isnull[i * k];
+ }
+
+ for (i = 0; i < k; i++)
+ {
+ /* prepare the sort function for the first dimension */
+ multi_sort_add_dimension(mss, i, combination[i], stats);
+
+ /* accumulate all the data into the array and sort it */
+ for (j = 0; j < numrows; j++)
+ {
+ items[j].values[i]
+ = heap_getattr(rows[j], attrs->values[combination[i]],
+ stats[combination[i]]->tupDesc,
+ &items[j].isnull[i]);
+ }
+ }
+
+ qsort_arg((void *) items, numrows, sizeof(SortItem),
+ multi_sort_compare, mss);
+
+ /* count number of distinct combinations */
+
+ f1 = 0;
+ cnt = 1;
+ d = 1;
+ for (i = 1; i < numrows; i++)
+ {
+ if (multi_sort_compare(&items[i], &items[i-1], mss) != 0)
+ {
+ if (cnt == 1)
+ f1 += 1;
+ else
+ {
+ nmultiple += 1;
+ summultiple += cnt;
+ }
+
+ d++;
+ cnt = 0;
+ }
+
+ cnt += 1;
+ }
+
+ if (cnt == 1)
+ f1 += 1;
+ else
+ {
+ nmultiple += 1;
+ summultiple += cnt;
+ }
+
+ return estimate_ndistinct(totalrows, numrows, d, f1);
+}
+
+MVNDistinct
+load_mv_ndistinct(Oid mvoid)
+{
+ bool isnull = false;
+ Datum ndist;
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ HeapTuple htup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(mvoid));
+
+#ifdef USE_ASSERT_CHECKING
+ Form_pg_mv_statistic mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
+ Assert(mvstat->ndist_enabled && mvstat->ndist_built);
+#endif
+
+ ndist = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_standist, &isnull);
+
+ Assert(!isnull);
+
+ ReleaseSysCache(htup);
+
+ return deserialize_mv_ndistinct(DatumGetByteaP(ndist));
+}
+
+/* The Duj1 estimator (already used in analyze.c). */
+static double
+estimate_ndistinct(double totalrows, int numrows, int d, int f1)
+{
+ double numer,
+ denom,
+ ndistinct;
+
+ numer = (double) numrows *(double) d;
+
+ denom = (double) (numrows - f1) +
+ (double) f1 * (double) numrows / totalrows;
+
+ ndistinct = numer / denom;
+
+ /* Clamp to sane range in case of roundoff error */
+ if (ndistinct < (double) d)
+ ndistinct = (double) d;
+
+ if (ndistinct > totalrows)
+ ndistinct = totalrows;
+
+ return floor(ndistinct + 0.5);
+}
+
+
+/*
+ * pg_ndistinct_in - input routine for type pg_ndistinct.
+ *
+ * pg_ndistinct is real enough to be a table column, but it has no operations
+ * of its own, and disallows input too
+ *
+ * XXX This is inspired by what pg_node_tree does.
+ */
+Datum
+pg_ndistinct_in(PG_FUNCTION_ARGS)
+{
+ /*
+ * pg_node_list stores the data in binary form and parsing text input is
+ * not needed, so disallow this.
+ */
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_ndistinct")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_ndistinct - output routine for type pg_ndistinct.
+ *
+ * histograms are serialized into a bytea value, so we simply call byteaout()
+ * to serialize the value into text. But it'd be nice to serialize that into
+ * a meaningful representation (e.g. for inspection by people).
+ */
+Datum
+pg_ndistinct_out(PG_FUNCTION_ARGS)
+{
+ int i, j;
+ char *ret;
+ StringInfoData str;
+
+ bytea *data = PG_GETARG_BYTEA_PP(0);
+
+ MVNDistinct ndist = deserialize_mv_ndistinct(data);
+
+ initStringInfo(&str);
+ appendStringInfoString(&str, "[");
+
+ for (i = 0; i < ndist->nitems; i++)
+ {
+ MVNDistinctItem item = ndist->items[i];
+
+ if (i > 0)
+ appendStringInfoString(&str, ", ");
+
+ appendStringInfoString(&str, "{");
+
+ for (j = 0; j < item.nattrs; j++)
+ {
+ if (j > 0)
+ appendStringInfoString(&str, ", ");
+
+ appendStringInfo(&str, "%d", item.attrs[j]);
+ }
+
+ appendStringInfo(&str, ", %f", item.ndistinct);
+
+ appendStringInfoString(&str, "}");
+ }
+
+ appendStringInfoString(&str, "]");
+
+ ret = pstrdup(str.data);
+ pfree(str.data);
+
+ PG_RETURN_CSTRING(ret);
+}
+
+/*
+ * pg_ndistinct_recv - binary input routine for type pg_ndistinct.
+ */
+Datum
+pg_ndistinct_recv(PG_FUNCTION_ARGS)
+{
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_ndistinct")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_ndistinct_send - binary output routine for type pg_ndistinct.
+ *
+ * XXX Histograms are serialized into a bytea value, so let's just send that.
+ */
+Datum
+pg_ndistinct_send(PG_FUNCTION_ARGS)
+{
+ return byteasend(fcinfo);
+}
+
+static int
+n_choose_k(int n, int k)
+{
+ int i, numer, denom;
+
+ Assert((n > 0) && (k > 0) && (n >= k));
+
+ numer = denom = 1;
+ for (i = 1; i <= k; i++)
+ {
+ numer *= (n - i + 1);
+ denom *= i;
+ }
+
+ Assert(numer % denom == 0);
+
+ return numer / denom;
+}
+
+static int
+num_combinations(int n)
+{
+ int k;
+ int ncombs = 0;
+
+ /* ignore combinations with a single column */
+ for (k = 2; k <= n; k++)
+ ncombs += n_choose_k(n, k);
+
+ return ncombs;
+}
+
+/*
+ * generate all combinations (k elements from n)
+ */
+static void
+generate_combinations_recurse(CombinationGenerator state,
+ int n, int index, int start, int *current)
+{
+ /* If we haven't filled all the elements, simply recurse. */
+ if (index < state->k)
+ {
+ int i;
+
+ /*
+ * The values have to be in ascending order, so make sure we start
+ * with the value passed by parameter.
+ */
+
+ for (i = start; i < n; i++)
+ {
+ current[index] = i;
+ generate_combinations_recurse(state, n, (index+1), (i+1), current);
+ }
+
+ return;
+ }
+ else
+ {
+ /* we got a correct combination */
+ state->combinations = (int*)repalloc(state->combinations,
+ state->k * (state->current + 1) * sizeof(int));
+ memcpy(&state->combinations[(state->k * state->current)],
+ current, state->k * sizeof(int));
+ state->current++;
+ }
+}
+
+/* generate all k-combinations of n elements */
+static void
+generate_combinations(CombinationGenerator state, int n)
+{
+ int *current = (int *) palloc0(sizeof(int) * state->k);
+
+ generate_combinations_recurse(state, n, 0, 0, current);
+
+ pfree(current);
+}
+
+/*
+ * initialize the generator of combinations, and prebuild them.
+ *
+ * This pre-builds all the combinations. We could also generate them in
+ * generator_next(), but this seems simpler.
+ */
+static CombinationGenerator
+generator_init(int2vector *attrs, int k)
+{
+ int n = attrs->dim1;
+ CombinationGenerator state;
+
+ Assert((n >= k) && (k > 0));
+
+ /* allocate the generator state as a single chunk of memory */
+ state = (CombinationGenerator) palloc0(sizeof(CombinationGeneratorData));
+ state->combinations = (int*)palloc(k * sizeof(int));
+
+ state->ncombinations = n_choose_k(n, k);
+ state->current = 0;
+ state->k = k;
+
+ /* now actually pre-generate all the combinations */
+ generate_combinations(state, n);
+
+ /* make sure we got the expected number of combinations */
+ Assert(state->current == state->ncombinations);
+
+ /* reset the number, so we start with the first one */
+ state->current = 0;
+
+ return state;
+}
+
+/* free the generator state */
+static void
+generator_free(CombinationGenerator state)
+{
+ /* we've allocated a single chunk, so just free it */
+ pfree(state);
+}
+
+/* generate next combination */
+static int *
+generator_next(CombinationGenerator state, int2vector *attrs)
+{
+ if (state->current == state->ncombinations)
+ return NULL;
+
+ return &state->combinations[state->k * state->current++];
+}
+
+/*
+ * serialize list of ndistinct items into a bytea
+ */
+bytea *
+serialize_mv_ndistinct(MVNDistinct ndistinct)
+{
+ int i;
+ bytea *output;
+ char *tmp;
+
+ /* we need to store nitems */
+ Size len = VARHDRSZ + offsetof(MVNDistinctData, items) +
+ ndistinct->nitems * offsetof(MVNDistinctItem, attrs);
+
+ /* and also include space for the actual attribute numbers */
+ for (i = 0; i < ndistinct->nitems; i++)
+ len += (sizeof(int) * ndistinct->items[i].nattrs);
+
+ output = (bytea *) palloc0(len);
+ SET_VARSIZE(output, len);
+
+ tmp = VARDATA(output);
+
+ ndistinct->magic = MVSTAT_NDISTINCT_MAGIC;
+ ndistinct->type = MVSTAT_NDISTINCT_TYPE_BASIC;
+
+ /* first, store the number of items */
+ memcpy(tmp, ndistinct, offsetof(MVNDistinctData, items));
+ tmp += offsetof(MVNDistinctData, items);
+
+ /* store number of attributes and attribute numbers for each ndistinct entry */
+ for (i = 0; i < ndistinct->nitems; i++)
+ {
+ MVNDistinctItem item = ndistinct->items[i];
+
+ memcpy(tmp, &item, offsetof(MVNDistinctItem, attrs));
+ tmp += offsetof(MVNDistinctItem, attrs);
+
+ memcpy(tmp, item.attrs, sizeof(int) * item.nattrs);
+ tmp += sizeof(int) * item.nattrs;
+
+ Assert(tmp <= ((char *) output + len));
+ }
+
+ return output;
+}
+
+/*
+ * Reads serialized ndistinct into MVNDistinct structure.
+ */
+MVNDistinct
+deserialize_mv_ndistinct(bytea *data)
+{
+ int i;
+ Size expected_size;
+ MVNDistinct ndistinct;
+ char *tmp;
+
+ if (data == NULL)
+ return NULL;
+
+ if (VARSIZE_ANY_EXHDR(data) < offsetof(MVNDistinctData, items))
+ elog(ERROR, "invalid MVNDistinct size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), offsetof(MVNDistinctData, items));
+
+ /* read the MVNDistinct header */
+ ndistinct = (MVNDistinct) palloc0(sizeof(MVNDistinctData));
+
+ /* initialize pointer to the data part (skip the varlena header) */
+ tmp = VARDATA_ANY(data);
+
+ /* get the header and perform basic sanity checks */
+ memcpy(ndistinct, tmp, offsetof(MVNDistinctData, items));
+ tmp += offsetof(MVNDistinctData, items);
+
+ if (ndistinct->magic != MVSTAT_NDISTINCT_MAGIC)
+ elog(ERROR, "invalid ndistinct magic %d (expected %dd)",
+ ndistinct->magic, MVSTAT_NDISTINCT_MAGIC);
+
+ if (ndistinct->type != MVSTAT_NDISTINCT_TYPE_BASIC)
+ elog(ERROR, "invalid ndistinct type %d (expected %dd)",
+ ndistinct->type, MVSTAT_NDISTINCT_TYPE_BASIC);
+
+ Assert(ndistinct->nitems > 0);
+
+ /* what minimum bytea size do we expect for those parameters */
+ expected_size = offsetof(MVNDistinctData, items) +
+ ndistinct->nitems * (offsetof(MVNDistinctItem, attrs) + sizeof(int) * 2);
+
+ if (VARSIZE_ANY_EXHDR(data) < expected_size)
+ elog(ERROR, "invalid dependencies size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ /* allocate space for the ndistinct items */
+ ndistinct = repalloc(ndistinct, offsetof(MVNDistinctData, items) +
+ (ndistinct->nitems * sizeof(MVNDistinctItem)));
+
+ for (i = 0; i < ndistinct->nitems; i++)
+ {
+ MVNDistinctItem *item = &ndistinct->items[i];
+
+ /* number of attributes */
+ memcpy(item, tmp, offsetof(MVNDistinctItem, attrs));
+ tmp += offsetof(MVNDistinctItem, attrs);
+
+ /* is the number of attributes valid? */
+ Assert((item->nattrs >= 2) && (item->nattrs <= MVSTATS_MAX_DIMENSIONS));
+
+ /* now that we know the number of attributes, allocate the attribute */
+ item->attrs = (int*)palloc0(item->nattrs * sizeof(int));
+
+ /* copy attribute numbers */
+ memcpy(item->attrs, tmp, sizeof(int) * item->nattrs);
+ tmp += sizeof(int) * item->nattrs;
+
+ /* still within the bytea */
+ Assert(tmp <= ((char *) data + VARSIZE_ANY(data)));
+ }
+
+ /* we should have consumed the whole bytea exactly */
+ Assert(tmp == ((char *) data + VARSIZE_ANY(data)));
+
+ return ndistinct;
+}
diff --git a/src/bin/psql/describe.c b/src/bin/psql/describe.c
index a582a37..1f1050b 100644
--- a/src/bin/psql/describe.c
+++ b/src/bin/psql/describe.c
@@ -2293,6 +2293,50 @@ describeOneTableDetails(const char *schemaname,
PQclear(result);
}
+ /* print any multivariate statistics */
+ if (pset.sversion >= 90600)
+ {
+ printfPQExpBuffer(&buf,
+ "SELECT oid, stanamespace::regnamespace AS nsp, staname, stakeys,\n"
+ " ndist_enabled,\n"
+ " ndist_built,\n"
+ " (SELECT string_agg(attname::text,', ')\n"
+ " FROM ((SELECT unnest(stakeys) AS attnum) s\n"
+ " JOIN pg_attribute a ON (starelid = a.attrelid and a.attnum = s.attnum))) AS attnums\n"
+ "FROM pg_mv_statistic stat WHERE starelid = '%s' ORDER BY 1;",
+ oid);
+
+ result = PSQLexec(buf.data);
+ if (!result)
+ goto error_return;
+ else
+ tuples = PQntuples(result);
+
+ if (tuples > 0)
+ {
+ printTableAddFooter(&cont, _("Statistics:"));
+ for (i = 0; i < tuples; i++)
+ {
+ printfPQExpBuffer(&buf, " ");
+
+ /* statistics name (qualified with namespace) */
+ appendPQExpBuffer(&buf, "\"%s.%s\" ",
+ PQgetvalue(result, i, 1),
+ PQgetvalue(result, i, 2));
+
+ /* options */
+ if (!strcmp(PQgetvalue(result, i, 4), "t"))
+ appendPQExpBuffer(&buf, "(dependencies)");
+
+ appendPQExpBuffer(&buf, " ON (%s)",
+ PQgetvalue(result, i, 6));
+
+ printTableAddFooter(&cont, buf.data);
+ }
+ }
+ PQclear(result);
+ }
+
/* print rules */
if (tableinfo.hasrules && tableinfo.relkind != 'm')
{
diff --git a/src/include/catalog/dependency.h b/src/include/catalog/dependency.h
index e8a302f..76c81a7 100644
--- a/src/include/catalog/dependency.h
+++ b/src/include/catalog/dependency.h
@@ -161,10 +161,11 @@ typedef enum ObjectClass
OCLASS_EXTENSION, /* pg_extension */
OCLASS_EVENT_TRIGGER, /* pg_event_trigger */
OCLASS_POLICY, /* pg_policy */
- OCLASS_TRANSFORM /* pg_transform */
+ OCLASS_TRANSFORM, /* pg_transform */
+ OCLASS_STATISTICS /* pg_mv_statistics */
} ObjectClass;
-#define LAST_OCLASS OCLASS_TRANSFORM
+#define LAST_OCLASS OCLASS_STATISTICS
/* flag bits for performDeletion/performMultipleDeletions: */
#define PERFORM_DELETION_INTERNAL 0x0001 /* internal action */
diff --git a/src/include/catalog/heap.h b/src/include/catalog/heap.h
index 0e4262f..3a3a200 100644
--- a/src/include/catalog/heap.h
+++ b/src/include/catalog/heap.h
@@ -119,6 +119,7 @@ extern void RemoveAttrDefault(Oid relid, AttrNumber attnum,
DropBehavior behavior, bool complain, bool internal);
extern void RemoveAttrDefaultById(Oid attrdefId);
extern void RemoveStatistics(Oid relid, AttrNumber attnum);
+extern void RemoveMVStatistics(Oid relid, AttrNumber attnum);
extern Form_pg_attribute SystemAttributeDefinition(AttrNumber attno,
bool relhasoids);
diff --git a/src/include/catalog/indexing.h b/src/include/catalog/indexing.h
index 293985d..7fb09a4 100644
--- a/src/include/catalog/indexing.h
+++ b/src/include/catalog/indexing.h
@@ -176,6 +176,13 @@ DECLARE_UNIQUE_INDEX(pg_largeobject_loid_pn_index, 2683, on pg_largeobject using
DECLARE_UNIQUE_INDEX(pg_largeobject_metadata_oid_index, 2996, on pg_largeobject_metadata using btree(oid oid_ops));
#define LargeObjectMetadataOidIndexId 2996
+DECLARE_UNIQUE_INDEX(pg_mv_statistic_oid_index, 3380, on pg_mv_statistic using btree(oid oid_ops));
+#define MvStatisticOidIndexId 3380
+DECLARE_UNIQUE_INDEX(pg_mv_statistic_name_index, 3997, on pg_mv_statistic using btree(staname name_ops, stanamespace oid_ops));
+#define MvStatisticNameIndexId 3997
+DECLARE_INDEX(pg_mv_statistic_relid_index, 3379, on pg_mv_statistic using btree(starelid oid_ops));
+#define MvStatisticRelidIndexId 3379
+
DECLARE_UNIQUE_INDEX(pg_namespace_nspname_index, 2684, on pg_namespace using btree(nspname name_ops));
#define NamespaceNameIndexId 2684
DECLARE_UNIQUE_INDEX(pg_namespace_oid_index, 2685, on pg_namespace using btree(oid oid_ops));
diff --git a/src/include/catalog/namespace.h b/src/include/catalog/namespace.h
index eee94d8..b9a4db1 100644
--- a/src/include/catalog/namespace.h
+++ b/src/include/catalog/namespace.h
@@ -141,6 +141,8 @@ extern Oid get_collation_oid(List *collname, bool missing_ok);
extern Oid get_conversion_oid(List *conname, bool missing_ok);
extern Oid FindDefaultConversionProc(int32 for_encoding, int32 to_encoding);
+extern Oid get_statistics_oid(List *names, bool missing_ok);
+
/* initialization & transaction cleanup code */
extern void InitializeSearchPath(void);
extern void AtEOXact_Namespace(bool isCommit, bool parallel);
diff --git a/src/include/catalog/pg_cast.h b/src/include/catalog/pg_cast.h
index 04d11c0..d54a0f9 100644
--- a/src/include/catalog/pg_cast.h
+++ b/src/include/catalog/pg_cast.h
@@ -254,6 +254,10 @@ DATA(insert ( 23 18 78 e f ));
/* pg_node_tree can be coerced to, but not from, text */
DATA(insert ( 194 25 0 i b ));
+/* pg_ndistinct can be coerced to, but not from, bytea and text */
+DATA(insert ( 3343 17 0 i b ));
+DATA(insert ( 3343 25 0 i i ));
+
/*
* Datetime category
*/
diff --git a/src/include/catalog/pg_mv_statistic.h b/src/include/catalog/pg_mv_statistic.h
new file mode 100644
index 0000000..fad80a3
--- /dev/null
+++ b/src/include/catalog/pg_mv_statistic.h
@@ -0,0 +1,78 @@
+/*-------------------------------------------------------------------------
+ *
+ * pg_mv_statistic.h
+ * definition of the system "multivariate statistic" relation (pg_mv_statistic)
+ * along with the relation's initial contents.
+ *
+ *
+ * Portions Copyright (c) 1996-2014, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/catalog/pg_mv_statistic.h
+ *
+ * NOTES
+ * the genbki.pl script reads this file and generates .bki
+ * information from the DATA() statements.
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef PG_MV_STATISTIC_H
+#define PG_MV_STATISTIC_H
+
+#include "catalog/genbki.h"
+
+/* ----------------
+ * pg_mv_statistic definition. cpp turns this into
+ * typedef struct FormData_pg_mv_statistic
+ * ----------------
+ */
+#define MvStatisticRelationId 3381
+
+CATALOG(pg_mv_statistic,3381)
+{
+ /* These fields form the unique key for the entry: */
+ Oid starelid; /* relation containing attributes */
+ NameData staname; /* statistics name */
+ Oid stanamespace; /* OID of namespace containing this statistics */
+ Oid staowner; /* statistics owner */
+
+ /* statistics requested to build */
+ bool ndist_enabled; /* build ndist coefficient? */
+
+ /* statistics that are available (if requested) */
+ bool ndist_built; /* ndistinct coeff built */
+
+ /*
+ * variable-length fields start here, but we allow direct access to
+ * stakeys
+ */
+ int2vector stakeys; /* array of column keys */
+
+#ifdef CATALOG_VARLEN
+ pg_ndistinct standist; /* ndistinct coeff (serialized) */
+#endif
+
+} FormData_pg_mv_statistic;
+
+/* ----------------
+ * Form_pg_mv_statistic corresponds to a pointer to a tuple with
+ * the format of pg_mv_statistic relation.
+ * ----------------
+ */
+typedef FormData_pg_mv_statistic *Form_pg_mv_statistic;
+
+/* ----------------
+ * compiler constants for pg_mv_statistic
+ * ----------------
+ */
+#define Natts_pg_mv_statistic 8
+#define Anum_pg_mv_statistic_starelid 1
+#define Anum_pg_mv_statistic_staname 2
+#define Anum_pg_mv_statistic_stanamespace 3
+#define Anum_pg_mv_statistic_staowner 4
+#define Anum_pg_mv_statistic_ndist_enabled 5
+#define Anum_pg_mv_statistic_ndist_built 6
+#define Anum_pg_mv_statistic_stakeys 7
+#define Anum_pg_mv_statistic_standist 8
+
+#endif /* PG_MV_STATISTIC_H */
diff --git a/src/include/catalog/pg_proc.h b/src/include/catalog/pg_proc.h
index a6cc2eb..8bf46e2 100644
--- a/src/include/catalog/pg_proc.h
+++ b/src/include/catalog/pg_proc.h
@@ -2722,6 +2722,15 @@ DESCR("current user privilege on any column by rel name");
DATA(insert OID = 3029 ( has_any_column_privilege PGNSP PGUID 12 10 0 0 0 f f f f t f s s 2 0 16 "26 25" _null_ _null_ _null_ _null_ _null_ has_any_column_privilege_id _null_ _null_ _null_ ));
DESCR("current user privilege on any column by rel oid");
+DATA(insert OID = 3344 ( pg_ndistinct_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 3343 "2275" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_in _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3345 ( pg_ndistinct_out PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 2275 "3343" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_out _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3346 ( pg_ndistinct_recv PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 3343 "2281" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_recv _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3347 ( pg_ndistinct_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "3343" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_send _null_ _null_ _null_ ));
+DESCR("I/O");
+
DATA(insert OID = 1928 ( pg_stat_get_numscans PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_numscans _null_ _null_ _null_ ));
DESCR("statistics: number of scans done for table/index");
DATA(insert OID = 1929 ( pg_stat_get_tuples_returned PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_tuples_returned _null_ _null_ _null_ ));
diff --git a/src/include/catalog/pg_type.h b/src/include/catalog/pg_type.h
index 162239c..e6ab642 100644
--- a/src/include/catalog/pg_type.h
+++ b/src/include/catalog/pg_type.h
@@ -364,6 +364,10 @@ DATA(insert OID = 194 ( pg_node_tree PGNSP PGUID -1 f b S f t \054 0 0 0 pg_node
DESCR("string representing an internal node tree");
#define PGNODETREEOID 194
+DATA(insert OID = 3343 ( pg_ndistinct PGNSP PGUID -1 f b S f t \054 0 0 0 pg_ndistinct_in pg_ndistinct_out pg_ndistinct_recv pg_ndistinct_send - - - i x f 0 -1 0 100 _null_ _null_ _null_ ));
+DESCR("multivariate ndistinct coefficients");
+#define PGNDISTINCTOID 3343
+
DATA(insert OID = 32 ( pg_ddl_command PGNSP PGUID SIZEOF_POINTER t p P f t \054 0 0 0 pg_ddl_command_in pg_ddl_command_out pg_ddl_command_recv pg_ddl_command_send - - - ALIGNOF_POINTER p f 0 -1 0 0 _null_ _null_ _null_ ));
DESCR("internal type for passing CollectedCommand");
#define PGDDLCOMMANDOID 32
diff --git a/src/include/catalog/toasting.h b/src/include/catalog/toasting.h
index b7a38ce..e1e4945 100644
--- a/src/include/catalog/toasting.h
+++ b/src/include/catalog/toasting.h
@@ -49,6 +49,7 @@ extern void BootstrapToastTable(char *relName,
DECLARE_TOAST(pg_attrdef, 2830, 2831);
DECLARE_TOAST(pg_constraint, 2832, 2833);
DECLARE_TOAST(pg_description, 2834, 2835);
+DECLARE_TOAST(pg_mv_statistic, 3439, 3440);
DECLARE_TOAST(pg_proc, 2836, 2837);
DECLARE_TOAST(pg_rewrite, 2838, 2839);
DECLARE_TOAST(pg_seclabel, 3598, 3599);
diff --git a/src/include/commands/defrem.h b/src/include/commands/defrem.h
index d790fbf..e5ef713 100644
--- a/src/include/commands/defrem.h
+++ b/src/include/commands/defrem.h
@@ -77,6 +77,10 @@ extern ObjectAddress DefineOperator(List *names, List *parameters);
extern void RemoveOperatorById(Oid operOid);
extern ObjectAddress AlterOperator(AlterOperatorStmt *stmt);
+/* commands/statscmds.c */
+extern ObjectAddress CreateStatistics(CreateStatsStmt *stmt);
+extern void RemoveStatisticsById(Oid statsOid);
+
/* commands/aggregatecmds.c */
extern ObjectAddress DefineAggregate(ParseState *pstate, List *name, List *args, bool oldstyle,
List *parameters);
diff --git a/src/include/nodes/nodes.h b/src/include/nodes/nodes.h
index 201f248..9790059 100644
--- a/src/include/nodes/nodes.h
+++ b/src/include/nodes/nodes.h
@@ -269,6 +269,7 @@ typedef enum NodeTag
T_PlaceHolderInfo,
T_MinMaxAggInfo,
T_PlannerParamItem,
+ T_MVStatisticInfo,
/*
* TAGS FOR MEMORY NODES (memnodes.h)
@@ -407,6 +408,7 @@ typedef enum NodeTag
T_CreateTransformStmt,
T_CreateAmStmt,
T_PartitionCmd,
+ T_CreateStatsStmt,
/*
* TAGS FOR PARSE TREE NODES (parsenodes.h)
diff --git a/src/include/nodes/parsenodes.h b/src/include/nodes/parsenodes.h
index 9d8ef77..13bfc52 100644
--- a/src/include/nodes/parsenodes.h
+++ b/src/include/nodes/parsenodes.h
@@ -603,6 +603,16 @@ typedef struct ColumnDef
int location; /* parse location, or -1 if none/unknown */
} ColumnDef;
+typedef struct CreateStatsStmt
+{
+ NodeTag type;
+ List *defnames; /* qualified name (list of Value strings) */
+ RangeVar *relation; /* relation to build statistics on */
+ List *keys; /* String nodes naming referenced column(s) */
+ bool if_not_exists; /* do nothing if statistics already exists */
+} CreateStatsStmt;
+
+
/*
* TableLikeClause - CREATE TABLE ( ... LIKE ... ) clause
*/
@@ -1516,6 +1526,7 @@ typedef enum ObjectType
OBJECT_RULE,
OBJECT_SCHEMA,
OBJECT_SEQUENCE,
+ OBJECT_STATISTICS,
OBJECT_TABCONSTRAINT,
OBJECT_TABLE,
OBJECT_TABLESPACE,
diff --git a/src/include/nodes/relation.h b/src/include/nodes/relation.h
index 3a1255a..3e9f930 100644
--- a/src/include/nodes/relation.h
+++ b/src/include/nodes/relation.h
@@ -520,6 +520,7 @@ typedef struct RelOptInfo
List *lateral_vars; /* LATERAL Vars and PHVs referenced by rel */
Relids lateral_referencers; /* rels that reference me laterally */
List *indexlist; /* list of IndexOptInfo */
+ List *mvstatlist; /* list of MVStatisticInfo */
BlockNumber pages; /* size estimates derived from pg_class */
double tuples;
double allvisfrac;
@@ -656,6 +657,32 @@ typedef struct ForeignKeyOptInfo
List *rinfos[INDEX_MAX_KEYS];
} ForeignKeyOptInfo;
+/*
+ * MVStatisticInfo
+ * Information about multivariate stats for planning/optimization
+ *
+ * This contains information about which columns are covered by the
+ * statistics (stakeys), which options were requested while adding the
+ * statistics (*_enabled), and which kinds of statistics were actually
+ * built and are available for the optimizer (*_built).
+ */
+typedef struct MVStatisticInfo
+{
+ NodeTag type;
+
+ Oid mvoid; /* OID of the statistics row */
+ RelOptInfo *rel; /* back-link to index's table */
+
+ /* enabled statistics */
+ bool ndist_enabled; /* ndistinct coefficient enabled */
+
+ /* built/available statistics */
+ bool ndist_built; /* ndistinct coefficient built */
+
+ /* columns in the statistics (attnums) */
+ int2vector *stakeys; /* attnums of the columns covered */
+
+} MVStatisticInfo;
/*
* EquivalenceClasses
diff --git a/src/include/utils/acl.h b/src/include/utils/acl.h
index faadebd..a1fbf3c 100644
--- a/src/include/utils/acl.h
+++ b/src/include/utils/acl.h
@@ -332,6 +332,7 @@ extern bool pg_foreign_data_wrapper_ownercheck(Oid srv_oid, Oid roleid);
extern bool pg_foreign_server_ownercheck(Oid srv_oid, Oid roleid);
extern bool pg_event_trigger_ownercheck(Oid et_oid, Oid roleid);
extern bool pg_extension_ownercheck(Oid ext_oid, Oid roleid);
+extern bool pg_statistics_ownercheck(Oid stat_oid, Oid roleid);
extern bool has_createrole_privilege(Oid roleid);
extern bool has_bypassrls_privilege(Oid roleid);
diff --git a/src/include/utils/builtins.h b/src/include/utils/builtins.h
index 7ed1623..de00edd 100644
--- a/src/include/utils/builtins.h
+++ b/src/include/utils/builtins.h
@@ -613,6 +613,10 @@ extern Datum pg_ddl_command_in(PG_FUNCTION_ARGS);
extern Datum pg_ddl_command_out(PG_FUNCTION_ARGS);
extern Datum pg_ddl_command_recv(PG_FUNCTION_ARGS);
extern Datum pg_ddl_command_send(PG_FUNCTION_ARGS);
+extern Datum pg_ndistinct_in(PG_FUNCTION_ARGS);
+extern Datum pg_ndistinct_out(PG_FUNCTION_ARGS);
+extern Datum pg_ndistinct_recv(PG_FUNCTION_ARGS);
+extern Datum pg_ndistinct_send(PG_FUNCTION_ARGS);
/* regexp.c */
extern Datum nameregexeq(PG_FUNCTION_ARGS);
diff --git a/src/include/utils/mvstats.h b/src/include/utils/mvstats.h
new file mode 100644
index 0000000..2937c78
--- /dev/null
+++ b/src/include/utils/mvstats.h
@@ -0,0 +1,60 @@
+/*-------------------------------------------------------------------------
+ *
+ * mvstats.h
+ * Multivariate statistics and selectivity estimation functions.
+ *
+ *
+ * Portions Copyright (c) 1996-2014, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/utils/mvstats.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef MVSTATS_H
+#define MVSTATS_H
+
+#include "fmgr.h"
+#include "commands/vacuum.h"
+
+#define MVSTATS_MAX_DIMENSIONS 8 /* max number of attributes */
+
+#define MVSTAT_NDISTINCT_MAGIC 0xA352BFA4 /* marks serialized bytea */
+#define MVSTAT_NDISTINCT_TYPE_BASIC 1 /* basic MCV list type */
+
+/* Multivariate distinct coefficients. */
+typedef struct MVNDistinctItem {
+ double ndistinct;
+ int nattrs;
+ int *attrs;
+} MVNDistinctItem;
+
+typedef struct MVNDistinctData {
+ uint32 magic; /* magic constant marker */
+ uint32 type; /* type of ndistinct (BASIC) */
+ int nitems;
+ MVNDistinctItem items[FLEXIBLE_ARRAY_MEMBER];
+} MVNDistinctData;
+
+typedef MVNDistinctData *MVNDistinct;
+
+
+MVNDistinct load_mv_ndistinct(Oid mvoid);
+
+bytea *serialize_mv_ndistinct(MVNDistinct ndistinct);
+
+/* deserialization of stats (serialization is private to analyze) */
+MVNDistinct deserialize_mv_ndistinct(bytea *data);
+
+
+MVNDistinct build_mv_ndistinct(double totalrows, int numrows, HeapTuple *rows,
+ int2vector *attrs, VacAttrStats **stats);
+
+void build_mv_stats(Relation onerel, double totalrows,
+ int numrows, HeapTuple *rows,
+ int natts, VacAttrStats **vacattrstats);
+
+void update_mv_stats(Oid relid, MVNDistinct ndistinct,
+ int2vector *attrs, VacAttrStats **stats);
+
+#endif
diff --git a/src/include/utils/rel.h b/src/include/utils/rel.h
index cd7ea1d..d89e1eb 100644
--- a/src/include/utils/rel.h
+++ b/src/include/utils/rel.h
@@ -91,6 +91,7 @@ typedef struct RelationData
bool rd_isvalid; /* relcache entry is valid */
char rd_indexvalid; /* state of rd_indexlist: 0 = not valid, 1 =
* valid, 2 = temporarily forced */
+ bool rd_mvstatvalid; /* state of rd_mvstatlist: true/false */
/*
* rd_createSubid is the ID of the highest subtransaction the rel has
@@ -134,6 +135,9 @@ typedef struct RelationData
Oid rd_oidindex; /* OID of unique index on OID, if any */
Oid rd_replidindex; /* OID of replica identity index, if any */
+ /* data managed by RelationGetMVStatList: */
+ List *rd_mvstatlist; /* list of OIDs of multivariate stats */
+
/* data managed by RelationGetIndexAttrBitmap: */
Bitmapset *rd_indexattr; /* identifies columns used in indexes */
Bitmapset *rd_keyattr; /* cols that can be ref'd by foreign keys */
diff --git a/src/include/utils/relcache.h b/src/include/utils/relcache.h
index 6ea7dd2..6f23593 100644
--- a/src/include/utils/relcache.h
+++ b/src/include/utils/relcache.h
@@ -39,6 +39,7 @@ extern void RelationClose(Relation relation);
*/
extern List *RelationGetFKeyList(Relation relation);
extern List *RelationGetIndexList(Relation relation);
+extern List *RelationGetMVStatList(Relation relation);
extern Oid RelationGetOidIndex(Relation relation);
extern Oid RelationGetReplicaIndex(Relation relation);
extern List *RelationGetIndexExpressions(Relation relation);
diff --git a/src/include/utils/syscache.h b/src/include/utils/syscache.h
index 4b7631e..cffb5dc 100644
--- a/src/include/utils/syscache.h
+++ b/src/include/utils/syscache.h
@@ -66,6 +66,8 @@ enum SysCacheIdentifier
INDEXRELID,
LANGNAME,
LANGOID,
+ MVSTATNAMENSP,
+ MVSTATOID,
NAMESPACENAME,
NAMESPACEOID,
OPERNAMENSP,
diff --git a/src/test/regress/expected/mv_ndistinct.out b/src/test/regress/expected/mv_ndistinct.out
new file mode 100644
index 0000000..5f55091
--- /dev/null
+++ b/src/test/regress/expected/mv_ndistinct.out
@@ -0,0 +1,117 @@
+-- data type passed by value
+CREATE TABLE ndistinct (
+ a INT,
+ b INT,
+ c INT,
+ d INT
+);
+-- unknown column
+CREATE STATISTICS s10 ON (unknown_column) FROM ndistinct;
+ERROR: column "unknown_column" referenced in statistics does not exist
+-- single column
+CREATE STATISTICS s10 ON (a) FROM ndistinct;
+ERROR: statistics require at least 2 columns
+-- single column, duplicated
+CREATE STATISTICS s10 ON (a,a) FROM ndistinct;
+ERROR: duplicate column name in statistics definition
+-- two columns, one duplicated
+CREATE STATISTICS s10 ON (a, a, b) FROM ndistinct;
+ERROR: duplicate column name in statistics definition
+-- correct command
+CREATE STATISTICS s10 ON (a, b, c) FROM ndistinct;
+-- perfectly correlated groups
+INSERT INTO ndistinct
+ SELECT i/100, i/100, i/100 FROM generate_series(1,10000) s(i);
+ANALYZE ndistinct;
+SELECT ndist_enabled, ndist_built, standist
+ FROM pg_mv_statistic WHERE starelid = 'ndistinct'::regclass;
+ ndist_enabled | ndist_built | standist
+---------------+-------------+-------------------------------------------------------------------------------------
+ t | t | [{0, 1, 101.000000}, {0, 2, 101.000000}, {1, 2, 101.000000}, {0, 1, 2, 101.000000}]
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b;
+ QUERY PLAN
+-----------------------------
+ HashAggregate
+ Group Key: a, b
+ -> Seq Scan on ndistinct
+(3 rows)
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c;
+ QUERY PLAN
+-----------------------------
+ HashAggregate
+ Group Key: a, b, c
+ -> Seq Scan on ndistinct
+(3 rows)
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c, d;
+ QUERY PLAN
+-----------------------------
+ HashAggregate
+ Group Key: a, b, c, d
+ -> Seq Scan on ndistinct
+(3 rows)
+
+TRUNCATE TABLE ndistinct;
+-- partially correlated groups
+INSERT INTO ndistinct
+ SELECT i/50, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE ndistinct;
+SELECT ndist_enabled, ndist_built, standist
+ FROM pg_mv_statistic WHERE starelid = 'ndistinct'::regclass;
+ ndist_enabled | ndist_built | standist
+---------------+-------------+-------------------------------------------------------------------------------------
+ t | t | [{0, 1, 201.000000}, {0, 2, 201.000000}, {1, 2, 101.000000}, {0, 1, 2, 201.000000}]
+(1 row)
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b;
+ QUERY PLAN
+---------------------------------------------------------------------
+ HashAggregate (cost=230.00..232.01 rows=201 width=16)
+ Group Key: a, b
+ -> Seq Scan on ndistinct (cost=0.00..155.00 rows=10000 width=8)
+(3 rows)
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c;
+ QUERY PLAN
+----------------------------------------------------------------------
+ HashAggregate (cost=255.00..257.01 rows=201 width=20)
+ Group Key: a, b, c
+ -> Seq Scan on ndistinct (cost=0.00..155.00 rows=10000 width=12)
+(3 rows)
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c, d;
+ QUERY PLAN
+----------------------------------------------------------------------
+ HashAggregate (cost=280.00..290.00 rows=1000 width=24)
+ Group Key: a, b, c, d
+ -> Seq Scan on ndistinct (cost=0.00..155.00 rows=10000 width=16)
+(3 rows)
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY b, c, d;
+ QUERY PLAN
+----------------------------------------------------------------------
+ HashAggregate (cost=255.00..265.00 rows=1000 width=20)
+ Group Key: b, c, d
+ -> Seq Scan on ndistinct (cost=0.00..155.00 rows=10000 width=12)
+(3 rows)
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, d;
+ QUERY PLAN
+---------------------------------------------------------------------
+ HashAggregate (cost=230.00..240.00 rows=1000 width=16)
+ Group Key: a, d
+ -> Seq Scan on ndistinct (cost=0.00..155.00 rows=10000 width=8)
+(3 rows)
+
+DROP TABLE ndistinct;
diff --git a/src/test/regress/expected/object_address.out b/src/test/regress/expected/object_address.out
index 7e81cdd..f85381a 100644
--- a/src/test/regress/expected/object_address.out
+++ b/src/test/regress/expected/object_address.out
@@ -36,6 +36,7 @@ ALTER DEFAULT PRIVILEGES FOR ROLE regress_addr_user REVOKE DELETE ON TABLES FROM
CREATE TRANSFORM FOR int LANGUAGE SQL (
FROM SQL WITH FUNCTION varchar_transform(internal),
TO SQL WITH FUNCTION int4recv(internal));
+CREATE STATISTICS addr_nsp.gentable_stat ON (a,b) FROM addr_nsp.gentable;
-- test some error cases
SELECT pg_get_object_address('stone', '{}', '{}');
ERROR: unrecognized object type "stone"
@@ -379,7 +380,8 @@ WITH objects (type, name, args) AS (VALUES
-- event trigger
('policy', '{addr_nsp, gentable, genpol}', '{}'),
('transform', '{int}', '{sql}'),
- ('access method', '{btree}', '{}')
+ ('access method', '{btree}', '{}'),
+ ('statistics', '{addr_nsp, gentable_stat}', '{}')
)
SELECT (pg_identify_object(addr1.classid, addr1.objid, addr1.subobjid)).*,
-- test roundtrip through pg_identify_object_as_address
@@ -427,13 +429,14 @@ SELECT (pg_identify_object(addr1.classid, addr1.objid, addr1.subobjid)).*,
trigger | | | t on addr_nsp.gentable | t
operator family | pg_catalog | integer_ops | pg_catalog.integer_ops USING btree | t
policy | | | genpol on addr_nsp.gentable | t
+ statistics | addr_nsp | gentable_stat | addr_nsp.gentable_stat | t
collation | pg_catalog | "default" | pg_catalog."default" | t
transform | | | for integer on language sql | t
text search dictionary | addr_nsp | addr_ts_dict | addr_nsp.addr_ts_dict | t
text search parser | addr_nsp | addr_ts_prs | addr_nsp.addr_ts_prs | t
text search configuration | addr_nsp | addr_ts_conf | addr_nsp.addr_ts_conf | t
text search template | addr_nsp | addr_ts_temp | addr_nsp.addr_ts_temp | t
-(42 rows)
+(43 rows)
---
--- Cleanup resources
diff --git a/src/test/regress/expected/opr_sanity.out b/src/test/regress/expected/opr_sanity.out
index 0bcec13..9a26205 100644
--- a/src/test/regress/expected/opr_sanity.out
+++ b/src/test/regress/expected/opr_sanity.out
@@ -817,11 +817,12 @@ WHERE c.castmethod = 'b' AND
text | character | 0 | i
character varying | character | 0 | i
pg_node_tree | text | 0 | i
+ pg_ndistinct | bytea | 0 | i
cidr | inet | 0 | i
xml | text | 0 | a
xml | character varying | 0 | a
xml | character | 0 | a
-(7 rows)
+(8 rows)
-- **************** pg_conversion ****************
-- Look for illegal values in pg_conversion fields.
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index e9cfadb..7599d6c 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1376,6 +1376,14 @@ pg_matviews| SELECT n.nspname AS schemaname,
LEFT JOIN pg_namespace n ON ((n.oid = c.relnamespace)))
LEFT JOIN pg_tablespace t ON ((t.oid = c.reltablespace)))
WHERE (c.relkind = 'm'::"char");
+pg_mv_stats| SELECT n.nspname AS schemaname,
+ c.relname AS tablename,
+ s.staname,
+ s.stakeys AS attnums,
+ length((s.standist)::text) AS ndistbytes
+ FROM ((pg_mv_statistic s
+ JOIN pg_class c ON ((c.oid = s.starelid)))
+ LEFT JOIN pg_namespace n ON ((n.oid = c.relnamespace)));
pg_policies| SELECT n.nspname AS schemaname,
c.relname AS tablename,
pol.polname AS policyname,
diff --git a/src/test/regress/expected/sanity_check.out b/src/test/regress/expected/sanity_check.out
index 7ad68c7..827f133 100644
--- a/src/test/regress/expected/sanity_check.out
+++ b/src/test/regress/expected/sanity_check.out
@@ -116,6 +116,7 @@ pg_init_privs|t
pg_language|t
pg_largeobject|t
pg_largeobject_metadata|t
+pg_mv_statistic|t
pg_namespace|t
pg_opclass|t
pg_operator|t
diff --git a/src/test/regress/expected/type_sanity.out b/src/test/regress/expected/type_sanity.out
index e5adfba..9356aaa 100644
--- a/src/test/regress/expected/type_sanity.out
+++ b/src/test/regress/expected/type_sanity.out
@@ -67,12 +67,13 @@ WHERE p1.typtype not in ('c','d','p') AND p1.typname NOT LIKE E'\\_%'
(SELECT 1 FROM pg_type as p2
WHERE p2.typname = ('_' || p1.typname)::name AND
p2.typelem = p1.oid and p1.typarray = p2.oid);
- oid | typname
------+--------------
- 194 | pg_node_tree
- 210 | smgr
- 705 | unknown
-(3 rows)
+ oid | typname
+------+--------------
+ 194 | pg_node_tree
+ 3343 | pg_ndistinct
+ 210 | smgr
+ 705 | unknown
+(4 rows)
-- Make sure typarray points to a varlena array type of our own base
SELECT p1.oid, p1.typname as basetype, p2.typname as arraytype,
diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule
index 8641769..b70b82f 100644
--- a/src/test/regress/parallel_schedule
+++ b/src/test/regress/parallel_schedule
@@ -113,3 +113,6 @@ test: event_trigger
# run stats by itself because its delay may be insufficient under heavy load
test: stats
+
+# run tests of multivariate stats
+test: mv_ndistinct
diff --git a/src/test/regress/serial_schedule b/src/test/regress/serial_schedule
index 835cf35..1c5bfaf 100644
--- a/src/test/regress/serial_schedule
+++ b/src/test/regress/serial_schedule
@@ -169,3 +169,4 @@ test: with
test: xml
test: event_trigger
test: stats
+test: mv_ndistinct
diff --git a/src/test/regress/sql/mv_ndistinct.sql b/src/test/regress/sql/mv_ndistinct.sql
new file mode 100644
index 0000000..5cef254
--- /dev/null
+++ b/src/test/regress/sql/mv_ndistinct.sql
@@ -0,0 +1,68 @@
+-- data type passed by value
+CREATE TABLE ndistinct (
+ a INT,
+ b INT,
+ c INT,
+ d INT
+);
+
+-- unknown column
+CREATE STATISTICS s10 ON (unknown_column) FROM ndistinct;
+
+-- single column
+CREATE STATISTICS s10 ON (a) FROM ndistinct;
+
+-- single column, duplicated
+CREATE STATISTICS s10 ON (a,a) FROM ndistinct;
+
+-- two columns, one duplicated
+CREATE STATISTICS s10 ON (a, a, b) FROM ndistinct;
+
+-- correct command
+CREATE STATISTICS s10 ON (a, b, c) FROM ndistinct;
+
+-- perfectly correlated groups
+INSERT INTO ndistinct
+ SELECT i/100, i/100, i/100 FROM generate_series(1,10000) s(i);
+
+ANALYZE ndistinct;
+
+SELECT ndist_enabled, ndist_built, standist
+ FROM pg_mv_statistic WHERE starelid = 'ndistinct'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b;
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c;
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c, d;
+
+TRUNCATE TABLE ndistinct;
+
+-- partially correlated groups
+INSERT INTO ndistinct
+ SELECT i/50, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE ndistinct;
+
+SELECT ndist_enabled, ndist_built, standist
+ FROM pg_mv_statistic WHERE starelid = 'ndistinct'::regclass;
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b;
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c;
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c, d;
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY b, c, d;
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, d;
+
+DROP TABLE ndistinct;
diff --git a/src/test/regress/sql/object_address.sql b/src/test/regress/sql/object_address.sql
index 7d1f93f..956bef3 100644
--- a/src/test/regress/sql/object_address.sql
+++ b/src/test/regress/sql/object_address.sql
@@ -39,6 +39,7 @@ ALTER DEFAULT PRIVILEGES FOR ROLE regress_addr_user REVOKE DELETE ON TABLES FROM
CREATE TRANSFORM FOR int LANGUAGE SQL (
FROM SQL WITH FUNCTION varchar_transform(internal),
TO SQL WITH FUNCTION int4recv(internal));
+CREATE STATISTICS addr_nsp.gentable_stat ON (a,b) FROM addr_nsp.gentable;
-- test some error cases
SELECT pg_get_object_address('stone', '{}', '{}');
@@ -169,7 +170,8 @@ WITH objects (type, name, args) AS (VALUES
-- event trigger
('policy', '{addr_nsp, gentable, genpol}', '{}'),
('transform', '{int}', '{sql}'),
- ('access method', '{btree}', '{}')
+ ('access method', '{btree}', '{}'),
+ ('statistics', '{addr_nsp, gentable_stat}', '{}')
)
SELECT (pg_identify_object(addr1.classid, addr1.objid, addr1.subobjid)).*,
-- test roundtrip through pg_identify_object_as_address
--
2.5.5
[binary/octet-stream] 0003-PATCH-functional-dependencies-only-the-ANALYZE-p-v22.patch (66.7K, ../../696aa95c-2411-9b2b-f36e-65b66bf47c88@2ndquadrant.com/4-0003-PATCH-functional-dependencies-only-the-ANALYZE-p-v22.patch)
download | inline diff:
From c3351cdc79390baadedcd4d72ec9c650ab60917b Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:36:25 +0200
Subject: [PATCH 3/9] PATCH: functional dependencies (only the ANALYZE part)
- implementation of soft functional dependencies (ANALYZE etc.)
- updates existing regression tests (new catalog etc.)
- new regression test for functional dependencies
- pg_ndistinct data type (varlena-based)
The algorithm detecting the dependencies is rather simple and probably
needs improvements, so that it detects more complicated dependencies,
and also validation of the math.
The patch introduces pg_dependencies, a new varlena data type for
storing serialized version of functional dependencies. This is similar
to what pg_ndistinct does for ndistinct coefficients.
---
doc/src/sgml/catalogs.sgml | 30 ++
doc/src/sgml/ref/create_statistics.sgml | 42 +-
src/backend/catalog/system_views.sql | 3 +-
src/backend/commands/statscmds.c | 37 +-
src/backend/nodes/copyfuncs.c | 1 +
src/backend/nodes/outfuncs.c | 2 +
src/backend/optimizer/util/plancat.c | 4 +-
src/backend/parser/gram.y | 14 +-
src/backend/utils/mvstats/Makefile | 2 +-
src/backend/utils/mvstats/README.dependencies | 118 +++++
src/backend/utils/mvstats/common.c | 22 +-
src/backend/utils/mvstats/dependencies.c | 632 ++++++++++++++++++++++++++
src/include/catalog/pg_cast.h | 4 +
src/include/catalog/pg_mv_statistic.h | 14 +-
src/include/catalog/pg_proc.h | 9 +
src/include/catalog/pg_type.h | 4 +
src/include/nodes/parsenodes.h | 1 +
src/include/nodes/relation.h | 2 +
src/include/utils/builtins.h | 4 +
src/include/utils/mvstats.h | 39 +-
src/test/regress/expected/mv_dependencies.out | 147 ++++++
src/test/regress/expected/mv_ndistinct.out | 10 +-
src/test/regress/expected/object_address.out | 2 +-
src/test/regress/expected/opr_sanity.out | 3 +-
src/test/regress/expected/rules.out | 3 +-
src/test/regress/expected/type_sanity.out | 7 +-
src/test/regress/parallel_schedule | 2 +-
src/test/regress/serial_schedule | 1 +
src/test/regress/sql/mv_dependencies.sql | 139 ++++++
src/test/regress/sql/mv_ndistinct.sql | 10 +-
src/test/regress/sql/object_address.sql | 2 +-
31 files changed, 1269 insertions(+), 41 deletions(-)
create mode 100644 src/backend/utils/mvstats/README.dependencies
create mode 100644 src/backend/utils/mvstats/dependencies.c
create mode 100644 src/test/regress/expected/mv_dependencies.out
create mode 100644 src/test/regress/sql/mv_dependencies.sql
diff --git a/doc/src/sgml/catalogs.sgml b/doc/src/sgml/catalogs.sgml
index b39ca69..716024b 100644
--- a/doc/src/sgml/catalogs.sgml
+++ b/doc/src/sgml/catalogs.sgml
@@ -4270,6 +4270,17 @@
</row>
<row>
+ <entry><structfield>deps_enabled</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, functional dependencies will be computed for the combination of
+ columns, covered by the statistics. This does not mean the dependencies
+ are already computed, though.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>ndist_built</structfield></entry>
<entry><type>bool</type></entry>
<entry></entry>
@@ -4280,6 +4291,16 @@
</row>
<row>
+ <entry><structfield>deps_built</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, functional depenedencies are already computed and available for
+ use during query estimation.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>stakeys</structfield></entry>
<entry><type>int2vector</type></entry>
<entry><literal><link linkend="catalog-pg-attribute"><structname>pg_attribute</structname></link>.attnum</literal></entry>
@@ -4299,6 +4320,15 @@
</entry>
</row>
+ <row>
+ <entry><structfield>stadeps</structfield></entry>
+ <entry><type>pg_dependencies</type></entry>
+ <entry></entry>
+ <entry>
+ Functional dependencies, serialized as <structname>pg_dependencies</> type.
+ </entry>
+ </row>
+
</tbody>
</tgroup>
</table>
diff --git a/doc/src/sgml/ref/create_statistics.sgml b/doc/src/sgml/ref/create_statistics.sgml
index 7fa118c..42adc38 100644
--- a/doc/src/sgml/ref/create_statistics.sgml
+++ b/doc/src/sgml/ref/create_statistics.sgml
@@ -21,8 +21,9 @@ PostgreSQL documentation
<refsynopsisdiv>
<synopsis>
-CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_name</replaceable> ON (
- <replaceable class="PARAMETER">column_name</replaceable>, <replaceable class="PARAMETER">column_name</replaceable> [, ...])
+CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_name</replaceable>
+ WITH ( <replaceable class="PARAMETER">option</replaceable> [= <replaceable class="PARAMETER">value</replaceable>] [, ... ] )
+ ON ( <replaceable class="PARAMETER">column_name</replaceable>, <replaceable class="PARAMETER">column_name</replaceable> [, ...])
FROM <replaceable class="PARAMETER">table_name</replaceable>
</synopsis>
@@ -100,6 +101,41 @@ CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_na
</variablelist>
+ <refsect2 id="SQL-CREATESTATISTICS-parameters">
+ <title id="SQL-CREATESTATISTICS-parameters-title">Parameters</title>
+
+ <indexterm zone="sql-createstatistics-parameters">
+ <primary>statistics parameters</primary>
+ </indexterm>
+
+ <para>
+ The <literal>WITH</> clause can specify <firstterm>options</>
+ for statistics. The currently available parameters are listed below.
+ </para>
+
+ <variablelist>
+
+ <varlistentry>
+ <term><literal>dependencies</> (<type>boolean</>)</term>
+ <listitem>
+ <para>
+ Enables functional dependencies for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><literal>ndistinct</> (<type>boolean</>)</term>
+ <listitem>
+ <para>
+ Enables ndistinct coefficients for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ </variablelist>
+
+ </refsect2>
</refsect1>
<refsect1 id="SQL-CREATESTATISTICS-examples">
@@ -120,7 +156,7 @@ CREATE TABLE t1 (
INSERT INTO t1 SELECT i/100, i/500
FROM generate_series(1,1000000) s(i);
-CREATE STATISTICS s1 ON (a, b) FROM t1;
+CREATE STATISTICS s1 WITH (dependencies) ON (a, b) FROM t1;
ANALYZE t1;
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 7eb356e..fc1b1a9 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -187,7 +187,8 @@ CREATE VIEW pg_mv_stats AS
C.relname AS tablename,
S.staname AS staname,
S.stakeys AS attnums,
- length(s.standist) AS ndistbytes
+ length(s.standist::bytea) AS ndistbytes,
+ length(S.stadeps::bytea) AS depsbytes
FROM (pg_mv_statistic S JOIN pg_class C ON (C.oid = S.starelid))
LEFT JOIN pg_namespace N ON (N.oid = C.relnamespace);
diff --git a/src/backend/commands/statscmds.c b/src/backend/commands/statscmds.c
index 8453dc4..ec41bbc 100644
--- a/src/backend/commands/statscmds.c
+++ b/src/backend/commands/statscmds.c
@@ -38,7 +38,9 @@ compare_int16(const void *a, const void *b)
}
/*
- * Implements the CREATE STATISTICS name ON (columns) FROM table
+ * Implements the CREATE STATISTICS command with syntax:
+ *
+ * CREATE STATISTICS name WITH (options) ON (columns) FROM table
*
* We do require that the types support sorting (ltopr), although some
* statistics might work with equality only.
@@ -72,6 +74,10 @@ CreateStatistics(CreateStatsStmt *stmt)
ObjectAddress parentobject,
childobject;
+ /* by default build nothing */
+ bool build_ndistinct = false,
+ build_dependencies = false;
+
Assert(IsA(stmt, CreateStatsStmt));
/* resolve the pieces of the name (namespace etc.) */
@@ -151,6 +157,31 @@ CreateStatistics(CreateStatsStmt *stmt)
(errcode(ERRCODE_UNDEFINED_COLUMN),
errmsg("duplicate column name in statistics definition")));
+ /*
+ * Parse the statistics options - currently only statistics types are
+ * recognized (ndistinct, dependencies).
+ */
+ foreach(l, stmt->options)
+ {
+ DefElem *opt = (DefElem *) lfirst(l);
+
+ if (strcmp(opt->defname, "ndistinct") == 0)
+ build_ndistinct = defGetBoolean(opt);
+ else if (strcmp(opt->defname, "dependencies") == 0)
+ build_dependencies = defGetBoolean(opt);
+ else
+ ereport(ERROR,
+ (errcode(ERRCODE_SYNTAX_ERROR),
+ errmsg("unrecognized STATISTICS option \"%s\"",
+ opt->defname)));
+ }
+
+ /* Make sure there's at least one statistics type specified. */
+ if (! (build_ndistinct || build_dependencies))
+ ereport(ERROR,
+ (errcode(ERRCODE_SYNTAX_ERROR),
+ errmsg("no statistics type (ndistinct, dependencies) requested")));
+
stakeys = buildint2vector(attnums, numcols);
/*
@@ -170,9 +201,11 @@ CreateStatistics(CreateStatsStmt *stmt)
values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(stakeys);
/* enabled statistics */
- values[Anum_pg_mv_statistic_ndist_enabled - 1] = BoolGetDatum(true);
+ values[Anum_pg_mv_statistic_ndist_enabled - 1] = BoolGetDatum(build_ndistinct);
+ values[Anum_pg_mv_statistic_deps_enabled - 1] = BoolGetDatum(build_dependencies);
nulls[Anum_pg_mv_statistic_standist - 1] = true;
+ nulls[Anum_pg_mv_statistic_stadeps - 1] = true;
/* insert the tuple into pg_mv_statistic */
mvstatrel = heap_open(MvStatisticRelationId, RowExclusiveLock);
diff --git a/src/backend/nodes/copyfuncs.c b/src/backend/nodes/copyfuncs.c
index 256f8c6..8791b1c 100644
--- a/src/backend/nodes/copyfuncs.c
+++ b/src/backend/nodes/copyfuncs.c
@@ -4260,6 +4260,7 @@ _copyCreateStatsStmt(const CreateStatsStmt *from)
COPY_NODE_FIELD(defnames);
COPY_NODE_FIELD(relation);
COPY_NODE_FIELD(keys);
+ COPY_NODE_FIELD(options);
COPY_SCALAR_FIELD(if_not_exists);
return newnode;
diff --git a/src/backend/nodes/outfuncs.c b/src/backend/nodes/outfuncs.c
index 1c2c200..26e504f 100644
--- a/src/backend/nodes/outfuncs.c
+++ b/src/backend/nodes/outfuncs.c
@@ -2180,9 +2180,11 @@ _outMVStatisticInfo(StringInfo str, const MVStatisticInfo *node)
/* enabled statistics */
WRITE_BOOL_FIELD(ndist_enabled);
+ WRITE_BOOL_FIELD(deps_enabled);
/* built/available statistics */
WRITE_BOOL_FIELD(ndist_built);
+ WRITE_BOOL_FIELD(deps_built);
}
static void
diff --git a/src/backend/optimizer/util/plancat.c b/src/backend/optimizer/util/plancat.c
index 16a90ea..91d4099 100644
--- a/src/backend/optimizer/util/plancat.c
+++ b/src/backend/optimizer/util/plancat.c
@@ -422,7 +422,7 @@ get_relation_info(PlannerInfo *root, Oid relationObjectId, bool inhparent,
mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
/* unavailable stats are not interesting for the planner */
- if (mvstat->ndist_built)
+ if (mvstat->deps_built || mvstat->ndist_built)
{
info = makeNode(MVStatisticInfo);
@@ -431,9 +431,11 @@ get_relation_info(PlannerInfo *root, Oid relationObjectId, bool inhparent,
/* enabled statistics */
info->ndist_enabled = mvstat->ndist_enabled;
+ info->deps_enabled = mvstat->deps_enabled;
/* built/available statistics */
info->ndist_built = mvstat->ndist_built;
+ info->deps_built = mvstat->deps_built;
/* stakeys */
adatum = SysCacheGetAttr(MVSTATOID, htup,
diff --git a/src/backend/parser/gram.y b/src/backend/parser/gram.y
index 3fd1fac..a4a965d 100644
--- a/src/backend/parser/gram.y
+++ b/src/backend/parser/gram.y
@@ -3722,21 +3722,23 @@ ExistingIndex: USING INDEX index_name { $$ = $3; }
*****************************************************************************/
-CreateStatsStmt: CREATE STATISTICS any_name ON '(' columnList ')' FROM qualified_name
+CreateStatsStmt: CREATE STATISTICS any_name opt_reloptions ON '(' columnList ')' FROM qualified_name
{
CreateStatsStmt *n = makeNode(CreateStatsStmt);
n->defnames = $3;
- n->relation = $9;
- n->keys = $6;
+ n->relation = $10;
+ n->keys = $7;
+ n->options = $4;
n->if_not_exists = false;
$$ = (Node *)n;
}
- | CREATE STATISTICS IF_P NOT EXISTS any_name ON '(' columnList ')' FROM qualified_name
+ | CREATE STATISTICS IF_P NOT EXISTS any_name opt_reloptions ON '(' columnList ')' FROM qualified_name
{
CreateStatsStmt *n = makeNode(CreateStatsStmt);
n->defnames = $6;
- n->relation = $12;
- n->keys = $9;
+ n->relation = $13;
+ n->keys = $10;
+ n->options = $7;
n->if_not_exists = true;
$$ = (Node *)n;
}
diff --git a/src/backend/utils/mvstats/Makefile b/src/backend/utils/mvstats/Makefile
index 7295d46..21fe7e5 100644
--- a/src/backend/utils/mvstats/Makefile
+++ b/src/backend/utils/mvstats/Makefile
@@ -12,6 +12,6 @@ subdir = src/backend/utils/mvstats
top_builddir = ../../../..
include $(top_builddir)/src/Makefile.global
-OBJS = common.o mvdist.o
+OBJS = common.o dependencies.o mvdist.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/utils/mvstats/README.dependencies b/src/backend/utils/mvstats/README.dependencies
new file mode 100644
index 0000000..86bee6e
--- /dev/null
+++ b/src/backend/utils/mvstats/README.dependencies
@@ -0,0 +1,118 @@
+Soft functional dependencies
+============================
+
+Functional dependencies are a concept well described in relational theory,
+particularly in definition of normalization and "normal forms". Wikipedia
+has a nice definition of a functional dependency [1]:
+
+ In a given table, an attribute Y is said to have a functional dependency
+ on a set of attributes X (written X -> Y) if and only if each X value is
+ associated with precisely one Y value. For example, in an "Employee"
+ table that includes the attributes "Employee ID" and "Employee Date of
+ Birth", the functional dependency
+
+ {Employee ID} -> {Employee Date of Birth}
+
+ would hold. It follows from the previous two sentences that each
+ {Employee ID} is associated with precisely one {Employee Date of Birth}.
+
+ [1] http://en.wikipedia.org/wiki/Database_normalization
+
+In practical terms, functional dependencies mean that a value in one column
+determines values in some other column. Consider for example this trivial
+table with two integer columns:
+
+ CREATE TABLE t (a INT, b INT)
+ AS SELECT i, i/10 FROM generate_series(1,100000) s(i);
+
+Clearly, knowledge of the value in column 'a' is sufficient to determine the
+value in column 'b', as it's simply (a/10). A more practical example may be
+addresses, where the knowledge of a ZIP code (usually) determines city. Larger
+cities may have multiple ZIP codes, so the dependency can't be reversed.
+
+Many datasets might be normalized not to contain such dependencies, but often
+it's not practical for various reasons. In some cases it's actually a conscious
+design choice to model the dataset in denormalized way, either because of
+performance or to make querying easier.
+
+
+soft dependencies
+-----------------
+
+Real-world data sets often contain data errors, either because of data entry
+mistakes (user mistyping the ZIP code) or perhaps issues in generating the
+data (e.g. a ZIP code mistakenly assigned to two cities in different states).
+
+A strict implementation would either ignore dependencies in such cases,
+rendering the approach mostly useless even for slightly noisy data sets, or
+result in sudden changes in behavior depending on minor differences between
+samples provided to ANALYZE.
+
+For this reason the statistics implementes "soft" functional dependencies,
+associating each functional dependency with a degree of validity (a number
+number between 0 and 1). This degree is then used to combine selectivities
+in a smooth manner.
+
+
+Mining dependencies (ANALYZE)
+-----------------------------
+
+The current algorithm is fairly simple - generate all possible functional
+dependencies, and for each one count the number of rows rows consistent it.
+Then use the fraction of rows (supporting/total) as the degree.
+
+To count the rows consistent with the dependency (a => b):
+
+ (a) Sort the data lexicographically, i.e. first by 'a' then 'b'.
+
+ (b) For each group of rows with the same 'a' value, count the number of
+ distinct values in 'b'.
+
+ (c) If there's a single distinct value in 'b', the rows are consistent with
+ the functional dependency. Otherwise they contradict it.
+
+The algorithm also requires a minimum size of the group to consider it
+consistent (currently 3 rows in the sample). Small groups make it less likely
+to break the consistency.
+
+
+Clause reduction (planner/optimizer)
+------------------------------------
+
+Apllying the functional dependencies is fairly simple - given a list of
+equality clauses, we compute selectivities of each clause and then use the
+degree to combine them using this formula
+
+ P(a=?,b=?) = P(a=?) * (d + (1-d) * P(b=?))
+
+Where 'd' is the degree of functional dependence (a=>b).
+
+With more than two equality clauses, this process happens recursively. For
+example for (a,b,c) we first use (a,b=>c) to break the computation into
+
+ P(a=?,b=?,c=?) = P(a=?,b=?) * (d + (1-d)*P(b=?))
+
+and then apply (a=>b) the same way on P(a=?,b=?).
+
+
+Consistecy of clauses
+---------------------
+
+Functional dependencies only express general dependencies between columns,
+without referencing particular values. This assumes that the equality clauses
+are in fact consistent with the functinal dependency, i.e. that given the a
+dependency (a=>b), the value in (b=?) clause is the value determined by (a=?).
+If that's not the case, the clauses are "inconsistent" with the functional
+dependency and the result will be over-estimation.
+
+This may happen for example when using conditions on ZIP and city name with
+mismatching values (ZIP for a different city), etc. In such case the result
+set will be empty, but we'll estimate the selectivity using the ZIP condition.
+
+In this case the default estimation based on AVIA principle happens to work
+better, but mostly by chance.
+
+This issue is the price for the simplicity of functional dependencies. If the
+application frequently constructs queries with clauses inconsistent with
+functional dependencies present in the data, the best solution is not to
+use functional dependencies, but one of the more complex types of statistics.
diff --git a/src/backend/utils/mvstats/common.c b/src/backend/utils/mvstats/common.c
index c4a2644..e1bed39 100644
--- a/src/backend/utils/mvstats/common.c
+++ b/src/backend/utils/mvstats/common.c
@@ -50,6 +50,7 @@ build_mv_stats(Relation onerel, double totalrows,
int j;
MVStatisticInfo *stat = (MVStatisticInfo *) lfirst(lc);
MVNDistinct ndistinct = NULL;
+ MVDependencies deps = NULL;
VacAttrStats **stats = NULL;
int numatts = 0;
@@ -86,8 +87,12 @@ build_mv_stats(Relation onerel, double totalrows,
if (stat->ndist_enabled)
ndistinct = build_mv_ndistinct(totalrows, numrows, rows, attrs, stats);
+ /* analyze functional dependencies between the columns */
+ if (stat->deps_enabled)
+ deps = build_mv_dependencies(numrows, rows, attrs, stats);
+
/* store the statistics in the catalog */
- update_mv_stats(stat->mvoid, ndistinct, attrs, stats);
+ update_mv_stats(stat->mvoid, ndistinct, deps, attrs, stats);
}
}
@@ -167,6 +172,8 @@ list_mv_stats(Oid relid)
info->stakeys = buildint2vector(stats->stakeys.values, stats->stakeys.dim1);
info->ndist_enabled = stats->ndist_enabled;
info->ndist_built = stats->ndist_built;
+ info->deps_enabled = stats->deps_enabled;
+ info->deps_built = stats->deps_built;
result = lappend(result, info);
}
@@ -184,7 +191,7 @@ list_mv_stats(Oid relid)
}
void
-update_mv_stats(Oid mvoid, MVNDistinct ndistinct,
+update_mv_stats(Oid mvoid, MVNDistinct ndistinct, MVDependencies dependencies,
int2vector *attrs, VacAttrStats **stats)
{
HeapTuple stup,
@@ -211,18 +218,29 @@ update_mv_stats(Oid mvoid, MVNDistinct ndistinct,
values[Anum_pg_mv_statistic_standist-1] = PointerGetDatum(data);
}
+ if (dependencies != NULL)
+ {
+ nulls[Anum_pg_mv_statistic_stadeps - 1] = false;
+ values[Anum_pg_mv_statistic_stadeps - 1]
+ = PointerGetDatum(serialize_mv_dependencies(dependencies));
+ }
+
/* always replace the value (either by bytea or NULL) */
replaces[Anum_pg_mv_statistic_standist - 1] = true;
+ replaces[Anum_pg_mv_statistic_stadeps - 1] = true;
/* always change the availability flags */
nulls[Anum_pg_mv_statistic_ndist_built - 1] = false;
+ nulls[Anum_pg_mv_statistic_deps_built - 1] = false;
nulls[Anum_pg_mv_statistic_stakeys - 1] = false;
/* use the new attnums, in case we removed some dropped ones */
replaces[Anum_pg_mv_statistic_ndist_built - 1] = true;
+ replaces[Anum_pg_mv_statistic_deps_built - 1] = true;
replaces[Anum_pg_mv_statistic_stakeys - 1] = true;
values[Anum_pg_mv_statistic_ndist_built - 1] = BoolGetDatum(ndistinct != NULL);
+ values[Anum_pg_mv_statistic_deps_built - 1] = BoolGetDatum(dependencies != NULL);
values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(attrs);
diff --git a/src/backend/utils/mvstats/dependencies.c b/src/backend/utils/mvstats/dependencies.c
new file mode 100644
index 0000000..03865ed
--- /dev/null
+++ b/src/backend/utils/mvstats/dependencies.c
@@ -0,0 +1,632 @@
+/*-------------------------------------------------------------------------
+ *
+ * dependencies.c
+ * POSTGRES multivariate functional dependencies
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/dependencies.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "common.h"
+
+#include "utils/bytea.h"
+#include "utils/lsyscache.h"
+
+/*
+ * Internal state for DependencyGenerator of dependencies. Dependencies are similar to
+ * k-permutations of n elements, except that the order does not matter for the
+ * first (k-1) elements. That is, (a,b=>c) and (b,a=>c) are equivalent.
+ */
+typedef struct DependencyGeneratorData
+{
+ int k; /* size of the dependency */
+ int current; /* next dependency to return (index) */
+ int ndependencies; /* number of dependencies generated */
+ int *dependencies; /* array of pre-generated dependencies */
+} DependencyGeneratorData;
+
+typedef DependencyGeneratorData *DependencyGenerator;
+
+static void
+generate_dependencies_recurse(DependencyGenerator state,
+ int n, int index, int start, int *current)
+{
+ /*
+ * The generator handles the first (k-1) elements differently from
+ * the last element.
+ */
+ if (index < (state->k - 1))
+ {
+ int i;
+
+ /*
+ * The first (k-1) values have to be in ascending order, which we
+ * generate recursively.
+ */
+
+ for (i = start; i < n; i++)
+ {
+ current[index] = i;
+ generate_dependencies_recurse(state, n, (index+1), (i+1), current);
+ }
+ }
+ else
+ {
+ int i;
+
+ /*
+ * the last element is the implied value, which does not respect the
+ * ascending order. We just need to check that the value is not in the
+ * first (k-1) elements.
+ */
+
+ for (i = 0; i < n; i++)
+ {
+ int j;
+ bool match = false;
+
+ current[index] = i;
+
+ for (j = 0; j < index; j++)
+ {
+ if (current[j] == i)
+ {
+ match = true;
+ break;
+ }
+ }
+
+ /*
+ * If the value is not found in the first part of the dependency,
+ * we're done.
+ */
+ if (! match)
+ {
+ state->dependencies
+ = (int*)repalloc(state->dependencies,
+ state->k * (state->ndependencies + 1) * sizeof(int));
+ memcpy(&state->dependencies[(state->k * state->ndependencies)],
+ current, state->k * sizeof(int));
+ state->ndependencies++;
+ }
+ }
+ }
+}
+
+/* generate all dependencies (k-permutations of n elements) */
+static void
+generate_dependencies(DependencyGenerator state, int n)
+{
+ int *current = (int *) palloc0(sizeof(int) * state->k);
+
+ generate_dependencies_recurse(state, n, 0, 0, current);
+
+ pfree(current);
+}
+
+/*
+ * initialize the DependencyGenerator of variations, and prebuild the variations
+ *
+ * This pre-builds all the variations. We could also generate them in
+ * DependencyGenerator_next(), but this seems simpler.
+ */
+static DependencyGenerator
+DependencyGenerator_init(int2vector *attrs, int k)
+{
+ int n = attrs->dim1;
+ DependencyGenerator state;
+
+ Assert((n >= k) && (k > 0));
+
+ /* allocate the DependencyGenerator state as a single chunk of memory */
+ state = (DependencyGenerator) palloc0(sizeof(DependencyGeneratorData));
+ state->dependencies = (int*)palloc(k * sizeof(int));
+
+ state->ndependencies = 0;
+ state->current = 0;
+ state->k = k;
+
+ /* now actually pre-generate all the variations */
+ generate_dependencies(state, n);
+
+ return state;
+}
+
+/* free the DependencyGenerator state */
+static void
+DependencyGenerator_free(DependencyGenerator state)
+{
+ /* we've allocated a single chunk, so just free it */
+ pfree(state);
+}
+
+/* generate next combination */
+static int *
+DependencyGenerator_next(DependencyGenerator state, int2vector *attrs)
+{
+ if (state->current == state->ndependencies)
+ return NULL;
+
+ return &state->dependencies[state->k * state->current++];
+}
+
+
+/*
+ * validates functional dependency on the data
+ *
+ * An actual work horse of detecting functional dependencies. Given a variation
+ * of k attributes, it checks that the first (k-1) are sufficient to determine
+ * the last one.
+ */
+static double
+dependency_degree(int numrows, HeapTuple *rows, int k, int *dependency,
+ VacAttrStats **stats, int2vector *attrs)
+{
+ int i,
+ j;
+ int nvalues = numrows * k;
+
+ /*
+ * XXX Maybe the threshold should be somehow related to the number of
+ * distinct values in the combination of columns we're analyzing. Assuming
+ * the distribution is uniform, we can estimate the average group size and
+ * use it as a threshold, similarly to what we do for MCV lists.
+ */
+ int min_group_size = 3;
+
+ /* number of groups supporting / contradicting the dependency */
+ int n_supporting = 0;
+ int n_contradicting = 0;
+
+ /* counters valid within a group */
+ int group_size = 0;
+ int n_violations = 0;
+
+ int n_supporting_rows = 0;
+ int n_contradicting_rows = 0;
+
+ /* sort info for all attributes columns */
+ MultiSortSupport mss = multi_sort_init(k);
+
+ /* data for the sort */
+ SortItem *items = (SortItem *) palloc0(numrows * sizeof(SortItem));
+ Datum *values = (Datum *) palloc0(sizeof(Datum) * nvalues);
+ bool *isnull = (bool *) palloc0(sizeof(bool) * nvalues);
+
+ /* fix the pointers to values/isnull */
+ for (i = 0; i < numrows; i++)
+ {
+ items[i].values = &values[i * k];
+ items[i].isnull = &isnull[i * k];
+ }
+
+ /*
+ * Verify the dependency (a,b,...)->z, using a rather simple algorithm:
+ *
+ * (a) sort the data lexicographically
+ *
+ * (b) split the data into groups by first (k-1) columns
+ *
+ * (c) for each group count different values in the last column
+ */
+
+ /* prepare the sort function for the first dimension, and SortItem array */
+ for (i = 0; i < k; i++)
+ {
+ multi_sort_add_dimension(mss, i, dependency[i], stats);
+
+ /* accumulate all the data for both columns into an array and sort it */
+ for (j = 0; j < numrows; j++)
+ {
+ items[j].values[i]
+ = heap_getattr(rows[j], attrs->values[dependency[i]],
+ stats[i]->tupDesc, &items[j].isnull[i]);
+ }
+ }
+
+ /* sort the items so that we can detect the groups */
+ qsort_arg((void *) items, numrows, sizeof(SortItem),
+ multi_sort_compare, mss);
+
+ /*
+ * Walk through the sorted array, split it into rows according to the
+ * first (k-1) columns. If there's a single value in the last column, we
+ * count the group as 'supporting' the functional dependency. Otherwise we
+ * count it as contradicting.
+ *
+ * We also require a group to have a minimum number of rows to be
+ * considered useful for supporting the dependency. Contradicting groups
+ * may be of any size, though.
+ *
+ * XXX The minimum size requirement makes it impossible to identify case
+ * when both columns are unique (or nearly unique), and therefore
+ * trivially functionally dependent.
+ */
+
+ /* start with the first row forming a group */
+ group_size = 1;
+
+ for (i = 1; i < numrows; i++)
+ {
+ /* end of the preceding group */
+ if (multi_sort_compare_dims(0, (k - 2), &items[i - 1], &items[i], mss) != 0)
+ {
+ /*
+ * If there is a single are no contradicting rows, count the group
+ * as supporting, otherwise contradicting.
+ */
+ if ((n_violations == 0) && (group_size >= min_group_size))
+ {
+ n_supporting += 1;
+ n_supporting_rows += group_size;
+ }
+ else if (n_violations > 0)
+ {
+ n_contradicting += 1;
+ n_contradicting_rows += group_size;
+ }
+
+ /* current values start a new group */
+ n_violations = 0;
+ group_size = 0;
+ }
+ /* first colums match, but the last one does not (so contradicting) */
+ else if (multi_sort_compare_dims((k - 1), (k - 1), &items[i - 1], &items[i], mss) != 0)
+ n_violations += 1;
+
+ group_size += 1;
+ }
+
+ /* handle the last group (just like above) */
+ if ((n_violations == 0) && (group_size >= min_group_size))
+ {
+ n_supporting += 1;
+ n_supporting_rows += group_size;
+ }
+ else if (n_violations)
+ {
+ n_contradicting += 1;
+ n_contradicting_rows += group_size;
+ }
+
+ pfree(items);
+ pfree(values);
+ pfree(isnull);
+ pfree(mss);
+
+ /* Compute the 'degree of validity' as (supporting/total). */
+ return (n_supporting_rows * 1.0 / numrows);
+}
+
+/*
+ * detects functional dependencies between groups of columns
+ *
+ * Generates all possible subsets of columns (variations) and checks if the
+ * last one is determined by the preceding ones. For example given 3 columns,
+ * there are 12 variations (6 for variations on 2 columns, 6 for 3 columns):
+ *
+ * two columns three columns
+ * ----------- -------------
+ * (a) -> c (a,b) -> c
+ * (b) -> c (b,a) -> c
+ * (a) -> b (a,c) -> b
+ * (c) -> b (c,a) -> b
+ * (c) -> a (c,b) -> a
+ * (b) -> a (b,c) -> a
+ */
+MVDependencies
+build_mv_dependencies(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats)
+{
+ int i;
+ int k;
+ int numattrs = attrs->dim1;
+
+ /* result */
+ MVDependencies dependencies = NULL;
+
+ Assert(numattrs >= 2);
+
+ /*
+ * We'll try build functional dependencies starting from the smallest ones
+ * covering jut 2 columns, to the largest ones, covering all columns
+ * included int the statistics. We start from the smallest ones because we
+ * want to be able to skip already implied ones.
+ */
+ for (k = 2; k <= numattrs; k++)
+ {
+ int *dependency; /* array with k elements */
+
+ /* prepare a DependencyGenerator of variation */
+ DependencyGenerator DependencyGenerator = DependencyGenerator_init(attrs, k);
+
+ /* generate all possible variations of k values (out of n) */
+ while ((dependency = DependencyGenerator_next(DependencyGenerator, attrs)))
+ {
+ double degree;
+ MVDependency d;
+
+ /* compute how valid the dependency seems */
+ degree = dependency_degree(numrows, rows, k, dependency, stats, attrs);
+
+ /* if the dependency seems entirely invalid, don't bother storing it */
+ if (degree == 0.0)
+ continue;
+
+ d = (MVDependency) palloc0(offsetof(MVDependencyData, attributes)
+ +k * sizeof(int));
+
+ /* copy the dependency (and keep the indexes into stakeys) */
+ d->degree = degree;
+ d->nattributes = k;
+ for (i = 0; i < k; i++)
+ d->attributes[i] = dependency[i];
+
+ /* initialize the list of dependencies */
+ if (dependencies == NULL)
+ {
+ dependencies
+ = (MVDependencies) palloc0(sizeof(MVDependenciesData));
+
+ dependencies->magic = MVSTAT_DEPS_MAGIC;
+ dependencies->type = MVSTAT_DEPS_TYPE_BASIC;
+ dependencies->ndeps = 0;
+ }
+
+ dependencies->ndeps++;
+ dependencies = (MVDependencies) repalloc(dependencies,
+ offsetof(MVDependenciesData, deps)
+ +dependencies->ndeps * sizeof(MVDependency));
+
+ dependencies->deps[dependencies->ndeps - 1] = d;
+ }
+
+ /* we're done with variations of k elements, so free the DependencyGenerator */
+ DependencyGenerator_free(DependencyGenerator);
+ }
+
+ return dependencies;
+}
+
+
+/*
+ * serialize list of dependencies into a bytea
+ */
+bytea *
+serialize_mv_dependencies(MVDependencies dependencies)
+{
+ int i;
+ bytea *output;
+ char *tmp;
+ Size len;
+
+ /* we need to store ndeps, with a number of attributes for each one */
+ len = VARHDRSZ + offsetof(MVDependenciesData, deps) +
+ dependencies->ndeps * offsetof(MVDependencyData, attributes);
+
+ /* and also include space for the actual attribute numbers and degrees */
+ for (i = 0; i < dependencies->ndeps; i++)
+ len += (sizeof(int16) * dependencies->deps[i]->nattributes);
+
+ output = (bytea *) palloc0(len);
+ SET_VARSIZE(output, len);
+
+ tmp = VARDATA(output);
+
+ /* first, store the number of dimensions / items */
+ memcpy(tmp, dependencies, offsetof(MVDependenciesData, deps));
+ tmp += offsetof(MVDependenciesData, deps);
+
+ /* store number of attributes and attribute numbers for each dependency */
+ for (i = 0; i < dependencies->ndeps; i++)
+ {
+ MVDependency d = dependencies->deps[i];
+
+ memcpy(tmp, d, offsetof(MVDependencyData, attributes));
+ tmp += offsetof(MVDependencyData, attributes);
+
+ memcpy(tmp, d->attributes, sizeof(int16) * d->nattributes);
+ tmp += sizeof(int16) * d->nattributes;
+
+ Assert(tmp <= ((char *) output + len));
+ }
+
+ return output;
+}
+
+/*
+ * Reads serialized dependencies into MVDependencies structure.
+ */
+MVDependencies
+deserialize_mv_dependencies(bytea *data)
+{
+ int i;
+ Size expected_size;
+ MVDependencies dependencies;
+ char *tmp;
+
+ if (data == NULL)
+ return NULL;
+
+ if (VARSIZE_ANY_EXHDR(data) < offsetof(MVDependenciesData, deps))
+ elog(ERROR, "invalid MVDependencies size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), offsetof(MVDependenciesData, deps));
+
+ /* read the MVDependencies header */
+ dependencies = (MVDependencies) palloc0(sizeof(MVDependenciesData));
+
+ /* initialize pointer to the data part (skip the varlena header) */
+ tmp = VARDATA_ANY(data);
+
+ /* get the header and perform basic sanity checks */
+ memcpy(dependencies, tmp, offsetof(MVDependenciesData, deps));
+ tmp += offsetof(MVDependenciesData, deps);
+
+ if (dependencies->magic != MVSTAT_DEPS_MAGIC)
+ elog(ERROR, "invalid dependency magic %d (expected %dd)",
+ dependencies->magic, MVSTAT_DEPS_MAGIC);
+
+ if (dependencies->type != MVSTAT_DEPS_TYPE_BASIC)
+ elog(ERROR, "invalid dependency type %d (expected %dd)",
+ dependencies->type, MVSTAT_DEPS_TYPE_BASIC);
+
+ Assert(dependencies->ndeps > 0);
+
+ /* what minimum bytea size do we expect for those parameters */
+ expected_size = offsetof(MVDependenciesData, deps) +
+ dependencies->ndeps * (offsetof(MVDependencyData, attributes) +
+ sizeof(int16) * 2);
+
+ if (VARSIZE_ANY_EXHDR(data) < expected_size)
+ elog(ERROR, "invalid dependencies size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ /* allocate space for the MCV items */
+ dependencies = repalloc(dependencies, offsetof(MVDependenciesData, deps)
+ +(dependencies->ndeps * sizeof(MVDependency)));
+
+ for (i = 0; i < dependencies->ndeps; i++)
+ {
+ double degree;
+ int k;
+ MVDependency d;
+
+ /* degree of validity */
+ memcpy(°ree, tmp, sizeof(double));
+ tmp += sizeof(double);
+
+ /* number of attributes */
+ memcpy(&k, tmp, sizeof(int));
+ tmp += sizeof(int);
+
+ /* is the number of attributes valid? */
+ Assert((k >= 2) && (k <= MVSTATS_MAX_DIMENSIONS));
+
+ /* now that we know the number of attributes, allocate the dependency */
+ d = (MVDependency) palloc0(offsetof(MVDependencyData, attributes) +
+ (k * sizeof(int)));
+
+ d->degree = degree;
+ d->nattributes = k;
+
+ /* copy attribute numbers */
+ memcpy(d->attributes, tmp, sizeof(int16) * d->nattributes);
+ tmp += sizeof(int16) * d->nattributes;
+
+ dependencies->deps[i] = d;
+
+ /* still within the bytea */
+ Assert(tmp <= ((char *) data + VARSIZE_ANY(data)));
+ }
+
+ /* we should have consumed the whole bytea exactly */
+ Assert(tmp == ((char *) data + VARSIZE_ANY(data)));
+
+ return dependencies;
+}
+
+/*
+ * pg_dependencies_in - input routine for type pg_dependencies.
+ *
+ * pg_dependencies is real enough to be a table column, but it has no operations
+ * of its own, and disallows input too
+ *
+ * XXX This is inspired by what pg_node_tree does.
+ */
+Datum
+pg_dependencies_in(PG_FUNCTION_ARGS)
+{
+ /*
+ * pg_node_list stores the data in binary form and parsing text input is
+ * not needed, so disallow this.
+ */
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_dependencies")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_dependencies - output routine for type pg_dependencies.
+ *
+ * histograms are serialized into a bytea value, so we simply call byteaout()
+ * to serialize the value into text. But it'd be nice to serialize that into
+ * a meaningful representation (e.g. for inspection by people).
+ */
+Datum
+pg_dependencies_out(PG_FUNCTION_ARGS)
+{
+ int i, j;
+ char *ret;
+ StringInfoData str;
+
+ bytea *data = PG_GETARG_BYTEA_PP(0);
+
+ MVDependencies dependencies = deserialize_mv_dependencies(data);
+
+ initStringInfo(&str);
+ appendStringInfoString(&str, "[");
+
+ for (i = 0; i < dependencies->ndeps; i++)
+ {
+ MVDependency dependency = dependencies->deps[i];
+
+ if (i > 0)
+ appendStringInfoString(&str, ", ");
+
+ appendStringInfoString(&str, "{");
+
+ for (j = 0; j < dependency->nattributes; j++)
+ {
+ if (j == dependency->nattributes-1)
+ appendStringInfoString(&str, " => ");
+ else if (j > 0)
+ appendStringInfoString(&str, ", ");
+
+ appendStringInfo(&str, "%d", dependency->attributes[j]);
+ }
+
+ appendStringInfo(&str, " : %f", dependency->degree);
+
+ appendStringInfoString(&str, "}");
+ }
+
+ appendStringInfoString(&str, "]");
+
+ ret = pstrdup(str.data);
+ pfree(str.data);
+
+ PG_RETURN_CSTRING(ret);
+}
+
+/*
+ * pg_dependencies_recv - binary input routine for type pg_dependencies.
+ */
+Datum
+pg_dependencies_recv(PG_FUNCTION_ARGS)
+{
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_dependencies")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_dependencies_send - binary output routine for type pg_dependencies.
+ *
+ * XXX Histograms are serialized into a bytea value, so let's just send that.
+ */
+Datum
+pg_dependencies_send(PG_FUNCTION_ARGS)
+{
+ return byteasend(fcinfo);
+}
diff --git a/src/include/catalog/pg_cast.h b/src/include/catalog/pg_cast.h
index d54a0f9..1ecf1cb 100644
--- a/src/include/catalog/pg_cast.h
+++ b/src/include/catalog/pg_cast.h
@@ -258,6 +258,10 @@ DATA(insert ( 194 25 0 i b ));
DATA(insert ( 3343 17 0 i b ));
DATA(insert ( 3343 25 0 i i ));
+/* pg_dependencies can be coerced to, but not from, bytea and text */
+DATA(insert ( 3353 17 0 i b ));
+DATA(insert ( 3353 25 0 i i ));
+
/*
* Datetime category
*/
diff --git a/src/include/catalog/pg_mv_statistic.h b/src/include/catalog/pg_mv_statistic.h
index fad80a3..e119cb7 100644
--- a/src/include/catalog/pg_mv_statistic.h
+++ b/src/include/catalog/pg_mv_statistic.h
@@ -38,9 +38,11 @@ CATALOG(pg_mv_statistic,3381)
/* statistics requested to build */
bool ndist_enabled; /* build ndist coefficient? */
+ bool deps_enabled; /* analyze dependencies? */
/* statistics that are available (if requested) */
bool ndist_built; /* ndistinct coeff built */
+ bool deps_built; /* dependencies were built */
/*
* variable-length fields start here, but we allow direct access to
@@ -50,6 +52,7 @@ CATALOG(pg_mv_statistic,3381)
#ifdef CATALOG_VARLEN
pg_ndistinct standist; /* ndistinct coeff (serialized) */
+ pg_dependencies stadeps; /* dependencies (serialized) */
#endif
} FormData_pg_mv_statistic;
@@ -65,14 +68,17 @@ typedef FormData_pg_mv_statistic *Form_pg_mv_statistic;
* compiler constants for pg_mv_statistic
* ----------------
*/
-#define Natts_pg_mv_statistic 8
+#define Natts_pg_mv_statistic 11
#define Anum_pg_mv_statistic_starelid 1
#define Anum_pg_mv_statistic_staname 2
#define Anum_pg_mv_statistic_stanamespace 3
#define Anum_pg_mv_statistic_staowner 4
#define Anum_pg_mv_statistic_ndist_enabled 5
-#define Anum_pg_mv_statistic_ndist_built 6
-#define Anum_pg_mv_statistic_stakeys 7
-#define Anum_pg_mv_statistic_standist 8
+#define Anum_pg_mv_statistic_deps_enabled 6
+#define Anum_pg_mv_statistic_ndist_built 7
+#define Anum_pg_mv_statistic_deps_built 8
+#define Anum_pg_mv_statistic_stakeys 9
+#define Anum_pg_mv_statistic_standist 10
+#define Anum_pg_mv_statistic_stadeps 11
#endif /* PG_MV_STATISTIC_H */
diff --git a/src/include/catalog/pg_proc.h b/src/include/catalog/pg_proc.h
index 8bf46e2..0c78361 100644
--- a/src/include/catalog/pg_proc.h
+++ b/src/include/catalog/pg_proc.h
@@ -2731,6 +2731,15 @@ DESCR("I/O");
DATA(insert OID = 3347 ( pg_ndistinct_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "3343" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_send _null_ _null_ _null_ ));
DESCR("I/O");
+DATA(insert OID = 3354 ( pg_dependencies_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 3353 "2275" _null_ _null_ _null_ _null_ _null_ pg_dependencies_in _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3355 ( pg_dependencies_out PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 2275 "3353" _null_ _null_ _null_ _null_ _null_ pg_dependencies_out _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3356 ( pg_dependencies_recv PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 3353 "2281" _null_ _null_ _null_ _null_ _null_ pg_dependencies_recv _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3357 ( pg_dependencies_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "3353" _null_ _null_ _null_ _null_ _null_ pg_dependencies_send _null_ _null_ _null_ ));
+DESCR("I/O");
+
DATA(insert OID = 1928 ( pg_stat_get_numscans PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_numscans _null_ _null_ _null_ ));
DESCR("statistics: number of scans done for table/index");
DATA(insert OID = 1929 ( pg_stat_get_tuples_returned PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_tuples_returned _null_ _null_ _null_ ));
diff --git a/src/include/catalog/pg_type.h b/src/include/catalog/pg_type.h
index e6ab642..250952b 100644
--- a/src/include/catalog/pg_type.h
+++ b/src/include/catalog/pg_type.h
@@ -368,6 +368,10 @@ DATA(insert OID = 3343 ( pg_ndistinct PGNSP PGUID -1 f b S f t \054 0 0 0 pg_nd
DESCR("multivariate ndistinct coefficients");
#define PGNDISTINCTOID 3343
+DATA(insert OID = 3353 ( pg_dependencies PGNSP PGUID -1 f b S f t \054 0 0 0 pg_dependencies_in pg_dependencies_out pg_dependencies_recv pg_dependencies_send - - - i x f 0 -1 0 100 _null_ _null_ _null_ ));
+DESCR("multivariate histogram");
+#define PGDEPENDENCIESOID 3353
+
DATA(insert OID = 32 ( pg_ddl_command PGNSP PGUID SIZEOF_POINTER t p P f t \054 0 0 0 pg_ddl_command_in pg_ddl_command_out pg_ddl_command_recv pg_ddl_command_send - - - ALIGNOF_POINTER p f 0 -1 0 0 _null_ _null_ _null_ ));
DESCR("internal type for passing CollectedCommand");
#define PGDDLCOMMANDOID 32
diff --git a/src/include/nodes/parsenodes.h b/src/include/nodes/parsenodes.h
index 13bfc52..5494cc4 100644
--- a/src/include/nodes/parsenodes.h
+++ b/src/include/nodes/parsenodes.h
@@ -609,6 +609,7 @@ typedef struct CreateStatsStmt
List *defnames; /* qualified name (list of Value strings) */
RangeVar *relation; /* relation to build statistics on */
List *keys; /* String nodes naming referenced column(s) */
+ List *options; /* list of DefElem nodes */
bool if_not_exists; /* do nothing if statistics already exists */
} CreateStatsStmt;
diff --git a/src/include/nodes/relation.h b/src/include/nodes/relation.h
index 3e9f930..8b7db72 100644
--- a/src/include/nodes/relation.h
+++ b/src/include/nodes/relation.h
@@ -674,9 +674,11 @@ typedef struct MVStatisticInfo
RelOptInfo *rel; /* back-link to index's table */
/* enabled statistics */
+ bool deps_enabled; /* functional dependencies enabled */
bool ndist_enabled; /* ndistinct coefficient enabled */
/* built/available statistics */
+ bool deps_built; /* functional dependencies built */
bool ndist_built; /* ndistinct coefficient built */
/* columns in the statistics (attnums) */
diff --git a/src/include/utils/builtins.h b/src/include/utils/builtins.h
index de00edd..4cb09e7 100644
--- a/src/include/utils/builtins.h
+++ b/src/include/utils/builtins.h
@@ -617,6 +617,10 @@ extern Datum pg_ndistinct_in(PG_FUNCTION_ARGS);
extern Datum pg_ndistinct_out(PG_FUNCTION_ARGS);
extern Datum pg_ndistinct_recv(PG_FUNCTION_ARGS);
extern Datum pg_ndistinct_send(PG_FUNCTION_ARGS);
+extern Datum pg_dependencies_in(PG_FUNCTION_ARGS);
+extern Datum pg_dependencies_out(PG_FUNCTION_ARGS);
+extern Datum pg_dependencies_recv(PG_FUNCTION_ARGS);
+extern Datum pg_dependencies_send(PG_FUNCTION_ARGS);
/* regexp.c */
extern Datum nameregexeq(PG_FUNCTION_ARGS);
diff --git a/src/include/utils/mvstats.h b/src/include/utils/mvstats.h
index 2937c78..227b9ff 100644
--- a/src/include/utils/mvstats.h
+++ b/src/include/utils/mvstats.h
@@ -39,22 +39,55 @@ typedef struct MVNDistinctData {
typedef MVNDistinctData *MVNDistinct;
+#define MVSTAT_DEPS_MAGIC 0xB4549A2C /* marks serialized bytea */
+#define MVSTAT_DEPS_TYPE_BASIC 1 /* basic dependencies type */
+
+/*
+ * Functional dependencies, tracking column-level relationships (values
+ * in one column determine values in another one).
+ */
+typedef struct MVDependencyData
+{
+ double degree; /* degree of validity (0-1) */
+ int nattributes; /* number of attributes */
+ int16 attributes[1]; /* attribute numbers */
+} MVDependencyData;
+
+typedef MVDependencyData *MVDependency;
+
+typedef struct MVDependenciesData
+{
+ uint32 magic; /* magic constant marker */
+ uint32 type; /* type of MV Dependencies (BASIC) */
+ int32 ndeps; /* number of dependencies */
+ MVDependency deps[1]; /* XXX why not a pointer? */
+} MVDependenciesData;
+
+typedef MVDependenciesData *MVDependencies;
+
+
+
MVNDistinct load_mv_ndistinct(Oid mvoid);
bytea *serialize_mv_ndistinct(MVNDistinct ndistinct);
+bytea *serialize_mv_dependencies(MVDependencies dependencies);
/* deserialization of stats (serialization is private to analyze) */
MVNDistinct deserialize_mv_ndistinct(bytea *data);
-
+MVDependencies deserialize_mv_dependencies(bytea *data);
MVNDistinct build_mv_ndistinct(double totalrows, int numrows, HeapTuple *rows,
- int2vector *attrs, VacAttrStats **stats);
+ int2vector *attrs, VacAttrStats **stats);
+
+MVDependencies build_mv_dependencies(int numrows, HeapTuple *rows,
+ int2vector *attrs,
+ VacAttrStats **stats);
void build_mv_stats(Relation onerel, double totalrows,
int numrows, HeapTuple *rows,
int natts, VacAttrStats **vacattrstats);
-void update_mv_stats(Oid relid, MVNDistinct ndistinct,
+void update_mv_stats(Oid relid, MVNDistinct ndistinct, MVDependencies dependencies,
int2vector *attrs, VacAttrStats **stats);
#endif
diff --git a/src/test/regress/expected/mv_dependencies.out b/src/test/regress/expected/mv_dependencies.out
new file mode 100644
index 0000000..d442a16
--- /dev/null
+++ b/src/test/regress/expected/mv_dependencies.out
@@ -0,0 +1,147 @@
+-- data type passed by value
+CREATE TABLE functional_dependencies (
+ a INT,
+ b INT,
+ c INT
+);
+-- unknown column
+CREATE STATISTICS s1 WITH (dependencies) ON (unknown_column) FROM functional_dependencies;
+ERROR: column "unknown_column" referenced in statistics does not exist
+-- single column
+CREATE STATISTICS s1 WITH (dependencies) ON (a) FROM functional_dependencies;
+ERROR: statistics require at least 2 columns
+-- single column, duplicated
+CREATE STATISTICS s1 WITH (dependencies) ON (a,a) FROM functional_dependencies;
+ERROR: duplicate column name in statistics definition
+-- two columns, one duplicated
+CREATE STATISTICS s1 WITH (dependencies) ON (a, a, b) FROM functional_dependencies;
+ERROR: duplicate column name in statistics definition
+-- correct command
+CREATE STATISTICS s1 WITH (dependencies) ON (a, b, c) FROM functional_dependencies;
+-- random data (no functional dependencies)
+INSERT INTO functional_dependencies
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+---------
+ t | f |
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 0.999900}, {0 => 2 : 0.999900}, {1 => 2 : 0.999900}, {0, 1 => 2 : 0.999900}, {0, 2 => 1 : 0.999900}]
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 0.999900}, {0 => 2 : 0.999900}, {1 => 2 : 0.494900}, {0, 1 => 2 : 0.999900}, {0, 2 => 1 : 0.999900}]
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 1.000000}, {0 => 2 : 1.000000}, {1 => 2 : 1.000000}, {0, 1 => 2 : 1.000000}, {0, 2 => 1 : 1.000000}]
+(1 row)
+
+DROP TABLE functional_dependencies;
+-- varlena type (text)
+CREATE TABLE functional_dependencies (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+CREATE STATISTICS s2 WITH (dependencies) ON (a, b, c) FROM functional_dependencies;
+-- random data (no functional dependencies)
+INSERT INTO functional_dependencies
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+---------
+ t | f |
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 0.999900}, {0 => 2 : 0.999900}, {1 => 2 : 0.999900}, {0, 1 => 2 : 0.999900}, {0, 2 => 1 : 0.999900}]
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 0.999900}, {0 => 2 : 0.999900}, {1 => 2 : 0.494900}, {0, 1 => 2 : 0.999900}, {0, 2 => 1 : 0.999900}]
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 1.000000}, {0 => 2 : 1.000000}, {1 => 2 : 1.000000}, {0, 1 => 2 : 1.000000}, {0, 2 => 1 : 1.000000}]
+(1 row)
+
+DROP TABLE functional_dependencies;
+-- NULL values (mix of int and text columns)
+CREATE TABLE functional_dependencies (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+CREATE STATISTICS s3 WITH (dependencies) ON (a, b, c, d) FROM functional_dependencies;
+INSERT INTO functional_dependencies
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
+ t | t | [{1 => 0 : 1.000000}, {2 => 0 : 1.000000}, {2 => 1 : 1.000000}, {3 => 0 : 1.000000}, {3 => 1 : 0.996700}, {0, 2 => 1 : 1.000000}, {0, 3 => 1 : 0.996700}, {1, 2 => 0 : 1.000000}, {1, 3 => 0 : 1.000000}, {2, 3 => 0 : 1.000000}, {2, 3 => 1 : 1.000000}, {0, 2, 3 => 1 : 1.000000}, {1, 2, 3 => 0 : 1.000000}]
+(1 row)
+
+DROP TABLE functional_dependencies;
diff --git a/src/test/regress/expected/mv_ndistinct.out b/src/test/regress/expected/mv_ndistinct.out
index 5f55091..06a7634 100644
--- a/src/test/regress/expected/mv_ndistinct.out
+++ b/src/test/regress/expected/mv_ndistinct.out
@@ -6,19 +6,19 @@ CREATE TABLE ndistinct (
d INT
);
-- unknown column
-CREATE STATISTICS s10 ON (unknown_column) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (unknown_column) FROM ndistinct;
ERROR: column "unknown_column" referenced in statistics does not exist
-- single column
-CREATE STATISTICS s10 ON (a) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a) FROM ndistinct;
ERROR: statistics require at least 2 columns
-- single column, duplicated
-CREATE STATISTICS s10 ON (a,a) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a,a) FROM ndistinct;
ERROR: duplicate column name in statistics definition
-- two columns, one duplicated
-CREATE STATISTICS s10 ON (a, a, b) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a, a, b) FROM ndistinct;
ERROR: duplicate column name in statistics definition
-- correct command
-CREATE STATISTICS s10 ON (a, b, c) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a, b, c) FROM ndistinct;
-- perfectly correlated groups
INSERT INTO ndistinct
SELECT i/100, i/100, i/100 FROM generate_series(1,10000) s(i);
diff --git a/src/test/regress/expected/object_address.out b/src/test/regress/expected/object_address.out
index f85381a..b9dd138 100644
--- a/src/test/regress/expected/object_address.out
+++ b/src/test/regress/expected/object_address.out
@@ -36,7 +36,7 @@ ALTER DEFAULT PRIVILEGES FOR ROLE regress_addr_user REVOKE DELETE ON TABLES FROM
CREATE TRANSFORM FOR int LANGUAGE SQL (
FROM SQL WITH FUNCTION varchar_transform(internal),
TO SQL WITH FUNCTION int4recv(internal));
-CREATE STATISTICS addr_nsp.gentable_stat ON (a,b) FROM addr_nsp.gentable;
+CREATE STATISTICS addr_nsp.gentable_stat WITH (ndistinct) ON (a,b) FROM addr_nsp.gentable;
-- test some error cases
SELECT pg_get_object_address('stone', '{}', '{}');
ERROR: unrecognized object type "stone"
diff --git a/src/test/regress/expected/opr_sanity.out b/src/test/regress/expected/opr_sanity.out
index 9a26205..db1cf8a 100644
--- a/src/test/regress/expected/opr_sanity.out
+++ b/src/test/regress/expected/opr_sanity.out
@@ -818,11 +818,12 @@ WHERE c.castmethod = 'b' AND
character varying | character | 0 | i
pg_node_tree | text | 0 | i
pg_ndistinct | bytea | 0 | i
+ pg_dependencies | bytea | 0 | i
cidr | inet | 0 | i
xml | text | 0 | a
xml | character varying | 0 | a
xml | character | 0 | a
-(8 rows)
+(9 rows)
-- **************** pg_conversion ****************
-- Look for illegal values in pg_conversion fields.
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 7599d6c..ea3f51d 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1380,7 +1380,8 @@ pg_mv_stats| SELECT n.nspname AS schemaname,
c.relname AS tablename,
s.staname,
s.stakeys AS attnums,
- length((s.standist)::text) AS ndistbytes
+ length((s.standist)::bytea) AS ndistbytes,
+ length((s.stadeps)::bytea) AS depsbytes
FROM ((pg_mv_statistic s
JOIN pg_class c ON ((c.oid = s.starelid)))
LEFT JOIN pg_namespace n ON ((n.oid = c.relnamespace)));
diff --git a/src/test/regress/expected/type_sanity.out b/src/test/regress/expected/type_sanity.out
index 9356aaa..8b849b9 100644
--- a/src/test/regress/expected/type_sanity.out
+++ b/src/test/regress/expected/type_sanity.out
@@ -67,13 +67,14 @@ WHERE p1.typtype not in ('c','d','p') AND p1.typname NOT LIKE E'\\_%'
(SELECT 1 FROM pg_type as p2
WHERE p2.typname = ('_' || p1.typname)::name AND
p2.typelem = p1.oid and p1.typarray = p2.oid);
- oid | typname
-------+--------------
+ oid | typname
+------+-----------------
194 | pg_node_tree
3343 | pg_ndistinct
+ 3353 | pg_dependencies
210 | smgr
705 | unknown
-(4 rows)
+(5 rows)
-- Make sure typarray points to a varlena array type of our own base
SELECT p1.oid, p1.typname as basetype, p2.typname as arraytype,
diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule
index b70b82f..ecb6f04 100644
--- a/src/test/regress/parallel_schedule
+++ b/src/test/regress/parallel_schedule
@@ -115,4 +115,4 @@ test: event_trigger
test: stats
# run tests of multivariate stats
-test: mv_ndistinct
+test: mv_ndistinct mv_dependencies
diff --git a/src/test/regress/serial_schedule b/src/test/regress/serial_schedule
index 1c5bfaf..fa0e993 100644
--- a/src/test/regress/serial_schedule
+++ b/src/test/regress/serial_schedule
@@ -170,3 +170,4 @@ test: xml
test: event_trigger
test: stats
test: mv_ndistinct
+test: mv_dependencies
diff --git a/src/test/regress/sql/mv_dependencies.sql b/src/test/regress/sql/mv_dependencies.sql
new file mode 100644
index 0000000..43df798
--- /dev/null
+++ b/src/test/regress/sql/mv_dependencies.sql
@@ -0,0 +1,139 @@
+-- data type passed by value
+CREATE TABLE functional_dependencies (
+ a INT,
+ b INT,
+ c INT
+);
+
+-- unknown column
+CREATE STATISTICS s1 WITH (dependencies) ON (unknown_column) FROM functional_dependencies;
+
+-- single column
+CREATE STATISTICS s1 WITH (dependencies) ON (a) FROM functional_dependencies;
+
+-- single column, duplicated
+CREATE STATISTICS s1 WITH (dependencies) ON (a,a) FROM functional_dependencies;
+
+-- two columns, one duplicated
+CREATE STATISTICS s1 WITH (dependencies) ON (a, a, b) FROM functional_dependencies;
+
+-- correct command
+CREATE STATISTICS s1 WITH (dependencies) ON (a, b, c) FROM functional_dependencies;
+
+-- random data (no functional dependencies)
+INSERT INTO functional_dependencies
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+DROP TABLE functional_dependencies;
+
+-- varlena type (text)
+CREATE TABLE functional_dependencies (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+
+CREATE STATISTICS s2 WITH (dependencies) ON (a, b, c) FROM functional_dependencies;
+
+-- random data (no functional dependencies)
+INSERT INTO functional_dependencies
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+DROP TABLE functional_dependencies;
+
+-- NULL values (mix of int and text columns)
+CREATE TABLE functional_dependencies (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+
+CREATE STATISTICS s3 WITH (dependencies) ON (a, b, c, d) FROM functional_dependencies;
+
+INSERT INTO functional_dependencies
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+DROP TABLE functional_dependencies;
diff --git a/src/test/regress/sql/mv_ndistinct.sql b/src/test/regress/sql/mv_ndistinct.sql
index 5cef254..43024ca 100644
--- a/src/test/regress/sql/mv_ndistinct.sql
+++ b/src/test/regress/sql/mv_ndistinct.sql
@@ -7,19 +7,19 @@ CREATE TABLE ndistinct (
);
-- unknown column
-CREATE STATISTICS s10 ON (unknown_column) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (unknown_column) FROM ndistinct;
-- single column
-CREATE STATISTICS s10 ON (a) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a) FROM ndistinct;
-- single column, duplicated
-CREATE STATISTICS s10 ON (a,a) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a,a) FROM ndistinct;
-- two columns, one duplicated
-CREATE STATISTICS s10 ON (a, a, b) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a, a, b) FROM ndistinct;
-- correct command
-CREATE STATISTICS s10 ON (a, b, c) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a, b, c) FROM ndistinct;
-- perfectly correlated groups
INSERT INTO ndistinct
diff --git a/src/test/regress/sql/object_address.sql b/src/test/regress/sql/object_address.sql
index 956bef3..bf3323f 100644
--- a/src/test/regress/sql/object_address.sql
+++ b/src/test/regress/sql/object_address.sql
@@ -39,7 +39,7 @@ ALTER DEFAULT PRIVILEGES FOR ROLE regress_addr_user REVOKE DELETE ON TABLES FROM
CREATE TRANSFORM FOR int LANGUAGE SQL (
FROM SQL WITH FUNCTION varchar_transform(internal),
TO SQL WITH FUNCTION int4recv(internal));
-CREATE STATISTICS addr_nsp.gentable_stat ON (a,b) FROM addr_nsp.gentable;
+CREATE STATISTICS addr_nsp.gentable_stat WITH (ndistinct) ON (a,b) FROM addr_nsp.gentable;
-- test some error cases
SELECT pg_get_object_address('stone', '{}', '{}');
--
2.5.5
[binary/octet-stream] 0004-PATCH-selectivity-estimation-using-functional-de-v22.patch (46.5K, ../../696aa95c-2411-9b2b-f36e-65b66bf47c88@2ndquadrant.com/5-0004-PATCH-selectivity-estimation-using-functional-de-v22.patch)
download | inline diff:
From 89af849280f99c6dab0d2b642677732461440ed0 Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:37:27 +0200
Subject: [PATCH 4/9] PATCH: selectivity estimation using functional
dependencies
Use functional dependencies to correct selectivity estimates of
equality clauses. For now this only works with regular WHERE
conditions, not join clauses etc.
Given two equality clauses
(a = 1) AND (b = 2)
we compute selectivity for each condition, and then combine them
using formula
P(a=1, b=2) = P(a=1) * [degree + (1 - degree) * P(b=2)]
where 'degree' of the functional dependence (a => b) is a number
between [0,1] measuring how much the knowledge of 'a' determines
the value of 'b'. For 'degree=0' this degrades to independence,
for 'degree=1' we get perfect functional dependency.
Estimates of more than two clauses are computed recursively, so
for example
(a = 1) AND (b = 2) AND (c = 3)
is first split into
P(a=1, b=2, c=3) = P(a=1, b=2) * [d + (1-d) * P(c=3)]
where 'd' is degree of (a,b => c) functional dependency. And then
the first part of the estimate is computed recursively:
P(a=1, b=2) = P(a=1) * [d + (1-d) * P(b=2)]
where 'd' is degree of (a => b) dependency.
The patch includes regression tests with functional dependencies
on several synthetic datasets (random, perfectly correlated, etc.)
---
doc/src/sgml/planstats.sgml | 178 +++++-
src/backend/optimizer/path/clausesel.c | 781 +++++++++++++++++++++++++-
src/backend/utils/mvstats/README.stats | 45 +-
src/backend/utils/mvstats/common.c | 1 +
src/backend/utils/mvstats/dependencies.c | 68 +++
src/include/utils/mvstats.h | 6 +-
src/test/regress/expected/mv_dependencies.out | 28 +-
src/test/regress/sql/mv_dependencies.sql | 19 +-
8 files changed, 1072 insertions(+), 54 deletions(-)
diff --git a/doc/src/sgml/planstats.sgml b/doc/src/sgml/planstats.sgml
index e9248b4..bdad2db 100644
--- a/doc/src/sgml/planstats.sgml
+++ b/doc/src/sgml/planstats.sgml
@@ -504,7 +504,7 @@ SELECT relpages, reltuples FROM pg_class WHERE relname = 't';
<programlisting>
EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1;
- QUERY PLAN
+ QUERY PLAN
-------------------------------------------------------------------------------------------------
Seq Scan on t (cost=0.00..170.00 rows=100 width=8) (actual time=0.031..2.870 rows=100 loops=1)
Filter: (a = 1)
@@ -527,7 +527,7 @@ EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1;
<programlisting>
EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
- QUERY PLAN
+ QUERY PLAN
-----------------------------------------------------------------------------------------------
Seq Scan on t (cost=0.00..195.00 rows=1 width=8) (actual time=0.033..3.006 rows=100 loops=1)
Filter: ((a = 1) AND (b = 1))
@@ -547,11 +547,11 @@ EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
<para>
Overestimates, i.e. errors in the opposite direction, are also possible.
Consider for example the following combination of range conditions, each
- matching
+ matching roughly half the rows.
<programlisting>
EXPLAIN ANALYZE SELECT * FROM t WHERE a <= 49 AND b > 49;
- QUERY PLAN
+ QUERY PLAN
------------------------------------------------------------------------------------------------
Seq Scan on t (cost=0.00..195.00 rows=2500 width=8) (actual time=1.607..1.607 rows=0 loops=1)
Filter: ((a <= 49) AND (b > 49))
@@ -587,6 +587,176 @@ EXPLAIN ANALYZE SELECT * FROM t WHERE a <= 49 AND b > 49;
sections.
</para>
+ <sect2 id="functional-dependencies">
+ <title>Functional Dependencies</title>
+
+ <para>
+ The simplest type of multivariate statistics are functional dependencies,
+ used in definitions of database normal forms. When simplified, saying that
+ <literal>b</> is functionally dependent on <literal>a</> means that
+ knowledge of value of <literal>a</> is sufficient to determine value of
+ <literal>b</>.
+ </para>
+
+ <para>
+ In normalized databases, only functional dependencies on primary keys
+ and super keys are allowed. In practice however many data sets are not
+ fully normalized, for example thanks to intentional denormalization for
+ performance reasons. The table <literal>t</> is an example of a data set
+ with functional dependencies. As <literal>a = b</> for all rows in the
+ table, <literal>a</> is functionally dependent on <literal>b</> and
+ <literal>b</> is functionally dependent on <literal>a</literal>.
+ </para>
+
+ <para>
+ Functional dependencies directly affect accuracy of the estimates, as
+ conditions on the dependent column(s) do not restrict the result set,
+ and are often redundant, causing underestimates. In the first example,
+ either <literal>a = 1</> or <literal>b = 1</> is sufficient (however see
+ <xref linkend="functional-dependencies-limitations">).
+ </para>
+
+ <para>
+ To inform the planner about the functional dependencies, or rather to
+ instruct it to search for them during <command>ANALYZE</>, we can use
+ the <command>CREATE STATISTICS</> command.
+
+<programlisting>
+CREATE STATISTICS s1 ON t (a,b) WITH (dependencies);
+ANALYZE t;
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
+ QUERY PLAN
+-------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=100 width=8) (actual time=0.095..3.118 rows=100 loops=1)
+ Filter: ((a = 1) AND (b = 1))
+ Rows Removed by Filter: 9900
+ Planning time: 0.367 ms
+ Execution time: 3.380 ms
+(5 rows)
+</programlisting>
+
+ As you can see, the estimate improved quite a bit, as the planner is now
+ aware of the functional dependencies and eliminates the second condition
+ when computing the estimates.
+ </para>
+
+ <para>
+ Let's inspect multivariate statistics on a table, as defined by
+ <command>CREATE STATISTICS</> and built by <command>ANALYZE</>. If you're
+ using <application>psql</>, the easiest way to list statistics on a table
+ is by using <command>\d</>.
+
+<programlisting>
+\d t
+ Table "public.t"
+ Column | Type | Modifiers
+--------+---------+-----------
+ a | integer |
+ b | integer |
+Statistics:
+ "public.s1" (dependencies) ON (a, b)
+</programlisting>
+
+ </para>
+
+ <para>
+ Similarly to per-column statistics, multivariate statistics are stored in
+ a system catalog called <structname>pg_mv_statistic</structname>, but
+ there is also a more convenient view <structname>pg_mv_stats</structname>.
+ To inspect the statistics <literal>s1</literal> defined above,
+ you may do this:
+
+<programlisting>
+SELECT tablename, staname, attnums, depsbytes, depsinfo
+ FROM pg_mv_stats WHERE staname = 's1';
+
+ tablename | staname | attnums | depsbytes | depsinfo
+-----------+---------+---------+-----------+----------------
+ t | s1 | 1 2 | 32 | dependencies=2
+(1 row)
+</programlisting>
+
+ This shows that the statistic is defined on table <structname>t</>,
+ <structfield>attnums</structfield> lists attribute numbers of columns
+ (references <structname>pg_attribute</structname>). It also shows
+ <command>ANALYZE</> found two functional dependencies, and size when
+ serialized into a <literal>bytea</> column. Inspecting the functional
+ dependencies is possible using <function>pg_mv_stats_dependencies_show</>
+ function.
+
+<programlisting>
+SELECT pg_mv_stats_dependencies_show(stadeps)
+ FROM pg_mv_statistic WHERE staname = 's1';
+
+ pg_mv_stats_dependencies_show
+-------------------------------
+ (1) => 2, (2) => 1
+(1 row)
+</programlisting>
+
+ Which confirms <literal>a</> is functionally dependent on <literal>b</> and
+ <literal>b</> is functionally dependent on <literal>a</literal>.
+ </para>
+
+ <para>
+ Now let's quickly discuss how this knowledge is applied when estimating
+ the selectivity. The planner walks through the conditions and attempts
+ to identify which conditions are already implied by other conditions,
+ and eliminates them (but only for the estimation, all conditions will be
+ checked on tuples during execution). In the example query, either of
+ the conditions may get eliminated, improving the estimate. This happens
+ in <function>clauselist_apply_dependencies</> in <filename>clausesel.c</>.
+ </para>
+
+ <sect3 id="functional-dependencies-limitations">
+ <title>Limitations of functional dependencies</title>
+
+ <para>
+ The first limitation of functional dependencies is that they only work
+ with simple equality conditions, comparing columns and constant values.
+ It's not possible to use them to eliminate equality conditions comparing
+ two columns or a column to an expression, range clauses, <literal>LIKE</>
+ or any other type of condition.
+ </para>
+
+ <para>
+ When eliminating the implied conditions, the planner assumes that the
+ conditions are compatible. Consider the following example, violating
+ this assumption:
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 10;
+ QUERY PLAN
+-----------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=100 width=8) (actual time=2.992..2.992 rows=0 loops=1)
+ Filter: ((a = 1) AND (b = 10))
+ Rows Removed by Filter: 10000
+ Planning time: 0.232 ms
+ Execution time: 3.033 ms
+(5 rows)
+</programlisting>
+
+ There are no rows with this combination of values, however the planner
+ is unable to verify whether the values match - it only knows that
+ the columns are functionally dependent.
+ </para>
+
+ <para>
+ This assumption is more about queries executed on the database - in many
+ cases it's actually satisfied (e.g. when the GUI only allows selecting
+ compatible values). But if that's not the case, functional dependencies
+ may not be a viable option.
+ </para>
+
+ <para>
+ For additional information about functional dependencies, see
+ <filename>src/backend/utils/mvstats/README.dependencies</>.
+ </para>
+
+ </sect3>
+
+ </sect2>
+
</sect1>
</chapter>
diff --git a/src/backend/optimizer/path/clausesel.c b/src/backend/optimizer/path/clausesel.c
index 02660c2..ec74b16 100644
--- a/src/backend/optimizer/path/clausesel.c
+++ b/src/backend/optimizer/path/clausesel.c
@@ -14,14 +14,19 @@
*/
#include "postgres.h"
+#include "access/sysattr.h"
+#include "catalog/pg_operator.h"
#include "nodes/makefuncs.h"
#include "optimizer/clauses.h"
#include "optimizer/cost.h"
#include "optimizer/pathnode.h"
#include "optimizer/plancat.h"
+#include "optimizer/var.h"
#include "utils/fmgroids.h"
#include "utils/lsyscache.h"
+#include "utils/mvstats.h"
#include "utils/selfuncs.h"
+#include "utils/typcache.h"
/*
@@ -41,6 +46,33 @@ typedef struct RangeQueryClause
static void addRangeClause(RangeQueryClause **rqlist, Node *clause,
bool varonleft, bool isLTsel, Selectivity s2);
+#define STATS_TYPE_FDEPS 0x01
+
+static bool clause_is_mv_compatible(Node *clause, Index relid, AttrNumber *attnum);
+
+static Bitmapset *collect_mv_attnums(List *clauses, Index relid);
+
+static int count_mv_attnums(List *clauses, Index relid);
+
+static int count_varnos(List *clauses, Index *relid);
+
+static MVStatisticInfo *choose_mv_statistics(List *mvstats, Bitmapset *attnums,
+ int types);
+
+static List *clauselist_mv_split(PlannerInfo *root, Index relid,
+ List *clauses, List **mvclauses,
+ MVStatisticInfo *mvstats, int types);
+
+static Selectivity clauselist_mv_selectivity_deps(PlannerInfo *root,
+ Index relid, List *clauses, MVStatisticInfo *mvstats,
+ Index varRelid, JoinType jointype, SpecialJoinInfo *sjinfo);
+
+static bool has_stats(List *stats, int type);
+
+static List *find_stats(PlannerInfo *root, Index relid);
+
+static bool stats_type_matches(MVStatisticInfo *stat, int type);
+
/****************************************************************************
* ROUTINES TO COMPUTE SELECTIVITIES
@@ -60,7 +92,19 @@ static void addRangeClause(RangeQueryClause **rqlist, Node *clause,
* subclauses. However, that's only right if the subclauses have independent
* probabilities, and in reality they are often NOT independent. So,
* we want to be smarter where we can.
-
+ *
+ * The first thing we try to do is applying multivariate statistics, in a way
+ * that intends to minimize the overhead when there are no multivariate stats
+ * on the relation. Thus we do several simple (and inexpensive) checks first,
+ * to verify that suitable multivariate statistics exist.
+ *
+ * If we identify such multivariate statistics apply, we try to apply them.
+ * Currently we only have (soft) functional dependencies, so we try to reduce
+ * the list of clauses.
+ *
+ * Then we remove the clauses estimated using multivariate stats, and process
+ * the rest of the clauses using the regular per-column stats.
+ *
* Currently, the only extra smarts we have is to recognize "range queries",
* such as "x > 34 AND x < 42". Clauses are recognized as possible range
* query components if they are restriction opclauses whose operators have
@@ -99,15 +143,81 @@ clauselist_selectivity(PlannerInfo *root,
RangeQueryClause *rqlist = NULL;
ListCell *l;
+ /* processing mv stats */
+ Oid relid = InvalidOid;
+
+ /* list of multivariate stats on the relation */
+ List *stats = NIL;
+
/*
- * If there's exactly one clause, then no use in trying to match up pairs,
- * so just go directly to clause_selectivity().
+ * If there's exactly one clause, then multivariate statistics is futile
+ * at this level (we might be able to apply them later if it's AND/OR
+ * clause). So just go directly to clause_selectivity().
*/
if (list_length(clauses) == 1)
return clause_selectivity(root, (Node *) linitial(clauses),
varRelid, jointype, sjinfo);
/*
+ * To fetch the statistics, we first need to determine the rel. Currently
+ * we only support estimates of simple restrictions referencing a single
+ * baserel (no join statistics). However set_baserel_size_estimates() sets
+ * varRelid=0 so we have to actually inspect the clauses by pull_varnos
+ * and see if there's just a single varno referenced.
+ *
+ * XXX Maybe there's a better way to find the relid?
+ */
+ if ((count_varnos(clauses, &relid) == 1) &&
+ ((varRelid == 0) || (varRelid == relid)))
+ stats = find_stats(root, relid);
+
+ /*
+ * Check that there are multivariate statistics usable for selectivity
+ * estimation, i.e. anything except ndistinct coefficients.
+ *
+ * Also check the number of attributes in clauses that might be estimated
+ * using those statistics, and that there are at least two such attributes.
+ * It may easily happen that we won't be able to estimate the clauses using
+ * the multivariate statistics anyway, but that requires a more expensive
+ * to verify (so the check check should be worth it).
+ *
+ * If there are no such stats or not enough attributes, don't waste time
+ * simply skip to estimation using the plain per-column stats.
+ */
+ if (has_stats(stats, STATS_TYPE_FDEPS) &&
+ (count_mv_attnums(clauses, relid) >= 2))
+ {
+ MVStatisticInfo *mvstat;
+ Bitmapset *mvattnums;
+
+ /* collect attributes from the compatible conditions */
+ mvattnums = collect_mv_attnums(clauses, relid);
+
+ /* and search for the statistic covering the most attributes */
+ mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_FDEPS);
+
+ if (mvstat != NULL) /* we have a matching stats */
+ {
+ /* clauses compatible with multi-variate stats */
+ List *mvclauses = NIL;
+
+ /* split the clauselist into regular and mv-clauses */
+ clauses = clauselist_mv_split(root, relid, clauses, &mvclauses,
+ mvstat, STATS_TYPE_FDEPS);
+
+ /* Empty list of clauses is a clear sign something went wrong. */
+ Assert(list_length(mvclauses));
+
+ /* we've chosen the histogram to match the clauses */
+ Assert(mvclauses != NIL);
+
+ /* compute the multivariate stats (dependencies) */
+ s1 *= clauselist_mv_selectivity_deps(root, relid, mvclauses, mvstat,
+ varRelid, jointype, sjinfo);
+ }
+ }
+
+ /*
* Initial scan over clauses. Anything that doesn't look like a potential
* rangequery clause gets multiplied into s1 and forgotten. Anything that
* does gets inserted into an rqlist entry.
@@ -763,3 +873,668 @@ clause_selectivity(PlannerInfo *root,
return s1;
}
+
+/*
+ * When applying functional dependencies, we start with the strongest ones
+ * strongest dependencies. That is, we select the dependency that:
+ *
+ * (a) has all attributes covered by the clauses
+ *
+ * (b) has the most attributes
+ *
+ * (c) has the higher degree of validity
+ *
+ * TODO Explain why we select the dependencies this way.
+ */
+static MVDependency
+find_strongest_dependency(MVStatisticInfo *mvstats, MVDependencies dependencies,
+ Bitmapset *attnums)
+{
+ int i;
+ MVDependency strongest = NULL;
+
+ /* number of attnums in clauses */
+ int nattnums = bms_num_members(attnums);
+
+ /*
+ * Iterate over the MVDependency items and find the strongest one from
+ * the fully-matched dependencies. We do the cheap checks first, before
+ * matching it against the attnums.
+ */
+ for (i = 0; i < dependencies->ndeps; i++)
+ {
+ MVDependency dependency = dependencies->deps[i];
+
+ /*
+ * Skip dependencies referencing more attributes than available clauses,
+ * as those can't be fully matched.
+ */
+ if (dependency->nattributes > nattnums)
+ continue;
+
+ /* We can skip dependencies on fewer attributes than the best one. */
+ if (strongest && (strongest->nattributes > dependency->nattributes))
+ continue;
+
+ /* And also weaker dependencies on the same number of attributes. */
+ if (strongest &&
+ (strongest->nattributes == dependency->nattributes) &&
+ (strongest->degree > dependency->degree))
+ continue;
+
+ /*
+ * Check that the dependency actually is fully covered by clauses.
+ * If the dependency is not fully matched by clauses, we can't use
+ * it for the estimation.
+ */
+ if (! dependency_is_fully_matched(dependency, attnums,
+ mvstats->stakeys->values))
+ continue;
+
+ /*
+ * We have a fully-matched dependency, and we already know it has to
+ * be stronger than the current one (otherwise we'd skip it before
+ * inspecting it at the very beginning.
+ */
+ strongest = dependency;
+ }
+
+ return strongest;
+}
+
+/*
+ * clauselist_mv_selectivity_deps
+ * estimate selectivity using functional dependencies
+ *
+ * Given equality clauses on attributes (a,b) we find the strongest dependency
+ * between them, i.e. either (a=>b) or (b=>a). Assuming (a=>b) is the selected
+ * dependency, we then combine the per-clause selectivities using the formula
+ *
+ * P(a,b) = P(a) * [f + (1-f)*P(b)]
+ *
+ * where 'f' is the degree of the dependency.
+ *
+ * With clauses on more than two attributes, the dependencies are applied
+ * recursively, starting with the widest/strongest dependencies. For example
+ * P(a,b,c) is first split like this:
+ *
+ * P(a,b,c) = P(a,b) * [f + (1-f)*P(c)]
+ *
+ * assuming (a,b=>c) is the strongest dependency.
+ */
+static Selectivity
+clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
+ List *clauses, MVStatisticInfo *mvstats,
+ Index varRelid, JoinType jointype,
+ SpecialJoinInfo *sjinfo)
+{
+ ListCell *lc;
+ Selectivity s1 = 1.0;
+ MVDependencies dependencies;
+
+ Assert(mvstats->deps_enabled && mvstats->deps_built);
+
+ /* load the dependency items stored in the statistics */
+ dependencies = load_mv_dependencies(mvstats->mvoid);
+
+ Assert(dependencies);
+
+ /*
+ * Apply the dependencies recursively, starting with the widest/strongest
+ * ones, and proceeding to the smaller/weaker ones. At the end of each
+ * round we factor in the selectivity of clauses on the implied attribute,
+ * and remove the clauses from the list.
+ */
+ while (true)
+ {
+ Selectivity s2 = 1.0;
+ Bitmapset *attnums;
+ MVDependency dependency;
+
+ /* clauses remaining after removing those on the "implied" attribute */
+ List *clauses_filtered = NIL;
+
+ attnums = collect_mv_attnums(clauses, relid);
+
+ /* no point in looking for dependencies with fewer than 2 attributes */
+ if (bms_num_members(attnums) < 2)
+ break;
+
+ /* the widest/strongest dependency, fully matched by clauses */
+ dependency = find_strongest_dependency(mvstats, dependencies, attnums);
+
+ /* if no suitable dependency was found, we're done */
+ if (! dependency)
+ break;
+
+ /*
+ * We found an applicable dependency, so find all the clauses on the
+ * implied attribute, so with dependency (a,b => c) we seach clauses
+ * on 'c'. We only really expect a single such clause, but in case
+ * there are more we simply multiply the selectivities as usual.
+ *
+ * XXX Maybe we should use the maximum, minimum or just error out?
+ */
+ foreach(lc, clauses)
+ {
+ AttrNumber attnum_clause = InvalidAttrNumber;
+ Node *clause = (Node *) lfirst(lc);
+
+ /*
+ * XXX We need the attnum referenced by the clause, and this is the
+ * easiest way to get it (but maybe not the best one). At this point
+ * we should only see equality clauses compatible with functional
+ * dependencies, so just error out if we stumble upon something else.
+ */
+ if (! clause_is_mv_compatible(clause, relid, &attnum_clause))
+ elog(ERROR, "clause not compatible with functional dependencies");
+
+ Assert(AttributeNumberIsValid(attnum_clause));
+
+ /*
+ * If the clause is not on the implied attribute, add it to the list
+ * of filtered clauses (for the next round) and continue with the
+ * next one.
+ */
+ if (! dependency_implies_attribute(dependency, attnum_clause,
+ mvstats->stakeys->values))
+ {
+ clauses_filtered = lappend(clauses_filtered, clause);
+ continue;
+ }
+
+ /*
+ * Otherwise compute selectivity of the clause, and multiply it with
+ * other clauses on the same attribute.
+ *
+ * XXX Not sure if we need to worry about multiple clauses, though.
+ * Those are all equality clauses, and if they reference different
+ * constants, that's not going to work.
+ */
+ s2 *= clause_selectivity(root, clause, varRelid, jointype, sjinfo);
+ }
+
+ /*
+ * Now factor in the selectivity for all the "implied" clauses into the
+ * final one, using this formula:
+ *
+ * P(a,b) = P(a) * (f + (1-f) * P(b))
+ *
+ * where 'f' is the degree of validity of the dependency.
+ */
+ s1 *= (dependency->degree + (1 - dependency->degree) * s2);
+
+ /* And only keep the filtered clauses for the next round. */
+ clauses = clauses_filtered;
+ }
+
+ /* And now simply multiply with selectivities of the remaining clauses. */
+ foreach (lc, clauses)
+ {
+ Node *clause = (Node *) lfirst(lc);
+
+ s1 *= clause_selectivity(root, clause, varRelid, jointype, sjinfo);
+ }
+
+ return s1;
+}
+
+/*
+ * Collect attributes from mv-compatible clauses.
+ */
+static Bitmapset *
+collect_mv_attnums(List *clauses, Index relid)
+{
+ Bitmapset *attnums = NULL;
+ ListCell *l;
+
+ /*
+ * Walk through the clauses and identify the ones we can estimate using
+ * multivariate stats, and remember the relid/columns. We'll then
+ * cross-check if we have suitable stats, and only if needed we'll split
+ * the clauses into multivariate and regular lists.
+ *
+ * For now we're only interested in RestrictInfo nodes with nested OpExpr,
+ * using either a range or equality.
+ */
+ foreach(l, clauses)
+ {
+ AttrNumber attnum;
+ Node *clause = (Node *) lfirst(l);
+
+ /* ignore the result for now - we only need the info */
+ if (clause_is_mv_compatible(clause, relid, &attnum))
+ attnums = bms_add_member(attnums, attnum);
+ }
+
+ /*
+ * If there are not at least two attributes referenced by the clause(s),
+ * we can throw everything out (as we'll revert to simple stats).
+ */
+ if (bms_num_members(attnums) <= 1)
+ {
+ if (attnums != NULL)
+ pfree(attnums);
+ attnums = NULL;
+ }
+
+ return attnums;
+}
+
+/*
+ * Count the number of attributes in clauses compatible with multivariate stats.
+ */
+static int
+count_mv_attnums(List *clauses, Index relid)
+{
+ int c;
+ Bitmapset *attnums = collect_mv_attnums(clauses, relid);
+
+ c = bms_num_members(attnums);
+
+ bms_free(attnums);
+
+ return c;
+}
+
+/*
+ * Count varnos referenced in the clauses, and if there's a single varno then
+ * return the index in 'relid'.
+ */
+static int
+count_varnos(List *clauses, Index *relid)
+{
+ int cnt;
+ Bitmapset *varnos = NULL;
+
+ varnos = pull_varnos((Node *) clauses);
+ cnt = bms_num_members(varnos);
+
+ /* if there's a single varno in the clauses, remember it */
+ if (bms_num_members(varnos) == 1)
+ *relid = bms_singleton_member(varnos);
+
+ bms_free(varnos);
+
+ return cnt;
+}
+
+static int
+count_attnums_covered_by_stats(MVStatisticInfo *info, Bitmapset *attnums)
+{
+ int i;
+ int matches = 0;
+ int2vector *attrs = info->stakeys;
+
+ /* count columns covered by the statistics */
+ for (i = 0; i < attrs->dim1; i++)
+ if (bms_is_member(attrs->values[i], attnums))
+ matches++;
+
+ return matches;
+}
+
+/*
+ * We're looking for statistics matching at least 2 attributes, referenced in
+ * clauses compatible with multivariate statistics. The current selection
+ * criteria is very simple - we choose the statistics referencing the most
+ * attributes.
+ *
+ * If there are multiple statistics referencing the same number of columns
+ * (from the clauses), the one with less source columns (as listed in the
+ * ADD STATISTICS when creating the statistics) wins. Else the first one wins.
+ *
+ * This is a very simple criteria, and has several weaknesses:
+ *
+ * (a) does not consider the accuracy of the statistics
+ *
+ * If there are two histograms built on the same set of columns, but one
+ * has 100 buckets and the other one has 1000 buckets (thus likely
+ * providing better estimates), this is not currently considered.
+ *
+ * (b) does not consider the type of statistics
+ *
+ * If there are three statistics - one containing just a MCV list, another
+ * one with just a histogram and a third one with both, we treat them equally.
+ *
+ * (c) does not consider the number of clauses
+ *
+ * As explained, only the number of referenced attributes counts, so if
+ * there are multiple clauses on a single attribute, this still counts as
+ * a single attribute.
+ *
+ * (d) does not consider type of condition
+ *
+ * Some clauses may work better with some statistics - for example equality
+ * clauses probably work better with MCV lists than with histograms. But
+ * IS [NOT] NULL conditions may often work better with histograms (thanks
+ * to NULL-buckets).
+ *
+ * So for example with five WHERE conditions
+ *
+ * WHERE (a = 1) AND (b = 1) AND (c = 1) AND (d = 1) AND (e = 1)
+ *
+ * and statistics on (a,b), (a,b,e) and (a,b,c,d), the last one will be selected
+ * as it references the most columns.
+ *
+ * Once we have selected the multivariate statistics, we split the list of
+ * clauses into two parts - conditions that are compatible with the selected
+ * stats, and conditions are estimated using simple statistics.
+ *
+ * From the example above, conditions
+ *
+ * (a = 1) AND (b = 1) AND (c = 1) AND (d = 1)
+ *
+ * will be estimated using the multivariate statistics (a,b,c,d) while the last
+ * condition (e = 1) will get estimated using the regular ones.
+ *
+ * There are various alternative selection criteria (e.g. counting conditions
+ * instead of just referenced attributes), but eventually the best option should
+ * be to combine multiple statistics. But that's much harder to do correctly.
+ *
+ * TODO: Select multiple statistics and combine them when computing the estimate.
+ *
+ * TODO: This will probably have to consider compatibility of clauses, because
+ * 'dependencies' will probably work only with equality clauses.
+ */
+static MVStatisticInfo *
+choose_mv_statistics(List *stats, Bitmapset *attnums, int types)
+{
+ ListCell *lc;
+
+ MVStatisticInfo *choice = NULL;
+
+ int current_matches = 2; /* goal #1: maximize */
+ int current_dims = (MVSTATS_MAX_DIMENSIONS + 1); /* goal #2: minimize */
+
+ /*
+ * Walk through the statistics (simple array with nmvstats elements) and
+ * for each one count the referenced attributes (encoded in the 'attnums'
+ * bitmap).
+ */
+ foreach(lc, stats)
+ {
+ MVStatisticInfo *info = (MVStatisticInfo *) lfirst(lc);
+
+ /* columns matching this statistics */
+ int matches = 0;
+
+ /* size (number of dimensions) of this statistics */
+ int numattrs = info->stakeys->dim1;
+
+ /* skip statistics not matching any of the requested types */
+ if (! (info->deps_built && (STATS_TYPE_FDEPS & types)))
+ continue;
+
+ /* count columns covered by the statistics */
+ matches = count_attnums_covered_by_stats(info, attnums);
+
+ /*
+ * Use this statistics when it increases the number of matched clauses
+ * or when it matches the same number of attributes but is smaller
+ * (in terms of number of attributes covered).
+ */
+ if ((matches > current_matches) ||
+ ((matches == current_matches) && (current_dims > numattrs)))
+ {
+ choice = info;
+ current_matches = matches;
+ current_dims = numattrs;
+ }
+ }
+
+ return choice;
+}
+
+
+/*
+ * clauselist_mv_split
+ * split the clause list into a part to be estimated using the provided
+ * statistics, and remaining clauses (estimated in some other way)
+ */
+static List *
+clauselist_mv_split(PlannerInfo *root, Index relid,
+ List *clauses, List **mvclauses,
+ MVStatisticInfo *mvstats, int types)
+{
+ int i;
+ ListCell *l;
+ List *non_mvclauses = NIL;
+
+ /* FIXME is there a better way to get info on int2vector? */
+ int2vector *attrs = mvstats->stakeys;
+ int numattrs = mvstats->stakeys->dim1;
+
+ Bitmapset *mvattnums = NULL;
+
+ /* build bitmap of attributes, so we can do bms_is_subset later */
+ for (i = 0; i < numattrs; i++)
+ mvattnums = bms_add_member(mvattnums, attrs->values[i]);
+
+ /* erase the list of mv-compatible clauses */
+ *mvclauses = NIL;
+
+ foreach(l, clauses)
+ {
+ bool match = false; /* by default not mv-compatible */
+ AttrNumber attnum = InvalidAttrNumber;
+ Node *clause = (Node *) lfirst(l);
+
+ if (clause_is_mv_compatible(clause, relid, &attnum))
+ {
+ /* are all the attributes part of the selected stats? */
+ if (bms_is_member(attnum, mvattnums))
+ match = true;
+ }
+
+ /*
+ * The clause matches the selected stats, so put it to the list of
+ * mv-compatible clauses. Otherwise, keep it in the list of 'regular'
+ * clauses (that may be selected later).
+ */
+ if (match)
+ *mvclauses = lappend(*mvclauses, clause);
+ else
+ non_mvclauses = lappend(non_mvclauses, clause);
+ }
+
+ /*
+ * Perform regular estimation using the clauses incompatible with the
+ * chosen histogram (or MV stats in general).
+ */
+ return non_mvclauses;
+
+}
+
+typedef struct
+{
+ Index varno; /* relid we're interested in */
+ Bitmapset *varattnos; /* attnums referenced by the clauses */
+} mv_compatible_context;
+
+/*
+ * Recursive walker that checks compatibility of the clause with multivariate
+ * statistics, and collects attnums from the Vars.
+ *
+ * XXX The original idea was to combine this with expression_tree_walker, but
+ * I've been unable to make that work - seems that does not quite allow
+ * checking the structure. Hence the explicit calls to the walker.
+ */
+static bool
+mv_compatible_walker(Node *node, mv_compatible_context *context)
+{
+ if (node == NULL)
+ return false;
+
+ if (IsA(node, RestrictInfo))
+ {
+ RestrictInfo *rinfo = (RestrictInfo *) node;
+
+ /* Pseudoconstants are not really interesting here. */
+ if (rinfo->pseudoconstant)
+ return true;
+
+ /* clauses referencing multiple varnos are incompatible */
+ if (bms_membership(rinfo->clause_relids) != BMS_SINGLETON)
+ return true;
+
+ /* check the clause inside the RestrictInfo */
+ return mv_compatible_walker((Node *) rinfo->clause, (void *) context);
+ }
+
+ if (IsA(node, Var))
+ {
+ Var *var = (Var *) node;
+
+ /*
+ * Also, the variable needs to reference the right relid (this might
+ * be unnecessary given the other checks, but let's be sure).
+ */
+ if (var->varno != context->varno)
+ return true;
+
+ /* Also skip system attributes (we don't allow stats on those). */
+ if (!AttrNumberIsForUserDefinedAttr(var->varattno))
+ return true;
+
+ /* Seems fine, so let's remember the attnum. */
+ context->varattnos = bms_add_member(context->varattnos, var->varattno);
+
+ return false;
+ }
+
+ /*
+ * And finally the operator expressions - we only allow simple expressions
+ * with two arguments, where one is a Var and the other is a constant, and
+ * it's a simple comparison (which we detect using estimator function).
+ */
+ if (is_opclause(node))
+ {
+ OpExpr *expr = (OpExpr *) node;
+ Var *var;
+ bool varonleft = true;
+ bool ok;
+
+ /*
+ * Only expressions with two arguments are considered compatible.
+ *
+ * XXX Possibly unnecessary (can OpExpr have different arg count?).
+ */
+ if (list_length(expr->args) != 2)
+ return true;
+
+ /* see if it actually has the right */
+ ok = (NumRelids((Node *) expr) == 1) &&
+ (is_pseudo_constant_clause(lsecond(expr->args)) ||
+ (varonleft = false,
+ is_pseudo_constant_clause(linitial(expr->args))));
+
+ /* unsupported structure (two variables or so) */
+ if (!ok)
+ return true;
+
+ /*
+ * If it's not a "<" or ">" or "=" operator, just ignore the clause.
+ * Otherwise note the relid and attnum for the variable. This uses the
+ * function for estimating selectivity, ont the operator directly (a
+ * bit awkward, but well ...).
+ */
+ switch (get_oprrest(expr->opno))
+ {
+ case F_EQSEL:
+
+ /* equality conditions are compatible with all statistics */
+ break;
+
+ default:
+
+ /* unknown estimator */
+ return true;
+ }
+
+ var = (varonleft) ? linitial(expr->args) : lsecond(expr->args);
+
+ return mv_compatible_walker((Node *) var, context);
+ }
+
+ /* Node not explicitly supported, so terminate */
+ return true;
+}
+
+/*
+ * Determines whether the clause is compatible with multivariate stats,
+ * and if it is, returns some additional information - varno (index
+ * into simple_rte_array) and a bitmap of attributes. This is then
+ * used to fetch related multivariate statistics.
+ *
+ * At this moment we only support basic conditions of the form
+ *
+ * variable OP constant
+ *
+ * where OP is one of [=,<,<=,>=,>] (which is however determined by
+ * looking at the associated function for estimating selectivity, just
+ * like with the single-dimensional case).
+ *
+ * TODO: Support 'OR clauses' - shouldn't be all that difficult to
+ * evaluate them using multivariate stats.
+ */
+static bool
+clause_is_mv_compatible(Node *clause, Index relid, AttrNumber *attnum)
+{
+ mv_compatible_context context;
+
+ context.varno = relid;
+ context.varattnos = NULL; /* no attnums */
+
+ if (mv_compatible_walker(clause, (void *) &context))
+ return false;
+
+ /* remember the newly collected attnums */
+ *attnum = bms_singleton_member(context.varattnos);
+
+ return true;
+}
+
+
+/*
+ * Check that the statistics matches at least one of the requested types.
+ */
+static bool
+stats_type_matches(MVStatisticInfo *stat, int type)
+{
+ if ((type & STATS_TYPE_FDEPS) && stat->deps_built)
+ return true;
+
+ return false;
+}
+
+/*
+ * Check that there are stats with at least one of the requested types.
+ */
+static bool
+has_stats(List *stats, int type)
+{
+ ListCell *s;
+
+ foreach(s, stats)
+ {
+ MVStatisticInfo *stat = (MVStatisticInfo *) lfirst(s);
+
+ /* terminate if we've found at least one matching statistics */
+ if (stats_type_matches(stat, type))
+ return true;
+ }
+
+ return false;
+}
+
+/*
+ * Lookups stats for the given baserel.
+ */
+static List *
+find_stats(PlannerInfo *root, Index relid)
+{
+ Assert(root->simple_rel_array[relid] != NULL);
+
+ return root->simple_rel_array[relid]->mvstatlist;
+}
diff --git a/src/backend/utils/mvstats/README.stats b/src/backend/utils/mvstats/README.stats
index 865a75d..251a468 100644
--- a/src/backend/utils/mvstats/README.stats
+++ b/src/backend/utils/mvstats/README.stats
@@ -8,48 +8,9 @@ not true, resulting in estimation errors.
Multivariate stats track different types of dependencies between the columns,
hopefully improving the estimates.
-
-Types of statistics
--------------------
-
-Currently we only have two kinds of multivariate statistics
-
- (a) soft functional dependencies (README.dependencies)
-
- (b) ndistinct coefficients
-
-
-Compatible clause types
------------------------
-
-Each type of statistics may be used to estimate some subset of clause types.
-
- (a) functional dependencies - equality clauses (AND), possibly IS NULL
-
-Currently only simple operator clauses (Var op Const) are supported, but it's
-possible to support more complex clause types, e.g. (Var op Var).
-
-
-Complex clauses
----------------
-
-We also support estimating more complex clauses - essentially AND/OR clauses
-with (Var op Const) as leaves, as long as all the referenced attributes are
-covered by a single statistics.
-
-For example this condition
-
- (a=1) AND ((b=2) OR ((c=3) AND (d=4)))
-
-may be estimated using statistics on (a,b,c,d). If we only have statistics on
-(b,c,d) we may estimate the second part, and estimate (a=1) using simple stats.
-
-If we only have statistics on (a,b,c) we can't apply it at all at this point,
-but it's worth pointing out clauselist_selectivity() works recursively and when
-handling the second part (the OR-clause), we'll be able to apply the statistics.
-
-Note: The multi-statistics estimation patch also makes it possible to pass some
-clauses as 'conditions' into the deeper parts of the expression tree.
+Currently we only have one kind of multivariate statistics - soft functional
+dependencies, and we use it to improve estimates of equality clauses. See
+README.dependencies for details.
Selectivity estimation
diff --git a/src/backend/utils/mvstats/common.c b/src/backend/utils/mvstats/common.c
index e1bed39..5f3fb55 100644
--- a/src/backend/utils/mvstats/common.c
+++ b/src/backend/utils/mvstats/common.c
@@ -300,6 +300,7 @@ compare_scalars_partition(const void *a, const void *b, void *arg)
return ApplySortComparator(da, false, db, false, ssup);
}
+
/* initialize multi-dimensional sort */
MultiSortSupport
multi_sort_init(int ndims)
diff --git a/src/backend/utils/mvstats/dependencies.c b/src/backend/utils/mvstats/dependencies.c
index 03865ed..de58302 100644
--- a/src/backend/utils/mvstats/dependencies.c
+++ b/src/backend/utils/mvstats/dependencies.c
@@ -320,6 +320,10 @@ dependency_degree(int numrows, HeapTuple *rows, int k, int *dependency,
* (c) -> b (c,a) -> b
* (c) -> a (c,b) -> a
* (b) -> a (b,c) -> a
+ *
+ * XXX Currently this builds redundant dependencies, becuse (a,b => c) and
+ * (b,a => c) is exactly the same thing, but both versions are generated
+ * and stored in the statistics.
*/
MVDependencies
build_mv_dependencies(int numrows, HeapTuple *rows, int2vector *attrs,
@@ -533,6 +537,70 @@ deserialize_mv_dependencies(bytea *data)
}
/*
+ * dependency_is_fully_matched
+ * checks that a functional dependency is fully matched given clauses on
+ * attributes (assuming the clauses are suitable equality clauses)
+ */
+bool
+dependency_is_fully_matched(MVDependency dependency, Bitmapset *attnums,
+ int16 *attmap)
+{
+ int j;
+
+ /*
+ * Check that the dependency actually is fully covered by clauses. We
+ * have to translate all attribute numbers, as those are referenced
+ */
+ for (j = 0; j < dependency->nattributes; j++)
+ {
+ int attnum = attmap[dependency->attributes[j]];
+
+ if (! bms_is_member(attnum, attnums))
+ return false;
+ }
+
+ return true;
+}
+
+/*
+ * dependency_implies_attribute
+ * check that the attnum matches is implied by the functional dependency
+ */
+bool
+dependency_implies_attribute(MVDependency dependency, AttrNumber attnum,
+ int16 *attmap)
+{
+ if (attnum == attmap[dependency->attributes[dependency->nattributes-1]])
+ return true;
+
+ return false;
+}
+
+MVDependencies
+load_mv_dependencies(Oid mvoid)
+{
+ bool isnull = false;
+ Datum deps;
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ HeapTuple htup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(mvoid));
+
+#ifdef USE_ASSERT_CHECKING
+ Form_pg_mv_statistic mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
+ Assert(mvstat->deps_enabled && mvstat->deps_built);
+#endif
+
+ deps = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_stadeps, &isnull);
+
+ Assert(!isnull);
+
+ ReleaseSysCache(htup);
+
+ return deserialize_mv_dependencies(DatumGetByteaP(deps));
+}
+
+/*
* pg_dependencies_in - input routine for type pg_dependencies.
*
* pg_dependencies is real enough to be a table column, but it has no operations
diff --git a/src/include/utils/mvstats.h b/src/include/utils/mvstats.h
index 227b9ff..3ad4e48 100644
--- a/src/include/utils/mvstats.h
+++ b/src/include/utils/mvstats.h
@@ -65,9 +65,13 @@ typedef struct MVDependenciesData
typedef MVDependenciesData *MVDependencies;
-
+bool dependency_implies_attribute(MVDependency dependency, AttrNumber attnum,
+ int16 *attmap);
+bool dependency_is_fully_matched(MVDependency dependency, Bitmapset *attnums,
+ int16 *attmap);
MVNDistinct load_mv_ndistinct(Oid mvoid);
+MVDependencies load_mv_dependencies(Oid mvoid);
bytea *serialize_mv_ndistinct(MVNDistinct ndistinct);
bytea *serialize_mv_dependencies(MVDependencies dependencies);
diff --git a/src/test/regress/expected/mv_dependencies.out b/src/test/regress/expected/mv_dependencies.out
index d442a16..cf57a67 100644
--- a/src/test/regress/expected/mv_dependencies.out
+++ b/src/test/regress/expected/mv_dependencies.out
@@ -55,8 +55,10 @@ SELECT deps_enabled, deps_built, stadeps
TRUNCATE functional_dependencies;
-- a => b, a => c, b => c
+-- check explain (expect bitmap index scan, not plain index scan)
INSERT INTO functional_dependencies
- SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ SELECT mod(i,400), mod(i,200), mod(i,100) FROM generate_series(1,30000) s(i);
+CREATE INDEX fdeps_idx ON functional_dependencies (a, b);
ANALYZE functional_dependencies;
SELECT deps_enabled, deps_built, stadeps
FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
@@ -65,6 +67,16 @@ SELECT deps_enabled, deps_built, stadeps
t | t | [{0 => 1 : 1.000000}, {0 => 2 : 1.000000}, {1 => 2 : 1.000000}, {0, 1 => 2 : 1.000000}, {0, 2 => 1 : 1.000000}]
(1 row)
+EXPLAIN (COSTS off)
+ SELECT * FROM functional_dependencies WHERE a = 10 AND b = 5;
+ QUERY PLAN
+---------------------------------------------
+ Bitmap Heap Scan on functional_dependencies
+ Recheck Cond: ((a = 10) AND (b = 5))
+ -> Bitmap Index Scan on fdeps_idx
+ Index Cond: ((a = 10) AND (b = 5))
+(4 rows)
+
DROP TABLE functional_dependencies;
-- varlena type (text)
CREATE TABLE functional_dependencies (
@@ -110,8 +122,10 @@ SELECT deps_enabled, deps_built, stadeps
TRUNCATE functional_dependencies;
-- a => b, a => c, b => c
+-- check explain (expect bitmap index scan, not plain index scan)
INSERT INTO functional_dependencies
- SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ SELECT mod(i,400), mod(i,200), mod(i,100) FROM generate_series(1,30000) s(i);
+CREATE INDEX fdeps_idx ON functional_dependencies (a, b);
ANALYZE functional_dependencies;
SELECT deps_enabled, deps_built, stadeps
FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
@@ -120,6 +134,16 @@ SELECT deps_enabled, deps_built, stadeps
t | t | [{0 => 1 : 1.000000}, {0 => 2 : 1.000000}, {1 => 2 : 1.000000}, {0, 1 => 2 : 1.000000}, {0, 2 => 1 : 1.000000}]
(1 row)
+EXPLAIN (COSTS off)
+ SELECT * FROM functional_dependencies WHERE a = '10' AND b = '5';
+ QUERY PLAN
+------------------------------------------------------------
+ Bitmap Heap Scan on functional_dependencies
+ Recheck Cond: ((a = '10'::text) AND (b = '5'::text))
+ -> Bitmap Index Scan on fdeps_idx
+ Index Cond: ((a = '10'::text) AND (b = '5'::text))
+(4 rows)
+
DROP TABLE functional_dependencies;
-- NULL values (mix of int and text columns)
CREATE TABLE functional_dependencies (
diff --git a/src/test/regress/sql/mv_dependencies.sql b/src/test/regress/sql/mv_dependencies.sql
index 43df798..49db649 100644
--- a/src/test/regress/sql/mv_dependencies.sql
+++ b/src/test/regress/sql/mv_dependencies.sql
@@ -53,13 +53,20 @@ SELECT deps_enabled, deps_built, stadeps
TRUNCATE functional_dependencies;
-- a => b, a => c, b => c
+-- check explain (expect bitmap index scan, not plain index scan)
INSERT INTO functional_dependencies
- SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ SELECT mod(i,400), mod(i,200), mod(i,100) FROM generate_series(1,30000) s(i);
+
+CREATE INDEX fdeps_idx ON functional_dependencies (a, b);
+
ANALYZE functional_dependencies;
SELECT deps_enabled, deps_built, stadeps
FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+EXPLAIN (COSTS off)
+ SELECT * FROM functional_dependencies WHERE a = 10 AND b = 5;
+
DROP TABLE functional_dependencies;
-- varlena type (text)
@@ -96,6 +103,7 @@ TRUNCATE functional_dependencies;
-- a => b, a => c
INSERT INTO functional_dependencies
SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+
ANALYZE functional_dependencies;
SELECT deps_enabled, deps_built, stadeps
@@ -104,13 +112,20 @@ SELECT deps_enabled, deps_built, stadeps
TRUNCATE functional_dependencies;
-- a => b, a => c, b => c
+-- check explain (expect bitmap index scan, not plain index scan)
INSERT INTO functional_dependencies
- SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ SELECT mod(i,400), mod(i,200), mod(i,100) FROM generate_series(1,30000) s(i);
+
+CREATE INDEX fdeps_idx ON functional_dependencies (a, b);
+
ANALYZE functional_dependencies;
SELECT deps_enabled, deps_built, stadeps
FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+EXPLAIN (COSTS off)
+ SELECT * FROM functional_dependencies WHERE a = '10' AND b = '5';
+
DROP TABLE functional_dependencies;
-- NULL values (mix of int and text columns)
--
2.5.5
[binary/octet-stream] 0005-PATCH-multivariate-MCV-lists-v22.patch (122.3K, ../../696aa95c-2411-9b2b-f36e-65b66bf47c88@2ndquadrant.com/6-0005-PATCH-multivariate-MCV-lists-v22.patch)
download | inline diff:
From 29eb7e89bf69361075230d7926553a035b0c841b Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:38:02 +0200
Subject: [PATCH 5/9] PATCH: multivariate MCV lists
- extends the pg_mv_statistic catalog (add 'mcv' fields)
- building the MCV lists during ANALYZE
- simple estimation while planning the queries
- pg_mcv_list data type (varlena-based)
Includes regression tests, mostly equal to regression tests for
functional dependencies.
A varlena-based data type for storing serialized MCV lists.
---
doc/src/sgml/catalogs.sgml | 30 +
doc/src/sgml/planstats.sgml | 157 ++++
doc/src/sgml/ref/create_statistics.sgml | 34 +
src/backend/catalog/system_views.sql | 4 +-
src/backend/commands/statscmds.c | 11 +-
src/backend/nodes/outfuncs.c | 2 +
src/backend/optimizer/path/clausesel.c | 636 +++++++++++++++-
src/backend/optimizer/util/plancat.c | 4 +-
src/backend/utils/mvstats/Makefile | 2 +-
src/backend/utils/mvstats/README.mcv | 137 ++++
src/backend/utils/mvstats/README.stats | 87 ++-
src/backend/utils/mvstats/common.c | 136 +++-
src/backend/utils/mvstats/common.h | 22 +-
src/backend/utils/mvstats/mcv.c | 1184 +++++++++++++++++++++++++++++
src/bin/psql/describe.c | 24 +-
src/include/catalog/pg_cast.h | 5 +
src/include/catalog/pg_mv_statistic.h | 18 +-
src/include/catalog/pg_proc.h | 14 +
src/include/catalog/pg_type.h | 4 +
src/include/nodes/relation.h | 6 +-
src/include/utils/builtins.h | 4 +
src/include/utils/mvstats.h | 67 +-
src/test/regress/expected/mv_mcv.out | 198 +++++
src/test/regress/expected/opr_sanity.out | 3 +-
src/test/regress/expected/rules.out | 4 +-
src/test/regress/expected/type_sanity.out | 3 +-
src/test/regress/parallel_schedule | 2 +-
src/test/regress/serial_schedule | 1 +
src/test/regress/sql/mv_mcv.sql | 169 ++++
29 files changed, 2898 insertions(+), 70 deletions(-)
create mode 100644 src/backend/utils/mvstats/README.mcv
create mode 100644 src/backend/utils/mvstats/mcv.c
create mode 100644 src/test/regress/expected/mv_mcv.out
create mode 100644 src/test/regress/sql/mv_mcv.sql
diff --git a/doc/src/sgml/catalogs.sgml b/doc/src/sgml/catalogs.sgml
index 716024b..b82ca13 100644
--- a/doc/src/sgml/catalogs.sgml
+++ b/doc/src/sgml/catalogs.sgml
@@ -4281,6 +4281,17 @@
</row>
<row>
+ <entry><structfield>mcv_enabled</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, MVC list will be computed for the combination of columns,
+ covered by the statistics. This does not mean the MCV list is already
+ computed, though.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>ndist_built</structfield></entry>
<entry><type>bool</type></entry>
<entry></entry>
@@ -4301,6 +4312,16 @@
</row>
<row>
+ <entry><structfield>mcv_built</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, MCV list is already computed and available for use during query
+ estimation.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>stakeys</structfield></entry>
<entry><type>int2vector</type></entry>
<entry><literal><link linkend="catalog-pg-attribute"><structname>pg_attribute</structname></link>.attnum</literal></entry>
@@ -4329,6 +4350,15 @@
</entry>
</row>
+ <row>
+ <entry><structfield>stamcv</structfield></entry>
+ <entry><type>pg_mcv_list</type></entry>
+ <entry></entry>
+ <entry>
+ MCV list, serialized as <structname>pg_mcv_list</> type.
+ </entry>
+ </row>
+
</tbody>
</tgroup>
</table>
diff --git a/doc/src/sgml/planstats.sgml b/doc/src/sgml/planstats.sgml
index bdad2db..1ee4293 100644
--- a/doc/src/sgml/planstats.sgml
+++ b/doc/src/sgml/planstats.sgml
@@ -757,6 +757,163 @@ EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 10;
</sect2>
+ <sect2 id="mcv-lists">
+ <title>MCV lists</title>
+
+ <para>
+ As explained in the previous section, functional dependencies are very
+ cheap and efficient type of statistics, but it has limitations due to the
+ global nature (only tracking column-level dependencies, not between values
+ stored in the columns).
+ </para>
+
+ <para>
+ This section introduces multivariate most-common values (<acronym>MCV</>)
+ lists, a direct generalization of the statistics introduced in
+ <xref linkend="row-estimation-examples">, that is not subject to this
+ limitation. It is however more expensive, both in terms of storage and
+ planning time.
+ </para>
+
+ <para>
+ Let's look at the example query from the previous section again, creating
+ a multivariate <acronym>MCV</> list on the columns (after dropping the
+ functional dependencies, to make sure the planner uses the newly created
+ <acronym>MCV</> list when computing the estimates).
+
+<programlisting>
+DROP STATISTICS s1;
+CREATE STATISTICS s2 ON t (a,b) WITH (mcv);
+ANALYZE t;
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
+ QUERY PLAN
+-------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=100 width=8) (actual time=0.036..3.011 rows=100 loops=1)
+ Filter: ((a = 1) AND (b = 1))
+ Rows Removed by Filter: 9900
+ Planning time: 0.188 ms
+ Execution time: 3.229 ms
+(5 rows)
+</programlisting>
+
+ The estimate is as accurate as with the functional dependencies, mostly
+ thanks to the table being a fairly small and having a simple distribution
+ with low number of distinct values. Before looking at the second query,
+ which was not handled by functional dependencies this well, let's inspect
+ the <acronym>MCV</> list a bit.
+ </para>
+
+ <para>
+ First, let's list statistics defined on a table using <command>\d</>
+ in <application>psql</>:
+
+<programlisting>
+\d t
+ Table "public.t"
+ Column | Type | Modifiers
+--------+---------+-----------
+ a | integer |
+ b | integer |
+Statistics:
+ "public.s2" (mcv) ON (a, b)
+</programlisting>
+
+ </para>
+
+ <para>
+ To inspect details of the <acronym>MCV</> statistics, we can look into the
+ <structname>pg_mv_stats</structname> view
+
+<programlisting>
+SELECT tablename, staname, attnums, mcvbytes, mcvinfo
+ FROM pg_mv_stats WHERE staname = 's2';
+ tablename | staname | attnums | mcvbytes | mcvinfo
+-----------+---------+---------+----------+------------
+ t | s2 | 1 2 | 2048 | nitems=100
+(1 row)
+</programlisting>
+
+ According to this, the statistics has 2kB when serialized into
+ a <literal>bytea</> value, and <command>ANALYZE</> found 100 distinct
+ combinations of values in the two columns.
+ </para>
+
+ <para>
+ Inspecting the contents of the MCV list is possible using
+ <function>pg_mv_mcv_items</> function.
+
+<programlisting>
+SELECT * FROM pg_mv_mcv_items((SELECT oid FROM pg_mv_statistic WHERE staname = 's2'));
+ index | values | nulls | frequency
+-------+---------+-------+-----------
+ 0 | {0,0} | {f,f} | 0.01
+ 1 | {1,1} | {f,f} | 0.01
+ 2 | {2,2} | {f,f} | 0.01
+...
+ 49 | {49,49} | {f,f} | 0.01
+ 50 | {50,0} | {f,f} | 0.01
+...
+ 97 | {97,47} | {f,f} | 0.01
+ 98 | {98,48} | {f,f} | 0.01
+ 99 | {99,49} | {f,f} | 0.01
+(100 rows)
+</programlisting>
+
+ Which confirms there are 100 distinct combinations of values in the two
+ columns, and all of them are equally likely (1% frequency for each).
+ Had there been any null values in either of the columns, this would be
+ identified in the <structfield>nulls</> column.
+ </para>
+
+ <para>
+ When estimating the selectivity, the planner applies all the conditions
+ on items in the <acronym>MCV</> list, and them sums the frequencies
+ of the matching ones. See <function>clauselist_mv_selectivity_mcvlist</>
+ in <filename>clausesel.c</> for details.
+ </para>
+
+ <para>
+ Compared to functional dependencies, <acronym>MCV</> lists have two major
+ advantages. Firstly, the list stores actual values, making it possible to
+ detect "incompatible" combinations.
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 10;
+ QUERY PLAN
+---------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=1 width=8) (actual time=2.823..2.823 rows=0 loops=1)
+ Filter: ((a = 1) AND (b = 10))
+ Rows Removed by Filter: 10000
+ Planning time: 0.268 ms
+ Execution time: 2.866 ms
+(5 rows)
+</programlisting>
+
+ Secondly, <acronym>MCV</> also handle a wide range of clause types, not
+ just equality clauses like functional dependencies. See for example the
+ example range query, presented earlier:
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a <= 49 AND b > 49;
+ QUERY PLAN
+---------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=1 width=8) (actual time=3.349..3.349 rows=0 loops=1)
+ Filter: ((a <= 49) AND (b > 49))
+ Rows Removed by Filter: 10000
+ Planning time: 0.163 ms
+ Execution time: 3.389 ms
+(5 rows)
+</programlisting>
+
+ </para>
+
+ <para>
+ For additional information about multivariate MCV lists, see
+ <filename>src/backend/utils/mvstats/README.mcv</>.
+ </para>
+
+ </sect2>
+
</sect1>
</chapter>
diff --git a/doc/src/sgml/ref/create_statistics.sgml b/doc/src/sgml/ref/create_statistics.sgml
index 42adc38..fc97b16 100644
--- a/doc/src/sgml/ref/create_statistics.sgml
+++ b/doc/src/sgml/ref/create_statistics.sgml
@@ -125,6 +125,15 @@ CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_na
</varlistentry>
<varlistentry>
+ <term><literal>mcv</> (<type>boolean</>)</term>
+ <listitem>
+ <para>
+ Enables MCV list for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
<term><literal>ndistinct</> (<type>boolean</>)</term>
<listitem>
<para>
@@ -168,6 +177,31 @@ EXPLAIN ANALYZE SELECT * FROM t1 WHERE (a = 1) AND (b = 1);
</programlisting>
</para>
+ <para>
+ Create table <structname>t2</> with two perfectly correlated columns
+ (containing identical data), and a MCV list on those columns:
+
+<programlisting>
+CREATE TABLE t2 (
+ a int,
+ b int
+);
+
+INSERT INTO t2 SELECT mod(i,100), mod(i,100)
+ FROM generate_series(1,1000000) s(i);
+
+CREATE STATISTICS s2 WITH (mcv) ON (a, b) FROM t2;
+
+ANALYZE t2;
+
+-- valid combination (found in MCV)
+EXPLAIN ANALYZE SELECT * FROM t2 WHERE (a = 1) AND (b = 1);
+
+-- invalid combination (not found in MCV)
+EXPLAIN ANALYZE SELECT * FROM t2 WHERE (a = 1) AND (b = 2);
+</programlisting>
+ </para>
+
</refsect1>
<refsect1>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index fc1b1a9..7ac3c4c 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -188,7 +188,9 @@ CREATE VIEW pg_mv_stats AS
S.staname AS staname,
S.stakeys AS attnums,
length(s.standist::bytea) AS ndistbytes,
- length(S.stadeps::bytea) AS depsbytes
+ length(S.stadeps::bytea) AS depsbytes,
+ length(S.stamcv::bytea) AS mcvbytes,
+ pg_mv_stats_mcvlist_info(S.stamcv) AS mcvinfo
FROM (pg_mv_statistic S JOIN pg_class C ON (C.oid = S.starelid))
LEFT JOIN pg_namespace N ON (N.oid = C.relnamespace);
diff --git a/src/backend/commands/statscmds.c b/src/backend/commands/statscmds.c
index ec41bbc..e428c69 100644
--- a/src/backend/commands/statscmds.c
+++ b/src/backend/commands/statscmds.c
@@ -76,7 +76,8 @@ CreateStatistics(CreateStatsStmt *stmt)
/* by default build nothing */
bool build_ndistinct = false,
- build_dependencies = false;
+ build_dependencies = false,
+ build_mcv = false;
Assert(IsA(stmt, CreateStatsStmt));
@@ -169,6 +170,8 @@ CreateStatistics(CreateStatsStmt *stmt)
build_ndistinct = defGetBoolean(opt);
else if (strcmp(opt->defname, "dependencies") == 0)
build_dependencies = defGetBoolean(opt);
+ else if (strcmp(opt->defname, "mcv") == 0)
+ build_mcv = defGetBoolean(opt);
else
ereport(ERROR,
(errcode(ERRCODE_SYNTAX_ERROR),
@@ -177,10 +180,10 @@ CreateStatistics(CreateStatsStmt *stmt)
}
/* Make sure there's at least one statistics type specified. */
- if (! (build_ndistinct || build_dependencies))
+ if (!(build_ndistinct || build_dependencies || build_mcv))
ereport(ERROR,
(errcode(ERRCODE_SYNTAX_ERROR),
- errmsg("no statistics type (ndistinct, dependencies) requested")));
+ errmsg("no statistics type (ndistinct, dependencies, mcv) requested")));
stakeys = buildint2vector(attnums, numcols);
@@ -203,9 +206,11 @@ CreateStatistics(CreateStatsStmt *stmt)
/* enabled statistics */
values[Anum_pg_mv_statistic_ndist_enabled - 1] = BoolGetDatum(build_ndistinct);
values[Anum_pg_mv_statistic_deps_enabled - 1] = BoolGetDatum(build_dependencies);
+ values[Anum_pg_mv_statistic_mcv_enabled - 1] = BoolGetDatum(build_mcv);
nulls[Anum_pg_mv_statistic_standist - 1] = true;
nulls[Anum_pg_mv_statistic_stadeps - 1] = true;
+ nulls[Anum_pg_mv_statistic_stamcv - 1] = true;
/* insert the tuple into pg_mv_statistic */
mvstatrel = heap_open(MvStatisticRelationId, RowExclusiveLock);
diff --git a/src/backend/nodes/outfuncs.c b/src/backend/nodes/outfuncs.c
index 26e504f..ea3db02 100644
--- a/src/backend/nodes/outfuncs.c
+++ b/src/backend/nodes/outfuncs.c
@@ -2181,10 +2181,12 @@ _outMVStatisticInfo(StringInfo str, const MVStatisticInfo *node)
/* enabled statistics */
WRITE_BOOL_FIELD(ndist_enabled);
WRITE_BOOL_FIELD(deps_enabled);
+ WRITE_BOOL_FIELD(mcv_enabled);
/* built/available statistics */
WRITE_BOOL_FIELD(ndist_built);
WRITE_BOOL_FIELD(deps_built);
+ WRITE_BOOL_FIELD(mcv_built);
}
static void
diff --git a/src/backend/optimizer/path/clausesel.c b/src/backend/optimizer/path/clausesel.c
index ec74b16..c422dd5 100644
--- a/src/backend/optimizer/path/clausesel.c
+++ b/src/backend/optimizer/path/clausesel.c
@@ -15,6 +15,7 @@
#include "postgres.h"
#include "access/sysattr.h"
+#include "catalog/pg_collation.h"
#include "catalog/pg_operator.h"
#include "nodes/makefuncs.h"
#include "optimizer/clauses.h"
@@ -47,12 +48,14 @@ static void addRangeClause(RangeQueryClause **rqlist, Node *clause,
bool varonleft, bool isLTsel, Selectivity s2);
#define STATS_TYPE_FDEPS 0x01
+#define STATS_TYPE_MCV 0x02
-static bool clause_is_mv_compatible(Node *clause, Index relid, AttrNumber *attnum);
+static bool clause_is_mv_compatible(Node *clause, Index relid, Bitmapset **attnums,
+ int type);
-static Bitmapset *collect_mv_attnums(List *clauses, Index relid);
+static Bitmapset *collect_mv_attnums(List *clauses, Index relid, int type);
-static int count_mv_attnums(List *clauses, Index relid);
+static int count_mv_attnums(List *clauses, Index relid, int type);
static int count_varnos(List *clauses, Index *relid);
@@ -63,10 +66,23 @@ static List *clauselist_mv_split(PlannerInfo *root, Index relid,
List *clauses, List **mvclauses,
MVStatisticInfo *mvstats, int types);
+static Selectivity clauselist_mv_selectivity(PlannerInfo *root,
+ List *clauses, MVStatisticInfo *mvstats);
+
static Selectivity clauselist_mv_selectivity_deps(PlannerInfo *root,
Index relid, List *clauses, MVStatisticInfo *mvstats,
Index varRelid, JoinType jointype, SpecialJoinInfo *sjinfo);
+static Selectivity clauselist_mv_selectivity_mcvlist(PlannerInfo *root,
+ List *clauses, MVStatisticInfo *mvstats,
+ bool *fullmatch, Selectivity *lowsel);
+
+static int update_match_bitmap_mcvlist(PlannerInfo *root, List *clauses,
+ int2vector *stakeys, MCVList mcvlist,
+ int nmatches, char *matches,
+ Selectivity *lowsel, bool *fullmatch,
+ bool is_or);
+
static bool has_stats(List *stats, int type);
static List *find_stats(PlannerInfo *root, Index relid);
@@ -74,6 +90,9 @@ static List *find_stats(PlannerInfo *root, Index relid);
static bool stats_type_matches(MVStatisticInfo *stat, int type);
+#define UPDATE_RESULT(m,r,isor) \
+ (m) = (isor) ? (Max(m,r)) : (Min(m,r))
+
/****************************************************************************
* ROUTINES TO COMPUTE SELECTIVITIES
****************************************************************************/
@@ -99,11 +118,13 @@ static bool stats_type_matches(MVStatisticInfo *stat, int type);
* to verify that suitable multivariate statistics exist.
*
* If we identify such multivariate statistics apply, we try to apply them.
- * Currently we only have (soft) functional dependencies, so we try to reduce
- * the list of clauses.
*
- * Then we remove the clauses estimated using multivariate stats, and process
- * the rest of the clauses using the regular per-column stats.
+ * First we try to reduce the list of clauses by applying (soft) functional
+ * dependencies, and then we try to estimate the selectivity of the reduced
+ * list of clauses using the multivariate MCV list.
+ *
+ * Finally we remove the portion of clauses estimated using multivariate stats,
+ * and process the rest of the clauses using the regular per-column stats.
*
* Currently, the only extra smarts we have is to recognize "range queries",
* such as "x > 34 AND x < 42". Clauses are recognized as possible range
@@ -173,7 +194,10 @@ clauselist_selectivity(PlannerInfo *root,
/*
* Check that there are multivariate statistics usable for selectivity
- * estimation, i.e. anything except ndistinct coefficients.
+ * estimation. We try to apply MCV lists first, because statistics
+ * tracking actual values tend to provide more reliable estimates than
+ * functional dependencies (which assume that the clauses are consistent
+ * with the statistics).
*
* Also check the number of attributes in clauses that might be estimated
* using those statistics, and that there are at least two such attributes.
@@ -184,14 +208,43 @@ clauselist_selectivity(PlannerInfo *root,
* If there are no such stats or not enough attributes, don't waste time
* simply skip to estimation using the plain per-column stats.
*/
+ if (has_stats(stats, STATS_TYPE_MCV) &&
+ (count_mv_attnums(clauses, relid, STATS_TYPE_MCV) >= 2))
+ {
+ /* collect attributes from the compatible conditions */
+ Bitmapset *mvattnums = collect_mv_attnums(clauses, relid,
+ STATS_TYPE_MCV);
+
+ /* and search for the statistic covering the most attributes */
+ MVStatisticInfo *mvstat = choose_mv_statistics(stats, mvattnums,
+ STATS_TYPE_MCV);
+
+ if (mvstat != NULL) /* we have a matching stats */
+ {
+ /* clauses compatible with multi-variate stats */
+ List *mvclauses = NIL;
+
+ /* split the clauselist into regular and mv-clauses */
+ clauses = clauselist_mv_split(root, relid, clauses, &mvclauses,
+ mvstat, STATS_TYPE_MCV);
+
+ /* we've chosen the histogram to match the clauses */
+ Assert(mvclauses != NIL);
+
+ /* compute the multivariate stats */
+ s1 *= clauselist_mv_selectivity(root, mvclauses, mvstat);
+ }
+ }
+
+ /* Now try to apply functional dependencies on the remaining clauses. */
if (has_stats(stats, STATS_TYPE_FDEPS) &&
- (count_mv_attnums(clauses, relid) >= 2))
+ (count_mv_attnums(clauses, relid, STATS_TYPE_FDEPS) >= 2))
{
MVStatisticInfo *mvstat;
Bitmapset *mvattnums;
/* collect attributes from the compatible conditions */
- mvattnums = collect_mv_attnums(clauses, relid);
+ mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_FDEPS);
/* and search for the statistic covering the most attributes */
mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_FDEPS);
@@ -994,7 +1047,7 @@ clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
/* clauses remaining after removing those on the "implied" attribute */
List *clauses_filtered = NIL;
- attnums = collect_mv_attnums(clauses, relid);
+ attnums = collect_mv_attnums(clauses, relid, STATS_TYPE_FDEPS);
/* no point in looking for dependencies with fewer than 2 attributes */
if (bms_num_members(attnums) < 2)
@@ -1017,7 +1070,7 @@ clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
*/
foreach(lc, clauses)
{
- AttrNumber attnum_clause = InvalidAttrNumber;
+ Bitmapset *attnums_clause = NULL;
Node *clause = (Node *) lfirst(lc);
/*
@@ -1026,17 +1079,20 @@ clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
* we should only see equality clauses compatible with functional
* dependencies, so just error out if we stumble upon something else.
*/
- if (! clause_is_mv_compatible(clause, relid, &attnum_clause))
+ if (! clause_is_mv_compatible(clause, relid, &attnums_clause,
+ STATS_TYPE_FDEPS))
elog(ERROR, "clause not compatible with functional dependencies");
- Assert(AttributeNumberIsValid(attnum_clause));
+ /* we also expect only simple equality clauses */
+ Assert(bms_num_members(attnums_clause) == 1);
/*
* If the clause is not on the implied attribute, add it to the list
* of filtered clauses (for the next round) and continue with the
* next one.
*/
- if (! dependency_implies_attribute(dependency, attnum_clause,
+ if (! dependency_implies_attribute(dependency,
+ bms_singleton_member(attnums_clause),
mvstats->stakeys->values))
{
clauses_filtered = lappend(clauses_filtered, clause);
@@ -1080,10 +1136,71 @@ clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
}
/*
+ * estimate selectivity of clauses using multivariate statistic
+ *
+ * Perform estimation of the clauses using a MCV list.
+ *
+ * This assumes all the clauses are compatible with the selected statistics
+ * (e.g. only reference columns covered by the statistics, use supported
+ * operator, etc.).
+ *
+ * TODO: We may support some additional conditions, most importantly those
+ * matching multiple columns (e.g. "a = b" or "a < b").
+ *
+ * TODO: Clamp the selectivity by min of the per-clause selectivities (i.e. the
+ * selectivity of the most restrictive clause), because that's the maximum
+ * we can ever get from ANDed list of clauses. This may probably prevent
+ * issues with hitting too many buckets and low precision histograms.
+ *
+ * TODO: We may remember the lowest frequency in the MCV list, and then later
+ * use it as a upper boundary for the selectivity (had there been a more
+ * frequent item, it'd be in the MCV list). This might improve cases with
+ * low-detail histograms.
+ *
+ * TODO: We may also derive some additional boundaries for the selectivity from
+ * the MCV list, because
+ *
+ * (a) if we have a "full equality condition" (one equality condition on
+ * each column of the statistic) and we found a match in the MCV list,
+ * then this is the final selectivity (and pretty accurate),
+ *
+ * (b) if we have a "full equality condition" and we haven't found a match
+ * in the MCV list, then the selectivity is below the lowest frequency
+ * found in the MCV list,
+ *
+ * TODO: When applying the clauses to the histogram/MCV list, we can do that
+ * from the most selective clauses first, because that'll eliminate the
+ * buckets/items sooner (so we'll be able to skip them without inspection,
+ * which is more expensive). But this requires really knowing the per-clause
+ * selectivities in advance, and that's not what we do now.
+ */
+static Selectivity
+clauselist_mv_selectivity(PlannerInfo *root, List *clauses, MVStatisticInfo *mvstats)
+{
+ bool fullmatch = false;
+
+ /*
+ * Lowest frequency in the MCV list (may be used as an upper bound for
+ * full equality conditions that did not match any MCV item).
+ */
+ Selectivity mcv_low = 0.0;
+
+ /*
+ * TODO: Evaluate simple 1D selectivities, use the smallest one as an
+ * upper bound, product as lower bound, and sort the clauses in ascending
+ * order by selectivity (to optimize the MCV/histogram evaluation).
+ */
+
+ /* Evaluate the MCV selectivity */
+ return clauselist_mv_selectivity_mcvlist(root, clauses, mvstats,
+ &fullmatch, &mcv_low);
+}
+
+/*
* Collect attributes from mv-compatible clauses.
*/
static Bitmapset *
-collect_mv_attnums(List *clauses, Index relid)
+collect_mv_attnums(List *clauses, Index relid, int types)
{
Bitmapset *attnums = NULL;
ListCell *l;
@@ -1099,12 +1216,10 @@ collect_mv_attnums(List *clauses, Index relid)
*/
foreach(l, clauses)
{
- AttrNumber attnum;
Node *clause = (Node *) lfirst(l);
- /* ignore the result for now - we only need the info */
- if (clause_is_mv_compatible(clause, relid, &attnum))
- attnums = bms_add_member(attnums, attnum);
+ /* ignore the result here - we only need the attnums */
+ clause_is_mv_compatible(clause, relid, &attnums, types);
}
/*
@@ -1125,10 +1240,10 @@ collect_mv_attnums(List *clauses, Index relid)
* Count the number of attributes in clauses compatible with multivariate stats.
*/
static int
-count_mv_attnums(List *clauses, Index relid)
+count_mv_attnums(List *clauses, Index relid, int type)
{
int c;
- Bitmapset *attnums = collect_mv_attnums(clauses, relid);
+ Bitmapset *attnums = collect_mv_attnums(clauses, relid, type);
c = bms_num_members(attnums);
@@ -1263,7 +1378,8 @@ choose_mv_statistics(List *stats, Bitmapset *attnums, int types)
int numattrs = info->stakeys->dim1;
/* skip statistics not matching any of the requested types */
- if (! (info->deps_built && (STATS_TYPE_FDEPS & types)))
+ if (! ((info->deps_built && (STATS_TYPE_FDEPS & types)) ||
+ (info->mcv_built && (STATS_TYPE_MCV & types))))
continue;
/* count columns covered by the statistics */
@@ -1317,13 +1433,13 @@ clauselist_mv_split(PlannerInfo *root, Index relid,
foreach(l, clauses)
{
bool match = false; /* by default not mv-compatible */
- AttrNumber attnum = InvalidAttrNumber;
+ Bitmapset *attnums = NULL;
Node *clause = (Node *) lfirst(l);
- if (clause_is_mv_compatible(clause, relid, &attnum))
+ if (clause_is_mv_compatible(clause, relid, &attnums, types))
{
/* are all the attributes part of the selected stats? */
- if (bms_is_member(attnum, mvattnums))
+ if (bms_is_subset(attnums, mvattnums))
match = true;
}
@@ -1348,6 +1464,7 @@ clauselist_mv_split(PlannerInfo *root, Index relid,
typedef struct
{
+ int types; /* types of statistics ? */
Index varno; /* relid we're interested in */
Bitmapset *varattnos; /* attnums referenced by the clauses */
} mv_compatible_context;
@@ -1382,6 +1499,49 @@ mv_compatible_walker(Node *node, mv_compatible_context *context)
return mv_compatible_walker((Node *) rinfo->clause, (void *) context);
}
+ if (or_clause(node) || and_clause(node) || not_clause(node))
+ {
+ /*
+ * AND/OR/NOT-clauses are supported if all sub-clauses are supported
+ *
+ * TODO: We might support mixed case, where some of the clauses are
+ * supported and some are not, and treat all supported subclauses as a
+ * single clause, compute it's selectivity using mv stats, and compute
+ * the total selectivity using the current algorithm.
+ *
+ * TODO: For RestrictInfo above an OR-clause, we might use the
+ * orclause with nested RestrictInfo - we won't have to call
+ * pull_varnos() for each clause, saving time.
+ *
+ * TODO: Perhaps this needs a bit more thought for functional
+ * dependencies? Those don't quite work for NOT cases.
+ */
+ BoolExpr *expr = (BoolExpr *) node;
+ ListCell *lc;
+
+ foreach(lc, expr->args)
+ {
+ if (mv_compatible_walker((Node *) lfirst(lc), context))
+ return true;
+ }
+
+ return false;
+ }
+
+ if (IsA(node, NullTest))
+ {
+ NullTest *nt = (NullTest *) node;
+
+ /*
+ * Only simple (Var IS NULL) expressions supported for now. Maybe we
+ * could use examine_variable to fix this?
+ */
+ if (!IsA(nt->arg, Var))
+ return true;
+
+ return mv_compatible_walker((Node *) (nt->arg), context);
+ }
+
if (IsA(node, Var))
{
Var *var = (Var *) node;
@@ -1442,10 +1602,18 @@ mv_compatible_walker(Node *node, mv_compatible_context *context)
switch (get_oprrest(expr->opno))
{
case F_EQSEL:
-
/* equality conditions are compatible with all statistics */
break;
+ case F_SCALARLTSEL:
+ case F_SCALARGTSEL:
+
+ /* not compatible with functional dependencies */
+ if (!(context->types & STATS_TYPE_MCV))
+ return true; /* terminate */
+
+ break;
+
default:
/* unknown estimator */
@@ -1479,10 +1647,11 @@ mv_compatible_walker(Node *node, mv_compatible_context *context)
* evaluate them using multivariate stats.
*/
static bool
-clause_is_mv_compatible(Node *clause, Index relid, AttrNumber *attnum)
+clause_is_mv_compatible(Node *clause, Index relid, Bitmapset **attnums, int types)
{
mv_compatible_context context;
+ context.types = types;
context.varno = relid;
context.varattnos = NULL; /* no attnums */
@@ -1490,7 +1659,7 @@ clause_is_mv_compatible(Node *clause, Index relid, AttrNumber *attnum)
return false;
/* remember the newly collected attnums */
- *attnum = bms_singleton_member(context.varattnos);
+ *attnums = bms_add_members(*attnums, context.varattnos);
return true;
}
@@ -1505,6 +1674,9 @@ stats_type_matches(MVStatisticInfo *stat, int type)
if ((type & STATS_TYPE_FDEPS) && stat->deps_built)
return true;
+ if ((type & STATS_TYPE_MCV) && stat->mcv_built)
+ return true;
+
return false;
}
@@ -1538,3 +1710,409 @@ find_stats(PlannerInfo *root, Index relid)
return root->simple_rel_array[relid]->mvstatlist;
}
+
+/*
+ * Estimate selectivity of clauses using a MCV list.
+ *
+ * If there's no MCV list for the stats, the function returns 0.0.
+ *
+ * While computing the estimate, the function checks whether all the
+ * columns were matched with an equality condition. If that's the case,
+ * we can skip processing the histogram, as there can be no rows in
+ * it with the same values - all the rows matching the condition are
+ * represented by the MCV item. This can only happen with equality
+ * on all the attributes.
+ *
+ * The algorithm works like this:
+ *
+ * 1) mark all items as 'match'
+ * 2) walk through all the clauses
+ * 3) for a particular clause, walk through all the items
+ * 4) skip items that are already 'no match'
+ * 5) check clause for items that still match
+ * 6) sum frequencies for items to get selectivity
+ *
+ * The function also returns the frequency of the least frequent item
+ * on the MCV list, which may be useful for clamping estimate from the
+ * histogram (all items not present in the MCV list are less frequent).
+ * This however seems useful only for cases with conditions on all
+ * attributes.
+ *
+ * TODO: This only handles AND-ed clauses, but it might work for OR-ed
+ * lists too - it just needs to reverse the logic a bit. I.e. start
+ * with 'no match' for all items, and mark the items as a match
+ * as the clauses are processed (and skip items that are 'match').
+ */
+static Selectivity
+clauselist_mv_selectivity_mcvlist(PlannerInfo *root, List *clauses,
+ MVStatisticInfo *mvstats, bool *fullmatch,
+ Selectivity *lowsel)
+{
+ int i;
+ Selectivity s = 0.0;
+ Selectivity u = 0.0;
+
+ MCVList mcvlist = NULL;
+ int nmatches = 0;
+
+ /* match/mismatch bitmap for each MCV item */
+ char *matches = NULL;
+
+ Assert(clauses != NIL);
+ Assert(list_length(clauses) >= 2);
+
+ /* there's no MCV list built yet */
+ if (!mvstats->mcv_built)
+ return 0.0;
+
+ mcvlist = load_mv_mcvlist(mvstats->mvoid);
+
+ Assert(mcvlist != NULL);
+ Assert(mcvlist->nitems > 0);
+
+ /* by default all the MCV items match the clauses fully */
+ matches = palloc0(sizeof(char) * mcvlist->nitems);
+ memset(matches, MVSTATS_MATCH_FULL, sizeof(char) * mcvlist->nitems);
+
+ /* number of matching MCV items */
+ nmatches = mcvlist->nitems;
+
+ nmatches = update_match_bitmap_mcvlist(root, clauses,
+ mvstats->stakeys, mcvlist,
+ nmatches, matches,
+ lowsel, fullmatch, false);
+
+ /* sum frequencies for all the matching MCV items */
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ /* used to 'scale' for MCV lists not covering all tuples */
+ u += mcvlist->items[i]->frequency;
+
+ if (matches[i] != MVSTATS_MATCH_NONE)
+ s += mcvlist->items[i]->frequency;
+ }
+
+ pfree(matches);
+ pfree(mcvlist);
+
+ return s * u;
+}
+
+/*
+ * Evaluate clauses using the MCV list, and update the match bitmap.
+ *
+ * The bitmap may be already partially set, so this is really a way to
+ * combine results of several clause lists - either when computing
+ * conditional probability P(A|B) or a combination of AND/OR clauses.
+ *
+ * TODO: This works with 'bitmap' where each bit is represented as a char,
+ * which is slightly wasteful. Instead, we could use a regular
+ * bitmap, reducing the size to ~1/8. Another thing is merging the
+ * bitmaps using & and |, which might be faster than min/max.
+ */
+static int
+update_match_bitmap_mcvlist(PlannerInfo *root, List *clauses,
+ int2vector *stakeys, MCVList mcvlist,
+ int nmatches, char *matches,
+ Selectivity *lowsel, bool *fullmatch,
+ bool is_or)
+{
+ int i;
+ ListCell *l;
+
+ Bitmapset *eqmatches = NULL; /* attributes with equality matches */
+
+ /* The bitmap may be partially built. */
+ Assert(nmatches >= 0);
+ Assert(nmatches <= mcvlist->nitems);
+ Assert(clauses != NIL);
+ Assert(list_length(clauses) >= 1);
+ Assert(mcvlist != NULL);
+ Assert(mcvlist->nitems > 0);
+
+ /* No possible matches (only works for AND-ded clauses) */
+ if (((nmatches == 0) && (!is_or)) ||
+ ((nmatches == mcvlist->nitems) && is_or))
+ return nmatches;
+
+ /*
+ * find the lowest frequency in the MCV list
+ *
+ * We need to do that here, because we do various tricks in the following
+ * code - skipping items already ruled out, etc.
+ *
+ * XXX A loop is necessary because the MCV list is not sorted by
+ * frequency.
+ */
+ *lowsel = 1.0;
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ MCVItem item = mcvlist->items[i];
+
+ if (item->frequency < *lowsel)
+ *lowsel = item->frequency;
+ }
+
+ /*
+ * Loop through the list of clauses, and for each of them evaluate all the
+ * MCV items not yet eliminated by the preceding clauses.
+ */
+ foreach(l, clauses)
+ {
+ Node *clause = (Node *) lfirst(l);
+
+ /* if it's a RestrictInfo, then extract the clause */
+ if (IsA(clause, RestrictInfo))
+ clause = (Node *) ((RestrictInfo *) clause)->clause;
+
+ /* if there are no remaining matches possible, we can stop */
+ if (((nmatches == 0) && (!is_or)) ||
+ ((nmatches == mcvlist->nitems) && is_or))
+ break;
+
+ /* it's either OpClause, or NullTest */
+ if (is_opclause(clause))
+ {
+ OpExpr *expr = (OpExpr *) clause;
+ bool varonleft = true;
+ bool ok;
+ FmgrInfo opproc;
+
+ /* get procedure computing operator selectivity */
+ RegProcedure oprrest = get_oprrest(expr->opno);
+
+ fmgr_info(get_opcode(expr->opno), &opproc);
+
+ ok = (NumRelids(clause) == 1) &&
+ (is_pseudo_constant_clause(lsecond(expr->args)) ||
+ (varonleft = false,
+ is_pseudo_constant_clause(linitial(expr->args))));
+
+ if (ok)
+ {
+
+ FmgrInfo gtproc;
+ Var *var = (varonleft) ? linitial(expr->args) : lsecond(expr->args);
+ Const *cst = (varonleft) ? lsecond(expr->args) : linitial(expr->args);
+ bool isgt = (!varonleft);
+
+ TypeCacheEntry *typecache
+ = lookup_type_cache(var->vartype, TYPECACHE_GT_OPR);
+
+ /* FIXME proper matching attribute to dimension */
+ int idx = mv_get_index(var->varattno, stakeys);
+
+ fmgr_info(get_opcode(typecache->gt_opr), >proc);
+
+ /*
+ * Walk through the MCV items and evaluate the current clause.
+ * We can skip items that were already ruled out, and
+ * terminate if there are no remaining MCV items that might
+ * possibly match.
+ */
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ bool mismatch = false;
+ MCVItem item = mcvlist->items[i];
+
+ /*
+ * If there are no more matches (AND) or no remaining
+ * unmatched items (OR), we can stop processing this
+ * clause.
+ */
+ if (((nmatches == 0) && (!is_or)) ||
+ ((nmatches == mcvlist->nitems) && is_or))
+ break;
+
+ /*
+ * For AND-lists, we can also mark NULL items as 'no
+ * match' (and then skip them). For OR-lists this is not
+ * possible.
+ */
+ if ((!is_or) && item->isnull[idx])
+ matches[i] = MVSTATS_MATCH_NONE;
+
+ /* skip MCV items that were already ruled out */
+ if ((!is_or) && (matches[i] == MVSTATS_MATCH_NONE))
+ continue;
+ else if (is_or && (matches[i] == MVSTATS_MATCH_FULL))
+ continue;
+
+ switch (oprrest)
+ {
+ case F_EQSEL:
+
+ /*
+ * We don't care about isgt in equality, because
+ * it does not matter whether it's (var = const)
+ * or (const = var).
+ */
+ mismatch = !DatumGetBool(FunctionCall2Coll(&opproc,
+ DEFAULT_COLLATION_OID,
+ cst->constvalue,
+ item->values[idx]));
+
+ if (!mismatch)
+ eqmatches = bms_add_member(eqmatches, idx);
+
+ break;
+
+ case F_SCALARLTSEL: /* column < constant */
+ case F_SCALARGTSEL: /* column > constant */
+
+ /*
+ * First check whether the constant is below the
+ * lower boundary (in that case we can skip the
+ * bucket, because there's no overlap).
+ */
+ if (isgt)
+ mismatch = !DatumGetBool(FunctionCall2Coll(&opproc,
+ DEFAULT_COLLATION_OID,
+ cst->constvalue,
+ item->values[idx]));
+ else
+ mismatch = !DatumGetBool(FunctionCall2Coll(&opproc,
+ DEFAULT_COLLATION_OID,
+ item->values[idx],
+ cst->constvalue));
+
+ break;
+ }
+
+ /*
+ * XXX The conditions on matches[i] are not needed, as we
+ * skip MCV items that can't become true/false, depending
+ * on the current flag. See beginning of the loop over MCV
+ * items.
+ */
+
+ if ((is_or) && (matches[i] == MVSTATS_MATCH_NONE) && (!mismatch))
+ {
+ /* OR - was MATCH_NONE, but will be MATCH_FULL */
+ matches[i] = MVSTATS_MATCH_FULL;
+ ++nmatches;
+ continue;
+ }
+ else if ((!is_or) && (matches[i] == MVSTATS_MATCH_FULL) && mismatch)
+ {
+ /* AND - was MATC_FULL, but will be MATCH_NONE */
+ matches[i] = MVSTATS_MATCH_NONE;
+ --nmatches;
+ continue;
+ }
+
+ }
+ }
+ }
+ else if (IsA(clause, NullTest))
+ {
+ NullTest *expr = (NullTest *) clause;
+ Var *var = (Var *) (expr->arg);
+
+ /* FIXME proper matching attribute to dimension */
+ int idx = mv_get_index(var->varattno, stakeys);
+
+ /*
+ * Walk through the MCV items and evaluate the current clause. We
+ * can skip items that were already ruled out, and terminate if
+ * there are no remaining MCV items that might possibly match.
+ */
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ MCVItem item = mcvlist->items[i];
+
+ /*
+ * if there are no more matches, we can stop processing this
+ * clause
+ */
+ if (nmatches == 0)
+ break;
+
+ /* skip MCV items that were already ruled out */
+ if (matches[i] == MVSTATS_MATCH_NONE)
+ continue;
+
+ /* if the clause mismatches the MCV item, set it as MATCH_NONE */
+ if (((expr->nulltesttype == IS_NULL) && (!item->isnull[idx])) ||
+ ((expr->nulltesttype == IS_NOT_NULL) && (item->isnull[idx])))
+ {
+ matches[i] = MVSTATS_MATCH_NONE;
+ --nmatches;
+ }
+ }
+ }
+ else if (or_clause(clause) || and_clause(clause))
+ {
+ /*
+ * AND/OR clause, with all clauses compatible with the selected MV
+ * stat
+ */
+
+ int i;
+ BoolExpr *orclause = ((BoolExpr *) clause);
+ List *orclauses = orclause->args;
+
+ /* match/mismatch bitmap for each MCV item */
+ int or_nmatches = 0;
+ char *or_matches = NULL;
+
+ Assert(orclauses != NIL);
+ Assert(list_length(orclauses) >= 2);
+
+ /* number of matching MCV items */
+ or_nmatches = mcvlist->nitems;
+
+ /* by default none of the MCV items matches the clauses */
+ or_matches = palloc0(sizeof(char) * or_nmatches);
+
+ if (or_clause(clause))
+ {
+ /* OR clauses assume nothing matches, initially */
+ memset(or_matches, MVSTATS_MATCH_NONE, sizeof(char) * or_nmatches);
+ or_nmatches = 0;
+ }
+ else
+ {
+ /* AND clauses assume nothing matches, initially */
+ memset(or_matches, MVSTATS_MATCH_FULL, sizeof(char) * or_nmatches);
+ }
+
+ /* build the match bitmap for the OR-clauses */
+ or_nmatches = update_match_bitmap_mcvlist(root, orclauses,
+ stakeys, mcvlist,
+ or_nmatches, or_matches,
+ lowsel, fullmatch, or_clause(clause));
+
+ /* merge the bitmap into the existing one */
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ /*
+ * Merge the result into the bitmap (Min for AND, Max for OR).
+ *
+ * FIXME this does not decrease the number of matches
+ */
+ UPDATE_RESULT(matches[i], or_matches[i], is_or);
+ }
+
+ pfree(or_matches);
+
+ }
+ else
+ {
+ elog(ERROR, "unknown clause type: %d", clause->type);
+ }
+ }
+
+ /*
+ * If all the columns were matched by equality, it's a full match. In this
+ * case there can be just a single MCV item, matching the clause (if there
+ * were two, both would match the other one).
+ */
+ *fullmatch = (bms_num_members(eqmatches) == mcvlist->ndimensions);
+
+ /* free the allocated pieces */
+ if (eqmatches)
+ pfree(eqmatches);
+
+ return nmatches;
+}
diff --git a/src/backend/optimizer/util/plancat.c b/src/backend/optimizer/util/plancat.c
index 91d4099..ff93ddb 100644
--- a/src/backend/optimizer/util/plancat.c
+++ b/src/backend/optimizer/util/plancat.c
@@ -422,7 +422,7 @@ get_relation_info(PlannerInfo *root, Oid relationObjectId, bool inhparent,
mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
/* unavailable stats are not interesting for the planner */
- if (mvstat->deps_built || mvstat->ndist_built)
+ if (mvstat->deps_built || mvstat->ndist_built || mvstat->mcv_built)
{
info = makeNode(MVStatisticInfo);
@@ -432,10 +432,12 @@ get_relation_info(PlannerInfo *root, Oid relationObjectId, bool inhparent,
/* enabled statistics */
info->ndist_enabled = mvstat->ndist_enabled;
info->deps_enabled = mvstat->deps_enabled;
+ info->mcv_enabled = mvstat->mcv_enabled;
/* built/available statistics */
info->ndist_built = mvstat->ndist_built;
info->deps_built = mvstat->deps_built;
+ info->mcv_built = mvstat->mcv_built;
/* stakeys */
adatum = SysCacheGetAttr(MVSTATOID, htup,
diff --git a/src/backend/utils/mvstats/Makefile b/src/backend/utils/mvstats/Makefile
index 21fe7e5..d5d47ba 100644
--- a/src/backend/utils/mvstats/Makefile
+++ b/src/backend/utils/mvstats/Makefile
@@ -12,6 +12,6 @@ subdir = src/backend/utils/mvstats
top_builddir = ../../../..
include $(top_builddir)/src/Makefile.global
-OBJS = common.o dependencies.o mvdist.o
+OBJS = common.o dependencies.o mcv.o mvdist.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/utils/mvstats/README.mcv b/src/backend/utils/mvstats/README.mcv
new file mode 100644
index 0000000..e93cfe4
--- /dev/null
+++ b/src/backend/utils/mvstats/README.mcv
@@ -0,0 +1,137 @@
+MCV lists
+=========
+
+Multivariate MCV (most-common values) lists are a straightforward extension of
+regular MCV list, tracking most frequent combinations of values for a group of
+attributes.
+
+This works particularly well for columns with a small number of distinct values,
+as the list may include all the combinations and approximate the distribution
+very accurately.
+
+For columns with large number of distinct values (e.g. those with continuous
+domains), the list will only track the most frequent combinations. If the
+distribution is mostly uniform (all combinations about equally frequent), the
+MCV list will be empty.
+
+Estimates of some clauses (e.g. equality) based on MCV lists are more accurate
+than when using histograms.
+
+Also, MCV lists don't necessarily require sorting of the values (the fact that
+we use sorting when building them is implementation detail), but even more
+importantly the ordering is not built into the approximation (while histograms
+are built on ordering). So MCV lists work well even for attributes where the
+ordering of the data type is disconnected from the meaning of the data. For
+example we know how to sort strings, but it's unlikely to make much sense for
+city names (or other label-like attributes).
+
+
+Selectivity estimation
+----------------------
+
+The estimation, implemented in clauselist_mv_selectivity_mcvlist(), is quite
+simple in principle - we need to identify MCV items matching all the clauses
+and sum frequencies of all those items.
+
+Currently MCV lists support estimation of the following clause types:
+
+ (a) equality clauses WHERE (a = 1) AND (b = 2)
+ (b) inequality clauses WHERE (a < 1) AND (b >= 2)
+ (c) NULL clauses WHERE (a IS NULL) AND (b IS NOT NULL)
+ (d) OR clauses WHERE (a < 1) OR (b >= 2)
+
+It's possible to add support for additional clauses, for example:
+
+ (e) multi-var clauses WHERE (a > b)
+
+and possibly others. These are tasks for the future, not yet implemented.
+
+
+Estimating equality clauses
+---------------------------
+
+When computing selectivity estimate for equality clauses
+
+ (a = 1) AND (b = 2)
+
+we can do this estimate pretty exactly assuming that two conditions are met:
+
+ (1) there's an equality condition on all attributes of the statistic
+
+ (2) we find a matching item in the MCV list
+
+In this case we know the MCV item represents all tuples matching the clauses,
+and the selectivity estimate is complete (i.e. we don't need to perform
+estimation using the histogram). This is what we call 'full match'.
+
+When only (1) holds, but there's no matching MCV item, we don't know whether
+there are no such rows or just are not very frequent. We can however use the
+frequency of the least frequent MCV item as an upper bound for the selectivity.
+
+For a combination of equality conditions (not full-match case) we can clamp the
+selectivity by the minimum of selectivities for each condition. For example if
+we know the number of distinct values for each column, we can use 1/ndistinct
+as a per-column estimate. Or rather 1/ndistinct + selectivity derived from the
+MCV list.
+
+We should also probably only use the 'residual ndistinct' by exluding the items
+included in the MCV list (and also residual frequency):
+
+ f = (1.0 - sum(MCV frequencies)) / (ndistinct - ndistinct(MCV list))
+
+but it's worth pointing out the ndistinct values are multi-variate for the
+columns referenced by the equality conditions.
+
+Note: Only the "full match" limit is currently implemented.
+
+
+Hashed MCV (not yet implemented)
+--------------------------------
+
+Regular MCV lists have to include actual values for each item, so if those items
+are large the list may be quite large. This is especially true for multi-variate
+MCV lists, although the current implementation partially mitigates this by
+performing de-duplicating the values before storing them on disk.
+
+It's possible to only store hashes (32-bit values) instead of the actual values,
+significantly reducing the space requirements. Obviously, this would only make
+the MCV lists useful for estimating equality conditions (assuming the 32-bit
+hashes make the collisions rare enough).
+
+This might also complicate matching the columns to available stats.
+
+
+TODO Consider implementing hashed MCV list, storing just 32-bit hashes instead
+ of the actual values. This type of MCV list will be useful only for
+ estimating equality clauses, and will reduce space requirements for large
+ varlena types (in such cases we usually only want equality anyway).
+
+TODO Currently there's no logic to consider building only a MCV list (and not
+ building the histogram at all), except for doing this decision manually in
+ ADD STATISTICS.
+
+
+Inspecting the MCV list
+-----------------------
+
+Inspecting the regular (per-attribute) MCV lists is trivial, as it's enough
+to select the columns from pg_stats - the data is encoded as anyarrays, so we
+simply get the text representation of the arrays.
+
+With multivariate MCV lits it's not that simple due to the possible mix of
+data types. It might be possible to produce similar array-like representation,
+but that'd unnecessarily complicate further processing and analysis of the MCV
+list. Instead, there's a SRF function providing values, frequencies etc.
+
+ SELECT * FROM pg_mv_mcv_items();
+
+It has two input parameters:
+
+ oid - OID of the MCV list (pg_mv_statistic.staoid)
+
+and produces a table with these columns:
+
+ - item ID (0...nitems-1)
+ - values (string array)
+ - nulls only (boolean array)
+ - frequency (double precision)
diff --git a/src/backend/utils/mvstats/README.stats b/src/backend/utils/mvstats/README.stats
index 251a468..57f5c8b 100644
--- a/src/backend/utils/mvstats/README.stats
+++ b/src/backend/utils/mvstats/README.stats
@@ -8,9 +8,50 @@ not true, resulting in estimation errors.
Multivariate stats track different types of dependencies between the columns,
hopefully improving the estimates.
-Currently we only have one kind of multivariate statistics - soft functional
-dependencies, and we use it to improve estimates of equality clauses. See
-README.dependencies for details.
+
+Types of statistics
+-------------------
+
+Currently we only have two kinds of multivariate statistics
+
+ (a) soft functional dependencies (README.dependencies)
+
+ (b) MCV lists (README.mcv)
+
+
+Compatible clause types
+-----------------------
+
+Each type of statistics may be used to estimate some subset of clause types.
+
+ (a) functional dependencies - equality clauses (AND), possibly IS NULL
+
+ (b) MCV list - equality and inequality clauses, IS [NOT] NULL, AND/OR
+
+Currently only simple operator clauses (Var op Const) are supported, but it's
+possible to support more complex clause types, e.g. (Var op Var).
+
+
+Complex clauses
+---------------
+
+We also support estimating more complex clauses - essentially AND/OR clauses
+with (Var op Const) as leaves, as long as all the referenced attributes are
+covered by a single statistics.
+
+For example this condition
+
+ (a=1) AND ((b=2) OR ((c=3) AND (d=4)))
+
+may be estimated using statistics on (a,b,c,d). If we only have statistics on
+(b,c,d) we may estimate the second part, and estimate (a=1) using simple stats.
+
+If we only have statistics on (a,b,c) we can't apply it at all at this point,
+but it's worth pointing out clauselist_selectivity() works recursively and when
+handling the second part (the OR-clause), we'll be able to apply the statistics.
+
+Note: The multi-statistics estimation patch also makes it possible to pass some
+clauses as 'conditions' into the deeper parts of the expression tree.
Selectivity estimation
@@ -23,21 +64,53 @@ When estimating selectivity, we aim to achieve several things:
(b) minimize the overhead, especially when no suitable multivariate stats
exist (so if you are not using multivariate stats, there's no overhead)
-This clauselist_selectivity() performs several inexpensive checks first, before
+Thus clauselist_selectivity() performs several inexpensive checks first, before
even attempting to do the more expensive estimation.
(1) check if there are multivariate stats on the relation
- (2) check there are at least two attributes referenced by clauses compatible
- with multivariate statistics (equality clauses for func. dependencies)
+ (2) check that there are functional dependencies on the table, and that
+ there are at least two attributes referenced by compatible clauses
+ (equality clauses for func. dependencies)
(3) perform reduction of equality clauses using func. dependencies
- (4) estimate the reduced list of clauses using regular statistics
+ (4) check that there are multivariate MCV lists on the table, and that
+ there are at least two attributes referenced by compatible clauses
+ (equalities, inequalities, etc.)
+
+ (5) find the best multivariate statistics (matching the most conditions)
+ and use it to compute the estimate
+
+ (6) estimate the remaining clauses (not estimated using multivariate stats)
+ using the regular per-column statistics
Whenever we find there are no suitable stats, we skip the expensive steps.
+Further (possibly crazy) ideas
+------------------------------
+
+Currently the clauses are only estimated using a single statistics, even if
+there are multiple candidate statistics - for example assume we have statistics
+on (a,b,c) and (b,c,d), and estimate conditions
+
+ (b = 1) AND (c = 2)
+
+Then both statistics may be used, but we only use one of them. Maybe we could
+use compute estimates using all candidate stats, and somehow aggregate them
+into the final estimate by using average or median.
+
+Some stats may give better estimates than others, but it's very difficult to say
+in advance which stats are the best (it depends on the number of buckets, number
+of additional columns not referenced in the clauses, type of condition etc.).
+
+But of course, this may result in expensive estimation (CPU-wise).
+
+So we might add a GUC to choose between a simple (single statistics) and thus
+multi-statistic estimation, possibly table-level parameter (ALTER TABLE ...).
+
+
Size of sample in ANALYZE
-------------------------
When performing ANALYZE, the number of rows to sample is determined as
diff --git a/src/backend/utils/mvstats/common.c b/src/backend/utils/mvstats/common.c
index 5f3fb55..5d9caa8 100644
--- a/src/backend/utils/mvstats/common.c
+++ b/src/backend/utils/mvstats/common.c
@@ -15,13 +15,13 @@
*/
#include "common.h"
+#include "utils/array.h"
static VacAttrStats **lookup_var_attr_stats(int2vector *attrs,
int natts, VacAttrStats **vacattrstats);
static List *list_mv_stats(Oid relid);
-
/*
* Compute requested multivariate stats, using the rows sampled for the
* plain (single-column) stats.
@@ -51,6 +51,8 @@ build_mv_stats(Relation onerel, double totalrows,
MVStatisticInfo *stat = (MVStatisticInfo *) lfirst(lc);
MVNDistinct ndistinct = NULL;
MVDependencies deps = NULL;
+ MCVList mcvlist = NULL;
+ int numrows_filtered = 0;
VacAttrStats **stats = NULL;
int numatts = 0;
@@ -91,8 +93,12 @@ build_mv_stats(Relation onerel, double totalrows,
if (stat->deps_enabled)
deps = build_mv_dependencies(numrows, rows, attrs, stats);
+ /* build the MCV list */
+ if (stat->mcv_enabled)
+ mcvlist = build_mv_mcvlist(numrows, rows, attrs, stats, &numrows_filtered);
+
/* store the statistics in the catalog */
- update_mv_stats(stat->mvoid, ndistinct, deps, attrs, stats);
+ update_mv_stats(stat->mvoid, ndistinct, deps, mcvlist, attrs, stats);
}
}
@@ -174,6 +180,8 @@ list_mv_stats(Oid relid)
info->ndist_built = stats->ndist_built;
info->deps_enabled = stats->deps_enabled;
info->deps_built = stats->deps_built;
+ info->mcv_enabled = stats->mcv_enabled;
+ info->mcv_built = stats->mcv_built;
result = lappend(result, info);
}
@@ -190,8 +198,56 @@ list_mv_stats(Oid relid)
return result;
}
+
+/*
+ * Find attnims of MV stats using the mvoid.
+ */
+int2vector *
+find_mv_attnums(Oid mvoid, Oid *relid)
+{
+ ArrayType *arr;
+ Datum adatum;
+ bool isnull;
+ HeapTuple htup;
+ int2vector *keys;
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ htup = SearchSysCache1(MVSTATOID,
+ ObjectIdGetDatum(mvoid));
+
+ /* XXX syscache contains OIDs of deleted stats (not invalidated) */
+ if (!HeapTupleIsValid(htup))
+ return NULL;
+
+ /* starelid */
+ adatum = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_starelid, &isnull);
+ Assert(!isnull);
+
+ *relid = DatumGetObjectId(adatum);
+
+ /* stakeys */
+ adatum = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_stakeys, &isnull);
+ Assert(!isnull);
+
+ arr = DatumGetArrayTypeP(adatum);
+
+ keys = buildint2vector((int16 *) ARR_DATA_PTR(arr),
+ ARR_DIMS(arr)[0]);
+ ReleaseSysCache(htup);
+
+ /*
+ * TODO maybe save the list into relcache, as in RelationGetIndexList
+ * (which was used as an inspiration of this one)?.
+ */
+
+ return keys;
+}
+
void
-update_mv_stats(Oid mvoid, MVNDistinct ndistinct, MVDependencies dependencies,
+update_mv_stats(Oid mvoid,
+ MVNDistinct ndistinct, MVDependencies dependencies, MCVList mcvlist,
int2vector *attrs, VacAttrStats **stats)
{
HeapTuple stup,
@@ -225,22 +281,36 @@ update_mv_stats(Oid mvoid, MVNDistinct ndistinct, MVDependencies dependencies,
= PointerGetDatum(serialize_mv_dependencies(dependencies));
}
+ if (mcvlist != NULL)
+ {
+ bytea *data = serialize_mv_mcvlist(mcvlist, attrs, stats);
+
+ nulls[Anum_pg_mv_statistic_stamcv - 1] = (data == NULL);
+ values[Anum_pg_mv_statistic_stamcv - 1] = PointerGetDatum(data);
+ }
+
/* always replace the value (either by bytea or NULL) */
replaces[Anum_pg_mv_statistic_standist - 1] = true;
replaces[Anum_pg_mv_statistic_stadeps - 1] = true;
+ replaces[Anum_pg_mv_statistic_stamcv - 1] = true;
/* always change the availability flags */
nulls[Anum_pg_mv_statistic_ndist_built - 1] = false;
nulls[Anum_pg_mv_statistic_deps_built - 1] = false;
+ nulls[Anum_pg_mv_statistic_mcv_built - 1] = false;
+
nulls[Anum_pg_mv_statistic_stakeys - 1] = false;
/* use the new attnums, in case we removed some dropped ones */
replaces[Anum_pg_mv_statistic_ndist_built - 1] = true;
replaces[Anum_pg_mv_statistic_deps_built - 1] = true;
+ replaces[Anum_pg_mv_statistic_mcv_built - 1] = true;
+
replaces[Anum_pg_mv_statistic_stakeys - 1] = true;
values[Anum_pg_mv_statistic_ndist_built - 1] = BoolGetDatum(ndistinct != NULL);
values[Anum_pg_mv_statistic_deps_built - 1] = BoolGetDatum(dependencies != NULL);
+ values[Anum_pg_mv_statistic_mcv_built - 1] = BoolGetDatum(mcvlist != NULL);
values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(attrs);
@@ -270,6 +340,23 @@ update_mv_stats(Oid mvoid, MVNDistinct ndistinct, MVDependencies dependencies,
heap_close(sd, RowExclusiveLock);
}
+
+int
+mv_get_index(AttrNumber varattno, int2vector *stakeys)
+{
+ int i,
+ idx = 0;
+
+ for (i = 0; i < stakeys->dim1; i++)
+ {
+ if (stakeys->values[i] < varattno)
+ idx += 1;
+ else
+ break;
+ }
+ return idx;
+}
+
/* multi-variate stats comparator */
/*
@@ -280,11 +367,15 @@ update_mv_stats(Oid mvoid, MVNDistinct ndistinct, MVDependencies dependencies,
int
compare_scalars_simple(const void *a, const void *b, void *arg)
{
- Datum da = *(Datum *) a;
- Datum db = *(Datum *) b;
- SortSupport ssup = (SortSupport) arg;
+ return compare_datums_simple(*(Datum *) a,
+ *(Datum *) b,
+ (SortSupport) arg);
+}
- return ApplySortComparator(da, false, db, false, ssup);
+int
+compare_datums_simple(Datum a, Datum b, SortSupport ssup)
+{
+ return ApplySortComparator(a, false, b, false, ssup);
}
/*
@@ -401,3 +492,34 @@ multi_sort_compare_dims(int start, int end,
return 0;
}
+
+/* simple counterpart to qsort_arg */
+void *
+bsearch_arg(const void *key, const void *base, size_t nmemb, size_t size,
+ int (*compar) (const void *, const void *, void *),
+ void *arg)
+{
+ size_t l,
+ u,
+ idx;
+ const void *p;
+ int comparison;
+
+ l = 0;
+ u = nmemb;
+ while (l < u)
+ {
+ idx = (l + u) / 2;
+ p = (void *) (((const char *) base) + (idx * size));
+ comparison = (*compar) (key, p, arg);
+
+ if (comparison < 0)
+ u = idx;
+ else if (comparison > 0)
+ l = idx + 1;
+ else
+ return (void *) p;
+ }
+
+ return NULL;
+}
diff --git a/src/backend/utils/mvstats/common.h b/src/backend/utils/mvstats/common.h
index e471c88..fe56f51 100644
--- a/src/backend/utils/mvstats/common.h
+++ b/src/backend/utils/mvstats/common.h
@@ -47,6 +47,15 @@ typedef struct
int tupno; /* position index for tuple it came from */
} ScalarItem;
+/* (de)serialization info */
+typedef struct DimensionInfo
+{
+ int nvalues; /* number of deduplicated values */
+ int nbytes; /* number of bytes (serialized) */
+ int typlen; /* pg_type.typlen */
+ bool typbyval; /* pg_type.typbyval */
+} DimensionInfo;
+
/* multi-sort */
typedef struct MultiSortSupportData
{
@@ -60,6 +69,7 @@ typedef struct SortItem
{
Datum *values;
bool *isnull;
+ int count;
} SortItem;
MultiSortSupport multi_sort_init(int ndims);
@@ -67,7 +77,7 @@ MultiSortSupport multi_sort_init(int ndims);
void multi_sort_add_dimension(MultiSortSupport mss, int sortdim,
int dim, VacAttrStats **vacattrstats);
-int multi_sort_compare(const void *a, const void *b, void *arg);
+int multi_sort_compare(const void *a, const void *b, void *arg);
int multi_sort_compare_dim(int dim, const SortItem *a,
const SortItem *b, MultiSortSupport mss);
@@ -76,5 +86,11 @@ int multi_sort_compare_dims(int start, int end, const SortItem *a,
const SortItem *b, MultiSortSupport mss);
/* comparators, used when constructing multivariate stats */
-int compare_scalars_simple(const void *a, const void *b, void *arg);
-int compare_scalars_partition(const void *a, const void *b, void *arg);
+int compare_datums_simple(Datum a, Datum b, SortSupport ssup);
+int compare_scalars_simple(const void *a, const void *b, void *arg);
+int compare_scalars_partition(const void *a, const void *b, void *arg);
+
+void *bsearch_arg(const void *key, const void *base,
+ size_t nmemb, size_t size,
+ int (*compar) (const void *, const void *, void *),
+ void *arg);
diff --git a/src/backend/utils/mvstats/mcv.c b/src/backend/utils/mvstats/mcv.c
new file mode 100644
index 0000000..c1c2409
--- /dev/null
+++ b/src/backend/utils/mvstats/mcv.c
@@ -0,0 +1,1184 @@
+/*-------------------------------------------------------------------------
+ *
+ * mcv.c
+ * POSTGRES multivariate MCV lists
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/mcv.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "fmgr.h"
+#include "funcapi.h"
+
+#include "utils/bytea.h"
+#include "utils/lsyscache.h"
+
+#include "common.h"
+
+/*
+ * Each serialized item needs to store (in this order):
+ *
+ * - indexes (ndim * sizeof(uint16))
+ * - null flags (ndim * sizeof(bool))
+ * - frequency (sizeof(double))
+ *
+ * So in total:
+ *
+ * ndim * (sizeof(uint16) + sizeof(bool)) + sizeof(double)
+ */
+#define ITEM_SIZE(ndims) \
+ (ndims * (sizeof(uint16) + sizeof(bool)) + sizeof(double))
+
+/* Macros for convenient access to parts of the serialized MCV item */
+#define ITEM_INDEXES(item) ((uint16*)item)
+#define ITEM_NULLS(item,ndims) ((bool*)(ITEM_INDEXES(item) + ndims))
+#define ITEM_FREQUENCY(item,ndims) ((double*)(ITEM_NULLS(item,ndims) + ndims))
+
+static MultiSortSupport build_mss(VacAttrStats **stats, int2vector *attrs);
+
+static SortItem *build_sorted_items(int numrows, HeapTuple *rows,
+ TupleDesc tdesc, MultiSortSupport mss,
+ int2vector *attrs);
+
+static SortItem *build_distinct_groups(int numrows, SortItem *items,
+ MultiSortSupport mss, int *ndistinct);
+
+static int count_distinct_groups(int numrows, SortItem *items,
+ MultiSortSupport mss);
+
+/*
+ * Builds MCV list from the set of sampled rows.
+ *
+ * The algorithm is quite simple:
+ *
+ * (1) sort the data (default collation, '<' for the data type)
+ *
+ * (2) count distinct groups, decide how many to keep
+ *
+ * (3) build the MCV list using the threshold determined in (2)
+ *
+ * (4) remove rows represented by the MCV from the sample
+ *
+ * The method also removes rows matching the MCV items from the input array,
+ * and passes the number of remaining rows (useful for building histograms)
+ * using the numrows_filtered parameter.
+ *
+ * FIXME: Single-dimensional MCV is sorted by frequency (descending). We should
+ * do that too, because when walking through the list we want to check
+ * the most frequent items first.
+ *
+ * TODO: We're using Datum (8B), even for data types (e.g. int4 or float4).
+ * Maybe we could save some space here, but the bytea compression should
+ * handle it just fine.
+ *
+ * TODO: This probably should not use the ndistinct directly (as computed from
+ * the table, but rather estimate the number of distinct values in the
+ * table), no?
+ */
+MCVList
+build_mv_mcvlist(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int *numrows_filtered)
+{
+ int i;
+ int numattrs = attrs->dim1;
+ int ndistinct = 0;
+ int mcv_threshold = 0;
+ int nitems = 0;
+
+ MCVList mcvlist = NULL;
+
+ /* comparator for all the columns */
+ MultiSortSupport mss = build_mss(stats, attrs);
+
+ /* sort the rows */
+ SortItem *items = build_sorted_items(numrows, rows, stats[0]->tupDesc,
+ mss, attrs);
+
+ /* transform the sorted rows into groups (sorted by frequency) */
+ SortItem *groups = build_distinct_groups(numrows, items, mss, &ndistinct);
+
+ /*
+ * Determine the minimum size of a group to be eligible for MCV list, and
+ * check how many groups actually pass that threshold. We use 1.25x the
+ * avarage group size, just like for regular statistics.
+ *
+ * But if we can fit all the distinct values in the MCV list (i.e. if
+ * there are less distinct groups than MVSTAT_MCVLIST_MAX_ITEMS), we'll
+ * require only 2 rows per group.
+ */
+ mcv_threshold = 1.25 * numrows / ndistinct;
+ mcv_threshold = (mcv_threshold < 4) ? 4 : mcv_threshold;
+
+ if (ndistinct <= MVSTAT_MCVLIST_MAX_ITEMS)
+ mcv_threshold = 2;
+
+ /* Walk through the groups and stop once we fall below the threshold. */
+ nitems = 0;
+ for (i = 0; i < ndistinct; i++)
+ {
+ if (groups[i].count < mcv_threshold)
+ break;
+
+ nitems++;
+ }
+
+ /* we know the number of MCV list items, so let's build the list */
+ if (nitems > 0)
+ {
+ /* allocate the MCV list structure, set parameters we know */
+ mcvlist = (MCVList) palloc0(sizeof(MCVListData));
+
+ mcvlist->magic = MVSTAT_MCV_MAGIC;
+ mcvlist->type = MVSTAT_MCV_TYPE_BASIC;
+ mcvlist->ndimensions = numattrs;
+ mcvlist->nitems = nitems;
+
+ /*
+ * Preallocate Datum/isnull arrays (not as a single chunk, as we will
+ * pass the result outside and thus it needs to be easy to pfree().
+ *
+ * XXX Although we're the only ones dealing with this.
+ */
+ mcvlist->items = (MCVItem *) palloc0(sizeof(MCVItem) * nitems);
+
+ for (i = 0; i < nitems; i++)
+ {
+ mcvlist->items[i] = (MCVItem) palloc0(sizeof(MCVItemData));
+ mcvlist->items[i]->values = (Datum *) palloc0(sizeof(Datum) * numattrs);
+ mcvlist->items[i]->isnull = (bool *) palloc0(sizeof(bool) * numattrs);
+ }
+
+ /* Copy the first chunk of groups into the result. */
+ for (i = 0; i < nitems; i++)
+ {
+ /* just pointer to the proper place in the list */
+ MCVItem item = mcvlist->items[i];
+
+ /* copy values from the _previous_ group (last item of) */
+ memcpy(item->values, groups[i].values, sizeof(Datum) * numattrs);
+ memcpy(item->isnull, groups[i].isnull, sizeof(bool) * numattrs);
+
+ /* and finally the group frequency */
+ item->frequency = (double) groups[i].count / numrows;
+ }
+
+ /* make sure the loops are consistent */
+ Assert(nitems == mcvlist->nitems);
+
+ /*
+ * Remove the rows matching the MCV list (i.e. keep only rows that are
+ * not represented by the MCV list). We will first sort the groups by
+ * the keys (not by count) and then use binary search.
+ */
+ if (nitems > ndistinct)
+ {
+ int i,
+ j;
+ int nfiltered = 0;
+
+ /* used for the searches */
+ SortItem key;
+
+ /* wfill this with data from the rows */
+ key.values = (Datum *) palloc0(numattrs * sizeof(Datum));
+ key.isnull = (bool *) palloc0(numattrs * sizeof(bool));
+
+ /*
+ * Sort the groups for bsearch_r (but only the items that actually
+ * made it to the MCV list).
+ */
+ qsort_arg((void *) groups, nitems, sizeof(SortItem),
+ multi_sort_compare, mss);
+
+ /* walk through the tuples, compare the values to MCV items */
+ for (i = 0; i < numrows; i++)
+ {
+ /* collect the key values from the row */
+ for (j = 0; j < numattrs; j++)
+ key.values[j]
+ = heap_getattr(rows[i], attrs->values[j],
+ stats[j]->tupDesc, &key.isnull[j]);
+
+ /* if not included in the MCV list, keep it in the array */
+ if (bsearch_arg(&key, groups, nitems, sizeof(SortItem),
+ multi_sort_compare, mss) == NULL)
+ rows[nfiltered++] = rows[i];
+ }
+
+ /* remember how many rows we actually kept */
+ *numrows_filtered = nfiltered;
+
+ /* free all the data used here */
+ pfree(key.values);
+ pfree(key.isnull);
+ }
+ else
+ /* the MCV list convers all the rows */
+ *numrows_filtered = 0;
+ }
+
+ pfree(items);
+ pfree(groups);
+
+ return mcvlist;
+}
+
+/* build MultiSortSupport for the attributes passed in attrs */
+static MultiSortSupport
+build_mss(VacAttrStats **stats, int2vector *attrs)
+{
+ int i;
+ int numattrs = attrs->dim1;
+
+ /* Sort by multiple columns (using array of SortSupport) */
+ MultiSortSupport mss = multi_sort_init(numattrs);
+
+ /* prepare the sort functions for all the attributes */
+ for (i = 0; i < numattrs; i++)
+ multi_sort_add_dimension(mss, i, i, stats);
+
+ return mss;
+}
+
+/* build sorted array of SortItem with values from rows */
+static SortItem *
+build_sorted_items(int numrows, HeapTuple *rows, TupleDesc tdesc,
+ MultiSortSupport mss, int2vector *attrs)
+{
+ int i,
+ j,
+ len;
+ int numattrs = attrs->dim1;
+ int nvalues = numrows * numattrs;
+
+ /*
+ * We won't allocate the arrays for each item independenly, but in one
+ * large chunk and then just set the pointers.
+ */
+ SortItem *items;
+ Datum *values;
+ bool *isnull;
+ char *ptr;
+
+ /* Compute the total amount of memory we need (both items and values). */
+ len = numrows * sizeof(SortItem) + nvalues * (sizeof(Datum) + sizeof(bool));
+
+ /* Allocate the memory and split it into the pieces. */
+ ptr = palloc0(len);
+
+ /* items to sort */
+ items = (SortItem *) ptr;
+ ptr += numrows * sizeof(SortItem);
+
+ /* values and null flags */
+ values = (Datum *) ptr;
+ ptr += nvalues * sizeof(Datum);
+
+ isnull = (bool *) ptr;
+ ptr += nvalues * sizeof(bool);
+
+ /* make sure we consumed the whole buffer exactly */
+ Assert((ptr - (char *) items) == len);
+
+ /* fix the pointers to Datum and bool arrays */
+ for (i = 0; i < numrows; i++)
+ {
+ items[i].values = &values[i * numattrs];
+ items[i].isnull = &isnull[i * numattrs];
+
+ /* load the values/null flags from sample rows */
+ for (j = 0; j < numattrs; j++)
+ {
+ items[i].values[j] = heap_getattr(rows[i],
+ attrs->values[j], /* attnum */
+ tdesc,
+ &items[i].isnull[j]); /* isnull */
+ }
+ }
+
+ /* do the sort, using the multi-sort */
+ qsort_arg((void *) items, numrows, sizeof(SortItem),
+ multi_sort_compare, mss);
+
+ return items;
+}
+
+/* count distinct combinations of SortItems in the array */
+static int
+count_distinct_groups(int numrows, SortItem *items, MultiSortSupport mss)
+{
+ int i;
+ int ndistinct;
+
+ ndistinct = 1;
+ for (i = 1; i < numrows; i++)
+ if (multi_sort_compare(&items[i], &items[i - 1], mss) != 0)
+ ndistinct += 1;
+
+ return ndistinct;
+}
+
+/* compares frequencies of the SortItem entries (in descending order) */
+static int
+compare_sort_item_count(const void *a, const void *b)
+{
+ SortItem *ia = (SortItem *) a;
+ SortItem *ib = (SortItem *) b;
+
+ if (ia->count == ib->count)
+ return 0;
+ else if (ia->count > ib->count)
+ return -1;
+
+ return 1;
+}
+
+/* builds SortItems for distinct groups and counts the matching items */
+static SortItem *
+build_distinct_groups(int numrows, SortItem *items, MultiSortSupport mss,
+ int *ndistinct)
+{
+ int i,
+ j;
+ int ngroups = count_distinct_groups(numrows, items, mss);
+
+ SortItem *groups = (SortItem *) palloc0(ngroups * sizeof(SortItem));
+
+ j = 0;
+ groups[0] = items[0];
+ groups[0].count = 1;
+
+ for (i = 1; i < numrows; i++)
+ {
+ if (multi_sort_compare(&items[i], &items[i - 1], mss) != 0)
+ groups[++j] = items[i];
+
+ groups[j].count++;
+ }
+
+ pg_qsort((void *) groups, ngroups, sizeof(SortItem),
+ compare_sort_item_count);
+
+ *ndistinct = ngroups;
+ return groups;
+}
+
+
+/* fetch the MCV list (as a bytea) from the pg_mv_statistic catalog */
+MCVList
+load_mv_mcvlist(Oid mvoid)
+{
+ bool isnull = false;
+ Datum mcvlist;
+
+#ifdef USE_ASSERT_CHECKING
+ Form_pg_mv_statistic mvstat;
+#endif
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ HeapTuple htup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(mvoid));
+
+ if (!HeapTupleIsValid(htup))
+ return NULL;
+
+#ifdef USE_ASSERT_CHECKING
+ mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
+ Assert(mvstat->mcv_enabled && mvstat->mcv_built);
+#endif
+
+ mcvlist = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_stamcv, &isnull);
+
+ Assert(!isnull);
+
+ ReleaseSysCache(htup);
+
+ return deserialize_mv_mcvlist(DatumGetByteaP(mcvlist));
+}
+
+/* print some basic info about the MCV list
+ *
+ * TODO: Add info about what part of the table this covers.
+ */
+Datum
+pg_mv_stats_mcvlist_info(PG_FUNCTION_ARGS)
+{
+ bytea *data = PG_GETARG_BYTEA_P(0);
+ char *result;
+
+ MCVList mcvlist = deserialize_mv_mcvlist(data);
+
+ result = palloc0(128);
+ snprintf(result, 128, "nitems=%d", mcvlist->nitems);
+
+ pfree(mcvlist);
+
+ PG_RETURN_TEXT_P(cstring_to_text(result));
+}
+
+/*
+ * serialize MCV list into a bytea value
+ *
+ *
+ * The basic algorithm is simple:
+ *
+ * (1) perform deduplication (for each attribute separately)
+ * (a) collect all (non-NULL) attribute values from all MCV items
+ * (b) sort the data (using 'lt' from VacAttrStats)
+ * (c) remove duplicate values from the array
+ *
+ * (2) serialize the arrays into a bytea value
+ *
+ * (3) process all MCV list items
+ * (a) replace values with indexes into the arrays
+ *
+ * Each attribute has to be processed separately, because we may be mixing
+ * different datatypes, with different sort operators, etc.
+ *
+ * We'll use uint16 values for the indexes in step (3), as we don't allow more
+ * than 8k MCV items, although that's mostly arbitrary limit. We might increase
+ * this to 65k and still fit into uint16.
+ *
+ * We don't really expect the serialization to save as much space as for
+ * histograms, because we are not doing any bucket splits (which is the source
+ * of high redundancy in histograms).
+ *
+ * TODO: Consider packing boolean flags (NULL) for each item into a single char
+ * (or a longer type) instead of using an array of bool items.
+ */
+bytea *
+serialize_mv_mcvlist(MCVList mcvlist, int2vector *attrs,
+ VacAttrStats **stats)
+{
+ int i,
+ j;
+ int ndims = mcvlist->ndimensions;
+ int itemsize = ITEM_SIZE(ndims);
+
+ SortSupport ssup;
+ DimensionInfo *info;
+
+ Size total_length;
+
+ /* allocate just once */
+ char *item = palloc0(itemsize);
+
+ /* serialized items (indexes into arrays, etc.) */
+ bytea *output;
+ char *data = NULL;
+
+ /* values per dimension (and number of non-NULL values) */
+ Datum **values = (Datum **) palloc0(sizeof(Datum *) * ndims);
+ int *counts = (int *) palloc0(sizeof(int) * ndims);
+
+ /*
+ * We'll include some rudimentary information about the attributes (type
+ * length, etc.), so that we don't have to look them up while
+ * deserializing the MCV list.
+ */
+ info = (DimensionInfo *) palloc0(sizeof(DimensionInfo) * ndims);
+
+ /* sort support data for all attributes included in the MCV list */
+ ssup = (SortSupport) palloc0(sizeof(SortSupportData) * ndims);
+
+ /* collect and deduplicate values for all attributes */
+ for (i = 0; i < ndims; i++)
+ {
+ int ndistinct;
+ StdAnalyzeData *tmp = (StdAnalyzeData *) stats[i]->extra_data;
+
+ /* copy important info about the data type (length, by-value) */
+ info[i].typlen = stats[i]->attrtype->typlen;
+ info[i].typbyval = stats[i]->attrtype->typbyval;
+
+ /* allocate space for values in the attribute and collect them */
+ values[i] = (Datum *) palloc0(sizeof(Datum) * mcvlist->nitems);
+
+ for (j = 0; j < mcvlist->nitems; j++)
+ {
+ /* skip NULL values - we don't need to serialize them */
+ if (mcvlist->items[j]->isnull[i])
+ continue;
+
+ values[i][counts[i]] = mcvlist->items[j]->values[i];
+ counts[i] += 1;
+ }
+
+ /* there are just NULL values in this dimension, we're done */
+ if (counts[i] == 0)
+ continue;
+
+ /* sort and deduplicate the data */
+ ssup[i].ssup_cxt = CurrentMemoryContext;
+ ssup[i].ssup_collation = DEFAULT_COLLATION_OID;
+ ssup[i].ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(tmp->ltopr, &ssup[i]);
+
+ qsort_arg(values[i], counts[i], sizeof(Datum),
+ compare_scalars_simple, &ssup[i]);
+
+ /*
+ * Walk through the array and eliminate duplicate values, but keep the
+ * ordering (so that we can do bsearch later). We know there's at
+ * least one item as (counts[i] != 0), so we can skip the first
+ * element.
+ */
+ ndistinct = 1; /* number of distinct values */
+ for (j = 1; j < counts[i]; j++)
+ {
+ /* if the value is the same as the previous one, we can skip it */
+ if (!compare_datums_simple(values[i][j - 1], values[i][j], &ssup[i]))
+ continue;
+
+ values[i][ndistinct] = values[i][j];
+ ndistinct += 1;
+ }
+
+ /* we must not exceed UINT16_MAX, as we use uint16 indexes */
+ Assert(ndistinct <= UINT16_MAX);
+
+ /*
+ * Store additional info about the attribute - number of deduplicated
+ * values, and also size of the serialized data. For fixed-length data
+ * types this is trivial to compute, for varwidth types we need to
+ * actually walk the array and sum the sizes.
+ */
+ info[i].nvalues = ndistinct;
+
+ if (info[i].typlen > 0) /* fixed-length data types */
+ info[i].nbytes = info[i].nvalues * info[i].typlen;
+ else if (info[i].typlen == -1) /* varlena */
+ {
+ info[i].nbytes = 0;
+ for (j = 0; j < info[i].nvalues; j++)
+ info[i].nbytes += VARSIZE_ANY(values[i][j]);
+ }
+ else if (info[i].typlen == -2) /* cstring */
+ {
+ info[i].nbytes = 0;
+ for (j = 0; j < info[i].nvalues; j++)
+ info[i].nbytes += strlen(DatumGetPointer(values[i][j]));
+ }
+
+ /* we know (count>0) so there must be some data */
+ Assert(info[i].nbytes > 0);
+ }
+
+ /*
+ * Now we can finally compute how much space we'll actually need for the
+ * serialized MCV list, as it contains these fields:
+ *
+ * - length (4B) for varlena - magic (4B) - type (4B) - ndimensions (4B) -
+ * nitems (4B) - info (ndim * sizeof(DimensionInfo) - arrays of values for
+ * each dimension - serialized items (nitems * itemsize)
+ *
+ * So the 'header' size is 20B + ndim * sizeof(DimensionInfo) and then we
+ * will place all the data (values + indexes).
+ */
+ total_length = (sizeof(int32) + offsetof(MCVListData, items)
+ +ndims * sizeof(DimensionInfo)
+ + mcvlist->nitems * itemsize);
+
+ for (i = 0; i < ndims; i++)
+ total_length += info[i].nbytes;
+
+ /* enforce arbitrary limit of 1MB */
+ if (total_length > (1024 * 1024))
+ elog(ERROR, "serialized MCV list exceeds 1MB (%ld)", total_length);
+
+ /* allocate space for the serialized MCV list, set header fields */
+ output = (bytea *) palloc0(total_length);
+ SET_VARSIZE(output, total_length);
+
+ /* 'data' points to the current position in the output buffer */
+ data = VARDATA(output);
+
+ /* MCV list header (number of items, ...) */
+ memcpy(data, mcvlist, offsetof(MCVListData, items));
+ data += offsetof(MCVListData, items);
+
+ /* information about the attributes */
+ memcpy(data, info, sizeof(DimensionInfo) * ndims);
+ data += sizeof(DimensionInfo) * ndims;
+
+ /* now serialize the deduplicated values for all attributes */
+ for (i = 0; i < ndims; i++)
+ {
+#ifdef USE_ASSERT_CHECKING
+ char *tmp = data; /* remember the starting point */
+#endif
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ Datum v = values[i][j];
+
+ if (info[i].typbyval) /* passed by value */
+ {
+ memcpy(data, &v, info[i].typlen);
+ data += info[i].typlen;
+ }
+ else if (info[i].typlen > 0) /* pased by reference */
+ {
+ memcpy(data, DatumGetPointer(v), info[i].typlen);
+ data += info[i].typlen;
+ }
+ else if (info[i].typlen == -1) /* varlena */
+ {
+ memcpy(data, DatumGetPointer(v), VARSIZE_ANY(v));
+ data += VARSIZE_ANY(v);
+ }
+ else if (info[i].typlen == -2) /* cstring */
+ {
+ memcpy(data, DatumGetPointer(v), strlen(DatumGetPointer(v)) + 1);
+ data += strlen(DatumGetPointer(v)) + 1; /* terminator */
+ }
+ }
+
+ /* make sure we got exactly the amount of data we expected */
+ Assert((data - tmp) == info[i].nbytes);
+ }
+
+ /* finally serialize the items, with uint16 indexes instead of the values */
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ MCVItem mcvitem = mcvlist->items[i];
+
+ /* don't write beyond the allocated space */
+ Assert(data <= (char *) output + total_length - itemsize);
+
+ /* reset the item (we only allocate it once and reuse it) */
+ memset(item, 0, itemsize);
+
+ for (j = 0; j < ndims; j++)
+ {
+ Datum *v = NULL;
+
+ /* do the lookup only for non-NULL values */
+ if (mcvlist->items[i]->isnull[j])
+ continue;
+
+ v = (Datum *) bsearch_arg(&mcvitem->values[j], values[j],
+ info[j].nvalues, sizeof(Datum),
+ compare_scalars_simple, &ssup[j]);
+
+ Assert(v != NULL); /* serialization or deduplication error */
+
+ /* compute index within the array */
+ ITEM_INDEXES(item)[j] = (v - values[j]);
+
+ /* check the index is within expected bounds */
+ Assert(ITEM_INDEXES(item)[j] >= 0);
+ Assert(ITEM_INDEXES(item)[j] < info[j].nvalues);
+ }
+
+ /* copy NULL and frequency flags into the item */
+ memcpy(ITEM_NULLS(item, ndims), mcvitem->isnull, sizeof(bool) * ndims);
+ memcpy(ITEM_FREQUENCY(item, ndims), &mcvitem->frequency, sizeof(double));
+
+ /* copy the serialized item into the array */
+ memcpy(data, item, itemsize);
+
+ data += itemsize;
+ }
+
+ /* at this point we expect to match the total_length exactly */
+ Assert((data - (char *) output) == total_length);
+
+ return output;
+}
+
+/*
+ * deserialize MCV list from the varlena value
+ *
+ *
+ * We deserialize the MCV list fully, because we don't expect there bo be a lot
+ * of duplicate values. But perhaps we should keep the MCV in serialized form
+ * just like histograms.
+ */
+MCVList
+deserialize_mv_mcvlist(bytea *data)
+{
+ int i,
+ j;
+ Size expected_size;
+ MCVList mcvlist;
+ char *tmp;
+
+ int ndims,
+ nitems,
+ itemsize;
+ DimensionInfo *info = NULL;
+
+ uint16 *indexes = NULL;
+ Datum **values = NULL;
+
+ /* local allocation buffer (used only for deserialization) */
+ int bufflen;
+ char *buff;
+ char *ptr;
+
+ /* buffer used for the result */
+ int rbufflen;
+ char *rbuff;
+ char *rptr;
+
+ if (data == NULL)
+ return NULL;
+
+ /* we can't deserialize the MCV if there's not even a complete header */
+ expected_size = offsetof(MCVListData, items);
+
+ if (VARSIZE_ANY_EXHDR(data) < expected_size)
+ elog(ERROR, "invalid MCV Size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), offsetof(MCVListData, items));
+
+ /* read the MCV list header */
+ mcvlist = (MCVList) palloc0(sizeof(MCVListData));
+
+ /* initialize pointer to the data part (skip the varlena header) */
+ tmp = VARDATA_ANY(data);
+
+ /* get the header and perform further sanity checks */
+ memcpy(mcvlist, tmp, offsetof(MCVListData, items));
+ tmp += offsetof(MCVListData, items);
+
+ if (mcvlist->magic != MVSTAT_MCV_MAGIC)
+ elog(ERROR, "invalid MCV magic %d (expected %dd)",
+ mcvlist->magic, MVSTAT_MCV_MAGIC);
+
+ if (mcvlist->type != MVSTAT_MCV_TYPE_BASIC)
+ elog(ERROR, "invalid MCV type %d (expected %dd)",
+ mcvlist->type, MVSTAT_MCV_TYPE_BASIC);
+
+ nitems = mcvlist->nitems;
+ ndims = mcvlist->ndimensions;
+ itemsize = ITEM_SIZE(ndims);
+
+ Assert((nitems > 0) && (nitems <= MVSTAT_MCVLIST_MAX_ITEMS));
+ Assert((ndims >= 2) && (ndims <= MVSTATS_MAX_DIMENSIONS));
+
+ /*
+ * Check amount of data including DimensionInfo for all dimensions and
+ * also the serialized items (including uint16 indexes). Also, walk
+ * through the dimension information and add it to the sum.
+ */
+ expected_size += ndims * sizeof(DimensionInfo) +
+ (nitems * itemsize);
+
+ /* check that we have at least the DimensionInfo records */
+ if (VARSIZE_ANY_EXHDR(data) < expected_size)
+ elog(ERROR, "invalid MCV size %ld (expected %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ info = (DimensionInfo *) (tmp);
+ tmp += ndims * sizeof(DimensionInfo);
+
+ /* account for the value arrays */
+ for (i = 0; i < ndims; i++)
+ {
+ Assert(info[i].nvalues >= 0);
+ Assert(info[i].nbytes >= 0);
+
+ expected_size += info[i].nbytes;
+ }
+
+ if (VARSIZE_ANY_EXHDR(data) != expected_size)
+ elog(ERROR, "invalid MCV size %ld (expected %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ /* looks OK - not corrupted or something */
+
+ /*
+ * Allocate one large chunk of memory for the intermediate data, needed
+ * only for deserializing the MCV list (and allocate densely to minimize
+ * the palloc overhead).
+ *
+ * Let's see how much space we'll actually need, and also include space
+ * for the array with pointers.
+ */
+ bufflen = sizeof(Datum *) * ndims; /* space for pointers */
+
+ for (i = 0; i < ndims; i++)
+ /* for full-size byval types, we reuse the serialized value */
+ if (!(info[i].typbyval && info[i].typlen == sizeof(Datum)))
+ bufflen += (sizeof(Datum) * info[i].nvalues);
+
+ buff = palloc0(bufflen);
+ ptr = buff;
+
+ values = (Datum **) buff;
+ ptr += (sizeof(Datum *) * ndims);
+
+ /*
+ * XXX This uses pointers to the original data array (the types not passed
+ * by value), so when someone frees the memory, e.g. by doing something
+ * like this:
+ *
+ * bytea * data = ... fetch the data from catalog ... MCVList mcvlist =
+ * deserialize_mcv_list(data); pfree(data);
+ *
+ * then 'mcvlist' references the freed memory. Should copy the pieces.
+ */
+ for (i = 0; i < ndims; i++)
+ {
+ if (info[i].typbyval)
+ {
+ /* passed by value / Datum - simply reuse the array */
+ if (info[i].typlen == sizeof(Datum))
+ {
+ values[i] = (Datum *) tmp;
+ tmp += info[i].nbytes;
+ }
+ else
+ {
+ values[i] = (Datum *) ptr;
+ ptr += (sizeof(Datum) * info[i].nvalues);
+
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ memcpy(&values[i][j], tmp, info[i].typlen);
+ tmp += info[i].typlen;
+ }
+ }
+ }
+ else
+ {
+ /* all the other types need a chunk of the buffer */
+ values[i] = (Datum *) ptr;
+ ptr += (sizeof(Datum) * info[i].nvalues);
+
+ /* pased by reference, but fixed length (name, tid, ...) */
+ if (info[i].typlen > 0)
+ {
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ values[i][j] = PointerGetDatum(tmp);
+ tmp += info[i].typlen;
+ }
+ }
+ else if (info[i].typlen == -1)
+ {
+ /* varlena */
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ values[i][j] = PointerGetDatum(tmp);
+ tmp += VARSIZE_ANY(tmp);
+ }
+ }
+ else if (info[i].typlen == -2)
+ {
+ /* cstring */
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ values[i][j] = PointerGetDatum(tmp);
+ tmp += (strlen(tmp) + 1); /* don't forget the \0 */
+ }
+ }
+ }
+ }
+
+ /* we should have exhausted the buffer exactly */
+ Assert((ptr - buff) == bufflen);
+
+ /* allocate space for all the MCV items in a single piece */
+ rbufflen = (sizeof(MCVItem) + sizeof(MCVItemData) +
+ sizeof(Datum) * ndims + sizeof(bool) * ndims) * nitems;
+
+ rbuff = palloc0(rbufflen);
+ rptr = rbuff;
+
+ mcvlist->items = (MCVItem *) rbuff;
+ rptr += (sizeof(MCVItem) * nitems);
+
+ for (i = 0; i < nitems; i++)
+ {
+ MCVItem item = (MCVItem) rptr;
+
+ rptr += (sizeof(MCVItemData));
+
+ item->values = (Datum *) rptr;
+ rptr += (sizeof(Datum) * ndims);
+
+ item->isnull = (bool *) rptr;
+ rptr += (sizeof(bool) * ndims);
+
+ /* just point to the right place */
+ indexes = ITEM_INDEXES(tmp);
+
+ memcpy(item->isnull, ITEM_NULLS(tmp, ndims), sizeof(bool) * ndims);
+ memcpy(&item->frequency, ITEM_FREQUENCY(tmp, ndims), sizeof(double));
+
+#ifdef ASSERT_CHECKING
+ for (j = 0; j < ndims; j++)
+ Assert(indexes[j] <= UINT16_MAX);
+#endif
+
+ /* translate the values */
+ for (j = 0; j < ndims; j++)
+ if (!item->isnull[j])
+ item->values[j] = values[j][indexes[j]];
+
+ mcvlist->items[i] = item;
+
+ tmp += ITEM_SIZE(ndims);
+
+ Assert(tmp <= (char *) data + VARSIZE_ANY(data));
+ }
+
+ /* check that we processed all the data */
+ Assert(tmp == (char *) data + VARSIZE_ANY(data));
+
+ /* release the temporary buffer */
+ pfree(buff);
+
+ return mcvlist;
+}
+
+/*
+ * SRF with details about buckets of a histogram:
+ *
+ * - item ID (0...nitems)
+ * - values (string array)
+ * - nulls only (boolean array)
+ * - frequency (double precision)
+ *
+ * The input is the OID of the statistics, and there are no rows returned if
+ * the statistics contains no histogram.
+ */
+PG_FUNCTION_INFO_V1(pg_mv_mcv_items);
+
+Datum
+pg_mv_mcv_items(PG_FUNCTION_ARGS)
+{
+ FuncCallContext *funcctx;
+ int call_cntr;
+ int max_calls;
+ TupleDesc tupdesc;
+ AttInMetadata *attinmeta;
+
+ /* stuff done only on the first call of the function */
+ if (SRF_IS_FIRSTCALL())
+ {
+ MemoryContext oldcontext;
+ MCVList mcvlist;
+
+ /* create a function context for cross-call persistence */
+ funcctx = SRF_FIRSTCALL_INIT();
+
+ /* switch to memory context appropriate for multiple function calls */
+ oldcontext = MemoryContextSwitchTo(funcctx->multi_call_memory_ctx);
+
+ mcvlist = load_mv_mcvlist(PG_GETARG_OID(0));
+
+ funcctx->user_fctx = mcvlist;
+
+ /* total number of tuples to be returned */
+ funcctx->max_calls = 0;
+ if (funcctx->user_fctx != NULL)
+ funcctx->max_calls = mcvlist->nitems;
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("function returning record called in context "
+ "that cannot accept type record")));
+
+ /* build metadata needed later to produce tuples from raw C-strings */
+ attinmeta = TupleDescGetAttInMetadata(tupdesc);
+ funcctx->attinmeta = attinmeta;
+
+ MemoryContextSwitchTo(oldcontext);
+ }
+
+ /* stuff done on every call of the function */
+ funcctx = SRF_PERCALL_SETUP();
+
+ call_cntr = funcctx->call_cntr;
+ max_calls = funcctx->max_calls;
+ attinmeta = funcctx->attinmeta;
+
+ if (call_cntr < max_calls) /* do when there is more left to send */
+ {
+ char **values;
+ HeapTuple tuple;
+ Datum result;
+ int2vector *stakeys;
+ Oid relid;
+
+ char *buff = palloc0(1024);
+ char *format;
+
+ int i;
+
+ Oid *outfuncs;
+ FmgrInfo *fmgrinfo;
+
+ MCVList mcvlist;
+ MCVItem item;
+
+ mcvlist = (MCVList) funcctx->user_fctx;
+
+ Assert(call_cntr < mcvlist->nitems);
+
+ item = mcvlist->items[call_cntr];
+
+ stakeys = find_mv_attnums(PG_GETARG_OID(0), &relid);
+
+ /*
+ * Prepare a values array for building the returned tuple. This should
+ * be an array of C strings which will be processed later by the type
+ * input functions.
+ */
+ values = (char **) palloc(4 * sizeof(char *));
+
+ values[0] = (char *) palloc(64 * sizeof(char));
+
+ /* arrays */
+ values[1] = (char *) palloc0(1024 * sizeof(char));
+ values[2] = (char *) palloc0(1024 * sizeof(char));
+
+ /* frequency */
+ values[3] = (char *) palloc(64 * sizeof(char));
+
+ outfuncs = (Oid *) palloc0(sizeof(Oid) * mcvlist->ndimensions);
+ fmgrinfo = (FmgrInfo *) palloc0(sizeof(FmgrInfo) * mcvlist->ndimensions);
+
+ for (i = 0; i < mcvlist->ndimensions; i++)
+ {
+ bool isvarlena;
+
+ getTypeOutputInfo(get_atttype(relid, stakeys->values[i]),
+ &outfuncs[i], &isvarlena);
+
+ fmgr_info(outfuncs[i], &fmgrinfo[i]);
+ }
+
+ snprintf(values[0], 64, "%d", call_cntr); /* item ID */
+
+ for (i = 0; i < mcvlist->ndimensions; i++)
+ {
+ Datum val,
+ valout;
+
+ format = "%s, %s";
+ if (i == 0)
+ format = "{%s%s";
+ else if (i == mcvlist->ndimensions - 1)
+ format = "%s, %s}";
+
+ if (item->isnull[i])
+ valout = CStringGetDatum("NULL");
+ else
+ {
+ val = item->values[i];
+ valout = FunctionCall1(&fmgrinfo[i], val);
+ }
+
+ snprintf(buff, 1024, format, values[1], DatumGetPointer(valout));
+ strncpy(values[1], buff, 1023);
+ buff[0] = '\0';
+
+ snprintf(buff, 1024, format, values[2], item->isnull[i] ? "t" : "f");
+ strncpy(values[2], buff, 1023);
+ buff[0] = '\0';
+ }
+
+ snprintf(values[3], 64, "%f", item->frequency); /* frequency */
+
+ /* build a tuple */
+ tuple = BuildTupleFromCStrings(attinmeta, values);
+
+ /* make the tuple into a datum */
+ result = HeapTupleGetDatum(tuple);
+
+ /* clean up (this is not really necessary) */
+ pfree(values[0]);
+ pfree(values[1]);
+ pfree(values[2]);
+ pfree(values[3]);
+
+ pfree(values);
+
+ SRF_RETURN_NEXT(funcctx, result);
+ }
+ else /* do when there is no more left */
+ {
+ SRF_RETURN_DONE(funcctx);
+ }
+}
+
+/*
+ * pg_mcv_list_in - input routine for type PG_MCV_LIST.
+ *
+ * pg_mcv_list is real enough to be a table column, but it has no operations
+ * of its own, and disallows input too
+ *
+ * XXX This is inspired by what pg_node_tree does.
+ */
+Datum
+pg_mcv_list_in(PG_FUNCTION_ARGS)
+{
+ /*
+ * pg_node_list stores the data in binary form and parsing text input is
+ * not needed, so disallow this.
+ */
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_mcv_list")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+
+/*
+ * pg_mcv_list_out - output routine for type PG_MCV_LIST.
+ *
+ * MCV lists are serialized into a bytea value, so we simply call byteaout()
+ * to serialize the value into text. But it'd be nice to serialize that into
+ * a meaningful representation (e.g. for inspection by people).
+ *
+ * FIXME not implemented yet, returning dummy value
+ */
+Datum
+pg_mcv_list_out(PG_FUNCTION_ARGS)
+{
+ return byteaout(fcinfo);
+}
+
+/*
+ * pg_mcv_list_recv - binary input routine for type PG_MCV_LIST.
+ */
+Datum
+pg_mcv_list_recv(PG_FUNCTION_ARGS)
+{
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_mcv_list")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_mcv_list_send - binary output routine for type PG_MCV_LIST.
+ *
+ * XXX MCV lists are serialized into a bytea value, so let's just send that.
+ */
+Datum
+pg_mcv_list_send(PG_FUNCTION_ARGS)
+{
+ return byteasend(fcinfo);
+}
diff --git a/src/bin/psql/describe.c b/src/bin/psql/describe.c
index 1f1050b..e2220a9 100644
--- a/src/bin/psql/describe.c
+++ b/src/bin/psql/describe.c
@@ -2298,8 +2298,8 @@ describeOneTableDetails(const char *schemaname,
{
printfPQExpBuffer(&buf,
"SELECT oid, stanamespace::regnamespace AS nsp, staname, stakeys,\n"
- " ndist_enabled,\n"
- " ndist_built,\n"
+ " ndist_enabled, deps_enabled, mcv_enabled,\n"
+ " ndist_built, deps_built, mcv_built,\n"
" (SELECT string_agg(attname::text,', ')\n"
" FROM ((SELECT unnest(stakeys) AS attnum) s\n"
" JOIN pg_attribute a ON (starelid = a.attrelid and a.attnum = s.attnum))) AS attnums\n"
@@ -2317,6 +2317,8 @@ describeOneTableDetails(const char *schemaname,
printTableAddFooter(&cont, _("Statistics:"));
for (i = 0; i < tuples; i++)
{
+ bool first = true;
+
printfPQExpBuffer(&buf, " ");
/* statistics name (qualified with namespace) */
@@ -2326,10 +2328,22 @@ describeOneTableDetails(const char *schemaname,
/* options */
if (!strcmp(PQgetvalue(result, i, 4), "t"))
- appendPQExpBuffer(&buf, "(dependencies)");
+ {
+ appendPQExpBuffer(&buf, "(dependencies");
+ first = false;
+ }
+
+ if (!strcmp(PQgetvalue(result, i, 5), "t"))
+ {
+ if (!first)
+ appendPQExpBuffer(&buf, ", mcv");
+ else
+ appendPQExpBuffer(&buf, "(mcv");
+ first = false;
+ }
- appendPQExpBuffer(&buf, " ON (%s)",
- PQgetvalue(result, i, 6));
+ appendPQExpBuffer(&buf, ") ON (%s)",
+ PQgetvalue(result, i, 9));
printTableAddFooter(&cont, buf.data);
}
diff --git a/src/include/catalog/pg_cast.h b/src/include/catalog/pg_cast.h
index 1ecf1cb..80aa9f5 100644
--- a/src/include/catalog/pg_cast.h
+++ b/src/include/catalog/pg_cast.h
@@ -262,6 +262,11 @@ DATA(insert ( 3343 25 0 i i ));
DATA(insert ( 3353 17 0 i b ));
DATA(insert ( 3353 25 0 i i ));
+/* pg_mcv_list can be coerced to, but not from, bytea and text */
+DATA(insert ( 441 17 0 i b ));
+DATA(insert ( 441 25 0 i i ));
+
+
/*
* Datetime category
*/
diff --git a/src/include/catalog/pg_mv_statistic.h b/src/include/catalog/pg_mv_statistic.h
index e119cb7..34049d6 100644
--- a/src/include/catalog/pg_mv_statistic.h
+++ b/src/include/catalog/pg_mv_statistic.h
@@ -39,10 +39,12 @@ CATALOG(pg_mv_statistic,3381)
/* statistics requested to build */
bool ndist_enabled; /* build ndist coefficient? */
bool deps_enabled; /* analyze dependencies? */
+ bool mcv_enabled; /* build MCV list? */
/* statistics that are available (if requested) */
bool ndist_built; /* ndistinct coeff built */
bool deps_built; /* dependencies were built */
+ bool mcv_built; /* MCV list was built */
/*
* variable-length fields start here, but we allow direct access to
@@ -53,6 +55,7 @@ CATALOG(pg_mv_statistic,3381)
#ifdef CATALOG_VARLEN
pg_ndistinct standist; /* ndistinct coeff (serialized) */
pg_dependencies stadeps; /* dependencies (serialized) */
+ pg_mcv_list stamcv; /* MCV list (serialized) */
#endif
} FormData_pg_mv_statistic;
@@ -68,17 +71,20 @@ typedef FormData_pg_mv_statistic *Form_pg_mv_statistic;
* compiler constants for pg_mv_statistic
* ----------------
*/
-#define Natts_pg_mv_statistic 11
+#define Natts_pg_mv_statistic 14
#define Anum_pg_mv_statistic_starelid 1
#define Anum_pg_mv_statistic_staname 2
#define Anum_pg_mv_statistic_stanamespace 3
#define Anum_pg_mv_statistic_staowner 4
#define Anum_pg_mv_statistic_ndist_enabled 5
#define Anum_pg_mv_statistic_deps_enabled 6
-#define Anum_pg_mv_statistic_ndist_built 7
-#define Anum_pg_mv_statistic_deps_built 8
-#define Anum_pg_mv_statistic_stakeys 9
-#define Anum_pg_mv_statistic_standist 10
-#define Anum_pg_mv_statistic_stadeps 11
+#define Anum_pg_mv_statistic_mcv_enabled 7
+#define Anum_pg_mv_statistic_ndist_built 8
+#define Anum_pg_mv_statistic_deps_built 9
+#define Anum_pg_mv_statistic_mcv_built 10
+#define Anum_pg_mv_statistic_stakeys 11
+#define Anum_pg_mv_statistic_standist 12
+#define Anum_pg_mv_statistic_stadeps 13
+#define Anum_pg_mv_statistic_stamcv 14
#endif /* PG_MV_STATISTIC_H */
diff --git a/src/include/catalog/pg_proc.h b/src/include/catalog/pg_proc.h
index 0c78361..6be685c 100644
--- a/src/include/catalog/pg_proc.h
+++ b/src/include/catalog/pg_proc.h
@@ -2722,6 +2722,11 @@ DESCR("current user privilege on any column by rel name");
DATA(insert OID = 3029 ( has_any_column_privilege PGNSP PGUID 12 10 0 0 0 f f f f t f s s 2 0 16 "26 25" _null_ _null_ _null_ _null_ _null_ has_any_column_privilege_id _null_ _null_ _null_ ));
DESCR("current user privilege on any column by rel oid");
+DATA(insert OID = 3376 ( pg_mv_stats_mcvlist_info PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 25 "441" _null_ _null_ _null_ _null_ _null_ pg_mv_stats_mcvlist_info _null_ _null_ _null_ ));
+DESCR("multi-variate statistics: MCV list info");
+DATA(insert OID = 3373 ( pg_mv_mcv_items PGNSP PGUID 12 1 1000 0 0 f f f f t t i s 1 0 2249 "26" "{26,23,1009,1000,701}" "{i,o,o,o,o}" "{oid,index,values,nulls,frequency}" _null_ _null_ pg_mv_mcv_items _null_ _null_ _null_ ));
+DESCR("details about MCV list items");
+
DATA(insert OID = 3344 ( pg_ndistinct_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 3343 "2275" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_in _null_ _null_ _null_ ));
DESCR("I/O");
DATA(insert OID = 3345 ( pg_ndistinct_out PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 2275 "3343" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_out _null_ _null_ _null_ ));
@@ -2740,6 +2745,15 @@ DESCR("I/O");
DATA(insert OID = 3357 ( pg_dependencies_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "3353" _null_ _null_ _null_ _null_ _null_ pg_dependencies_send _null_ _null_ _null_ ));
DESCR("I/O");
+DATA(insert OID = 442 ( pg_mcv_list_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 441 "2275" _null_ _null_ _null_ _null_ _null_ pg_mcv_list_in _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 443 ( pg_mcv_list_out PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 2275 "441" _null_ _null_ _null_ _null_ _null_ pg_mcv_list_out _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 444 ( pg_mcv_list_recv PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 441 "2281" _null_ _null_ _null_ _null_ _null_ pg_mcv_list_recv _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 445 ( pg_mcv_list_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "441" _null_ _null_ _null_ _null_ _null_ pg_mcv_list_send _null_ _null_ _null_ ));
+DESCR("I/O");
+
DATA(insert OID = 1928 ( pg_stat_get_numscans PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_numscans _null_ _null_ _null_ ));
DESCR("statistics: number of scans done for table/index");
DATA(insert OID = 1929 ( pg_stat_get_tuples_returned PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_tuples_returned _null_ _null_ _null_ ));
diff --git a/src/include/catalog/pg_type.h b/src/include/catalog/pg_type.h
index 250952b..4621703 100644
--- a/src/include/catalog/pg_type.h
+++ b/src/include/catalog/pg_type.h
@@ -372,6 +372,10 @@ DATA(insert OID = 3353 ( pg_dependencies PGNSP PGUID -1 f b S f t \054 0 0 0 pg
DESCR("multivariate histogram");
#define PGDEPENDENCIESOID 3353
+DATA(insert OID = 441 ( pg_mcv_list PGNSP PGUID -1 f b S f t \054 0 0 0 pg_mcv_list_in pg_mcv_list_out pg_mcv_list_recv pg_mcv_list_send - - - i x f 0 -1 0 100 _null_ _null_ _null_ ));
+DESCR("multivariate MCV list");
+#define PGMCVLISTOID 441
+
DATA(insert OID = 32 ( pg_ddl_command PGNSP PGUID SIZEOF_POINTER t p P f t \054 0 0 0 pg_ddl_command_in pg_ddl_command_out pg_ddl_command_recv pg_ddl_command_send - - - ALIGNOF_POINTER p f 0 -1 0 0 _null_ _null_ _null_ ));
DESCR("internal type for passing CollectedCommand");
#define PGDDLCOMMANDOID 32
diff --git a/src/include/nodes/relation.h b/src/include/nodes/relation.h
index 8b7db72..b37835c 100644
--- a/src/include/nodes/relation.h
+++ b/src/include/nodes/relation.h
@@ -674,12 +674,14 @@ typedef struct MVStatisticInfo
RelOptInfo *rel; /* back-link to index's table */
/* enabled statistics */
- bool deps_enabled; /* functional dependencies enabled */
bool ndist_enabled; /* ndistinct coefficient enabled */
+ bool deps_enabled; /* functional dependencies enabled */
+ bool mcv_enabled; /* MCV list enabled */
/* built/available statistics */
- bool deps_built; /* functional dependencies built */
bool ndist_built; /* ndistinct coefficient built */
+ bool deps_built; /* functional dependencies built */
+ bool mcv_built; /* MCV list built */
/* columns in the statistics (attnums) */
int2vector *stakeys; /* attnums of the columns covered */
diff --git a/src/include/utils/builtins.h b/src/include/utils/builtins.h
index 4cb09e7..fe95bad 100644
--- a/src/include/utils/builtins.h
+++ b/src/include/utils/builtins.h
@@ -621,6 +621,10 @@ extern Datum pg_dependencies_in(PG_FUNCTION_ARGS);
extern Datum pg_dependencies_out(PG_FUNCTION_ARGS);
extern Datum pg_dependencies_recv(PG_FUNCTION_ARGS);
extern Datum pg_dependencies_send(PG_FUNCTION_ARGS);
+extern Datum pg_mcv_list_in(PG_FUNCTION_ARGS);
+extern Datum pg_mcv_list_out(PG_FUNCTION_ARGS);
+extern Datum pg_mcv_list_recv(PG_FUNCTION_ARGS);
+extern Datum pg_mcv_list_send(PG_FUNCTION_ARGS);
/* regexp.c */
extern Datum nameregexeq(PG_FUNCTION_ARGS);
diff --git a/src/include/utils/mvstats.h b/src/include/utils/mvstats.h
index 3ad4e48..b17fcba 100644
--- a/src/include/utils/mvstats.h
+++ b/src/include/utils/mvstats.h
@@ -17,6 +17,14 @@
#include "fmgr.h"
#include "commands/vacuum.h"
+/*
+ * Degree of how much MCV item matches a clause.
+ * This is then considered when computing the selectivity.
+ */
+#define MVSTATS_MATCH_NONE 0 /* no match at all */
+#define MVSTATS_MATCH_PARTIAL 1 /* partial match */
+#define MVSTATS_MATCH_FULL 2 /* full match */
+
#define MVSTATS_MAX_DIMENSIONS 8 /* max number of attributes */
#define MVSTAT_NDISTINCT_MAGIC 0xA352BFA4 /* marks serialized bytea */
@@ -65,6 +73,42 @@ typedef struct MVDependenciesData
typedef MVDependenciesData *MVDependencies;
+
+/* used to flag stats serialized to bytea */
+#define MVSTAT_MCV_MAGIC 0xE1A651C2 /* marks serialized bytea */
+#define MVSTAT_MCV_TYPE_BASIC 1 /* basic MCV list type */
+
+/* max items in MCV list (mostly arbitrary number */
+#define MVSTAT_MCVLIST_MAX_ITEMS 8192
+
+/*
+ * Multivariate MCV (most-common value) lists
+ *
+ * A straight-forward extension of MCV items - i.e. a list (array) of
+ * combinations of attribute values, together with a frequency and
+ * null flags.
+ */
+typedef struct MCVItemData
+{
+ double frequency; /* frequency of this combination */
+ bool *isnull; /* lags of NULL values (up to 32 columns) */
+ Datum *values; /* variable-length (ndimensions) */
+} MCVItemData;
+
+typedef MCVItemData *MCVItem;
+
+/* multivariate MCV list - essentally an array of MCV items */
+typedef struct MCVListData
+{
+ uint32 magic; /* magic constant marker */
+ uint32 type; /* type of MCV list (BASIC) */
+ uint32 ndimensions; /* number of dimensions */
+ uint32 nitems; /* number of MCV items in the array */
+ MCVItem *items; /* array of MCV items */
+} MCVListData;
+
+typedef MCVListData *MCVList;
+
bool dependency_implies_attribute(MVDependency dependency, AttrNumber attnum,
int16 *attmap);
bool dependency_is_fully_matched(MVDependency dependency, Bitmapset *attnums,
@@ -72,13 +116,30 @@ bool dependency_is_fully_matched(MVDependency dependency, Bitmapset *attnums,
MVNDistinct load_mv_ndistinct(Oid mvoid);
MVDependencies load_mv_dependencies(Oid mvoid);
+MCVList load_mv_mcvlist(Oid mvoid);
bytea *serialize_mv_ndistinct(MVNDistinct ndistinct);
bytea *serialize_mv_dependencies(MVDependencies dependencies);
+bytea *serialize_mv_mcvlist(MCVList mcvlist, int2vector *attrs,
+ VacAttrStats **stats);
/* deserialization of stats (serialization is private to analyze) */
MVNDistinct deserialize_mv_ndistinct(bytea *data);
MVDependencies deserialize_mv_dependencies(bytea *data);
+MCVList deserialize_mv_mcvlist(bytea *data);
+
+/*
+ * Returns index of the attribute number within the vector (i.e. a
+ * dimension within the stats).
+ */
+int mv_get_index(AttrNumber varattno, int2vector *stakeys);
+
+int2vector *find_mv_attnums(Oid mvoid, Oid *relid);
+
+/* functions for inspecting the statistics */
+extern Datum pg_mv_stats_mcvlist_info(PG_FUNCTION_ARGS);
+extern Datum pg_mv_mcvlist_items(PG_FUNCTION_ARGS);
+
MVNDistinct build_mv_ndistinct(double totalrows, int numrows, HeapTuple *rows,
int2vector *attrs, VacAttrStats **stats);
@@ -87,11 +148,15 @@ MVDependencies build_mv_dependencies(int numrows, HeapTuple *rows,
int2vector *attrs,
VacAttrStats **stats);
+MCVList build_mv_mcvlist(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int *numrows_filtered);
+
void build_mv_stats(Relation onerel, double totalrows,
int numrows, HeapTuple *rows,
int natts, VacAttrStats **vacattrstats);
-void update_mv_stats(Oid relid, MVNDistinct ndistinct, MVDependencies dependencies,
+void update_mv_stats(Oid relid,
+ MVNDistinct ndistinct, MVDependencies dependencies, MCVList mcvlist,
int2vector *attrs, VacAttrStats **stats);
#endif
diff --git a/src/test/regress/expected/mv_mcv.out b/src/test/regress/expected/mv_mcv.out
new file mode 100644
index 0000000..d8ba619
--- /dev/null
+++ b/src/test/regress/expected/mv_mcv.out
@@ -0,0 +1,198 @@
+-- data type passed by value
+CREATE TABLE mcv_list (
+ a INT,
+ b INT,
+ c INT
+);
+-- unknown column
+CREATE STATISTICS s4 WITH (mcv) ON (unknown_column) FROM mcv_list;
+ERROR: column "unknown_column" referenced in statistics does not exist
+-- single column
+CREATE STATISTICS s4 WITH (mcv) ON (a) FROM mcv_list;
+ERROR: statistics require at least 2 columns
+-- single column, duplicated
+CREATE STATISTICS s4 WITH (mcv) ON (a, a) FROM mcv_list;
+ERROR: duplicate column name in statistics definition
+-- two columns, one duplicated
+CREATE STATISTICS s4 WITH (mcv) ON (a, a, b) FROM mcv_list;
+ERROR: duplicate column name in statistics definition
+-- unknown option
+CREATE STATISTICS s4 WITH (unknown_option) ON (a, b, c) FROM mcv_list;
+ERROR: unrecognized STATISTICS option "unknown_option"
+-- correct command
+CREATE STATISTICS s4 WITH (mcv) ON (a, b, c) FROM mcv_list;
+-- random data
+INSERT INTO mcv_list
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | f |
+(1 row)
+
+TRUNCATE mcv_list;
+-- a => b, a => c, b => c
+INSERT INTO mcv_list
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=1000
+(1 row)
+
+TRUNCATE mcv_list;
+-- a => b, a => c
+INSERT INTO mcv_list
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=1000
+(1 row)
+
+TRUNCATE mcv_list;
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mcv_list
+ SELECT i/100, i/200, i/400 FROM generate_series(1,10000) s(i);
+CREATE INDEX mcv_idx ON mcv_list (a, b);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=100
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a = 10 AND b = 5;
+ QUERY PLAN
+--------------------------------------------
+ Bitmap Heap Scan on mcv_list
+ Recheck Cond: ((a = 10) AND (b = 5))
+ -> Bitmap Index Scan on mcv_idx
+ Index Cond: ((a = 10) AND (b = 5))
+(4 rows)
+
+DROP TABLE mcv_list;
+-- varlena type (text)
+CREATE TABLE mcv_list (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+CREATE STATISTICS s5 WITH (mcv) ON (a, b, c) FROM mcv_list;
+-- random data
+INSERT INTO mcv_list
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | f |
+(1 row)
+
+TRUNCATE mcv_list;
+-- a => b, a => c, b => c
+INSERT INTO mcv_list
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=1000
+(1 row)
+
+TRUNCATE mcv_list;
+-- a => b, a => c
+INSERT INTO mcv_list
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=1000
+(1 row)
+
+TRUNCATE mcv_list;
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mcv_list
+ SELECT i/100, i/200, i/400 FROM generate_series(1,10000) s(i);
+CREATE INDEX mcv_idx ON mcv_list (a, b);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=100
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a = '10' AND b = '5';
+ QUERY PLAN
+------------------------------------------------------------
+ Bitmap Heap Scan on mcv_list
+ Recheck Cond: ((a = '10'::text) AND (b = '5'::text))
+ -> Bitmap Index Scan on mcv_idx
+ Index Cond: ((a = '10'::text) AND (b = '5'::text))
+(4 rows)
+
+TRUNCATE mcv_list;
+-- check explain (expect bitmap index scan, not plain index scan) with NULLs
+INSERT INTO mcv_list
+ SELECT
+ (CASE WHEN i/100 = 0 THEN NULL ELSE i/100 END),
+ (CASE WHEN i/200 = 0 THEN NULL ELSE i/200 END),
+ (CASE WHEN i/400 = 0 THEN NULL ELSE i/400 END)
+ FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=100
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a IS NULL AND b IS NULL;
+ QUERY PLAN
+---------------------------------------------------
+ Bitmap Heap Scan on mcv_list
+ Recheck Cond: ((a IS NULL) AND (b IS NULL))
+ -> Bitmap Index Scan on mcv_idx
+ Index Cond: ((a IS NULL) AND (b IS NULL))
+(4 rows)
+
+DROP TABLE mcv_list;
+-- NULL values (mix of int and text columns)
+CREATE TABLE mcv_list (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+CREATE STATISTICS s6 WITH (mcv) ON (a, b, c, d) FROM mcv_list;
+INSERT INTO mcv_list
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=1200
+(1 row)
+
+DROP TABLE mcv_list;
diff --git a/src/test/regress/expected/opr_sanity.out b/src/test/regress/expected/opr_sanity.out
index db1cf8a..9969c10 100644
--- a/src/test/regress/expected/opr_sanity.out
+++ b/src/test/regress/expected/opr_sanity.out
@@ -819,11 +819,12 @@ WHERE c.castmethod = 'b' AND
pg_node_tree | text | 0 | i
pg_ndistinct | bytea | 0 | i
pg_dependencies | bytea | 0 | i
+ pg_mcv_list | bytea | 0 | i
cidr | inet | 0 | i
xml | text | 0 | a
xml | character varying | 0 | a
xml | character | 0 | a
-(9 rows)
+(10 rows)
-- **************** pg_conversion ****************
-- Look for illegal values in pg_conversion fields.
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index ea3f51d..d14d864 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1381,7 +1381,9 @@ pg_mv_stats| SELECT n.nspname AS schemaname,
s.staname,
s.stakeys AS attnums,
length((s.standist)::bytea) AS ndistbytes,
- length((s.stadeps)::bytea) AS depsbytes
+ length((s.stadeps)::bytea) AS depsbytes,
+ length((s.stamcv)::bytea) AS mcvbytes,
+ pg_mv_stats_mcvlist_info(s.stamcv) AS mcvinfo
FROM ((pg_mv_statistic s
JOIN pg_class c ON ((c.oid = s.starelid)))
LEFT JOIN pg_namespace n ON ((n.oid = c.relnamespace)));
diff --git a/src/test/regress/expected/type_sanity.out b/src/test/regress/expected/type_sanity.out
index 8b849b9..b810e71 100644
--- a/src/test/regress/expected/type_sanity.out
+++ b/src/test/regress/expected/type_sanity.out
@@ -72,9 +72,10 @@ WHERE p1.typtype not in ('c','d','p') AND p1.typname NOT LIKE E'\\_%'
194 | pg_node_tree
3343 | pg_ndistinct
3353 | pg_dependencies
+ 441 | pg_mcv_list
210 | smgr
705 | unknown
-(5 rows)
+(6 rows)
-- Make sure typarray points to a varlena array type of our own base
SELECT p1.oid, p1.typname as basetype, p2.typname as arraytype,
diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule
index ecb6f04..d6c3cc0 100644
--- a/src/test/regress/parallel_schedule
+++ b/src/test/regress/parallel_schedule
@@ -115,4 +115,4 @@ test: event_trigger
test: stats
# run tests of multivariate stats
-test: mv_ndistinct mv_dependencies
+test: mv_ndistinct mv_dependencies mv_mcv
diff --git a/src/test/regress/serial_schedule b/src/test/regress/serial_schedule
index fa0e993..2394d74 100644
--- a/src/test/regress/serial_schedule
+++ b/src/test/regress/serial_schedule
@@ -171,3 +171,4 @@ test: event_trigger
test: stats
test: mv_ndistinct
test: mv_dependencies
+test: mv_mcv
diff --git a/src/test/regress/sql/mv_mcv.sql b/src/test/regress/sql/mv_mcv.sql
new file mode 100644
index 0000000..693288f
--- /dev/null
+++ b/src/test/regress/sql/mv_mcv.sql
@@ -0,0 +1,169 @@
+-- data type passed by value
+CREATE TABLE mcv_list (
+ a INT,
+ b INT,
+ c INT
+);
+
+-- unknown column
+CREATE STATISTICS s4 WITH (mcv) ON (unknown_column) FROM mcv_list;
+
+-- single column
+CREATE STATISTICS s4 WITH (mcv) ON (a) FROM mcv_list;
+
+-- single column, duplicated
+CREATE STATISTICS s4 WITH (mcv) ON (a, a) FROM mcv_list;
+
+-- two columns, one duplicated
+CREATE STATISTICS s4 WITH (mcv) ON (a, a, b) FROM mcv_list;
+
+-- unknown option
+CREATE STATISTICS s4 WITH (unknown_option) ON (a, b, c) FROM mcv_list;
+
+-- correct command
+CREATE STATISTICS s4 WITH (mcv) ON (a, b, c) FROM mcv_list;
+
+-- random data
+INSERT INTO mcv_list
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- a => b, a => c, b => c
+INSERT INTO mcv_list
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- a => b, a => c
+INSERT INTO mcv_list
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mcv_list
+ SELECT i/100, i/200, i/400 FROM generate_series(1,10000) s(i);
+CREATE INDEX mcv_idx ON mcv_list (a, b);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a = 10 AND b = 5;
+
+DROP TABLE mcv_list;
+
+-- varlena type (text)
+CREATE TABLE mcv_list (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+
+CREATE STATISTICS s5 WITH (mcv) ON (a, b, c) FROM mcv_list;
+
+-- random data
+INSERT INTO mcv_list
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- a => b, a => c, b => c
+INSERT INTO mcv_list
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- a => b, a => c
+INSERT INTO mcv_list
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mcv_list
+ SELECT i/100, i/200, i/400 FROM generate_series(1,10000) s(i);
+CREATE INDEX mcv_idx ON mcv_list (a, b);
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a = '10' AND b = '5';
+
+TRUNCATE mcv_list;
+
+-- check explain (expect bitmap index scan, not plain index scan) with NULLs
+INSERT INTO mcv_list
+ SELECT
+ (CASE WHEN i/100 = 0 THEN NULL ELSE i/100 END),
+ (CASE WHEN i/200 = 0 THEN NULL ELSE i/200 END),
+ (CASE WHEN i/400 = 0 THEN NULL ELSE i/400 END)
+ FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a IS NULL AND b IS NULL;
+
+DROP TABLE mcv_list;
+
+-- NULL values (mix of int and text columns)
+CREATE TABLE mcv_list (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+
+CREATE STATISTICS s6 WITH (mcv) ON (a, b, c, d) FROM mcv_list;
+
+INSERT INTO mcv_list
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+DROP TABLE mcv_list;
--
2.5.5
[binary/octet-stream] 0006-PATCH-multivariate-histograms-v22.patch (149.8K, ../../696aa95c-2411-9b2b-f36e-65b66bf47c88@2ndquadrant.com/7-0006-PATCH-multivariate-histograms-v22.patch)
download | inline diff:
From 1330a200b59b7f0f77374e8414f615a818e7ab6f Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:38:35 +0200
Subject: [PATCH 6/9] PATCH: multivariate histograms
- extends the pg_mv_statistic catalog (add 'hist' fields)
- building the histograms during ANALYZE
- simple estimation while planning the queries
- pg_histogram data type (varlena-based)
Includes regression tests mostly equal to those for functional
dependencies / MCV lists.
A new varlena-based data type for storing serialized histograms.
---
doc/src/sgml/catalogs.sgml | 30 +
doc/src/sgml/planstats.sgml | 125 ++
doc/src/sgml/ref/create_statistics.sgml | 35 +
src/backend/catalog/system_views.sql | 4 +-
src/backend/commands/statscmds.c | 11 +-
src/backend/nodes/outfuncs.c | 2 +
src/backend/optimizer/path/clausesel.c | 606 +++++++-
src/backend/optimizer/util/plancat.c | 4 +-
src/backend/utils/mvstats/Makefile | 2 +-
src/backend/utils/mvstats/README.histogram | 299 ++++
src/backend/utils/mvstats/README.stats | 2 +
src/backend/utils/mvstats/common.c | 31 +-
src/backend/utils/mvstats/common.h | 8 +-
src/backend/utils/mvstats/histogram.c | 2123 ++++++++++++++++++++++++++++
src/bin/psql/describe.c | 15 +-
src/include/catalog/pg_cast.h | 3 +
src/include/catalog/pg_mv_statistic.h | 22 +-
src/include/catalog/pg_proc.h | 13 +
src/include/catalog/pg_type.h | 4 +
src/include/nodes/relation.h | 2 +
src/include/utils/builtins.h | 4 +
src/include/utils/mvstats.h | 128 +-
src/test/regress/expected/mv_histogram.out | 198 +++
src/test/regress/expected/opr_sanity.out | 3 +-
src/test/regress/expected/rules.out | 4 +-
src/test/regress/expected/type_sanity.out | 3 +-
src/test/regress/parallel_schedule | 2 +-
src/test/regress/serial_schedule | 1 +
src/test/regress/sql/mv_histogram.sql | 167 +++
29 files changed, 3802 insertions(+), 49 deletions(-)
create mode 100644 src/backend/utils/mvstats/README.histogram
create mode 100644 src/backend/utils/mvstats/histogram.c
create mode 100644 src/test/regress/expected/mv_histogram.out
create mode 100644 src/test/regress/sql/mv_histogram.sql
diff --git a/doc/src/sgml/catalogs.sgml b/doc/src/sgml/catalogs.sgml
index b82ca13..f422a92 100644
--- a/doc/src/sgml/catalogs.sgml
+++ b/doc/src/sgml/catalogs.sgml
@@ -4292,6 +4292,17 @@
</row>
<row>
+ <entry><structfield>hist_enabled</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, histogram will be computed for the combination of columns,
+ covered by the statistics. This does not mean the histogram is already
+ computed, though.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>ndist_built</structfield></entry>
<entry><type>bool</type></entry>
<entry></entry>
@@ -4322,6 +4333,16 @@
</row>
<row>
+ <entry><structfield>hist_built</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, histogram is already computed and available for use during query
+ estimation.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>stakeys</structfield></entry>
<entry><type>int2vector</type></entry>
<entry><literal><link linkend="catalog-pg-attribute"><structname>pg_attribute</structname></link>.attnum</literal></entry>
@@ -4359,6 +4380,15 @@
</entry>
</row>
+ <row>
+ <entry><structfield>stahist</structfield></entry>
+ <entry><type>pg_histogram</type></entry>
+ <entry></entry>
+ <entry>
+ Histogram, serialized as <structname>pg_histogram</> type.
+ </entry>
+ </row>
+
</tbody>
</tgroup>
</table>
diff --git a/doc/src/sgml/planstats.sgml b/doc/src/sgml/planstats.sgml
index 1ee4293..3c80966 100644
--- a/doc/src/sgml/planstats.sgml
+++ b/doc/src/sgml/planstats.sgml
@@ -914,6 +914,131 @@ EXPLAIN ANALYZE SELECT * FROM t WHERE a <= 49 AND b > 49;
</sect2>
+ <sect2 id="mv-histograms">
+ <title>Histograms</title>
+
+ <para>
+ <acronym>MCV</> lists, introduced in the previous section, work very well
+ for low-cardinality columns (i.e. columns with only very few distinct
+ values), and for columns with a few very frequent values (and possibly
+ many rare ones). Histograms, a generalization of per-column histograms
+ briefly described in <xref linkend="row-estimation-examples">, are meant
+ to address the other cases, i.e. high-cardinality columns, particularly
+ when there are no frequent values.
+ </para>
+
+ <para>
+ Although the example data we've used so far is not a very good match, we
+ can try creating a histogram instead of the <acronym>MCV</> list. With the
+ histogram in place, you may get a plan like this:
+
+<programlisting>
+DROP STATISTICS s2;
+CREATE STATISTICS s3 ON t (a,b) WITH (histogram);
+ANALYZE t;
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
+ QUERY PLAN
+-------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=100 width=8) (actual time=0.035..2.967 rows=100 loops=1)
+ Filter: ((a = 1) AND (b = 1))
+ Rows Removed by Filter: 9900
+ Planning time: 0.227 ms
+ Execution time: 3.189 ms
+(5 rows)
+</programlisting>
+
+ Which seems quite accurate, however for other combinations of values the
+ results may be much worse, as illustrated by the following query
+
+<programlisting>
+ QUERY PLAN
+-----------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=100 width=8) (actual time=2.771..2.771 rows=0 loops=1)
+ Filter: ((a = 1) AND (b = 10))
+ Rows Removed by Filter: 10000
+ Planning time: 0.179 ms
+ Execution time: 2.812 ms
+(5 rows)
+</programlisting>
+
+ This is due to histograms tracking ranges of values, not individual values.
+ That means it's only possible say whether a bucket may contain items
+ matching the conditions, but it's unclear how many such tuples there
+ actually are in the bucket. Moreover, for larger tables only a small subset
+ of rows gets sampled by <command>ANALYZE</>, causing small variations in
+ the shape of buckets.
+ </para>
+
+ <para>
+ To inspect details of the histogram, we can look into the
+ <structname>pg_mv_stats</> view
+
+<programlisting>
+SELECT tablename, staname, attnums, histbytes, histinfo
+ FROM pg_mv_stats WHERE staname = 's3';
+ tablename | staname | attnums | histbytes | histinfo
+-----------+---------+---------+-----------+-------------
+ t | s3 | 1 2 | 1928 | nbuckets=64
+(1 row)
+</programlisting>
+
+ This shows the histogram has 64 buckets, but as we know there are 100
+ distinct combinations of values in the two columns. This means there are
+ buckets containing multiple combinations, causing the inaccuracy.
+ </para>
+
+ <para>
+ Similarly to <acronym>MCV</> lists, we can inspect histogram contents
+ using a function called <function>pg_mv_histogram_buckets</>.
+
+<programlisting>
+test=# SELECT * FROM pg_mv_histogram_buckets((SELECT oid FROM pg_mv_statistic WHERE staname = 's3'), 0);
+ index | minvals | maxvals | nullsonly | mininclusive | maxinclusive | frequency | density | bucket_volume
+-------+---------+---------+-----------+--------------+--------------+-----------+----------+---------------
+ 0 | {0,0} | {3,1} | {f,f} | {t,t} | {f,f} | 0.01 | 1.68 | 0.005952
+ 1 | {50,0} | {51,3} | {f,f} | {t,t} | {f,f} | 0.01 | 1.12 | 0.008929
+ 2 | {0,25} | {26,31} | {f,f} | {t,t} | {f,f} | 0.01 | 0.28 | 0.035714
+...
+ 61 | {60,0} | {99,12} | {f,f} | {t,t} | {t,f} | 0.02 | 0.124444 | 0.160714
+ 62 | {34,35} | {37,49} | {f,f} | {t,t} | {t,t} | 0.02 | 0.96 | 0.020833
+ 63 | {84,35} | {87,49} | {f,f} | {t,t} | {t,t} | 0.02 | 0.96 | 0.020833
+(64 rows)
+</programlisting>
+
+ Which confirms there are 64 buckets, with frequencies ranging between 1%
+ and 2%. The <structfield>minvals</> and <structfield>maxvals</> show the
+ bucket boundaries, <structfield>nullsonly</> shows which columns contain
+ only null values (in the given bucket).
+ </para>
+
+ <para>
+ Similarly to <acronym>MCV</> lists, the planner applies all conditions to
+ the buckets, and sums the frequencies of the matching ones. For details,
+ see <function>clauselist_mv_selectivity_histogram</> function in
+ <filename>clausesel.c</>.
+ </para>
+
+ <para>
+ It's also possible to build <acronym>MCV</> lists and a histogram, in which
+ case <command>ANALYZE</> will build a <acronym>MCV</> lists with the most
+ frequent values, and a histogram on the remaining part of the sample.
+
+<programlisting>
+DROP STATISTICS s3;
+CREATE STATISTICS s4 ON t (a,b) WITH (mcv, histogram);
+</programlisting>
+
+ In this case the <acronym>MCV</> list and histogram are treated as a single
+ composed statistics.
+ </para>
+
+ <para>
+ For additional information about multivariate histograms, see
+ <filename>src/backend/utils/mvstats/README.histogram</>.
+ </para>
+
+ </sect2>
+
</sect1>
</chapter>
diff --git a/doc/src/sgml/ref/create_statistics.sgml b/doc/src/sgml/ref/create_statistics.sgml
index fc97b16..5b1f1ca 100644
--- a/doc/src/sgml/ref/create_statistics.sgml
+++ b/doc/src/sgml/ref/create_statistics.sgml
@@ -125,6 +125,15 @@ CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_na
</varlistentry>
<varlistentry>
+ <term><literal>histogram</> (<type>boolean</>)</term>
+ <listitem>
+ <para>
+ Enables histogram for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
<term><literal>mcv</> (<type>boolean</>)</term>
<listitem>
<para>
@@ -202,6 +211,32 @@ EXPLAIN ANALYZE SELECT * FROM t2 WHERE (a = 1) AND (b = 2);
</programlisting>
</para>
+ <para>
+ Create table <structname>t3</> with two strongly correlated columns, and
+ a histogram on those two columns:
+
+<programlisting>
+CREATE TABLE t3 (
+ a float,
+ b float
+);
+
+INSERT INTO t3 SELECT mod(i,1000), mod(i,1000) + 50 * (r - 0.5) FROM (
+ SELECT i, random() r FROM generate_series(1,1000000) s(i)
+ ) foo;
+
+CREATE STATISTICS s3 WITH (histogram) ON (a, b) FROM t3;
+
+ANALYZE t3;
+
+-- small overlap
+EXPLAIN ANALYZE SELECT * FROM t3 WHERE (a < 500) AND (b > 500);
+
+-- no overlap
+EXPLAIN ANALYZE SELECT * FROM t3 WHERE (a < 400) AND (b > 600);
+</programlisting>
+ </para>
+
</refsect1>
<refsect1>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 7ac3c4c..7c3532c 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -190,7 +190,9 @@ CREATE VIEW pg_mv_stats AS
length(s.standist::bytea) AS ndistbytes,
length(S.stadeps::bytea) AS depsbytes,
length(S.stamcv::bytea) AS mcvbytes,
- pg_mv_stats_mcvlist_info(S.stamcv) AS mcvinfo
+ pg_mv_stats_mcvlist_info(S.stamcv) AS mcvinfo,
+ length(S.stahist::bytea) AS histbytes,
+ pg_mv_stats_histogram_info(S.stahist) AS histinfo
FROM (pg_mv_statistic S JOIN pg_class C ON (C.oid = S.starelid))
LEFT JOIN pg_namespace N ON (N.oid = C.relnamespace);
diff --git a/src/backend/commands/statscmds.c b/src/backend/commands/statscmds.c
index e428c69..6c001ed 100644
--- a/src/backend/commands/statscmds.c
+++ b/src/backend/commands/statscmds.c
@@ -77,7 +77,8 @@ CreateStatistics(CreateStatsStmt *stmt)
/* by default build nothing */
bool build_ndistinct = false,
build_dependencies = false,
- build_mcv = false;
+ build_mcv = false,
+ build_histogram = false;
Assert(IsA(stmt, CreateStatsStmt));
@@ -172,6 +173,8 @@ CreateStatistics(CreateStatsStmt *stmt)
build_dependencies = defGetBoolean(opt);
else if (strcmp(opt->defname, "mcv") == 0)
build_mcv = defGetBoolean(opt);
+ else if (strcmp(opt->defname, "histogram") == 0)
+ build_histogram = defGetBoolean(opt);
else
ereport(ERROR,
(errcode(ERRCODE_SYNTAX_ERROR),
@@ -180,10 +183,10 @@ CreateStatistics(CreateStatsStmt *stmt)
}
/* Make sure there's at least one statistics type specified. */
- if (!(build_ndistinct || build_dependencies || build_mcv))
+ if (!(build_ndistinct || build_dependencies || build_mcv || build_histogram))
ereport(ERROR,
(errcode(ERRCODE_SYNTAX_ERROR),
- errmsg("no statistics type (ndistinct, dependencies, mcv) requested")));
+ errmsg("no statistics type (ndistinct, dependencies, mcv, histogram) requested")));
stakeys = buildint2vector(attnums, numcols);
@@ -207,10 +210,12 @@ CreateStatistics(CreateStatsStmt *stmt)
values[Anum_pg_mv_statistic_ndist_enabled - 1] = BoolGetDatum(build_ndistinct);
values[Anum_pg_mv_statistic_deps_enabled - 1] = BoolGetDatum(build_dependencies);
values[Anum_pg_mv_statistic_mcv_enabled - 1] = BoolGetDatum(build_mcv);
+ values[Anum_pg_mv_statistic_hist_enabled - 1] = BoolGetDatum(build_histogram);
nulls[Anum_pg_mv_statistic_standist - 1] = true;
nulls[Anum_pg_mv_statistic_stadeps - 1] = true;
nulls[Anum_pg_mv_statistic_stamcv - 1] = true;
+ nulls[Anum_pg_mv_statistic_stahist - 1] = true;
/* insert the tuple into pg_mv_statistic */
mvstatrel = heap_open(MvStatisticRelationId, RowExclusiveLock);
diff --git a/src/backend/nodes/outfuncs.c b/src/backend/nodes/outfuncs.c
index ea3db02..901a328 100644
--- a/src/backend/nodes/outfuncs.c
+++ b/src/backend/nodes/outfuncs.c
@@ -2182,11 +2182,13 @@ _outMVStatisticInfo(StringInfo str, const MVStatisticInfo *node)
WRITE_BOOL_FIELD(ndist_enabled);
WRITE_BOOL_FIELD(deps_enabled);
WRITE_BOOL_FIELD(mcv_enabled);
+ WRITE_BOOL_FIELD(hist_enabled);
/* built/available statistics */
WRITE_BOOL_FIELD(ndist_built);
WRITE_BOOL_FIELD(deps_built);
WRITE_BOOL_FIELD(mcv_built);
+ WRITE_BOOL_FIELD(hist_built);
}
static void
diff --git a/src/backend/optimizer/path/clausesel.c b/src/backend/optimizer/path/clausesel.c
index c422dd5..0243b4d 100644
--- a/src/backend/optimizer/path/clausesel.c
+++ b/src/backend/optimizer/path/clausesel.c
@@ -49,6 +49,7 @@ static void addRangeClause(RangeQueryClause **rqlist, Node *clause,
#define STATS_TYPE_FDEPS 0x01
#define STATS_TYPE_MCV 0x02
+#define STATS_TYPE_HIST 0x04
static bool clause_is_mv_compatible(Node *clause, Index relid, Bitmapset **attnums,
int type);
@@ -77,12 +78,21 @@ static Selectivity clauselist_mv_selectivity_mcvlist(PlannerInfo *root,
List *clauses, MVStatisticInfo *mvstats,
bool *fullmatch, Selectivity *lowsel);
+static Selectivity clauselist_mv_selectivity_histogram(PlannerInfo *root,
+ List *clauses, MVStatisticInfo *mvstats);
+
static int update_match_bitmap_mcvlist(PlannerInfo *root, List *clauses,
int2vector *stakeys, MCVList mcvlist,
int nmatches, char *matches,
Selectivity *lowsel, bool *fullmatch,
bool is_or);
+static int update_match_bitmap_histogram(PlannerInfo *root, List *clauses,
+ int2vector *stakeys,
+ MVSerializedHistogram mvhist,
+ int nmatches, char *matches,
+ bool is_or);
+
static bool has_stats(List *stats, int type);
static List *find_stats(PlannerInfo *root, Index relid);
@@ -93,6 +103,7 @@ static bool stats_type_matches(MVStatisticInfo *stat, int type);
#define UPDATE_RESULT(m,r,isor) \
(m) = (isor) ? (Max(m,r)) : (Min(m,r))
+
/****************************************************************************
* ROUTINES TO COMPUTE SELECTIVITIES
****************************************************************************/
@@ -121,7 +132,7 @@ static bool stats_type_matches(MVStatisticInfo *stat, int type);
*
* First we try to reduce the list of clauses by applying (soft) functional
* dependencies, and then we try to estimate the selectivity of the reduced
- * list of clauses using the multivariate MCV list.
+ * list of clauses using the multivariate MCV list and histograms.
*
* Finally we remove the portion of clauses estimated using multivariate stats,
* and process the rest of the clauses using the regular per-column stats.
@@ -208,16 +219,17 @@ clauselist_selectivity(PlannerInfo *root,
* If there are no such stats or not enough attributes, don't waste time
* simply skip to estimation using the plain per-column stats.
*/
- if (has_stats(stats, STATS_TYPE_MCV) &&
- (count_mv_attnums(clauses, relid, STATS_TYPE_MCV) >= 2))
+ if (has_stats(stats, STATS_TYPE_MCV | STATS_TYPE_HIST) &&
+ (count_mv_attnums(clauses, relid,
+ STATS_TYPE_MCV | STATS_TYPE_HIST) >= 2))
{
/* collect attributes from the compatible conditions */
Bitmapset *mvattnums = collect_mv_attnums(clauses, relid,
- STATS_TYPE_MCV);
+ STATS_TYPE_MCV | STATS_TYPE_HIST);
/* and search for the statistic covering the most attributes */
MVStatisticInfo *mvstat = choose_mv_statistics(stats, mvattnums,
- STATS_TYPE_MCV);
+ STATS_TYPE_MCV | STATS_TYPE_HIST);
if (mvstat != NULL) /* we have a matching stats */
{
@@ -226,7 +238,7 @@ clauselist_selectivity(PlannerInfo *root,
/* split the clauselist into regular and mv-clauses */
clauses = clauselist_mv_split(root, relid, clauses, &mvclauses,
- mvstat, STATS_TYPE_MCV);
+ mvstat, STATS_TYPE_MCV | STATS_TYPE_HIST);
/* we've chosen the histogram to match the clauses */
Assert(mvclauses != NIL);
@@ -1178,6 +1190,8 @@ static Selectivity
clauselist_mv_selectivity(PlannerInfo *root, List *clauses, MVStatisticInfo *mvstats)
{
bool fullmatch = false;
+ Selectivity s1 = 0.0,
+ s2 = 0.0;
/*
* Lowest frequency in the MCV list (may be used as an upper bound for
@@ -1191,9 +1205,26 @@ clauselist_mv_selectivity(PlannerInfo *root, List *clauses, MVStatisticInfo *mvs
* order by selectivity (to optimize the MCV/histogram evaluation).
*/
- /* Evaluate the MCV selectivity */
- return clauselist_mv_selectivity_mcvlist(root, clauses, mvstats,
- &fullmatch, &mcv_low);
+ /* Evaluate the MCV first. */
+ s1 = clauselist_mv_selectivity_mcvlist(root, clauses, mvstats,
+ &fullmatch, &mcv_low);
+
+ /*
+ * If we got a full equality match on the MCV list, we're done (and the
+ * estimate is pretty good).
+ */
+ if (fullmatch && (s1 > 0.0))
+ return s1;
+
+ /*
+ * TODO if (fullmatch) without matching MCV item, use the mcv_low
+ * selectivity as upper bound
+ */
+
+ s2 = clauselist_mv_selectivity_histogram(root, clauses, mvstats);
+
+ /* TODO clamp to <= 1.0 (or more strictly, when possible) */
+ return s1 + s2;
}
/*
@@ -1379,7 +1410,8 @@ choose_mv_statistics(List *stats, Bitmapset *attnums, int types)
/* skip statistics not matching any of the requested types */
if (! ((info->deps_built && (STATS_TYPE_FDEPS & types)) ||
- (info->mcv_built && (STATS_TYPE_MCV & types))))
+ (info->mcv_built && (STATS_TYPE_MCV & types)) ||
+ (info->hist_built && (STATS_TYPE_HIST & types))))
continue;
/* count columns covered by the statistics */
@@ -1609,7 +1641,7 @@ mv_compatible_walker(Node *node, mv_compatible_context *context)
case F_SCALARGTSEL:
/* not compatible with functional dependencies */
- if (!(context->types & STATS_TYPE_MCV))
+ if (!(context->types & (STATS_TYPE_MCV | STATS_TYPE_HIST)))
return true; /* terminate */
break;
@@ -1677,6 +1709,9 @@ stats_type_matches(MVStatisticInfo *stat, int type)
if ((type & STATS_TYPE_MCV) && stat->mcv_built)
return true;
+ if ((type & STATS_TYPE_HIST) && stat->hist_built)
+ return true;
+
return false;
}
@@ -1695,6 +1730,9 @@ has_stats(List *stats, int type)
/* terminate if we've found at least one matching statistics */
if (stats_type_matches(stat, type))
return true;
+
+ if ((type & STATS_TYPE_HIST) && stat->hist_built)
+ return true;
}
return false;
@@ -1725,12 +1763,12 @@ find_stats(PlannerInfo *root, Index relid)
*
* The algorithm works like this:
*
- * 1) mark all items as 'match'
- * 2) walk through all the clauses
- * 3) for a particular clause, walk through all the items
- * 4) skip items that are already 'no match'
- * 5) check clause for items that still match
- * 6) sum frequencies for items to get selectivity
+ * 1) mark all items as 'match'
+ * 2) walk through all the clauses
+ * 3) for a particular clause, walk through all the items
+ * 4) skip items that are already 'no match'
+ * 5) check clause for items that still match
+ * 6) sum frequencies for items to get selectivity
*
* The function also returns the frequency of the least frequent item
* on the MCV list, which may be useful for clamping estimate from the
@@ -2116,3 +2154,537 @@ update_match_bitmap_mcvlist(PlannerInfo *root, List *clauses,
return nmatches;
}
+
+/*
+ * Estimate selectivity of clauses using a histogram.
+ *
+ * If there's no histogram for the stats, the function returns 0.0.
+ *
+ * The general idea of this method is similar to how MCV lists are
+ * processed, except that this introduces the concept of a partial
+ * match (MCV only works with full match / mismatch).
+ *
+ * The algorithm works like this:
+ *
+ * 1) mark all buckets as 'full match'
+ * 2) walk through all the clauses
+ * 3) for a particular clause, walk through all the buckets
+ * 4) skip buckets that are already 'no match'
+ * 5) check clause for buckets that still match (at least partially)
+ * 6) sum frequencies for buckets to get selectivity
+ *
+ * Unlike MCV lists, histograms have a concept of a partial match. In
+ * that case we use 1/2 the bucket, to minimize the average error. The
+ * MV histograms are usually less detailed than the per-column ones,
+ * meaning the sum is often quite high (thanks to combining a lot of
+ * "partially hit" buckets).
+ *
+ * Maybe we could use per-bucket information with number of distinct
+ * values it contains (for each dimension), and then use that to correct
+ * the estimate (so with 10 distinct values, we'd use 1/10 of the bucket
+ * frequency). We might also scale the value depending on the actual
+ * ndistinct estimate (not just the values observed in the sample).
+ *
+ * Another option would be to multiply the selectivities, i.e. if we get
+ * 'partial match' for a bucket for multiple conditions, we might use
+ * 0.5^k (where k is the number of conditions), instead of 0.5. This
+ * probably does not minimize the average error, though.
+ *
+ * TODO: This might use a similar shortcut to MCV lists - count buckets
+ * marked as partial/full match, and terminate once this drop to 0.
+ * Not sure if it's really worth it - for MCV lists a situation like
+ * this is not uncommon, but for histograms it's not that clear.
+ */
+static Selectivity
+clauselist_mv_selectivity_histogram(PlannerInfo *root, List *clauses,
+ MVStatisticInfo *mvstats)
+{
+ int i;
+ Selectivity s = 0.0;
+ Selectivity u = 0.0;
+
+ int nmatches = 0;
+ char *matches = NULL;
+
+ MVSerializedHistogram mvhist = NULL;
+
+ /* there's no histogram */
+ if (!mvstats->hist_built)
+ return 0.0;
+
+ /* There may be no histogram in the stats (check hist_built flag) */
+ mvhist = load_mv_histogram(mvstats->mvoid);
+
+ Assert(mvhist != NULL);
+ Assert(clauses != NIL);
+ Assert(list_length(clauses) >= 2);
+
+ /*
+ * Bitmap of bucket matches (mismatch, partial, full). by default all
+ * buckets fully match (and we'll eliminate them).
+ */
+ matches = palloc0(sizeof(char) * mvhist->nbuckets);
+ memset(matches, MVSTATS_MATCH_FULL, sizeof(char) * mvhist->nbuckets);
+
+ nmatches = mvhist->nbuckets;
+
+ /* build the match bitmap */
+ update_match_bitmap_histogram(root, clauses,
+ mvstats->stakeys, mvhist,
+ nmatches, matches, false);
+
+ /* now, walk through the buckets and sum the selectivities */
+ for (i = 0; i < mvhist->nbuckets; i++)
+ {
+ /*
+ * Find out what part of the data is covered by the histogram, so that
+ * we can 'scale' the selectivity properly (e.g. when only 50% of the
+ * sample got into the histogram, and the rest is in a MCV list).
+ *
+ * TODO This might be handled by keeping a global "frequency" for the
+ * whole histogram, which might save us some time spent accessing the
+ * not-matching part of the histogram. Although it's likely in a
+ * cache, so it's very fast.
+ */
+ u += mvhist->buckets[i]->ntuples;
+
+ if (matches[i] == MVSTATS_MATCH_FULL)
+ s += mvhist->buckets[i]->ntuples;
+ else if (matches[i] == MVSTATS_MATCH_PARTIAL)
+ s += 0.5 * mvhist->buckets[i]->ntuples;
+ }
+
+#ifdef DEBUG_MVHIST
+ debug_histogram_matches(mvhist, matches);
+#endif
+
+ /* release the allocated bitmap and deserialized histogram */
+ pfree(matches);
+ pfree(mvhist);
+
+ return s * u;
+}
+
+/* cached result of bucket boundary comparison for a single dimension */
+
+#define HIST_CACHE_NOT_FOUND 0x00
+#define HIST_CACHE_FALSE 0x01
+#define HIST_CACHE_TRUE 0x03
+#define HIST_CACHE_MASK 0x02
+
+static char
+bucket_contains_value(FmgrInfo ltproc, Datum constvalue,
+ Datum min_value, Datum max_value,
+ int min_index, int max_index,
+ bool min_include, bool max_include,
+ char *callcache)
+{
+ bool a,
+ b;
+
+ char min_cached = callcache[min_index];
+ char max_cached = callcache[max_index];
+
+ /*
+ * First some quick checks on equality - if any of the boundaries equals,
+ * we have a partial match (so no need to call the comparator).
+ */
+ if (((min_value == constvalue) && (min_include)) ||
+ ((max_value == constvalue) && (max_include)))
+ return MVSTATS_MATCH_PARTIAL;
+
+ /* Keep the values 0/1 because of the XOR at the end. */
+ a = ((min_cached & HIST_CACHE_MASK) >> 1);
+ b = ((max_cached & HIST_CACHE_MASK) >> 1);
+
+ /*
+ * If result for the bucket lower bound not in cache, evaluate the
+ * function and store the result in the cache.
+ */
+ if (!min_cached)
+ {
+ a = DatumGetBool(FunctionCall2Coll(<proc,
+ DEFAULT_COLLATION_OID,
+ constvalue, min_value));
+ /* remember the result */
+ callcache[min_index] = (a) ? HIST_CACHE_TRUE : HIST_CACHE_FALSE;
+ }
+
+ /* And do the same for the upper bound. */
+ if (!max_cached)
+ {
+ b = DatumGetBool(FunctionCall2Coll(<proc,
+ DEFAULT_COLLATION_OID,
+ constvalue, max_value));
+ /* remember the result */
+ callcache[max_index] = (b) ? HIST_CACHE_TRUE : HIST_CACHE_FALSE;
+ }
+
+ return (a ^ b) ? MVSTATS_MATCH_PARTIAL : MVSTATS_MATCH_NONE;
+}
+
+static char
+bucket_is_smaller_than_value(FmgrInfo opproc, Datum constvalue,
+ Datum min_value, Datum max_value,
+ int min_index, int max_index,
+ bool min_include, bool max_include,
+ char *callcache, bool isgt)
+{
+ char min_cached = callcache[min_index];
+ char max_cached = callcache[max_index];
+
+ /* Keep the values 0/1 because of the XOR at the end. */
+ bool a = ((min_cached & HIST_CACHE_MASK) >> 1);
+ bool b = ((max_cached & HIST_CACHE_MASK) >> 1);
+
+ if (!min_cached)
+ {
+ a = DatumGetBool(FunctionCall2Coll(&opproc,
+ DEFAULT_COLLATION_OID,
+ min_value,
+ constvalue));
+ /* remember the result */
+ callcache[min_index] = (a) ? HIST_CACHE_TRUE : HIST_CACHE_FALSE;
+ }
+
+ if (!max_cached)
+ {
+ b = DatumGetBool(FunctionCall2Coll(&opproc,
+ DEFAULT_COLLATION_OID,
+ max_value,
+ constvalue));
+ /* remember the result */
+ callcache[max_index] = (b) ? HIST_CACHE_TRUE : HIST_CACHE_FALSE;
+ }
+
+ /*
+ * Now, we need to combine both results into the final answer, and we need
+ * to be careful about the 'isgt' variable which kinda inverts the
+ * meaning.
+ *
+ * First, we handle the case when each boundary returns different results.
+ * In that case the outcome can only be 'partial' match.
+ */
+ if (a != b)
+ return MVSTATS_MATCH_PARTIAL;
+
+ /*
+ * When the results are the same, then it depends on the 'isgt' value.
+ * There are four options:
+ *
+ * isgt=false a=b=true => full match isgt=false a=b=false => empty
+ * isgt=true a=b=true => empty isgt=true a=b=false => full match
+ *
+ * We'll cheat a bit, because we know that (a=b) so we'll use just one of
+ * them.
+ */
+ if (isgt)
+ return (!a) ? MVSTATS_MATCH_FULL : MVSTATS_MATCH_NONE;
+ else
+ return (a) ? MVSTATS_MATCH_FULL : MVSTATS_MATCH_NONE;
+}
+
+/*
+ * Evaluate clauses using the histogram, and update the match bitmap.
+ *
+ * The bitmap may be already partially set, so this is really a way to
+ * combine results of several clause lists - either when computing
+ * conditional probability P(A|B) or a combination of AND/OR clauses.
+ *
+ * Note: This is not a simple bitmap in the sense that there are more
+ * than two possible values for each item - no match, partial
+ * match and full match. So we need 2 bits per item.
+ *
+ * TODO: This works with 'bitmap' where each item is represented as a
+ * char, which is slightly wasteful. Instead, we could use a bitmap
+ * with 2 bits per item, reducing the size to ~1/4. By using values
+ * 0, 1 and 3 (instead of 0, 1 and 2), the operations (merging etc.)
+ * might be performed just like for simple bitmap by using & and |,
+ * which might be faster than min/max.
+ */
+static int
+update_match_bitmap_histogram(PlannerInfo *root, List *clauses,
+ int2vector *stakeys,
+ MVSerializedHistogram mvhist,
+ int nmatches, char *matches,
+ bool is_or)
+{
+ int i;
+ ListCell *l;
+
+ /*
+ * Used for caching function calls, only once per deduplicated value.
+ *
+ * We know may have up to (2 * nbuckets) values per dimension. It's
+ * probably overkill, but let's allocate that once for all clauses, to
+ * minimize overhead.
+ *
+ * Also, we only need two bits per value, but this allocates byte per
+ * value. Might be worth optimizing.
+ *
+ * 0x00 - not yet called 0x01 - called, result is 'false' 0x03 - called,
+ * result is 'true'
+ */
+ char *callcache = palloc(mvhist->nbuckets);
+
+ Assert(mvhist != NULL);
+ Assert(mvhist->nbuckets > 0);
+ Assert(nmatches >= 0);
+ Assert(nmatches <= mvhist->nbuckets);
+
+ Assert(clauses != NIL);
+ Assert(list_length(clauses) >= 1);
+
+ /* loop through the clauses and do the estimation */
+ foreach(l, clauses)
+ {
+ Node *clause = (Node *) lfirst(l);
+
+ /* if it's a RestrictInfo, then extract the clause */
+ if (IsA(clause, RestrictInfo))
+ clause = (Node *) ((RestrictInfo *) clause)->clause;
+
+ /* it's either OpClause, or NullTest */
+ if (is_opclause(clause))
+ {
+ OpExpr *expr = (OpExpr *) clause;
+ bool varonleft = true;
+ bool ok;
+
+ FmgrInfo opproc; /* operator */
+
+ fmgr_info(get_opcode(expr->opno), &opproc);
+
+ /* reset the cache (per clause) */
+ memset(callcache, 0, mvhist->nbuckets);
+
+ ok = (NumRelids(clause) == 1) &&
+ (is_pseudo_constant_clause(lsecond(expr->args)) ||
+ (varonleft = false,
+ is_pseudo_constant_clause(linitial(expr->args))));
+
+ if (ok)
+ {
+ FmgrInfo ltproc;
+ RegProcedure oprrest = get_oprrest(expr->opno);
+
+ Var *var = (varonleft) ? linitial(expr->args) : lsecond(expr->args);
+ Const *cst = (varonleft) ? lsecond(expr->args) : linitial(expr->args);
+ bool isgt = (!varonleft);
+
+ TypeCacheEntry *typecache
+ = lookup_type_cache(var->vartype, TYPECACHE_LT_OPR);
+
+ /* lookup dimension for the attribute */
+ int idx = mv_get_index(var->varattno, stakeys);
+
+ fmgr_info(get_opcode(typecache->lt_opr), <proc);
+
+ /*
+ * Check this for all buckets that still have "true" in the
+ * bitmap
+ *
+ * We already know the clauses use suitable operators (because
+ * that's how we filtered them).
+ */
+ for (i = 0; i < mvhist->nbuckets; i++)
+ {
+ char res = MVSTATS_MATCH_NONE;
+
+ MVSerializedBucket bucket = mvhist->buckets[i];
+
+ /* histogram boundaries */
+ Datum minval,
+ maxval;
+ bool mininclude,
+ maxinclude;
+ int minidx,
+ maxidx;
+
+ /*
+ * For AND-lists, we can also mark NULL buckets as 'no
+ * match' (and then skip them). For OR-lists this is not
+ * possible.
+ */
+ if ((!is_or) && bucket->nullsonly[idx])
+ matches[i] = MVSTATS_MATCH_NONE;
+
+ /*
+ * Skip buckets that were already eliminated - this is
+ * impotant considering how we update the info (we only
+ * lower the match). We can't really do anything about the
+ * MATCH_PARTIAL buckets.
+ */
+ if ((!is_or) && (matches[i] == MVSTATS_MATCH_NONE))
+ continue;
+ else if (is_or && (matches[i] == MVSTATS_MATCH_FULL))
+ continue;
+
+ /* lookup the values and cache of function calls */
+ minidx = bucket->min[idx];
+ maxidx = bucket->max[idx];
+
+ minval = mvhist->values[idx][bucket->min[idx]];
+ maxval = mvhist->values[idx][bucket->max[idx]];
+
+ mininclude = bucket->min_inclusive[idx];
+ maxinclude = bucket->max_inclusive[idx];
+
+ /*
+ * TODO Maybe it's possible to add here a similar
+ * optimization as for the MCV lists:
+ *
+ * (nmatches == 0) && AND-list => all eliminated (FALSE)
+ * (nmatches == N) && OR-list => all eliminated (TRUE)
+ *
+ * But it's more complex because of the partial matches.
+ */
+
+ /*
+ * If it's not a "<" or ">" or "=" operator, just ignore
+ * the clause. Otherwise note the relid and attnum for the
+ * variable.
+ *
+ * TODO I'm really unsure the handling of 'isgt' flag
+ * (that is, clauses with reverse order of
+ * variable/constant) is correct. I wouldn't be surprised
+ * if there was some mixup. Using the lt/gt operators
+ * instead of messing with the opproc could make it
+ * simpler. It would however be using a different operator
+ * than the query, although it's not any shadier than
+ * using the selectivity function as is done currently.
+ */
+ switch (oprrest)
+ {
+ case F_SCALARLTSEL: /* Var < Const */
+ case F_SCALARGTSEL: /* Var > Const */
+
+ res = bucket_is_smaller_than_value(opproc, cst->constvalue,
+ minval, maxval,
+ minidx, maxidx,
+ mininclude, maxinclude,
+ callcache, isgt);
+ break;
+
+ case F_EQSEL:
+
+ /*
+ * We only check whether the value is within the
+ * bucket, using the lt operator, and we also
+ * check for equality with the boundaries.
+ */
+
+ res = bucket_contains_value(ltproc, cst->constvalue,
+ minval, maxval,
+ minidx, maxidx,
+ mininclude, maxinclude,
+ callcache);
+ break;
+ }
+
+ UPDATE_RESULT(matches[i], res, is_or);
+
+ }
+ }
+ }
+ else if (IsA(clause, NullTest))
+ {
+ NullTest *expr = (NullTest *) clause;
+ Var *var = (Var *) (expr->arg);
+
+ /* FIXME proper matching attribute to dimension */
+ int idx = mv_get_index(var->varattno, stakeys);
+
+ /*
+ * Walk through the buckets and evaluate the current clause. We
+ * can skip items that were already ruled out, and terminate if
+ * there are no remaining buckets that might possibly match.
+ */
+ for (i = 0; i < mvhist->nbuckets; i++)
+ {
+ MVSerializedBucket bucket = mvhist->buckets[i];
+
+ /*
+ * Skip buckets that were already eliminated - this is
+ * impotant considering how we update the info (we only lower
+ * the match)
+ */
+ if ((!is_or) && (matches[i] == MVSTATS_MATCH_NONE))
+ continue;
+ else if (is_or && (matches[i] == MVSTATS_MATCH_FULL))
+ continue;
+
+ /* if the clause mismatches the bucket, set it as MATCH_NONE */
+ if ((expr->nulltesttype == IS_NULL)
+ && (!bucket->nullsonly[idx]))
+ UPDATE_RESULT(matches[i], MVSTATS_MATCH_NONE, is_or);
+
+ else if ((expr->nulltesttype == IS_NOT_NULL) &&
+ (bucket->nullsonly[idx]))
+ UPDATE_RESULT(matches[i], MVSTATS_MATCH_NONE, is_or);
+ }
+ }
+ else if (or_clause(clause) || and_clause(clause))
+ {
+ /*
+ * AND/OR clause, with all clauses compatible with the selected MV
+ * stat
+ */
+
+ int i;
+ BoolExpr *orclause = ((BoolExpr *) clause);
+ List *orclauses = orclause->args;
+
+ /* match/mismatch bitmap for each bucket */
+ int or_nmatches = 0;
+ char *or_matches = NULL;
+
+ Assert(orclauses != NIL);
+ Assert(list_length(orclauses) >= 2);
+
+ /* number of matching buckets */
+ or_nmatches = mvhist->nbuckets;
+
+ /* by default none of the buckets matches the clauses */
+ or_matches = palloc0(sizeof(char) * or_nmatches);
+
+ if (or_clause(clause))
+ {
+ /* OR clauses assume nothing matches, initially */
+ memset(or_matches, MVSTATS_MATCH_NONE, sizeof(char) * or_nmatches);
+ or_nmatches = 0;
+ }
+ else
+ {
+ /* AND clauses assume nothing matches, initially */
+ memset(or_matches, MVSTATS_MATCH_FULL, sizeof(char) * or_nmatches);
+ }
+
+ /* build the match bitmap for the OR-clauses */
+ or_nmatches = update_match_bitmap_histogram(root, orclauses,
+ stakeys, mvhist,
+ or_nmatches, or_matches, or_clause(clause));
+
+ /* merge the bitmap into the existing one */
+ for (i = 0; i < mvhist->nbuckets; i++)
+ {
+ /*
+ * Merge the result into the bitmap (Min for AND, Max for OR).
+ *
+ * FIXME this does not decrease the number of matches
+ */
+ UPDATE_RESULT(matches[i], or_matches[i], is_or);
+ }
+
+ pfree(or_matches);
+
+ }
+ else
+ elog(ERROR, "unknown clause type: %d", clause->type);
+ }
+
+ /* free the call cache */
+ pfree(callcache);
+
+ return nmatches;
+}
diff --git a/src/backend/optimizer/util/plancat.c b/src/backend/optimizer/util/plancat.c
index ff93ddb..f7615f5 100644
--- a/src/backend/optimizer/util/plancat.c
+++ b/src/backend/optimizer/util/plancat.c
@@ -422,7 +422,7 @@ get_relation_info(PlannerInfo *root, Oid relationObjectId, bool inhparent,
mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
/* unavailable stats are not interesting for the planner */
- if (mvstat->deps_built || mvstat->ndist_built || mvstat->mcv_built)
+ if (mvstat->deps_built || mvstat->ndist_built || mvstat->mcv_built || mvstat->hist_built)
{
info = makeNode(MVStatisticInfo);
@@ -433,11 +433,13 @@ get_relation_info(PlannerInfo *root, Oid relationObjectId, bool inhparent,
info->ndist_enabled = mvstat->ndist_enabled;
info->deps_enabled = mvstat->deps_enabled;
info->mcv_enabled = mvstat->mcv_enabled;
+ info->hist_enabled = mvstat->hist_enabled;
/* built/available statistics */
info->ndist_built = mvstat->ndist_built;
info->deps_built = mvstat->deps_built;
info->mcv_built = mvstat->mcv_built;
+ info->hist_built = mvstat->hist_built;
/* stakeys */
adatum = SysCacheGetAttr(MVSTATOID, htup,
diff --git a/src/backend/utils/mvstats/Makefile b/src/backend/utils/mvstats/Makefile
index d5d47ba..d4b88e9 100644
--- a/src/backend/utils/mvstats/Makefile
+++ b/src/backend/utils/mvstats/Makefile
@@ -12,6 +12,6 @@ subdir = src/backend/utils/mvstats
top_builddir = ../../../..
include $(top_builddir)/src/Makefile.global
-OBJS = common.o dependencies.o mcv.o mvdist.o
+OBJS = common.o dependencies.o histogram.o mcv.o mvdist.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/utils/mvstats/README.histogram b/src/backend/utils/mvstats/README.histogram
new file mode 100644
index 0000000..a182fa3
--- /dev/null
+++ b/src/backend/utils/mvstats/README.histogram
@@ -0,0 +1,299 @@
+Multivariate histograms
+=======================
+
+Histograms on individual attributes consist of buckets represented by ranges,
+covering the domain of the attribute. That is, each bucket is a [min,max]
+interval, and contains all values in this range. The histogram is built in such
+a way that all buckets have about the same frequency.
+
+Multivariate histograms are an extension into n-dimensional space - the buckets
+are n-dimensional intervals (i.e. n-dimensional rectagles), covering the domain
+of the combination of attributes. That is, each bucket has a vector of lower
+and upper boundaries, denoted min[i] and max[i] (where i = 1..n).
+
+In addition to the boundaries, each bucket tracks additional info:
+
+ * frequency (fraction of tuples in the bucket)
+ * whether the boundaries are inclusive or exclusive
+ * whether the dimension contains only NULL values
+ * number of distinct values in each dimension (for building only)
+
+It's possible that in the future we'll multiple histogram types, with different
+features. We do however expect all the types to share the same representation
+(buckets as ranges) and only differ in how we build them.
+
+The current implementation builds non-overlapping buckets, that may not be true
+for some histogram types and the code should not rely on this assumption. There
+are interesting types of histograms (or algorithms) with overlapping buckets.
+
+When used on low-cardinality data, histograms usually perform considerably worse
+than MCV lists (which are a good fit for this kind of data). This is especially
+true on label-like values, where ordering of the values is mostly unrelated to
+meaning of the data, as proper ordering is crucial for histograms.
+
+On high-cardinality data the histograms are usually a better choice, because MCV
+lists can't represent the distribution accurately enough.
+
+
+Selectivity estimation
+----------------------
+
+The estimation is implemented in clauselist_mv_selectivity_histogram(), and
+works very similarly to clauselist_mv_selectivity_mcvlist.
+
+The main difference is that while MCV lists support exact matches, histograms
+often result in approximate matches - e.g. with equality we can only say if
+the constant would be part of the bucket, but not whether it really is there
+or what fraction of the bucket it corresponds to. In this case we rely on
+some defaults just like in the per-column histograms.
+
+The current implementation uses histograms to estimates those types of clauses
+(think of WHERE conditions):
+
+ (a) equality clauses WHERE (a = 1) AND (b = 2)
+ (b) inequality clauses WHERE (a < 1) AND (b >= 2)
+ (c) NULL clauses WHERE (a IS NULL) AND (b IS NOT NULL)
+ (d) OR-clauses WHERE (a = 1) OR (b = 2)
+
+Similarly to MCV lists, it's possible to add support for additional types of
+clauses, for example:
+
+ (e) multi-var clauses WHERE (a > b)
+
+and so on. These are tasks for the future, not yet implemented.
+
+
+When evaluating a clause on a bucket, we may get one of three results:
+
+ (a) FULL_MATCH - The bucket definitely matches the clause.
+
+ (b) PARTIAL_MATCH - The bucket matches the clause, but not necessarily all
+ the tuples it represents.
+
+ (c) NO_MATCH - The bucket definitely does not match the clause.
+
+This may be illustrated using a range [1, 5], which is essentially a 1-D bucket.
+With clause
+
+ WHERE (a < 10) => FULL_MATCH (all range values are below
+ 10, so the whole bucket matches)
+
+ WHERE (a < 3) => PARTIAL_MATCH (there may be values matching
+ the clause, but we don't know how many)
+
+ WHERE (a < 0) => NO_MATCH (the whole range is above 1, so
+ no values from the bucket can match)
+
+Some clauses may produce only some of those results - for example equality
+clauses may never produce FULL_MATCH as we always hit only part of the bucket
+(we can't match both boundaries at the same time). This results in less accurate
+estimates compared to MCV lists, where we can hit a MCV items exactly (there's
+no PARTIAL match in MCV).
+
+There are also clauses that may not produce any PARTIAL_MATCH results. A nice
+example of that is 'IS [NOT] NULL' clause, which either matches the bucket
+completely (FULL_MATCH) or not at all (NO_MATCH), thanks to how the NULL-buckets
+are constructed.
+
+Computing the total selectivity estimate is trivial - simply sum selectivities
+from all the FULL_MATCH and PARTIAL_MATCH buckets (but for buckets marked with
+PARTIAL_MATCH, multiply the frequency by 0.5 to minimize the average error).
+
+
+Building a histogram
+---------------------
+
+The algorithm of building a histogram in general is quite simple:
+
+ (a) create an initial bucket (containing all sample rows)
+
+ (b) create NULL buckets (by splitting the initial bucket)
+
+ (c) repeat
+
+ (1) choose bucket to split next
+
+ (2) terminate if no bucket that might be split found, or if we've
+ reached the maximum number of buckets (16384)
+
+ (3) choose dimension to partition the bucket by
+
+ (4) partition the bucket by the selected dimension
+
+The main complexity is hidden in steps (c.1) and (c.3), i.e. how we choose the
+bucket and dimension for the split, as discussed in the next section.
+
+
+Partitioning criteria
+---------------------
+
+Similarly to one-dimensional histograms, we want to produce buckets with roughly
+the same frequency.
+
+We also need to produce "regular" buckets, because buckets with one dimension
+much longer than the others are very likely to match a lot of conditions (which
+increases error, even if the bucket frequency is very low).
+
+This is especially important when handling OR-clauses, because in that case each
+clause may add buckets independently. With AND-clauses all the clauses have to
+match each bucket, which makes this issue somewhat less concenrning.
+
+To achieve this, we choose the largest bucket (containing the most sample rows),
+but we only choose buckets that can actually be split (have at least 3 different
+combinations of values).
+
+Then we choose the "longest" dimension of the bucket, which is computed by using
+the distinct values in the sample as a measure.
+
+For details see functions select_bucket_to_partition() and partition_bucket(),
+which also includes further discussion.
+
+
+The current limit on number of buckets (16384) is mostly arbitrary, but chosen
+so that it guarantees we don't exceed the number of distinct values indexable by
+uint16 in any of the dimensions. In practice we could handle more buckets as we
+index each dimension separately and the splits should use the dimensions evenly.
+
+Also, histograms this large (with 16k values in multiple dimensions) would be
+quite expensive to build and process, so the 16k limit is rather reasonable.
+
+The actual number of buckets is also related to statistics target, because we
+require MIN_BUCKET_ROWS (10) tuples per bucket before a split, so we can't have
+more than (2 * 300 * target / 10) buckets. For the default target (100) this
+evaluates to ~6k.
+
+
+NULL handling (create_null_buckets)
+-----------------------------------
+
+When building histograms on a single attribute, we first filter out NULL values.
+In the multivariate case, we can't really do that because the rows may contain
+a mix of NULL and non-NULL values in different columns (so we can't simply
+filter all of them out).
+
+For this reason, the histograms are built in a way so that for each bucket, each
+dimension only contains only NULL or non-NULL values. Building the NULL-buckets
+happens as the first step in the build, by the create_null_buckets() function.
+The number of NULL buckets, as produced by this function, has a clear upper
+boundary (2^N) where N is the number of dimensions (attributes the histogram is
+built on). Or rather 2^K where K is the number of attributes that are not marked
+as not-NULL.
+
+The buckets with NULL dimensions are then subject to the same build algorithm
+(i.e. may be split into smaller buckets) just like any other bucket, but may
+only be split by non-NULL dimension.
+
+
+Serialization
+-------------
+
+To store the histogram in pg_mv_statistic table, it is serialized into a more
+efficient form. We also use the representation for estimation, i.e. we don't
+fully deserialize the histogram.
+
+For example the boundary values are deduplicated to minimize the required space.
+How much redundancy is there, actually? Let's assume there are no NULL values,
+so we start with a single bucket - in that case we have 2*N boundaries. Each
+time we split a bucket we introduce one new value (in the "middle" of one of
+the dimensions), and keep boundries for all the other dimensions. So after K
+splits, we have up to
+
+ 2*N + K
+
+unique boundary values (we may have fewe values, if the same value is used for
+several splits). But after K splits we do have (K+1) buckets, so
+
+ (K+1) * 2 * N
+
+boundary values. Using e.g. N=4 and K=999, we arrive to those numbers:
+
+ 2*N + K = 1007
+ (K+1) * 2 * N = 8000
+
+wich means a lot of redundancy. It's somewhat counter-intuitive that the number
+of distinct values does not really depend on the number of dimensions (except
+for the initial bucket, but that's negligible compared to the total).
+
+By deduplicating the values and replacing them with 16-bit indexes (uint16), we
+reduce the required space to
+
+ 1007 * 8 + 8000 * 2 ~= 24kB
+
+which is significantly less than 64kB required for the 'raw' histogram (assuming
+the values are 8B).
+
+While the bytea compression (pglz) might achieve the same reduction of space,
+the deduplicated representation is used to optimize the estimation by caching
+results of function calls for already visited values. This significantly
+reduces the number of calls to (often quite expensive) operators.
+
+Note: Of course, this reasoning only holds for histograms built by the algorithm
+that simply splits the buckets in half. Other histograms types (e.g. containing
+overlapping buckets) may behave differently and require different serialization.
+
+Serialized histograms are marked with 'magic' constant, to make it easier to
+check the bytea value really is a serialized histogram.
+
+
+varlena compression
+-------------------
+
+This serialization may however disable automatic varlena compression, the array
+of unique values is placed at the beginning of the serialized form. Which is
+exactly the chunk used by pglz to check if the data is compressible, and it
+will probably decide it's not very compressible. This is similar to the issue
+we had with JSONB initially.
+
+Maybe storing buckets first would make it work, as the buckets may be better
+compressible.
+
+On the other hand the serialization is actually a context-aware compression,
+usually compressing to ~30% (or even less, with large data types). So the lack
+of additional pglz compression may be acceptable.
+
+
+Deserialization
+---------------
+
+The deserialization is not a perfect inverse of the serialization, as we keep
+the deduplicated arrays. This reduces the amount of memory and also allows
+optimizations during estimation (e.g. we can cache results for the distinct
+values, saving expensive function calls).
+
+
+Inspecting the histogram
+------------------------
+
+Inspecting the regular (per-attribute) histograms is trivial, as it's enough
+to select the columns from pg_stats - the data is encoded as anyarray, so we
+simply get the text representation of the array.
+
+With multivariate histograms it's not that simple due to the possible mix of
+data types in the histogram. It might be possible to produce similar array-like
+text representation, but that'd unnecessarily complicate further processing
+and analysis of the histogram. Instead, there's a SRF function that allows
+access to lower/upper boundaries, frequencies etc.
+
+ SELECT * FROM pg_mv_histogram_buckets();
+
+It has two input parameters:
+
+ oid - OID of the histogram (pg_mv_statistic.staoid)
+ otype - type of output
+
+and produces a table with these columns:
+
+ - bucket ID (0...nbuckets-1)
+ - lower bucket boundaries (string array)
+ - upper bucket boundaries (string array)
+ - nulls only dimensions (boolean array)
+ - lower boundary inclusive (boolean array)
+ - upper boundary includive (boolean array)
+ - frequency (double precision)
+
+The 'otype' accepts three values, determining what will be returned in the
+lower/upper boundary arrays:
+
+ - 0 - values stored in the histogram, encoded as text
+ - 1 - indexes into the deduplicated arrays
+ - 2 - idnexes into the deduplicated arrays, scaled to [0,1]
diff --git a/src/backend/utils/mvstats/README.stats b/src/backend/utils/mvstats/README.stats
index 57f5c8b..6072b29 100644
--- a/src/backend/utils/mvstats/README.stats
+++ b/src/backend/utils/mvstats/README.stats
@@ -18,6 +18,8 @@ Currently we only have two kinds of multivariate statistics
(b) MCV lists (README.mcv)
+ (c) multivariate histograms (README.histogram)
+
Compatible clause types
-----------------------
diff --git a/src/backend/utils/mvstats/common.c b/src/backend/utils/mvstats/common.c
index 5d9caa8..5246012 100644
--- a/src/backend/utils/mvstats/common.c
+++ b/src/backend/utils/mvstats/common.c
@@ -13,6 +13,7 @@
*
*-------------------------------------------------------------------------
*/
+#include "postgres.h"
#include "common.h"
#include "utils/array.h"
@@ -52,7 +53,8 @@ build_mv_stats(Relation onerel, double totalrows,
MVNDistinct ndistinct = NULL;
MVDependencies deps = NULL;
MCVList mcvlist = NULL;
- int numrows_filtered = 0;
+ MVHistogram histogram = NULL;
+ int numrows_filtered = numrows;
VacAttrStats **stats = NULL;
int numatts = 0;
@@ -97,8 +99,12 @@ build_mv_stats(Relation onerel, double totalrows,
if (stat->mcv_enabled)
mcvlist = build_mv_mcvlist(numrows, rows, attrs, stats, &numrows_filtered);
- /* store the statistics in the catalog */
- update_mv_stats(stat->mvoid, ndistinct, deps, mcvlist, attrs, stats);
+ /* build a multivariate histogram on the columns */
+ if ((numrows_filtered > 0) && (stat->hist_enabled))
+ histogram = build_mv_histogram(numrows_filtered, rows, attrs, stats, numrows);
+
+ /* store the histogram / MCV list in the catalog */
+ update_mv_stats(stat->mvoid, ndistinct, deps, mcvlist, histogram, attrs, stats);
}
}
@@ -182,6 +188,8 @@ list_mv_stats(Oid relid)
info->deps_built = stats->deps_built;
info->mcv_enabled = stats->mcv_enabled;
info->mcv_built = stats->mcv_built;
+ info->hist_enabled = stats->hist_enabled;
+ info->hist_built = stats->hist_built;
result = lappend(result, info);
}
@@ -198,7 +206,6 @@ list_mv_stats(Oid relid)
return result;
}
-
/*
* Find attnims of MV stats using the mvoid.
*/
@@ -247,7 +254,8 @@ find_mv_attnums(Oid mvoid, Oid *relid)
void
update_mv_stats(Oid mvoid,
- MVNDistinct ndistinct, MVDependencies dependencies, MCVList mcvlist,
+ MVNDistinct ndistinct, MVDependencies dependencies,
+ MCVList mcvlist, MVHistogram histogram,
int2vector *attrs, VacAttrStats **stats)
{
HeapTuple stup,
@@ -289,15 +297,26 @@ update_mv_stats(Oid mvoid,
values[Anum_pg_mv_statistic_stamcv - 1] = PointerGetDatum(data);
}
+ if (histogram != NULL)
+ {
+ bytea *data = serialize_mv_histogram(histogram, attrs, stats);
+
+ nulls[Anum_pg_mv_statistic_stahist - 1] = (data == NULL);
+ values[Anum_pg_mv_statistic_stahist - 1]
+ = PointerGetDatum(data);
+ }
+
/* always replace the value (either by bytea or NULL) */
replaces[Anum_pg_mv_statistic_standist - 1] = true;
replaces[Anum_pg_mv_statistic_stadeps - 1] = true;
replaces[Anum_pg_mv_statistic_stamcv - 1] = true;
+ replaces[Anum_pg_mv_statistic_stahist - 1] = true;
/* always change the availability flags */
nulls[Anum_pg_mv_statistic_ndist_built - 1] = false;
nulls[Anum_pg_mv_statistic_deps_built - 1] = false;
nulls[Anum_pg_mv_statistic_mcv_built - 1] = false;
+ nulls[Anum_pg_mv_statistic_hist_built - 1] = false;
nulls[Anum_pg_mv_statistic_stakeys - 1] = false;
@@ -305,12 +324,14 @@ update_mv_stats(Oid mvoid,
replaces[Anum_pg_mv_statistic_ndist_built - 1] = true;
replaces[Anum_pg_mv_statistic_deps_built - 1] = true;
replaces[Anum_pg_mv_statistic_mcv_built - 1] = true;
+ replaces[Anum_pg_mv_statistic_hist_built - 1] = true;
replaces[Anum_pg_mv_statistic_stakeys - 1] = true;
values[Anum_pg_mv_statistic_ndist_built - 1] = BoolGetDatum(ndistinct != NULL);
values[Anum_pg_mv_statistic_deps_built - 1] = BoolGetDatum(dependencies != NULL);
values[Anum_pg_mv_statistic_mcv_built - 1] = BoolGetDatum(mcvlist != NULL);
+ values[Anum_pg_mv_statistic_hist_built - 1] = BoolGetDatum(histogram != NULL);
values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(attrs);
diff --git a/src/backend/utils/mvstats/common.h b/src/backend/utils/mvstats/common.h
index fe56f51..96c0317 100644
--- a/src/backend/utils/mvstats/common.h
+++ b/src/backend/utils/mvstats/common.h
@@ -77,7 +77,7 @@ MultiSortSupport multi_sort_init(int ndims);
void multi_sort_add_dimension(MultiSortSupport mss, int sortdim,
int dim, VacAttrStats **vacattrstats);
-int multi_sort_compare(const void *a, const void *b, void *arg);
+int multi_sort_compare(const void *a, const void *b, void *arg);
int multi_sort_compare_dim(int dim, const SortItem *a,
const SortItem *b, MultiSortSupport mss);
@@ -86,9 +86,9 @@ int multi_sort_compare_dims(int start, int end, const SortItem *a,
const SortItem *b, MultiSortSupport mss);
/* comparators, used when constructing multivariate stats */
-int compare_datums_simple(Datum a, Datum b, SortSupport ssup);
-int compare_scalars_simple(const void *a, const void *b, void *arg);
-int compare_scalars_partition(const void *a, const void *b, void *arg);
+int compare_datums_simple(Datum a, Datum b, SortSupport ssup);
+int compare_scalars_simple(const void *a, const void *b, void *arg);
+int compare_scalars_partition(const void *a, const void *b, void *arg);
void *bsearch_arg(const void *key, const void *base,
size_t nmemb, size_t size,
diff --git a/src/backend/utils/mvstats/histogram.c b/src/backend/utils/mvstats/histogram.c
new file mode 100644
index 0000000..fc0c9c2
--- /dev/null
+++ b/src/backend/utils/mvstats/histogram.c
@@ -0,0 +1,2123 @@
+/*-------------------------------------------------------------------------
+ *
+ * histogram.c
+ * POSTGRES multivariate histograms
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/histogram.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "fmgr.h"
+#include "funcapi.h"
+
+#include "utils/bytea.h"
+#include "utils/lsyscache.h"
+
+#include "common.h"
+#include <math.h>
+
+
+static MVBucket create_initial_mv_bucket(int numrows, HeapTuple *rows,
+ int2vector *attrs,
+ VacAttrStats **stats);
+
+static MVBucket select_bucket_to_partition(int nbuckets, MVBucket *buckets);
+
+static MVBucket partition_bucket(MVBucket bucket, int2vector *attrs,
+ VacAttrStats **stats,
+ int *ndistvalues, Datum **distvalues);
+
+static MVBucket copy_mv_bucket(MVBucket bucket, uint32 ndimensions);
+
+static void update_bucket_ndistinct(MVBucket bucket, int2vector *attrs,
+ VacAttrStats **stats);
+
+static void update_dimension_ndistinct(MVBucket bucket, int dimension,
+ int2vector *attrs,
+ VacAttrStats **stats,
+ bool update_boundaries);
+
+static void create_null_buckets(MVHistogram histogram, int bucket_idx,
+ int2vector *attrs, VacAttrStats **stats);
+
+static Datum *build_ndistinct(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int i, int *nvals);
+
+/*
+ * Each serialized bucket needs to store (in this order):
+ *
+ * - number of tuples (float)
+ * - number of distinct (float)
+ * - min inclusive flags (ndim * sizeof(bool))
+ * - max inclusive flags (ndim * sizeof(bool))
+ * - null dimension flags (ndim * sizeof(bool))
+ * - min boundary indexes (2 * ndim * sizeof(uint16))
+ * - max boundary indexes (2 * ndim * sizeof(uint16))
+ *
+ * So in total:
+ *
+ * ndim * (4 * sizeof(uint16) + 3 * sizeof(bool)) + (2 * sizeof(float))
+ */
+#define BUCKET_SIZE(ndims) \
+ (ndims * (4 * sizeof(uint16) + 3 * sizeof(bool)) + sizeof(float))
+
+/* pointers into a flat serialized bucket of BUCKET_SIZE(n) bytes */
+#define BUCKET_NTUPLES(b) (*(float*)b)
+#define BUCKET_MIN_INCL(b,n) ((bool*)(b + sizeof(float)))
+#define BUCKET_MAX_INCL(b,n) (BUCKET_MIN_INCL(b,n) + n)
+#define BUCKET_NULLS_ONLY(b,n) (BUCKET_MAX_INCL(b,n) + n)
+#define BUCKET_MIN_INDEXES(b,n) ((uint16*)(BUCKET_NULLS_ONLY(b,n) + n))
+#define BUCKET_MAX_INDEXES(b,n) ((BUCKET_MIN_INDEXES(b,n) + n))
+
+/* can't split bucket with less than 10 rows */
+#define MIN_BUCKET_ROWS 10
+
+/*
+ * Data used while building the histogram.
+ */
+typedef struct HistogramBuildData
+{
+
+ float ndistinct; /* frequency of distinct values */
+
+ HeapTuple *rows; /* aray of sample rows */
+ uint32 numrows; /* number of sample rows (array size) */
+
+ /*
+ * Number of distinct values in each dimension. This is used when building
+ * the histogram (and is not serialized/deserialized).
+ */
+ uint32 *ndistincts;
+
+} HistogramBuildData;
+
+typedef HistogramBuildData *HistogramBuild;
+
+/*
+ * builds a multivariate algorithm
+ *
+ * The build algorithm is iterative - initially a single bucket containing all
+ * the sample rows is formed, and then repeatedly split into smaller buckets.
+ * In each step the largest bucket (in some sense) is chosen to be split next.
+ *
+ * The criteria for selecting the largest bucket (and the dimension for the
+ * split) needs to be elaborate enough to produce buckets of roughly the same
+ * size, and also regular shape (not very long in one dimension).
+ *
+ * The current algorithm works like this:
+ *
+ * build NULL-buckets (create_null_buckets)
+ *
+ * while [maximum number of buckets not reached]
+ *
+ * choose bucket to partition (largest bucket)
+ * if no bucket to partition
+ * terminate the algorithm
+ *
+ * choose bucket dimension to partition (largest dimension)
+ * split the bucket into two buckets
+ *
+ * See the discussion at select_bucket_to_partition and partition_bucket for
+ * more details about the algorithm.
+ */
+MVHistogram
+build_mv_histogram(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int numrows_total)
+{
+ int i;
+ int numattrs = attrs->dim1;
+
+ int *ndistvalues;
+ Datum **distvalues;
+
+ MVHistogram histogram;
+
+ HeapTuple *rows_copy = (HeapTuple *) palloc0(numrows * sizeof(HeapTuple));
+
+ memcpy(rows_copy, rows, sizeof(HeapTuple) * numrows);
+
+ Assert((numattrs >= 2) && (numattrs <= MVSTATS_MAX_DIMENSIONS));
+
+ /* build histogram header */
+
+ histogram = (MVHistogram) palloc0(sizeof(MVHistogramData));
+
+ histogram->magic = MVSTAT_HIST_MAGIC;
+ histogram->type = MVSTAT_HIST_TYPE_BASIC;
+
+ histogram->nbuckets = 1;
+ histogram->ndimensions = numattrs;
+
+ /* create max buckets (better than repalloc for short-lived objects) */
+ histogram->buckets
+ = (MVBucket *) palloc0(MVSTAT_HIST_MAX_BUCKETS * sizeof(MVBucket));
+
+ /* create the initial bucket, covering the whole sample set */
+ histogram->buckets[0]
+ = create_initial_mv_bucket(numrows, rows_copy, attrs, stats);
+
+ /*
+ * Collect info on distinct values in each dimension (used later to select
+ * dimension to partition).
+ */
+ ndistvalues = (int *) palloc0(sizeof(int) * numattrs);
+ distvalues = (Datum **) palloc0(sizeof(Datum *) * numattrs);
+
+ for (i = 0; i < numattrs; i++)
+ distvalues[i] = build_ndistinct(numrows, rows, attrs, stats, i,
+ &ndistvalues[i]);
+
+ /*
+ * Split the initial bucket into buckets that don't mix NULL and non-NULL
+ * values in a single dimension.
+ */
+ create_null_buckets(histogram, 0, attrs, stats);
+
+ /*
+ * Do the actual histogram build - select a bucket and split it.
+ */
+ while (histogram->nbuckets < MVSTAT_HIST_MAX_BUCKETS)
+ {
+ MVBucket bucket = select_bucket_to_partition(histogram->nbuckets,
+ histogram->buckets);
+
+ /* no buckets eligible for partitioning */
+ if (bucket == NULL)
+ break;
+
+ /* we modify the bucket in-place and add one new bucket */
+ histogram->buckets[histogram->nbuckets++]
+ = partition_bucket(bucket, attrs, stats, ndistvalues, distvalues);
+ }
+
+ /* finalize the histogram build - compute the frequencies etc. */
+ for (i = 0; i < histogram->nbuckets; i++)
+ {
+ HistogramBuild build_data
+ = ((HistogramBuild) histogram->buckets[i]->build_data);
+
+ /*
+ * The frequency has to be computed from the whole sample, in case
+ * some of the rows were used for MCV.
+ *
+ * XXX Perhaps this should simply compute frequency with respect to
+ * the local freuquency, and then factor-in the MCV later.
+ *
+ * FIXME The 'ntuples' sounds a bit inappropriate for frequency.
+ */
+ histogram->buckets[i]->ntuples
+ = (build_data->numrows * 1.0) / numrows_total;
+ }
+
+ return histogram;
+}
+
+/* build array of distinct values for a single attribute */
+static Datum *
+build_ndistinct(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int i, int *nvals)
+{
+ int j;
+ int nvalues,
+ ndistinct;
+ Datum *values,
+ *distvalues;
+
+ SortSupportData ssup;
+ StdAnalyzeData *mystats = (StdAnalyzeData *) stats[i]->extra_data;
+
+ /* initialize sort support, etc. */
+ memset(&ssup, 0, sizeof(ssup));
+ ssup.ssup_cxt = CurrentMemoryContext;
+
+ /* We always use the default collation for statistics */
+ ssup.ssup_collation = DEFAULT_COLLATION_OID;
+ ssup.ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(mystats->ltopr, &ssup);
+
+ nvalues = 0;
+ values = (Datum *) palloc0(sizeof(Datum) * numrows);
+
+ /* collect values from the sample rows, ignore NULLs */
+ for (j = 0; j < numrows; j++)
+ {
+ Datum value;
+ bool isnull;
+
+ /*
+ * remember the index of the sample row, to make the partitioning
+ * simpler
+ */
+ value = heap_getattr(rows[j], attrs->values[i],
+ stats[i]->tupDesc, &isnull);
+
+ if (isnull)
+ continue;
+
+ values[nvalues++] = value;
+ }
+
+ /* if no non-NULL values were found, free the memory and terminate */
+ if (nvalues == 0)
+ {
+ pfree(values);
+ return NULL;
+ }
+
+ /* sort the array of values using the SortSupport */
+ qsort_arg((void *) values, nvalues, sizeof(Datum),
+ compare_scalars_simple, (void *) &ssup);
+
+ /* count the distinct values first, and allocate just enough memory */
+ ndistinct = 1;
+ for (j = 1; j < nvalues; j++)
+ if (compare_scalars_simple(&values[j], &values[j - 1], &ssup) != 0)
+ ndistinct += 1;
+
+ distvalues = (Datum *) palloc0(sizeof(Datum) * ndistinct);
+
+ /* now collect distinct values into the array */
+ distvalues[0] = values[0];
+ ndistinct = 1;
+
+ for (j = 1; j < nvalues; j++)
+ {
+ if (compare_scalars_simple(&values[j], &values[j - 1], &ssup) != 0)
+ {
+ distvalues[ndistinct] = values[j];
+ ndistinct += 1;
+ }
+ }
+
+ pfree(values);
+
+ *nvals = ndistinct;
+ return distvalues;
+}
+
+/* fetch the histogram (as a bytea) from the pg_mv_statistic catalog */
+MVSerializedHistogram
+load_mv_histogram(Oid mvoid)
+{
+ bool isnull = false;
+ Datum histogram;
+
+#ifdef USE_ASSERT_CHECKING
+ Form_pg_mv_statistic mvstat;
+#endif
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ HeapTuple htup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(mvoid));
+
+ if (!HeapTupleIsValid(htup))
+ return NULL;
+
+#ifdef USE_ASSERT_CHECKING
+ mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
+ Assert(mvstat->hist_enabled && mvstat->hist_built);
+#endif
+
+ histogram = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_stahist, &isnull);
+
+ Assert(!isnull);
+
+ ReleaseSysCache(htup);
+
+ return deserialize_mv_histogram(DatumGetByteaP(histogram));
+}
+
+/* print some basic info about the histogram */
+Datum
+pg_mv_stats_histogram_info(PG_FUNCTION_ARGS)
+{
+ bytea *data = PG_GETARG_BYTEA_P(0);
+ char *result;
+
+ MVSerializedHistogram hist = deserialize_mv_histogram(data);
+
+ result = palloc0(128);
+ snprintf(result, 128, "nbuckets=%d", hist->nbuckets);
+
+ PG_RETURN_TEXT_P(cstring_to_text(result));
+}
+
+/*
+ * Serialize the MV histogram into a bytea value. The basic algorithm is quite
+ * simple, and mostly mimincs the MCV serialization:
+ *
+ * (1) perform deduplication for each attribute (separately)
+ *
+ * (a) collect all (non-NULL) attribute values from all buckets
+ * (b) sort the data (using 'lt' from VacAttrStats)
+ * (c) remove duplicate values from the array
+ *
+ * (2) serialize the arrays into a bytea value
+ *
+ * (3) process all buckets
+ *
+ * (a) replace min/max values with indexes into the arrays
+ *
+ * Each attribute has to be processed separately, as we're mixing different
+ * datatypes, and we we need to use the right operators to compare/sort them.
+ * We're also mixing pass-by-value and pass-by-ref types, and so on.
+ *
+ *
+ * FIXME This probably leaks memory, or at least uses it inefficiently
+ * (many small palloc() calls instead of a large one).
+ *
+ * TODO Consider packing boolean flags (NULL) for each item into 'char' or
+ * a longer type (instead of using an array of bool items).
+ */
+bytea *
+serialize_mv_histogram(MVHistogram histogram, int2vector *attrs,
+ VacAttrStats **stats)
+{
+ int i = 0,
+ j = 0;
+ Size total_length = 0;
+
+ bytea *output = NULL;
+ char *data = NULL;
+
+ DimensionInfo *info;
+ SortSupport ssup;
+
+ int nbuckets = histogram->nbuckets;
+ int ndims = histogram->ndimensions;
+
+ /* allocated for serialized bucket data */
+ int bucketsize = BUCKET_SIZE(ndims);
+ char *bucket = palloc0(bucketsize);
+
+ /* values per dimension (and number of non-NULL values) */
+ Datum **values = (Datum **) palloc0(sizeof(Datum *) * ndims);
+ int *counts = (int *) palloc0(sizeof(int) * ndims);
+
+ /* info about dimensions (for deserialize) */
+ info = (DimensionInfo *) palloc0(sizeof(DimensionInfo) * ndims);
+
+ /* sort support data */
+ ssup = (SortSupport) palloc0(sizeof(SortSupportData) * ndims);
+
+ /* collect and deduplicate values for each dimension separately */
+ for (i = 0; i < ndims; i++)
+ {
+ int count;
+ StdAnalyzeData *tmp = (StdAnalyzeData *) stats[i]->extra_data;
+
+ /* keep important info about the data type */
+ info[i].typlen = stats[i]->attrtype->typlen;
+ info[i].typbyval = stats[i]->attrtype->typbyval;
+
+ /*
+ * Allocate space for all min/max values, including NULLs (we won't
+ * use them, but we don't know how many are there), and then collect
+ * all non-NULL values.
+ */
+ values[i] = (Datum *) palloc0(sizeof(Datum) * nbuckets * 2);
+
+ for (j = 0; j < histogram->nbuckets; j++)
+ {
+ /* skip buckets where this dimension is NULL-only */
+ if (!histogram->buckets[j]->nullsonly[i])
+ {
+ values[i][counts[i]] = histogram->buckets[j]->min[i];
+ counts[i] += 1;
+
+ values[i][counts[i]] = histogram->buckets[j]->max[i];
+ counts[i] += 1;
+ }
+ }
+
+ /* there are just NULL values in this dimension */
+ if (counts[i] == 0)
+ continue;
+
+ /* sort and deduplicate */
+ ssup[i].ssup_cxt = CurrentMemoryContext;
+ ssup[i].ssup_collation = DEFAULT_COLLATION_OID;
+ ssup[i].ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(tmp->ltopr, &ssup[i]);
+
+ qsort_arg(values[i], counts[i], sizeof(Datum),
+ compare_scalars_simple, &ssup[i]);
+
+ /*
+ * Walk through the array and eliminate duplicitate values, but keep
+ * the ordering (so that we can do bsearch later). We know there's at
+ * least 1 item, so we can skip the first element.
+ */
+ count = 1; /* number of deduplicated items */
+ for (j = 1; j < counts[i]; j++)
+ {
+ /* if it's different from the previous value, we need to keep it */
+ if (compare_datums_simple(values[i][j - 1], values[i][j], &ssup[i]) != 0)
+ {
+ /* XXX: not needed if (count == j) */
+ values[i][count] = values[i][j];
+ count += 1;
+ }
+ }
+
+ /* make sure we fit into uint16 */
+ Assert(count <= UINT16_MAX);
+
+ /* keep info about the deduplicated count */
+ info[i].nvalues = count;
+
+ /* compute size of the serialized data */
+ if (info[i].typlen > 0)
+ /* byval or byref, but with fixed length (name, tid, ...) */
+ info[i].nbytes = info[i].nvalues * info[i].typlen;
+ else if (info[i].typlen == -1)
+ /* varlena, so just use VARSIZE_ANY */
+ for (j = 0; j < info[i].nvalues; j++)
+ info[i].nbytes += VARSIZE_ANY(values[i][j]);
+ else if (info[i].typlen == -2)
+ /* cstring, so simply strlen */
+ for (j = 0; j < info[i].nvalues; j++)
+ info[i].nbytes += strlen(DatumGetPointer(values[i][j]));
+ else
+ elog(ERROR, "unknown data type typbyval=%d typlen=%d",
+ info[i].typbyval, info[i].typlen);
+ }
+
+ /*
+ * Now we finally know how much space we'll need for the serialized
+ * histogram, as it contains these fields:
+ *
+ * - length (4B) for varlena - magic (4B) - type (4B) - ndimensions (4B) -
+ * nbuckets (4B) - info (ndim * sizeof(DimensionInfo) - arrays of values
+ * for each dimension - serialized buckets (nbuckets * bucketsize)
+ *
+ * So the 'header' size is 20B + ndim * sizeof(DimensionInfo) and then
+ * we'll place the data (and buckets).
+ */
+ total_length = (sizeof(int32) + offsetof(MVHistogramData, buckets)
+ +ndims * sizeof(DimensionInfo)
+ + nbuckets * bucketsize);
+
+ /* account for the deduplicated data */
+ for (i = 0; i < ndims; i++)
+ total_length += info[i].nbytes;
+
+ /* enforce arbitrary limit of 1MB */
+ if (total_length > (1024 * 1024))
+ elog(ERROR, "serialized histogram exceeds 1MB (%ld > %d)",
+ total_length, (1024 * 1024));
+
+ /* allocate space for the serialized histogram list, set header */
+ output = (bytea *) palloc0(total_length);
+ SET_VARSIZE(output, total_length);
+
+ /* we'll use 'data' to keep track of the place to write data */
+ data = VARDATA(output);
+
+ memcpy(data, histogram, offsetof(MVHistogramData, buckets));
+ data += offsetof(MVHistogramData, buckets);
+
+ memcpy(data, info, sizeof(DimensionInfo) * ndims);
+ data += sizeof(DimensionInfo) * ndims;
+
+ /* serialize the deduplicated values for all attributes */
+ for (i = 0; i < ndims; i++)
+ {
+#ifdef USE_ASSERT_CHECKING
+ char *tmp = data;
+#endif
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ Datum v = values[i][j];
+
+ if (info[i].typbyval) /* passed by value */
+ {
+ memcpy(data, &v, info[i].typlen);
+ data += info[i].typlen;
+ }
+ else if (info[i].typlen > 0) /* pased by reference */
+ {
+ memcpy(data, DatumGetPointer(v), info[i].typlen);
+ data += info[i].typlen;
+ }
+ else if (info[i].typlen == -1) /* varlena */
+ {
+ memcpy(data, DatumGetPointer(v), VARSIZE_ANY(v));
+ data += VARSIZE_ANY(values[i][j]);
+ }
+ else if (info[i].typlen == -2) /* cstring */
+ {
+ memcpy(data, DatumGetPointer(v), strlen(DatumGetPointer(v)) + 1);
+ data += strlen(DatumGetPointer(v)) + 1;
+ }
+ }
+
+ /* make sure we got exactly the amount of data we expected */
+ Assert((data - tmp) == info[i].nbytes);
+ }
+
+ /* finally serialize the items, with uint16 indexes instead of the values */
+ for (i = 0; i < nbuckets; i++)
+ {
+ /* don't write beyond the allocated space */
+ Assert(data <= (char *) output + total_length - bucketsize);
+
+ /* reset the values for each item */
+ memset(bucket, 0, bucketsize);
+
+ BUCKET_NTUPLES(bucket) = histogram->buckets[i]->ntuples;
+
+ for (j = 0; j < ndims; j++)
+ {
+ /* do the lookup only for non-NULL values */
+ if (!histogram->buckets[i]->nullsonly[j])
+ {
+ uint16 idx;
+ Datum *v = NULL;
+
+ /* min boundary */
+ v = (Datum *) bsearch_arg(&histogram->buckets[i]->min[j],
+ values[j], info[j].nvalues, sizeof(Datum),
+ compare_scalars_simple, &ssup[j]);
+
+ Assert(v != NULL); /* serialization or deduplication
+ * error */
+
+ /* compute index within the array */
+ idx = (v - values[j]);
+
+ Assert((idx >= 0) && (idx < info[j].nvalues));
+
+ BUCKET_MIN_INDEXES(bucket, ndims)[j] = idx;
+
+ /* max boundary */
+ v = (Datum *) bsearch_arg(&histogram->buckets[i]->max[j],
+ values[j], info[j].nvalues, sizeof(Datum),
+ compare_scalars_simple, &ssup[j]);
+
+ Assert(v != NULL); /* serialization or deduplication
+ * error */
+
+ /* compute index within the array */
+ idx = (v - values[j]);
+
+ Assert((idx >= 0) && (idx < info[j].nvalues));
+
+ BUCKET_MAX_INDEXES(bucket, ndims)[j] = idx;
+ }
+ }
+
+ /* copy flags (nulls, min/max inclusive) */
+ memcpy(BUCKET_NULLS_ONLY(bucket, ndims),
+ histogram->buckets[i]->nullsonly, sizeof(bool) * ndims);
+
+ memcpy(BUCKET_MIN_INCL(bucket, ndims),
+ histogram->buckets[i]->min_inclusive, sizeof(bool) * ndims);
+
+ memcpy(BUCKET_MAX_INCL(bucket, ndims),
+ histogram->buckets[i]->max_inclusive, sizeof(bool) * ndims);
+
+ /* copy the item into the array */
+ memcpy(data, bucket, bucketsize);
+
+ data += bucketsize;
+ }
+
+ /* at this point we expect to match the total_length exactly */
+ Assert((data - (char *) output) == total_length);
+
+ /* free the values/counts arrays here */
+ pfree(counts);
+ pfree(info);
+ pfree(ssup);
+
+ for (i = 0; i < ndims; i++)
+ pfree(values[i]);
+
+ pfree(values);
+
+ return output;
+}
+
+/*
+ * Returns histogram in a partially-serialized form (keeps the boundary values
+ * deduplicated, so that it's possible to optimize the estimation part by
+ * caching function call results between buckets etc.).
+ */
+MVSerializedHistogram
+deserialize_mv_histogram(bytea *data)
+{
+ int i = 0,
+ j = 0;
+
+ Size expected_size;
+ char *tmp = NULL;
+
+ MVSerializedHistogram histogram;
+ DimensionInfo *info;
+
+ int nbuckets;
+ int ndims;
+ int bucketsize;
+
+ /* temporary deserialization buffer */
+ int bufflen;
+ char *buff;
+ char *ptr;
+
+ if (data == NULL)
+ return NULL;
+
+ if (VARSIZE_ANY_EXHDR(data) < offsetof(MVSerializedHistogramData, buckets))
+ elog(ERROR, "invalid histogram size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), offsetof(MVSerializedHistogramData, buckets));
+
+ /* read the histogram header */
+ histogram
+ = (MVSerializedHistogram) palloc(sizeof(MVSerializedHistogramData));
+
+ /* initialize pointer to the data part (skip the varlena header) */
+ tmp = VARDATA_ANY(data);
+
+ /* get the header and perform basic sanity checks */
+ memcpy(histogram, tmp, offsetof(MVSerializedHistogramData, buckets));
+ tmp += offsetof(MVSerializedHistogramData, buckets);
+
+ if (histogram->magic != MVSTAT_HIST_MAGIC)
+ elog(ERROR, "invalid histogram magic %d (expected %dd)",
+ histogram->magic, MVSTAT_HIST_MAGIC);
+
+ if (histogram->type != MVSTAT_HIST_TYPE_BASIC)
+ elog(ERROR, "invalid histogram type %d (expected %dd)",
+ histogram->type, MVSTAT_HIST_TYPE_BASIC);
+
+ nbuckets = histogram->nbuckets;
+ ndims = histogram->ndimensions;
+ bucketsize = BUCKET_SIZE(ndims);
+
+ Assert((nbuckets > 0) && (nbuckets <= MVSTAT_HIST_MAX_BUCKETS));
+ Assert((ndims >= 2) && (ndims <= MVSTATS_MAX_DIMENSIONS));
+
+ /*
+ * What size do we expect with those parameters (it's incomplete, as we
+ * yet have to count the array sizes (from DimensionInfo records).
+ */
+ expected_size = offsetof(MVSerializedHistogramData, buckets) +
+ ndims * sizeof(DimensionInfo) +
+ (nbuckets * bucketsize);
+
+ /* check that we have at least the DimensionInfo records */
+ if (VARSIZE_ANY_EXHDR(data) < expected_size)
+ elog(ERROR, "invalid histogram size %ld (expected %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ info = (DimensionInfo *) (tmp);
+ tmp += ndims * sizeof(DimensionInfo);
+
+ /* account for the value arrays */
+ for (i = 0; i < ndims; i++)
+ expected_size += info[i].nbytes;
+
+ if (VARSIZE_ANY_EXHDR(data) != expected_size)
+ elog(ERROR, "invalid histogram size %ld (expected %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ /* looks OK - not corrupted or something */
+
+ /* a single buffer for all the values and counts */
+ bufflen = (sizeof(int) + sizeof(Datum *)) * ndims;
+
+ for (i = 0; i < ndims; i++)
+ /* don't allocate space for byval types, matching Datum */
+ if (!(info[i].typbyval && (info[i].typlen == sizeof(Datum))))
+ bufflen += (sizeof(Datum) * info[i].nvalues);
+
+ /* also, include space for the result, tracking the buckets */
+ bufflen += nbuckets * (
+ sizeof(MVSerializedBucket) + /* bucket pointer */
+ sizeof(MVSerializedBucketData)); /* bucket data */
+
+ buff = palloc0(bufflen);
+ ptr = buff;
+
+ histogram->nvalues = (int *) ptr;
+ ptr += (sizeof(int) * ndims);
+
+ histogram->values = (Datum **) ptr;
+ ptr += (sizeof(Datum *) * ndims);
+
+ /*
+ * FIXME This uses pointers to the original data array (the types not
+ * passed by value), so when someone frees the memory, e.g. by doing
+ * something like this:
+ *
+ * bytea * data = ... fetch the data from catalog ... MCVList mcvlist =
+ * deserialize_mcv_list(data); pfree(data);
+ *
+ * then 'mcvlist' references the freed memory. This needs to copy the
+ * pieces.
+ *
+ * TODO same as in MCV deserialization / consider moving to common.c
+ */
+ for (i = 0; i < ndims; i++)
+ {
+ histogram->nvalues[i] = info[i].nvalues;
+
+ if (info[i].typbyval)
+ {
+ /* passed by value / Datum - simply reuse the array */
+ if (info[i].typlen == sizeof(Datum))
+ {
+ histogram->values[i] = (Datum *) tmp;
+ tmp += info[i].nbytes;
+ }
+ else
+ {
+ histogram->values[i] = (Datum *) ptr;
+ ptr += (sizeof(Datum) * info[i].nvalues);
+
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ memcpy(&histogram->values[i][j], tmp, info[i].typlen);
+ tmp += info[i].typlen;
+ }
+ }
+ }
+ else
+ {
+ /* all the other types need a chunk of the buffer */
+ histogram->values[i] = (Datum *) ptr;
+ ptr += (sizeof(Datum) * info[i].nvalues);
+
+ if (info[i].typlen > 0)
+ {
+ /* pased by reference, but fixed length (name, tid, ...) */
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ histogram->values[i][j] = PointerGetDatum(tmp);
+ tmp += info[i].typlen;
+ }
+ }
+ else if (info[i].typlen == -1)
+ {
+ /* varlena */
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ histogram->values[i][j] = PointerGetDatum(tmp);
+ tmp += VARSIZE_ANY(tmp);
+ }
+ }
+ else if (info[i].typlen == -2)
+ {
+ /* cstring */
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ histogram->values[i][j] = PointerGetDatum(tmp);
+ tmp += (strlen(tmp) + 1); /* don't forget the \0 */
+ }
+ }
+ }
+ }
+
+ histogram->buckets = (MVSerializedBucket *) ptr;
+ ptr += (sizeof(MVSerializedBucket) * nbuckets);
+
+ for (i = 0; i < nbuckets; i++)
+ {
+ MVSerializedBucket bucket = (MVSerializedBucket) ptr;
+
+ ptr += sizeof(MVSerializedBucketData);
+
+ bucket->ntuples = BUCKET_NTUPLES(tmp);
+ bucket->nullsonly = BUCKET_NULLS_ONLY(tmp, ndims);
+ bucket->min_inclusive = BUCKET_MIN_INCL(tmp, ndims);
+ bucket->max_inclusive = BUCKET_MAX_INCL(tmp, ndims);
+
+ bucket->min = BUCKET_MIN_INDEXES(tmp, ndims);
+ bucket->max = BUCKET_MAX_INDEXES(tmp, ndims);
+
+ histogram->buckets[i] = bucket;
+
+ Assert(tmp <= (char *) data + VARSIZE_ANY(data));
+
+ tmp += bucketsize;
+ }
+
+ /* at this point we expect to match the total_length exactly */
+ Assert((tmp - VARDATA(data)) == expected_size);
+
+ /* we should exhaust the output buffer exactly */
+ Assert((ptr - buff) == bufflen);
+
+ return histogram;
+}
+
+/*
+ * Build the initial bucket, which will be then split into smaller ones.
+ */
+static MVBucket
+create_initial_mv_bucket(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats)
+{
+ int i;
+ int numattrs = attrs->dim1;
+ HistogramBuild data = NULL;
+
+ /* TODO allocate bucket as a single piece, including all the fields. */
+ MVBucket bucket = (MVBucket) palloc0(sizeof(MVBucketData));
+
+ Assert(numrows > 0);
+ Assert(rows != NULL);
+ Assert((numattrs >= 2) && (numattrs <= MVSTATS_MAX_DIMENSIONS));
+
+ /* allocate the per-dimension arrays */
+
+ /* flags for null-only dimensions */
+ bucket->nullsonly = (bool *) palloc0(numattrs * sizeof(bool));
+
+ /* inclusiveness boundaries - lower/upper bounds */
+ bucket->min_inclusive = (bool *) palloc0(numattrs * sizeof(bool));
+ bucket->max_inclusive = (bool *) palloc0(numattrs * sizeof(bool));
+
+ /* lower/upper boundaries */
+ bucket->min = (Datum *) palloc0(numattrs * sizeof(Datum));
+ bucket->max = (Datum *) palloc0(numattrs * sizeof(Datum));
+
+ /* build-data */
+ data = (HistogramBuild) palloc0(sizeof(HistogramBuildData));
+
+ /* number of distinct values (per dimension) */
+ data->ndistincts = (uint32 *) palloc0(numattrs * sizeof(uint32));
+
+ /* all the sample rows fall into the initial bucket */
+ data->numrows = numrows;
+ data->rows = rows;
+
+ bucket->build_data = data;
+
+ /*
+ * Update the number of ndistinct combinations in the bucket (which we use
+ * when selecting bucket to partition), and then number of distinct values
+ * for each partition (which we use when choosing which dimension to
+ * split).
+ */
+ update_bucket_ndistinct(bucket, attrs, stats);
+
+ /* Update ndistinct (and also set min/max) for all dimensions. */
+ for (i = 0; i < numattrs; i++)
+ update_dimension_ndistinct(bucket, i, attrs, stats, true);
+
+ return bucket;
+}
+
+/*
+ * Choose the bucket to partition next.
+ *
+ * The current criteria is rather simple, chosen so that the algorithm produces
+ * buckets with about equal frequency and regular size. We select the bucket
+ * with the highest number of distinct values, and then split it by the longest
+ * dimension.
+ *
+ * The distinct values are uniformly mapped to [0,1] interval, and this is used
+ * to compute length of the value range.
+ *
+ * NOTE: This is not the same array used for deduplication, as this contains
+ * values for all the tuples from the sample, not just the boundary values.
+ *
+ * Returns either pointer to the bucket selected to be partitioned, or NULL if
+ * there are no buckets that may be split (e.g. if all buckets are too small
+ * or contain too few distinct values).
+ *
+ *
+ * Tricky example
+ * --------------
+ *
+ * Consider this table:
+ *
+ * CREATE TABLE t AS SELECT i AS a, i AS b
+ * FROM generate_series(1,1000000) s(i);
+ *
+ * CREATE STATISTICS s1 ON t (a,b) WITH (histogram);
+ *
+ * ANALYZE t;
+ *
+ * It's a very specific (and perhaps artificial) example, because every bucket
+ * always has exactly the same number of distinct values in all dimensions,
+ * which makes the partitioning tricky.
+ *
+ * Then:
+ *
+ * SELECT * FROM t WHERE (a < 100) AND (b < 100);
+ *
+ * is estimated to return ~120 rows, while in reality it returns only 99.
+ *
+ * QUERY PLAN
+ * -------------------------------------------------------------
+ * Seq Scan on t (cost=0.00..19425.00 rows=117 width=8)
+ * (actual time=0.129..82.776 rows=99 loops=1)
+ * Filter: ((a < 100) AND (b < 100))
+ * Rows Removed by Filter: 999901
+ * Planning time: 1.286 ms
+ * Execution time: 82.984 ms
+ * (5 rows)
+ *
+ * So this estimate is reasonably close. Let's change the query to OR clause:
+ *
+ * SELECT * FROM t WHERE (a < 100) OR (b < 100);
+ *
+ * QUERY PLAN
+ * -------------------------------------------------------------
+ * Seq Scan on t (cost=0.00..19425.00 rows=8100 width=8)
+ * (actual time=0.145..99.910 rows=99 loops=1)
+ * Filter: ((a < 100) OR (b < 100))
+ * Rows Removed by Filter: 999901
+ * Planning time: 1.578 ms
+ * Execution time: 100.132 ms
+ * (5 rows)
+ *
+ * That's clearly a much worse estimate. This happens because the histogram
+ * contains buckets like this:
+ *
+ * bucket 592 [3 30310] [30134 30593] => [0.000233]
+ *
+ * i.e. the length of "a" dimension is (30310-3)=30307, while the length of "b"
+ * is (30593-30134)=459. So the "b" dimension is much narrower than "a".
+ * Of course, there are also buckets where "b" is the wider dimension.
+ *
+ * This is partially mitigated by selecting the "longest" dimension but that
+ * only happens after we already selected the bucket. So if we never select the
+ * bucket, this optimization does not apply.
+ *
+ * The other reason why this particular example behaves so poorly is due to the
+ * way we actually split the selected bucket. We do attempt to divide the bucket
+ * into two parts containing about the same number of tuples, but that does not
+ * too well when most of the tuples is squashed on one side of the bucket.
+ *
+ * For example for columns with data on the diagonal (i.e. when a=b), we end up
+ * with a narrow bucket on the diagonal and a huge bucket overing the remaining
+ * part (with much lower density).
+ *
+ * So perhaps we need two partitioning strategies - one aiming to split buckets
+ * with high frequency (number of sampled rows), the other aiming to split
+ * "large" buckets. And alternating between them, somehow.
+ *
+ * TODO Consider using similar lower boundary for row count as for simple
+ * histograms, i.e. 300 tuples per bucket.
+ */
+static MVBucket
+select_bucket_to_partition(int nbuckets, MVBucket *buckets)
+{
+ int i;
+ int numrows = 0;
+ MVBucket bucket = NULL;
+
+ for (i = 0; i < nbuckets; i++)
+ {
+ HistogramBuild data = (HistogramBuild) buckets[i]->build_data;
+
+ /* if the number of rows is higher, use this bucket */
+ if ((data->ndistinct > 2) &&
+ (data->numrows > numrows) &&
+ (data->numrows >= MIN_BUCKET_ROWS))
+ {
+ bucket = buckets[i];
+ numrows = data->numrows;
+ }
+ }
+
+ /* may be NULL if there are not buckets with (ndistinct>1) */
+ return bucket;
+}
+
+/*
+ * A simple bucket partitioning implementation - we choose the longest bucket
+ * dimension, measured using the array of distinct values built at the very
+ * beginning of the build.
+ *
+ * We map all the distinct values to a [0,1] interval, uniformly distributed,
+ * and then use this to measure length. It's essentially a number of distinct
+ * values within the range, normalized to [0,1].
+ *
+ * Then we choose a 'middle' value splitting the bucket into two parts with
+ * roughly the same frequency.
+ *
+ * This splits the bucket by tweaking the existing one, and returning the new
+ * bucket (essentially shrinking the existing one in-place and returning the
+ * other "half" as a new bucket). The caller is responsible for adding the new
+ * bucket into the list of buckets.
+ *
+ * There are multiple histogram options, centered around the partitioning
+ * criteria, specifying both how to choose a bucket and the dimension most in
+ * need of a split. For a nice summary and general overview, see "rK-Hist : an
+ * R-Tree based histogram for multi-dimensional selectivity estimation" thesis
+ * by J. A. Lopez, Concordia University, p.34-37 (and possibly p. 32-34 for
+ * explanation of the terms).
+ *
+ * It requires care to prevent splitting only one dimension and not splitting
+ * another one at all (which might happen easily in case of strongly dependent
+ * columns - e.g. y=x). The current algorithm minimizes this, but may still
+ * happen for perfectly dependent examples (when all the dimensions have equal
+ * length, the first one will be selected).
+ *
+ * TODO Should probably consider statistics target for the columns (e.g.
+ * to split dimensions with higher statistics target more frequently).
+ */
+static MVBucket
+partition_bucket(MVBucket bucket, int2vector *attrs,
+ VacAttrStats **stats,
+ int *ndistvalues, Datum **distvalues)
+{
+ int i;
+ int dimension;
+ int numattrs = attrs->dim1;
+
+ Datum split_value;
+ MVBucket new_bucket;
+ HistogramBuild new_data;
+
+ /* needed for sort, when looking for the split value */
+ bool isNull;
+ int nvalues = 0;
+ HistogramBuild data = (HistogramBuild) bucket->build_data;
+ StdAnalyzeData *mystats = NULL;
+ ScalarItem *values = (ScalarItem *) palloc0(data->numrows * sizeof(ScalarItem));
+ SortSupportData ssup;
+
+ int nrows = 1; /* number of rows below current value */
+ double delta;
+
+ /* needed when splitting the values */
+ HeapTuple *oldrows = data->rows;
+ int oldnrows = data->numrows;
+
+ /*
+ * We can't split buckets with a single distinct value (this also
+ * disqualifies NULL-only dimensions). Also, there has to be multiple
+ * sample rows (otherwise, how could there be more distinct values).
+ */
+ Assert(data->ndistinct > 1);
+ Assert(data->numrows > 1);
+ Assert((numattrs >= 2) && (numattrs <= MVSTATS_MAX_DIMENSIONS));
+
+ /* Look for the next dimension to split. */
+ delta = 0.0;
+ dimension = -1;
+
+ for (i = 0; i < numattrs; i++)
+ {
+ Datum *a,
+ *b;
+
+ mystats = (StdAnalyzeData *) stats[i]->extra_data;
+
+ /* initialize sort support, etc. */
+ memset(&ssup, 0, sizeof(ssup));
+ ssup.ssup_cxt = CurrentMemoryContext;
+
+ /* We always use the default collation for statistics */
+ ssup.ssup_collation = DEFAULT_COLLATION_OID;
+ ssup.ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(mystats->ltopr, &ssup);
+
+ /* can't split NULL-only dimension */
+ if (bucket->nullsonly[i])
+ continue;
+
+ /* can't split dimension with a single ndistinct value */
+ if (data->ndistincts[i] <= 1)
+ continue;
+
+ /* search for min boundary in the distinct list */
+ a = (Datum *) bsearch_arg(&bucket->min[i],
+ distvalues[i], ndistvalues[i],
+ sizeof(Datum), compare_scalars_simple, &ssup);
+
+ b = (Datum *) bsearch_arg(&bucket->max[i],
+ distvalues[i], ndistvalues[i],
+ sizeof(Datum), compare_scalars_simple, &ssup);
+
+ /* if this dimension is 'larger' then partition by it */
+ if (((b - a) * 1.0 / ndistvalues[i]) > delta)
+ {
+ delta = ((b - a) * 1.0 / ndistvalues[i]);
+ dimension = i;
+ }
+ }
+
+ /*
+ * If we haven't found a dimension here, we've done something wrong in
+ * select_bucket_to_partition.
+ */
+ Assert(dimension != -1);
+
+ /*
+ * Walk through the selected dimension, collect and sort the values and
+ * then choose the value to use as the new boundary.
+ */
+ mystats = (StdAnalyzeData *) stats[dimension]->extra_data;
+
+ /* initialize sort support, etc. */
+ memset(&ssup, 0, sizeof(ssup));
+ ssup.ssup_cxt = CurrentMemoryContext;
+
+ /* We always use the default collation for statistics */
+ ssup.ssup_collation = DEFAULT_COLLATION_OID;
+ ssup.ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(mystats->ltopr, &ssup);
+
+ for (i = 0; i < data->numrows; i++)
+ {
+ /*
+ * remember the index of the sample row, to make the partitioning
+ * simpler
+ */
+ values[nvalues].value = heap_getattr(data->rows[i], attrs->values[dimension],
+ stats[dimension]->tupDesc, &isNull);
+ values[nvalues].tupno = i;
+
+ /* no NULL values allowed here (we never split null-only dimension) */
+ Assert(!isNull);
+
+ nvalues++;
+ }
+
+ /* sort the array of values */
+ qsort_arg((void *) values, nvalues, sizeof(ScalarItem),
+ compare_scalars_partition, (void *) &ssup);
+
+ /*
+ * We know there are bucket->ndistincts[dimension] distinct values in this
+ * dimension, and we want to split this into half, so walk through the
+ * array and stop once we see (ndistinct/2) values.
+ *
+ * We always choose the "next" value, i.e. (n/2+1)-th distinct value, and
+ * use it as an exclusive upper boundary (and inclusive lower boundary).
+ *
+ * TODO Maybe we should use "average" of the two middle distinct values
+ * (at least for even distinct counts), but that would require being able
+ * to do an average (which does not work for non-numeric types).
+ *
+ * TODO Another option is to look for a split that'd give about 50% tuples
+ * (not distinct values) in each partition. That might work better when
+ * there are a few very frequent values, and many rare ones.
+ */
+ delta = fabs(data->numrows);
+ split_value = values[0].value;
+
+ for (i = 1; i < data->numrows; i++)
+ {
+ if (values[i].value != values[i - 1].value)
+ {
+ /* are we closer to splitting the bucket in half? */
+ if (fabs(i - data->numrows / 2.0) < delta)
+ {
+ /* let's assume we'll use this value for the split */
+ split_value = values[i].value;
+ delta = fabs(i - data->numrows / 2.0);
+ nrows = i;
+ }
+ }
+ }
+
+ Assert(nrows > 0);
+ Assert(nrows < data->numrows);
+
+ /*
+ * create the new bucket as a (incomplete) copy of the one being
+ * partitioned.
+ */
+ new_bucket = copy_mv_bucket(bucket, numattrs);
+ new_data = (HistogramBuild) new_bucket->build_data;
+
+ /*
+ * Do the actual split of the chosen dimension, using the split value as
+ * the upper bound for the existing bucket, and lower bound for the new
+ * one.
+ */
+ bucket->max[dimension] = split_value;
+ new_bucket->min[dimension] = split_value;
+
+ /*
+ * We also treat only one side of the new boundary as inclusive, in the
+ * bucket where it happens to be the upper boundary. We never set the
+ * min_inclusive[] to false anywhere, but we set it to true anyway.
+ */
+ bucket->max_inclusive[dimension] = false;
+ new_bucket->min_inclusive[dimension] = true;
+
+ /*
+ * Redistribute the sample tuples using the 'ScalarItem->tupno' index. We
+ * know 'nrows' rows should remain in the original bucket and the rest
+ * goes to the new one.
+ */
+
+ data->rows = (HeapTuple *) palloc0(nrows * sizeof(HeapTuple));
+ new_data->rows = (HeapTuple *) palloc0((oldnrows - nrows) * sizeof(HeapTuple));
+
+ data->numrows = nrows;
+ new_data->numrows = (oldnrows - nrows);
+
+ /*
+ * The first nrows should go to the first bucket, the rest should go to
+ * the new one. Use the tupno field to get the actual HeapTuple row from
+ * the original array of sample rows.
+ */
+ for (i = 0; i < nrows; i++)
+ memcpy(&data->rows[i], &oldrows[values[i].tupno], sizeof(HeapTuple));
+
+ for (i = nrows; i < oldnrows; i++)
+ memcpy(&new_data->rows[i - nrows], &oldrows[values[i].tupno], sizeof(HeapTuple));
+
+ /* update ndistinct values for the buckets (total and per dimension) */
+ update_bucket_ndistinct(bucket, attrs, stats);
+ update_bucket_ndistinct(new_bucket, attrs, stats);
+
+ /*
+ * TODO We don't need to do this for the dimension we used for split,
+ * because we know how many distinct values went to each partition.
+ */
+ for (i = 0; i < numattrs; i++)
+ {
+ update_dimension_ndistinct(bucket, i, attrs, stats, false);
+ update_dimension_ndistinct(new_bucket, i, attrs, stats, false);
+ }
+
+ pfree(oldrows);
+ pfree(values);
+
+ return new_bucket;
+}
+
+/*
+ * Copy a histogram bucket. The copy does not include the build-time data, i.e.
+ * sampled rows etc.
+ */
+static MVBucket
+copy_mv_bucket(MVBucket bucket, uint32 ndimensions)
+{
+ /* TODO allocate as a single piece (including all the fields) */
+ MVBucket new_bucket = (MVBucket) palloc0(sizeof(MVBucketData));
+ HistogramBuild data = (HistogramBuild) palloc0(sizeof(HistogramBuildData));
+
+ /*
+ * Copy only the attributes that will stay the same after the split, and
+ * we'll recompute the rest after the split.
+ */
+
+ /* allocate the per-dimension arrays */
+ new_bucket->nullsonly = (bool *) palloc0(ndimensions * sizeof(bool));
+
+ /* inclusiveness boundaries - lower/upper bounds */
+ new_bucket->min_inclusive = (bool *) palloc0(ndimensions * sizeof(bool));
+ new_bucket->max_inclusive = (bool *) palloc0(ndimensions * sizeof(bool));
+
+ /* lower/upper boundaries */
+ new_bucket->min = (Datum *) palloc0(ndimensions * sizeof(Datum));
+ new_bucket->max = (Datum *) palloc0(ndimensions * sizeof(Datum));
+
+ /* copy data */
+ memcpy(new_bucket->nullsonly, bucket->nullsonly, ndimensions * sizeof(bool));
+
+ memcpy(new_bucket->min_inclusive, bucket->min_inclusive, ndimensions * sizeof(bool));
+ memcpy(new_bucket->min, bucket->min, ndimensions * sizeof(Datum));
+
+ memcpy(new_bucket->max_inclusive, bucket->max_inclusive, ndimensions * sizeof(bool));
+ memcpy(new_bucket->max, bucket->max, ndimensions * sizeof(Datum));
+
+ /* allocate and copy the interesting part of the build data */
+ data->ndistincts = (uint32 *) palloc0(ndimensions * sizeof(uint32));
+
+ new_bucket->build_data = data;
+
+ return new_bucket;
+}
+
+/*
+ * Counts the number of distinct values in the bucket. This just copies the
+ * Datum values into a simple array, and sorts them using memcmp-based
+ * comparator. That means it only works for pass-by-value data types (assuming
+ * they don't use collations etc.)
+ */
+static void
+update_bucket_ndistinct(MVBucket bucket, int2vector *attrs, VacAttrStats **stats)
+{
+ int i,
+ j;
+ int numattrs = attrs->dim1;
+
+ HistogramBuild data = (HistogramBuild) bucket->build_data;
+ int numrows = data->numrows;
+
+ MultiSortSupport mss = multi_sort_init(numattrs);
+
+ /*
+ * We could collect this while walking through all the attributes above
+ * (this way we have to call heap_getattr twice).
+ */
+ SortItem *items = (SortItem *) palloc0(numrows * sizeof(SortItem));
+ Datum *values = (Datum *) palloc0(numrows * sizeof(Datum) * numattrs);
+ bool *isnull = (bool *) palloc0(numrows * sizeof(bool) * numattrs);
+
+ for (i = 0; i < numrows; i++)
+ {
+ items[i].values = &values[i * numattrs];
+ items[i].isnull = &isnull[i * numattrs];
+ }
+
+ /* prepare the sort function for the first dimension */
+ for (i = 0; i < numattrs; i++)
+ multi_sort_add_dimension(mss, i, i, stats);
+
+ /* collect the values */
+ for (i = 0; i < numrows; i++)
+ for (j = 0; j < numattrs; j++)
+ items[i].values[j]
+ = heap_getattr(data->rows[i], attrs->values[j],
+ stats[j]->tupDesc, &items[i].isnull[j]);
+
+ qsort_arg((void *) items, numrows, sizeof(SortItem),
+ multi_sort_compare, mss);
+
+ data->ndistinct = 1;
+
+ for (i = 1; i < numrows; i++)
+ if (multi_sort_compare(&items[i], &items[i - 1], mss) != 0)
+ data->ndistinct += 1;
+
+ pfree(items);
+ pfree(values);
+ pfree(isnull);
+}
+
+/*
+ * Count distinct values per bucket dimension.
+ */
+static void
+update_dimension_ndistinct(MVBucket bucket, int dimension, int2vector *attrs,
+ VacAttrStats **stats, bool update_boundaries)
+{
+ int j;
+ int nvalues = 0;
+ bool isNull;
+ HistogramBuild data = (HistogramBuild) bucket->build_data;
+ Datum *values = (Datum *) palloc0(data->numrows * sizeof(Datum));
+ SortSupportData ssup;
+
+ StdAnalyzeData *mystats = (StdAnalyzeData *) stats[dimension]->extra_data;
+
+ /* we may already know this is a NULL-only dimension */
+ if (bucket->nullsonly[dimension])
+ data->ndistincts[dimension] = 1;
+
+ memset(&ssup, 0, sizeof(ssup));
+ ssup.ssup_cxt = CurrentMemoryContext;
+
+ /* We always use the default collation for statistics */
+ ssup.ssup_collation = DEFAULT_COLLATION_OID;
+ ssup.ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(mystats->ltopr, &ssup);
+
+ for (j = 0; j < data->numrows; j++)
+ {
+ values[nvalues] = heap_getattr(data->rows[j], attrs->values[dimension],
+ stats[dimension]->tupDesc, &isNull);
+
+ /* ignore NULL values */
+ if (!isNull)
+ nvalues++;
+ }
+
+ /* there's always at least 1 distinct value (may be NULL) */
+ data->ndistincts[dimension] = 1;
+
+ /*
+ * if there are only NULL values in the column, mark it so and continue
+ * with the next one
+ */
+ if (nvalues == 0)
+ {
+ pfree(values);
+ bucket->nullsonly[dimension] = true;
+ return;
+ }
+
+ /* sort the array (pass-by-value datum */
+ qsort_arg((void *) values, nvalues, sizeof(Datum),
+ compare_scalars_simple, (void *) &ssup);
+
+ /*
+ * Update min/max boundaries to the smallest bounding box. Generally, this
+ * needs to be done only when constructing the initial bucket.
+ */
+ if (update_boundaries)
+ {
+ /* store the min/max values */
+ bucket->min[dimension] = values[0];
+ bucket->min_inclusive[dimension] = true;
+
+ bucket->max[dimension] = values[nvalues - 1];
+ bucket->max_inclusive[dimension] = true;
+ }
+
+ /*
+ * Walk through the array and count distinct values by comparing
+ * succeeding values.
+ *
+ * FIXME This only works for pass-by-value types (i.e. not VARCHARs etc.).
+ * Although thanks to the deduplication it might work even for those types
+ * (equal values will get the same item in the deduplicated array).
+ */
+ for (j = 1; j < nvalues; j++)
+ {
+ if (values[j] != values[j - 1])
+ data->ndistincts[dimension] += 1;
+ }
+
+ pfree(values);
+}
+
+/*
+ * A properly built histogram must not contain buckets mixing NULL and non-NULL
+ * values in a single dimension. Each dimension may either be marked as 'nulls
+ * only', and thus containing only NULL values, or it must not contain any NULL
+ * values.
+ *
+ * Therefore, if the sample contains NULL values in any of the columns, it's
+ * necessary to build those NULL-buckets. This is done in an iterative way
+ * using this algorithm, operating on a single bucket:
+ *
+ * (1) Check that all dimensions are well-formed (not mixing NULL and
+ * non-NULL values).
+ *
+ * (2) If all dimensions are well-formed, terminate.
+ *
+ * (3) If the dimension contains only NULL values, but is not marked as
+ * NULL-only, mark it as NULL-only and run the algorithm again (on
+ * this bucket).
+ *
+ * (4) If the dimension mixes NULL and non-NULL values, split the bucket
+ * into two parts - one with NULL values, one with non-NULL values
+ * (replacing the current one). Then run the algorithm on both buckets.
+ *
+ * This is executed in a recursive manner, but the number of executions should
+ * be quite low - limited by the number of NULL-buckets. Also, in each branch
+ * the number of nested calls is limited by the number of dimensions
+ * (attributes) of the histogram.
+ *
+ * At the end, there should be buckets with no mixed dimensions. The number of
+ * buckets produced by this algorithm is rather limited - with N dimensions,
+ * there may be only 2^N such buckets (each dimension may be either NULL or
+ * non-NULL). So with 8 dimensions (current value of MVSTATS_MAX_DIMENSIONS)
+ * there may be only 256 such buckets.
+ *
+ * After this, a 'regular' bucket-split algorithm shall run, further optimizing
+ * the histogram.
+ */
+static void
+create_null_buckets(MVHistogram histogram, int bucket_idx,
+ int2vector *attrs, VacAttrStats **stats)
+{
+ int i,
+ j;
+ int null_dim = -1;
+ int null_count = 0;
+ bool null_found = false;
+ MVBucket bucket,
+ null_bucket;
+ int null_idx,
+ curr_idx;
+ HistogramBuild data,
+ null_data;
+
+ /* remember original values from the bucket */
+ int numrows;
+ HeapTuple *oldrows = NULL;
+
+ Assert(bucket_idx < histogram->nbuckets);
+ Assert(histogram->ndimensions == attrs->dim1);
+
+ bucket = histogram->buckets[bucket_idx];
+ data = (HistogramBuild) bucket->build_data;
+
+ numrows = data->numrows;
+ oldrows = data->rows;
+
+ /*
+ * Walk through all rows / dimensions, and stop once we find NULL in a
+ * dimension not yet marked as NULL-only.
+ */
+ for (i = 0; i < data->numrows; i++)
+ {
+ /*
+ * FIXME We don't need to start from the first attribute here - we can
+ * start from the last known dimension.
+ */
+ for (j = 0; j < histogram->ndimensions; j++)
+ {
+ /* Is this a NULL-only dimension? If yes, skip. */
+ if (bucket->nullsonly[j])
+ continue;
+
+ /* found a NULL in that dimension? */
+ if (heap_attisnull(data->rows[i], attrs->values[j]))
+ {
+ null_found = true;
+ null_dim = j;
+ break;
+ }
+ }
+
+ /* terminate if we found attribute with NULL values */
+ if (null_found)
+ break;
+ }
+
+ /* no regular dimension contains NULL values => we're done */
+ if (!null_found)
+ return;
+
+ /* walk through the rows again, count NULL values in 'null_dim' */
+ for (i = 0; i < data->numrows; i++)
+ {
+ if (heap_attisnull(data->rows[i], attrs->values[null_dim]))
+ null_count += 1;
+ }
+
+ Assert(null_count <= data->numrows);
+
+ /*
+ * If (null_count == numrows) the dimension already is NULL-only, but is
+ * not yet marked like that. It's enough to mark it and repeat the process
+ * recursively (until we run out of dimensions).
+ */
+ if (null_count == data->numrows)
+ {
+ bucket->nullsonly[null_dim] = true;
+ create_null_buckets(histogram, bucket_idx, attrs, stats);
+ return;
+ }
+
+ /*
+ * We have to split the bucket into two - one with NULL values in the
+ * dimension, one with non-NULL values. We don't need to sort the data or
+ * anything, but otherwise it's similar to what partition_bucket() does.
+ */
+
+ /* create bucket with NULL-only dimension 'dim' */
+ null_bucket = copy_mv_bucket(bucket, histogram->ndimensions);
+ null_data = (HistogramBuild) null_bucket->build_data;
+
+ /* remember the current array info */
+ oldrows = data->rows;
+ numrows = data->numrows;
+
+ /* we'll keep non-NULL values in the current bucket */
+ data->numrows = (numrows - null_count);
+ data->rows
+ = (HeapTuple *) palloc0(data->numrows * sizeof(HeapTuple));
+
+ /* and the NULL values will go to the new one */
+ null_data->numrows = null_count;
+ null_data->rows
+ = (HeapTuple *) palloc0(null_data->numrows * sizeof(HeapTuple));
+
+ /* mark the dimension as NULL-only (in the new bucket) */
+ null_bucket->nullsonly[null_dim] = true;
+
+ /* walk through the sample rows and distribute them accordingly */
+ null_idx = 0;
+ curr_idx = 0;
+ for (i = 0; i < numrows; i++)
+ {
+ if (heap_attisnull(oldrows[i], attrs->values[null_dim]))
+ /* NULL => copy to the new bucket */
+ memcpy(&null_data->rows[null_idx++], &oldrows[i],
+ sizeof(HeapTuple));
+ else
+ memcpy(&data->rows[curr_idx++], &oldrows[i],
+ sizeof(HeapTuple));
+ }
+
+ /* update ndistinct values for the buckets (total and per dimension) */
+ update_bucket_ndistinct(bucket, attrs, stats);
+ update_bucket_ndistinct(null_bucket, attrs, stats);
+
+ /*
+ * TODO We don't need to do this for the dimension we used for split,
+ * because we know how many distinct values went to each bucket (NULL is
+ * not a value, so NULL buckets get 0, and the other bucket got all the
+ * distinct values).
+ */
+ for (i = 0; i < histogram->ndimensions; i++)
+ {
+ update_dimension_ndistinct(bucket, i, attrs, stats, false);
+ update_dimension_ndistinct(null_bucket, i, attrs, stats, false);
+ }
+
+ pfree(oldrows);
+
+ /* add the NULL bucket to the histogram */
+ histogram->buckets[histogram->nbuckets++] = null_bucket;
+
+ /*
+ * And now run the function recursively on both buckets (the new one
+ * first, because the call may change number of buckets, and it's used as
+ * an index).
+ */
+ create_null_buckets(histogram, (histogram->nbuckets - 1), attrs, stats);
+ create_null_buckets(histogram, bucket_idx, attrs, stats);
+}
+
+/*
+ * SRF with details about buckets of a histogram:
+ *
+ * - bucket ID (0...nbuckets)
+ * - min values (string array)
+ * - max values (string array)
+ * - nulls only (boolean array)
+ * - min inclusive flags (boolean array)
+ * - max inclusive flags (boolean array)
+ * - frequency (double precision)
+ *
+ * The input is the OID of the statistics, and there are no rows returned if the
+ * statistics contains no histogram (or if there's no statistics for the OID).
+ *
+ * The second parameter (type) determines what values will be returned
+ * in the (minvals,maxvals). There are three possible values:
+ *
+ * 0 (actual values)
+ * -----------------
+ * - prints actual values
+ * - using the output function of the data type (as string)
+ * - handy for investigating the histogram
+ *
+ * 1 (distinct index)
+ * ------------------
+ * - prints index of the distinct value (into the serialized array)
+ * - makes it easier to spot neighbor buckets, etc.
+ * - handy for plotting the histogram
+ *
+ * 2 (normalized distinct index)
+ * -----------------------------
+ * - prints index of the distinct value, but normalized into [0,1]
+ * - similar to 1, but shows how 'long' the bucket range is
+ * - handy for plotting the histogram
+ *
+ * When plotting the histogram, be careful as the (1) and (2) options skew the
+ * lengths by distributing the distinct values uniformly. For data types
+ * without a clear meaning of 'distance' (e.g. strings) that is not a big deal,
+ * but for numbers it may be confusing.
+ */
+PG_FUNCTION_INFO_V1(pg_mv_histogram_buckets);
+
+#define OUTPUT_FORMAT_RAW 0
+#define OUTPUT_FORMAT_INDEXES 1
+#define OUTPUT_FORMAT_DISTINCT 2
+
+Datum
+pg_mv_histogram_buckets(PG_FUNCTION_ARGS)
+{
+ FuncCallContext *funcctx;
+ int call_cntr;
+ int max_calls;
+ TupleDesc tupdesc;
+ AttInMetadata *attinmeta;
+
+ Oid mvoid = PG_GETARG_OID(0);
+ int otype = PG_GETARG_INT32(1);
+
+ if ((otype < 0) || (otype > 2))
+ elog(ERROR, "invalid output type specified");
+
+ /* stuff done only on the first call of the function */
+ if (SRF_IS_FIRSTCALL())
+ {
+ MemoryContext oldcontext;
+ MVSerializedHistogram histogram;
+
+ /* create a function context for cross-call persistence */
+ funcctx = SRF_FIRSTCALL_INIT();
+
+ /* switch to memory context appropriate for multiple function calls */
+ oldcontext = MemoryContextSwitchTo(funcctx->multi_call_memory_ctx);
+
+ histogram = load_mv_histogram(mvoid);
+
+ funcctx->user_fctx = histogram;
+
+ /* total number of tuples to be returned */
+ funcctx->max_calls = 0;
+ if (funcctx->user_fctx != NULL)
+ funcctx->max_calls = histogram->nbuckets;
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("function returning record called in context "
+ "that cannot accept type record")));
+
+ /*
+ * generate attribute metadata needed later to produce tuples from raw
+ * C strings
+ */
+ attinmeta = TupleDescGetAttInMetadata(tupdesc);
+ funcctx->attinmeta = attinmeta;
+
+ MemoryContextSwitchTo(oldcontext);
+ }
+
+ /* stuff done on every call of the function */
+ funcctx = SRF_PERCALL_SETUP();
+
+ call_cntr = funcctx->call_cntr;
+ max_calls = funcctx->max_calls;
+ attinmeta = funcctx->attinmeta;
+
+ if (call_cntr < max_calls) /* do when there is more left to send */
+ {
+ char **values;
+ HeapTuple tuple;
+ Datum result;
+ int2vector *stakeys;
+ Oid relid;
+ double bucket_volume = 1.0;
+ StringInfo bufs;
+
+ char *format;
+ int i;
+
+ Oid *outfuncs;
+ FmgrInfo *fmgrinfo;
+
+ MVSerializedHistogram histogram;
+ MVSerializedBucket bucket;
+
+ histogram = (MVSerializedHistogram) funcctx->user_fctx;
+
+ Assert(call_cntr < histogram->nbuckets);
+
+ bucket = histogram->buckets[call_cntr];
+
+ stakeys = find_mv_attnums(mvoid, &relid);
+
+ /*
+ * The scalar values will be formatted directly, using snprintf.
+ *
+ * The 'array' values will be formatted through StringInfo.
+ */
+ values = (char **) palloc0(9 * sizeof(char *));
+ bufs = (StringInfo) palloc0(9 * sizeof(StringInfoData));
+
+ values[0] = (char *) palloc(64 * sizeof(char));
+
+ initStringInfo(&bufs[1]); /* lower boundaries */
+ initStringInfo(&bufs[2]); /* upper boundaries */
+ initStringInfo(&bufs[3]); /* nulls-only */
+ initStringInfo(&bufs[4]); /* lower inclusive */
+ initStringInfo(&bufs[5]); /* upper inclusive */
+
+ values[6] = (char *) palloc(64 * sizeof(char));
+ values[7] = (char *) palloc(64 * sizeof(char));
+ values[8] = (char *) palloc(64 * sizeof(char));
+
+ /* we need to do this only when printing the actual values */
+ outfuncs = (Oid *) palloc0(sizeof(Oid) * histogram->ndimensions);
+ fmgrinfo = (FmgrInfo *) palloc0(sizeof(FmgrInfo) * histogram->ndimensions);
+
+ /*
+ * lookup output functions for all histogram dimensions
+ *
+ * XXX This might be one in the first call and stored in user_fctx.
+ */
+ for (i = 0; i < histogram->ndimensions; i++)
+ {
+ bool isvarlena;
+
+ getTypeOutputInfo(get_atttype(relid, stakeys->values[i]),
+ &outfuncs[i], &isvarlena);
+
+ fmgr_info(outfuncs[i], &fmgrinfo[i]);
+ }
+
+ snprintf(values[0], 64, "%d", call_cntr); /* bucket ID */
+
+ /*
+ * for the arrays of lower/upper boundaries, formated according to
+ * otype
+ */
+ for (i = 0; i < histogram->ndimensions; i++)
+ {
+ Datum *vals = histogram->values[i];
+
+ uint16 minidx = bucket->min[i];
+ uint16 maxidx = bucket->max[i];
+
+ /*
+ * compute bucket volume, using distinct values as a measure
+ *
+ * XXX Not really sure what to do for NULL dimensions here, so
+ * let's simply count them as '1'.
+ */
+ bucket_volume
+ *= (double) (maxidx - minidx + 1) / (histogram->nvalues[i] - 1);
+
+ if (i == 0)
+ format = "{%s"; /* fist dimension */
+ else if (i < (histogram->ndimensions - 1))
+ format = ", %s"; /* medium dimensions */
+ else
+ format = ", %s}"; /* last dimension */
+
+ appendStringInfo(&bufs[3], format, bucket->nullsonly[i] ? "t" : "f");
+ appendStringInfo(&bufs[4], format, bucket->min_inclusive[i] ? "t" : "f");
+ appendStringInfo(&bufs[5], format, bucket->max_inclusive[i] ? "t" : "f");
+
+ /*
+ * for NULL-only dimension, simply put there the NULL and
+ * continue
+ */
+ if (bucket->nullsonly[i])
+ {
+ if (i == 0)
+ format = "{%s";
+ else if (i < (histogram->ndimensions - 1))
+ format = ", %s";
+ else
+ format = ", %s}";
+
+ appendStringInfo(&bufs[1], format, "NULL");
+ appendStringInfo(&bufs[2], format, "NULL");
+
+ continue;
+ }
+
+ /* otherwise we really need to format the value */
+ switch (otype)
+ {
+ case OUTPUT_FORMAT_RAW: /* actual boundary values */
+
+ if (i == 0)
+ format = "{%s";
+ else if (i < (histogram->ndimensions - 1))
+ format = ", %s";
+ else
+ format = ", %s}";
+
+ appendStringInfo(&bufs[1], format,
+ FunctionCall1(&fmgrinfo[i], vals[minidx]));
+
+ appendStringInfo(&bufs[2], format,
+ FunctionCall1(&fmgrinfo[i], vals[maxidx]));
+
+ break;
+
+ case OUTPUT_FORMAT_INDEXES: /* indexes into deduplicated
+ * arrays */
+
+ if (i == 0)
+ format = "{%d";
+ else if (i < (histogram->ndimensions - 1))
+ format = ", %d";
+ else
+ format = ", %d}";
+
+ appendStringInfo(&bufs[1], format, minidx);
+
+ appendStringInfo(&bufs[2], format, maxidx);
+
+ break;
+
+ case OUTPUT_FORMAT_DISTINCT: /* distinct arrays as measure */
+
+ if (i == 0)
+ format = "{%f";
+ else if (i < (histogram->ndimensions - 1))
+ format = ", %f";
+ else
+ format = ", %f}";
+
+ appendStringInfo(&bufs[1], format,
+ (minidx * 1.0 / (histogram->nvalues[i] - 1)));
+
+ appendStringInfo(&bufs[2], format,
+ (maxidx * 1.0 / (histogram->nvalues[i] - 1)));
+
+ break;
+
+ default:
+ elog(ERROR, "unknown output type: %d", otype);
+ }
+ }
+
+ values[1] = bufs[1].data;
+ values[2] = bufs[2].data;
+ values[3] = bufs[3].data;
+ values[4] = bufs[4].data;
+ values[5] = bufs[5].data;
+
+ snprintf(values[6], 64, "%f", bucket->ntuples); /* frequency */
+ snprintf(values[7], 64, "%f", bucket->ntuples / bucket_volume); /* density */
+ snprintf(values[8], 64, "%f", bucket_volume); /* volume (as a
+ * fraction) */
+
+ /* build a tuple */
+ tuple = BuildTupleFromCStrings(attinmeta, values);
+
+ /* make the tuple into a datum */
+ result = HeapTupleGetDatum(tuple);
+
+ /* clean up (this is not really necessary) */
+ pfree(values[0]);
+ pfree(values[6]);
+ pfree(values[7]);
+ pfree(values[8]);
+
+ resetStringInfo(&bufs[1]);
+ resetStringInfo(&bufs[2]);
+ resetStringInfo(&bufs[3]);
+ resetStringInfo(&bufs[4]);
+ resetStringInfo(&bufs[5]);
+
+ pfree(bufs);
+ pfree(values);
+
+ SRF_RETURN_NEXT(funcctx, result);
+ }
+ else /* do when there is no more left */
+ {
+ SRF_RETURN_DONE(funcctx);
+ }
+}
+
+/*
+ * pg_histogram_in - input routine for type pg_histogram.
+ *
+ * pg_histogram is real enough to be a table column, but it has no operations
+ * of its own, and disallows input too
+ *
+ * XXX This is inspired by what pg_node_tree does.
+ */
+Datum
+pg_histogram_in(PG_FUNCTION_ARGS)
+{
+ /*
+ * pg_node_list stores the data in binary form and parsing text input is
+ * not needed, so disallow this.
+ */
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_histogram")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_histogram - output routine for type PG_HISTOGRAM.
+ *
+ * histograms are serialized into a bytea value, so we simply call byteaout()
+ * to serialize the value into text. But it'd be nice to serialize that into
+ * a meaningful representation (e.g. for inspection by people).
+ *
+ * FIXME not implemented yet, returning dummy value
+ */
+Datum
+pg_histogram_out(PG_FUNCTION_ARGS)
+{
+ return byteaout(fcinfo);
+}
+
+/*
+ * pg_histogram_recv - binary input routine for type pg_histogram.
+ */
+Datum
+pg_histogram_recv(PG_FUNCTION_ARGS)
+{
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_histogram")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_histogram_send - binary output routine for type pg_histogram.
+ *
+ * XXX Histograms are serialized into a bytea value, so let's just send that.
+ */
+Datum
+pg_histogram_send(PG_FUNCTION_ARGS)
+{
+ return byteasend(fcinfo);
+}
+
+#ifdef DEBUG_MVHIST
+/*
+ * prints debugging info about matched histogram buckets (full/partial)
+ *
+ * XXX Currently works only for INT data type.
+ */
+void
+debug_histogram_matches(MVSerializedHistogram mvhist, char *matches)
+{
+ int i,
+ j;
+
+ float ffull = 0,
+ fpartial = 0;
+ int nfull = 0,
+ npartial = 0;
+
+ StringInfoData buf;
+
+ initStringInfo(&buf);
+
+ for (i = 0; i < mvhist->nbuckets; i++)
+ {
+ MVSerializedBucket bucket = mvhist->buckets[i];
+
+ if (!matches[i])
+ continue;
+
+ /* increment the counters */
+ nfull += (matches[i] == MVSTATS_MATCH_FULL) ? 1 : 0;
+ npartial += (matches[i] == MVSTATS_MATCH_PARTIAL) ? 1 : 0;
+
+ /* and also update the frequencies */
+ ffull += (matches[i] == MVSTATS_MATCH_FULL) ? bucket->ntuples : 0;
+ fpartial += (matches[i] == MVSTATS_MATCH_PARTIAL) ? bucket->ntuples : 0;
+
+ resetStringInfo(&buf);
+
+ /* build ranges for all the dimentions */
+ for (j = 0; j < mvhist->ndimensions; j++)
+ {
+ appendStringInfo(&buf, '[%d %d]',
+ DatumGetInt32(mvhist->values[j][bucket->min[j]]),
+ DatumGetInt32(mvhist->values[j][bucket->max[j]]));
+ }
+
+ elog(WARNING, "bucket %d %s => %d [%f]", i, buf.data, matches[i], bucket->ntuples);
+ }
+
+ elog(WARNING, "full=%f partial=%f (%f)", ffull, fpartial, (ffull + 0.5 * fpartial));
+}
+#endif
diff --git a/src/bin/psql/describe.c b/src/bin/psql/describe.c
index e2220a9..4246f2a 100644
--- a/src/bin/psql/describe.c
+++ b/src/bin/psql/describe.c
@@ -2298,8 +2298,8 @@ describeOneTableDetails(const char *schemaname,
{
printfPQExpBuffer(&buf,
"SELECT oid, stanamespace::regnamespace AS nsp, staname, stakeys,\n"
- " ndist_enabled, deps_enabled, mcv_enabled,\n"
- " ndist_built, deps_built, mcv_built,\n"
+ " ndist_enabled, deps_enabled, mcv_enabled, hist_enabled,\n"
+ " ndist_built, deps_built, mcv_built, hist_built,\n"
" (SELECT string_agg(attname::text,', ')\n"
" FROM ((SELECT unnest(stakeys) AS attnum) s\n"
" JOIN pg_attribute a ON (starelid = a.attrelid and a.attnum = s.attnum))) AS attnums\n"
@@ -2342,8 +2342,17 @@ describeOneTableDetails(const char *schemaname,
first = false;
}
+ if (!strcmp(PQgetvalue(result, i, 6), "t"))
+ {
+ if (!first)
+ appendPQExpBuffer(&buf, ", histogram");
+ else
+ appendPQExpBuffer(&buf, "(histogram");
+ first = false;
+ }
+
appendPQExpBuffer(&buf, ") ON (%s)",
- PQgetvalue(result, i, 9));
+ PQgetvalue(result, i, 12));
printTableAddFooter(&cont, buf.data);
}
diff --git a/src/include/catalog/pg_cast.h b/src/include/catalog/pg_cast.h
index 80aa9f5..a6bcbb2 100644
--- a/src/include/catalog/pg_cast.h
+++ b/src/include/catalog/pg_cast.h
@@ -266,6 +266,9 @@ DATA(insert ( 3353 25 0 i i ));
DATA(insert ( 441 17 0 i b ));
DATA(insert ( 441 25 0 i i ));
+/* pg_histogram can be coerced to, but not from, bytea */
+DATA(insert ( 774 17 0 i b ));
+
/*
* Datetime category
diff --git a/src/include/catalog/pg_mv_statistic.h b/src/include/catalog/pg_mv_statistic.h
index 34049d6..d30d3cd9 100644
--- a/src/include/catalog/pg_mv_statistic.h
+++ b/src/include/catalog/pg_mv_statistic.h
@@ -40,11 +40,13 @@ CATALOG(pg_mv_statistic,3381)
bool ndist_enabled; /* build ndist coefficient? */
bool deps_enabled; /* analyze dependencies? */
bool mcv_enabled; /* build MCV list? */
+ bool hist_enabled; /* build histogram? */
/* statistics that are available (if requested) */
bool ndist_built; /* ndistinct coeff built */
bool deps_built; /* dependencies were built */
bool mcv_built; /* MCV list was built */
+ bool hist_built; /* histogram was built */
/*
* variable-length fields start here, but we allow direct access to
@@ -56,6 +58,7 @@ CATALOG(pg_mv_statistic,3381)
pg_ndistinct standist; /* ndistinct coeff (serialized) */
pg_dependencies stadeps; /* dependencies (serialized) */
pg_mcv_list stamcv; /* MCV list (serialized) */
+ pg_histogram stahist; /* MV histogram (serialized) */
#endif
} FormData_pg_mv_statistic;
@@ -71,7 +74,7 @@ typedef FormData_pg_mv_statistic *Form_pg_mv_statistic;
* compiler constants for pg_mv_statistic
* ----------------
*/
-#define Natts_pg_mv_statistic 14
+#define Natts_pg_mv_statistic 17
#define Anum_pg_mv_statistic_starelid 1
#define Anum_pg_mv_statistic_staname 2
#define Anum_pg_mv_statistic_stanamespace 3
@@ -79,12 +82,15 @@ typedef FormData_pg_mv_statistic *Form_pg_mv_statistic;
#define Anum_pg_mv_statistic_ndist_enabled 5
#define Anum_pg_mv_statistic_deps_enabled 6
#define Anum_pg_mv_statistic_mcv_enabled 7
-#define Anum_pg_mv_statistic_ndist_built 8
-#define Anum_pg_mv_statistic_deps_built 9
-#define Anum_pg_mv_statistic_mcv_built 10
-#define Anum_pg_mv_statistic_stakeys 11
-#define Anum_pg_mv_statistic_standist 12
-#define Anum_pg_mv_statistic_stadeps 13
-#define Anum_pg_mv_statistic_stamcv 14
+#define Anum_pg_mv_statistic_hist_enabled 8
+#define Anum_pg_mv_statistic_ndist_built 9
+#define Anum_pg_mv_statistic_deps_built 10
+#define Anum_pg_mv_statistic_mcv_built 11
+#define Anum_pg_mv_statistic_hist_built 12
+#define Anum_pg_mv_statistic_stakeys 13
+#define Anum_pg_mv_statistic_standist 14
+#define Anum_pg_mv_statistic_stadeps 15
+#define Anum_pg_mv_statistic_stamcv 16
+#define Anum_pg_mv_statistic_stahist 17
#endif /* PG_MV_STATISTIC_H */
diff --git a/src/include/catalog/pg_proc.h b/src/include/catalog/pg_proc.h
index 6be685c..6c9ad2a 100644
--- a/src/include/catalog/pg_proc.h
+++ b/src/include/catalog/pg_proc.h
@@ -2726,6 +2726,10 @@ DATA(insert OID = 3376 ( pg_mv_stats_mcvlist_info PGNSP PGUID 12 1 0 0 0 f f f
DESCR("multi-variate statistics: MCV list info");
DATA(insert OID = 3373 ( pg_mv_mcv_items PGNSP PGUID 12 1 1000 0 0 f f f f t t i s 1 0 2249 "26" "{26,23,1009,1000,701}" "{i,o,o,o,o}" "{oid,index,values,nulls,frequency}" _null_ _null_ pg_mv_mcv_items _null_ _null_ _null_ ));
DESCR("details about MCV list items");
+DATA(insert OID = 3375 ( pg_mv_stats_histogram_info PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 25 "774" _null_ _null_ _null_ _null_ _null_ pg_mv_stats_histogram_info _null_ _null_ _null_ ));
+DESCR("multi-variate statistics: histogram info");
+DATA(insert OID = 3374 ( pg_mv_histogram_buckets PGNSP PGUID 12 1 1000 0 0 f f f f t t i s 2 0 2249 "26 23" "{26,23,23,1009,1009,1000,1000,1000,701,701,701}" "{i,i,o,o,o,o,o,o,o,o,o}" "{oid,otype,index,minvals,maxvals,nullsonly,mininclusive,maxinclusive,frequency,density,bucket_volume}" _null_ _null_ pg_mv_histogram_buckets _null_ _null_ _null_ ));
+DESCR("details about histogram buckets");
DATA(insert OID = 3344 ( pg_ndistinct_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 3343 "2275" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_in _null_ _null_ _null_ ));
DESCR("I/O");
@@ -2754,6 +2758,15 @@ DESCR("I/O");
DATA(insert OID = 445 ( pg_mcv_list_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "441" _null_ _null_ _null_ _null_ _null_ pg_mcv_list_send _null_ _null_ _null_ ));
DESCR("I/O");
+DATA(insert OID = 775 ( pg_histogram_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 774 "2275" _null_ _null_ _null_ _null_ _null_ pg_histogram_in _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 776 ( pg_histogram_out PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 2275 "774" _null_ _null_ _null_ _null_ _null_ pg_histogram_out _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 777 ( pg_histogram_recv PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 774 "2281" _null_ _null_ _null_ _null_ _null_ pg_histogram_recv _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 778 ( pg_histogram_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "774" _null_ _null_ _null_ _null_ _null_ pg_histogram_send _null_ _null_ _null_ ));
+DESCR("I/O");
+
DATA(insert OID = 1928 ( pg_stat_get_numscans PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_numscans _null_ _null_ _null_ ));
DESCR("statistics: number of scans done for table/index");
DATA(insert OID = 1929 ( pg_stat_get_tuples_returned PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_tuples_returned _null_ _null_ _null_ ));
diff --git a/src/include/catalog/pg_type.h b/src/include/catalog/pg_type.h
index 4621703..fdd7ba6 100644
--- a/src/include/catalog/pg_type.h
+++ b/src/include/catalog/pg_type.h
@@ -376,6 +376,10 @@ DATA(insert OID = 441 ( pg_mcv_list PGNSP PGUID -1 f b S f t \054 0 0 0 pg_mcv_
DESCR("multivariate MCV list");
#define PGMCVLISTOID 441
+DATA(insert OID = 774 ( pg_histogram PGNSP PGUID -1 f b S f t \054 0 0 0 pg_histogram_in pg_histogram_out pg_histogram_recv pg_histogram_send - - - i x f 0 -1 0 100 _null_ _null_ _null_ ));
+DESCR("multivariate histogram");
+#define PGHISTOGRAMOID 774
+
DATA(insert OID = 32 ( pg_ddl_command PGNSP PGUID SIZEOF_POINTER t p P f t \054 0 0 0 pg_ddl_command_in pg_ddl_command_out pg_ddl_command_recv pg_ddl_command_send - - - ALIGNOF_POINTER p f 0 -1 0 0 _null_ _null_ _null_ ));
DESCR("internal type for passing CollectedCommand");
#define PGDDLCOMMANDOID 32
diff --git a/src/include/nodes/relation.h b/src/include/nodes/relation.h
index b37835c..0e6ab3e 100644
--- a/src/include/nodes/relation.h
+++ b/src/include/nodes/relation.h
@@ -677,11 +677,13 @@ typedef struct MVStatisticInfo
bool ndist_enabled; /* ndistinct coefficient enabled */
bool deps_enabled; /* functional dependencies enabled */
bool mcv_enabled; /* MCV list enabled */
+ bool hist_enabled; /* histogram enabled */
/* built/available statistics */
bool ndist_built; /* ndistinct coefficient built */
bool deps_built; /* functional dependencies built */
bool mcv_built; /* MCV list built */
+ bool hist_built; /* histogram built */
/* columns in the statistics (attnums) */
int2vector *stakeys; /* attnums of the columns covered */
diff --git a/src/include/utils/builtins.h b/src/include/utils/builtins.h
index fe95bad..3db64f0 100644
--- a/src/include/utils/builtins.h
+++ b/src/include/utils/builtins.h
@@ -625,6 +625,10 @@ extern Datum pg_mcv_list_in(PG_FUNCTION_ARGS);
extern Datum pg_mcv_list_out(PG_FUNCTION_ARGS);
extern Datum pg_mcv_list_recv(PG_FUNCTION_ARGS);
extern Datum pg_mcv_list_send(PG_FUNCTION_ARGS);
+extern Datum pg_histogram_in(PG_FUNCTION_ARGS);
+extern Datum pg_histogram_out(PG_FUNCTION_ARGS);
+extern Datum pg_histogram_recv(PG_FUNCTION_ARGS);
+extern Datum pg_histogram_send(PG_FUNCTION_ARGS);
/* regexp.c */
extern Datum nameregexeq(PG_FUNCTION_ARGS);
diff --git a/src/include/utils/mvstats.h b/src/include/utils/mvstats.h
index b17fcba..21aaef7 100644
--- a/src/include/utils/mvstats.h
+++ b/src/include/utils/mvstats.h
@@ -18,7 +18,7 @@
#include "commands/vacuum.h"
/*
- * Degree of how much MCV item matches a clause.
+ * Degree of how much MCV item / histogram bucket matches a clause.
* This is then considered when computing the selectivity.
*/
#define MVSTATS_MATCH_NONE 0 /* no match at all */
@@ -114,19 +114,133 @@ bool dependency_implies_attribute(MVDependency dependency, AttrNumber attnum,
bool dependency_is_fully_matched(MVDependency dependency, Bitmapset *attnums,
int16 *attmap);
+/* used to flag stats serialized to bytea */
+#define MVSTAT_HIST_MAGIC 0x7F8C5670 /* marks serialized bytea */
+#define MVSTAT_HIST_TYPE_BASIC 1 /* basic histogram type */
+
+/* max buckets in a histogram (mostly arbitrary number */
+#define MVSTAT_HIST_MAX_BUCKETS 16384
+
+/*
+ * Multivariate histograms
+ */
+typedef struct MVBucketData
+{
+
+ /* Frequencies of this bucket. */
+ float ntuples; /* frequency of tuples tuples */
+
+ /*
+ * Information about dimensions being NULL-only. Not yet used.
+ */
+ bool *nullsonly;
+
+ /* lower boundaries - values and information about the inequalities */
+ Datum *min;
+ bool *min_inclusive;
+
+ /* upper boundaries - values and information about the inequalities */
+ Datum *max;
+ bool *max_inclusive;
+
+ /* used when building the histogram (not serialized/deserialized) */
+ void *build_data;
+
+} MVBucketData;
+
+typedef MVBucketData *MVBucket;
+
+
+typedef struct MVHistogramData
+{
+
+ uint32 magic; /* magic constant marker */
+ uint32 type; /* type of histogram (BASIC) */
+ uint32 nbuckets; /* number of buckets (buckets array) */
+ uint32 ndimensions; /* number of dimensions */
+
+ MVBucket *buckets; /* array of buckets */
+
+} MVHistogramData;
+
+typedef MVHistogramData *MVHistogram;
+
+/*
+ * Histogram in a partially serialized form, with deduplicated boundary
+ * values etc.
+ *
+ * TODO add more detailed description here
+ */
+
+typedef struct MVSerializedBucketData
+{
+
+ /* Frequencies of this bucket. */
+ float ntuples; /* frequency of tuples tuples */
+
+ /*
+ * Information about dimensions being NULL-only. Not yet used.
+ */
+ bool *nullsonly;
+
+ /* lower boundaries - values and information about the inequalities */
+ uint16 *min;
+ bool *min_inclusive;
+
+ /*
+ * indexes of upper boundaries - values and information about the
+ * inequalities (exclusive vs. inclusive)
+ */
+ uint16 *max;
+ bool *max_inclusive;
+
+} MVSerializedBucketData;
+
+typedef MVSerializedBucketData *MVSerializedBucket;
+
+typedef struct MVSerializedHistogramData
+{
+
+ uint32 magic; /* magic constant marker */
+ uint32 type; /* type of histogram (BASIC) */
+ uint32 nbuckets; /* number of buckets (buckets array) */
+ uint32 ndimensions; /* number of dimensions */
+
+ /*
+ * keep this the same with MVHistogramData, because of deserialization
+ * (same offset)
+ */
+ MVSerializedBucket *buckets; /* array of buckets */
+
+ /*
+ * serialized boundary values, one array per dimension, deduplicated (the
+ * min/max indexes point into these arrays)
+ */
+ int *nvalues;
+ Datum **values;
+
+} MVSerializedHistogramData;
+
+typedef MVSerializedHistogramData *MVSerializedHistogram;
+
+
MVNDistinct load_mv_ndistinct(Oid mvoid);
MVDependencies load_mv_dependencies(Oid mvoid);
MCVList load_mv_mcvlist(Oid mvoid);
+MVSerializedHistogram load_mv_histogram(Oid mvoid);
bytea *serialize_mv_ndistinct(MVNDistinct ndistinct);
bytea *serialize_mv_dependencies(MVDependencies dependencies);
bytea *serialize_mv_mcvlist(MCVList mcvlist, int2vector *attrs,
VacAttrStats **stats);
+bytea *serialize_mv_histogram(MVHistogram histogram, int2vector *attrs,
+ VacAttrStats **stats);
/* deserialization of stats (serialization is private to analyze) */
MVNDistinct deserialize_mv_ndistinct(bytea *data);
MVDependencies deserialize_mv_dependencies(bytea *data);
MCVList deserialize_mv_mcvlist(bytea *data);
+MVSerializedHistogram deserialize_mv_histogram(bytea * data);
/*
* Returns index of the attribute number within the vector (i.e. a
@@ -139,6 +253,8 @@ int2vector *find_mv_attnums(Oid mvoid, Oid *relid);
/* functions for inspecting the statistics */
extern Datum pg_mv_stats_mcvlist_info(PG_FUNCTION_ARGS);
extern Datum pg_mv_mcvlist_items(PG_FUNCTION_ARGS);
+extern Datum pg_mv_stats_histogram_info(PG_FUNCTION_ARGS);
+extern Datum pg_mv_histogram_buckets(PG_FUNCTION_ARGS);
MVNDistinct build_mv_ndistinct(double totalrows, int numrows, HeapTuple *rows,
@@ -151,12 +267,20 @@ MVDependencies build_mv_dependencies(int numrows, HeapTuple *rows,
MCVList build_mv_mcvlist(int numrows, HeapTuple *rows, int2vector *attrs,
VacAttrStats **stats, int *numrows_filtered);
+MVHistogram build_mv_histogram(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int numrows_total);
+
void build_mv_stats(Relation onerel, double totalrows,
int numrows, HeapTuple *rows,
int natts, VacAttrStats **vacattrstats);
void update_mv_stats(Oid relid,
- MVNDistinct ndistinct, MVDependencies dependencies, MCVList mcvlist,
+ MVNDistinct ndistinct, MVDependencies dependencies,
+ MCVList mcvlist, MVHistogram histogram,
int2vector *attrs, VacAttrStats **stats);
+#ifdef DEBUG_MVHIST
+extern void debug_histogram_matches(MVSerializedHistogram mvhist, char *matches);
+#endif
+
#endif
diff --git a/src/test/regress/expected/mv_histogram.out b/src/test/regress/expected/mv_histogram.out
new file mode 100644
index 0000000..16410ce
--- /dev/null
+++ b/src/test/regress/expected/mv_histogram.out
@@ -0,0 +1,198 @@
+-- data type passed by value
+CREATE TABLE mv_histogram (
+ a INT,
+ b INT,
+ c INT
+);
+-- unknown column
+CREATE STATISTICS s7 WITH (histogram) ON (unknown_column) FROM mv_histogram;
+ERROR: column "unknown_column" referenced in statistics does not exist
+-- single column
+CREATE STATISTICS s7 WITH (histogram) ON (a) FROM mv_histogram;
+ERROR: statistics require at least 2 columns
+-- single column, duplicated
+CREATE STATISTICS s7 WITH (histogram) ON (a, a) FROM mv_histogram;
+ERROR: duplicate column name in statistics definition
+-- two columns, one duplicated
+CREATE STATISTICS s7 WITH (histogram) ON (a, a, b) FROM mv_histogram;
+ERROR: duplicate column name in statistics definition
+-- unknown option
+CREATE STATISTICS s7 WITH (unknown_option) ON (a, b, c) FROM mv_histogram;
+ERROR: unrecognized STATISTICS option "unknown_option"
+-- correct command
+CREATE STATISTICS s7 WITH (histogram) ON (a, b, c) FROM mv_histogram;
+-- random data (no functional dependencies)
+INSERT INTO mv_histogram
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- a => b, a => c, b => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- a => b, a => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mv_histogram
+ SELECT i/100, i/200, i/400 FROM generate_series(1,30000) s(i);
+CREATE INDEX hist_idx ON mv_histogram (a, b);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a = 10 AND b = 5;
+ QUERY PLAN
+--------------------------------------------
+ Bitmap Heap Scan on mv_histogram
+ Recheck Cond: ((a = 10) AND (b = 5))
+ -> Bitmap Index Scan on hist_idx
+ Index Cond: ((a = 10) AND (b = 5))
+(4 rows)
+
+DROP TABLE mv_histogram;
+-- varlena type (text)
+CREATE TABLE mv_histogram (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+CREATE STATISTICS s8 WITH (histogram) ON (a, b, c) FROM mv_histogram;
+-- random data (no functional dependencies)
+INSERT INTO mv_histogram
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- a => b, a => c, b => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- a => b, a => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mv_histogram
+ SELECT i/100, i/200, i/400 FROM generate_series(1,30000) s(i);
+CREATE INDEX hist_idx ON mv_histogram (a, b);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a = '10' AND b = '5';
+ QUERY PLAN
+------------------------------------------------------------
+ Bitmap Heap Scan on mv_histogram
+ Recheck Cond: ((a = '10'::text) AND (b = '5'::text))
+ -> Bitmap Index Scan on hist_idx
+ Index Cond: ((a = '10'::text) AND (b = '5'::text))
+(4 rows)
+
+TRUNCATE mv_histogram;
+-- check explain (expect bitmap index scan, not plain index scan) with NULLs
+INSERT INTO mv_histogram
+ SELECT
+ (CASE WHEN i/100 = 0 THEN NULL ELSE i/100 END),
+ (CASE WHEN i/200 = 0 THEN NULL ELSE i/200 END),
+ (CASE WHEN i/400 = 0 THEN NULL ELSE i/400 END)
+ FROM generate_series(1,30000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a IS NULL AND b IS NULL;
+ QUERY PLAN
+---------------------------------------------------
+ Bitmap Heap Scan on mv_histogram
+ Recheck Cond: ((a IS NULL) AND (b IS NULL))
+ -> Bitmap Index Scan on hist_idx
+ Index Cond: ((a IS NULL) AND (b IS NULL))
+(4 rows)
+
+DROP TABLE mv_histogram;
+-- NULL values (mix of int and text columns)
+CREATE TABLE mv_histogram (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+CREATE STATISTICS s9 WITH (histogram) ON (a, b, c, d) FROM mv_histogram;
+INSERT INTO mv_histogram
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+DROP TABLE mv_histogram;
diff --git a/src/test/regress/expected/opr_sanity.out b/src/test/regress/expected/opr_sanity.out
index 9969c10..a9d8163 100644
--- a/src/test/regress/expected/opr_sanity.out
+++ b/src/test/regress/expected/opr_sanity.out
@@ -820,11 +820,12 @@ WHERE c.castmethod = 'b' AND
pg_ndistinct | bytea | 0 | i
pg_dependencies | bytea | 0 | i
pg_mcv_list | bytea | 0 | i
+ pg_histogram | bytea | 0 | i
cidr | inet | 0 | i
xml | text | 0 | a
xml | character varying | 0 | a
xml | character | 0 | a
-(10 rows)
+(11 rows)
-- **************** pg_conversion ****************
-- Look for illegal values in pg_conversion fields.
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index d14d864..759fc76 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1383,7 +1383,9 @@ pg_mv_stats| SELECT n.nspname AS schemaname,
length((s.standist)::bytea) AS ndistbytes,
length((s.stadeps)::bytea) AS depsbytes,
length((s.stamcv)::bytea) AS mcvbytes,
- pg_mv_stats_mcvlist_info(s.stamcv) AS mcvinfo
+ pg_mv_stats_mcvlist_info(s.stamcv) AS mcvinfo,
+ length((s.stahist)::bytea) AS histbytes,
+ pg_mv_stats_histogram_info(s.stahist) AS histinfo
FROM ((pg_mv_statistic s
JOIN pg_class c ON ((c.oid = s.starelid)))
LEFT JOIN pg_namespace n ON ((n.oid = c.relnamespace)));
diff --git a/src/test/regress/expected/type_sanity.out b/src/test/regress/expected/type_sanity.out
index b810e71..2e0c3b8 100644
--- a/src/test/regress/expected/type_sanity.out
+++ b/src/test/regress/expected/type_sanity.out
@@ -73,9 +73,10 @@ WHERE p1.typtype not in ('c','d','p') AND p1.typname NOT LIKE E'\\_%'
3343 | pg_ndistinct
3353 | pg_dependencies
441 | pg_mcv_list
+ 774 | pg_histogram
210 | smgr
705 | unknown
-(6 rows)
+(7 rows)
-- Make sure typarray points to a varlena array type of our own base
SELECT p1.oid, p1.typname as basetype, p2.typname as arraytype,
diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule
index d6c3cc0..8b750fc 100644
--- a/src/test/regress/parallel_schedule
+++ b/src/test/regress/parallel_schedule
@@ -115,4 +115,4 @@ test: event_trigger
test: stats
# run tests of multivariate stats
-test: mv_ndistinct mv_dependencies mv_mcv
+test: mv_ndistinct mv_dependencies mv_mcv mv_histogram
diff --git a/src/test/regress/serial_schedule b/src/test/regress/serial_schedule
index 2394d74..ff47035 100644
--- a/src/test/regress/serial_schedule
+++ b/src/test/regress/serial_schedule
@@ -172,3 +172,4 @@ test: stats
test: mv_ndistinct
test: mv_dependencies
test: mv_mcv
+test: mv_histogram
diff --git a/src/test/regress/sql/mv_histogram.sql b/src/test/regress/sql/mv_histogram.sql
new file mode 100644
index 0000000..55197cb
--- /dev/null
+++ b/src/test/regress/sql/mv_histogram.sql
@@ -0,0 +1,167 @@
+-- data type passed by value
+CREATE TABLE mv_histogram (
+ a INT,
+ b INT,
+ c INT
+);
+
+-- unknown column
+CREATE STATISTICS s7 WITH (histogram) ON (unknown_column) FROM mv_histogram;
+
+-- single column
+CREATE STATISTICS s7 WITH (histogram) ON (a) FROM mv_histogram;
+
+-- single column, duplicated
+CREATE STATISTICS s7 WITH (histogram) ON (a, a) FROM mv_histogram;
+
+-- two columns, one duplicated
+CREATE STATISTICS s7 WITH (histogram) ON (a, a, b) FROM mv_histogram;
+
+-- unknown option
+CREATE STATISTICS s7 WITH (unknown_option) ON (a, b, c) FROM mv_histogram;
+
+-- correct command
+CREATE STATISTICS s7 WITH (histogram) ON (a, b, c) FROM mv_histogram;
+
+-- random data (no functional dependencies)
+INSERT INTO mv_histogram
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- a => b, a => c, b => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- a => b, a => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mv_histogram
+ SELECT i/100, i/200, i/400 FROM generate_series(1,30000) s(i);
+CREATE INDEX hist_idx ON mv_histogram (a, b);
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a = 10 AND b = 5;
+
+DROP TABLE mv_histogram;
+
+-- varlena type (text)
+CREATE TABLE mv_histogram (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+
+CREATE STATISTICS s8 WITH (histogram) ON (a, b, c) FROM mv_histogram;
+
+-- random data (no functional dependencies)
+INSERT INTO mv_histogram
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- a => b, a => c, b => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- a => b, a => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mv_histogram
+ SELECT i/100, i/200, i/400 FROM generate_series(1,30000) s(i);
+CREATE INDEX hist_idx ON mv_histogram (a, b);
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a = '10' AND b = '5';
+
+TRUNCATE mv_histogram;
+
+-- check explain (expect bitmap index scan, not plain index scan) with NULLs
+INSERT INTO mv_histogram
+ SELECT
+ (CASE WHEN i/100 = 0 THEN NULL ELSE i/100 END),
+ (CASE WHEN i/200 = 0 THEN NULL ELSE i/200 END),
+ (CASE WHEN i/400 = 0 THEN NULL ELSE i/400 END)
+ FROM generate_series(1,30000) s(i);
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a IS NULL AND b IS NULL;
+
+DROP TABLE mv_histogram;
+
+-- NULL values (mix of int and text columns)
+CREATE TABLE mv_histogram (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+
+CREATE STATISTICS s9 WITH (histogram) ON (a, b, c, d) FROM mv_histogram;
+
+INSERT INTO mv_histogram
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+DROP TABLE mv_histogram;
--
2.5.5
[binary/octet-stream] 0007-WIP-use-ndistinct-for-selectivity-estimation-in--v22.patch (14.6K, ../../696aa95c-2411-9b2b-f36e-65b66bf47c88@2ndquadrant.com/8-0007-WIP-use-ndistinct-for-selectivity-estimation-in--v22.patch)
download | inline diff:
From 965892c97cb950e8b2bf56118165ae42144a34cc Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Thu, 27 Oct 2016 15:24:42 +0200
Subject: [PATCH 7/9] WIP: use ndistinct for selectivity estimation in
clausesel.c
---
src/backend/optimizer/path/clausesel.c | 382 ++++++++++++++++++++++++++-------
1 file changed, 299 insertions(+), 83 deletions(-)
diff --git a/src/backend/optimizer/path/clausesel.c b/src/backend/optimizer/path/clausesel.c
index 0243b4d..c35b914 100644
--- a/src/backend/optimizer/path/clausesel.c
+++ b/src/backend/optimizer/path/clausesel.c
@@ -47,9 +47,10 @@ typedef struct RangeQueryClause
static void addRangeClause(RangeQueryClause **rqlist, Node *clause,
bool varonleft, bool isLTsel, Selectivity s2);
-#define STATS_TYPE_FDEPS 0x01
-#define STATS_TYPE_MCV 0x02
-#define STATS_TYPE_HIST 0x04
+#define STATS_TYPE_NDIST 0x01
+#define STATS_TYPE_FDEPS 0x02
+#define STATS_TYPE_MCV 0x04
+#define STATS_TYPE_HIST 0x08
static bool clause_is_mv_compatible(Node *clause, Index relid, Bitmapset **attnums,
int type);
@@ -70,6 +71,10 @@ static List *clauselist_mv_split(PlannerInfo *root, Index relid,
static Selectivity clauselist_mv_selectivity(PlannerInfo *root,
List *clauses, MVStatisticInfo *mvstats);
+static Selectivity clauselist_mv_selectivity_ndist(PlannerInfo *root,
+ Index relid, List *clauses, MVStatisticInfo *mvstats,
+ Index varRelid, JoinType jointype, SpecialJoinInfo *sjinfo);
+
static Selectivity clauselist_mv_selectivity_deps(PlannerInfo *root,
Index relid, List *clauses, MVStatisticInfo *mvstats,
Index varRelid, JoinType jointype, SpecialJoinInfo *sjinfo);
@@ -282,6 +287,37 @@ clauselist_selectivity(PlannerInfo *root,
}
}
+ /* And finally, try to use ndistinct coefficients. */
+ if (has_stats(stats, STATS_TYPE_NDIST) &&
+ (count_mv_attnums(clauses, relid, STATS_TYPE_NDIST) >= 2))
+ {
+ MVStatisticInfo *mvstat;
+ Bitmapset *mvattnums;
+
+ /* collect attributes from the compatible conditions */
+ mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_NDIST);
+
+ /* and search for the statistic covering the most attributes */
+ mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_NDIST);
+
+ if (mvstat != NULL) /* we have a matching stats */
+ {
+ /* clauses compatible with multi-variate stats */
+ List *mvclauses = NIL;
+
+ /* split the clauselist into regular and mv-clauses */
+ clauses = clauselist_mv_split(root, relid, clauses, &mvclauses,
+ mvstat, STATS_TYPE_NDIST);
+
+ /* we've chosen the histogram to match the clauses */
+ Assert(mvclauses != NIL);
+
+ /* compute the multivariate stats (dependencies) */
+ s1 *= clauselist_mv_selectivity_ndist(root, relid, mvclauses, mvstat,
+ varRelid, jointype, sjinfo);
+ }
+ }
+
/*
* Initial scan over clauses. Anything that doesn't look like a potential
* rangequery clause gets multiplied into s1 and forgotten. Anything that
@@ -939,6 +975,261 @@ clause_selectivity(PlannerInfo *root,
return s1;
}
+
+/*
+ * estimate selectivity of clauses using multivariate statistic
+ *
+ * Perform estimation of the clauses using a MCV list.
+ *
+ * This assumes all the clauses are compatible with the selected statistics
+ * (e.g. only reference columns covered by the statistics, use supported
+ * operator, etc.).
+ *
+ * TODO: We may support some additional conditions, most importantly those
+ * matching multiple columns (e.g. "a = b" or "a < b").
+ *
+ * TODO: Clamp the selectivity by min of the per-clause selectivities (i.e. the
+ * selectivity of the most restrictive clause), because that's the maximum
+ * we can ever get from ANDed list of clauses. This may probably prevent
+ * issues with hitting too many buckets and low precision histograms.
+ *
+ * TODO: We may remember the lowest frequency in the MCV list, and then later
+ * use it as a upper boundary for the selectivity (had there been a more
+ * frequent item, it'd be in the MCV list). This might improve cases with
+ * low-detail histograms.
+ *
+ * TODO: We may also derive some additional boundaries for the selectivity from
+ * the MCV list, because
+ *
+ * (a) if we have a "full equality condition" (one equality condition on
+ * each column of the statistic) and we found a match in the MCV list,
+ * then this is the final selectivity (and pretty accurate),
+ *
+ * (b) if we have a "full equality condition" and we haven't found a match
+ * in the MCV list, then the selectivity is below the lowest frequency
+ * found in the MCV list,
+ *
+ * TODO: When applying the clauses to the histogram/MCV list, we can do that
+ * from the most selective clauses first, because that'll eliminate the
+ * buckets/items sooner (so we'll be able to skip them without inspection,
+ * which is more expensive). But this requires really knowing the per-clause
+ * selectivities in advance, and that's not what we do now.
+ */
+static Selectivity
+clauselist_mv_selectivity(PlannerInfo *root, List *clauses, MVStatisticInfo *mvstats)
+{
+ bool fullmatch = false;
+ Selectivity s1 = 0.0,
+ s2 = 0.0;
+
+ /*
+ * Lowest frequency in the MCV list (may be used as an upper bound for
+ * full equality conditions that did not match any MCV item).
+ */
+ Selectivity mcv_low = 0.0;
+
+ /*
+ * TODO: Evaluate simple 1D selectivities, use the smallest one as an
+ * upper bound, product as lower bound, and sort the clauses in ascending
+ * order by selectivity (to optimize the MCV/histogram evaluation).
+ */
+
+ /* Evaluate the MCV first. */
+ s1 = clauselist_mv_selectivity_mcvlist(root, clauses, mvstats,
+ &fullmatch, &mcv_low);
+
+ /*
+ * If we got a full equality match on the MCV list, we're done (and the
+ * estimate is pretty good).
+ */
+ if (fullmatch && (s1 > 0.0))
+ return s1;
+
+ /*
+ * TODO if (fullmatch) without matching MCV item, use the mcv_low
+ * selectivity as upper bound
+ */
+
+ s2 = clauselist_mv_selectivity_histogram(root, clauses, mvstats);
+
+ /* TODO clamp to <= 1.0 (or more strictly, when possible) */
+ return s1 + s2;
+}
+
+static MVNDistinctItem *
+find_widest_ndistinct_item(MVNDistinct ndistinct, Bitmapset *attnums,
+ int16 *attmap)
+{
+ int i;
+ MVNDistinctItem *widest = NULL;
+
+ /* number of attnums in clauses */
+ int nattnums = bms_num_members(attnums);
+
+ /* with less than two attributes, we can bail out right away */
+ if (nattnums < 2)
+ return NULL;
+
+ /*
+ * Iterate over the MVNDistinctItem items and find the widest one from
+ * those fully-matched by clasuse.
+ */
+ for (i = 0; i < ndistinct->nitems; i++)
+ {
+ int j;
+ bool full_match = true;
+ MVNDistinctItem *item = &ndistinct->items[i];
+
+ /*
+ * Skip items referencing more attributes than available clauses,
+ * as those can't be fully matched.
+ */
+ if (item->nattrs > nattnums)
+ continue;
+
+ /* We can skip items with fewer attributes than the best one. */
+ if (widest && (widest->nattrs >= item->nattrs))
+ continue;
+
+ /*
+ * Check that the item actually is fully covered by clauses. We
+ * have to translate all attribute numbers.
+ */
+ for (j = 0; j < item->nattrs; j++)
+ {
+ int attnum = attmap[item->attrs[j]];
+
+ if (! bms_is_member(attnum, attnums))
+ {
+ full_match = false;
+ break;
+ }
+ }
+
+ /*
+ * If the item is not fully matched by clauses, we can't use
+ * it for the estimation.
+ */
+ if (! full_match)
+ continue;
+
+ /*
+ * We have a fully-matched item, and we already know it has to
+ * be wider than the current one (otherwise we'd skip it before
+ * inspecting it at the very beginning).
+ */
+ widest = item;
+ }
+
+ return widest;
+}
+
+static bool
+attnum_in_ndistinct_item(MVNDistinctItem *item, int attnum, int16 *attmap)
+{
+ int j;
+
+ for (j = 0; j < item->nattrs; j++)
+ {
+ if (attnum == attmap[item->attrs[j]])
+ return true;
+ }
+
+ return false;
+}
+
+static Selectivity
+clauselist_mv_selectivity_ndist(PlannerInfo *root, Index relid,
+ List *clauses, MVStatisticInfo *mvstats,
+ Index varRelid, JoinType jointype,
+ SpecialJoinInfo *sjinfo)
+{
+ ListCell *lc;
+ Selectivity s1 = 1.0;
+ MVNDistinct ndistinct;
+ MVNDistinctItem *item;
+ Bitmapset *attnums;
+ List *clauses_filtered = NIL;
+
+ /* we should only get here if the statistics includes ndistinct */
+ Assert(mvstats->ndist_enabled && mvstats->ndist_built);
+
+ /* load the ndistinct items stored in the statistics */
+ ndistinct = load_mv_ndistinct(mvstats->mvoid);
+
+ /* collect attnums in the clauses */
+ attnums = collect_mv_attnums(clauses, relid, STATS_TYPE_NDIST);
+
+ Assert(bms_num_members(attnums) >= 2);
+
+ /*
+ * Search for the widest ndistinct item (covering the most clauses), and
+ * then use it to estimate the number of entries.
+ */
+ item = find_widest_ndistinct_item(ndistinct, attnums,
+ mvstats->stakeys->values);
+
+ if (item)
+ {
+ /*
+ * We have an applicable item, so identify all covered clauses, and
+ * remove them from the list of clauses.
+ */
+ foreach(lc, clauses)
+ {
+ Bitmapset *attnums_clause = NULL;
+ Node *clause = (Node *) lfirst(lc);
+
+ /*
+ * XXX We need the attnum referenced by the clause, and this is the
+ * easiest way to get it (but maybe not the best one). At this point
+ * we should only see equality clauses, so just error out if we
+ * stumble upon something else.
+ */
+ if (! clause_is_mv_compatible(clause, relid, &attnums_clause,
+ STATS_TYPE_NDIST))
+ elog(ERROR, "clause not compatible with ndistinct stats");
+
+ /*
+ * We also expect only simple equality clauses, with a single Var.
+ *
+ * XXX This checks the number of attnums, not the number of Vars,
+ * but clause_is_mv_compatible only accepts (Var=Const) clauses.
+ */
+ Assert(bms_num_members(attnums_clause) == 1);
+
+ /*
+ * If the clause matches the selected ndistinct item, add it to
+ * the list of ndistinct clauses.
+ */
+ if (!attnum_in_ndistinct_item(item,
+ bms_singleton_member(attnums_clause),
+ mvstats->stakeys->values))
+ clauses_filtered = lappend(clauses_filtered, clause);
+ }
+
+ /* Compute selectivity using the ndistinct item. */
+ s1 *= (1.0 / item->ndistinct);
+
+ /*
+ * Throw away the clauses matched by the ndistinct, so that we don't
+ * estimate them twice.
+ */
+ clauses = clauses_filtered;
+ }
+
+ /* And now simply multiply with selectivities of the remaining clauses. */
+ foreach (lc, clauses)
+ {
+ Node *clause = (Node *) lfirst(lc);
+
+ s1 *= clause_selectivity(root, clause, varRelid, jointype, sjinfo);
+ }
+
+ return s1;
+}
+
+
/*
* When applying functional dependencies, we start with the strongest ones
* strongest dependencies. That is, we select the dependency that:
@@ -1147,85 +1438,6 @@ clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
return s1;
}
-/*
- * estimate selectivity of clauses using multivariate statistic
- *
- * Perform estimation of the clauses using a MCV list.
- *
- * This assumes all the clauses are compatible with the selected statistics
- * (e.g. only reference columns covered by the statistics, use supported
- * operator, etc.).
- *
- * TODO: We may support some additional conditions, most importantly those
- * matching multiple columns (e.g. "a = b" or "a < b").
- *
- * TODO: Clamp the selectivity by min of the per-clause selectivities (i.e. the
- * selectivity of the most restrictive clause), because that's the maximum
- * we can ever get from ANDed list of clauses. This may probably prevent
- * issues with hitting too many buckets and low precision histograms.
- *
- * TODO: We may remember the lowest frequency in the MCV list, and then later
- * use it as a upper boundary for the selectivity (had there been a more
- * frequent item, it'd be in the MCV list). This might improve cases with
- * low-detail histograms.
- *
- * TODO: We may also derive some additional boundaries for the selectivity from
- * the MCV list, because
- *
- * (a) if we have a "full equality condition" (one equality condition on
- * each column of the statistic) and we found a match in the MCV list,
- * then this is the final selectivity (and pretty accurate),
- *
- * (b) if we have a "full equality condition" and we haven't found a match
- * in the MCV list, then the selectivity is below the lowest frequency
- * found in the MCV list,
- *
- * TODO: When applying the clauses to the histogram/MCV list, we can do that
- * from the most selective clauses first, because that'll eliminate the
- * buckets/items sooner (so we'll be able to skip them without inspection,
- * which is more expensive). But this requires really knowing the per-clause
- * selectivities in advance, and that's not what we do now.
- */
-static Selectivity
-clauselist_mv_selectivity(PlannerInfo *root, List *clauses, MVStatisticInfo *mvstats)
-{
- bool fullmatch = false;
- Selectivity s1 = 0.0,
- s2 = 0.0;
-
- /*
- * Lowest frequency in the MCV list (may be used as an upper bound for
- * full equality conditions that did not match any MCV item).
- */
- Selectivity mcv_low = 0.0;
-
- /*
- * TODO: Evaluate simple 1D selectivities, use the smallest one as an
- * upper bound, product as lower bound, and sort the clauses in ascending
- * order by selectivity (to optimize the MCV/histogram evaluation).
- */
-
- /* Evaluate the MCV first. */
- s1 = clauselist_mv_selectivity_mcvlist(root, clauses, mvstats,
- &fullmatch, &mcv_low);
-
- /*
- * If we got a full equality match on the MCV list, we're done (and the
- * estimate is pretty good).
- */
- if (fullmatch && (s1 > 0.0))
- return s1;
-
- /*
- * TODO if (fullmatch) without matching MCV item, use the mcv_low
- * selectivity as upper bound
- */
-
- s2 = clauselist_mv_selectivity_histogram(root, clauses, mvstats);
-
- /* TODO clamp to <= 1.0 (or more strictly, when possible) */
- return s1 + s2;
-}
/*
* Collect attributes from mv-compatible clauses.
@@ -1409,7 +1621,8 @@ choose_mv_statistics(List *stats, Bitmapset *attnums, int types)
int numattrs = info->stakeys->dim1;
/* skip statistics not matching any of the requested types */
- if (! ((info->deps_built && (STATS_TYPE_FDEPS & types)) ||
+ if (! ((info->ndist_built && (STATS_TYPE_NDIST & types)) ||
+ (info->deps_built && (STATS_TYPE_FDEPS & types)) ||
(info->mcv_built && (STATS_TYPE_MCV & types)) ||
(info->hist_built && (STATS_TYPE_HIST & types))))
continue;
@@ -1703,6 +1916,9 @@ clause_is_mv_compatible(Node *clause, Index relid, Bitmapset **attnums, int type
static bool
stats_type_matches(MVStatisticInfo *stat, int type)
{
+ if ((type & STATS_TYPE_NDIST) && stat->ndist_built)
+ return true;
+
if ((type & STATS_TYPE_FDEPS) && stat->deps_built)
return true;
--
2.5.5
[binary/octet-stream] 0008-WIP-allow-using-multiple-statistics-in-clauselis-v22.patch (9.3K, ../../696aa95c-2411-9b2b-f36e-65b66bf47c88@2ndquadrant.com/9-0008-WIP-allow-using-multiple-statistics-in-clauselis-v22.patch)
download | inline diff:
From 6a4c76169f37c0e24efd3d3786cec1132258fa2f Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Fri, 28 Oct 2016 17:03:09 +0200
Subject: [PATCH 8/9] WIP: allow using multiple statistics in
clauselist_selectivity
---
src/backend/optimizer/path/clausesel.c | 31 +++++++-----
src/test/regress/expected/mv_statistics.out | 78 +++++++++++++++++++++++++++++
src/test/regress/parallel_schedule | 2 +-
src/test/regress/serial_schedule | 1 +
src/test/regress/sql/mv_statistics.sql | 60 ++++++++++++++++++++++
5 files changed, 159 insertions(+), 13 deletions(-)
create mode 100644 src/test/regress/expected/mv_statistics.out
create mode 100644 src/test/regress/sql/mv_statistics.sql
diff --git a/src/backend/optimizer/path/clausesel.c b/src/backend/optimizer/path/clausesel.c
index c35b914..e8b214f 100644
--- a/src/backend/optimizer/path/clausesel.c
+++ b/src/backend/optimizer/path/clausesel.c
@@ -228,15 +228,16 @@ clauselist_selectivity(PlannerInfo *root,
(count_mv_attnums(clauses, relid,
STATS_TYPE_MCV | STATS_TYPE_HIST) >= 2))
{
+ Bitmapset *mvattnums;
+ MVStatisticInfo *mvstat;
+
/* collect attributes from the compatible conditions */
- Bitmapset *mvattnums = collect_mv_attnums(clauses, relid,
- STATS_TYPE_MCV | STATS_TYPE_HIST);
+ mvattnums = collect_mv_attnums(clauses, relid,
+ STATS_TYPE_MCV | STATS_TYPE_HIST);
/* and search for the statistic covering the most attributes */
- MVStatisticInfo *mvstat = choose_mv_statistics(stats, mvattnums,
- STATS_TYPE_MCV | STATS_TYPE_HIST);
-
- if (mvstat != NULL) /* we have a matching stats */
+ while ((mvstat = choose_mv_statistics(stats, mvattnums,
+ STATS_TYPE_MCV | STATS_TYPE_HIST)))
{
/* clauses compatible with multi-variate stats */
List *mvclauses = NIL;
@@ -250,6 +251,10 @@ clauselist_selectivity(PlannerInfo *root,
/* compute the multivariate stats */
s1 *= clauselist_mv_selectivity(root, mvclauses, mvstat);
+
+ /* update the bitmap if attnums using the remaining clauses) */
+ mvattnums = collect_mv_attnums(clauses, relid,
+ STATS_TYPE_MCV | STATS_TYPE_HIST);
}
}
@@ -264,9 +269,7 @@ clauselist_selectivity(PlannerInfo *root,
mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_FDEPS);
/* and search for the statistic covering the most attributes */
- mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_FDEPS);
-
- if (mvstat != NULL) /* we have a matching stats */
+ while ((mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_FDEPS)))
{
/* clauses compatible with multi-variate stats */
List *mvclauses = NIL;
@@ -284,6 +287,9 @@ clauselist_selectivity(PlannerInfo *root,
/* compute the multivariate stats (dependencies) */
s1 *= clauselist_mv_selectivity_deps(root, relid, mvclauses, mvstat,
varRelid, jointype, sjinfo);
+
+ /* update the bitmap if attnums using the remaining clauses) */
+ mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_FDEPS);
}
}
@@ -298,9 +304,7 @@ clauselist_selectivity(PlannerInfo *root,
mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_NDIST);
/* and search for the statistic covering the most attributes */
- mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_NDIST);
-
- if (mvstat != NULL) /* we have a matching stats */
+ while ((mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_NDIST)))
{
/* clauses compatible with multi-variate stats */
List *mvclauses = NIL;
@@ -315,6 +319,9 @@ clauselist_selectivity(PlannerInfo *root,
/* compute the multivariate stats (dependencies) */
s1 *= clauselist_mv_selectivity_ndist(root, relid, mvclauses, mvstat,
varRelid, jointype, sjinfo);
+
+ /* collect attributes from the compatible conditions */
+ mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_NDIST);
}
}
diff --git a/src/test/regress/expected/mv_statistics.out b/src/test/regress/expected/mv_statistics.out
new file mode 100644
index 0000000..7eb6f2e
--- /dev/null
+++ b/src/test/regress/expected/mv_statistics.out
@@ -0,0 +1,78 @@
+-- data type passed by value
+CREATE TABLE multi_stats (
+ a INT,
+ b INT,
+ c INT,
+ d INT,
+ e INT,
+ f INT,
+ g INT,
+ h INT
+);
+-- MCV list on (a,b)
+CREATE STATISTICS m1 WITH (mcv) ON (a, b) FROM multi_stats;
+-- histogram on (c,d)
+CREATE STATISTICS m2 WITH (histogram) ON (c, d) FROM multi_stats;
+-- functional dependencies on (e,f)
+CREATE STATISTICS m3 WITH (dependencies) ON (e, f) FROM multi_stats;
+-- ndistinct coefficients on (g,h)
+CREATE STATISTICS m4 WITH (ndistinct) ON (g, h) FROM multi_stats;
+-- perfectly correlated groups
+INSERT INTO multi_stats
+SELECT
+ i, i/2, -- MCV
+ i, i + j, -- histogram
+ k, k/2, -- dependencies
+ l/5, l/10 -- ndistinct
+FROM (
+ SELECT
+ mod(x, 13) AS i,
+ mod(x, 17) AS j,
+ mod(x, 11) AS k,
+ mod(x, 51) AS l
+ FROM generate_series(1,30000) AS s(x)
+) foo;
+ANALYZE multi_stats;
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (c >= 3) AND (d <= 10);
+ QUERY PLAN
+----------------------------------------------------------------
+ Seq Scan on multi_stats (cost=0.00..821.00 rows=413 width=32)
+ Filter: ((c >= 3) AND (d <= 10) AND (a = 8) AND (b = 4))
+(2 rows)
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (e = 10) AND (f = 5);
+ QUERY PLAN
+----------------------------------------------------------------
+ Seq Scan on multi_stats (cost=0.00..821.00 rows=210 width=32)
+ Filter: ((a = 8) AND (b = 4) AND (e = 10) AND (f = 5))
+(2 rows)
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (e = 10) AND (f = 5);
+ QUERY PLAN
+----------------------------------------------------------------
+ Seq Scan on multi_stats (cost=0.00..821.00 rows=210 width=32)
+ Filter: ((a = 8) AND (b = 4) AND (e = 10) AND (f = 5))
+(2 rows)
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (g = 2) AND (h = 1);
+ QUERY PLAN
+----------------------------------------------------------------
+ Seq Scan on multi_stats (cost=0.00..821.00 rows=210 width=32)
+ Filter: ((a = 8) AND (b = 4) AND (g = 2) AND (h = 1))
+(2 rows)
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND
+ (c >= 3) AND (d <= 10) AND
+ (e = 10) AND (f = 5);
+ QUERY PLAN
+-------------------------------------------------------------------------------------
+ Seq Scan on multi_stats (cost=0.00..971.00 rows=37 width=32)
+ Filter: ((c >= 3) AND (d <= 10) AND (a = 8) AND (b = 4) AND (e = 10) AND (f = 5))
+(2 rows)
+
+DROP TABLE multi_stats;
diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule
index 8b750fc..b220707 100644
--- a/src/test/regress/parallel_schedule
+++ b/src/test/regress/parallel_schedule
@@ -115,4 +115,4 @@ test: event_trigger
test: stats
# run tests of multivariate stats
-test: mv_ndistinct mv_dependencies mv_mcv mv_histogram
+test: mv_ndistinct mv_dependencies mv_mcv mv_histogram mv_statistics
diff --git a/src/test/regress/serial_schedule b/src/test/regress/serial_schedule
index ff47035..9a8de27 100644
--- a/src/test/regress/serial_schedule
+++ b/src/test/regress/serial_schedule
@@ -173,3 +173,4 @@ test: mv_ndistinct
test: mv_dependencies
test: mv_mcv
test: mv_histogram
+test: mv_statistics
diff --git a/src/test/regress/sql/mv_statistics.sql b/src/test/regress/sql/mv_statistics.sql
new file mode 100644
index 0000000..cd12ad0
--- /dev/null
+++ b/src/test/regress/sql/mv_statistics.sql
@@ -0,0 +1,60 @@
+-- data type passed by value
+CREATE TABLE multi_stats (
+ a INT,
+ b INT,
+ c INT,
+ d INT,
+ e INT,
+ f INT,
+ g INT,
+ h INT
+);
+
+-- MCV list on (a,b)
+CREATE STATISTICS m1 WITH (mcv) ON (a, b) FROM multi_stats;
+
+-- histogram on (c,d)
+CREATE STATISTICS m2 WITH (histogram) ON (c, d) FROM multi_stats;
+
+-- functional dependencies on (e,f)
+CREATE STATISTICS m3 WITH (dependencies) ON (e, f) FROM multi_stats;
+
+-- ndistinct coefficients on (g,h)
+CREATE STATISTICS m4 WITH (ndistinct) ON (g, h) FROM multi_stats;
+
+-- perfectly correlated groups
+INSERT INTO multi_stats
+SELECT
+ i, i/2, -- MCV
+ i, i + j, -- histogram
+ k, k/2, -- dependencies
+ l/5, l/10 -- ndistinct
+FROM (
+ SELECT
+ mod(x, 13) AS i,
+ mod(x, 17) AS j,
+ mod(x, 11) AS k,
+ mod(x, 51) AS l
+ FROM generate_series(1,30000) AS s(x)
+) foo;
+
+ANALYZE multi_stats;
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (c >= 3) AND (d <= 10);
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (e = 10) AND (f = 5);
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (e = 10) AND (f = 5);
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (g = 2) AND (h = 1);
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND
+ (c >= 3) AND (d <= 10) AND
+ (e = 10) AND (f = 5);
+
+DROP TABLE multi_stats;
--
2.5.5
[binary/octet-stream] 0009-WIP-psql-tab-completion-basics-v22.patch (3.1K, ../../696aa95c-2411-9b2b-f36e-65b66bf47c88@2ndquadrant.com/10-0009-WIP-psql-tab-completion-basics-v22.patch)
download | inline diff:
From a1be139710453d0a799eb5773cbe3598ffcecfed Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Fri, 28 Oct 2016 20:46:37 +0200
Subject: [PATCH 9/9] WIP: psql tab-completion basics
---
src/bin/psql/tab-complete.c | 30 ++++++++++++++++++++++++++++--
1 file changed, 28 insertions(+), 2 deletions(-)
diff --git a/src/bin/psql/tab-complete.c b/src/bin/psql/tab-complete.c
index 02c8d60..2d11aad 100644
--- a/src/bin/psql/tab-complete.c
+++ b/src/bin/psql/tab-complete.c
@@ -448,6 +448,21 @@ static const SchemaQuery Query_for_list_of_foreign_tables = {
NULL
};
+static const SchemaQuery Query_for_list_of_statistics = {
+ /* catname */
+ "pg_catalog.pg_mv_statistic s",
+ /* selcondition */
+ NULL,
+ /* viscondition */
+ NULL,
+ /* namespace */
+ "s.stanamespace",
+ /* result */
+ "pg_catalog.quote_ident(s.staname)",
+ /* qualresult */
+ NULL
+};
+
static const SchemaQuery Query_for_list_of_tables = {
/* catname */
"pg_catalog.pg_class c",
@@ -965,6 +980,7 @@ static const pgsql_thing_t words_after_create[] = {
{"SCHEMA", Query_for_list_of_schemas},
{"SEQUENCE", NULL, &Query_for_list_of_sequences},
{"SERVER", Query_for_list_of_servers},
+ {"STATISTICS", NULL, &Query_for_list_of_statistics},
{"TABLE", NULL, &Query_for_list_of_tables},
{"TABLESPACE", Query_for_list_of_tablespaces},
{"TEMP", NULL, NULL, THING_NO_DROP}, /* for CREATE TEMP TABLE ... */
@@ -1407,8 +1423,8 @@ psql_completion(const char *text, int start, int end)
{"AGGREGATE", "COLLATION", "CONVERSION", "DATABASE", "DEFAULT PRIVILEGES", "DOMAIN",
"EVENT TRIGGER", "EXTENSION", "FOREIGN DATA WRAPPER", "FOREIGN TABLE", "FUNCTION",
"GROUP", "INDEX", "LANGUAGE", "LARGE OBJECT", "MATERIALIZED VIEW", "OPERATOR",
- "POLICY", "ROLE", "RULE", "SCHEMA", "SERVER", "SEQUENCE", "SYSTEM", "TABLE",
- "TABLESPACE", "TEXT SEARCH", "TRIGGER", "TYPE",
+ "POLICY", "ROLE", "RULE", "SCHEMA", "SERVER", "SEQUENCE", "STATISTICS", "SYSTEM",
+ "TABLE", "TABLESPACE", "TEXT SEARCH", "TRIGGER", "TYPE",
"USER", "USER MAPPING FOR", "VIEW", NULL};
COMPLETE_WITH_LIST(list_ALTER);
@@ -1692,6 +1708,10 @@ psql_completion(const char *text, int start, int end)
else if (Matches5("ALTER", "RULE", MatchAny, "ON", MatchAny))
COMPLETE_WITH_CONST("RENAME TO");
+ /* ALTER STATISTICS <name> */
+ else if (Matches3("ALTER", "STATISTICS", MatchAny))
+ COMPLETE_WITH_LIST3("OWNER TO", "RENAME TO", "SET SCHEMA");
+
/* ALTER TRIGGER <name>, add ON */
else if (Matches3("ALTER", "TRIGGER", MatchAny))
COMPLETE_WITH_CONST("ON");
@@ -2257,6 +2277,12 @@ psql_completion(const char *text, int start, int end)
else if (Matches3("CREATE", "SERVER", MatchAny))
COMPLETE_WITH_LIST3("TYPE", "VERSION", "FOREIGN DATA WRAPPER");
+/* CREATE STATISTICS <name> */
+ else if (Matches3("CREATE", "STATISTICS", MatchAny))
+ COMPLETE_WITH_LIST2("WITH", "ON");
+ else if (Matches4("CREATE", "STATISTICS", MatchAny, "ON|WITH"))
+ COMPLETE_WITH_CONST("(");
+
/* CREATE TABLE --- is allowed inside CREATE SCHEMA, so use TailMatches */
/* Complete "CREATE TEMP/TEMPORARY" with the possible temp objects */
else if (TailMatches2("CREATE", "TEMP|TEMPORARY"))
--
2.5.5
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-04 14:21 Dilip Kumar <dilipbalaut@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
3 siblings, 1 reply; 70+ messages in thread
From: Dilip Kumar @ 2017-01-04 14:21 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Wed, Jan 4, 2017 at 8:05 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> Attached is v22 of the patch series, rebased to current master and fixing
> the reported bug. I haven't made any other changes - the issues reported by
> Petr are mostly minor, so I've decided to wait a bit more for (hopefully)
> other reviews.
v22 fixes the problem, I reported. In my test, I observed that group
by estimation is much better with ndistinct stat.
Here is one example:
postgres=# explain analyze select p_brand, p_type, p_size from part
group by p_brand, p_type, p_size;
QUERY PLAN
-----------------------------------------------------------------------------------------------------------------------
HashAggregate (cost=37992.00..38992.00 rows=100000 width=36) (actual
time=953.359..1011.302 rows=186607 loops=1)
Group Key: p_brand, p_type, p_size
-> Seq Scan on part (cost=0.00..30492.00 rows=1000000 width=36)
(actual time=0.013..163.672 rows=1000000 loops=1)
Planning time: 0.194 ms
Execution time: 1020.776 ms
(5 rows)
postgres=# CREATE STATISTICS s2 WITH (ndistinct) on (p_brand, p_type,
p_size) from part;
CREATE STATISTICS
postgres=# analyze part;
ANALYZE
postgres=# explain analyze select p_brand, p_type, p_size from part
group by p_brand, p_type, p_size;
QUERY PLAN
-----------------------------------------------------------------------------------------------------------------------
HashAggregate (cost=37992.00..39622.46 rows=163046 width=36) (actual
time=935.162..992.944 rows=186607 loops=1)
Group Key: p_brand, p_type, p_size
-> Seq Scan on part (cost=0.00..30492.00 rows=1000000 width=36)
(actual time=0.013..156.746 rows=1000000 loops=1)
Planning time: 0.308 ms
Execution time: 1001.889 ms
In above example,
Without MVStat-> estimated: 100000 Actual: 186607
With MVStat-> estimated: 163046 Actual: 186607
--
Regards,
Dilip Kumar
EnterpriseDB: http://www.enterprisedb.com
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-04 21:57 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Dilip Kumar <dilipbalaut@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Tomas Vondra @ 2017-01-04 21:57 UTC (permalink / raw)
To: Dilip Kumar <dilipbalaut@gmail.com>; +Cc: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 01/04/2017 03:21 PM, Dilip Kumar wrote:
> On Wed, Jan 4, 2017 at 8:05 AM, Tomas Vondra
> <tomas.vondra@2ndquadrant.com> wrote:
>> Attached is v22 of the patch series, rebased to current master and fixing
>> the reported bug. I haven't made any other changes - the issues reported by
>> Petr are mostly minor, so I've decided to wait a bit more for (hopefully)
>> other reviews.
>
> v22 fixes the problem, I reported. In my test, I observed that group
> by estimation is much better with ndistinct stat.
>
> Here is one example:
>
> postgres=# explain analyze select p_brand, p_type, p_size from part
> group by p_brand, p_type, p_size;
> QUERY PLAN
> -----------------------------------------------------------------------------------------------------------------------
> HashAggregate (cost=37992.00..38992.00 rows=100000 width=36) (actual
> time=953.359..1011.302 rows=186607 loops=1)
> Group Key: p_brand, p_type, p_size
> -> Seq Scan on part (cost=0.00..30492.00 rows=1000000 width=36)
> (actual time=0.013..163.672 rows=1000000 loops=1)
> Planning time: 0.194 ms
> Execution time: 1020.776 ms
> (5 rows)
>
> postgres=# CREATE STATISTICS s2 WITH (ndistinct) on (p_brand, p_type,
> p_size) from part;
> CREATE STATISTICS
> postgres=# analyze part;
> ANALYZE
> postgres=# explain analyze select p_brand, p_type, p_size from part
> group by p_brand, p_type, p_size;
> QUERY PLAN
> -----------------------------------------------------------------------------------------------------------------------
> HashAggregate (cost=37992.00..39622.46 rows=163046 width=36) (actual
> time=935.162..992.944 rows=186607 loops=1)
> Group Key: p_brand, p_type, p_size
> -> Seq Scan on part (cost=0.00..30492.00 rows=1000000 width=36)
> (actual time=0.013..156.746 rows=1000000 loops=1)
> Planning time: 0.308 ms
> Execution time: 1001.889 ms
>
> In above example,
> Without MVStat-> estimated: 100000 Actual: 186607
> With MVStat-> estimated: 163046 Actual: 186607
>
Thanks. Those plans match my experiments with the TPC-H data set,
although I've been playing with the smallest scale (1GB).
It's not very difficult to make the estimation error arbitrary large,
e.g. by using perfectly correlated (identical) columns.
regard
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-25 05:55 Michael Paquier <michael.paquier@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
3 siblings, 1 reply; 70+ messages in thread
From: Michael Paquier @ 2017-01-25 05:55 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Wed, Jan 4, 2017 at 11:35 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> On 01/03/2017 05:22 PM, Tomas Vondra wrote:
>>
>> On 01/03/2017 02:42 PM, Dilip Kumar wrote:
>
> ...
>>>
>>> I think it should be easily reproducible, in case it's not I can send
>>> call stack or core dump.
>>>
>>
>> Thanks for the report. It was trivial to reproduce and it turned out to
>> be a fairly simple bug. Will send a new version of the patch soon.
>>
>
> Attached is v22 of the patch series, rebased to current master and fixing
> the reported bug. I haven't made any other changes - the issues reported by
> Petr are mostly minor, so I've decided to wait a bit more for (hopefully)
> other reviews.
And nothing has happened since. Are there people willing to review
this patch and help it proceed? As this patch is quite large, I am not
sure if it is fit to join the last CF. Thoughts?
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-25 12:56 Alvaro Herrera <alvherre@2ndquadrant.com>
parent: Michael Paquier <michael.paquier@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Alvaro Herrera @ 2017-01-25 12:56 UTC (permalink / raw)
To: Michael Paquier <michael.paquier@gmail.com>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Michael Paquier wrote:
> On Wed, Jan 4, 2017 at 11:35 AM, Tomas Vondra
> <tomas.vondra@2ndquadrant.com> wrote:
> > Attached is v22 of the patch series, rebased to current master and fixing
> > the reported bug. I haven't made any other changes - the issues reported by
> > Petr are mostly minor, so I've decided to wait a bit more for (hopefully)
> > other reviews.
>
> And nothing has happened since. Are there people willing to review
> this patch and help it proceed?
I am going to grab this patch as committer.
> As this patch is quite large, I am not sure if it is fit to join the
> last CF. Thoughts?
All patches, regardless of size, are welcome to join any commitfest.
The last commitfest is not different in that regard. The rule I
remember is that patches may not arrive *for the first time* in the last
commitfest. This patch has already seen a lot of work in previous
commitfests, so it's fine.
--
Álvaro Herrera https://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-25 21:43 Michael Paquier <michael.paquier@gmail.com>
parent: Alvaro Herrera <alvherre@2ndquadrant.com>
0 siblings, 1 reply; 70+ messages in thread
From: Michael Paquier @ 2017-01-25 21:43 UTC (permalink / raw)
To: Alvaro Herrera <alvherre@2ndquadrant.com>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Wed, Jan 25, 2017 at 9:56 PM, Alvaro Herrera
<alvherre@2ndquadrant.com> wrote:
> Michael Paquier wrote:
>> And nothing has happened since. Are there people willing to review
>> this patch and help it proceed?
>
> I am going to grab this patch as committer.
Thanks, that's good to know.
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-26 09:03 Ideriha, Takeshi <ideriha.takeshi@jp.fujitsu.com>
parent: Michael Paquier <michael.paquier@gmail.com>
0 siblings, 2 replies; 70+ messages in thread
From: Ideriha, Takeshi @ 2017-01-26 09:03 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Michael Paquier <michael.paquier@gmail.com>; Alvaro Herrera <alvherre@2ndquadrant.com>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Hi
When you have time, could you rebase the pathes?
Some patches cannot be applied to the current HEAD.
0001 patch can be applied but the following 0002 patch cannot be.
I've just started reading your patch (mainly docs and README, not yet source code.)
Though these are minor things, I've found some typos or mistakes in the document and README.
>+ statistics on the table. The statistics will be created in the in the
>+ current database. The statistics will be owned by the user issuing
Regarding line 629 at 0002-PATCH-shared-infrastructure-and-ndistinct-coeffi-v22.patch,
there is a double "in the".
>+ knowledge of a value in the first column is sufficient for detemining the
>+ value in the other column. Then functional dependencies are built on those
Regarding line 701 at 0002-PATCH,
"determining" is mistakenly spelled "detemining".
>@@ -0,0 +1,98 @@
>+Multivariate statististics
>+==========================
Regarding line 2415 at 0002-PATCH, "statististics" should be statistics
>+ <refnamediv>
>+ <refname>CREATE STATISTICS</refname>
>+ <refpurpose>define a new statistics</refpurpose>
>+ </refnamediv>
>+ <refnamediv>
>+ <refname>DROP STATISTICS</refname>
>+ <refpurpose>remove a statistics</refpurpose>
>+ </refnamediv>
Regarding line 612 and 771 at 0002-PATCH,
I assume saying "multiple statistics" explicitly is easier to understand to users
since these commands don't for the statistics we already have in the pg_statistics in my understanding.
>+ [1] http://en.wikipedia.org/wiki/Database_normalization
Regarding line 386 at 0003-PATCH, is it better to change this link to this one:
https://en.wikipedia.org/wiki/Functional_dependency ?
README.dependencies cites directly above link.
Though I pointed out these typoes and so on,
I believe these feedback are less priority compared to the source code itself.
So please work on my feedback if you have time.
regards,
Ideriha Takeshi
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-26 09:43 Dilip Kumar <dilipbalaut@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
0 siblings, 1 reply; 70+ messages in thread
From: Dilip Kumar @ 2017-01-26 09:43 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Thu, Jan 5, 2017 at 3:27 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> Thanks. Those plans match my experiments with the TPC-H data set, although
> I've been playing with the smallest scale (1GB).
>
> It's not very difficult to make the estimation error arbitrary large, e.g.
> by using perfectly correlated (identical) columns.
I have done an initial review for ndistint and histogram patches,
there are few review comments.
ndistinct
---------
1. Duplicate statistics:
postgres=# create statistics s with (ndistinct) on (a,c) from t;
2017-01-07 16:21:54.575 IST [63817] ERROR: duplicate key value
violates unique constraint "pg_mv_statistic_name_index"
2017-01-07 16:21:54.575 IST [63817] DETAIL: Key (staname,
stanamespace)=(s, 2200) already exists.
2017-01-07 16:21:54.575 IST [63817] STATEMENT: create statistics s
with (ndistinct) on (a,c) from t;
ERROR: duplicate key value violates unique constraint
"pg_mv_statistic_name_index"
DETAIL: Key (staname, stanamespace)=(s, 2200) already exists.
For duplicate statistics, I think we can check the existence of the
statistics and give more meaningful error code something statistics
"s" already exist.
2. Typo
+ /*
+ * Sort the attnums, which makes detecting duplicies somewhat
+ * easier, and it does not hurt (it does not affect the efficiency,
+ * onlike for indexes, for example).
+ */
/onlike/unlike
3. Typo
/*
* Find attnims of MV stats using the mvoid.
*/
int2vector *
find_mv_attnums(Oid mvoid, Oid *relid)
/attnims/attnums
histograms
--------------
+ if (matches[i] == MVSTATS_MATCH_FULL)
+ s += mvhist->buckets[i]->ntuples;
+ else if (matches[i] == MVSTATS_MATCH_PARTIAL)
+ s += 0.5 * mvhist->buckets[i]->ntuples;
Isn't it will be better that take some percentage of the bucket based
on the number of distinct element for partial matching buckets.
+static int
+update_match_bitmap_histogram(PlannerInfo *root, List *clauses,
+ int2vector *stakeys,
+ MVSerializedHistogram mvhist,
+ int nmatches, char *matches,
+ bool is_or)
+{
+ int i;
For each clause we are processing all the buckets, can't we use some
data structure which can make multi-dimensions information searching
faster.
Something like HTree, RTree, Maybe storing histogram in these formats
will be difficult?
--
Regards,
Dilip Kumar
EnterpriseDB: http://www.enterprisedb.com
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-26 11:01 Kyotaro HORIGUCHI <horiguchi.kyotaro@lab.ntt.co.jp>
parent: Ideriha, Takeshi <ideriha.takeshi@jp.fujitsu.com>
1 sibling, 1 reply; 70+ messages in thread
From: Kyotaro HORIGUCHI @ 2017-01-26 11:01 UTC (permalink / raw)
To: ideriha.takeshi@jp.fujitsu.com; +Cc: tomas.vondra@2ndquadrant.com; dilipbalaut@gmail.com; Langote_Amit_f8@lab.ntt.co.jp; dean.a.rasheed@gmail.com; hlinnaka@iki.fi; robertmhaas@gmail.com; ishii@postgresql.org; david@pgmasters.net; michael.paquier@gmail.com; alvherre@2ndquadrant.com; tgl@sss.pgh.pa.us; petr@2ndquadrant.com; jeff.janes@gmail.com; pgsql-hackers
Hello, I'll return on this since this should welcome more eyeballs.
At Thu, 26 Jan 2017 09:03:10 +0000, "Ideriha, Takeshi" <ideriha.takeshi@jp.fujitsu.com> wrote in <4E72940DA2BF16479384A86D54D0988A565822A9@G01JPEXMBKW04>
> Hi
>
> When you have time, could you rebase the pathes?
> Some patches cannot be applied to the current HEAD.
For those who are willing to look this,
352a24a1f9d6f7d4abb1175bfd22acc358f43140 breaks this. So just
before it can accept this patches cleanly.
> 0001 patch can be applied but the following 0002 patch cannot be.
>
> I've just started reading your patch (mainly docs and README, not yet source code.)
>
> Though these are minor things, I've found some typos or mistakes in the document and README.
>
> >+ statistics on the table. The statistics will be created in the in the
> >+ current database. The statistics will be owned by the user issuing
>
> Regarding line 629 at 0002-PATCH-shared-infrastructure-and-ndistinct-coeffi-v22.patch,
> there is a double "in the".
>
> >+ knowledge of a value in the first column is sufficient for detemining the
> >+ value in the other column. Then functional dependencies are built on those
>
> Regarding line 701 at 0002-PATCH,
> "determining" is mistakenly spelled "detemining".
>
>
> >@@ -0,0 +1,98 @@
> >+Multivariate statististics
> >+==========================
>
> Regarding line 2415 at 0002-PATCH, "statististics" should be statistics
>
>
> >+ <refnamediv>
> >+ <refname>CREATE STATISTICS</refname>
> >+ <refpurpose>define a new statistics</refpurpose>
> >+ </refnamediv>
>
> >+ <refnamediv>
> >+ <refname>DROP STATISTICS</refname>
> >+ <refpurpose>remove a statistics</refpurpose>
> >+ </refnamediv>
>
> Regarding line 612 and 771 at 0002-PATCH,
> I assume saying "multiple statistics" explicitly is easier to understand to users
> since these commands don't for the statistics we already have in the pg_statistics in my understanding.
>
> >+ [1] http://en.wikipedia.org/wiki/Database_normalization
>
> Regarding line 386 at 0003-PATCH, is it better to change this link to this one:
> https://en.wikipedia.org/wiki/Functional_dependency ?
> README.dependencies cites directly above link.
>
> Though I pointed out these typoes and so on,
> I believe these feedback are less priority compared to the source code itself.
>
> So please work on my feedback if you have time.
README.dependencies
> dependencies, and for each one count the number of rows rows consistent it.
"of rows rows consistent it" => "or rows consistent with it"?
> are in fact consistent with the functinal dependency, i.e. that given the a
"that given the a" => "that given a" ?
dependencies.c:
dependency_dgree():
- The k is assumed larger than 1. I think assertion is required.
- "/* end of the preceding group */" seems to be better if it
is just after the "if (multi_sort.." currently just after it.
- The following comment seems mis-edited.
> * If there is a single are no contradicting rows, count the group
> * as supporting, otherwise contradicting.
maybe this would be like the following? The varialbe counting
the first "contradiction" is named "n_violations". This seems
somewhat confusing.
> * If there are no violating rows up to here, count the group
> * as supporting, otherwise contradicting.
- "/* first columns match, but the last one does not"
else if (multi_sort_compare_dims((k - 1), (k - 1), ...
The above comparison should use multi_sort_compare_dim, not
dims
- This function counts "n_contradicting_rows" but it is not
referenced. Anyway n_contradicting_rows = numrows -
n_supporing_rows so it and n_contradicting seem
unncecessary.
build_mv_dependencies():
- In the commnet,
"* covering jut 2 columns, to the largest ones, covering all columns"
"* included int the statistics. We start from the smallest ones because we"
l1: "jut" => "just", l2: "int" => "in"
mvstats.h:
- struct MVDependencyData/ MVDependenciesData
The varialbe length member at the last of the structs should
be defined using FLEXIBLE_ARRAY_MEMBER, from the convention.
- I'm not sure how much it impacts performance, but some
struct members seems to have a bit too wide types. For
example, MVDepedenciesData.type is of int32 but it can have
only '1' for now and it won't be two-digits. Also ndeps
cannot be so large.
common.c:
multi_sort_compare_dims needs comment.
general:
This patch uses int16 as the type of attrubute number but it
might be better to use AttrNumber for the purpose.
(Specifically it seems defined as the type for an attribute
index but also used as the varialbe for number of attributes)
Sorry for the random comment in advance. I'll learn this further.
regards,
--
Kyotaro Horiguchi
NTT Open Source Software Center
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-30 16:12 Alvaro Herrera <alvherre@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
3 siblings, 1 reply; 70+ messages in thread
From: Alvaro Herrera @ 2017-01-30 16:12 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Tomas Vondra wrote:
> On 01/03/2017 05:22 PM, Tomas Vondra wrote:
> > On 01/03/2017 02:42 PM, Dilip Kumar wrote:
> ...
> > > I think it should be easily reproducible, in case it's not I can send
> > > call stack or core dump.
> > >
> >
> > Thanks for the report. It was trivial to reproduce and it turned out to
> > be a fairly simple bug. Will send a new version of the patch soon.
> >
>
> Attached is v22 of the patch series, rebased to current master and fixing
> the reported bug. I haven't made any other changes - the issues reported by
> Petr are mostly minor, so I've decided to wait a bit more for (hopefully)
> other reviews.
Hmm. So we have a catalog pg_mv_statistics which stores two things:
1. the configuration regarding mvstats that have been requested by user
via CREATE/ALTER STATISTICS
2. the actual values captured from the above, via ANALYZE
I think this conflates two things that really are separate, given their
different timings and usage patterns. This decision is causing the
catalog to have columns enabled/built flags for each set of stats
requested, which looks a bit odd. In particular, the fact that you have
to heap_update the catalog in order to add more stuff as it's built
looks inconvenient.
Have you thought about having the "requested" bits be separate from the
actual computed values? Something like
pg_mv_statistics
starelid
staname
stanamespace
staowner -- all the above as currently
staenabled array of "char" {d,f,s}
stakeys
// no CATALOG_VARLEN here
where each char in the staenabled array has a #define and indicates one
type, "ndistinct", "functional dep", "selectivity" etc.
The actual values computed by ANALYZE would live in a catalog like:
pg_mv_statistics_values
stvstaid -- OID of the corresponding pg_mv_statistics row. Needed?
stvrelid -- same as starelid
stvkeys -- same as stakeys
#ifdef CATALOG_VARLEN
stvkind 'd' or 'f' or 's', etc
stvvalue the bytea blob
#endif
I think that would be simpler, both conceptually and in terms of code.
The other angle to consider is planner-side: how does the planner gets
to the values? I think as far as the planner goes, the first catalog
doesn't matter at all, because a statistics type that has been enabled
but not computed is not interesting at all; planner only cares about the
values in the second catalog (this is why I added stvkeys). Currently
you're just caching a single pg_mv_statistics row in get_relation_info
(and only if any of the "built" flags is set), which is simple. With my
proposed change, you'd need to keep multiple pg_mv_statistics_values
rows.
But maybe you already tried something like what I propose and there's a
reason not to do it?
--
Álvaro Herrera https://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-30 16:55 Alvaro Herrera <alvherre@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
3 siblings, 1 reply; 70+ messages in thread
From: Alvaro Herrera @ 2017-01-30 16:55 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Minor nitpicks:
Let me suggest to use get_attnum() in CreateStatistics instead of
SearchSysCacheAttName for each column. Also, we use type AttrNumber for
attribute numbers rather than int16. Finally in the same function you
have an erroneous ERRCODE_UNDEFINED_COLUMN which should be
ERRCODE_DUPLICATE_COLUMN in the loop that searches for duplicates.
May I suggest that compare_int16 be named attnum_cmp (just to be
consistent with other qsort comparators) and look like
return *((const AttrNumber *) a) - *((const AttrNumber *) b);
instead of memcmp?
--
Álvaro Herrera https://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-30 19:12 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Kyotaro HORIGUCHI <horiguchi.kyotaro@lab.ntt.co.jp>
0 siblings, 3 replies; 70+ messages in thread
From: Tomas Vondra @ 2017-01-30 19:12 UTC (permalink / raw)
To: Kyotaro HORIGUCHI <horiguchi.kyotaro@lab.ntt.co.jp>; ideriha.takeshi@jp.fujitsu.com; +Cc: dilipbalaut@gmail.com; Langote_Amit_f8@lab.ntt.co.jp; dean.a.rasheed@gmail.com; hlinnaka@iki.fi; robertmhaas@gmail.com; ishii@postgresql.org; david@pgmasters.net; michael.paquier@gmail.com; alvherre@2ndquadrant.com; tgl@sss.pgh.pa.us; petr@2ndquadrant.com; jeff.janes@gmail.com; pgsql-hackers
Hi everyone,
thanks for the reviews! Attached is v23 of the patch series, addressing
most of the points raised in the reviews.
A quick summary of the changes (I'll respond to the other threads for
points that deserve a bit more detailed discussion):
0) Rebase to current master. The main culprit was the pesky logical
replication patch committed a week ago, because SUBSCRIPTION and
STATISTICS are right next to each other in gram.y, various switches etc.
1) Many typos, mentioned by all the reviewers.
2) I've added a short explanation (in alter_table.sgml) of how ALTER
TABLE ... DROP COLUMN handles multivariate statistics, i.e. that those
are only dropped if there would be a single remaining column.
3) I've reworded 'thoroughly' to 'in more detail' in planstats.sgml, to
make Petr happy ;-)
4) Added missing comments to get_statistics_oid, RelationGetMVStatList,
update_mv_stats, ndistinct_for_combination. Also update_mv_stats() was
not used outside common.c, so I've made it static and removed the
prototype from mvstats.h.
5) I've changed 'statistics does not exist' to 'statistics do not exist'
on a number of places.
6) Removed XXX about checking for duplicates in CreateStatistics. I
agree with Petr that we shouldn't do such checks, as we're not doing
that for other objects (e.g. indexes).
7) I've moved moved the code loading statistics from get_relation_info
into a new function get_relation_statistics, to get rid of the
if (true)
{
...
}
block, which was there due to mimicking how index details are loaded
without having hasindex-like flag. I like this better than merging the
block into get_relation_info directly.
8) I've changed 'a statistics' to 'multivariate statistics' on a few
places in sgml docs, to make it clear it's not referring to the
'regular' statistics (e.g. at CREATE/DROP STATISTICS, mentioned by
Ideriha Takeshi).
9) I've changed the link in README.dependencies to
https://en.wikipedia.org/wiki/Functional_dependency as proposed by
Ideriha Takeshi. I'm pretty sure the wiki page about database
normalization, referenced by the original link, included a nice
functional dependency example some time ago, but it seems to have
changed and the new link is better.
But perhaps it's not a good idea to link to wikipedia, as the pages
clearly change quite significantly?
10) The CREATE STATISTICS now reports a nice 'already exists' message,
instead of the 'duplicate key', pointed out by Dilip.
11) MVNDistinctItem/MVNDistinctData now use FLEXIBLE_ARRAY_MEMBER for
the array, just like the other structs.
On 01/26/2017 12:01 PM, Kyotaro HORIGUCHI wrote:
> dependencies.c:
>
> dependency_dgree():
>
> - The k is assumed larger than 1. I think assertion is required.
>
> - "/* end of the preceding group */" seems to be better if it
> is just after the "if (multi_sort.." currently just after it.
>
> - The following comment seems mis-edited.
> > * If there is a single are no contradicting rows, count the group
> > * as supporting, otherwise contradicting.
>
> maybe this would be like the following? The varialbe counting
> the first "contradiction" is named "n_violations". This seems
> somewhat confusing.
>
> > * If there are no violating rows up to here, count the group
> > * as supporting, otherwise contradicting.
>
> - "/* first columns match, but the last one does not"
> else if (multi_sort_compare_dims((k - 1), (k - 1), ...
>
> The above comparison should use multi_sort_compare_dim, not
> dims
>
> - This function counts "n_contradicting_rows" but it is not
> referenced. Anyway n_contradicting_rows = numrows -
> n_supporing_rows so it and n_contradicting seem
> unncecessary.
>
Yes, absolutely. This was clearly unnecessary remainder of the original
implementation, and I failed to clean it up after adopting Dean's idea
of continuous dependency degree.
I've also reworked the method a bit, moving handling of the last group
into the main loop (instead of doing that separately right after the
loop, which I think was a bit ugly anyway). Can you check if you're
happy with the code & comments now?
>
> mvstats.h:
>
> - struct MVDependencyData/ MVDependenciesData
>
> The varialbe length member at the last of the structs should
> be defined using FLEXIBLE_ARRAY_MEMBER, from the convention.
>
Yes, fixed. The other structures already used that macro, but I failed
to notice MVDependencyData/ MVDependenciesData need that fix too.
>
> - I'm not sure how much it impacts performance, but some
> struct members seems to have a bit too wide types. For
> example, MVDepedenciesData.type is of int32 but it can have
> only '1' for now and it won't be two-digits. Also ndeps
> cannot be so large.
>
I doubt the impact on performance is measurable, particularly for the
global fields (e.g. nbuckets is tiny compared to the space needed for
the buckets themselves).
But I think you're right we shouldn't use fields wider than actually
needed (e.g. using uint32 for nbuckets is a bit insane, and uint16 would
be just fine). It's not just a matter of performance, but also a way to
document expected values etc.
I'll go through the fields and use smaller data types where appropriate.
>
> general:
> This patch uses int16 as the type of attrubute number but it
> might be better to use AttrNumber for the purpose.
> (Specifically it seems defined as the type for an attribute
> index but also used as the varialbe for number of attributes)
>
Agreed. Will check with the struct members.
>
> Sorry for the random comment in advance. I'll learn this further.
>
Thanks for the review!
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
Attachments:
[binary/octet-stream] 0001-teach-pull_-varno-varattno-_walker-about-Restric-v23.patch (1.4K, ../../1fa982dc-ffb4-eef7-eb55-d8cf7d84f1c4@2ndquadrant.com/2-0001-teach-pull_-varno-varattno-_walker-about-Restric-v23.patch)
download | inline diff:
From 5dfca15a5db7cabd9145c76715cb9aea396ec83f Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:35:05 +0200
Subject: [PATCH 1/9] teach pull_(varno|varattno)_walker about RestrictInfo
otherwise pull_varnos fails when processing OR clauses
---
src/backend/optimizer/util/var.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
diff --git a/src/backend/optimizer/util/var.c b/src/backend/optimizer/util/var.c
index cf326ae..7056bcd 100644
--- a/src/backend/optimizer/util/var.c
+++ b/src/backend/optimizer/util/var.c
@@ -196,6 +196,13 @@ pull_varnos_walker(Node *node, pull_varnos_context *context)
context->sublevels_up--;
return result;
}
+ if (IsA(node, RestrictInfo))
+ {
+ RestrictInfo *rinfo = (RestrictInfo*)node;
+ context->varnos = bms_add_members(context->varnos,
+ rinfo->clause_relids);
+ return false;
+ }
return expression_tree_walker(node, pull_varnos_walker,
(void *) context);
}
@@ -244,6 +251,15 @@ pull_varattnos_walker(Node *node, pull_varattnos_context *context)
return false;
}
+ if (IsA(node, RestrictInfo))
+ {
+ RestrictInfo *rinfo = (RestrictInfo *)node;
+
+ return expression_tree_walker((Node*)rinfo->clause,
+ pull_varattnos_walker,
+ (void*) context);
+ }
+
/* Should not find an unplanned subquery */
Assert(!IsA(node, Query));
--
2.5.5
[binary/octet-stream] 0002-PATCH-shared-infrastructure-and-ndistinct-coeffi-v23.patch (141.9K, ../../1fa982dc-ffb4-eef7-eb55-d8cf7d84f1c4@2ndquadrant.com/3-0002-PATCH-shared-infrastructure-and-ndistinct-coeffi-v23.patch)
download | inline diff:
From 7f32d263f8b6e6aea95d6544fd10e46177346fb3 Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:35:47 +0200
Subject: [PATCH 2/9] PATCH: shared infrastructure and ndistinct coefficients
Basic infrastructure shared by all kinds of multivariate stats, most
importantly:
- adds a new system catalog (pg_mv_statistic)
- CREATE STATISTICS name ON (columns) FROM table
- DROP STATISTICS name
- ALTER STATISTICS ... OWNER TO / SET SCHEMA / RENAME
- implementation of ndistinct coefficients (the simplest type of
multivariate statistics)
- computing ndistinct coefficients during ANALYZE
- updates existing regression tests (new catalog etc.)
- modifies estimate_num_groups() to use ndistinct if available
The current implementation requires a valid 'ltopr' for the columns, so
that we can sort the sample rows in various ways, both in this patch
and other kinds of statistics. Maybe this restriction could be relaxed
in the future, requiring just 'eqopr' in case of stats not sorting the
data (e.g. functional dependencies and MCV lists).
Some of the stats implemented in follow-up patches (e.g. functional
dependencies and MCV list with limited functionality) might be made
to work with hashes of the values. That would save a lot of space
for storing the statistics, and it would be sufficient for estimating
equality conditions.
creating statistics
-------------------
Statistics are created by CREATE STATISTICS command, with this syntax:
CREATE STATISTICS statistics_name ON (columns) FROM table
where 'statistics_name' may be a fully-qualified name (i.e. specifying
a schema). It's expected that we'll eventually add support for join
statistics, referencing tables that may be located in different schemas,
so we can't make the name unique per-table (like constraints), and we
can't just pick one of the table schemas.
dropping statistics
-------------------
The statistics may be dropped automatically using DROP STATISTICS.
After ALTER TABLE ... DROP COLUMN, statistics referencing are:
(a) dropped, if the statistics would reference only one column
(b) retained, but modified on the next ANALYZE
The goal of the lazy cleanup is not to disrupt the optimizer, but
arguably this is over-engineering and it might also made work just
like for indexes by simply dropping all dependent statistics on
ALTER TABLE ... DROP COLUMN. If the user wants to minimize impact,
the smaller statistics needs to be created explicitly in advance.
This also adds a simple list of statistics to \d in psql.
ndistinct coefficients
----------------------
The patch only implements a very simple type of statistics, tracking
the number of groups for different combinations of columns. For
example given columns (a,b,c) the statistics will estimate number
of distinct combinations of values in (a,b), (a,c), (b,c) and (a,b,c).
This is then used in estimate_num_groups() for estimating cardinality
of GROUP BY and similar clauses.
pg_ndistinct data type
----------------------
The patch introduces pg_ndistinct, a new varlena data type used for
serialized version of ndistinct coefficients. Internally it's just
a bytea value, but it allows us to control casting, input/output
and so on. It's somewhat inspired by pg_node_tree.
---
doc/src/sgml/catalogs.sgml | 108 +++++
doc/src/sgml/planstats.sgml | 141 +++++++
doc/src/sgml/ref/allfiles.sgml | 3 +
doc/src/sgml/ref/alter_statistics.sgml | 115 ++++++
doc/src/sgml/ref/alter_table.sgml | 8 +-
doc/src/sgml/ref/create_statistics.sgml | 152 +++++++
doc/src/sgml/ref/drop_statistics.sgml | 91 ++++
doc/src/sgml/reference.sgml | 3 +
src/backend/catalog/Makefile | 1 +
src/backend/catalog/aclchk.c | 27 ++
src/backend/catalog/dependency.c | 11 +-
src/backend/catalog/heap.c | 101 +++++
src/backend/catalog/namespace.c | 56 +++
src/backend/catalog/objectaddress.c | 53 +++
src/backend/catalog/system_views.sql | 10 +
src/backend/commands/Makefile | 6 +-
src/backend/commands/alter.c | 3 +
src/backend/commands/analyze.c | 8 +
src/backend/commands/dropcmds.c | 4 +
src/backend/commands/event_trigger.c | 3 +
src/backend/commands/statscmds.c | 259 ++++++++++++
src/backend/nodes/copyfuncs.c | 16 +
src/backend/nodes/outfuncs.c | 18 +
src/backend/optimizer/util/plancat.c | 72 +++-
src/backend/parser/gram.y | 58 ++-
src/backend/tcop/utility.c | 12 +
src/backend/utils/Makefile | 2 +-
src/backend/utils/adt/selfuncs.c | 168 +++++++-
src/backend/utils/cache/relcache.c | 78 ++++
src/backend/utils/cache/syscache.c | 23 ++
src/backend/utils/mvstats/Makefile | 17 +
src/backend/utils/mvstats/README.ndistinct | 22 +
src/backend/utils/mvstats/README.stats | 98 +++++
src/backend/utils/mvstats/common.c | 391 ++++++++++++++++++
src/backend/utils/mvstats/common.h | 80 ++++
src/backend/utils/mvstats/mvdist.c | 597 +++++++++++++++++++++++++++
src/bin/psql/describe.c | 44 ++
src/include/catalog/dependency.h | 5 +-
src/include/catalog/heap.h | 1 +
src/include/catalog/indexing.h | 7 +
src/include/catalog/namespace.h | 2 +
src/include/catalog/pg_cast.h | 4 +
src/include/catalog/pg_mv_statistic.h | 78 ++++
src/include/catalog/pg_proc.h | 9 +
src/include/catalog/pg_type.h | 4 +
src/include/catalog/toasting.h | 1 +
src/include/commands/defrem.h | 4 +
src/include/nodes/nodes.h | 2 +
src/include/nodes/parsenodes.h | 11 +
src/include/nodes/relation.h | 27 ++
src/include/utils/acl.h | 1 +
src/include/utils/builtins.h | 6 +
src/include/utils/mvstats.h | 57 +++
src/include/utils/rel.h | 4 +
src/include/utils/relcache.h | 1 +
src/include/utils/syscache.h | 2 +
src/test/regress/expected/mv_ndistinct.out | 117 ++++++
src/test/regress/expected/object_address.out | 7 +-
src/test/regress/expected/opr_sanity.out | 3 +-
src/test/regress/expected/rules.out | 8 +
src/test/regress/expected/sanity_check.out | 1 +
src/test/regress/expected/type_sanity.out | 11 +-
src/test/regress/parallel_schedule | 3 +
src/test/regress/serial_schedule | 1 +
src/test/regress/sql/mv_ndistinct.sql | 68 +++
src/test/regress/sql/object_address.sql | 4 +-
66 files changed, 3286 insertions(+), 22 deletions(-)
create mode 100644 doc/src/sgml/ref/alter_statistics.sgml
create mode 100644 doc/src/sgml/ref/create_statistics.sgml
create mode 100644 doc/src/sgml/ref/drop_statistics.sgml
create mode 100644 src/backend/commands/statscmds.c
create mode 100644 src/backend/utils/mvstats/Makefile
create mode 100644 src/backend/utils/mvstats/README.ndistinct
create mode 100644 src/backend/utils/mvstats/README.stats
create mode 100644 src/backend/utils/mvstats/common.c
create mode 100644 src/backend/utils/mvstats/common.h
create mode 100644 src/backend/utils/mvstats/mvdist.c
create mode 100644 src/include/catalog/pg_mv_statistic.h
create mode 100644 src/include/utils/mvstats.h
create mode 100644 src/test/regress/expected/mv_ndistinct.out
create mode 100644 src/test/regress/sql/mv_ndistinct.sql
diff --git a/doc/src/sgml/catalogs.sgml b/doc/src/sgml/catalogs.sgml
index 086fafc..2a7bd6c 100644
--- a/doc/src/sgml/catalogs.sgml
+++ b/doc/src/sgml/catalogs.sgml
@@ -201,6 +201,11 @@
</row>
<row>
+ <entry><link linkend="catalog-pg-mv-statistic"><structname>pg_mv_statistic</structname></link></entry>
+ <entry>multivariate statistics</entry>
+ </row>
+
+ <row>
<entry><link linkend="catalog-pg-namespace"><structname>pg_namespace</structname></link></entry>
<entry>schemas</entry>
</row>
@@ -4211,6 +4216,109 @@
</table>
</sect1>
+ <sect1 id="catalog-pg-mv-statistic">
+ <title><structname>pg_mv_statistic</structname></title>
+
+ <indexterm zone="catalog-pg-mv-statistic">
+ <primary>pg_mv_statistic</primary>
+ </indexterm>
+
+ <para>
+ The catalog <structname>pg_mv_statistic</structname>
+ holds multivariate statistics about combinations of columns.
+ </para>
+
+ <table>
+ <title><structname>pg_mv_statistic</> Columns</title>
+
+ <tgroup cols="4">
+ <thead>
+ <row>
+ <entry>Name</entry>
+ <entry>Type</entry>
+ <entry>References</entry>
+ <entry>Description</entry>
+ </row>
+ </thead>
+
+ <tbody>
+
+ <row>
+ <entry><structfield>starelid</structfield></entry>
+ <entry><type>oid</type></entry>
+ <entry><literal><link linkend="catalog-pg-class"><structname>pg_class</structname></link>.oid</literal></entry>
+ <entry>The table that the described columns belongs to</entry>
+ </row>
+
+ <row>
+ <entry><structfield>staname</structfield></entry>
+ <entry><type>name</type></entry>
+ <entry></entry>
+ <entry>Name of the statistic.</entry>
+ </row>
+
+ <row>
+ <entry><structfield>stanamespace</structfield></entry>
+ <entry><type>oid</type></entry>
+ <entry><literal><link linkend="catalog-pg-namespace"><structname>pg_namespace</structname></link>.oid</literal></entry>
+ <entry>
+ The OID of the namespace that contains this statistic
+ </entry>
+ </row>
+
+ <row>
+ <entry><structfield>staowner</structfield></entry>
+ <entry><type>oid</type></entry>
+ <entry><literal><link linkend="catalog-pg-authid"><structname>pg_authid</structname></link>.oid</literal></entry>
+ <entry>Owner of the statistic</entry>
+ </row>
+
+ <row>
+ <entry><structfield>ndist_enabled</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, ndistinct coefficients will be computed for the combination of
+ columns, covered by the statistics. This does not mean the coefficients
+ are already computed, though.
+ </entry>
+ </row>
+
+ <row>
+ <entry><structfield>ndist_built</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, ndistinct coefficients are already computed and available for
+ use during query estimation.
+ </entry>
+ </row>
+
+ <row>
+ <entry><structfield>stakeys</structfield></entry>
+ <entry><type>int2vector</type></entry>
+ <entry><literal><link linkend="catalog-pg-attribute"><structname>pg_attribute</structname></link>.attnum</literal></entry>
+ <entry>
+ This is an array of values that indicate which table columns this
+ statistic covers. For example a value of <literal>1 3</literal> would
+ mean that the first and the third table columns make up the statistic key.
+ </entry>
+ </row>
+
+ <row>
+ <entry><structfield>standist</structfield></entry>
+ <entry><type>pg_ndistinct</type></entry>
+ <entry></entry>
+ <entry>
+ Ndistict coefficients, serialized as <structname>pg_ndistinct</> type.
+ </entry>
+ </row>
+
+ </tbody>
+ </tgroup>
+ </table>
+ </sect1>
+
<sect1 id="catalog-pg-namespace">
<title><structname>pg_namespace</structname></title>
diff --git a/doc/src/sgml/planstats.sgml b/doc/src/sgml/planstats.sgml
index b73c66b..d5b975d 100644
--- a/doc/src/sgml/planstats.sgml
+++ b/doc/src/sgml/planstats.sgml
@@ -448,4 +448,145 @@ rows = (outer_cardinality * inner_cardinality) * selectivity
</sect1>
+ <sect1 id="multivariate-statistics">
+ <title>Multivariate Statistics</title>
+
+ <indexterm zone="multivariate-statistics">
+ <primary>multivariate statistics</primary>
+ <secondary>planner</secondary>
+ </indexterm>
+
+ <para>
+ The examples presented in <xref linkend="row-estimation-examples"> used
+ statistics about individual columns to compute selectivity estimates.
+ When estimating conditions on multiple columns, the planner assumes
+ independence of the conditions and multiplies the selectivities. When the
+ columns are correlated, the independence assumption is violated, and the
+ estimates may be off by several orders of magnitude, resulting in poor
+ plan choices.
+ </para>
+
+ <para>
+ The examples presented below demonstrate such estimation errors on simple
+ data sets, and also how to resolve them by creating multivariate statistics
+ using <command>CREATE STATISTICS</> command.
+ </para>
+
+ <para>
+ Let's start with a very simple data set - a table with two columns,
+ containing exactly the same values:
+
+<programlisting>
+CREATE TABLE t (a INT, b INT);
+INSERT INTO t SELECT i % 100, i % 100 FROM generate_series(1, 10000) s(i);
+ANALYZE t;
+</programlisting>
+
+ As explained in <xref linkend="planner-stats">, the planner can determine
+ cardinality of <structname>t</structname> using the number of pages and
+ rows is looked up in <structname>pg_class</structname>:
+
+<programlisting>
+SELECT relpages, reltuples FROM pg_class WHERE relname = 't';
+
+ relpages | reltuples
+----------+-----------
+ 45 | 10000
+</programlisting>
+
+ The data distribution is very simple - there are only 100 distinct values
+ in each column, uniformly distributed.
+ </para>
+
+ <para>
+ The following example shows the result of estimating a <literal>WHERE</>
+ condition on the <structfield>a</> column:
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1;
+ QUERY PLAN
+-------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..170.00 rows=100 width=8) (actual time=0.031..2.870 rows=100 loops=1)
+ Filter: (a = 1)
+ Rows Removed by Filter: 9900
+ Planning time: 0.092 ms
+ Execution time: 3.103 ms
+(5 rows)
+</programlisting>
+
+ The planner examines the condition and computes the estimate using
+ <function>eqsel</>, the selectivity function for <literal>=</>, and
+ statistics stored in the <structname>pg_stats</> table. In this case
+ the planner estimates the condition matches 1% rows, and by comparing
+ the estimated and actual number of rows, we see that the estimate is
+ very accurate (in fact exact, as the table is very small).
+ </para>
+
+ <para>
+ Adding a condition on the second column results in the following plan:
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
+ QUERY PLAN
+-----------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=1 width=8) (actual time=0.033..3.006 rows=100 loops=1)
+ Filter: ((a = 1) AND (b = 1))
+ Rows Removed by Filter: 9900
+ Planning time: 0.121 ms
+ Execution time: 3.220 ms
+(5 rows)
+</programlisting>
+
+ The planner estimates the selectivity for each condition individually,
+ arriving to the 1% estimates as above, and then multiplies them, getting
+ the final 0.01% estimate. The plan however shows that this results in
+ a significant underestimate, as the actual number of rows matching the
+ conditions is two orders of magnitude higher than estimated.
+ </para>
+
+ <para>
+ Overestimates, i.e. errors in the opposite direction, are also possible.
+ Consider for example the following combination of range conditions, each
+ matching
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a <= 49 AND b > 49;
+ QUERY PLAN
+------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=2500 width=8) (actual time=1.607..1.607 rows=0 loops=1)
+ Filter: ((a <= 49) AND (b > 49))
+ Rows Removed by Filter: 10000
+ Planning time: 0.050 ms
+ Execution time: 1.623 ms
+(5 rows)
+</programlisting>
+
+ The planner examines both <literal>WHERE</> clauses and estimates them
+ using the <function>scalarltsel</> and <function>scalargtsel</> functions,
+ specified as the selectivity functions matching the <literal><=</> and
+ <literal>></literal> operators. Both conditions match 50% of the
+ table, and assuming independence the planner multiplies them to compute
+ the total estimate of 25%. However as the explain output shows, the actual
+ number of rows is 0, because the columns are correlated and the conditions
+ contradict each other.
+ </para>
+
+ <para>
+ Both estimation errors are caused by violation of the independence
+ assumption, as the two columns contain exactly the same values, and are
+ therefore perfectly correlated. Providing additional information about
+ correlation between columns is the purpose of multivariate statistics,
+ and the rest of this section explains in more detail how the planner
+ leverages them to improve estimates.
+ </para>
+
+ <para>
+ For additional details about multivariate statistics, see
+ <filename>src/backend/utils/mvstats/README.stats</>. There are additional
+ <literal>READMEs</> for each type of statistics, mentioned in the following
+ sections.
+ </para>
+
+ </sect1>
+
</chapter>
diff --git a/doc/src/sgml/ref/allfiles.sgml b/doc/src/sgml/ref/allfiles.sgml
index 0d09f81..a49da6d 100644
--- a/doc/src/sgml/ref/allfiles.sgml
+++ b/doc/src/sgml/ref/allfiles.sgml
@@ -34,6 +34,7 @@ Complete list of usable sgml source files in this directory.
<!ENTITY alterSequence SYSTEM "alter_sequence.sgml">
<!ENTITY alterSubscription SYSTEM "alter_subscription.sgml">
<!ENTITY alterSystem SYSTEM "alter_system.sgml">
+<!ENTITY alterStatistics SYSTEM "alter_statistics.sgml">
<!ENTITY alterTable SYSTEM "alter_table.sgml">
<!ENTITY alterTableSpace SYSTEM "alter_tablespace.sgml">
<!ENTITY alterTSConfig SYSTEM "alter_tsconfig.sgml">
@@ -80,6 +81,7 @@ Complete list of usable sgml source files in this directory.
<!ENTITY createSchema SYSTEM "create_schema.sgml">
<!ENTITY createSequence SYSTEM "create_sequence.sgml">
<!ENTITY createServer SYSTEM "create_server.sgml">
+<!ENTITY createStatistics SYSTEM "create_statistics.sgml">
<!ENTITY createSubscription SYSTEM "create_subscription.sgml">
<!ENTITY createTable SYSTEM "create_table.sgml">
<!ENTITY createTableAs SYSTEM "create_table_as.sgml">
@@ -126,6 +128,7 @@ Complete list of usable sgml source files in this directory.
<!ENTITY dropSchema SYSTEM "drop_schema.sgml">
<!ENTITY dropSequence SYSTEM "drop_sequence.sgml">
<!ENTITY dropServer SYSTEM "drop_server.sgml">
+<!ENTITY dropStatistics SYSTEM "drop_statistics.sgml">
<!ENTITY dropSubscription SYSTEM "drop_subscription.sgml">
<!ENTITY dropTable SYSTEM "drop_table.sgml">
<!ENTITY dropTableSpace SYSTEM "drop_tablespace.sgml">
diff --git a/doc/src/sgml/ref/alter_statistics.sgml b/doc/src/sgml/ref/alter_statistics.sgml
new file mode 100644
index 0000000..3f477cb
--- /dev/null
+++ b/doc/src/sgml/ref/alter_statistics.sgml
@@ -0,0 +1,115 @@
+<!--
+doc/src/sgml/ref/alter_statistics.sgml
+PostgreSQL documentation
+-->
+
+<refentry id="SQL-ALTERSTATISTICS">
+ <indexterm zone="sql-alterstatistics">
+ <primary>ALTER STATISTICS</primary>
+ </indexterm>
+
+ <refmeta>
+ <refentrytitle>ALTER STATISTICS</refentrytitle>
+ <manvolnum>7</manvolnum>
+ <refmiscinfo>SQL - Language Statements</refmiscinfo>
+ </refmeta>
+
+ <refnamediv>
+ <refname>ALTER STATISTICS</refname>
+ <refpurpose>
+ change the definition of a multivariate statistics
+ </refpurpose>
+ </refnamediv>
+
+ <refsynopsisdiv>
+<synopsis>
+ALTER STATISTICS <replaceable class="parameter">name</replaceable> OWNER TO { <replaceable class="PARAMETER">new_owner</replaceable> | CURRENT_USER | SESSION_USER }
+ALTER STATISTICS <replaceable class="parameter">name</replaceable> RENAME TO <replaceable class="parameter">new_name</replaceable>
+ALTER STATISTICS <replaceable class="parameter">name</replaceable> SET SCHEMA <replaceable class="parameter">new_schema</replaceable>
+</synopsis>
+ </refsynopsisdiv>
+
+ <refsect1>
+ <title>Description</title>
+
+ <para>
+ <command>ALTER STATISTICS</command> changes the parameters of an existing
+ multivariate statistics. Any parameters not specifically set in the
+ <command>ALTER STATISTICS</command> command retain their prior settings.
+ </para>
+
+ <para>
+ You must own the statistics to use <command>ALTER STATISTICS</>.
+ To change a statistics' schema, you must also have <literal>CREATE</>
+ privilege on the new schema.
+ To alter the owner, you must also be a direct or indirect member of the new
+ owning role, and that role must have <literal>CREATE</literal> privilege on
+ the statistics' schema. (These restrictions enforce that altering the owner
+ doesn't do anything you couldn't do by dropping and recreating the statistics.
+ However, a superuser can alter ownership of any statistics anyway.)
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>Parameters</title>
+
+ <para>
+ <variablelist>
+ <varlistentry>
+ <term><replaceable class="parameter">name</replaceable></term>
+ <listitem>
+ <para>
+ The name (optionally schema-qualified) of a statistics to be altered.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="PARAMETER">new_owner</replaceable></term>
+ <listitem>
+ <para>
+ The user name of the new owner of the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="parameter">new_name</replaceable></term>
+ <listitem>
+ <para>
+ The new name for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="parameter">new_schema</replaceable></term>
+ <listitem>
+ <para>
+ The new schema for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ </variablelist>
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>Compatibility</title>
+
+ <para>
+ There's no <command>ALTER STATISTICS</command> command in the SQL standard.
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>See Also</title>
+
+ <simplelist type="inline">
+ <member><xref linkend="sql-createstatistics"></member>
+ <member><xref linkend="sql-dropstatistics"></member>
+ </simplelist>
+ </refsect1>
+
+</refentry>
diff --git a/doc/src/sgml/ref/alter_table.sgml b/doc/src/sgml/ref/alter_table.sgml
index da431f8..9ce1079 100644
--- a/doc/src/sgml/ref/alter_table.sgml
+++ b/doc/src/sgml/ref/alter_table.sgml
@@ -119,9 +119,11 @@ ALTER TABLE [ IF EXISTS ] <replaceable class="PARAMETER">name</replaceable>
<para>
This form drops a column from a table. Indexes and
table constraints involving the column will be automatically
- dropped as well. You will need to say <literal>CASCADE</> if
- anything outside the table depends on the column, for example,
- foreign key references or views.
+ dropped as well. Multivariate statistics referencing the column will
+ be dropped only if there would remain a single non-dropped column.
+ You will need to say <literal>CASCADE</> if anything outside the table
+ depends on the column, for example, foreign key references or views.
+
If <literal>IF EXISTS</literal> is specified and the column
does not exist, no error is thrown. In this case a notice
is issued instead.
diff --git a/doc/src/sgml/ref/create_statistics.sgml b/doc/src/sgml/ref/create_statistics.sgml
new file mode 100644
index 0000000..9f6a65c
--- /dev/null
+++ b/doc/src/sgml/ref/create_statistics.sgml
@@ -0,0 +1,152 @@
+<!--
+doc/src/sgml/ref/create_statistics.sgml
+PostgreSQL documentation
+-->
+
+<refentry id="SQL-CREATESTATISTICS">
+ <indexterm zone="sql-createstatistics">
+ <primary>CREATE STATISTICS</primary>
+ </indexterm>
+
+ <refmeta>
+ <refentrytitle>CREATE STATISTICS</refentrytitle>
+ <manvolnum>7</manvolnum>
+ <refmiscinfo>SQL - Language Statements</refmiscinfo>
+ </refmeta>
+
+ <refnamediv>
+ <refname>CREATE STATISTICS</refname>
+ <refpurpose>define multivariate statistics</refpurpose>
+ </refnamediv>
+
+ <refsynopsisdiv>
+<synopsis>
+CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_name</replaceable> ON (
+ <replaceable class="PARAMETER">column_name</replaceable>, <replaceable class="PARAMETER">column_name</replaceable> [, ...])
+ FROM <replaceable class="PARAMETER">table_name</replaceable>
+</synopsis>
+
+ </refsynopsisdiv>
+
+ <refsect1 id="SQL-CREATESTATISTICS-description">
+ <title>Description</title>
+
+ <para>
+ <command>CREATE STATISTICS</command> will create a new multivariate
+ statistics on the table. The statistics will be created in the current
+ database and will be owned by the user issuing the command.
+ </para>
+
+ <para>
+ If a schema name is given (for example, <literal>CREATE STATISTICS
+ myschema.mystat ...</>) then the statistics is created in the specified
+ schema. Otherwise it is created in the current schema. The name of
+ the table must be distinct from the name of any other statistics in the
+ same schema.
+ </para>
+
+ <para>
+ To be able to create a table, you must have <literal>USAGE</literal>
+ privilege on all column types or the type in the <literal>OF</literal>
+ clause, respectively.
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>Parameters</title>
+
+ <variablelist>
+
+ <varlistentry>
+ <term><literal>IF NOT EXISTS</></term>
+ <listitem>
+ <para>
+ Do not throw an error if a statistics with the same name already exists.
+ A notice is issued in this case. Note that there is no guarantee that
+ the existing statistics is anything like the one that would have been
+ created.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="PARAMETER">statistics_name</replaceable></term>
+ <listitem>
+ <para>
+ The name (optionally schema-qualified) of the statistics to be created.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="PARAMETER">table_name</replaceable></term>
+ <listitem>
+ <para>
+ The name (optionally schema-qualified) of the table the statistics should
+ be created on.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="PARAMETER">column_name</replaceable></term>
+ <listitem>
+ <para>
+ The name of a column to be included in the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ </variablelist>
+
+ </refsect1>
+
+ <refsect1 id="SQL-CREATESTATISTICS-examples">
+ <title>Examples</title>
+
+ <para>
+ Create table <structname>t1</> with two functionally dependent columns, i.e.
+ knowledge of a value in the first column is sufficient for determining the
+ value in the other column. Then functional dependencies are built on those
+ columns:
+
+<programlisting>
+CREATE TABLE t1 (
+ a int,
+ b int
+);
+
+INSERT INTO t1 SELECT i/100, i/500
+ FROM generate_series(1,1000000) s(i);
+
+CREATE STATISTICS s1 ON (a, b) FROM t1;
+
+ANALYZE t1;
+
+-- valid combination of values
+EXPLAIN ANALYZE SELECT * FROM t1 WHERE (a = 1) AND (b = 0);
+
+-- invalid combination of values
+EXPLAIN ANALYZE SELECT * FROM t1 WHERE (a = 1) AND (b = 1);
+</programlisting>
+ </para>
+
+ </refsect1>
+
+ <refsect1>
+ <title>Compatibility</title>
+
+ <para>
+ There's no <command>CREATE STATISTICS</command> command in the SQL standard.
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>See Also</title>
+
+ <simplelist type="inline">
+ <member><xref linkend="sql-alterstatistics"></member>
+ <member><xref linkend="sql-dropstatistics"></member>
+ </simplelist>
+ </refsect1>
+</refentry>
diff --git a/doc/src/sgml/ref/drop_statistics.sgml b/doc/src/sgml/ref/drop_statistics.sgml
new file mode 100644
index 0000000..3e73d10
--- /dev/null
+++ b/doc/src/sgml/ref/drop_statistics.sgml
@@ -0,0 +1,91 @@
+<!--
+doc/src/sgml/ref/drop_statistics.sgml
+PostgreSQL documentation
+-->
+
+<refentry id="SQL-DROPSTATISTICS">
+ <indexterm zone="sql-dropstatistics">
+ <primary>DROP STATISTICS</primary>
+ </indexterm>
+
+ <refmeta>
+ <refentrytitle>DROP STATISTICS</refentrytitle>
+ <manvolnum>7</manvolnum>
+ <refmiscinfo>SQL - Language Statements</refmiscinfo>
+ </refmeta>
+
+ <refnamediv>
+ <refname>DROP STATISTICS</refname>
+ <refpurpose>remove multivariate statistics</refpurpose>
+ </refnamediv>
+
+ <refsynopsisdiv>
+<synopsis>
+DROP STATISTICS [ IF EXISTS ] <replaceable class="PARAMETER">name</replaceable> [, ...]
+</synopsis>
+ </refsynopsisdiv>
+
+ <refsect1>
+ <title>Description</title>
+
+ <para>
+ <command>DROP STATISTICS</command> removes statistics from the database.
+ Only the statistics owner, the schema owner, and superuser can drop a
+ statistics.
+ </para>
+
+ </refsect1>
+
+ <refsect1>
+ <title>Parameters</title>
+
+ <variablelist>
+ <varlistentry>
+ <term><literal>IF EXISTS</literal></term>
+ <listitem>
+ <para>
+ Do not throw an error if the statistics do not exist. A notice is
+ issued in this case.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="PARAMETER">name</replaceable></term>
+ <listitem>
+ <para>
+ The name (optionally schema-qualified) of the statistics to drop.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ </variablelist>
+ </refsect1>
+
+ <refsect1>
+ <title>Examples</title>
+
+ <para>
+ ...
+ </para>
+
+ </refsect1>
+
+ <refsect1>
+ <title>Compatibility</title>
+
+ <para>
+ There's no <command>DROP STATISTICS</command> command in the SQL standard.
+ </para>
+ </refsect1>
+
+ <refsect1>
+ <title>See Also</title>
+
+ <simplelist type="inline">
+ <member><xref linkend="sql-alterstatistics"></member>
+ <member><xref linkend="sql-createstatistics"></member>
+ </simplelist>
+ </refsect1>
+
+</refentry>
diff --git a/doc/src/sgml/reference.sgml b/doc/src/sgml/reference.sgml
index 34007d3..abc3f42 100644
--- a/doc/src/sgml/reference.sgml
+++ b/doc/src/sgml/reference.sgml
@@ -60,6 +60,7 @@
&alterSchema;
&alterSequence;
&alterServer;
+ &alterStatistics;
&alterSubscription;
&alterSystem;
&alterTable;
@@ -108,6 +109,7 @@
&createSchema;
&createSequence;
&createServer;
+ &createStatistics;
&createSubscription;
&createTable;
&createTableAs;
@@ -154,6 +156,7 @@
&dropSchema;
&dropSequence;
&dropServer;
+ &dropStatistics;
&dropSubscription;
&dropTable;
&dropTableSpace;
diff --git a/src/backend/catalog/Makefile b/src/backend/catalog/Makefile
index 3136858..5a891a0 100644
--- a/src/backend/catalog/Makefile
+++ b/src/backend/catalog/Makefile
@@ -33,6 +33,7 @@ POSTGRES_BKI_SRCS = $(addprefix $(top_srcdir)/src/include/catalog/,\
pg_attrdef.h pg_constraint.h pg_inherits.h pg_index.h pg_operator.h \
pg_opfamily.h pg_opclass.h pg_am.h pg_amop.h pg_amproc.h \
pg_language.h pg_largeobject_metadata.h pg_largeobject.h pg_aggregate.h \
+ pg_mv_statistic.h \
pg_statistic.h pg_rewrite.h pg_trigger.h pg_event_trigger.h pg_description.h \
pg_cast.h pg_enum.h pg_namespace.h pg_conversion.h pg_depend.h \
pg_database.h pg_db_role_setting.h pg_tablespace.h pg_pltemplate.h \
diff --git a/src/backend/catalog/aclchk.c b/src/backend/catalog/aclchk.c
index f4df6df..d0e5b2e 100644
--- a/src/backend/catalog/aclchk.c
+++ b/src/backend/catalog/aclchk.c
@@ -40,6 +40,7 @@
#include "catalog/pg_language.h"
#include "catalog/pg_largeobject.h"
#include "catalog/pg_largeobject_metadata.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_operator.h"
@@ -5133,6 +5134,32 @@ pg_subscription_ownercheck(Oid sub_oid, Oid roleid)
}
/*
+ * Ownership check for a multivariate statistics (specified by OID).
+ */
+bool
+pg_statistics_ownercheck(Oid stat_oid, Oid roleid)
+{
+ HeapTuple tuple;
+ Oid ownerId;
+
+ /* Superusers bypass all permission checking. */
+ if (superuser_arg(roleid))
+ return true;
+
+ tuple = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(stat_oid));
+ if (!HeapTupleIsValid(tuple))
+ ereport(ERROR,
+ (errcode(ERRCODE_UNDEFINED_OBJECT),
+ errmsg("statistics with OID %u do not exist", stat_oid)));
+
+ ownerId = ((Form_pg_mv_statistic) GETSTRUCT(tuple))->staowner;
+
+ ReleaseSysCache(tuple);
+
+ return has_privs_of_role(roleid, ownerId);
+}
+
+/*
* Check whether specified role has CREATEROLE privilege (or is a superuser)
*
* Note: roles do not have owners per se; instead we use this test in
diff --git a/src/backend/catalog/dependency.c b/src/backend/catalog/dependency.c
index 1c43af6..7083270 100644
--- a/src/backend/catalog/dependency.c
+++ b/src/backend/catalog/dependency.c
@@ -42,6 +42,7 @@
#include "catalog/pg_init_privs.h"
#include "catalog/pg_language.h"
#include "catalog/pg_largeobject.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_operator.h"
@@ -171,7 +172,8 @@ static const Oid object_classes[] = {
PublicationRelationId, /* OCLASS_PUBLICATION */
PublicationRelRelationId, /* OCLASS_PUBLICATION_REL */
SubscriptionRelationId, /* OCLASS_SUBSCRIPTION */
- TransformRelationId /* OCLASS_TRANSFORM */
+ TransformRelationId, /* OCLASS_TRANSFORM */
+ MvStatisticRelationId /* OCLASS_STATISTICS */
};
@@ -1263,6 +1265,10 @@ doDeletion(const ObjectAddress *object, int flags)
DropTransformById(object->objectId);
break;
+ case OCLASS_STATISTICS:
+ RemoveStatisticsById(object->objectId);
+ break;
+
default:
elog(ERROR, "unrecognized object class: %u",
object->classId);
@@ -2430,6 +2436,9 @@ getObjectClass(const ObjectAddress *object)
case TransformRelationId:
return OCLASS_TRANSFORM;
+
+ case MvStatisticRelationId:
+ return OCLASS_STATISTICS;
}
/* shouldn't get here */
diff --git a/src/backend/catalog/heap.c b/src/backend/catalog/heap.c
index 7ce9115..470f7ad 100644
--- a/src/backend/catalog/heap.c
+++ b/src/backend/catalog/heap.c
@@ -48,6 +48,7 @@
#include "catalog/pg_constraint_fn.h"
#include "catalog/pg_foreign_table.h"
#include "catalog/pg_inherits.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_partitioned_table.h"
@@ -1614,7 +1615,10 @@ RemoveAttributeById(Oid relid, AttrNumber attnum)
heap_close(attr_rel, RowExclusiveLock);
if (attnum > 0)
+ {
RemoveStatistics(relid, attnum);
+ RemoveMVStatistics(relid, attnum);
+ }
relation_close(rel, NoLock);
}
@@ -1865,6 +1869,11 @@ heap_drop_with_catalog(Oid relid)
RemoveStatistics(relid, 0);
/*
+ * delete multi-variate statistics
+ */
+ RemoveMVStatistics(relid, 0);
+
+ /*
* delete attribute tuples
*/
DeleteAttributeTuples(relid);
@@ -2783,6 +2792,98 @@ RemoveStatistics(Oid relid, AttrNumber attnum)
/*
+ * RemoveMVStatistics --- remove entries in pg_mv_statistic for a rel
+ *
+ * If attnum is zero, remove all entries for rel; else remove only the one(s)
+ * for that column.
+ */
+void
+RemoveMVStatistics(Oid relid, AttrNumber attnum)
+{
+ Relation pgmvstatistic;
+ TupleDesc tupdesc = NULL;
+ SysScanDesc scan;
+ ScanKeyData key;
+ HeapTuple tuple;
+
+ /*
+ * When dropping a column, we'll drop statistics with a single remaining
+ * (undropped column). To do that, we need the tuple descriptor.
+ *
+ * We already have the relation locked (as we're running ALTER TABLE ...
+ * DROP COLUMN), so we'll just get the descriptor here.
+ */
+ if (attnum != 0)
+ {
+ Relation rel = relation_open(relid, NoLock);
+
+ /* multivariate stats are supported on tables and matviews */
+ if (rel->rd_rel->relkind == RELKIND_RELATION ||
+ rel->rd_rel->relkind == RELKIND_MATVIEW)
+ tupdesc = RelationGetDescr(rel);
+
+ relation_close(rel, NoLock);
+ }
+
+ if (tupdesc == NULL)
+ return;
+
+ pgmvstatistic = heap_open(MvStatisticRelationId, RowExclusiveLock);
+
+ ScanKeyInit(&key,
+ Anum_pg_mv_statistic_starelid,
+ BTEqualStrategyNumber, F_OIDEQ,
+ ObjectIdGetDatum(relid));
+
+ scan = systable_beginscan(pgmvstatistic,
+ MvStatisticRelidIndexId,
+ true, NULL, 1, &key);
+
+ /* we must loop even when attnum != 0, in case of inherited stats */
+ while (HeapTupleIsValid(tuple = systable_getnext(scan)))
+ {
+ bool delete = true;
+
+ if (attnum != 0)
+ {
+ Datum adatum;
+ bool isnull;
+ int i;
+ int ncolumns = 0;
+ ArrayType *arr;
+ int16 *attnums;
+
+ /* get the columns */
+ adatum = SysCacheGetAttr(MVSTATOID, tuple,
+ Anum_pg_mv_statistic_stakeys, &isnull);
+ Assert(!isnull);
+
+ arr = DatumGetArrayTypeP(adatum);
+ attnums = (int16 *) ARR_DATA_PTR(arr);
+
+ for (i = 0; i < ARR_DIMS(arr)[0]; i++)
+ {
+ /* count the column unless it's has been / is being dropped */
+ if ((!tupdesc->attrs[attnums[i] - 1]->attisdropped) &&
+ (attnums[i] != attnum))
+ ncolumns += 1;
+ }
+
+ /* delete if there are less than two attributes */
+ delete = (ncolumns < 2);
+ }
+
+ if (delete)
+ simple_heap_delete(pgmvstatistic, &tuple->t_self);
+ }
+
+ systable_endscan(scan);
+
+ heap_close(pgmvstatistic, RowExclusiveLock);
+}
+
+
+/*
* RelationTruncateIndexes - truncate all indexes associated
* with the heap relation to zero tuples.
*
diff --git a/src/backend/catalog/namespace.c b/src/backend/catalog/namespace.c
index a38da30..6cd7d36 100644
--- a/src/backend/catalog/namespace.c
+++ b/src/backend/catalog/namespace.c
@@ -4236,3 +4236,59 @@ pg_is_other_temp_schema(PG_FUNCTION_ARGS)
PG_RETURN_BOOL(isOtherTempNamespace(oid));
}
+
+/*
+ * get_statistics_oid - find a statistics by possibly qualified name
+ *
+ * If not found, returns InvalidOid if missing_ok, else throws error
+ */
+Oid
+get_statistics_oid(List *names, bool missing_ok)
+{
+ char *schemaname;
+ char *stats_name;
+ Oid namespaceId;
+ Oid stats_oid = InvalidOid;
+ ListCell *l;
+
+ /* deconstruct the name list */
+ DeconstructQualifiedName(names, &schemaname, &stats_name);
+
+ if (schemaname)
+ {
+ /* use exact schema given */
+ namespaceId = LookupExplicitNamespace(schemaname, missing_ok);
+ if (missing_ok && !OidIsValid(namespaceId))
+ stats_oid = InvalidOid;
+ else
+ stats_oid = GetSysCacheOid2(MVSTATNAMENSP,
+ PointerGetDatum(stats_name),
+ ObjectIdGetDatum(namespaceId));
+ }
+ else
+ {
+ /* search for it in search path */
+ recomputeNamespacePath();
+
+ foreach(l, activeSearchPath)
+ {
+ namespaceId = lfirst_oid(l);
+
+ if (namespaceId == myTempNamespace)
+ continue; /* do not look in temp namespace */
+ stats_oid = GetSysCacheOid2(MVSTATNAMENSP,
+ PointerGetDatum(stats_name),
+ ObjectIdGetDatum(namespaceId));
+ if (OidIsValid(stats_oid))
+ break;
+ }
+ }
+
+ if (!OidIsValid(stats_oid) && !missing_ok)
+ ereport(ERROR,
+ (errcode(ERRCODE_UNDEFINED_OBJECT),
+ errmsg("statistics \"%s\" do not exist",
+ NameListToString(names))));
+
+ return stats_oid;
+}
diff --git a/src/backend/catalog/objectaddress.c b/src/backend/catalog/objectaddress.c
index 2a38792..d390dc1 100644
--- a/src/backend/catalog/objectaddress.c
+++ b/src/backend/catalog/objectaddress.c
@@ -39,6 +39,7 @@
#include "catalog/pg_language.h"
#include "catalog/pg_largeobject.h"
#include "catalog/pg_largeobject_metadata.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_opfamily.h"
@@ -478,6 +479,18 @@ static const ObjectPropertyType ObjectProperty[] =
InvalidAttrNumber,
-1,
true
+ },
+ {
+ MvStatisticRelationId,
+ MvStatisticOidIndexId,
+ MVSTATOID,
+ MVSTATNAMENSP,
+ Anum_pg_mv_statistic_staname,
+ Anum_pg_mv_statistic_stanamespace,
+ Anum_pg_mv_statistic_staowner,
+ InvalidAttrNumber, /* no ACL (same as relation) */
+ -1, /* no ACL */
+ true
}
};
@@ -696,6 +709,10 @@ static const struct object_type_map
/* OCLASS_TRANSFORM */
{
"transform", OBJECT_TRANSFORM
+ },
+ /* OBJECT_STATISTICS */
+ {
+ "statistics", OBJECT_STATISTICS
}
};
@@ -980,6 +997,11 @@ get_object_address(ObjectType objtype, List *objname, List *objargs,
address = get_object_address_defacl(objname, objargs,
missing_ok);
break;
+ case OBJECT_STATISTICS:
+ address.classId = MvStatisticRelationId;
+ address.objectId = get_statistics_oid(objname, missing_ok);
+ address.objectSubId = 0;
+ break;
default:
elog(ERROR, "unrecognized objtype: %d", (int) objtype);
/* placate compiler, in case it thinks elog might return */
@@ -2361,6 +2383,10 @@ check_object_ownership(Oid roleid, ObjectType objtype, ObjectAddress address,
(errcode(ERRCODE_INSUFFICIENT_PRIVILEGE),
errmsg("must be superuser")));
break;
+ case OBJECT_STATISTICS:
+ if (!pg_statistics_ownercheck(address.objectId, roleid))
+ aclcheck_error_type(ACLCHECK_NOT_OWNER, address.objectId);
+ break;
default:
elog(ERROR, "unrecognized object type: %d",
(int) objtype);
@@ -3848,6 +3874,10 @@ getObjectTypeDescription(const ObjectAddress *object)
appendStringInfoString(&buffer, "subscription");
break;
+ case OCLASS_STATISTICS:
+ appendStringInfoString(&buffer, "statistics");
+ break;
+
default:
appendStringInfo(&buffer, "unrecognized %u", object->classId);
break;
@@ -4871,6 +4901,29 @@ getObjectIdentityParts(const ObjectAddress *object,
break;
}
+ case OCLASS_STATISTICS:
+ {
+ HeapTuple tup;
+ Form_pg_mv_statistic formStatistic;
+ char *schema;
+
+ tup = SearchSysCache1(MVSTATOID,
+ ObjectIdGetDatum(object->objectId));
+ if (!HeapTupleIsValid(tup))
+ elog(ERROR, "cache lookup failed for statistics %u",
+ object->objectId);
+ formStatistic = (Form_pg_mv_statistic) GETSTRUCT(tup);
+ schema = get_namespace_name_or_temp(formStatistic->stanamespace);
+ appendStringInfoString(&buffer,
+ quote_qualified_identifier(schema,
+ NameStr(formStatistic->staname)));
+ if (objname)
+ *objname = list_make2(schema,
+ pstrdup(NameStr(formStatistic->staname)));
+ ReleaseSysCache(tup);
+ }
+ break;
+
default:
appendStringInfo(&buffer, "unrecognized object %u %u %d",
object->classId,
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 4dfedf8..00ab440 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -181,6 +181,16 @@ CREATE OR REPLACE VIEW pg_sequences AS
WHERE NOT pg_is_other_temp_schema(N.oid)
AND relkind = 'S';
+CREATE VIEW pg_mv_stats AS
+ SELECT
+ N.nspname AS schemaname,
+ C.relname AS tablename,
+ S.staname AS staname,
+ S.stakeys AS attnums,
+ length(s.standist) AS ndistbytes
+ FROM (pg_mv_statistic S JOIN pg_class C ON (C.oid = S.starelid))
+ LEFT JOIN pg_namespace N ON (N.oid = C.relnamespace);
+
CREATE VIEW pg_stats WITH (security_barrier) AS
SELECT
nspname AS schemaname,
diff --git a/src/backend/commands/Makefile b/src/backend/commands/Makefile
index e0fab38..4a6c99e 100644
--- a/src/backend/commands/Makefile
+++ b/src/backend/commands/Makefile
@@ -18,8 +18,8 @@ OBJS = amcmds.o aggregatecmds.o alter.o analyze.o async.o cluster.o comment.o \
event_trigger.o explain.o extension.o foreigncmds.o functioncmds.o \
indexcmds.o lockcmds.o matview.o operatorcmds.o opclasscmds.o \
policy.o portalcmds.o prepare.o proclang.o publicationcmds.o \
- schemacmds.o seclabel.o sequence.o subscriptioncmds.o tablecmds.o \
- tablespace.o trigger.o tsearchcmds.o typecmds.o user.o vacuum.o \
- vacuumlazy.o variable.o view.o
+ schemacmds.o seclabel.o sequence.o statscmds.o subscriptioncmds.o \
+ tablecmds.o tablespace.o trigger.o tsearchcmds.o typecmds.o user.o \
+ vacuum.o vacuumlazy.o variable.o view.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/commands/alter.c b/src/backend/commands/alter.c
index 768fcc8..e2d1243 100644
--- a/src/backend/commands/alter.c
+++ b/src/backend/commands/alter.c
@@ -361,6 +361,7 @@ ExecRenameStmt(RenameStmt *stmt)
case OBJECT_OPCLASS:
case OBJECT_OPFAMILY:
case OBJECT_LANGUAGE:
+ case OBJECT_STATISTICS:
case OBJECT_TSCONFIGURATION:
case OBJECT_TSDICTIONARY:
case OBJECT_TSPARSER:
@@ -475,6 +476,7 @@ ExecAlterObjectSchemaStmt(AlterObjectSchemaStmt *stmt,
case OBJECT_OPERATOR:
case OBJECT_OPCLASS:
case OBJECT_OPFAMILY:
+ case OBJECT_STATISTICS:
case OBJECT_TSCONFIGURATION:
case OBJECT_TSDICTIONARY:
case OBJECT_TSPARSER:
@@ -790,6 +792,7 @@ ExecAlterOwnerStmt(AlterOwnerStmt *stmt)
case OBJECT_OPERATOR:
case OBJECT_OPCLASS:
case OBJECT_OPFAMILY:
+ case OBJECT_STATISTICS:
case OBJECT_TABLESPACE:
case OBJECT_TSDICTIONARY:
case OBJECT_TSCONFIGURATION:
diff --git a/src/backend/commands/analyze.c b/src/backend/commands/analyze.c
index c9f6afe..27eaf79 100644
--- a/src/backend/commands/analyze.c
+++ b/src/backend/commands/analyze.c
@@ -17,6 +17,7 @@
#include <math.h>
#include "access/multixact.h"
+#include "access/sysattr.h"
#include "access/transam.h"
#include "access/tupconvert.h"
#include "access/tuptoaster.h"
@@ -27,6 +28,7 @@
#include "catalog/indexing.h"
#include "catalog/pg_collation.h"
#include "catalog/pg_inherits_fn.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "commands/dbcommands.h"
#include "commands/tablecmds.h"
@@ -45,10 +47,13 @@
#include "storage/procarray.h"
#include "utils/acl.h"
#include "utils/attoptcache.h"
+#include "utils/builtins.h"
#include "utils/datum.h"
+#include "utils/fmgroids.h"
#include "utils/guc.h"
#include "utils/lsyscache.h"
#include "utils/memutils.h"
+#include "utils/mvstats.h"
#include "utils/pg_rusage.h"
#include "utils/sampling.h"
#include "utils/sortsupport.h"
@@ -559,6 +564,9 @@ do_analyze_rel(Relation onerel, int options, VacuumParams *params,
update_attstats(RelationGetRelid(Irel[ind]), false,
thisdata->attr_cnt, thisdata->vacattrstats);
}
+
+ /* Build multivariate stats (if there are any). */
+ build_mv_stats(onerel, totalrows, numrows, rows, attr_cnt, vacattrstats);
}
/*
diff --git a/src/backend/commands/dropcmds.c b/src/backend/commands/dropcmds.c
index ff3108c..8fd4269 100644
--- a/src/backend/commands/dropcmds.c
+++ b/src/backend/commands/dropcmds.c
@@ -294,6 +294,10 @@ does_not_exist_skipping(ObjectType objtype, List *objname, List *objargs)
msg = gettext_noop("schema \"%s\" does not exist, skipping");
name = NameListToString(objname);
break;
+ case OBJECT_STATISTICS:
+ msg = gettext_noop("statistics \"%s\" do not exist, skipping");
+ name = NameListToString(objname);
+ break;
case OBJECT_TSPARSER:
if (!schema_does_not_exist_skipping(objname, &msg, &name))
{
diff --git a/src/backend/commands/event_trigger.c b/src/backend/commands/event_trigger.c
index 8125537..763a45a 100644
--- a/src/backend/commands/event_trigger.c
+++ b/src/backend/commands/event_trigger.c
@@ -112,6 +112,7 @@ static event_trigger_support_data event_trigger_support[] = {
{"SCHEMA", true},
{"SEQUENCE", true},
{"SERVER", true},
+ {"STATISTICS", true},
{"SUBSCRIPTION", true},
{"TABLE", true},
{"TABLESPACE", false},
@@ -1111,6 +1112,7 @@ EventTriggerSupportsObjectType(ObjectType obtype)
case OBJECT_SCHEMA:
case OBJECT_SEQUENCE:
case OBJECT_SUBSCRIPTION:
+ case OBJECT_STATISTICS:
case OBJECT_TABCONSTRAINT:
case OBJECT_TABLE:
case OBJECT_TRANSFORM:
@@ -1176,6 +1178,7 @@ EventTriggerSupportsObjectClass(ObjectClass objclass)
case OCLASS_PUBLICATION:
case OCLASS_PUBLICATION_REL:
case OCLASS_SUBSCRIPTION:
+ case OCLASS_STATISTICS:
return true;
}
diff --git a/src/backend/commands/statscmds.c b/src/backend/commands/statscmds.c
new file mode 100644
index 0000000..bde7e4b
--- /dev/null
+++ b/src/backend/commands/statscmds.c
@@ -0,0 +1,259 @@
+/*-------------------------------------------------------------------------
+ *
+ * statscmds.c
+ * Commands for creating and altering multivariate statistics
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/commands/statscmds.c
+ *
+ *-------------------------------------------------------------------------
+ */
+#include "postgres.h"
+
+#include "access/relscan.h"
+#include "catalog/dependency.h"
+#include "catalog/indexing.h"
+#include "catalog/namespace.h"
+#include "catalog/pg_mv_statistic.h"
+#include "catalog/pg_namespace.h"
+#include "commands/defrem.h"
+#include "miscadmin.h"
+#include "utils/builtins.h"
+#include "utils/inval.h"
+#include "utils/memutils.h"
+#include "utils/mvstats.h"
+#include "utils/rel.h"
+#include "utils/syscache.h"
+
+
+/* used for sorting the attnums in ExecCreateStatistics */
+static int
+compare_int16(const void *a, const void *b)
+{
+ return memcmp(a, b, sizeof(int16));
+}
+
+/*
+ * Implements the CREATE STATISTICS name ON (columns) FROM table
+ *
+ * We do require that the types support sorting (ltopr), although some
+ * statistics might work with equality only.
+ */
+ObjectAddress
+CreateStatistics(CreateStatsStmt *stmt)
+{
+ int i;
+ ListCell *l;
+ int16 attnums[MVSTATS_MAX_DIMENSIONS];
+ int numcols = 0;
+ ObjectAddress address = InvalidObjectAddress;
+ char *namestr;
+ NameData staname;
+ Oid statoid;
+ Oid namespaceId;
+
+ HeapTuple htup;
+ Datum values[Natts_pg_mv_statistic];
+ bool nulls[Natts_pg_mv_statistic];
+ int2vector *stakeys;
+ Relation mvstatrel;
+ Relation rel;
+ Oid relid;
+ ObjectAddress parentobject,
+ childobject;
+
+ Assert(IsA(stmt, CreateStatsStmt));
+
+ /* resolve the pieces of the name (namespace etc.) */
+ namespaceId = QualifiedNameGetCreationNamespace(stmt->defnames, &namestr);
+ namestrcpy(&staname, namestr);
+
+ /*
+ * If if_not_exists was given and the statistics already exists, bail out.
+ */
+ if (SearchSysCacheExists2(MVSTATNAMENSP,
+ PointerGetDatum(&staname),
+ ObjectIdGetDatum(namespaceId)))
+ {
+ if (stmt->if_not_exists)
+ {
+ ereport(NOTICE,
+ (errcode(ERRCODE_DUPLICATE_OBJECT),
+ errmsg("statistics \"%s\" already exist, skipping",
+ namestr)));
+ return InvalidObjectAddress;
+ }
+
+ ereport(ERROR,
+ (errcode(ERRCODE_DUPLICATE_OBJECT),
+ errmsg("statistics \"%s\" already exist", namestr)));
+ }
+
+ rel = heap_openrv(stmt->relation, AccessExclusiveLock);
+ relid = RelationGetRelid(rel);
+
+ /*
+ * Transform column names to array of attnums. While doing that, we
+ * also enforce the maximum number of keys.
+ */
+ foreach(l, stmt->keys)
+ {
+ char *attname = strVal(lfirst(l));
+ HeapTuple atttuple;
+
+ atttuple = SearchSysCacheAttName(relid, attname);
+
+ if (!HeapTupleIsValid(atttuple))
+ ereport(ERROR,
+ (errcode(ERRCODE_UNDEFINED_COLUMN),
+ errmsg("column \"%s\" referenced in statistics does not exist",
+ attname)));
+
+ /* more than MVSTATS_MAX_DIMENSIONS columns not allowed */
+ if (numcols >= MVSTATS_MAX_DIMENSIONS)
+ ereport(ERROR,
+ (errcode(ERRCODE_TOO_MANY_COLUMNS),
+ errmsg("cannot have more than %d keys in statistics",
+ MVSTATS_MAX_DIMENSIONS)));
+
+ attnums[numcols] = ((Form_pg_attribute) GETSTRUCT(atttuple))->attnum;
+ ReleaseSysCache(atttuple);
+ numcols++;
+ }
+
+ /*
+ * Check that at least two columns were specified in the statement.
+ * The upper bound was already checked in the loop above.
+ */
+ if (numcols < 2)
+ ereport(ERROR,
+ (errcode(ERRCODE_TOO_MANY_COLUMNS),
+ errmsg("statistics require at least 2 columns")));
+
+ /*
+ * Sort the attnums, which makes detecting duplicies somewhat
+ * easier, and it does not hurt (it does not affect the efficiency,
+ * unlike for indexes, for example).
+ */
+ qsort(attnums, numcols, sizeof(int16), compare_int16);
+
+ /*
+ * Look for duplicities in the list of columns. The attnums are sorted
+ * so just check consecutive elements.
+ */
+ for (i = 1; i < numcols; i++)
+ if (attnums[i] == attnums[i-1])
+ ereport(ERROR,
+ (errcode(ERRCODE_UNDEFINED_COLUMN),
+ errmsg("duplicate column name in statistics definition")));
+
+ stakeys = buildint2vector(attnums, numcols);
+
+ /*
+ * Everything seems fine, so let's build the pg_mv_statistic entry.
+ * At this point we obviously only have the keys and options.
+ */
+
+ memset(values, 0, sizeof(values));
+ memset(nulls, false, sizeof(nulls));
+
+ /* metadata */
+ values[Anum_pg_mv_statistic_starelid - 1] = ObjectIdGetDatum(relid);
+ values[Anum_pg_mv_statistic_staname - 1] = NameGetDatum(&staname);
+ values[Anum_pg_mv_statistic_stanamespace - 1] = ObjectIdGetDatum(namespaceId);
+ values[Anum_pg_mv_statistic_staowner - 1] = ObjectIdGetDatum(GetUserId());
+
+ values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(stakeys);
+
+ /* enabled statistics */
+ values[Anum_pg_mv_statistic_ndist_enabled - 1] = BoolGetDatum(true);
+
+ nulls[Anum_pg_mv_statistic_standist - 1] = true;
+
+ /* insert the tuple into pg_mv_statistic */
+ mvstatrel = heap_open(MvStatisticRelationId, RowExclusiveLock);
+
+ htup = heap_form_tuple(mvstatrel->rd_att, values, nulls);
+
+ simple_heap_insert(mvstatrel, htup);
+
+ CatalogUpdateIndexes(mvstatrel, htup);
+
+ statoid = HeapTupleGetOid(htup);
+
+ heap_freetuple(htup);
+
+ /*
+ * Add a dependency on a table, so that stats get dropped on DROP TABLE.
+ */
+ ObjectAddressSet(parentobject, RelationRelationId, relid);
+ ObjectAddressSet(childobject, MvStatisticRelationId, statoid);
+
+ recordDependencyOn(&childobject, &parentobject, DEPENDENCY_AUTO);
+
+ /*
+ * Also add dependency on the schema (to drop statistics on DROP SCHEMA).
+ * This is not handled automatically by DROP TABLE because statistics have
+ * their own schema.
+ */
+ ObjectAddressSet(parentobject, NamespaceRelationId, namespaceId);
+
+ recordDependencyOn(&childobject, &parentobject, DEPENDENCY_AUTO);
+
+ heap_close(mvstatrel, RowExclusiveLock);
+
+ relation_close(rel, NoLock);
+
+ /*
+ * Invalidate relcache so that others see the new statistics.
+ */
+ CacheInvalidateRelcache(rel);
+
+ ObjectAddressSet(address, MvStatisticRelationId, statoid);
+
+ return address;
+}
+
+
+/*
+ * Implements the DROP STATISTICS
+ *
+ * DROP STATISTICS stats_name
+ */
+void
+RemoveStatisticsById(Oid statsOid)
+{
+ Relation relation;
+ Oid relid;
+ Relation rel;
+ HeapTuple tup;
+ Form_pg_mv_statistic mvstat;
+
+ /*
+ * Delete the pg_proc tuple.
+ */
+ relation = heap_open(MvStatisticRelationId, RowExclusiveLock);
+
+ tup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(statsOid));
+
+ if (!HeapTupleIsValid(tup)) /* should not happen */
+ elog(ERROR, "cache lookup failed for statistics %u", statsOid);
+
+ mvstat = (Form_pg_mv_statistic) GETSTRUCT(tup);
+ relid = mvstat->starelid;
+
+ rel = heap_open(relid, AccessExclusiveLock);
+
+ simple_heap_delete(relation, &tup->t_self);
+
+ CacheInvalidateRelcache(rel);
+
+ ReleaseSysCache(tup);
+
+ heap_close(relation, RowExclusiveLock);
+ heap_close(rel, NoLock);
+}
diff --git a/src/backend/nodes/copyfuncs.c b/src/backend/nodes/copyfuncs.c
index 30d733e..dc42be0 100644
--- a/src/backend/nodes/copyfuncs.c
+++ b/src/backend/nodes/copyfuncs.c
@@ -4349,6 +4349,19 @@ _copyDropSubscriptionStmt(const DropSubscriptionStmt *from)
return newnode;
}
+static CreateStatsStmt *
+_copyCreateStatsStmt(const CreateStatsStmt *from)
+{
+ CreateStatsStmt *newnode = makeNode(CreateStatsStmt);
+
+ COPY_NODE_FIELD(defnames);
+ COPY_NODE_FIELD(relation);
+ COPY_NODE_FIELD(keys);
+ COPY_SCALAR_FIELD(if_not_exists);
+
+ return newnode;
+}
+
/* ****************************************************************
* pg_list.h copy functions
* ****************************************************************
@@ -5272,6 +5285,9 @@ copyObject(const void *from)
case T_CommonTableExpr:
retval = _copyCommonTableExpr(from);
break;
+ case T_CreateStatsStmt:
+ retval = _copyCreateStatsStmt(from);
+ break;
case T_FuncWithArgs:
retval = _copyFuncWithArgs(from);
break;
diff --git a/src/backend/nodes/outfuncs.c b/src/backend/nodes/outfuncs.c
index 1560ac3..57cc0b4 100644
--- a/src/backend/nodes/outfuncs.c
+++ b/src/backend/nodes/outfuncs.c
@@ -2193,6 +2193,21 @@ _outForeignKeyOptInfo(StringInfo str, const ForeignKeyOptInfo *node)
}
static void
+_outMVStatisticInfo(StringInfo str, const MVStatisticInfo *node)
+{
+ WRITE_NODE_TYPE("MVSTATISTICINFO");
+
+ /* NB: this isn't a complete set of fields */
+ WRITE_OID_FIELD(mvoid);
+
+ /* enabled statistics */
+ WRITE_BOOL_FIELD(ndist_enabled);
+
+ /* built/available statistics */
+ WRITE_BOOL_FIELD(ndist_built);
+}
+
+static void
_outEquivalenceClass(StringInfo str, const EquivalenceClass *node)
{
/*
@@ -3799,6 +3814,9 @@ outNode(StringInfo str, const void *obj)
case T_PlannerParamItem:
_outPlannerParamItem(str, obj);
break;
+ case T_MVStatisticInfo:
+ _outMVStatisticInfo(str, obj);
+ break;
case T_ExtensibleNode:
_outExtensibleNode(str, obj);
diff --git a/src/backend/optimizer/util/plancat.c b/src/backend/optimizer/util/plancat.c
index 7836e6b..fc9ad93 100644
--- a/src/backend/optimizer/util/plancat.c
+++ b/src/backend/optimizer/util/plancat.c
@@ -29,6 +29,7 @@
#include "catalog/heap.h"
#include "catalog/partition.h"
#include "catalog/pg_am.h"
+#include "catalog/pg_mv_statistic.h"
#include "foreign/fdwapi.h"
#include "miscadmin.h"
#include "nodes/makefuncs.h"
@@ -41,7 +42,9 @@
#include "parser/parsetree.h"
#include "rewrite/rewriteManip.h"
#include "storage/bufmgr.h"
+#include "utils/builtins.h"
#include "utils/lsyscache.h"
+#include "utils/syscache.h"
#include "utils/rel.h"
#include "utils/snapmgr.h"
@@ -63,7 +66,7 @@ static List *get_relation_constraints(PlannerInfo *root,
bool include_notnull);
static List *build_index_tlist(PlannerInfo *root, IndexOptInfo *index,
Relation heapRelation);
-
+static List *get_relation_statistics(RelOptInfo *rel, Relation relation);
/*
* get_relation_info -
@@ -397,6 +400,8 @@ get_relation_info(PlannerInfo *root, Oid relationObjectId, bool inhparent,
rel->indexlist = indexinfos;
+ rel->mvstatlist = get_relation_statistics(rel, relation);
+
/* Grab foreign-table info using the relcache, while we have it */
if (relation->rd_rel->relkind == RELKIND_FOREIGN_TABLE)
{
@@ -1250,6 +1255,71 @@ get_relation_constraints(PlannerInfo *root,
return result;
}
+/*
+ * get_relation_statistics
+ *
+ * Retrieve multivariate statistics defined on the table.
+ *
+ * Returns a List (possibly empty) of MVStatisticInfo objects describing
+ * the statistics. Only attributes needed for selecting statistics are
+ * retrieved (columns covered by the statistics, etc.).
+ */
+static List *
+get_relation_statistics(RelOptInfo *rel, Relation relation)
+{
+ List *mvstatoidlist;
+ ListCell *l;
+ List *stainfos = NIL;
+
+ mvstatoidlist = RelationGetMVStatList(relation);
+
+ foreach(l, mvstatoidlist)
+ {
+ ArrayType *arr;
+ Datum adatum;
+ bool isnull;
+ Oid mvoid = lfirst_oid(l);
+ Form_pg_mv_statistic mvstat;
+ MVStatisticInfo *info;
+
+ HeapTuple htup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(mvoid));
+
+ mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
+
+ /* unavailable stats are not interesting for the planner */
+ if (mvstat->ndist_built)
+ {
+ info = makeNode(MVStatisticInfo);
+
+ info->mvoid = mvoid;
+ info->rel = rel;
+
+ /* enabled statistics */
+ info->ndist_enabled = mvstat->ndist_enabled;
+
+ /* built/available statistics */
+ info->ndist_built = mvstat->ndist_built;
+
+ /* stakeys */
+ adatum = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_stakeys, &isnull);
+ Assert(!isnull);
+
+ arr = DatumGetArrayTypeP(adatum);
+
+ info->stakeys = buildint2vector((int16 *) ARR_DATA_PTR(arr),
+ ARR_DIMS(arr)[0]);
+
+ stainfos = lcons(info, stainfos);
+ }
+
+ ReleaseSysCache(htup);
+ }
+
+ list_free(mvstatoidlist);
+
+ return stainfos;
+}
/*
* relation_excluded_by_constraints
diff --git a/src/backend/parser/gram.y b/src/backend/parser/gram.y
index a4edea0..475a8a6 100644
--- a/src/backend/parser/gram.y
+++ b/src/backend/parser/gram.y
@@ -257,7 +257,7 @@ static Node *makeRecursiveViewSelect(char *relname, List *aliases, Node *query);
ConstraintsSetStmt CopyStmt CreateAsStmt CreateCastStmt
CreateDomainStmt CreateExtensionStmt CreateGroupStmt CreateOpClassStmt
CreateOpFamilyStmt AlterOpFamilyStmt CreatePLangStmt
- CreateSchemaStmt CreateSeqStmt CreateStmt CreateTableSpaceStmt
+ CreateSchemaStmt CreateSeqStmt CreateStmt CreateStatsStmt CreateTableSpaceStmt
CreateFdwStmt CreateForeignServerStmt CreateForeignTableStmt
CreateAssertStmt CreateTransformStmt CreateTrigStmt CreateEventTrigStmt
CreateUserStmt CreateUserMappingStmt CreateRoleStmt CreatePolicyStmt
@@ -867,6 +867,7 @@ stmt :
| CreateSeqStmt
| CreateStmt
| CreateSubscriptionStmt
+ | CreateStatsStmt
| CreateTableSpaceStmt
| CreateTransformStmt
| CreateTrigStmt
@@ -3747,6 +3748,34 @@ OptConsTableSpace: USING INDEX TABLESPACE name { $$ = $4; }
ExistingIndex: USING INDEX index_name { $$ = $3; }
;
+/*****************************************************************************
+ *
+ * QUERY :
+ * CREATE STATISTICS stats_name ON relname (columns) WITH (options)
+ *
+ *****************************************************************************/
+
+
+CreateStatsStmt: CREATE STATISTICS any_name ON '(' columnList ')' FROM qualified_name
+ {
+ CreateStatsStmt *n = makeNode(CreateStatsStmt);
+ n->defnames = $3;
+ n->relation = $9;
+ n->keys = $6;
+ n->if_not_exists = false;
+ $$ = (Node *)n;
+ }
+ | CREATE STATISTICS IF_P NOT EXISTS any_name ON '(' columnList ')' FROM qualified_name
+ {
+ CreateStatsStmt *n = makeNode(CreateStatsStmt);
+ n->defnames = $6;
+ n->relation = $12;
+ n->keys = $9;
+ n->if_not_exists = true;
+ $$ = (Node *)n;
+ }
+ ;
+
/*****************************************************************************
*
@@ -6090,6 +6119,7 @@ drop_type: TABLE { $$ = OBJECT_TABLE; }
| TEXT_P SEARCH TEMPLATE { $$ = OBJECT_TSTEMPLATE; }
| TEXT_P SEARCH CONFIGURATION { $$ = OBJECT_TSCONFIGURATION; }
| PUBLICATION { $$ = OBJECT_PUBLICATION; }
+ | STATISTICS { $$ = OBJECT_STATISTICS; }
;
any_name_list:
@@ -8475,6 +8505,15 @@ RenameStmt: ALTER AGGREGATE aggregate_with_argtypes RENAME TO name
n->missing_ok = false;
$$ = (Node *)n;
}
+ | ALTER STATISTICS any_name RENAME TO name
+ {
+ RenameStmt *n = makeNode(RenameStmt);
+ n->renameType = OBJECT_STATISTICS;
+ n->object = $3;
+ n->newname = $6;
+ n->missing_ok = false;
+ $$ = (Node *)n;
+ }
;
opt_column: COLUMN { $$ = COLUMN; }
@@ -8760,6 +8799,15 @@ AlterObjectSchemaStmt:
n->missing_ok = false;
$$ = (Node *)n;
}
+ | ALTER STATISTICS any_name SET SCHEMA name
+ {
+ AlterObjectSchemaStmt *n = makeNode(AlterObjectSchemaStmt);
+ n->objectType = OBJECT_STATISTICS;
+ n->object = $3;
+ n->newschema = $6;
+ n->missing_ok = false;
+ $$ = (Node *)n;
+ }
;
/*****************************************************************************
@@ -8966,6 +9014,14 @@ AlterOwnerStmt: ALTER AGGREGATE aggregate_with_argtypes OWNER TO RoleSpec
n->newowner = $6;
$$ = (Node *)n;
}
+ | ALTER STATISTICS name OWNER TO RoleSpec
+ {
+ AlterOwnerStmt *n = makeNode(AlterOwnerStmt);
+ n->objectType = OBJECT_STATISTICS;
+ n->object = list_make1(makeString($3));
+ n->newowner = $6;
+ $$ = (Node *)n;
+ }
;
diff --git a/src/backend/tcop/utility.c b/src/backend/tcop/utility.c
index 5d3be38..6f2371c 100644
--- a/src/backend/tcop/utility.c
+++ b/src/backend/tcop/utility.c
@@ -1621,6 +1621,10 @@ ProcessUtilitySlow(ParseState *pstate,
commandCollected = true;
break;
+ case T_CreateStatsStmt: /* CREATE STATISTICS */
+ address = CreateStatistics((CreateStatsStmt *) parsetree);
+ break;
+
default:
elog(ERROR, "unrecognized node type: %d",
(int) nodeTag(parsetree));
@@ -1986,6 +1990,8 @@ AlterObjectTypeCommandTag(ObjectType objtype)
break;
case OBJECT_SUBSCRIPTION:
tag = "ALTER SUBSCRIPTION";
+ case OBJECT_STATISTICS:
+ tag = "ALTER STATISTICS";
break;
default:
tag = "???";
@@ -2280,6 +2286,8 @@ CreateCommandTag(Node *parsetree)
break;
case OBJECT_PUBLICATION:
tag = "DROP PUBLICATION";
+ case OBJECT_STATISTICS:
+ tag = "DROP STATISTICS";
break;
default:
tag = "???";
@@ -2679,6 +2687,10 @@ CreateCommandTag(Node *parsetree)
tag = "EXECUTE";
break;
+ case T_CreateStatsStmt:
+ tag = "CREATE STATISTICS";
+ break;
+
case T_DeallocateStmt:
{
DeallocateStmt *stmt = (DeallocateStmt *) parsetree;
diff --git a/src/backend/utils/Makefile b/src/backend/utils/Makefile
index 2e35ca5..0eb2331 100644
--- a/src/backend/utils/Makefile
+++ b/src/backend/utils/Makefile
@@ -9,7 +9,7 @@ top_builddir = ../../..
include $(top_builddir)/src/Makefile.global
OBJS = fmgrtab.o
-SUBDIRS = adt cache error fmgr hash init mb misc mmgr resowner sort time
+SUBDIRS = adt cache error fmgr hash init mb misc mmgr mvstats resowner sort time
# location of Catalog.pm
catalogdir = $(top_srcdir)/src/backend/catalog
diff --git a/src/backend/utils/adt/selfuncs.c b/src/backend/utils/adt/selfuncs.c
index fa32e9e..7774c44 100644
--- a/src/backend/utils/adt/selfuncs.c
+++ b/src/backend/utils/adt/selfuncs.c
@@ -132,6 +132,7 @@
#include "utils/fmgroids.h"
#include "utils/index_selfuncs.h"
#include "utils/lsyscache.h"
+#include "utils/mvstats.h"
#include "utils/nabstime.h"
#include "utils/pg_locale.h"
#include "utils/rel.h"
@@ -207,6 +208,8 @@ static Const *string_to_const(const char *str, Oid datatype);
static Const *string_to_bytea_const(const char *str, size_t str_len);
static List *add_predicate_to_quals(IndexOptInfo *index, List *indexQuals);
+static double find_ndistinct(PlannerInfo *root, RelOptInfo *rel, List *varinfos,
+ bool *found);
/*
* eqsel - Selectivity of "=" for any data types.
@@ -3436,12 +3439,26 @@ estimate_num_groups(PlannerInfo *root, List *groupExprs, double input_rows,
* don't know by how much. We should never clamp to less than the
* largest ndistinct value for any of the Vars, though, since
* there will surely be at least that many groups.
+ *
+ * However we don't need to do this if we have ndistinct stats on
+ * the columns - in that case we can simply use the coefficient
+ * to get the (probably way more accurate) estimate.
+ *
+ * XXX Might benefit from some refactoring, mixing the ndistinct
+ * coefficients and clamp seems a bit unfortunate.
*/
double clamp = rel->tuples;
if (relvarcount > 1)
{
- clamp *= 0.1;
+ bool found;
+ double ndist = find_ndistinct(root, rel, varinfos, &found);
+
+ if (found)
+ reldistinct = ndist;
+ else
+ clamp *= 0.1;
+
if (clamp < relmaxndistinct)
{
clamp = relmaxndistinct;
@@ -3450,6 +3467,7 @@ estimate_num_groups(PlannerInfo *root, List *groupExprs, double input_rows,
clamp = rel->tuples;
}
}
+
if (reldistinct > clamp)
reldistinct = clamp;
@@ -7600,3 +7618,151 @@ brincostestimate(PlannerInfo *root, IndexPath *path, double loop_count,
/* XXX what about pages_per_range? */
}
+
+/*
+ * Find applicable ndistinct statistics and compute the coefficient to
+ * correct the estimate (simply a product of per-column ndistincts).
+ *
+ * XXX Currently we only look for a perfect match, i.e. a single ndistinct
+ * estimate exactly matching all the columns of the statistics. This may be
+ * a bit problematic as adding a column (not covered by the ndistinct stats)
+ * will prevent us from using the stats entirely. So instead this needs to
+ * estimate the covered attributes, and then combine that with the extra
+ * attributes somehow (probably the old way).
+ */
+static double
+find_ndistinct(PlannerInfo *root, RelOptInfo *rel, List *varinfos, bool *found)
+{
+ ListCell *lc;
+ Bitmapset *attnums = NULL;
+ VariableStatData vardata;
+
+ /* assume we haven't found any suitable ndistinct statistics */
+ *found = false;
+
+ /* bail out immediately if the table has no multivariate statistics */
+ if (!rel->mvstatlist)
+ return 0.0;
+
+ foreach(lc, varinfos)
+ {
+ GroupVarInfo *varinfo = (GroupVarInfo *) lfirst(lc);
+
+ if (varinfo->rel != rel)
+ continue;
+
+ /* FIXME handle expressions in general only */
+
+ /*
+ * examine the variable (or expression) so that we know which
+ * attribute we're dealing with - we need this for matching the
+ * ndistinct coefficient
+ *
+ * FIXME probably might remember this from estimate_num_groups
+ */
+ examine_variable(root, varinfo->var, 0, &vardata);
+
+ if (HeapTupleIsValid(vardata.statsTuple))
+ {
+ Form_pg_statistic stats
+ = (Form_pg_statistic) GETSTRUCT(vardata.statsTuple);
+
+ attnums = bms_add_member(attnums, stats->staattnum);
+
+ ReleaseVariableStats(vardata);
+ }
+ }
+
+ /* look for a matching ndistinct statistics */
+ foreach (lc, rel->mvstatlist)
+ {
+ int i, k;
+ bool matches;
+ MVStatisticInfo *info = (MVStatisticInfo *)lfirst(lc);
+
+ /* skip statistics without ndistinct coefficient built */
+ if (!info->ndist_built)
+ continue;
+
+ /*
+ * Only ndistinct stats covering all Vars are acceptable, which can't
+ * happen if the statistics has fewer attributes than we have Vars.
+ */
+ if (bms_num_members(attnums) > info->stakeys->dim1)
+ continue;
+
+ /* check that all Vars are covered by the statistic */
+ matches = true; /* assume match until we find unmatched attribute */
+ k = -1;
+ while ((k = bms_next_member(attnums, k)) >= 0)
+ {
+ bool attr_found = false;
+ for (i = 0; i < info->stakeys->dim1; i++)
+ {
+ if (info->stakeys->values[i] == k)
+ {
+ attr_found = true;
+ break;
+ }
+ }
+
+ /* found attribute not covered by this ndistinct stats, skip */
+ if (!attr_found)
+ {
+ matches = false;
+ break;
+ }
+ }
+
+ if (! matches)
+ continue;
+
+ /* hey, this statistics matches! great, let's extract the value */
+ *found = true;
+
+ {
+ int j;
+ MVNDistinct stat = load_mv_ndistinct(info->mvoid);
+
+ for (j = 0; j < stat->nitems; j++)
+ {
+ bool item_matches = true;
+ MVNDistinctItem * item = &stat->items[j];
+
+ /* not the right item (different number of attributes) */
+ if (item->nattrs != bms_num_members(attnums))
+ continue;
+
+ /* check the attribute numbers */
+ k = -1;
+ while ((k = bms_next_member(attnums, k)) >= 0)
+ {
+ bool attr_found = false;
+ for (i = 0; i < item->nattrs; i++)
+ {
+ if (info->stakeys->values[item->attrs[i]] == k)
+ {
+ attr_found = true;
+ break;
+ }
+ }
+
+ if (! attr_found)
+ {
+ item_matches = false;
+ break;
+ }
+ }
+
+ if (! item_matches)
+ continue;
+
+ return item->ndistinct;
+ }
+ }
+ }
+
+ Assert(!(*found));
+
+ return 0.0;
+}
diff --git a/src/backend/utils/cache/relcache.c b/src/backend/utils/cache/relcache.c
index 26ff7e1..1316104 100644
--- a/src/backend/utils/cache/relcache.c
+++ b/src/backend/utils/cache/relcache.c
@@ -49,6 +49,7 @@
#include "catalog/pg_auth_members.h"
#include "catalog/pg_constraint.h"
#include "catalog/pg_database.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_partitioned_table.h"
@@ -4452,6 +4453,81 @@ RelationGetIndexList(Relation relation)
}
/*
+ * RelationGetMVStatList -- get a list of OIDs of statistics on this relation
+ *
+ * The statistics list is created only if someone requests it, in a way
+ * similar to RelationGetIndexList(). We scan pg_mv_statistic to find
+ * relevant statistics, and add the list to the relcache entry so that we
+ * won't have to compute it again. Note that shared cache inval of a
+ * relcache entry will delete the old list and set rd_mvstatvalid to 0,
+ * so that we must recompute the statistics list on next request. This
+ * handles creation or deletion of a statistic.
+ *
+ * The returned list is guaranteed to be sorted in order by OID, although
+ * this is not currently needed.
+ *
+ * Since shared cache inval causes the relcache's copy of the list to go away,
+ * we return a copy of the list palloc'd in the caller's context. The caller
+ * may list_free() the returned list after scanning it. This is necessary
+ * since the caller will typically be doing syscache lookups on the relevant
+ * statistics, and syscache lookup could cause SI messages to be processed!
+ */
+List *
+RelationGetMVStatList(Relation relation)
+{
+ Relation indrel;
+ SysScanDesc indscan;
+ ScanKeyData skey;
+ HeapTuple htup;
+ List *result;
+ List *oldlist;
+ MemoryContext oldcxt;
+
+ /* Quick exit if we already computed the list. */
+ if (relation->rd_mvstatvalid != 0)
+ return list_copy(relation->rd_mvstatlist);
+
+ /*
+ * We build the list we intend to return (in the caller's context) while
+ * doing the scan. After successfully completing the scan, we copy that
+ * list into the relcache entry. This avoids cache-context memory leakage
+ * if we get some sort of error partway through.
+ */
+ result = NIL;
+
+ /* Prepare to scan pg_index for entries having indrelid = this rel. */
+ ScanKeyInit(&skey,
+ Anum_pg_mv_statistic_starelid,
+ BTEqualStrategyNumber, F_OIDEQ,
+ ObjectIdGetDatum(RelationGetRelid(relation)));
+
+ indrel = heap_open(MvStatisticRelationId, AccessShareLock);
+ indscan = systable_beginscan(indrel, MvStatisticRelidIndexId, true,
+ NULL, 1, &skey);
+
+ while (HeapTupleIsValid(htup = systable_getnext(indscan)))
+ /* TODO maybe include only already built statistics? */
+ result = insert_ordered_oid(result, HeapTupleGetOid(htup));
+
+ systable_endscan(indscan);
+
+ heap_close(indrel, AccessShareLock);
+
+ /* Now save a copy of the completed list in the relcache entry. */
+ oldcxt = MemoryContextSwitchTo(CacheMemoryContext);
+ oldlist = relation->rd_mvstatlist;
+ relation->rd_mvstatlist = list_copy(result);
+
+ relation->rd_mvstatvalid = true;
+ MemoryContextSwitchTo(oldcxt);
+
+ /* Don't leak the old list, if there is one */
+ list_free(oldlist);
+
+ return result;
+}
+
+/*
* insert_ordered_oid
* Insert a new Oid into a sorted list of Oids, preserving ordering
*
@@ -5531,6 +5607,8 @@ load_relcache_init_file(bool shared)
rel->rd_pkattr = NULL;
rel->rd_idattr = NULL;
rel->rd_pubactions = NULL;
+ rel->rd_mvstatvalid = false;
+ rel->rd_mvstatlist = NIL;
rel->rd_createSubid = InvalidSubTransactionId;
rel->rd_newRelfilenodeSubid = InvalidSubTransactionId;
rel->rd_amcache = NULL;
diff --git a/src/backend/utils/cache/syscache.c b/src/backend/utils/cache/syscache.c
index bdfaa0c..fbd1885 100644
--- a/src/backend/utils/cache/syscache.c
+++ b/src/backend/utils/cache/syscache.c
@@ -44,6 +44,7 @@
#include "catalog/pg_foreign_server.h"
#include "catalog/pg_foreign_table.h"
#include "catalog/pg_language.h"
+#include "catalog/pg_mv_statistic.h"
#include "catalog/pg_namespace.h"
#include "catalog/pg_opclass.h"
#include "catalog/pg_operator.h"
@@ -507,6 +508,28 @@ static const struct cachedesc cacheinfo[] = {
},
4
},
+ {MvStatisticRelationId, /* MVSTATNAMENSP */
+ MvStatisticNameIndexId,
+ 2,
+ {
+ Anum_pg_mv_statistic_staname,
+ Anum_pg_mv_statistic_stanamespace,
+ 0,
+ 0
+ },
+ 4
+ },
+ {MvStatisticRelationId, /* MVSTATOID */
+ MvStatisticOidIndexId,
+ 1,
+ {
+ ObjectIdAttributeNumber,
+ 0,
+ 0,
+ 0
+ },
+ 4
+ },
{NamespaceRelationId, /* NAMESPACENAME */
NamespaceNameIndexId,
1,
diff --git a/src/backend/utils/mvstats/Makefile b/src/backend/utils/mvstats/Makefile
new file mode 100644
index 0000000..7295d46
--- /dev/null
+++ b/src/backend/utils/mvstats/Makefile
@@ -0,0 +1,17 @@
+#-------------------------------------------------------------------------
+#
+# Makefile--
+# Makefile for utils/mvstats
+#
+# IDENTIFICATION
+# src/backend/utils/mvstats/Makefile
+#
+#-------------------------------------------------------------------------
+
+subdir = src/backend/utils/mvstats
+top_builddir = ../../../..
+include $(top_builddir)/src/Makefile.global
+
+OBJS = common.o mvdist.o
+
+include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/utils/mvstats/README.ndistinct b/src/backend/utils/mvstats/README.ndistinct
new file mode 100644
index 0000000..9365b17
--- /dev/null
+++ b/src/backend/utils/mvstats/README.ndistinct
@@ -0,0 +1,22 @@
+ndistinct coefficients
+======================
+
+Estimating number of groups in a combination of columns (e.g. for GROUP BY)
+is tricky, and the estimation error is often significant.
+
+The ndistinct coefficients address this by storing ndistinct estimates not
+only for individual columns, but also for (all) combinations of columns.
+So for example given three columns (a,b,c) the statistics will estimate
+ndistinct for (a,b), (a,c), (b,c) and (a,b,c). The per-column estimates
+are already available in pg_statistic.
+
+
+GROUP BY estimation (estimate_num_groups)
+-----------------------------------------
+
+Although ndistinct coefficient might be used for selectivity estimation
+(of equality conditions in WHERE clause), that is not implemented at this
+point.
+
+Instead, ndistinct coefficients are only used in estimate_num_groups() to
+estimate grouped queries.
diff --git a/src/backend/utils/mvstats/README.stats b/src/backend/utils/mvstats/README.stats
new file mode 100644
index 0000000..30d60d6
--- /dev/null
+++ b/src/backend/utils/mvstats/README.stats
@@ -0,0 +1,98 @@
+Multivariate statistics
+=======================
+
+When estimating various quantities (e.g. condition selectivities) the default
+approach relies on the assumption of independence. In practice that's often
+not true, resulting in estimation errors.
+
+Multivariate stats track different types of dependencies between the columns,
+hopefully improving the estimates.
+
+
+Types of statistics
+-------------------
+
+Currently we only have two kinds of multivariate statistics
+
+ (a) soft functional dependencies (README.dependencies)
+
+ (b) ndistinct coefficients
+
+
+Compatible clause types
+-----------------------
+
+Each type of statistics may be used to estimate some subset of clause types.
+
+ (a) functional dependencies - equality clauses (AND), possibly IS NULL
+
+Currently only simple operator clauses (Var op Const) are supported, but it's
+possible to support more complex clause types, e.g. (Var op Var).
+
+
+Complex clauses
+---------------
+
+We also support estimating more complex clauses - essentially AND/OR clauses
+with (Var op Const) as leaves, as long as all the referenced attributes are
+covered by a single statistics.
+
+For example this condition
+
+ (a=1) AND ((b=2) OR ((c=3) AND (d=4)))
+
+may be estimated using statistics on (a,b,c,d). If we only have statistics on
+(b,c,d) we may estimate the second part, and estimate (a=1) using simple stats.
+
+If we only have statistics on (a,b,c) we can't apply it at all at this point,
+but it's worth pointing out clauselist_selectivity() works recursively and when
+handling the second part (the OR-clause), we'll be able to apply the statistics.
+
+Note: The multi-statistics estimation patch also makes it possible to pass some
+clauses as 'conditions' into the deeper parts of the expression tree.
+
+
+Selectivity estimation
+----------------------
+
+When estimating selectivity, we aim to achieve several things:
+
+ (a) maximize the estimate accuracy
+
+ (b) minimize the overhead, especially when no suitable multivariate stats
+ exist (so if you are not using multivariate stats, there's no overhead)
+
+This clauselist_selectivity() performs several inexpensive checks first, before
+even attempting to do the more expensive estimation.
+
+ (1) check if there are multivariate stats on the relation
+
+ (2) check there are at least two attributes referenced by clauses compatible
+ with multivariate statistics (equality clauses for func. dependencies)
+
+ (3) perform reduction of equality clauses using func. dependencies
+
+ (4) estimate the reduced list of clauses using regular statistics
+
+Whenever we find there are no suitable stats, we skip the expensive steps.
+
+
+Size of sample in ANALYZE
+-------------------------
+When performing ANALYZE, the number of rows to sample is determined as
+
+ (300 * statistics_target)
+
+That works reasonably well for statistics on individual columns, but perhaps
+it's not enough for multivariate statistics. Papers analyzing estimation errors
+all use samples proportional to the table (usually finding that 1-3% of the
+table is enough to build accurate stats).
+
+The requested accuracy (number of MCV items or histogram bins) should also
+be considered when determining the sample size, and in multivariate statistics
+those are not necessarily limited by statistics_target.
+
+This however merits further discussion, because collecting the sample is quite
+expensive and increasing it further would make ANALYZE even more painful.
+Judging by the experiments with the current implementation, the fixed size
+seems to work reasonably well for now, so we leave this as a future work.
diff --git a/src/backend/utils/mvstats/common.c b/src/backend/utils/mvstats/common.c
new file mode 100644
index 0000000..7d2f3f3
--- /dev/null
+++ b/src/backend/utils/mvstats/common.c
@@ -0,0 +1,391 @@
+/*-------------------------------------------------------------------------
+ *
+ * common.c
+ * POSTGRES multivariate statistics
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/common.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "common.h"
+
+static VacAttrStats **lookup_var_attr_stats(int2vector *attrs,
+ int natts, VacAttrStats **vacattrstats);
+
+static List *list_mv_stats(Oid relid);
+
+static void update_mv_stats(Oid relid, MVNDistinct ndistinct,
+ int2vector *attrs, VacAttrStats **stats);
+
+
+/*
+ * Compute requested multivariate stats, using the rows sampled for the
+ * plain (single-column) stats.
+ *
+ * This fetches a list of stats from pg_mv_statistic, computes the stats
+ * and serializes them back into the catalog (as bytea values).
+ */
+void
+build_mv_stats(Relation onerel, double totalrows,
+ int numrows, HeapTuple *rows,
+ int natts, VacAttrStats **vacattrstats)
+{
+ ListCell *lc;
+ List *mvstats;
+
+ TupleDesc tupdesc = RelationGetDescr(onerel);
+
+ /*
+ * Fetch defined MV groups from pg_mv_statistic, and then compute the MV
+ * statistics (histograms for now).
+ */
+ mvstats = list_mv_stats(RelationGetRelid(onerel));
+
+ foreach(lc, mvstats)
+ {
+ int j;
+ MVStatisticInfo *stat = (MVStatisticInfo *) lfirst(lc);
+ MVNDistinct ndistinct = NULL;
+
+ VacAttrStats **stats = NULL;
+ int numatts = 0;
+
+ /* int2 vector of attnums the stats should be computed on */
+ int2vector *attrs = stat->stakeys;
+
+ /* see how many of the columns are not dropped */
+ for (j = 0; j < attrs->dim1; j++)
+ if (!tupdesc->attrs[attrs->values[j] - 1]->attisdropped)
+ numatts += 1;
+
+ /* if there are dropped attributes, build a filtered int2vector */
+ if (numatts != attrs->dim1)
+ {
+ int16 *tmp = palloc0(numatts * sizeof(int16));
+ int attnum = 0;
+
+ for (j = 0; j < attrs->dim1; j++)
+ if (!tupdesc->attrs[attrs->values[j] - 1]->attisdropped)
+ tmp[attnum++] = attrs->values[j];
+
+ pfree(attrs);
+ attrs = buildint2vector(tmp, numatts);
+ }
+
+ /* filter only the interesting vacattrstats records */
+ stats = lookup_var_attr_stats(attrs, natts, vacattrstats);
+
+ /* check allowed number of dimensions */
+ Assert((attrs->dim1 >= 2) && (attrs->dim1 <= MVSTATS_MAX_DIMENSIONS));
+
+ /* compute ndistinct coefficients */
+ if (stat->ndist_enabled)
+ ndistinct = build_mv_ndistinct(totalrows, numrows, rows, attrs, stats);
+
+ /* store the statistics in the catalog */
+ update_mv_stats(stat->mvoid, ndistinct, attrs, stats);
+ }
+}
+
+/*
+ * Lookup the VacAttrStats info for the selected columns, with indexes
+ * matching the attrs vector (to make it easy to work with when
+ * computing multivariate stats).
+ */
+static VacAttrStats **
+lookup_var_attr_stats(int2vector *attrs, int natts, VacAttrStats **vacattrstats)
+{
+ int i,
+ j;
+ int numattrs = attrs->dim1;
+ VacAttrStats **stats = (VacAttrStats **) palloc0(numattrs * sizeof(VacAttrStats *));
+
+ /* lookup VacAttrStats info for the requested columns (same attnum) */
+ for (i = 0; i < numattrs; i++)
+ {
+ stats[i] = NULL;
+ for (j = 0; j < natts; j++)
+ {
+ if (attrs->values[i] == vacattrstats[j]->tupattnum)
+ {
+ stats[i] = vacattrstats[j];
+ break;
+ }
+ }
+
+ /*
+ * Check that we found the info, that the attnum matches and that
+ * there's the requested 'lt' operator and that the type is
+ * 'passed-by-value'.
+ */
+ Assert(stats[i] != NULL);
+ Assert(stats[i]->tupattnum == attrs->values[i]);
+
+ /*
+ * FIXME This is rather ugly way to check for 'ltopr' (which is
+ * defined for 'scalar' attributes).
+ */
+ Assert(((StdAnalyzeData *) stats[i]->extra_data)->ltopr != InvalidOid);
+ }
+
+ return stats;
+}
+
+/*
+ * Fetch list of MV stats defined on a table, without the actual data
+ * for histograms, MCV lists etc.
+ */
+static List *
+list_mv_stats(Oid relid)
+{
+ Relation indrel;
+ SysScanDesc indscan;
+ ScanKeyData skey;
+ HeapTuple htup;
+ List *result = NIL;
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ ScanKeyInit(&skey,
+ Anum_pg_mv_statistic_starelid,
+ BTEqualStrategyNumber, F_OIDEQ,
+ ObjectIdGetDatum(relid));
+
+ indrel = heap_open(MvStatisticRelationId, AccessShareLock);
+ indscan = systable_beginscan(indrel, MvStatisticRelidIndexId, true,
+ NULL, 1, &skey);
+
+ while (HeapTupleIsValid(htup = systable_getnext(indscan)))
+ {
+ MVStatisticInfo *info = makeNode(MVStatisticInfo);
+ Form_pg_mv_statistic stats = (Form_pg_mv_statistic) GETSTRUCT(htup);
+
+ info->mvoid = HeapTupleGetOid(htup);
+ info->stakeys = buildint2vector(stats->stakeys.values, stats->stakeys.dim1);
+ info->ndist_enabled = stats->ndist_enabled;
+ info->ndist_built = stats->ndist_built;
+
+ result = lappend(result, info);
+ }
+
+ systable_endscan(indscan);
+
+ heap_close(indrel, AccessShareLock);
+
+ /*
+ * TODO maybe save the list into relcache, as in RelationGetIndexList
+ * (which was used as an inspiration of this one)?.
+ */
+
+ return result;
+}
+
+/*
+ * update_mv_stats
+ * Serializes the statistics and stores them into the pg_mv_statistic tuple.
+ */
+static void
+update_mv_stats(Oid mvoid, MVNDistinct ndistinct,
+ int2vector *attrs, VacAttrStats **stats)
+{
+ HeapTuple stup,
+ oldtup;
+ Datum values[Natts_pg_mv_statistic];
+ bool nulls[Natts_pg_mv_statistic];
+ bool replaces[Natts_pg_mv_statistic];
+
+ Relation sd = heap_open(MvStatisticRelationId, RowExclusiveLock);
+
+ memset(nulls, 1, Natts_pg_mv_statistic * sizeof(bool));
+ memset(replaces, 0, Natts_pg_mv_statistic * sizeof(bool));
+ memset(values, 0, Natts_pg_mv_statistic * sizeof(Datum));
+
+ /*
+ * Construct a new pg_mv_statistic tuple - replace only the histogram and
+ * MCV list, depending whether it actually was computed.
+ */
+ if (ndistinct != NULL)
+ {
+ bytea *data = serialize_mv_ndistinct(ndistinct);
+
+ nulls[Anum_pg_mv_statistic_standist -1] = (data == NULL);
+ values[Anum_pg_mv_statistic_standist-1] = PointerGetDatum(data);
+ }
+
+ /* always replace the value (either by bytea or NULL) */
+ replaces[Anum_pg_mv_statistic_standist - 1] = true;
+
+ /* always change the availability flags */
+ nulls[Anum_pg_mv_statistic_ndist_built - 1] = false;
+ nulls[Anum_pg_mv_statistic_stakeys - 1] = false;
+
+ /* use the new attnums, in case we removed some dropped ones */
+ replaces[Anum_pg_mv_statistic_ndist_built - 1] = true;
+ replaces[Anum_pg_mv_statistic_stakeys - 1] = true;
+
+ values[Anum_pg_mv_statistic_ndist_built - 1] = BoolGetDatum(ndistinct != NULL);
+
+ values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(attrs);
+
+ /* Is there already a pg_mv_statistic tuple for this attribute? */
+ oldtup = SearchSysCache1(MVSTATOID,
+ ObjectIdGetDatum(mvoid));
+
+ if (HeapTupleIsValid(oldtup))
+ {
+ /* Yes, replace it */
+ stup = heap_modify_tuple(oldtup,
+ RelationGetDescr(sd),
+ values,
+ nulls,
+ replaces);
+ ReleaseSysCache(oldtup);
+ simple_heap_update(sd, &stup->t_self, stup);
+ }
+ else
+ elog(ERROR, "invalid pg_mv_statistic record (oid=%d)", mvoid);
+
+ /* update indexes too */
+ CatalogUpdateIndexes(sd, stup);
+
+ heap_freetuple(stup);
+
+ heap_close(sd, RowExclusiveLock);
+}
+
+/* multi-variate stats comparator */
+
+/*
+ * qsort_arg comparator for sorting Datums (MV stats)
+ *
+ * This does not maintain the tupnoLink array.
+ */
+int
+compare_scalars_simple(const void *a, const void *b, void *arg)
+{
+ Datum da = *(Datum *) a;
+ Datum db = *(Datum *) b;
+ SortSupport ssup = (SortSupport) arg;
+
+ return ApplySortComparator(da, false, db, false, ssup);
+}
+
+/*
+ * qsort_arg comparator for sorting data when partitioning a MV bucket
+ */
+int
+compare_scalars_partition(const void *a, const void *b, void *arg)
+{
+ Datum da = ((ScalarItem *) a)->value;
+ Datum db = ((ScalarItem *) b)->value;
+ SortSupport ssup = (SortSupport) arg;
+
+ return ApplySortComparator(da, false, db, false, ssup);
+}
+
+/* initialize multi-dimensional sort */
+MultiSortSupport
+multi_sort_init(int ndims)
+{
+ MultiSortSupport mss;
+
+ Assert(ndims >= 2);
+
+ mss = (MultiSortSupport) palloc0(offsetof(MultiSortSupportData, ssup)
+ +sizeof(SortSupportData) * ndims);
+
+ mss->ndims = ndims;
+
+ return mss;
+}
+
+/*
+ * add sort into for dimension 'dim' (index into vacattrstats) to mss,
+ * at the position 'sortattr'
+ */
+void
+multi_sort_add_dimension(MultiSortSupport mss, int sortdim,
+ int dim, VacAttrStats **vacattrstats)
+{
+ /* first, lookup StdAnalyzeData for the dimension (attribute) */
+ SortSupportData ssup;
+ StdAnalyzeData *tmp = (StdAnalyzeData *) vacattrstats[dim]->extra_data;
+
+ Assert(mss != NULL);
+ Assert(sortdim < mss->ndims);
+
+ /* initialize sort support, etc. */
+ memset(&ssup, 0, sizeof(ssup));
+ ssup.ssup_cxt = CurrentMemoryContext;
+
+ /* We always use the default collation for statistics */
+ ssup.ssup_collation = DEFAULT_COLLATION_OID;
+ ssup.ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(tmp->ltopr, &ssup);
+
+ mss->ssup[sortdim] = ssup;
+}
+
+/* compare all the dimensions in the selected order */
+int
+multi_sort_compare(const void *a, const void *b, void *arg)
+{
+ int i;
+ SortItem *ia = (SortItem *) a;
+ SortItem *ib = (SortItem *) b;
+
+ MultiSortSupport mss = (MultiSortSupport) arg;
+
+ for (i = 0; i < mss->ndims; i++)
+ {
+ int compare;
+
+ compare = ApplySortComparator(ia->values[i], ia->isnull[i],
+ ib->values[i], ib->isnull[i],
+ &mss->ssup[i]);
+
+ if (compare != 0)
+ return compare;
+
+ }
+
+ /* equal by default */
+ return 0;
+}
+
+/* compare selected dimension */
+int
+multi_sort_compare_dim(int dim, const SortItem *a, const SortItem *b,
+ MultiSortSupport mss)
+{
+ return ApplySortComparator(a->values[dim], a->isnull[dim],
+ b->values[dim], b->isnull[dim],
+ &mss->ssup[dim]);
+}
+
+int
+multi_sort_compare_dims(int start, int end,
+ const SortItem *a, const SortItem *b,
+ MultiSortSupport mss)
+{
+ int dim;
+
+ for (dim = start; dim <= end; dim++)
+ {
+ int r = ApplySortComparator(a->values[dim], a->isnull[dim],
+ b->values[dim], b->isnull[dim],
+ &mss->ssup[dim]);
+
+ if (r != 0)
+ return r;
+ }
+
+ return 0;
+}
diff --git a/src/backend/utils/mvstats/common.h b/src/backend/utils/mvstats/common.h
new file mode 100644
index 0000000..e471c88
--- /dev/null
+++ b/src/backend/utils/mvstats/common.h
@@ -0,0 +1,80 @@
+/*-------------------------------------------------------------------------
+ *
+ * common.h
+ * POSTGRES multivariate statistics
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/common.h
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "access/sysattr.h"
+#include "access/tuptoaster.h"
+#include "catalog/indexing.h"
+#include "catalog/pg_collation.h"
+#include "catalog/pg_mv_statistic.h"
+#include "foreign/fdwapi.h"
+#include "postmaster/autovacuum.h"
+#include "storage/lmgr.h"
+#include "utils/builtins.h"
+#include "utils/datum.h"
+#include "utils/fmgroids.h"
+#include "utils/mvstats.h"
+#include "utils/sortsupport.h"
+#include "utils/syscache.h"
+
+
+/* FIXME private structure copied from analyze.c */
+
+typedef struct
+{
+ Oid eqopr; /* '=' operator for datatype, if any */
+ Oid eqfunc; /* and associated function */
+ Oid ltopr; /* '<' operator for datatype, if any */
+} StdAnalyzeData;
+
+typedef struct
+{
+ Datum value; /* a data value */
+ int tupno; /* position index for tuple it came from */
+} ScalarItem;
+
+/* multi-sort */
+typedef struct MultiSortSupportData
+{
+ int ndims; /* number of dimensions supported by the */
+ SortSupportData ssup[1]; /* sort support data for each dimension */
+} MultiSortSupportData;
+
+typedef MultiSortSupportData *MultiSortSupport;
+
+typedef struct SortItem
+{
+ Datum *values;
+ bool *isnull;
+} SortItem;
+
+MultiSortSupport multi_sort_init(int ndims);
+
+void multi_sort_add_dimension(MultiSortSupport mss, int sortdim,
+ int dim, VacAttrStats **vacattrstats);
+
+int multi_sort_compare(const void *a, const void *b, void *arg);
+
+int multi_sort_compare_dim(int dim, const SortItem *a,
+ const SortItem *b, MultiSortSupport mss);
+
+int multi_sort_compare_dims(int start, int end, const SortItem *a,
+ const SortItem *b, MultiSortSupport mss);
+
+/* comparators, used when constructing multivariate stats */
+int compare_scalars_simple(const void *a, const void *b, void *arg);
+int compare_scalars_partition(const void *a, const void *b, void *arg);
diff --git a/src/backend/utils/mvstats/mvdist.c b/src/backend/utils/mvstats/mvdist.c
new file mode 100644
index 0000000..188bf99
--- /dev/null
+++ b/src/backend/utils/mvstats/mvdist.c
@@ -0,0 +1,597 @@
+/*-------------------------------------------------------------------------
+ *
+ * mvdist.c
+ * POSTGRES multivariate distinct coefficients
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/mvdist.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include <math.h>
+
+#include "common.h"
+#include "utils/bytea.h"
+#include "utils/lsyscache.h"
+
+static double estimate_ndistinct(double totalrows, int numrows, int d, int f1);
+
+/* internal state for generator of k-combinations of n elements */
+typedef struct CombinationGeneratorData
+{
+
+ int k; /* size of the combination */
+ int current; /* index of the next combination to return */
+
+ int ncombinations; /* number of combinations (size of array) */
+ int *combinations; /* array of pre-built combinations */
+
+} CombinationGeneratorData;
+
+typedef CombinationGeneratorData *CombinationGenerator;
+
+/* generator API */
+static CombinationGenerator generator_init(int2vector *attrs, int k);
+static void generator_free(CombinationGenerator state);
+static int *generator_next(CombinationGenerator state, int2vector *attrs);
+
+static int n_choose_k(int n, int k);
+static int num_combinations(int n);
+static double ndistinct_for_combination(double totalrows, int numrows,
+ HeapTuple *rows, int2vector *attrs, VacAttrStats **stats,
+ int k, int *combination);
+
+/*
+ * Compute ndistinct coefficient for the combination of attributes. This
+ * computes the ndistinct estimate using the same estimator used in analyze.c
+ * and then computes the coefficient.
+ */
+MVNDistinct
+build_mv_ndistinct(double totalrows, int numrows, HeapTuple *rows,
+ int2vector *attrs, VacAttrStats **stats)
+{
+ int i, k;
+ int numattrs = attrs->dim1;
+ int numcombs = num_combinations(numattrs);
+
+ MVNDistinct result;
+
+ result = palloc0(offsetof(MVNDistinctData, items) +
+ numcombs * sizeof(MVNDistinctItem));
+
+ result->nitems = numcombs;
+
+ i = 0;
+ for (k = 2; k <= numattrs; k++)
+ {
+ int * combination;
+ CombinationGenerator generator;
+
+ generator = generator_init(attrs, k);
+
+ while ((combination = generator_next(generator, attrs)))
+ {
+ MVNDistinctItem *item = &result->items[i++];
+
+ item->nattrs = k;
+ item->ndistinct = ndistinct_for_combination(totalrows, numrows, rows,
+ attrs, stats, k, combination);
+
+ item->attrs = palloc(k * sizeof(int));
+ memcpy(item->attrs, combination, k * sizeof(int));
+
+ /* must not overflow the output array */
+ Assert(i <= result->nitems);
+ }
+
+ generator_free(generator);
+ }
+
+ /* must consume exactly the whole output array */
+ Assert(i == result->nitems);
+
+ return result;
+}
+
+/*
+ * ndistinct_for_combination
+ * Estimates number of distinct values in a combination of columns.
+ *
+ * This uses the same ndistinct estimator as compute_scalar_stats() in
+ * ANALYZE, i.e.
+ *
+ * n*d / (n - f1 + f1*n/N)
+ *
+ * except that instead of values in a single column we are dealing with
+ * combination of multiple columns.
+ */
+static double
+ndistinct_for_combination(double totalrows, int numrows, HeapTuple *rows,
+ int2vector *attrs, VacAttrStats **stats,
+ int k, int *combination)
+{
+ int i, j;
+ int f1, cnt, d;
+ int nmultiple, summultiple;
+ MultiSortSupport mss = multi_sort_init(k);
+
+ /*
+ * It's possible to sort the sample rows directly, but this seemed
+ * somehow simpler / less error prone. Another option would be to
+ * allocate the arrays for each SortItem separately, but that'd be
+ * significant overhead (not just CPU, but especially memory bloat).
+ */
+ SortItem * items = (SortItem*)palloc0(numrows * sizeof(SortItem));
+
+ Datum *values = (Datum*)palloc0(sizeof(Datum) * numrows * k);
+ bool *isnull = (bool*)palloc0(sizeof(bool) * numrows * k);
+
+ Assert((k >= 2) && (k <= attrs->dim1));
+
+ for (i = 0; i < numrows; i++)
+ {
+ items[i].values = &values[i * k];
+ items[i].isnull = &isnull[i * k];
+ }
+
+ for (i = 0; i < k; i++)
+ {
+ /* prepare the sort function for the first dimension */
+ multi_sort_add_dimension(mss, i, combination[i], stats);
+
+ /* accumulate all the data into the array and sort it */
+ for (j = 0; j < numrows; j++)
+ {
+ items[j].values[i]
+ = heap_getattr(rows[j], attrs->values[combination[i]],
+ stats[combination[i]]->tupDesc,
+ &items[j].isnull[i]);
+ }
+ }
+
+ qsort_arg((void *) items, numrows, sizeof(SortItem),
+ multi_sort_compare, mss);
+
+ /* count number of distinct combinations */
+
+ f1 = 0;
+ cnt = 1;
+ d = 1;
+ for (i = 1; i < numrows; i++)
+ {
+ if (multi_sort_compare(&items[i], &items[i-1], mss) != 0)
+ {
+ if (cnt == 1)
+ f1 += 1;
+ else
+ {
+ nmultiple += 1;
+ summultiple += cnt;
+ }
+
+ d++;
+ cnt = 0;
+ }
+
+ cnt += 1;
+ }
+
+ if (cnt == 1)
+ f1 += 1;
+ else
+ {
+ nmultiple += 1;
+ summultiple += cnt;
+ }
+
+ return estimate_ndistinct(totalrows, numrows, d, f1);
+}
+
+MVNDistinct
+load_mv_ndistinct(Oid mvoid)
+{
+ bool isnull = false;
+ Datum ndist;
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ HeapTuple htup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(mvoid));
+
+#ifdef USE_ASSERT_CHECKING
+ Form_pg_mv_statistic mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
+ Assert(mvstat->ndist_enabled && mvstat->ndist_built);
+#endif
+
+ ndist = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_standist, &isnull);
+
+ Assert(!isnull);
+
+ ReleaseSysCache(htup);
+
+ return deserialize_mv_ndistinct(DatumGetByteaP(ndist));
+}
+
+/* The Duj1 estimator (already used in analyze.c). */
+static double
+estimate_ndistinct(double totalrows, int numrows, int d, int f1)
+{
+ double numer,
+ denom,
+ ndistinct;
+
+ numer = (double) numrows *(double) d;
+
+ denom = (double) (numrows - f1) +
+ (double) f1 * (double) numrows / totalrows;
+
+ ndistinct = numer / denom;
+
+ /* Clamp to sane range in case of roundoff error */
+ if (ndistinct < (double) d)
+ ndistinct = (double) d;
+
+ if (ndistinct > totalrows)
+ ndistinct = totalrows;
+
+ return floor(ndistinct + 0.5);
+}
+
+
+/*
+ * pg_ndistinct_in - input routine for type pg_ndistinct.
+ *
+ * pg_ndistinct is real enough to be a table column, but it has no operations
+ * of its own, and disallows input too
+ *
+ * XXX This is inspired by what pg_node_tree does.
+ */
+Datum
+pg_ndistinct_in(PG_FUNCTION_ARGS)
+{
+ /*
+ * pg_node_list stores the data in binary form and parsing text input is
+ * not needed, so disallow this.
+ */
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_ndistinct")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_ndistinct - output routine for type pg_ndistinct.
+ *
+ * histograms are serialized into a bytea value, so we simply call byteaout()
+ * to serialize the value into text. But it'd be nice to serialize that into
+ * a meaningful representation (e.g. for inspection by people).
+ */
+Datum
+pg_ndistinct_out(PG_FUNCTION_ARGS)
+{
+ int i, j;
+ char *ret;
+ StringInfoData str;
+
+ bytea *data = PG_GETARG_BYTEA_PP(0);
+
+ MVNDistinct ndist = deserialize_mv_ndistinct(data);
+
+ initStringInfo(&str);
+ appendStringInfoString(&str, "[");
+
+ for (i = 0; i < ndist->nitems; i++)
+ {
+ MVNDistinctItem item = ndist->items[i];
+
+ if (i > 0)
+ appendStringInfoString(&str, ", ");
+
+ appendStringInfoString(&str, "{");
+
+ for (j = 0; j < item.nattrs; j++)
+ {
+ if (j > 0)
+ appendStringInfoString(&str, ", ");
+
+ appendStringInfo(&str, "%d", item.attrs[j]);
+ }
+
+ appendStringInfo(&str, ", %f", item.ndistinct);
+
+ appendStringInfoString(&str, "}");
+ }
+
+ appendStringInfoString(&str, "]");
+
+ ret = pstrdup(str.data);
+ pfree(str.data);
+
+ PG_RETURN_CSTRING(ret);
+}
+
+/*
+ * pg_ndistinct_recv - binary input routine for type pg_ndistinct.
+ */
+Datum
+pg_ndistinct_recv(PG_FUNCTION_ARGS)
+{
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_ndistinct")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_ndistinct_send - binary output routine for type pg_ndistinct.
+ *
+ * XXX Histograms are serialized into a bytea value, so let's just send that.
+ */
+Datum
+pg_ndistinct_send(PG_FUNCTION_ARGS)
+{
+ return byteasend(fcinfo);
+}
+
+static int
+n_choose_k(int n, int k)
+{
+ int i, numer, denom;
+
+ Assert((n > 0) && (k > 0) && (n >= k));
+
+ numer = denom = 1;
+ for (i = 1; i <= k; i++)
+ {
+ numer *= (n - i + 1);
+ denom *= i;
+ }
+
+ Assert(numer % denom == 0);
+
+ return numer / denom;
+}
+
+static int
+num_combinations(int n)
+{
+ int k;
+ int ncombs = 0;
+
+ /* ignore combinations with a single column */
+ for (k = 2; k <= n; k++)
+ ncombs += n_choose_k(n, k);
+
+ return ncombs;
+}
+
+/*
+ * generate all combinations (k elements from n)
+ */
+static void
+generate_combinations_recurse(CombinationGenerator state,
+ int n, int index, int start, int *current)
+{
+ /* If we haven't filled all the elements, simply recurse. */
+ if (index < state->k)
+ {
+ int i;
+
+ /*
+ * The values have to be in ascending order, so make sure we start
+ * with the value passed by parameter.
+ */
+
+ for (i = start; i < n; i++)
+ {
+ current[index] = i;
+ generate_combinations_recurse(state, n, (index+1), (i+1), current);
+ }
+
+ return;
+ }
+ else
+ {
+ /* we got a correct combination */
+ state->combinations = (int*)repalloc(state->combinations,
+ state->k * (state->current + 1) * sizeof(int));
+ memcpy(&state->combinations[(state->k * state->current)],
+ current, state->k * sizeof(int));
+ state->current++;
+ }
+}
+
+/* generate all k-combinations of n elements */
+static void
+generate_combinations(CombinationGenerator state, int n)
+{
+ int *current = (int *) palloc0(sizeof(int) * state->k);
+
+ generate_combinations_recurse(state, n, 0, 0, current);
+
+ pfree(current);
+}
+
+/*
+ * initialize the generator of combinations, and prebuild them.
+ *
+ * This pre-builds all the combinations. We could also generate them in
+ * generator_next(), but this seems simpler.
+ */
+static CombinationGenerator
+generator_init(int2vector *attrs, int k)
+{
+ int n = attrs->dim1;
+ CombinationGenerator state;
+
+ Assert((n >= k) && (k > 0));
+
+ /* allocate the generator state as a single chunk of memory */
+ state = (CombinationGenerator) palloc0(sizeof(CombinationGeneratorData));
+ state->combinations = (int*)palloc(k * sizeof(int));
+
+ state->ncombinations = n_choose_k(n, k);
+ state->current = 0;
+ state->k = k;
+
+ /* now actually pre-generate all the combinations */
+ generate_combinations(state, n);
+
+ /* make sure we got the expected number of combinations */
+ Assert(state->current == state->ncombinations);
+
+ /* reset the number, so we start with the first one */
+ state->current = 0;
+
+ return state;
+}
+
+/* free the generator state */
+static void
+generator_free(CombinationGenerator state)
+{
+ /* we've allocated a single chunk, so just free it */
+ pfree(state);
+}
+
+/* generate next combination */
+static int *
+generator_next(CombinationGenerator state, int2vector *attrs)
+{
+ if (state->current == state->ncombinations)
+ return NULL;
+
+ return &state->combinations[state->k * state->current++];
+}
+
+/*
+ * serialize list of ndistinct items into a bytea
+ */
+bytea *
+serialize_mv_ndistinct(MVNDistinct ndistinct)
+{
+ int i;
+ bytea *output;
+ char *tmp;
+
+ /* we need to store nitems */
+ Size len = VARHDRSZ + offsetof(MVNDistinctData, items) +
+ ndistinct->nitems * offsetof(MVNDistinctItem, attrs);
+
+ /* and also include space for the actual attribute numbers */
+ for (i = 0; i < ndistinct->nitems; i++)
+ len += (sizeof(int) * ndistinct->items[i].nattrs);
+
+ output = (bytea *) palloc0(len);
+ SET_VARSIZE(output, len);
+
+ tmp = VARDATA(output);
+
+ ndistinct->magic = MVSTAT_NDISTINCT_MAGIC;
+ ndistinct->type = MVSTAT_NDISTINCT_TYPE_BASIC;
+
+ /* first, store the number of items */
+ memcpy(tmp, ndistinct, offsetof(MVNDistinctData, items));
+ tmp += offsetof(MVNDistinctData, items);
+
+ /* store number of attributes and attribute numbers for each ndistinct entry */
+ for (i = 0; i < ndistinct->nitems; i++)
+ {
+ MVNDistinctItem item = ndistinct->items[i];
+
+ memcpy(tmp, &item, offsetof(MVNDistinctItem, attrs));
+ tmp += offsetof(MVNDistinctItem, attrs);
+
+ memcpy(tmp, item.attrs, sizeof(int) * item.nattrs);
+ tmp += sizeof(int) * item.nattrs;
+
+ Assert(tmp <= ((char *) output + len));
+ }
+
+ return output;
+}
+
+/*
+ * Reads serialized ndistinct into MVNDistinct structure.
+ */
+MVNDistinct
+deserialize_mv_ndistinct(bytea *data)
+{
+ int i;
+ Size expected_size;
+ MVNDistinct ndistinct;
+ char *tmp;
+
+ if (data == NULL)
+ return NULL;
+
+ if (VARSIZE_ANY_EXHDR(data) < offsetof(MVNDistinctData, items))
+ elog(ERROR, "invalid MVNDistinct size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), offsetof(MVNDistinctData, items));
+
+ /* read the MVNDistinct header */
+ ndistinct = (MVNDistinct) palloc0(sizeof(MVNDistinctData));
+
+ /* initialize pointer to the data part (skip the varlena header) */
+ tmp = VARDATA_ANY(data);
+
+ /* get the header and perform basic sanity checks */
+ memcpy(ndistinct, tmp, offsetof(MVNDistinctData, items));
+ tmp += offsetof(MVNDistinctData, items);
+
+ if (ndistinct->magic != MVSTAT_NDISTINCT_MAGIC)
+ elog(ERROR, "invalid ndistinct magic %d (expected %dd)",
+ ndistinct->magic, MVSTAT_NDISTINCT_MAGIC);
+
+ if (ndistinct->type != MVSTAT_NDISTINCT_TYPE_BASIC)
+ elog(ERROR, "invalid ndistinct type %d (expected %dd)",
+ ndistinct->type, MVSTAT_NDISTINCT_TYPE_BASIC);
+
+ Assert(ndistinct->nitems > 0);
+
+ /* what minimum bytea size do we expect for those parameters */
+ expected_size = offsetof(MVNDistinctData, items) +
+ ndistinct->nitems * (offsetof(MVNDistinctItem, attrs) + sizeof(int) * 2);
+
+ if (VARSIZE_ANY_EXHDR(data) < expected_size)
+ elog(ERROR, "invalid dependencies size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ /* allocate space for the ndistinct items */
+ ndistinct = repalloc(ndistinct, offsetof(MVNDistinctData, items) +
+ (ndistinct->nitems * sizeof(MVNDistinctItem)));
+
+ for (i = 0; i < ndistinct->nitems; i++)
+ {
+ MVNDistinctItem *item = &ndistinct->items[i];
+
+ /* number of attributes */
+ memcpy(item, tmp, offsetof(MVNDistinctItem, attrs));
+ tmp += offsetof(MVNDistinctItem, attrs);
+
+ /* is the number of attributes valid? */
+ Assert((item->nattrs >= 2) && (item->nattrs <= MVSTATS_MAX_DIMENSIONS));
+
+ /* now that we know the number of attributes, allocate the attribute */
+ item->attrs = (int*)palloc0(item->nattrs * sizeof(int));
+
+ /* copy attribute numbers */
+ memcpy(item->attrs, tmp, sizeof(int) * item->nattrs);
+ tmp += sizeof(int) * item->nattrs;
+
+ /* still within the bytea */
+ Assert(tmp <= ((char *) data + VARSIZE_ANY(data)));
+ }
+
+ /* we should have consumed the whole bytea exactly */
+ Assert(tmp == ((char *) data + VARSIZE_ANY(data)));
+
+ return ndistinct;
+}
diff --git a/src/bin/psql/describe.c b/src/bin/psql/describe.c
index c501168..e7d5b51 100644
--- a/src/bin/psql/describe.c
+++ b/src/bin/psql/describe.c
@@ -2293,6 +2293,50 @@ describeOneTableDetails(const char *schemaname,
PQclear(result);
}
+ /* print any multivariate statistics */
+ if (pset.sversion >= 90600)
+ {
+ printfPQExpBuffer(&buf,
+ "SELECT oid, stanamespace::regnamespace AS nsp, staname, stakeys,\n"
+ " ndist_enabled,\n"
+ " ndist_built,\n"
+ " (SELECT string_agg(attname::text,', ')\n"
+ " FROM ((SELECT unnest(stakeys) AS attnum) s\n"
+ " JOIN pg_attribute a ON (starelid = a.attrelid and a.attnum = s.attnum))) AS attnums\n"
+ "FROM pg_mv_statistic stat WHERE starelid = '%s' ORDER BY 1;",
+ oid);
+
+ result = PSQLexec(buf.data);
+ if (!result)
+ goto error_return;
+ else
+ tuples = PQntuples(result);
+
+ if (tuples > 0)
+ {
+ printTableAddFooter(&cont, _("Statistics:"));
+ for (i = 0; i < tuples; i++)
+ {
+ printfPQExpBuffer(&buf, " ");
+
+ /* statistics name (qualified with namespace) */
+ appendPQExpBuffer(&buf, "\"%s.%s\" ",
+ PQgetvalue(result, i, 1),
+ PQgetvalue(result, i, 2));
+
+ /* options */
+ if (!strcmp(PQgetvalue(result, i, 4), "t"))
+ appendPQExpBuffer(&buf, "(dependencies)");
+
+ appendPQExpBuffer(&buf, " ON (%s)",
+ PQgetvalue(result, i, 6));
+
+ printTableAddFooter(&cont, buf.data);
+ }
+ }
+ PQclear(result);
+ }
+
/* print rules */
if (tableinfo.hasrules && tableinfo.relkind != 'm')
{
diff --git a/src/include/catalog/dependency.h b/src/include/catalog/dependency.h
index 10759c7..86acca4 100644
--- a/src/include/catalog/dependency.h
+++ b/src/include/catalog/dependency.h
@@ -164,10 +164,11 @@ typedef enum ObjectClass
OCLASS_PUBLICATION, /* pg_publication */
OCLASS_PUBLICATION_REL, /* pg_publication_rel */
OCLASS_SUBSCRIPTION, /* pg_subscription */
- OCLASS_TRANSFORM /* pg_transform */
+ OCLASS_TRANSFORM, /* pg_transform */
+ OCLASS_STATISTICS /* pg_mv_statistics */
} ObjectClass;
-#define LAST_OCLASS OCLASS_TRANSFORM
+#define LAST_OCLASS OCLASS_STATISTICS
/* flag bits for performDeletion/performMultipleDeletions: */
#define PERFORM_DELETION_INTERNAL 0x0001 /* internal action */
diff --git a/src/include/catalog/heap.h b/src/include/catalog/heap.h
index 1187797..2d2a8c9 100644
--- a/src/include/catalog/heap.h
+++ b/src/include/catalog/heap.h
@@ -119,6 +119,7 @@ extern void RemoveAttrDefault(Oid relid, AttrNumber attnum,
DropBehavior behavior, bool complain, bool internal);
extern void RemoveAttrDefaultById(Oid attrdefId);
extern void RemoveStatistics(Oid relid, AttrNumber attnum);
+extern void RemoveMVStatistics(Oid relid, AttrNumber attnum);
extern Form_pg_attribute SystemAttributeDefinition(AttrNumber attno,
bool relhasoids);
diff --git a/src/include/catalog/indexing.h b/src/include/catalog/indexing.h
index a3635a4..e938300 100644
--- a/src/include/catalog/indexing.h
+++ b/src/include/catalog/indexing.h
@@ -176,6 +176,13 @@ DECLARE_UNIQUE_INDEX(pg_largeobject_loid_pn_index, 2683, on pg_largeobject using
DECLARE_UNIQUE_INDEX(pg_largeobject_metadata_oid_index, 2996, on pg_largeobject_metadata using btree(oid oid_ops));
#define LargeObjectMetadataOidIndexId 2996
+DECLARE_UNIQUE_INDEX(pg_mv_statistic_oid_index, 3380, on pg_mv_statistic using btree(oid oid_ops));
+#define MvStatisticOidIndexId 3380
+DECLARE_UNIQUE_INDEX(pg_mv_statistic_name_index, 3997, on pg_mv_statistic using btree(staname name_ops, stanamespace oid_ops));
+#define MvStatisticNameIndexId 3997
+DECLARE_INDEX(pg_mv_statistic_relid_index, 3379, on pg_mv_statistic using btree(starelid oid_ops));
+#define MvStatisticRelidIndexId 3379
+
DECLARE_UNIQUE_INDEX(pg_namespace_nspname_index, 2684, on pg_namespace using btree(nspname name_ops));
#define NamespaceNameIndexId 2684
DECLARE_UNIQUE_INDEX(pg_namespace_oid_index, 2685, on pg_namespace using btree(oid oid_ops));
diff --git a/src/include/catalog/namespace.h b/src/include/catalog/namespace.h
index dbeb25b..35e0e2b 100644
--- a/src/include/catalog/namespace.h
+++ b/src/include/catalog/namespace.h
@@ -141,6 +141,8 @@ extern Oid get_collation_oid(List *collname, bool missing_ok);
extern Oid get_conversion_oid(List *conname, bool missing_ok);
extern Oid FindDefaultConversionProc(int32 for_encoding, int32 to_encoding);
+extern Oid get_statistics_oid(List *names, bool missing_ok);
+
/* initialization & transaction cleanup code */
extern void InitializeSearchPath(void);
extern void AtEOXact_Namespace(bool isCommit, bool parallel);
diff --git a/src/include/catalog/pg_cast.h b/src/include/catalog/pg_cast.h
index 80a40ab..bf39d43 100644
--- a/src/include/catalog/pg_cast.h
+++ b/src/include/catalog/pg_cast.h
@@ -254,6 +254,10 @@ DATA(insert ( 23 18 78 e f ));
/* pg_node_tree can be coerced to, but not from, text */
DATA(insert ( 194 25 0 i b ));
+/* pg_ndistinct can be coerced to, but not from, bytea and text */
+DATA(insert ( 3353 17 0 i b ));
+DATA(insert ( 3353 25 0 i i ));
+
/*
* Datetime category
*/
diff --git a/src/include/catalog/pg_mv_statistic.h b/src/include/catalog/pg_mv_statistic.h
new file mode 100644
index 0000000..fad80a3
--- /dev/null
+++ b/src/include/catalog/pg_mv_statistic.h
@@ -0,0 +1,78 @@
+/*-------------------------------------------------------------------------
+ *
+ * pg_mv_statistic.h
+ * definition of the system "multivariate statistic" relation (pg_mv_statistic)
+ * along with the relation's initial contents.
+ *
+ *
+ * Portions Copyright (c) 1996-2014, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/catalog/pg_mv_statistic.h
+ *
+ * NOTES
+ * the genbki.pl script reads this file and generates .bki
+ * information from the DATA() statements.
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef PG_MV_STATISTIC_H
+#define PG_MV_STATISTIC_H
+
+#include "catalog/genbki.h"
+
+/* ----------------
+ * pg_mv_statistic definition. cpp turns this into
+ * typedef struct FormData_pg_mv_statistic
+ * ----------------
+ */
+#define MvStatisticRelationId 3381
+
+CATALOG(pg_mv_statistic,3381)
+{
+ /* These fields form the unique key for the entry: */
+ Oid starelid; /* relation containing attributes */
+ NameData staname; /* statistics name */
+ Oid stanamespace; /* OID of namespace containing this statistics */
+ Oid staowner; /* statistics owner */
+
+ /* statistics requested to build */
+ bool ndist_enabled; /* build ndist coefficient? */
+
+ /* statistics that are available (if requested) */
+ bool ndist_built; /* ndistinct coeff built */
+
+ /*
+ * variable-length fields start here, but we allow direct access to
+ * stakeys
+ */
+ int2vector stakeys; /* array of column keys */
+
+#ifdef CATALOG_VARLEN
+ pg_ndistinct standist; /* ndistinct coeff (serialized) */
+#endif
+
+} FormData_pg_mv_statistic;
+
+/* ----------------
+ * Form_pg_mv_statistic corresponds to a pointer to a tuple with
+ * the format of pg_mv_statistic relation.
+ * ----------------
+ */
+typedef FormData_pg_mv_statistic *Form_pg_mv_statistic;
+
+/* ----------------
+ * compiler constants for pg_mv_statistic
+ * ----------------
+ */
+#define Natts_pg_mv_statistic 8
+#define Anum_pg_mv_statistic_starelid 1
+#define Anum_pg_mv_statistic_staname 2
+#define Anum_pg_mv_statistic_stanamespace 3
+#define Anum_pg_mv_statistic_staowner 4
+#define Anum_pg_mv_statistic_ndist_enabled 5
+#define Anum_pg_mv_statistic_ndist_built 6
+#define Anum_pg_mv_statistic_stakeys 7
+#define Anum_pg_mv_statistic_standist 8
+
+#endif /* PG_MV_STATISTIC_H */
diff --git a/src/include/catalog/pg_proc.h b/src/include/catalog/pg_proc.h
index 31c828a..940a991 100644
--- a/src/include/catalog/pg_proc.h
+++ b/src/include/catalog/pg_proc.h
@@ -2726,6 +2726,15 @@ DESCR("current user privilege on any column by rel name");
DATA(insert OID = 3029 ( has_any_column_privilege PGNSP PGUID 12 10 0 0 0 f f f f t f s s 2 0 16 "26 25" _null_ _null_ _null_ _null_ _null_ has_any_column_privilege_id _null_ _null_ _null_ ));
DESCR("current user privilege on any column by rel oid");
+DATA(insert OID = 3354 ( pg_ndistinct_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 3353 "2275" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_in _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3355 ( pg_ndistinct_out PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 2275 "3353" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_out _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3356 ( pg_ndistinct_recv PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 3353 "2281" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_recv _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3357 ( pg_ndistinct_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "3353" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_send _null_ _null_ _null_ ));
+DESCR("I/O");
+
DATA(insert OID = 1928 ( pg_stat_get_numscans PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_numscans _null_ _null_ _null_ ));
DESCR("statistics: number of scans done for table/index");
DATA(insert OID = 1929 ( pg_stat_get_tuples_returned PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_tuples_returned _null_ _null_ _null_ ));
diff --git a/src/include/catalog/pg_type.h b/src/include/catalog/pg_type.h
index 6e4c65e..9c9caf3 100644
--- a/src/include/catalog/pg_type.h
+++ b/src/include/catalog/pg_type.h
@@ -364,6 +364,10 @@ DATA(insert OID = 194 ( pg_node_tree PGNSP PGUID -1 f b S f t \054 0 0 0 pg_node
DESCR("string representing an internal node tree");
#define PGNODETREEOID 194
+DATA(insert OID = 3353 ( pg_ndistinct PGNSP PGUID -1 f b S f t \054 0 0 0 pg_ndistinct_in pg_ndistinct_out pg_ndistinct_recv pg_ndistinct_send - - - i x f 0 -1 0 100 _null_ _null_ _null_ ));
+DESCR("multivariate ndistinct coefficients");
+#define PGNDISTINCTOID 3353
+
DATA(insert OID = 32 ( pg_ddl_command PGNSP PGUID SIZEOF_POINTER t p P f t \054 0 0 0 pg_ddl_command_in pg_ddl_command_out pg_ddl_command_recv pg_ddl_command_send - - - ALIGNOF_POINTER p f 0 -1 0 0 _null_ _null_ _null_ ));
DESCR("internal type for passing CollectedCommand");
#define PGDDLCOMMANDOID 32
diff --git a/src/include/catalog/toasting.h b/src/include/catalog/toasting.h
index db7f145..37a2f7a 100644
--- a/src/include/catalog/toasting.h
+++ b/src/include/catalog/toasting.h
@@ -49,6 +49,7 @@ extern void BootstrapToastTable(char *relName,
DECLARE_TOAST(pg_attrdef, 2830, 2831);
DECLARE_TOAST(pg_constraint, 2832, 2833);
DECLARE_TOAST(pg_description, 2834, 2835);
+DECLARE_TOAST(pg_mv_statistic, 3439, 3440);
DECLARE_TOAST(pg_proc, 2836, 2837);
DECLARE_TOAST(pg_rewrite, 2838, 2839);
DECLARE_TOAST(pg_seclabel, 3598, 3599);
diff --git a/src/include/commands/defrem.h b/src/include/commands/defrem.h
index 8740cee..c323e81 100644
--- a/src/include/commands/defrem.h
+++ b/src/include/commands/defrem.h
@@ -77,6 +77,10 @@ extern ObjectAddress DefineOperator(List *names, List *parameters);
extern void RemoveOperatorById(Oid operOid);
extern ObjectAddress AlterOperator(AlterOperatorStmt *stmt);
+/* commands/statscmds.c */
+extern ObjectAddress CreateStatistics(CreateStatsStmt *stmt);
+extern void RemoveStatisticsById(Oid statsOid);
+
/* commands/aggregatecmds.c */
extern ObjectAddress DefineAggregate(ParseState *pstate, List *name, List *args, bool oldstyle,
List *parameters);
diff --git a/src/include/nodes/nodes.h b/src/include/nodes/nodes.h
index 95dd8ba..e828b43 100644
--- a/src/include/nodes/nodes.h
+++ b/src/include/nodes/nodes.h
@@ -272,6 +272,7 @@ typedef enum NodeTag
T_PlaceHolderInfo,
T_MinMaxAggInfo,
T_PlannerParamItem,
+ T_MVStatisticInfo,
/*
* TAGS FOR MEMORY NODES (memnodes.h)
@@ -416,6 +417,7 @@ typedef enum NodeTag
T_CreateSubscriptionStmt,
T_AlterSubscriptionStmt,
T_DropSubscriptionStmt,
+ T_CreateStatsStmt,
/*
* TAGS FOR PARSE TREE NODES (parsenodes.h)
diff --git a/src/include/nodes/parsenodes.h b/src/include/nodes/parsenodes.h
index 07a8436..18e1dd1 100644
--- a/src/include/nodes/parsenodes.h
+++ b/src/include/nodes/parsenodes.h
@@ -611,6 +611,16 @@ typedef struct ColumnDef
int location; /* parse location, or -1 if none/unknown */
} ColumnDef;
+typedef struct CreateStatsStmt
+{
+ NodeTag type;
+ List *defnames; /* qualified name (list of Value strings) */
+ RangeVar *relation; /* relation to build statistics on */
+ List *keys; /* String nodes naming referenced column(s) */
+ bool if_not_exists; /* do nothing if statistics already exists */
+} CreateStatsStmt;
+
+
/*
* TableLikeClause - CREATE TABLE ( ... LIKE ... ) clause
*/
@@ -1554,6 +1564,7 @@ typedef enum ObjectType
OBJECT_SCHEMA,
OBJECT_SEQUENCE,
OBJECT_SUBSCRIPTION,
+ OBJECT_STATISTICS,
OBJECT_TABCONSTRAINT,
OBJECT_TABLE,
OBJECT_TABLESPACE,
diff --git a/src/include/nodes/relation.h b/src/include/nodes/relation.h
index 643be54..7a55151 100644
--- a/src/include/nodes/relation.h
+++ b/src/include/nodes/relation.h
@@ -525,6 +525,7 @@ typedef struct RelOptInfo
List *lateral_vars; /* LATERAL Vars and PHVs referenced by rel */
Relids lateral_referencers; /* rels that reference me laterally */
List *indexlist; /* list of IndexOptInfo */
+ List *mvstatlist; /* list of MVStatisticInfo */
BlockNumber pages; /* size estimates derived from pg_class */
double tuples;
double allvisfrac;
@@ -663,6 +664,32 @@ typedef struct ForeignKeyOptInfo
List *rinfos[INDEX_MAX_KEYS];
} ForeignKeyOptInfo;
+/*
+ * MVStatisticInfo
+ * Information about multivariate stats for planning/optimization
+ *
+ * This contains information about which columns are covered by the
+ * statistics (stakeys), which options were requested while adding the
+ * statistics (*_enabled), and which kinds of statistics were actually
+ * built and are available for the optimizer (*_built).
+ */
+typedef struct MVStatisticInfo
+{
+ NodeTag type;
+
+ Oid mvoid; /* OID of the statistics row */
+ RelOptInfo *rel; /* back-link to index's table */
+
+ /* enabled statistics */
+ bool ndist_enabled; /* ndistinct coefficient enabled */
+
+ /* built/available statistics */
+ bool ndist_built; /* ndistinct coefficient built */
+
+ /* columns in the statistics (attnums) */
+ int2vector *stakeys; /* attnums of the columns covered */
+
+} MVStatisticInfo;
/*
* EquivalenceClasses
diff --git a/src/include/utils/acl.h b/src/include/utils/acl.h
index 686141b..4368adc 100644
--- a/src/include/utils/acl.h
+++ b/src/include/utils/acl.h
@@ -322,6 +322,7 @@ extern bool pg_event_trigger_ownercheck(Oid et_oid, Oid roleid);
extern bool pg_extension_ownercheck(Oid ext_oid, Oid roleid);
extern bool pg_publication_ownercheck(Oid pub_oid, Oid roleid);
extern bool pg_subscription_ownercheck(Oid sub_oid, Oid roleid);
+extern bool pg_statistics_ownercheck(Oid stat_oid, Oid roleid);
extern bool has_createrole_privilege(Oid roleid);
extern bool has_bypassrls_privilege(Oid roleid);
diff --git a/src/include/utils/builtins.h b/src/include/utils/builtins.h
index 5bdca82..262ee94 100644
--- a/src/include/utils/builtins.h
+++ b/src/include/utils/builtins.h
@@ -68,6 +68,12 @@ extern int float8_cmp_internal(float8 a, float8 b);
extern oidvector *buildoidvector(const Oid *oids, int n);
extern Oid oidparse(Node *node);
+/* mvdist.c */
+extern Datum pg_ndistinct_in(PG_FUNCTION_ARGS);
+extern Datum pg_ndistinct_out(PG_FUNCTION_ARGS);
+extern Datum pg_ndistinct_recv(PG_FUNCTION_ARGS);
+extern Datum pg_ndistinct_send(PG_FUNCTION_ARGS);
+
/* regexp.c */
extern char *regexp_fixed_prefix(text *text_re, bool case_insensitive,
Oid collation, bool *exact);
diff --git a/src/include/utils/mvstats.h b/src/include/utils/mvstats.h
new file mode 100644
index 0000000..0660c59
--- /dev/null
+++ b/src/include/utils/mvstats.h
@@ -0,0 +1,57 @@
+/*-------------------------------------------------------------------------
+ *
+ * mvstats.h
+ * Multivariate statistics and selectivity estimation functions.
+ *
+ *
+ * Portions Copyright (c) 1996-2014, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ * src/include/utils/mvstats.h
+ *
+ *-------------------------------------------------------------------------
+ */
+#ifndef MVSTATS_H
+#define MVSTATS_H
+
+#include "fmgr.h"
+#include "commands/vacuum.h"
+
+#define MVSTATS_MAX_DIMENSIONS 8 /* max number of attributes */
+
+#define MVSTAT_NDISTINCT_MAGIC 0xA352BFA4 /* marks serialized bytea */
+#define MVSTAT_NDISTINCT_TYPE_BASIC 1 /* basic MCV list type */
+
+/* Multivariate distinct coefficients. */
+typedef struct MVNDistinctItem {
+ double ndistinct;
+ int nattrs;
+ int *attrs;
+} MVNDistinctItem;
+
+typedef struct MVNDistinctData {
+ uint32 magic; /* magic constant marker */
+ uint32 type; /* type of ndistinct (BASIC) */
+ int nitems;
+ MVNDistinctItem items[FLEXIBLE_ARRAY_MEMBER];
+} MVNDistinctData;
+
+typedef MVNDistinctData *MVNDistinct;
+
+
+MVNDistinct load_mv_ndistinct(Oid mvoid);
+
+bytea *serialize_mv_ndistinct(MVNDistinct ndistinct);
+
+/* deserialization of stats (serialization is private to analyze) */
+MVNDistinct deserialize_mv_ndistinct(bytea *data);
+
+
+MVNDistinct build_mv_ndistinct(double totalrows, int numrows, HeapTuple *rows,
+ int2vector *attrs, VacAttrStats **stats);
+
+void build_mv_stats(Relation onerel, double totalrows,
+ int numrows, HeapTuple *rows,
+ int natts, VacAttrStats **vacattrstats);
+
+#endif
diff --git a/src/include/utils/rel.h b/src/include/utils/rel.h
index a617a7c..121d4c9 100644
--- a/src/include/utils/rel.h
+++ b/src/include/utils/rel.h
@@ -92,6 +92,7 @@ typedef struct RelationData
bool rd_isvalid; /* relcache entry is valid */
char rd_indexvalid; /* state of rd_indexlist: 0 = not valid, 1 =
* valid, 2 = temporarily forced */
+ bool rd_mvstatvalid; /* state of rd_mvstatlist: true/false */
/*
* rd_createSubid is the ID of the highest subtransaction the rel has
@@ -136,6 +137,9 @@ typedef struct RelationData
Oid rd_pkindex; /* OID of primary key, if any */
Oid rd_replidindex; /* OID of replica identity index, if any */
+ /* data managed by RelationGetMVStatList: */
+ List *rd_mvstatlist; /* list of OIDs of multivariate stats */
+
/* data managed by RelationGetIndexAttrBitmap: */
Bitmapset *rd_indexattr; /* identifies columns used in indexes */
Bitmapset *rd_keyattr; /* cols that can be ref'd by foreign keys */
diff --git a/src/include/utils/relcache.h b/src/include/utils/relcache.h
index da36b67..bee390e 100644
--- a/src/include/utils/relcache.h
+++ b/src/include/utils/relcache.h
@@ -39,6 +39,7 @@ extern void RelationClose(Relation relation);
*/
extern List *RelationGetFKeyList(Relation relation);
extern List *RelationGetIndexList(Relation relation);
+extern List *RelationGetMVStatList(Relation relation);
extern Oid RelationGetOidIndex(Relation relation);
extern Oid RelationGetPrimaryKeyIndex(Relation relation);
extern Oid RelationGetReplicaIndex(Relation relation);
diff --git a/src/include/utils/syscache.h b/src/include/utils/syscache.h
index 66f60d2..d4ebbf7 100644
--- a/src/include/utils/syscache.h
+++ b/src/include/utils/syscache.h
@@ -66,6 +66,8 @@ enum SysCacheIdentifier
INDEXRELID,
LANGNAME,
LANGOID,
+ MVSTATNAMENSP,
+ MVSTATOID,
NAMESPACENAME,
NAMESPACEOID,
OPERNAMENSP,
diff --git a/src/test/regress/expected/mv_ndistinct.out b/src/test/regress/expected/mv_ndistinct.out
new file mode 100644
index 0000000..5f55091
--- /dev/null
+++ b/src/test/regress/expected/mv_ndistinct.out
@@ -0,0 +1,117 @@
+-- data type passed by value
+CREATE TABLE ndistinct (
+ a INT,
+ b INT,
+ c INT,
+ d INT
+);
+-- unknown column
+CREATE STATISTICS s10 ON (unknown_column) FROM ndistinct;
+ERROR: column "unknown_column" referenced in statistics does not exist
+-- single column
+CREATE STATISTICS s10 ON (a) FROM ndistinct;
+ERROR: statistics require at least 2 columns
+-- single column, duplicated
+CREATE STATISTICS s10 ON (a,a) FROM ndistinct;
+ERROR: duplicate column name in statistics definition
+-- two columns, one duplicated
+CREATE STATISTICS s10 ON (a, a, b) FROM ndistinct;
+ERROR: duplicate column name in statistics definition
+-- correct command
+CREATE STATISTICS s10 ON (a, b, c) FROM ndistinct;
+-- perfectly correlated groups
+INSERT INTO ndistinct
+ SELECT i/100, i/100, i/100 FROM generate_series(1,10000) s(i);
+ANALYZE ndistinct;
+SELECT ndist_enabled, ndist_built, standist
+ FROM pg_mv_statistic WHERE starelid = 'ndistinct'::regclass;
+ ndist_enabled | ndist_built | standist
+---------------+-------------+-------------------------------------------------------------------------------------
+ t | t | [{0, 1, 101.000000}, {0, 2, 101.000000}, {1, 2, 101.000000}, {0, 1, 2, 101.000000}]
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b;
+ QUERY PLAN
+-----------------------------
+ HashAggregate
+ Group Key: a, b
+ -> Seq Scan on ndistinct
+(3 rows)
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c;
+ QUERY PLAN
+-----------------------------
+ HashAggregate
+ Group Key: a, b, c
+ -> Seq Scan on ndistinct
+(3 rows)
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c, d;
+ QUERY PLAN
+-----------------------------
+ HashAggregate
+ Group Key: a, b, c, d
+ -> Seq Scan on ndistinct
+(3 rows)
+
+TRUNCATE TABLE ndistinct;
+-- partially correlated groups
+INSERT INTO ndistinct
+ SELECT i/50, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE ndistinct;
+SELECT ndist_enabled, ndist_built, standist
+ FROM pg_mv_statistic WHERE starelid = 'ndistinct'::regclass;
+ ndist_enabled | ndist_built | standist
+---------------+-------------+-------------------------------------------------------------------------------------
+ t | t | [{0, 1, 201.000000}, {0, 2, 201.000000}, {1, 2, 101.000000}, {0, 1, 2, 201.000000}]
+(1 row)
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b;
+ QUERY PLAN
+---------------------------------------------------------------------
+ HashAggregate (cost=230.00..232.01 rows=201 width=16)
+ Group Key: a, b
+ -> Seq Scan on ndistinct (cost=0.00..155.00 rows=10000 width=8)
+(3 rows)
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c;
+ QUERY PLAN
+----------------------------------------------------------------------
+ HashAggregate (cost=255.00..257.01 rows=201 width=20)
+ Group Key: a, b, c
+ -> Seq Scan on ndistinct (cost=0.00..155.00 rows=10000 width=12)
+(3 rows)
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c, d;
+ QUERY PLAN
+----------------------------------------------------------------------
+ HashAggregate (cost=280.00..290.00 rows=1000 width=24)
+ Group Key: a, b, c, d
+ -> Seq Scan on ndistinct (cost=0.00..155.00 rows=10000 width=16)
+(3 rows)
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY b, c, d;
+ QUERY PLAN
+----------------------------------------------------------------------
+ HashAggregate (cost=255.00..265.00 rows=1000 width=20)
+ Group Key: b, c, d
+ -> Seq Scan on ndistinct (cost=0.00..155.00 rows=10000 width=12)
+(3 rows)
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, d;
+ QUERY PLAN
+---------------------------------------------------------------------
+ HashAggregate (cost=230.00..240.00 rows=1000 width=16)
+ Group Key: a, d
+ -> Seq Scan on ndistinct (cost=0.00..155.00 rows=10000 width=8)
+(3 rows)
+
+DROP TABLE ndistinct;
diff --git a/src/test/regress/expected/object_address.out b/src/test/regress/expected/object_address.out
index ec5ada9..2b5c022 100644
--- a/src/test/regress/expected/object_address.out
+++ b/src/test/regress/expected/object_address.out
@@ -38,6 +38,7 @@ CREATE TRANSFORM FOR int LANGUAGE SQL (
TO SQL WITH FUNCTION int4recv(internal));
CREATE PUBLICATION addr_pub FOR TABLE addr_nsp.gentable;
CREATE SUBSCRIPTION addr_sub CONNECTION '' PUBLICATION bar WITH (DISABLED, NOCREATE SLOT);
+CREATE STATISTICS addr_nsp.gentable_stat ON (a,b) FROM addr_nsp.gentable;
-- test some error cases
SELECT pg_get_object_address('stone', '{}', '{}');
ERROR: unrecognized object type "stone"
@@ -399,7 +400,8 @@ WITH objects (type, name, args) AS (VALUES
('access method', '{btree}', '{}'),
('publication', '{addr_pub}', '{}'),
('publication relation', '{addr_nsp, gentable}', '{addr_pub}'),
- ('subscription', '{addr_sub}', '{}')
+ ('subscription', '{addr_sub}', '{}'),
+ ('statistics', '{addr_nsp, gentable_stat}', '{}')
)
SELECT (pg_identify_object(addr1.classid, addr1.objid, addr1.subobjid)).*,
-- test roundtrip through pg_identify_object_as_address
@@ -447,6 +449,7 @@ SELECT (pg_identify_object(addr1.classid, addr1.objid, addr1.subobjid)).*,
trigger | | | t on addr_nsp.gentable | t
operator family | pg_catalog | integer_ops | pg_catalog.integer_ops USING btree | t
policy | | | genpol on addr_nsp.gentable | t
+ statistics | addr_nsp | gentable_stat | addr_nsp.gentable_stat | t
collation | pg_catalog | "default" | pg_catalog."default" | t
transform | | | for integer on language sql | t
text search dictionary | addr_nsp | addr_ts_dict | addr_nsp.addr_ts_dict | t
@@ -456,7 +459,7 @@ SELECT (pg_identify_object(addr1.classid, addr1.objid, addr1.subobjid)).*,
subscription | | addr_sub | addr_sub | t
publication | | addr_pub | addr_pub | t
publication relation | | | gentable in publication addr_pub | t
-(45 rows)
+(46 rows)
---
--- Cleanup resources
diff --git a/src/test/regress/expected/opr_sanity.out b/src/test/regress/expected/opr_sanity.out
index 0bcec13..9a26205 100644
--- a/src/test/regress/expected/opr_sanity.out
+++ b/src/test/regress/expected/opr_sanity.out
@@ -817,11 +817,12 @@ WHERE c.castmethod = 'b' AND
text | character | 0 | i
character varying | character | 0 | i
pg_node_tree | text | 0 | i
+ pg_ndistinct | bytea | 0 | i
cidr | inet | 0 | i
xml | text | 0 | a
xml | character varying | 0 | a
xml | character | 0 | a
-(7 rows)
+(8 rows)
-- **************** pg_conversion ****************
-- Look for illegal values in pg_conversion fields.
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 60abcad..2c54779 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1376,6 +1376,14 @@ pg_matviews| SELECT n.nspname AS schemaname,
LEFT JOIN pg_namespace n ON ((n.oid = c.relnamespace)))
LEFT JOIN pg_tablespace t ON ((t.oid = c.reltablespace)))
WHERE (c.relkind = 'm'::"char");
+pg_mv_stats| SELECT n.nspname AS schemaname,
+ c.relname AS tablename,
+ s.staname,
+ s.stakeys AS attnums,
+ length((s.standist)::text) AS ndistbytes
+ FROM ((pg_mv_statistic s
+ JOIN pg_class c ON ((c.oid = s.starelid)))
+ LEFT JOIN pg_namespace n ON ((n.oid = c.relnamespace)));
pg_policies| SELECT n.nspname AS schemaname,
c.relname AS tablename,
pol.polname AS policyname,
diff --git a/src/test/regress/expected/sanity_check.out b/src/test/regress/expected/sanity_check.out
index 0af013f..9d6bd18 100644
--- a/src/test/regress/expected/sanity_check.out
+++ b/src/test/regress/expected/sanity_check.out
@@ -116,6 +116,7 @@ pg_init_privs|t
pg_language|t
pg_largeobject|t
pg_largeobject_metadata|t
+pg_mv_statistic|t
pg_namespace|t
pg_opclass|t
pg_operator|t
diff --git a/src/test/regress/expected/type_sanity.out b/src/test/regress/expected/type_sanity.out
index 312d290..6281cef 100644
--- a/src/test/regress/expected/type_sanity.out
+++ b/src/test/regress/expected/type_sanity.out
@@ -67,11 +67,12 @@ WHERE p1.typtype not in ('c','d','p') AND p1.typname NOT LIKE E'\\_%'
(SELECT 1 FROM pg_type as p2
WHERE p2.typname = ('_' || p1.typname)::name AND
p2.typelem = p1.oid and p1.typarray = p2.oid);
- oid | typname
------+--------------
- 194 | pg_node_tree
- 210 | smgr
-(2 rows)
+ oid | typname
+------+--------------
+ 194 | pg_node_tree
+ 3353 | pg_ndistinct
+ 210 | smgr
+(3 rows)
-- Make sure typarray points to a varlena array type of our own base
SELECT p1.oid, p1.typname as basetype, p2.typname as arraytype,
diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule
index e9b2bad..0273ea6 100644
--- a/src/test/regress/parallel_schedule
+++ b/src/test/regress/parallel_schedule
@@ -116,3 +116,6 @@ test: event_trigger
# run stats by itself because its delay may be insufficient under heavy load
test: stats
+
+# run tests of multivariate stats
+test: mv_ndistinct
diff --git a/src/test/regress/serial_schedule b/src/test/regress/serial_schedule
index 7cdc0f6..f7f3a14 100644
--- a/src/test/regress/serial_schedule
+++ b/src/test/regress/serial_schedule
@@ -171,3 +171,4 @@ test: with
test: xml
test: event_trigger
test: stats
+test: mv_ndistinct
diff --git a/src/test/regress/sql/mv_ndistinct.sql b/src/test/regress/sql/mv_ndistinct.sql
new file mode 100644
index 0000000..5cef254
--- /dev/null
+++ b/src/test/regress/sql/mv_ndistinct.sql
@@ -0,0 +1,68 @@
+-- data type passed by value
+CREATE TABLE ndistinct (
+ a INT,
+ b INT,
+ c INT,
+ d INT
+);
+
+-- unknown column
+CREATE STATISTICS s10 ON (unknown_column) FROM ndistinct;
+
+-- single column
+CREATE STATISTICS s10 ON (a) FROM ndistinct;
+
+-- single column, duplicated
+CREATE STATISTICS s10 ON (a,a) FROM ndistinct;
+
+-- two columns, one duplicated
+CREATE STATISTICS s10 ON (a, a, b) FROM ndistinct;
+
+-- correct command
+CREATE STATISTICS s10 ON (a, b, c) FROM ndistinct;
+
+-- perfectly correlated groups
+INSERT INTO ndistinct
+ SELECT i/100, i/100, i/100 FROM generate_series(1,10000) s(i);
+
+ANALYZE ndistinct;
+
+SELECT ndist_enabled, ndist_built, standist
+ FROM pg_mv_statistic WHERE starelid = 'ndistinct'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b;
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c;
+
+EXPLAIN (COSTS off)
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c, d;
+
+TRUNCATE TABLE ndistinct;
+
+-- partially correlated groups
+INSERT INTO ndistinct
+ SELECT i/50, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE ndistinct;
+
+SELECT ndist_enabled, ndist_built, standist
+ FROM pg_mv_statistic WHERE starelid = 'ndistinct'::regclass;
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b;
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c;
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, b, c, d;
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY b, c, d;
+
+EXPLAIN
+ SELECT COUNT(*) FROM ndistinct GROUP BY a, d;
+
+DROP TABLE ndistinct;
diff --git a/src/test/regress/sql/object_address.sql b/src/test/regress/sql/object_address.sql
index e658ea3..791b942 100644
--- a/src/test/regress/sql/object_address.sql
+++ b/src/test/regress/sql/object_address.sql
@@ -41,6 +41,7 @@ CREATE TRANSFORM FOR int LANGUAGE SQL (
TO SQL WITH FUNCTION int4recv(internal));
CREATE PUBLICATION addr_pub FOR TABLE addr_nsp.gentable;
CREATE SUBSCRIPTION addr_sub CONNECTION '' PUBLICATION bar WITH (DISABLED, NOCREATE SLOT);
+CREATE STATISTICS addr_nsp.gentable_stat ON (a,b) FROM addr_nsp.gentable;
-- test some error cases
SELECT pg_get_object_address('stone', '{}', '{}');
@@ -179,7 +180,8 @@ WITH objects (type, name, args) AS (VALUES
('access method', '{btree}', '{}'),
('publication', '{addr_pub}', '{}'),
('publication relation', '{addr_nsp, gentable}', '{addr_pub}'),
- ('subscription', '{addr_sub}', '{}')
+ ('subscription', '{addr_sub}', '{}'),
+ ('statistics', '{addr_nsp, gentable_stat}', '{}')
)
SELECT (pg_identify_object(addr1.classid, addr1.objid, addr1.subobjid)).*,
-- test roundtrip through pg_identify_object_as_address
--
2.5.5
[binary/octet-stream] 0003-PATCH-functional-dependencies-only-the-ANALYZE-p-v23.patch (67.0K, ../../1fa982dc-ffb4-eef7-eb55-d8cf7d84f1c4@2ndquadrant.com/4-0003-PATCH-functional-dependencies-only-the-ANALYZE-p-v23.patch)
download | inline diff:
From e04f7a0b43dc914d5b661723e1a4a14abc1df4ef Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:36:25 +0200
Subject: [PATCH 3/9] PATCH: functional dependencies (only the ANALYZE part)
- implementation of soft functional dependencies (ANALYZE etc.)
- updates existing regression tests (new catalog etc.)
- new regression test for functional dependencies
- pg_ndistinct data type (varlena-based)
The algorithm detecting the dependencies is rather simple and probably
needs improvements, so that it detects more complicated dependencies,
and also validation of the math.
The patch introduces pg_dependencies, a new varlena data type for
storing serialized version of functional dependencies. This is similar
to what pg_ndistinct does for ndistinct coefficients.
---
doc/src/sgml/catalogs.sgml | 30 ++
doc/src/sgml/ref/create_statistics.sgml | 42 +-
src/backend/catalog/system_views.sql | 3 +-
src/backend/commands/statscmds.c | 37 +-
src/backend/nodes/copyfuncs.c | 1 +
src/backend/nodes/outfuncs.c | 2 +
src/backend/optimizer/util/plancat.c | 4 +-
src/backend/parser/gram.y | 14 +-
src/backend/utils/mvstats/Makefile | 2 +-
src/backend/utils/mvstats/README.dependencies | 118 +++++
src/backend/utils/mvstats/common.c | 26 +-
src/backend/utils/mvstats/dependencies.c | 622 ++++++++++++++++++++++++++
src/include/catalog/pg_cast.h | 4 +
src/include/catalog/pg_mv_statistic.h | 14 +-
src/include/catalog/pg_proc.h | 9 +
src/include/catalog/pg_type.h | 4 +
src/include/nodes/parsenodes.h | 1 +
src/include/nodes/relation.h | 2 +
src/include/utils/builtins.h | 4 +
src/include/utils/mvstats.h | 37 +-
src/test/regress/expected/mv_dependencies.out | 147 ++++++
src/test/regress/expected/mv_ndistinct.out | 10 +-
src/test/regress/expected/object_address.out | 2 +-
src/test/regress/expected/opr_sanity.out | 3 +-
src/test/regress/expected/rules.out | 3 +-
src/test/regress/expected/type_sanity.out | 7 +-
src/test/regress/parallel_schedule | 2 +-
src/test/regress/serial_schedule | 1 +
src/test/regress/sql/mv_dependencies.sql | 139 ++++++
src/test/regress/sql/mv_ndistinct.sql | 10 +-
src/test/regress/sql/object_address.sql | 2 +-
31 files changed, 1261 insertions(+), 41 deletions(-)
create mode 100644 src/backend/utils/mvstats/README.dependencies
create mode 100644 src/backend/utils/mvstats/dependencies.c
create mode 100644 src/test/regress/expected/mv_dependencies.out
create mode 100644 src/test/regress/sql/mv_dependencies.sql
diff --git a/doc/src/sgml/catalogs.sgml b/doc/src/sgml/catalogs.sgml
index 2a7bd6c..852f573 100644
--- a/doc/src/sgml/catalogs.sgml
+++ b/doc/src/sgml/catalogs.sgml
@@ -4285,6 +4285,17 @@
</row>
<row>
+ <entry><structfield>deps_enabled</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, functional dependencies will be computed for the combination of
+ columns, covered by the statistics. This does not mean the dependencies
+ are already computed, though.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>ndist_built</structfield></entry>
<entry><type>bool</type></entry>
<entry></entry>
@@ -4295,6 +4306,16 @@
</row>
<row>
+ <entry><structfield>deps_built</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, functional depenedencies are already computed and available for
+ use during query estimation.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>stakeys</structfield></entry>
<entry><type>int2vector</type></entry>
<entry><literal><link linkend="catalog-pg-attribute"><structname>pg_attribute</structname></link>.attnum</literal></entry>
@@ -4314,6 +4335,15 @@
</entry>
</row>
+ <row>
+ <entry><structfield>stadeps</structfield></entry>
+ <entry><type>pg_dependencies</type></entry>
+ <entry></entry>
+ <entry>
+ Functional dependencies, serialized as <structname>pg_dependencies</> type.
+ </entry>
+ </row>
+
</tbody>
</tgroup>
</table>
diff --git a/doc/src/sgml/ref/create_statistics.sgml b/doc/src/sgml/ref/create_statistics.sgml
index 9f6a65c..eaa39ee 100644
--- a/doc/src/sgml/ref/create_statistics.sgml
+++ b/doc/src/sgml/ref/create_statistics.sgml
@@ -21,8 +21,9 @@ PostgreSQL documentation
<refsynopsisdiv>
<synopsis>
-CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_name</replaceable> ON (
- <replaceable class="PARAMETER">column_name</replaceable>, <replaceable class="PARAMETER">column_name</replaceable> [, ...])
+CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_name</replaceable>
+ WITH ( <replaceable class="PARAMETER">option</replaceable> [= <replaceable class="PARAMETER">value</replaceable>] [, ... ] )
+ ON ( <replaceable class="PARAMETER">column_name</replaceable>, <replaceable class="PARAMETER">column_name</replaceable> [, ...])
FROM <replaceable class="PARAMETER">table_name</replaceable>
</synopsis>
@@ -99,6 +100,41 @@ CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_na
</variablelist>
+ <refsect2 id="SQL-CREATESTATISTICS-parameters">
+ <title id="SQL-CREATESTATISTICS-parameters-title">Parameters</title>
+
+ <indexterm zone="sql-createstatistics-parameters">
+ <primary>statistics parameters</primary>
+ </indexterm>
+
+ <para>
+ The <literal>WITH</> clause can specify <firstterm>options</>
+ for statistics. The currently available parameters are listed below.
+ </para>
+
+ <variablelist>
+
+ <varlistentry>
+ <term><literal>dependencies</> (<type>boolean</>)</term>
+ <listitem>
+ <para>
+ Enables functional dependencies for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><literal>ndistinct</> (<type>boolean</>)</term>
+ <listitem>
+ <para>
+ Enables ndistinct coefficients for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ </variablelist>
+
+ </refsect2>
</refsect1>
<refsect1 id="SQL-CREATESTATISTICS-examples">
@@ -119,7 +155,7 @@ CREATE TABLE t1 (
INSERT INTO t1 SELECT i/100, i/500
FROM generate_series(1,1000000) s(i);
-CREATE STATISTICS s1 ON (a, b) FROM t1;
+CREATE STATISTICS s1 WITH (dependencies) ON (a, b) FROM t1;
ANALYZE t1;
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 00ab440..216ece5 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -187,7 +187,8 @@ CREATE VIEW pg_mv_stats AS
C.relname AS tablename,
S.staname AS staname,
S.stakeys AS attnums,
- length(s.standist) AS ndistbytes
+ length(s.standist::bytea) AS ndistbytes,
+ length(S.stadeps::bytea) AS depsbytes
FROM (pg_mv_statistic S JOIN pg_class C ON (C.oid = S.starelid))
LEFT JOIN pg_namespace N ON (N.oid = C.relnamespace);
diff --git a/src/backend/commands/statscmds.c b/src/backend/commands/statscmds.c
index bde7e4b..af4f4d3 100644
--- a/src/backend/commands/statscmds.c
+++ b/src/backend/commands/statscmds.c
@@ -38,7 +38,9 @@ compare_int16(const void *a, const void *b)
}
/*
- * Implements the CREATE STATISTICS name ON (columns) FROM table
+ * Implements the CREATE STATISTICS command with syntax:
+ *
+ * CREATE STATISTICS name WITH (options) ON (columns) FROM table
*
* We do require that the types support sorting (ltopr), although some
* statistics might work with equality only.
@@ -66,6 +68,10 @@ CreateStatistics(CreateStatsStmt *stmt)
ObjectAddress parentobject,
childobject;
+ /* by default build nothing */
+ bool build_ndistinct = false,
+ build_dependencies = false;
+
Assert(IsA(stmt, CreateStatsStmt));
/* resolve the pieces of the name (namespace etc.) */
@@ -151,6 +157,31 @@ CreateStatistics(CreateStatsStmt *stmt)
(errcode(ERRCODE_UNDEFINED_COLUMN),
errmsg("duplicate column name in statistics definition")));
+ /*
+ * Parse the statistics options - currently only statistics types are
+ * recognized (ndistinct, dependencies).
+ */
+ foreach(l, stmt->options)
+ {
+ DefElem *opt = (DefElem *) lfirst(l);
+
+ if (strcmp(opt->defname, "ndistinct") == 0)
+ build_ndistinct = defGetBoolean(opt);
+ else if (strcmp(opt->defname, "dependencies") == 0)
+ build_dependencies = defGetBoolean(opt);
+ else
+ ereport(ERROR,
+ (errcode(ERRCODE_SYNTAX_ERROR),
+ errmsg("unrecognized STATISTICS option \"%s\"",
+ opt->defname)));
+ }
+
+ /* Make sure there's at least one statistics type specified. */
+ if (! (build_ndistinct || build_dependencies))
+ ereport(ERROR,
+ (errcode(ERRCODE_SYNTAX_ERROR),
+ errmsg("no statistics type (ndistinct, dependencies) requested")));
+
stakeys = buildint2vector(attnums, numcols);
/*
@@ -170,9 +201,11 @@ CreateStatistics(CreateStatsStmt *stmt)
values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(stakeys);
/* enabled statistics */
- values[Anum_pg_mv_statistic_ndist_enabled - 1] = BoolGetDatum(true);
+ values[Anum_pg_mv_statistic_ndist_enabled - 1] = BoolGetDatum(build_ndistinct);
+ values[Anum_pg_mv_statistic_deps_enabled - 1] = BoolGetDatum(build_dependencies);
nulls[Anum_pg_mv_statistic_standist - 1] = true;
+ nulls[Anum_pg_mv_statistic_stadeps - 1] = true;
/* insert the tuple into pg_mv_statistic */
mvstatrel = heap_open(MvStatisticRelationId, RowExclusiveLock);
diff --git a/src/backend/nodes/copyfuncs.c b/src/backend/nodes/copyfuncs.c
index dc42be0..6e465a7 100644
--- a/src/backend/nodes/copyfuncs.c
+++ b/src/backend/nodes/copyfuncs.c
@@ -4357,6 +4357,7 @@ _copyCreateStatsStmt(const CreateStatsStmt *from)
COPY_NODE_FIELD(defnames);
COPY_NODE_FIELD(relation);
COPY_NODE_FIELD(keys);
+ COPY_NODE_FIELD(options);
COPY_SCALAR_FIELD(if_not_exists);
return newnode;
diff --git a/src/backend/nodes/outfuncs.c b/src/backend/nodes/outfuncs.c
index 57cc0b4..c72473b 100644
--- a/src/backend/nodes/outfuncs.c
+++ b/src/backend/nodes/outfuncs.c
@@ -2202,9 +2202,11 @@ _outMVStatisticInfo(StringInfo str, const MVStatisticInfo *node)
/* enabled statistics */
WRITE_BOOL_FIELD(ndist_enabled);
+ WRITE_BOOL_FIELD(deps_enabled);
/* built/available statistics */
WRITE_BOOL_FIELD(ndist_built);
+ WRITE_BOOL_FIELD(deps_built);
}
static void
diff --git a/src/backend/optimizer/util/plancat.c b/src/backend/optimizer/util/plancat.c
index fc9ad93..8129143 100644
--- a/src/backend/optimizer/util/plancat.c
+++ b/src/backend/optimizer/util/plancat.c
@@ -1287,7 +1287,7 @@ get_relation_statistics(RelOptInfo *rel, Relation relation)
mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
/* unavailable stats are not interesting for the planner */
- if (mvstat->ndist_built)
+ if (mvstat->deps_built || mvstat->ndist_built)
{
info = makeNode(MVStatisticInfo);
@@ -1296,9 +1296,11 @@ get_relation_statistics(RelOptInfo *rel, Relation relation)
/* enabled statistics */
info->ndist_enabled = mvstat->ndist_enabled;
+ info->deps_enabled = mvstat->deps_enabled;
/* built/available statistics */
info->ndist_built = mvstat->ndist_built;
+ info->deps_built = mvstat->deps_built;
/* stakeys */
adatum = SysCacheGetAttr(MVSTATOID, htup,
diff --git a/src/backend/parser/gram.y b/src/backend/parser/gram.y
index 475a8a6..f61765f 100644
--- a/src/backend/parser/gram.y
+++ b/src/backend/parser/gram.y
@@ -3756,21 +3756,23 @@ ExistingIndex: USING INDEX index_name { $$ = $3; }
*****************************************************************************/
-CreateStatsStmt: CREATE STATISTICS any_name ON '(' columnList ')' FROM qualified_name
+CreateStatsStmt: CREATE STATISTICS any_name opt_reloptions ON '(' columnList ')' FROM qualified_name
{
CreateStatsStmt *n = makeNode(CreateStatsStmt);
n->defnames = $3;
- n->relation = $9;
- n->keys = $6;
+ n->relation = $10;
+ n->keys = $7;
+ n->options = $4;
n->if_not_exists = false;
$$ = (Node *)n;
}
- | CREATE STATISTICS IF_P NOT EXISTS any_name ON '(' columnList ')' FROM qualified_name
+ | CREATE STATISTICS IF_P NOT EXISTS any_name opt_reloptions ON '(' columnList ')' FROM qualified_name
{
CreateStatsStmt *n = makeNode(CreateStatsStmt);
n->defnames = $6;
- n->relation = $12;
- n->keys = $9;
+ n->relation = $13;
+ n->keys = $10;
+ n->options = $7;
n->if_not_exists = true;
$$ = (Node *)n;
}
diff --git a/src/backend/utils/mvstats/Makefile b/src/backend/utils/mvstats/Makefile
index 7295d46..21fe7e5 100644
--- a/src/backend/utils/mvstats/Makefile
+++ b/src/backend/utils/mvstats/Makefile
@@ -12,6 +12,6 @@ subdir = src/backend/utils/mvstats
top_builddir = ../../../..
include $(top_builddir)/src/Makefile.global
-OBJS = common.o mvdist.o
+OBJS = common.o dependencies.o mvdist.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/utils/mvstats/README.dependencies b/src/backend/utils/mvstats/README.dependencies
new file mode 100644
index 0000000..908f094
--- /dev/null
+++ b/src/backend/utils/mvstats/README.dependencies
@@ -0,0 +1,118 @@
+Soft functional dependencies
+============================
+
+Functional dependencies are a concept well described in relational theory,
+particularly in definition of normalization and "normal forms". Wikipedia
+has a nice definition of a functional dependency [1]:
+
+ In a given table, an attribute Y is said to have a functional dependency
+ on a set of attributes X (written X -> Y) if and only if each X value is
+ associated with precisely one Y value. For example, in an "Employee"
+ table that includes the attributes "Employee ID" and "Employee Date of
+ Birth", the functional dependency
+
+ {Employee ID} -> {Employee Date of Birth}
+
+ would hold. It follows from the previous two sentences that each
+ {Employee ID} is associated with precisely one {Employee Date of Birth}.
+
+ [1] https://en.wikipedia.org/wiki/Functional_dependency
+
+In practical terms, functional dependencies mean that a value in one column
+determines values in some other column. Consider for example this trivial
+table with two integer columns:
+
+ CREATE TABLE t (a INT, b INT)
+ AS SELECT i, i/10 FROM generate_series(1,100000) s(i);
+
+Clearly, knowledge of the value in column 'a' is sufficient to determine the
+value in column 'b', as it's simply (a/10). A more practical example may be
+addresses, where the knowledge of a ZIP code (usually) determines city. Larger
+cities may have multiple ZIP codes, so the dependency can't be reversed.
+
+Many datasets might be normalized not to contain such dependencies, but often
+it's not practical for various reasons. In some cases it's actually a conscious
+design choice to model the dataset in denormalized way, either because of
+performance or to make querying easier.
+
+
+soft dependencies
+-----------------
+
+Real-world data sets often contain data errors, either because of data entry
+mistakes (user mistyping the ZIP code) or perhaps issues in generating the
+data (e.g. a ZIP code mistakenly assigned to two cities in different states).
+
+A strict implementation would either ignore dependencies in such cases,
+rendering the approach mostly useless even for slightly noisy data sets, or
+result in sudden changes in behavior depending on minor differences between
+samples provided to ANALYZE.
+
+For this reason the statistics implementes "soft" functional dependencies,
+associating each functional dependency with a degree of validity (a number
+number between 0 and 1). This degree is then used to combine selectivities
+in a smooth manner.
+
+
+Mining dependencies (ANALYZE)
+-----------------------------
+
+The current algorithm is fairly simple - generate all possible functional
+dependencies, and for each one count the number of rows rows consistent it.
+Then use the fraction of rows (supporting/total) as the degree.
+
+To count the rows consistent with the dependency (a => b):
+
+ (a) Sort the data lexicographically, i.e. first by 'a' then 'b'.
+
+ (b) For each group of rows with the same 'a' value, count the number of
+ distinct values in 'b'.
+
+ (c) If there's a single distinct value in 'b', the rows are consistent with
+ the functional dependency. Otherwise they contradict it.
+
+The algorithm also requires a minimum size of the group to consider it
+consistent (currently 3 rows in the sample). Small groups make it less likely
+to break the consistency.
+
+
+Clause reduction (planner/optimizer)
+------------------------------------
+
+Apllying the functional dependencies is fairly simple - given a list of
+equality clauses, we compute selectivities of each clause and then use the
+degree to combine them using this formula
+
+ P(a=?,b=?) = P(a=?) * (d + (1-d) * P(b=?))
+
+Where 'd' is the degree of functional dependence (a=>b).
+
+With more than two equality clauses, this process happens recursively. For
+example for (a,b,c) we first use (a,b=>c) to break the computation into
+
+ P(a=?,b=?,c=?) = P(a=?,b=?) * (d + (1-d)*P(b=?))
+
+and then apply (a=>b) the same way on P(a=?,b=?).
+
+
+Consistecy of clauses
+---------------------
+
+Functional dependencies only express general dependencies between columns,
+without referencing particular values. This assumes that the equality clauses
+are in fact consistent with the functinal dependency, i.e. that given a
+dependency (a=>b), the value in (b=?) clause is the value determined by (a=?).
+If that's not the case, the clauses are "inconsistent" with the functional
+dependency and the result will be over-estimation.
+
+This may happen for example when using conditions on ZIP and city name with
+mismatching values (ZIP for a different city), etc. In such case the result
+set will be empty, but we'll estimate the selectivity using the ZIP condition.
+
+In this case the default estimation based on AVIA principle happens to work
+better, but mostly by chance.
+
+This issue is the price for the simplicity of functional dependencies. If the
+application frequently constructs queries with clauses inconsistent with
+functional dependencies present in the data, the best solution is not to
+use functional dependencies, but one of the more complex types of statistics.
diff --git a/src/backend/utils/mvstats/common.c b/src/backend/utils/mvstats/common.c
index 7d2f3f3..4b570a1 100644
--- a/src/backend/utils/mvstats/common.c
+++ b/src/backend/utils/mvstats/common.c
@@ -21,7 +21,8 @@ static VacAttrStats **lookup_var_attr_stats(int2vector *attrs,
static List *list_mv_stats(Oid relid);
-static void update_mv_stats(Oid relid, MVNDistinct ndistinct,
+static void update_mv_stats(Oid relid,
+ MVNDistinct ndistinct, MVDependencies dependencies,
int2vector *attrs, VacAttrStats **stats);
@@ -53,6 +54,7 @@ build_mv_stats(Relation onerel, double totalrows,
int j;
MVStatisticInfo *stat = (MVStatisticInfo *) lfirst(lc);
MVNDistinct ndistinct = NULL;
+ MVDependencies deps = NULL;
VacAttrStats **stats = NULL;
int numatts = 0;
@@ -89,8 +91,12 @@ build_mv_stats(Relation onerel, double totalrows,
if (stat->ndist_enabled)
ndistinct = build_mv_ndistinct(totalrows, numrows, rows, attrs, stats);
+ /* analyze functional dependencies between the columns */
+ if (stat->deps_enabled)
+ deps = build_mv_dependencies(numrows, rows, attrs, stats);
+
/* store the statistics in the catalog */
- update_mv_stats(stat->mvoid, ndistinct, attrs, stats);
+ update_mv_stats(stat->mvoid, ndistinct, deps, attrs, stats);
}
}
@@ -170,6 +176,8 @@ list_mv_stats(Oid relid)
info->stakeys = buildint2vector(stats->stakeys.values, stats->stakeys.dim1);
info->ndist_enabled = stats->ndist_enabled;
info->ndist_built = stats->ndist_built;
+ info->deps_enabled = stats->deps_enabled;
+ info->deps_built = stats->deps_built;
result = lappend(result, info);
}
@@ -191,7 +199,7 @@ list_mv_stats(Oid relid)
* Serializes the statistics and stores them into the pg_mv_statistic tuple.
*/
static void
-update_mv_stats(Oid mvoid, MVNDistinct ndistinct,
+update_mv_stats(Oid mvoid, MVNDistinct ndistinct, MVDependencies dependencies,
int2vector *attrs, VacAttrStats **stats)
{
HeapTuple stup,
@@ -218,18 +226,29 @@ update_mv_stats(Oid mvoid, MVNDistinct ndistinct,
values[Anum_pg_mv_statistic_standist-1] = PointerGetDatum(data);
}
+ if (dependencies != NULL)
+ {
+ nulls[Anum_pg_mv_statistic_stadeps - 1] = false;
+ values[Anum_pg_mv_statistic_stadeps - 1]
+ = PointerGetDatum(serialize_mv_dependencies(dependencies));
+ }
+
/* always replace the value (either by bytea or NULL) */
replaces[Anum_pg_mv_statistic_standist - 1] = true;
+ replaces[Anum_pg_mv_statistic_stadeps - 1] = true;
/* always change the availability flags */
nulls[Anum_pg_mv_statistic_ndist_built - 1] = false;
+ nulls[Anum_pg_mv_statistic_deps_built - 1] = false;
nulls[Anum_pg_mv_statistic_stakeys - 1] = false;
/* use the new attnums, in case we removed some dropped ones */
replaces[Anum_pg_mv_statistic_ndist_built - 1] = true;
+ replaces[Anum_pg_mv_statistic_deps_built - 1] = true;
replaces[Anum_pg_mv_statistic_stakeys - 1] = true;
values[Anum_pg_mv_statistic_ndist_built - 1] = BoolGetDatum(ndistinct != NULL);
+ values[Anum_pg_mv_statistic_deps_built - 1] = BoolGetDatum(dependencies != NULL);
values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(attrs);
@@ -370,6 +389,7 @@ multi_sort_compare_dim(int dim, const SortItem *a, const SortItem *b,
&mss->ssup[dim]);
}
+/* compare all the dimensions in a given range (inclusive) */
int
multi_sort_compare_dims(int start, int end,
const SortItem *a, const SortItem *b,
diff --git a/src/backend/utils/mvstats/dependencies.c b/src/backend/utils/mvstats/dependencies.c
new file mode 100644
index 0000000..c6390e2
--- /dev/null
+++ b/src/backend/utils/mvstats/dependencies.c
@@ -0,0 +1,622 @@
+/*-------------------------------------------------------------------------
+ *
+ * dependencies.c
+ * POSTGRES multivariate functional dependencies
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/dependencies.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "common.h"
+
+#include "utils/bytea.h"
+#include "utils/lsyscache.h"
+
+/*
+ * Internal state for DependencyGenerator of dependencies. Dependencies are similar to
+ * k-permutations of n elements, except that the order does not matter for the
+ * first (k-1) elements. That is, (a,b=>c) and (b,a=>c) are equivalent.
+ */
+typedef struct DependencyGeneratorData
+{
+ int k; /* size of the dependency */
+ int current; /* next dependency to return (index) */
+ int ndependencies; /* number of dependencies generated */
+ int *dependencies; /* array of pre-generated dependencies */
+} DependencyGeneratorData;
+
+typedef DependencyGeneratorData *DependencyGenerator;
+
+static void
+generate_dependencies_recurse(DependencyGenerator state,
+ int n, int index, int start, int *current)
+{
+ /*
+ * The generator handles the first (k-1) elements differently from
+ * the last element.
+ */
+ if (index < (state->k - 1))
+ {
+ int i;
+
+ /*
+ * The first (k-1) values have to be in ascending order, which we
+ * generate recursively.
+ */
+
+ for (i = start; i < n; i++)
+ {
+ current[index] = i;
+ generate_dependencies_recurse(state, n, (index+1), (i+1), current);
+ }
+ }
+ else
+ {
+ int i;
+
+ /*
+ * the last element is the implied value, which does not respect the
+ * ascending order. We just need to check that the value is not in the
+ * first (k-1) elements.
+ */
+
+ for (i = 0; i < n; i++)
+ {
+ int j;
+ bool match = false;
+
+ current[index] = i;
+
+ for (j = 0; j < index; j++)
+ {
+ if (current[j] == i)
+ {
+ match = true;
+ break;
+ }
+ }
+
+ /*
+ * If the value is not found in the first part of the dependency,
+ * we're done.
+ */
+ if (! match)
+ {
+ state->dependencies
+ = (int*)repalloc(state->dependencies,
+ state->k * (state->ndependencies + 1) * sizeof(int));
+ memcpy(&state->dependencies[(state->k * state->ndependencies)],
+ current, state->k * sizeof(int));
+ state->ndependencies++;
+ }
+ }
+ }
+}
+
+/* generate all dependencies (k-permutations of n elements) */
+static void
+generate_dependencies(DependencyGenerator state, int n)
+{
+ int *current = (int *) palloc0(sizeof(int) * state->k);
+
+ generate_dependencies_recurse(state, n, 0, 0, current);
+
+ pfree(current);
+}
+
+/*
+ * initialize the DependencyGenerator of variations, and prebuild the variations
+ *
+ * This pre-builds all the variations. We could also generate them in
+ * DependencyGenerator_next(), but this seems simpler.
+ */
+static DependencyGenerator
+DependencyGenerator_init(int2vector *attrs, int k)
+{
+ int n = attrs->dim1;
+ DependencyGenerator state;
+
+ Assert((n >= k) && (k > 0));
+
+ /* allocate the DependencyGenerator state as a single chunk of memory */
+ state = (DependencyGenerator) palloc0(sizeof(DependencyGeneratorData));
+ state->dependencies = (int*)palloc(k * sizeof(int));
+
+ state->ndependencies = 0;
+ state->current = 0;
+ state->k = k;
+
+ /* now actually pre-generate all the variations */
+ generate_dependencies(state, n);
+
+ return state;
+}
+
+/* free the DependencyGenerator state */
+static void
+DependencyGenerator_free(DependencyGenerator state)
+{
+ /* we've allocated a single chunk, so just free it */
+ pfree(state);
+}
+
+/* generate next combination */
+static int *
+DependencyGenerator_next(DependencyGenerator state, int2vector *attrs)
+{
+ if (state->current == state->ndependencies)
+ return NULL;
+
+ return &state->dependencies[state->k * state->current++];
+}
+
+
+/*
+ * validates functional dependency on the data
+ *
+ * An actual work horse of detecting functional dependencies. Given a variation
+ * of k attributes, it checks that the first (k-1) are sufficient to determine
+ * the last one.
+ */
+static double
+dependency_degree(int numrows, HeapTuple *rows, int k, int *dependency,
+ VacAttrStats **stats, int2vector *attrs)
+{
+ int i,
+ j;
+ int nvalues = numrows * k;
+ MultiSortSupport mss;
+ SortItem *items;
+ Datum *values;
+ bool *isnull;
+
+ /*
+ * XXX Maybe the threshold should be somehow related to the number of
+ * distinct values in the combination of columns we're analyzing. Assuming
+ * the distribution is uniform, we can estimate the average group size and
+ * use it as a threshold, similarly to what we do for MCV lists.
+ */
+ int min_group_size = 3;
+
+ /* counters valid within a group */
+ int group_size = 0;
+ int n_violations = 0;
+
+ /* total number of rows supporting (consistent with) the dependency */
+ int n_supporting_rows = 0;
+
+ /* Make sure we have at least two input attributes. */
+ Assert(k >= 2);
+
+ /* sort info for all attributes columns */
+ mss = multi_sort_init(k);
+
+ /* data for the sort */
+ items = (SortItem *) palloc0(numrows * sizeof(SortItem));
+ values = (Datum *) palloc0(sizeof(Datum) * nvalues);
+ isnull = (bool *) palloc0(sizeof(bool) * nvalues);
+
+ /* fix the pointers to values/isnull */
+ for (i = 0; i < numrows; i++)
+ {
+ items[i].values = &values[i * k];
+ items[i].isnull = &isnull[i * k];
+ }
+
+ /*
+ * Verify the dependency (a,b,...)->z, using a rather simple algorithm:
+ *
+ * (a) sort the data lexicographically
+ *
+ * (b) split the data into groups by first (k-1) columns
+ *
+ * (c) for each group count different values in the last column
+ */
+
+ /* prepare the sort function for the first dimension, and SortItem array */
+ for (i = 0; i < k; i++)
+ {
+ multi_sort_add_dimension(mss, i, dependency[i], stats);
+
+ /* accumulate all the data for both columns into an array and sort it */
+ for (j = 0; j < numrows; j++)
+ {
+ items[j].values[i]
+ = heap_getattr(rows[j], attrs->values[dependency[i]],
+ stats[i]->tupDesc, &items[j].isnull[i]);
+ }
+ }
+
+ /* sort the items so that we can detect the groups */
+ qsort_arg((void *) items, numrows, sizeof(SortItem),
+ multi_sort_compare, mss);
+
+ /*
+ * Walk through the sorted array, split it into rows according to the
+ * first (k-1) columns. If there's a single value in the last column, we
+ * count the group as 'supporting' the functional dependency. Otherwise we
+ * count it as contradicting.
+ *
+ * We also require a group to have a minimum number of rows to be
+ * considered useful for supporting the dependency. Contradicting groups
+ * may be of any size, though.
+ *
+ * XXX The minimum size requirement makes it impossible to identify case
+ * when both columns are unique (or nearly unique), and therefore
+ * trivially functionally dependent.
+ */
+
+ /* start with the first row forming a group */
+ group_size = 1;
+
+ for (i = 1; i <= numrows; i++)
+ {
+ /*
+ * Check if the group ended, which may be either because we processed
+ * all the items (i==numrows), or because the i-th item is not equal
+ * to the preceding one.
+ */
+ if ((i == numrows) ||
+ (multi_sort_compare_dims(0, (k - 2), &items[i - 1], &items[i], mss) != 0))
+ {
+ /*
+ * Do accounting for the preceding group, and reset counters.
+ *
+ * If there were no contradicting rows in the group, count the
+ * rows as supporting.
+ */
+ if ((n_violations == 0) && (group_size >= min_group_size))
+ n_supporting_rows += group_size;
+
+ /* current values start a new group */
+ n_violations = 0;
+ group_size = 0;
+ }
+ /* first colums match, but the last one does not (so contradicting) */
+ else if (multi_sort_compare_dim((k - 1), &items[i - 1], &items[i], mss) != 0)
+ n_violations += 1;
+
+ group_size += 1;
+ }
+
+ pfree(items);
+ pfree(values);
+ pfree(isnull);
+ pfree(mss);
+
+ /* Compute the 'degree of validity' as (supporting/total). */
+ return (n_supporting_rows * 1.0 / numrows);
+}
+
+/*
+ * detects functional dependencies between groups of columns
+ *
+ * Generates all possible subsets of columns (variations) and checks if the
+ * last one is determined by the preceding ones. For example given 3 columns,
+ * there are 12 variations (6 for variations on 2 columns, 6 for 3 columns):
+ *
+ * two columns three columns
+ * ----------- -------------
+ * (a) -> c (a,b) -> c
+ * (b) -> c (b,a) -> c
+ * (a) -> b (a,c) -> b
+ * (c) -> b (c,a) -> b
+ * (c) -> a (c,b) -> a
+ * (b) -> a (b,c) -> a
+ */
+MVDependencies
+build_mv_dependencies(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats)
+{
+ int i;
+ int k;
+ int numattrs = attrs->dim1;
+
+ /* result */
+ MVDependencies dependencies = NULL;
+
+ Assert(numattrs >= 2);
+
+ /*
+ * We'll try build functional dependencies starting from the smallest ones
+ * covering just 2 columns, to the largest ones, covering all columns
+ * included int the statistics. We start from the smallest ones because we
+ * want to be able to skip already implied ones.
+ */
+ for (k = 2; k <= numattrs; k++)
+ {
+ int *dependency; /* array with k elements */
+
+ /* prepare a DependencyGenerator of variation */
+ DependencyGenerator DependencyGenerator = DependencyGenerator_init(attrs, k);
+
+ /* generate all possible variations of k values (out of n) */
+ while ((dependency = DependencyGenerator_next(DependencyGenerator, attrs)))
+ {
+ double degree;
+ MVDependency d;
+
+ /* compute how valid the dependency seems */
+ degree = dependency_degree(numrows, rows, k, dependency, stats, attrs);
+
+ /* if the dependency seems entirely invalid, don't bother storing it */
+ if (degree == 0.0)
+ continue;
+
+ d = (MVDependency) palloc0(offsetof(MVDependencyData, attributes)
+ +k * sizeof(int));
+
+ /* copy the dependency (and keep the indexes into stakeys) */
+ d->degree = degree;
+ d->nattributes = k;
+ for (i = 0; i < k; i++)
+ d->attributes[i] = dependency[i];
+
+ /* initialize the list of dependencies */
+ if (dependencies == NULL)
+ {
+ dependencies
+ = (MVDependencies) palloc0(sizeof(MVDependenciesData));
+
+ dependencies->magic = MVSTAT_DEPS_MAGIC;
+ dependencies->type = MVSTAT_DEPS_TYPE_BASIC;
+ dependencies->ndeps = 0;
+ }
+
+ dependencies->ndeps++;
+ dependencies = (MVDependencies) repalloc(dependencies,
+ offsetof(MVDependenciesData, deps)
+ +dependencies->ndeps * sizeof(MVDependency));
+
+ dependencies->deps[dependencies->ndeps - 1] = d;
+ }
+
+ /* we're done with variations of k elements, so free the DependencyGenerator */
+ DependencyGenerator_free(DependencyGenerator);
+ }
+
+ return dependencies;
+}
+
+
+/*
+ * serialize list of dependencies into a bytea
+ */
+bytea *
+serialize_mv_dependencies(MVDependencies dependencies)
+{
+ int i;
+ bytea *output;
+ char *tmp;
+ Size len;
+
+ /* we need to store ndeps, with a number of attributes for each one */
+ len = VARHDRSZ + offsetof(MVDependenciesData, deps) +
+ dependencies->ndeps * offsetof(MVDependencyData, attributes);
+
+ /* and also include space for the actual attribute numbers and degrees */
+ for (i = 0; i < dependencies->ndeps; i++)
+ len += (sizeof(int16) * dependencies->deps[i]->nattributes);
+
+ output = (bytea *) palloc0(len);
+ SET_VARSIZE(output, len);
+
+ tmp = VARDATA(output);
+
+ /* first, store the number of dimensions / items */
+ memcpy(tmp, dependencies, offsetof(MVDependenciesData, deps));
+ tmp += offsetof(MVDependenciesData, deps);
+
+ /* store number of attributes and attribute numbers for each dependency */
+ for (i = 0; i < dependencies->ndeps; i++)
+ {
+ MVDependency d = dependencies->deps[i];
+
+ memcpy(tmp, d, offsetof(MVDependencyData, attributes));
+ tmp += offsetof(MVDependencyData, attributes);
+
+ memcpy(tmp, d->attributes, sizeof(int16) * d->nattributes);
+ tmp += sizeof(int16) * d->nattributes;
+
+ Assert(tmp <= ((char *) output + len));
+ }
+
+ return output;
+}
+
+/*
+ * Reads serialized dependencies into MVDependencies structure.
+ */
+MVDependencies
+deserialize_mv_dependencies(bytea *data)
+{
+ int i;
+ Size expected_size;
+ MVDependencies dependencies;
+ char *tmp;
+
+ if (data == NULL)
+ return NULL;
+
+ if (VARSIZE_ANY_EXHDR(data) < offsetof(MVDependenciesData, deps))
+ elog(ERROR, "invalid MVDependencies size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), offsetof(MVDependenciesData, deps));
+
+ /* read the MVDependencies header */
+ dependencies = (MVDependencies) palloc0(sizeof(MVDependenciesData));
+
+ /* initialize pointer to the data part (skip the varlena header) */
+ tmp = VARDATA_ANY(data);
+
+ /* get the header and perform basic sanity checks */
+ memcpy(dependencies, tmp, offsetof(MVDependenciesData, deps));
+ tmp += offsetof(MVDependenciesData, deps);
+
+ if (dependencies->magic != MVSTAT_DEPS_MAGIC)
+ elog(ERROR, "invalid dependency magic %d (expected %dd)",
+ dependencies->magic, MVSTAT_DEPS_MAGIC);
+
+ if (dependencies->type != MVSTAT_DEPS_TYPE_BASIC)
+ elog(ERROR, "invalid dependency type %d (expected %dd)",
+ dependencies->type, MVSTAT_DEPS_TYPE_BASIC);
+
+ Assert(dependencies->ndeps > 0);
+
+ /* what minimum bytea size do we expect for those parameters */
+ expected_size = offsetof(MVDependenciesData, deps) +
+ dependencies->ndeps * (offsetof(MVDependencyData, attributes) +
+ sizeof(int16) * 2);
+
+ if (VARSIZE_ANY_EXHDR(data) < expected_size)
+ elog(ERROR, "invalid dependencies size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ /* allocate space for the MCV items */
+ dependencies = repalloc(dependencies, offsetof(MVDependenciesData, deps)
+ +(dependencies->ndeps * sizeof(MVDependency)));
+
+ for (i = 0; i < dependencies->ndeps; i++)
+ {
+ double degree;
+ int k;
+ MVDependency d;
+
+ /* degree of validity */
+ memcpy(°ree, tmp, sizeof(double));
+ tmp += sizeof(double);
+
+ /* number of attributes */
+ memcpy(&k, tmp, sizeof(int));
+ tmp += sizeof(int);
+
+ /* is the number of attributes valid? */
+ Assert((k >= 2) && (k <= MVSTATS_MAX_DIMENSIONS));
+
+ /* now that we know the number of attributes, allocate the dependency */
+ d = (MVDependency) palloc0(offsetof(MVDependencyData, attributes) +
+ (k * sizeof(int)));
+
+ d->degree = degree;
+ d->nattributes = k;
+
+ /* copy attribute numbers */
+ memcpy(d->attributes, tmp, sizeof(int16) * d->nattributes);
+ tmp += sizeof(int16) * d->nattributes;
+
+ dependencies->deps[i] = d;
+
+ /* still within the bytea */
+ Assert(tmp <= ((char *) data + VARSIZE_ANY(data)));
+ }
+
+ /* we should have consumed the whole bytea exactly */
+ Assert(tmp == ((char *) data + VARSIZE_ANY(data)));
+
+ return dependencies;
+}
+
+/*
+ * pg_dependencies_in - input routine for type pg_dependencies.
+ *
+ * pg_dependencies is real enough to be a table column, but it has no operations
+ * of its own, and disallows input too
+ *
+ * XXX This is inspired by what pg_node_tree does.
+ */
+Datum
+pg_dependencies_in(PG_FUNCTION_ARGS)
+{
+ /*
+ * pg_node_list stores the data in binary form and parsing text input is
+ * not needed, so disallow this.
+ */
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_dependencies")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_dependencies - output routine for type pg_dependencies.
+ *
+ * histograms are serialized into a bytea value, so we simply call byteaout()
+ * to serialize the value into text. But it'd be nice to serialize that into
+ * a meaningful representation (e.g. for inspection by people).
+ */
+Datum
+pg_dependencies_out(PG_FUNCTION_ARGS)
+{
+ int i, j;
+ char *ret;
+ StringInfoData str;
+
+ bytea *data = PG_GETARG_BYTEA_PP(0);
+
+ MVDependencies dependencies = deserialize_mv_dependencies(data);
+
+ initStringInfo(&str);
+ appendStringInfoString(&str, "[");
+
+ for (i = 0; i < dependencies->ndeps; i++)
+ {
+ MVDependency dependency = dependencies->deps[i];
+
+ if (i > 0)
+ appendStringInfoString(&str, ", ");
+
+ appendStringInfoString(&str, "{");
+
+ for (j = 0; j < dependency->nattributes; j++)
+ {
+ if (j == dependency->nattributes-1)
+ appendStringInfoString(&str, " => ");
+ else if (j > 0)
+ appendStringInfoString(&str, ", ");
+
+ appendStringInfo(&str, "%d", dependency->attributes[j]);
+ }
+
+ appendStringInfo(&str, " : %f", dependency->degree);
+
+ appendStringInfoString(&str, "}");
+ }
+
+ appendStringInfoString(&str, "]");
+
+ ret = pstrdup(str.data);
+ pfree(str.data);
+
+ PG_RETURN_CSTRING(ret);
+}
+
+/*
+ * pg_dependencies_recv - binary input routine for type pg_dependencies.
+ */
+Datum
+pg_dependencies_recv(PG_FUNCTION_ARGS)
+{
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_dependencies")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_dependencies_send - binary output routine for type pg_dependencies.
+ *
+ * XXX Histograms are serialized into a bytea value, so let's just send that.
+ */
+Datum
+pg_dependencies_send(PG_FUNCTION_ARGS)
+{
+ return byteasend(fcinfo);
+}
diff --git a/src/include/catalog/pg_cast.h b/src/include/catalog/pg_cast.h
index bf39d43..22fa4b8 100644
--- a/src/include/catalog/pg_cast.h
+++ b/src/include/catalog/pg_cast.h
@@ -258,6 +258,10 @@ DATA(insert ( 194 25 0 i b ));
DATA(insert ( 3353 17 0 i b ));
DATA(insert ( 3353 25 0 i i ));
+/* pg_dependencies can be coerced to, but not from, bytea and text */
+DATA(insert ( 3358 17 0 i b ));
+DATA(insert ( 3358 25 0 i i ));
+
/*
* Datetime category
*/
diff --git a/src/include/catalog/pg_mv_statistic.h b/src/include/catalog/pg_mv_statistic.h
index fad80a3..e119cb7 100644
--- a/src/include/catalog/pg_mv_statistic.h
+++ b/src/include/catalog/pg_mv_statistic.h
@@ -38,9 +38,11 @@ CATALOG(pg_mv_statistic,3381)
/* statistics requested to build */
bool ndist_enabled; /* build ndist coefficient? */
+ bool deps_enabled; /* analyze dependencies? */
/* statistics that are available (if requested) */
bool ndist_built; /* ndistinct coeff built */
+ bool deps_built; /* dependencies were built */
/*
* variable-length fields start here, but we allow direct access to
@@ -50,6 +52,7 @@ CATALOG(pg_mv_statistic,3381)
#ifdef CATALOG_VARLEN
pg_ndistinct standist; /* ndistinct coeff (serialized) */
+ pg_dependencies stadeps; /* dependencies (serialized) */
#endif
} FormData_pg_mv_statistic;
@@ -65,14 +68,17 @@ typedef FormData_pg_mv_statistic *Form_pg_mv_statistic;
* compiler constants for pg_mv_statistic
* ----------------
*/
-#define Natts_pg_mv_statistic 8
+#define Natts_pg_mv_statistic 11
#define Anum_pg_mv_statistic_starelid 1
#define Anum_pg_mv_statistic_staname 2
#define Anum_pg_mv_statistic_stanamespace 3
#define Anum_pg_mv_statistic_staowner 4
#define Anum_pg_mv_statistic_ndist_enabled 5
-#define Anum_pg_mv_statistic_ndist_built 6
-#define Anum_pg_mv_statistic_stakeys 7
-#define Anum_pg_mv_statistic_standist 8
+#define Anum_pg_mv_statistic_deps_enabled 6
+#define Anum_pg_mv_statistic_ndist_built 7
+#define Anum_pg_mv_statistic_deps_built 8
+#define Anum_pg_mv_statistic_stakeys 9
+#define Anum_pg_mv_statistic_standist 10
+#define Anum_pg_mv_statistic_stadeps 11
#endif /* PG_MV_STATISTIC_H */
diff --git a/src/include/catalog/pg_proc.h b/src/include/catalog/pg_proc.h
index 940a991..b1f7b75 100644
--- a/src/include/catalog/pg_proc.h
+++ b/src/include/catalog/pg_proc.h
@@ -2735,6 +2735,15 @@ DESCR("I/O");
DATA(insert OID = 3357 ( pg_ndistinct_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "3353" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_send _null_ _null_ _null_ ));
DESCR("I/O");
+DATA(insert OID = 3359 ( pg_dependencies_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 3358 "2275" _null_ _null_ _null_ _null_ _null_ pg_dependencies_in _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3360 ( pg_dependencies_out PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 2275 "3358" _null_ _null_ _null_ _null_ _null_ pg_dependencies_out _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3361 ( pg_dependencies_recv PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 3358 "2281" _null_ _null_ _null_ _null_ _null_ pg_dependencies_recv _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 3362 ( pg_dependencies_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "3358" _null_ _null_ _null_ _null_ _null_ pg_dependencies_send _null_ _null_ _null_ ));
+DESCR("I/O");
+
DATA(insert OID = 1928 ( pg_stat_get_numscans PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_numscans _null_ _null_ _null_ ));
DESCR("statistics: number of scans done for table/index");
DATA(insert OID = 1929 ( pg_stat_get_tuples_returned PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_tuples_returned _null_ _null_ _null_ ));
diff --git a/src/include/catalog/pg_type.h b/src/include/catalog/pg_type.h
index 9c9caf3..da637d4 100644
--- a/src/include/catalog/pg_type.h
+++ b/src/include/catalog/pg_type.h
@@ -368,6 +368,10 @@ DATA(insert OID = 3353 ( pg_ndistinct PGNSP PGUID -1 f b S f t \054 0 0 0 pg_nd
DESCR("multivariate ndistinct coefficients");
#define PGNDISTINCTOID 3353
+DATA(insert OID = 3358 ( pg_dependencies PGNSP PGUID -1 f b S f t \054 0 0 0 pg_dependencies_in pg_dependencies_out pg_dependencies_recv pg_dependencies_send - - - i x f 0 -1 0 100 _null_ _null_ _null_ ));
+DESCR("multivariate histogram");
+#define PGDEPENDENCIESOID 3358
+
DATA(insert OID = 32 ( pg_ddl_command PGNSP PGUID SIZEOF_POINTER t p P f t \054 0 0 0 pg_ddl_command_in pg_ddl_command_out pg_ddl_command_recv pg_ddl_command_send - - - ALIGNOF_POINTER p f 0 -1 0 0 _null_ _null_ _null_ ));
DESCR("internal type for passing CollectedCommand");
#define PGDDLCOMMANDOID 32
diff --git a/src/include/nodes/parsenodes.h b/src/include/nodes/parsenodes.h
index 18e1dd1..fe4b93a 100644
--- a/src/include/nodes/parsenodes.h
+++ b/src/include/nodes/parsenodes.h
@@ -617,6 +617,7 @@ typedef struct CreateStatsStmt
List *defnames; /* qualified name (list of Value strings) */
RangeVar *relation; /* relation to build statistics on */
List *keys; /* String nodes naming referenced column(s) */
+ List *options; /* list of DefElem nodes */
bool if_not_exists; /* do nothing if statistics already exists */
} CreateStatsStmt;
diff --git a/src/include/nodes/relation.h b/src/include/nodes/relation.h
index 7a55151..56957e8 100644
--- a/src/include/nodes/relation.h
+++ b/src/include/nodes/relation.h
@@ -681,9 +681,11 @@ typedef struct MVStatisticInfo
RelOptInfo *rel; /* back-link to index's table */
/* enabled statistics */
+ bool deps_enabled; /* functional dependencies enabled */
bool ndist_enabled; /* ndistinct coefficient enabled */
/* built/available statistics */
+ bool deps_built; /* functional dependencies built */
bool ndist_built; /* ndistinct coefficient built */
/* columns in the statistics (attnums) */
diff --git a/src/include/utils/builtins.h b/src/include/utils/builtins.h
index 262ee94..9ffd80c 100644
--- a/src/include/utils/builtins.h
+++ b/src/include/utils/builtins.h
@@ -73,6 +73,10 @@ extern Datum pg_ndistinct_in(PG_FUNCTION_ARGS);
extern Datum pg_ndistinct_out(PG_FUNCTION_ARGS);
extern Datum pg_ndistinct_recv(PG_FUNCTION_ARGS);
extern Datum pg_ndistinct_send(PG_FUNCTION_ARGS);
+extern Datum pg_dependencies_in(PG_FUNCTION_ARGS);
+extern Datum pg_dependencies_out(PG_FUNCTION_ARGS);
+extern Datum pg_dependencies_recv(PG_FUNCTION_ARGS);
+extern Datum pg_dependencies_send(PG_FUNCTION_ARGS);
/* regexp.c */
extern char *regexp_fixed_prefix(text *text_re, bool case_insensitive,
diff --git a/src/include/utils/mvstats.h b/src/include/utils/mvstats.h
index 0660c59..e5a49bf 100644
--- a/src/include/utils/mvstats.h
+++ b/src/include/utils/mvstats.h
@@ -39,16 +39,49 @@ typedef struct MVNDistinctData {
typedef MVNDistinctData *MVNDistinct;
+#define MVSTAT_DEPS_MAGIC 0xB4549A2C /* marks serialized bytea */
+#define MVSTAT_DEPS_TYPE_BASIC 1 /* basic dependencies type */
+
+/*
+ * Functional dependencies, tracking column-level relationships (values
+ * in one column determine values in another one).
+ */
+typedef struct MVDependencyData
+{
+ double degree; /* degree of validity (0-1) */
+ int nattributes; /* number of attributes */
+ int16 attributes[FLEXIBLE_ARRAY_MEMBER]; /* attribute numbers */
+} MVDependencyData;
+
+typedef MVDependencyData *MVDependency;
+
+typedef struct MVDependenciesData
+{
+ uint32 magic; /* magic constant marker */
+ uint32 type; /* type of MV Dependencies (BASIC) */
+ int32 ndeps; /* number of dependencies */
+ MVDependency deps[FLEXIBLE_ARRAY_MEMBER]; /* dependencies */
+} MVDependenciesData;
+
+typedef MVDependenciesData *MVDependencies;
+
+
+
MVNDistinct load_mv_ndistinct(Oid mvoid);
bytea *serialize_mv_ndistinct(MVNDistinct ndistinct);
+bytea *serialize_mv_dependencies(MVDependencies dependencies);
/* deserialization of stats (serialization is private to analyze) */
MVNDistinct deserialize_mv_ndistinct(bytea *data);
-
+MVDependencies deserialize_mv_dependencies(bytea *data);
MVNDistinct build_mv_ndistinct(double totalrows, int numrows, HeapTuple *rows,
- int2vector *attrs, VacAttrStats **stats);
+ int2vector *attrs, VacAttrStats **stats);
+
+MVDependencies build_mv_dependencies(int numrows, HeapTuple *rows,
+ int2vector *attrs,
+ VacAttrStats **stats);
void build_mv_stats(Relation onerel, double totalrows,
int numrows, HeapTuple *rows,
diff --git a/src/test/regress/expected/mv_dependencies.out b/src/test/regress/expected/mv_dependencies.out
new file mode 100644
index 0000000..d442a16
--- /dev/null
+++ b/src/test/regress/expected/mv_dependencies.out
@@ -0,0 +1,147 @@
+-- data type passed by value
+CREATE TABLE functional_dependencies (
+ a INT,
+ b INT,
+ c INT
+);
+-- unknown column
+CREATE STATISTICS s1 WITH (dependencies) ON (unknown_column) FROM functional_dependencies;
+ERROR: column "unknown_column" referenced in statistics does not exist
+-- single column
+CREATE STATISTICS s1 WITH (dependencies) ON (a) FROM functional_dependencies;
+ERROR: statistics require at least 2 columns
+-- single column, duplicated
+CREATE STATISTICS s1 WITH (dependencies) ON (a,a) FROM functional_dependencies;
+ERROR: duplicate column name in statistics definition
+-- two columns, one duplicated
+CREATE STATISTICS s1 WITH (dependencies) ON (a, a, b) FROM functional_dependencies;
+ERROR: duplicate column name in statistics definition
+-- correct command
+CREATE STATISTICS s1 WITH (dependencies) ON (a, b, c) FROM functional_dependencies;
+-- random data (no functional dependencies)
+INSERT INTO functional_dependencies
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+---------
+ t | f |
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 0.999900}, {0 => 2 : 0.999900}, {1 => 2 : 0.999900}, {0, 1 => 2 : 0.999900}, {0, 2 => 1 : 0.999900}]
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 0.999900}, {0 => 2 : 0.999900}, {1 => 2 : 0.494900}, {0, 1 => 2 : 0.999900}, {0, 2 => 1 : 0.999900}]
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 1.000000}, {0 => 2 : 1.000000}, {1 => 2 : 1.000000}, {0, 1 => 2 : 1.000000}, {0, 2 => 1 : 1.000000}]
+(1 row)
+
+DROP TABLE functional_dependencies;
+-- varlena type (text)
+CREATE TABLE functional_dependencies (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+CREATE STATISTICS s2 WITH (dependencies) ON (a, b, c) FROM functional_dependencies;
+-- random data (no functional dependencies)
+INSERT INTO functional_dependencies
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+---------
+ t | f |
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 0.999900}, {0 => 2 : 0.999900}, {1 => 2 : 0.999900}, {0, 1 => 2 : 0.999900}, {0, 2 => 1 : 0.999900}]
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 0.999900}, {0 => 2 : 0.999900}, {1 => 2 : 0.494900}, {0, 1 => 2 : 0.999900}, {0, 2 => 1 : 0.999900}]
+(1 row)
+
+TRUNCATE functional_dependencies;
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------
+ t | t | [{0 => 1 : 1.000000}, {0 => 2 : 1.000000}, {1 => 2 : 1.000000}, {0, 1 => 2 : 1.000000}, {0, 2 => 1 : 1.000000}]
+(1 row)
+
+DROP TABLE functional_dependencies;
+-- NULL values (mix of int and text columns)
+CREATE TABLE functional_dependencies (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+CREATE STATISTICS s3 WITH (dependencies) ON (a, b, c, d) FROM functional_dependencies;
+INSERT INTO functional_dependencies
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+ deps_enabled | deps_built | stadeps
+--------------+------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
+ t | t | [{1 => 0 : 1.000000}, {2 => 0 : 1.000000}, {2 => 1 : 1.000000}, {3 => 0 : 1.000000}, {3 => 1 : 0.996700}, {0, 2 => 1 : 1.000000}, {0, 3 => 1 : 0.996700}, {1, 2 => 0 : 1.000000}, {1, 3 => 0 : 1.000000}, {2, 3 => 0 : 1.000000}, {2, 3 => 1 : 1.000000}, {0, 2, 3 => 1 : 1.000000}, {1, 2, 3 => 0 : 1.000000}]
+(1 row)
+
+DROP TABLE functional_dependencies;
diff --git a/src/test/regress/expected/mv_ndistinct.out b/src/test/regress/expected/mv_ndistinct.out
index 5f55091..06a7634 100644
--- a/src/test/regress/expected/mv_ndistinct.out
+++ b/src/test/regress/expected/mv_ndistinct.out
@@ -6,19 +6,19 @@ CREATE TABLE ndistinct (
d INT
);
-- unknown column
-CREATE STATISTICS s10 ON (unknown_column) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (unknown_column) FROM ndistinct;
ERROR: column "unknown_column" referenced in statistics does not exist
-- single column
-CREATE STATISTICS s10 ON (a) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a) FROM ndistinct;
ERROR: statistics require at least 2 columns
-- single column, duplicated
-CREATE STATISTICS s10 ON (a,a) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a,a) FROM ndistinct;
ERROR: duplicate column name in statistics definition
-- two columns, one duplicated
-CREATE STATISTICS s10 ON (a, a, b) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a, a, b) FROM ndistinct;
ERROR: duplicate column name in statistics definition
-- correct command
-CREATE STATISTICS s10 ON (a, b, c) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a, b, c) FROM ndistinct;
-- perfectly correlated groups
INSERT INTO ndistinct
SELECT i/100, i/100, i/100 FROM generate_series(1,10000) s(i);
diff --git a/src/test/regress/expected/object_address.out b/src/test/regress/expected/object_address.out
index 2b5c022..f574554 100644
--- a/src/test/regress/expected/object_address.out
+++ b/src/test/regress/expected/object_address.out
@@ -38,7 +38,7 @@ CREATE TRANSFORM FOR int LANGUAGE SQL (
TO SQL WITH FUNCTION int4recv(internal));
CREATE PUBLICATION addr_pub FOR TABLE addr_nsp.gentable;
CREATE SUBSCRIPTION addr_sub CONNECTION '' PUBLICATION bar WITH (DISABLED, NOCREATE SLOT);
-CREATE STATISTICS addr_nsp.gentable_stat ON (a,b) FROM addr_nsp.gentable;
+CREATE STATISTICS addr_nsp.gentable_stat WITH (ndistinct) ON (a,b) FROM addr_nsp.gentable;
-- test some error cases
SELECT pg_get_object_address('stone', '{}', '{}');
ERROR: unrecognized object type "stone"
diff --git a/src/test/regress/expected/opr_sanity.out b/src/test/regress/expected/opr_sanity.out
index 9a26205..db1cf8a 100644
--- a/src/test/regress/expected/opr_sanity.out
+++ b/src/test/regress/expected/opr_sanity.out
@@ -818,11 +818,12 @@ WHERE c.castmethod = 'b' AND
character varying | character | 0 | i
pg_node_tree | text | 0 | i
pg_ndistinct | bytea | 0 | i
+ pg_dependencies | bytea | 0 | i
cidr | inet | 0 | i
xml | text | 0 | a
xml | character varying | 0 | a
xml | character | 0 | a
-(8 rows)
+(9 rows)
-- **************** pg_conversion ****************
-- Look for illegal values in pg_conversion fields.
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 2c54779..39179a6 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1380,7 +1380,8 @@ pg_mv_stats| SELECT n.nspname AS schemaname,
c.relname AS tablename,
s.staname,
s.stakeys AS attnums,
- length((s.standist)::text) AS ndistbytes
+ length((s.standist)::bytea) AS ndistbytes,
+ length((s.stadeps)::bytea) AS depsbytes
FROM ((pg_mv_statistic s
JOIN pg_class c ON ((c.oid = s.starelid)))
LEFT JOIN pg_namespace n ON ((n.oid = c.relnamespace)));
diff --git a/src/test/regress/expected/type_sanity.out b/src/test/regress/expected/type_sanity.out
index 6281cef..b0b40ca 100644
--- a/src/test/regress/expected/type_sanity.out
+++ b/src/test/regress/expected/type_sanity.out
@@ -67,12 +67,13 @@ WHERE p1.typtype not in ('c','d','p') AND p1.typname NOT LIKE E'\\_%'
(SELECT 1 FROM pg_type as p2
WHERE p2.typname = ('_' || p1.typname)::name AND
p2.typelem = p1.oid and p1.typarray = p2.oid);
- oid | typname
-------+--------------
+ oid | typname
+------+-----------------
194 | pg_node_tree
3353 | pg_ndistinct
+ 3358 | pg_dependencies
210 | smgr
-(3 rows)
+(4 rows)
-- Make sure typarray points to a varlena array type of our own base
SELECT p1.oid, p1.typname as basetype, p2.typname as arraytype,
diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule
index 0273ea6..fda9166 100644
--- a/src/test/regress/parallel_schedule
+++ b/src/test/regress/parallel_schedule
@@ -118,4 +118,4 @@ test: event_trigger
test: stats
# run tests of multivariate stats
-test: mv_ndistinct
+test: mv_ndistinct mv_dependencies
diff --git a/src/test/regress/serial_schedule b/src/test/regress/serial_schedule
index f7f3a14..90d74d2 100644
--- a/src/test/regress/serial_schedule
+++ b/src/test/regress/serial_schedule
@@ -172,3 +172,4 @@ test: xml
test: event_trigger
test: stats
test: mv_ndistinct
+test: mv_dependencies
diff --git a/src/test/regress/sql/mv_dependencies.sql b/src/test/regress/sql/mv_dependencies.sql
new file mode 100644
index 0000000..43df798
--- /dev/null
+++ b/src/test/regress/sql/mv_dependencies.sql
@@ -0,0 +1,139 @@
+-- data type passed by value
+CREATE TABLE functional_dependencies (
+ a INT,
+ b INT,
+ c INT
+);
+
+-- unknown column
+CREATE STATISTICS s1 WITH (dependencies) ON (unknown_column) FROM functional_dependencies;
+
+-- single column
+CREATE STATISTICS s1 WITH (dependencies) ON (a) FROM functional_dependencies;
+
+-- single column, duplicated
+CREATE STATISTICS s1 WITH (dependencies) ON (a,a) FROM functional_dependencies;
+
+-- two columns, one duplicated
+CREATE STATISTICS s1 WITH (dependencies) ON (a, a, b) FROM functional_dependencies;
+
+-- correct command
+CREATE STATISTICS s1 WITH (dependencies) ON (a, b, c) FROM functional_dependencies;
+
+-- random data (no functional dependencies)
+INSERT INTO functional_dependencies
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+DROP TABLE functional_dependencies;
+
+-- varlena type (text)
+CREATE TABLE functional_dependencies (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+
+CREATE STATISTICS s2 WITH (dependencies) ON (a, b, c) FROM functional_dependencies;
+
+-- random data (no functional dependencies)
+INSERT INTO functional_dependencies
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c
+INSERT INTO functional_dependencies
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+TRUNCATE functional_dependencies;
+
+-- a => b, a => c, b => c
+INSERT INTO functional_dependencies
+ SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+DROP TABLE functional_dependencies;
+
+-- NULL values (mix of int and text columns)
+CREATE TABLE functional_dependencies (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+
+CREATE STATISTICS s3 WITH (dependencies) ON (a, b, c, d) FROM functional_dependencies;
+
+INSERT INTO functional_dependencies
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+
+ANALYZE functional_dependencies;
+
+SELECT deps_enabled, deps_built, stadeps
+ FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+
+DROP TABLE functional_dependencies;
diff --git a/src/test/regress/sql/mv_ndistinct.sql b/src/test/regress/sql/mv_ndistinct.sql
index 5cef254..43024ca 100644
--- a/src/test/regress/sql/mv_ndistinct.sql
+++ b/src/test/regress/sql/mv_ndistinct.sql
@@ -7,19 +7,19 @@ CREATE TABLE ndistinct (
);
-- unknown column
-CREATE STATISTICS s10 ON (unknown_column) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (unknown_column) FROM ndistinct;
-- single column
-CREATE STATISTICS s10 ON (a) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a) FROM ndistinct;
-- single column, duplicated
-CREATE STATISTICS s10 ON (a,a) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a,a) FROM ndistinct;
-- two columns, one duplicated
-CREATE STATISTICS s10 ON (a, a, b) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a, a, b) FROM ndistinct;
-- correct command
-CREATE STATISTICS s10 ON (a, b, c) FROM ndistinct;
+CREATE STATISTICS s10 WITH (ndistinct) ON (a, b, c) FROM ndistinct;
-- perfectly correlated groups
INSERT INTO ndistinct
diff --git a/src/test/regress/sql/object_address.sql b/src/test/regress/sql/object_address.sql
index 791b942..902599b 100644
--- a/src/test/regress/sql/object_address.sql
+++ b/src/test/regress/sql/object_address.sql
@@ -41,7 +41,7 @@ CREATE TRANSFORM FOR int LANGUAGE SQL (
TO SQL WITH FUNCTION int4recv(internal));
CREATE PUBLICATION addr_pub FOR TABLE addr_nsp.gentable;
CREATE SUBSCRIPTION addr_sub CONNECTION '' PUBLICATION bar WITH (DISABLED, NOCREATE SLOT);
-CREATE STATISTICS addr_nsp.gentable_stat ON (a,b) FROM addr_nsp.gentable;
+CREATE STATISTICS addr_nsp.gentable_stat WITH (ndistinct) ON (a,b) FROM addr_nsp.gentable;
-- test some error cases
SELECT pg_get_object_address('stone', '{}', '{}');
--
2.5.5
[binary/octet-stream] 0004-PATCH-selectivity-estimation-using-functional-de-v23.patch (46.5K, ../../1fa982dc-ffb4-eef7-eb55-d8cf7d84f1c4@2ndquadrant.com/5-0004-PATCH-selectivity-estimation-using-functional-de-v23.patch)
download | inline diff:
From a5badc43aa37d249c562a4605478bb7c897b76f6 Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:37:27 +0200
Subject: [PATCH 4/9] PATCH: selectivity estimation using functional
dependencies
Use functional dependencies to correct selectivity estimates of
equality clauses. For now this only works with regular WHERE
conditions, not join clauses etc.
Given two equality clauses
(a = 1) AND (b = 2)
we compute selectivity for each condition, and then combine them
using formula
P(a=1, b=2) = P(a=1) * [degree + (1 - degree) * P(b=2)]
where 'degree' of the functional dependence (a => b) is a number
between [0,1] measuring how much the knowledge of 'a' determines
the value of 'b'. For 'degree=0' this degrades to independence,
for 'degree=1' we get perfect functional dependency.
Estimates of more than two clauses are computed recursively, so
for example
(a = 1) AND (b = 2) AND (c = 3)
is first split into
P(a=1, b=2, c=3) = P(a=1, b=2) * [d + (1-d) * P(c=3)]
where 'd' is degree of (a,b => c) functional dependency. And then
the first part of the estimate is computed recursively:
P(a=1, b=2) = P(a=1) * [d + (1-d) * P(b=2)]
where 'd' is degree of (a => b) dependency.
The patch includes regression tests with functional dependencies
on several synthetic datasets (random, perfectly correlated, etc.)
---
doc/src/sgml/planstats.sgml | 178 +++++-
src/backend/optimizer/path/clausesel.c | 781 +++++++++++++++++++++++++-
src/backend/utils/mvstats/README.stats | 45 +-
src/backend/utils/mvstats/common.c | 1 +
src/backend/utils/mvstats/dependencies.c | 68 +++
src/include/utils/mvstats.h | 6 +-
src/test/regress/expected/mv_dependencies.out | 28 +-
src/test/regress/sql/mv_dependencies.sql | 19 +-
8 files changed, 1072 insertions(+), 54 deletions(-)
diff --git a/doc/src/sgml/planstats.sgml b/doc/src/sgml/planstats.sgml
index d5b975d..5436c8a 100644
--- a/doc/src/sgml/planstats.sgml
+++ b/doc/src/sgml/planstats.sgml
@@ -504,7 +504,7 @@ SELECT relpages, reltuples FROM pg_class WHERE relname = 't';
<programlisting>
EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1;
- QUERY PLAN
+ QUERY PLAN
-------------------------------------------------------------------------------------------------
Seq Scan on t (cost=0.00..170.00 rows=100 width=8) (actual time=0.031..2.870 rows=100 loops=1)
Filter: (a = 1)
@@ -527,7 +527,7 @@ EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1;
<programlisting>
EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
- QUERY PLAN
+ QUERY PLAN
-----------------------------------------------------------------------------------------------
Seq Scan on t (cost=0.00..195.00 rows=1 width=8) (actual time=0.033..3.006 rows=100 loops=1)
Filter: ((a = 1) AND (b = 1))
@@ -547,11 +547,11 @@ EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
<para>
Overestimates, i.e. errors in the opposite direction, are also possible.
Consider for example the following combination of range conditions, each
- matching
+ matching roughly half the rows.
<programlisting>
EXPLAIN ANALYZE SELECT * FROM t WHERE a <= 49 AND b > 49;
- QUERY PLAN
+ QUERY PLAN
------------------------------------------------------------------------------------------------
Seq Scan on t (cost=0.00..195.00 rows=2500 width=8) (actual time=1.607..1.607 rows=0 loops=1)
Filter: ((a <= 49) AND (b > 49))
@@ -587,6 +587,176 @@ EXPLAIN ANALYZE SELECT * FROM t WHERE a <= 49 AND b > 49;
sections.
</para>
+ <sect2 id="functional-dependencies">
+ <title>Functional Dependencies</title>
+
+ <para>
+ The simplest type of multivariate statistics are functional dependencies,
+ used in definitions of database normal forms. When simplified, saying that
+ <literal>b</> is functionally dependent on <literal>a</> means that
+ knowledge of value of <literal>a</> is sufficient to determine value of
+ <literal>b</>.
+ </para>
+
+ <para>
+ In normalized databases, only functional dependencies on primary keys
+ and super keys are allowed. In practice however many data sets are not
+ fully normalized, for example thanks to intentional denormalization for
+ performance reasons. The table <literal>t</> is an example of a data set
+ with functional dependencies. As <literal>a = b</> for all rows in the
+ table, <literal>a</> is functionally dependent on <literal>b</> and
+ <literal>b</> is functionally dependent on <literal>a</literal>.
+ </para>
+
+ <para>
+ Functional dependencies directly affect accuracy of the estimates, as
+ conditions on the dependent column(s) do not restrict the result set,
+ and are often redundant, causing underestimates. In the first example,
+ either <literal>a = 1</> or <literal>b = 1</> is sufficient (however see
+ <xref linkend="functional-dependencies-limitations">).
+ </para>
+
+ <para>
+ To inform the planner about the functional dependencies, or rather to
+ instruct it to search for them during <command>ANALYZE</>, we can use
+ the <command>CREATE STATISTICS</> command.
+
+<programlisting>
+CREATE STATISTICS s1 ON t (a,b) WITH (dependencies);
+ANALYZE t;
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
+ QUERY PLAN
+-------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=100 width=8) (actual time=0.095..3.118 rows=100 loops=1)
+ Filter: ((a = 1) AND (b = 1))
+ Rows Removed by Filter: 9900
+ Planning time: 0.367 ms
+ Execution time: 3.380 ms
+(5 rows)
+</programlisting>
+
+ As you can see, the estimate improved quite a bit, as the planner is now
+ aware of the functional dependencies and eliminates the second condition
+ when computing the estimates.
+ </para>
+
+ <para>
+ Let's inspect multivariate statistics on a table, as defined by
+ <command>CREATE STATISTICS</> and built by <command>ANALYZE</>. If you're
+ using <application>psql</>, the easiest way to list statistics on a table
+ is by using <command>\d</>.
+
+<programlisting>
+\d t
+ Table "public.t"
+ Column | Type | Modifiers
+--------+---------+-----------
+ a | integer |
+ b | integer |
+Statistics:
+ "public.s1" (dependencies) ON (a, b)
+</programlisting>
+
+ </para>
+
+ <para>
+ Similarly to per-column statistics, multivariate statistics are stored in
+ a system catalog called <structname>pg_mv_statistic</structname>, but
+ there is also a more convenient view <structname>pg_mv_stats</structname>.
+ To inspect the statistics <literal>s1</literal> defined above,
+ you may do this:
+
+<programlisting>
+SELECT tablename, staname, attnums, depsbytes, depsinfo
+ FROM pg_mv_stats WHERE staname = 's1';
+
+ tablename | staname | attnums | depsbytes | depsinfo
+-----------+---------+---------+-----------+----------------
+ t | s1 | 1 2 | 32 | dependencies=2
+(1 row)
+</programlisting>
+
+ This shows that the statistic is defined on table <structname>t</>,
+ <structfield>attnums</structfield> lists attribute numbers of columns
+ (references <structname>pg_attribute</structname>). It also shows
+ <command>ANALYZE</> found two functional dependencies, and size when
+ serialized into a <literal>bytea</> column. Inspecting the functional
+ dependencies is possible using <function>pg_mv_stats_dependencies_show</>
+ function.
+
+<programlisting>
+SELECT pg_mv_stats_dependencies_show(stadeps)
+ FROM pg_mv_statistic WHERE staname = 's1';
+
+ pg_mv_stats_dependencies_show
+-------------------------------
+ (1) => 2, (2) => 1
+(1 row)
+</programlisting>
+
+ Which confirms <literal>a</> is functionally dependent on <literal>b</> and
+ <literal>b</> is functionally dependent on <literal>a</literal>.
+ </para>
+
+ <para>
+ Now let's quickly discuss how this knowledge is applied when estimating
+ the selectivity. The planner walks through the conditions and attempts
+ to identify which conditions are already implied by other conditions,
+ and eliminates them (but only for the estimation, all conditions will be
+ checked on tuples during execution). In the example query, either of
+ the conditions may get eliminated, improving the estimate. This happens
+ in <function>clauselist_apply_dependencies</> in <filename>clausesel.c</>.
+ </para>
+
+ <sect3 id="functional-dependencies-limitations">
+ <title>Limitations of functional dependencies</title>
+
+ <para>
+ The first limitation of functional dependencies is that they only work
+ with simple equality conditions, comparing columns and constant values.
+ It's not possible to use them to eliminate equality conditions comparing
+ two columns or a column to an expression, range clauses, <literal>LIKE</>
+ or any other type of condition.
+ </para>
+
+ <para>
+ When eliminating the implied conditions, the planner assumes that the
+ conditions are compatible. Consider the following example, violating
+ this assumption:
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 10;
+ QUERY PLAN
+-----------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=100 width=8) (actual time=2.992..2.992 rows=0 loops=1)
+ Filter: ((a = 1) AND (b = 10))
+ Rows Removed by Filter: 10000
+ Planning time: 0.232 ms
+ Execution time: 3.033 ms
+(5 rows)
+</programlisting>
+
+ There are no rows with this combination of values, however the planner
+ is unable to verify whether the values match - it only knows that
+ the columns are functionally dependent.
+ </para>
+
+ <para>
+ This assumption is more about queries executed on the database - in many
+ cases it's actually satisfied (e.g. when the GUI only allows selecting
+ compatible values). But if that's not the case, functional dependencies
+ may not be a viable option.
+ </para>
+
+ <para>
+ For additional information about functional dependencies, see
+ <filename>src/backend/utils/mvstats/README.dependencies</>.
+ </para>
+
+ </sect3>
+
+ </sect2>
+
</sect1>
</chapter>
diff --git a/src/backend/optimizer/path/clausesel.c b/src/backend/optimizer/path/clausesel.c
index af2934a..cc79282 100644
--- a/src/backend/optimizer/path/clausesel.c
+++ b/src/backend/optimizer/path/clausesel.c
@@ -14,14 +14,19 @@
*/
#include "postgres.h"
+#include "access/sysattr.h"
+#include "catalog/pg_operator.h"
#include "nodes/makefuncs.h"
#include "optimizer/clauses.h"
#include "optimizer/cost.h"
#include "optimizer/pathnode.h"
#include "optimizer/plancat.h"
+#include "optimizer/var.h"
#include "utils/fmgroids.h"
#include "utils/lsyscache.h"
+#include "utils/mvstats.h"
#include "utils/selfuncs.h"
+#include "utils/typcache.h"
/*
@@ -41,6 +46,33 @@ typedef struct RangeQueryClause
static void addRangeClause(RangeQueryClause **rqlist, Node *clause,
bool varonleft, bool isLTsel, Selectivity s2);
+#define STATS_TYPE_FDEPS 0x01
+
+static bool clause_is_mv_compatible(Node *clause, Index relid, AttrNumber *attnum);
+
+static Bitmapset *collect_mv_attnums(List *clauses, Index relid);
+
+static int count_mv_attnums(List *clauses, Index relid);
+
+static int count_varnos(List *clauses, Index *relid);
+
+static MVStatisticInfo *choose_mv_statistics(List *mvstats, Bitmapset *attnums,
+ int types);
+
+static List *clauselist_mv_split(PlannerInfo *root, Index relid,
+ List *clauses, List **mvclauses,
+ MVStatisticInfo *mvstats, int types);
+
+static Selectivity clauselist_mv_selectivity_deps(PlannerInfo *root,
+ Index relid, List *clauses, MVStatisticInfo *mvstats,
+ Index varRelid, JoinType jointype, SpecialJoinInfo *sjinfo);
+
+static bool has_stats(List *stats, int type);
+
+static List *find_stats(PlannerInfo *root, Index relid);
+
+static bool stats_type_matches(MVStatisticInfo *stat, int type);
+
/****************************************************************************
* ROUTINES TO COMPUTE SELECTIVITIES
@@ -60,7 +92,19 @@ static void addRangeClause(RangeQueryClause **rqlist, Node *clause,
* subclauses. However, that's only right if the subclauses have independent
* probabilities, and in reality they are often NOT independent. So,
* we want to be smarter where we can.
-
+ *
+ * The first thing we try to do is applying multivariate statistics, in a way
+ * that intends to minimize the overhead when there are no multivariate stats
+ * on the relation. Thus we do several simple (and inexpensive) checks first,
+ * to verify that suitable multivariate statistics exist.
+ *
+ * If we identify such multivariate statistics apply, we try to apply them.
+ * Currently we only have (soft) functional dependencies, so we try to reduce
+ * the list of clauses.
+ *
+ * Then we remove the clauses estimated using multivariate stats, and process
+ * the rest of the clauses using the regular per-column stats.
+ *
* Currently, the only extra smarts we have is to recognize "range queries",
* such as "x > 34 AND x < 42". Clauses are recognized as possible range
* query components if they are restriction opclauses whose operators have
@@ -99,15 +143,81 @@ clauselist_selectivity(PlannerInfo *root,
RangeQueryClause *rqlist = NULL;
ListCell *l;
+ /* processing mv stats */
+ Oid relid = InvalidOid;
+
+ /* list of multivariate stats on the relation */
+ List *stats = NIL;
+
/*
- * If there's exactly one clause, then no use in trying to match up pairs,
- * so just go directly to clause_selectivity().
+ * If there's exactly one clause, then multivariate statistics is futile
+ * at this level (we might be able to apply them later if it's AND/OR
+ * clause). So just go directly to clause_selectivity().
*/
if (list_length(clauses) == 1)
return clause_selectivity(root, (Node *) linitial(clauses),
varRelid, jointype, sjinfo);
/*
+ * To fetch the statistics, we first need to determine the rel. Currently
+ * we only support estimates of simple restrictions referencing a single
+ * baserel (no join statistics). However set_baserel_size_estimates() sets
+ * varRelid=0 so we have to actually inspect the clauses by pull_varnos
+ * and see if there's just a single varno referenced.
+ *
+ * XXX Maybe there's a better way to find the relid?
+ */
+ if ((count_varnos(clauses, &relid) == 1) &&
+ ((varRelid == 0) || (varRelid == relid)))
+ stats = find_stats(root, relid);
+
+ /*
+ * Check that there are multivariate statistics usable for selectivity
+ * estimation, i.e. anything except ndistinct coefficients.
+ *
+ * Also check the number of attributes in clauses that might be estimated
+ * using those statistics, and that there are at least two such attributes.
+ * It may easily happen that we won't be able to estimate the clauses using
+ * the multivariate statistics anyway, but that requires a more expensive
+ * to verify (so the check check should be worth it).
+ *
+ * If there are no such stats or not enough attributes, don't waste time
+ * simply skip to estimation using the plain per-column stats.
+ */
+ if (has_stats(stats, STATS_TYPE_FDEPS) &&
+ (count_mv_attnums(clauses, relid) >= 2))
+ {
+ MVStatisticInfo *mvstat;
+ Bitmapset *mvattnums;
+
+ /* collect attributes from the compatible conditions */
+ mvattnums = collect_mv_attnums(clauses, relid);
+
+ /* and search for the statistic covering the most attributes */
+ mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_FDEPS);
+
+ if (mvstat != NULL) /* we have a matching stats */
+ {
+ /* clauses compatible with multi-variate stats */
+ List *mvclauses = NIL;
+
+ /* split the clauselist into regular and mv-clauses */
+ clauses = clauselist_mv_split(root, relid, clauses, &mvclauses,
+ mvstat, STATS_TYPE_FDEPS);
+
+ /* Empty list of clauses is a clear sign something went wrong. */
+ Assert(list_length(mvclauses));
+
+ /* we've chosen the histogram to match the clauses */
+ Assert(mvclauses != NIL);
+
+ /* compute the multivariate stats (dependencies) */
+ s1 *= clauselist_mv_selectivity_deps(root, relid, mvclauses, mvstat,
+ varRelid, jointype, sjinfo);
+ }
+ }
+
+ /*
* Initial scan over clauses. Anything that doesn't look like a potential
* rangequery clause gets multiplied into s1 and forgotten. Anything that
* does gets inserted into an rqlist entry.
@@ -763,3 +873,668 @@ clause_selectivity(PlannerInfo *root,
return s1;
}
+
+/*
+ * When applying functional dependencies, we start with the strongest ones
+ * strongest dependencies. That is, we select the dependency that:
+ *
+ * (a) has all attributes covered by the clauses
+ *
+ * (b) has the most attributes
+ *
+ * (c) has the higher degree of validity
+ *
+ * TODO Explain why we select the dependencies this way.
+ */
+static MVDependency
+find_strongest_dependency(MVStatisticInfo *mvstats, MVDependencies dependencies,
+ Bitmapset *attnums)
+{
+ int i;
+ MVDependency strongest = NULL;
+
+ /* number of attnums in clauses */
+ int nattnums = bms_num_members(attnums);
+
+ /*
+ * Iterate over the MVDependency items and find the strongest one from
+ * the fully-matched dependencies. We do the cheap checks first, before
+ * matching it against the attnums.
+ */
+ for (i = 0; i < dependencies->ndeps; i++)
+ {
+ MVDependency dependency = dependencies->deps[i];
+
+ /*
+ * Skip dependencies referencing more attributes than available clauses,
+ * as those can't be fully matched.
+ */
+ if (dependency->nattributes > nattnums)
+ continue;
+
+ /* We can skip dependencies on fewer attributes than the best one. */
+ if (strongest && (strongest->nattributes > dependency->nattributes))
+ continue;
+
+ /* And also weaker dependencies on the same number of attributes. */
+ if (strongest &&
+ (strongest->nattributes == dependency->nattributes) &&
+ (strongest->degree > dependency->degree))
+ continue;
+
+ /*
+ * Check that the dependency actually is fully covered by clauses.
+ * If the dependency is not fully matched by clauses, we can't use
+ * it for the estimation.
+ */
+ if (! dependency_is_fully_matched(dependency, attnums,
+ mvstats->stakeys->values))
+ continue;
+
+ /*
+ * We have a fully-matched dependency, and we already know it has to
+ * be stronger than the current one (otherwise we'd skip it before
+ * inspecting it at the very beginning.
+ */
+ strongest = dependency;
+ }
+
+ return strongest;
+}
+
+/*
+ * clauselist_mv_selectivity_deps
+ * estimate selectivity using functional dependencies
+ *
+ * Given equality clauses on attributes (a,b) we find the strongest dependency
+ * between them, i.e. either (a=>b) or (b=>a). Assuming (a=>b) is the selected
+ * dependency, we then combine the per-clause selectivities using the formula
+ *
+ * P(a,b) = P(a) * [f + (1-f)*P(b)]
+ *
+ * where 'f' is the degree of the dependency.
+ *
+ * With clauses on more than two attributes, the dependencies are applied
+ * recursively, starting with the widest/strongest dependencies. For example
+ * P(a,b,c) is first split like this:
+ *
+ * P(a,b,c) = P(a,b) * [f + (1-f)*P(c)]
+ *
+ * assuming (a,b=>c) is the strongest dependency.
+ */
+static Selectivity
+clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
+ List *clauses, MVStatisticInfo *mvstats,
+ Index varRelid, JoinType jointype,
+ SpecialJoinInfo *sjinfo)
+{
+ ListCell *lc;
+ Selectivity s1 = 1.0;
+ MVDependencies dependencies;
+
+ Assert(mvstats->deps_enabled && mvstats->deps_built);
+
+ /* load the dependency items stored in the statistics */
+ dependencies = load_mv_dependencies(mvstats->mvoid);
+
+ Assert(dependencies);
+
+ /*
+ * Apply the dependencies recursively, starting with the widest/strongest
+ * ones, and proceeding to the smaller/weaker ones. At the end of each
+ * round we factor in the selectivity of clauses on the implied attribute,
+ * and remove the clauses from the list.
+ */
+ while (true)
+ {
+ Selectivity s2 = 1.0;
+ Bitmapset *attnums;
+ MVDependency dependency;
+
+ /* clauses remaining after removing those on the "implied" attribute */
+ List *clauses_filtered = NIL;
+
+ attnums = collect_mv_attnums(clauses, relid);
+
+ /* no point in looking for dependencies with fewer than 2 attributes */
+ if (bms_num_members(attnums) < 2)
+ break;
+
+ /* the widest/strongest dependency, fully matched by clauses */
+ dependency = find_strongest_dependency(mvstats, dependencies, attnums);
+
+ /* if no suitable dependency was found, we're done */
+ if (! dependency)
+ break;
+
+ /*
+ * We found an applicable dependency, so find all the clauses on the
+ * implied attribute, so with dependency (a,b => c) we seach clauses
+ * on 'c'. We only really expect a single such clause, but in case
+ * there are more we simply multiply the selectivities as usual.
+ *
+ * XXX Maybe we should use the maximum, minimum or just error out?
+ */
+ foreach(lc, clauses)
+ {
+ AttrNumber attnum_clause = InvalidAttrNumber;
+ Node *clause = (Node *) lfirst(lc);
+
+ /*
+ * XXX We need the attnum referenced by the clause, and this is the
+ * easiest way to get it (but maybe not the best one). At this point
+ * we should only see equality clauses compatible with functional
+ * dependencies, so just error out if we stumble upon something else.
+ */
+ if (! clause_is_mv_compatible(clause, relid, &attnum_clause))
+ elog(ERROR, "clause not compatible with functional dependencies");
+
+ Assert(AttributeNumberIsValid(attnum_clause));
+
+ /*
+ * If the clause is not on the implied attribute, add it to the list
+ * of filtered clauses (for the next round) and continue with the
+ * next one.
+ */
+ if (! dependency_implies_attribute(dependency, attnum_clause,
+ mvstats->stakeys->values))
+ {
+ clauses_filtered = lappend(clauses_filtered, clause);
+ continue;
+ }
+
+ /*
+ * Otherwise compute selectivity of the clause, and multiply it with
+ * other clauses on the same attribute.
+ *
+ * XXX Not sure if we need to worry about multiple clauses, though.
+ * Those are all equality clauses, and if they reference different
+ * constants, that's not going to work.
+ */
+ s2 *= clause_selectivity(root, clause, varRelid, jointype, sjinfo);
+ }
+
+ /*
+ * Now factor in the selectivity for all the "implied" clauses into the
+ * final one, using this formula:
+ *
+ * P(a,b) = P(a) * (f + (1-f) * P(b))
+ *
+ * where 'f' is the degree of validity of the dependency.
+ */
+ s1 *= (dependency->degree + (1 - dependency->degree) * s2);
+
+ /* And only keep the filtered clauses for the next round. */
+ clauses = clauses_filtered;
+ }
+
+ /* And now simply multiply with selectivities of the remaining clauses. */
+ foreach (lc, clauses)
+ {
+ Node *clause = (Node *) lfirst(lc);
+
+ s1 *= clause_selectivity(root, clause, varRelid, jointype, sjinfo);
+ }
+
+ return s1;
+}
+
+/*
+ * Collect attributes from mv-compatible clauses.
+ */
+static Bitmapset *
+collect_mv_attnums(List *clauses, Index relid)
+{
+ Bitmapset *attnums = NULL;
+ ListCell *l;
+
+ /*
+ * Walk through the clauses and identify the ones we can estimate using
+ * multivariate stats, and remember the relid/columns. We'll then
+ * cross-check if we have suitable stats, and only if needed we'll split
+ * the clauses into multivariate and regular lists.
+ *
+ * For now we're only interested in RestrictInfo nodes with nested OpExpr,
+ * using either a range or equality.
+ */
+ foreach(l, clauses)
+ {
+ AttrNumber attnum;
+ Node *clause = (Node *) lfirst(l);
+
+ /* ignore the result for now - we only need the info */
+ if (clause_is_mv_compatible(clause, relid, &attnum))
+ attnums = bms_add_member(attnums, attnum);
+ }
+
+ /*
+ * If there are not at least two attributes referenced by the clause(s),
+ * we can throw everything out (as we'll revert to simple stats).
+ */
+ if (bms_num_members(attnums) <= 1)
+ {
+ if (attnums != NULL)
+ pfree(attnums);
+ attnums = NULL;
+ }
+
+ return attnums;
+}
+
+/*
+ * Count the number of attributes in clauses compatible with multivariate stats.
+ */
+static int
+count_mv_attnums(List *clauses, Index relid)
+{
+ int c;
+ Bitmapset *attnums = collect_mv_attnums(clauses, relid);
+
+ c = bms_num_members(attnums);
+
+ bms_free(attnums);
+
+ return c;
+}
+
+/*
+ * Count varnos referenced in the clauses, and if there's a single varno then
+ * return the index in 'relid'.
+ */
+static int
+count_varnos(List *clauses, Index *relid)
+{
+ int cnt;
+ Bitmapset *varnos = NULL;
+
+ varnos = pull_varnos((Node *) clauses);
+ cnt = bms_num_members(varnos);
+
+ /* if there's a single varno in the clauses, remember it */
+ if (bms_num_members(varnos) == 1)
+ *relid = bms_singleton_member(varnos);
+
+ bms_free(varnos);
+
+ return cnt;
+}
+
+static int
+count_attnums_covered_by_stats(MVStatisticInfo *info, Bitmapset *attnums)
+{
+ int i;
+ int matches = 0;
+ int2vector *attrs = info->stakeys;
+
+ /* count columns covered by the statistics */
+ for (i = 0; i < attrs->dim1; i++)
+ if (bms_is_member(attrs->values[i], attnums))
+ matches++;
+
+ return matches;
+}
+
+/*
+ * We're looking for statistics matching at least 2 attributes, referenced in
+ * clauses compatible with multivariate statistics. The current selection
+ * criteria is very simple - we choose the statistics referencing the most
+ * attributes.
+ *
+ * If there are multiple statistics referencing the same number of columns
+ * (from the clauses), the one with less source columns (as listed in the
+ * ADD STATISTICS when creating the statistics) wins. Else the first one wins.
+ *
+ * This is a very simple criteria, and has several weaknesses:
+ *
+ * (a) does not consider the accuracy of the statistics
+ *
+ * If there are two histograms built on the same set of columns, but one
+ * has 100 buckets and the other one has 1000 buckets (thus likely
+ * providing better estimates), this is not currently considered.
+ *
+ * (b) does not consider the type of statistics
+ *
+ * If there are three statistics - one containing just a MCV list, another
+ * one with just a histogram and a third one with both, we treat them equally.
+ *
+ * (c) does not consider the number of clauses
+ *
+ * As explained, only the number of referenced attributes counts, so if
+ * there are multiple clauses on a single attribute, this still counts as
+ * a single attribute.
+ *
+ * (d) does not consider type of condition
+ *
+ * Some clauses may work better with some statistics - for example equality
+ * clauses probably work better with MCV lists than with histograms. But
+ * IS [NOT] NULL conditions may often work better with histograms (thanks
+ * to NULL-buckets).
+ *
+ * So for example with five WHERE conditions
+ *
+ * WHERE (a = 1) AND (b = 1) AND (c = 1) AND (d = 1) AND (e = 1)
+ *
+ * and statistics on (a,b), (a,b,e) and (a,b,c,d), the last one will be selected
+ * as it references the most columns.
+ *
+ * Once we have selected the multivariate statistics, we split the list of
+ * clauses into two parts - conditions that are compatible with the selected
+ * stats, and conditions are estimated using simple statistics.
+ *
+ * From the example above, conditions
+ *
+ * (a = 1) AND (b = 1) AND (c = 1) AND (d = 1)
+ *
+ * will be estimated using the multivariate statistics (a,b,c,d) while the last
+ * condition (e = 1) will get estimated using the regular ones.
+ *
+ * There are various alternative selection criteria (e.g. counting conditions
+ * instead of just referenced attributes), but eventually the best option should
+ * be to combine multiple statistics. But that's much harder to do correctly.
+ *
+ * TODO: Select multiple statistics and combine them when computing the estimate.
+ *
+ * TODO: This will probably have to consider compatibility of clauses, because
+ * 'dependencies' will probably work only with equality clauses.
+ */
+static MVStatisticInfo *
+choose_mv_statistics(List *stats, Bitmapset *attnums, int types)
+{
+ ListCell *lc;
+
+ MVStatisticInfo *choice = NULL;
+
+ int current_matches = 2; /* goal #1: maximize */
+ int current_dims = (MVSTATS_MAX_DIMENSIONS + 1); /* goal #2: minimize */
+
+ /*
+ * Walk through the statistics (simple array with nmvstats elements) and
+ * for each one count the referenced attributes (encoded in the 'attnums'
+ * bitmap).
+ */
+ foreach(lc, stats)
+ {
+ MVStatisticInfo *info = (MVStatisticInfo *) lfirst(lc);
+
+ /* columns matching this statistics */
+ int matches = 0;
+
+ /* size (number of dimensions) of this statistics */
+ int numattrs = info->stakeys->dim1;
+
+ /* skip statistics not matching any of the requested types */
+ if (! (info->deps_built && (STATS_TYPE_FDEPS & types)))
+ continue;
+
+ /* count columns covered by the statistics */
+ matches = count_attnums_covered_by_stats(info, attnums);
+
+ /*
+ * Use this statistics when it increases the number of matched clauses
+ * or when it matches the same number of attributes but is smaller
+ * (in terms of number of attributes covered).
+ */
+ if ((matches > current_matches) ||
+ ((matches == current_matches) && (current_dims > numattrs)))
+ {
+ choice = info;
+ current_matches = matches;
+ current_dims = numattrs;
+ }
+ }
+
+ return choice;
+}
+
+
+/*
+ * clauselist_mv_split
+ * split the clause list into a part to be estimated using the provided
+ * statistics, and remaining clauses (estimated in some other way)
+ */
+static List *
+clauselist_mv_split(PlannerInfo *root, Index relid,
+ List *clauses, List **mvclauses,
+ MVStatisticInfo *mvstats, int types)
+{
+ int i;
+ ListCell *l;
+ List *non_mvclauses = NIL;
+
+ /* FIXME is there a better way to get info on int2vector? */
+ int2vector *attrs = mvstats->stakeys;
+ int numattrs = mvstats->stakeys->dim1;
+
+ Bitmapset *mvattnums = NULL;
+
+ /* build bitmap of attributes, so we can do bms_is_subset later */
+ for (i = 0; i < numattrs; i++)
+ mvattnums = bms_add_member(mvattnums, attrs->values[i]);
+
+ /* erase the list of mv-compatible clauses */
+ *mvclauses = NIL;
+
+ foreach(l, clauses)
+ {
+ bool match = false; /* by default not mv-compatible */
+ AttrNumber attnum = InvalidAttrNumber;
+ Node *clause = (Node *) lfirst(l);
+
+ if (clause_is_mv_compatible(clause, relid, &attnum))
+ {
+ /* are all the attributes part of the selected stats? */
+ if (bms_is_member(attnum, mvattnums))
+ match = true;
+ }
+
+ /*
+ * The clause matches the selected stats, so put it to the list of
+ * mv-compatible clauses. Otherwise, keep it in the list of 'regular'
+ * clauses (that may be selected later).
+ */
+ if (match)
+ *mvclauses = lappend(*mvclauses, clause);
+ else
+ non_mvclauses = lappend(non_mvclauses, clause);
+ }
+
+ /*
+ * Perform regular estimation using the clauses incompatible with the
+ * chosen histogram (or MV stats in general).
+ */
+ return non_mvclauses;
+
+}
+
+typedef struct
+{
+ Index varno; /* relid we're interested in */
+ Bitmapset *varattnos; /* attnums referenced by the clauses */
+} mv_compatible_context;
+
+/*
+ * Recursive walker that checks compatibility of the clause with multivariate
+ * statistics, and collects attnums from the Vars.
+ *
+ * XXX The original idea was to combine this with expression_tree_walker, but
+ * I've been unable to make that work - seems that does not quite allow
+ * checking the structure. Hence the explicit calls to the walker.
+ */
+static bool
+mv_compatible_walker(Node *node, mv_compatible_context *context)
+{
+ if (node == NULL)
+ return false;
+
+ if (IsA(node, RestrictInfo))
+ {
+ RestrictInfo *rinfo = (RestrictInfo *) node;
+
+ /* Pseudoconstants are not really interesting here. */
+ if (rinfo->pseudoconstant)
+ return true;
+
+ /* clauses referencing multiple varnos are incompatible */
+ if (bms_membership(rinfo->clause_relids) != BMS_SINGLETON)
+ return true;
+
+ /* check the clause inside the RestrictInfo */
+ return mv_compatible_walker((Node *) rinfo->clause, (void *) context);
+ }
+
+ if (IsA(node, Var))
+ {
+ Var *var = (Var *) node;
+
+ /*
+ * Also, the variable needs to reference the right relid (this might
+ * be unnecessary given the other checks, but let's be sure).
+ */
+ if (var->varno != context->varno)
+ return true;
+
+ /* Also skip system attributes (we don't allow stats on those). */
+ if (!AttrNumberIsForUserDefinedAttr(var->varattno))
+ return true;
+
+ /* Seems fine, so let's remember the attnum. */
+ context->varattnos = bms_add_member(context->varattnos, var->varattno);
+
+ return false;
+ }
+
+ /*
+ * And finally the operator expressions - we only allow simple expressions
+ * with two arguments, where one is a Var and the other is a constant, and
+ * it's a simple comparison (which we detect using estimator function).
+ */
+ if (is_opclause(node))
+ {
+ OpExpr *expr = (OpExpr *) node;
+ Var *var;
+ bool varonleft = true;
+ bool ok;
+
+ /*
+ * Only expressions with two arguments are considered compatible.
+ *
+ * XXX Possibly unnecessary (can OpExpr have different arg count?).
+ */
+ if (list_length(expr->args) != 2)
+ return true;
+
+ /* see if it actually has the right */
+ ok = (NumRelids((Node *) expr) == 1) &&
+ (is_pseudo_constant_clause(lsecond(expr->args)) ||
+ (varonleft = false,
+ is_pseudo_constant_clause(linitial(expr->args))));
+
+ /* unsupported structure (two variables or so) */
+ if (!ok)
+ return true;
+
+ /*
+ * If it's not a "<" or ">" or "=" operator, just ignore the clause.
+ * Otherwise note the relid and attnum for the variable. This uses the
+ * function for estimating selectivity, ont the operator directly (a
+ * bit awkward, but well ...).
+ */
+ switch (get_oprrest(expr->opno))
+ {
+ case F_EQSEL:
+
+ /* equality conditions are compatible with all statistics */
+ break;
+
+ default:
+
+ /* unknown estimator */
+ return true;
+ }
+
+ var = (varonleft) ? linitial(expr->args) : lsecond(expr->args);
+
+ return mv_compatible_walker((Node *) var, context);
+ }
+
+ /* Node not explicitly supported, so terminate */
+ return true;
+}
+
+/*
+ * Determines whether the clause is compatible with multivariate stats,
+ * and if it is, returns some additional information - varno (index
+ * into simple_rte_array) and a bitmap of attributes. This is then
+ * used to fetch related multivariate statistics.
+ *
+ * At this moment we only support basic conditions of the form
+ *
+ * variable OP constant
+ *
+ * where OP is one of [=,<,<=,>=,>] (which is however determined by
+ * looking at the associated function for estimating selectivity, just
+ * like with the single-dimensional case).
+ *
+ * TODO: Support 'OR clauses' - shouldn't be all that difficult to
+ * evaluate them using multivariate stats.
+ */
+static bool
+clause_is_mv_compatible(Node *clause, Index relid, AttrNumber *attnum)
+{
+ mv_compatible_context context;
+
+ context.varno = relid;
+ context.varattnos = NULL; /* no attnums */
+
+ if (mv_compatible_walker(clause, (void *) &context))
+ return false;
+
+ /* remember the newly collected attnums */
+ *attnum = bms_singleton_member(context.varattnos);
+
+ return true;
+}
+
+
+/*
+ * Check that the statistics matches at least one of the requested types.
+ */
+static bool
+stats_type_matches(MVStatisticInfo *stat, int type)
+{
+ if ((type & STATS_TYPE_FDEPS) && stat->deps_built)
+ return true;
+
+ return false;
+}
+
+/*
+ * Check that there are stats with at least one of the requested types.
+ */
+static bool
+has_stats(List *stats, int type)
+{
+ ListCell *s;
+
+ foreach(s, stats)
+ {
+ MVStatisticInfo *stat = (MVStatisticInfo *) lfirst(s);
+
+ /* terminate if we've found at least one matching statistics */
+ if (stats_type_matches(stat, type))
+ return true;
+ }
+
+ return false;
+}
+
+/*
+ * Lookups stats for the given baserel.
+ */
+static List *
+find_stats(PlannerInfo *root, Index relid)
+{
+ Assert(root->simple_rel_array[relid] != NULL);
+
+ return root->simple_rel_array[relid]->mvstatlist;
+}
diff --git a/src/backend/utils/mvstats/README.stats b/src/backend/utils/mvstats/README.stats
index 30d60d6..814f39c 100644
--- a/src/backend/utils/mvstats/README.stats
+++ b/src/backend/utils/mvstats/README.stats
@@ -8,48 +8,9 @@ not true, resulting in estimation errors.
Multivariate stats track different types of dependencies between the columns,
hopefully improving the estimates.
-
-Types of statistics
--------------------
-
-Currently we only have two kinds of multivariate statistics
-
- (a) soft functional dependencies (README.dependencies)
-
- (b) ndistinct coefficients
-
-
-Compatible clause types
------------------------
-
-Each type of statistics may be used to estimate some subset of clause types.
-
- (a) functional dependencies - equality clauses (AND), possibly IS NULL
-
-Currently only simple operator clauses (Var op Const) are supported, but it's
-possible to support more complex clause types, e.g. (Var op Var).
-
-
-Complex clauses
----------------
-
-We also support estimating more complex clauses - essentially AND/OR clauses
-with (Var op Const) as leaves, as long as all the referenced attributes are
-covered by a single statistics.
-
-For example this condition
-
- (a=1) AND ((b=2) OR ((c=3) AND (d=4)))
-
-may be estimated using statistics on (a,b,c,d). If we only have statistics on
-(b,c,d) we may estimate the second part, and estimate (a=1) using simple stats.
-
-If we only have statistics on (a,b,c) we can't apply it at all at this point,
-but it's worth pointing out clauselist_selectivity() works recursively and when
-handling the second part (the OR-clause), we'll be able to apply the statistics.
-
-Note: The multi-statistics estimation patch also makes it possible to pass some
-clauses as 'conditions' into the deeper parts of the expression tree.
+Currently we only have one kind of multivariate statistics - soft functional
+dependencies, and we use it to improve estimates of equality clauses. See
+README.dependencies for details.
Selectivity estimation
diff --git a/src/backend/utils/mvstats/common.c b/src/backend/utils/mvstats/common.c
index 4b570a1..39e3b92 100644
--- a/src/backend/utils/mvstats/common.c
+++ b/src/backend/utils/mvstats/common.c
@@ -308,6 +308,7 @@ compare_scalars_partition(const void *a, const void *b, void *arg)
return ApplySortComparator(da, false, db, false, ssup);
}
+
/* initialize multi-dimensional sort */
MultiSortSupport
multi_sort_init(int ndims)
diff --git a/src/backend/utils/mvstats/dependencies.c b/src/backend/utils/mvstats/dependencies.c
index c6390e2..6bca03b 100644
--- a/src/backend/utils/mvstats/dependencies.c
+++ b/src/backend/utils/mvstats/dependencies.c
@@ -310,6 +310,10 @@ dependency_degree(int numrows, HeapTuple *rows, int k, int *dependency,
* (c) -> b (c,a) -> b
* (c) -> a (c,b) -> a
* (b) -> a (b,c) -> a
+ *
+ * XXX Currently this builds redundant dependencies, becuse (a,b => c) and
+ * (b,a => c) is exactly the same thing, but both versions are generated
+ * and stored in the statistics.
*/
MVDependencies
build_mv_dependencies(int numrows, HeapTuple *rows, int2vector *attrs,
@@ -523,6 +527,70 @@ deserialize_mv_dependencies(bytea *data)
}
/*
+ * dependency_is_fully_matched
+ * checks that a functional dependency is fully matched given clauses on
+ * attributes (assuming the clauses are suitable equality clauses)
+ */
+bool
+dependency_is_fully_matched(MVDependency dependency, Bitmapset *attnums,
+ int16 *attmap)
+{
+ int j;
+
+ /*
+ * Check that the dependency actually is fully covered by clauses. We
+ * have to translate all attribute numbers, as those are referenced
+ */
+ for (j = 0; j < dependency->nattributes; j++)
+ {
+ int attnum = attmap[dependency->attributes[j]];
+
+ if (! bms_is_member(attnum, attnums))
+ return false;
+ }
+
+ return true;
+}
+
+/*
+ * dependency_implies_attribute
+ * check that the attnum matches is implied by the functional dependency
+ */
+bool
+dependency_implies_attribute(MVDependency dependency, AttrNumber attnum,
+ int16 *attmap)
+{
+ if (attnum == attmap[dependency->attributes[dependency->nattributes-1]])
+ return true;
+
+ return false;
+}
+
+MVDependencies
+load_mv_dependencies(Oid mvoid)
+{
+ bool isnull = false;
+ Datum deps;
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ HeapTuple htup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(mvoid));
+
+#ifdef USE_ASSERT_CHECKING
+ Form_pg_mv_statistic mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
+ Assert(mvstat->deps_enabled && mvstat->deps_built);
+#endif
+
+ deps = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_stadeps, &isnull);
+
+ Assert(!isnull);
+
+ ReleaseSysCache(htup);
+
+ return deserialize_mv_dependencies(DatumGetByteaP(deps));
+}
+
+/*
* pg_dependencies_in - input routine for type pg_dependencies.
*
* pg_dependencies is real enough to be a table column, but it has no operations
diff --git a/src/include/utils/mvstats.h b/src/include/utils/mvstats.h
index e5a49bf..b230747 100644
--- a/src/include/utils/mvstats.h
+++ b/src/include/utils/mvstats.h
@@ -65,9 +65,13 @@ typedef struct MVDependenciesData
typedef MVDependenciesData *MVDependencies;
-
+bool dependency_implies_attribute(MVDependency dependency, AttrNumber attnum,
+ int16 *attmap);
+bool dependency_is_fully_matched(MVDependency dependency, Bitmapset *attnums,
+ int16 *attmap);
MVNDistinct load_mv_ndistinct(Oid mvoid);
+MVDependencies load_mv_dependencies(Oid mvoid);
bytea *serialize_mv_ndistinct(MVNDistinct ndistinct);
bytea *serialize_mv_dependencies(MVDependencies dependencies);
diff --git a/src/test/regress/expected/mv_dependencies.out b/src/test/regress/expected/mv_dependencies.out
index d442a16..cf57a67 100644
--- a/src/test/regress/expected/mv_dependencies.out
+++ b/src/test/regress/expected/mv_dependencies.out
@@ -55,8 +55,10 @@ SELECT deps_enabled, deps_built, stadeps
TRUNCATE functional_dependencies;
-- a => b, a => c, b => c
+-- check explain (expect bitmap index scan, not plain index scan)
INSERT INTO functional_dependencies
- SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ SELECT mod(i,400), mod(i,200), mod(i,100) FROM generate_series(1,30000) s(i);
+CREATE INDEX fdeps_idx ON functional_dependencies (a, b);
ANALYZE functional_dependencies;
SELECT deps_enabled, deps_built, stadeps
FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
@@ -65,6 +67,16 @@ SELECT deps_enabled, deps_built, stadeps
t | t | [{0 => 1 : 1.000000}, {0 => 2 : 1.000000}, {1 => 2 : 1.000000}, {0, 1 => 2 : 1.000000}, {0, 2 => 1 : 1.000000}]
(1 row)
+EXPLAIN (COSTS off)
+ SELECT * FROM functional_dependencies WHERE a = 10 AND b = 5;
+ QUERY PLAN
+---------------------------------------------
+ Bitmap Heap Scan on functional_dependencies
+ Recheck Cond: ((a = 10) AND (b = 5))
+ -> Bitmap Index Scan on fdeps_idx
+ Index Cond: ((a = 10) AND (b = 5))
+(4 rows)
+
DROP TABLE functional_dependencies;
-- varlena type (text)
CREATE TABLE functional_dependencies (
@@ -110,8 +122,10 @@ SELECT deps_enabled, deps_built, stadeps
TRUNCATE functional_dependencies;
-- a => b, a => c, b => c
+-- check explain (expect bitmap index scan, not plain index scan)
INSERT INTO functional_dependencies
- SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ SELECT mod(i,400), mod(i,200), mod(i,100) FROM generate_series(1,30000) s(i);
+CREATE INDEX fdeps_idx ON functional_dependencies (a, b);
ANALYZE functional_dependencies;
SELECT deps_enabled, deps_built, stadeps
FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
@@ -120,6 +134,16 @@ SELECT deps_enabled, deps_built, stadeps
t | t | [{0 => 1 : 1.000000}, {0 => 2 : 1.000000}, {1 => 2 : 1.000000}, {0, 1 => 2 : 1.000000}, {0, 2 => 1 : 1.000000}]
(1 row)
+EXPLAIN (COSTS off)
+ SELECT * FROM functional_dependencies WHERE a = '10' AND b = '5';
+ QUERY PLAN
+------------------------------------------------------------
+ Bitmap Heap Scan on functional_dependencies
+ Recheck Cond: ((a = '10'::text) AND (b = '5'::text))
+ -> Bitmap Index Scan on fdeps_idx
+ Index Cond: ((a = '10'::text) AND (b = '5'::text))
+(4 rows)
+
DROP TABLE functional_dependencies;
-- NULL values (mix of int and text columns)
CREATE TABLE functional_dependencies (
diff --git a/src/test/regress/sql/mv_dependencies.sql b/src/test/regress/sql/mv_dependencies.sql
index 43df798..49db649 100644
--- a/src/test/regress/sql/mv_dependencies.sql
+++ b/src/test/regress/sql/mv_dependencies.sql
@@ -53,13 +53,20 @@ SELECT deps_enabled, deps_built, stadeps
TRUNCATE functional_dependencies;
-- a => b, a => c, b => c
+-- check explain (expect bitmap index scan, not plain index scan)
INSERT INTO functional_dependencies
- SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ SELECT mod(i,400), mod(i,200), mod(i,100) FROM generate_series(1,30000) s(i);
+
+CREATE INDEX fdeps_idx ON functional_dependencies (a, b);
+
ANALYZE functional_dependencies;
SELECT deps_enabled, deps_built, stadeps
FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+EXPLAIN (COSTS off)
+ SELECT * FROM functional_dependencies WHERE a = 10 AND b = 5;
+
DROP TABLE functional_dependencies;
-- varlena type (text)
@@ -96,6 +103,7 @@ TRUNCATE functional_dependencies;
-- a => b, a => c
INSERT INTO functional_dependencies
SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+
ANALYZE functional_dependencies;
SELECT deps_enabled, deps_built, stadeps
@@ -104,13 +112,20 @@ SELECT deps_enabled, deps_built, stadeps
TRUNCATE functional_dependencies;
-- a => b, a => c, b => c
+-- check explain (expect bitmap index scan, not plain index scan)
INSERT INTO functional_dependencies
- SELECT i/10000, i/20000, i/40000 FROM generate_series(1,1000000) s(i);
+ SELECT mod(i,400), mod(i,200), mod(i,100) FROM generate_series(1,30000) s(i);
+
+CREATE INDEX fdeps_idx ON functional_dependencies (a, b);
+
ANALYZE functional_dependencies;
SELECT deps_enabled, deps_built, stadeps
FROM pg_mv_statistic WHERE starelid = 'functional_dependencies'::regclass;
+EXPLAIN (COSTS off)
+ SELECT * FROM functional_dependencies WHERE a = '10' AND b = '5';
+
DROP TABLE functional_dependencies;
-- NULL values (mix of int and text columns)
--
2.5.5
[binary/octet-stream] 0005-PATCH-multivariate-MCV-lists-v23.patch (122.3K, ../../1fa982dc-ffb4-eef7-eb55-d8cf7d84f1c4@2ndquadrant.com/6-0005-PATCH-multivariate-MCV-lists-v23.patch)
download | inline diff:
From 699bc7bb78f1fc1a37225f2f69439ee39ee6adcf Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:38:02 +0200
Subject: [PATCH 5/9] PATCH: multivariate MCV lists
- extends the pg_mv_statistic catalog (add 'mcv' fields)
- building the MCV lists during ANALYZE
- simple estimation while planning the queries
- pg_mcv_list data type (varlena-based)
Includes regression tests, mostly equal to regression tests for
functional dependencies.
A varlena-based data type for storing serialized MCV lists.
---
doc/src/sgml/catalogs.sgml | 30 +
doc/src/sgml/planstats.sgml | 157 ++++
doc/src/sgml/ref/create_statistics.sgml | 34 +
src/backend/catalog/system_views.sql | 4 +-
src/backend/commands/statscmds.c | 11 +-
src/backend/nodes/outfuncs.c | 2 +
src/backend/optimizer/path/clausesel.c | 636 +++++++++++++++-
src/backend/optimizer/util/plancat.c | 4 +-
src/backend/utils/mvstats/Makefile | 2 +-
src/backend/utils/mvstats/README.mcv | 137 ++++
src/backend/utils/mvstats/README.stats | 87 ++-
src/backend/utils/mvstats/common.c | 136 +++-
src/backend/utils/mvstats/common.h | 22 +-
src/backend/utils/mvstats/mcv.c | 1184 +++++++++++++++++++++++++++++
src/bin/psql/describe.c | 24 +-
src/include/catalog/pg_cast.h | 5 +
src/include/catalog/pg_mv_statistic.h | 18 +-
src/include/catalog/pg_proc.h | 14 +
src/include/catalog/pg_type.h | 4 +
src/include/nodes/relation.h | 6 +-
src/include/utils/builtins.h | 4 +
src/include/utils/mvstats.h | 64 ++
src/test/regress/expected/mv_mcv.out | 198 +++++
src/test/regress/expected/opr_sanity.out | 3 +-
src/test/regress/expected/rules.out | 4 +-
src/test/regress/expected/type_sanity.out | 3 +-
src/test/regress/parallel_schedule | 2 +-
src/test/regress/serial_schedule | 1 +
src/test/regress/sql/mv_mcv.sql | 169 ++++
29 files changed, 2896 insertions(+), 69 deletions(-)
create mode 100644 src/backend/utils/mvstats/README.mcv
create mode 100644 src/backend/utils/mvstats/mcv.c
create mode 100644 src/test/regress/expected/mv_mcv.out
create mode 100644 src/test/regress/sql/mv_mcv.sql
diff --git a/doc/src/sgml/catalogs.sgml b/doc/src/sgml/catalogs.sgml
index 852f573..bca03e9 100644
--- a/doc/src/sgml/catalogs.sgml
+++ b/doc/src/sgml/catalogs.sgml
@@ -4296,6 +4296,17 @@
</row>
<row>
+ <entry><structfield>mcv_enabled</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, MVC list will be computed for the combination of columns,
+ covered by the statistics. This does not mean the MCV list is already
+ computed, though.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>ndist_built</structfield></entry>
<entry><type>bool</type></entry>
<entry></entry>
@@ -4316,6 +4327,16 @@
</row>
<row>
+ <entry><structfield>mcv_built</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, MCV list is already computed and available for use during query
+ estimation.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>stakeys</structfield></entry>
<entry><type>int2vector</type></entry>
<entry><literal><link linkend="catalog-pg-attribute"><structname>pg_attribute</structname></link>.attnum</literal></entry>
@@ -4344,6 +4365,15 @@
</entry>
</row>
+ <row>
+ <entry><structfield>stamcv</structfield></entry>
+ <entry><type>pg_mcv_list</type></entry>
+ <entry></entry>
+ <entry>
+ MCV list, serialized as <structname>pg_mcv_list</> type.
+ </entry>
+ </row>
+
</tbody>
</tgroup>
</table>
diff --git a/doc/src/sgml/planstats.sgml b/doc/src/sgml/planstats.sgml
index 5436c8a..57f9441 100644
--- a/doc/src/sgml/planstats.sgml
+++ b/doc/src/sgml/planstats.sgml
@@ -757,6 +757,163 @@ EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 10;
</sect2>
+ <sect2 id="mcv-lists">
+ <title>MCV lists</title>
+
+ <para>
+ As explained in the previous section, functional dependencies are very
+ cheap and efficient type of statistics, but it has limitations due to the
+ global nature (only tracking column-level dependencies, not between values
+ stored in the columns).
+ </para>
+
+ <para>
+ This section introduces multivariate most-common values (<acronym>MCV</>)
+ lists, a direct generalization of the statistics introduced in
+ <xref linkend="row-estimation-examples">, that is not subject to this
+ limitation. It is however more expensive, both in terms of storage and
+ planning time.
+ </para>
+
+ <para>
+ Let's look at the example query from the previous section again, creating
+ a multivariate <acronym>MCV</> list on the columns (after dropping the
+ functional dependencies, to make sure the planner uses the newly created
+ <acronym>MCV</> list when computing the estimates).
+
+<programlisting>
+DROP STATISTICS s1;
+CREATE STATISTICS s2 ON t (a,b) WITH (mcv);
+ANALYZE t;
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
+ QUERY PLAN
+-------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=100 width=8) (actual time=0.036..3.011 rows=100 loops=1)
+ Filter: ((a = 1) AND (b = 1))
+ Rows Removed by Filter: 9900
+ Planning time: 0.188 ms
+ Execution time: 3.229 ms
+(5 rows)
+</programlisting>
+
+ The estimate is as accurate as with the functional dependencies, mostly
+ thanks to the table being a fairly small and having a simple distribution
+ with low number of distinct values. Before looking at the second query,
+ which was not handled by functional dependencies this well, let's inspect
+ the <acronym>MCV</> list a bit.
+ </para>
+
+ <para>
+ First, let's list statistics defined on a table using <command>\d</>
+ in <application>psql</>:
+
+<programlisting>
+\d t
+ Table "public.t"
+ Column | Type | Modifiers
+--------+---------+-----------
+ a | integer |
+ b | integer |
+Statistics:
+ "public.s2" (mcv) ON (a, b)
+</programlisting>
+
+ </para>
+
+ <para>
+ To inspect details of the <acronym>MCV</> statistics, we can look into the
+ <structname>pg_mv_stats</structname> view
+
+<programlisting>
+SELECT tablename, staname, attnums, mcvbytes, mcvinfo
+ FROM pg_mv_stats WHERE staname = 's2';
+ tablename | staname | attnums | mcvbytes | mcvinfo
+-----------+---------+---------+----------+------------
+ t | s2 | 1 2 | 2048 | nitems=100
+(1 row)
+</programlisting>
+
+ According to this, the statistics has 2kB when serialized into
+ a <literal>bytea</> value, and <command>ANALYZE</> found 100 distinct
+ combinations of values in the two columns.
+ </para>
+
+ <para>
+ Inspecting the contents of the MCV list is possible using
+ <function>pg_mv_mcv_items</> function.
+
+<programlisting>
+SELECT * FROM pg_mv_mcv_items((SELECT oid FROM pg_mv_statistic WHERE staname = 's2'));
+ index | values | nulls | frequency
+-------+---------+-------+-----------
+ 0 | {0,0} | {f,f} | 0.01
+ 1 | {1,1} | {f,f} | 0.01
+ 2 | {2,2} | {f,f} | 0.01
+...
+ 49 | {49,49} | {f,f} | 0.01
+ 50 | {50,0} | {f,f} | 0.01
+...
+ 97 | {97,47} | {f,f} | 0.01
+ 98 | {98,48} | {f,f} | 0.01
+ 99 | {99,49} | {f,f} | 0.01
+(100 rows)
+</programlisting>
+
+ Which confirms there are 100 distinct combinations of values in the two
+ columns, and all of them are equally likely (1% frequency for each).
+ Had there been any null values in either of the columns, this would be
+ identified in the <structfield>nulls</> column.
+ </para>
+
+ <para>
+ When estimating the selectivity, the planner applies all the conditions
+ on items in the <acronym>MCV</> list, and them sums the frequencies
+ of the matching ones. See <function>clauselist_mv_selectivity_mcvlist</>
+ in <filename>clausesel.c</> for details.
+ </para>
+
+ <para>
+ Compared to functional dependencies, <acronym>MCV</> lists have two major
+ advantages. Firstly, the list stores actual values, making it possible to
+ detect "incompatible" combinations.
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 10;
+ QUERY PLAN
+---------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=1 width=8) (actual time=2.823..2.823 rows=0 loops=1)
+ Filter: ((a = 1) AND (b = 10))
+ Rows Removed by Filter: 10000
+ Planning time: 0.268 ms
+ Execution time: 2.866 ms
+(5 rows)
+</programlisting>
+
+ Secondly, <acronym>MCV</> also handle a wide range of clause types, not
+ just equality clauses like functional dependencies. See for example the
+ example range query, presented earlier:
+
+<programlisting>
+EXPLAIN ANALYZE SELECT * FROM t WHERE a <= 49 AND b > 49;
+ QUERY PLAN
+---------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=1 width=8) (actual time=3.349..3.349 rows=0 loops=1)
+ Filter: ((a <= 49) AND (b > 49))
+ Rows Removed by Filter: 10000
+ Planning time: 0.163 ms
+ Execution time: 3.389 ms
+(5 rows)
+</programlisting>
+
+ </para>
+
+ <para>
+ For additional information about multivariate MCV lists, see
+ <filename>src/backend/utils/mvstats/README.mcv</>.
+ </para>
+
+ </sect2>
+
</sect1>
</chapter>
diff --git a/doc/src/sgml/ref/create_statistics.sgml b/doc/src/sgml/ref/create_statistics.sgml
index eaa39ee..e95d8d3 100644
--- a/doc/src/sgml/ref/create_statistics.sgml
+++ b/doc/src/sgml/ref/create_statistics.sgml
@@ -124,6 +124,15 @@ CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_na
</varlistentry>
<varlistentry>
+ <term><literal>mcv</> (<type>boolean</>)</term>
+ <listitem>
+ <para>
+ Enables MCV list for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
<term><literal>ndistinct</> (<type>boolean</>)</term>
<listitem>
<para>
@@ -167,6 +176,31 @@ EXPLAIN ANALYZE SELECT * FROM t1 WHERE (a = 1) AND (b = 1);
</programlisting>
</para>
+ <para>
+ Create table <structname>t2</> with two perfectly correlated columns
+ (containing identical data), and a MCV list on those columns:
+
+<programlisting>
+CREATE TABLE t2 (
+ a int,
+ b int
+);
+
+INSERT INTO t2 SELECT mod(i,100), mod(i,100)
+ FROM generate_series(1,1000000) s(i);
+
+CREATE STATISTICS s2 WITH (mcv) ON (a, b) FROM t2;
+
+ANALYZE t2;
+
+-- valid combination (found in MCV)
+EXPLAIN ANALYZE SELECT * FROM t2 WHERE (a = 1) AND (b = 1);
+
+-- invalid combination (not found in MCV)
+EXPLAIN ANALYZE SELECT * FROM t2 WHERE (a = 1) AND (b = 2);
+</programlisting>
+ </para>
+
</refsect1>
<refsect1>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 216ece5..d4d9c24 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -188,7 +188,9 @@ CREATE VIEW pg_mv_stats AS
S.staname AS staname,
S.stakeys AS attnums,
length(s.standist::bytea) AS ndistbytes,
- length(S.stadeps::bytea) AS depsbytes
+ length(S.stadeps::bytea) AS depsbytes,
+ length(S.stamcv::bytea) AS mcvbytes,
+ pg_mv_stats_mcvlist_info(S.stamcv) AS mcvinfo
FROM (pg_mv_statistic S JOIN pg_class C ON (C.oid = S.starelid))
LEFT JOIN pg_namespace N ON (N.oid = C.relnamespace);
diff --git a/src/backend/commands/statscmds.c b/src/backend/commands/statscmds.c
index af4f4d3..ef05745 100644
--- a/src/backend/commands/statscmds.c
+++ b/src/backend/commands/statscmds.c
@@ -70,7 +70,8 @@ CreateStatistics(CreateStatsStmt *stmt)
/* by default build nothing */
bool build_ndistinct = false,
- build_dependencies = false;
+ build_dependencies = false,
+ build_mcv = false;
Assert(IsA(stmt, CreateStatsStmt));
@@ -169,6 +170,8 @@ CreateStatistics(CreateStatsStmt *stmt)
build_ndistinct = defGetBoolean(opt);
else if (strcmp(opt->defname, "dependencies") == 0)
build_dependencies = defGetBoolean(opt);
+ else if (strcmp(opt->defname, "mcv") == 0)
+ build_mcv = defGetBoolean(opt);
else
ereport(ERROR,
(errcode(ERRCODE_SYNTAX_ERROR),
@@ -177,10 +180,10 @@ CreateStatistics(CreateStatsStmt *stmt)
}
/* Make sure there's at least one statistics type specified. */
- if (! (build_ndistinct || build_dependencies))
+ if (!(build_ndistinct || build_dependencies || build_mcv))
ereport(ERROR,
(errcode(ERRCODE_SYNTAX_ERROR),
- errmsg("no statistics type (ndistinct, dependencies) requested")));
+ errmsg("no statistics type (ndistinct, dependencies, mcv) requested")));
stakeys = buildint2vector(attnums, numcols);
@@ -203,9 +206,11 @@ CreateStatistics(CreateStatsStmt *stmt)
/* enabled statistics */
values[Anum_pg_mv_statistic_ndist_enabled - 1] = BoolGetDatum(build_ndistinct);
values[Anum_pg_mv_statistic_deps_enabled - 1] = BoolGetDatum(build_dependencies);
+ values[Anum_pg_mv_statistic_mcv_enabled - 1] = BoolGetDatum(build_mcv);
nulls[Anum_pg_mv_statistic_standist - 1] = true;
nulls[Anum_pg_mv_statistic_stadeps - 1] = true;
+ nulls[Anum_pg_mv_statistic_stamcv - 1] = true;
/* insert the tuple into pg_mv_statistic */
mvstatrel = heap_open(MvStatisticRelationId, RowExclusiveLock);
diff --git a/src/backend/nodes/outfuncs.c b/src/backend/nodes/outfuncs.c
index c72473b..a9cc9ad 100644
--- a/src/backend/nodes/outfuncs.c
+++ b/src/backend/nodes/outfuncs.c
@@ -2203,10 +2203,12 @@ _outMVStatisticInfo(StringInfo str, const MVStatisticInfo *node)
/* enabled statistics */
WRITE_BOOL_FIELD(ndist_enabled);
WRITE_BOOL_FIELD(deps_enabled);
+ WRITE_BOOL_FIELD(mcv_enabled);
/* built/available statistics */
WRITE_BOOL_FIELD(ndist_built);
WRITE_BOOL_FIELD(deps_built);
+ WRITE_BOOL_FIELD(mcv_built);
}
static void
diff --git a/src/backend/optimizer/path/clausesel.c b/src/backend/optimizer/path/clausesel.c
index cc79282..abdbc5b 100644
--- a/src/backend/optimizer/path/clausesel.c
+++ b/src/backend/optimizer/path/clausesel.c
@@ -15,6 +15,7 @@
#include "postgres.h"
#include "access/sysattr.h"
+#include "catalog/pg_collation.h"
#include "catalog/pg_operator.h"
#include "nodes/makefuncs.h"
#include "optimizer/clauses.h"
@@ -47,12 +48,14 @@ static void addRangeClause(RangeQueryClause **rqlist, Node *clause,
bool varonleft, bool isLTsel, Selectivity s2);
#define STATS_TYPE_FDEPS 0x01
+#define STATS_TYPE_MCV 0x02
-static bool clause_is_mv_compatible(Node *clause, Index relid, AttrNumber *attnum);
+static bool clause_is_mv_compatible(Node *clause, Index relid, Bitmapset **attnums,
+ int type);
-static Bitmapset *collect_mv_attnums(List *clauses, Index relid);
+static Bitmapset *collect_mv_attnums(List *clauses, Index relid, int type);
-static int count_mv_attnums(List *clauses, Index relid);
+static int count_mv_attnums(List *clauses, Index relid, int type);
static int count_varnos(List *clauses, Index *relid);
@@ -63,10 +66,23 @@ static List *clauselist_mv_split(PlannerInfo *root, Index relid,
List *clauses, List **mvclauses,
MVStatisticInfo *mvstats, int types);
+static Selectivity clauselist_mv_selectivity(PlannerInfo *root,
+ List *clauses, MVStatisticInfo *mvstats);
+
static Selectivity clauselist_mv_selectivity_deps(PlannerInfo *root,
Index relid, List *clauses, MVStatisticInfo *mvstats,
Index varRelid, JoinType jointype, SpecialJoinInfo *sjinfo);
+static Selectivity clauselist_mv_selectivity_mcvlist(PlannerInfo *root,
+ List *clauses, MVStatisticInfo *mvstats,
+ bool *fullmatch, Selectivity *lowsel);
+
+static int update_match_bitmap_mcvlist(PlannerInfo *root, List *clauses,
+ int2vector *stakeys, MCVList mcvlist,
+ int nmatches, char *matches,
+ Selectivity *lowsel, bool *fullmatch,
+ bool is_or);
+
static bool has_stats(List *stats, int type);
static List *find_stats(PlannerInfo *root, Index relid);
@@ -74,6 +90,9 @@ static List *find_stats(PlannerInfo *root, Index relid);
static bool stats_type_matches(MVStatisticInfo *stat, int type);
+#define UPDATE_RESULT(m,r,isor) \
+ (m) = (isor) ? (Max(m,r)) : (Min(m,r))
+
/****************************************************************************
* ROUTINES TO COMPUTE SELECTIVITIES
****************************************************************************/
@@ -99,11 +118,13 @@ static bool stats_type_matches(MVStatisticInfo *stat, int type);
* to verify that suitable multivariate statistics exist.
*
* If we identify such multivariate statistics apply, we try to apply them.
- * Currently we only have (soft) functional dependencies, so we try to reduce
- * the list of clauses.
*
- * Then we remove the clauses estimated using multivariate stats, and process
- * the rest of the clauses using the regular per-column stats.
+ * First we try to reduce the list of clauses by applying (soft) functional
+ * dependencies, and then we try to estimate the selectivity of the reduced
+ * list of clauses using the multivariate MCV list.
+ *
+ * Finally we remove the portion of clauses estimated using multivariate stats,
+ * and process the rest of the clauses using the regular per-column stats.
*
* Currently, the only extra smarts we have is to recognize "range queries",
* such as "x > 34 AND x < 42". Clauses are recognized as possible range
@@ -173,7 +194,10 @@ clauselist_selectivity(PlannerInfo *root,
/*
* Check that there are multivariate statistics usable for selectivity
- * estimation, i.e. anything except ndistinct coefficients.
+ * estimation. We try to apply MCV lists first, because statistics
+ * tracking actual values tend to provide more reliable estimates than
+ * functional dependencies (which assume that the clauses are consistent
+ * with the statistics).
*
* Also check the number of attributes in clauses that might be estimated
* using those statistics, and that there are at least two such attributes.
@@ -184,14 +208,43 @@ clauselist_selectivity(PlannerInfo *root,
* If there are no such stats or not enough attributes, don't waste time
* simply skip to estimation using the plain per-column stats.
*/
+ if (has_stats(stats, STATS_TYPE_MCV) &&
+ (count_mv_attnums(clauses, relid, STATS_TYPE_MCV) >= 2))
+ {
+ /* collect attributes from the compatible conditions */
+ Bitmapset *mvattnums = collect_mv_attnums(clauses, relid,
+ STATS_TYPE_MCV);
+
+ /* and search for the statistic covering the most attributes */
+ MVStatisticInfo *mvstat = choose_mv_statistics(stats, mvattnums,
+ STATS_TYPE_MCV);
+
+ if (mvstat != NULL) /* we have a matching stats */
+ {
+ /* clauses compatible with multi-variate stats */
+ List *mvclauses = NIL;
+
+ /* split the clauselist into regular and mv-clauses */
+ clauses = clauselist_mv_split(root, relid, clauses, &mvclauses,
+ mvstat, STATS_TYPE_MCV);
+
+ /* we've chosen the histogram to match the clauses */
+ Assert(mvclauses != NIL);
+
+ /* compute the multivariate stats */
+ s1 *= clauselist_mv_selectivity(root, mvclauses, mvstat);
+ }
+ }
+
+ /* Now try to apply functional dependencies on the remaining clauses. */
if (has_stats(stats, STATS_TYPE_FDEPS) &&
- (count_mv_attnums(clauses, relid) >= 2))
+ (count_mv_attnums(clauses, relid, STATS_TYPE_FDEPS) >= 2))
{
MVStatisticInfo *mvstat;
Bitmapset *mvattnums;
/* collect attributes from the compatible conditions */
- mvattnums = collect_mv_attnums(clauses, relid);
+ mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_FDEPS);
/* and search for the statistic covering the most attributes */
mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_FDEPS);
@@ -994,7 +1047,7 @@ clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
/* clauses remaining after removing those on the "implied" attribute */
List *clauses_filtered = NIL;
- attnums = collect_mv_attnums(clauses, relid);
+ attnums = collect_mv_attnums(clauses, relid, STATS_TYPE_FDEPS);
/* no point in looking for dependencies with fewer than 2 attributes */
if (bms_num_members(attnums) < 2)
@@ -1017,7 +1070,7 @@ clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
*/
foreach(lc, clauses)
{
- AttrNumber attnum_clause = InvalidAttrNumber;
+ Bitmapset *attnums_clause = NULL;
Node *clause = (Node *) lfirst(lc);
/*
@@ -1026,17 +1079,20 @@ clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
* we should only see equality clauses compatible with functional
* dependencies, so just error out if we stumble upon something else.
*/
- if (! clause_is_mv_compatible(clause, relid, &attnum_clause))
+ if (! clause_is_mv_compatible(clause, relid, &attnums_clause,
+ STATS_TYPE_FDEPS))
elog(ERROR, "clause not compatible with functional dependencies");
- Assert(AttributeNumberIsValid(attnum_clause));
+ /* we also expect only simple equality clauses */
+ Assert(bms_num_members(attnums_clause) == 1);
/*
* If the clause is not on the implied attribute, add it to the list
* of filtered clauses (for the next round) and continue with the
* next one.
*/
- if (! dependency_implies_attribute(dependency, attnum_clause,
+ if (! dependency_implies_attribute(dependency,
+ bms_singleton_member(attnums_clause),
mvstats->stakeys->values))
{
clauses_filtered = lappend(clauses_filtered, clause);
@@ -1080,10 +1136,71 @@ clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
}
/*
+ * estimate selectivity of clauses using multivariate statistic
+ *
+ * Perform estimation of the clauses using a MCV list.
+ *
+ * This assumes all the clauses are compatible with the selected statistics
+ * (e.g. only reference columns covered by the statistics, use supported
+ * operator, etc.).
+ *
+ * TODO: We may support some additional conditions, most importantly those
+ * matching multiple columns (e.g. "a = b" or "a < b").
+ *
+ * TODO: Clamp the selectivity by min of the per-clause selectivities (i.e. the
+ * selectivity of the most restrictive clause), because that's the maximum
+ * we can ever get from ANDed list of clauses. This may probably prevent
+ * issues with hitting too many buckets and low precision histograms.
+ *
+ * TODO: We may remember the lowest frequency in the MCV list, and then later
+ * use it as a upper boundary for the selectivity (had there been a more
+ * frequent item, it'd be in the MCV list). This might improve cases with
+ * low-detail histograms.
+ *
+ * TODO: We may also derive some additional boundaries for the selectivity from
+ * the MCV list, because
+ *
+ * (a) if we have a "full equality condition" (one equality condition on
+ * each column of the statistic) and we found a match in the MCV list,
+ * then this is the final selectivity (and pretty accurate),
+ *
+ * (b) if we have a "full equality condition" and we haven't found a match
+ * in the MCV list, then the selectivity is below the lowest frequency
+ * found in the MCV list,
+ *
+ * TODO: When applying the clauses to the histogram/MCV list, we can do that
+ * from the most selective clauses first, because that'll eliminate the
+ * buckets/items sooner (so we'll be able to skip them without inspection,
+ * which is more expensive). But this requires really knowing the per-clause
+ * selectivities in advance, and that's not what we do now.
+ */
+static Selectivity
+clauselist_mv_selectivity(PlannerInfo *root, List *clauses, MVStatisticInfo *mvstats)
+{
+ bool fullmatch = false;
+
+ /*
+ * Lowest frequency in the MCV list (may be used as an upper bound for
+ * full equality conditions that did not match any MCV item).
+ */
+ Selectivity mcv_low = 0.0;
+
+ /*
+ * TODO: Evaluate simple 1D selectivities, use the smallest one as an
+ * upper bound, product as lower bound, and sort the clauses in ascending
+ * order by selectivity (to optimize the MCV/histogram evaluation).
+ */
+
+ /* Evaluate the MCV selectivity */
+ return clauselist_mv_selectivity_mcvlist(root, clauses, mvstats,
+ &fullmatch, &mcv_low);
+}
+
+/*
* Collect attributes from mv-compatible clauses.
*/
static Bitmapset *
-collect_mv_attnums(List *clauses, Index relid)
+collect_mv_attnums(List *clauses, Index relid, int types)
{
Bitmapset *attnums = NULL;
ListCell *l;
@@ -1099,12 +1216,10 @@ collect_mv_attnums(List *clauses, Index relid)
*/
foreach(l, clauses)
{
- AttrNumber attnum;
Node *clause = (Node *) lfirst(l);
- /* ignore the result for now - we only need the info */
- if (clause_is_mv_compatible(clause, relid, &attnum))
- attnums = bms_add_member(attnums, attnum);
+ /* ignore the result here - we only need the attnums */
+ clause_is_mv_compatible(clause, relid, &attnums, types);
}
/*
@@ -1125,10 +1240,10 @@ collect_mv_attnums(List *clauses, Index relid)
* Count the number of attributes in clauses compatible with multivariate stats.
*/
static int
-count_mv_attnums(List *clauses, Index relid)
+count_mv_attnums(List *clauses, Index relid, int type)
{
int c;
- Bitmapset *attnums = collect_mv_attnums(clauses, relid);
+ Bitmapset *attnums = collect_mv_attnums(clauses, relid, type);
c = bms_num_members(attnums);
@@ -1263,7 +1378,8 @@ choose_mv_statistics(List *stats, Bitmapset *attnums, int types)
int numattrs = info->stakeys->dim1;
/* skip statistics not matching any of the requested types */
- if (! (info->deps_built && (STATS_TYPE_FDEPS & types)))
+ if (! ((info->deps_built && (STATS_TYPE_FDEPS & types)) ||
+ (info->mcv_built && (STATS_TYPE_MCV & types))))
continue;
/* count columns covered by the statistics */
@@ -1317,13 +1433,13 @@ clauselist_mv_split(PlannerInfo *root, Index relid,
foreach(l, clauses)
{
bool match = false; /* by default not mv-compatible */
- AttrNumber attnum = InvalidAttrNumber;
+ Bitmapset *attnums = NULL;
Node *clause = (Node *) lfirst(l);
- if (clause_is_mv_compatible(clause, relid, &attnum))
+ if (clause_is_mv_compatible(clause, relid, &attnums, types))
{
/* are all the attributes part of the selected stats? */
- if (bms_is_member(attnum, mvattnums))
+ if (bms_is_subset(attnums, mvattnums))
match = true;
}
@@ -1348,6 +1464,7 @@ clauselist_mv_split(PlannerInfo *root, Index relid,
typedef struct
{
+ int types; /* types of statistics ? */
Index varno; /* relid we're interested in */
Bitmapset *varattnos; /* attnums referenced by the clauses */
} mv_compatible_context;
@@ -1382,6 +1499,49 @@ mv_compatible_walker(Node *node, mv_compatible_context *context)
return mv_compatible_walker((Node *) rinfo->clause, (void *) context);
}
+ if (or_clause(node) || and_clause(node) || not_clause(node))
+ {
+ /*
+ * AND/OR/NOT-clauses are supported if all sub-clauses are supported
+ *
+ * TODO: We might support mixed case, where some of the clauses are
+ * supported and some are not, and treat all supported subclauses as a
+ * single clause, compute it's selectivity using mv stats, and compute
+ * the total selectivity using the current algorithm.
+ *
+ * TODO: For RestrictInfo above an OR-clause, we might use the
+ * orclause with nested RestrictInfo - we won't have to call
+ * pull_varnos() for each clause, saving time.
+ *
+ * TODO: Perhaps this needs a bit more thought for functional
+ * dependencies? Those don't quite work for NOT cases.
+ */
+ BoolExpr *expr = (BoolExpr *) node;
+ ListCell *lc;
+
+ foreach(lc, expr->args)
+ {
+ if (mv_compatible_walker((Node *) lfirst(lc), context))
+ return true;
+ }
+
+ return false;
+ }
+
+ if (IsA(node, NullTest))
+ {
+ NullTest *nt = (NullTest *) node;
+
+ /*
+ * Only simple (Var IS NULL) expressions supported for now. Maybe we
+ * could use examine_variable to fix this?
+ */
+ if (!IsA(nt->arg, Var))
+ return true;
+
+ return mv_compatible_walker((Node *) (nt->arg), context);
+ }
+
if (IsA(node, Var))
{
Var *var = (Var *) node;
@@ -1442,10 +1602,18 @@ mv_compatible_walker(Node *node, mv_compatible_context *context)
switch (get_oprrest(expr->opno))
{
case F_EQSEL:
-
/* equality conditions are compatible with all statistics */
break;
+ case F_SCALARLTSEL:
+ case F_SCALARGTSEL:
+
+ /* not compatible with functional dependencies */
+ if (!(context->types & STATS_TYPE_MCV))
+ return true; /* terminate */
+
+ break;
+
default:
/* unknown estimator */
@@ -1479,10 +1647,11 @@ mv_compatible_walker(Node *node, mv_compatible_context *context)
* evaluate them using multivariate stats.
*/
static bool
-clause_is_mv_compatible(Node *clause, Index relid, AttrNumber *attnum)
+clause_is_mv_compatible(Node *clause, Index relid, Bitmapset **attnums, int types)
{
mv_compatible_context context;
+ context.types = types;
context.varno = relid;
context.varattnos = NULL; /* no attnums */
@@ -1490,7 +1659,7 @@ clause_is_mv_compatible(Node *clause, Index relid, AttrNumber *attnum)
return false;
/* remember the newly collected attnums */
- *attnum = bms_singleton_member(context.varattnos);
+ *attnums = bms_add_members(*attnums, context.varattnos);
return true;
}
@@ -1505,6 +1674,9 @@ stats_type_matches(MVStatisticInfo *stat, int type)
if ((type & STATS_TYPE_FDEPS) && stat->deps_built)
return true;
+ if ((type & STATS_TYPE_MCV) && stat->mcv_built)
+ return true;
+
return false;
}
@@ -1538,3 +1710,409 @@ find_stats(PlannerInfo *root, Index relid)
return root->simple_rel_array[relid]->mvstatlist;
}
+
+/*
+ * Estimate selectivity of clauses using a MCV list.
+ *
+ * If there's no MCV list for the stats, the function returns 0.0.
+ *
+ * While computing the estimate, the function checks whether all the
+ * columns were matched with an equality condition. If that's the case,
+ * we can skip processing the histogram, as there can be no rows in
+ * it with the same values - all the rows matching the condition are
+ * represented by the MCV item. This can only happen with equality
+ * on all the attributes.
+ *
+ * The algorithm works like this:
+ *
+ * 1) mark all items as 'match'
+ * 2) walk through all the clauses
+ * 3) for a particular clause, walk through all the items
+ * 4) skip items that are already 'no match'
+ * 5) check clause for items that still match
+ * 6) sum frequencies for items to get selectivity
+ *
+ * The function also returns the frequency of the least frequent item
+ * on the MCV list, which may be useful for clamping estimate from the
+ * histogram (all items not present in the MCV list are less frequent).
+ * This however seems useful only for cases with conditions on all
+ * attributes.
+ *
+ * TODO: This only handles AND-ed clauses, but it might work for OR-ed
+ * lists too - it just needs to reverse the logic a bit. I.e. start
+ * with 'no match' for all items, and mark the items as a match
+ * as the clauses are processed (and skip items that are 'match').
+ */
+static Selectivity
+clauselist_mv_selectivity_mcvlist(PlannerInfo *root, List *clauses,
+ MVStatisticInfo *mvstats, bool *fullmatch,
+ Selectivity *lowsel)
+{
+ int i;
+ Selectivity s = 0.0;
+ Selectivity u = 0.0;
+
+ MCVList mcvlist = NULL;
+ int nmatches = 0;
+
+ /* match/mismatch bitmap for each MCV item */
+ char *matches = NULL;
+
+ Assert(clauses != NIL);
+ Assert(list_length(clauses) >= 2);
+
+ /* there's no MCV list built yet */
+ if (!mvstats->mcv_built)
+ return 0.0;
+
+ mcvlist = load_mv_mcvlist(mvstats->mvoid);
+
+ Assert(mcvlist != NULL);
+ Assert(mcvlist->nitems > 0);
+
+ /* by default all the MCV items match the clauses fully */
+ matches = palloc0(sizeof(char) * mcvlist->nitems);
+ memset(matches, MVSTATS_MATCH_FULL, sizeof(char) * mcvlist->nitems);
+
+ /* number of matching MCV items */
+ nmatches = mcvlist->nitems;
+
+ nmatches = update_match_bitmap_mcvlist(root, clauses,
+ mvstats->stakeys, mcvlist,
+ nmatches, matches,
+ lowsel, fullmatch, false);
+
+ /* sum frequencies for all the matching MCV items */
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ /* used to 'scale' for MCV lists not covering all tuples */
+ u += mcvlist->items[i]->frequency;
+
+ if (matches[i] != MVSTATS_MATCH_NONE)
+ s += mcvlist->items[i]->frequency;
+ }
+
+ pfree(matches);
+ pfree(mcvlist);
+
+ return s * u;
+}
+
+/*
+ * Evaluate clauses using the MCV list, and update the match bitmap.
+ *
+ * The bitmap may be already partially set, so this is really a way to
+ * combine results of several clause lists - either when computing
+ * conditional probability P(A|B) or a combination of AND/OR clauses.
+ *
+ * TODO: This works with 'bitmap' where each bit is represented as a char,
+ * which is slightly wasteful. Instead, we could use a regular
+ * bitmap, reducing the size to ~1/8. Another thing is merging the
+ * bitmaps using & and |, which might be faster than min/max.
+ */
+static int
+update_match_bitmap_mcvlist(PlannerInfo *root, List *clauses,
+ int2vector *stakeys, MCVList mcvlist,
+ int nmatches, char *matches,
+ Selectivity *lowsel, bool *fullmatch,
+ bool is_or)
+{
+ int i;
+ ListCell *l;
+
+ Bitmapset *eqmatches = NULL; /* attributes with equality matches */
+
+ /* The bitmap may be partially built. */
+ Assert(nmatches >= 0);
+ Assert(nmatches <= mcvlist->nitems);
+ Assert(clauses != NIL);
+ Assert(list_length(clauses) >= 1);
+ Assert(mcvlist != NULL);
+ Assert(mcvlist->nitems > 0);
+
+ /* No possible matches (only works for AND-ded clauses) */
+ if (((nmatches == 0) && (!is_or)) ||
+ ((nmatches == mcvlist->nitems) && is_or))
+ return nmatches;
+
+ /*
+ * find the lowest frequency in the MCV list
+ *
+ * We need to do that here, because we do various tricks in the following
+ * code - skipping items already ruled out, etc.
+ *
+ * XXX A loop is necessary because the MCV list is not sorted by
+ * frequency.
+ */
+ *lowsel = 1.0;
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ MCVItem item = mcvlist->items[i];
+
+ if (item->frequency < *lowsel)
+ *lowsel = item->frequency;
+ }
+
+ /*
+ * Loop through the list of clauses, and for each of them evaluate all the
+ * MCV items not yet eliminated by the preceding clauses.
+ */
+ foreach(l, clauses)
+ {
+ Node *clause = (Node *) lfirst(l);
+
+ /* if it's a RestrictInfo, then extract the clause */
+ if (IsA(clause, RestrictInfo))
+ clause = (Node *) ((RestrictInfo *) clause)->clause;
+
+ /* if there are no remaining matches possible, we can stop */
+ if (((nmatches == 0) && (!is_or)) ||
+ ((nmatches == mcvlist->nitems) && is_or))
+ break;
+
+ /* it's either OpClause, or NullTest */
+ if (is_opclause(clause))
+ {
+ OpExpr *expr = (OpExpr *) clause;
+ bool varonleft = true;
+ bool ok;
+ FmgrInfo opproc;
+
+ /* get procedure computing operator selectivity */
+ RegProcedure oprrest = get_oprrest(expr->opno);
+
+ fmgr_info(get_opcode(expr->opno), &opproc);
+
+ ok = (NumRelids(clause) == 1) &&
+ (is_pseudo_constant_clause(lsecond(expr->args)) ||
+ (varonleft = false,
+ is_pseudo_constant_clause(linitial(expr->args))));
+
+ if (ok)
+ {
+
+ FmgrInfo gtproc;
+ Var *var = (varonleft) ? linitial(expr->args) : lsecond(expr->args);
+ Const *cst = (varonleft) ? lsecond(expr->args) : linitial(expr->args);
+ bool isgt = (!varonleft);
+
+ TypeCacheEntry *typecache
+ = lookup_type_cache(var->vartype, TYPECACHE_GT_OPR);
+
+ /* FIXME proper matching attribute to dimension */
+ int idx = mv_get_index(var->varattno, stakeys);
+
+ fmgr_info(get_opcode(typecache->gt_opr), >proc);
+
+ /*
+ * Walk through the MCV items and evaluate the current clause.
+ * We can skip items that were already ruled out, and
+ * terminate if there are no remaining MCV items that might
+ * possibly match.
+ */
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ bool mismatch = false;
+ MCVItem item = mcvlist->items[i];
+
+ /*
+ * If there are no more matches (AND) or no remaining
+ * unmatched items (OR), we can stop processing this
+ * clause.
+ */
+ if (((nmatches == 0) && (!is_or)) ||
+ ((nmatches == mcvlist->nitems) && is_or))
+ break;
+
+ /*
+ * For AND-lists, we can also mark NULL items as 'no
+ * match' (and then skip them). For OR-lists this is not
+ * possible.
+ */
+ if ((!is_or) && item->isnull[idx])
+ matches[i] = MVSTATS_MATCH_NONE;
+
+ /* skip MCV items that were already ruled out */
+ if ((!is_or) && (matches[i] == MVSTATS_MATCH_NONE))
+ continue;
+ else if (is_or && (matches[i] == MVSTATS_MATCH_FULL))
+ continue;
+
+ switch (oprrest)
+ {
+ case F_EQSEL:
+
+ /*
+ * We don't care about isgt in equality, because
+ * it does not matter whether it's (var = const)
+ * or (const = var).
+ */
+ mismatch = !DatumGetBool(FunctionCall2Coll(&opproc,
+ DEFAULT_COLLATION_OID,
+ cst->constvalue,
+ item->values[idx]));
+
+ if (!mismatch)
+ eqmatches = bms_add_member(eqmatches, idx);
+
+ break;
+
+ case F_SCALARLTSEL: /* column < constant */
+ case F_SCALARGTSEL: /* column > constant */
+
+ /*
+ * First check whether the constant is below the
+ * lower boundary (in that case we can skip the
+ * bucket, because there's no overlap).
+ */
+ if (isgt)
+ mismatch = !DatumGetBool(FunctionCall2Coll(&opproc,
+ DEFAULT_COLLATION_OID,
+ cst->constvalue,
+ item->values[idx]));
+ else
+ mismatch = !DatumGetBool(FunctionCall2Coll(&opproc,
+ DEFAULT_COLLATION_OID,
+ item->values[idx],
+ cst->constvalue));
+
+ break;
+ }
+
+ /*
+ * XXX The conditions on matches[i] are not needed, as we
+ * skip MCV items that can't become true/false, depending
+ * on the current flag. See beginning of the loop over MCV
+ * items.
+ */
+
+ if ((is_or) && (matches[i] == MVSTATS_MATCH_NONE) && (!mismatch))
+ {
+ /* OR - was MATCH_NONE, but will be MATCH_FULL */
+ matches[i] = MVSTATS_MATCH_FULL;
+ ++nmatches;
+ continue;
+ }
+ else if ((!is_or) && (matches[i] == MVSTATS_MATCH_FULL) && mismatch)
+ {
+ /* AND - was MATC_FULL, but will be MATCH_NONE */
+ matches[i] = MVSTATS_MATCH_NONE;
+ --nmatches;
+ continue;
+ }
+
+ }
+ }
+ }
+ else if (IsA(clause, NullTest))
+ {
+ NullTest *expr = (NullTest *) clause;
+ Var *var = (Var *) (expr->arg);
+
+ /* FIXME proper matching attribute to dimension */
+ int idx = mv_get_index(var->varattno, stakeys);
+
+ /*
+ * Walk through the MCV items and evaluate the current clause. We
+ * can skip items that were already ruled out, and terminate if
+ * there are no remaining MCV items that might possibly match.
+ */
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ MCVItem item = mcvlist->items[i];
+
+ /*
+ * if there are no more matches, we can stop processing this
+ * clause
+ */
+ if (nmatches == 0)
+ break;
+
+ /* skip MCV items that were already ruled out */
+ if (matches[i] == MVSTATS_MATCH_NONE)
+ continue;
+
+ /* if the clause mismatches the MCV item, set it as MATCH_NONE */
+ if (((expr->nulltesttype == IS_NULL) && (!item->isnull[idx])) ||
+ ((expr->nulltesttype == IS_NOT_NULL) && (item->isnull[idx])))
+ {
+ matches[i] = MVSTATS_MATCH_NONE;
+ --nmatches;
+ }
+ }
+ }
+ else if (or_clause(clause) || and_clause(clause))
+ {
+ /*
+ * AND/OR clause, with all clauses compatible with the selected MV
+ * stat
+ */
+
+ int i;
+ BoolExpr *orclause = ((BoolExpr *) clause);
+ List *orclauses = orclause->args;
+
+ /* match/mismatch bitmap for each MCV item */
+ int or_nmatches = 0;
+ char *or_matches = NULL;
+
+ Assert(orclauses != NIL);
+ Assert(list_length(orclauses) >= 2);
+
+ /* number of matching MCV items */
+ or_nmatches = mcvlist->nitems;
+
+ /* by default none of the MCV items matches the clauses */
+ or_matches = palloc0(sizeof(char) * or_nmatches);
+
+ if (or_clause(clause))
+ {
+ /* OR clauses assume nothing matches, initially */
+ memset(or_matches, MVSTATS_MATCH_NONE, sizeof(char) * or_nmatches);
+ or_nmatches = 0;
+ }
+ else
+ {
+ /* AND clauses assume nothing matches, initially */
+ memset(or_matches, MVSTATS_MATCH_FULL, sizeof(char) * or_nmatches);
+ }
+
+ /* build the match bitmap for the OR-clauses */
+ or_nmatches = update_match_bitmap_mcvlist(root, orclauses,
+ stakeys, mcvlist,
+ or_nmatches, or_matches,
+ lowsel, fullmatch, or_clause(clause));
+
+ /* merge the bitmap into the existing one */
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ /*
+ * Merge the result into the bitmap (Min for AND, Max for OR).
+ *
+ * FIXME this does not decrease the number of matches
+ */
+ UPDATE_RESULT(matches[i], or_matches[i], is_or);
+ }
+
+ pfree(or_matches);
+
+ }
+ else
+ {
+ elog(ERROR, "unknown clause type: %d", clause->type);
+ }
+ }
+
+ /*
+ * If all the columns were matched by equality, it's a full match. In this
+ * case there can be just a single MCV item, matching the clause (if there
+ * were two, both would match the other one).
+ */
+ *fullmatch = (bms_num_members(eqmatches) == mcvlist->ndimensions);
+
+ /* free the allocated pieces */
+ if (eqmatches)
+ pfree(eqmatches);
+
+ return nmatches;
+}
diff --git a/src/backend/optimizer/util/plancat.c b/src/backend/optimizer/util/plancat.c
index 8129143..9dd4e83 100644
--- a/src/backend/optimizer/util/plancat.c
+++ b/src/backend/optimizer/util/plancat.c
@@ -1287,7 +1287,7 @@ get_relation_statistics(RelOptInfo *rel, Relation relation)
mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
/* unavailable stats are not interesting for the planner */
- if (mvstat->deps_built || mvstat->ndist_built)
+ if (mvstat->deps_built || mvstat->ndist_built || mvstat->mcv_built)
{
info = makeNode(MVStatisticInfo);
@@ -1297,10 +1297,12 @@ get_relation_statistics(RelOptInfo *rel, Relation relation)
/* enabled statistics */
info->ndist_enabled = mvstat->ndist_enabled;
info->deps_enabled = mvstat->deps_enabled;
+ info->mcv_enabled = mvstat->mcv_enabled;
/* built/available statistics */
info->ndist_built = mvstat->ndist_built;
info->deps_built = mvstat->deps_built;
+ info->mcv_built = mvstat->mcv_built;
/* stakeys */
adatum = SysCacheGetAttr(MVSTATOID, htup,
diff --git a/src/backend/utils/mvstats/Makefile b/src/backend/utils/mvstats/Makefile
index 21fe7e5..d5d47ba 100644
--- a/src/backend/utils/mvstats/Makefile
+++ b/src/backend/utils/mvstats/Makefile
@@ -12,6 +12,6 @@ subdir = src/backend/utils/mvstats
top_builddir = ../../../..
include $(top_builddir)/src/Makefile.global
-OBJS = common.o dependencies.o mvdist.o
+OBJS = common.o dependencies.o mcv.o mvdist.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/utils/mvstats/README.mcv b/src/backend/utils/mvstats/README.mcv
new file mode 100644
index 0000000..e93cfe4
--- /dev/null
+++ b/src/backend/utils/mvstats/README.mcv
@@ -0,0 +1,137 @@
+MCV lists
+=========
+
+Multivariate MCV (most-common values) lists are a straightforward extension of
+regular MCV list, tracking most frequent combinations of values for a group of
+attributes.
+
+This works particularly well for columns with a small number of distinct values,
+as the list may include all the combinations and approximate the distribution
+very accurately.
+
+For columns with large number of distinct values (e.g. those with continuous
+domains), the list will only track the most frequent combinations. If the
+distribution is mostly uniform (all combinations about equally frequent), the
+MCV list will be empty.
+
+Estimates of some clauses (e.g. equality) based on MCV lists are more accurate
+than when using histograms.
+
+Also, MCV lists don't necessarily require sorting of the values (the fact that
+we use sorting when building them is implementation detail), but even more
+importantly the ordering is not built into the approximation (while histograms
+are built on ordering). So MCV lists work well even for attributes where the
+ordering of the data type is disconnected from the meaning of the data. For
+example we know how to sort strings, but it's unlikely to make much sense for
+city names (or other label-like attributes).
+
+
+Selectivity estimation
+----------------------
+
+The estimation, implemented in clauselist_mv_selectivity_mcvlist(), is quite
+simple in principle - we need to identify MCV items matching all the clauses
+and sum frequencies of all those items.
+
+Currently MCV lists support estimation of the following clause types:
+
+ (a) equality clauses WHERE (a = 1) AND (b = 2)
+ (b) inequality clauses WHERE (a < 1) AND (b >= 2)
+ (c) NULL clauses WHERE (a IS NULL) AND (b IS NOT NULL)
+ (d) OR clauses WHERE (a < 1) OR (b >= 2)
+
+It's possible to add support for additional clauses, for example:
+
+ (e) multi-var clauses WHERE (a > b)
+
+and possibly others. These are tasks for the future, not yet implemented.
+
+
+Estimating equality clauses
+---------------------------
+
+When computing selectivity estimate for equality clauses
+
+ (a = 1) AND (b = 2)
+
+we can do this estimate pretty exactly assuming that two conditions are met:
+
+ (1) there's an equality condition on all attributes of the statistic
+
+ (2) we find a matching item in the MCV list
+
+In this case we know the MCV item represents all tuples matching the clauses,
+and the selectivity estimate is complete (i.e. we don't need to perform
+estimation using the histogram). This is what we call 'full match'.
+
+When only (1) holds, but there's no matching MCV item, we don't know whether
+there are no such rows or just are not very frequent. We can however use the
+frequency of the least frequent MCV item as an upper bound for the selectivity.
+
+For a combination of equality conditions (not full-match case) we can clamp the
+selectivity by the minimum of selectivities for each condition. For example if
+we know the number of distinct values for each column, we can use 1/ndistinct
+as a per-column estimate. Or rather 1/ndistinct + selectivity derived from the
+MCV list.
+
+We should also probably only use the 'residual ndistinct' by exluding the items
+included in the MCV list (and also residual frequency):
+
+ f = (1.0 - sum(MCV frequencies)) / (ndistinct - ndistinct(MCV list))
+
+but it's worth pointing out the ndistinct values are multi-variate for the
+columns referenced by the equality conditions.
+
+Note: Only the "full match" limit is currently implemented.
+
+
+Hashed MCV (not yet implemented)
+--------------------------------
+
+Regular MCV lists have to include actual values for each item, so if those items
+are large the list may be quite large. This is especially true for multi-variate
+MCV lists, although the current implementation partially mitigates this by
+performing de-duplicating the values before storing them on disk.
+
+It's possible to only store hashes (32-bit values) instead of the actual values,
+significantly reducing the space requirements. Obviously, this would only make
+the MCV lists useful for estimating equality conditions (assuming the 32-bit
+hashes make the collisions rare enough).
+
+This might also complicate matching the columns to available stats.
+
+
+TODO Consider implementing hashed MCV list, storing just 32-bit hashes instead
+ of the actual values. This type of MCV list will be useful only for
+ estimating equality clauses, and will reduce space requirements for large
+ varlena types (in such cases we usually only want equality anyway).
+
+TODO Currently there's no logic to consider building only a MCV list (and not
+ building the histogram at all), except for doing this decision manually in
+ ADD STATISTICS.
+
+
+Inspecting the MCV list
+-----------------------
+
+Inspecting the regular (per-attribute) MCV lists is trivial, as it's enough
+to select the columns from pg_stats - the data is encoded as anyarrays, so we
+simply get the text representation of the arrays.
+
+With multivariate MCV lits it's not that simple due to the possible mix of
+data types. It might be possible to produce similar array-like representation,
+but that'd unnecessarily complicate further processing and analysis of the MCV
+list. Instead, there's a SRF function providing values, frequencies etc.
+
+ SELECT * FROM pg_mv_mcv_items();
+
+It has two input parameters:
+
+ oid - OID of the MCV list (pg_mv_statistic.staoid)
+
+and produces a table with these columns:
+
+ - item ID (0...nitems-1)
+ - values (string array)
+ - nulls only (boolean array)
+ - frequency (double precision)
diff --git a/src/backend/utils/mvstats/README.stats b/src/backend/utils/mvstats/README.stats
index 814f39c..8d3d268 100644
--- a/src/backend/utils/mvstats/README.stats
+++ b/src/backend/utils/mvstats/README.stats
@@ -8,9 +8,50 @@ not true, resulting in estimation errors.
Multivariate stats track different types of dependencies between the columns,
hopefully improving the estimates.
-Currently we only have one kind of multivariate statistics - soft functional
-dependencies, and we use it to improve estimates of equality clauses. See
-README.dependencies for details.
+
+Types of statistics
+-------------------
+
+Currently we only have two kinds of multivariate statistics
+
+ (a) soft functional dependencies (README.dependencies)
+
+ (b) MCV lists (README.mcv)
+
+
+Compatible clause types
+-----------------------
+
+Each type of statistics may be used to estimate some subset of clause types.
+
+ (a) functional dependencies - equality clauses (AND), possibly IS NULL
+
+ (b) MCV list - equality and inequality clauses, IS [NOT] NULL, AND/OR
+
+Currently only simple operator clauses (Var op Const) are supported, but it's
+possible to support more complex clause types, e.g. (Var op Var).
+
+
+Complex clauses
+---------------
+
+We also support estimating more complex clauses - essentially AND/OR clauses
+with (Var op Const) as leaves, as long as all the referenced attributes are
+covered by a single statistics.
+
+For example this condition
+
+ (a=1) AND ((b=2) OR ((c=3) AND (d=4)))
+
+may be estimated using statistics on (a,b,c,d). If we only have statistics on
+(b,c,d) we may estimate the second part, and estimate (a=1) using simple stats.
+
+If we only have statistics on (a,b,c) we can't apply it at all at this point,
+but it's worth pointing out clauselist_selectivity() works recursively and when
+handling the second part (the OR-clause), we'll be able to apply the statistics.
+
+Note: The multi-statistics estimation patch also makes it possible to pass some
+clauses as 'conditions' into the deeper parts of the expression tree.
Selectivity estimation
@@ -23,21 +64,53 @@ When estimating selectivity, we aim to achieve several things:
(b) minimize the overhead, especially when no suitable multivariate stats
exist (so if you are not using multivariate stats, there's no overhead)
-This clauselist_selectivity() performs several inexpensive checks first, before
+Thus clauselist_selectivity() performs several inexpensive checks first, before
even attempting to do the more expensive estimation.
(1) check if there are multivariate stats on the relation
- (2) check there are at least two attributes referenced by clauses compatible
- with multivariate statistics (equality clauses for func. dependencies)
+ (2) check that there are functional dependencies on the table, and that
+ there are at least two attributes referenced by compatible clauses
+ (equality clauses for func. dependencies)
(3) perform reduction of equality clauses using func. dependencies
- (4) estimate the reduced list of clauses using regular statistics
+ (4) check that there are multivariate MCV lists on the table, and that
+ there are at least two attributes referenced by compatible clauses
+ (equalities, inequalities, etc.)
+
+ (5) find the best multivariate statistics (matching the most conditions)
+ and use it to compute the estimate
+
+ (6) estimate the remaining clauses (not estimated using multivariate stats)
+ using the regular per-column statistics
Whenever we find there are no suitable stats, we skip the expensive steps.
+Further (possibly crazy) ideas
+------------------------------
+
+Currently the clauses are only estimated using a single statistics, even if
+there are multiple candidate statistics - for example assume we have statistics
+on (a,b,c) and (b,c,d), and estimate conditions
+
+ (b = 1) AND (c = 2)
+
+Then both statistics may be used, but we only use one of them. Maybe we could
+use compute estimates using all candidate stats, and somehow aggregate them
+into the final estimate by using average or median.
+
+Some stats may give better estimates than others, but it's very difficult to say
+in advance which stats are the best (it depends on the number of buckets, number
+of additional columns not referenced in the clauses, type of condition etc.).
+
+But of course, this may result in expensive estimation (CPU-wise).
+
+So we might add a GUC to choose between a simple (single statistics) and thus
+multi-statistic estimation, possibly table-level parameter (ALTER TABLE ...).
+
+
Size of sample in ANALYZE
-------------------------
When performing ANALYZE, the number of rows to sample is determined as
diff --git a/src/backend/utils/mvstats/common.c b/src/backend/utils/mvstats/common.c
index 39e3b92..fc8eae2 100644
--- a/src/backend/utils/mvstats/common.c
+++ b/src/backend/utils/mvstats/common.c
@@ -15,6 +15,7 @@
*/
#include "common.h"
+#include "utils/array.h"
static VacAttrStats **lookup_var_attr_stats(int2vector *attrs,
int natts, VacAttrStats **vacattrstats);
@@ -23,9 +24,9 @@ static List *list_mv_stats(Oid relid);
static void update_mv_stats(Oid relid,
MVNDistinct ndistinct, MVDependencies dependencies,
+ MCVList mcvlist,
int2vector *attrs, VacAttrStats **stats);
-
/*
* Compute requested multivariate stats, using the rows sampled for the
* plain (single-column) stats.
@@ -55,6 +56,8 @@ build_mv_stats(Relation onerel, double totalrows,
MVStatisticInfo *stat = (MVStatisticInfo *) lfirst(lc);
MVNDistinct ndistinct = NULL;
MVDependencies deps = NULL;
+ MCVList mcvlist = NULL;
+ int numrows_filtered = 0;
VacAttrStats **stats = NULL;
int numatts = 0;
@@ -95,8 +98,12 @@ build_mv_stats(Relation onerel, double totalrows,
if (stat->deps_enabled)
deps = build_mv_dependencies(numrows, rows, attrs, stats);
+ /* build the MCV list */
+ if (stat->mcv_enabled)
+ mcvlist = build_mv_mcvlist(numrows, rows, attrs, stats, &numrows_filtered);
+
/* store the statistics in the catalog */
- update_mv_stats(stat->mvoid, ndistinct, deps, attrs, stats);
+ update_mv_stats(stat->mvoid, ndistinct, deps, mcvlist, attrs, stats);
}
}
@@ -178,6 +185,8 @@ list_mv_stats(Oid relid)
info->ndist_built = stats->ndist_built;
info->deps_enabled = stats->deps_enabled;
info->deps_built = stats->deps_built;
+ info->mcv_enabled = stats->mcv_enabled;
+ info->mcv_built = stats->mcv_built;
result = lappend(result, info);
}
@@ -195,11 +204,58 @@ list_mv_stats(Oid relid)
}
/*
+ * Find attnums of MV stats using the mvoid.
+ */
+int2vector *
+find_mv_attnums(Oid mvoid, Oid *relid)
+{
+ ArrayType *arr;
+ Datum adatum;
+ bool isnull;
+ HeapTuple htup;
+ int2vector *keys;
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ htup = SearchSysCache1(MVSTATOID,
+ ObjectIdGetDatum(mvoid));
+
+ /* XXX syscache contains OIDs of deleted stats (not invalidated) */
+ if (!HeapTupleIsValid(htup))
+ return NULL;
+
+ /* starelid */
+ adatum = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_starelid, &isnull);
+ Assert(!isnull);
+
+ *relid = DatumGetObjectId(adatum);
+
+ /* stakeys */
+ adatum = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_stakeys, &isnull);
+ Assert(!isnull);
+
+ arr = DatumGetArrayTypeP(adatum);
+
+ keys = buildint2vector((int16 *) ARR_DATA_PTR(arr),
+ ARR_DIMS(arr)[0]);
+ ReleaseSysCache(htup);
+
+ /*
+ * TODO maybe save the list into relcache, as in RelationGetIndexList
+ * (which was used as an inspiration of this one)?.
+ */
+
+ return keys;
+}
+
+/*
* update_mv_stats
* Serializes the statistics and stores them into the pg_mv_statistic tuple.
*/
static void
-update_mv_stats(Oid mvoid, MVNDistinct ndistinct, MVDependencies dependencies,
+update_mv_stats(Oid mvoid,
+ MVNDistinct ndistinct, MVDependencies dependencies, MCVList mcvlist,
int2vector *attrs, VacAttrStats **stats)
{
HeapTuple stup,
@@ -233,22 +289,36 @@ update_mv_stats(Oid mvoid, MVNDistinct ndistinct, MVDependencies dependencies,
= PointerGetDatum(serialize_mv_dependencies(dependencies));
}
+ if (mcvlist != NULL)
+ {
+ bytea *data = serialize_mv_mcvlist(mcvlist, attrs, stats);
+
+ nulls[Anum_pg_mv_statistic_stamcv - 1] = (data == NULL);
+ values[Anum_pg_mv_statistic_stamcv - 1] = PointerGetDatum(data);
+ }
+
/* always replace the value (either by bytea or NULL) */
replaces[Anum_pg_mv_statistic_standist - 1] = true;
replaces[Anum_pg_mv_statistic_stadeps - 1] = true;
+ replaces[Anum_pg_mv_statistic_stamcv - 1] = true;
/* always change the availability flags */
nulls[Anum_pg_mv_statistic_ndist_built - 1] = false;
nulls[Anum_pg_mv_statistic_deps_built - 1] = false;
+ nulls[Anum_pg_mv_statistic_mcv_built - 1] = false;
+
nulls[Anum_pg_mv_statistic_stakeys - 1] = false;
/* use the new attnums, in case we removed some dropped ones */
replaces[Anum_pg_mv_statistic_ndist_built - 1] = true;
replaces[Anum_pg_mv_statistic_deps_built - 1] = true;
+ replaces[Anum_pg_mv_statistic_mcv_built - 1] = true;
+
replaces[Anum_pg_mv_statistic_stakeys - 1] = true;
values[Anum_pg_mv_statistic_ndist_built - 1] = BoolGetDatum(ndistinct != NULL);
values[Anum_pg_mv_statistic_deps_built - 1] = BoolGetDatum(dependencies != NULL);
+ values[Anum_pg_mv_statistic_mcv_built - 1] = BoolGetDatum(mcvlist != NULL);
values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(attrs);
@@ -278,6 +348,23 @@ update_mv_stats(Oid mvoid, MVNDistinct ndistinct, MVDependencies dependencies,
heap_close(sd, RowExclusiveLock);
}
+
+int
+mv_get_index(AttrNumber varattno, int2vector *stakeys)
+{
+ int i,
+ idx = 0;
+
+ for (i = 0; i < stakeys->dim1; i++)
+ {
+ if (stakeys->values[i] < varattno)
+ idx += 1;
+ else
+ break;
+ }
+ return idx;
+}
+
/* multi-variate stats comparator */
/*
@@ -288,11 +375,15 @@ update_mv_stats(Oid mvoid, MVNDistinct ndistinct, MVDependencies dependencies,
int
compare_scalars_simple(const void *a, const void *b, void *arg)
{
- Datum da = *(Datum *) a;
- Datum db = *(Datum *) b;
- SortSupport ssup = (SortSupport) arg;
+ return compare_datums_simple(*(Datum *) a,
+ *(Datum *) b,
+ (SortSupport) arg);
+}
- return ApplySortComparator(da, false, db, false, ssup);
+int
+compare_datums_simple(Datum a, Datum b, SortSupport ssup)
+{
+ return ApplySortComparator(a, false, b, false, ssup);
}
/*
@@ -410,3 +501,34 @@ multi_sort_compare_dims(int start, int end,
return 0;
}
+
+/* simple counterpart to qsort_arg */
+void *
+bsearch_arg(const void *key, const void *base, size_t nmemb, size_t size,
+ int (*compar) (const void *, const void *, void *),
+ void *arg)
+{
+ size_t l,
+ u,
+ idx;
+ const void *p;
+ int comparison;
+
+ l = 0;
+ u = nmemb;
+ while (l < u)
+ {
+ idx = (l + u) / 2;
+ p = (void *) (((const char *) base) + (idx * size));
+ comparison = (*compar) (key, p, arg);
+
+ if (comparison < 0)
+ u = idx;
+ else if (comparison > 0)
+ l = idx + 1;
+ else
+ return (void *) p;
+ }
+
+ return NULL;
+}
diff --git a/src/backend/utils/mvstats/common.h b/src/backend/utils/mvstats/common.h
index e471c88..fe56f51 100644
--- a/src/backend/utils/mvstats/common.h
+++ b/src/backend/utils/mvstats/common.h
@@ -47,6 +47,15 @@ typedef struct
int tupno; /* position index for tuple it came from */
} ScalarItem;
+/* (de)serialization info */
+typedef struct DimensionInfo
+{
+ int nvalues; /* number of deduplicated values */
+ int nbytes; /* number of bytes (serialized) */
+ int typlen; /* pg_type.typlen */
+ bool typbyval; /* pg_type.typbyval */
+} DimensionInfo;
+
/* multi-sort */
typedef struct MultiSortSupportData
{
@@ -60,6 +69,7 @@ typedef struct SortItem
{
Datum *values;
bool *isnull;
+ int count;
} SortItem;
MultiSortSupport multi_sort_init(int ndims);
@@ -67,7 +77,7 @@ MultiSortSupport multi_sort_init(int ndims);
void multi_sort_add_dimension(MultiSortSupport mss, int sortdim,
int dim, VacAttrStats **vacattrstats);
-int multi_sort_compare(const void *a, const void *b, void *arg);
+int multi_sort_compare(const void *a, const void *b, void *arg);
int multi_sort_compare_dim(int dim, const SortItem *a,
const SortItem *b, MultiSortSupport mss);
@@ -76,5 +86,11 @@ int multi_sort_compare_dims(int start, int end, const SortItem *a,
const SortItem *b, MultiSortSupport mss);
/* comparators, used when constructing multivariate stats */
-int compare_scalars_simple(const void *a, const void *b, void *arg);
-int compare_scalars_partition(const void *a, const void *b, void *arg);
+int compare_datums_simple(Datum a, Datum b, SortSupport ssup);
+int compare_scalars_simple(const void *a, const void *b, void *arg);
+int compare_scalars_partition(const void *a, const void *b, void *arg);
+
+void *bsearch_arg(const void *key, const void *base,
+ size_t nmemb, size_t size,
+ int (*compar) (const void *, const void *, void *),
+ void *arg);
diff --git a/src/backend/utils/mvstats/mcv.c b/src/backend/utils/mvstats/mcv.c
new file mode 100644
index 0000000..c1c2409
--- /dev/null
+++ b/src/backend/utils/mvstats/mcv.c
@@ -0,0 +1,1184 @@
+/*-------------------------------------------------------------------------
+ *
+ * mcv.c
+ * POSTGRES multivariate MCV lists
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/mcv.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "fmgr.h"
+#include "funcapi.h"
+
+#include "utils/bytea.h"
+#include "utils/lsyscache.h"
+
+#include "common.h"
+
+/*
+ * Each serialized item needs to store (in this order):
+ *
+ * - indexes (ndim * sizeof(uint16))
+ * - null flags (ndim * sizeof(bool))
+ * - frequency (sizeof(double))
+ *
+ * So in total:
+ *
+ * ndim * (sizeof(uint16) + sizeof(bool)) + sizeof(double)
+ */
+#define ITEM_SIZE(ndims) \
+ (ndims * (sizeof(uint16) + sizeof(bool)) + sizeof(double))
+
+/* Macros for convenient access to parts of the serialized MCV item */
+#define ITEM_INDEXES(item) ((uint16*)item)
+#define ITEM_NULLS(item,ndims) ((bool*)(ITEM_INDEXES(item) + ndims))
+#define ITEM_FREQUENCY(item,ndims) ((double*)(ITEM_NULLS(item,ndims) + ndims))
+
+static MultiSortSupport build_mss(VacAttrStats **stats, int2vector *attrs);
+
+static SortItem *build_sorted_items(int numrows, HeapTuple *rows,
+ TupleDesc tdesc, MultiSortSupport mss,
+ int2vector *attrs);
+
+static SortItem *build_distinct_groups(int numrows, SortItem *items,
+ MultiSortSupport mss, int *ndistinct);
+
+static int count_distinct_groups(int numrows, SortItem *items,
+ MultiSortSupport mss);
+
+/*
+ * Builds MCV list from the set of sampled rows.
+ *
+ * The algorithm is quite simple:
+ *
+ * (1) sort the data (default collation, '<' for the data type)
+ *
+ * (2) count distinct groups, decide how many to keep
+ *
+ * (3) build the MCV list using the threshold determined in (2)
+ *
+ * (4) remove rows represented by the MCV from the sample
+ *
+ * The method also removes rows matching the MCV items from the input array,
+ * and passes the number of remaining rows (useful for building histograms)
+ * using the numrows_filtered parameter.
+ *
+ * FIXME: Single-dimensional MCV is sorted by frequency (descending). We should
+ * do that too, because when walking through the list we want to check
+ * the most frequent items first.
+ *
+ * TODO: We're using Datum (8B), even for data types (e.g. int4 or float4).
+ * Maybe we could save some space here, but the bytea compression should
+ * handle it just fine.
+ *
+ * TODO: This probably should not use the ndistinct directly (as computed from
+ * the table, but rather estimate the number of distinct values in the
+ * table), no?
+ */
+MCVList
+build_mv_mcvlist(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int *numrows_filtered)
+{
+ int i;
+ int numattrs = attrs->dim1;
+ int ndistinct = 0;
+ int mcv_threshold = 0;
+ int nitems = 0;
+
+ MCVList mcvlist = NULL;
+
+ /* comparator for all the columns */
+ MultiSortSupport mss = build_mss(stats, attrs);
+
+ /* sort the rows */
+ SortItem *items = build_sorted_items(numrows, rows, stats[0]->tupDesc,
+ mss, attrs);
+
+ /* transform the sorted rows into groups (sorted by frequency) */
+ SortItem *groups = build_distinct_groups(numrows, items, mss, &ndistinct);
+
+ /*
+ * Determine the minimum size of a group to be eligible for MCV list, and
+ * check how many groups actually pass that threshold. We use 1.25x the
+ * avarage group size, just like for regular statistics.
+ *
+ * But if we can fit all the distinct values in the MCV list (i.e. if
+ * there are less distinct groups than MVSTAT_MCVLIST_MAX_ITEMS), we'll
+ * require only 2 rows per group.
+ */
+ mcv_threshold = 1.25 * numrows / ndistinct;
+ mcv_threshold = (mcv_threshold < 4) ? 4 : mcv_threshold;
+
+ if (ndistinct <= MVSTAT_MCVLIST_MAX_ITEMS)
+ mcv_threshold = 2;
+
+ /* Walk through the groups and stop once we fall below the threshold. */
+ nitems = 0;
+ for (i = 0; i < ndistinct; i++)
+ {
+ if (groups[i].count < mcv_threshold)
+ break;
+
+ nitems++;
+ }
+
+ /* we know the number of MCV list items, so let's build the list */
+ if (nitems > 0)
+ {
+ /* allocate the MCV list structure, set parameters we know */
+ mcvlist = (MCVList) palloc0(sizeof(MCVListData));
+
+ mcvlist->magic = MVSTAT_MCV_MAGIC;
+ mcvlist->type = MVSTAT_MCV_TYPE_BASIC;
+ mcvlist->ndimensions = numattrs;
+ mcvlist->nitems = nitems;
+
+ /*
+ * Preallocate Datum/isnull arrays (not as a single chunk, as we will
+ * pass the result outside and thus it needs to be easy to pfree().
+ *
+ * XXX Although we're the only ones dealing with this.
+ */
+ mcvlist->items = (MCVItem *) palloc0(sizeof(MCVItem) * nitems);
+
+ for (i = 0; i < nitems; i++)
+ {
+ mcvlist->items[i] = (MCVItem) palloc0(sizeof(MCVItemData));
+ mcvlist->items[i]->values = (Datum *) palloc0(sizeof(Datum) * numattrs);
+ mcvlist->items[i]->isnull = (bool *) palloc0(sizeof(bool) * numattrs);
+ }
+
+ /* Copy the first chunk of groups into the result. */
+ for (i = 0; i < nitems; i++)
+ {
+ /* just pointer to the proper place in the list */
+ MCVItem item = mcvlist->items[i];
+
+ /* copy values from the _previous_ group (last item of) */
+ memcpy(item->values, groups[i].values, sizeof(Datum) * numattrs);
+ memcpy(item->isnull, groups[i].isnull, sizeof(bool) * numattrs);
+
+ /* and finally the group frequency */
+ item->frequency = (double) groups[i].count / numrows;
+ }
+
+ /* make sure the loops are consistent */
+ Assert(nitems == mcvlist->nitems);
+
+ /*
+ * Remove the rows matching the MCV list (i.e. keep only rows that are
+ * not represented by the MCV list). We will first sort the groups by
+ * the keys (not by count) and then use binary search.
+ */
+ if (nitems > ndistinct)
+ {
+ int i,
+ j;
+ int nfiltered = 0;
+
+ /* used for the searches */
+ SortItem key;
+
+ /* wfill this with data from the rows */
+ key.values = (Datum *) palloc0(numattrs * sizeof(Datum));
+ key.isnull = (bool *) palloc0(numattrs * sizeof(bool));
+
+ /*
+ * Sort the groups for bsearch_r (but only the items that actually
+ * made it to the MCV list).
+ */
+ qsort_arg((void *) groups, nitems, sizeof(SortItem),
+ multi_sort_compare, mss);
+
+ /* walk through the tuples, compare the values to MCV items */
+ for (i = 0; i < numrows; i++)
+ {
+ /* collect the key values from the row */
+ for (j = 0; j < numattrs; j++)
+ key.values[j]
+ = heap_getattr(rows[i], attrs->values[j],
+ stats[j]->tupDesc, &key.isnull[j]);
+
+ /* if not included in the MCV list, keep it in the array */
+ if (bsearch_arg(&key, groups, nitems, sizeof(SortItem),
+ multi_sort_compare, mss) == NULL)
+ rows[nfiltered++] = rows[i];
+ }
+
+ /* remember how many rows we actually kept */
+ *numrows_filtered = nfiltered;
+
+ /* free all the data used here */
+ pfree(key.values);
+ pfree(key.isnull);
+ }
+ else
+ /* the MCV list convers all the rows */
+ *numrows_filtered = 0;
+ }
+
+ pfree(items);
+ pfree(groups);
+
+ return mcvlist;
+}
+
+/* build MultiSortSupport for the attributes passed in attrs */
+static MultiSortSupport
+build_mss(VacAttrStats **stats, int2vector *attrs)
+{
+ int i;
+ int numattrs = attrs->dim1;
+
+ /* Sort by multiple columns (using array of SortSupport) */
+ MultiSortSupport mss = multi_sort_init(numattrs);
+
+ /* prepare the sort functions for all the attributes */
+ for (i = 0; i < numattrs; i++)
+ multi_sort_add_dimension(mss, i, i, stats);
+
+ return mss;
+}
+
+/* build sorted array of SortItem with values from rows */
+static SortItem *
+build_sorted_items(int numrows, HeapTuple *rows, TupleDesc tdesc,
+ MultiSortSupport mss, int2vector *attrs)
+{
+ int i,
+ j,
+ len;
+ int numattrs = attrs->dim1;
+ int nvalues = numrows * numattrs;
+
+ /*
+ * We won't allocate the arrays for each item independenly, but in one
+ * large chunk and then just set the pointers.
+ */
+ SortItem *items;
+ Datum *values;
+ bool *isnull;
+ char *ptr;
+
+ /* Compute the total amount of memory we need (both items and values). */
+ len = numrows * sizeof(SortItem) + nvalues * (sizeof(Datum) + sizeof(bool));
+
+ /* Allocate the memory and split it into the pieces. */
+ ptr = palloc0(len);
+
+ /* items to sort */
+ items = (SortItem *) ptr;
+ ptr += numrows * sizeof(SortItem);
+
+ /* values and null flags */
+ values = (Datum *) ptr;
+ ptr += nvalues * sizeof(Datum);
+
+ isnull = (bool *) ptr;
+ ptr += nvalues * sizeof(bool);
+
+ /* make sure we consumed the whole buffer exactly */
+ Assert((ptr - (char *) items) == len);
+
+ /* fix the pointers to Datum and bool arrays */
+ for (i = 0; i < numrows; i++)
+ {
+ items[i].values = &values[i * numattrs];
+ items[i].isnull = &isnull[i * numattrs];
+
+ /* load the values/null flags from sample rows */
+ for (j = 0; j < numattrs; j++)
+ {
+ items[i].values[j] = heap_getattr(rows[i],
+ attrs->values[j], /* attnum */
+ tdesc,
+ &items[i].isnull[j]); /* isnull */
+ }
+ }
+
+ /* do the sort, using the multi-sort */
+ qsort_arg((void *) items, numrows, sizeof(SortItem),
+ multi_sort_compare, mss);
+
+ return items;
+}
+
+/* count distinct combinations of SortItems in the array */
+static int
+count_distinct_groups(int numrows, SortItem *items, MultiSortSupport mss)
+{
+ int i;
+ int ndistinct;
+
+ ndistinct = 1;
+ for (i = 1; i < numrows; i++)
+ if (multi_sort_compare(&items[i], &items[i - 1], mss) != 0)
+ ndistinct += 1;
+
+ return ndistinct;
+}
+
+/* compares frequencies of the SortItem entries (in descending order) */
+static int
+compare_sort_item_count(const void *a, const void *b)
+{
+ SortItem *ia = (SortItem *) a;
+ SortItem *ib = (SortItem *) b;
+
+ if (ia->count == ib->count)
+ return 0;
+ else if (ia->count > ib->count)
+ return -1;
+
+ return 1;
+}
+
+/* builds SortItems for distinct groups and counts the matching items */
+static SortItem *
+build_distinct_groups(int numrows, SortItem *items, MultiSortSupport mss,
+ int *ndistinct)
+{
+ int i,
+ j;
+ int ngroups = count_distinct_groups(numrows, items, mss);
+
+ SortItem *groups = (SortItem *) palloc0(ngroups * sizeof(SortItem));
+
+ j = 0;
+ groups[0] = items[0];
+ groups[0].count = 1;
+
+ for (i = 1; i < numrows; i++)
+ {
+ if (multi_sort_compare(&items[i], &items[i - 1], mss) != 0)
+ groups[++j] = items[i];
+
+ groups[j].count++;
+ }
+
+ pg_qsort((void *) groups, ngroups, sizeof(SortItem),
+ compare_sort_item_count);
+
+ *ndistinct = ngroups;
+ return groups;
+}
+
+
+/* fetch the MCV list (as a bytea) from the pg_mv_statistic catalog */
+MCVList
+load_mv_mcvlist(Oid mvoid)
+{
+ bool isnull = false;
+ Datum mcvlist;
+
+#ifdef USE_ASSERT_CHECKING
+ Form_pg_mv_statistic mvstat;
+#endif
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ HeapTuple htup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(mvoid));
+
+ if (!HeapTupleIsValid(htup))
+ return NULL;
+
+#ifdef USE_ASSERT_CHECKING
+ mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
+ Assert(mvstat->mcv_enabled && mvstat->mcv_built);
+#endif
+
+ mcvlist = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_stamcv, &isnull);
+
+ Assert(!isnull);
+
+ ReleaseSysCache(htup);
+
+ return deserialize_mv_mcvlist(DatumGetByteaP(mcvlist));
+}
+
+/* print some basic info about the MCV list
+ *
+ * TODO: Add info about what part of the table this covers.
+ */
+Datum
+pg_mv_stats_mcvlist_info(PG_FUNCTION_ARGS)
+{
+ bytea *data = PG_GETARG_BYTEA_P(0);
+ char *result;
+
+ MCVList mcvlist = deserialize_mv_mcvlist(data);
+
+ result = palloc0(128);
+ snprintf(result, 128, "nitems=%d", mcvlist->nitems);
+
+ pfree(mcvlist);
+
+ PG_RETURN_TEXT_P(cstring_to_text(result));
+}
+
+/*
+ * serialize MCV list into a bytea value
+ *
+ *
+ * The basic algorithm is simple:
+ *
+ * (1) perform deduplication (for each attribute separately)
+ * (a) collect all (non-NULL) attribute values from all MCV items
+ * (b) sort the data (using 'lt' from VacAttrStats)
+ * (c) remove duplicate values from the array
+ *
+ * (2) serialize the arrays into a bytea value
+ *
+ * (3) process all MCV list items
+ * (a) replace values with indexes into the arrays
+ *
+ * Each attribute has to be processed separately, because we may be mixing
+ * different datatypes, with different sort operators, etc.
+ *
+ * We'll use uint16 values for the indexes in step (3), as we don't allow more
+ * than 8k MCV items, although that's mostly arbitrary limit. We might increase
+ * this to 65k and still fit into uint16.
+ *
+ * We don't really expect the serialization to save as much space as for
+ * histograms, because we are not doing any bucket splits (which is the source
+ * of high redundancy in histograms).
+ *
+ * TODO: Consider packing boolean flags (NULL) for each item into a single char
+ * (or a longer type) instead of using an array of bool items.
+ */
+bytea *
+serialize_mv_mcvlist(MCVList mcvlist, int2vector *attrs,
+ VacAttrStats **stats)
+{
+ int i,
+ j;
+ int ndims = mcvlist->ndimensions;
+ int itemsize = ITEM_SIZE(ndims);
+
+ SortSupport ssup;
+ DimensionInfo *info;
+
+ Size total_length;
+
+ /* allocate just once */
+ char *item = palloc0(itemsize);
+
+ /* serialized items (indexes into arrays, etc.) */
+ bytea *output;
+ char *data = NULL;
+
+ /* values per dimension (and number of non-NULL values) */
+ Datum **values = (Datum **) palloc0(sizeof(Datum *) * ndims);
+ int *counts = (int *) palloc0(sizeof(int) * ndims);
+
+ /*
+ * We'll include some rudimentary information about the attributes (type
+ * length, etc.), so that we don't have to look them up while
+ * deserializing the MCV list.
+ */
+ info = (DimensionInfo *) palloc0(sizeof(DimensionInfo) * ndims);
+
+ /* sort support data for all attributes included in the MCV list */
+ ssup = (SortSupport) palloc0(sizeof(SortSupportData) * ndims);
+
+ /* collect and deduplicate values for all attributes */
+ for (i = 0; i < ndims; i++)
+ {
+ int ndistinct;
+ StdAnalyzeData *tmp = (StdAnalyzeData *) stats[i]->extra_data;
+
+ /* copy important info about the data type (length, by-value) */
+ info[i].typlen = stats[i]->attrtype->typlen;
+ info[i].typbyval = stats[i]->attrtype->typbyval;
+
+ /* allocate space for values in the attribute and collect them */
+ values[i] = (Datum *) palloc0(sizeof(Datum) * mcvlist->nitems);
+
+ for (j = 0; j < mcvlist->nitems; j++)
+ {
+ /* skip NULL values - we don't need to serialize them */
+ if (mcvlist->items[j]->isnull[i])
+ continue;
+
+ values[i][counts[i]] = mcvlist->items[j]->values[i];
+ counts[i] += 1;
+ }
+
+ /* there are just NULL values in this dimension, we're done */
+ if (counts[i] == 0)
+ continue;
+
+ /* sort and deduplicate the data */
+ ssup[i].ssup_cxt = CurrentMemoryContext;
+ ssup[i].ssup_collation = DEFAULT_COLLATION_OID;
+ ssup[i].ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(tmp->ltopr, &ssup[i]);
+
+ qsort_arg(values[i], counts[i], sizeof(Datum),
+ compare_scalars_simple, &ssup[i]);
+
+ /*
+ * Walk through the array and eliminate duplicate values, but keep the
+ * ordering (so that we can do bsearch later). We know there's at
+ * least one item as (counts[i] != 0), so we can skip the first
+ * element.
+ */
+ ndistinct = 1; /* number of distinct values */
+ for (j = 1; j < counts[i]; j++)
+ {
+ /* if the value is the same as the previous one, we can skip it */
+ if (!compare_datums_simple(values[i][j - 1], values[i][j], &ssup[i]))
+ continue;
+
+ values[i][ndistinct] = values[i][j];
+ ndistinct += 1;
+ }
+
+ /* we must not exceed UINT16_MAX, as we use uint16 indexes */
+ Assert(ndistinct <= UINT16_MAX);
+
+ /*
+ * Store additional info about the attribute - number of deduplicated
+ * values, and also size of the serialized data. For fixed-length data
+ * types this is trivial to compute, for varwidth types we need to
+ * actually walk the array and sum the sizes.
+ */
+ info[i].nvalues = ndistinct;
+
+ if (info[i].typlen > 0) /* fixed-length data types */
+ info[i].nbytes = info[i].nvalues * info[i].typlen;
+ else if (info[i].typlen == -1) /* varlena */
+ {
+ info[i].nbytes = 0;
+ for (j = 0; j < info[i].nvalues; j++)
+ info[i].nbytes += VARSIZE_ANY(values[i][j]);
+ }
+ else if (info[i].typlen == -2) /* cstring */
+ {
+ info[i].nbytes = 0;
+ for (j = 0; j < info[i].nvalues; j++)
+ info[i].nbytes += strlen(DatumGetPointer(values[i][j]));
+ }
+
+ /* we know (count>0) so there must be some data */
+ Assert(info[i].nbytes > 0);
+ }
+
+ /*
+ * Now we can finally compute how much space we'll actually need for the
+ * serialized MCV list, as it contains these fields:
+ *
+ * - length (4B) for varlena - magic (4B) - type (4B) - ndimensions (4B) -
+ * nitems (4B) - info (ndim * sizeof(DimensionInfo) - arrays of values for
+ * each dimension - serialized items (nitems * itemsize)
+ *
+ * So the 'header' size is 20B + ndim * sizeof(DimensionInfo) and then we
+ * will place all the data (values + indexes).
+ */
+ total_length = (sizeof(int32) + offsetof(MCVListData, items)
+ +ndims * sizeof(DimensionInfo)
+ + mcvlist->nitems * itemsize);
+
+ for (i = 0; i < ndims; i++)
+ total_length += info[i].nbytes;
+
+ /* enforce arbitrary limit of 1MB */
+ if (total_length > (1024 * 1024))
+ elog(ERROR, "serialized MCV list exceeds 1MB (%ld)", total_length);
+
+ /* allocate space for the serialized MCV list, set header fields */
+ output = (bytea *) palloc0(total_length);
+ SET_VARSIZE(output, total_length);
+
+ /* 'data' points to the current position in the output buffer */
+ data = VARDATA(output);
+
+ /* MCV list header (number of items, ...) */
+ memcpy(data, mcvlist, offsetof(MCVListData, items));
+ data += offsetof(MCVListData, items);
+
+ /* information about the attributes */
+ memcpy(data, info, sizeof(DimensionInfo) * ndims);
+ data += sizeof(DimensionInfo) * ndims;
+
+ /* now serialize the deduplicated values for all attributes */
+ for (i = 0; i < ndims; i++)
+ {
+#ifdef USE_ASSERT_CHECKING
+ char *tmp = data; /* remember the starting point */
+#endif
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ Datum v = values[i][j];
+
+ if (info[i].typbyval) /* passed by value */
+ {
+ memcpy(data, &v, info[i].typlen);
+ data += info[i].typlen;
+ }
+ else if (info[i].typlen > 0) /* pased by reference */
+ {
+ memcpy(data, DatumGetPointer(v), info[i].typlen);
+ data += info[i].typlen;
+ }
+ else if (info[i].typlen == -1) /* varlena */
+ {
+ memcpy(data, DatumGetPointer(v), VARSIZE_ANY(v));
+ data += VARSIZE_ANY(v);
+ }
+ else if (info[i].typlen == -2) /* cstring */
+ {
+ memcpy(data, DatumGetPointer(v), strlen(DatumGetPointer(v)) + 1);
+ data += strlen(DatumGetPointer(v)) + 1; /* terminator */
+ }
+ }
+
+ /* make sure we got exactly the amount of data we expected */
+ Assert((data - tmp) == info[i].nbytes);
+ }
+
+ /* finally serialize the items, with uint16 indexes instead of the values */
+ for (i = 0; i < mcvlist->nitems; i++)
+ {
+ MCVItem mcvitem = mcvlist->items[i];
+
+ /* don't write beyond the allocated space */
+ Assert(data <= (char *) output + total_length - itemsize);
+
+ /* reset the item (we only allocate it once and reuse it) */
+ memset(item, 0, itemsize);
+
+ for (j = 0; j < ndims; j++)
+ {
+ Datum *v = NULL;
+
+ /* do the lookup only for non-NULL values */
+ if (mcvlist->items[i]->isnull[j])
+ continue;
+
+ v = (Datum *) bsearch_arg(&mcvitem->values[j], values[j],
+ info[j].nvalues, sizeof(Datum),
+ compare_scalars_simple, &ssup[j]);
+
+ Assert(v != NULL); /* serialization or deduplication error */
+
+ /* compute index within the array */
+ ITEM_INDEXES(item)[j] = (v - values[j]);
+
+ /* check the index is within expected bounds */
+ Assert(ITEM_INDEXES(item)[j] >= 0);
+ Assert(ITEM_INDEXES(item)[j] < info[j].nvalues);
+ }
+
+ /* copy NULL and frequency flags into the item */
+ memcpy(ITEM_NULLS(item, ndims), mcvitem->isnull, sizeof(bool) * ndims);
+ memcpy(ITEM_FREQUENCY(item, ndims), &mcvitem->frequency, sizeof(double));
+
+ /* copy the serialized item into the array */
+ memcpy(data, item, itemsize);
+
+ data += itemsize;
+ }
+
+ /* at this point we expect to match the total_length exactly */
+ Assert((data - (char *) output) == total_length);
+
+ return output;
+}
+
+/*
+ * deserialize MCV list from the varlena value
+ *
+ *
+ * We deserialize the MCV list fully, because we don't expect there bo be a lot
+ * of duplicate values. But perhaps we should keep the MCV in serialized form
+ * just like histograms.
+ */
+MCVList
+deserialize_mv_mcvlist(bytea *data)
+{
+ int i,
+ j;
+ Size expected_size;
+ MCVList mcvlist;
+ char *tmp;
+
+ int ndims,
+ nitems,
+ itemsize;
+ DimensionInfo *info = NULL;
+
+ uint16 *indexes = NULL;
+ Datum **values = NULL;
+
+ /* local allocation buffer (used only for deserialization) */
+ int bufflen;
+ char *buff;
+ char *ptr;
+
+ /* buffer used for the result */
+ int rbufflen;
+ char *rbuff;
+ char *rptr;
+
+ if (data == NULL)
+ return NULL;
+
+ /* we can't deserialize the MCV if there's not even a complete header */
+ expected_size = offsetof(MCVListData, items);
+
+ if (VARSIZE_ANY_EXHDR(data) < expected_size)
+ elog(ERROR, "invalid MCV Size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), offsetof(MCVListData, items));
+
+ /* read the MCV list header */
+ mcvlist = (MCVList) palloc0(sizeof(MCVListData));
+
+ /* initialize pointer to the data part (skip the varlena header) */
+ tmp = VARDATA_ANY(data);
+
+ /* get the header and perform further sanity checks */
+ memcpy(mcvlist, tmp, offsetof(MCVListData, items));
+ tmp += offsetof(MCVListData, items);
+
+ if (mcvlist->magic != MVSTAT_MCV_MAGIC)
+ elog(ERROR, "invalid MCV magic %d (expected %dd)",
+ mcvlist->magic, MVSTAT_MCV_MAGIC);
+
+ if (mcvlist->type != MVSTAT_MCV_TYPE_BASIC)
+ elog(ERROR, "invalid MCV type %d (expected %dd)",
+ mcvlist->type, MVSTAT_MCV_TYPE_BASIC);
+
+ nitems = mcvlist->nitems;
+ ndims = mcvlist->ndimensions;
+ itemsize = ITEM_SIZE(ndims);
+
+ Assert((nitems > 0) && (nitems <= MVSTAT_MCVLIST_MAX_ITEMS));
+ Assert((ndims >= 2) && (ndims <= MVSTATS_MAX_DIMENSIONS));
+
+ /*
+ * Check amount of data including DimensionInfo for all dimensions and
+ * also the serialized items (including uint16 indexes). Also, walk
+ * through the dimension information and add it to the sum.
+ */
+ expected_size += ndims * sizeof(DimensionInfo) +
+ (nitems * itemsize);
+
+ /* check that we have at least the DimensionInfo records */
+ if (VARSIZE_ANY_EXHDR(data) < expected_size)
+ elog(ERROR, "invalid MCV size %ld (expected %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ info = (DimensionInfo *) (tmp);
+ tmp += ndims * sizeof(DimensionInfo);
+
+ /* account for the value arrays */
+ for (i = 0; i < ndims; i++)
+ {
+ Assert(info[i].nvalues >= 0);
+ Assert(info[i].nbytes >= 0);
+
+ expected_size += info[i].nbytes;
+ }
+
+ if (VARSIZE_ANY_EXHDR(data) != expected_size)
+ elog(ERROR, "invalid MCV size %ld (expected %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ /* looks OK - not corrupted or something */
+
+ /*
+ * Allocate one large chunk of memory for the intermediate data, needed
+ * only for deserializing the MCV list (and allocate densely to minimize
+ * the palloc overhead).
+ *
+ * Let's see how much space we'll actually need, and also include space
+ * for the array with pointers.
+ */
+ bufflen = sizeof(Datum *) * ndims; /* space for pointers */
+
+ for (i = 0; i < ndims; i++)
+ /* for full-size byval types, we reuse the serialized value */
+ if (!(info[i].typbyval && info[i].typlen == sizeof(Datum)))
+ bufflen += (sizeof(Datum) * info[i].nvalues);
+
+ buff = palloc0(bufflen);
+ ptr = buff;
+
+ values = (Datum **) buff;
+ ptr += (sizeof(Datum *) * ndims);
+
+ /*
+ * XXX This uses pointers to the original data array (the types not passed
+ * by value), so when someone frees the memory, e.g. by doing something
+ * like this:
+ *
+ * bytea * data = ... fetch the data from catalog ... MCVList mcvlist =
+ * deserialize_mcv_list(data); pfree(data);
+ *
+ * then 'mcvlist' references the freed memory. Should copy the pieces.
+ */
+ for (i = 0; i < ndims; i++)
+ {
+ if (info[i].typbyval)
+ {
+ /* passed by value / Datum - simply reuse the array */
+ if (info[i].typlen == sizeof(Datum))
+ {
+ values[i] = (Datum *) tmp;
+ tmp += info[i].nbytes;
+ }
+ else
+ {
+ values[i] = (Datum *) ptr;
+ ptr += (sizeof(Datum) * info[i].nvalues);
+
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ memcpy(&values[i][j], tmp, info[i].typlen);
+ tmp += info[i].typlen;
+ }
+ }
+ }
+ else
+ {
+ /* all the other types need a chunk of the buffer */
+ values[i] = (Datum *) ptr;
+ ptr += (sizeof(Datum) * info[i].nvalues);
+
+ /* pased by reference, but fixed length (name, tid, ...) */
+ if (info[i].typlen > 0)
+ {
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ values[i][j] = PointerGetDatum(tmp);
+ tmp += info[i].typlen;
+ }
+ }
+ else if (info[i].typlen == -1)
+ {
+ /* varlena */
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ values[i][j] = PointerGetDatum(tmp);
+ tmp += VARSIZE_ANY(tmp);
+ }
+ }
+ else if (info[i].typlen == -2)
+ {
+ /* cstring */
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ values[i][j] = PointerGetDatum(tmp);
+ tmp += (strlen(tmp) + 1); /* don't forget the \0 */
+ }
+ }
+ }
+ }
+
+ /* we should have exhausted the buffer exactly */
+ Assert((ptr - buff) == bufflen);
+
+ /* allocate space for all the MCV items in a single piece */
+ rbufflen = (sizeof(MCVItem) + sizeof(MCVItemData) +
+ sizeof(Datum) * ndims + sizeof(bool) * ndims) * nitems;
+
+ rbuff = palloc0(rbufflen);
+ rptr = rbuff;
+
+ mcvlist->items = (MCVItem *) rbuff;
+ rptr += (sizeof(MCVItem) * nitems);
+
+ for (i = 0; i < nitems; i++)
+ {
+ MCVItem item = (MCVItem) rptr;
+
+ rptr += (sizeof(MCVItemData));
+
+ item->values = (Datum *) rptr;
+ rptr += (sizeof(Datum) * ndims);
+
+ item->isnull = (bool *) rptr;
+ rptr += (sizeof(bool) * ndims);
+
+ /* just point to the right place */
+ indexes = ITEM_INDEXES(tmp);
+
+ memcpy(item->isnull, ITEM_NULLS(tmp, ndims), sizeof(bool) * ndims);
+ memcpy(&item->frequency, ITEM_FREQUENCY(tmp, ndims), sizeof(double));
+
+#ifdef ASSERT_CHECKING
+ for (j = 0; j < ndims; j++)
+ Assert(indexes[j] <= UINT16_MAX);
+#endif
+
+ /* translate the values */
+ for (j = 0; j < ndims; j++)
+ if (!item->isnull[j])
+ item->values[j] = values[j][indexes[j]];
+
+ mcvlist->items[i] = item;
+
+ tmp += ITEM_SIZE(ndims);
+
+ Assert(tmp <= (char *) data + VARSIZE_ANY(data));
+ }
+
+ /* check that we processed all the data */
+ Assert(tmp == (char *) data + VARSIZE_ANY(data));
+
+ /* release the temporary buffer */
+ pfree(buff);
+
+ return mcvlist;
+}
+
+/*
+ * SRF with details about buckets of a histogram:
+ *
+ * - item ID (0...nitems)
+ * - values (string array)
+ * - nulls only (boolean array)
+ * - frequency (double precision)
+ *
+ * The input is the OID of the statistics, and there are no rows returned if
+ * the statistics contains no histogram.
+ */
+PG_FUNCTION_INFO_V1(pg_mv_mcv_items);
+
+Datum
+pg_mv_mcv_items(PG_FUNCTION_ARGS)
+{
+ FuncCallContext *funcctx;
+ int call_cntr;
+ int max_calls;
+ TupleDesc tupdesc;
+ AttInMetadata *attinmeta;
+
+ /* stuff done only on the first call of the function */
+ if (SRF_IS_FIRSTCALL())
+ {
+ MemoryContext oldcontext;
+ MCVList mcvlist;
+
+ /* create a function context for cross-call persistence */
+ funcctx = SRF_FIRSTCALL_INIT();
+
+ /* switch to memory context appropriate for multiple function calls */
+ oldcontext = MemoryContextSwitchTo(funcctx->multi_call_memory_ctx);
+
+ mcvlist = load_mv_mcvlist(PG_GETARG_OID(0));
+
+ funcctx->user_fctx = mcvlist;
+
+ /* total number of tuples to be returned */
+ funcctx->max_calls = 0;
+ if (funcctx->user_fctx != NULL)
+ funcctx->max_calls = mcvlist->nitems;
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("function returning record called in context "
+ "that cannot accept type record")));
+
+ /* build metadata needed later to produce tuples from raw C-strings */
+ attinmeta = TupleDescGetAttInMetadata(tupdesc);
+ funcctx->attinmeta = attinmeta;
+
+ MemoryContextSwitchTo(oldcontext);
+ }
+
+ /* stuff done on every call of the function */
+ funcctx = SRF_PERCALL_SETUP();
+
+ call_cntr = funcctx->call_cntr;
+ max_calls = funcctx->max_calls;
+ attinmeta = funcctx->attinmeta;
+
+ if (call_cntr < max_calls) /* do when there is more left to send */
+ {
+ char **values;
+ HeapTuple tuple;
+ Datum result;
+ int2vector *stakeys;
+ Oid relid;
+
+ char *buff = palloc0(1024);
+ char *format;
+
+ int i;
+
+ Oid *outfuncs;
+ FmgrInfo *fmgrinfo;
+
+ MCVList mcvlist;
+ MCVItem item;
+
+ mcvlist = (MCVList) funcctx->user_fctx;
+
+ Assert(call_cntr < mcvlist->nitems);
+
+ item = mcvlist->items[call_cntr];
+
+ stakeys = find_mv_attnums(PG_GETARG_OID(0), &relid);
+
+ /*
+ * Prepare a values array for building the returned tuple. This should
+ * be an array of C strings which will be processed later by the type
+ * input functions.
+ */
+ values = (char **) palloc(4 * sizeof(char *));
+
+ values[0] = (char *) palloc(64 * sizeof(char));
+
+ /* arrays */
+ values[1] = (char *) palloc0(1024 * sizeof(char));
+ values[2] = (char *) palloc0(1024 * sizeof(char));
+
+ /* frequency */
+ values[3] = (char *) palloc(64 * sizeof(char));
+
+ outfuncs = (Oid *) palloc0(sizeof(Oid) * mcvlist->ndimensions);
+ fmgrinfo = (FmgrInfo *) palloc0(sizeof(FmgrInfo) * mcvlist->ndimensions);
+
+ for (i = 0; i < mcvlist->ndimensions; i++)
+ {
+ bool isvarlena;
+
+ getTypeOutputInfo(get_atttype(relid, stakeys->values[i]),
+ &outfuncs[i], &isvarlena);
+
+ fmgr_info(outfuncs[i], &fmgrinfo[i]);
+ }
+
+ snprintf(values[0], 64, "%d", call_cntr); /* item ID */
+
+ for (i = 0; i < mcvlist->ndimensions; i++)
+ {
+ Datum val,
+ valout;
+
+ format = "%s, %s";
+ if (i == 0)
+ format = "{%s%s";
+ else if (i == mcvlist->ndimensions - 1)
+ format = "%s, %s}";
+
+ if (item->isnull[i])
+ valout = CStringGetDatum("NULL");
+ else
+ {
+ val = item->values[i];
+ valout = FunctionCall1(&fmgrinfo[i], val);
+ }
+
+ snprintf(buff, 1024, format, values[1], DatumGetPointer(valout));
+ strncpy(values[1], buff, 1023);
+ buff[0] = '\0';
+
+ snprintf(buff, 1024, format, values[2], item->isnull[i] ? "t" : "f");
+ strncpy(values[2], buff, 1023);
+ buff[0] = '\0';
+ }
+
+ snprintf(values[3], 64, "%f", item->frequency); /* frequency */
+
+ /* build a tuple */
+ tuple = BuildTupleFromCStrings(attinmeta, values);
+
+ /* make the tuple into a datum */
+ result = HeapTupleGetDatum(tuple);
+
+ /* clean up (this is not really necessary) */
+ pfree(values[0]);
+ pfree(values[1]);
+ pfree(values[2]);
+ pfree(values[3]);
+
+ pfree(values);
+
+ SRF_RETURN_NEXT(funcctx, result);
+ }
+ else /* do when there is no more left */
+ {
+ SRF_RETURN_DONE(funcctx);
+ }
+}
+
+/*
+ * pg_mcv_list_in - input routine for type PG_MCV_LIST.
+ *
+ * pg_mcv_list is real enough to be a table column, but it has no operations
+ * of its own, and disallows input too
+ *
+ * XXX This is inspired by what pg_node_tree does.
+ */
+Datum
+pg_mcv_list_in(PG_FUNCTION_ARGS)
+{
+ /*
+ * pg_node_list stores the data in binary form and parsing text input is
+ * not needed, so disallow this.
+ */
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_mcv_list")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+
+/*
+ * pg_mcv_list_out - output routine for type PG_MCV_LIST.
+ *
+ * MCV lists are serialized into a bytea value, so we simply call byteaout()
+ * to serialize the value into text. But it'd be nice to serialize that into
+ * a meaningful representation (e.g. for inspection by people).
+ *
+ * FIXME not implemented yet, returning dummy value
+ */
+Datum
+pg_mcv_list_out(PG_FUNCTION_ARGS)
+{
+ return byteaout(fcinfo);
+}
+
+/*
+ * pg_mcv_list_recv - binary input routine for type PG_MCV_LIST.
+ */
+Datum
+pg_mcv_list_recv(PG_FUNCTION_ARGS)
+{
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_mcv_list")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_mcv_list_send - binary output routine for type PG_MCV_LIST.
+ *
+ * XXX MCV lists are serialized into a bytea value, so let's just send that.
+ */
+Datum
+pg_mcv_list_send(PG_FUNCTION_ARGS)
+{
+ return byteasend(fcinfo);
+}
diff --git a/src/bin/psql/describe.c b/src/bin/psql/describe.c
index e7d5b51..db74d93 100644
--- a/src/bin/psql/describe.c
+++ b/src/bin/psql/describe.c
@@ -2298,8 +2298,8 @@ describeOneTableDetails(const char *schemaname,
{
printfPQExpBuffer(&buf,
"SELECT oid, stanamespace::regnamespace AS nsp, staname, stakeys,\n"
- " ndist_enabled,\n"
- " ndist_built,\n"
+ " ndist_enabled, deps_enabled, mcv_enabled,\n"
+ " ndist_built, deps_built, mcv_built,\n"
" (SELECT string_agg(attname::text,', ')\n"
" FROM ((SELECT unnest(stakeys) AS attnum) s\n"
" JOIN pg_attribute a ON (starelid = a.attrelid and a.attnum = s.attnum))) AS attnums\n"
@@ -2317,6 +2317,8 @@ describeOneTableDetails(const char *schemaname,
printTableAddFooter(&cont, _("Statistics:"));
for (i = 0; i < tuples; i++)
{
+ bool first = true;
+
printfPQExpBuffer(&buf, " ");
/* statistics name (qualified with namespace) */
@@ -2326,10 +2328,22 @@ describeOneTableDetails(const char *schemaname,
/* options */
if (!strcmp(PQgetvalue(result, i, 4), "t"))
- appendPQExpBuffer(&buf, "(dependencies)");
+ {
+ appendPQExpBuffer(&buf, "(dependencies");
+ first = false;
+ }
+
+ if (!strcmp(PQgetvalue(result, i, 5), "t"))
+ {
+ if (!first)
+ appendPQExpBuffer(&buf, ", mcv");
+ else
+ appendPQExpBuffer(&buf, "(mcv");
+ first = false;
+ }
- appendPQExpBuffer(&buf, " ON (%s)",
- PQgetvalue(result, i, 6));
+ appendPQExpBuffer(&buf, ") ON (%s)",
+ PQgetvalue(result, i, 9));
printTableAddFooter(&cont, buf.data);
}
diff --git a/src/include/catalog/pg_cast.h b/src/include/catalog/pg_cast.h
index 22fa4b8..80d8ea2 100644
--- a/src/include/catalog/pg_cast.h
+++ b/src/include/catalog/pg_cast.h
@@ -262,6 +262,11 @@ DATA(insert ( 3353 25 0 i i ));
DATA(insert ( 3358 17 0 i b ));
DATA(insert ( 3358 25 0 i i ));
+/* pg_mcv_list can be coerced to, but not from, bytea and text */
+DATA(insert ( 441 17 0 i b ));
+DATA(insert ( 441 25 0 i i ));
+
+
/*
* Datetime category
*/
diff --git a/src/include/catalog/pg_mv_statistic.h b/src/include/catalog/pg_mv_statistic.h
index e119cb7..34049d6 100644
--- a/src/include/catalog/pg_mv_statistic.h
+++ b/src/include/catalog/pg_mv_statistic.h
@@ -39,10 +39,12 @@ CATALOG(pg_mv_statistic,3381)
/* statistics requested to build */
bool ndist_enabled; /* build ndist coefficient? */
bool deps_enabled; /* analyze dependencies? */
+ bool mcv_enabled; /* build MCV list? */
/* statistics that are available (if requested) */
bool ndist_built; /* ndistinct coeff built */
bool deps_built; /* dependencies were built */
+ bool mcv_built; /* MCV list was built */
/*
* variable-length fields start here, but we allow direct access to
@@ -53,6 +55,7 @@ CATALOG(pg_mv_statistic,3381)
#ifdef CATALOG_VARLEN
pg_ndistinct standist; /* ndistinct coeff (serialized) */
pg_dependencies stadeps; /* dependencies (serialized) */
+ pg_mcv_list stamcv; /* MCV list (serialized) */
#endif
} FormData_pg_mv_statistic;
@@ -68,17 +71,20 @@ typedef FormData_pg_mv_statistic *Form_pg_mv_statistic;
* compiler constants for pg_mv_statistic
* ----------------
*/
-#define Natts_pg_mv_statistic 11
+#define Natts_pg_mv_statistic 14
#define Anum_pg_mv_statistic_starelid 1
#define Anum_pg_mv_statistic_staname 2
#define Anum_pg_mv_statistic_stanamespace 3
#define Anum_pg_mv_statistic_staowner 4
#define Anum_pg_mv_statistic_ndist_enabled 5
#define Anum_pg_mv_statistic_deps_enabled 6
-#define Anum_pg_mv_statistic_ndist_built 7
-#define Anum_pg_mv_statistic_deps_built 8
-#define Anum_pg_mv_statistic_stakeys 9
-#define Anum_pg_mv_statistic_standist 10
-#define Anum_pg_mv_statistic_stadeps 11
+#define Anum_pg_mv_statistic_mcv_enabled 7
+#define Anum_pg_mv_statistic_ndist_built 8
+#define Anum_pg_mv_statistic_deps_built 9
+#define Anum_pg_mv_statistic_mcv_built 10
+#define Anum_pg_mv_statistic_stakeys 11
+#define Anum_pg_mv_statistic_standist 12
+#define Anum_pg_mv_statistic_stadeps 13
+#define Anum_pg_mv_statistic_stamcv 14
#endif /* PG_MV_STATISTIC_H */
diff --git a/src/include/catalog/pg_proc.h b/src/include/catalog/pg_proc.h
index b1f7b75..7cf1e5a 100644
--- a/src/include/catalog/pg_proc.h
+++ b/src/include/catalog/pg_proc.h
@@ -2726,6 +2726,11 @@ DESCR("current user privilege on any column by rel name");
DATA(insert OID = 3029 ( has_any_column_privilege PGNSP PGUID 12 10 0 0 0 f f f f t f s s 2 0 16 "26 25" _null_ _null_ _null_ _null_ _null_ has_any_column_privilege_id _null_ _null_ _null_ ));
DESCR("current user privilege on any column by rel oid");
+DATA(insert OID = 3376 ( pg_mv_stats_mcvlist_info PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 25 "441" _null_ _null_ _null_ _null_ _null_ pg_mv_stats_mcvlist_info _null_ _null_ _null_ ));
+DESCR("multi-variate statistics: MCV list info");
+DATA(insert OID = 3373 ( pg_mv_mcv_items PGNSP PGUID 12 1 1000 0 0 f f f f t t i s 1 0 2249 "26" "{26,23,1009,1000,701}" "{i,o,o,o,o}" "{oid,index,values,nulls,frequency}" _null_ _null_ pg_mv_mcv_items _null_ _null_ _null_ ));
+DESCR("details about MCV list items");
+
DATA(insert OID = 3354 ( pg_ndistinct_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 3353 "2275" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_in _null_ _null_ _null_ ));
DESCR("I/O");
DATA(insert OID = 3355 ( pg_ndistinct_out PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 2275 "3353" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_out _null_ _null_ _null_ ));
@@ -2744,6 +2749,15 @@ DESCR("I/O");
DATA(insert OID = 3362 ( pg_dependencies_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "3358" _null_ _null_ _null_ _null_ _null_ pg_dependencies_send _null_ _null_ _null_ ));
DESCR("I/O");
+DATA(insert OID = 442 ( pg_mcv_list_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 441 "2275" _null_ _null_ _null_ _null_ _null_ pg_mcv_list_in _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 443 ( pg_mcv_list_out PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 2275 "441" _null_ _null_ _null_ _null_ _null_ pg_mcv_list_out _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 444 ( pg_mcv_list_recv PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 441 "2281" _null_ _null_ _null_ _null_ _null_ pg_mcv_list_recv _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 445 ( pg_mcv_list_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "441" _null_ _null_ _null_ _null_ _null_ pg_mcv_list_send _null_ _null_ _null_ ));
+DESCR("I/O");
+
DATA(insert OID = 1928 ( pg_stat_get_numscans PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_numscans _null_ _null_ _null_ ));
DESCR("statistics: number of scans done for table/index");
DATA(insert OID = 1929 ( pg_stat_get_tuples_returned PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_tuples_returned _null_ _null_ _null_ ));
diff --git a/src/include/catalog/pg_type.h b/src/include/catalog/pg_type.h
index da637d4..fbac135 100644
--- a/src/include/catalog/pg_type.h
+++ b/src/include/catalog/pg_type.h
@@ -372,6 +372,10 @@ DATA(insert OID = 3358 ( pg_dependencies PGNSP PGUID -1 f b S f t \054 0 0 0 pg
DESCR("multivariate histogram");
#define PGDEPENDENCIESOID 3358
+DATA(insert OID = 441 ( pg_mcv_list PGNSP PGUID -1 f b S f t \054 0 0 0 pg_mcv_list_in pg_mcv_list_out pg_mcv_list_recv pg_mcv_list_send - - - i x f 0 -1 0 100 _null_ _null_ _null_ ));
+DESCR("multivariate MCV list");
+#define PGMCVLISTOID 441
+
DATA(insert OID = 32 ( pg_ddl_command PGNSP PGUID SIZEOF_POINTER t p P f t \054 0 0 0 pg_ddl_command_in pg_ddl_command_out pg_ddl_command_recv pg_ddl_command_send - - - ALIGNOF_POINTER p f 0 -1 0 0 _null_ _null_ _null_ ));
DESCR("internal type for passing CollectedCommand");
#define PGDDLCOMMANDOID 32
diff --git a/src/include/nodes/relation.h b/src/include/nodes/relation.h
index 56957e8..d912827 100644
--- a/src/include/nodes/relation.h
+++ b/src/include/nodes/relation.h
@@ -681,12 +681,14 @@ typedef struct MVStatisticInfo
RelOptInfo *rel; /* back-link to index's table */
/* enabled statistics */
- bool deps_enabled; /* functional dependencies enabled */
bool ndist_enabled; /* ndistinct coefficient enabled */
+ bool deps_enabled; /* functional dependencies enabled */
+ bool mcv_enabled; /* MCV list enabled */
/* built/available statistics */
- bool deps_built; /* functional dependencies built */
bool ndist_built; /* ndistinct coefficient built */
+ bool deps_built; /* functional dependencies built */
+ bool mcv_built; /* MCV list built */
/* columns in the statistics (attnums) */
int2vector *stakeys; /* attnums of the columns covered */
diff --git a/src/include/utils/builtins.h b/src/include/utils/builtins.h
index 9ffd80c..9ed080a 100644
--- a/src/include/utils/builtins.h
+++ b/src/include/utils/builtins.h
@@ -77,6 +77,10 @@ extern Datum pg_dependencies_in(PG_FUNCTION_ARGS);
extern Datum pg_dependencies_out(PG_FUNCTION_ARGS);
extern Datum pg_dependencies_recv(PG_FUNCTION_ARGS);
extern Datum pg_dependencies_send(PG_FUNCTION_ARGS);
+extern Datum pg_mcv_list_in(PG_FUNCTION_ARGS);
+extern Datum pg_mcv_list_out(PG_FUNCTION_ARGS);
+extern Datum pg_mcv_list_recv(PG_FUNCTION_ARGS);
+extern Datum pg_mcv_list_send(PG_FUNCTION_ARGS);
/* regexp.c */
extern char *regexp_fixed_prefix(text *text_re, bool case_insensitive,
diff --git a/src/include/utils/mvstats.h b/src/include/utils/mvstats.h
index b230747..0c4f621 100644
--- a/src/include/utils/mvstats.h
+++ b/src/include/utils/mvstats.h
@@ -17,6 +17,14 @@
#include "fmgr.h"
#include "commands/vacuum.h"
+/*
+ * Degree of how much MCV item matches a clause.
+ * This is then considered when computing the selectivity.
+ */
+#define MVSTATS_MATCH_NONE 0 /* no match at all */
+#define MVSTATS_MATCH_PARTIAL 1 /* partial match */
+#define MVSTATS_MATCH_FULL 2 /* full match */
+
#define MVSTATS_MAX_DIMENSIONS 8 /* max number of attributes */
#define MVSTAT_NDISTINCT_MAGIC 0xA352BFA4 /* marks serialized bytea */
@@ -65,6 +73,42 @@ typedef struct MVDependenciesData
typedef MVDependenciesData *MVDependencies;
+
+/* used to flag stats serialized to bytea */
+#define MVSTAT_MCV_MAGIC 0xE1A651C2 /* marks serialized bytea */
+#define MVSTAT_MCV_TYPE_BASIC 1 /* basic MCV list type */
+
+/* max items in MCV list (mostly arbitrary number */
+#define MVSTAT_MCVLIST_MAX_ITEMS 8192
+
+/*
+ * Multivariate MCV (most-common value) lists
+ *
+ * A straight-forward extension of MCV items - i.e. a list (array) of
+ * combinations of attribute values, together with a frequency and
+ * null flags.
+ */
+typedef struct MCVItemData
+{
+ double frequency; /* frequency of this combination */
+ bool *isnull; /* lags of NULL values (up to 32 columns) */
+ Datum *values; /* variable-length (ndimensions) */
+} MCVItemData;
+
+typedef MCVItemData *MCVItem;
+
+/* multivariate MCV list - essentally an array of MCV items */
+typedef struct MCVListData
+{
+ uint32 magic; /* magic constant marker */
+ uint32 type; /* type of MCV list (BASIC) */
+ uint32 ndimensions; /* number of dimensions */
+ uint32 nitems; /* number of MCV items in the array */
+ MCVItem *items; /* array of MCV items */
+} MCVListData;
+
+typedef MCVListData *MCVList;
+
bool dependency_implies_attribute(MVDependency dependency, AttrNumber attnum,
int16 *attmap);
bool dependency_is_fully_matched(MVDependency dependency, Bitmapset *attnums,
@@ -72,13 +116,30 @@ bool dependency_is_fully_matched(MVDependency dependency, Bitmapset *attnums,
MVNDistinct load_mv_ndistinct(Oid mvoid);
MVDependencies load_mv_dependencies(Oid mvoid);
+MCVList load_mv_mcvlist(Oid mvoid);
bytea *serialize_mv_ndistinct(MVNDistinct ndistinct);
bytea *serialize_mv_dependencies(MVDependencies dependencies);
+bytea *serialize_mv_mcvlist(MCVList mcvlist, int2vector *attrs,
+ VacAttrStats **stats);
/* deserialization of stats (serialization is private to analyze) */
MVNDistinct deserialize_mv_ndistinct(bytea *data);
MVDependencies deserialize_mv_dependencies(bytea *data);
+MCVList deserialize_mv_mcvlist(bytea *data);
+
+/*
+ * Returns index of the attribute number within the vector (i.e. a
+ * dimension within the stats).
+ */
+int mv_get_index(AttrNumber varattno, int2vector *stakeys);
+
+int2vector *find_mv_attnums(Oid mvoid, Oid *relid);
+
+/* functions for inspecting the statistics */
+extern Datum pg_mv_stats_mcvlist_info(PG_FUNCTION_ARGS);
+extern Datum pg_mv_mcvlist_items(PG_FUNCTION_ARGS);
+
MVNDistinct build_mv_ndistinct(double totalrows, int numrows, HeapTuple *rows,
int2vector *attrs, VacAttrStats **stats);
@@ -87,6 +148,9 @@ MVDependencies build_mv_dependencies(int numrows, HeapTuple *rows,
int2vector *attrs,
VacAttrStats **stats);
+MCVList build_mv_mcvlist(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int *numrows_filtered);
+
void build_mv_stats(Relation onerel, double totalrows,
int numrows, HeapTuple *rows,
int natts, VacAttrStats **vacattrstats);
diff --git a/src/test/regress/expected/mv_mcv.out b/src/test/regress/expected/mv_mcv.out
new file mode 100644
index 0000000..d8ba619
--- /dev/null
+++ b/src/test/regress/expected/mv_mcv.out
@@ -0,0 +1,198 @@
+-- data type passed by value
+CREATE TABLE mcv_list (
+ a INT,
+ b INT,
+ c INT
+);
+-- unknown column
+CREATE STATISTICS s4 WITH (mcv) ON (unknown_column) FROM mcv_list;
+ERROR: column "unknown_column" referenced in statistics does not exist
+-- single column
+CREATE STATISTICS s4 WITH (mcv) ON (a) FROM mcv_list;
+ERROR: statistics require at least 2 columns
+-- single column, duplicated
+CREATE STATISTICS s4 WITH (mcv) ON (a, a) FROM mcv_list;
+ERROR: duplicate column name in statistics definition
+-- two columns, one duplicated
+CREATE STATISTICS s4 WITH (mcv) ON (a, a, b) FROM mcv_list;
+ERROR: duplicate column name in statistics definition
+-- unknown option
+CREATE STATISTICS s4 WITH (unknown_option) ON (a, b, c) FROM mcv_list;
+ERROR: unrecognized STATISTICS option "unknown_option"
+-- correct command
+CREATE STATISTICS s4 WITH (mcv) ON (a, b, c) FROM mcv_list;
+-- random data
+INSERT INTO mcv_list
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | f |
+(1 row)
+
+TRUNCATE mcv_list;
+-- a => b, a => c, b => c
+INSERT INTO mcv_list
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=1000
+(1 row)
+
+TRUNCATE mcv_list;
+-- a => b, a => c
+INSERT INTO mcv_list
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=1000
+(1 row)
+
+TRUNCATE mcv_list;
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mcv_list
+ SELECT i/100, i/200, i/400 FROM generate_series(1,10000) s(i);
+CREATE INDEX mcv_idx ON mcv_list (a, b);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=100
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a = 10 AND b = 5;
+ QUERY PLAN
+--------------------------------------------
+ Bitmap Heap Scan on mcv_list
+ Recheck Cond: ((a = 10) AND (b = 5))
+ -> Bitmap Index Scan on mcv_idx
+ Index Cond: ((a = 10) AND (b = 5))
+(4 rows)
+
+DROP TABLE mcv_list;
+-- varlena type (text)
+CREATE TABLE mcv_list (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+CREATE STATISTICS s5 WITH (mcv) ON (a, b, c) FROM mcv_list;
+-- random data
+INSERT INTO mcv_list
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | f |
+(1 row)
+
+TRUNCATE mcv_list;
+-- a => b, a => c, b => c
+INSERT INTO mcv_list
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=1000
+(1 row)
+
+TRUNCATE mcv_list;
+-- a => b, a => c
+INSERT INTO mcv_list
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=1000
+(1 row)
+
+TRUNCATE mcv_list;
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mcv_list
+ SELECT i/100, i/200, i/400 FROM generate_series(1,10000) s(i);
+CREATE INDEX mcv_idx ON mcv_list (a, b);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=100
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a = '10' AND b = '5';
+ QUERY PLAN
+------------------------------------------------------------
+ Bitmap Heap Scan on mcv_list
+ Recheck Cond: ((a = '10'::text) AND (b = '5'::text))
+ -> Bitmap Index Scan on mcv_idx
+ Index Cond: ((a = '10'::text) AND (b = '5'::text))
+(4 rows)
+
+TRUNCATE mcv_list;
+-- check explain (expect bitmap index scan, not plain index scan) with NULLs
+INSERT INTO mcv_list
+ SELECT
+ (CASE WHEN i/100 = 0 THEN NULL ELSE i/100 END),
+ (CASE WHEN i/200 = 0 THEN NULL ELSE i/200 END),
+ (CASE WHEN i/400 = 0 THEN NULL ELSE i/400 END)
+ FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=100
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a IS NULL AND b IS NULL;
+ QUERY PLAN
+---------------------------------------------------
+ Bitmap Heap Scan on mcv_list
+ Recheck Cond: ((a IS NULL) AND (b IS NULL))
+ -> Bitmap Index Scan on mcv_idx
+ Index Cond: ((a IS NULL) AND (b IS NULL))
+(4 rows)
+
+DROP TABLE mcv_list;
+-- NULL values (mix of int and text columns)
+CREATE TABLE mcv_list (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+CREATE STATISTICS s6 WITH (mcv) ON (a, b, c, d) FROM mcv_list;
+INSERT INTO mcv_list
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+ mcv_enabled | mcv_built | pg_mv_stats_mcvlist_info
+-------------+-----------+--------------------------
+ t | t | nitems=1200
+(1 row)
+
+DROP TABLE mcv_list;
diff --git a/src/test/regress/expected/opr_sanity.out b/src/test/regress/expected/opr_sanity.out
index db1cf8a..9969c10 100644
--- a/src/test/regress/expected/opr_sanity.out
+++ b/src/test/regress/expected/opr_sanity.out
@@ -819,11 +819,12 @@ WHERE c.castmethod = 'b' AND
pg_node_tree | text | 0 | i
pg_ndistinct | bytea | 0 | i
pg_dependencies | bytea | 0 | i
+ pg_mcv_list | bytea | 0 | i
cidr | inet | 0 | i
xml | text | 0 | a
xml | character varying | 0 | a
xml | character | 0 | a
-(9 rows)
+(10 rows)
-- **************** pg_conversion ****************
-- Look for illegal values in pg_conversion fields.
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 39179a6..2e3c40e 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1381,7 +1381,9 @@ pg_mv_stats| SELECT n.nspname AS schemaname,
s.staname,
s.stakeys AS attnums,
length((s.standist)::bytea) AS ndistbytes,
- length((s.stadeps)::bytea) AS depsbytes
+ length((s.stadeps)::bytea) AS depsbytes,
+ length((s.stamcv)::bytea) AS mcvbytes,
+ pg_mv_stats_mcvlist_info(s.stamcv) AS mcvinfo
FROM ((pg_mv_statistic s
JOIN pg_class c ON ((c.oid = s.starelid)))
LEFT JOIN pg_namespace n ON ((n.oid = c.relnamespace)));
diff --git a/src/test/regress/expected/type_sanity.out b/src/test/regress/expected/type_sanity.out
index b0b40ca..dde15b9 100644
--- a/src/test/regress/expected/type_sanity.out
+++ b/src/test/regress/expected/type_sanity.out
@@ -72,8 +72,9 @@ WHERE p1.typtype not in ('c','d','p') AND p1.typname NOT LIKE E'\\_%'
194 | pg_node_tree
3353 | pg_ndistinct
3358 | pg_dependencies
+ 441 | pg_mcv_list
210 | smgr
-(4 rows)
+(5 rows)
-- Make sure typarray points to a varlena array type of our own base
SELECT p1.oid, p1.typname as basetype, p2.typname as arraytype,
diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule
index fda9166..d805840 100644
--- a/src/test/regress/parallel_schedule
+++ b/src/test/regress/parallel_schedule
@@ -118,4 +118,4 @@ test: event_trigger
test: stats
# run tests of multivariate stats
-test: mv_ndistinct mv_dependencies
+test: mv_ndistinct mv_dependencies mv_mcv
diff --git a/src/test/regress/serial_schedule b/src/test/regress/serial_schedule
index 90d74d2..72c6acd 100644
--- a/src/test/regress/serial_schedule
+++ b/src/test/regress/serial_schedule
@@ -173,3 +173,4 @@ test: event_trigger
test: stats
test: mv_ndistinct
test: mv_dependencies
+test: mv_mcv
diff --git a/src/test/regress/sql/mv_mcv.sql b/src/test/regress/sql/mv_mcv.sql
new file mode 100644
index 0000000..693288f
--- /dev/null
+++ b/src/test/regress/sql/mv_mcv.sql
@@ -0,0 +1,169 @@
+-- data type passed by value
+CREATE TABLE mcv_list (
+ a INT,
+ b INT,
+ c INT
+);
+
+-- unknown column
+CREATE STATISTICS s4 WITH (mcv) ON (unknown_column) FROM mcv_list;
+
+-- single column
+CREATE STATISTICS s4 WITH (mcv) ON (a) FROM mcv_list;
+
+-- single column, duplicated
+CREATE STATISTICS s4 WITH (mcv) ON (a, a) FROM mcv_list;
+
+-- two columns, one duplicated
+CREATE STATISTICS s4 WITH (mcv) ON (a, a, b) FROM mcv_list;
+
+-- unknown option
+CREATE STATISTICS s4 WITH (unknown_option) ON (a, b, c) FROM mcv_list;
+
+-- correct command
+CREATE STATISTICS s4 WITH (mcv) ON (a, b, c) FROM mcv_list;
+
+-- random data
+INSERT INTO mcv_list
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- a => b, a => c, b => c
+INSERT INTO mcv_list
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- a => b, a => c
+INSERT INTO mcv_list
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mcv_list
+ SELECT i/100, i/200, i/400 FROM generate_series(1,10000) s(i);
+CREATE INDEX mcv_idx ON mcv_list (a, b);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a = 10 AND b = 5;
+
+DROP TABLE mcv_list;
+
+-- varlena type (text)
+CREATE TABLE mcv_list (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+
+CREATE STATISTICS s5 WITH (mcv) ON (a, b, c) FROM mcv_list;
+
+-- random data
+INSERT INTO mcv_list
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- a => b, a => c, b => c
+INSERT INTO mcv_list
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- a => b, a => c
+INSERT INTO mcv_list
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+TRUNCATE mcv_list;
+
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mcv_list
+ SELECT i/100, i/200, i/400 FROM generate_series(1,10000) s(i);
+CREATE INDEX mcv_idx ON mcv_list (a, b);
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a = '10' AND b = '5';
+
+TRUNCATE mcv_list;
+
+-- check explain (expect bitmap index scan, not plain index scan) with NULLs
+INSERT INTO mcv_list
+ SELECT
+ (CASE WHEN i/100 = 0 THEN NULL ELSE i/100 END),
+ (CASE WHEN i/200 = 0 THEN NULL ELSE i/200 END),
+ (CASE WHEN i/400 = 0 THEN NULL ELSE i/400 END)
+ FROM generate_series(1,10000) s(i);
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mcv_list WHERE a IS NULL AND b IS NULL;
+
+DROP TABLE mcv_list;
+
+-- NULL values (mix of int and text columns)
+CREATE TABLE mcv_list (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+
+CREATE STATISTICS s6 WITH (mcv) ON (a, b, c, d) FROM mcv_list;
+
+INSERT INTO mcv_list
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+
+ANALYZE mcv_list;
+
+SELECT mcv_enabled, mcv_built, pg_mv_stats_mcvlist_info(stamcv)
+ FROM pg_mv_statistic WHERE starelid = 'mcv_list'::regclass;
+
+DROP TABLE mcv_list;
--
2.5.5
[binary/octet-stream] 0006-PATCH-multivariate-histograms-v23.patch (149.7K, ../../1fa982dc-ffb4-eef7-eb55-d8cf7d84f1c4@2ndquadrant.com/7-0006-PATCH-multivariate-histograms-v23.patch)
download | inline diff:
From 83696f24eceb2c9d7a71ffb74171c30bc0c3727a Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Sun, 23 Oct 2016 17:38:35 +0200
Subject: [PATCH 6/9] PATCH: multivariate histograms
- extends the pg_mv_statistic catalog (add 'hist' fields)
- building the histograms during ANALYZE
- simple estimation while planning the queries
- pg_histogram data type (varlena-based)
Includes regression tests mostly equal to those for functional
dependencies / MCV lists.
A new varlena-based data type for storing serialized histograms.
---
doc/src/sgml/catalogs.sgml | 30 +
doc/src/sgml/planstats.sgml | 125 ++
doc/src/sgml/ref/create_statistics.sgml | 35 +
src/backend/catalog/system_views.sql | 4 +-
src/backend/commands/statscmds.c | 11 +-
src/backend/nodes/outfuncs.c | 2 +
src/backend/optimizer/path/clausesel.c | 606 +++++++-
src/backend/optimizer/util/plancat.c | 4 +-
src/backend/utils/mvstats/Makefile | 2 +-
src/backend/utils/mvstats/README.histogram | 299 ++++
src/backend/utils/mvstats/README.stats | 2 +
src/backend/utils/mvstats/common.c | 32 +-
src/backend/utils/mvstats/common.h | 8 +-
src/backend/utils/mvstats/histogram.c | 2123 ++++++++++++++++++++++++++++
src/bin/psql/describe.c | 15 +-
src/include/catalog/pg_cast.h | 3 +
src/include/catalog/pg_mv_statistic.h | 22 +-
src/include/catalog/pg_proc.h | 13 +
src/include/catalog/pg_type.h | 4 +
src/include/nodes/relation.h | 2 +
src/include/utils/builtins.h | 4 +
src/include/utils/mvstats.h | 125 +-
src/test/regress/expected/mv_histogram.out | 198 +++
src/test/regress/expected/opr_sanity.out | 3 +-
src/test/regress/expected/rules.out | 4 +-
src/test/regress/expected/type_sanity.out | 3 +-
src/test/regress/parallel_schedule | 2 +-
src/test/regress/serial_schedule | 1 +
src/test/regress/sql/mv_histogram.sql | 167 +++
29 files changed, 3801 insertions(+), 48 deletions(-)
create mode 100644 src/backend/utils/mvstats/README.histogram
create mode 100644 src/backend/utils/mvstats/histogram.c
create mode 100644 src/test/regress/expected/mv_histogram.out
create mode 100644 src/test/regress/sql/mv_histogram.sql
diff --git a/doc/src/sgml/catalogs.sgml b/doc/src/sgml/catalogs.sgml
index bca03e9..be34e24 100644
--- a/doc/src/sgml/catalogs.sgml
+++ b/doc/src/sgml/catalogs.sgml
@@ -4307,6 +4307,17 @@
</row>
<row>
+ <entry><structfield>hist_enabled</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, histogram will be computed for the combination of columns,
+ covered by the statistics. This does not mean the histogram is already
+ computed, though.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>ndist_built</structfield></entry>
<entry><type>bool</type></entry>
<entry></entry>
@@ -4337,6 +4348,16 @@
</row>
<row>
+ <entry><structfield>hist_built</structfield></entry>
+ <entry><type>bool</type></entry>
+ <entry></entry>
+ <entry>
+ If true, histogram is already computed and available for use during query
+ estimation.
+ </entry>
+ </row>
+
+ <row>
<entry><structfield>stakeys</structfield></entry>
<entry><type>int2vector</type></entry>
<entry><literal><link linkend="catalog-pg-attribute"><structname>pg_attribute</structname></link>.attnum</literal></entry>
@@ -4374,6 +4395,15 @@
</entry>
</row>
+ <row>
+ <entry><structfield>stahist</structfield></entry>
+ <entry><type>pg_histogram</type></entry>
+ <entry></entry>
+ <entry>
+ Histogram, serialized as <structname>pg_histogram</> type.
+ </entry>
+ </row>
+
</tbody>
</tgroup>
</table>
diff --git a/doc/src/sgml/planstats.sgml b/doc/src/sgml/planstats.sgml
index 57f9441..2896b04 100644
--- a/doc/src/sgml/planstats.sgml
+++ b/doc/src/sgml/planstats.sgml
@@ -914,6 +914,131 @@ EXPLAIN ANALYZE SELECT * FROM t WHERE a <= 49 AND b > 49;
</sect2>
+ <sect2 id="mv-histograms">
+ <title>Histograms</title>
+
+ <para>
+ <acronym>MCV</> lists, introduced in the previous section, work very well
+ for low-cardinality columns (i.e. columns with only very few distinct
+ values), and for columns with a few very frequent values (and possibly
+ many rare ones). Histograms, a generalization of per-column histograms
+ briefly described in <xref linkend="row-estimation-examples">, are meant
+ to address the other cases, i.e. high-cardinality columns, particularly
+ when there are no frequent values.
+ </para>
+
+ <para>
+ Although the example data we've used so far is not a very good match, we
+ can try creating a histogram instead of the <acronym>MCV</> list. With the
+ histogram in place, you may get a plan like this:
+
+<programlisting>
+DROP STATISTICS s2;
+CREATE STATISTICS s3 ON t (a,b) WITH (histogram);
+ANALYZE t;
+EXPLAIN ANALYZE SELECT * FROM t WHERE a = 1 AND b = 1;
+ QUERY PLAN
+-------------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=100 width=8) (actual time=0.035..2.967 rows=100 loops=1)
+ Filter: ((a = 1) AND (b = 1))
+ Rows Removed by Filter: 9900
+ Planning time: 0.227 ms
+ Execution time: 3.189 ms
+(5 rows)
+</programlisting>
+
+ Which seems quite accurate, however for other combinations of values the
+ results may be much worse, as illustrated by the following query
+
+<programlisting>
+ QUERY PLAN
+-----------------------------------------------------------------------------------------------
+ Seq Scan on t (cost=0.00..195.00 rows=100 width=8) (actual time=2.771..2.771 rows=0 loops=1)
+ Filter: ((a = 1) AND (b = 10))
+ Rows Removed by Filter: 10000
+ Planning time: 0.179 ms
+ Execution time: 2.812 ms
+(5 rows)
+</programlisting>
+
+ This is due to histograms tracking ranges of values, not individual values.
+ That means it's only possible say whether a bucket may contain items
+ matching the conditions, but it's unclear how many such tuples there
+ actually are in the bucket. Moreover, for larger tables only a small subset
+ of rows gets sampled by <command>ANALYZE</>, causing small variations in
+ the shape of buckets.
+ </para>
+
+ <para>
+ To inspect details of the histogram, we can look into the
+ <structname>pg_mv_stats</> view
+
+<programlisting>
+SELECT tablename, staname, attnums, histbytes, histinfo
+ FROM pg_mv_stats WHERE staname = 's3';
+ tablename | staname | attnums | histbytes | histinfo
+-----------+---------+---------+-----------+-------------
+ t | s3 | 1 2 | 1928 | nbuckets=64
+(1 row)
+</programlisting>
+
+ This shows the histogram has 64 buckets, but as we know there are 100
+ distinct combinations of values in the two columns. This means there are
+ buckets containing multiple combinations, causing the inaccuracy.
+ </para>
+
+ <para>
+ Similarly to <acronym>MCV</> lists, we can inspect histogram contents
+ using a function called <function>pg_mv_histogram_buckets</>.
+
+<programlisting>
+test=# SELECT * FROM pg_mv_histogram_buckets((SELECT oid FROM pg_mv_statistic WHERE staname = 's3'), 0);
+ index | minvals | maxvals | nullsonly | mininclusive | maxinclusive | frequency | density | bucket_volume
+-------+---------+---------+-----------+--------------+--------------+-----------+----------+---------------
+ 0 | {0,0} | {3,1} | {f,f} | {t,t} | {f,f} | 0.01 | 1.68 | 0.005952
+ 1 | {50,0} | {51,3} | {f,f} | {t,t} | {f,f} | 0.01 | 1.12 | 0.008929
+ 2 | {0,25} | {26,31} | {f,f} | {t,t} | {f,f} | 0.01 | 0.28 | 0.035714
+...
+ 61 | {60,0} | {99,12} | {f,f} | {t,t} | {t,f} | 0.02 | 0.124444 | 0.160714
+ 62 | {34,35} | {37,49} | {f,f} | {t,t} | {t,t} | 0.02 | 0.96 | 0.020833
+ 63 | {84,35} | {87,49} | {f,f} | {t,t} | {t,t} | 0.02 | 0.96 | 0.020833
+(64 rows)
+</programlisting>
+
+ Which confirms there are 64 buckets, with frequencies ranging between 1%
+ and 2%. The <structfield>minvals</> and <structfield>maxvals</> show the
+ bucket boundaries, <structfield>nullsonly</> shows which columns contain
+ only null values (in the given bucket).
+ </para>
+
+ <para>
+ Similarly to <acronym>MCV</> lists, the planner applies all conditions to
+ the buckets, and sums the frequencies of the matching ones. For details,
+ see <function>clauselist_mv_selectivity_histogram</> function in
+ <filename>clausesel.c</>.
+ </para>
+
+ <para>
+ It's also possible to build <acronym>MCV</> lists and a histogram, in which
+ case <command>ANALYZE</> will build a <acronym>MCV</> lists with the most
+ frequent values, and a histogram on the remaining part of the sample.
+
+<programlisting>
+DROP STATISTICS s3;
+CREATE STATISTICS s4 ON t (a,b) WITH (mcv, histogram);
+</programlisting>
+
+ In this case the <acronym>MCV</> list and histogram are treated as a single
+ composed statistics.
+ </para>
+
+ <para>
+ For additional information about multivariate histograms, see
+ <filename>src/backend/utils/mvstats/README.histogram</>.
+ </para>
+
+ </sect2>
+
</sect1>
</chapter>
diff --git a/doc/src/sgml/ref/create_statistics.sgml b/doc/src/sgml/ref/create_statistics.sgml
index e95d8d3..de419d2 100644
--- a/doc/src/sgml/ref/create_statistics.sgml
+++ b/doc/src/sgml/ref/create_statistics.sgml
@@ -124,6 +124,15 @@ CREATE STATISTICS [ IF NOT EXISTS ] <replaceable class="PARAMETER">statistics_na
</varlistentry>
<varlistentry>
+ <term><literal>histogram</> (<type>boolean</>)</term>
+ <listitem>
+ <para>
+ Enables histogram for the statistics.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
<term><literal>mcv</> (<type>boolean</>)</term>
<listitem>
<para>
@@ -201,6 +210,32 @@ EXPLAIN ANALYZE SELECT * FROM t2 WHERE (a = 1) AND (b = 2);
</programlisting>
</para>
+ <para>
+ Create table <structname>t3</> with two strongly correlated columns, and
+ a histogram on those two columns:
+
+<programlisting>
+CREATE TABLE t3 (
+ a float,
+ b float
+);
+
+INSERT INTO t3 SELECT mod(i,1000), mod(i,1000) + 50 * (r - 0.5) FROM (
+ SELECT i, random() r FROM generate_series(1,1000000) s(i)
+ ) foo;
+
+CREATE STATISTICS s3 WITH (histogram) ON (a, b) FROM t3;
+
+ANALYZE t3;
+
+-- small overlap
+EXPLAIN ANALYZE SELECT * FROM t3 WHERE (a < 500) AND (b > 500);
+
+-- no overlap
+EXPLAIN ANALYZE SELECT * FROM t3 WHERE (a < 400) AND (b > 600);
+</programlisting>
+ </para>
+
</refsect1>
<refsect1>
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index d4d9c24..2501455 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -190,7 +190,9 @@ CREATE VIEW pg_mv_stats AS
length(s.standist::bytea) AS ndistbytes,
length(S.stadeps::bytea) AS depsbytes,
length(S.stamcv::bytea) AS mcvbytes,
- pg_mv_stats_mcvlist_info(S.stamcv) AS mcvinfo
+ pg_mv_stats_mcvlist_info(S.stamcv) AS mcvinfo,
+ length(S.stahist::bytea) AS histbytes,
+ pg_mv_stats_histogram_info(S.stahist) AS histinfo
FROM (pg_mv_statistic S JOIN pg_class C ON (C.oid = S.starelid))
LEFT JOIN pg_namespace N ON (N.oid = C.relnamespace);
diff --git a/src/backend/commands/statscmds.c b/src/backend/commands/statscmds.c
index ef05745..2e91b0c 100644
--- a/src/backend/commands/statscmds.c
+++ b/src/backend/commands/statscmds.c
@@ -71,7 +71,8 @@ CreateStatistics(CreateStatsStmt *stmt)
/* by default build nothing */
bool build_ndistinct = false,
build_dependencies = false,
- build_mcv = false;
+ build_mcv = false,
+ build_histogram = false;
Assert(IsA(stmt, CreateStatsStmt));
@@ -172,6 +173,8 @@ CreateStatistics(CreateStatsStmt *stmt)
build_dependencies = defGetBoolean(opt);
else if (strcmp(opt->defname, "mcv") == 0)
build_mcv = defGetBoolean(opt);
+ else if (strcmp(opt->defname, "histogram") == 0)
+ build_histogram = defGetBoolean(opt);
else
ereport(ERROR,
(errcode(ERRCODE_SYNTAX_ERROR),
@@ -180,10 +183,10 @@ CreateStatistics(CreateStatsStmt *stmt)
}
/* Make sure there's at least one statistics type specified. */
- if (!(build_ndistinct || build_dependencies || build_mcv))
+ if (!(build_ndistinct || build_dependencies || build_mcv || build_histogram))
ereport(ERROR,
(errcode(ERRCODE_SYNTAX_ERROR),
- errmsg("no statistics type (ndistinct, dependencies, mcv) requested")));
+ errmsg("no statistics type (ndistinct, dependencies, mcv, histogram) requested")));
stakeys = buildint2vector(attnums, numcols);
@@ -207,10 +210,12 @@ CreateStatistics(CreateStatsStmt *stmt)
values[Anum_pg_mv_statistic_ndist_enabled - 1] = BoolGetDatum(build_ndistinct);
values[Anum_pg_mv_statistic_deps_enabled - 1] = BoolGetDatum(build_dependencies);
values[Anum_pg_mv_statistic_mcv_enabled - 1] = BoolGetDatum(build_mcv);
+ values[Anum_pg_mv_statistic_hist_enabled - 1] = BoolGetDatum(build_histogram);
nulls[Anum_pg_mv_statistic_standist - 1] = true;
nulls[Anum_pg_mv_statistic_stadeps - 1] = true;
nulls[Anum_pg_mv_statistic_stamcv - 1] = true;
+ nulls[Anum_pg_mv_statistic_stahist - 1] = true;
/* insert the tuple into pg_mv_statistic */
mvstatrel = heap_open(MvStatisticRelationId, RowExclusiveLock);
diff --git a/src/backend/nodes/outfuncs.c b/src/backend/nodes/outfuncs.c
index a9cc9ad..27dbe76 100644
--- a/src/backend/nodes/outfuncs.c
+++ b/src/backend/nodes/outfuncs.c
@@ -2204,11 +2204,13 @@ _outMVStatisticInfo(StringInfo str, const MVStatisticInfo *node)
WRITE_BOOL_FIELD(ndist_enabled);
WRITE_BOOL_FIELD(deps_enabled);
WRITE_BOOL_FIELD(mcv_enabled);
+ WRITE_BOOL_FIELD(hist_enabled);
/* built/available statistics */
WRITE_BOOL_FIELD(ndist_built);
WRITE_BOOL_FIELD(deps_built);
WRITE_BOOL_FIELD(mcv_built);
+ WRITE_BOOL_FIELD(hist_built);
}
static void
diff --git a/src/backend/optimizer/path/clausesel.c b/src/backend/optimizer/path/clausesel.c
index abdbc5b..fddbcc4 100644
--- a/src/backend/optimizer/path/clausesel.c
+++ b/src/backend/optimizer/path/clausesel.c
@@ -49,6 +49,7 @@ static void addRangeClause(RangeQueryClause **rqlist, Node *clause,
#define STATS_TYPE_FDEPS 0x01
#define STATS_TYPE_MCV 0x02
+#define STATS_TYPE_HIST 0x04
static bool clause_is_mv_compatible(Node *clause, Index relid, Bitmapset **attnums,
int type);
@@ -77,12 +78,21 @@ static Selectivity clauselist_mv_selectivity_mcvlist(PlannerInfo *root,
List *clauses, MVStatisticInfo *mvstats,
bool *fullmatch, Selectivity *lowsel);
+static Selectivity clauselist_mv_selectivity_histogram(PlannerInfo *root,
+ List *clauses, MVStatisticInfo *mvstats);
+
static int update_match_bitmap_mcvlist(PlannerInfo *root, List *clauses,
int2vector *stakeys, MCVList mcvlist,
int nmatches, char *matches,
Selectivity *lowsel, bool *fullmatch,
bool is_or);
+static int update_match_bitmap_histogram(PlannerInfo *root, List *clauses,
+ int2vector *stakeys,
+ MVSerializedHistogram mvhist,
+ int nmatches, char *matches,
+ bool is_or);
+
static bool has_stats(List *stats, int type);
static List *find_stats(PlannerInfo *root, Index relid);
@@ -93,6 +103,7 @@ static bool stats_type_matches(MVStatisticInfo *stat, int type);
#define UPDATE_RESULT(m,r,isor) \
(m) = (isor) ? (Max(m,r)) : (Min(m,r))
+
/****************************************************************************
* ROUTINES TO COMPUTE SELECTIVITIES
****************************************************************************/
@@ -121,7 +132,7 @@ static bool stats_type_matches(MVStatisticInfo *stat, int type);
*
* First we try to reduce the list of clauses by applying (soft) functional
* dependencies, and then we try to estimate the selectivity of the reduced
- * list of clauses using the multivariate MCV list.
+ * list of clauses using the multivariate MCV list and histograms.
*
* Finally we remove the portion of clauses estimated using multivariate stats,
* and process the rest of the clauses using the regular per-column stats.
@@ -208,16 +219,17 @@ clauselist_selectivity(PlannerInfo *root,
* If there are no such stats or not enough attributes, don't waste time
* simply skip to estimation using the plain per-column stats.
*/
- if (has_stats(stats, STATS_TYPE_MCV) &&
- (count_mv_attnums(clauses, relid, STATS_TYPE_MCV) >= 2))
+ if (has_stats(stats, STATS_TYPE_MCV | STATS_TYPE_HIST) &&
+ (count_mv_attnums(clauses, relid,
+ STATS_TYPE_MCV | STATS_TYPE_HIST) >= 2))
{
/* collect attributes from the compatible conditions */
Bitmapset *mvattnums = collect_mv_attnums(clauses, relid,
- STATS_TYPE_MCV);
+ STATS_TYPE_MCV | STATS_TYPE_HIST);
/* and search for the statistic covering the most attributes */
MVStatisticInfo *mvstat = choose_mv_statistics(stats, mvattnums,
- STATS_TYPE_MCV);
+ STATS_TYPE_MCV | STATS_TYPE_HIST);
if (mvstat != NULL) /* we have a matching stats */
{
@@ -226,7 +238,7 @@ clauselist_selectivity(PlannerInfo *root,
/* split the clauselist into regular and mv-clauses */
clauses = clauselist_mv_split(root, relid, clauses, &mvclauses,
- mvstat, STATS_TYPE_MCV);
+ mvstat, STATS_TYPE_MCV | STATS_TYPE_HIST);
/* we've chosen the histogram to match the clauses */
Assert(mvclauses != NIL);
@@ -1178,6 +1190,8 @@ static Selectivity
clauselist_mv_selectivity(PlannerInfo *root, List *clauses, MVStatisticInfo *mvstats)
{
bool fullmatch = false;
+ Selectivity s1 = 0.0,
+ s2 = 0.0;
/*
* Lowest frequency in the MCV list (may be used as an upper bound for
@@ -1191,9 +1205,26 @@ clauselist_mv_selectivity(PlannerInfo *root, List *clauses, MVStatisticInfo *mvs
* order by selectivity (to optimize the MCV/histogram evaluation).
*/
- /* Evaluate the MCV selectivity */
- return clauselist_mv_selectivity_mcvlist(root, clauses, mvstats,
- &fullmatch, &mcv_low);
+ /* Evaluate the MCV first. */
+ s1 = clauselist_mv_selectivity_mcvlist(root, clauses, mvstats,
+ &fullmatch, &mcv_low);
+
+ /*
+ * If we got a full equality match on the MCV list, we're done (and the
+ * estimate is pretty good).
+ */
+ if (fullmatch && (s1 > 0.0))
+ return s1;
+
+ /*
+ * TODO if (fullmatch) without matching MCV item, use the mcv_low
+ * selectivity as upper bound
+ */
+
+ s2 = clauselist_mv_selectivity_histogram(root, clauses, mvstats);
+
+ /* TODO clamp to <= 1.0 (or more strictly, when possible) */
+ return s1 + s2;
}
/*
@@ -1379,7 +1410,8 @@ choose_mv_statistics(List *stats, Bitmapset *attnums, int types)
/* skip statistics not matching any of the requested types */
if (! ((info->deps_built && (STATS_TYPE_FDEPS & types)) ||
- (info->mcv_built && (STATS_TYPE_MCV & types))))
+ (info->mcv_built && (STATS_TYPE_MCV & types)) ||
+ (info->hist_built && (STATS_TYPE_HIST & types))))
continue;
/* count columns covered by the statistics */
@@ -1609,7 +1641,7 @@ mv_compatible_walker(Node *node, mv_compatible_context *context)
case F_SCALARGTSEL:
/* not compatible with functional dependencies */
- if (!(context->types & STATS_TYPE_MCV))
+ if (!(context->types & (STATS_TYPE_MCV | STATS_TYPE_HIST)))
return true; /* terminate */
break;
@@ -1677,6 +1709,9 @@ stats_type_matches(MVStatisticInfo *stat, int type)
if ((type & STATS_TYPE_MCV) && stat->mcv_built)
return true;
+ if ((type & STATS_TYPE_HIST) && stat->hist_built)
+ return true;
+
return false;
}
@@ -1695,6 +1730,9 @@ has_stats(List *stats, int type)
/* terminate if we've found at least one matching statistics */
if (stats_type_matches(stat, type))
return true;
+
+ if ((type & STATS_TYPE_HIST) && stat->hist_built)
+ return true;
}
return false;
@@ -1725,12 +1763,12 @@ find_stats(PlannerInfo *root, Index relid)
*
* The algorithm works like this:
*
- * 1) mark all items as 'match'
- * 2) walk through all the clauses
- * 3) for a particular clause, walk through all the items
- * 4) skip items that are already 'no match'
- * 5) check clause for items that still match
- * 6) sum frequencies for items to get selectivity
+ * 1) mark all items as 'match'
+ * 2) walk through all the clauses
+ * 3) for a particular clause, walk through all the items
+ * 4) skip items that are already 'no match'
+ * 5) check clause for items that still match
+ * 6) sum frequencies for items to get selectivity
*
* The function also returns the frequency of the least frequent item
* on the MCV list, which may be useful for clamping estimate from the
@@ -2116,3 +2154,537 @@ update_match_bitmap_mcvlist(PlannerInfo *root, List *clauses,
return nmatches;
}
+
+/*
+ * Estimate selectivity of clauses using a histogram.
+ *
+ * If there's no histogram for the stats, the function returns 0.0.
+ *
+ * The general idea of this method is similar to how MCV lists are
+ * processed, except that this introduces the concept of a partial
+ * match (MCV only works with full match / mismatch).
+ *
+ * The algorithm works like this:
+ *
+ * 1) mark all buckets as 'full match'
+ * 2) walk through all the clauses
+ * 3) for a particular clause, walk through all the buckets
+ * 4) skip buckets that are already 'no match'
+ * 5) check clause for buckets that still match (at least partially)
+ * 6) sum frequencies for buckets to get selectivity
+ *
+ * Unlike MCV lists, histograms have a concept of a partial match. In
+ * that case we use 1/2 the bucket, to minimize the average error. The
+ * MV histograms are usually less detailed than the per-column ones,
+ * meaning the sum is often quite high (thanks to combining a lot of
+ * "partially hit" buckets).
+ *
+ * Maybe we could use per-bucket information with number of distinct
+ * values it contains (for each dimension), and then use that to correct
+ * the estimate (so with 10 distinct values, we'd use 1/10 of the bucket
+ * frequency). We might also scale the value depending on the actual
+ * ndistinct estimate (not just the values observed in the sample).
+ *
+ * Another option would be to multiply the selectivities, i.e. if we get
+ * 'partial match' for a bucket for multiple conditions, we might use
+ * 0.5^k (where k is the number of conditions), instead of 0.5. This
+ * probably does not minimize the average error, though.
+ *
+ * TODO: This might use a similar shortcut to MCV lists - count buckets
+ * marked as partial/full match, and terminate once this drop to 0.
+ * Not sure if it's really worth it - for MCV lists a situation like
+ * this is not uncommon, but for histograms it's not that clear.
+ */
+static Selectivity
+clauselist_mv_selectivity_histogram(PlannerInfo *root, List *clauses,
+ MVStatisticInfo *mvstats)
+{
+ int i;
+ Selectivity s = 0.0;
+ Selectivity u = 0.0;
+
+ int nmatches = 0;
+ char *matches = NULL;
+
+ MVSerializedHistogram mvhist = NULL;
+
+ /* there's no histogram */
+ if (!mvstats->hist_built)
+ return 0.0;
+
+ /* There may be no histogram in the stats (check hist_built flag) */
+ mvhist = load_mv_histogram(mvstats->mvoid);
+
+ Assert(mvhist != NULL);
+ Assert(clauses != NIL);
+ Assert(list_length(clauses) >= 2);
+
+ /*
+ * Bitmap of bucket matches (mismatch, partial, full). by default all
+ * buckets fully match (and we'll eliminate them).
+ */
+ matches = palloc0(sizeof(char) * mvhist->nbuckets);
+ memset(matches, MVSTATS_MATCH_FULL, sizeof(char) * mvhist->nbuckets);
+
+ nmatches = mvhist->nbuckets;
+
+ /* build the match bitmap */
+ update_match_bitmap_histogram(root, clauses,
+ mvstats->stakeys, mvhist,
+ nmatches, matches, false);
+
+ /* now, walk through the buckets and sum the selectivities */
+ for (i = 0; i < mvhist->nbuckets; i++)
+ {
+ /*
+ * Find out what part of the data is covered by the histogram, so that
+ * we can 'scale' the selectivity properly (e.g. when only 50% of the
+ * sample got into the histogram, and the rest is in a MCV list).
+ *
+ * TODO This might be handled by keeping a global "frequency" for the
+ * whole histogram, which might save us some time spent accessing the
+ * not-matching part of the histogram. Although it's likely in a
+ * cache, so it's very fast.
+ */
+ u += mvhist->buckets[i]->ntuples;
+
+ if (matches[i] == MVSTATS_MATCH_FULL)
+ s += mvhist->buckets[i]->ntuples;
+ else if (matches[i] == MVSTATS_MATCH_PARTIAL)
+ s += 0.5 * mvhist->buckets[i]->ntuples;
+ }
+
+#ifdef DEBUG_MVHIST
+ debug_histogram_matches(mvhist, matches);
+#endif
+
+ /* release the allocated bitmap and deserialized histogram */
+ pfree(matches);
+ pfree(mvhist);
+
+ return s * u;
+}
+
+/* cached result of bucket boundary comparison for a single dimension */
+
+#define HIST_CACHE_NOT_FOUND 0x00
+#define HIST_CACHE_FALSE 0x01
+#define HIST_CACHE_TRUE 0x03
+#define HIST_CACHE_MASK 0x02
+
+static char
+bucket_contains_value(FmgrInfo ltproc, Datum constvalue,
+ Datum min_value, Datum max_value,
+ int min_index, int max_index,
+ bool min_include, bool max_include,
+ char *callcache)
+{
+ bool a,
+ b;
+
+ char min_cached = callcache[min_index];
+ char max_cached = callcache[max_index];
+
+ /*
+ * First some quick checks on equality - if any of the boundaries equals,
+ * we have a partial match (so no need to call the comparator).
+ */
+ if (((min_value == constvalue) && (min_include)) ||
+ ((max_value == constvalue) && (max_include)))
+ return MVSTATS_MATCH_PARTIAL;
+
+ /* Keep the values 0/1 because of the XOR at the end. */
+ a = ((min_cached & HIST_CACHE_MASK) >> 1);
+ b = ((max_cached & HIST_CACHE_MASK) >> 1);
+
+ /*
+ * If result for the bucket lower bound not in cache, evaluate the
+ * function and store the result in the cache.
+ */
+ if (!min_cached)
+ {
+ a = DatumGetBool(FunctionCall2Coll(<proc,
+ DEFAULT_COLLATION_OID,
+ constvalue, min_value));
+ /* remember the result */
+ callcache[min_index] = (a) ? HIST_CACHE_TRUE : HIST_CACHE_FALSE;
+ }
+
+ /* And do the same for the upper bound. */
+ if (!max_cached)
+ {
+ b = DatumGetBool(FunctionCall2Coll(<proc,
+ DEFAULT_COLLATION_OID,
+ constvalue, max_value));
+ /* remember the result */
+ callcache[max_index] = (b) ? HIST_CACHE_TRUE : HIST_CACHE_FALSE;
+ }
+
+ return (a ^ b) ? MVSTATS_MATCH_PARTIAL : MVSTATS_MATCH_NONE;
+}
+
+static char
+bucket_is_smaller_than_value(FmgrInfo opproc, Datum constvalue,
+ Datum min_value, Datum max_value,
+ int min_index, int max_index,
+ bool min_include, bool max_include,
+ char *callcache, bool isgt)
+{
+ char min_cached = callcache[min_index];
+ char max_cached = callcache[max_index];
+
+ /* Keep the values 0/1 because of the XOR at the end. */
+ bool a = ((min_cached & HIST_CACHE_MASK) >> 1);
+ bool b = ((max_cached & HIST_CACHE_MASK) >> 1);
+
+ if (!min_cached)
+ {
+ a = DatumGetBool(FunctionCall2Coll(&opproc,
+ DEFAULT_COLLATION_OID,
+ min_value,
+ constvalue));
+ /* remember the result */
+ callcache[min_index] = (a) ? HIST_CACHE_TRUE : HIST_CACHE_FALSE;
+ }
+
+ if (!max_cached)
+ {
+ b = DatumGetBool(FunctionCall2Coll(&opproc,
+ DEFAULT_COLLATION_OID,
+ max_value,
+ constvalue));
+ /* remember the result */
+ callcache[max_index] = (b) ? HIST_CACHE_TRUE : HIST_CACHE_FALSE;
+ }
+
+ /*
+ * Now, we need to combine both results into the final answer, and we need
+ * to be careful about the 'isgt' variable which kinda inverts the
+ * meaning.
+ *
+ * First, we handle the case when each boundary returns different results.
+ * In that case the outcome can only be 'partial' match.
+ */
+ if (a != b)
+ return MVSTATS_MATCH_PARTIAL;
+
+ /*
+ * When the results are the same, then it depends on the 'isgt' value.
+ * There are four options:
+ *
+ * isgt=false a=b=true => full match isgt=false a=b=false => empty
+ * isgt=true a=b=true => empty isgt=true a=b=false => full match
+ *
+ * We'll cheat a bit, because we know that (a=b) so we'll use just one of
+ * them.
+ */
+ if (isgt)
+ return (!a) ? MVSTATS_MATCH_FULL : MVSTATS_MATCH_NONE;
+ else
+ return (a) ? MVSTATS_MATCH_FULL : MVSTATS_MATCH_NONE;
+}
+
+/*
+ * Evaluate clauses using the histogram, and update the match bitmap.
+ *
+ * The bitmap may be already partially set, so this is really a way to
+ * combine results of several clause lists - either when computing
+ * conditional probability P(A|B) or a combination of AND/OR clauses.
+ *
+ * Note: This is not a simple bitmap in the sense that there are more
+ * than two possible values for each item - no match, partial
+ * match and full match. So we need 2 bits per item.
+ *
+ * TODO: This works with 'bitmap' where each item is represented as a
+ * char, which is slightly wasteful. Instead, we could use a bitmap
+ * with 2 bits per item, reducing the size to ~1/4. By using values
+ * 0, 1 and 3 (instead of 0, 1 and 2), the operations (merging etc.)
+ * might be performed just like for simple bitmap by using & and |,
+ * which might be faster than min/max.
+ */
+static int
+update_match_bitmap_histogram(PlannerInfo *root, List *clauses,
+ int2vector *stakeys,
+ MVSerializedHistogram mvhist,
+ int nmatches, char *matches,
+ bool is_or)
+{
+ int i;
+ ListCell *l;
+
+ /*
+ * Used for caching function calls, only once per deduplicated value.
+ *
+ * We know may have up to (2 * nbuckets) values per dimension. It's
+ * probably overkill, but let's allocate that once for all clauses, to
+ * minimize overhead.
+ *
+ * Also, we only need two bits per value, but this allocates byte per
+ * value. Might be worth optimizing.
+ *
+ * 0x00 - not yet called 0x01 - called, result is 'false' 0x03 - called,
+ * result is 'true'
+ */
+ char *callcache = palloc(mvhist->nbuckets);
+
+ Assert(mvhist != NULL);
+ Assert(mvhist->nbuckets > 0);
+ Assert(nmatches >= 0);
+ Assert(nmatches <= mvhist->nbuckets);
+
+ Assert(clauses != NIL);
+ Assert(list_length(clauses) >= 1);
+
+ /* loop through the clauses and do the estimation */
+ foreach(l, clauses)
+ {
+ Node *clause = (Node *) lfirst(l);
+
+ /* if it's a RestrictInfo, then extract the clause */
+ if (IsA(clause, RestrictInfo))
+ clause = (Node *) ((RestrictInfo *) clause)->clause;
+
+ /* it's either OpClause, or NullTest */
+ if (is_opclause(clause))
+ {
+ OpExpr *expr = (OpExpr *) clause;
+ bool varonleft = true;
+ bool ok;
+
+ FmgrInfo opproc; /* operator */
+
+ fmgr_info(get_opcode(expr->opno), &opproc);
+
+ /* reset the cache (per clause) */
+ memset(callcache, 0, mvhist->nbuckets);
+
+ ok = (NumRelids(clause) == 1) &&
+ (is_pseudo_constant_clause(lsecond(expr->args)) ||
+ (varonleft = false,
+ is_pseudo_constant_clause(linitial(expr->args))));
+
+ if (ok)
+ {
+ FmgrInfo ltproc;
+ RegProcedure oprrest = get_oprrest(expr->opno);
+
+ Var *var = (varonleft) ? linitial(expr->args) : lsecond(expr->args);
+ Const *cst = (varonleft) ? lsecond(expr->args) : linitial(expr->args);
+ bool isgt = (!varonleft);
+
+ TypeCacheEntry *typecache
+ = lookup_type_cache(var->vartype, TYPECACHE_LT_OPR);
+
+ /* lookup dimension for the attribute */
+ int idx = mv_get_index(var->varattno, stakeys);
+
+ fmgr_info(get_opcode(typecache->lt_opr), <proc);
+
+ /*
+ * Check this for all buckets that still have "true" in the
+ * bitmap
+ *
+ * We already know the clauses use suitable operators (because
+ * that's how we filtered them).
+ */
+ for (i = 0; i < mvhist->nbuckets; i++)
+ {
+ char res = MVSTATS_MATCH_NONE;
+
+ MVSerializedBucket bucket = mvhist->buckets[i];
+
+ /* histogram boundaries */
+ Datum minval,
+ maxval;
+ bool mininclude,
+ maxinclude;
+ int minidx,
+ maxidx;
+
+ /*
+ * For AND-lists, we can also mark NULL buckets as 'no
+ * match' (and then skip them). For OR-lists this is not
+ * possible.
+ */
+ if ((!is_or) && bucket->nullsonly[idx])
+ matches[i] = MVSTATS_MATCH_NONE;
+
+ /*
+ * Skip buckets that were already eliminated - this is
+ * impotant considering how we update the info (we only
+ * lower the match). We can't really do anything about the
+ * MATCH_PARTIAL buckets.
+ */
+ if ((!is_or) && (matches[i] == MVSTATS_MATCH_NONE))
+ continue;
+ else if (is_or && (matches[i] == MVSTATS_MATCH_FULL))
+ continue;
+
+ /* lookup the values and cache of function calls */
+ minidx = bucket->min[idx];
+ maxidx = bucket->max[idx];
+
+ minval = mvhist->values[idx][bucket->min[idx]];
+ maxval = mvhist->values[idx][bucket->max[idx]];
+
+ mininclude = bucket->min_inclusive[idx];
+ maxinclude = bucket->max_inclusive[idx];
+
+ /*
+ * TODO Maybe it's possible to add here a similar
+ * optimization as for the MCV lists:
+ *
+ * (nmatches == 0) && AND-list => all eliminated (FALSE)
+ * (nmatches == N) && OR-list => all eliminated (TRUE)
+ *
+ * But it's more complex because of the partial matches.
+ */
+
+ /*
+ * If it's not a "<" or ">" or "=" operator, just ignore
+ * the clause. Otherwise note the relid and attnum for the
+ * variable.
+ *
+ * TODO I'm really unsure the handling of 'isgt' flag
+ * (that is, clauses with reverse order of
+ * variable/constant) is correct. I wouldn't be surprised
+ * if there was some mixup. Using the lt/gt operators
+ * instead of messing with the opproc could make it
+ * simpler. It would however be using a different operator
+ * than the query, although it's not any shadier than
+ * using the selectivity function as is done currently.
+ */
+ switch (oprrest)
+ {
+ case F_SCALARLTSEL: /* Var < Const */
+ case F_SCALARGTSEL: /* Var > Const */
+
+ res = bucket_is_smaller_than_value(opproc, cst->constvalue,
+ minval, maxval,
+ minidx, maxidx,
+ mininclude, maxinclude,
+ callcache, isgt);
+ break;
+
+ case F_EQSEL:
+
+ /*
+ * We only check whether the value is within the
+ * bucket, using the lt operator, and we also
+ * check for equality with the boundaries.
+ */
+
+ res = bucket_contains_value(ltproc, cst->constvalue,
+ minval, maxval,
+ minidx, maxidx,
+ mininclude, maxinclude,
+ callcache);
+ break;
+ }
+
+ UPDATE_RESULT(matches[i], res, is_or);
+
+ }
+ }
+ }
+ else if (IsA(clause, NullTest))
+ {
+ NullTest *expr = (NullTest *) clause;
+ Var *var = (Var *) (expr->arg);
+
+ /* FIXME proper matching attribute to dimension */
+ int idx = mv_get_index(var->varattno, stakeys);
+
+ /*
+ * Walk through the buckets and evaluate the current clause. We
+ * can skip items that were already ruled out, and terminate if
+ * there are no remaining buckets that might possibly match.
+ */
+ for (i = 0; i < mvhist->nbuckets; i++)
+ {
+ MVSerializedBucket bucket = mvhist->buckets[i];
+
+ /*
+ * Skip buckets that were already eliminated - this is
+ * impotant considering how we update the info (we only lower
+ * the match)
+ */
+ if ((!is_or) && (matches[i] == MVSTATS_MATCH_NONE))
+ continue;
+ else if (is_or && (matches[i] == MVSTATS_MATCH_FULL))
+ continue;
+
+ /* if the clause mismatches the bucket, set it as MATCH_NONE */
+ if ((expr->nulltesttype == IS_NULL)
+ && (!bucket->nullsonly[idx]))
+ UPDATE_RESULT(matches[i], MVSTATS_MATCH_NONE, is_or);
+
+ else if ((expr->nulltesttype == IS_NOT_NULL) &&
+ (bucket->nullsonly[idx]))
+ UPDATE_RESULT(matches[i], MVSTATS_MATCH_NONE, is_or);
+ }
+ }
+ else if (or_clause(clause) || and_clause(clause))
+ {
+ /*
+ * AND/OR clause, with all clauses compatible with the selected MV
+ * stat
+ */
+
+ int i;
+ BoolExpr *orclause = ((BoolExpr *) clause);
+ List *orclauses = orclause->args;
+
+ /* match/mismatch bitmap for each bucket */
+ int or_nmatches = 0;
+ char *or_matches = NULL;
+
+ Assert(orclauses != NIL);
+ Assert(list_length(orclauses) >= 2);
+
+ /* number of matching buckets */
+ or_nmatches = mvhist->nbuckets;
+
+ /* by default none of the buckets matches the clauses */
+ or_matches = palloc0(sizeof(char) * or_nmatches);
+
+ if (or_clause(clause))
+ {
+ /* OR clauses assume nothing matches, initially */
+ memset(or_matches, MVSTATS_MATCH_NONE, sizeof(char) * or_nmatches);
+ or_nmatches = 0;
+ }
+ else
+ {
+ /* AND clauses assume nothing matches, initially */
+ memset(or_matches, MVSTATS_MATCH_FULL, sizeof(char) * or_nmatches);
+ }
+
+ /* build the match bitmap for the OR-clauses */
+ or_nmatches = update_match_bitmap_histogram(root, orclauses,
+ stakeys, mvhist,
+ or_nmatches, or_matches, or_clause(clause));
+
+ /* merge the bitmap into the existing one */
+ for (i = 0; i < mvhist->nbuckets; i++)
+ {
+ /*
+ * Merge the result into the bitmap (Min for AND, Max for OR).
+ *
+ * FIXME this does not decrease the number of matches
+ */
+ UPDATE_RESULT(matches[i], or_matches[i], is_or);
+ }
+
+ pfree(or_matches);
+
+ }
+ else
+ elog(ERROR, "unknown clause type: %d", clause->type);
+ }
+
+ /* free the call cache */
+ pfree(callcache);
+
+ return nmatches;
+}
diff --git a/src/backend/optimizer/util/plancat.c b/src/backend/optimizer/util/plancat.c
index 9dd4e83..c804e13 100644
--- a/src/backend/optimizer/util/plancat.c
+++ b/src/backend/optimizer/util/plancat.c
@@ -1287,7 +1287,7 @@ get_relation_statistics(RelOptInfo *rel, Relation relation)
mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
/* unavailable stats are not interesting for the planner */
- if (mvstat->deps_built || mvstat->ndist_built || mvstat->mcv_built)
+ if (mvstat->deps_built || mvstat->ndist_built || mvstat->mcv_built || mvstat->hist_built)
{
info = makeNode(MVStatisticInfo);
@@ -1298,11 +1298,13 @@ get_relation_statistics(RelOptInfo *rel, Relation relation)
info->ndist_enabled = mvstat->ndist_enabled;
info->deps_enabled = mvstat->deps_enabled;
info->mcv_enabled = mvstat->mcv_enabled;
+ info->hist_enabled = mvstat->hist_enabled;
/* built/available statistics */
info->ndist_built = mvstat->ndist_built;
info->deps_built = mvstat->deps_built;
info->mcv_built = mvstat->mcv_built;
+ info->hist_built = mvstat->hist_built;
/* stakeys */
adatum = SysCacheGetAttr(MVSTATOID, htup,
diff --git a/src/backend/utils/mvstats/Makefile b/src/backend/utils/mvstats/Makefile
index d5d47ba..d4b88e9 100644
--- a/src/backend/utils/mvstats/Makefile
+++ b/src/backend/utils/mvstats/Makefile
@@ -12,6 +12,6 @@ subdir = src/backend/utils/mvstats
top_builddir = ../../../..
include $(top_builddir)/src/Makefile.global
-OBJS = common.o dependencies.o mcv.o mvdist.o
+OBJS = common.o dependencies.o histogram.o mcv.o mvdist.o
include $(top_srcdir)/src/backend/common.mk
diff --git a/src/backend/utils/mvstats/README.histogram b/src/backend/utils/mvstats/README.histogram
new file mode 100644
index 0000000..a182fa3
--- /dev/null
+++ b/src/backend/utils/mvstats/README.histogram
@@ -0,0 +1,299 @@
+Multivariate histograms
+=======================
+
+Histograms on individual attributes consist of buckets represented by ranges,
+covering the domain of the attribute. That is, each bucket is a [min,max]
+interval, and contains all values in this range. The histogram is built in such
+a way that all buckets have about the same frequency.
+
+Multivariate histograms are an extension into n-dimensional space - the buckets
+are n-dimensional intervals (i.e. n-dimensional rectagles), covering the domain
+of the combination of attributes. That is, each bucket has a vector of lower
+and upper boundaries, denoted min[i] and max[i] (where i = 1..n).
+
+In addition to the boundaries, each bucket tracks additional info:
+
+ * frequency (fraction of tuples in the bucket)
+ * whether the boundaries are inclusive or exclusive
+ * whether the dimension contains only NULL values
+ * number of distinct values in each dimension (for building only)
+
+It's possible that in the future we'll multiple histogram types, with different
+features. We do however expect all the types to share the same representation
+(buckets as ranges) and only differ in how we build them.
+
+The current implementation builds non-overlapping buckets, that may not be true
+for some histogram types and the code should not rely on this assumption. There
+are interesting types of histograms (or algorithms) with overlapping buckets.
+
+When used on low-cardinality data, histograms usually perform considerably worse
+than MCV lists (which are a good fit for this kind of data). This is especially
+true on label-like values, where ordering of the values is mostly unrelated to
+meaning of the data, as proper ordering is crucial for histograms.
+
+On high-cardinality data the histograms are usually a better choice, because MCV
+lists can't represent the distribution accurately enough.
+
+
+Selectivity estimation
+----------------------
+
+The estimation is implemented in clauselist_mv_selectivity_histogram(), and
+works very similarly to clauselist_mv_selectivity_mcvlist.
+
+The main difference is that while MCV lists support exact matches, histograms
+often result in approximate matches - e.g. with equality we can only say if
+the constant would be part of the bucket, but not whether it really is there
+or what fraction of the bucket it corresponds to. In this case we rely on
+some defaults just like in the per-column histograms.
+
+The current implementation uses histograms to estimates those types of clauses
+(think of WHERE conditions):
+
+ (a) equality clauses WHERE (a = 1) AND (b = 2)
+ (b) inequality clauses WHERE (a < 1) AND (b >= 2)
+ (c) NULL clauses WHERE (a IS NULL) AND (b IS NOT NULL)
+ (d) OR-clauses WHERE (a = 1) OR (b = 2)
+
+Similarly to MCV lists, it's possible to add support for additional types of
+clauses, for example:
+
+ (e) multi-var clauses WHERE (a > b)
+
+and so on. These are tasks for the future, not yet implemented.
+
+
+When evaluating a clause on a bucket, we may get one of three results:
+
+ (a) FULL_MATCH - The bucket definitely matches the clause.
+
+ (b) PARTIAL_MATCH - The bucket matches the clause, but not necessarily all
+ the tuples it represents.
+
+ (c) NO_MATCH - The bucket definitely does not match the clause.
+
+This may be illustrated using a range [1, 5], which is essentially a 1-D bucket.
+With clause
+
+ WHERE (a < 10) => FULL_MATCH (all range values are below
+ 10, so the whole bucket matches)
+
+ WHERE (a < 3) => PARTIAL_MATCH (there may be values matching
+ the clause, but we don't know how many)
+
+ WHERE (a < 0) => NO_MATCH (the whole range is above 1, so
+ no values from the bucket can match)
+
+Some clauses may produce only some of those results - for example equality
+clauses may never produce FULL_MATCH as we always hit only part of the bucket
+(we can't match both boundaries at the same time). This results in less accurate
+estimates compared to MCV lists, where we can hit a MCV items exactly (there's
+no PARTIAL match in MCV).
+
+There are also clauses that may not produce any PARTIAL_MATCH results. A nice
+example of that is 'IS [NOT] NULL' clause, which either matches the bucket
+completely (FULL_MATCH) or not at all (NO_MATCH), thanks to how the NULL-buckets
+are constructed.
+
+Computing the total selectivity estimate is trivial - simply sum selectivities
+from all the FULL_MATCH and PARTIAL_MATCH buckets (but for buckets marked with
+PARTIAL_MATCH, multiply the frequency by 0.5 to minimize the average error).
+
+
+Building a histogram
+---------------------
+
+The algorithm of building a histogram in general is quite simple:
+
+ (a) create an initial bucket (containing all sample rows)
+
+ (b) create NULL buckets (by splitting the initial bucket)
+
+ (c) repeat
+
+ (1) choose bucket to split next
+
+ (2) terminate if no bucket that might be split found, or if we've
+ reached the maximum number of buckets (16384)
+
+ (3) choose dimension to partition the bucket by
+
+ (4) partition the bucket by the selected dimension
+
+The main complexity is hidden in steps (c.1) and (c.3), i.e. how we choose the
+bucket and dimension for the split, as discussed in the next section.
+
+
+Partitioning criteria
+---------------------
+
+Similarly to one-dimensional histograms, we want to produce buckets with roughly
+the same frequency.
+
+We also need to produce "regular" buckets, because buckets with one dimension
+much longer than the others are very likely to match a lot of conditions (which
+increases error, even if the bucket frequency is very low).
+
+This is especially important when handling OR-clauses, because in that case each
+clause may add buckets independently. With AND-clauses all the clauses have to
+match each bucket, which makes this issue somewhat less concenrning.
+
+To achieve this, we choose the largest bucket (containing the most sample rows),
+but we only choose buckets that can actually be split (have at least 3 different
+combinations of values).
+
+Then we choose the "longest" dimension of the bucket, which is computed by using
+the distinct values in the sample as a measure.
+
+For details see functions select_bucket_to_partition() and partition_bucket(),
+which also includes further discussion.
+
+
+The current limit on number of buckets (16384) is mostly arbitrary, but chosen
+so that it guarantees we don't exceed the number of distinct values indexable by
+uint16 in any of the dimensions. In practice we could handle more buckets as we
+index each dimension separately and the splits should use the dimensions evenly.
+
+Also, histograms this large (with 16k values in multiple dimensions) would be
+quite expensive to build and process, so the 16k limit is rather reasonable.
+
+The actual number of buckets is also related to statistics target, because we
+require MIN_BUCKET_ROWS (10) tuples per bucket before a split, so we can't have
+more than (2 * 300 * target / 10) buckets. For the default target (100) this
+evaluates to ~6k.
+
+
+NULL handling (create_null_buckets)
+-----------------------------------
+
+When building histograms on a single attribute, we first filter out NULL values.
+In the multivariate case, we can't really do that because the rows may contain
+a mix of NULL and non-NULL values in different columns (so we can't simply
+filter all of them out).
+
+For this reason, the histograms are built in a way so that for each bucket, each
+dimension only contains only NULL or non-NULL values. Building the NULL-buckets
+happens as the first step in the build, by the create_null_buckets() function.
+The number of NULL buckets, as produced by this function, has a clear upper
+boundary (2^N) where N is the number of dimensions (attributes the histogram is
+built on). Or rather 2^K where K is the number of attributes that are not marked
+as not-NULL.
+
+The buckets with NULL dimensions are then subject to the same build algorithm
+(i.e. may be split into smaller buckets) just like any other bucket, but may
+only be split by non-NULL dimension.
+
+
+Serialization
+-------------
+
+To store the histogram in pg_mv_statistic table, it is serialized into a more
+efficient form. We also use the representation for estimation, i.e. we don't
+fully deserialize the histogram.
+
+For example the boundary values are deduplicated to minimize the required space.
+How much redundancy is there, actually? Let's assume there are no NULL values,
+so we start with a single bucket - in that case we have 2*N boundaries. Each
+time we split a bucket we introduce one new value (in the "middle" of one of
+the dimensions), and keep boundries for all the other dimensions. So after K
+splits, we have up to
+
+ 2*N + K
+
+unique boundary values (we may have fewe values, if the same value is used for
+several splits). But after K splits we do have (K+1) buckets, so
+
+ (K+1) * 2 * N
+
+boundary values. Using e.g. N=4 and K=999, we arrive to those numbers:
+
+ 2*N + K = 1007
+ (K+1) * 2 * N = 8000
+
+wich means a lot of redundancy. It's somewhat counter-intuitive that the number
+of distinct values does not really depend on the number of dimensions (except
+for the initial bucket, but that's negligible compared to the total).
+
+By deduplicating the values and replacing them with 16-bit indexes (uint16), we
+reduce the required space to
+
+ 1007 * 8 + 8000 * 2 ~= 24kB
+
+which is significantly less than 64kB required for the 'raw' histogram (assuming
+the values are 8B).
+
+While the bytea compression (pglz) might achieve the same reduction of space,
+the deduplicated representation is used to optimize the estimation by caching
+results of function calls for already visited values. This significantly
+reduces the number of calls to (often quite expensive) operators.
+
+Note: Of course, this reasoning only holds for histograms built by the algorithm
+that simply splits the buckets in half. Other histograms types (e.g. containing
+overlapping buckets) may behave differently and require different serialization.
+
+Serialized histograms are marked with 'magic' constant, to make it easier to
+check the bytea value really is a serialized histogram.
+
+
+varlena compression
+-------------------
+
+This serialization may however disable automatic varlena compression, the array
+of unique values is placed at the beginning of the serialized form. Which is
+exactly the chunk used by pglz to check if the data is compressible, and it
+will probably decide it's not very compressible. This is similar to the issue
+we had with JSONB initially.
+
+Maybe storing buckets first would make it work, as the buckets may be better
+compressible.
+
+On the other hand the serialization is actually a context-aware compression,
+usually compressing to ~30% (or even less, with large data types). So the lack
+of additional pglz compression may be acceptable.
+
+
+Deserialization
+---------------
+
+The deserialization is not a perfect inverse of the serialization, as we keep
+the deduplicated arrays. This reduces the amount of memory and also allows
+optimizations during estimation (e.g. we can cache results for the distinct
+values, saving expensive function calls).
+
+
+Inspecting the histogram
+------------------------
+
+Inspecting the regular (per-attribute) histograms is trivial, as it's enough
+to select the columns from pg_stats - the data is encoded as anyarray, so we
+simply get the text representation of the array.
+
+With multivariate histograms it's not that simple due to the possible mix of
+data types in the histogram. It might be possible to produce similar array-like
+text representation, but that'd unnecessarily complicate further processing
+and analysis of the histogram. Instead, there's a SRF function that allows
+access to lower/upper boundaries, frequencies etc.
+
+ SELECT * FROM pg_mv_histogram_buckets();
+
+It has two input parameters:
+
+ oid - OID of the histogram (pg_mv_statistic.staoid)
+ otype - type of output
+
+and produces a table with these columns:
+
+ - bucket ID (0...nbuckets-1)
+ - lower bucket boundaries (string array)
+ - upper bucket boundaries (string array)
+ - nulls only dimensions (boolean array)
+ - lower boundary inclusive (boolean array)
+ - upper boundary includive (boolean array)
+ - frequency (double precision)
+
+The 'otype' accepts three values, determining what will be returned in the
+lower/upper boundary arrays:
+
+ - 0 - values stored in the histogram, encoded as text
+ - 1 - indexes into the deduplicated arrays
+ - 2 - idnexes into the deduplicated arrays, scaled to [0,1]
diff --git a/src/backend/utils/mvstats/README.stats b/src/backend/utils/mvstats/README.stats
index 8d3d268..9cc1c3e 100644
--- a/src/backend/utils/mvstats/README.stats
+++ b/src/backend/utils/mvstats/README.stats
@@ -18,6 +18,8 @@ Currently we only have two kinds of multivariate statistics
(b) MCV lists (README.mcv)
+ (c) multivariate histograms (README.histogram)
+
Compatible clause types
-----------------------
diff --git a/src/backend/utils/mvstats/common.c b/src/backend/utils/mvstats/common.c
index fc8eae2..82f4e4a 100644
--- a/src/backend/utils/mvstats/common.c
+++ b/src/backend/utils/mvstats/common.c
@@ -13,6 +13,7 @@
*
*-------------------------------------------------------------------------
*/
+#include "postgres.h"
#include "common.h"
#include "utils/array.h"
@@ -24,7 +25,7 @@ static List *list_mv_stats(Oid relid);
static void update_mv_stats(Oid relid,
MVNDistinct ndistinct, MVDependencies dependencies,
- MCVList mcvlist,
+ MCVList mcvlist, MVHistogram histogram,
int2vector *attrs, VacAttrStats **stats);
/*
@@ -57,7 +58,8 @@ build_mv_stats(Relation onerel, double totalrows,
MVNDistinct ndistinct = NULL;
MVDependencies deps = NULL;
MCVList mcvlist = NULL;
- int numrows_filtered = 0;
+ MVHistogram histogram = NULL;
+ int numrows_filtered = numrows;
VacAttrStats **stats = NULL;
int numatts = 0;
@@ -102,8 +104,12 @@ build_mv_stats(Relation onerel, double totalrows,
if (stat->mcv_enabled)
mcvlist = build_mv_mcvlist(numrows, rows, attrs, stats, &numrows_filtered);
- /* store the statistics in the catalog */
- update_mv_stats(stat->mvoid, ndistinct, deps, mcvlist, attrs, stats);
+ /* build a multivariate histogram on the columns */
+ if ((numrows_filtered > 0) && (stat->hist_enabled))
+ histogram = build_mv_histogram(numrows_filtered, rows, attrs, stats, numrows);
+
+ /* store the histogram / MCV list in the catalog */
+ update_mv_stats(stat->mvoid, ndistinct, deps, mcvlist, histogram, attrs, stats);
}
}
@@ -187,6 +193,8 @@ list_mv_stats(Oid relid)
info->deps_built = stats->deps_built;
info->mcv_enabled = stats->mcv_enabled;
info->mcv_built = stats->mcv_built;
+ info->hist_enabled = stats->hist_enabled;
+ info->hist_built = stats->hist_built;
result = lappend(result, info);
}
@@ -255,7 +263,8 @@ find_mv_attnums(Oid mvoid, Oid *relid)
*/
static void
update_mv_stats(Oid mvoid,
- MVNDistinct ndistinct, MVDependencies dependencies, MCVList mcvlist,
+ MVNDistinct ndistinct, MVDependencies dependencies,
+ MCVList mcvlist, MVHistogram histogram,
int2vector *attrs, VacAttrStats **stats)
{
HeapTuple stup,
@@ -297,15 +306,26 @@ update_mv_stats(Oid mvoid,
values[Anum_pg_mv_statistic_stamcv - 1] = PointerGetDatum(data);
}
+ if (histogram != NULL)
+ {
+ bytea *data = serialize_mv_histogram(histogram, attrs, stats);
+
+ nulls[Anum_pg_mv_statistic_stahist - 1] = (data == NULL);
+ values[Anum_pg_mv_statistic_stahist - 1]
+ = PointerGetDatum(data);
+ }
+
/* always replace the value (either by bytea or NULL) */
replaces[Anum_pg_mv_statistic_standist - 1] = true;
replaces[Anum_pg_mv_statistic_stadeps - 1] = true;
replaces[Anum_pg_mv_statistic_stamcv - 1] = true;
+ replaces[Anum_pg_mv_statistic_stahist - 1] = true;
/* always change the availability flags */
nulls[Anum_pg_mv_statistic_ndist_built - 1] = false;
nulls[Anum_pg_mv_statistic_deps_built - 1] = false;
nulls[Anum_pg_mv_statistic_mcv_built - 1] = false;
+ nulls[Anum_pg_mv_statistic_hist_built - 1] = false;
nulls[Anum_pg_mv_statistic_stakeys - 1] = false;
@@ -313,12 +333,14 @@ update_mv_stats(Oid mvoid,
replaces[Anum_pg_mv_statistic_ndist_built - 1] = true;
replaces[Anum_pg_mv_statistic_deps_built - 1] = true;
replaces[Anum_pg_mv_statistic_mcv_built - 1] = true;
+ replaces[Anum_pg_mv_statistic_hist_built - 1] = true;
replaces[Anum_pg_mv_statistic_stakeys - 1] = true;
values[Anum_pg_mv_statistic_ndist_built - 1] = BoolGetDatum(ndistinct != NULL);
values[Anum_pg_mv_statistic_deps_built - 1] = BoolGetDatum(dependencies != NULL);
values[Anum_pg_mv_statistic_mcv_built - 1] = BoolGetDatum(mcvlist != NULL);
+ values[Anum_pg_mv_statistic_hist_built - 1] = BoolGetDatum(histogram != NULL);
values[Anum_pg_mv_statistic_stakeys - 1] = PointerGetDatum(attrs);
diff --git a/src/backend/utils/mvstats/common.h b/src/backend/utils/mvstats/common.h
index fe56f51..96c0317 100644
--- a/src/backend/utils/mvstats/common.h
+++ b/src/backend/utils/mvstats/common.h
@@ -77,7 +77,7 @@ MultiSortSupport multi_sort_init(int ndims);
void multi_sort_add_dimension(MultiSortSupport mss, int sortdim,
int dim, VacAttrStats **vacattrstats);
-int multi_sort_compare(const void *a, const void *b, void *arg);
+int multi_sort_compare(const void *a, const void *b, void *arg);
int multi_sort_compare_dim(int dim, const SortItem *a,
const SortItem *b, MultiSortSupport mss);
@@ -86,9 +86,9 @@ int multi_sort_compare_dims(int start, int end, const SortItem *a,
const SortItem *b, MultiSortSupport mss);
/* comparators, used when constructing multivariate stats */
-int compare_datums_simple(Datum a, Datum b, SortSupport ssup);
-int compare_scalars_simple(const void *a, const void *b, void *arg);
-int compare_scalars_partition(const void *a, const void *b, void *arg);
+int compare_datums_simple(Datum a, Datum b, SortSupport ssup);
+int compare_scalars_simple(const void *a, const void *b, void *arg);
+int compare_scalars_partition(const void *a, const void *b, void *arg);
void *bsearch_arg(const void *key, const void *base,
size_t nmemb, size_t size,
diff --git a/src/backend/utils/mvstats/histogram.c b/src/backend/utils/mvstats/histogram.c
new file mode 100644
index 0000000..fc0c9c2
--- /dev/null
+++ b/src/backend/utils/mvstats/histogram.c
@@ -0,0 +1,2123 @@
+/*-------------------------------------------------------------------------
+ *
+ * histogram.c
+ * POSTGRES multivariate histograms
+ *
+ *
+ * Portions Copyright (c) 1996-2015, PostgreSQL Global Development Group
+ * Portions Copyright (c) 1994, Regents of the University of California
+ *
+ *
+ * IDENTIFICATION
+ * src/backend/utils/mvstats/histogram.c
+ *
+ *-------------------------------------------------------------------------
+ */
+
+#include "postgres.h"
+
+#include "fmgr.h"
+#include "funcapi.h"
+
+#include "utils/bytea.h"
+#include "utils/lsyscache.h"
+
+#include "common.h"
+#include <math.h>
+
+
+static MVBucket create_initial_mv_bucket(int numrows, HeapTuple *rows,
+ int2vector *attrs,
+ VacAttrStats **stats);
+
+static MVBucket select_bucket_to_partition(int nbuckets, MVBucket *buckets);
+
+static MVBucket partition_bucket(MVBucket bucket, int2vector *attrs,
+ VacAttrStats **stats,
+ int *ndistvalues, Datum **distvalues);
+
+static MVBucket copy_mv_bucket(MVBucket bucket, uint32 ndimensions);
+
+static void update_bucket_ndistinct(MVBucket bucket, int2vector *attrs,
+ VacAttrStats **stats);
+
+static void update_dimension_ndistinct(MVBucket bucket, int dimension,
+ int2vector *attrs,
+ VacAttrStats **stats,
+ bool update_boundaries);
+
+static void create_null_buckets(MVHistogram histogram, int bucket_idx,
+ int2vector *attrs, VacAttrStats **stats);
+
+static Datum *build_ndistinct(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int i, int *nvals);
+
+/*
+ * Each serialized bucket needs to store (in this order):
+ *
+ * - number of tuples (float)
+ * - number of distinct (float)
+ * - min inclusive flags (ndim * sizeof(bool))
+ * - max inclusive flags (ndim * sizeof(bool))
+ * - null dimension flags (ndim * sizeof(bool))
+ * - min boundary indexes (2 * ndim * sizeof(uint16))
+ * - max boundary indexes (2 * ndim * sizeof(uint16))
+ *
+ * So in total:
+ *
+ * ndim * (4 * sizeof(uint16) + 3 * sizeof(bool)) + (2 * sizeof(float))
+ */
+#define BUCKET_SIZE(ndims) \
+ (ndims * (4 * sizeof(uint16) + 3 * sizeof(bool)) + sizeof(float))
+
+/* pointers into a flat serialized bucket of BUCKET_SIZE(n) bytes */
+#define BUCKET_NTUPLES(b) (*(float*)b)
+#define BUCKET_MIN_INCL(b,n) ((bool*)(b + sizeof(float)))
+#define BUCKET_MAX_INCL(b,n) (BUCKET_MIN_INCL(b,n) + n)
+#define BUCKET_NULLS_ONLY(b,n) (BUCKET_MAX_INCL(b,n) + n)
+#define BUCKET_MIN_INDEXES(b,n) ((uint16*)(BUCKET_NULLS_ONLY(b,n) + n))
+#define BUCKET_MAX_INDEXES(b,n) ((BUCKET_MIN_INDEXES(b,n) + n))
+
+/* can't split bucket with less than 10 rows */
+#define MIN_BUCKET_ROWS 10
+
+/*
+ * Data used while building the histogram.
+ */
+typedef struct HistogramBuildData
+{
+
+ float ndistinct; /* frequency of distinct values */
+
+ HeapTuple *rows; /* aray of sample rows */
+ uint32 numrows; /* number of sample rows (array size) */
+
+ /*
+ * Number of distinct values in each dimension. This is used when building
+ * the histogram (and is not serialized/deserialized).
+ */
+ uint32 *ndistincts;
+
+} HistogramBuildData;
+
+typedef HistogramBuildData *HistogramBuild;
+
+/*
+ * builds a multivariate algorithm
+ *
+ * The build algorithm is iterative - initially a single bucket containing all
+ * the sample rows is formed, and then repeatedly split into smaller buckets.
+ * In each step the largest bucket (in some sense) is chosen to be split next.
+ *
+ * The criteria for selecting the largest bucket (and the dimension for the
+ * split) needs to be elaborate enough to produce buckets of roughly the same
+ * size, and also regular shape (not very long in one dimension).
+ *
+ * The current algorithm works like this:
+ *
+ * build NULL-buckets (create_null_buckets)
+ *
+ * while [maximum number of buckets not reached]
+ *
+ * choose bucket to partition (largest bucket)
+ * if no bucket to partition
+ * terminate the algorithm
+ *
+ * choose bucket dimension to partition (largest dimension)
+ * split the bucket into two buckets
+ *
+ * See the discussion at select_bucket_to_partition and partition_bucket for
+ * more details about the algorithm.
+ */
+MVHistogram
+build_mv_histogram(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int numrows_total)
+{
+ int i;
+ int numattrs = attrs->dim1;
+
+ int *ndistvalues;
+ Datum **distvalues;
+
+ MVHistogram histogram;
+
+ HeapTuple *rows_copy = (HeapTuple *) palloc0(numrows * sizeof(HeapTuple));
+
+ memcpy(rows_copy, rows, sizeof(HeapTuple) * numrows);
+
+ Assert((numattrs >= 2) && (numattrs <= MVSTATS_MAX_DIMENSIONS));
+
+ /* build histogram header */
+
+ histogram = (MVHistogram) palloc0(sizeof(MVHistogramData));
+
+ histogram->magic = MVSTAT_HIST_MAGIC;
+ histogram->type = MVSTAT_HIST_TYPE_BASIC;
+
+ histogram->nbuckets = 1;
+ histogram->ndimensions = numattrs;
+
+ /* create max buckets (better than repalloc for short-lived objects) */
+ histogram->buckets
+ = (MVBucket *) palloc0(MVSTAT_HIST_MAX_BUCKETS * sizeof(MVBucket));
+
+ /* create the initial bucket, covering the whole sample set */
+ histogram->buckets[0]
+ = create_initial_mv_bucket(numrows, rows_copy, attrs, stats);
+
+ /*
+ * Collect info on distinct values in each dimension (used later to select
+ * dimension to partition).
+ */
+ ndistvalues = (int *) palloc0(sizeof(int) * numattrs);
+ distvalues = (Datum **) palloc0(sizeof(Datum *) * numattrs);
+
+ for (i = 0; i < numattrs; i++)
+ distvalues[i] = build_ndistinct(numrows, rows, attrs, stats, i,
+ &ndistvalues[i]);
+
+ /*
+ * Split the initial bucket into buckets that don't mix NULL and non-NULL
+ * values in a single dimension.
+ */
+ create_null_buckets(histogram, 0, attrs, stats);
+
+ /*
+ * Do the actual histogram build - select a bucket and split it.
+ */
+ while (histogram->nbuckets < MVSTAT_HIST_MAX_BUCKETS)
+ {
+ MVBucket bucket = select_bucket_to_partition(histogram->nbuckets,
+ histogram->buckets);
+
+ /* no buckets eligible for partitioning */
+ if (bucket == NULL)
+ break;
+
+ /* we modify the bucket in-place and add one new bucket */
+ histogram->buckets[histogram->nbuckets++]
+ = partition_bucket(bucket, attrs, stats, ndistvalues, distvalues);
+ }
+
+ /* finalize the histogram build - compute the frequencies etc. */
+ for (i = 0; i < histogram->nbuckets; i++)
+ {
+ HistogramBuild build_data
+ = ((HistogramBuild) histogram->buckets[i]->build_data);
+
+ /*
+ * The frequency has to be computed from the whole sample, in case
+ * some of the rows were used for MCV.
+ *
+ * XXX Perhaps this should simply compute frequency with respect to
+ * the local freuquency, and then factor-in the MCV later.
+ *
+ * FIXME The 'ntuples' sounds a bit inappropriate for frequency.
+ */
+ histogram->buckets[i]->ntuples
+ = (build_data->numrows * 1.0) / numrows_total;
+ }
+
+ return histogram;
+}
+
+/* build array of distinct values for a single attribute */
+static Datum *
+build_ndistinct(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int i, int *nvals)
+{
+ int j;
+ int nvalues,
+ ndistinct;
+ Datum *values,
+ *distvalues;
+
+ SortSupportData ssup;
+ StdAnalyzeData *mystats = (StdAnalyzeData *) stats[i]->extra_data;
+
+ /* initialize sort support, etc. */
+ memset(&ssup, 0, sizeof(ssup));
+ ssup.ssup_cxt = CurrentMemoryContext;
+
+ /* We always use the default collation for statistics */
+ ssup.ssup_collation = DEFAULT_COLLATION_OID;
+ ssup.ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(mystats->ltopr, &ssup);
+
+ nvalues = 0;
+ values = (Datum *) palloc0(sizeof(Datum) * numrows);
+
+ /* collect values from the sample rows, ignore NULLs */
+ for (j = 0; j < numrows; j++)
+ {
+ Datum value;
+ bool isnull;
+
+ /*
+ * remember the index of the sample row, to make the partitioning
+ * simpler
+ */
+ value = heap_getattr(rows[j], attrs->values[i],
+ stats[i]->tupDesc, &isnull);
+
+ if (isnull)
+ continue;
+
+ values[nvalues++] = value;
+ }
+
+ /* if no non-NULL values were found, free the memory and terminate */
+ if (nvalues == 0)
+ {
+ pfree(values);
+ return NULL;
+ }
+
+ /* sort the array of values using the SortSupport */
+ qsort_arg((void *) values, nvalues, sizeof(Datum),
+ compare_scalars_simple, (void *) &ssup);
+
+ /* count the distinct values first, and allocate just enough memory */
+ ndistinct = 1;
+ for (j = 1; j < nvalues; j++)
+ if (compare_scalars_simple(&values[j], &values[j - 1], &ssup) != 0)
+ ndistinct += 1;
+
+ distvalues = (Datum *) palloc0(sizeof(Datum) * ndistinct);
+
+ /* now collect distinct values into the array */
+ distvalues[0] = values[0];
+ ndistinct = 1;
+
+ for (j = 1; j < nvalues; j++)
+ {
+ if (compare_scalars_simple(&values[j], &values[j - 1], &ssup) != 0)
+ {
+ distvalues[ndistinct] = values[j];
+ ndistinct += 1;
+ }
+ }
+
+ pfree(values);
+
+ *nvals = ndistinct;
+ return distvalues;
+}
+
+/* fetch the histogram (as a bytea) from the pg_mv_statistic catalog */
+MVSerializedHistogram
+load_mv_histogram(Oid mvoid)
+{
+ bool isnull = false;
+ Datum histogram;
+
+#ifdef USE_ASSERT_CHECKING
+ Form_pg_mv_statistic mvstat;
+#endif
+
+ /* Prepare to scan pg_mv_statistic for entries having indrelid = this rel. */
+ HeapTuple htup = SearchSysCache1(MVSTATOID, ObjectIdGetDatum(mvoid));
+
+ if (!HeapTupleIsValid(htup))
+ return NULL;
+
+#ifdef USE_ASSERT_CHECKING
+ mvstat = (Form_pg_mv_statistic) GETSTRUCT(htup);
+ Assert(mvstat->hist_enabled && mvstat->hist_built);
+#endif
+
+ histogram = SysCacheGetAttr(MVSTATOID, htup,
+ Anum_pg_mv_statistic_stahist, &isnull);
+
+ Assert(!isnull);
+
+ ReleaseSysCache(htup);
+
+ return deserialize_mv_histogram(DatumGetByteaP(histogram));
+}
+
+/* print some basic info about the histogram */
+Datum
+pg_mv_stats_histogram_info(PG_FUNCTION_ARGS)
+{
+ bytea *data = PG_GETARG_BYTEA_P(0);
+ char *result;
+
+ MVSerializedHistogram hist = deserialize_mv_histogram(data);
+
+ result = palloc0(128);
+ snprintf(result, 128, "nbuckets=%d", hist->nbuckets);
+
+ PG_RETURN_TEXT_P(cstring_to_text(result));
+}
+
+/*
+ * Serialize the MV histogram into a bytea value. The basic algorithm is quite
+ * simple, and mostly mimincs the MCV serialization:
+ *
+ * (1) perform deduplication for each attribute (separately)
+ *
+ * (a) collect all (non-NULL) attribute values from all buckets
+ * (b) sort the data (using 'lt' from VacAttrStats)
+ * (c) remove duplicate values from the array
+ *
+ * (2) serialize the arrays into a bytea value
+ *
+ * (3) process all buckets
+ *
+ * (a) replace min/max values with indexes into the arrays
+ *
+ * Each attribute has to be processed separately, as we're mixing different
+ * datatypes, and we we need to use the right operators to compare/sort them.
+ * We're also mixing pass-by-value and pass-by-ref types, and so on.
+ *
+ *
+ * FIXME This probably leaks memory, or at least uses it inefficiently
+ * (many small palloc() calls instead of a large one).
+ *
+ * TODO Consider packing boolean flags (NULL) for each item into 'char' or
+ * a longer type (instead of using an array of bool items).
+ */
+bytea *
+serialize_mv_histogram(MVHistogram histogram, int2vector *attrs,
+ VacAttrStats **stats)
+{
+ int i = 0,
+ j = 0;
+ Size total_length = 0;
+
+ bytea *output = NULL;
+ char *data = NULL;
+
+ DimensionInfo *info;
+ SortSupport ssup;
+
+ int nbuckets = histogram->nbuckets;
+ int ndims = histogram->ndimensions;
+
+ /* allocated for serialized bucket data */
+ int bucketsize = BUCKET_SIZE(ndims);
+ char *bucket = palloc0(bucketsize);
+
+ /* values per dimension (and number of non-NULL values) */
+ Datum **values = (Datum **) palloc0(sizeof(Datum *) * ndims);
+ int *counts = (int *) palloc0(sizeof(int) * ndims);
+
+ /* info about dimensions (for deserialize) */
+ info = (DimensionInfo *) palloc0(sizeof(DimensionInfo) * ndims);
+
+ /* sort support data */
+ ssup = (SortSupport) palloc0(sizeof(SortSupportData) * ndims);
+
+ /* collect and deduplicate values for each dimension separately */
+ for (i = 0; i < ndims; i++)
+ {
+ int count;
+ StdAnalyzeData *tmp = (StdAnalyzeData *) stats[i]->extra_data;
+
+ /* keep important info about the data type */
+ info[i].typlen = stats[i]->attrtype->typlen;
+ info[i].typbyval = stats[i]->attrtype->typbyval;
+
+ /*
+ * Allocate space for all min/max values, including NULLs (we won't
+ * use them, but we don't know how many are there), and then collect
+ * all non-NULL values.
+ */
+ values[i] = (Datum *) palloc0(sizeof(Datum) * nbuckets * 2);
+
+ for (j = 0; j < histogram->nbuckets; j++)
+ {
+ /* skip buckets where this dimension is NULL-only */
+ if (!histogram->buckets[j]->nullsonly[i])
+ {
+ values[i][counts[i]] = histogram->buckets[j]->min[i];
+ counts[i] += 1;
+
+ values[i][counts[i]] = histogram->buckets[j]->max[i];
+ counts[i] += 1;
+ }
+ }
+
+ /* there are just NULL values in this dimension */
+ if (counts[i] == 0)
+ continue;
+
+ /* sort and deduplicate */
+ ssup[i].ssup_cxt = CurrentMemoryContext;
+ ssup[i].ssup_collation = DEFAULT_COLLATION_OID;
+ ssup[i].ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(tmp->ltopr, &ssup[i]);
+
+ qsort_arg(values[i], counts[i], sizeof(Datum),
+ compare_scalars_simple, &ssup[i]);
+
+ /*
+ * Walk through the array and eliminate duplicitate values, but keep
+ * the ordering (so that we can do bsearch later). We know there's at
+ * least 1 item, so we can skip the first element.
+ */
+ count = 1; /* number of deduplicated items */
+ for (j = 1; j < counts[i]; j++)
+ {
+ /* if it's different from the previous value, we need to keep it */
+ if (compare_datums_simple(values[i][j - 1], values[i][j], &ssup[i]) != 0)
+ {
+ /* XXX: not needed if (count == j) */
+ values[i][count] = values[i][j];
+ count += 1;
+ }
+ }
+
+ /* make sure we fit into uint16 */
+ Assert(count <= UINT16_MAX);
+
+ /* keep info about the deduplicated count */
+ info[i].nvalues = count;
+
+ /* compute size of the serialized data */
+ if (info[i].typlen > 0)
+ /* byval or byref, but with fixed length (name, tid, ...) */
+ info[i].nbytes = info[i].nvalues * info[i].typlen;
+ else if (info[i].typlen == -1)
+ /* varlena, so just use VARSIZE_ANY */
+ for (j = 0; j < info[i].nvalues; j++)
+ info[i].nbytes += VARSIZE_ANY(values[i][j]);
+ else if (info[i].typlen == -2)
+ /* cstring, so simply strlen */
+ for (j = 0; j < info[i].nvalues; j++)
+ info[i].nbytes += strlen(DatumGetPointer(values[i][j]));
+ else
+ elog(ERROR, "unknown data type typbyval=%d typlen=%d",
+ info[i].typbyval, info[i].typlen);
+ }
+
+ /*
+ * Now we finally know how much space we'll need for the serialized
+ * histogram, as it contains these fields:
+ *
+ * - length (4B) for varlena - magic (4B) - type (4B) - ndimensions (4B) -
+ * nbuckets (4B) - info (ndim * sizeof(DimensionInfo) - arrays of values
+ * for each dimension - serialized buckets (nbuckets * bucketsize)
+ *
+ * So the 'header' size is 20B + ndim * sizeof(DimensionInfo) and then
+ * we'll place the data (and buckets).
+ */
+ total_length = (sizeof(int32) + offsetof(MVHistogramData, buckets)
+ +ndims * sizeof(DimensionInfo)
+ + nbuckets * bucketsize);
+
+ /* account for the deduplicated data */
+ for (i = 0; i < ndims; i++)
+ total_length += info[i].nbytes;
+
+ /* enforce arbitrary limit of 1MB */
+ if (total_length > (1024 * 1024))
+ elog(ERROR, "serialized histogram exceeds 1MB (%ld > %d)",
+ total_length, (1024 * 1024));
+
+ /* allocate space for the serialized histogram list, set header */
+ output = (bytea *) palloc0(total_length);
+ SET_VARSIZE(output, total_length);
+
+ /* we'll use 'data' to keep track of the place to write data */
+ data = VARDATA(output);
+
+ memcpy(data, histogram, offsetof(MVHistogramData, buckets));
+ data += offsetof(MVHistogramData, buckets);
+
+ memcpy(data, info, sizeof(DimensionInfo) * ndims);
+ data += sizeof(DimensionInfo) * ndims;
+
+ /* serialize the deduplicated values for all attributes */
+ for (i = 0; i < ndims; i++)
+ {
+#ifdef USE_ASSERT_CHECKING
+ char *tmp = data;
+#endif
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ Datum v = values[i][j];
+
+ if (info[i].typbyval) /* passed by value */
+ {
+ memcpy(data, &v, info[i].typlen);
+ data += info[i].typlen;
+ }
+ else if (info[i].typlen > 0) /* pased by reference */
+ {
+ memcpy(data, DatumGetPointer(v), info[i].typlen);
+ data += info[i].typlen;
+ }
+ else if (info[i].typlen == -1) /* varlena */
+ {
+ memcpy(data, DatumGetPointer(v), VARSIZE_ANY(v));
+ data += VARSIZE_ANY(values[i][j]);
+ }
+ else if (info[i].typlen == -2) /* cstring */
+ {
+ memcpy(data, DatumGetPointer(v), strlen(DatumGetPointer(v)) + 1);
+ data += strlen(DatumGetPointer(v)) + 1;
+ }
+ }
+
+ /* make sure we got exactly the amount of data we expected */
+ Assert((data - tmp) == info[i].nbytes);
+ }
+
+ /* finally serialize the items, with uint16 indexes instead of the values */
+ for (i = 0; i < nbuckets; i++)
+ {
+ /* don't write beyond the allocated space */
+ Assert(data <= (char *) output + total_length - bucketsize);
+
+ /* reset the values for each item */
+ memset(bucket, 0, bucketsize);
+
+ BUCKET_NTUPLES(bucket) = histogram->buckets[i]->ntuples;
+
+ for (j = 0; j < ndims; j++)
+ {
+ /* do the lookup only for non-NULL values */
+ if (!histogram->buckets[i]->nullsonly[j])
+ {
+ uint16 idx;
+ Datum *v = NULL;
+
+ /* min boundary */
+ v = (Datum *) bsearch_arg(&histogram->buckets[i]->min[j],
+ values[j], info[j].nvalues, sizeof(Datum),
+ compare_scalars_simple, &ssup[j]);
+
+ Assert(v != NULL); /* serialization or deduplication
+ * error */
+
+ /* compute index within the array */
+ idx = (v - values[j]);
+
+ Assert((idx >= 0) && (idx < info[j].nvalues));
+
+ BUCKET_MIN_INDEXES(bucket, ndims)[j] = idx;
+
+ /* max boundary */
+ v = (Datum *) bsearch_arg(&histogram->buckets[i]->max[j],
+ values[j], info[j].nvalues, sizeof(Datum),
+ compare_scalars_simple, &ssup[j]);
+
+ Assert(v != NULL); /* serialization or deduplication
+ * error */
+
+ /* compute index within the array */
+ idx = (v - values[j]);
+
+ Assert((idx >= 0) && (idx < info[j].nvalues));
+
+ BUCKET_MAX_INDEXES(bucket, ndims)[j] = idx;
+ }
+ }
+
+ /* copy flags (nulls, min/max inclusive) */
+ memcpy(BUCKET_NULLS_ONLY(bucket, ndims),
+ histogram->buckets[i]->nullsonly, sizeof(bool) * ndims);
+
+ memcpy(BUCKET_MIN_INCL(bucket, ndims),
+ histogram->buckets[i]->min_inclusive, sizeof(bool) * ndims);
+
+ memcpy(BUCKET_MAX_INCL(bucket, ndims),
+ histogram->buckets[i]->max_inclusive, sizeof(bool) * ndims);
+
+ /* copy the item into the array */
+ memcpy(data, bucket, bucketsize);
+
+ data += bucketsize;
+ }
+
+ /* at this point we expect to match the total_length exactly */
+ Assert((data - (char *) output) == total_length);
+
+ /* free the values/counts arrays here */
+ pfree(counts);
+ pfree(info);
+ pfree(ssup);
+
+ for (i = 0; i < ndims; i++)
+ pfree(values[i]);
+
+ pfree(values);
+
+ return output;
+}
+
+/*
+ * Returns histogram in a partially-serialized form (keeps the boundary values
+ * deduplicated, so that it's possible to optimize the estimation part by
+ * caching function call results between buckets etc.).
+ */
+MVSerializedHistogram
+deserialize_mv_histogram(bytea *data)
+{
+ int i = 0,
+ j = 0;
+
+ Size expected_size;
+ char *tmp = NULL;
+
+ MVSerializedHistogram histogram;
+ DimensionInfo *info;
+
+ int nbuckets;
+ int ndims;
+ int bucketsize;
+
+ /* temporary deserialization buffer */
+ int bufflen;
+ char *buff;
+ char *ptr;
+
+ if (data == NULL)
+ return NULL;
+
+ if (VARSIZE_ANY_EXHDR(data) < offsetof(MVSerializedHistogramData, buckets))
+ elog(ERROR, "invalid histogram size %ld (expected at least %ld)",
+ VARSIZE_ANY_EXHDR(data), offsetof(MVSerializedHistogramData, buckets));
+
+ /* read the histogram header */
+ histogram
+ = (MVSerializedHistogram) palloc(sizeof(MVSerializedHistogramData));
+
+ /* initialize pointer to the data part (skip the varlena header) */
+ tmp = VARDATA_ANY(data);
+
+ /* get the header and perform basic sanity checks */
+ memcpy(histogram, tmp, offsetof(MVSerializedHistogramData, buckets));
+ tmp += offsetof(MVSerializedHistogramData, buckets);
+
+ if (histogram->magic != MVSTAT_HIST_MAGIC)
+ elog(ERROR, "invalid histogram magic %d (expected %dd)",
+ histogram->magic, MVSTAT_HIST_MAGIC);
+
+ if (histogram->type != MVSTAT_HIST_TYPE_BASIC)
+ elog(ERROR, "invalid histogram type %d (expected %dd)",
+ histogram->type, MVSTAT_HIST_TYPE_BASIC);
+
+ nbuckets = histogram->nbuckets;
+ ndims = histogram->ndimensions;
+ bucketsize = BUCKET_SIZE(ndims);
+
+ Assert((nbuckets > 0) && (nbuckets <= MVSTAT_HIST_MAX_BUCKETS));
+ Assert((ndims >= 2) && (ndims <= MVSTATS_MAX_DIMENSIONS));
+
+ /*
+ * What size do we expect with those parameters (it's incomplete, as we
+ * yet have to count the array sizes (from DimensionInfo records).
+ */
+ expected_size = offsetof(MVSerializedHistogramData, buckets) +
+ ndims * sizeof(DimensionInfo) +
+ (nbuckets * bucketsize);
+
+ /* check that we have at least the DimensionInfo records */
+ if (VARSIZE_ANY_EXHDR(data) < expected_size)
+ elog(ERROR, "invalid histogram size %ld (expected %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ info = (DimensionInfo *) (tmp);
+ tmp += ndims * sizeof(DimensionInfo);
+
+ /* account for the value arrays */
+ for (i = 0; i < ndims; i++)
+ expected_size += info[i].nbytes;
+
+ if (VARSIZE_ANY_EXHDR(data) != expected_size)
+ elog(ERROR, "invalid histogram size %ld (expected %ld)",
+ VARSIZE_ANY_EXHDR(data), expected_size);
+
+ /* looks OK - not corrupted or something */
+
+ /* a single buffer for all the values and counts */
+ bufflen = (sizeof(int) + sizeof(Datum *)) * ndims;
+
+ for (i = 0; i < ndims; i++)
+ /* don't allocate space for byval types, matching Datum */
+ if (!(info[i].typbyval && (info[i].typlen == sizeof(Datum))))
+ bufflen += (sizeof(Datum) * info[i].nvalues);
+
+ /* also, include space for the result, tracking the buckets */
+ bufflen += nbuckets * (
+ sizeof(MVSerializedBucket) + /* bucket pointer */
+ sizeof(MVSerializedBucketData)); /* bucket data */
+
+ buff = palloc0(bufflen);
+ ptr = buff;
+
+ histogram->nvalues = (int *) ptr;
+ ptr += (sizeof(int) * ndims);
+
+ histogram->values = (Datum **) ptr;
+ ptr += (sizeof(Datum *) * ndims);
+
+ /*
+ * FIXME This uses pointers to the original data array (the types not
+ * passed by value), so when someone frees the memory, e.g. by doing
+ * something like this:
+ *
+ * bytea * data = ... fetch the data from catalog ... MCVList mcvlist =
+ * deserialize_mcv_list(data); pfree(data);
+ *
+ * then 'mcvlist' references the freed memory. This needs to copy the
+ * pieces.
+ *
+ * TODO same as in MCV deserialization / consider moving to common.c
+ */
+ for (i = 0; i < ndims; i++)
+ {
+ histogram->nvalues[i] = info[i].nvalues;
+
+ if (info[i].typbyval)
+ {
+ /* passed by value / Datum - simply reuse the array */
+ if (info[i].typlen == sizeof(Datum))
+ {
+ histogram->values[i] = (Datum *) tmp;
+ tmp += info[i].nbytes;
+ }
+ else
+ {
+ histogram->values[i] = (Datum *) ptr;
+ ptr += (sizeof(Datum) * info[i].nvalues);
+
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ memcpy(&histogram->values[i][j], tmp, info[i].typlen);
+ tmp += info[i].typlen;
+ }
+ }
+ }
+ else
+ {
+ /* all the other types need a chunk of the buffer */
+ histogram->values[i] = (Datum *) ptr;
+ ptr += (sizeof(Datum) * info[i].nvalues);
+
+ if (info[i].typlen > 0)
+ {
+ /* pased by reference, but fixed length (name, tid, ...) */
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ histogram->values[i][j] = PointerGetDatum(tmp);
+ tmp += info[i].typlen;
+ }
+ }
+ else if (info[i].typlen == -1)
+ {
+ /* varlena */
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ histogram->values[i][j] = PointerGetDatum(tmp);
+ tmp += VARSIZE_ANY(tmp);
+ }
+ }
+ else if (info[i].typlen == -2)
+ {
+ /* cstring */
+ for (j = 0; j < info[i].nvalues; j++)
+ {
+ /* just point into the array */
+ histogram->values[i][j] = PointerGetDatum(tmp);
+ tmp += (strlen(tmp) + 1); /* don't forget the \0 */
+ }
+ }
+ }
+ }
+
+ histogram->buckets = (MVSerializedBucket *) ptr;
+ ptr += (sizeof(MVSerializedBucket) * nbuckets);
+
+ for (i = 0; i < nbuckets; i++)
+ {
+ MVSerializedBucket bucket = (MVSerializedBucket) ptr;
+
+ ptr += sizeof(MVSerializedBucketData);
+
+ bucket->ntuples = BUCKET_NTUPLES(tmp);
+ bucket->nullsonly = BUCKET_NULLS_ONLY(tmp, ndims);
+ bucket->min_inclusive = BUCKET_MIN_INCL(tmp, ndims);
+ bucket->max_inclusive = BUCKET_MAX_INCL(tmp, ndims);
+
+ bucket->min = BUCKET_MIN_INDEXES(tmp, ndims);
+ bucket->max = BUCKET_MAX_INDEXES(tmp, ndims);
+
+ histogram->buckets[i] = bucket;
+
+ Assert(tmp <= (char *) data + VARSIZE_ANY(data));
+
+ tmp += bucketsize;
+ }
+
+ /* at this point we expect to match the total_length exactly */
+ Assert((tmp - VARDATA(data)) == expected_size);
+
+ /* we should exhaust the output buffer exactly */
+ Assert((ptr - buff) == bufflen);
+
+ return histogram;
+}
+
+/*
+ * Build the initial bucket, which will be then split into smaller ones.
+ */
+static MVBucket
+create_initial_mv_bucket(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats)
+{
+ int i;
+ int numattrs = attrs->dim1;
+ HistogramBuild data = NULL;
+
+ /* TODO allocate bucket as a single piece, including all the fields. */
+ MVBucket bucket = (MVBucket) palloc0(sizeof(MVBucketData));
+
+ Assert(numrows > 0);
+ Assert(rows != NULL);
+ Assert((numattrs >= 2) && (numattrs <= MVSTATS_MAX_DIMENSIONS));
+
+ /* allocate the per-dimension arrays */
+
+ /* flags for null-only dimensions */
+ bucket->nullsonly = (bool *) palloc0(numattrs * sizeof(bool));
+
+ /* inclusiveness boundaries - lower/upper bounds */
+ bucket->min_inclusive = (bool *) palloc0(numattrs * sizeof(bool));
+ bucket->max_inclusive = (bool *) palloc0(numattrs * sizeof(bool));
+
+ /* lower/upper boundaries */
+ bucket->min = (Datum *) palloc0(numattrs * sizeof(Datum));
+ bucket->max = (Datum *) palloc0(numattrs * sizeof(Datum));
+
+ /* build-data */
+ data = (HistogramBuild) palloc0(sizeof(HistogramBuildData));
+
+ /* number of distinct values (per dimension) */
+ data->ndistincts = (uint32 *) palloc0(numattrs * sizeof(uint32));
+
+ /* all the sample rows fall into the initial bucket */
+ data->numrows = numrows;
+ data->rows = rows;
+
+ bucket->build_data = data;
+
+ /*
+ * Update the number of ndistinct combinations in the bucket (which we use
+ * when selecting bucket to partition), and then number of distinct values
+ * for each partition (which we use when choosing which dimension to
+ * split).
+ */
+ update_bucket_ndistinct(bucket, attrs, stats);
+
+ /* Update ndistinct (and also set min/max) for all dimensions. */
+ for (i = 0; i < numattrs; i++)
+ update_dimension_ndistinct(bucket, i, attrs, stats, true);
+
+ return bucket;
+}
+
+/*
+ * Choose the bucket to partition next.
+ *
+ * The current criteria is rather simple, chosen so that the algorithm produces
+ * buckets with about equal frequency and regular size. We select the bucket
+ * with the highest number of distinct values, and then split it by the longest
+ * dimension.
+ *
+ * The distinct values are uniformly mapped to [0,1] interval, and this is used
+ * to compute length of the value range.
+ *
+ * NOTE: This is not the same array used for deduplication, as this contains
+ * values for all the tuples from the sample, not just the boundary values.
+ *
+ * Returns either pointer to the bucket selected to be partitioned, or NULL if
+ * there are no buckets that may be split (e.g. if all buckets are too small
+ * or contain too few distinct values).
+ *
+ *
+ * Tricky example
+ * --------------
+ *
+ * Consider this table:
+ *
+ * CREATE TABLE t AS SELECT i AS a, i AS b
+ * FROM generate_series(1,1000000) s(i);
+ *
+ * CREATE STATISTICS s1 ON t (a,b) WITH (histogram);
+ *
+ * ANALYZE t;
+ *
+ * It's a very specific (and perhaps artificial) example, because every bucket
+ * always has exactly the same number of distinct values in all dimensions,
+ * which makes the partitioning tricky.
+ *
+ * Then:
+ *
+ * SELECT * FROM t WHERE (a < 100) AND (b < 100);
+ *
+ * is estimated to return ~120 rows, while in reality it returns only 99.
+ *
+ * QUERY PLAN
+ * -------------------------------------------------------------
+ * Seq Scan on t (cost=0.00..19425.00 rows=117 width=8)
+ * (actual time=0.129..82.776 rows=99 loops=1)
+ * Filter: ((a < 100) AND (b < 100))
+ * Rows Removed by Filter: 999901
+ * Planning time: 1.286 ms
+ * Execution time: 82.984 ms
+ * (5 rows)
+ *
+ * So this estimate is reasonably close. Let's change the query to OR clause:
+ *
+ * SELECT * FROM t WHERE (a < 100) OR (b < 100);
+ *
+ * QUERY PLAN
+ * -------------------------------------------------------------
+ * Seq Scan on t (cost=0.00..19425.00 rows=8100 width=8)
+ * (actual time=0.145..99.910 rows=99 loops=1)
+ * Filter: ((a < 100) OR (b < 100))
+ * Rows Removed by Filter: 999901
+ * Planning time: 1.578 ms
+ * Execution time: 100.132 ms
+ * (5 rows)
+ *
+ * That's clearly a much worse estimate. This happens because the histogram
+ * contains buckets like this:
+ *
+ * bucket 592 [3 30310] [30134 30593] => [0.000233]
+ *
+ * i.e. the length of "a" dimension is (30310-3)=30307, while the length of "b"
+ * is (30593-30134)=459. So the "b" dimension is much narrower than "a".
+ * Of course, there are also buckets where "b" is the wider dimension.
+ *
+ * This is partially mitigated by selecting the "longest" dimension but that
+ * only happens after we already selected the bucket. So if we never select the
+ * bucket, this optimization does not apply.
+ *
+ * The other reason why this particular example behaves so poorly is due to the
+ * way we actually split the selected bucket. We do attempt to divide the bucket
+ * into two parts containing about the same number of tuples, but that does not
+ * too well when most of the tuples is squashed on one side of the bucket.
+ *
+ * For example for columns with data on the diagonal (i.e. when a=b), we end up
+ * with a narrow bucket on the diagonal and a huge bucket overing the remaining
+ * part (with much lower density).
+ *
+ * So perhaps we need two partitioning strategies - one aiming to split buckets
+ * with high frequency (number of sampled rows), the other aiming to split
+ * "large" buckets. And alternating between them, somehow.
+ *
+ * TODO Consider using similar lower boundary for row count as for simple
+ * histograms, i.e. 300 tuples per bucket.
+ */
+static MVBucket
+select_bucket_to_partition(int nbuckets, MVBucket *buckets)
+{
+ int i;
+ int numrows = 0;
+ MVBucket bucket = NULL;
+
+ for (i = 0; i < nbuckets; i++)
+ {
+ HistogramBuild data = (HistogramBuild) buckets[i]->build_data;
+
+ /* if the number of rows is higher, use this bucket */
+ if ((data->ndistinct > 2) &&
+ (data->numrows > numrows) &&
+ (data->numrows >= MIN_BUCKET_ROWS))
+ {
+ bucket = buckets[i];
+ numrows = data->numrows;
+ }
+ }
+
+ /* may be NULL if there are not buckets with (ndistinct>1) */
+ return bucket;
+}
+
+/*
+ * A simple bucket partitioning implementation - we choose the longest bucket
+ * dimension, measured using the array of distinct values built at the very
+ * beginning of the build.
+ *
+ * We map all the distinct values to a [0,1] interval, uniformly distributed,
+ * and then use this to measure length. It's essentially a number of distinct
+ * values within the range, normalized to [0,1].
+ *
+ * Then we choose a 'middle' value splitting the bucket into two parts with
+ * roughly the same frequency.
+ *
+ * This splits the bucket by tweaking the existing one, and returning the new
+ * bucket (essentially shrinking the existing one in-place and returning the
+ * other "half" as a new bucket). The caller is responsible for adding the new
+ * bucket into the list of buckets.
+ *
+ * There are multiple histogram options, centered around the partitioning
+ * criteria, specifying both how to choose a bucket and the dimension most in
+ * need of a split. For a nice summary and general overview, see "rK-Hist : an
+ * R-Tree based histogram for multi-dimensional selectivity estimation" thesis
+ * by J. A. Lopez, Concordia University, p.34-37 (and possibly p. 32-34 for
+ * explanation of the terms).
+ *
+ * It requires care to prevent splitting only one dimension and not splitting
+ * another one at all (which might happen easily in case of strongly dependent
+ * columns - e.g. y=x). The current algorithm minimizes this, but may still
+ * happen for perfectly dependent examples (when all the dimensions have equal
+ * length, the first one will be selected).
+ *
+ * TODO Should probably consider statistics target for the columns (e.g.
+ * to split dimensions with higher statistics target more frequently).
+ */
+static MVBucket
+partition_bucket(MVBucket bucket, int2vector *attrs,
+ VacAttrStats **stats,
+ int *ndistvalues, Datum **distvalues)
+{
+ int i;
+ int dimension;
+ int numattrs = attrs->dim1;
+
+ Datum split_value;
+ MVBucket new_bucket;
+ HistogramBuild new_data;
+
+ /* needed for sort, when looking for the split value */
+ bool isNull;
+ int nvalues = 0;
+ HistogramBuild data = (HistogramBuild) bucket->build_data;
+ StdAnalyzeData *mystats = NULL;
+ ScalarItem *values = (ScalarItem *) palloc0(data->numrows * sizeof(ScalarItem));
+ SortSupportData ssup;
+
+ int nrows = 1; /* number of rows below current value */
+ double delta;
+
+ /* needed when splitting the values */
+ HeapTuple *oldrows = data->rows;
+ int oldnrows = data->numrows;
+
+ /*
+ * We can't split buckets with a single distinct value (this also
+ * disqualifies NULL-only dimensions). Also, there has to be multiple
+ * sample rows (otherwise, how could there be more distinct values).
+ */
+ Assert(data->ndistinct > 1);
+ Assert(data->numrows > 1);
+ Assert((numattrs >= 2) && (numattrs <= MVSTATS_MAX_DIMENSIONS));
+
+ /* Look for the next dimension to split. */
+ delta = 0.0;
+ dimension = -1;
+
+ for (i = 0; i < numattrs; i++)
+ {
+ Datum *a,
+ *b;
+
+ mystats = (StdAnalyzeData *) stats[i]->extra_data;
+
+ /* initialize sort support, etc. */
+ memset(&ssup, 0, sizeof(ssup));
+ ssup.ssup_cxt = CurrentMemoryContext;
+
+ /* We always use the default collation for statistics */
+ ssup.ssup_collation = DEFAULT_COLLATION_OID;
+ ssup.ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(mystats->ltopr, &ssup);
+
+ /* can't split NULL-only dimension */
+ if (bucket->nullsonly[i])
+ continue;
+
+ /* can't split dimension with a single ndistinct value */
+ if (data->ndistincts[i] <= 1)
+ continue;
+
+ /* search for min boundary in the distinct list */
+ a = (Datum *) bsearch_arg(&bucket->min[i],
+ distvalues[i], ndistvalues[i],
+ sizeof(Datum), compare_scalars_simple, &ssup);
+
+ b = (Datum *) bsearch_arg(&bucket->max[i],
+ distvalues[i], ndistvalues[i],
+ sizeof(Datum), compare_scalars_simple, &ssup);
+
+ /* if this dimension is 'larger' then partition by it */
+ if (((b - a) * 1.0 / ndistvalues[i]) > delta)
+ {
+ delta = ((b - a) * 1.0 / ndistvalues[i]);
+ dimension = i;
+ }
+ }
+
+ /*
+ * If we haven't found a dimension here, we've done something wrong in
+ * select_bucket_to_partition.
+ */
+ Assert(dimension != -1);
+
+ /*
+ * Walk through the selected dimension, collect and sort the values and
+ * then choose the value to use as the new boundary.
+ */
+ mystats = (StdAnalyzeData *) stats[dimension]->extra_data;
+
+ /* initialize sort support, etc. */
+ memset(&ssup, 0, sizeof(ssup));
+ ssup.ssup_cxt = CurrentMemoryContext;
+
+ /* We always use the default collation for statistics */
+ ssup.ssup_collation = DEFAULT_COLLATION_OID;
+ ssup.ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(mystats->ltopr, &ssup);
+
+ for (i = 0; i < data->numrows; i++)
+ {
+ /*
+ * remember the index of the sample row, to make the partitioning
+ * simpler
+ */
+ values[nvalues].value = heap_getattr(data->rows[i], attrs->values[dimension],
+ stats[dimension]->tupDesc, &isNull);
+ values[nvalues].tupno = i;
+
+ /* no NULL values allowed here (we never split null-only dimension) */
+ Assert(!isNull);
+
+ nvalues++;
+ }
+
+ /* sort the array of values */
+ qsort_arg((void *) values, nvalues, sizeof(ScalarItem),
+ compare_scalars_partition, (void *) &ssup);
+
+ /*
+ * We know there are bucket->ndistincts[dimension] distinct values in this
+ * dimension, and we want to split this into half, so walk through the
+ * array and stop once we see (ndistinct/2) values.
+ *
+ * We always choose the "next" value, i.e. (n/2+1)-th distinct value, and
+ * use it as an exclusive upper boundary (and inclusive lower boundary).
+ *
+ * TODO Maybe we should use "average" of the two middle distinct values
+ * (at least for even distinct counts), but that would require being able
+ * to do an average (which does not work for non-numeric types).
+ *
+ * TODO Another option is to look for a split that'd give about 50% tuples
+ * (not distinct values) in each partition. That might work better when
+ * there are a few very frequent values, and many rare ones.
+ */
+ delta = fabs(data->numrows);
+ split_value = values[0].value;
+
+ for (i = 1; i < data->numrows; i++)
+ {
+ if (values[i].value != values[i - 1].value)
+ {
+ /* are we closer to splitting the bucket in half? */
+ if (fabs(i - data->numrows / 2.0) < delta)
+ {
+ /* let's assume we'll use this value for the split */
+ split_value = values[i].value;
+ delta = fabs(i - data->numrows / 2.0);
+ nrows = i;
+ }
+ }
+ }
+
+ Assert(nrows > 0);
+ Assert(nrows < data->numrows);
+
+ /*
+ * create the new bucket as a (incomplete) copy of the one being
+ * partitioned.
+ */
+ new_bucket = copy_mv_bucket(bucket, numattrs);
+ new_data = (HistogramBuild) new_bucket->build_data;
+
+ /*
+ * Do the actual split of the chosen dimension, using the split value as
+ * the upper bound for the existing bucket, and lower bound for the new
+ * one.
+ */
+ bucket->max[dimension] = split_value;
+ new_bucket->min[dimension] = split_value;
+
+ /*
+ * We also treat only one side of the new boundary as inclusive, in the
+ * bucket where it happens to be the upper boundary. We never set the
+ * min_inclusive[] to false anywhere, but we set it to true anyway.
+ */
+ bucket->max_inclusive[dimension] = false;
+ new_bucket->min_inclusive[dimension] = true;
+
+ /*
+ * Redistribute the sample tuples using the 'ScalarItem->tupno' index. We
+ * know 'nrows' rows should remain in the original bucket and the rest
+ * goes to the new one.
+ */
+
+ data->rows = (HeapTuple *) palloc0(nrows * sizeof(HeapTuple));
+ new_data->rows = (HeapTuple *) palloc0((oldnrows - nrows) * sizeof(HeapTuple));
+
+ data->numrows = nrows;
+ new_data->numrows = (oldnrows - nrows);
+
+ /*
+ * The first nrows should go to the first bucket, the rest should go to
+ * the new one. Use the tupno field to get the actual HeapTuple row from
+ * the original array of sample rows.
+ */
+ for (i = 0; i < nrows; i++)
+ memcpy(&data->rows[i], &oldrows[values[i].tupno], sizeof(HeapTuple));
+
+ for (i = nrows; i < oldnrows; i++)
+ memcpy(&new_data->rows[i - nrows], &oldrows[values[i].tupno], sizeof(HeapTuple));
+
+ /* update ndistinct values for the buckets (total and per dimension) */
+ update_bucket_ndistinct(bucket, attrs, stats);
+ update_bucket_ndistinct(new_bucket, attrs, stats);
+
+ /*
+ * TODO We don't need to do this for the dimension we used for split,
+ * because we know how many distinct values went to each partition.
+ */
+ for (i = 0; i < numattrs; i++)
+ {
+ update_dimension_ndistinct(bucket, i, attrs, stats, false);
+ update_dimension_ndistinct(new_bucket, i, attrs, stats, false);
+ }
+
+ pfree(oldrows);
+ pfree(values);
+
+ return new_bucket;
+}
+
+/*
+ * Copy a histogram bucket. The copy does not include the build-time data, i.e.
+ * sampled rows etc.
+ */
+static MVBucket
+copy_mv_bucket(MVBucket bucket, uint32 ndimensions)
+{
+ /* TODO allocate as a single piece (including all the fields) */
+ MVBucket new_bucket = (MVBucket) palloc0(sizeof(MVBucketData));
+ HistogramBuild data = (HistogramBuild) palloc0(sizeof(HistogramBuildData));
+
+ /*
+ * Copy only the attributes that will stay the same after the split, and
+ * we'll recompute the rest after the split.
+ */
+
+ /* allocate the per-dimension arrays */
+ new_bucket->nullsonly = (bool *) palloc0(ndimensions * sizeof(bool));
+
+ /* inclusiveness boundaries - lower/upper bounds */
+ new_bucket->min_inclusive = (bool *) palloc0(ndimensions * sizeof(bool));
+ new_bucket->max_inclusive = (bool *) palloc0(ndimensions * sizeof(bool));
+
+ /* lower/upper boundaries */
+ new_bucket->min = (Datum *) palloc0(ndimensions * sizeof(Datum));
+ new_bucket->max = (Datum *) palloc0(ndimensions * sizeof(Datum));
+
+ /* copy data */
+ memcpy(new_bucket->nullsonly, bucket->nullsonly, ndimensions * sizeof(bool));
+
+ memcpy(new_bucket->min_inclusive, bucket->min_inclusive, ndimensions * sizeof(bool));
+ memcpy(new_bucket->min, bucket->min, ndimensions * sizeof(Datum));
+
+ memcpy(new_bucket->max_inclusive, bucket->max_inclusive, ndimensions * sizeof(bool));
+ memcpy(new_bucket->max, bucket->max, ndimensions * sizeof(Datum));
+
+ /* allocate and copy the interesting part of the build data */
+ data->ndistincts = (uint32 *) palloc0(ndimensions * sizeof(uint32));
+
+ new_bucket->build_data = data;
+
+ return new_bucket;
+}
+
+/*
+ * Counts the number of distinct values in the bucket. This just copies the
+ * Datum values into a simple array, and sorts them using memcmp-based
+ * comparator. That means it only works for pass-by-value data types (assuming
+ * they don't use collations etc.)
+ */
+static void
+update_bucket_ndistinct(MVBucket bucket, int2vector *attrs, VacAttrStats **stats)
+{
+ int i,
+ j;
+ int numattrs = attrs->dim1;
+
+ HistogramBuild data = (HistogramBuild) bucket->build_data;
+ int numrows = data->numrows;
+
+ MultiSortSupport mss = multi_sort_init(numattrs);
+
+ /*
+ * We could collect this while walking through all the attributes above
+ * (this way we have to call heap_getattr twice).
+ */
+ SortItem *items = (SortItem *) palloc0(numrows * sizeof(SortItem));
+ Datum *values = (Datum *) palloc0(numrows * sizeof(Datum) * numattrs);
+ bool *isnull = (bool *) palloc0(numrows * sizeof(bool) * numattrs);
+
+ for (i = 0; i < numrows; i++)
+ {
+ items[i].values = &values[i * numattrs];
+ items[i].isnull = &isnull[i * numattrs];
+ }
+
+ /* prepare the sort function for the first dimension */
+ for (i = 0; i < numattrs; i++)
+ multi_sort_add_dimension(mss, i, i, stats);
+
+ /* collect the values */
+ for (i = 0; i < numrows; i++)
+ for (j = 0; j < numattrs; j++)
+ items[i].values[j]
+ = heap_getattr(data->rows[i], attrs->values[j],
+ stats[j]->tupDesc, &items[i].isnull[j]);
+
+ qsort_arg((void *) items, numrows, sizeof(SortItem),
+ multi_sort_compare, mss);
+
+ data->ndistinct = 1;
+
+ for (i = 1; i < numrows; i++)
+ if (multi_sort_compare(&items[i], &items[i - 1], mss) != 0)
+ data->ndistinct += 1;
+
+ pfree(items);
+ pfree(values);
+ pfree(isnull);
+}
+
+/*
+ * Count distinct values per bucket dimension.
+ */
+static void
+update_dimension_ndistinct(MVBucket bucket, int dimension, int2vector *attrs,
+ VacAttrStats **stats, bool update_boundaries)
+{
+ int j;
+ int nvalues = 0;
+ bool isNull;
+ HistogramBuild data = (HistogramBuild) bucket->build_data;
+ Datum *values = (Datum *) palloc0(data->numrows * sizeof(Datum));
+ SortSupportData ssup;
+
+ StdAnalyzeData *mystats = (StdAnalyzeData *) stats[dimension]->extra_data;
+
+ /* we may already know this is a NULL-only dimension */
+ if (bucket->nullsonly[dimension])
+ data->ndistincts[dimension] = 1;
+
+ memset(&ssup, 0, sizeof(ssup));
+ ssup.ssup_cxt = CurrentMemoryContext;
+
+ /* We always use the default collation for statistics */
+ ssup.ssup_collation = DEFAULT_COLLATION_OID;
+ ssup.ssup_nulls_first = false;
+
+ PrepareSortSupportFromOrderingOp(mystats->ltopr, &ssup);
+
+ for (j = 0; j < data->numrows; j++)
+ {
+ values[nvalues] = heap_getattr(data->rows[j], attrs->values[dimension],
+ stats[dimension]->tupDesc, &isNull);
+
+ /* ignore NULL values */
+ if (!isNull)
+ nvalues++;
+ }
+
+ /* there's always at least 1 distinct value (may be NULL) */
+ data->ndistincts[dimension] = 1;
+
+ /*
+ * if there are only NULL values in the column, mark it so and continue
+ * with the next one
+ */
+ if (nvalues == 0)
+ {
+ pfree(values);
+ bucket->nullsonly[dimension] = true;
+ return;
+ }
+
+ /* sort the array (pass-by-value datum */
+ qsort_arg((void *) values, nvalues, sizeof(Datum),
+ compare_scalars_simple, (void *) &ssup);
+
+ /*
+ * Update min/max boundaries to the smallest bounding box. Generally, this
+ * needs to be done only when constructing the initial bucket.
+ */
+ if (update_boundaries)
+ {
+ /* store the min/max values */
+ bucket->min[dimension] = values[0];
+ bucket->min_inclusive[dimension] = true;
+
+ bucket->max[dimension] = values[nvalues - 1];
+ bucket->max_inclusive[dimension] = true;
+ }
+
+ /*
+ * Walk through the array and count distinct values by comparing
+ * succeeding values.
+ *
+ * FIXME This only works for pass-by-value types (i.e. not VARCHARs etc.).
+ * Although thanks to the deduplication it might work even for those types
+ * (equal values will get the same item in the deduplicated array).
+ */
+ for (j = 1; j < nvalues; j++)
+ {
+ if (values[j] != values[j - 1])
+ data->ndistincts[dimension] += 1;
+ }
+
+ pfree(values);
+}
+
+/*
+ * A properly built histogram must not contain buckets mixing NULL and non-NULL
+ * values in a single dimension. Each dimension may either be marked as 'nulls
+ * only', and thus containing only NULL values, or it must not contain any NULL
+ * values.
+ *
+ * Therefore, if the sample contains NULL values in any of the columns, it's
+ * necessary to build those NULL-buckets. This is done in an iterative way
+ * using this algorithm, operating on a single bucket:
+ *
+ * (1) Check that all dimensions are well-formed (not mixing NULL and
+ * non-NULL values).
+ *
+ * (2) If all dimensions are well-formed, terminate.
+ *
+ * (3) If the dimension contains only NULL values, but is not marked as
+ * NULL-only, mark it as NULL-only and run the algorithm again (on
+ * this bucket).
+ *
+ * (4) If the dimension mixes NULL and non-NULL values, split the bucket
+ * into two parts - one with NULL values, one with non-NULL values
+ * (replacing the current one). Then run the algorithm on both buckets.
+ *
+ * This is executed in a recursive manner, but the number of executions should
+ * be quite low - limited by the number of NULL-buckets. Also, in each branch
+ * the number of nested calls is limited by the number of dimensions
+ * (attributes) of the histogram.
+ *
+ * At the end, there should be buckets with no mixed dimensions. The number of
+ * buckets produced by this algorithm is rather limited - with N dimensions,
+ * there may be only 2^N such buckets (each dimension may be either NULL or
+ * non-NULL). So with 8 dimensions (current value of MVSTATS_MAX_DIMENSIONS)
+ * there may be only 256 such buckets.
+ *
+ * After this, a 'regular' bucket-split algorithm shall run, further optimizing
+ * the histogram.
+ */
+static void
+create_null_buckets(MVHistogram histogram, int bucket_idx,
+ int2vector *attrs, VacAttrStats **stats)
+{
+ int i,
+ j;
+ int null_dim = -1;
+ int null_count = 0;
+ bool null_found = false;
+ MVBucket bucket,
+ null_bucket;
+ int null_idx,
+ curr_idx;
+ HistogramBuild data,
+ null_data;
+
+ /* remember original values from the bucket */
+ int numrows;
+ HeapTuple *oldrows = NULL;
+
+ Assert(bucket_idx < histogram->nbuckets);
+ Assert(histogram->ndimensions == attrs->dim1);
+
+ bucket = histogram->buckets[bucket_idx];
+ data = (HistogramBuild) bucket->build_data;
+
+ numrows = data->numrows;
+ oldrows = data->rows;
+
+ /*
+ * Walk through all rows / dimensions, and stop once we find NULL in a
+ * dimension not yet marked as NULL-only.
+ */
+ for (i = 0; i < data->numrows; i++)
+ {
+ /*
+ * FIXME We don't need to start from the first attribute here - we can
+ * start from the last known dimension.
+ */
+ for (j = 0; j < histogram->ndimensions; j++)
+ {
+ /* Is this a NULL-only dimension? If yes, skip. */
+ if (bucket->nullsonly[j])
+ continue;
+
+ /* found a NULL in that dimension? */
+ if (heap_attisnull(data->rows[i], attrs->values[j]))
+ {
+ null_found = true;
+ null_dim = j;
+ break;
+ }
+ }
+
+ /* terminate if we found attribute with NULL values */
+ if (null_found)
+ break;
+ }
+
+ /* no regular dimension contains NULL values => we're done */
+ if (!null_found)
+ return;
+
+ /* walk through the rows again, count NULL values in 'null_dim' */
+ for (i = 0; i < data->numrows; i++)
+ {
+ if (heap_attisnull(data->rows[i], attrs->values[null_dim]))
+ null_count += 1;
+ }
+
+ Assert(null_count <= data->numrows);
+
+ /*
+ * If (null_count == numrows) the dimension already is NULL-only, but is
+ * not yet marked like that. It's enough to mark it and repeat the process
+ * recursively (until we run out of dimensions).
+ */
+ if (null_count == data->numrows)
+ {
+ bucket->nullsonly[null_dim] = true;
+ create_null_buckets(histogram, bucket_idx, attrs, stats);
+ return;
+ }
+
+ /*
+ * We have to split the bucket into two - one with NULL values in the
+ * dimension, one with non-NULL values. We don't need to sort the data or
+ * anything, but otherwise it's similar to what partition_bucket() does.
+ */
+
+ /* create bucket with NULL-only dimension 'dim' */
+ null_bucket = copy_mv_bucket(bucket, histogram->ndimensions);
+ null_data = (HistogramBuild) null_bucket->build_data;
+
+ /* remember the current array info */
+ oldrows = data->rows;
+ numrows = data->numrows;
+
+ /* we'll keep non-NULL values in the current bucket */
+ data->numrows = (numrows - null_count);
+ data->rows
+ = (HeapTuple *) palloc0(data->numrows * sizeof(HeapTuple));
+
+ /* and the NULL values will go to the new one */
+ null_data->numrows = null_count;
+ null_data->rows
+ = (HeapTuple *) palloc0(null_data->numrows * sizeof(HeapTuple));
+
+ /* mark the dimension as NULL-only (in the new bucket) */
+ null_bucket->nullsonly[null_dim] = true;
+
+ /* walk through the sample rows and distribute them accordingly */
+ null_idx = 0;
+ curr_idx = 0;
+ for (i = 0; i < numrows; i++)
+ {
+ if (heap_attisnull(oldrows[i], attrs->values[null_dim]))
+ /* NULL => copy to the new bucket */
+ memcpy(&null_data->rows[null_idx++], &oldrows[i],
+ sizeof(HeapTuple));
+ else
+ memcpy(&data->rows[curr_idx++], &oldrows[i],
+ sizeof(HeapTuple));
+ }
+
+ /* update ndistinct values for the buckets (total and per dimension) */
+ update_bucket_ndistinct(bucket, attrs, stats);
+ update_bucket_ndistinct(null_bucket, attrs, stats);
+
+ /*
+ * TODO We don't need to do this for the dimension we used for split,
+ * because we know how many distinct values went to each bucket (NULL is
+ * not a value, so NULL buckets get 0, and the other bucket got all the
+ * distinct values).
+ */
+ for (i = 0; i < histogram->ndimensions; i++)
+ {
+ update_dimension_ndistinct(bucket, i, attrs, stats, false);
+ update_dimension_ndistinct(null_bucket, i, attrs, stats, false);
+ }
+
+ pfree(oldrows);
+
+ /* add the NULL bucket to the histogram */
+ histogram->buckets[histogram->nbuckets++] = null_bucket;
+
+ /*
+ * And now run the function recursively on both buckets (the new one
+ * first, because the call may change number of buckets, and it's used as
+ * an index).
+ */
+ create_null_buckets(histogram, (histogram->nbuckets - 1), attrs, stats);
+ create_null_buckets(histogram, bucket_idx, attrs, stats);
+}
+
+/*
+ * SRF with details about buckets of a histogram:
+ *
+ * - bucket ID (0...nbuckets)
+ * - min values (string array)
+ * - max values (string array)
+ * - nulls only (boolean array)
+ * - min inclusive flags (boolean array)
+ * - max inclusive flags (boolean array)
+ * - frequency (double precision)
+ *
+ * The input is the OID of the statistics, and there are no rows returned if the
+ * statistics contains no histogram (or if there's no statistics for the OID).
+ *
+ * The second parameter (type) determines what values will be returned
+ * in the (minvals,maxvals). There are three possible values:
+ *
+ * 0 (actual values)
+ * -----------------
+ * - prints actual values
+ * - using the output function of the data type (as string)
+ * - handy for investigating the histogram
+ *
+ * 1 (distinct index)
+ * ------------------
+ * - prints index of the distinct value (into the serialized array)
+ * - makes it easier to spot neighbor buckets, etc.
+ * - handy for plotting the histogram
+ *
+ * 2 (normalized distinct index)
+ * -----------------------------
+ * - prints index of the distinct value, but normalized into [0,1]
+ * - similar to 1, but shows how 'long' the bucket range is
+ * - handy for plotting the histogram
+ *
+ * When plotting the histogram, be careful as the (1) and (2) options skew the
+ * lengths by distributing the distinct values uniformly. For data types
+ * without a clear meaning of 'distance' (e.g. strings) that is not a big deal,
+ * but for numbers it may be confusing.
+ */
+PG_FUNCTION_INFO_V1(pg_mv_histogram_buckets);
+
+#define OUTPUT_FORMAT_RAW 0
+#define OUTPUT_FORMAT_INDEXES 1
+#define OUTPUT_FORMAT_DISTINCT 2
+
+Datum
+pg_mv_histogram_buckets(PG_FUNCTION_ARGS)
+{
+ FuncCallContext *funcctx;
+ int call_cntr;
+ int max_calls;
+ TupleDesc tupdesc;
+ AttInMetadata *attinmeta;
+
+ Oid mvoid = PG_GETARG_OID(0);
+ int otype = PG_GETARG_INT32(1);
+
+ if ((otype < 0) || (otype > 2))
+ elog(ERROR, "invalid output type specified");
+
+ /* stuff done only on the first call of the function */
+ if (SRF_IS_FIRSTCALL())
+ {
+ MemoryContext oldcontext;
+ MVSerializedHistogram histogram;
+
+ /* create a function context for cross-call persistence */
+ funcctx = SRF_FIRSTCALL_INIT();
+
+ /* switch to memory context appropriate for multiple function calls */
+ oldcontext = MemoryContextSwitchTo(funcctx->multi_call_memory_ctx);
+
+ histogram = load_mv_histogram(mvoid);
+
+ funcctx->user_fctx = histogram;
+
+ /* total number of tuples to be returned */
+ funcctx->max_calls = 0;
+ if (funcctx->user_fctx != NULL)
+ funcctx->max_calls = histogram->nbuckets;
+
+ /* Build a tuple descriptor for our result type */
+ if (get_call_result_type(fcinfo, NULL, &tupdesc) != TYPEFUNC_COMPOSITE)
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("function returning record called in context "
+ "that cannot accept type record")));
+
+ /*
+ * generate attribute metadata needed later to produce tuples from raw
+ * C strings
+ */
+ attinmeta = TupleDescGetAttInMetadata(tupdesc);
+ funcctx->attinmeta = attinmeta;
+
+ MemoryContextSwitchTo(oldcontext);
+ }
+
+ /* stuff done on every call of the function */
+ funcctx = SRF_PERCALL_SETUP();
+
+ call_cntr = funcctx->call_cntr;
+ max_calls = funcctx->max_calls;
+ attinmeta = funcctx->attinmeta;
+
+ if (call_cntr < max_calls) /* do when there is more left to send */
+ {
+ char **values;
+ HeapTuple tuple;
+ Datum result;
+ int2vector *stakeys;
+ Oid relid;
+ double bucket_volume = 1.0;
+ StringInfo bufs;
+
+ char *format;
+ int i;
+
+ Oid *outfuncs;
+ FmgrInfo *fmgrinfo;
+
+ MVSerializedHistogram histogram;
+ MVSerializedBucket bucket;
+
+ histogram = (MVSerializedHistogram) funcctx->user_fctx;
+
+ Assert(call_cntr < histogram->nbuckets);
+
+ bucket = histogram->buckets[call_cntr];
+
+ stakeys = find_mv_attnums(mvoid, &relid);
+
+ /*
+ * The scalar values will be formatted directly, using snprintf.
+ *
+ * The 'array' values will be formatted through StringInfo.
+ */
+ values = (char **) palloc0(9 * sizeof(char *));
+ bufs = (StringInfo) palloc0(9 * sizeof(StringInfoData));
+
+ values[0] = (char *) palloc(64 * sizeof(char));
+
+ initStringInfo(&bufs[1]); /* lower boundaries */
+ initStringInfo(&bufs[2]); /* upper boundaries */
+ initStringInfo(&bufs[3]); /* nulls-only */
+ initStringInfo(&bufs[4]); /* lower inclusive */
+ initStringInfo(&bufs[5]); /* upper inclusive */
+
+ values[6] = (char *) palloc(64 * sizeof(char));
+ values[7] = (char *) palloc(64 * sizeof(char));
+ values[8] = (char *) palloc(64 * sizeof(char));
+
+ /* we need to do this only when printing the actual values */
+ outfuncs = (Oid *) palloc0(sizeof(Oid) * histogram->ndimensions);
+ fmgrinfo = (FmgrInfo *) palloc0(sizeof(FmgrInfo) * histogram->ndimensions);
+
+ /*
+ * lookup output functions for all histogram dimensions
+ *
+ * XXX This might be one in the first call and stored in user_fctx.
+ */
+ for (i = 0; i < histogram->ndimensions; i++)
+ {
+ bool isvarlena;
+
+ getTypeOutputInfo(get_atttype(relid, stakeys->values[i]),
+ &outfuncs[i], &isvarlena);
+
+ fmgr_info(outfuncs[i], &fmgrinfo[i]);
+ }
+
+ snprintf(values[0], 64, "%d", call_cntr); /* bucket ID */
+
+ /*
+ * for the arrays of lower/upper boundaries, formated according to
+ * otype
+ */
+ for (i = 0; i < histogram->ndimensions; i++)
+ {
+ Datum *vals = histogram->values[i];
+
+ uint16 minidx = bucket->min[i];
+ uint16 maxidx = bucket->max[i];
+
+ /*
+ * compute bucket volume, using distinct values as a measure
+ *
+ * XXX Not really sure what to do for NULL dimensions here, so
+ * let's simply count them as '1'.
+ */
+ bucket_volume
+ *= (double) (maxidx - minidx + 1) / (histogram->nvalues[i] - 1);
+
+ if (i == 0)
+ format = "{%s"; /* fist dimension */
+ else if (i < (histogram->ndimensions - 1))
+ format = ", %s"; /* medium dimensions */
+ else
+ format = ", %s}"; /* last dimension */
+
+ appendStringInfo(&bufs[3], format, bucket->nullsonly[i] ? "t" : "f");
+ appendStringInfo(&bufs[4], format, bucket->min_inclusive[i] ? "t" : "f");
+ appendStringInfo(&bufs[5], format, bucket->max_inclusive[i] ? "t" : "f");
+
+ /*
+ * for NULL-only dimension, simply put there the NULL and
+ * continue
+ */
+ if (bucket->nullsonly[i])
+ {
+ if (i == 0)
+ format = "{%s";
+ else if (i < (histogram->ndimensions - 1))
+ format = ", %s";
+ else
+ format = ", %s}";
+
+ appendStringInfo(&bufs[1], format, "NULL");
+ appendStringInfo(&bufs[2], format, "NULL");
+
+ continue;
+ }
+
+ /* otherwise we really need to format the value */
+ switch (otype)
+ {
+ case OUTPUT_FORMAT_RAW: /* actual boundary values */
+
+ if (i == 0)
+ format = "{%s";
+ else if (i < (histogram->ndimensions - 1))
+ format = ", %s";
+ else
+ format = ", %s}";
+
+ appendStringInfo(&bufs[1], format,
+ FunctionCall1(&fmgrinfo[i], vals[minidx]));
+
+ appendStringInfo(&bufs[2], format,
+ FunctionCall1(&fmgrinfo[i], vals[maxidx]));
+
+ break;
+
+ case OUTPUT_FORMAT_INDEXES: /* indexes into deduplicated
+ * arrays */
+
+ if (i == 0)
+ format = "{%d";
+ else if (i < (histogram->ndimensions - 1))
+ format = ", %d";
+ else
+ format = ", %d}";
+
+ appendStringInfo(&bufs[1], format, minidx);
+
+ appendStringInfo(&bufs[2], format, maxidx);
+
+ break;
+
+ case OUTPUT_FORMAT_DISTINCT: /* distinct arrays as measure */
+
+ if (i == 0)
+ format = "{%f";
+ else if (i < (histogram->ndimensions - 1))
+ format = ", %f";
+ else
+ format = ", %f}";
+
+ appendStringInfo(&bufs[1], format,
+ (minidx * 1.0 / (histogram->nvalues[i] - 1)));
+
+ appendStringInfo(&bufs[2], format,
+ (maxidx * 1.0 / (histogram->nvalues[i] - 1)));
+
+ break;
+
+ default:
+ elog(ERROR, "unknown output type: %d", otype);
+ }
+ }
+
+ values[1] = bufs[1].data;
+ values[2] = bufs[2].data;
+ values[3] = bufs[3].data;
+ values[4] = bufs[4].data;
+ values[5] = bufs[5].data;
+
+ snprintf(values[6], 64, "%f", bucket->ntuples); /* frequency */
+ snprintf(values[7], 64, "%f", bucket->ntuples / bucket_volume); /* density */
+ snprintf(values[8], 64, "%f", bucket_volume); /* volume (as a
+ * fraction) */
+
+ /* build a tuple */
+ tuple = BuildTupleFromCStrings(attinmeta, values);
+
+ /* make the tuple into a datum */
+ result = HeapTupleGetDatum(tuple);
+
+ /* clean up (this is not really necessary) */
+ pfree(values[0]);
+ pfree(values[6]);
+ pfree(values[7]);
+ pfree(values[8]);
+
+ resetStringInfo(&bufs[1]);
+ resetStringInfo(&bufs[2]);
+ resetStringInfo(&bufs[3]);
+ resetStringInfo(&bufs[4]);
+ resetStringInfo(&bufs[5]);
+
+ pfree(bufs);
+ pfree(values);
+
+ SRF_RETURN_NEXT(funcctx, result);
+ }
+ else /* do when there is no more left */
+ {
+ SRF_RETURN_DONE(funcctx);
+ }
+}
+
+/*
+ * pg_histogram_in - input routine for type pg_histogram.
+ *
+ * pg_histogram is real enough to be a table column, but it has no operations
+ * of its own, and disallows input too
+ *
+ * XXX This is inspired by what pg_node_tree does.
+ */
+Datum
+pg_histogram_in(PG_FUNCTION_ARGS)
+{
+ /*
+ * pg_node_list stores the data in binary form and parsing text input is
+ * not needed, so disallow this.
+ */
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_histogram")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_histogram - output routine for type PG_HISTOGRAM.
+ *
+ * histograms are serialized into a bytea value, so we simply call byteaout()
+ * to serialize the value into text. But it'd be nice to serialize that into
+ * a meaningful representation (e.g. for inspection by people).
+ *
+ * FIXME not implemented yet, returning dummy value
+ */
+Datum
+pg_histogram_out(PG_FUNCTION_ARGS)
+{
+ return byteaout(fcinfo);
+}
+
+/*
+ * pg_histogram_recv - binary input routine for type pg_histogram.
+ */
+Datum
+pg_histogram_recv(PG_FUNCTION_ARGS)
+{
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot accept a value of type %s", "pg_histogram")));
+
+ PG_RETURN_VOID(); /* keep compiler quiet */
+}
+
+/*
+ * pg_histogram_send - binary output routine for type pg_histogram.
+ *
+ * XXX Histograms are serialized into a bytea value, so let's just send that.
+ */
+Datum
+pg_histogram_send(PG_FUNCTION_ARGS)
+{
+ return byteasend(fcinfo);
+}
+
+#ifdef DEBUG_MVHIST
+/*
+ * prints debugging info about matched histogram buckets (full/partial)
+ *
+ * XXX Currently works only for INT data type.
+ */
+void
+debug_histogram_matches(MVSerializedHistogram mvhist, char *matches)
+{
+ int i,
+ j;
+
+ float ffull = 0,
+ fpartial = 0;
+ int nfull = 0,
+ npartial = 0;
+
+ StringInfoData buf;
+
+ initStringInfo(&buf);
+
+ for (i = 0; i < mvhist->nbuckets; i++)
+ {
+ MVSerializedBucket bucket = mvhist->buckets[i];
+
+ if (!matches[i])
+ continue;
+
+ /* increment the counters */
+ nfull += (matches[i] == MVSTATS_MATCH_FULL) ? 1 : 0;
+ npartial += (matches[i] == MVSTATS_MATCH_PARTIAL) ? 1 : 0;
+
+ /* and also update the frequencies */
+ ffull += (matches[i] == MVSTATS_MATCH_FULL) ? bucket->ntuples : 0;
+ fpartial += (matches[i] == MVSTATS_MATCH_PARTIAL) ? bucket->ntuples : 0;
+
+ resetStringInfo(&buf);
+
+ /* build ranges for all the dimentions */
+ for (j = 0; j < mvhist->ndimensions; j++)
+ {
+ appendStringInfo(&buf, '[%d %d]',
+ DatumGetInt32(mvhist->values[j][bucket->min[j]]),
+ DatumGetInt32(mvhist->values[j][bucket->max[j]]));
+ }
+
+ elog(WARNING, "bucket %d %s => %d [%f]", i, buf.data, matches[i], bucket->ntuples);
+ }
+
+ elog(WARNING, "full=%f partial=%f (%f)", ffull, fpartial, (ffull + 0.5 * fpartial));
+}
+#endif
diff --git a/src/bin/psql/describe.c b/src/bin/psql/describe.c
index db74d93..cf73aec 100644
--- a/src/bin/psql/describe.c
+++ b/src/bin/psql/describe.c
@@ -2298,8 +2298,8 @@ describeOneTableDetails(const char *schemaname,
{
printfPQExpBuffer(&buf,
"SELECT oid, stanamespace::regnamespace AS nsp, staname, stakeys,\n"
- " ndist_enabled, deps_enabled, mcv_enabled,\n"
- " ndist_built, deps_built, mcv_built,\n"
+ " ndist_enabled, deps_enabled, mcv_enabled, hist_enabled,\n"
+ " ndist_built, deps_built, mcv_built, hist_built,\n"
" (SELECT string_agg(attname::text,', ')\n"
" FROM ((SELECT unnest(stakeys) AS attnum) s\n"
" JOIN pg_attribute a ON (starelid = a.attrelid and a.attnum = s.attnum))) AS attnums\n"
@@ -2342,8 +2342,17 @@ describeOneTableDetails(const char *schemaname,
first = false;
}
+ if (!strcmp(PQgetvalue(result, i, 6), "t"))
+ {
+ if (!first)
+ appendPQExpBuffer(&buf, ", histogram");
+ else
+ appendPQExpBuffer(&buf, "(histogram");
+ first = false;
+ }
+
appendPQExpBuffer(&buf, ") ON (%s)",
- PQgetvalue(result, i, 9));
+ PQgetvalue(result, i, 12));
printTableAddFooter(&cont, buf.data);
}
diff --git a/src/include/catalog/pg_cast.h b/src/include/catalog/pg_cast.h
index 80d8ea2..f62ba50 100644
--- a/src/include/catalog/pg_cast.h
+++ b/src/include/catalog/pg_cast.h
@@ -266,6 +266,9 @@ DATA(insert ( 3358 25 0 i i ));
DATA(insert ( 441 17 0 i b ));
DATA(insert ( 441 25 0 i i ));
+/* pg_histogram can be coerced to, but not from, bytea */
+DATA(insert ( 774 17 0 i b ));
+
/*
* Datetime category
diff --git a/src/include/catalog/pg_mv_statistic.h b/src/include/catalog/pg_mv_statistic.h
index 34049d6..d30d3cd9 100644
--- a/src/include/catalog/pg_mv_statistic.h
+++ b/src/include/catalog/pg_mv_statistic.h
@@ -40,11 +40,13 @@ CATALOG(pg_mv_statistic,3381)
bool ndist_enabled; /* build ndist coefficient? */
bool deps_enabled; /* analyze dependencies? */
bool mcv_enabled; /* build MCV list? */
+ bool hist_enabled; /* build histogram? */
/* statistics that are available (if requested) */
bool ndist_built; /* ndistinct coeff built */
bool deps_built; /* dependencies were built */
bool mcv_built; /* MCV list was built */
+ bool hist_built; /* histogram was built */
/*
* variable-length fields start here, but we allow direct access to
@@ -56,6 +58,7 @@ CATALOG(pg_mv_statistic,3381)
pg_ndistinct standist; /* ndistinct coeff (serialized) */
pg_dependencies stadeps; /* dependencies (serialized) */
pg_mcv_list stamcv; /* MCV list (serialized) */
+ pg_histogram stahist; /* MV histogram (serialized) */
#endif
} FormData_pg_mv_statistic;
@@ -71,7 +74,7 @@ typedef FormData_pg_mv_statistic *Form_pg_mv_statistic;
* compiler constants for pg_mv_statistic
* ----------------
*/
-#define Natts_pg_mv_statistic 14
+#define Natts_pg_mv_statistic 17
#define Anum_pg_mv_statistic_starelid 1
#define Anum_pg_mv_statistic_staname 2
#define Anum_pg_mv_statistic_stanamespace 3
@@ -79,12 +82,15 @@ typedef FormData_pg_mv_statistic *Form_pg_mv_statistic;
#define Anum_pg_mv_statistic_ndist_enabled 5
#define Anum_pg_mv_statistic_deps_enabled 6
#define Anum_pg_mv_statistic_mcv_enabled 7
-#define Anum_pg_mv_statistic_ndist_built 8
-#define Anum_pg_mv_statistic_deps_built 9
-#define Anum_pg_mv_statistic_mcv_built 10
-#define Anum_pg_mv_statistic_stakeys 11
-#define Anum_pg_mv_statistic_standist 12
-#define Anum_pg_mv_statistic_stadeps 13
-#define Anum_pg_mv_statistic_stamcv 14
+#define Anum_pg_mv_statistic_hist_enabled 8
+#define Anum_pg_mv_statistic_ndist_built 9
+#define Anum_pg_mv_statistic_deps_built 10
+#define Anum_pg_mv_statistic_mcv_built 11
+#define Anum_pg_mv_statistic_hist_built 12
+#define Anum_pg_mv_statistic_stakeys 13
+#define Anum_pg_mv_statistic_standist 14
+#define Anum_pg_mv_statistic_stadeps 15
+#define Anum_pg_mv_statistic_stamcv 16
+#define Anum_pg_mv_statistic_stahist 17
#endif /* PG_MV_STATISTIC_H */
diff --git a/src/include/catalog/pg_proc.h b/src/include/catalog/pg_proc.h
index 7cf1e5a..653bf1a 100644
--- a/src/include/catalog/pg_proc.h
+++ b/src/include/catalog/pg_proc.h
@@ -2730,6 +2730,10 @@ DATA(insert OID = 3376 ( pg_mv_stats_mcvlist_info PGNSP PGUID 12 1 0 0 0 f f f
DESCR("multi-variate statistics: MCV list info");
DATA(insert OID = 3373 ( pg_mv_mcv_items PGNSP PGUID 12 1 1000 0 0 f f f f t t i s 1 0 2249 "26" "{26,23,1009,1000,701}" "{i,o,o,o,o}" "{oid,index,values,nulls,frequency}" _null_ _null_ pg_mv_mcv_items _null_ _null_ _null_ ));
DESCR("details about MCV list items");
+DATA(insert OID = 3375 ( pg_mv_stats_histogram_info PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 25 "774" _null_ _null_ _null_ _null_ _null_ pg_mv_stats_histogram_info _null_ _null_ _null_ ));
+DESCR("multi-variate statistics: histogram info");
+DATA(insert OID = 3374 ( pg_mv_histogram_buckets PGNSP PGUID 12 1 1000 0 0 f f f f t t i s 2 0 2249 "26 23" "{26,23,23,1009,1009,1000,1000,1000,701,701,701}" "{i,i,o,o,o,o,o,o,o,o,o}" "{oid,otype,index,minvals,maxvals,nullsonly,mininclusive,maxinclusive,frequency,density,bucket_volume}" _null_ _null_ pg_mv_histogram_buckets _null_ _null_ _null_ ));
+DESCR("details about histogram buckets");
DATA(insert OID = 3354 ( pg_ndistinct_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 3353 "2275" _null_ _null_ _null_ _null_ _null_ pg_ndistinct_in _null_ _null_ _null_ ));
DESCR("I/O");
@@ -2758,6 +2762,15 @@ DESCR("I/O");
DATA(insert OID = 445 ( pg_mcv_list_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "441" _null_ _null_ _null_ _null_ _null_ pg_mcv_list_send _null_ _null_ _null_ ));
DESCR("I/O");
+DATA(insert OID = 775 ( pg_histogram_in PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 774 "2275" _null_ _null_ _null_ _null_ _null_ pg_histogram_in _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 776 ( pg_histogram_out PGNSP PGUID 12 1 0 0 0 f f f f t f i s 1 0 2275 "774" _null_ _null_ _null_ _null_ _null_ pg_histogram_out _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 777 ( pg_histogram_recv PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 774 "2281" _null_ _null_ _null_ _null_ _null_ pg_histogram_recv _null_ _null_ _null_ ));
+DESCR("I/O");
+DATA(insert OID = 778 ( pg_histogram_send PGNSP PGUID 12 1 0 0 0 f f f f t f s s 1 0 17 "774" _null_ _null_ _null_ _null_ _null_ pg_histogram_send _null_ _null_ _null_ ));
+DESCR("I/O");
+
DATA(insert OID = 1928 ( pg_stat_get_numscans PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_numscans _null_ _null_ _null_ ));
DESCR("statistics: number of scans done for table/index");
DATA(insert OID = 1929 ( pg_stat_get_tuples_returned PGNSP PGUID 12 1 0 0 0 f f f f t f s r 1 0 20 "26" _null_ _null_ _null_ _null_ _null_ pg_stat_get_tuples_returned _null_ _null_ _null_ ));
diff --git a/src/include/catalog/pg_type.h b/src/include/catalog/pg_type.h
index fbac135..7133862 100644
--- a/src/include/catalog/pg_type.h
+++ b/src/include/catalog/pg_type.h
@@ -376,6 +376,10 @@ DATA(insert OID = 441 ( pg_mcv_list PGNSP PGUID -1 f b S f t \054 0 0 0 pg_mcv_
DESCR("multivariate MCV list");
#define PGMCVLISTOID 441
+DATA(insert OID = 774 ( pg_histogram PGNSP PGUID -1 f b S f t \054 0 0 0 pg_histogram_in pg_histogram_out pg_histogram_recv pg_histogram_send - - - i x f 0 -1 0 100 _null_ _null_ _null_ ));
+DESCR("multivariate histogram");
+#define PGHISTOGRAMOID 774
+
DATA(insert OID = 32 ( pg_ddl_command PGNSP PGUID SIZEOF_POINTER t p P f t \054 0 0 0 pg_ddl_command_in pg_ddl_command_out pg_ddl_command_recv pg_ddl_command_send - - - ALIGNOF_POINTER p f 0 -1 0 0 _null_ _null_ _null_ ));
DESCR("internal type for passing CollectedCommand");
#define PGDDLCOMMANDOID 32
diff --git a/src/include/nodes/relation.h b/src/include/nodes/relation.h
index d912827..f99f547 100644
--- a/src/include/nodes/relation.h
+++ b/src/include/nodes/relation.h
@@ -684,11 +684,13 @@ typedef struct MVStatisticInfo
bool ndist_enabled; /* ndistinct coefficient enabled */
bool deps_enabled; /* functional dependencies enabled */
bool mcv_enabled; /* MCV list enabled */
+ bool hist_enabled; /* histogram enabled */
/* built/available statistics */
bool ndist_built; /* ndistinct coefficient built */
bool deps_built; /* functional dependencies built */
bool mcv_built; /* MCV list built */
+ bool hist_built; /* histogram built */
/* columns in the statistics (attnums) */
int2vector *stakeys; /* attnums of the columns covered */
diff --git a/src/include/utils/builtins.h b/src/include/utils/builtins.h
index 9ed080a..1c7925b 100644
--- a/src/include/utils/builtins.h
+++ b/src/include/utils/builtins.h
@@ -81,6 +81,10 @@ extern Datum pg_mcv_list_in(PG_FUNCTION_ARGS);
extern Datum pg_mcv_list_out(PG_FUNCTION_ARGS);
extern Datum pg_mcv_list_recv(PG_FUNCTION_ARGS);
extern Datum pg_mcv_list_send(PG_FUNCTION_ARGS);
+extern Datum pg_histogram_in(PG_FUNCTION_ARGS);
+extern Datum pg_histogram_out(PG_FUNCTION_ARGS);
+extern Datum pg_histogram_recv(PG_FUNCTION_ARGS);
+extern Datum pg_histogram_send(PG_FUNCTION_ARGS);
/* regexp.c */
extern char *regexp_fixed_prefix(text *text_re, bool case_insensitive,
diff --git a/src/include/utils/mvstats.h b/src/include/utils/mvstats.h
index 0c4f621..5d8c024 100644
--- a/src/include/utils/mvstats.h
+++ b/src/include/utils/mvstats.h
@@ -18,7 +18,7 @@
#include "commands/vacuum.h"
/*
- * Degree of how much MCV item matches a clause.
+ * Degree of how much MCV item / histogram bucket matches a clause.
* This is then considered when computing the selectivity.
*/
#define MVSTATS_MATCH_NONE 0 /* no match at all */
@@ -114,19 +114,133 @@ bool dependency_implies_attribute(MVDependency dependency, AttrNumber attnum,
bool dependency_is_fully_matched(MVDependency dependency, Bitmapset *attnums,
int16 *attmap);
+/* used to flag stats serialized to bytea */
+#define MVSTAT_HIST_MAGIC 0x7F8C5670 /* marks serialized bytea */
+#define MVSTAT_HIST_TYPE_BASIC 1 /* basic histogram type */
+
+/* max buckets in a histogram (mostly arbitrary number */
+#define MVSTAT_HIST_MAX_BUCKETS 16384
+
+/*
+ * Multivariate histograms
+ */
+typedef struct MVBucketData
+{
+
+ /* Frequencies of this bucket. */
+ float ntuples; /* frequency of tuples tuples */
+
+ /*
+ * Information about dimensions being NULL-only. Not yet used.
+ */
+ bool *nullsonly;
+
+ /* lower boundaries - values and information about the inequalities */
+ Datum *min;
+ bool *min_inclusive;
+
+ /* upper boundaries - values and information about the inequalities */
+ Datum *max;
+ bool *max_inclusive;
+
+ /* used when building the histogram (not serialized/deserialized) */
+ void *build_data;
+
+} MVBucketData;
+
+typedef MVBucketData *MVBucket;
+
+
+typedef struct MVHistogramData
+{
+
+ uint32 magic; /* magic constant marker */
+ uint32 type; /* type of histogram (BASIC) */
+ uint32 nbuckets; /* number of buckets (buckets array) */
+ uint32 ndimensions; /* number of dimensions */
+
+ MVBucket *buckets; /* array of buckets */
+
+} MVHistogramData;
+
+typedef MVHistogramData *MVHistogram;
+
+/*
+ * Histogram in a partially serialized form, with deduplicated boundary
+ * values etc.
+ *
+ * TODO add more detailed description here
+ */
+
+typedef struct MVSerializedBucketData
+{
+
+ /* Frequencies of this bucket. */
+ float ntuples; /* frequency of tuples tuples */
+
+ /*
+ * Information about dimensions being NULL-only. Not yet used.
+ */
+ bool *nullsonly;
+
+ /* lower boundaries - values and information about the inequalities */
+ uint16 *min;
+ bool *min_inclusive;
+
+ /*
+ * indexes of upper boundaries - values and information about the
+ * inequalities (exclusive vs. inclusive)
+ */
+ uint16 *max;
+ bool *max_inclusive;
+
+} MVSerializedBucketData;
+
+typedef MVSerializedBucketData *MVSerializedBucket;
+
+typedef struct MVSerializedHistogramData
+{
+
+ uint32 magic; /* magic constant marker */
+ uint32 type; /* type of histogram (BASIC) */
+ uint32 nbuckets; /* number of buckets (buckets array) */
+ uint32 ndimensions; /* number of dimensions */
+
+ /*
+ * keep this the same with MVHistogramData, because of deserialization
+ * (same offset)
+ */
+ MVSerializedBucket *buckets; /* array of buckets */
+
+ /*
+ * serialized boundary values, one array per dimension, deduplicated (the
+ * min/max indexes point into these arrays)
+ */
+ int *nvalues;
+ Datum **values;
+
+} MVSerializedHistogramData;
+
+typedef MVSerializedHistogramData *MVSerializedHistogram;
+
+
MVNDistinct load_mv_ndistinct(Oid mvoid);
MVDependencies load_mv_dependencies(Oid mvoid);
MCVList load_mv_mcvlist(Oid mvoid);
+MVSerializedHistogram load_mv_histogram(Oid mvoid);
bytea *serialize_mv_ndistinct(MVNDistinct ndistinct);
bytea *serialize_mv_dependencies(MVDependencies dependencies);
bytea *serialize_mv_mcvlist(MCVList mcvlist, int2vector *attrs,
VacAttrStats **stats);
+bytea *serialize_mv_histogram(MVHistogram histogram, int2vector *attrs,
+ VacAttrStats **stats);
/* deserialization of stats (serialization is private to analyze) */
MVNDistinct deserialize_mv_ndistinct(bytea *data);
MVDependencies deserialize_mv_dependencies(bytea *data);
MCVList deserialize_mv_mcvlist(bytea *data);
+MVSerializedHistogram deserialize_mv_histogram(bytea * data);
/*
* Returns index of the attribute number within the vector (i.e. a
@@ -139,6 +253,8 @@ int2vector *find_mv_attnums(Oid mvoid, Oid *relid);
/* functions for inspecting the statistics */
extern Datum pg_mv_stats_mcvlist_info(PG_FUNCTION_ARGS);
extern Datum pg_mv_mcvlist_items(PG_FUNCTION_ARGS);
+extern Datum pg_mv_stats_histogram_info(PG_FUNCTION_ARGS);
+extern Datum pg_mv_histogram_buckets(PG_FUNCTION_ARGS);
MVNDistinct build_mv_ndistinct(double totalrows, int numrows, HeapTuple *rows,
@@ -151,8 +267,15 @@ MVDependencies build_mv_dependencies(int numrows, HeapTuple *rows,
MCVList build_mv_mcvlist(int numrows, HeapTuple *rows, int2vector *attrs,
VacAttrStats **stats, int *numrows_filtered);
+MVHistogram build_mv_histogram(int numrows, HeapTuple *rows, int2vector *attrs,
+ VacAttrStats **stats, int numrows_total);
+
void build_mv_stats(Relation onerel, double totalrows,
int numrows, HeapTuple *rows,
int natts, VacAttrStats **vacattrstats);
+#ifdef DEBUG_MVHIST
+extern void debug_histogram_matches(MVSerializedHistogram mvhist, char *matches);
+#endif
+
#endif
diff --git a/src/test/regress/expected/mv_histogram.out b/src/test/regress/expected/mv_histogram.out
new file mode 100644
index 0000000..16410ce
--- /dev/null
+++ b/src/test/regress/expected/mv_histogram.out
@@ -0,0 +1,198 @@
+-- data type passed by value
+CREATE TABLE mv_histogram (
+ a INT,
+ b INT,
+ c INT
+);
+-- unknown column
+CREATE STATISTICS s7 WITH (histogram) ON (unknown_column) FROM mv_histogram;
+ERROR: column "unknown_column" referenced in statistics does not exist
+-- single column
+CREATE STATISTICS s7 WITH (histogram) ON (a) FROM mv_histogram;
+ERROR: statistics require at least 2 columns
+-- single column, duplicated
+CREATE STATISTICS s7 WITH (histogram) ON (a, a) FROM mv_histogram;
+ERROR: duplicate column name in statistics definition
+-- two columns, one duplicated
+CREATE STATISTICS s7 WITH (histogram) ON (a, a, b) FROM mv_histogram;
+ERROR: duplicate column name in statistics definition
+-- unknown option
+CREATE STATISTICS s7 WITH (unknown_option) ON (a, b, c) FROM mv_histogram;
+ERROR: unrecognized STATISTICS option "unknown_option"
+-- correct command
+CREATE STATISTICS s7 WITH (histogram) ON (a, b, c) FROM mv_histogram;
+-- random data (no functional dependencies)
+INSERT INTO mv_histogram
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- a => b, a => c, b => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- a => b, a => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mv_histogram
+ SELECT i/100, i/200, i/400 FROM generate_series(1,30000) s(i);
+CREATE INDEX hist_idx ON mv_histogram (a, b);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a = 10 AND b = 5;
+ QUERY PLAN
+--------------------------------------------
+ Bitmap Heap Scan on mv_histogram
+ Recheck Cond: ((a = 10) AND (b = 5))
+ -> Bitmap Index Scan on hist_idx
+ Index Cond: ((a = 10) AND (b = 5))
+(4 rows)
+
+DROP TABLE mv_histogram;
+-- varlena type (text)
+CREATE TABLE mv_histogram (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+CREATE STATISTICS s8 WITH (histogram) ON (a, b, c) FROM mv_histogram;
+-- random data (no functional dependencies)
+INSERT INTO mv_histogram
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- a => b, a => c, b => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- a => b, a => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+TRUNCATE mv_histogram;
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mv_histogram
+ SELECT i/100, i/200, i/400 FROM generate_series(1,30000) s(i);
+CREATE INDEX hist_idx ON mv_histogram (a, b);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a = '10' AND b = '5';
+ QUERY PLAN
+------------------------------------------------------------
+ Bitmap Heap Scan on mv_histogram
+ Recheck Cond: ((a = '10'::text) AND (b = '5'::text))
+ -> Bitmap Index Scan on hist_idx
+ Index Cond: ((a = '10'::text) AND (b = '5'::text))
+(4 rows)
+
+TRUNCATE mv_histogram;
+-- check explain (expect bitmap index scan, not plain index scan) with NULLs
+INSERT INTO mv_histogram
+ SELECT
+ (CASE WHEN i/100 = 0 THEN NULL ELSE i/100 END),
+ (CASE WHEN i/200 = 0 THEN NULL ELSE i/200 END),
+ (CASE WHEN i/400 = 0 THEN NULL ELSE i/400 END)
+ FROM generate_series(1,30000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a IS NULL AND b IS NULL;
+ QUERY PLAN
+---------------------------------------------------
+ Bitmap Heap Scan on mv_histogram
+ Recheck Cond: ((a IS NULL) AND (b IS NULL))
+ -> Bitmap Index Scan on hist_idx
+ Index Cond: ((a IS NULL) AND (b IS NULL))
+(4 rows)
+
+DROP TABLE mv_histogram;
+-- NULL values (mix of int and text columns)
+CREATE TABLE mv_histogram (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+CREATE STATISTICS s9 WITH (histogram) ON (a, b, c, d) FROM mv_histogram;
+INSERT INTO mv_histogram
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+ hist_enabled | hist_built
+--------------+------------
+ t | t
+(1 row)
+
+DROP TABLE mv_histogram;
diff --git a/src/test/regress/expected/opr_sanity.out b/src/test/regress/expected/opr_sanity.out
index 9969c10..a9d8163 100644
--- a/src/test/regress/expected/opr_sanity.out
+++ b/src/test/regress/expected/opr_sanity.out
@@ -820,11 +820,12 @@ WHERE c.castmethod = 'b' AND
pg_ndistinct | bytea | 0 | i
pg_dependencies | bytea | 0 | i
pg_mcv_list | bytea | 0 | i
+ pg_histogram | bytea | 0 | i
cidr | inet | 0 | i
xml | text | 0 | a
xml | character varying | 0 | a
xml | character | 0 | a
-(10 rows)
+(11 rows)
-- **************** pg_conversion ****************
-- Look for illegal values in pg_conversion fields.
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 2e3c40e..27e903c 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1383,7 +1383,9 @@ pg_mv_stats| SELECT n.nspname AS schemaname,
length((s.standist)::bytea) AS ndistbytes,
length((s.stadeps)::bytea) AS depsbytes,
length((s.stamcv)::bytea) AS mcvbytes,
- pg_mv_stats_mcvlist_info(s.stamcv) AS mcvinfo
+ pg_mv_stats_mcvlist_info(s.stamcv) AS mcvinfo,
+ length((s.stahist)::bytea) AS histbytes,
+ pg_mv_stats_histogram_info(s.stahist) AS histinfo
FROM ((pg_mv_statistic s
JOIN pg_class c ON ((c.oid = s.starelid)))
LEFT JOIN pg_namespace n ON ((n.oid = c.relnamespace)));
diff --git a/src/test/regress/expected/type_sanity.out b/src/test/regress/expected/type_sanity.out
index dde15b9..4d3c4d7 100644
--- a/src/test/regress/expected/type_sanity.out
+++ b/src/test/regress/expected/type_sanity.out
@@ -73,8 +73,9 @@ WHERE p1.typtype not in ('c','d','p') AND p1.typname NOT LIKE E'\\_%'
3353 | pg_ndistinct
3358 | pg_dependencies
441 | pg_mcv_list
+ 774 | pg_histogram
210 | smgr
-(5 rows)
+(6 rows)
-- Make sure typarray points to a varlena array type of our own base
SELECT p1.oid, p1.typname as basetype, p2.typname as arraytype,
diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule
index d805840..36dd618 100644
--- a/src/test/regress/parallel_schedule
+++ b/src/test/regress/parallel_schedule
@@ -118,4 +118,4 @@ test: event_trigger
test: stats
# run tests of multivariate stats
-test: mv_ndistinct mv_dependencies mv_mcv
+test: mv_ndistinct mv_dependencies mv_mcv mv_histogram
diff --git a/src/test/regress/serial_schedule b/src/test/regress/serial_schedule
index 72c6acd..34f5467 100644
--- a/src/test/regress/serial_schedule
+++ b/src/test/regress/serial_schedule
@@ -174,3 +174,4 @@ test: stats
test: mv_ndistinct
test: mv_dependencies
test: mv_mcv
+test: mv_histogram
diff --git a/src/test/regress/sql/mv_histogram.sql b/src/test/regress/sql/mv_histogram.sql
new file mode 100644
index 0000000..55197cb
--- /dev/null
+++ b/src/test/regress/sql/mv_histogram.sql
@@ -0,0 +1,167 @@
+-- data type passed by value
+CREATE TABLE mv_histogram (
+ a INT,
+ b INT,
+ c INT
+);
+
+-- unknown column
+CREATE STATISTICS s7 WITH (histogram) ON (unknown_column) FROM mv_histogram;
+
+-- single column
+CREATE STATISTICS s7 WITH (histogram) ON (a) FROM mv_histogram;
+
+-- single column, duplicated
+CREATE STATISTICS s7 WITH (histogram) ON (a, a) FROM mv_histogram;
+
+-- two columns, one duplicated
+CREATE STATISTICS s7 WITH (histogram) ON (a, a, b) FROM mv_histogram;
+
+-- unknown option
+CREATE STATISTICS s7 WITH (unknown_option) ON (a, b, c) FROM mv_histogram;
+
+-- correct command
+CREATE STATISTICS s7 WITH (histogram) ON (a, b, c) FROM mv_histogram;
+
+-- random data (no functional dependencies)
+INSERT INTO mv_histogram
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- a => b, a => c, b => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- a => b, a => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mv_histogram
+ SELECT i/100, i/200, i/400 FROM generate_series(1,30000) s(i);
+CREATE INDEX hist_idx ON mv_histogram (a, b);
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a = 10 AND b = 5;
+
+DROP TABLE mv_histogram;
+
+-- varlena type (text)
+CREATE TABLE mv_histogram (
+ a TEXT,
+ b TEXT,
+ c TEXT
+);
+
+CREATE STATISTICS s8 WITH (histogram) ON (a, b, c) FROM mv_histogram;
+
+-- random data (no functional dependencies)
+INSERT INTO mv_histogram
+ SELECT mod(i, 111), mod(i, 123), mod(i, 23) FROM generate_series(1,10000) s(i);
+
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- a => b, a => c, b => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/100, i/200 FROM generate_series(1,10000) s(i);
+
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- a => b, a => c
+INSERT INTO mv_histogram
+ SELECT i/10, i/150, i/200 FROM generate_series(1,10000) s(i);
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+TRUNCATE mv_histogram;
+
+-- check explain (expect bitmap index scan, not plain index scan)
+INSERT INTO mv_histogram
+ SELECT i/100, i/200, i/400 FROM generate_series(1,30000) s(i);
+CREATE INDEX hist_idx ON mv_histogram (a, b);
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a = '10' AND b = '5';
+
+TRUNCATE mv_histogram;
+
+-- check explain (expect bitmap index scan, not plain index scan) with NULLs
+INSERT INTO mv_histogram
+ SELECT
+ (CASE WHEN i/100 = 0 THEN NULL ELSE i/100 END),
+ (CASE WHEN i/200 = 0 THEN NULL ELSE i/200 END),
+ (CASE WHEN i/400 = 0 THEN NULL ELSE i/400 END)
+ FROM generate_series(1,30000) s(i);
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+EXPLAIN (COSTS off)
+ SELECT * FROM mv_histogram WHERE a IS NULL AND b IS NULL;
+
+DROP TABLE mv_histogram;
+
+-- NULL values (mix of int and text columns)
+CREATE TABLE mv_histogram (
+ a INT,
+ b TEXT,
+ c INT,
+ d TEXT
+);
+
+CREATE STATISTICS s9 WITH (histogram) ON (a, b, c, d) FROM mv_histogram;
+
+INSERT INTO mv_histogram
+ SELECT
+ mod(i, 100),
+ (CASE WHEN mod(i, 200) = 0 THEN NULL ELSE mod(i,200) END),
+ mod(i, 400),
+ (CASE WHEN mod(i, 300) = 0 THEN NULL ELSE mod(i,600) END)
+ FROM generate_series(1,10000) s(i);
+
+ANALYZE mv_histogram;
+
+SELECT hist_enabled, hist_built
+ FROM pg_mv_statistic WHERE starelid = 'mv_histogram'::regclass;
+
+DROP TABLE mv_histogram;
--
2.5.5
[binary/octet-stream] 0007-WIP-use-ndistinct-for-selectivity-estimation-in--v23.patch (14.6K, ../../1fa982dc-ffb4-eef7-eb55-d8cf7d84f1c4@2ndquadrant.com/8-0007-WIP-use-ndistinct-for-selectivity-estimation-in--v23.patch)
download | inline diff:
From 20a213b27d337bed72256589f0267f30a7724fe8 Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Thu, 27 Oct 2016 15:24:42 +0200
Subject: [PATCH 7/9] WIP: use ndistinct for selectivity estimation in
clausesel.c
---
src/backend/optimizer/path/clausesel.c | 382 ++++++++++++++++++++++++++-------
1 file changed, 299 insertions(+), 83 deletions(-)
diff --git a/src/backend/optimizer/path/clausesel.c b/src/backend/optimizer/path/clausesel.c
index fddbcc4..da5c340 100644
--- a/src/backend/optimizer/path/clausesel.c
+++ b/src/backend/optimizer/path/clausesel.c
@@ -47,9 +47,10 @@ typedef struct RangeQueryClause
static void addRangeClause(RangeQueryClause **rqlist, Node *clause,
bool varonleft, bool isLTsel, Selectivity s2);
-#define STATS_TYPE_FDEPS 0x01
-#define STATS_TYPE_MCV 0x02
-#define STATS_TYPE_HIST 0x04
+#define STATS_TYPE_NDIST 0x01
+#define STATS_TYPE_FDEPS 0x02
+#define STATS_TYPE_MCV 0x04
+#define STATS_TYPE_HIST 0x08
static bool clause_is_mv_compatible(Node *clause, Index relid, Bitmapset **attnums,
int type);
@@ -70,6 +71,10 @@ static List *clauselist_mv_split(PlannerInfo *root, Index relid,
static Selectivity clauselist_mv_selectivity(PlannerInfo *root,
List *clauses, MVStatisticInfo *mvstats);
+static Selectivity clauselist_mv_selectivity_ndist(PlannerInfo *root,
+ Index relid, List *clauses, MVStatisticInfo *mvstats,
+ Index varRelid, JoinType jointype, SpecialJoinInfo *sjinfo);
+
static Selectivity clauselist_mv_selectivity_deps(PlannerInfo *root,
Index relid, List *clauses, MVStatisticInfo *mvstats,
Index varRelid, JoinType jointype, SpecialJoinInfo *sjinfo);
@@ -282,6 +287,37 @@ clauselist_selectivity(PlannerInfo *root,
}
}
+ /* And finally, try to use ndistinct coefficients. */
+ if (has_stats(stats, STATS_TYPE_NDIST) &&
+ (count_mv_attnums(clauses, relid, STATS_TYPE_NDIST) >= 2))
+ {
+ MVStatisticInfo *mvstat;
+ Bitmapset *mvattnums;
+
+ /* collect attributes from the compatible conditions */
+ mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_NDIST);
+
+ /* and search for the statistic covering the most attributes */
+ mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_NDIST);
+
+ if (mvstat != NULL) /* we have a matching stats */
+ {
+ /* clauses compatible with multi-variate stats */
+ List *mvclauses = NIL;
+
+ /* split the clauselist into regular and mv-clauses */
+ clauses = clauselist_mv_split(root, relid, clauses, &mvclauses,
+ mvstat, STATS_TYPE_NDIST);
+
+ /* we've chosen the histogram to match the clauses */
+ Assert(mvclauses != NIL);
+
+ /* compute the multivariate stats (dependencies) */
+ s1 *= clauselist_mv_selectivity_ndist(root, relid, mvclauses, mvstat,
+ varRelid, jointype, sjinfo);
+ }
+ }
+
/*
* Initial scan over clauses. Anything that doesn't look like a potential
* rangequery clause gets multiplied into s1 and forgotten. Anything that
@@ -939,6 +975,261 @@ clause_selectivity(PlannerInfo *root,
return s1;
}
+
+/*
+ * estimate selectivity of clauses using multivariate statistic
+ *
+ * Perform estimation of the clauses using a MCV list.
+ *
+ * This assumes all the clauses are compatible with the selected statistics
+ * (e.g. only reference columns covered by the statistics, use supported
+ * operator, etc.).
+ *
+ * TODO: We may support some additional conditions, most importantly those
+ * matching multiple columns (e.g. "a = b" or "a < b").
+ *
+ * TODO: Clamp the selectivity by min of the per-clause selectivities (i.e. the
+ * selectivity of the most restrictive clause), because that's the maximum
+ * we can ever get from ANDed list of clauses. This may probably prevent
+ * issues with hitting too many buckets and low precision histograms.
+ *
+ * TODO: We may remember the lowest frequency in the MCV list, and then later
+ * use it as a upper boundary for the selectivity (had there been a more
+ * frequent item, it'd be in the MCV list). This might improve cases with
+ * low-detail histograms.
+ *
+ * TODO: We may also derive some additional boundaries for the selectivity from
+ * the MCV list, because
+ *
+ * (a) if we have a "full equality condition" (one equality condition on
+ * each column of the statistic) and we found a match in the MCV list,
+ * then this is the final selectivity (and pretty accurate),
+ *
+ * (b) if we have a "full equality condition" and we haven't found a match
+ * in the MCV list, then the selectivity is below the lowest frequency
+ * found in the MCV list,
+ *
+ * TODO: When applying the clauses to the histogram/MCV list, we can do that
+ * from the most selective clauses first, because that'll eliminate the
+ * buckets/items sooner (so we'll be able to skip them without inspection,
+ * which is more expensive). But this requires really knowing the per-clause
+ * selectivities in advance, and that's not what we do now.
+ */
+static Selectivity
+clauselist_mv_selectivity(PlannerInfo *root, List *clauses, MVStatisticInfo *mvstats)
+{
+ bool fullmatch = false;
+ Selectivity s1 = 0.0,
+ s2 = 0.0;
+
+ /*
+ * Lowest frequency in the MCV list (may be used as an upper bound for
+ * full equality conditions that did not match any MCV item).
+ */
+ Selectivity mcv_low = 0.0;
+
+ /*
+ * TODO: Evaluate simple 1D selectivities, use the smallest one as an
+ * upper bound, product as lower bound, and sort the clauses in ascending
+ * order by selectivity (to optimize the MCV/histogram evaluation).
+ */
+
+ /* Evaluate the MCV first. */
+ s1 = clauselist_mv_selectivity_mcvlist(root, clauses, mvstats,
+ &fullmatch, &mcv_low);
+
+ /*
+ * If we got a full equality match on the MCV list, we're done (and the
+ * estimate is pretty good).
+ */
+ if (fullmatch && (s1 > 0.0))
+ return s1;
+
+ /*
+ * TODO if (fullmatch) without matching MCV item, use the mcv_low
+ * selectivity as upper bound
+ */
+
+ s2 = clauselist_mv_selectivity_histogram(root, clauses, mvstats);
+
+ /* TODO clamp to <= 1.0 (or more strictly, when possible) */
+ return s1 + s2;
+}
+
+static MVNDistinctItem *
+find_widest_ndistinct_item(MVNDistinct ndistinct, Bitmapset *attnums,
+ int16 *attmap)
+{
+ int i;
+ MVNDistinctItem *widest = NULL;
+
+ /* number of attnums in clauses */
+ int nattnums = bms_num_members(attnums);
+
+ /* with less than two attributes, we can bail out right away */
+ if (nattnums < 2)
+ return NULL;
+
+ /*
+ * Iterate over the MVNDistinctItem items and find the widest one from
+ * those fully-matched by clasuse.
+ */
+ for (i = 0; i < ndistinct->nitems; i++)
+ {
+ int j;
+ bool full_match = true;
+ MVNDistinctItem *item = &ndistinct->items[i];
+
+ /*
+ * Skip items referencing more attributes than available clauses,
+ * as those can't be fully matched.
+ */
+ if (item->nattrs > nattnums)
+ continue;
+
+ /* We can skip items with fewer attributes than the best one. */
+ if (widest && (widest->nattrs >= item->nattrs))
+ continue;
+
+ /*
+ * Check that the item actually is fully covered by clauses. We
+ * have to translate all attribute numbers.
+ */
+ for (j = 0; j < item->nattrs; j++)
+ {
+ int attnum = attmap[item->attrs[j]];
+
+ if (! bms_is_member(attnum, attnums))
+ {
+ full_match = false;
+ break;
+ }
+ }
+
+ /*
+ * If the item is not fully matched by clauses, we can't use
+ * it for the estimation.
+ */
+ if (! full_match)
+ continue;
+
+ /*
+ * We have a fully-matched item, and we already know it has to
+ * be wider than the current one (otherwise we'd skip it before
+ * inspecting it at the very beginning).
+ */
+ widest = item;
+ }
+
+ return widest;
+}
+
+static bool
+attnum_in_ndistinct_item(MVNDistinctItem *item, int attnum, int16 *attmap)
+{
+ int j;
+
+ for (j = 0; j < item->nattrs; j++)
+ {
+ if (attnum == attmap[item->attrs[j]])
+ return true;
+ }
+
+ return false;
+}
+
+static Selectivity
+clauselist_mv_selectivity_ndist(PlannerInfo *root, Index relid,
+ List *clauses, MVStatisticInfo *mvstats,
+ Index varRelid, JoinType jointype,
+ SpecialJoinInfo *sjinfo)
+{
+ ListCell *lc;
+ Selectivity s1 = 1.0;
+ MVNDistinct ndistinct;
+ MVNDistinctItem *item;
+ Bitmapset *attnums;
+ List *clauses_filtered = NIL;
+
+ /* we should only get here if the statistics includes ndistinct */
+ Assert(mvstats->ndist_enabled && mvstats->ndist_built);
+
+ /* load the ndistinct items stored in the statistics */
+ ndistinct = load_mv_ndistinct(mvstats->mvoid);
+
+ /* collect attnums in the clauses */
+ attnums = collect_mv_attnums(clauses, relid, STATS_TYPE_NDIST);
+
+ Assert(bms_num_members(attnums) >= 2);
+
+ /*
+ * Search for the widest ndistinct item (covering the most clauses), and
+ * then use it to estimate the number of entries.
+ */
+ item = find_widest_ndistinct_item(ndistinct, attnums,
+ mvstats->stakeys->values);
+
+ if (item)
+ {
+ /*
+ * We have an applicable item, so identify all covered clauses, and
+ * remove them from the list of clauses.
+ */
+ foreach(lc, clauses)
+ {
+ Bitmapset *attnums_clause = NULL;
+ Node *clause = (Node *) lfirst(lc);
+
+ /*
+ * XXX We need the attnum referenced by the clause, and this is the
+ * easiest way to get it (but maybe not the best one). At this point
+ * we should only see equality clauses, so just error out if we
+ * stumble upon something else.
+ */
+ if (! clause_is_mv_compatible(clause, relid, &attnums_clause,
+ STATS_TYPE_NDIST))
+ elog(ERROR, "clause not compatible with ndistinct stats");
+
+ /*
+ * We also expect only simple equality clauses, with a single Var.
+ *
+ * XXX This checks the number of attnums, not the number of Vars,
+ * but clause_is_mv_compatible only accepts (Var=Const) clauses.
+ */
+ Assert(bms_num_members(attnums_clause) == 1);
+
+ /*
+ * If the clause matches the selected ndistinct item, add it to
+ * the list of ndistinct clauses.
+ */
+ if (!attnum_in_ndistinct_item(item,
+ bms_singleton_member(attnums_clause),
+ mvstats->stakeys->values))
+ clauses_filtered = lappend(clauses_filtered, clause);
+ }
+
+ /* Compute selectivity using the ndistinct item. */
+ s1 *= (1.0 / item->ndistinct);
+
+ /*
+ * Throw away the clauses matched by the ndistinct, so that we don't
+ * estimate them twice.
+ */
+ clauses = clauses_filtered;
+ }
+
+ /* And now simply multiply with selectivities of the remaining clauses. */
+ foreach (lc, clauses)
+ {
+ Node *clause = (Node *) lfirst(lc);
+
+ s1 *= clause_selectivity(root, clause, varRelid, jointype, sjinfo);
+ }
+
+ return s1;
+}
+
+
/*
* When applying functional dependencies, we start with the strongest ones
* strongest dependencies. That is, we select the dependency that:
@@ -1147,85 +1438,6 @@ clauselist_mv_selectivity_deps(PlannerInfo *root, Index relid,
return s1;
}
-/*
- * estimate selectivity of clauses using multivariate statistic
- *
- * Perform estimation of the clauses using a MCV list.
- *
- * This assumes all the clauses are compatible with the selected statistics
- * (e.g. only reference columns covered by the statistics, use supported
- * operator, etc.).
- *
- * TODO: We may support some additional conditions, most importantly those
- * matching multiple columns (e.g. "a = b" or "a < b").
- *
- * TODO: Clamp the selectivity by min of the per-clause selectivities (i.e. the
- * selectivity of the most restrictive clause), because that's the maximum
- * we can ever get from ANDed list of clauses. This may probably prevent
- * issues with hitting too many buckets and low precision histograms.
- *
- * TODO: We may remember the lowest frequency in the MCV list, and then later
- * use it as a upper boundary for the selectivity (had there been a more
- * frequent item, it'd be in the MCV list). This might improve cases with
- * low-detail histograms.
- *
- * TODO: We may also derive some additional boundaries for the selectivity from
- * the MCV list, because
- *
- * (a) if we have a "full equality condition" (one equality condition on
- * each column of the statistic) and we found a match in the MCV list,
- * then this is the final selectivity (and pretty accurate),
- *
- * (b) if we have a "full equality condition" and we haven't found a match
- * in the MCV list, then the selectivity is below the lowest frequency
- * found in the MCV list,
- *
- * TODO: When applying the clauses to the histogram/MCV list, we can do that
- * from the most selective clauses first, because that'll eliminate the
- * buckets/items sooner (so we'll be able to skip them without inspection,
- * which is more expensive). But this requires really knowing the per-clause
- * selectivities in advance, and that's not what we do now.
- */
-static Selectivity
-clauselist_mv_selectivity(PlannerInfo *root, List *clauses, MVStatisticInfo *mvstats)
-{
- bool fullmatch = false;
- Selectivity s1 = 0.0,
- s2 = 0.0;
-
- /*
- * Lowest frequency in the MCV list (may be used as an upper bound for
- * full equality conditions that did not match any MCV item).
- */
- Selectivity mcv_low = 0.0;
-
- /*
- * TODO: Evaluate simple 1D selectivities, use the smallest one as an
- * upper bound, product as lower bound, and sort the clauses in ascending
- * order by selectivity (to optimize the MCV/histogram evaluation).
- */
-
- /* Evaluate the MCV first. */
- s1 = clauselist_mv_selectivity_mcvlist(root, clauses, mvstats,
- &fullmatch, &mcv_low);
-
- /*
- * If we got a full equality match on the MCV list, we're done (and the
- * estimate is pretty good).
- */
- if (fullmatch && (s1 > 0.0))
- return s1;
-
- /*
- * TODO if (fullmatch) without matching MCV item, use the mcv_low
- * selectivity as upper bound
- */
-
- s2 = clauselist_mv_selectivity_histogram(root, clauses, mvstats);
-
- /* TODO clamp to <= 1.0 (or more strictly, when possible) */
- return s1 + s2;
-}
/*
* Collect attributes from mv-compatible clauses.
@@ -1409,7 +1621,8 @@ choose_mv_statistics(List *stats, Bitmapset *attnums, int types)
int numattrs = info->stakeys->dim1;
/* skip statistics not matching any of the requested types */
- if (! ((info->deps_built && (STATS_TYPE_FDEPS & types)) ||
+ if (! ((info->ndist_built && (STATS_TYPE_NDIST & types)) ||
+ (info->deps_built && (STATS_TYPE_FDEPS & types)) ||
(info->mcv_built && (STATS_TYPE_MCV & types)) ||
(info->hist_built && (STATS_TYPE_HIST & types))))
continue;
@@ -1703,6 +1916,9 @@ clause_is_mv_compatible(Node *clause, Index relid, Bitmapset **attnums, int type
static bool
stats_type_matches(MVStatisticInfo *stat, int type)
{
+ if ((type & STATS_TYPE_NDIST) && stat->ndist_built)
+ return true;
+
if ((type & STATS_TYPE_FDEPS) && stat->deps_built)
return true;
--
2.5.5
[binary/octet-stream] 0008-WIP-allow-using-multiple-statistics-in-clauselis-v23.patch (9.3K, ../../1fa982dc-ffb4-eef7-eb55-d8cf7d84f1c4@2ndquadrant.com/9-0008-WIP-allow-using-multiple-statistics-in-clauselis-v23.patch)
download | inline diff:
From 4899debf66176fba3199f28d2147a8d113c0cbbc Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Fri, 28 Oct 2016 17:03:09 +0200
Subject: [PATCH 8/9] WIP: allow using multiple statistics in
clauselist_selectivity
---
src/backend/optimizer/path/clausesel.c | 31 +++++++-----
src/test/regress/expected/mv_statistics.out | 78 +++++++++++++++++++++++++++++
src/test/regress/parallel_schedule | 2 +-
src/test/regress/serial_schedule | 1 +
src/test/regress/sql/mv_statistics.sql | 60 ++++++++++++++++++++++
5 files changed, 159 insertions(+), 13 deletions(-)
create mode 100644 src/test/regress/expected/mv_statistics.out
create mode 100644 src/test/regress/sql/mv_statistics.sql
diff --git a/src/backend/optimizer/path/clausesel.c b/src/backend/optimizer/path/clausesel.c
index da5c340..c449c96 100644
--- a/src/backend/optimizer/path/clausesel.c
+++ b/src/backend/optimizer/path/clausesel.c
@@ -228,15 +228,16 @@ clauselist_selectivity(PlannerInfo *root,
(count_mv_attnums(clauses, relid,
STATS_TYPE_MCV | STATS_TYPE_HIST) >= 2))
{
+ Bitmapset *mvattnums;
+ MVStatisticInfo *mvstat;
+
/* collect attributes from the compatible conditions */
- Bitmapset *mvattnums = collect_mv_attnums(clauses, relid,
- STATS_TYPE_MCV | STATS_TYPE_HIST);
+ mvattnums = collect_mv_attnums(clauses, relid,
+ STATS_TYPE_MCV | STATS_TYPE_HIST);
/* and search for the statistic covering the most attributes */
- MVStatisticInfo *mvstat = choose_mv_statistics(stats, mvattnums,
- STATS_TYPE_MCV | STATS_TYPE_HIST);
-
- if (mvstat != NULL) /* we have a matching stats */
+ while ((mvstat = choose_mv_statistics(stats, mvattnums,
+ STATS_TYPE_MCV | STATS_TYPE_HIST)))
{
/* clauses compatible with multi-variate stats */
List *mvclauses = NIL;
@@ -250,6 +251,10 @@ clauselist_selectivity(PlannerInfo *root,
/* compute the multivariate stats */
s1 *= clauselist_mv_selectivity(root, mvclauses, mvstat);
+
+ /* update the bitmap if attnums using the remaining clauses) */
+ mvattnums = collect_mv_attnums(clauses, relid,
+ STATS_TYPE_MCV | STATS_TYPE_HIST);
}
}
@@ -264,9 +269,7 @@ clauselist_selectivity(PlannerInfo *root,
mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_FDEPS);
/* and search for the statistic covering the most attributes */
- mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_FDEPS);
-
- if (mvstat != NULL) /* we have a matching stats */
+ while ((mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_FDEPS)))
{
/* clauses compatible with multi-variate stats */
List *mvclauses = NIL;
@@ -284,6 +287,9 @@ clauselist_selectivity(PlannerInfo *root,
/* compute the multivariate stats (dependencies) */
s1 *= clauselist_mv_selectivity_deps(root, relid, mvclauses, mvstat,
varRelid, jointype, sjinfo);
+
+ /* update the bitmap if attnums using the remaining clauses) */
+ mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_FDEPS);
}
}
@@ -298,9 +304,7 @@ clauselist_selectivity(PlannerInfo *root,
mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_NDIST);
/* and search for the statistic covering the most attributes */
- mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_NDIST);
-
- if (mvstat != NULL) /* we have a matching stats */
+ while ((mvstat = choose_mv_statistics(stats, mvattnums, STATS_TYPE_NDIST)))
{
/* clauses compatible with multi-variate stats */
List *mvclauses = NIL;
@@ -315,6 +319,9 @@ clauselist_selectivity(PlannerInfo *root,
/* compute the multivariate stats (dependencies) */
s1 *= clauselist_mv_selectivity_ndist(root, relid, mvclauses, mvstat,
varRelid, jointype, sjinfo);
+
+ /* collect attributes from the compatible conditions */
+ mvattnums = collect_mv_attnums(clauses, relid, STATS_TYPE_NDIST);
}
}
diff --git a/src/test/regress/expected/mv_statistics.out b/src/test/regress/expected/mv_statistics.out
new file mode 100644
index 0000000..7eb6f2e
--- /dev/null
+++ b/src/test/regress/expected/mv_statistics.out
@@ -0,0 +1,78 @@
+-- data type passed by value
+CREATE TABLE multi_stats (
+ a INT,
+ b INT,
+ c INT,
+ d INT,
+ e INT,
+ f INT,
+ g INT,
+ h INT
+);
+-- MCV list on (a,b)
+CREATE STATISTICS m1 WITH (mcv) ON (a, b) FROM multi_stats;
+-- histogram on (c,d)
+CREATE STATISTICS m2 WITH (histogram) ON (c, d) FROM multi_stats;
+-- functional dependencies on (e,f)
+CREATE STATISTICS m3 WITH (dependencies) ON (e, f) FROM multi_stats;
+-- ndistinct coefficients on (g,h)
+CREATE STATISTICS m4 WITH (ndistinct) ON (g, h) FROM multi_stats;
+-- perfectly correlated groups
+INSERT INTO multi_stats
+SELECT
+ i, i/2, -- MCV
+ i, i + j, -- histogram
+ k, k/2, -- dependencies
+ l/5, l/10 -- ndistinct
+FROM (
+ SELECT
+ mod(x, 13) AS i,
+ mod(x, 17) AS j,
+ mod(x, 11) AS k,
+ mod(x, 51) AS l
+ FROM generate_series(1,30000) AS s(x)
+) foo;
+ANALYZE multi_stats;
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (c >= 3) AND (d <= 10);
+ QUERY PLAN
+----------------------------------------------------------------
+ Seq Scan on multi_stats (cost=0.00..821.00 rows=413 width=32)
+ Filter: ((c >= 3) AND (d <= 10) AND (a = 8) AND (b = 4))
+(2 rows)
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (e = 10) AND (f = 5);
+ QUERY PLAN
+----------------------------------------------------------------
+ Seq Scan on multi_stats (cost=0.00..821.00 rows=210 width=32)
+ Filter: ((a = 8) AND (b = 4) AND (e = 10) AND (f = 5))
+(2 rows)
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (e = 10) AND (f = 5);
+ QUERY PLAN
+----------------------------------------------------------------
+ Seq Scan on multi_stats (cost=0.00..821.00 rows=210 width=32)
+ Filter: ((a = 8) AND (b = 4) AND (e = 10) AND (f = 5))
+(2 rows)
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (g = 2) AND (h = 1);
+ QUERY PLAN
+----------------------------------------------------------------
+ Seq Scan on multi_stats (cost=0.00..821.00 rows=210 width=32)
+ Filter: ((a = 8) AND (b = 4) AND (g = 2) AND (h = 1))
+(2 rows)
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND
+ (c >= 3) AND (d <= 10) AND
+ (e = 10) AND (f = 5);
+ QUERY PLAN
+-------------------------------------------------------------------------------------
+ Seq Scan on multi_stats (cost=0.00..971.00 rows=37 width=32)
+ Filter: ((c >= 3) AND (d <= 10) AND (a = 8) AND (b = 4) AND (e = 10) AND (f = 5))
+(2 rows)
+
+DROP TABLE multi_stats;
diff --git a/src/test/regress/parallel_schedule b/src/test/regress/parallel_schedule
index 36dd618..bd4a294 100644
--- a/src/test/regress/parallel_schedule
+++ b/src/test/regress/parallel_schedule
@@ -118,4 +118,4 @@ test: event_trigger
test: stats
# run tests of multivariate stats
-test: mv_ndistinct mv_dependencies mv_mcv mv_histogram
+test: mv_ndistinct mv_dependencies mv_mcv mv_histogram mv_statistics
diff --git a/src/test/regress/serial_schedule b/src/test/regress/serial_schedule
index 34f5467..54cc854 100644
--- a/src/test/regress/serial_schedule
+++ b/src/test/regress/serial_schedule
@@ -175,3 +175,4 @@ test: mv_ndistinct
test: mv_dependencies
test: mv_mcv
test: mv_histogram
+test: mv_statistics
diff --git a/src/test/regress/sql/mv_statistics.sql b/src/test/regress/sql/mv_statistics.sql
new file mode 100644
index 0000000..cd12ad0
--- /dev/null
+++ b/src/test/regress/sql/mv_statistics.sql
@@ -0,0 +1,60 @@
+-- data type passed by value
+CREATE TABLE multi_stats (
+ a INT,
+ b INT,
+ c INT,
+ d INT,
+ e INT,
+ f INT,
+ g INT,
+ h INT
+);
+
+-- MCV list on (a,b)
+CREATE STATISTICS m1 WITH (mcv) ON (a, b) FROM multi_stats;
+
+-- histogram on (c,d)
+CREATE STATISTICS m2 WITH (histogram) ON (c, d) FROM multi_stats;
+
+-- functional dependencies on (e,f)
+CREATE STATISTICS m3 WITH (dependencies) ON (e, f) FROM multi_stats;
+
+-- ndistinct coefficients on (g,h)
+CREATE STATISTICS m4 WITH (ndistinct) ON (g, h) FROM multi_stats;
+
+-- perfectly correlated groups
+INSERT INTO multi_stats
+SELECT
+ i, i/2, -- MCV
+ i, i + j, -- histogram
+ k, k/2, -- dependencies
+ l/5, l/10 -- ndistinct
+FROM (
+ SELECT
+ mod(x, 13) AS i,
+ mod(x, 17) AS j,
+ mod(x, 11) AS k,
+ mod(x, 51) AS l
+ FROM generate_series(1,30000) AS s(x)
+) foo;
+
+ANALYZE multi_stats;
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (c >= 3) AND (d <= 10);
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (e = 10) AND (f = 5);
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (e = 10) AND (f = 5);
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND (g = 2) AND (h = 1);
+
+EXPLAIN SELECT * FROM multi_stats
+ WHERE (a = 8) AND (b = 4) AND
+ (c >= 3) AND (d <= 10) AND
+ (e = 10) AND (f = 5);
+
+DROP TABLE multi_stats;
--
2.5.5
[binary/octet-stream] 0009-WIP-psql-tab-completion-basics-v23.patch (3.0K, ../../1fa982dc-ffb4-eef7-eb55-d8cf7d84f1c4@2ndquadrant.com/10-0009-WIP-psql-tab-completion-basics-v23.patch)
download | inline diff:
From 9b9ce038dca938cc78c935448fb331d83a6d2f6c Mon Sep 17 00:00:00 2001
From: Tomas Vondra <tomas@pgaddict.com>
Date: Fri, 28 Oct 2016 20:46:37 +0200
Subject: [PATCH 9/9] WIP: psql tab-completion basics
---
src/bin/psql/tab-complete.c | 30 ++++++++++++++++++++++++++++--
1 file changed, 28 insertions(+), 2 deletions(-)
diff --git a/src/bin/psql/tab-complete.c b/src/bin/psql/tab-complete.c
index d6fffcf..8e804d1 100644
--- a/src/bin/psql/tab-complete.c
+++ b/src/bin/psql/tab-complete.c
@@ -448,6 +448,21 @@ static const SchemaQuery Query_for_list_of_foreign_tables = {
NULL
};
+static const SchemaQuery Query_for_list_of_statistics = {
+ /* catname */
+ "pg_catalog.pg_mv_statistic s",
+ /* selcondition */
+ NULL,
+ /* viscondition */
+ NULL,
+ /* namespace */
+ "s.stanamespace",
+ /* result */
+ "pg_catalog.quote_ident(s.staname)",
+ /* qualresult */
+ NULL
+};
+
static const SchemaQuery Query_for_list_of_tables = {
/* catname */
"pg_catalog.pg_class c",
@@ -966,6 +981,7 @@ static const pgsql_thing_t words_after_create[] = {
{"SCHEMA", Query_for_list_of_schemas},
{"SEQUENCE", NULL, &Query_for_list_of_sequences},
{"SERVER", Query_for_list_of_servers},
+ {"STATISTICS", NULL, &Query_for_list_of_statistics},
{"SUBSCRIPTION", NULL, NULL},
{"TABLE", NULL, &Query_for_list_of_tables},
{"TABLESPACE", Query_for_list_of_tablespaces},
@@ -1410,8 +1426,8 @@ psql_completion(const char *text, int start, int end)
"EVENT TRIGGER", "EXTENSION", "FOREIGN DATA WRAPPER", "FOREIGN TABLE", "FUNCTION",
"GROUP", "INDEX", "LANGUAGE", "LARGE OBJECT", "MATERIALIZED VIEW", "OPERATOR",
"POLICY", "PUBLICATION", "ROLE", "RULE", "SCHEMA", "SERVER", "SEQUENCE",
- "SUBSCRIPTION", "SYSTEM", "TABLE", "TABLESPACE", "TEXT SEARCH", "TRIGGER", "TYPE",
- "USER", "USER MAPPING FOR", "VIEW", NULL};
+ "STATISTICS", "SUBSCRIPTION", "SYSTEM", "TABLE", "TABLESPACE", "TEXT SEARCH",
+ "TRIGGER", "TYPE", "USER", "USER MAPPING FOR", "VIEW", NULL};
COMPLETE_WITH_LIST(list_ALTER);
}
@@ -1713,6 +1729,10 @@ psql_completion(const char *text, int start, int end)
else if (Matches5("ALTER", "RULE", MatchAny, "ON", MatchAny))
COMPLETE_WITH_CONST("RENAME TO");
+ /* ALTER STATISTICS <name> */
+ else if (Matches3("ALTER", "STATISTICS", MatchAny))
+ COMPLETE_WITH_LIST3("OWNER TO", "RENAME TO", "SET SCHEMA");
+
/* ALTER TRIGGER <name>, add ON */
else if (Matches3("ALTER", "TRIGGER", MatchAny))
COMPLETE_WITH_CONST("ON");
@@ -2292,6 +2312,12 @@ psql_completion(const char *text, int start, int end)
else if (Matches3("CREATE", "SERVER", MatchAny))
COMPLETE_WITH_LIST3("TYPE", "VERSION", "FOREIGN DATA WRAPPER");
+/* CREATE STATISTICS <name> */
+ else if (Matches3("CREATE", "STATISTICS", MatchAny))
+ COMPLETE_WITH_LIST2("WITH", "ON");
+ else if (Matches4("CREATE", "STATISTICS", MatchAny, "ON|WITH"))
+ COMPLETE_WITH_CONST("(");
+
/* CREATE TABLE --- is allowed inside CREATE SCHEMA, so use TailMatches */
/* Complete "CREATE TEMP/TEMPORARY" with the possible temp objects */
else if (TailMatches2("CREATE", "TEMP|TEMPORARY"))
--
2.5.5
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-30 19:33 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Dilip Kumar <dilipbalaut@gmail.com>
0 siblings, 0 replies; 70+ messages in thread
From: Tomas Vondra @ 2017-01-30 19:33 UTC (permalink / raw)
To: Dilip Kumar <dilipbalaut@gmail.com>; +Cc: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Álvaro Herrera <alvherre@2ndquadrant.com>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 01/26/2017 10:43 AM, Dilip Kumar wrote:
>
> histograms
> --------------
> + if (matches[i] == MVSTATS_MATCH_FULL)
> + s += mvhist->buckets[i]->ntuples;
> + else if (matches[i] == MVSTATS_MATCH_PARTIAL)
> + s += 0.5 * mvhist->buckets[i]->ntuples;
>
> Isn't it will be better that take some percentage of the bucket based
> on the number of distinct element for partial matching buckets.
>
I don't think so, for the same reason why ineq_histogram_selectivity()
in selfuncs.c uses
binfrac = 0.5;
for partial bucket matches - it provides minimum average error. Even if
we knew the number of distinct items in the bucket, we have no idea what
the distribution within the bucket looks like. Maybe 99% of the bucket
are covered by a single distinct value, maybe all the items are squashed
on one side of the bucket, etc.
Moreover we don't really know the number of distinct values in the
bucket - we only know the number of distinct items in the sample, and
only while building the histogram. I don't think it makes much sense to
estimate the number of distinct items in a bucket, because the buckets
contain only very few rows so the estimates would be wildly inaccurate.
>
> +static int
> +update_match_bitmap_histogram(PlannerInfo *root, List *clauses,
> + int2vector *stakeys,
> + MVSerializedHistogram mvhist,
> + int nmatches, char *matches,
> + bool is_or)
> +{
> + int i;
>
> For each clause we are processing all the buckets, can't we use some
> data structure which can make multi-dimensions information searching
> faster.
>
No, we're not processing all buckets for each clause. We're' only
processing buckets that were not "ruled out" by preceding clauses.
That's the whole point of the bitmap.
For example for condition (a=1) AND (b=2), the code will first evaluate
(a=1) on all buckets, and then (b=2) but only on buckets where (a=1) was
evaluated as true. Similarly for OR clauses.
>
> Something like HTree, RTree, Maybe storing histogram in these formats
> will be difficult?
>
Maybe, but I don't want to do that in the first version. I'm not opposed
to doing that in the future, if we find out the v1 histograms are not
efficient (I don't think we will, based on tests I did while working on
the patch). Support for other histogram implementations is pretty much
why there is 'type' field in the struct.
For now I think we should stick with the simple implementation.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-30 19:36 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Ideriha, Takeshi <ideriha.takeshi@jp.fujitsu.com>
1 sibling, 0 replies; 70+ messages in thread
From: Tomas Vondra @ 2017-01-30 19:36 UTC (permalink / raw)
To: Ideriha, Takeshi <ideriha.takeshi@jp.fujitsu.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Michael Paquier <michael.paquier@gmail.com>; Alvaro Herrera <alvherre@2ndquadrant.com>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Hello,
On 01/26/2017 10:03 AM, Ideriha, Takeshi wrote:
>
> Though I pointed out these typoes and so on,
> I believe these feedback are less priority compared to the source code itself.
>
> So please work on my feedback if you have time.
>
I think getting the comments (and docs in general) right is just as
important as the code. So thank you for your review!
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-30 19:40 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Alvaro Herrera <alvherre@2ndquadrant.com>
0 siblings, 0 replies; 70+ messages in thread
From: Tomas Vondra @ 2017-01-30 19:40 UTC (permalink / raw)
To: Alvaro Herrera <alvherre@2ndquadrant.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 01/30/2017 05:55 PM, Alvaro Herrera wrote:
> Minor nitpicks:
>
> Let me suggest to use get_attnum() in CreateStatistics instead of
> SearchSysCacheAttName for each column. Also, we use type AttrNumber for
> attribute numbers rather than int16. Finally in the same function you
> have an erroneous ERRCODE_UNDEFINED_COLUMN which should be
> ERRCODE_DUPLICATE_COLUMN in the loop that searches for duplicates.
>
> May I suggest that compare_int16 be named attnum_cmp (just to be
> consistent with other qsort comparators) and look like
> return *((const AttrNumber *) a) - *((const AttrNumber *) b);
> instead of memcmp?
>
Yes, I think this is pretty much what Kyotaro-san pointed out in his
review. I'll go through the patch and make sure the correct data types
are used.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-30 20:00 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Alvaro Herrera <alvherre@2ndquadrant.com>
0 siblings, 1 reply; 70+ messages in thread
From: Tomas Vondra @ 2017-01-30 20:00 UTC (permalink / raw)
To: Alvaro Herrera <alvherre@2ndquadrant.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 01/30/2017 05:12 PM, Alvaro Herrera wrote:
>
> Hmm. So we have a catalog pg_mv_statistics which stores two things:
> 1. the configuration regarding mvstats that have been requested by user
> via CREATE/ALTER STATISTICS
> 2. the actual values captured from the above, via ANALYZE
>
> I think this conflates two things that really are separate, given their
> different timings and usage patterns. This decision is causing the
> catalog to have columns enabled/built flags for each set of stats
> requested, which looks a bit odd. In particular, the fact that you have
> to heap_update the catalog in order to add more stuff as it's built
> looks inconvenient.
>
> Have you thought about having the "requested" bits be separate from the
> actual computed values? Something like
>
> pg_mv_statistics
> starelid
> staname
> stanamespace
> staowner -- all the above as currently
> staenabled array of "char" {d,f,s}
> stakeys
> // no CATALOG_VARLEN here
>
> where each char in the staenabled array has a #define and indicates one
> type, "ndistinct", "functional dep", "selectivity" etc.
>
> The actual values computed by ANALYZE would live in a catalog like:
>
> pg_mv_statistics_values
> stvstaid -- OID of the corresponding pg_mv_statistics row. Needed?
Definitely needed. How else would you know which MCV list and histogram
belong together? This works just like in pg_statistic - when both MCV
and histograms are enabled for the statistic, we first build MCV list,
then histogram on remaining rows. So we need to pair them.
> stvrelid -- same as starelid
> stvkeys -- same as stakeys
> #ifdef CATALOG_VARLEN
> stvkind 'd' or 'f' or 's', etc
> stvvalue the bytea blob
> #endif
>
> I think that would be simpler, both conceptually and in terms of code.
I think the main issue here is that it throws away the special data
types (pg_histogram, pg_mcv, pg_ndistinct, pg_dependencies), which I
think is a neat idea and would like to keep it. This would throw that
away, making everything bytea again. I don't like that.
>
> The other angle to consider is planner-side: how does the planner gets
> to the values? I think as far as the planner goes, the first catalog
> doesn't matter at all, because a statistics type that has been enabled
> but not computed is not interesting at all; planner only cares about the
> values in the second catalog (this is why I added stvkeys). Currently
> you're just caching a single pg_mv_statistics row in get_relation_info
> (and only if any of the "built" flags is set), which is simple. With my
> proposed change, you'd need to keep multiple pg_mv_statistics_values
> rows.
>
> But maybe you already tried something like what I propose and there's a
> reason not to do it?
>
Honestly, I don't see how this improves the situation. We still need to
cache data for exactly one catalog, so how is that simpler?
The way I see it, it actually makes things more complicated, because now
we have two catalogs to manage instead of one (e.g. when doing DROP
STATISTICS, or after ALTER TABLE ... DROP COLUMN).
The 'built' flags may be easily replaced with a check if the bytea-like
columns are NULL, and the 'enabled' columns may be replaced by the array
of char, just like you proposed.
That'd give us a single catalog looking like this:
pg_mv_statistics
starelid
staname
stanamespace
staowner -- all the above as currently
staenabled array of "char" {d,f,s}
stakeys
stadeps (dependencies)
standist (ndistinct coefficients)
stamcv (MCV list)
stahist (histogram)
Which is probably a better / simpler structure than the current one.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-30 20:37 Alvaro Herrera <alvherre@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
0 siblings, 1 reply; 70+ messages in thread
From: Alvaro Herrera @ 2017-01-30 20:37 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
Tomas Vondra wrote:
> The 'built' flags may be easily replaced with a check if the bytea-like
> columns are NULL, and the 'enabled' columns may be replaced by the array of
> char, just like you proposed.
>
> That'd give us a single catalog looking like this:
>
> pg_mv_statistics
> starelid
> staname
> stanamespace
> staowner -- all the above as currently
> staenabled array of "char" {d,f,s}
> stakeys
> stadeps (dependencies)
> standist (ndistinct coefficients)
> stamcv (MCV list)
> stahist (histogram)
>
> Which is probably a better / simpler structure than the current one.
Looks good to me. I don't think we need to keep the names very short --
I would propose "standistinct", "stahistogram", "stadependencies".
Thanks,
--
Álvaro Herrera https://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-30 21:57 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Alvaro Herrera <alvherre@2ndquadrant.com>
0 siblings, 2 replies; 70+ messages in thread
From: Tomas Vondra @ 2017-01-30 21:57 UTC (permalink / raw)
To: Alvaro Herrera <alvherre@2ndquadrant.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 01/30/2017 09:37 PM, Alvaro Herrera wrote:
> Tomas Vondra wrote:
>
>> The 'built' flags may be easily replaced with a check if the bytea-like
>> columns are NULL, and the 'enabled' columns may be replaced by the array of
>> char, just like you proposed.
>>
>> That'd give us a single catalog looking like this:
>>
>> pg_mv_statistics
>> starelid
>> staname
>> stanamespace
>> staowner -- all the above as currently
>> staenabled array of "char" {d,f,s}
>> stakeys
>> stadeps (dependencies)
>> standist (ndistinct coefficients)
>> stamcv (MCV list)
>> stahist (histogram)
>>
>> Which is probably a better / simpler structure than the current one.
>
> Looks good to me. I don't think we need to keep the names very short --
> I would propose "standistinct", "stahistogram", "stadependencies".
>
Yeah, I got annoyed by the short names too.
This however reminds me that perhaps pg_mv_statistic is not the best
name. I know others proposed pg_statistic_ext (and pg_stats_ext), and
while I wasn't a big fan initially, I think it's a better name. People
generally don't know what 'multivariate' means, while 'extended' is
better known (e.g. because Oracle uses it for similar stuff).
So I think I'll switch to that name too.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-31 04:21 Michael Paquier <michael.paquier@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
1 sibling, 0 replies; 70+ messages in thread
From: Michael Paquier @ 2017-01-31 04:21 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Alvaro Herrera <alvherre@2ndquadrant.com>; Dilip Kumar <dilipbalaut@gmail.com>; Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On Tue, Jan 31, 2017 at 6:57 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> This however reminds me that perhaps pg_mv_statistic is not the best name. I
> know others proposed pg_statistic_ext (and pg_stats_ext), and while I wasn't
> a big fan initially, I think it's a better name. People generally don't know
> what 'multivariate' means, while 'extended' is better known (e.g. because
> Oracle uses it for similar stuff).
>
> So I think I'll switch to that name too.
I have moved this patch to the next CF, with Álvaro as reviewer.
--
Michael
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-31 06:52 Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
1 sibling, 1 reply; 70+ messages in thread
From: Amit Langote @ 2017-01-31 06:52 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Alvaro Herrera <alvherre@2ndquadrant.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 2017/01/31 6:57, Tomas Vondra wrote:
> On 01/30/2017 09:37 PM, Alvaro Herrera wrote:
>> Looks good to me. I don't think we need to keep the names very short --
>> I would propose "standistinct", "stahistogram", "stadependencies".
>>
>
> Yeah, I got annoyed by the short names too.
>
> This however reminds me that perhaps pg_mv_statistic is not the best name.
> I know others proposed pg_statistic_ext (and pg_stats_ext), and while I
> wasn't a big fan initially, I think it's a better name. People generally
> don't know what 'multivariate' means, while 'extended' is better known
> (e.g. because Oracle uses it for similar stuff).
>
> So I think I'll switch to that name too.
+1 to pg_statistics_ext. Maybe, even pg_statistics_extended, however
being that verbose may not be warranted.
Thanks,
Amit
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-01-31 16:10 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>
0 siblings, 0 replies; 70+ messages in thread
From: Tomas Vondra @ 2017-01-31 16:10 UTC (permalink / raw)
To: Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>; Alvaro Herrera <alvherre@2ndquadrant.com>; +Cc: Dilip Kumar <dilipbalaut@gmail.com>; Dean Rasheed <dean.a.rasheed@gmail.com>; Heikki Linnakangas <hlinnaka@iki.fi>; Michael Paquier <michael.paquier@gmail.com>; Robert Haas <robertmhaas@gmail.com>; Tatsuo Ishii <ishii@postgresql.org>; David Steele <david@pgmasters.net>; Tom Lane <tgl@sss.pgh.pa.us>; Petr Jelinek <petr@2ndquadrant.com>; Jeff Janes <jeff.janes@gmail.com>; pgsql-hackers
On 01/31/2017 07:52 AM, Amit Langote wrote:
> On 2017/01/31 6:57, Tomas Vondra wrote:
>> On 01/30/2017 09:37 PM, Alvaro Herrera wrote:
>>> Looks good to me. I don't think we need to keep the names very short --
>>> I would propose "standistinct", "stahistogram", "stadependencies".
>>>
>>
>> Yeah, I got annoyed by the short names too.
>>
>> This however reminds me that perhaps pg_mv_statistic is not the best name.
>> I know others proposed pg_statistic_ext (and pg_stats_ext), and while I
>> wasn't a big fan initially, I think it's a better name. People generally
>> don't know what 'multivariate' means, while 'extended' is better known
>> (e.g. because Oracle uses it for similar stuff).
>>
>> So I think I'll switch to that name too.
>
> +1 to pg_statistics_ext. Maybe, even pg_statistics_extended, however
> being that verbose may not be warranted.
>
Yeah, I think pg_statistic_extended / pg_stats_extended seems fine.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-01 22:52 Alvaro Herrera <alvherre@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
2 siblings, 1 reply; 70+ messages in thread
From: Alvaro Herrera @ 2017-02-01 22:52 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: pgsql-hackers
Still looking at 0002.
pg_ndistinct_in disallows input, claiming that pg_node_tree does the
same thing. But pg_node_tree does it for security reasons: you could
crash the backend if you supplied a malicious value. I don't think that
applies to pg_ndistinct_in. Perhaps it will be useful to inject fake
stats at some point, so why not allow it? It shouldn't be complicated
(though it does require writing some additional code, so perhaps that's
one reason we don't want to allow input of these values).
The comment on top of pg_ndistinct_out is missing the "_out"; also it
talks about histograms, which is not what this is about.
In the same function, a trivial point you don't need to pstrdup() the
.data out of a stringinfo; it's already palloc'ed in the right context
-- just PG_RETURN_CSTRING(str.data) and forget about "ret". Saves you
one line.
Nearby, some auxiliary functions such as n_choose_k and num_combinations
are not documented. What it is that they do? I'd move these at the end
of the file, keeping the important entry points at the top of the file.
I see this patch has a estimate_ndistinct() which claims to be a re-
implementation of code already in analyze.c, but it is actually a lot
simpler than what analyze.c does. I've been wondering if it'd be a good
idea to use some of this code so that some routines are moved out of
analyze.c; good implementations of statistics-related functions would
live in src/backend/statistics/ where they can be used both by analyze.c
and your new mvstats stuff. (More generally I am beginning to wonder if
the new directory should be just src/backend/statistics.)
common.h does not belong in src/backend/utils/mvstats; IMO it should be
called src/include/utils/mvstat.h. Also, it must not include
postgres.h, and it probably doesn't need most of the #includes it has;
those are better put into whatever include it. It definitely needs a
guarding #ifdef MVSTATS_H around its whole content too. An include file
is not just a way to avoid #includes in other files; it is supposed to
be a minimally invasive way of exporting the structs and functions
implemented in some file into other files. So it must be kept minimal.
psql/tab-complete.c compares the wrong version number (9.6 instead of
10).
Is it important to have a cast from pg_ndistinct to bytea? I think
it's odd that outputting it as bytea yields something completely
different than as text. (The bytea is not human readable and cannot be
used for future input, so what is the point?)
In another subthread you seem to have surrendered to the opinion that
the new catalog should be called pg_statistics_ext, just in case in the
future we come up with additional things to put on it. However, given
its schema, with a "starelid / stakeys", is it sensible to think that
we're going to get anything other than something that involves multiple
variables? Maybe it should just be "pg_statistics_multivar" and if
something else comes along we create another catalog with an appropriate
schema. Heck, how does this catalog serve the purpose of cross-table
statistics in the first place, given that it has room to record a single
relid only? Are you thinking that in the future you'd change starelid
into an oidvector column?
The comment in gram.y about the CREATE STATISTICS is at odds with what
is actually allowed by the grammar.
I think the name of a statistics is only useful to DROP/ALTER it, right?
I wonder why it's useful that statistics belongs in a schema. Perhaps
it should be a global object? I suppose the name collisions would
become bothersome if you have many mvstats.
--
Álvaro Herrera https://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-02 08:59 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Alvaro Herrera <alvherre@2ndquadrant.com>
0 siblings, 2 replies; 70+ messages in thread
From: Tomas Vondra @ 2017-02-02 08:59 UTC (permalink / raw)
To: Alvaro Herrera <alvherre@2ndquadrant.com>; +Cc: pgsql-hackers
On 02/01/2017 11:52 PM, Alvaro Herrera wrote:
> Still looking at 0002.
>
> pg_ndistinct_in disallows input, claiming that pg_node_tree does the
> same thing. But pg_node_tree does it for security reasons: you could
> crash the backend if you supplied a malicious value. I don't think
> that applies to pg_ndistinct_in. Perhaps it will be useful to inject
> fake stats at some point, so why not allow it? It shouldn't be
> complicated (though it does require writing some additional code, so
> perhaps that's one reason we don't want to allow input of these
> values).
>
Yes, I haven't written the code, and I'm not sure it's a very practical
way to inject custom statistics. But if we decide to allow that in the
future, we can probably add the code.
There's a subtle difference between pg_node_tree and the data types for
statistics - pg_node_tree stores the value as a string (matching the
nodeToString output), so the _in function is fairly simple. Of course,
stringToNode() assumes safe input, which is why the input is disabled.
OTOH the statistics are stored in an optimized binary format, allowing
to use the value directly (without having to do expensive parsing etc).
I was thinking that the easiest way to add support for _in would be to
add a bunch of Nodes for the statistics, along with in/out functions,
but keeping the internal binary representation. But that'll be tricky to
do in a safe way - even if those nodes are coded in a very defensive
ways, I'd bet there'll be ways to inject unsafe nodes.
So I'm OK with not having the _in for now. If needed, it's possible to
construct the statistics as a bytea using a bit of C code. That's at
least obviously unsafe, as anything written in C, touching the memory.
> The comment on top of pg_ndistinct_out is missing the "_out"; also it
> talks about histograms, which is not what this is about.
>
OK, will fix.
> In the same function, a trivial point you don't need to pstrdup() the
> .data out of a stringinfo; it's already palloc'ed in the right context
> -- just PG_RETURN_CSTRING(str.data) and forget about "ret". Saves you
> one line.
>
Will fix too.
> Nearby, some auxiliary functions such as n_choose_k and
> num_combinations are not documented. What it is that they do? I'd
> move these at the end of the file, keeping the important entry points
> at the top of the file.
I'd say n-choose-k is pretty widely known term from combinatorics. The
comment would essentially say just 'this is n-choose-k' which seems
rather pointless. So as much as I dislike the self-documenting code,
this actually seems like a good case of that.
> I see this patch has a estimate_ndistinct() which claims to be a re-
> implementation of code already in analyze.c, but it is actually a lot
> simpler than what analyze.c does. I've been wondering if it'd be a good
> idea to use some of this code so that some routines are moved out of
> analyze.c; good implementations of statistics-related functions would
> live in src/backend/statistics/ where they can be used both by analyze.c
> and your new mvstats stuff. (More generally I am beginning to wonder if
> the new directory should be just src/backend/statistics.)
>
I'll look into that. I have to check if I ignored some assumptions or
corner cases the analyze.c deals with.
> common.h does not belong in src/backend/utils/mvstats; IMO it should be
> called src/include/utils/mvstat.h. Also, it must not include
> postgres.h, and it probably doesn't need most of the #includes it has;
> those are better put into whatever include it. It definitely needs a
> guarding #ifdef MVSTATS_H around its whole content too. An include file
> is not just a way to avoid #includes in other files; it is supposed to
> be a minimally invasive way of exporting the structs and functions
> implemented in some file into other files. So it must be kept minimal.
>
Will do.
> psql/tab-complete.c compares the wrong version number (9.6 instead of
> 10).
>
> Is it important to have a cast from pg_ndistinct to bytea? I think
> it's odd that outputting it as bytea yields something completely
> different than as text. (The bytea is not human readable and cannot be
> used for future input, so what is the point?)
>
Because it internally is a bytea, and it seems useful to have the
ability to inspect the bytea value directly (e.g. to see the length of
the bytea and not the string output).
>
> In another subthread you seem to have surrendered to the opinion that
> the new catalog should be called pg_statistics_ext, just in case in the
> future we come up with additional things to put on it. However, given
> its schema, with a "starelid / stakeys", is it sensible to think that
> we're going to get anything other than something that involves multiple
> variables? Maybe it should just be "pg_statistics_multivar" and if
> something else comes along we create another catalog with an appropriate
> schema. Heck, how does this catalog serve the purpose of cross-table
> statistics in the first place, given that it has room to record a single
> relid only? Are you thinking that in the future you'd change starelid
> into an oidvector column?
>
Yes, I think the starelid will turn into OID vector. The reason why I
haven't done that in the current version of the catalog is to keep it
simple. Supporting join statistics will require tracking OID for each
attribute, because those will be from multiple relations. It'll also
require tracking "join condition" and so on.
We've designed the CREATED STATISTICS syntax to support this extension,
but I'm strongly against complicating the catalogs at this point.
> The comment in gram.y about the CREATE STATISTICS is at odds with what
> is actually allowed by the grammar.
>
Which comment?
> I think the name of a statistics is only useful to DROP/ALTER it, right?
> I wonder why it's useful that statistics belongs in a schema. Perhaps
> it should be a global object? I suppose the name collisions would
> become bothersome if you have many mvstats.
>
I think it shouldn't be a global object. I consider them to be a part of
a schema (just like indexes, for example). Imagine you have a
multi-tenant database, with using exactly the same (tables/indexes)
schema, but keept in different schemas. Why shouldn't it be possible to
also use the same set of statistics for each tenant?
T.
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-04 00:03 Robert Haas <robertmhaas@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
1 sibling, 0 replies; 70+ messages in thread
From: Robert Haas @ 2017-02-04 00:03 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Alvaro Herrera <alvherre@2ndquadrant.com>; pgsql-hackers
On Thu, Feb 2, 2017 at 3:59 AM, Tomas Vondra
<tomas.vondra@2ndquadrant.com> wrote:
> There's a subtle difference between pg_node_tree and the data types for
> statistics - pg_node_tree stores the value as a string (matching the
> nodeToString output), so the _in function is fairly simple. Of course,
> stringToNode() assumes safe input, which is why the input is disabled.
>
> OTOH the statistics are stored in an optimized binary format, allowing to
> use the value directly (without having to do expensive parsing etc).
>
> I was thinking that the easiest way to add support for _in would be to add a
> bunch of Nodes for the statistics, along with in/out functions, but keeping
> the internal binary representation. But that'll be tricky to do in a safe
> way - even if those nodes are coded in a very defensive ways, I'd bet
> there'll be ways to inject unsafe nodes.
>
> So I'm OK with not having the _in for now. If needed, it's possible to
> construct the statistics as a bytea using a bit of C code. That's at least
> obviously unsafe, as anything written in C, touching the memory.
Since these data types are already special-purpose, I don't really see
why it would be desirable to entangle them with the existing code for
serializing and deserializing Nodes. Whether or not it's absolutely
necessary for these types to have input functions, it seems at least
possible that it would be useful, and it becomes much less likely that
we can make that work if it's piggybacking on stringToNode().
--
Robert Haas
EnterpriseDB: http://www.enterprisedb.com
The Enterprise PostgreSQL Company
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-06 21:26 Alvaro Herrera <alvherre@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
1 sibling, 1 reply; 70+ messages in thread
From: Alvaro Herrera @ 2017-02-06 21:26 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: pgsql-hackers
Tomas Vondra wrote:
> On 02/01/2017 11:52 PM, Alvaro Herrera wrote:
> > Nearby, some auxiliary functions such as n_choose_k and
> > num_combinations are not documented. What it is that they do? I'd
> > move these at the end of the file, keeping the important entry points
> > at the top of the file.
>
> I'd say n-choose-k is pretty widely known term from combinatorics. The
> comment would essentially say just 'this is n-choose-k' which seems rather
> pointless. So as much as I dislike the self-documenting code, this actually
> seems like a good case of that.
Actually, we do have such comments all over the place. I knew this as
"n sobre k", so the english name doesn't immediately ring a bell with me
until I look it up; I think the function comment could just say
"n_choose_k -- this function returns the binomial coefficient".
> > I see this patch has a estimate_ndistinct() which claims to be a re-
> > implementation of code already in analyze.c, but it is actually a lot
> > simpler than what analyze.c does. I've been wondering if it'd be a good
> > idea to use some of this code so that some routines are moved out of
> > analyze.c; good implementations of statistics-related functions would
> > live in src/backend/statistics/ where they can be used both by analyze.c
> > and your new mvstats stuff. (More generally I am beginning to wonder if
> > the new directory should be just src/backend/statistics.)
>
> I'll look into that. I have to check if I ignored some assumptions or corner
> cases the analyze.c deals with.
Maybe it's not terribly important to refactor analyze.c from the get go,
but let's give the subdir a more general name. Hence my vote for having
the subdir be "statistics" instead of "mvstats".
> > In another subthread you seem to have surrendered to the opinion that
> > the new catalog should be called pg_statistics_ext, just in case in the
> > future we come up with additional things to put on it. However, given
> > its schema, with a "starelid / stakeys", is it sensible to think that
> > we're going to get anything other than something that involves multiple
> > variables? Maybe it should just be "pg_statistics_multivar" and if
> > something else comes along we create another catalog with an appropriate
> > schema. Heck, how does this catalog serve the purpose of cross-table
> > statistics in the first place, given that it has room to record a single
> > relid only? Are you thinking that in the future you'd change starelid
> > into an oidvector column?
>
> Yes, I think the starelid will turn into OID vector. The reason why I
> haven't done that in the current version of the catalog is to keep it
> simple.
OK -- as long as we know what the way forward is, I'm good. Still, my
main point was that even if we have multiple rels, this catalog will be
about having multivariate statistics, and not different kinds of
statistical data. I would keep pg_mv_statistics, really.
> > The comment in gram.y about the CREATE STATISTICS is at odds with what
> > is actually allowed by the grammar.
>
> Which comment?
This one:
* CREATE STATISTICS stats_name ON relname (columns) WITH (options)
the production actually says:
CREATE STATISTICS any_name ON '(' columnList ')' FROM qualified_name
> > I think the name of a statistics is only useful to DROP/ALTER it, right?
> > I wonder why it's useful that statistics belongs in a schema. Perhaps
> > it should be a global object? I suppose the name collisions would
> > become bothersome if you have many mvstats.
>
> I think it shouldn't be a global object. I consider them to be a part of a
> schema (just like indexes, for example). Imagine you have a multi-tenant
> database, with using exactly the same (tables/indexes) schema, but keept in
> different schemas. Why shouldn't it be possible to also use the same set of
> statistics for each tenant?
True. Suggestion withdrawn.
--
Álvaro Herrera https://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-06 22:11 Alvaro Herrera <alvherre@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
2 siblings, 0 replies; 70+ messages in thread
From: Alvaro Herrera @ 2017-02-06 22:11 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Kyotaro HORIGUCHI <horiguchi.kyotaro@lab.ntt.co.jp>; ideriha.takeshi@jp.fujitsu.com; dilipbalaut@gmail.com; Langote_Amit_f8@lab.ntt.co.jp; dean.a.rasheed@gmail.com; hlinnaka@iki.fi; robertmhaas@gmail.com; ishii@postgresql.org; david@pgmasters.net; michael.paquier@gmail.com; tgl@sss.pgh.pa.us; petr@2ndquadrant.com; jeff.janes@gmail.com; pgsql-hackers
Looking at 0003, I notice that gram.y is changed to add a WITH ( .. )
clause. If it's not specified, an error is raised. If you create
stats with (ndistinct) then you can't alter it later to add
"dependencies" or whatever; unless I misunderstand, you have to drop the
statistics and create another one. Probably in a forthcoming patch we
should have ALTER support to add a stats type.
Also, why isn't the default to build everything, rather than nothing?
BTW, almost everything in the backend could be inside "utils/", so let's
not do that -- let's just create src/backend/statistics/ for all your
code.
Here a few notes while reading README.dependencies -- some typos, two
questions.
diff --git a/src/backend/utils/mvstats/README.dependencies b/src/backend/utils/mvstats/README.dependencies
index 908f094..7f3ed3d 100644
--- a/src/backend/utils/mvstats/README.dependencies
+++ b/src/backend/utils/mvstats/README.dependencies
@@ -36,7 +36,7 @@ design choice to model the dataset in denormalized way, either because of
performance or to make querying easier.
-soft dependencies
+Soft dependencies
-----------------
Real-world data sets often contain data errors, either because of data entry
@@ -48,7 +48,7 @@ rendering the approach mostly useless even for slightly noisy data sets, or
result in sudden changes in behavior depending on minor differences between
samples provided to ANALYZE.
-For this reason the statistics implementes "soft" functional dependencies,
+For this reason the statistics implements "soft" functional dependencies,
associating each functional dependency with a degree of validity (a number
number between 0 and 1). This degree is then used to combine selectivities
in a smooth manner.
@@ -75,6 +75,7 @@ The algorithm also requires a minimum size of the group to consider it
consistent (currently 3 rows in the sample). Small groups make it less likely
to break the consistency.
+## What is it that we store in the catalog?
Clause reduction (planner/optimizer)
------------------------------------
@@ -95,12 +96,12 @@ example for (a,b,c) we first use (a,b=>c) to break the computation into
and then apply (a=>b) the same way on P(a=?,b=?).
-Consistecy of clauses
+Consistency of clauses
---------------------
Functional dependencies only express general dependencies between columns,
without referencing particular values. This assumes that the equality clauses
-are in fact consistent with the functinal dependency, i.e. that given a
+are in fact consistent with the functional dependency, i.e. that given a
dependency (a=>b), the value in (b=?) clause is the value determined by (a=?).
If that's not the case, the clauses are "inconsistent" with the functional
dependency and the result will be over-estimation.
@@ -111,6 +112,7 @@ set will be empty, but we'll estimate the selectivity using the ZIP condition.
In this case the default estimation based on AVIA principle happens to work
better, but mostly by chance.
+## what is AVIA principle?
This issue is the price for the simplicity of functional dependencies. If the
application frequently constructs queries with clauses inconsistent with
--
Álvaro Herrera https://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-07 00:38 Alvaro Herrera <alvherre@2ndquadrant.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
2 siblings, 0 replies; 70+ messages in thread
From: Alvaro Herrera @ 2017-02-07 00:38 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: Kyotaro HORIGUCHI <horiguchi.kyotaro@lab.ntt.co.jp>; ideriha.takeshi@jp.fujitsu.com; dilipbalaut@gmail.com; Langote_Amit_f8@lab.ntt.co.jp; dean.a.rasheed@gmail.com; hlinnaka@iki.fi; robertmhaas@gmail.com; ishii@postgresql.org; david@pgmasters.net; michael.paquier@gmail.com; tgl@sss.pgh.pa.us; petr@2ndquadrant.com; jeff.janes@gmail.com; pgsql-hackers
Still about 0003. dependencies.c comment at the top of the file should
contain some details about what is it implementing and a general
description of the algorithm and data structures. As before, it's best
to have the main entry point build_mv_dependencies at the top, the other
public functions, keeping the internal routines at the bottom of the
file. That eases code study for future readers. (Minimizing number of
function prototypes is not a goal.)
What is MVSTAT_DEPS_TYPE_BASIC? Is "functional dependencies" really
BASIC? I wonder if it should be TYPE_FUNCTIONAL_DEPS or something.
As with pg_ndistinct_out, there's no need to pstrdup(str.data), as it's
already palloc'ed in the right context.
--
Álvaro Herrera https://www.2ndQuadrant.com/
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-08 15:23 Dean Rasheed <dean.a.rasheed@gmail.com>
parent: Alvaro Herrera <alvherre@2ndquadrant.com>
0 siblings, 1 reply; 70+ messages in thread
From: Dean Rasheed @ 2017-02-08 15:23 UTC (permalink / raw)
To: Alvaro Herrera <alvherre@2ndquadrant.com>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; pgsql-hackers
On 6 February 2017 at 21:26, Alvaro Herrera <alvherre@2ndquadrant.com> wrote:
> Tomas Vondra wrote:
>> On 02/01/2017 11:52 PM, Alvaro Herrera wrote:
>
>> > Nearby, some auxiliary functions such as n_choose_k and
>> > num_combinations are not documented. What it is that they do? I'd
>> > move these at the end of the file, keeping the important entry points
>> > at the top of the file.
>>
>> I'd say n-choose-k is pretty widely known term from combinatorics. The
>> comment would essentially say just 'this is n-choose-k' which seems rather
>> pointless. So as much as I dislike the self-documenting code, this actually
>> seems like a good case of that.
>
> Actually, we do have such comments all over the place. I knew this as
> "n sobre k", so the english name doesn't immediately ring a bell with me
> until I look it up; I think the function comment could just say
> "n_choose_k -- this function returns the binomial coefficient".
>
One of the things you have to watch out for when writing code to
compute binomial coefficients is integer overflow, since the numerator
and denominator get large very quickly. For example, the current code
will overflow for n=13, k=12, which really isn't that large.
This can be avoided by computing the product in reverse and using a
larger datatype like a 64-bit integer to store a single intermediate
result. The point about multiplying the terms in reverse is that it
guarantees that each intermediate result is an exact integer (a
smaller binomial coefficient), so there is no need to track separate
numerators and denominators, and you avoid huge intermediate
factorials. Here's what that looks like in psuedo-code:
binomial(int n, int k):
# Save computational effort by using the symmetry of the binomial
# coefficients
k = min(k, n-k);
# Compute the result using binomial(n, k) = binomial(n-1, k-1) * n / k,
# starting from binomial(n-k, 0) = 1, and computing the sequence
# binomial(n-k+1, 1), binomial(n-k+2, 2), ...
#
# Note that each intermediate result is an exact integer.
int64 result = 1;
for (int i = 1; i <= k; i++)
{
result = (result * (n-k+i)) / i;
if (result > INT_MAX) Raise overflow error
}
return (int) result;
Note also that I think num_combinations(n) is just an expensive way of
calculating 2^n - n - 1.
Regards,
Dean
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-08 16:09 David Fetter <david@fetter.org>
parent: Dean Rasheed <dean.a.rasheed@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: David Fetter @ 2017-02-08 16:09 UTC (permalink / raw)
To: Dean Rasheed <dean.a.rasheed@gmail.com>; +Cc: Alvaro Herrera <alvherre@2ndquadrant.com>; Tomas Vondra <tomas.vondra@2ndquadrant.com>; pgsql-hackers
On Wed, Feb 08, 2017 at 03:23:25PM +0000, Dean Rasheed wrote:
> On 6 February 2017 at 21:26, Alvaro Herrera <alvherre@2ndquadrant.com> wrote:
> > Tomas Vondra wrote:
> >> On 02/01/2017 11:52 PM, Alvaro Herrera wrote:
> >
> >> > Nearby, some auxiliary functions such as n_choose_k and
> >> > num_combinations are not documented. What it is that they do? I'd
> >> > move these at the end of the file, keeping the important entry points
> >> > at the top of the file.
> >>
> >> I'd say n-choose-k is pretty widely known term from combinatorics. The
> >> comment would essentially say just 'this is n-choose-k' which seems rather
> >> pointless. So as much as I dislike the self-documenting code, this actually
> >> seems like a good case of that.
> >
> > Actually, we do have such comments all over the place. I knew this as
> > "n sobre k", so the english name doesn't immediately ring a bell with me
> > until I look it up; I think the function comment could just say
> > "n_choose_k -- this function returns the binomial coefficient".
>
> One of the things you have to watch out for when writing code to
> compute binomial coefficients is integer overflow, since the numerator
> and denominator get large very quickly. For example, the current code
> will overflow for n=13, k=12, which really isn't that large.
>
> This can be avoided by computing the product in reverse and using a
> larger datatype like a 64-bit integer to store a single intermediate
> result. The point about multiplying the terms in reverse is that it
> guarantees that each intermediate result is an exact integer (a
> smaller binomial coefficient), so there is no need to track separate
> numerators and denominators, and you avoid huge intermediate
> factorials. Here's what that looks like in psuedo-code:
>
> binomial(int n, int k):
> # Save computational effort by using the symmetry of the binomial
> # coefficients
> k = min(k, n-k);
>
> # Compute the result using binomial(n, k) = binomial(n-1, k-1) * n / k,
> # starting from binomial(n-k, 0) = 1, and computing the sequence
> # binomial(n-k+1, 1), binomial(n-k+2, 2), ...
> #
> # Note that each intermediate result is an exact integer.
> int64 result = 1;
> for (int i = 1; i <= k; i++)
> {
> result = (result * (n-k+i)) / i;
> if (result > INT_MAX) Raise overflow error
> }
> return (int) result;
>
>
> Note also that I think num_combinations(n) is just an expensive way of
> calculating 2^n - n - 1.
Combinations are n!/(k! * (n-k)!), so computing those is more
along the lines of:
unsigned long long
choose(unsigned long long n, unsigned long long k) {
if (k > n) {
return 0;
}
unsigned long long r = 1;
for (unsigned long long d = 1; d <= k; ++d) {
r *= n--;
r /= d;
}
return r;
}
which greatly reduces the chance of overflow.
Best,
David.
--
David Fetter <david(at)fetter(dot)org> http://fetter.org/
Phone: +1 415 235 3778 AIM: dfetter666 Yahoo!: dfetter
Skype: davidfetter XMPP: david(dot)fetter(at)gmail(dot)com
Remember to vote!
Consider donating to Postgres: http://www.postgresql.org/about/donate
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-08 18:40 Dean Rasheed <dean.a.rasheed@gmail.com>
parent: David Fetter <david@fetter.org>
0 siblings, 1 reply; 70+ messages in thread
From: Dean Rasheed @ 2017-02-08 18:40 UTC (permalink / raw)
To: David Fetter <david@fetter.org>; +Cc: Alvaro Herrera <alvherre@2ndquadrant.com>; Tomas Vondra <tomas.vondra@2ndquadrant.com>; pgsql-hackers
On 8 February 2017 at 16:09, David Fetter <david@fetter.org> wrote:
> Combinations are n!/(k! * (n-k)!), so computing those is more
> along the lines of:
>
> unsigned long long
> choose(unsigned long long n, unsigned long long k) {
> if (k > n) {
> return 0;
> }
> unsigned long long r = 1;
> for (unsigned long long d = 1; d <= k; ++d) {
> r *= n--;
> r /= d;
> }
> return r;
> }
>
> which greatly reduces the chance of overflow.
>
Hmm, but that doesn't actually prevent overflows, since it can
overflow in the multiplication step, and there is no protection
against that.
In the algorithm I presented, the inputs and the intermediate result
are kept below INT_MAX, so the multiplication step cannot overflow the
64-bit integer, and it will only raise an overflow error if the actual
result won't fit in a 32-bit int. Actually a crucial part of that,
which I failed to mention previously, is the first step replacing k
with min(k, n-k). This is necessary for inputs like (100,99), which
should return 100, and which must be computed as 100 choose 1, not 100
choose 99, otherwise it will overflow internally before getting to the
final result.
Regards,
Dean
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-11 01:17 Tomas Vondra <tomas.vondra@2ndquadrant.com>
parent: Dean Rasheed <dean.a.rasheed@gmail.com>
0 siblings, 1 reply; 70+ messages in thread
From: Tomas Vondra @ 2017-02-11 01:17 UTC (permalink / raw)
To: Dean Rasheed <dean.a.rasheed@gmail.com>; David Fetter <david@fetter.org>; +Cc: Alvaro Herrera <alvherre@2ndquadrant.com>; pgsql-hackers
On 02/08/2017 07:40 PM, Dean Rasheed wrote:
> On 8 February 2017 at 16:09, David Fetter <david@fetter.org> wrote:
>> Combinations are n!/(k! * (n-k)!), so computing those is more
>> along the lines of:
>>
>> unsigned long long
>> choose(unsigned long long n, unsigned long long k) {
>> if (k > n) {
>> return 0;
>> }
>> unsigned long long r = 1;
>> for (unsigned long long d = 1; d <= k; ++d) {
>> r *= n--;
>> r /= d;
>> }
>> return r;
>> }
>>
>> which greatly reduces the chance of overflow.
>>
>
> Hmm, but that doesn't actually prevent overflows, since it can
> overflow in the multiplication step, and there is no protection
> against that.
>
> In the algorithm I presented, the inputs and the intermediate result
> are kept below INT_MAX, so the multiplication step cannot overflow the
> 64-bit integer, and it will only raise an overflow error if the actual
> result won't fit in a 32-bit int. Actually a crucial part of that,
> which I failed to mention previously, is the first step replacing k
> with min(k, n-k). This is necessary for inputs like (100,99), which
> should return 100, and which must be computed as 100 choose 1, not 100
> choose 99, otherwise it will overflow internally before getting to the
> final result.
>
Thanks for the feedback, I'll fix this. I've allowed myself to be a bit
sloppy because the number of attributes in the statistics is currently
limited to 8, so the overflows are currently not an issue. But it
doesn't hurt to make it future-proof, in case we change that mostly
artificial limit sometime in the future.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-12 10:35 Dean Rasheed <dean.a.rasheed@gmail.com>
parent: Tomas Vondra <tomas.vondra@2ndquadrant.com>
0 siblings, 1 reply; 70+ messages in thread
From: Dean Rasheed @ 2017-02-12 10:35 UTC (permalink / raw)
To: Tomas Vondra <tomas.vondra@2ndquadrant.com>; +Cc: David Fetter <david@fetter.org>; Alvaro Herrera <alvherre@2ndquadrant.com>; pgsql-hackers
On 11 February 2017 at 01:17, Tomas Vondra <tomas.vondra@2ndquadrant.com> wrote:
> Thanks for the feedback, I'll fix this. I've allowed myself to be a bit
> sloppy because the number of attributes in the statistics is currently
> limited to 8, so the overflows are currently not an issue. But it doesn't
> hurt to make it future-proof, in case we change that mostly artificial limit
> sometime in the future.
>
Ah right, so it can't overflow at present, but it's neater to have an
overflow-proof algorithm.
Thinking about the exactness of the division steps is quite
interesting. Actually, the order of the multiplying factors doesn't
matter as long as the divisors are in increasing order. So in both my
proposal:
result = 1
for (i = 1; i <= k; i++)
result = (result * (n-k+i)) / i;
and David's proposal, which is equivalent but has the multiplying
factors in the opposite order, equivalent to:
result = 1
for (i = 1; i <= k; i++)
result = (result * (n-i+1)) / i;
the divisions are exact at each step. The first time through the loop
it divides by 1 which is trivially exact. The second time it divides
by 2, having multiplied by 2 consecutive factors, one of which is
therefore guaranteed to be divisible by 2. The third time it divides
by 3, having multiplied by 3 consecutive factors, one of which is
therefore guaranteed to be divisible by 3, and so on.
My approach originally seemed more logical to me because of the way it
derives from the recurrence relation binomial(n, k) = binomial(n-1,
k-1) * n / k, but they both work fine as long as they have suitable
overflow checks.
It's also interesting that descriptions of this algorithm tend to talk
about setting k to min(k, n-k) at the start as an optimisation step,
as I did in fact, whereas it's actually more than that -- it helps
prevent unnecessary intermediate overflows when k > n/2. Of course,
that's not a worry for the current use of this function, but it's good
to have a robust algorithm.
Regards,
Dean
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
* Re: multivariate statistics (v19)
@ 2017-02-12 19:42 David Fetter <david@fetter.org>
parent: Dean Rasheed <dean.a.rasheed@gmail.com>
0 siblings, 0 replies; 70+ messages in thread
From: David Fetter @ 2017-02-12 19:42 UTC (permalink / raw)
To: Dean Rasheed <dean.a.rasheed@gmail.com>; +Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>; Alvaro Herrera <alvherre@2ndquadrant.com>; pgsql-hackers
On Sun, Feb 12, 2017 at 10:35:04AM +0000, Dean Rasheed wrote:
> On 11 February 2017 at 01:17, Tomas Vondra <tomas.vondra@2ndquadrant.com> wrote:
> > Thanks for the feedback, I'll fix this. I've allowed myself to be a bit
> > sloppy because the number of attributes in the statistics is currently
> > limited to 8, so the overflows are currently not an issue. But it doesn't
> > hurt to make it future-proof, in case we change that mostly artificial limit
> > sometime in the future.
> >
>
> Ah right, so it can't overflow at present, but it's neater to have an
> overflow-proof algorithm.
>
> Thinking about the exactness of the division steps is quite
> interesting. Actually, the order of the multiplying factors doesn't
> matter as long as the divisors are in increasing order. So in both my
> proposal:
>
> result = 1
> for (i = 1; i <= k; i++)
> result = (result * (n-k+i)) / i;
>
> and David's proposal, which is equivalent but has the multiplying
> factors in the opposite order, equivalent to:
>
> result = 1
> for (i = 1; i <= k; i++)
> result = (result * (n-i+1)) / i;
>
> the divisions are exact at each step. The first time through the loop
> it divides by 1 which is trivially exact. The second time it divides
> by 2, having multiplied by 2 consecutive factors, one of which is
> therefore guaranteed to be divisible by 2. The third time it divides
> by 3, having multiplied by 3 consecutive factors, one of which is
> therefore guaranteed to be divisible by 3, and so on.
Right. You know you can use integer division, which make sense as
permutations of discrete sets are always integers.
> My approach originally seemed more logical to me because of the way it
> derives from the recurrence relation binomial(n, k) = binomial(n-1,
> k-1) * n / k, but they both work fine as long as they have suitable
> overflow checks.
Right. We could even cache those checks (sorry) based on data type
limits by architecture and OS if performance on those operations ever
matters that much.
> It's also interesting that descriptions of this algorithm tend to
> talk about setting k to min(k, n-k) at the start as an optimisation
> step, as I did in fact, whereas it's actually more than that -- it
> helps prevent unnecessary intermediate overflows when k > n/2. Of
> course, that's not a worry for the current use of this function, but
> it's good to have a robust algorithm.
Indeed. :)
Best,
David.
--
David Fetter <david(at)fetter(dot)org> http://fetter.org/
Phone: +1 415 235 3778 AIM: dfetter666 Yahoo!: dfetter
Skype: davidfetter XMPP: david(dot)fetter(at)gmail(dot)com
Remember to vote!
Consider donating to Postgres: http://www.postgresql.org/about/donate
--
Sent via pgsql-hackers mailing list (pgsql-hackers@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-hackers
^ permalink raw reply [nested|flat] 70+ messages in thread
end of thread, other threads:[~2017-02-12 19:42 UTC | newest]
Thread overview: 70+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2016-08-03 01:58 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-08-05 04:24 ` Michael Paquier <michael.paquier@gmail.com>
2016-08-05 17:38 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-08-05 22:21 ` Michael Paquier <michael.paquier@gmail.com>
2016-08-10 04:41 ` Michael Paquier <michael.paquier@gmail.com>
2016-08-10 11:33 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-08-10 11:50 ` Petr Jelinek <petr@2ndquadrant.com>
2016-08-10 12:24 ` Michael Paquier <michael.paquier@gmail.com>
2016-08-10 18:09 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-08-10 12:23 ` Michael Paquier <michael.paquier@gmail.com>
2016-08-10 18:34 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-08-11 05:55 ` Michael Paquier <michael.paquier@gmail.com>
2016-08-15 20:50 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-08-23 17:03 ` Robert Haas <robertmhaas@gmail.com>
2016-08-30 06:54 ` Michael Paquier <michael.paquier@gmail.com>
2016-09-12 14:08 ` Dean Rasheed <dean.a.rasheed@gmail.com>
2016-09-13 22:01 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-09-30 11:10 ` Heikki Linnakangas <hlinnaka@iki.fi>
2016-10-03 01:46 ` Michael Paquier <michael.paquier@gmail.com>
2016-10-03 11:25 ` Heikki Linnakangas <hlinnaka@iki.fi>
2016-10-04 03:25 ` Michael Paquier <michael.paquier@gmail.com>
2016-10-04 07:37 ` Dean Rasheed <dean.a.rasheed@gmail.com>
2016-10-04 08:51 ` Gavin Flower <GavinFlower@archidevsys.co.nz>
2016-10-04 07:49 ` Dean Rasheed <dean.a.rasheed@gmail.com>
2016-10-04 08:15 ` Heikki Linnakangas <hlinnaka@iki.fi>
2016-10-04 09:21 ` Dean Rasheed <dean.a.rasheed@gmail.com>
2016-10-11 03:39 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-10-29 19:23 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-12-12 11:26 ` Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>
2016-12-12 21:50 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-12-30 13:05 ` Petr Jelinek <petr.jelinek@2ndquadrant.com>
2016-12-30 13:12 ` Petr Jelinek <petr.jelinek@2ndquadrant.com>
2017-01-03 21:55 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-01-03 13:42 ` Dilip Kumar <dilipbalaut@gmail.com>
2017-01-03 16:22 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-01-04 02:35 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-01-04 14:21 ` Dilip Kumar <dilipbalaut@gmail.com>
2017-01-04 21:57 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-01-26 09:43 ` Dilip Kumar <dilipbalaut@gmail.com>
2017-01-30 19:33 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-01-25 05:55 ` Michael Paquier <michael.paquier@gmail.com>
2017-01-25 12:56 ` Alvaro Herrera <alvherre@2ndquadrant.com>
2017-01-25 21:43 ` Michael Paquier <michael.paquier@gmail.com>
2017-01-26 09:03 ` Ideriha, Takeshi <ideriha.takeshi@jp.fujitsu.com>
2017-01-26 11:01 ` Kyotaro HORIGUCHI <horiguchi.kyotaro@lab.ntt.co.jp>
2017-01-30 19:12 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-02-01 22:52 ` Alvaro Herrera <alvherre@2ndquadrant.com>
2017-02-02 08:59 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-02-04 00:03 ` Robert Haas <robertmhaas@gmail.com>
2017-02-06 21:26 ` Alvaro Herrera <alvherre@2ndquadrant.com>
2017-02-08 15:23 ` Dean Rasheed <dean.a.rasheed@gmail.com>
2017-02-08 16:09 ` David Fetter <david@fetter.org>
2017-02-08 18:40 ` Dean Rasheed <dean.a.rasheed@gmail.com>
2017-02-11 01:17 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-02-12 10:35 ` Dean Rasheed <dean.a.rasheed@gmail.com>
2017-02-12 19:42 ` David Fetter <david@fetter.org>
2017-02-06 22:11 ` Alvaro Herrera <alvherre@2ndquadrant.com>
2017-02-07 00:38 ` Alvaro Herrera <alvherre@2ndquadrant.com>
2017-01-30 19:36 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-01-30 16:12 ` Alvaro Herrera <alvherre@2ndquadrant.com>
2017-01-30 20:00 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-01-30 20:37 ` Alvaro Herrera <alvherre@2ndquadrant.com>
2017-01-30 21:57 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-01-31 04:21 ` Michael Paquier <michael.paquier@gmail.com>
2017-01-31 06:52 ` Amit Langote <Langote_Amit_f8@lab.ntt.co.jp>
2017-01-31 16:10 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2017-01-30 16:55 ` Alvaro Herrera <alvherre@2ndquadrant.com>
2017-01-30 19:40 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
2016-08-10 13:29 ` Ants Aasma <ants.aasma@eesti.ee>
2016-08-10 18:07 ` Tomas Vondra <tomas.vondra@2ndquadrant.com>
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox