From: Tom Lane <tgl@sss.pgh.pa.us>
To: Dimitrios Apostolou <jimis@gmx.net>
cc: pgsql-general@lists.postgresql.org
Subject: Re: SELECT DISTINCT chooses parallel seqscan instead of indexscan on huge table with 1000 partitions
In-reply-to: <559b0e40-63e6-fa9a-6b03-d1eba10f30f8@gmx.net>
References: <7886a68f-b466-2131-1747-f69f0fb71a37@gmx.net> <69077f15-4125-2d63-733f-21ce6eac4f01@gmx.net> <559b0e40-63e6-fa9a-6b03-d1eba10f30f8@gmx.net>
Comments: In-reply-to Dimitrios Apostolou <jimis@gmx.net>
	message dated "Fri, 10 May 2024 21:35:57 +0200"
MIME-Version: 1.0
Content-Type: text/plain; charset="us-ascii"
Content-ID: <1629462.1715372568.1@sss.pgh.pa.us>
Date: Fri, 10 May 2024 16:22:48 -0400
Message-ID: <1629463.1715372568@sss.pgh.pa.us>
Archived-At: <https://www.postgresql.org/message-id/1629463.1715372568%40sss.pgh.pa.us>
Precedence: bulk

Dimitrios Apostolou <jimis@gmx.net> writes:
> Further digging into this simple query, if I force the non-parallel plan
> by setting max_parallel_workers_per_gather TO 0, I see that the query
> planner comes up with a cost much higher:

>   Limit  (cost=363.84..1134528847.47 rows=10 width=4)
>     ->  Unique  (cost=363.84..22690570036.41 rows=200 width=4)
>           ->  Append  (cost=363.84..22527480551.58 rows=65235793929 width=4)
> ...

> The total cost on the 1st line (cost=363.84..1134528847.47) has a much
> higher upper limit than the total cost when
> max_parallel_workers_per_gather is 4 (cost=853891608.79..853891608.99).
> This explains the planner's choice. But I wonder why the cost estimation
> is so far away from reality.

I'd say the blame lies with that (probably-default) estimate of
just 200 distinct rows.  That means the planner expects to have
to read about 5% (10/200) of the tables to get the result, and
that's making fast-start plans look bad.

Possibly an explicit ANALYZE on the partitioned table would help.

			regards, tom lane