From: Andy Colson <andy@squeakycode.net>
To: Greg Smith <greg@2ndquadrant.com>
Cc: James Mansion <james@mansionfamily.plus.com>
Cc: Robert Haas <robertmhaas@gmail.com>
Cc: Claudio Freire <klaussfreire@gmail.com>
Cc: Tomas Vondra <tv@fuzzy.cz>
Cc: pgsql-performance@postgresql.org
Subject: Re: Performance
Date: Fri, 29 Apr 2011 15:23:25 -0500
Message-ID: <4DBB1E3D.2090002@squeakycode.net> (raw)
In-Reply-To: <4DBB09B5.80108@2ndquadrant.com>
References: <DF2D5436-117C-4D02-9B3C-A55723B7DDE1@darkstatic.com>
<20110412171855.GA14292@tux>
<FC3A3A2B-3ECB-41BA-8F94-356D6FED3695@darkstatic.com>
<4DA496FA.3070908@fuzzy.cz>
<8F22D592-23C1-4A3C-94A5-48363332ADD3@darkstatic.com>
<4DA4BFA2.5060601@fuzzy.cz>
<B93DFBBB-DA56-4044-A508-0B7E4A2CFD28@darkstatic.com>
<4DA4D40A.4010200@fuzzy.cz>
<A0B9339E-9D58-4842-A206-050B773360B6@darkstatic.com>
<4DA56DA8020000250003C783@gw.wicourts.gov>
<BANLkTinJqXqErAtQg--1MOaw2LJNFZmwOg@mail.gmail.com>
<BANLkTimyWkoX8Dj=4CKAjhY82ibru-An7g@mail.gmail.com>
<4DA6216E.9020907@fuzzy.cz>
<BANLkTim_m-oz3UMXEKUdFjWJ6uuAza79cw@mail.gmail.com>
<4DA63125.5070106@fuzzy.cz>
<BANLkTikYtnzTfS8Yjc-3ap9kysXLLpJwmg@mail.gmail.com>
<D2F0DB51-7D42-4573-99CA-B72E28440B16@gmail.com>
<4DBA75DC.6070506@mansionfamily.plus.com>
<4DBB09B5.80108@2ndquadrant.com>
On 4/29/2011 1:55 PM, Greg Smith wrote:
> James Mansion wrote:
>> Does the server know which IO it thinks is sequential, and which it
>> thinks is random? Could it not time the IOs (perhaps optionally) and
>> at least keep some sort of statistics of the actual observed times?
>
> It makes some assumptions based on what the individual query nodes are
> doing. Sequential scans are obviously sequential; index lookupss random;
> bitmap index scans random.
>
> The "measure the I/O and determine cache state from latency profile" has
> been tried, I believe it was Greg Stark who ran a good experiment of
> that a few years ago. Based on the difficulties of figuring out what
> you're actually going to with that data, I don't think the idea will
> ever go anywhere. There are some really nasty feedback loops possible in
> all these approaches for better modeling what's in cache, and this one
> suffers the worst from that possibility. If for example you discover
> that accessing index blocks is slow, you might avoid using them in favor
> of a measured fast sequential scan. Once you've fallen into that local
> minimum, you're stuck there. Since you never access the index blocks,
> they'll never get into RAM so that accessing them becomes fast--even
> though doing that once might be much more efficient, long-term, than
> avoiding the index.
>
> There are also some severe query plan stability issues with this idea
> beyond this. The idea that your plan might vary based on execution
> latency, that the system load going up can make query plans alter with
> it, is terrifying for a production server.
>
How about if the stats were kept, but had no affect on plans, or
optimizer or anything else.
It would be a diag tool. When someone wrote the list saying "AH! It
used the wrong index!". You could say, "please post your config
settings, and the stats from 'select * from pg_stats_something'"
We (or, you really) could compare the seq_page_cost and random_page_cost
from the config to the stats collected by PG and determine they are way
off... and you should edit your config a little and restart PG.
-Andy
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-performance@postgresql.org
Cc: andy@squeakycode.net, greg@2ndquadrant.com, james@mansionfamily.plus.com, robertmhaas@gmail.com, klaussfreire@gmail.com, tv@fuzzy.cz
Subject: Re: Performance
In-Reply-To: <4DBB1E3D.2090002@squeakycode.net>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox