pg.ddx.io  pgsql-hackers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Jeff Davis <pgsql@j-davis.com>
To: Justin Pryzby <pryzby@telsasoft.com>
To: Andres Freund <andres@anarazel.de>
Cc: Tomas Vondra <tomas.vondra@2ndquadrant.com>
Cc: pgsql-hackers@lists.postgresql.org, Jeff Janes <jeff.janes@gmail.com>
Subject: Re: explain HashAggregate to report bucket and memory stats
Date: Fri, 13 Mar 2020 10:15:46 -0700
Message-ID: <f95260dc8cd3b5a4de85d9f89b5bae19dfbc4c13.camel@j-davis.com> (raw)
In-Reply-To: <20200306213310.GM684@telsasoft.com>
References: <20200103161925.GM12066@telsasoft.com>
	<20200203145301.53mozz7gcdaklnjc@alap3.anarazel.de>
	<20200216000220.GF31889@telsasoft.com>
	<20200216175307.GJ31889@telsasoft.com>
	<20200219201037.GA21433@telsasoft.com>
	<20200222215335.bu4kky7el4nccls5@development>
	<20200225033532.GX31889@telsasoft.com>
	<20200301154540.GL29456@telsasoft.com>
	<20200306175859.d56ohskarwldyrrw@alap3.anarazel.de>
	<20200306213310.GM684@telsasoft.com>


+		/* hashtable->entrysize includes additionalsize */
+		hashtable->instrument.space_peak_hash = Max(
+			hashtable->instrument.space_peak_hash,
+			hashtable->hashtab->size *
sizeof(TupleHashEntryData));
+
+		hashtable->instrument.space_peak_tuples = Max(
+			hashtable->instrument.space_peak_tuples,
+				hashtable->hashtab->members *
hashtable->entrysize);

I think, in general, we should avoid estimates/projections for
reporting and try to get at a real number, like
MemoryContextMemAllocated(). (Aside: I may want to tweak exactly what
that function reports so that it doesn't count the unused portion of
the last block.)

For instance, the report is still not accurate, because it doesn't
account for pass-by-ref transition state values.

To use memory-context-based reporting, it's hard to make the stats a
part of the tuple hash table, because the tuple hash table doesn't own
the memory contexts (they are passed in). It's also hard to make it
per-hashtable (e.g. for grouping sets), unless we put each grouping set
in its own memory context.

Also, is there a reason you report two different memory values
(hashtable and tuples)? I don't object, but it seems like a little too
much detail.

Regards,
	Jeff Davis







view thread (24+ messages)  latest in thread

Message-ID: <f95260dc8cd3b5a4de85d9f89b5bae19dfbc4c13.camel@j-davis.com>
Permalink:  ../f95260dc8cd3b5a4de85d9f89b5bae19dfbc4c13.camel@j-davis.com/
Also on:    postgresql.org/message-id/f95260dc8cd3b5a4de85d9f89b5bae19dfbc4c13.camel@j-davis.com

 · 

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: pgsql@j-davis.com, pryzby@telsasoft.com, andres@anarazel.de, tomas.vondra@2ndquadrant.com, jeff.janes@gmail.com
  Subject: Re: explain HashAggregate to report bucket and memory stats
  In-Reply-To: <f95260dc8cd3b5a4de85d9f89b5bae19dfbc4c13.camel@j-davis.com>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox