pg.ddx.io  pgsql-hackers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Tomas Vondra <tomas.vondra@2ndquadrant.com>
To: Dilip Kumar <dilipbalaut@gmail.com>
Cc: Robert Haas <robertmhaas@gmail.com>
Cc: Alexander Korotkov <a.korotkov@postgrespro.ru>
Cc: David Steele <david@pgmasters.net>
Cc: Ildus Kurbangaliev <i.kurbangaliev@gmail.com>
Cc: Dmitry Dolgov <9erthalion6@gmail.com>
Cc: PostgreSQL Developers <pgsql-hackers@lists.postgresql.org>
Subject: Re: [HACKERS] Custom compression methods
Date: Mon, 5 Oct 2020 00:07:13 +0200
Message-ID: <20201004220713.6vlmm2e3amlz2dil@development> (raw)
In-Reply-To: <CAFiTN-v2_h2+nPv6A-XLrs+E-ms970Xr78wGYv8A9fmDb1P0AQ@mail.gmail.com>
References: <20190228164415.62bd4107@ildus-work.localdomain>
	<53ab954a-bf44-2bff-fbe0-29f093899604@pgmasters.net>
	<CAPpHfdsQtby_G0HxPMrDA9G=swhC4o0sbm9hZ35A_uP8uw2TEQ@mail.gmail.com>
	<CA+TgmobSDVgUage9qQ5P_=F_9jaMkCgyKxUQGtFQU7oN4kX-AA@mail.gmail.com>
	<CAFiTN-uUpX3ck=K0mLEk-G_kUQY=SNOTeqdaNRR9FMdQrHKebw@mail.gmail.com>
	<CAFiTN-viudxb0kGwtwsXhSCAaNZ4OUYcP-7PM+tr-rEbgGeHQw@mail.gmail.com>
	<CA+TgmoYYeMv=eJqf2JF8VWDPe=S6rYfH78DOxzE-Sa2kQHhiPw@mail.gmail.com>
	<CAFiTN-vcbfy5ScKVUp16c1N_wzP0RL6EkPBAg_Jm3eDK0ftO5Q@mail.gmail.com>
	<CAFiTN-t+0y5xnPx+sSvverfXPk9E6OMUkdfgMg7mcTTbs8rokQ@mail.gmail.com>
	<CAFiTN-v2_h2+nPv6A-XLrs+E-ms970Xr78wGYv8A9fmDb1P0AQ@mail.gmail.com>

Hi,

I took a look at this patch after a long time, and done a bit of a
review+testing. I haven't re-read the whole thread since 2017 so some of
the following comments might be mistaken - sorry about that :-(


1) The "cmapi.h" naming seems unnecessarily short. I'd suggest using
simply compression or something like that. I see little reason to
shorten "compression" to "cm", or to prefix files with "cm_". For
example compression/cm_zlib.c might just be compression/zlib.c.


2) I see index_form_tuple does this:

     Datum  cvalue = toast_compress_datum(untoasted_values[i],
                                          DefaultCompressionMethod);

which seems wrong - why shouldn't the indexes use the same compression
method as the underlying table?


3) dumpTableSchema in pg_dump.c does this:

     switch (tbinfo->attcompression[j])
     {
         case 'p':
             cmname = "pglz";
         case 'z':
             cmname = "zlib";
     }

which is broken as it's missing break, so 'p' will produce 'zlib'.


4) The name ExecCompareCompressionMethod is somewhat misleading, as the
functions is not merely comparing compression methods - it also
recompresses the data.


5) CheckCompressionMethodsPreserved should document what the return
value is (true when new list contains all old values, thus not requiring
a rewrite). Maybe "Compare" would be a better name?


6) The new field in ColumnDef is missing a comment.


7) It's not clear to me what "partial list" in the PRESERVE docs means.

+ which of them should be kept on the column. Without PRESERVE or partial
+ list of compression methods the table will be rewritten.


8) The initial synopsis in alter_table.sgml includes the PRESERVE
syntax, but then later in the page it's omitted (yet the section talks
about the keyword).


9) attcompression ...

The main issue I see is what the patch does with attcompression. Instead
of just using it to store a the compression method, it's also used to
store the preserved compression methods. And using NameData to store
this seems wrong too - if we really want to store this info, the correct
way is either using text[] or inventing charvector or similar.

But to me this seems very much like a misuse of attcompression to track
dependencies on compression methods, necessary because we don't have a
separate catalog listing compression methods. If we had that, I think we
could simply add dependencies between attributes and that catalog.

Moreover, having the catalog would allow adding compression methods
(from extensions etc) instead of just having a list of hard-coded
compression methods. Which seems like a strange limitation, considering
this thread is called "custom compression methods".


10) compression parameters?

I wonder if we could/should allow parameters, like compression level
(and maybe other stuff, depending on the compression method). PG13
allowed that for opclasses, so perhaps we should allow it here too.


11) pg_column_compression

When specifying compression method not present in attcompression, we get
this error message and hint:

   test=# alter table t alter COLUMN a set compression "pglz" preserve (zlib);
   ERROR:  "zlib" compression access method cannot be preserved
   HINT:  use "pg_column_compression" function for list of compression methods

but there is no pg_column_compression function, so the hint is wrong.


regards

-- 
Tomas Vondra                  http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services





view thread (430+ messages)  latest in thread

Message-ID: <20201004220713.6vlmm2e3amlz2dil@development>
Permalink:  ../20201004220713.6vlmm2e3amlz2dil@development/
Also on:    postgresql.org/message-id/20201004220713.6vlmm2e3amlz2dil@development

 · 

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: tomas.vondra@2ndquadrant.com, dilipbalaut@gmail.com, robertmhaas@gmail.com, a.korotkov@postgrespro.ru, david@pgmasters.net, i.kurbangaliev@gmail.com, 9erthalion6@gmail.com, pgsql-hackers@lists.postgresql.org
  Subject: Re: [HACKERS] Custom compression methods
  In-Reply-To: <20201004220713.6vlmm2e3amlz2dil@development>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox