From: Tomas Vondra <tomas.vondra@2ndquadrant.com>
To: Robert Haas <robertmhaas@gmail.com>
Cc: Ildus Kurbangaliev <i.kurbangaliev@postgrespro.ru>
Cc: pgsql-hackers@postgresql.org <pgsql-hackers@postgresql.org>
Subject: Re: [HACKERS] Custom compression methods
Date: Thu, 14 Dec 2017 18:23:30 +0100
Message-ID: <f86f1a2f-9510-e712-2a53-bfc17f6d2414@2ndquadrant.com> (raw)
In-Reply-To: <CA+TgmoYszJaQivv1eG4JAV3S19t61Y68KLJPNiRUABNWurQANA@mail.gmail.com>
References: <20170907194236.4cefce96@wp.localdomain>
<29527031-c837-59e1-760d-677bb33d6b0f@2ndquadrant.com>
<CA+Tgmoax3Hz3ZLupxXofurFbpgKBEwe8qd7N65C7G6xoE=xQTw@mail.gmail.com>
<eefe010e-1bac-03b2-2ca9-e94985c5faa2@2ndquadrant.com>
<CA+TgmoY0OnO8cfKEarPb6AXHGTBh9EeW93cemXhmWfcH0xg+xw@mail.gmail.com>
<b1e047de-67d5-2fd8-ad9e-93434497ad91@2ndquadrant.com>
<CA+TgmoYszJaQivv1eG4JAV3S19t61Y68KLJPNiRUABNWurQANA@mail.gmail.com>
On 12/14/2017 04:21 PM, Robert Haas wrote:
> On Wed, Dec 13, 2017 at 5:10 AM, Tomas Vondra
> <tomas.vondra@2ndquadrant.com> wrote:
>>> 2. If several data types can benefit from a similar approach, it has
>>> to be separately implemented for each one.
>>
>> I don't think the current solution improves that, though. If you
>> want to exploit internal features of individual data types, it
>> pretty much requires code customized to every such data type.
>>
>> For example you can't take the tsvector compression and just slap
>> it on tsquery, because it relies on knowledge of internal tsvector
>> structure. So you need separate implementations anyway.
>
> I don't think that's necessarily true. Certainly, it's true that
> *if* tsvector compression depends on knowledge of internal tsvector
> structure, *then* that you can't use the implementation for anything
> else (this, by the way, means that there needs to be some way for a
> compression method to reject being applied to a column of a data
> type it doesn't like).
I believe such dependency (on implementation details) is pretty much the
main benefit of datatype-aware compression methods. If you don't rely on
such assumption, then I'd say it's a general-purpose compression method.
> However, it seems possible to imagine compression algorithms that can
> work for a variety of data types, too. There might be a compression
> algorithm that is theoretically a general-purpose algorithm but has
> features which are particularly well-suited to, say, JSON or XML
> data, because it looks for word boundaries to decide on what strings
> to insert into the compression dictionary.
>
Can you give an example of such algorithm? Because I haven't seen such
example, and I find arguments based on hypothetical compression methods
somewhat suspicious.
FWIW I'm not against considering such compression methods, but OTOH it
may not be such a great primary use case to drive the overall design.
regards
--
Tomas Vondra http://www.2ndQuadrant.com
PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-hackers@postgresql.org
Cc: tomas.vondra@2ndquadrant.com, robertmhaas@gmail.com, i.kurbangaliev@postgrespro.ru
Subject: Re: [HACKERS] Custom compression methods
In-Reply-To: <f86f1a2f-9510-e712-2a53-bfc17f6d2414@2ndquadrant.com>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox