Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtp (Exim 4.84_2) (envelope-from ) id 1eP408-0000XM-5C for pgsql-hackers@arkaria.postgresql.org; Wed, 13 Dec 2017 10:10:56 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.84_2) (envelope-from ) id 1eP407-0004rH-PG for pgsql-hackers@arkaria.postgresql.org; Wed, 13 Dec 2017 10:10:55 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_CBC_SHA384:256) (Exim 4.84_2) (envelope-from ) id 1eP407-0004r7-Iq for pgsql-hackers@lists.postgresql.org; Wed, 13 Dec 2017 10:10:55 +0000 Received: from mail-wm0-x242.google.com ([2a00:1450:400c:c09::242]) by magus.postgresql.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_CBC_SHA1:256) (Exim 4.89) (envelope-from ) id 1eP404-0003Sa-D9 for pgsql-hackers@postgresql.org; Wed, 13 Dec 2017 10:10:55 +0000 Received: by mail-wm0-x242.google.com with SMTP id g130so21118351wme.0 for ; Wed, 13 Dec 2017 02:10:51 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=2ndquadrant-com.20150623.gappssmtp.com; s=20150623; h=subject:to:cc:references:from:message-id:date:user-agent :mime-version:in-reply-to:content-language:content-transfer-encoding; bh=n4qeS3+qwFufohFvmUkcP17Z1SbJZJreVY7YID2qYtc=; b=TrmczXLpgTPpffK1CBuuhczqSbXTg9mI7id29KFlvqODSuhenzqaNz4u7Z0PDl134B wtkQ7kEC8Tpernt3BGVeoC0JhMk9GwznT74A+m8xHSLKyNdMFnPVo9ymyzIASxsR3AkD Iv2lABQTYsdj9lJnfnZ/+1MnaOONzhEsoc3aGDGnvKi6X7pvnepBokTABkUb1md6yMr8 5etQbzH5m/9kPhbqH+beQcL8WWkl73sN6Q4cbJg9+/gOuAIuzpwUwJCUCsRtwRaf+ORN h0pIeYh1AYcA0jlDh9utkZaR458s6A00SlIWc+iFywJiVsQVj4O6dKJrAmnW2sqCx/q8 KULw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:subject:to:cc:references:from:message-id:date :user-agent:mime-version:in-reply-to:content-language :content-transfer-encoding; bh=n4qeS3+qwFufohFvmUkcP17Z1SbJZJreVY7YID2qYtc=; b=gQSwkMvtvtRc3asbdQVRd3Nax1aUFlZ4Ze+uY/Q2mfcm2QzR+3F14y6vca0fD0j97q cgbxwoFat1WypBz06Im2i99SvIQakzOjDqsjBU01rh4poVEQvjif283K3AJE12z/MvHR XYIev2ZQjkzbAaZrrgqQoHqskaf7vEfCpHp9yA6SkbDKHTothscdNjnL+QAc8eJT3ywF dus2zwZCBQBCp24b3tBphiiMmr/O5gS7QU9be71Le92TXyQwCgW5rMLMyQfzOzBm0NuU lIeQVDbvrnTQRPvNwdTZLhXFoGB2gRYdTyYW3rjp7qRhF+q7ODap8XmRpg0S5P1ReLul 57Cg== X-Gm-Message-State: AKGB3mIqJPQdMqucYP/1BI2qwIfXmVOd0Kmb4tcsABdXEKY49cveV0qH ZTQf1fBHxLaiU3U3/cAU/EjCuO1BiLAyQASS+d2CsBT8xh/q1pEtCjTNLoxHqx2uHU0PoO6lL1B wrB0U/O5aT8/NPhLukPgme1XwuMxXgCzDt81xiu2S0PpCGshhnx88bEiuplGKT3m+iqKDh4CYKF nFHGJwZbl6 X-Google-Smtp-Source: ACJfBouYBS/owPgIhfDZU7WguDWgXqZXR1REqaIMtCu+U2TELdVrmhkeDgGM5x2QGeQX1rGCYtF7NA== X-Received: by 10.80.240.17 with SMTP id r17mr6826921edl.57.1513159850523; Wed, 13 Dec 2017 02:10:50 -0800 (PST) Received: from [10.137.2.19] (ip-78-102-97-226.net.upcbroadband.cz. [78.102.97.226]) by smtp.gmail.com with ESMTPSA id q3sm1035577edd.61.2017.12.13.02.10.49 (version=TLS1_2 cipher=ECDHE-RSA-AES128-GCM-SHA256 bits=128/128); Wed, 13 Dec 2017 02:10:49 -0800 (PST) Subject: Re: [HACKERS] Custom compression methods To: Robert Haas Cc: Ildus Kurbangaliev , "pgsql-hackers@postgresql.org" References: <20170907194236.4cefce96@wp.localdomain> <29527031-c837-59e1-760d-677bb33d6b0f@2ndquadrant.com> From: Tomas Vondra Message-ID: Date: Wed, 13 Dec 2017 11:10:46 +0100 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101 Thunderbird/52.3.0 MIME-Version: 1.0 In-Reply-To: Content-Type: text/plain; charset=utf-8 Content-Language: en-US Content-Transfer-Encoding: 7bit List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Precedence: bulk On 12/13/2017 01:54 AM, Robert Haas wrote: > On Tue, Dec 12, 2017 at 5:07 PM, Tomas Vondra > wrote: >>> I definitely think there's a place for compression built right into >>> the data type. I'm still happy about commit >>> 145343534c153d1e6c3cff1fa1855787684d9a38 -- although really, more >>> needs to be done there. But that type of improvement and what is >>> proposed here are basically orthogonal. Having either one is good; >>> having both is better. >>> >> Why orthogonal? > > I mean, they are different things. Data types are already free to > invent more compact representations, and that does not preclude > applying pglz to the result. > >> For example, why couldn't (or shouldn't) the tsvector compression be >> done by tsvector code itself? Why should we be doing that at the varlena >> level (so that the tsvector code does not even know about it)? > > We could do that, but then: > > 1. The compression algorithm would be hard-coded into the system > rather than changeable. Pluggability has some value. > Sure. I agree extensibility of pretty much all parts is a significant asset of the project. > 2. If several data types can benefit from a similar approach, it has > to be separately implemented for each one. > I don't think the current solution improves that, though. If you want to exploit internal features of individual data types, it pretty much requires code customized to every such data type. For example you can't take the tsvector compression and just slap it on tsquery, because it relies on knowledge of internal tsvector structure. So you need separate implementations anyway. > 3. Compression is only applied to large-ish values. If you are just > making the data type representation more compact, you probably want to > apply the new representation to all values. If you are compressing in > the sense that the original data gets smaller but harder to interpret, > then you probably only want to apply the technique where the value is > already pretty wide, and maybe respect the user's configured storage > attributes. TOAST knows about some of that kind of stuff. > Good point. One such parameter that I really miss is compression level. I can imagine tuning it through CREATE COMPRESSION METHOD, but it does not seem quite possible with compression happening in a datatype. >> It seems to me the main reason is that tsvector actually does not allow >> us to do that, as there's no good way to distinguish the different >> internal format (e.g. by storing a flag or format version in some sort >> of header, etc.). > > That is also a potential problem, although I suspect it is possible to > work around it somehow for most data types. It might be annoying, > though. > >>> I think there may also be a place for declaring that a particular data >>> type has a "privileged" type of TOAST compression; if you use that >>> kind of compression for that data type, the data type will do smart >>> things, and if not, it will have to decompress in more cases. But I >>> think this infrastructure makes that kind of thing easier, not harder. >> >> I don't quite understand how that would be done. Isn't TOAST meant to be >> entirely transparent for the datatypes? I can imagine custom TOAST >> compression (which is pretty much what the patch does, after all), but I >> don't see how the datatype could do anything smart about it, because it >> has no idea which particular compression was used. And considering the >> OIDs of the compression methods do change, I'm not sure that's fixable. > > I don't think TOAST needs to be entirely transparent for the > datatypes. We've already dipped our toe in the water by allowing some > operations on "short" varlenas, and there's really nothing to prevent > a given datatype from going further. The OID problem you mentioned > would presumably be solved by hard-coding the OIDs for any built-in, > privileged compression methods. > Stupid question, but what do you mean by "short" varlenas? regards -- Tomas Vondra http://www.2ndQuadrant.com PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services