From: Alexandr Popov <a.popov@postgrespro.ru>
To: Anastasia Lubennikova <a.lubennikova@postgrespro.ru>
To: David Steele <david@pgmasters.net>
To: pgsql-hackers@postgresql.org
Subject: Re: [WIP] Effective storage of duplicates in B-tree index.
Date: Wed, 23 Mar 2016 18:30:14 +0300
Message-ID: <56F2B686.9070602@postgrespro.ru> (raw)
In-Reply-To: <56EC38A9.9030303@postgrespro.ru>
References: <55E4051B.7020209@postgrespro.ru>
<56AA2081.1080001@postgrespro.ru>
<CAA-aLv68Ptrnk2HpHSg9wYR0c+D2AyTB+XrS671brxqQz3dH3w@mail.gmail.com>
<56AA3E06.8040006@postgrespro.ru>
<CAA-aLv5aQezrfh2iJD97R1sQ6nqMsR2QZ_YAzNfAOkkMgO8CFQ@mail.gmail.com>
<56AB6D30.2040900@postgrespro.ru>
<20160129184733.2ca9026a@fujitsu>
<CAA-aLv7pmO41G78jfLsftsCvkXZw6dL8HZoCX9ute6nQxJv2tw@mail.gmail.com>
<56AB9866.6050207@postgrespro.ru>
<CAM3SWZQ3_PLQCH4w7uQ8q_f2t4HEseKTr2n0rQ5pxA18OeRTJw@mail.gmail.com>
<56C5FCE1.1090509@postgrespro.ru>
<56C5FF80.5050905@postgrespro.ru>
<56E6B64E.6000101@pgmasters.net>
<56E84834.2070003@postgrespro.ru>
<56EC38A9.9030303@postgrespro.ru>
List-Unsubscribe: <mailto:majordomo@postgresql.org?body=unsub%20pgsql-hackers>
On 18.03.2016 20:19, Anastasia Lubennikova wrote:
> Please, find the new version of the patch attached. Now it has WAL
> functionality.
>
> Detailed description of the feature you can find in README draft
> https://goo.gl/50O8Q0
>
> This patch is pretty complicated, so I ask everyone, who interested in
> this feature,
> to help with reviewing and testing it. I will be grateful for any
> feedback.
> But please, don't complain about code style, it is still work in
> progress.
>
> Next things I'm going to do:
> 1. More debugging and testing. I'm going to attach in next message
> couple of sql scripts for testing.
> 2. Fix NULLs processing
> 3. Add a flag into pg_index, that allows to enable/disable compression
> for each particular index.
> 4. Recheck locking considerations. I tried to write code as less
> invasive as possible, but we need to make sure that algorithm is still
> correct.
> 5. Change BTMaxItemSize
> 6. Bring back microvacuum functionality.
>
Hi, hackers.
It's my first review, so do not be strict to me.
I have tested this patch on the next table:
create table message
(
id serial,
usr_id integer,
text text
);
CREATE INDEX message_usr_id ON message (usr_id);
The table has 10000000 records.
I found the following:
The less unique keys the less size of the table.
Next 2 tablas demonstrates it.
New B-tree
Count of unique keys (usr_id), index“s size , time of creation
10000000 ;"214 MB" ;"00:00:34.193441"
3333333 ;"214 MB" ;"00:00:45.731173"
2000000 ;"129 MB" ;"00:00:41.445876"
1000000 ;"129 MB" ;"00:00:38.455616"
100000 ;"86 MB" ;"00:00:40.887626"
10000 ;"79 MB" ;"00:00:47.199774"
Old B-tree
Count of unique keys (usr_id), index“s size , time of creation
10000000 ;"214 MB" ;"00:00:35.043677"
3333333 ;"286 MB" ;"00:00:40.922845"
2000000 ;"300 MB" ;"00:00:46.454846"
1000000 ;"278 MB" ;"00:00:42.323525"
100000 ;"287 MB" ;"00:00:47.438132"
10000 ;"280 MB" ;"00:01:00.307873"
I inserted data randomly and sequentially, it did not influence the
index's size.
Time of select, insert and update random rows is not changed. It is
great, but certainly it needs some more detailed study.
Alexander Popov
Postgres Professional: http://www.postgrespro.com
The Russian Postgres Company
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-hackers@postgresql.org
Cc: a.popov@postgrespro.ru, a.lubennikova@postgrespro.ru, david@pgmasters.net
Subject: Re: [WIP] Effective storage of duplicates in B-tree index.
In-Reply-To: <56F2B686.9070602@postgrespro.ru>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox