agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedFrom: Antonin Houska <ah@cybertec.at>
To: Kirill Reshke <reshkekirill@gmail.com>
Cc: Alvaro Herrera <alvherre@alvh.no-ip.org>
Cc: Pavel Stehule <pavel.stehule@gmail.com>
Cc: Michael Paquier <michael@paquier.xyz>
Cc: PostgreSQL Hackers <pgsql-hackers@postgresql.org>
Subject: Re: why there is not VACUUM FULL CONCURRENTLY?
Date: Fri, 02 Aug 2024 08:09:26 +0200
Message-ID: <1788.1722578966@antos> (raw)
In-Reply-To: <CALdSSPig+9SVK34EN33v6-iFh17FFNLxW0cpHToX=miNLDqodg@mail.gmail.com>
References: <202401311004.2yky72qydzxn@alvherre.pgsql>
<82651.1720540558@antos>
<CALdSSPig+9SVK34EN33v6-iFh17FFNLxW0cpHToX=miNLDqodg@mail.gmail.com>
Kirill Reshke <reshkekirill@gmail.com> wrote:
> What is the size of the biggest relation successfully vacuumed
> via pg_squeeze?
> Looks like in case of big relartion or high insertion load,
> replication may lag and never catch up...
Users reports problems rather than successes, so I don't know. 400 GB was
reported in [1] but it's possible that the table size for this test was
determined based on available disk space.
I think that the amount of data changes performed during the "squeezing"
matters more than the table size. In [2] one user reported "thounsands of
UPSERTs per second", but the amount of data also depends on row size, which he
didn't mention.
pg_squeeze gives up if it fails to catch up a few times. The first version of
my patch does not check this, I'll add the corresponding code in the next
version.
> However, in general, the 3rd patch is really big, very hard to
> comprehend. Please consider splitting this into smaller (and
> reviewable) pieces.
I'll try to move some preparation steps into separate diffs, but not sure if
that will make the main diff much smaller. I prefer self-contained patches, as
also explained in [3].
> Also, we obviously need more tests on this. Both tap-test and
> regression tests I suppose.
Sure. The next version will use the injection points to test if "concurrent
data changes" are processed correctly.
> One more thing is about pg_squeeze background workers. They act in an
> autovacuum-like fashion, aren't they? Maybe we can support this kind
> of relation processing in core too?
Maybe later. Even just adding the CONCURRENTLY option to CLUSTER and VACUUM
FULL requires quite some effort.
[1] https://github.com/cybertec-postgresql/pg_squeeze/issues/51
[2]
https://github.com/cybertec-postgresql/pg_squeeze/issues/21#issuecomment-514495369
[3] http://peter.eisentraut.org/blog/2024/05/14/when-to-split-patches-for-postgresql
--
Antonin Houska
Web: https://www.cybertec-postgresql.com
view thread (88+ messages) latest in thread
Message-ID: <1788.1722578966@antos>
Permalink: ../1788.1722578966@antos/
Also on: postgresql.org/message-id/1788.1722578966@antos
reply
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-hackers@postgresql.org
Cc: ah@cybertec.at, reshkekirill@gmail.com, alvherre@alvh.no-ip.org, pavel.stehule@gmail.com, michael@paquier.xyz
Subject: Re: why there is not VACUUM FULL CONCURRENTLY?
In-Reply-To: <1788.1722578966@antos>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox