agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedFrom: Antonin Houska <ah@cybertec.at>
To: Marcos Pegoraro <marcos@f10.com.br>
Cc: Alvaro Herrera <alvherre@alvh.no-ip.org>
Cc: Michael Banck <mbanck@gmx.net>
Cc: Junwang Zhao <zhjwpku@gmail.com>
Cc: Kirill Reshke <reshkekirill@gmail.com>
Cc: Pavel Stehule <pavel.stehule@gmail.com>
Cc: Michael Paquier <michael@paquier.xyz>
Cc: PostgreSQL Hackers <pgsql-hackers@postgresql.org>
Subject: Re: why there is not VACUUM FULL CONCURRENTLY?
Date: Wed, 26 Feb 2025 09:48:08 +0100
Message-ID: <127361.1740559688@localhost> (raw)
In-Reply-To: <CAB-JLwYSy+oquVkiwubhqs-rHfjvF6p5PRmQHRg4_1vtaL_E7A@mail.gmail.com>
References: <202501301529.ejggbtao2skr@alvherre.pgsql>
<63871.1739984882@localhost>
<CAB-JLwYSy+oquVkiwubhqs-rHfjvF6p5PRmQHRg4_1vtaL_E7A@mail.gmail.com>
Marcos Pegoraro <marcos@f10.com.br> wrote:
> You mentioned fillfactor only on cluster notes, would be good to mention it
> on refsynopsisdiv, I think.
ok, I've added a note to the first paragraph.
Attached here is the REPACK command as well as the patch set that adds the
CONCURRENTLY option. The new symbols have been renamed so they resemble REPACK
rather than CLUSTER.
Please note that 0008 is a new part which makes the setting wal_leve=logical
unnecessary.
--
Antonin Houska
Web: https://www.cybertec-postgresql.com
From 8da24b040720e5a30517076825760c7516c42bfa Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 26 Feb 2025 09:17:20 +0100
Subject: [PATCH 4/9] Add CONCURRENTLY option to REPACK command.
The REPACK command copies the relation data into a new file, creates new
indexes and eventually swaps the files. To make sure that the old file does
not change during the copying, the relation is locked in an exclusive mode,
which prevents applications from both reading and writing. (To keep the data
consistent, we'd only need to prevent the applications from writing, but even
reading needs to be blocked before we can swap the files - otherwise some
applications could continue using the old file. Since we cannot get stronger
lock without releasing the weaker one first, we acquire the exclusive lock in
the beginning and keep it till the end of the processing.)
This patch introduces an alternative workflow, which only requires the
exclusive lock when the relation (and index) files are being swapped.
(Supposedly, the swapping should be pretty fast.) On the other hand, when we
copy the data to the new file, we allow applications to read from the relation
and even write into it.
First, we scan the relation using a "historic snapshot", and insert all the
tuples satisfying this snapshot into the new file. Note that, before creating
that snapshot, we need to make sure that all the other backends treat the
relation as a system catalog: in particular, they must log information on new
command IDs (CIDs). We achieve that by adding the relation ID into a shared
hash table and waiting until all the transactions currently writing into the
table (i.e. transactions possibly not aware of the new entry) have finished.
Second, logical decoding is used to capture the data changes done by
applications during the copying (i.e. changes that do not satisfy the historic
snapshot mentioned above), and those are applied to the new file before we
acquire the exclusive lock we need to swap the files. (Of course, more data
changes can take place while we are waiting for the lock - these will be
applied to the new file after we have acquired the lock, before we swap the
files.)
While copying the data into the new file, we hold a lock that prevents
applications from changing the relation tuple descriptor (tuples inserted into
the old file must fit into the new file). However, as we have to release that
lock before getting the exclusive one, it's possible that someone adds or
drops a column, or changes the data type of an existing one. Therefore we have
to check the tuple descriptor before we swap the files. If we find out that
the tuple descriptor changed, ERROR is raised and all the changes are rolled
back. Since a lot of effort can be wasted in such a case, the ALTER TABLE
command also tries to check if REPACK CONCURRENTLY is running on the same
relation, and raises an ERROR if it is.
Like the existing implementation of REPACK, the variant with the CONCURRENTLY
option also requires an extra space for the new relation and index files
(which coexist with the old files for some time). In addition, the
CONCURRENTLY option might introduce a lag in releasing WAL segments for
archiving / recycling. This is due to the decoding of the data changes done by
application concurrently. However, this lag should not be more than a single
WAL segment.
---
doc/src/sgml/monitoring.sgml | 65 +-
doc/src/sgml/ref/repack.sgml | 116 +-
src/Makefile | 1 +
src/backend/access/heap/heapam.c | 8 +-
src/backend/access/heap/heapam_handler.c | 145 +-
src/backend/access/heap/heapam_visibility.c | 30 +-
src/backend/catalog/index.c | 43 +-
src/backend/catalog/system_views.sql | 30 +-
src/backend/commands/cluster.c | 2667 ++++++++++++++++-
src/backend/commands/matview.c | 2 +-
src/backend/commands/tablecmds.c | 11 +
src/backend/commands/vacuum.c | 12 +-
src/backend/meson.build | 1 +
src/backend/parser/gram.y | 17 +-
src/backend/replication/logical/decode.c | 24 +
src/backend/replication/logical/snapbuild.c | 20 +
.../replication/pgoutput_repack/Makefile | 32 +
.../replication/pgoutput_repack/meson.build | 18 +
.../pgoutput_repack/pgoutput_repack.c | 286 ++
src/backend/storage/ipc/ipci.c | 3 +
src/backend/tcop/utility.c | 10 +
src/backend/utils/activity/backend_progress.c | 16 +
.../utils/activity/wait_event_names.txt | 1 +
src/backend/utils/cache/inval.c | 21 +
src/backend/utils/cache/relcache.c | 5 +
src/backend/utils/time/snapmgr.c | 3 +-
src/bin/psql/tab-complete.in.c | 24 +-
src/include/access/heapam.h | 4 +
src/include/access/tableam.h | 10 +
src/include/catalog/index.h | 3 +
src/include/commands/cluster.h | 93 +-
src/include/commands/progress.h | 17 +-
src/include/nodes/parsenodes.h | 1 +
src/include/replication/snapbuild.h | 1 +
src/include/storage/lockdefs.h | 5 +-
src/include/storage/lwlocklist.h | 1 +
src/include/utils/backend_progress.h | 3 +-
src/include/utils/inval.h | 2 +
src/include/utils/rel.h | 7 +-
src/include/utils/snapmgr.h | 2 +
src/test/regress/expected/rules.out | 29 +-
41 files changed, 3536 insertions(+), 253 deletions(-)
create mode 100644 src/backend/replication/pgoutput_repack/Makefile
create mode 100644 src/backend/replication/pgoutput_repack/meson.build
create mode 100644 src/backend/replication/pgoutput_repack/pgoutput_repack.c
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 58e1becf02..8d73c01c55 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -5780,14 +5780,35 @@ FROM pg_stat_get_backend_idset() AS backendid;
<row>
<entry role="catalog_table_entry"><para role="column_definition">
- <structfield>heap_tuples_written</structfield> <type>bigint</type>
+ <structfield>heap_tuples_inserted</structfield> <type>bigint</type>
</para>
<para>
- Number of heap tuples written.
+ Number of heap tuples inserted.
This counter only advances when the phase is
<literal>seq scanning heap</literal>,
- <literal>index scanning heap</literal>
- or <literal>writing new heap</literal>.
+ <literal>index scanning heap</literal>,
+ <literal>writing new heap</literal>
+ or <literal>catch-up</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_tuples_updated</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap tuples updated.
+ This counter only advances when the phase is <literal>catch-up</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_tuples_deleted</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap tuples deleted.
+ This counter only advances when the phase is <literal>catch-up</literal>.
</para></entry>
</row>
@@ -6003,14 +6024,35 @@ FROM pg_stat_get_backend_idset() AS backendid;
<row>
<entry role="catalog_table_entry"><para role="column_definition">
- <structfield>heap_tuples_written</structfield> <type>bigint</type>
+ <structfield>heap_tuples_inserted</structfield> <type>bigint</type>
</para>
<para>
- Number of heap tuples written.
+ Number of heap tuples inserted.
This counter only advances when the phase is
<literal>seq scanning heap</literal>,
- <literal>index scanning heap</literal>
- or <literal>writing new heap</literal>.
+ <literal>index scanning heap</literal>,
+ <literal>writing new heap</literal>
+ or <literal>catch-up</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_tuples_updated</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap tuples updated.
+ This counter only advances when the phase is <literal>catch-up</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_tuples_deleted</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap tuples deleted.
+ This counter only advances when the phase is <literal>catch-up</literal>.
</para></entry>
</row>
@@ -6091,6 +6133,13 @@ FROM pg_stat_get_backend_idset() AS backendid;
<command>REPACK</command> is currently writing the new heap.
</entry>
</row>
+ <row>
+ <entry><literal>catch-up</literal></entry>
+ <entry>
+ <command>REPACK</command> is currently processing the DML commands that
+ other transactions executed during any of the preceding phase.
+ </entry>
+ </row>
<row>
<entry><literal>swapping relation files</literal></entry>
<entry>
diff --git a/doc/src/sgml/ref/repack.sgml b/doc/src/sgml/ref/repack.sgml
index 84f3c3e3f2..9ee640e351 100644
--- a/doc/src/sgml/ref/repack.sgml
+++ b/doc/src/sgml/ref/repack.sgml
@@ -22,6 +22,7 @@ PostgreSQL documentation
<refsynopsisdiv>
<synopsis>
REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <replaceable class="parameter">table_name</replaceable> [ USING INDEX<replaceable class="parameter">index_name</replaceable> ] ]
+REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] CONCURRENTLY <replaceable class="parameter">table_name</replaceable> [ USING INDEX<replaceable class="parameter">index_name</replaceable> ]
<phrase>where <replaceable class="parameter">option</replaceable> can be one of:</phrase>
@@ -48,7 +49,8 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <re
processes every table and materialized view in the current database that
the current user has the <literal>MAINTAIN</literal> privilege on. This
form of <command>REPACK</command> cannot be executed inside a transaction
- block.
+ block. Also, this form is not allowed if
+ the <literal>CONCURRENTLY</literal> option is used.
</para>
<para>
@@ -61,7 +63,8 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <re
When a table is being repacked, an <literal>ACCESS EXCLUSIVE</literal> lock
is acquired on it. This prevents any other database operations (both reads
and writes) from operating on the table until the <command>REPACK</command>
- is finished.
+ is finished. If you want to keep the table accessible during the repacking,
+ consider using the <literal>CONCURRENTLY</literal> option.
</para>
<refsect2 id="sql-repack-notes-on-clustering" xreflabel="Notes on Clustering">
@@ -160,6 +163,115 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <re
</listitem>
</varlistentry>
+ <varlistentry>
+ <term><literal>CONCURRENTLY</literal></term>
+ <listitem>
+ <para>
+ Allow other transactions to use the table while it is being repacked.
+ </para>
+
+ <para>
+ Internally, <command>REPACK</command> copies the contents of the table
+ (ignoring dead tuples) into a new file, sorted by the specified index,
+ and also creates a new file for each index. Then it swaps the old and
+ new files for the table and all the indexes, and deletes the old
+ files. The <literal>ACCESS EXCLUSIVE</literal> lock is needed to make
+ sure that the old files do not change during the processing because the
+ changes would get lost due to the swap.
+ </para>
+
+ <para>
+ With the <literal>CONCURRENTLY</literal> option, the <literal>ACCESS
+ EXCLUSIVE</literal> lock is only acquired to swap the table and index
+ files. The data changes that took place during the creation of the new
+ table and index files are captured using logical decoding
+ (<xref linkend="logicaldecoding"/>) and applied before
+ the <literal>ACCESS EXCLUSIVE</literal> lock is requested. Thus the lock
+ is typically held only for the time needed to swap the files, which
+ should be pretty short.
+ </para>
+
+ <para>
+ Note that <command>REPACK</command> with the
+ the <literal>CONCURRENTLY</literal> option does not try to order the
+ rows inserted into the table after the repacking started. Also
+ note <command>REPACK</command> might fail to complete due to DDL
+ commands executed on the table by other transactions during the
+ repacking.
+ </para>
+
+ <note>
+ <para>
+ In addition to the temporary space requirements explained in
+ <xref linkend="sql-repack-notes-on-resources"/>,
+ the <literal>CONCURRENTLY</literal> option can add to the usage of
+ temporary space a bit more. The reason is that other transactions can
+ perform DML operations which cannot be applied to the new file until
+ <command>REPACK</command> has copied all the tuples from the old
+ file. Thus the tuples inserted into the old file during the copying are
+ also stored in separately in a temporary file, so they can eventually
+ be applied to the new file.
+ </para>
+
+ <para>
+ Furthermore, the data changes performed during the copying are
+ extracted from <link linkend="wal">write-ahead log</link> (WAL), and
+ this extraction (decoding) only takes place when certain amount of WAL
+ has been written. Therefore, WAL removal can be delayed by this
+ threshold. Currently the threshold is equal to the value of
+ the <link linkend="guc-wal-segment-size"><varname>wal_segment_size</varname></link>
+ configuration parameter.
+ </para>
+ </note>
+
+ <para>
+ The <literal>CONCURRENTLY</literal> option cannot be used in the
+ following cases:
+
+ <itemizedlist>
+ <listitem>
+ <para>
+ The table is <literal>UNLOGGED</literal>.
+ </para>
+ </listitem>
+
+ <listitem>
+ <para>
+ The table is partitioned.
+ </para>
+ </listitem>
+
+ <listitem>
+ <para>
+ The table is a system catalog or a <acronym>TOAST</acronym> table.
+ </para>
+ </listitem>
+
+ <listitem>
+ <para>
+ <command>REPACK</command> is executed inside a transaction block.
+ </para>
+ </listitem>
+
+ <listitem>
+ <para>
+ The <link linkend="guc-wal-level"><varname>wal_level</varname></link>
+ configuration parameter is less than <literal>logical</literal>.
+ </para>
+ </listitem>
+
+ <listitem>
+ <para>
+ The <link linkend="guc-max-replication-slots"><varname>max_replication_slots</varname></link>
+ configuration parameter does not allow for creation of an additional
+ replication slot.
+ </para>
+ </listitem>
+ </itemizedlist>
+ </para>
+ </listitem>
+ </varlistentry>
+
<varlistentry>
<term><literal>VERBOSE</literal></term>
<listitem>
diff --git a/src/Makefile b/src/Makefile
index 2f31a2f20a..b18c9a14ff 100644
--- a/src/Makefile
+++ b/src/Makefile
@@ -23,6 +23,7 @@ SUBDIRS = \
interfaces \
backend/replication/libpqwalreceiver \
backend/replication/pgoutput \
+ backend/replication/pgoutput_repack \
fe_utils \
bin \
pl \
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index fa7935a0ed..cb856a74ee 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -2093,8 +2093,14 @@ heap_insert(Relation relation, HeapTuple tup, CommandId cid,
/*
* If this is a catalog, we need to transmit combo CIDs to properly
* decode, so log that as well.
+ *
+ * For the main heap (as opposed to TOAST), we only receive
+ * HEAP_INSERT_NO_LOGICAL when doing REPACK CONCURRENTLY, in which
+ * case the visibility information does not change. Therefore, there's
+ * no need to update the decoding snapshot.
*/
- if (RelationIsAccessibleInLogicalDecoding(relation))
+ if ((options & HEAP_INSERT_NO_LOGICAL) == 0 &&
+ RelationIsAccessibleInLogicalDecoding(relation))
log_heap_new_cid(relation, heaptup);
/*
diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index 5c3cab8bc2..b2bfd05dc9 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -33,6 +33,7 @@
#include "catalog/index.h"
#include "catalog/storage.h"
#include "catalog/storage_xlog.h"
+#include "commands/cluster.h"
#include "commands/progress.h"
#include "executor/executor.h"
#include "miscadmin.h"
@@ -53,6 +54,9 @@ static void reform_and_rewrite_tuple(HeapTuple tuple,
static bool SampleHeapTupleVisible(TableScanDesc scan, Buffer buffer,
HeapTuple tuple,
OffsetNumber tupoffset);
+static HeapTuple accept_tuple_for_concurrent_copy(HeapTuple tuple,
+ Snapshot snapshot,
+ Buffer buffer);
static BlockNumber heapam_scan_get_blocks_done(HeapScanDesc hscan);
@@ -681,6 +685,8 @@ static void
heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
Relation OldIndex, bool use_sort,
TransactionId OldestXmin,
+ Snapshot snapshot,
+ LogicalDecodingContext *decoding_ctx,
TransactionId *xid_cutoff,
MultiXactId *multi_cutoff,
double *num_tuples,
@@ -701,6 +707,8 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
bool *isnull;
BufferHeapTupleTableSlot *hslot;
BlockNumber prev_cblock = InvalidBlockNumber;
+ bool concurrent = snapshot != NULL;
+ XLogRecPtr end_of_wal_prev = GetFlushRecPtr(NULL);
/* Remember if it's a system catalog */
is_system_catalog = IsSystemRelation(OldHeap);
@@ -779,8 +787,10 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
for (;;)
{
HeapTuple tuple;
+ bool tuple_copied = false;
Buffer buf;
bool isdead;
+ HTSV_Result vis;
CHECK_FOR_INTERRUPTS();
@@ -835,7 +845,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
LockBuffer(buf, BUFFER_LOCK_SHARE);
- switch (HeapTupleSatisfiesVacuum(tuple, OldestXmin, buf))
+ switch ((vis = HeapTupleSatisfiesVacuum(tuple, OldestXmin, buf)))
{
case HEAPTUPLE_DEAD:
/* Definitely dead */
@@ -851,14 +861,15 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
case HEAPTUPLE_INSERT_IN_PROGRESS:
/*
- * Since we hold exclusive lock on the relation, normally the
- * only way to see this is if it was inserted earlier in our
- * own transaction. However, it can happen in system
+ * As long as we hold exclusive lock on the relation, normally
+ * the only way to see this is if it was inserted earlier in
+ * our own transaction. However, it can happen in system
* catalogs, since we tend to release write lock before commit
- * there. Give a warning if neither case applies; but in any
- * case we had better copy it.
+ * there. Also, there's no exclusive lock during concurrent
+ * processing. Give a warning if neither case applies; but in
+ * any case we had better copy it.
*/
- if (!is_system_catalog &&
+ if (!is_system_catalog && !concurrent &&
!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetXmin(tuple->t_data)))
elog(WARNING, "concurrent insert in progress within table \"%s\"",
RelationGetRelationName(OldHeap));
@@ -870,7 +881,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
/*
* Similar situation to INSERT_IN_PROGRESS case.
*/
- if (!is_system_catalog &&
+ if (!is_system_catalog && !concurrent &&
!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetUpdateXid(tuple->t_data)))
elog(WARNING, "concurrent delete in progress within table \"%s\"",
RelationGetRelationName(OldHeap));
@@ -884,8 +895,6 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
break;
}
- LockBuffer(buf, BUFFER_LOCK_UNLOCK);
-
if (isdead)
{
*tups_vacuumed += 1;
@@ -896,9 +905,47 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
*tups_vacuumed += 1;
*tups_recently_dead -= 1;
}
+
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
continue;
}
+ if (concurrent)
+ {
+ /*
+ * Ignore concurrent changes now, they'll be processed later via
+ * logical decoding.
+ *
+ * INSERT_IN_PROGRESS is rejected right away because our snapshot
+ * should represent a point in time which should precede (or be
+ * equal to) the state of transactions as it was when the
+ * "SatisfiesVacuum" test was performed. Thus
+ * accept_tuple_for_concurrent_copy() should not consider the
+ * tuple inserted.
+ */
+ if (vis == HEAPTUPLE_INSERT_IN_PROGRESS)
+ tuple = NULL;
+ else
+ tuple = accept_tuple_for_concurrent_copy(tuple, snapshot,
+ buf);
+ /* Tuple not suitable for the new heap? */
+ if (tuple == NULL)
+ {
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+ continue;
+ }
+
+ /* Remember that we have to free the tuple eventually. */
+ tuple_copied = true;
+ }
+
+ /*
+ * In the concurrent case, we have a copy of the tuple, so we don't
+ * worry whether the source tuple will be deleted / updated after we
+ * release the lock.
+ */
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+
*num_tuples += 1;
if (tuplesort != NULL)
{
@@ -915,7 +962,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
{
const int ct_index[] = {
PROGRESS_REPACK_HEAP_TUPLES_SCANNED,
- PROGRESS_REPACK_HEAP_TUPLES_WRITTEN
+ PROGRESS_REPACK_HEAP_TUPLES_INSERTED
};
int64 ct_val[2];
@@ -930,6 +977,33 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
ct_val[1] = *num_tuples;
pgstat_progress_update_multi_param(2, ct_index, ct_val);
}
+ if (tuple_copied)
+ heap_freetuple(tuple);
+
+ /*
+ * Process the WAL produced by the load, as well as by other
+ * transactions, so that the replication slot can advance and WAL does
+ * not pile up. Use wal_segment_size as a threshold so that we do not
+ * introduce the decoding overhead too often.
+ *
+ * Of course, we must not apply the changes until the initial load has
+ * completed.
+ *
+ * Note that our insertions into the new table should not be decoded
+ * as we (intentionally) do not write the logical decoding specific
+ * information to WAL.
+ */
+ if (concurrent)
+ {
+ XLogRecPtr end_of_wal;
+
+ end_of_wal = GetFlushRecPtr(NULL);
+ if ((end_of_wal - end_of_wal_prev) > wal_segment_size)
+ {
+ repack_decode_concurrent_changes(decoding_ctx, end_of_wal);
+ end_of_wal_prev = end_of_wal;
+ }
+ }
}
if (indexScan != NULL)
@@ -973,7 +1047,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
values, isnull,
rwstate);
/* Report n_tuples */
- pgstat_progress_update_param(PROGRESS_REPACK_HEAP_TUPLES_WRITTEN,
+ pgstat_progress_update_param(PROGRESS_REPACK_HEAP_TUPLES_INSERTED,
n_tuples);
}
@@ -2626,6 +2700,53 @@ SampleHeapTupleVisible(TableScanDesc scan, Buffer buffer,
}
}
+/*
+ * Return copy of 'tuple' if it has been inserted according to 'snapshot', or
+ * NULL if the insertion took place in the future. If the tuple is already
+ * marked as deleted or updated by a transaction that 'snapshot' still
+ * considers running, clear the deletion / update XID in the header of the
+ * copied tuple. This way the returned tuple is suitable for insertion into
+ * the new heap.
+ */
+static HeapTuple
+accept_tuple_for_concurrent_copy(HeapTuple tuple, Snapshot snapshot,
+ Buffer buffer)
+{
+ HeapTuple result;
+
+ Assert(snapshot->snapshot_type == SNAPSHOT_MVCC);
+
+ /*
+ * First, check if the tuple insertion is visible by our snapshot.
+ */
+ if (!HeapTupleMVCCInserted(tuple, snapshot, buffer))
+ return NULL;
+
+ result = heap_copytuple(tuple);
+
+ /*
+ * If the tuple was deleted / updated but our snapshot still sees it, we
+ * need to keep it. In that case, clear the information that indicates the
+ * deletion / update. Otherwise the tuple chain would stay incomplete (as
+ * we will reject the new tuple above), and the delete / update would fail
+ * if executed later during logical decoding.
+ */
+ if (TransactionIdIsNormal(HeapTupleHeaderGetRawXmax(result->t_data)) &&
+ HeapTupleMVCCNotDeleted(result, snapshot, buffer))
+ {
+ /* TODO More work needed here?*/
+ result->t_data->t_infomask |= HEAP_XMAX_INVALID;
+ HeapTupleHeaderSetXmax(result->t_data, 0);
+ }
+
+ /*
+ * Accept the tuple even if our snapshot considers it deleted - older
+ * snapshots can still see the tuple, while the decoded transactions
+ * should not try to update / delete it again.
+ */
+ return result;
+}
+
/* ------------------------------------------------------------------------
* Definition of the heap table access method.
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index e146605bd5..d9be93aadc 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -955,16 +955,31 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
* did TransactionIdIsInProgress in each call --- to no avail, as long as the
* inserting/deleting transaction was still running --- which was more cycles
* and more contention on ProcArrayLock.
+ *
+ * The checks are split into two functions, HeapTupleMVCCInserted() and
+ * HeapTupleMVCCNotDeleted(), because they are also useful separately.
*/
static bool
HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
Buffer buffer)
{
- HeapTupleHeader tuple = htup->t_data;
-
Assert(ItemPointerIsValid(&htup->t_self));
Assert(htup->t_tableOid != InvalidOid);
+ return HeapTupleMVCCInserted(htup, snapshot, buffer) &&
+ HeapTupleMVCCNotDeleted(htup, snapshot, buffer);
+}
+
+/*
+ * HeapTupleMVCCInserted
+ * True iff heap tuple was successfully inserted for the given MVCC
+ * snapshot.
+ */
+bool
+HeapTupleMVCCInserted(HeapTuple htup, Snapshot snapshot, Buffer buffer)
+{
+ HeapTupleHeader tuple = htup->t_data;
+
if (!HeapTupleHeaderXminCommitted(tuple))
{
if (HeapTupleHeaderXminInvalid(tuple))
@@ -1073,6 +1088,17 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
}
/* by here, the inserting transaction has committed */
+ return true;
+}
+
+/*
+ * HeapTupleMVCCNotDeleted
+ * True iff heap tuple was not deleted for the given MVCC snapshot.
+ */
+bool
+HeapTupleMVCCNotDeleted(HeapTuple htup, Snapshot snapshot, Buffer buffer)
+{
+ HeapTupleHeader tuple = htup->t_data;
if (tuple->t_infomask & HEAP_XMAX_INVALID) /* xid invalid or aborted */
return true;
diff --git a/src/backend/catalog/index.c b/src/backend/catalog/index.c
index c84f67059a..39b121c0b8 100644
--- a/src/backend/catalog/index.c
+++ b/src/backend/catalog/index.c
@@ -1417,22 +1417,7 @@ index_concurrently_create_copy(Relation heapRelation, Oid oldIndexId,
for (int i = 0; i < newInfo->ii_NumIndexAttrs; i++)
opclassOptions[i] = get_attoptions(oldIndexId, i + 1);
- /* Extract statistic targets for each attribute */
- stattargets = palloc0_array(NullableDatum, newInfo->ii_NumIndexAttrs);
- for (int i = 0; i < newInfo->ii_NumIndexAttrs; i++)
- {
- HeapTuple tp;
- Datum dat;
-
- tp = SearchSysCache2(ATTNUM, ObjectIdGetDatum(oldIndexId), Int16GetDatum(i + 1));
- if (!HeapTupleIsValid(tp))
- elog(ERROR, "cache lookup failed for attribute %d of relation %u",
- i + 1, oldIndexId);
- dat = SysCacheGetAttr(ATTNUM, tp, Anum_pg_attribute_attstattarget, &isnull);
- ReleaseSysCache(tp);
- stattargets[i].value = dat;
- stattargets[i].isnull = isnull;
- }
+ stattargets = get_index_stattargets(oldIndexId, newInfo);
/*
* Now create the new index.
@@ -1471,6 +1456,32 @@ index_concurrently_create_copy(Relation heapRelation, Oid oldIndexId,
return newIndexId;
}
+NullableDatum *
+get_index_stattargets(Oid indexid, IndexInfo *indInfo)
+{
+ NullableDatum *stattargets;
+
+ /* Extract statistic targets for each attribute */
+ stattargets = palloc0_array(NullableDatum, indInfo->ii_NumIndexAttrs);
+ for (int i = 0; i < indInfo->ii_NumIndexAttrs; i++)
+ {
+ HeapTuple tp;
+ Datum dat;
+ bool isnull;
+
+ tp = SearchSysCache2(ATTNUM, ObjectIdGetDatum(indexid), Int16GetDatum(i + 1));
+ if (!HeapTupleIsValid(tp))
+ elog(ERROR, "cache lookup failed for attribute %d of relation %u",
+ i + 1, indexid);
+ dat = SysCacheGetAttr(ATTNUM, tp, Anum_pg_attribute_attstattarget, &isnull);
+ ReleaseSysCache(tp);
+ stattargets[i].value = dat;
+ stattargets[i].isnull = isnull;
+ }
+
+ return stattargets;
+}
+
/*
* index_concurrently_build
*
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index b8209b2acd..c301d83d9b 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1249,16 +1249,17 @@ CREATE VIEW pg_stat_progress_cluster AS
WHEN 2 THEN 'index scanning heap'
WHEN 3 THEN 'sorting tuples'
WHEN 4 THEN 'writing new heap'
- WHEN 5 THEN 'swapping relation files'
- WHEN 6 THEN 'rebuilding index'
- WHEN 7 THEN 'performing final cleanup'
+ -- 5 is 'catch-up', but that should not appear here.
+ WHEN 6 THEN 'swapping relation files'
+ WHEN 7 THEN 'rebuilding index'
+ WHEN 8 THEN 'performing final cleanup'
END AS phase,
CAST(S.param3 AS oid) AS cluster_index_relid,
S.param4 AS heap_tuples_scanned,
S.param5 AS heap_tuples_written,
- S.param6 AS heap_blks_total,
- S.param7 AS heap_blks_scanned,
- S.param8 AS index_rebuild_count
+ S.param8 AS heap_blks_total,
+ S.param9 AS heap_blks_scanned,
+ S.param10 AS index_rebuild_count
FROM pg_stat_get_progress_info('CLUSTER') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
@@ -1275,16 +1276,19 @@ CREATE VIEW pg_stat_progress_repack AS
WHEN 2 THEN 'index scanning heap'
WHEN 3 THEN 'sorting tuples'
WHEN 4 THEN 'writing new heap'
- WHEN 5 THEN 'swapping relation files'
- WHEN 6 THEN 'rebuilding index'
- WHEN 7 THEN 'performing final cleanup'
+ WHEN 5 THEN 'catch-up'
+ WHEN 6 THEN 'swapping relation files'
+ WHEN 7 THEN 'rebuilding index'
+ WHEN 8 THEN 'performing final cleanup'
END AS phase,
CAST(S.param3 AS oid) AS repack_index_relid,
S.param4 AS heap_tuples_scanned,
- S.param5 AS heap_tuples_written,
- S.param6 AS heap_blks_total,
- S.param7 AS heap_blks_scanned,
- S.param8 AS index_rebuild_count
+ S.param5 AS heap_tuples_inserted,
+ S.param6 AS heap_tuples_updated,
+ S.param7 AS heap_tuples_deleted,
+ S.param8 AS heap_blks_total,
+ S.param9 AS heap_blks_scanned,
+ S.param10 AS index_rebuild_count
FROM pg_stat_get_progress_info('REPACK') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
diff --git a/src/backend/commands/cluster.c b/src/backend/commands/cluster.c
index d0f2588a97..592ff6041b 100644
--- a/src/backend/commands/cluster.c
+++ b/src/backend/commands/cluster.c
@@ -25,6 +25,10 @@
#include "access/toast_internals.h"
#include "access/transam.h"
#include "access/xact.h"
+#include "access/xlog.h"
+#include "access/xlog_internal.h"
+#include "access/xloginsert.h"
+#include "access/xlogutils.h"
#include "catalog/catalog.h"
#include "catalog/dependency.h"
#include "catalog/heap.h"
@@ -32,6 +36,7 @@
#include "catalog/namespace.h"
#include "catalog/objectaccess.h"
#include "catalog/pg_am.h"
+#include "catalog/pg_control.h"
#include "catalog/pg_inherits.h"
#include "catalog/toasting.h"
#include "commands/cluster.h"
@@ -39,10 +44,15 @@
#include "commands/progress.h"
#include "commands/tablecmds.h"
#include "commands/vacuum.h"
+#include "executor/executor.h"
#include "miscadmin.h"
#include "optimizer/optimizer.h"
#include "pgstat.h"
+#include "replication/decode.h"
+#include "replication/logical.h"
+#include "replication/snapbuild.h"
#include "storage/bufmgr.h"
+#include "storage/ipc.h"
#include "storage/lmgr.h"
#include "storage/predicate.h"
#include "utils/acl.h"
@@ -76,14 +86,96 @@ typedef struct
((cmd) == CLUSTER_COMMAND_REPACK ? \
"repack" : "vacuum"))
+/*
+ * The following definitions are used for concurrent processing.
+ */
+
+/*
+ * OID of the table being repacked by this backend.
+ */
+static Oid repacked_rel = InvalidOid;
+/* The same for its TOAST relation. */
+static Oid repacked_rel_toast = InvalidOid;
+
+/*
+ * The locators are used to avoid logical decoding of data that we do not need
+ * for our table.
+ */
+RelFileLocator repacked_rel_locator = {.relNumber = InvalidOid};
+RelFileLocator repacked_rel_toast_locator = {.relNumber = InvalidOid};
+
+#define REPACK_CONCURRENT_IN_PROGRESS_MSG \
+ "relation \"%s\" is already being processed by REPACK CONCURRENTLY"
+
+/*
+ * Everything we need to call ExecInsertIndexTuples().
+ */
+typedef struct IndexInsertState
+{
+ ResultRelInfo *rri;
+ EState *estate;
+ ExprContext *econtext;
+
+ Relation ident_index;
+} IndexInsertState;
+
+/*
+ * Catalog information to check if another backend changed the relation in
+ * such a way that makes CLUSTE CONCURRENTLY unable to continue. Such changes
+ * are possible because cluster_rel() has to release its lock on the relation
+ * in order to acquire AccessExclusiveLock that it needs to swap the relation
+ * files.
+ *
+ * The most obvious problem is that the tuple descriptor has changed, since
+ * then the tuples we try to insert into the new storage are not guaranteed to
+ * fit into the storage.
+ *
+ * Another problem is relfilenode changed by another backend. It's not
+ * necessarily a correctness issue (e.g. when the other backend ran
+ * cluster_rel()), but it's safer for us to terminate the table processing in
+ * such cases. However, this information is also needs to be checked during
+ * logical decoding, so we store it in global variables repacked_rel_locator
+ * and repacked_rel_toast_locator above.
+ *
+ * Where possible, commands which might change the relation in an incompatible
+ * way should check if REPACK CONCURRENTLY is running, before they start to do
+ * the actual changes (see is_concurrent_repack_in_progress()). Anything else
+ * must be caught by check_catalog_changes(), which uses this structure.
+ */
+typedef struct CatalogState
+{
+ /* Tuple descriptor of the relation. */
+ TupleDesc tupdesc;
+
+ /* The number of indexes tracked. */
+ int ninds;
+ /* The index OIDs. */
+ Oid *ind_oids;
+ /* The index tuple descriptors. */
+ TupleDesc *ind_tupdescs;
+
+ /* The following are copies of the corresponding fields of pg_class. */
+ char relpersistence;
+ char replident;
+
+ /* rd_replidindex */
+ Oid replidindex;
+} CatalogState;
+
+/* The WAL segment being decoded. */
+static XLogSegNo repack_current_segment = 0;
+
static void cluster_multiple_rels(List *rtcs, ClusterParams *params,
- ClusterCommand cmd);
+ ClusterCommand cmd, LOCKMODE lockmode,
+ bool isTopLevel);
static void rebuild_relation(Relation OldHeap, Relation index, bool verbose,
- ClusterCommand cmd);
+ ClusterCommand cmd, bool concurrent);
static void copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
+ Snapshot snapshot, LogicalDecodingContext *decoding_ctx,
bool verbose, ClusterCommand cmd,
bool *pSwapToastByContent,
- TransactionId *pFreezeXid, MultiXactId *pCutoffMulti);
+ TransactionId *pFreezeXid,
+ MultiXactId *pCutoffMulti);
static List *get_tables_to_cluster(MemoryContext cluster_context);
static List *get_tables_to_repack(MemoryContext repack_context);
static List *get_tables_to_cluster_partitioned(MemoryContext cluster_context,
@@ -91,8 +183,91 @@ static List *get_tables_to_cluster_partitioned(MemoryContext cluster_context,
ClusterCommand cmd);
static bool cluster_is_permitted_for_relation(Oid relid, Oid userid,
ClusterCommand cmd);
+static void begin_concurrent_repack(Relation *rel_p, Relation *index_p,
+ bool *entered_p);
+static void end_concurrent_repack(bool error);
+static void cluster_before_shmem_exit_callback(int code, Datum arg);
+static CatalogState *get_catalog_state(Relation rel);
+static void free_catalog_state(CatalogState *state);
+static void check_catalog_changes(Relation rel, CatalogState *cat_state);
+static LogicalDecodingContext *setup_logical_decoding(Oid relid,
+ const char *slotname,
+ TupleDesc tupdesc);
+static HeapTuple get_changed_tuple(char *change);
+static void apply_concurrent_changes(RepackDecodingState *dstate,
+ Relation rel, ScanKey key, int nkeys,
+ IndexInsertState *iistate);
+static void apply_concurrent_insert(Relation rel, ConcurrentChange *change,
+ HeapTuple tup, IndexInsertState *iistate,
+ TupleTableSlot *index_slot);
+static void apply_concurrent_update(Relation rel, HeapTuple tup,
+ HeapTuple tup_target,
+ ConcurrentChange *change,
+ IndexInsertState *iistate,
+ TupleTableSlot *index_slot);
+static void apply_concurrent_delete(Relation rel, HeapTuple tup_target,
+ ConcurrentChange *change);
+static HeapTuple find_target_tuple(Relation rel, ScanKey key, int nkeys,
+ HeapTuple tup_key,
+ IndexInsertState *iistate,
+ TupleTableSlot *ident_slot,
+ IndexScanDesc *scan_p);
+static void process_concurrent_changes(LogicalDecodingContext *ctx,
+ XLogRecPtr end_of_wal,
+ Relation rel_dst,
+ Relation rel_src,
+ ScanKey ident_key,
+ int ident_key_nentries,
+ IndexInsertState *iistate);
+static IndexInsertState *get_index_insert_state(Relation relation,
+ Oid ident_index_id);
+static ScanKey build_identity_key(Oid ident_idx_oid, Relation rel_src,
+ int *nentries);
+static void free_index_insert_state(IndexInsertState *iistate);
+static void cleanup_logical_decoding(LogicalDecodingContext *ctx);
+static void rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
+ Relation cl_index,
+ CatalogState *cat_state,
+ LogicalDecodingContext *ctx,
+ bool swap_toast_by_content,
+ TransactionId frozenXid,
+ MultiXactId cutoffMulti);
+static List *build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes);
+
+/*
+ * Use this API when relation needs to be unlocked, closed and re-opened. If
+ * the relation got dropped while being unlocked, raise ERROR that mentions
+ * the relation name rather than OID.
+ */
+typedef struct RelReopenInfo
+{
+ /*
+ * The relation to be closed. Pointer to the value is stored here so that
+ * the user gets his reference updated automatically on re-opening.
+ *
+ * When calling unlock_and_close_relations(), 'relid' can be passed
+ * instead of 'rel_p' when the caller only needs to gather information for
+ * subsequent opening.
+ */
+ Relation *rel_p;
+ Oid relid;
+
+ char relkind;
+ LOCKMODE lockmode_orig; /* The existing lock mode */
+ LOCKMODE lockmode_new; /* The lock mode after the relation is
+ * re-opened */
+
+ char *relname; /* Relation name, initialized automatically. */
+} RelReopenInfo;
+
+static void init_rel_reopen_info(RelReopenInfo *rri, Relation *rel_p,
+ Oid relid, LOCKMODE lockmode_orig,
+ LOCKMODE lockmode_new);
+static void unlock_and_close_relations(RelReopenInfo *rels, int nrel);
+static void reopen_relations(RelReopenInfo *rels, int nrel);
static Relation process_single_relation(RangeVar *relation, char *indexname,
- ClusterCommand cmd,
+ ClusterCommand cmd, LOCKMODE lockmode,
+ bool isTopLevel,
ClusterParams *params,
Oid *indexOid_p);
@@ -151,8 +326,9 @@ cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel)
if (stmt->relation != NULL)
{
rel = process_single_relation(stmt->relation, stmt->indexname,
- CLUSTER_COMMAND_CLUSTER, ¶ms,
- &indexOid);
+ CLUSTER_COMMAND_CLUSTER,
+ AccessExclusiveLock, isTopLevel,
+ ¶ms, &indexOid);
if (rel == NULL)
return;
}
@@ -202,7 +378,8 @@ cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel)
}
/* Do the job. */
- cluster_multiple_rels(rtcs, ¶ms, CLUSTER_COMMAND_CLUSTER);
+ cluster_multiple_rels(rtcs, ¶ms, CLUSTER_COMMAND_CLUSTER,
+ AccessExclusiveLock, isTopLevel);
/* Start a new transaction for the cleanup work. */
StartTransactionCommand();
@@ -219,8 +396,8 @@ cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel)
* return.
*/
static void
-cluster_multiple_rels(List *rtcs, ClusterParams *params,
- ClusterCommand cmd)
+cluster_multiple_rels(List *rtcs, ClusterParams *params, ClusterCommand cmd,
+ LOCKMODE lockmode, bool isTopLevel)
{
ListCell *lc;
@@ -240,10 +417,10 @@ cluster_multiple_rels(List *rtcs, ClusterParams *params,
/* functions in indexes may want a snapshot set */
PushActiveSnapshot(GetTransactionSnapshot());
- rel = table_open(rtc->tableOid, AccessExclusiveLock);
+ rel = table_open(rtc->tableOid, lockmode);
/* Process this table */
- cluster_rel(rel, rtc->indexOid, params, cmd);
+ cluster_rel(rel, rtc->indexOid, params, cmd, isTopLevel);
/* cluster_rel closes the relation, but keeps lock */
PopActiveSnapshot();
@@ -267,12 +444,18 @@ cluster_multiple_rels(List *rtcs, ClusterParams *params,
* instead of index order. This is the new implementation of VACUUM FULL,
* and error messages should refer to the operation as VACUUM not CLUSTER.
*
+ * Note that, in the concurrent case, the function releases the lock at some
+ * point, in order to get AccessExclusiveLock for the final steps (i.e. to
+ * swap the relation files). To make things simpler, the caller should expect
+ * OldHeap to be closed on return, regardless CLUOPT_CONCURRENT. (The
+ * AccessExclusiveLock is kept till the end of the transaction.)
+ *
* 'cmd' indicates which commands is being executed. REPACK should be the only
* caller of this function in the future.
*/
void
cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
- ClusterCommand cmd)
+ ClusterCommand cmd, bool isTopLevel)
{
Oid tableOid = RelationGetRelid(OldHeap);
Oid save_userid;
@@ -282,8 +465,53 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
bool recheck = ((params->options & CLUOPT_RECHECK) != 0);
Relation index;
const char *cmd_str = CLUSTER_COMMAND_STR(cmd);
+ bool concurrent = ((params->options & CLUOPT_CONCURRENT) != 0);
+ LOCKMODE lmode;
+ bool entered, success;
+
+ /*
+ * Check that the correct lock is held. The lock mode is
+ * AccessExclusiveLock for normal processing and ShareUpdateExclusiveLock
+ * for concurrent processing (so that SELECT, INSERT, UPDATE and DELETE
+ * commands work, but cluster_rel() cannot be called concurrently for the
+ * same relation).
+ */
+ lmode = !concurrent ? AccessExclusiveLock : ShareUpdateExclusiveLock;
+
+ /*
+ * Skip the relation if it's being processed concurrently. In such a case,
+ * we cannot rely on a lock because the other backend needs to release it
+ * temporarily at some point.
+ *
+ * This check should not take place until we have a lock that prevents
+ * another backend from starting VREPACK CONCURRENTLY after our check.
+ */
+ Assert(CheckRelationLockedByMe(OldHeap, lmode, false));
+ if (is_concurrent_repack_in_progress(tableOid))
+ {
+ ereport(NOTICE,
+ (errmsg(REPACK_CONCURRENT_IN_PROGRESS_MSG,
+ RelationGetRelationName(OldHeap))));
+ table_close(OldHeap, lmode);
+ return;
+ }
+
+ /* There are specific requirements on concurrent processing. */
+ if (concurrent)
+ {
+ /*
+ * Make sure we have no XID assigned, otherwise call of
+ * setup_logical_decoding() can cause a deadlock.
+ *
+ * The existence of transaction block actually does not imply that XID
+ * was already assigned, but it very likely is. We might want to check
+ * the result of GetCurrentTransactionIdIfAny() instead, but that
+ * would be less clear from user's perspective.
+ */
+ PreventInTransactionBlock(isTopLevel, "REPACK CONCURRENTLY");
- Assert(CheckRelationLockedByMe(OldHeap, AccessExclusiveLock, false));
+ can_repack_concurrently(OldHeap);
+ }
/* Check for user-requested abort. */
CHECK_FOR_INTERRUPTS();
@@ -333,7 +561,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
/* Check that the user still has privileges for the relation */
if (!cluster_is_permitted_for_relation(tableOid, save_userid, cmd))
{
- relation_close(OldHeap, AccessExclusiveLock);
+ relation_close(OldHeap, lmode);
goto out;
}
@@ -348,7 +576,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
*/
if (RELATION_IS_OTHER_TEMP(OldHeap))
{
- relation_close(OldHeap, AccessExclusiveLock);
+ relation_close(OldHeap, lmode);
goto out;
}
@@ -359,7 +587,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
*/
if (!SearchSysCacheExists1(RELOID, ObjectIdGetDatum(indexOid)))
{
- relation_close(OldHeap, AccessExclusiveLock);
+ relation_close(OldHeap, lmode);
goto out;
}
@@ -370,7 +598,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
if ((params->options & CLUOPT_RECHECK_ISCLUSTERED) != 0 &&
!get_index_isclustered(indexOid))
{
- relation_close(OldHeap, AccessExclusiveLock);
+ relation_close(OldHeap, lmode);
goto out;
}
}
@@ -390,6 +618,11 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
ereport(ERROR,
(errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
errmsg("cannot %s a shared catalog", cmd_str)));
+ /*
+ * The CONCURRENTLY case should have been rejected earlier because it does
+ * not support system catalogs.
+ */
+ Assert(!(OldHeap->rd_rel->relisshared && concurrent));
/*
* Don't process temp tables of other backends ... their local buffer
@@ -411,8 +644,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
if (OidIsValid(indexOid))
{
/* verify the index is good and lock it */
- check_index_is_clusterable(OldHeap, indexOid, AccessExclusiveLock,
- cmd);
+ check_index_is_clusterable(OldHeap, indexOid, lmode, cmd);
/* also open it */
index = index_open(indexOid, NoLock);
}
@@ -429,7 +661,8 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
if (OldHeap->rd_rel->relkind == RELKIND_MATVIEW &&
!RelationIsPopulated(OldHeap))
{
- relation_close(OldHeap, AccessExclusiveLock);
+ index_close(index, lmode);
+ relation_close(OldHeap, lmode);
goto out;
}
@@ -442,11 +675,42 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
* invalid, because we move tuples around. Promote them to relation
* locks. Predicate locks on indexes will be promoted when they are
* reindexed.
+ *
+ * During concurrent processing, the heap as well as its indexes stay in
+ * operation, so we postpone this step until they are locked using
+ * AccessExclusiveLock near the end of the processing.
*/
- TransferPredicateLocksToHeapRelation(OldHeap);
+ if (!concurrent)
+ TransferPredicateLocksToHeapRelation(OldHeap);
/* rebuild_relation does all the dirty work */
- rebuild_relation(OldHeap, index, verbose, cmd);
+ entered = false;
+ success = false;
+ PG_TRY();
+ {
+ /*
+ * For concurrent processing, make sure other transactions treat this
+ * table as if it was a system / user catalog, and WAL the relevant
+ * additional information. ERROR is raised if another backend is
+ * processing the same table.
+ */
+ if (concurrent)
+ {
+ Relation *index_p = index ? &index : NULL;
+
+ begin_concurrent_repack(&OldHeap, index_p, &entered);
+ }
+
+ rebuild_relation(OldHeap, index, verbose, cmd, concurrent);
+ success = true;
+ }
+ PG_FINALLY();
+ {
+ if (concurrent && entered)
+ end_concurrent_repack(!success);
+ }
+ PG_END_TRY();
+
/* rebuild_relation closes OldHeap, and index if valid */
out:
@@ -595,19 +859,86 @@ mark_index_clustered(Relation rel, Oid indexOid, bool is_internal)
table_close(pg_index, RowExclusiveLock);
}
+/*
+ * Check if the CONCURRENTLY option is legal for the relation.
+ */
+void
+can_repack_concurrently(Relation rel)
+{
+ char relpersistence, replident;
+ Oid ident_idx;
+
+ /* Data changes in system relations are not logically decoded. */
+ if (IsCatalogRelation(rel))
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot repack relation \"%s\"",
+ RelationGetRelationName(rel)),
+ errhint("REPACK CONCURRENTLY is not supported for catalog relations.")));
+
+ /*
+ * reorderbuffer.c does not seem to handle processing of TOAST relation
+ * alone.
+ */
+ if (IsToastRelation(rel))
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot repack relation \"%s\"",
+ RelationGetRelationName(rel)),
+ errhint("REPACK (CONCURRENTLY) is not supported for TOAST relations, unless the main relation is repacked too.")));
+
+ relpersistence = rel->rd_rel->relpersistence;
+ if (relpersistence != RELPERSISTENCE_PERMANENT)
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("cannot repack relation \"%s\"",
+ RelationGetRelationName(rel)),
+ errhint("REPACK (CONCURRENTLY) is only allowed for permanent relations.")));
+
+ /* With NOTHING, WAL does not contain the old tuple. */
+ replident = rel->rd_rel->relreplident;
+ if (replident == REPLICA_IDENTITY_NOTHING)
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("cannot repack relation \"%s\"",
+ RelationGetRelationName(rel)),
+ errhint("Relation \"%s\" has insufficient replication identity.",
+ RelationGetRelationName(rel))));
+
+ /*
+ * Identity index is not set if the replica identity is FULL, but PK might
+ * exist in such a case.
+ */
+ ident_idx = RelationGetReplicaIndex(rel);
+ if (!OidIsValid(ident_idx) && OidIsValid(rel->rd_pkindex))
+ ident_idx = rel->rd_pkindex;
+ if (!OidIsValid(ident_idx))
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("cannot process relation \"%s\"",
+ RelationGetRelationName(rel)),
+ (errhint("Relation \"%s\" has no identity index.",
+ RelationGetRelationName(rel)))));
+}
+
/*
* rebuild_relation: rebuild an existing relation in index or physical order
*
- * OldHeap: table to rebuild.
+ * OldHeap: table to rebuild. See cluster_rel() for comments on the required
+ * lock strength.
+ *
* index: index to cluster by, or NULL to rewrite in physical order.
*
- * On entry, heap and index (if one is given) must be open, and
- * AccessExclusiveLock held on them.
- * On exit, they are closed, but locks on them are not released.
+ * On entry, heap and index (if one is given) must be open, and the
+ * appropriate lock held on them (AccessExclusiveLock for exclusive processing
+ * and ShareUpdateExclusiveLock for concurrent processing)..
+ *
+ * On exit, they are closed, but still locked with AccessExclusiveLock (The
+ * function handles the lock upgrade if 'concurrent' is true.)
*/
static void
rebuild_relation(Relation OldHeap, Relation index, bool verbose,
- ClusterCommand cmd)
+ ClusterCommand cmd, bool concurrent)
{
Oid tableOid = RelationGetRelid(OldHeap);
Oid accessMethod = OldHeap->rd_rel->relam;
@@ -615,13 +946,81 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
Oid OIDNewHeap;
Relation NewHeap;
char relpersistence;
- bool is_system_catalog;
bool swap_toast_by_content;
TransactionId frozenXid;
MultiXactId cutoffMulti;
+ NameData slotname;
+ LogicalDecodingContext *ctx = NULL;
+ Snapshot snapshot = NULL;
+ CatalogState *cat_state = NULL;
+ LOCKMODE lmode;
+
+ lmode = !concurrent ? AccessExclusiveLock : ShareUpdateExclusiveLock;
+
+ Assert(CheckRelationLockedByMe(OldHeap, lmode, false) &&
+ (index == NULL || CheckRelationLockedByMe(index, lmode, false)));
+
+ if (concurrent)
+ {
+ TupleDesc tupdesc;
+ RelReopenInfo rri[2];
+ int nrel;
+
+ /*
+ * REPACK CONCURRENTLY is not allowed in a transaction block, so this
+ * should never fire.
+ */
+ Assert(GetTopTransactionIdIfAny() == InvalidTransactionId);
- Assert(CheckRelationLockedByMe(OldHeap, AccessExclusiveLock, false) &&
- (index == NULL || CheckRelationLockedByMe(index, AccessExclusiveLock, false)));
+ /*
+ * A single backend should not execute multiple REPACK commands at a
+ * time, so use PID to make the slot unique.
+ */
+ snprintf(NameStr(slotname), NAMEDATALEN, "repack_%d", MyProcPid);
+
+ /*
+ * Gather catalog information so that we can check later if the old
+ * relation has not changed while unlocked.
+ *
+ * Since this function also checks if the relation can be processed,
+ * it's important to call it before we spend notable amount of time to
+ * setup the logical decoding. Not sure though if it's necessary to do
+ * it even earlier.
+ */
+ cat_state = get_catalog_state(OldHeap);
+
+ tupdesc = CreateTupleDescCopy(RelationGetDescr(OldHeap));
+
+ /*
+ * Unlock the relation (and possibly the clustering index) to avoid
+ * deadlock because setup_logical_decoding() will wait for all the
+ * running transactions (with XID assigned) to finish. Some of those
+ * transactions might be waiting for a lock on our relation.
+ */
+ nrel = 0;
+ init_rel_reopen_info(&rri[nrel++], &OldHeap, InvalidOid,
+ ShareUpdateExclusiveLock,
+ ShareUpdateExclusiveLock);
+ if (index)
+ init_rel_reopen_info(&rri[nrel++], &index, InvalidOid,
+ ShareUpdateExclusiveLock,
+ ShareUpdateExclusiveLock);
+ unlock_and_close_relations(rri, nrel);
+
+ /* Prepare to capture the concurrent data changes. */
+ ctx = setup_logical_decoding(tableOid, NameStr(slotname), tupdesc);
+
+ /* Lock the table (and index) again. */
+ reopen_relations(rri, nrel);
+
+ /*
+ * Check if a 'tupdesc' could have changed while the relation was
+ * unlocked.
+ */
+ check_catalog_changes(OldHeap, cat_state);
+
+ snapshot = SnapBuildInitialSnapshotForRepack(ctx->snapshot_builder);
+ }
if (index)
/* Mark the correct index as clustered */
@@ -629,7 +1028,6 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
/* Remember info about rel before closing OldHeap */
relpersistence = OldHeap->rd_rel->relpersistence;
- is_system_catalog = IsSystemRelation(OldHeap);
/*
* Create the transient table that will receive the re-ordered data.
@@ -645,30 +1043,51 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
NewHeap = table_open(OIDNewHeap, NoLock);
/* Copy the heap data into the new table in the desired order */
- copy_table_data(NewHeap, OldHeap, index, verbose, cmd,
- &swap_toast_by_content, &frozenXid, &cutoffMulti);
+ copy_table_data(NewHeap, OldHeap, index, snapshot, ctx, verbose,
+ cmd, &swap_toast_by_content, &frozenXid, &cutoffMulti);
+ if (concurrent)
+ {
+ rebuild_relation_finish_concurrent(NewHeap, OldHeap, index,
+ cat_state, ctx,
+ swap_toast_by_content,
+ frozenXid, cutoffMulti);
+
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_FINAL_CLEANUP);
+
+ /* Done with decoding. */
+ FreeSnapshot(snapshot);
+ free_catalog_state(cat_state);
+ cleanup_logical_decoding(ctx);
+ ReplicationSlotRelease();
+ ReplicationSlotDrop(NameStr(slotname), false);
+ }
+ else
+ {
+ bool is_system_catalog = IsSystemRelation(OldHeap);
- /* Close relcache entries, but keep lock until transaction commit */
- table_close(OldHeap, NoLock);
- if (index)
- index_close(index, NoLock);
+ /* Close relcache entries, but keep lock until transaction commit */
+ table_close(OldHeap, NoLock);
+ if (index)
+ index_close(index, NoLock);
- /*
- * Close the new relation so it can be dropped as soon as the storage is
- * swapped. The relation is not visible to others, so no need to unlock it
- * explicitly.
- */
- table_close(NewHeap, NoLock);
+ /*
+ * Close the new relation so it can be dropped as soon as the storage
+ * is swapped. The relation is not visible to others, so no need to
+ * unlock it explicitly.
+ */
+ table_close(NewHeap, NoLock);
- /*
- * Swap the physical files of the target and transient tables, then
- * rebuild the target's indexes and throw away the transient table.
- */
- finish_heap_swap(tableOid, OIDNewHeap, is_system_catalog,
- swap_toast_by_content, false, true,
- frozenXid, cutoffMulti,
- relpersistence);
+ /*
+ * Swap the physical files of the target and transient tables, then
+ * rebuild the target's indexes and throw away the transient table.
+ */
+ finish_heap_swap(tableOid, OIDNewHeap, is_system_catalog,
+ swap_toast_by_content, false, true, true,
+ frozenXid, cutoffMulti,
+ relpersistence);
+ }
}
@@ -803,14 +1222,18 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
/*
* Do the physical copying of table data.
*
+ * 'snapshot' and 'decoding_ctx': see table_relation_copy_for_cluster(). Pass
+ * iff concurrent processing is required.
+ *
* There are three output parameters:
* *pSwapToastByContent is set true if toast tables must be swapped by content.
* *pFreezeXid receives the TransactionId used as freeze cutoff point.
* *pCutoffMulti receives the MultiXactId used as a cutoff point.
*/
static void
-copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verbose,
- ClusterCommand cmd, bool *pSwapToastByContent,
+copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
+ Snapshot snapshot, LogicalDecodingContext *decoding_ctx,
+ bool verbose, ClusterCommand cmd, bool *pSwapToastByContent,
TransactionId *pFreezeXid, MultiXactId *pCutoffMulti)
{
Relation relRelation;
@@ -829,6 +1252,7 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
const char *cmd_str = CLUSTER_COMMAND_STR(cmd);
PGRUsage ru0;
char *nspname;
+ bool concurrent = snapshot != NULL;
pg_rusage_init(&ru0);
@@ -855,8 +1279,12 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
*
* We don't need to open the toast relation here, just lock it. The lock
* will be held till end of transaction.
+ *
+ * In the REPACK CONCURRENTLY case, the lock does not help because we need
+ * to release it temporarily at some point. Instead, we expect VACUUM /
+ * CLUSTER to skip tables which are present in RepackedRelsHash.
*/
- if (OldHeap->rd_rel->reltoastrelid)
+ if (OldHeap->rd_rel->reltoastrelid && !concurrent)
LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
/*
@@ -932,8 +1360,48 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
* provided, else plain seqscan.
*/
if (OldIndex != NULL && OldIndex->rd_rel->relam == BTREE_AM_OID)
+ {
+ ResourceOwner oldowner = NULL;
+ ResourceOwner resowner = NULL;
+
+ /*
+ * In the CONCURRENT case, use a dedicated resource owner so we don't
+ * leave any additional locks behind us that we cannot release easily.
+ */
+ if (concurrent)
+ {
+ Assert(CheckRelationLockedByMe(OldHeap, ShareUpdateExclusiveLock,
+ false));
+ Assert(CheckRelationLockedByMe(OldIndex, ShareUpdateExclusiveLock,
+ false));
+
+ resowner = ResourceOwnerCreate(CurrentResourceOwner,
+ "plan_cluster_use_sort");
+ oldowner = CurrentResourceOwner;
+ CurrentResourceOwner = resowner;
+ }
+
use_sort = plan_cluster_use_sort(RelationGetRelid(OldHeap),
RelationGetRelid(OldIndex));
+
+ if (concurrent)
+ {
+ CurrentResourceOwner = oldowner;
+
+ /*
+ * We are primarily concerned about locks, but if the planner
+ * happened to allocate any other resources, we should release
+ * them too because we're going to delete the whole resowner.
+ */
+ ResourceOwnerRelease(resowner, RESOURCE_RELEASE_BEFORE_LOCKS,
+ false, false);
+ ResourceOwnerRelease(resowner, RESOURCE_RELEASE_LOCKS,
+ false, false);
+ ResourceOwnerRelease(resowner, RESOURCE_RELEASE_AFTER_LOCKS,
+ false, false);
+ ResourceOwnerDelete(resowner);
+ }
+ }
else
use_sort = false;
@@ -965,7 +1433,9 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
* values (e.g. because the AM doesn't use freezing).
*/
table_relation_copy_for_cluster(OldHeap, NewHeap, OldIndex, use_sort,
- cutoffs.OldestXmin, &cutoffs.FreezeLimit,
+ cutoffs.OldestXmin, snapshot,
+ decoding_ctx,
+ &cutoffs.FreezeLimit,
&cutoffs.MultiXactCutoff,
&num_tuples, &tups_vacuumed,
&tups_recently_dead);
@@ -974,7 +1444,11 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
*pFreezeXid = cutoffs.FreezeLimit;
*pCutoffMulti = cutoffs.MultiXactCutoff;
- /* Reset rd_toastoid just to be tidy --- it shouldn't be looked at again */
+ /*
+ * Reset rd_toastoid just to be tidy --- it shouldn't be looked at
+ * again. In the CONCURRENTLY case, we need to set it again before
+ * applying the concurrent changes.
+ */
NewHeap->rd_toastoid = InvalidOid;
num_pages = RelationGetNumberOfBlocks(NewHeap);
@@ -1427,14 +1901,13 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
bool swap_toast_by_content,
bool check_constraints,
bool is_internal,
+ bool reindex,
TransactionId frozenXid,
MultiXactId cutoffMulti,
char newrelpersistence)
{
ObjectAddress object;
Oid mapped_tables[4];
- int reindex_flags;
- ReindexParams reindex_params = {0};
int i;
/* Report that we are now swapping relation files */
@@ -1460,39 +1933,46 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
if (is_system_catalog)
CacheInvalidateCatalog(OIDOldHeap);
- /*
- * Rebuild each index on the relation (but not the toast table, which is
- * all-new at this point). It is important to do this before the DROP
- * step because if we are processing a system catalog that will be used
- * during DROP, we want to have its indexes available. There is no
- * advantage to the other order anyway because this is all transactional,
- * so no chance to reclaim disk space before commit. We do not need a
- * final CommandCounterIncrement() because reindex_relation does it.
- *
- * Note: because index_build is called via reindex_relation, it will never
- * set indcheckxmin true for the indexes. This is OK even though in some
- * sense we are building new indexes rather than rebuilding existing ones,
- * because the new heap won't contain any HOT chains at all, let alone
- * broken ones, so it can't be necessary to set indcheckxmin.
- */
- reindex_flags = REINDEX_REL_SUPPRESS_INDEX_USE;
- if (check_constraints)
- reindex_flags |= REINDEX_REL_CHECK_CONSTRAINTS;
+ if (reindex)
+ {
+ int reindex_flags;
+ ReindexParams reindex_params = {0};
- /*
- * Ensure that the indexes have the same persistence as the parent
- * relation.
- */
- if (newrelpersistence == RELPERSISTENCE_UNLOGGED)
- reindex_flags |= REINDEX_REL_FORCE_INDEXES_UNLOGGED;
- else if (newrelpersistence == RELPERSISTENCE_PERMANENT)
- reindex_flags |= REINDEX_REL_FORCE_INDEXES_PERMANENT;
+ /*
+ * Rebuild each index on the relation (but not the toast table, which
+ * is all-new at this point). It is important to do this before the
+ * DROP step because if we are processing a system catalog that will
+ * be used during DROP, we want to have its indexes available. There
+ * is no advantage to the other order anyway because this is all
+ * transactional, so no chance to reclaim disk space before commit.
+ * We do not need a final CommandCounterIncrement() because
+ * reindex_relation does it.
+ *
+ * Note: because index_build is called via reindex_relation, it will never
+ * set indcheckxmin true for the indexes. This is OK even though in some
+ * sense we are building new indexes rather than rebuilding existing ones,
+ * because the new heap won't contain any HOT chains at all, let alone
+ * broken ones, so it can't be necessary to set indcheckxmin.
+ */
+ reindex_flags = REINDEX_REL_SUPPRESS_INDEX_USE;
+ if (check_constraints)
+ reindex_flags |= REINDEX_REL_CHECK_CONSTRAINTS;
- /* Report that we are now reindexing relations */
- pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
- PROGRESS_REPACK_PHASE_REBUILD_INDEX);
+ /*
+ * Ensure that the indexes have the same persistence as the parent
+ * relation.
+ */
+ if (newrelpersistence == RELPERSISTENCE_UNLOGGED)
+ reindex_flags |= REINDEX_REL_FORCE_INDEXES_UNLOGGED;
+ else if (newrelpersistence == RELPERSISTENCE_PERMANENT)
+ reindex_flags |= REINDEX_REL_FORCE_INDEXES_PERMANENT;
+
+ /* Report that we are now reindexing relations */
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_REBUILD_INDEX);
- reindex_relation(NULL, OIDOldHeap, reindex_flags, &reindex_params);
+ reindex_relation(NULL, OIDOldHeap, reindex_flags, &reindex_params);
+ }
/* Report that we are now doing clean up */
pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
@@ -1804,89 +2284,1975 @@ cluster_is_permitted_for_relation(Oid relid, Oid userid, ClusterCommand cmd)
return false;
}
+#define REPL_PLUGIN_NAME "pgoutput_repack"
+
/*
- * REPACK is intended to be a replacement of both CLUSTER and VACUUM FULL.
+ * Each relation being processed by REPACK CONCURRENTLY must be in the
+ * repackedRels hashtable.
*/
+typedef struct RepackedRel
+{
+ Oid relid;
+ Oid dbid;
+} RepackedRel;
+
+static HTAB *RepackedRelsHash = NULL;
+
+/* Maximum number of entries in the hashtable. */
+static int maxRepackedRels = 0;
+
+Size
+RepackShmemSize(void)
+{
+ /*
+ * A replication slot is needed for the processing, so use this GUC to
+ * allocate memory for the hashtable.
+ */
+ maxRepackedRels = max_replication_slots;
+
+ return hash_estimate_size(maxRepackedRels, sizeof(RepackedRel));
+}
+
void
-repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
+RepackShmemInit(void)
{
- ListCell *lc;
- ClusterParams params = {0};
- bool verbose = false;
- Relation rel = NULL;
- Oid indexOid = InvalidOid;
- MemoryContext repack_context;
- List *rtcs;
+ HASHCTL info;
- /* Parse option list */
- foreach(lc, stmt->params)
- {
- DefElem *opt = (DefElem *) lfirst(lc);
+ info.keysize = sizeof(RepackedRel);
+ info.entrysize = info.keysize;
- if (strcmp(opt->defname, "verbose") == 0)
- verbose = defGetBoolean(opt);
- else
- ereport(ERROR,
- (errcode(ERRCODE_SYNTAX_ERROR),
- errmsg("unrecognized REPACK option \"%s\"",
- opt->defname),
- parser_errposition(pstate, opt->location)));
+ RepackedRelsHash = ShmemInitHash("Repacked Relations",
+ maxRepackedRels,
+ maxRepackedRels,
+ &info,
+ HASH_ELEM | HASH_BLOBS);
+}
+
+/*
+ * Call this function before REPACK CONCURRENTLY starts to setup logical
+ * decoding. It makes sure that other users of the table put enough
+ * information into WAL.
+ *
+ * The point is that on various places we expect that the table we're
+ * processing is treated like a system catalog. For example, we need to be
+ * able to scan it using a "historic snapshot" anytime during the processing
+ * (as opposed to scanning only at the start point of the decoding, logical
+ * replication does during initial table synchronization), in order to apply
+ * concurrent UPDATE / DELETE commands.
+ *
+ * Since we need to close and reopen the relation here, the 'rel_p' and
+ * 'index_p' arguments are in/out.
+ *
+ * 'enter_p' receives a bool value telling whether relation OID was entered
+ * into the hashtable or not.
+ */
+static void
+begin_concurrent_repack(Relation *rel_p, Relation *index_p,
+ bool *entered_p)
+{
+ Relation rel = *rel_p;
+ Oid relid, toastrelid;
+ RepackedRel key, *entry;
+ bool found;
+ RelReopenInfo rri[2];
+ int nrel;
+ static bool before_shmem_exit_callback_setup = false;
+
+ relid = RelationGetRelid(rel);
+
+ /*
+ * Make sure that we do not leave an entry in RepackedRelsHash if exiting
+ * due to FATAL.
+ */
+ if (!before_shmem_exit_callback_setup)
+ {
+ before_shmem_exit(cluster_before_shmem_exit_callback, 0);
+ before_shmem_exit_callback_setup = true;
}
- params.options = (verbose ? CLUOPT_VERBOSE : 0);
+ memset(&key, 0, sizeof(key));
+ key.relid = relid;
+ key.dbid = MyDatabaseId;
- if (stmt->relation != NULL)
+ *entered_p = false;
+ LWLockAcquire(RepackedRelsLock, LW_EXCLUSIVE);
+ entry = (RepackedRel *)
+ hash_search(RepackedRelsHash, &key, HASH_ENTER_NULL, &found);
+ if (found)
{
- rel = process_single_relation(stmt->relation, stmt->indexname,
- CLUSTER_COMMAND_REPACK, ¶ms,
- &indexOid);
- if (rel == NULL)
- return;
+ /*
+ * Since REPACK CONCURRENTLY takes ShareRowExclusiveLock, a conflict
+ * should occur much earlier. However that lock may be released
+ * temporarily, see below. Anyway, we should complain whatever the
+ * reason of the conflict might be.
+ */
+ ereport(ERROR,
+ (errmsg(REPACK_CONCURRENT_IN_PROGRESS_MSG,
+ RelationGetRelationName(rel))));
}
+ if (entry == NULL)
+ ereport(ERROR,
+ (errmsg("too many requests for REPACK CONCURRENTLY at a time")),
+ (errhint("Please consider increasing the \"max_replication_slots\" configuration parameter.")));
/*
- * By here, we know we are in a multi-table situation. In order to avoid
- * holding locks for too long, we want to process each table in its own
- * transaction. This forces us to disallow running inside a user
- * transaction block.
+ * Even if the insertion of TOAST relid should fail below, the caller has
+ * to do cleanup.
*/
- PreventInTransactionBlock(isTopLevel, "REPACK");
+ *entered_p = true;
- /* Also, we need a memory context to hold our list of relations */
- repack_context = AllocSetContextCreate(PortalContext,
- "Repack",
- ALLOCSET_DEFAULT_SIZES);
+ /*
+ * Enable the callback to remove the entry in case of exit. We should not
+ * do this earlier, otherwise an attempt to insert already existing entry
+ * could make us remove that entry (inserted by another backend) during
+ * ERROR handling.
+ */
+ Assert(!OidIsValid(repacked_rel));
+ repacked_rel = relid;
- params.options |= CLUOPT_RECHECK;
- if (rel != NULL)
+ /*
+ * TOAST relation is not accessed using historic snapshot, but we enter it
+ * here to protect it from being VACUUMed by another backend. (Lock does
+ * not help in the CONCURRENTLY case because cannot hold it continuously
+ * till the end of the transaction.) See the comments on locking TOAST
+ * relation in copy_table_data().
+ */
+ toastrelid = rel->rd_rel->reltoastrelid;
+ if (OidIsValid(toastrelid))
{
- Oid relid;
- bool rel_is_index;
+ key.relid = toastrelid;
+ entry = (RepackedRel *)
+ hash_search(RepackedRelsHash, &key, HASH_ENTER_NULL, &found);
+ if (found)
+ /*
+ * If we could enter the main fork the TOAST should succeed
+ * too. Nevertheless, check.
+ */
+ ereport(ERROR,
+ (errmsg("TOAST relation of \"%s\" is already being processed by REPACK CONCURRENTLY",
+ RelationGetRelationName(rel))));
+ if (entry == NULL)
+ ereport(ERROR,
+ (errmsg("too many requests for REPACK CONCURRENTLY at a time")),
+ (errhint("Please consider increasing the \"max_replication_slots\" configuration parameter.")));
- Assert(rel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE);
+ Assert(!OidIsValid(repacked_rel_toast));
+ repacked_rel_toast = toastrelid;
+ }
+ LWLockRelease(RepackedRelsLock);
- if (OidIsValid(indexOid))
- {
- relid = indexOid;
- rel_is_index = true;
- }
- else
- {
- relid = RelationGetRelid(rel);
- rel_is_index = false;
- }
- rtcs = get_tables_to_cluster_partitioned(repack_context, relid,
- rel_is_index,
- CLUSTER_COMMAND_REPACK);
+ /*
+ * Make sure that other backends are aware of the new hash entry.
+ *
+ * Besides sending the invalidation message, we need to force re-opening
+ * of the relation, which includes the actual invalidation (and thus
+ * checking of our hashtable on the next access).
+ */
+ CacheInvalidateRelcacheImmediate(rel);
+ /*
+ * Since the hashtable only needs to be checked by write transactions,
+ * lock the relation in a mode that conflicts with any DML command. (The
+ * reading transactions are supposed to close the relation before opening
+ * it with higher lock.) Once we have the relation (and its index) locked,
+ * we unlock it immediately and then re-lock using the original mode.
+ */
+ nrel = 0;
+ init_rel_reopen_info(&rri[nrel++], rel_p, InvalidOid,
+ ShareUpdateExclusiveLock, ShareLock);
+ if (index_p)
+ {
+ /*
+ * Another transaction might want to open both the relation and the
+ * index. If it already has the relation lock and is waiting for the
+ * index lock, we should release the index lock, otherwise our request
+ * for ShareLock on the relation can end up in a deadlock.
+ */
+ init_rel_reopen_info(&rri[nrel++], index_p, InvalidOid,
+ ShareUpdateExclusiveLock, ShareLock);
+ }
+ unlock_and_close_relations(rri, nrel);
+ /*
+ * XXX It's not strictly necessary to lock the index here, but it's
+ * probably not worth teaching the "reopen API" about this special case.
+ */
+ reopen_relations(rri, nrel);
+
+ /* Switch back to the original lock. */
+ nrel = 0;
+ init_rel_reopen_info(&rri[nrel++], rel_p, InvalidOid,
+ ShareLock, ShareUpdateExclusiveLock);
+ if (index_p)
+ init_rel_reopen_info(&rri[nrel++], index_p, InvalidOid,
+ ShareLock, ShareUpdateExclusiveLock);
+ unlock_and_close_relations(rri, nrel);
+ reopen_relations(rri, nrel);
+ /* Make sure the reopened relcache entry is used, not the old one. */
+ rel = *rel_p;
+
+ /* Avoid logical decoding of other relations by this backend. */
+ repacked_rel_locator = rel->rd_locator;
+ if (OidIsValid(toastrelid))
+ {
+ Relation toastrel;
+
+ /* Avoid logical decoding of other TOAST relations. */
+ toastrel = table_open(toastrelid, AccessShareLock);
+ repacked_rel_toast_locator = toastrel->rd_locator;
+ table_close(toastrel, AccessShareLock);
+ }
+}
+
+/*
+ * Call this when done with REPACK CONCURRENTLY.
+ *
+ * 'error' tells whether the function is being called in order to handle
+ * error.
+ */
+static void
+end_concurrent_repack(bool error)
+{
+ RepackedRel key;
+ RepackedRel *entry = NULL, *entry_toast = NULL;
+ Oid relid = repacked_rel;
+ Oid toastrelid = repacked_rel_toast;
+
+ /* Remove the relation from the hash if we managed to insert one. */
+ if (OidIsValid(repacked_rel))
+ {
+ memset(&key, 0, sizeof(key));
+ key.relid = repacked_rel;
+ key.dbid = MyDatabaseId;
+ LWLockAcquire(RepackedRelsLock, LW_EXCLUSIVE);
+ entry = hash_search(RepackedRelsHash, &key, HASH_REMOVE, NULL);
+
+ /*
+ * By clearing this variable we also disable
+ * cluster_before_shmem_exit_callback().
+ */
+ repacked_rel = InvalidOid;
+ }
+
+ /* Remove the TOAST relation if there is one. */
+ if (OidIsValid(repacked_rel_toast))
+ {
+ key.relid = repacked_rel_toast;
+ entry_toast = hash_search(RepackedRelsHash, &key, HASH_REMOVE,
+ NULL);
+
+ repacked_rel_toast = InvalidOid;
+ }
+ LWLockRelease(RepackedRelsLock);
+
+ /* Restore normal function of logical decoding. */
+ repacked_rel_locator.relNumber = InvalidOid;
+ repacked_rel_toast_locator.relNumber = InvalidOid;
+
+ /*
+ * On normal completion (!error), we should not really fail to remove the
+ * entry. But if it wasn't there for any reason, raise ERROR to make sure
+ * the transaction is aborted: if other transactions, while changing the
+ * contents of the relation, didn't know that REPACK CONCURRENTLY was in
+ * progress, they could have missed to WAL enough information, and thus we
+ * could have produced an inconsistent table contents.
+ *
+ * On the other hand, if we are already handling an error, there's no
+ * reason to worry about inconsistent contents of the new storage because
+ * the transaction is going to be rolled back anyway. Furthermore, by
+ * raising ERROR here we'd shadow the original error.
+ */
+ if (!error)
+ {
+ char *relname;
+
+ if (OidIsValid(relid) && entry == NULL)
+ {
+ relname = get_rel_name(relid);
+ if (!relname)
+ ereport(ERROR,
+ (errmsg("cache lookup failed for relation %u",
+ relid)));
+
+ ereport(ERROR,
+ (errmsg("relation \"%s\" not found among repacked relations",
+ relname)));
+ }
+
+ /*
+ * Likewise, the TOAST relation should not have disappeared.
+ */
+ if (OidIsValid(toastrelid) && entry_toast == NULL)
+ {
+ relname = get_rel_name(key.relid);
+ if (!relname)
+ ereport(ERROR,
+ (errmsg("cache lookup failed for relation %u",
+ key.relid)));
+
+ ereport(ERROR,
+ (errmsg("relation \"%s\" not found among repacked relations",
+ relname)));
+ }
+ }
+
+ /*
+ * Note: unlike begin_concurrent_repack(), here we do not lock/unlock the
+ * relation: 1) On normal completion, the caller is already holding
+ * AccessExclusiveLock (till the end of the transaction), 2) on ERROR /
+ * FATAL, we try to do the cleanup asap, but the worst case is that other
+ * backends will write unnecessary information to WAL until they close the
+ * relation.
+ */
+}
+
+/*
+ * A wrapper to call end_concurrent_repack() as a before_shmem_exit callback.
+ */
+static void
+cluster_before_shmem_exit_callback(int code, Datum arg)
+{
+ if (OidIsValid(repacked_rel) || OidIsValid(repacked_rel_toast))
+ end_concurrent_repack(true);
+}
+
+/*
+ * Check if relation is currently being processed by REPACK CONCURRENTLY.
+ */
+bool
+is_concurrent_repack_in_progress(Oid relid)
+{
+ RepackedRel key, *entry;
+
+ memset(&key, 0, sizeof(key));
+ key.relid = relid;
+ key.dbid = MyDatabaseId;
+
+ LWLockAcquire(RepackedRelsLock, LW_SHARED);
+ entry = (RepackedRel *)
+ hash_search(RepackedRelsHash, &key, HASH_FIND, NULL);
+ LWLockRelease(RepackedRelsLock);
+
+ return entry != NULL;
+}
+
+/*
+ * Check if REPACK CONCURRENTLY is already running for given relation, and if
+ * so, raise ERROR. The problem is that cluster_rel() needs to release its
+ * lock on the relation temporarily at some point, so our lock alone does not
+ * help. Commands that might break what cluster_rel() is doing should call
+ * this function first.
+ *
+ * Return without checking if lockmode allows for race conditions which would
+ * make the result meaningless. In that case, cluster_rel() itself should
+ * throw ERROR if the relation was changed by us in an incompatible
+ * way. However, if it managed to do most of its work by then, a lot of CPU
+ * time might be wasted.
+ */
+void
+check_for_concurrent_repack(Oid relid, LOCKMODE lockmode)
+{
+ /*
+ * If the caller does not have a lock that conflicts with
+ * ShareUpdateExclusiveLock, the check makes little sense because REPACK
+ * CONCURRENTLY can start anytime after the check.
+ */
+ if (lockmode < ShareUpdateExclusiveLock)
+ return;
+
+ /*
+ * The caller has a lock which conflicts with REPACK CONCURRENTLY, so if
+ * that's not running now, it cannot start until the caller's transaction
+ * has completed.
+ */
+ if (is_concurrent_repack_in_progress(relid))
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg(REPACK_CONCURRENT_IN_PROGRESS_MSG,
+ get_rel_name(relid))));
+
+}
+
+/*
+ * Check if relation is eligible for REPACK CONCURRENTLY and retrieve the
+ * catalog state to be passed later to check_catalog_changes.
+ *
+ * Caller is supposed to hold (at least) ShareUpdateExclusiveLock on the
+ * relation.
+ */
+static CatalogState *
+get_catalog_state(Relation rel)
+{
+ CatalogState *result = palloc_object(CatalogState);
+ List *ind_oids;
+ ListCell *lc;
+ int ninds, i;
+ char relpersistence = rel->rd_rel->relpersistence;
+ char replident = rel->rd_rel->relreplident;
+ Oid ident_idx = RelationGetReplicaIndex(rel);
+ TupleDesc td_src = RelationGetDescr(rel);
+
+ /*
+ * While gathering the catalog information, check if there is a reason not
+ * to proceed.
+ *
+ * This function was already called, but the relation was unlocked since
+ * (see begin_concurrent_repack()). check_catalog_changes() should catch
+ * any "disruptive" changes in the future.
+ */
+ can_repack_concurrently(rel);
+
+ /* No index should be dropped while we are checking it. */
+ Assert(CheckRelationLockedByMe(rel, ShareUpdateExclusiveLock, true));
+
+ ind_oids = RelationGetIndexList(rel);
+ result->ninds = ninds = list_length(ind_oids);
+ result->ind_oids = palloc_array(Oid, ninds);
+ result->ind_tupdescs = palloc_array(TupleDesc, ninds);
+ i = 0;
+ foreach(lc, ind_oids)
+ {
+ Oid ind_oid = lfirst_oid(lc);
+ Relation index;
+ TupleDesc td_ind_src, td_ind_dst;
+
+ /*
+ * Weaker lock should be o.k. for the index, but this one should not
+ * break anything either.
+ */
+ index = index_open(ind_oid, ShareUpdateExclusiveLock);
+
+ result->ind_oids[i] = RelationGetRelid(index);
+ td_ind_src = RelationGetDescr(index);
+ td_ind_dst = palloc(TupleDescSize(td_ind_src));
+ TupleDescCopy(td_ind_dst, td_ind_src);
+ result->ind_tupdescs[i] = td_ind_dst;
+ i++;
+
+ index_close(index, ShareUpdateExclusiveLock);
+ }
+
+ /* Fill-in the relation info. */
+ result->tupdesc = palloc(TupleDescSize(td_src));
+ TupleDescCopy(result->tupdesc, td_src);
+ result->relpersistence = relpersistence;
+ result->replident = replident;
+ result->replidindex = ident_idx;
+
+ return result;
+}
+
+static void
+free_catalog_state(CatalogState *state)
+{
+ /* We are only interested in indexes. */
+ if (state->ninds == 0)
+ return;
+
+ for (int i = 0; i < state->ninds; i++)
+ FreeTupleDesc(state->ind_tupdescs[i]);
+
+ FreeTupleDesc(state->tupdesc);
+ pfree(state->ind_oids);
+ pfree(state->ind_tupdescs);
+ pfree(state);
+}
+
+/*
+ * Raise ERROR if 'rel' changed in a way that does not allow further
+ * processing of REPACK CONCURRENTLY.
+ *
+ * Besides the relation's tuple descriptor, it's important to check indexes:
+ * concurrent change of index definition (can it happen in other way than
+ * dropping and re-creating the index, accidentally with the same OID?) can be
+ * a problem because we may already have the new index built. If an index was
+ * created or dropped concurrently, we'd fail to swap the index storage. In
+ * any case, we prefer to check the indexes early to get an explicit error
+ * message about the mismatch. Furthermore, the earlier we detect the change,
+ * the fewer CPU cycles we waste.
+ *
+ * Note that we do not check constraints because the transaction which changed
+ * them must have ensured that the existing tuples satisfy the new
+ * constraints. If any DML commands were necessary for that, we will simply
+ * decode them from WAL and apply them to the new storage.
+ *
+ * Caller is supposed to hold (at least) ShareUpdateExclusiveLock on the
+ * relation.
+ */
+static void
+check_catalog_changes(Relation rel, CatalogState *cat_state)
+{
+ Oid reltoastrelid = rel->rd_rel->reltoastrelid;
+ List *ind_oids;
+ ListCell *lc;
+ LOCKMODE lockmode;
+ Oid ident_idx;
+ TupleDesc td, td_cp;
+
+ /* First, check the relation info. */
+
+ /* TOAST is not easy to change, but check. */
+ if (reltoastrelid != repacked_rel_toast)
+ ereport(ERROR,
+ errmsg("TOAST relation of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel)));
+
+ /*
+ * Likewise, check_for_concurrent_repack() should prevent others from
+ * changing the relation file concurrently, but it's our responsibility to
+ * avoid data loss. (The original locators are stored outside cat_state,
+ * but the check belongs to this function.)
+ */
+ if (!RelFileLocatorEquals(rel->rd_locator, repacked_rel_locator))
+ ereport(ERROR,
+ (errmsg("file of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel))));
+ if (OidIsValid(reltoastrelid))
+ {
+ Relation toastrel;
+
+ toastrel = table_open(reltoastrelid, AccessShareLock);
+ if (!RelFileLocatorEquals(toastrel->rd_locator,
+ repacked_rel_toast_locator))
+ ereport(ERROR,
+ (errmsg("file of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(toastrel))));
+ table_close(toastrel, AccessShareLock);
+ }
+
+ if (rel->rd_rel->relpersistence != cat_state->relpersistence)
+ ereport(ERROR,
+ errmsg("persistence of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel)));
+
+ if (cat_state->replident != rel->rd_rel->relreplident)
+ ereport(ERROR,
+ errmsg("replica identity of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel)));
+
+ ident_idx = RelationGetReplicaIndex(rel);
+ if (ident_idx == InvalidOid && rel->rd_pkindex != InvalidOid)
+ ident_idx = rel->rd_pkindex;
+ if (cat_state->replidindex != ident_idx)
+ ereport(ERROR,
+ errmsg("identity index of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel)));
+
+ /*
+ * As cat_state contains a copy (which has the constraint info cleared),
+ * create a temporary copy for the comparison.
+ */
+ td = RelationGetDescr(rel);
+ td_cp = palloc(TupleDescSize(td));
+ TupleDescCopy(td_cp, td);
+ if (!equalTupleDescs(cat_state->tupdesc, td_cp))
+ ereport(ERROR,
+ errmsg("definition of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel)));
+ FreeTupleDesc(td_cp);
+
+ /* Now we are only interested in indexes. */
+ if (cat_state->ninds == 0)
+ return;
+
+ /* No index should be dropped while we are checking the relation. */
+ lockmode = ShareUpdateExclusiveLock;
+ Assert(CheckRelationLockedByMe(rel, lockmode, true));
+
+ ind_oids = RelationGetIndexList(rel);
+ if (list_length(ind_oids) != cat_state->ninds)
+ goto failed_index;
+
+ foreach(lc, ind_oids)
+ {
+ Oid ind_oid = lfirst_oid(lc);
+ int i;
+ TupleDesc tupdesc;
+ Relation index;
+
+ /* Find the index in cat_state. */
+ for (i = 0; i < cat_state->ninds; i++)
+ {
+ if (cat_state->ind_oids[i] == ind_oid)
+ break;
+ }
+ /*
+ * OID not found, i.e. the index was replaced by another one. XXX
+ * Should we yet try to find if an index having the desired tuple
+ * descriptor exists? Or should we always look for the tuple
+ * descriptor and not use OIDs at all?
+ */
+ if (i == cat_state->ninds)
+ goto failed_index;
+
+ /* Check the tuple descriptor. */
+ index = try_index_open(ind_oid, lockmode);
+ if (index == NULL)
+ goto failed_index;
+ tupdesc = RelationGetDescr(index);
+ if (!equalTupleDescs(cat_state->ind_tupdescs[i], tupdesc))
+ goto failed_index;
+ index_close(index, lockmode);
+ }
+
+ return;
+
+failed_index:
+ ereport(ERROR,
+ (errmsg("index(es) of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel))));
+}
+
+/*
+ * This function is much like pg_create_logical_replication_slot() except that
+ * the new slot is neither released (if anyone else could read changes from
+ * our slot, we could miss changes other backends do while we copy the
+ * existing data into temporary table), nor persisted (it's easier to handle
+ * crash by restarting all the work from scratch).
+ */
+static LogicalDecodingContext *
+setup_logical_decoding(Oid relid, const char *slotname, TupleDesc tupdesc)
+{
+ LogicalDecodingContext *ctx;
+ RepackDecodingState *dstate;
+
+ /*
+ * Check if we can use logical decoding.
+ */
+ CheckSlotPermissions();
+ CheckLogicalDecodingRequirements();
+
+ /* RS_TEMPORARY so that the slot gets cleaned up on ERROR. */
+ ReplicationSlotCreate(slotname, true, RS_TEMPORARY, false, false, false);
+
+ /*
+ * Neither prepare_write nor do_write callback nor update_progress is
+ * useful for us.
+ *
+ * Regarding the value of need_full_snapshot, we pass false because the
+ * table we are processing is present in RepackedRelsHash and therefore,
+ * regarding logical decoding, treated like a catalog.
+ */
+ ctx = CreateInitDecodingContext(REPL_PLUGIN_NAME,
+ NIL,
+ false,
+ InvalidXLogRecPtr,
+ XL_ROUTINE(.page_read = read_local_xlog_page,
+ .segment_open = wal_segment_open,
+ .segment_close = wal_segment_close),
+ NULL, NULL, NULL);
+
+ /*
+ * We don't have control on setting fast_forward, so at least check it.
+ */
+ Assert(!ctx->fast_forward);
+
+ DecodingContextFindStartpoint(ctx);
+
+ /* Some WAL records should have been read. */
+ Assert(ctx->reader->EndRecPtr != InvalidXLogRecPtr);
+
+ XLByteToSeg(ctx->reader->EndRecPtr, repack_current_segment,
+ wal_segment_size);
+
+ /*
+ * Setup structures to store decoded changes.
+ */
+ dstate = palloc0(sizeof(RepackDecodingState));
+ dstate->relid = relid;
+ dstate->tstore = tuplestore_begin_heap(false, false,
+ maintenance_work_mem);
+ dstate->tupdesc = tupdesc;
+
+ /* Initialize the descriptor to store the changes ... */
+ dstate->tupdesc_change = CreateTemplateTupleDesc(1);
+
+ TupleDescInitEntry(dstate->tupdesc_change, 1, NULL, BYTEAOID, -1, 0);
+ /* ... as well as the corresponding slot. */
+ dstate->tsslot = MakeSingleTupleTableSlot(dstate->tupdesc_change,
+ &TTSOpsMinimalTuple);
+
+ dstate->resowner = ResourceOwnerCreate(CurrentResourceOwner,
+ "logical decoding");
+
+ ctx->output_writer_private = dstate;
+ return ctx;
+}
+
+/*
+ * Retrieve tuple from ConcurrentChange structure.
+ *
+ * The input data starts with the structure but it might not be appropriately
+ * aligned.
+ */
+static HeapTuple
+get_changed_tuple(char *change)
+{
+ HeapTupleData tup_data;
+ HeapTuple result;
+ char *src;
+
+ /*
+ * Ensure alignment before accessing the fields. (This is why we can't use
+ * heap_copytuple() instead of this function.)
+ */
+ src = change + offsetof(ConcurrentChange, tup_data);
+ memcpy(&tup_data, src, sizeof(HeapTupleData));
+
+ result = (HeapTuple) palloc(HEAPTUPLESIZE + tup_data.t_len);
+ memcpy(result, &tup_data, sizeof(HeapTupleData));
+ result->t_data = (HeapTupleHeader) ((char *) result + HEAPTUPLESIZE);
+ src = change + SizeOfConcurrentChange;
+ memcpy(result->t_data, src, result->t_len);
+
+ return result;
+}
+
+/*
+ * Decode logical changes from the WAL sequence up to end_of_wal.
+ */
+void
+repack_decode_concurrent_changes(LogicalDecodingContext *ctx,
+ XLogRecPtr end_of_wal)
+{
+ RepackDecodingState *dstate;
+ ResourceOwner resowner_old;
+ PgBackendProgress progress;
+
+ /*
+ * Invalidate the "present" cache before moving to "(recent) history".
+ */
+ InvalidateSystemCaches();
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+ resowner_old = CurrentResourceOwner;
+ CurrentResourceOwner = dstate->resowner;
+
+ /*
+ * reorderbuffer.c uses internal subtransaction, whose abort ends the
+ * command progress reporting. Save the status here so we can restore when
+ * done with the decoding.
+ */
+ memcpy(&progress, &MyBEEntry->st_progress, sizeof(PgBackendProgress));
+
+ PG_TRY();
+ {
+ while (ctx->reader->EndRecPtr < end_of_wal)
+ {
+ XLogRecord *record;
+ XLogSegNo segno_new;
+ char *errm = NULL;
+ XLogRecPtr end_lsn;
+
+ record = XLogReadRecord(ctx->reader, &errm);
+ if (errm)
+ elog(ERROR, "%s", errm);
+
+ if (record != NULL)
+ LogicalDecodingProcessRecord(ctx, ctx->reader);
+
+ /*
+ * If WAL segment boundary has been crossed, inform the decoding
+ * system that the catalog_xmin can advance. (We can confirm more
+ * often, but a filling a single WAL segment should not take much
+ * time.)
+ */
+ end_lsn = ctx->reader->EndRecPtr;
+ XLByteToSeg(end_lsn, segno_new, wal_segment_size);
+ if (segno_new != repack_current_segment)
+ {
+ LogicalConfirmReceivedLocation(end_lsn);
+ elog(DEBUG1, "REPACK: confirmed receive location %X/%X",
+ (uint32) (end_lsn >> 32), (uint32) end_lsn);
+ repack_current_segment = segno_new;
+ }
+
+ CHECK_FOR_INTERRUPTS();
+ }
+ InvalidateSystemCaches();
+ CurrentResourceOwner = resowner_old;
+ }
+ PG_CATCH();
+ {
+ InvalidateSystemCaches();
+ CurrentResourceOwner = resowner_old;
+ PG_RE_THROW();
+ }
+ PG_END_TRY();
+
+ /* Restore the progress reporting status. */
+ pgstat_progress_restore_state(&progress);
+}
+
+/*
+ * Apply changes that happened during the initial load.
+ *
+ * Scan key is passed by caller, so it does not have to be constructed
+ * multiple times. Key entries have all fields initialized, except for
+ * sk_argument.
+ */
+static void
+apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
+ ScanKey key, int nkeys, IndexInsertState *iistate)
+{
+ TupleTableSlot *index_slot, *ident_slot;
+ HeapTuple tup_old = NULL;
+
+ if (dstate->nchanges == 0)
+ return;
+
+ /* TupleTableSlot is needed to pass the tuple to ExecInsertIndexTuples(). */
+ index_slot = MakeSingleTupleTableSlot(dstate->tupdesc, &TTSOpsHeapTuple);
+ iistate->econtext->ecxt_scantuple = index_slot;
+
+ /* A slot to fetch tuples from identity index. */
+ ident_slot = table_slot_create(rel, NULL);
+
+ while (tuplestore_gettupleslot(dstate->tstore, true, false,
+ dstate->tsslot))
+ {
+ bool shouldFree;
+ HeapTuple tup_change,
+ tup,
+ tup_exist;
+ char *change_raw, *src;
+ ConcurrentChange change;
+ bool isnull[1];
+ Datum values[1];
+
+ CHECK_FOR_INTERRUPTS();
+
+ /* Get the change from the single-column tuple. */
+ tup_change = ExecFetchSlotHeapTuple(dstate->tsslot, false, &shouldFree);
+ heap_deform_tuple(tup_change, dstate->tupdesc_change, values, isnull);
+ Assert(!isnull[0]);
+
+ /* Make sure we access aligned data. */
+ change_raw = (char *) DatumGetByteaP(values[0]);
+ src = (char *) VARDATA(change_raw);
+ memcpy(&change, src, SizeOfConcurrentChange);
+
+ /* TRUNCATE change contains no tuple, so process it separately. */
+ if (change.kind == CHANGE_TRUNCATE)
+ {
+ /*
+ * All the things that ExecuteTruncateGuts() does (such as firing
+ * triggers or handling the DROP_CASCADE behavior) should have
+ * taken place on the source relation. Thus we only do the actual
+ * truncation of the new relation (and its indexes).
+ */
+ heap_truncate_one_rel(rel);
+
+ pfree(tup_change);
+ continue;
+ }
+
+ /*
+ * Extract the tuple from the change. The tuple is copied here because
+ * it might be assigned to 'tup_old', in which case it needs to
+ * survive into the next iteration.
+ */
+ tup = get_changed_tuple(src);
+
+ if (change.kind == CHANGE_UPDATE_OLD)
+ {
+ Assert(tup_old == NULL);
+ tup_old = tup;
+ }
+ else if (change.kind == CHANGE_INSERT)
+ {
+ Assert(tup_old == NULL);
+
+ apply_concurrent_insert(rel, &change, tup, iistate, index_slot);
+
+ pfree(tup);
+ }
+ else if (change.kind == CHANGE_UPDATE_NEW ||
+ change.kind == CHANGE_DELETE)
+ {
+ IndexScanDesc ind_scan = NULL;
+ HeapTuple tup_key;
+
+ if (change.kind == CHANGE_UPDATE_NEW)
+ {
+ tup_key = tup_old != NULL ? tup_old : tup;
+ }
+ else
+ {
+ Assert(tup_old == NULL);
+ tup_key = tup;
+ }
+
+ /*
+ * Find the tuple to be updated or deleted.
+ */
+ tup_exist = find_target_tuple(rel, key, nkeys, tup_key,
+ iistate, ident_slot, &ind_scan);
+ if (tup_exist == NULL)
+ elog(ERROR, "Failed to find target tuple");
+
+ if (change.kind == CHANGE_UPDATE_NEW)
+ apply_concurrent_update(rel, tup, tup_exist, &change, iistate,
+ index_slot);
+ else
+ apply_concurrent_delete(rel, tup_exist, &change);
+
+ if (tup_old != NULL)
+ {
+ pfree(tup_old);
+ tup_old = NULL;
+ }
+
+ pfree(tup);
+ index_endscan(ind_scan);
+ }
+ else
+ elog(ERROR, "Unrecognized kind of change: %d", change.kind);
+
+ /* If there's any change, make it visible to the next iteration. */
+ if (change.kind != CHANGE_UPDATE_OLD)
+ {
+ CommandCounterIncrement();
+ UpdateActiveSnapshotCommandId();
+ }
+
+ /* TTSOpsMinimalTuple has .get_heap_tuple==NULL. */
+ Assert(shouldFree);
+ pfree(tup_change);
+ }
+
+ tuplestore_clear(dstate->tstore);
+ dstate->nchanges = 0;
+
+ /* Cleanup. */
+ ExecDropSingleTupleTableSlot(index_slot);
+ ExecDropSingleTupleTableSlot(ident_slot);
+}
+
+static void
+apply_concurrent_insert(Relation rel, ConcurrentChange *change, HeapTuple tup,
+ IndexInsertState *iistate, TupleTableSlot *index_slot)
+{
+ List *recheck;
+
+
+ heap_insert(rel, tup, GetCurrentCommandId(true), HEAP_INSERT_NO_LOGICAL, NULL);
+
+ /*
+ * Update indexes.
+ *
+ * In case functions in the index need the active snapshot and caller
+ * hasn't set one.
+ */
+ ExecStoreHeapTuple(tup, index_slot, false);
+ recheck = ExecInsertIndexTuples(iistate->rri,
+ index_slot,
+ iistate->estate,
+ false, /* update */
+ false, /* noDupErr */
+ NULL, /* specConflict */
+ NIL, /* arbiterIndexes */
+ false /* onlySummarizing */
+ );
+
+ /*
+ * If recheck is required, it must have been preformed on the source
+ * relation by now. (All the logical changes we process here are already
+ * committed.)
+ */
+ list_free(recheck);
+
+ pgstat_progress_incr_param(PROGRESS_REPACK_HEAP_TUPLES_INSERTED, 1);
+}
+
+static void
+apply_concurrent_update(Relation rel, HeapTuple tup, HeapTuple tup_target,
+ ConcurrentChange *change, IndexInsertState *iistate,
+ TupleTableSlot *index_slot)
+{
+ List *recheck;
+ TU_UpdateIndexes update_indexes;
+
+ /*
+ * Write the new tuple into the new heap. ('tup' gets the TID assigned
+ * here.)
+ */
+ simple_heap_update(rel, &tup_target->t_self, tup, &update_indexes);
+
+ ExecStoreHeapTuple(tup, index_slot, false);
+
+ if (update_indexes != TU_None)
+ {
+ recheck = ExecInsertIndexTuples(iistate->rri,
+ index_slot,
+ iistate->estate,
+ true, /* update */
+ false, /* noDupErr */
+ NULL, /* specConflict */
+ NIL, /* arbiterIndexes */
+ /* onlySummarizing */
+ update_indexes == TU_Summarizing);
+ list_free(recheck);
+ }
+
+ pgstat_progress_incr_param(PROGRESS_REPACK_HEAP_TUPLES_UPDATED, 1);
+}
+
+static void
+apply_concurrent_delete(Relation rel, HeapTuple tup_target,
+ ConcurrentChange *change)
+{
+ simple_heap_delete(rel, &tup_target->t_self);
+
+ pgstat_progress_incr_param(PROGRESS_REPACK_HEAP_TUPLES_DELETED, 1);
+}
+
+/*
+ * Find the tuple to be updated or deleted.
+ *
+ * 'key' is a pre-initialized scan key, into which the function will put the
+ * key values.
+ *
+ * 'tup_key' is a tuple containing the key values for the scan.
+ *
+ * On exit,'*scan_p' contains the scan descriptor used. The caller must close
+ * it when he no longer needs the tuple returned.
+ */
+static HeapTuple
+find_target_tuple(Relation rel, ScanKey key, int nkeys, HeapTuple tup_key,
+ IndexInsertState *iistate,
+ TupleTableSlot *ident_slot, IndexScanDesc *scan_p)
+{
+ IndexScanDesc scan;
+ Form_pg_index ident_form;
+ int2vector *ident_indkey;
+ HeapTuple result = NULL;
+
+ scan = index_beginscan(rel, iistate->ident_index, GetActiveSnapshot(),
+ nkeys, 0);
+ *scan_p = scan;
+ index_rescan(scan, key, nkeys, NULL, 0);
+
+ /* Info needed to retrieve key values from heap tuple. */
+ ident_form = iistate->ident_index->rd_index;
+ ident_indkey = &ident_form->indkey;
+
+ /* Use the incoming tuple to finalize the scan key. */
+ for (int i = 0; i < scan->numberOfKeys; i++)
+ {
+ ScanKey entry;
+ bool isnull;
+ int16 attno_heap;
+
+ entry = &scan->keyData[i];
+ attno_heap = ident_indkey->values[i];
+ entry->sk_argument = heap_getattr(tup_key,
+ attno_heap,
+ rel->rd_att,
+ &isnull);
+ Assert(!isnull);
+ }
+ if (index_getnext_slot(scan, ForwardScanDirection, ident_slot))
+ {
+ bool shouldFree;
+
+ result = ExecFetchSlotHeapTuple(ident_slot, false, &shouldFree);
+ /* TTSOpsBufferHeapTuple has .get_heap_tuple != NULL. */
+ Assert(!shouldFree);
+ }
+
+ return result;
+}
+
+/*
+ * Decode and apply concurrent changes.
+ *
+ * Pass rel_src iff its reltoastrelid is needed.
+ */
+static void
+process_concurrent_changes(LogicalDecodingContext *ctx, XLogRecPtr end_of_wal,
+ Relation rel_dst, Relation rel_src, ScanKey ident_key,
+ int ident_key_nentries, IndexInsertState *iistate)
+{
+ RepackDecodingState *dstate;
+
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_CATCH_UP);
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+
+ repack_decode_concurrent_changes(ctx, end_of_wal);
+
+ if (dstate->nchanges == 0)
+ return;
+
+ PG_TRY();
+ {
+ /*
+ * Make sure that TOAST values can eventually be accessed via the old
+ * relation - see comment in copy_table_data().
+ */
+ if (rel_src)
+ rel_dst->rd_toastoid = rel_src->rd_rel->reltoastrelid;
+
+ apply_concurrent_changes(dstate, rel_dst, ident_key,
+ ident_key_nentries, iistate);
+ }
+ PG_FINALLY();
+ {
+ if (rel_src)
+ rel_dst->rd_toastoid = InvalidOid;
+ }
+ PG_END_TRY();
+}
+
+static IndexInsertState *
+get_index_insert_state(Relation relation, Oid ident_index_id)
+{
+ EState *estate;
+ int i;
+ IndexInsertState *result;
+
+ result = (IndexInsertState *) palloc0(sizeof(IndexInsertState));
+ estate = CreateExecutorState();
+ result->econtext = GetPerTupleExprContext(estate);
+
+ result->rri = (ResultRelInfo *) palloc(sizeof(ResultRelInfo));
+ InitResultRelInfo(result->rri, relation, 0, 0, 0);
+ ExecOpenIndices(result->rri, false);
+
+ /*
+ * Find the relcache entry of the identity index so that we spend no extra
+ * effort to open / close it.
+ */
+ for (i = 0; i < result->rri->ri_NumIndices; i++)
+ {
+ Relation ind_rel;
+
+ ind_rel = result->rri->ri_IndexRelationDescs[i];
+ if (ind_rel->rd_id == ident_index_id)
+ result->ident_index = ind_rel;
+ }
+ if (result->ident_index == NULL)
+ elog(ERROR, "Failed to open identity index");
+
+ /* Only initialize fields needed by ExecInsertIndexTuples(). */
+ result->estate = estate;
+
+ return result;
+}
+
+/*
+ * Build scan key to process logical changes.
+ */
+static ScanKey
+build_identity_key(Oid ident_idx_oid, Relation rel_src, int *nentries)
+{
+ Relation ident_idx_rel;
+ Form_pg_index ident_idx;
+ int n,
+ i;
+ ScanKey result;
+
+ Assert(OidIsValid(ident_idx_oid));
+ ident_idx_rel = index_open(ident_idx_oid, AccessShareLock);
+ ident_idx = ident_idx_rel->rd_index;
+ n = ident_idx->indnatts;
+ result = (ScanKey) palloc(sizeof(ScanKeyData) * n);
+ for (i = 0; i < n; i++)
+ {
+ ScanKey entry;
+ int16 relattno;
+ Form_pg_attribute att;
+ Oid opfamily,
+ opcintype,
+ opno,
+ opcode;
+
+ entry = &result[i];
+ relattno = ident_idx->indkey.values[i];
+ if (relattno >= 1)
+ {
+ TupleDesc desc;
+
+ desc = rel_src->rd_att;
+ att = TupleDescAttr(desc, relattno - 1);
+ }
+ else
+ elog(ERROR, "Unexpected attribute number %d in index", relattno);
+
+ opfamily = ident_idx_rel->rd_opfamily[i];
+ opcintype = ident_idx_rel->rd_opcintype[i];
+ opno = get_opfamily_member(opfamily, opcintype, opcintype,
+ BTEqualStrategyNumber);
+
+ if (!OidIsValid(opno))
+ elog(ERROR, "Failed to find = operator for type %u", opcintype);
+
+ opcode = get_opcode(opno);
+ if (!OidIsValid(opcode))
+ elog(ERROR, "Failed to find = operator for operator %u", opno);
+
+ /* Initialize everything but argument. */
+ ScanKeyInit(entry,
+ i + 1,
+ BTEqualStrategyNumber, opcode,
+ (Datum) NULL);
+ entry->sk_collation = att->attcollation;
+ }
+ index_close(ident_idx_rel, AccessShareLock);
+
+ *nentries = n;
+ return result;
+}
+
+static void
+free_index_insert_state(IndexInsertState *iistate)
+{
+ ExecCloseIndices(iistate->rri);
+ FreeExecutorState(iistate->estate);
+ pfree(iistate->rri);
+ pfree(iistate);
+}
+
+static void
+cleanup_logical_decoding(LogicalDecodingContext *ctx)
+{
+ RepackDecodingState *dstate;
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+
+ ExecDropSingleTupleTableSlot(dstate->tsslot);
+ FreeTupleDesc(dstate->tupdesc_change);
+ FreeTupleDesc(dstate->tupdesc);
+ tuplestore_end(dstate->tstore);
+
+ FreeDecodingContext(ctx);
+}
+
+/*
+ * The final steps of rebuild_relation() for concurrent processing.
+ *
+ * On entry, NewHeap is locked in AccessExclusiveLock mode. OldHeap and its
+ * clustering index (if one is passed) are still locked in a mode that allows
+ * concurrent data changes. On exit, both tables and their indexes are closed,
+ * but locked in AccessExclusiveLock mode.
+ */
+static void
+rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
+ Relation cl_index,
+ CatalogState *cat_state,
+ LogicalDecodingContext *ctx,
+ bool swap_toast_by_content,
+ TransactionId frozenXid,
+ MultiXactId cutoffMulti)
+{
+ LOCKMODE lockmode_old PG_USED_FOR_ASSERTS_ONLY;
+ List *ind_oids_new;
+ Oid old_table_oid = RelationGetRelid(OldHeap);
+ Oid new_table_oid = RelationGetRelid(NewHeap);
+ List *ind_oids_old = RelationGetIndexList(OldHeap);
+ ListCell *lc, *lc2;
+ char relpersistence;
+ bool is_system_catalog;
+ Oid ident_idx_old, ident_idx_new;
+ IndexInsertState *iistate;
+ ScanKey ident_key;
+ int ident_key_nentries;
+ XLogRecPtr wal_insert_ptr, end_of_wal;
+ char dummy_rec_data = '\0';
+ RelReopenInfo *rri = NULL;
+ int nrel;
+ Relation *ind_refs_all, *ind_refs_p;
+
+ /* Like in cluster_rel(). */
+ lockmode_old = ShareUpdateExclusiveLock;
+ Assert(CheckRelationLockedByMe(OldHeap, lockmode_old, false));
+ Assert(cl_index == NULL ||
+ CheckRelationLockedByMe(cl_index, lockmode_old, false));
+ /* This is expected from the caller. */
+ Assert(CheckRelationLockedByMe(NewHeap, AccessExclusiveLock, false));
+
+ ident_idx_old = RelationGetReplicaIndex(OldHeap);
+
+ /*
+ * Unlike the exclusive case, we build new indexes for the new relation
+ * rather than swapping the storage and reindexing the old relation. The
+ * point is that the index build can take some time, so we do it before we
+ * get AccessExclusiveLock on the old heap and therefore we cannot swap
+ * the heap storage yet.
+ *
+ * index_create() will lock the new indexes using AccessExclusiveLock
+ * creation - no need to change that.
+ */
+ ind_oids_new = build_new_indexes(NewHeap, OldHeap, ind_oids_old);
+
+ /*
+ * Processing shouldn't start w/o valid identity index.
+ */
+ Assert(OidIsValid(ident_idx_old));
+
+ /* Find "identity index" on the new relation. */
+ ident_idx_new = InvalidOid;
+ forboth(lc, ind_oids_old, lc2, ind_oids_new)
+ {
+ Oid ind_old = lfirst_oid(lc);
+ Oid ind_new = lfirst_oid(lc2);
+
+ if (ident_idx_old == ind_old)
+ {
+ ident_idx_new = ind_new;
+ break;
+ }
+ }
+ if (!OidIsValid(ident_idx_new))
+ /*
+ * Should not happen, given our lock on the old relation.
+ */
+ ereport(ERROR,
+ (errmsg("Identity index missing on the new relation")));
+
+ /* Executor state to update indexes. */
+ iistate = get_index_insert_state(NewHeap, ident_idx_new);
+
+ /*
+ * Build scan key that we'll use to look for rows to be updated / deleted
+ * during logical decoding.
+ */
+ ident_key = build_identity_key(ident_idx_new, OldHeap, &ident_key_nentries);
+
+ /*
+ * Flush all WAL records inserted so far (possibly except for the last
+ * incomplete page, see GetInsertRecPtr), to minimize the amount of data
+ * we need to flush while holding exclusive lock on the source table.
+ */
+ wal_insert_ptr = GetInsertRecPtr();
+ XLogFlush(wal_insert_ptr);
+ end_of_wal = GetFlushRecPtr(NULL);
+
+ /*
+ * Apply concurrent changes first time, to minimize the time we need to
+ * hold AccessExclusiveLock. (Quite some amount of WAL could have been
+ * written during the data copying and index creation.)
+ */
+ process_concurrent_changes(ctx, end_of_wal, NewHeap,
+ swap_toast_by_content ? OldHeap : NULL,
+ ident_key, ident_key_nentries, iistate);
+
+ /*
+ * Release the locks that allowed concurrent data changes, in order to
+ * acquire the AccessExclusiveLock.
+ */
+ nrel = 0;
+ /*
+ * We unlock the old relation (and its clustering index), but then we will
+ * lock the relation and *all* its indexes because we want to swap their
+ * storage.
+ *
+ * (NewHeap is already locked, as well as its indexes.)
+ */
+ rri = palloc_array(RelReopenInfo, 1 + list_length(ind_oids_old));
+ init_rel_reopen_info(&rri[nrel++], &OldHeap, InvalidOid,
+ ShareUpdateExclusiveLock, AccessExclusiveLock);
+ /* References to the re-opened indexes will be stored in this array. */
+ ind_refs_all = palloc_array(Relation, list_length(ind_oids_old));
+ ind_refs_p = ind_refs_all;
+ /* The clustering index is a special case. */
+ if (cl_index)
+ {
+ *ind_refs_p = cl_index;
+ init_rel_reopen_info(&rri[nrel], ind_refs_p, InvalidOid,
+ ShareUpdateExclusiveLock, AccessExclusiveLock);
+ nrel++;
+ ind_refs_p++;
+ }
+ /*
+ * Initialize also the entries for the other indexes (currently unlocked)
+ * because we will have to lock them.
+ */
+ foreach(lc, ind_oids_old)
+ {
+ Oid ind_oid;
+
+ ind_oid = lfirst_oid(lc);
+ /* Clustering index is already in the array, or there is none. */
+ if (cl_index && RelationGetRelid(cl_index) == ind_oid)
+ continue;
+
+ Assert(nrel < (1 + list_length(ind_oids_old)));
+
+ *ind_refs_p = NULL;
+ init_rel_reopen_info(&rri[nrel],
+ /*
+ * In this special case we do not have the
+ * relcache reference, use OID instead.
+ */
+ ind_refs_p,
+ ind_oid,
+ NoLock, /* Nothing to unlock. */
+ AccessExclusiveLock);
+
+ nrel++;
+ ind_refs_p++;
+ }
+ /* Perform the actual unlocking and re-locking. */
+ unlock_and_close_relations(rri, nrel);
+ reopen_relations(rri, nrel);
+
+ /*
+ * In addition, lock the OldHeap's TOAST relation that we skipped for the
+ * CONCURRENTLY option in copy_table_data(). This lock will be needed to
+ * swap the relation files.
+ */
+ if (OidIsValid(OldHeap->rd_rel->reltoastrelid))
+ LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
+
+ /*
+ * Check if the new indexes match the old ones, i.e. no changes occurred
+ * while OldHeap was unlocked.
+ *
+ * XXX It's probably not necessary to check the relation tuple descriptor
+ * here because the logical decoding was already active when we released
+ * the lock, and thus the corresponding data changes won't be lost.
+ * However processing of those changes might take a lot of time.
+ */
+ check_catalog_changes(OldHeap, cat_state);
+
+ /*
+ * Tuples and pages of the old heap will be gone, but the heap will stay.
+ */
+ TransferPredicateLocksToHeapRelation(OldHeap);
+ /* The same for indexes. */
+ for (int i = 0; i < (nrel - 1); i++)
+ {
+ Relation index = ind_refs_all[i];
+
+ TransferPredicateLocksToHeapRelation(index);
+
+ /*
+ * References to indexes on the old relation are not needed anymore,
+ * however locks stay till the end of the transaction.
+ */
+ index_close(index, NoLock);
+ }
+ pfree(ind_refs_all);
+
+ /*
+ * Flush anything we see in WAL, to make sure that all changes committed
+ * while we were waiting for the exclusive lock are available for
+ * decoding. This should not be necessary if all backends had
+ * synchronous_commit set, but we can't rely on this setting.
+ *
+ * Unfortunately, GetInsertRecPtr() may lag behind the actual insert
+ * position, and GetLastImportantRecPtr() points at the start of the last
+ * record rather than at the end. Thus the simplest way to determine the
+ * insert position is to insert a dummy record and use its LSN.
+ *
+ * XXX Consider using GetLastImportantRecPtr() and adding the size of the
+ * last record (plus the total size of all the page headers the record
+ * spans)?
+ */
+ XLogBeginInsert();
+ XLogRegisterData(&dummy_rec_data, 1);
+ wal_insert_ptr = XLogInsert(RM_XLOG_ID, XLOG_NOOP);
+ XLogFlush(wal_insert_ptr);
+ end_of_wal = GetFlushRecPtr(NULL);
+
+ /* Apply the concurrent changes again. */
+ process_concurrent_changes(ctx, end_of_wal, NewHeap,
+ swap_toast_by_content ? OldHeap : NULL,
+ ident_key, ident_key_nentries, iistate);
+
+ /* Remember info about rel before closing OldHeap */
+ relpersistence = OldHeap->rd_rel->relpersistence;
+ is_system_catalog = IsSystemRelation(OldHeap);
+
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_SWAP_REL_FILES);
+
+ forboth(lc, ind_oids_old, lc2, ind_oids_new)
+ {
+ Oid ind_old = lfirst_oid(lc);
+ Oid ind_new = lfirst_oid(lc2);
+ Oid mapped_tables[4];
+
+ /* Zero out possible results from swapped_relation_files */
+ memset(mapped_tables, 0, sizeof(mapped_tables));
+
+ swap_relation_files(ind_old, ind_new,
+ (old_table_oid == RelationRelationId),
+ swap_toast_by_content,
+ true,
+ InvalidTransactionId,
+ InvalidMultiXactId,
+ mapped_tables);
+
+#ifdef USE_ASSERT_CHECKING
+ /*
+ * Concurrent processing is not supported for system relations, so
+ * there should be no mapped tables.
+ */
+ for (int i = 0; i < 4; i++)
+ Assert(mapped_tables[i] == 0);
+#endif
+ }
+
+ /* The new indexes must be visible for deletion. */
+ CommandCounterIncrement();
+
+ /* Close the old heap but keep lock until transaction commit. */
+ table_close(OldHeap, NoLock);
+ /* Close the new heap. (We didn't have to open its indexes). */
+ table_close(NewHeap, NoLock);
+
+ /* Cleanup what we don't need anymore. (And close the identity index.) */
+ pfree(ident_key);
+ free_index_insert_state(iistate);
+
+ /*
+ * Swap the relations and their TOAST relations and TOAST indexes. This
+ * also drops the new relation and its indexes.
+ *
+ * (System catalogs are currently not supported.)
+ */
+ Assert(!is_system_catalog);
+ finish_heap_swap(old_table_oid, new_table_oid,
+ is_system_catalog,
+ swap_toast_by_content,
+ false, true, false,
+ frozenXid, cutoffMulti,
+ relpersistence);
+
+ pfree(rri);
+}
+
+/*
+ * Build indexes on NewHeap according to those on OldHeap.
+ *
+ * OldIndexes is the list of index OIDs on OldHeap.
+ *
+ * A list of OIDs of the corresponding indexes created on NewHeap is
+ * returned. The order of items does match, so we can use these arrays to swap
+ * index storage.
+ */
+static List *
+build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes)
+{
+ StringInfo ind_name;
+ ListCell *lc;
+ List *result = NIL;
+
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_REBUILD_INDEX);
+
+ ind_name = makeStringInfo();
+
+ foreach(lc, OldIndexes)
+ {
+ Oid ind_oid,
+ ind_oid_new,
+ tbsp_oid;
+ Relation ind;
+ IndexInfo *ind_info;
+ int i,
+ heap_col_id;
+ List *colnames;
+ int16 indnatts;
+ Oid *collations,
+ *opclasses;
+ HeapTuple tup;
+ bool isnull;
+ Datum d;
+ oidvector *oidvec;
+ int2vector *int2vec;
+ size_t oid_arr_size;
+ size_t int2_arr_size;
+ int16 *indoptions;
+ text *reloptions = NULL;
+ bits16 flags;
+ Datum *opclassOptions;
+ NullableDatum *stattargets;
+
+ ind_oid = lfirst_oid(lc);
+ ind = index_open(ind_oid, AccessShareLock);
+ ind_info = BuildIndexInfo(ind);
+
+ tbsp_oid = ind->rd_rel->reltablespace;
+ /*
+ * Index name really doesn't matter, we'll eventually use only their
+ * storage. Just make them unique within the table.
+ */
+ resetStringInfo(ind_name);
+ appendStringInfo(ind_name, "ind_%d",
+ list_cell_number(OldIndexes, lc));
+
+ flags = 0;
+ if (ind->rd_index->indisprimary)
+ flags |= INDEX_CREATE_IS_PRIMARY;
+
+ colnames = NIL;
+ indnatts = ind->rd_index->indnatts;
+ oid_arr_size = sizeof(Oid) * indnatts;
+ int2_arr_size = sizeof(int16) * indnatts;
+
+ collations = (Oid *) palloc(oid_arr_size);
+ for (i = 0; i < indnatts; i++)
+ {
+ char *colname;
+
+ heap_col_id = ind->rd_index->indkey.values[i];
+ if (heap_col_id > 0)
+ {
+ Form_pg_attribute att;
+
+ /* Normal attribute. */
+ att = TupleDescAttr(OldHeap->rd_att, heap_col_id - 1);
+ colname = pstrdup(NameStr(att->attname));
+ collations[i] = att->attcollation;
+ }
+ else if (heap_col_id == 0)
+ {
+ HeapTuple tuple;
+ Form_pg_attribute att;
+
+ /*
+ * Expression column is not present in relcache. What we need
+ * here is an attribute of the *index* relation.
+ */
+ tuple = SearchSysCache2(ATTNUM,
+ ObjectIdGetDatum(ind_oid),
+ Int16GetDatum(i + 1));
+ if (!HeapTupleIsValid(tuple))
+ elog(ERROR,
+ "cache lookup failed for attribute %d of relation %u",
+ i + 1, ind_oid);
+ att = (Form_pg_attribute) GETSTRUCT(tuple);
+ colname = pstrdup(NameStr(att->attname));
+ collations[i] = att->attcollation;
+ ReleaseSysCache(tuple);
+ }
+ else
+ elog(ERROR, "Unexpected column number: %d",
+ heap_col_id);
+
+ colnames = lappend(colnames, colname);
+ }
+
+ /*
+ * Special effort needed for variable length attributes of
+ * Form_pg_index.
+ */
+ tup = SearchSysCache1(INDEXRELID, ObjectIdGetDatum(ind_oid));
+ if (!HeapTupleIsValid(tup))
+ elog(ERROR, "cache lookup failed for index %u", ind_oid);
+ d = SysCacheGetAttr(INDEXRELID, tup, Anum_pg_index_indclass, &isnull);
+ Assert(!isnull);
+ oidvec = (oidvector *) DatumGetPointer(d);
+ opclasses = (Oid *) palloc(oid_arr_size);
+ memcpy(opclasses, oidvec->values, oid_arr_size);
+
+ d = SysCacheGetAttr(INDEXRELID, tup, Anum_pg_index_indoption,
+ &isnull);
+ Assert(!isnull);
+ int2vec = (int2vector *) DatumGetPointer(d);
+ indoptions = (int16 *) palloc(int2_arr_size);
+ memcpy(indoptions, int2vec->values, int2_arr_size);
+ ReleaseSysCache(tup);
+
+ tup = SearchSysCache1(RELOID, ObjectIdGetDatum(ind_oid));
+ if (!HeapTupleIsValid(tup))
+ elog(ERROR, "cache lookup failed for index relation %u", ind_oid);
+ d = SysCacheGetAttr(RELOID, tup, Anum_pg_class_reloptions, &isnull);
+ reloptions = !isnull ? DatumGetTextPCopy(d) : NULL;
+ ReleaseSysCache(tup);
+
+ opclassOptions = palloc0(sizeof(Datum) * ind_info->ii_NumIndexAttrs);
+ for (i = 0; i < ind_info->ii_NumIndexAttrs; i++)
+ opclassOptions[i] = get_attoptions(ind_oid, i + 1);
+
+ stattargets = get_index_stattargets(ind_oid, ind_info);
+
+ /*
+ * Neither parentIndexRelid nor parentConstraintId needs to be passed
+ * since the new catalog entries (pg_constraint, pg_inherits) would
+ * eventually be dropped. Therefore there's no need to record valid
+ * dependency on parents.
+ */
+ ind_oid_new = index_create(NewHeap,
+ ind_name->data,
+ InvalidOid,
+ InvalidOid, /* parentIndexRelid */
+ InvalidOid, /* parentConstraintId */
+ InvalidOid,
+ ind_info,
+ colnames,
+ ind->rd_rel->relam,
+ tbsp_oid,
+ collations,
+ opclasses,
+ opclassOptions,
+ indoptions,
+ stattargets,
+ PointerGetDatum(reloptions),
+ flags, /* flags */
+ 0, /* constr_flags */
+ false, /* allow_system_table_mods */
+ false, /* is_internal */
+ NULL /* constraintId */
+ );
+ result = lappend_oid(result, ind_oid_new);
+
+ index_close(ind, AccessShareLock);
+ list_free_deep(colnames);
+ pfree(collations);
+ pfree(opclasses);
+ pfree(indoptions);
+ if (reloptions)
+ pfree(reloptions);
+ }
+
+ return result;
+}
+
+static void
+init_rel_reopen_info(RelReopenInfo *rri, Relation *rel_p, Oid relid,
+ LOCKMODE lockmode_orig, LOCKMODE lockmode_new)
+{
+ rri->rel_p = rel_p;
+ rri->relid = relid;
+ rri->lockmode_orig = lockmode_orig;
+ rri->lockmode_new = lockmode_new;
+}
+
+/*
+ * Unlock and close relations specified by items of the 'rels' array. 'nrels'
+ * is the number of items.
+ *
+ * Information needed to (re)open the relations (or to issue meaningful ERROR)
+ * is added to the array items.
+ */
+static void
+unlock_and_close_relations(RelReopenInfo *rels, int nrel)
+{
+ int i;
+ RelReopenInfo *rri;
+
+ /*
+ * First, retrieve the information that we will need for re-opening.
+ *
+ * We could close (and unlock) each relation as soon as we have gathered
+ * the related information, but then we would have to be careful not to
+ * unlock the table until we have the info on all its indexes. (Once we
+ * unlock the table, any index can be dropped, and thus we can fail to get
+ * the name we want to report if re-opening fails.) It seem simpler to
+ * separate the work into two iterations.
+ */
+ for (i = 0; i < nrel; i++)
+ {
+ Relation rel;
+
+ rri = &rels[i];
+ rel = *rri->rel_p;
+
+ if (rel)
+ {
+ Assert(CheckRelationLockedByMe(rel, rri->lockmode_orig, false));
+ Assert(!OidIsValid(rri->relid));
+
+ rri->relid = RelationGetRelid(rel);
+ rri->relkind = rel->rd_rel->relkind;
+ rri->relname = pstrdup(RelationGetRelationName(rel));
+ }
+ else
+ {
+ Assert(OidIsValid(rri->relid));
+
+ rri->relname = get_rel_name(rri->relid);
+ rri->relkind = get_rel_relkind(rri->relid);
+ }
+ }
+
+ /* Second, close the relations. */
+ for (i = 0; i < nrel; i++)
+ {
+ Relation rel;
+
+ rri = &rels[i];
+ rel = *rri->rel_p;
+
+ /* Close the relation if the caller passed one. */
+ if (rel)
+ {
+ if (rri->relkind == RELKIND_RELATION)
+ table_close(rel, rri->lockmode_orig);
+ else
+ {
+ Assert(rri->relkind == RELKIND_INDEX);
+
+ index_close(rel, rri->lockmode_orig);
+ }
+ }
+ }
+}
+
+/*
+ * Re-open the relations closed previously by unlock_and_close_relations().
+ */
+static void
+reopen_relations(RelReopenInfo *rels, int nrel)
+{
+ for (int i = 0; i < nrel; i++)
+ {
+ RelReopenInfo *rri = &rels[i];
+ Relation rel;
+
+ if (rri->relkind == RELKIND_RELATION)
+ {
+ rel = try_table_open(rri->relid, rri->lockmode_new);
+ }
+ else
+ {
+ Assert(rri->relkind == RELKIND_INDEX);
+
+ rel = try_index_open(rri->relid, rri->lockmode_new);
+ }
+
+ if (rel == NULL)
+ {
+ const char *kind_str;
+
+ kind_str = (rri->relkind == RELKIND_RELATION) ? "table" : "index";
+ ereport(ERROR,
+ (errmsg("could not open \%s \"%s\"", kind_str,
+ rri->relname),
+ errhint("The %s could have been dropped by another transaction.",
+ kind_str)));
+ }
+ *rri->rel_p = rel;
+
+ pfree(rri->relname);
+ }
+}
+
+/*
+ * REPACK is intended to be a replacement of both CLUSTER and VACUUM FULL.
+ */
+void
+repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
+{
+ ListCell *lc;
+ ClusterParams params = {0};
+ bool verbose = false;
+ Relation rel = NULL;
+ Oid indexOid = InvalidOid;
+ MemoryContext repack_context;
+ List *rtcs;
+ LOCKMODE lockmode;
+
+ /* Parse option list */
+ foreach(lc, stmt->params)
+ {
+ DefElem *opt = (DefElem *) lfirst(lc);
+
+ if (strcmp(opt->defname, "verbose") == 0)
+ verbose = defGetBoolean(opt);
+ else
+ ereport(ERROR,
+ (errcode(ERRCODE_SYNTAX_ERROR),
+ errmsg("unrecognized REPACK option \"%s\"",
+ opt->defname),
+ parser_errposition(pstate, opt->location)));
+ }
+
+ params.options =
+ (verbose ? CLUOPT_VERBOSE : 0) |
+ (stmt->concurrent ? CLUOPT_CONCURRENT : 0);
+
+ /*
+ * Determine the lock mode expected by cluster_rel().
+ *
+ * In the exclusive case, we obtain AccessExclusiveLock right away to
+ * avoid lock-upgrade hazard in the single-transaction case. In the
+ * CONCURRENTLY case, the AccessExclusiveLock will only be used at the end
+ * of processing, supposedly for very short time. Until then, we'll have
+ * to unlock the relation temporarily, so there's no lock-upgrade hazard.
+ */
+ lockmode = (params.options & CLUOPT_CONCURRENT) == 0 ?
+ AccessExclusiveLock : ShareUpdateExclusiveLock;
+
+ if (stmt->relation != NULL)
+ {
+ rel = process_single_relation(stmt->relation, stmt->indexname,
+ CLUSTER_COMMAND_REPACK, lockmode,
+ isTopLevel, ¶ms, &indexOid);
+ if (rel == NULL)
+ return;
+ }
+
+ /*
+ * By here, we know we are in a multi-table situation.
+ *
+ * Concurrent processing is currently considered rather special (e.g. in
+ * terms of resources consumed) so it is not performed in bulk.
+ */
+ if (params.options & CLUOPT_CONCURRENT)
+ {
+ if (rel != NULL)
+ {
+ Assert(rel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE);
+ ereport(ERROR,
+ (errmsg("REPACK CONCURRENTLY not supported for partitioned tables"),
+ errhint("Consider running the command for individual partitions.")));
+ }
+ else
+ ereport(ERROR,
+ (errmsg("REPACK CONCURRENTLY requires explicit table name")));
+ }
+
+ /*
+ * In order to avoid holding locks for too long, we want to process each
+ * table in its own transaction. This forces us to disallow running
+ * inside a user transaction block.
+ */
+ PreventInTransactionBlock(isTopLevel, "REPACK");
+
+ /* Also, we need a memory context to hold our list of relations */
+ repack_context = AllocSetContextCreate(PortalContext,
+ "Repack",
+ ALLOCSET_DEFAULT_SIZES);
+
+ params.options |= CLUOPT_RECHECK;
+ if (rel != NULL)
+ {
+ Oid relid;
+ bool rel_is_index;
+
+ Assert(rel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE);
+ /* See the ereport() above. */
+ Assert((params.options & CLUOPT_CONCURRENT) == 0);
+
+ if (OidIsValid(indexOid))
+ {
+ relid = indexOid;
+ rel_is_index = true;
+ }
+ else
+ {
+ relid = RelationGetRelid(rel);
+ rel_is_index = false;
+ }
+ rtcs = get_tables_to_cluster_partitioned(repack_context, relid,
+ rel_is_index,
+ CLUSTER_COMMAND_REPACK);
/* close relation, releasing lock on parent table */
- table_close(rel, AccessExclusiveLock);
+ table_close(rel, lockmode);
}
else
rtcs = get_tables_to_repack(repack_context);
/* Do the job. */
- cluster_multiple_rels(rtcs, ¶ms, CLUSTER_COMMAND_REPACK);
+ cluster_multiple_rels(rtcs, ¶ms, CLUSTER_COMMAND_REPACK, lockmode,
+ isTopLevel);
+
/* Start a new transaction for the cleanup work. */
StartTransactionCommand();
@@ -1904,7 +4270,8 @@ repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
*/
static Relation
process_single_relation(RangeVar *relation, char *indexname,
- ClusterCommand cmd, ClusterParams *params,
+ ClusterCommand cmd, LOCKMODE lockmode,
+ bool isTopLevel, ClusterParams *params,
Oid *indexOid_p)
{
Relation rel;
@@ -1914,12 +4281,10 @@ process_single_relation(RangeVar *relation, char *indexname,
Oid tableOid;
/*
- * Find, lock, and check permissions on the table. We obtain
- * AccessExclusiveLock right away to avoid lock-upgrade hazard in the
- * single-transaction case.
+ * Find, lock, and check permissions on the table.
*/
tableOid = RangeVarGetRelidExtended(relation,
- AccessExclusiveLock,
+ lockmode,
0,
RangeVarCallbackMaintainsTable,
NULL);
@@ -1973,7 +4338,7 @@ process_single_relation(RangeVar *relation, char *indexname,
/* For non-partitioned tables, do what we came here to do. */
if (rel->rd_rel->relkind != RELKIND_PARTITIONED_TABLE)
{
- cluster_rel(rel, indexOid, params, cmd);
+ cluster_rel(rel, indexOid, params, cmd, isTopLevel);
/* cluster_rel closes the relation, but keeps lock */
return NULL;
diff --git a/src/backend/commands/matview.c b/src/backend/commands/matview.c
index 0bfbc5ca6d..eae34fbe6c 100644
--- a/src/backend/commands/matview.c
+++ b/src/backend/commands/matview.c
@@ -906,7 +906,7 @@ refresh_by_match_merge(Oid matviewOid, Oid tempOid, Oid relowner,
static void
refresh_by_heap_swap(Oid matviewOid, Oid OIDNewHeap, char relpersistence)
{
- finish_heap_swap(matviewOid, OIDNewHeap, false, false, true, true,
+ finish_heap_swap(matviewOid, OIDNewHeap, false, false, true, true, true,
RecentXmin, ReadNextMultiXactId(), relpersistence);
}
diff --git a/src/backend/commands/tablecmds.c b/src/backend/commands/tablecmds.c
index 901cb321c3..364f2b6a81 100644
--- a/src/backend/commands/tablecmds.c
+++ b/src/backend/commands/tablecmds.c
@@ -4527,6 +4527,16 @@ AlterTableInternal(Oid relid, List *cmds, bool recurse)
rel = relation_open(relid, lockmode);
+ /*
+ * If lockmode allows, check if REPACK CONCURRENTLY is in progress. If
+ * lockmode is too weak, cluster_rel() should detect incompatible DDLs
+ * executed by us.
+ *
+ * XXX We might skip the changes for DDLs which do not change the tuple
+ * descriptor.
+ */
+ check_for_concurrent_repack(relid, lockmode);
+
EventTriggerAlterTableRelid(relid);
ATController(NULL, rel, cmds, recurse, lockmode, NULL);
@@ -5960,6 +5970,7 @@ ATRewriteTables(AlterTableStmt *parsetree, List **wqueue, LOCKMODE lockmode,
finish_heap_swap(tab->relid, OIDNewHeap,
false, false, true,
!OidIsValid(tab->newTableSpace),
+ true,
RecentXmin,
ReadNextMultiXactId(),
persistence);
diff --git a/src/backend/commands/vacuum.c b/src/backend/commands/vacuum.c
index 59dddcd31f..30e1bb5719 100644
--- a/src/backend/commands/vacuum.c
+++ b/src/backend/commands/vacuum.c
@@ -123,7 +123,7 @@ static void vac_truncate_clog(TransactionId frozenXID,
TransactionId lastSaneFrozenXid,
MultiXactId lastSaneMinMulti);
static bool vacuum_rel(Oid relid, RangeVar *relation, VacuumParams *params,
- BufferAccessStrategy bstrategy);
+ BufferAccessStrategy bstrategy, bool isTopLevel);
static double compute_parallel_delay(void);
static VacOptValue get_vacoptval_from_boolean(DefElem *def);
static bool vac_tid_reaped(ItemPointer itemptr, void *state);
@@ -633,7 +633,8 @@ vacuum(List *relations, VacuumParams *params, BufferAccessStrategy bstrategy,
if (params->options & VACOPT_VACUUM)
{
- if (!vacuum_rel(vrel->oid, vrel->relation, params, bstrategy))
+ if (!vacuum_rel(vrel->oid, vrel->relation, params, bstrategy,
+ isTopLevel))
continue;
}
@@ -1989,7 +1990,7 @@ vac_truncate_clog(TransactionId frozenXID,
*/
static bool
vacuum_rel(Oid relid, RangeVar *relation, VacuumParams *params,
- BufferAccessStrategy bstrategy)
+ BufferAccessStrategy bstrategy, bool isTopLevel)
{
LOCKMODE lmode;
Relation rel;
@@ -2249,7 +2250,7 @@ vacuum_rel(Oid relid, RangeVar *relation, VacuumParams *params,
/* VACUUM FULL is now a variant of CLUSTER; see cluster.c */
cluster_rel(rel, InvalidOid, &cluster_params,
- CLUSTER_COMMAND_VACUUM);
+ CLUSTER_COMMAND_VACUUM, isTopLevel);
/* cluster_rel closes the relation, but keeps lock */
rel = NULL;
@@ -2295,7 +2296,8 @@ vacuum_rel(Oid relid, RangeVar *relation, VacuumParams *params,
toast_vacuum_params.options |= VACOPT_PROCESS_MAIN;
toast_vacuum_params.toast_parent = relid;
- vacuum_rel(toast_relid, NULL, &toast_vacuum_params, bstrategy);
+ vacuum_rel(toast_relid, NULL, &toast_vacuum_params, bstrategy,
+ isTopLevel);
}
/*
diff --git a/src/backend/meson.build b/src/backend/meson.build
index 2b0db21480..50aa385a58 100644
--- a/src/backend/meson.build
+++ b/src/backend/meson.build
@@ -194,5 +194,6 @@ pg_test_mod_args = pg_mod_args + {
subdir('jit/llvm')
subdir('replication/libpqwalreceiver')
subdir('replication/pgoutput')
+subdir('replication/pgoutput_repack')
subdir('snowball')
subdir('utils/mb/conversion_procs')
diff --git a/src/backend/parser/gram.y b/src/backend/parser/gram.y
index 8b4c226495..5c937b7db8 100644
--- a/src/backend/parser/gram.y
+++ b/src/backend/parser/gram.y
@@ -11874,27 +11874,30 @@ cluster_index_specification:
*
* QUERY:
* REPACK [ (options) ] [ <qualified_name> [ USING INDEX <index_name> ] ]
+ * REPACK [ (options) ] CONCURRENTLY <qualified_name> [ USING INDEX <index_name> ]
*
*****************************************************************************/
RepackStmt:
- REPACK qualified_name repack_index_specification
+ REPACK opt_concurrently qualified_name repack_index_specification
{
RepackStmt *n = makeNode(RepackStmt);
- n->relation = $2;
- n->indexname = $3;
+ n->concurrent = $2;
+ n->relation = $3;
+ n->indexname = $4;
n->params = NIL;
$$ = (Node *) n;
}
- | REPACK '(' utility_option_list ')' qualified_name repack_index_specification
+ | REPACK '(' utility_option_list ')' opt_concurrently qualified_name repack_index_specification
{
RepackStmt *n = makeNode(RepackStmt);
- n->relation = $5;
- n->indexname = $6;
n->params = $3;
+ n->concurrent = $5;
+ n->relation = $6;
+ n->indexname = $7;
$$ = (Node *) n;
}
@@ -11905,6 +11908,7 @@ RepackStmt:
n->relation = NULL;
n->indexname = NULL;
n->params = NIL;
+ n->concurrent = false;
$$ = (Node *) n;
}
@@ -11915,6 +11919,7 @@ RepackStmt:
n->relation = NULL;
n->indexname = NULL;
n->params = $3;
+ n->concurrent = false;
$$ = (Node *) n;
}
;
diff --git a/src/backend/replication/logical/decode.c b/src/backend/replication/logical/decode.c
index 24d88f368d..a6df190747 100644
--- a/src/backend/replication/logical/decode.c
+++ b/src/backend/replication/logical/decode.c
@@ -33,6 +33,7 @@
#include "access/xlogreader.h"
#include "access/xlogrecord.h"
#include "catalog/pg_control.h"
+#include "commands/cluster.h"
#include "replication/decode.h"
#include "replication/logical.h"
#include "replication/message.h"
@@ -467,6 +468,29 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
TransactionId xid = XLogRecGetXid(buf->record);
SnapBuild *builder = ctx->snapshot_builder;
+ /*
+ * Check if REPACK CONCURRENTLY is being performed by this backend. If so,
+ * only decode data changes of the table that it is processing, and the
+ * changes of its TOAST relation.
+ *
+ * (TOAST locator should not be set unless the main is.)
+ */
+ Assert(!OidIsValid(repacked_rel_toast_locator.relNumber) ||
+ OidIsValid(repacked_rel_locator.relNumber));
+
+ if (OidIsValid(repacked_rel_locator.relNumber))
+ {
+ XLogReaderState *r = buf->record;
+ RelFileLocator locator;
+
+ /* Not all records contain the block. */
+ if (XLogRecGetBlockTagExtended(r, 0, &locator, NULL, NULL, NULL) &&
+ !RelFileLocatorEquals(locator, repacked_rel_locator) &&
+ (!OidIsValid(repacked_rel_toast_locator.relNumber) ||
+ !RelFileLocatorEquals(locator, repacked_rel_toast_locator)))
+ return;
+ }
+
ReorderBufferProcessXid(ctx->reorder, xid, buf->origptr);
/*
diff --git a/src/backend/replication/logical/snapbuild.c b/src/backend/replication/logical/snapbuild.c
index 8c83ff6feb..c54a1277cc 100644
--- a/src/backend/replication/logical/snapbuild.c
+++ b/src/backend/replication/logical/snapbuild.c
@@ -486,6 +486,26 @@ SnapBuildInitialSnapshot(SnapBuild *builder)
return SnapBuildMVCCFromHistoric(snap, true);
}
+/*
+ * Build an MVCC snapshot for the initial data load performed by REPACK
+ * CONCURRENTLY command.
+ *
+ * The snapshot will only be used to scan one particular relation, which is
+ * treated like a catalog (therefore ->building_full_snapshot is not
+ * important), and the caller should already have a replication slot setup (so
+ * we do not set MyProc->xmin). XXX Do we yet need to add some restrictions?
+ */
+Snapshot
+SnapBuildInitialSnapshotForRepack(SnapBuild *builder)
+{
+ Snapshot snap;
+
+ Assert(builder->state == SNAPBUILD_CONSISTENT);
+
+ snap = SnapBuildBuildSnapshot(builder);
+ return SnapBuildMVCCFromHistoric(snap, false);
+}
+
/*
* Turn a historic MVCC snapshot into an ordinary MVCC snapshot.
*
diff --git a/src/backend/replication/pgoutput_repack/Makefile b/src/backend/replication/pgoutput_repack/Makefile
new file mode 100644
index 0000000000..4efeb713b7
--- /dev/null
+++ b/src/backend/replication/pgoutput_repack/Makefile
@@ -0,0 +1,32 @@
+#-------------------------------------------------------------------------
+#
+# Makefile--
+# Makefile for src/backend/replication/pgoutput_repack
+#
+# IDENTIFICATION
+# src/backend/replication/pgoutput_repack
+#
+#-------------------------------------------------------------------------
+
+subdir = src/backend/replication/pgoutput_repack
+top_builddir = ../../../..
+include $(top_builddir)/src/Makefile.global
+
+OBJS = \
+ $(WIN32RES) \
+ pgoutput_repack.o
+PGFILEDESC = "pgoutput_repack - logical replication output plugin for REPACK command"
+NAME = pgoutput_repack
+
+all: all-shared-lib
+
+include $(top_srcdir)/src/Makefile.shlib
+
+install: all installdirs install-lib
+
+installdirs: installdirs-lib
+
+uninstall: uninstall-lib
+
+clean distclean: clean-lib
+ rm -f $(OBJS)
diff --git a/src/backend/replication/pgoutput_repack/meson.build b/src/backend/replication/pgoutput_repack/meson.build
new file mode 100644
index 0000000000..133e865a4a
--- /dev/null
+++ b/src/backend/replication/pgoutput_repack/meson.build
@@ -0,0 +1,18 @@
+# Copyright (c) 2022-2024, PostgreSQL Global Development Group
+
+pgoutput_repack_sources = files(
+ 'pgoutput_repack.c',
+)
+
+if host_system == 'windows'
+ pgoutput_repack_sources += rc_lib_gen.process(win32ver_rc, extra_args: [
+ '--NAME', 'pgoutput_repack',
+ '--FILEDESC', 'pgoutput_repack - logical replication output plugin for REPACK command',])
+endif
+
+pgoutput_repack = shared_module('pgoutput_repack',
+ pgoutput_repack_sources,
+ kwargs: pg_mod_args,
+)
+
+backend_targets += pgoutput_repack
diff --git a/src/backend/replication/pgoutput_repack/pgoutput_repack.c b/src/backend/replication/pgoutput_repack/pgoutput_repack.c
new file mode 100644
index 0000000000..1ef9b3cbfd
--- /dev/null
+++ b/src/backend/replication/pgoutput_repack/pgoutput_repack.c
@@ -0,0 +1,286 @@
+/*-------------------------------------------------------------------------
+ *
+ * pgoutput_cluster.c
+ * Logical Replication output plugin for REPACK command
+ *
+ * Copyright (c) 2012-2024, PostgreSQL Global Development Group
+ *
+ * IDENTIFICATION
+ * src/backend/replication/pgoutput_cluster/pgoutput_cluster.c
+ *
+ *-------------------------------------------------------------------------
+ */
+#include "postgres.h"
+
+#include "access/heaptoast.h"
+#include "commands/cluster.h"
+#include "replication/snapbuild.h"
+
+PG_MODULE_MAGIC;
+
+static void plugin_startup(LogicalDecodingContext *ctx,
+ OutputPluginOptions *opt, bool is_init);
+static void plugin_shutdown(LogicalDecodingContext *ctx);
+static void plugin_begin_txn(LogicalDecodingContext *ctx,
+ ReorderBufferTXN *txn);
+static void plugin_commit_txn(LogicalDecodingContext *ctx,
+ ReorderBufferTXN *txn, XLogRecPtr commit_lsn);
+static void plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
+ Relation rel, ReorderBufferChange *change);
+static void plugin_truncate(struct LogicalDecodingContext *ctx,
+ ReorderBufferTXN *txn, int nrelations,
+ Relation relations[],
+ ReorderBufferChange *change);
+static void store_change(LogicalDecodingContext *ctx,
+ ConcurrentChangeKind kind, HeapTuple tuple);
+
+void
+_PG_output_plugin_init(OutputPluginCallbacks *cb)
+{
+ AssertVariableIsOfType(&_PG_output_plugin_init, LogicalOutputPluginInit);
+
+ cb->startup_cb = plugin_startup;
+ cb->begin_cb = plugin_begin_txn;
+ cb->change_cb = plugin_change;
+ cb->truncate_cb = plugin_truncate;
+ cb->commit_cb = plugin_commit_txn;
+ cb->shutdown_cb = plugin_shutdown;
+}
+
+
+/* initialize this plugin */
+static void
+plugin_startup(LogicalDecodingContext *ctx, OutputPluginOptions *opt,
+ bool is_init)
+{
+ ctx->output_plugin_private = NULL;
+
+ /* Probably unnecessary, as we don't use the SQL interface ... */
+ opt->output_type = OUTPUT_PLUGIN_BINARY_OUTPUT;
+
+ if (ctx->output_plugin_options != NIL)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("This plugin does not expect any options")));
+ }
+}
+
+static void
+plugin_shutdown(LogicalDecodingContext *ctx)
+{
+}
+
+/*
+ * As we don't release the slot during processing of particular table, there's
+ * no room for SQL interface, even for debugging purposes. Therefore we need
+ * neither OutputPluginPrepareWrite() nor OutputPluginWrite() in the plugin
+ * callbacks. (Although we might want to write custom callbacks, this API
+ * seems to be unnecessarily generic for our purposes.)
+ */
+
+/* BEGIN callback */
+static void
+plugin_begin_txn(LogicalDecodingContext *ctx, ReorderBufferTXN *txn)
+{
+}
+
+/* COMMIT callback */
+static void
+plugin_commit_txn(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
+ XLogRecPtr commit_lsn)
+{
+}
+
+/*
+ * Callback for individual changed tuples
+ */
+static void
+plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
+ Relation relation, ReorderBufferChange *change)
+{
+ RepackDecodingState *dstate;
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+
+ /* Only interested in one particular relation. */
+ if (relation->rd_id != dstate->relid)
+ return;
+
+ /* Decode entry depending on its type */
+ switch (change->action)
+ {
+ case REORDER_BUFFER_CHANGE_INSERT:
+ {
+ HeapTuple newtuple;
+
+ newtuple = change->data.tp.newtuple != NULL ?
+ change->data.tp.newtuple : NULL;
+
+ /*
+ * Identity checks in the main function should have made this
+ * impossible.
+ */
+ if (newtuple == NULL)
+ elog(ERROR, "Incomplete insert info.");
+
+ store_change(ctx, CHANGE_INSERT, newtuple);
+ }
+ break;
+ case REORDER_BUFFER_CHANGE_UPDATE:
+ {
+ HeapTuple oldtuple,
+ newtuple;
+
+ oldtuple = change->data.tp.oldtuple != NULL ?
+ change->data.tp.oldtuple : NULL;
+ newtuple = change->data.tp.newtuple != NULL ?
+ change->data.tp.newtuple : NULL;
+
+ if (newtuple == NULL)
+ elog(ERROR, "Incomplete update info.");
+
+ if (oldtuple != NULL)
+ store_change(ctx, CHANGE_UPDATE_OLD, oldtuple);
+
+ store_change(ctx, CHANGE_UPDATE_NEW, newtuple);
+ }
+ break;
+ case REORDER_BUFFER_CHANGE_DELETE:
+ {
+ HeapTuple oldtuple;
+
+ oldtuple = change->data.tp.oldtuple ?
+ change->data.tp.oldtuple : NULL;
+
+ if (oldtuple == NULL)
+ elog(ERROR, "Incomplete delete info.");
+
+ store_change(ctx, CHANGE_DELETE, oldtuple);
+ }
+ break;
+ default:
+ /* Should not come here */
+ Assert(false);
+ break;
+ }
+}
+
+static void
+plugin_truncate(struct LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
+ int nrelations, Relation relations[],
+ ReorderBufferChange *change)
+{
+ RepackDecodingState *dstate;
+ int i;
+ Relation relation = NULL;
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+
+ /* Find the relation we are processing. */
+ for (i = 0; i < nrelations; i++)
+ {
+ relation = relations[i];
+
+ if (RelationGetRelid(relation) == dstate->relid)
+ break;
+ }
+
+ /* Is this truncation of another relation? */
+ if (i == nrelations)
+ return;
+
+ store_change(ctx, CHANGE_TRUNCATE, NULL);
+}
+
+/* Store concurrent data change. */
+static void
+store_change(LogicalDecodingContext *ctx, ConcurrentChangeKind kind,
+ HeapTuple tuple)
+{
+ RepackDecodingState *dstate;
+ char *change_raw;
+ ConcurrentChange change;
+ bool flattened = false;
+ Size size;
+ Datum values[1];
+ bool isnull[1];
+ char *dst, *dst_start;
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+
+ size = MAXALIGN(VARHDRSZ) + SizeOfConcurrentChange;
+
+ if (tuple)
+ {
+ /*
+ * ReorderBufferCommit() stores the TOAST chunks in its private memory
+ * context and frees them after having called
+ * apply_change(). Therefore we need flat copy (including TOAST) that
+ * we eventually copy into the memory context which is available to
+ * decode_concurrent_changes().
+ */
+ if (HeapTupleHasExternal(tuple))
+ {
+ /*
+ * toast_flatten_tuple_to_datum() might be more convenient but we
+ * don't want the decompression it does.
+ */
+ tuple = toast_flatten_tuple(tuple, dstate->tupdesc);
+ flattened = true;
+ }
+
+ size += tuple->t_len;
+ }
+
+ /* XXX Isn't there any function / macro to do this? */
+ if (size >= 0x3FFFFFFF)
+ elog(ERROR, "Change is too big.");
+
+ /* Construct the change. */
+ change_raw = (char *) palloc0(size);
+ SET_VARSIZE(change_raw, size);
+ /*
+ * Since the varlena alignment might not be sufficient for the structure,
+ * set the fields in a local instance and remember where it should
+ * eventually be copied.
+ */
+ change.kind = kind;
+ dst_start = (char *) VARDATA(change_raw);
+
+ /* No other information is needed for TRUNCATE. */
+ if (change.kind == CHANGE_TRUNCATE)
+ {
+ memcpy(dst_start, &change, SizeOfConcurrentChange);
+ goto store;
+ }
+
+ /*
+ * Copy the tuple.
+ *
+ * CAUTION: change->tup_data.t_data must be fixed on retrieval!
+ */
+ memcpy(&change.tup_data, tuple, sizeof(HeapTupleData));
+ dst = dst_start + SizeOfConcurrentChange;
+ memcpy(dst, tuple->t_data, tuple->t_len);
+
+ /* The data has been copied. */
+ if (flattened)
+ pfree(tuple);
+
+store:
+ /* Copy the structure so it can be stored. */
+ memcpy(dst_start, &change, SizeOfConcurrentChange);
+
+ /* Store as tuple of 1 bytea column. */
+ values[0] = PointerGetDatum(change_raw);
+ isnull[0] = false;
+ tuplestore_putvalues(dstate->tstore, dstate->tupdesc_change,
+ values, isnull);
+
+ /* Accounting. */
+ dstate->nchanges++;
+
+ /* Cleanup. */
+ pfree(change_raw);
+}
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 174eed7036..07e477d279 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -25,6 +25,7 @@
#include "access/xlogprefetcher.h"
#include "access/xlogrecovery.h"
#include "commands/async.h"
+#include "commands/cluster.h"
#include "miscadmin.h"
#include "pgstat.h"
#include "postmaster/autovacuum.h"
@@ -148,6 +149,7 @@ CalculateShmemSize(int *num_semaphores)
size = add_size(size, WaitEventCustomShmemSize());
size = add_size(size, InjectionPointShmemSize());
size = add_size(size, SlotSyncShmemSize());
+ size = add_size(size, RepackShmemSize());
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -340,6 +342,7 @@ CreateOrAttachShmemStructs(void)
StatsShmemInit();
WaitEventCustomShmemInit();
InjectionPointShmemInit();
+ RepackShmemInit();
}
/*
diff --git a/src/backend/tcop/utility.c b/src/backend/tcop/utility.c
index bf3ba3c2ae..4ee4c47487 100644
--- a/src/backend/tcop/utility.c
+++ b/src/backend/tcop/utility.c
@@ -1307,6 +1307,16 @@ ProcessUtilitySlow(ParseState *pstate,
lockmode = AlterTableGetLockLevel(atstmt->cmds);
relid = AlterTableLookupRelation(atstmt, lockmode);
+ /*
+ * If lockmode allows, check if REPACK CONCURRENT is in
+ * progress. If lockmode is too weak, cluster_rel() should
+ * detect incompatible DDLs executed by us.
+ *
+ * XXX We might skip the changes for DDLs which do not
+ * change the tuple descriptor.
+ */
+ check_for_concurrent_repack(relid, lockmode);
+
if (OidIsValid(relid))
{
AlterTableUtilityContext atcontext;
diff --git a/src/backend/utils/activity/backend_progress.c b/src/backend/utils/activity/backend_progress.c
index eebc968193..e2c84baba9 100644
--- a/src/backend/utils/activity/backend_progress.c
+++ b/src/backend/utils/activity/backend_progress.c
@@ -162,3 +162,19 @@ pgstat_progress_end_command(void)
beentry->st_progress.command_target = InvalidOid;
PGSTAT_END_WRITE_ACTIVITY(beentry);
}
+
+void
+pgstat_progress_restore_state(PgBackendProgress *backup)
+{
+ volatile PgBackendStatus *beentry = MyBEEntry;
+
+ if (!beentry || !pgstat_track_activities)
+ return;
+
+ PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
+ beentry->st_progress.command = backup->command;
+ beentry->st_progress.command_target = backup->command_target;
+ memcpy(MyBEEntry->st_progress.param, backup->param,
+ sizeof(beentry->st_progress.param));
+ PGSTAT_END_WRITE_ACTIVITY(beentry);
+}
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index e199f07162..5a0097d53b 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -346,6 +346,7 @@ WALSummarizer "Waiting to read or update WAL summarization state."
DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
+RepackedRels "Waiting to read or update information on tables being repacked concurrently."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/cache/inval.c b/src/backend/utils/cache/inval.c
index 700ccb6df9..cb92ddb1e3 100644
--- a/src/backend/utils/cache/inval.c
+++ b/src/backend/utils/cache/inval.c
@@ -1569,6 +1569,27 @@ CacheInvalidateRelcache(Relation relation)
databaseId, relationId);
}
+/*
+ * CacheInvalidateRelcacheImmediate
+ * Send invalidation message for the specified relation's relcache entry.
+ *
+ * Currently this is used in REPACK CONCURRENTLY, to make sure that other
+ * backends are aware that the command is being executed for the relation.
+ */
+void
+CacheInvalidateRelcacheImmediate(Relation relation)
+{
+ SharedInvalidationMessage msg;
+
+ msg.rc.id = SHAREDINVALRELCACHE_ID;
+ msg.rc.dbId = MyDatabaseId;
+ msg.rc.relId = RelationGetRelid(relation);
+ /* check AddCatcacheInvalidationMessage() for an explanation */
+ VALGRIND_MAKE_MEM_DEFINED(&msg, sizeof(msg));
+
+ SendSharedInvalidMessages(&msg, 1);
+}
+
/*
* CacheInvalidateRelcacheAll
* Register invalidation of the whole relcache at the end of command.
diff --git a/src/backend/utils/cache/relcache.c b/src/backend/utils/cache/relcache.c
index 398114373e..1273149178 100644
--- a/src/backend/utils/cache/relcache.c
+++ b/src/backend/utils/cache/relcache.c
@@ -64,6 +64,7 @@
#include "catalog/pg_type.h"
#include "catalog/schemapg.h"
#include "catalog/storage.h"
+#include "commands/cluster.h"
#include "commands/policy.h"
#include "commands/publicationcmds.h"
#include "commands/trigger.h"
@@ -1249,6 +1250,10 @@ retry:
/* make sure relation is marked as having no open file yet */
relation->rd_smgr = NULL;
+ /* Is REPACK CONCURRENTLY in progress? */
+ relation->rd_repack_concurrent =
+ is_concurrent_repack_in_progress(targetRelId);
+
/*
* now we can free the memory allocated for pg_class_tuple
*/
diff --git a/src/backend/utils/time/snapmgr.c b/src/backend/utils/time/snapmgr.c
index 42bded373b..103d1249bb 100644
--- a/src/backend/utils/time/snapmgr.c
+++ b/src/backend/utils/time/snapmgr.c
@@ -154,7 +154,6 @@ static List *exportedSnapshots = NIL;
/* Prototypes for local functions */
static void UnregisterSnapshotNoOwner(Snapshot snapshot);
-static void FreeSnapshot(Snapshot snapshot);
static void SnapshotResetXmin(void);
/* ResourceOwner callbacks to track snapshot references */
@@ -587,7 +586,7 @@ CopySnapshot(Snapshot snapshot)
* FreeSnapshot
* Free the memory associated with a snapshot.
*/
-static void
+void
FreeSnapshot(Snapshot snapshot)
{
Assert(snapshot->regd_count == 0);
diff --git a/src/bin/psql/tab-complete.in.c b/src/bin/psql/tab-complete.in.c
index 72338fffb2..9aefd06481 100644
--- a/src/bin/psql/tab-complete.in.c
+++ b/src/bin/psql/tab-complete.in.c
@@ -4910,18 +4910,26 @@ match_previous_words(int pattern_id,
}
/* REPACK */
- else if (Matches("REPACK"))
+ else if (Matches("REPACK") || Matches("REPACK", "(*)"))
+ COMPLETE_WITH_SCHEMA_QUERY_PLUS(Query_for_list_of_clusterables,
+ "CONCURRENTLY");
+ else if (Matches("REPACK", "CONCURRENTLY"))
COMPLETE_WITH_SCHEMA_QUERY(Query_for_list_of_clusterables);
- else if (Matches("REPACK", "(*)"))
+ else if (Matches("REPACK", "(*)", "CONCURRENTLY"))
COMPLETE_WITH_SCHEMA_QUERY(Query_for_list_of_clusterables);
- /* If we have REPACK <sth>, then add "USING INDEX" */
- else if (Matches("REPACK", MatchAnyExcept("(")))
+ /* If we have REPACK [ CONCURRENTLY ] <sth>, then add "USING INDEX" */
+ else if (Matches("REPACK", MatchAnyExcept("(|CONCURRENTLY")) ||
+ Matches("REPACK", "CONCURRENTLY", MatchAnyExcept("(")))
COMPLETE_WITH("USING INDEX");
- /* If we have REPACK (*) <sth>, then add "USING INDEX" */
- else if (Matches("REPACK", "(*)", MatchAny))
+ /* If we have REPACK (*) [ CONCURRENTLY ] <sth>, then add "USING INDEX" */
+ else if (Matches("REPACK", "(*)", MatchAnyExcept("CONCURRENTLY")) ||
+ Matches("REPACK", "(*)", "CONCURRENTLY", MatchAnyExcept("(")))
COMPLETE_WITH("USING INDEX");
- /* If we have REPACK <sth> USING, then add the index as well */
- else if (Matches("REPACK", MatchAny, "USING", "INDEX"))
+ /*
+ * Complete ... [ (*) ] [ CONCURRENTLY ] <sth> USING INDEX, with a list of
+ * indexes for <sth>.
+ */
+ else if (TailMatches(MatchAnyExcept("(|CONCURRENTLY"), "USING", "INDEX"))
{
set_completion_reference(prev3_wd);
COMPLETE_WITH_SCHEMA_QUERY(Query_for_index_of_table);
diff --git a/src/include/access/heapam.h b/src/include/access/heapam.h
index 1640d9c32f..bdeb2f8354 100644
--- a/src/include/access/heapam.h
+++ b/src/include/access/heapam.h
@@ -421,6 +421,10 @@ extern HTSV_Result HeapTupleSatisfiesVacuumHorizon(HeapTuple htup, Buffer buffer
TransactionId *dead_after);
extern void HeapTupleSetHintBits(HeapTupleHeader tuple, Buffer buffer,
uint16 infomask, TransactionId xid);
+extern bool HeapTupleMVCCInserted(HeapTuple htup, Snapshot snapshot,
+ Buffer buffer);
+extern bool HeapTupleMVCCNotDeleted(HeapTuple htup, Snapshot snapshot,
+ Buffer buffer);
extern bool HeapTupleHeaderIsOnlyLocked(HeapTupleHeader tuple);
extern bool HeapTupleIsSurelyDead(HeapTuple htup,
struct GlobalVisState *vistest);
diff --git a/src/include/access/tableam.h b/src/include/access/tableam.h
index 131c050c15..aa3190986a 100644
--- a/src/include/access/tableam.h
+++ b/src/include/access/tableam.h
@@ -21,6 +21,7 @@
#include "access/sdir.h"
#include "access/xact.h"
#include "executor/tuptable.h"
+#include "replication/logical.h"
#include "storage/read_stream.h"
#include "utils/rel.h"
#include "utils/snapshot.h"
@@ -630,6 +631,8 @@ typedef struct TableAmRoutine
Relation OldIndex,
bool use_sort,
TransactionId OldestXmin,
+ Snapshot snapshot,
+ LogicalDecodingContext *decoding_ctx,
TransactionId *xid_cutoff,
MultiXactId *multi_cutoff,
double *num_tuples,
@@ -1673,6 +1676,10 @@ table_relation_copy_data(Relation rel, const RelFileLocator *newrlocator)
* not needed for the relation's AM
* - *xid_cutoff - ditto
* - *multi_cutoff - ditto
+ * - snapshot - if != NULL, ignore data changes done by transactions that this
+ * (MVCC) snapshot considers still in-progress or in the future.
+ * - decoding_ctx - logical decoding context, to capture concurrent data
+ * changes.
*
* Output parameters:
* - *xid_cutoff - rel's new relfrozenxid value, may be invalid
@@ -1685,6 +1692,8 @@ table_relation_copy_for_cluster(Relation OldTable, Relation NewTable,
Relation OldIndex,
bool use_sort,
TransactionId OldestXmin,
+ Snapshot snapshot,
+ LogicalDecodingContext *decoding_ctx,
TransactionId *xid_cutoff,
MultiXactId *multi_cutoff,
double *num_tuples,
@@ -1693,6 +1702,7 @@ table_relation_copy_for_cluster(Relation OldTable, Relation NewTable,
{
OldTable->rd_tableam->relation_copy_for_cluster(OldTable, NewTable, OldIndex,
use_sort, OldestXmin,
+ snapshot, decoding_ctx,
xid_cutoff, multi_cutoff,
num_tuples, tups_vacuumed,
tups_recently_dead);
diff --git a/src/include/catalog/index.h b/src/include/catalog/index.h
index 4daa8bef5e..66431cc19e 100644
--- a/src/include/catalog/index.h
+++ b/src/include/catalog/index.h
@@ -100,6 +100,9 @@ extern Oid index_concurrently_create_copy(Relation heapRelation,
Oid tablespaceOid,
const char *newName);
+extern NullableDatum *get_index_stattargets(Oid indexid,
+ IndexInfo *indInfo);
+
extern void index_concurrently_build(Oid heapRelationId,
Oid indexRelationId);
diff --git a/src/include/commands/cluster.h b/src/include/commands/cluster.h
index c2976905e4..6fb5f5509c 100644
--- a/src/include/commands/cluster.h
+++ b/src/include/commands/cluster.h
@@ -13,10 +13,15 @@
#ifndef CLUSTER_H
#define CLUSTER_H
+#include "nodes/execnodes.h"
#include "nodes/parsenodes.h"
#include "parser/parse_node.h"
+#include "replication/logical.h"
#include "storage/lock.h"
+#include "storage/relfilelocator.h"
#include "utils/relcache.h"
+#include "utils/resowner.h"
+#include "utils/tuplestore.h"
/* flag bits for ClusterParams->options */
@@ -24,6 +29,7 @@
#define CLUOPT_RECHECK 0x02 /* recheck relation state */
#define CLUOPT_RECHECK_ISCLUSTERED 0x04 /* recheck relation state for
* indisclustered */
+#define CLUOPT_CONCURRENT 0x08 /* allow concurrent data changes */
/* options for CLUSTER */
typedef struct ClusterParams
@@ -46,14 +52,91 @@ typedef enum ClusterCommand
CLUSTER_COMMAND_VACUUM
} ClusterCommand;
+/*
+ * The following definitions are used by REPACK CONCURRENTLY.
+ */
+
+extern RelFileLocator repacked_rel_locator;
+extern RelFileLocator repacked_rel_toast_locator;
+
+typedef enum
+{
+ CHANGE_INSERT,
+ CHANGE_UPDATE_OLD,
+ CHANGE_UPDATE_NEW,
+ CHANGE_DELETE,
+ CHANGE_TRUNCATE
+} ConcurrentChangeKind;
+
+typedef struct ConcurrentChange
+{
+ /* See the enum above. */
+ ConcurrentChangeKind kind;
+
+ /*
+ * The actual tuple.
+ *
+ * The tuple data follows the ConcurrentChange structure. Before use make
+ * sure the tuple is correctly aligned (ConcurrentChange can be stored as
+ * bytea) and that tuple->t_data is fixed.
+ */
+ HeapTupleData tup_data;
+} ConcurrentChange;
+
+#define SizeOfConcurrentChange (offsetof(ConcurrentChange, tup_data) + \
+ sizeof(HeapTupleData))
+
+/*
+ * Logical decoding state.
+ *
+ * Here we store the data changes that we decode from WAL while the table
+ * contents is being copied to a new storage. Also the necessary metadata
+ * needed to apply these changes to the table is stored here.
+ */
+typedef struct RepackDecodingState
+{
+ /* The relation whose changes we're decoding. */
+ Oid relid;
+
+ /*
+ * Decoded changes are stored here. Although we try to avoid excessive
+ * batches, it can happen that the changes need to be stored to disk. The
+ * tuplestore does this transparently.
+ */
+ Tuplestorestate *tstore;
+
+ /* The current number of changes in tstore. */
+ double nchanges;
+
+ /*
+ * Descriptor to store the ConcurrentChange structure serialized (bytea).
+ * We can't store the tuple directly because tuplestore only supports
+ * minimum tuple and we may need to transfer OID system column from the
+ * output plugin. Also we need to transfer the change kind, so it's better
+ * to put everything in the structure than to use 2 tuplestores "in
+ * parallel".
+ */
+ TupleDesc tupdesc_change;
+
+ /* Tuple descriptor needed to update indexes. */
+ TupleDesc tupdesc;
+
+ /* Slot to retrieve data from tstore. */
+ TupleTableSlot *tsslot;
+
+ ResourceOwner resowner;
+} RepackDecodingState;
+
extern void cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel);
extern void cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
- ClusterCommand cmd);
+ ClusterCommand cmd, bool isTopLevel);
extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid,
LOCKMODE lockmode,
ClusterCommand cmd);
extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal);
-
+extern void can_repack_concurrently(Relation rel);
+extern void repack_decode_concurrent_changes(LogicalDecodingContext *ctx,
+ XLogRecPtr end_of_wal);
extern Oid make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
char relpersistence, LOCKMODE lockmode);
extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
@@ -61,9 +144,15 @@ extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
bool swap_toast_by_content,
bool check_constraints,
bool is_internal,
+ bool reindex,
TransactionId frozenXid,
MultiXactId cutoffMulti,
char newrelpersistence);
+extern Size RepackShmemSize(void);
+extern void RepackShmemInit(void);
+extern bool is_concurrent_repack_in_progress(Oid relid);
+extern void check_for_concurrent_repack(Oid relid, LOCKMODE lockmode);
+
extern void repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel);
#endif /* CLUSTER_H */
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 7644267e14..6b1b1a4c1a 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -67,10 +67,12 @@
#define PROGRESS_REPACK_PHASE 1
#define PROGRESS_REPACK_INDEX_RELID 2
#define PROGRESS_REPACK_HEAP_TUPLES_SCANNED 3
-#define PROGRESS_REPACK_HEAP_TUPLES_WRITTEN 4
-#define PROGRESS_REPACK_TOTAL_HEAP_BLKS 5
-#define PROGRESS_REPACK_HEAP_BLKS_SCANNED 6
-#define PROGRESS_REPACK_INDEX_REBUILD_COUNT 7
+#define PROGRESS_REPACK_HEAP_TUPLES_INSERTED 4
+#define PROGRESS_REPACK_HEAP_TUPLES_UPDATED 5
+#define PROGRESS_REPACK_HEAP_TUPLES_DELETED 6
+#define PROGRESS_REPACK_TOTAL_HEAP_BLKS 7
+#define PROGRESS_REPACK_HEAP_BLKS_SCANNED 8
+#define PROGRESS_REPACK_INDEX_REBUILD_COUNT 9
/*
* Phases of repack (as advertised via PROGRESS_REPACK_PHASE).
@@ -83,9 +85,10 @@
#define PROGRESS_REPACK_PHASE_INDEX_SCAN_HEAP 2
#define PROGRESS_REPACK_PHASE_SORT_TUPLES 3
#define PROGRESS_REPACK_PHASE_WRITE_NEW_HEAP 4
-#define PROGRESS_REPACK_PHASE_SWAP_REL_FILES 5
-#define PROGRESS_REPACK_PHASE_REBUILD_INDEX 6
-#define PROGRESS_REPACK_PHASE_FINAL_CLEANUP 7
+#define PROGRESS_REPACK_PHASE_CATCH_UP 5
+#define PROGRESS_REPACK_PHASE_SWAP_REL_FILES 6
+#define PROGRESS_REPACK_PHASE_REBUILD_INDEX 8
+#define PROGRESS_REPACK_PHASE_FINAL_CLEANUP 8
/* Commands of PROGRESS_REPACK */
#define PROGRESS_REPACK_COMMAND_REPACK 1
diff --git a/src/include/nodes/parsenodes.h b/src/include/nodes/parsenodes.h
index 03ed0450df..d6053cab9f 100644
--- a/src/include/nodes/parsenodes.h
+++ b/src/include/nodes/parsenodes.h
@@ -3924,6 +3924,7 @@ typedef struct RepackStmt
RangeVar *relation; /* relation being repacked */
char *indexname; /* order tuples by this index */
List *params; /* list of DefElem nodes */
+ bool concurrent; /* allow concurrent access? */
} RepackStmt;
diff --git a/src/include/replication/snapbuild.h b/src/include/replication/snapbuild.h
index 6d4d2d1814..802fc4b082 100644
--- a/src/include/replication/snapbuild.h
+++ b/src/include/replication/snapbuild.h
@@ -73,6 +73,7 @@ extern void FreeSnapshotBuilder(SnapBuild *builder);
extern void SnapBuildSnapDecRefcount(Snapshot snap);
extern Snapshot SnapBuildInitialSnapshot(SnapBuild *builder);
+extern Snapshot SnapBuildInitialSnapshotForRepack(SnapBuild *builder);
extern Snapshot SnapBuildMVCCFromHistoric(Snapshot snapshot, bool in_place);
extern const char *SnapBuildExportSnapshot(SnapBuild *builder);
extern void SnapBuildClearExportedSnapshot(void);
diff --git a/src/include/storage/lockdefs.h b/src/include/storage/lockdefs.h
index 7f3ba0352f..b0d81b736d 100644
--- a/src/include/storage/lockdefs.h
+++ b/src/include/storage/lockdefs.h
@@ -36,8 +36,9 @@ typedef int LOCKMODE;
#define AccessShareLock 1 /* SELECT */
#define RowShareLock 2 /* SELECT FOR UPDATE/FOR SHARE */
#define RowExclusiveLock 3 /* INSERT, UPDATE, DELETE */
-#define ShareUpdateExclusiveLock 4 /* VACUUM (non-FULL), ANALYZE, CREATE
- * INDEX CONCURRENTLY */
+#define ShareUpdateExclusiveLock 4 /* VACUUM (non-exclusive), ANALYZE, CREATE
+ * INDEX CONCURRENTLY, REPACK
+ * CONCURRENTLY */
#define ShareLock 5 /* CREATE INDEX (WITHOUT CONCURRENTLY) */
#define ShareRowExclusiveLock 6 /* like EXCLUSIVE MODE, but allows ROW
* SHARE */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index cf56545238..f07973b458 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -83,3 +83,4 @@ PG_LWLOCK(49, WALSummarizer)
PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
+PG_LWLOCK(54, RepackedRels)
diff --git a/src/include/utils/backend_progress.h b/src/include/utils/backend_progress.h
index 2f1de46d05..cea341276b 100644
--- a/src/include/utils/backend_progress.h
+++ b/src/include/utils/backend_progress.h
@@ -36,7 +36,7 @@ typedef enum ProgressCommandType
/*
* Any command which wishes can advertise that it is running by setting
- * command, command_target, and param[]. command_target should be the OID of
+ * ommand, command_target, and param[]. command_target should be the OID of
* the relation which the command targets (we assume there's just one, as this
* is meant for utility commands), but the meaning of each element in the
* param array is command-specific.
@@ -56,6 +56,7 @@ extern void pgstat_progress_parallel_incr_param(int index, int64 incr);
extern void pgstat_progress_update_multi_param(int nparam, const int *index,
const int64 *val);
extern void pgstat_progress_end_command(void);
+extern void pgstat_progress_restore_state(PgBackendProgress *backup);
#endif /* BACKEND_PROGRESS_H */
diff --git a/src/include/utils/inval.h b/src/include/utils/inval.h
index 40658ba2ff..6b2faed672 100644
--- a/src/include/utils/inval.h
+++ b/src/include/utils/inval.h
@@ -49,6 +49,8 @@ extern void CacheInvalidateCatalog(Oid catalogId);
extern void CacheInvalidateRelcache(Relation relation);
+extern void CacheInvalidateRelcacheImmediate(Relation relation);
+
extern void CacheInvalidateRelcacheAll(void);
extern void CacheInvalidateRelcacheByTuple(HeapTuple classTuple);
diff --git a/src/include/utils/rel.h b/src/include/utils/rel.h
index db3e504c3d..741b29226d 100644
--- a/src/include/utils/rel.h
+++ b/src/include/utils/rel.h
@@ -253,6 +253,9 @@ typedef struct RelationData
bool pgstat_enabled; /* should relation stats be counted */
/* use "struct" here to avoid needing to include pgstat.h: */
struct PgStat_TableStatus *pgstat_info; /* statistics collection area */
+
+ /* Is REPACK CONCURRENTLY being performed on this relation? */
+ bool rd_repack_concurrent;
} RelationData;
@@ -691,7 +694,9 @@ RelationCloseSmgr(Relation relation)
#define RelationIsAccessibleInLogicalDecoding(relation) \
(XLogLogicalInfoActive() && \
RelationNeedsWAL(relation) && \
- (IsCatalogRelation(relation) || RelationIsUsedAsCatalogTable(relation)))
+ (IsCatalogRelation(relation) || \
+ RelationIsUsedAsCatalogTable(relation) || \
+ (relation)->rd_repack_concurrent))
/*
* RelationIsLogicallyLogged
diff --git a/src/include/utils/snapmgr.h b/src/include/utils/snapmgr.h
index 147b190210..5eeabdc6c4 100644
--- a/src/include/utils/snapmgr.h
+++ b/src/include/utils/snapmgr.h
@@ -61,6 +61,8 @@ extern Snapshot GetLatestSnapshot(void);
extern void SnapshotSetCommandId(CommandId curcid);
extern Snapshot CopySnapshot(Snapshot snapshot);
+extern void FreeSnapshot(Snapshot snapshot);
+
extern Snapshot GetCatalogSnapshot(Oid relid);
extern Snapshot GetNonHistoricCatalogSnapshot(Oid relid);
extern void InvalidateCatalogSnapshot(void);
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 50d87af2fd..587c0c85b0 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1969,17 +1969,17 @@ pg_stat_progress_cluster| SELECT s.pid,
WHEN 2 THEN 'index scanning heap'::text
WHEN 3 THEN 'sorting tuples'::text
WHEN 4 THEN 'writing new heap'::text
- WHEN 5 THEN 'swapping relation files'::text
- WHEN 6 THEN 'rebuilding index'::text
- WHEN 7 THEN 'performing final cleanup'::text
+ WHEN 6 THEN 'swapping relation files'::text
+ WHEN 7 THEN 'rebuilding index'::text
+ WHEN 8 THEN 'performing final cleanup'::text
ELSE NULL::text
END AS phase,
(s.param3)::oid AS cluster_index_relid,
s.param4 AS heap_tuples_scanned,
s.param5 AS heap_tuples_written,
- s.param6 AS heap_blks_total,
- s.param7 AS heap_blks_scanned,
- s.param8 AS index_rebuild_count
+ s.param8 AS heap_blks_total,
+ s.param9 AS heap_blks_scanned,
+ s.param10 AS index_rebuild_count
FROM (pg_stat_get_progress_info('CLUSTER'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_progress_copy| SELECT s.pid,
@@ -2055,17 +2055,20 @@ pg_stat_progress_repack| SELECT s.pid,
WHEN 2 THEN 'index scanning heap'::text
WHEN 3 THEN 'sorting tuples'::text
WHEN 4 THEN 'writing new heap'::text
- WHEN 5 THEN 'swapping relation files'::text
- WHEN 6 THEN 'rebuilding index'::text
- WHEN 7 THEN 'performing final cleanup'::text
+ WHEN 5 THEN 'catch-up'::text
+ WHEN 6 THEN 'swapping relation files'::text
+ WHEN 7 THEN 'rebuilding index'::text
+ WHEN 8 THEN 'performing final cleanup'::text
ELSE NULL::text
END AS phase,
(s.param3)::oid AS repack_index_relid,
s.param4 AS heap_tuples_scanned,
- s.param5 AS heap_tuples_written,
- s.param6 AS heap_blks_total,
- s.param7 AS heap_blks_scanned,
- s.param8 AS index_rebuild_count
+ s.param5 AS heap_tuples_inserted,
+ s.param6 AS heap_tuples_updated,
+ s.param7 AS heap_tuples_deleted,
+ s.param8 AS heap_blks_total,
+ s.param9 AS heap_blks_scanned,
+ s.param10 AS index_rebuild_count
FROM (pg_stat_get_progress_info('REPACK'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_progress_vacuum| SELECT s.pid,
--
2.43.5
Attachments:
[text/x-diff] v08-0001-Add-REPACK-command.patch (89.7K, ../127361.1740559688@localhost/2-v08-0001-Add-REPACK-command.patch)
download | inline diff:
From 2f83bde90c324ce515455a03135bcde34dd70760 Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 26 Feb 2025 09:17:20 +0100
Subject: [PATCH 1/9] Add REPACK command.
The existing CLUSTER command as well as VACUUM with the FULL option both
reclaim unused space by rewriting table. Now that we want to enhance this
functionality (in particular, by adding a new option CONCURRENTLY), we should
enhance both commands because they are both implemented by the same function
(cluster.c:cluster_rel). However, adding the same option to two different
commands is not very user-friendly. Therefore it was decided to create a new
command and to declare both CLUSTER command and the FULL option of VACUUM
deprecated. Future enhancements to this rewriting code will only affect the
new command.
Like CLUSTER, the REPACK command reorders the table according to the specified
index. Unlike CLUSTER, REPACK does not require the index: if only table is
specified, the command acts as VACUUM FULL. As we don't want to remove CLUSTER
and VACUUM FULL yet, there are three callers of the cluster_rel() function
now: REPACK, CLUSTER and VACUUM FULL. When we need to distinguish who is
calling this function (mostly for logging, but also for progress reporting),
we can no longer use the OID of the clustering index: both REPACK and VACUUM
FULL can pass InvalidOid. Therefore, this patch introduces a new enumeration
type ClusterCommand, and adds an argument of this type to the cluster_rel()
function and to all the functions that need to distinguish the caller.
Like CLUSTER and VACUUM FULL, the REPACK COMMAND without arguments processes
all the tables on which the current user has the MAINTAIN privilege.
A new view pg_stat_progress_repack view is added to monitor the progress of
REPACK. Currently it displays the same information as pg_stat_progress_cluster
(except that column names might differ), but it'll also display the status of
the REPACK CONCURRENTLY command in the future, so the view definitions will
eventually diverge.
Regarding user documentation, the patch moves the information on clustering
from cluster.sgml to the new file repack.sgml. cluster.sgml now contains a
link that points to the related section of repack.sgml. A note on deprecation
and a link to repack.sgml are added to both cluster.sgml and vacuum.sgml.
---
doc/src/sgml/monitoring.sgml | 230 +++++++++++
doc/src/sgml/ref/allfiles.sgml | 1 +
doc/src/sgml/ref/cluster.sgml | 79 +---
doc/src/sgml/ref/repack.sgml | 254 ++++++++++++
doc/src/sgml/ref/vacuum.sgml | 8 +
doc/src/sgml/reference.sgml | 1 +
src/backend/access/heap/heapam_handler.c | 32 +-
src/backend/catalog/index.c | 2 +-
src/backend/catalog/system_views.sql | 27 ++
src/backend/commands/cluster.c | 496 +++++++++++++++++------
src/backend/commands/tablecmds.c | 3 +-
src/backend/commands/vacuum.c | 3 +-
src/backend/parser/gram.y | 63 ++-
src/backend/tcop/utility.c | 9 +
src/backend/utils/adt/pgstatfuncs.c | 2 +
src/bin/psql/tab-complete.in.c | 31 +-
src/include/commands/cluster.h | 22 +-
src/include/commands/progress.h | 60 ++-
src/include/nodes/parsenodes.h | 13 +
src/include/parser/kwlist.h | 1 +
src/include/tcop/cmdtaglist.h | 1 +
src/include/utils/backend_progress.h | 1 +
src/test/regress/expected/cluster.out | 180 ++++++++
src/test/regress/expected/rules.out | 27 ++
src/test/regress/sql/cluster.sql | 73 ++++
src/tools/pgindent/typedefs.list | 2 +
26 files changed, 1385 insertions(+), 236 deletions(-)
create mode 100644 doc/src/sgml/ref/repack.sgml
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 9178f1d34e..58e1becf02 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -400,6 +400,14 @@ postgres 27093 0.0 0.0 30096 2752 ? Ss 11:34 0:00 postgres: ser
</entry>
</row>
+ <row>
+ <entry><structname>pg_stat_progress_repack</structname><indexterm><primary>pg_stat_progress_repack</primary></indexterm></entry>
+ <entry>One row for each backend running
+ <command>REPACK</command>, showing current progress. See
+ <xref linkend="repack-progress-reporting"/>.
+ </entry>
+ </row>
+
<row>
<entry><structname>pg_stat_progress_basebackup</structname><indexterm><primary>pg_stat_progress_basebackup</primary></indexterm></entry>
<entry>One row for each WAL sender process streaming a base backup,
@@ -5885,6 +5893,228 @@ FROM pg_stat_get_backend_idset() AS backendid;
</table>
</sect2>
+ <sect2 id="repack-progress-reporting">
+ <title>REPACK Progress Reporting</title>
+
+ <indexterm>
+ <primary>pg_stat_progress_repack</primary>
+ </indexterm>
+
+ <para>
+ Whenever <command>REPACK</command> is running,
+ the <structname>pg_stat_progress_repack</structname> view will contain a
+ row for each backend that is currently running the command. The tables
+ below describe the information that will be reported and provide
+ information about how to interpret it.
+ </para>
+
+ <table id="pg-stat-progress-repack-view" xreflabel="pg_stat_progress_repack">
+ <title><structname>pg_stat_progress_repack</structname> View</title>
+ <tgroup cols="1">
+ <thead>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ Column Type
+ </para>
+ <para>
+ Description
+ </para></entry>
+ </row>
+ </thead>
+
+ <tbody>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>pid</structfield> <type>integer</type>
+ </para>
+ <para>
+ Process ID of backend.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>datid</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the database to which this backend is connected.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>datname</structfield> <type>name</type>
+ </para>
+ <para>
+ Name of the database to which this backend is connected.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>relid</structfield> <type>oid</type>
+ </para>
+ <para>
+ OID of the table being repacked.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>command</structfield> <type>text</type>
+ </para>
+ <para>
+ The command that is running. Currently, the only value
+ is <literal>REPACK</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>phase</structfield> <type>text</type>
+ </para>
+ <para>
+ Current processing phase. See <xref linkend="repack-phases"/>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>repack_index_relid</structfield> <type>oid</type>
+ </para>
+ <para>
+ If the table is being scanned using an index, this is the OID of the
+ index being used; otherwise, it is zero.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_tuples_scanned</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap tuples scanned.
+ This counter only advances when the phase is
+ <literal>seq scanning heap</literal>,
+ <literal>index scanning heap</literal>
+ or <literal>writing new heap</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_tuples_written</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap tuples written.
+ This counter only advances when the phase is
+ <literal>seq scanning heap</literal>,
+ <literal>index scanning heap</literal>
+ or <literal>writing new heap</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_blks_total</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Total number of heap blocks in the table. This number is reported
+ as of the beginning of <literal>seq scanning heap</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_blks_scanned</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap blocks scanned. This counter only advances when the
+ phase is <literal>seq scanning heap</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_rebuild_count</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of indexes rebuilt. This counter only advances when the phase
+ is <literal>rebuilding index</literal>.
+ </para></entry>
+ </row>
+ </tbody>
+ </tgroup>
+ </table>
+
+ <table id="repack-phases">
+ <title>REPACK Phases</title>
+ <tgroup cols="2">
+ <colspec colname="col1" colwidth="1*"/>
+ <colspec colname="col2" colwidth="2*"/>
+ <thead>
+ <row>
+ <entry>Phase</entry>
+ <entry>Description</entry>
+ </row>
+ </thead>
+
+ <tbody>
+ <row>
+ <entry><literal>initializing</literal></entry>
+ <entry>
+ The command is preparing to begin scanning the heap. This phase is
+ expected to be very brief.
+ </entry>
+ </row>
+ <row>
+ <entry><literal>seq scanning heap</literal></entry>
+ <entry>
+ The command is currently scanning the table using a sequential scan.
+ </entry>
+ </row>
+ <row>
+ <entry><literal>index scanning heap</literal></entry>
+ <entry>
+ <command>REPACK</command> is currently scanning the table using an index scan.
+ </entry>
+ </row>
+ <row>
+ <entry><literal>sorting tuples</literal></entry>
+ <entry>
+ <command>REPACK</command> is currently sorting tuples.
+ </entry>
+ </row>
+ <row>
+ <entry><literal>writing new heap</literal></entry>
+ <entry>
+ <command>REPACK</command> is currently writing the new heap.
+ </entry>
+ </row>
+ <row>
+ <entry><literal>swapping relation files</literal></entry>
+ <entry>
+ The command is currently swapping newly-built files into place.
+ </entry>
+ </row>
+ <row>
+ <entry><literal>rebuilding index</literal></entry>
+ <entry>
+ The command is currently rebuilding an index.
+ </entry>
+ </row>
+ <row>
+ <entry><literal>performing final cleanup</literal></entry>
+ <entry>
+ The command is performing final cleanup. When this phase is
+ completed, <command>REPACK</command> will end.
+ </entry>
+ </row>
+ </tbody>
+ </tgroup>
+ </table>
+ </sect2>
+
<sect2 id="copy-progress-reporting">
<title>COPY Progress Reporting</title>
diff --git a/doc/src/sgml/ref/allfiles.sgml b/doc/src/sgml/ref/allfiles.sgml
index f5be638867..c0ef654fcb 100644
--- a/doc/src/sgml/ref/allfiles.sgml
+++ b/doc/src/sgml/ref/allfiles.sgml
@@ -167,6 +167,7 @@ Complete list of usable sgml source files in this directory.
<!ENTITY refreshMaterializedView SYSTEM "refresh_materialized_view.sgml">
<!ENTITY reindex SYSTEM "reindex.sgml">
<!ENTITY releaseSavepoint SYSTEM "release_savepoint.sgml">
+<!ENTITY repack SYSTEM "repack.sgml">
<!ENTITY reset SYSTEM "reset.sgml">
<!ENTITY revoke SYSTEM "revoke.sgml">
<!ENTITY rollback SYSTEM "rollback.sgml">
diff --git a/doc/src/sgml/ref/cluster.sgml b/doc/src/sgml/ref/cluster.sgml
index 8811f169ea..54bb2362c8 100644
--- a/doc/src/sgml/ref/cluster.sgml
+++ b/doc/src/sgml/ref/cluster.sgml
@@ -42,17 +42,23 @@ CLUSTER [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <r
<replaceable class="parameter">table_name</replaceable>.
</para>
- <para>
- When a table is clustered, it is physically reordered
- based on the index information. Clustering is a one-time operation:
- when the table is subsequently updated, the changes are
- not clustered. That is, no attempt is made to store new or
- updated rows according to their index order. (If one wishes, one can
- periodically recluster by issuing the command again. Also, setting
- the table's <literal>fillfactor</literal> storage parameter to less than
- 100% can aid in preserving cluster ordering during updates, since updated
- rows are kept on the same page if enough space is available there.)
- </para>
+ <warning>
+ <para>
+ The <command>CLUSTER</command> command is deprecated in favor of
+ <xref linkend="sql-repack"/>.
+ </para>
+ </warning>
+
+ <note>
+ <para>
+ <xref linkend="sql-repack-notes-on-clustering"/> explain how clustering
+ works, whether it is initiated by <command>CLUSTER</command> or
+ by <command>REPACK</command>. The notable difference between the two is
+ that <command>REPACK</command> does not remember the index used last
+ time. Thus if you don't specify an index, <command>REPACK</command>
+ rewrites the table but does not try to cluster it.
+ </para>
+ </note>
<para>
When a table is clustered, <productname>PostgreSQL</productname>
@@ -136,63 +142,12 @@ CLUSTER [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <r
on the table.
</para>
- <para>
- In cases where you are accessing single rows randomly
- within a table, the actual order of the data in the
- table is unimportant. However, if you tend to access some
- data more than others, and there is an index that groups
- them together, you will benefit from using <command>CLUSTER</command>.
- If you are requesting a range of indexed values from a table, or a
- single indexed value that has multiple rows that match,
- <command>CLUSTER</command> will help because once the index identifies the
- table page for the first row that matches, all other rows
- that match are probably already on the same table page,
- and so you save disk accesses and speed up the query.
- </para>
-
- <para>
- <command>CLUSTER</command> can re-sort the table using either an index scan
- on the specified index, or (if the index is a b-tree) a sequential
- scan followed by sorting. It will attempt to choose the method that
- will be faster, based on planner cost parameters and available statistical
- information.
- </para>
-
<para>
While <command>CLUSTER</command> is running, the <xref
linkend="guc-search-path"/> is temporarily changed to <literal>pg_catalog,
pg_temp</literal>.
</para>
- <para>
- When an index scan is used, a temporary copy of the table is created that
- contains the table data in the index order. Temporary copies of each
- index on the table are created as well. Therefore, you need free space on
- disk at least equal to the sum of the table size and the index sizes.
- </para>
-
- <para>
- When a sequential scan and sort is used, a temporary sort file is
- also created, so that the peak temporary space requirement is as much
- as double the table size, plus the index sizes. This method is often
- faster than the index scan method, but if the disk space requirement is
- intolerable, you can disable this choice by temporarily setting <xref
- linkend="guc-enable-sort"/> to <literal>off</literal>.
- </para>
-
- <para>
- It is advisable to set <xref linkend="guc-maintenance-work-mem"/> to
- a reasonably large value (but not more than the amount of RAM you can
- dedicate to the <command>CLUSTER</command> operation) before clustering.
- </para>
-
- <para>
- Because the planner records statistics about the ordering of
- tables, it is advisable to run <link linkend="sql-analyze"><command>ANALYZE</command></link>
- on the newly clustered table.
- Otherwise, the planner might make poor choices of query plans.
- </para>
-
<para>
Because <command>CLUSTER</command> remembers which indexes are clustered,
one can cluster the tables one wants clustered manually the first time,
diff --git a/doc/src/sgml/ref/repack.sgml b/doc/src/sgml/ref/repack.sgml
new file mode 100644
index 0000000000..84f3c3e3f2
--- /dev/null
+++ b/doc/src/sgml/ref/repack.sgml
@@ -0,0 +1,254 @@
+<!--
+doc/src/sgml/ref/repack.sgml
+PostgreSQL documentation
+-->
+
+<refentry id="sql-repack">
+ <indexterm zone="sql-repack">
+ <primary>REPACK</primary>
+ </indexterm>
+
+ <refmeta>
+ <refentrytitle>REPACK</refentrytitle>
+ <manvolnum>7</manvolnum>
+ <refmiscinfo>SQL - Language Statements</refmiscinfo>
+ </refmeta>
+
+ <refnamediv>
+ <refname>REPACK</refname>
+ <refpurpose>cluster a table according to an index</refpurpose>
+ </refnamediv>
+
+ <refsynopsisdiv>
+<synopsis>
+REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <replaceable class="parameter">table_name</replaceable> [ USING INDEX<replaceable class="parameter">index_name</replaceable> ] ]
+
+<phrase>where <replaceable class="parameter">option</replaceable> can be one of:</phrase>
+
+ VERBOSE [ <replaceable class="parameter">boolean</replaceable> ]
+</synopsis>
+ </refsynopsisdiv>
+
+ <refsect1>
+ <title>Description</title>
+
+ <para>
+ <command>REPACK</command> reclaims storage occupied by dead
+ tuples. Unlike <command>VACUUM</command>, it does so by rewriting the
+ entire contents of the table specified
+ by <replaceable class="parameter">table_name</replaceable> into a new disk
+ file with no extra space (except for the space guaranteed by
+ the <literal>fillfactor</literal> storage parameter), allowing unused space
+ to be returned to the operating system.
+ </para>
+
+ <para>
+ Without
+ a <replaceable class="parameter">table_name</replaceable>, <command>REPACK</command>
+ processes every table and materialized view in the current database that
+ the current user has the <literal>MAINTAIN</literal> privilege on. This
+ form of <command>REPACK</command> cannot be executed inside a transaction
+ block.
+ </para>
+
+ <para>
+ If <replaceable class="parameter">index_name</replaceable> is specified,
+ the table is clustered by this index. Please see the notes on clustering
+ below.
+ </para>
+
+ <para>
+ When a table is being repacked, an <literal>ACCESS EXCLUSIVE</literal> lock
+ is acquired on it. This prevents any other database operations (both reads
+ and writes) from operating on the table until the <command>REPACK</command>
+ is finished.
+ </para>
+
+ <refsect2 id="sql-repack-notes-on-clustering" xreflabel="Notes on Clustering">
+ <title>Notes on Clustering</title>
+
+ <para>
+ When a table is clustered, it is physically reordered based on the index
+ information. Clustering is a one-time operation: when the table is
+ subsequently updated, the changes are not clustered. That is, no attempt
+ is made to store new or updated rows according to their index order. (If
+ one wishes, one can periodically recluster by issuing the command again.
+ Also, setting the table's <literal>fillfactor</literal> storage parameter
+ to less than 100% can aid in preserving cluster ordering during updates,
+ since updated rows are kept on the same page if enough space is available
+ there.)
+ </para>
+
+ <para>
+ In cases where you are accessing single rows randomly within a table, the
+ actual order of the data in the table is unimportant. However, if you tend
+ to access some data more than others, and there is an index that groups
+ them together, you will benefit from using <command>REPACK</command>. If
+ you are requesting a range of indexed values from a table, or a single
+ indexed value that has multiple rows that match,
+ <command>REPACK</command> will help because once the index identifies the
+ table page for the first row that matches, all other rows that match are
+ probably already on the same table page, and so you save disk accesses and
+ speed up the query.
+ </para>
+
+ <para>
+ <command>REPACK</command> can re-sort the table using either an index scan
+ on the specified index (if the index is a b-tree), or a sequential scan
+ followed by sorting. It will attempt to choose the method that will be
+ faster, based on planner cost parameters and available statistical
+ information.
+ </para>
+
+ <para>
+ Because the planner records statistics about the ordering of tables, it is
+ advisable to
+ run <link linkend="sql-analyze"><command>ANALYZE</command></link> on the
+ newly repacked table. Otherwise, the planner might make poor choices of
+ query plans.
+ </para>
+ </refsect2>
+
+ <refsect2 id="sql-repack-notes-on-resources" xreflabel="Notes on Resources">
+ <title>Notes on Resources</title>
+
+ <para>
+ When an index scan or a sequential scan without sort is used, a temporary
+ copy of the table is created that contains the table data in the index
+ order. Temporary copies of each index on the table are created as well.
+ Therefore, you need free space on disk at least equal to the sum of the
+ table size and the index sizes.
+ </para>
+
+ <para>
+ When a sequential scan and sort is used, a temporary sort file is also
+ created, so that the peak temporary space requirement is as much as double
+ the table size, plus the index sizes. This method is often faster than
+ the index scan method, but if the disk space requirement is intolerable,
+ you can disable this choice by temporarily setting
+ <xref linkend="guc-enable-sort"/> to <literal>off</literal>.
+ </para>
+
+ <para>
+ It is advisable to set <xref linkend="guc-maintenance-work-mem"/> to a
+ reasonably large value (but not more than the amount of RAM you can
+ dedicate to the <command>REPACK</command> operation) before repacking.
+ </para>
+ </refsect2>
+
+ </refsect1>
+
+ <refsect1>
+ <title>Parameters</title>
+
+ <variablelist>
+ <varlistentry>
+ <term><replaceable class="parameter">table_name</replaceable></term>
+ <listitem>
+ <para>
+ The name (possibly schema-qualified) of a table.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="parameter">index_name</replaceable></term>
+ <listitem>
+ <para>
+ The name of an index.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><literal>VERBOSE</literal></term>
+ <listitem>
+ <para>
+ Prints a progress report as each table is clustered
+ at <literal>INFO</literal> level.
+ </para>
+ </listitem>
+ </varlistentry>
+
+ <varlistentry>
+ <term><replaceable class="parameter">boolean</replaceable></term>
+ <listitem>
+ <para>
+ Specifies whether the selected option should be turned on or off.
+ You can write <literal>TRUE</literal>, <literal>ON</literal>, or
+ <literal>1</literal> to enable the option, and <literal>FALSE</literal>,
+ <literal>OFF</literal>, or <literal>0</literal> to disable it. The
+ <replaceable class="parameter">boolean</replaceable> value can also
+ be omitted, in which case <literal>TRUE</literal> is assumed.
+ </para>
+ </listitem>
+ </varlistentry>
+ </variablelist>
+ </refsect1>
+
+ <refsect1>
+ <title>Notes</title>
+
+ <para>
+ To repack a table, one must have the <literal>MAINTAIN</literal> privilege
+ on the table.
+ </para>
+
+ <para>
+ While <command>REPACK</command> is running, the <xref
+ linkend="guc-search-path"/> is temporarily changed to <literal>pg_catalog,
+ pg_temp</literal>.
+ </para>
+
+ <para>
+ Each backend running <command>REPACK</command> will report its progress
+ in the <structname>pg_stat_progress_repack</structname> view. See
+ <xref linkend="repack-progress-reporting"/> for details.
+ </para>
+
+ <para>
+ Repacking a partitioned table repacks each of its partitions. If an index
+ is specified, each partition is clustered using the partition of that
+ index. <command>REPACK</command> on a partitioned table cannot be executed
+ inside a transaction block.
+ </para>
+
+ </refsect1>
+
+ <refsect1>
+ <title>Examples</title>
+
+ <para>
+ Repack the table <literal>employees</literal>:
+<programlisting>
+REPACK employees;
+</programlisting>
+ </para>
+
+
+ <para>
+ Cluster the table <literal>employees</literal> on the basis of its
+ index <literal>employees_ind</literal>:
+<programlisting>
+REPACK employees USING INDEX employees_ind;
+</programlisting>
+ </para>
+
+ <para>
+ Repack all tables in the database on which you have
+ the <literal>MAINTAIN</literal> privilege:
+<programlisting>
+REPACK;
+</programlisting></para>
+ </refsect1>
+
+ <refsect1>
+ <title>Compatibility</title>
+
+ <para>
+ There is no <command>REPACK</command> statement in the SQL standard.
+ </para>
+
+ </refsect1>
+
+</refentry>
diff --git a/doc/src/sgml/ref/vacuum.sgml b/doc/src/sgml/ref/vacuum.sgml
index 971b1237d4..2b5a5d0ac4 100644
--- a/doc/src/sgml/ref/vacuum.sgml
+++ b/doc/src/sgml/ref/vacuum.sgml
@@ -98,6 +98,14 @@ VACUUM [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <re
<varlistentry>
<term><literal>FULL</literal></term>
<listitem>
+
+ <warning>
+ <para>
+ The <command>FULL</command> parameter is deprecated in favor of
+ <xref linkend="sql-repack"/>.
+ </para>
+ </warning>
+
<para>
Selects <quote>full</quote> vacuum, which can reclaim more
space, but takes much longer and exclusively locks the table.
diff --git a/doc/src/sgml/reference.sgml b/doc/src/sgml/reference.sgml
index ff85ace83f..229912d35b 100644
--- a/doc/src/sgml/reference.sgml
+++ b/doc/src/sgml/reference.sgml
@@ -195,6 +195,7 @@
&refreshMaterializedView;
&reindex;
&releaseSavepoint;
+ &repack;
&reset;
&revoke;
&rollback;
diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index e78682c3ce..5c3cab8bc2 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -737,13 +737,13 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
if (OldIndex != NULL && !use_sort)
{
const int ci_index[] = {
- PROGRESS_CLUSTER_PHASE,
- PROGRESS_CLUSTER_INDEX_RELID
+ PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_INDEX_RELID
};
int64 ci_val[2];
/* Set phase and OIDOldIndex to columns */
- ci_val[0] = PROGRESS_CLUSTER_PHASE_INDEX_SCAN_HEAP;
+ ci_val[0] = PROGRESS_REPACK_PHASE_INDEX_SCAN_HEAP;
ci_val[1] = RelationGetRelid(OldIndex);
pgstat_progress_update_multi_param(2, ci_index, ci_val);
@@ -755,15 +755,15 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
else
{
/* In scan-and-sort mode and also VACUUM FULL, set phase */
- pgstat_progress_update_param(PROGRESS_CLUSTER_PHASE,
- PROGRESS_CLUSTER_PHASE_SEQ_SCAN_HEAP);
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_SEQ_SCAN_HEAP);
tableScan = table_beginscan(OldHeap, SnapshotAny, 0, (ScanKey) NULL);
heapScan = (HeapScanDesc) tableScan;
indexScan = NULL;
/* Set total heap blocks */
- pgstat_progress_update_param(PROGRESS_CLUSTER_TOTAL_HEAP_BLKS,
+ pgstat_progress_update_param(PROGRESS_REPACK_TOTAL_HEAP_BLKS,
heapScan->rs_nblocks);
}
@@ -805,7 +805,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
* is manually updated to the correct value when the table
* scan finishes.
*/
- pgstat_progress_update_param(PROGRESS_CLUSTER_HEAP_BLKS_SCANNED,
+ pgstat_progress_update_param(PROGRESS_REPACK_HEAP_BLKS_SCANNED,
heapScan->rs_nblocks);
break;
}
@@ -821,7 +821,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
*/
if (prev_cblock != heapScan->rs_cblock)
{
- pgstat_progress_update_param(PROGRESS_CLUSTER_HEAP_BLKS_SCANNED,
+ pgstat_progress_update_param(PROGRESS_REPACK_HEAP_BLKS_SCANNED,
(heapScan->rs_cblock +
heapScan->rs_nblocks -
heapScan->rs_startblock
@@ -908,14 +908,14 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
* In scan-and-sort mode, report increase in number of tuples
* scanned
*/
- pgstat_progress_update_param(PROGRESS_CLUSTER_HEAP_TUPLES_SCANNED,
+ pgstat_progress_update_param(PROGRESS_REPACK_HEAP_TUPLES_SCANNED,
*num_tuples);
}
else
{
const int ct_index[] = {
- PROGRESS_CLUSTER_HEAP_TUPLES_SCANNED,
- PROGRESS_CLUSTER_HEAP_TUPLES_WRITTEN
+ PROGRESS_REPACK_HEAP_TUPLES_SCANNED,
+ PROGRESS_REPACK_HEAP_TUPLES_WRITTEN
};
int64 ct_val[2];
@@ -948,14 +948,14 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
double n_tuples = 0;
/* Report that we are now sorting tuples */
- pgstat_progress_update_param(PROGRESS_CLUSTER_PHASE,
- PROGRESS_CLUSTER_PHASE_SORT_TUPLES);
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_SORT_TUPLES);
tuplesort_performsort(tuplesort);
/* Report that we are now writing new heap */
- pgstat_progress_update_param(PROGRESS_CLUSTER_PHASE,
- PROGRESS_CLUSTER_PHASE_WRITE_NEW_HEAP);
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_WRITE_NEW_HEAP);
for (;;)
{
@@ -973,7 +973,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
values, isnull,
rwstate);
/* Report n_tuples */
- pgstat_progress_update_param(PROGRESS_CLUSTER_HEAP_TUPLES_WRITTEN,
+ pgstat_progress_update_param(PROGRESS_REPACK_HEAP_TUPLES_WRITTEN,
n_tuples);
}
diff --git a/src/backend/catalog/index.c b/src/backend/catalog/index.c
index f37b990c81..c84f67059a 100644
--- a/src/backend/catalog/index.c
+++ b/src/backend/catalog/index.c
@@ -4051,7 +4051,7 @@ reindex_relation(const ReindexStmt *stmt, Oid relid, int flags,
Assert(!ReindexIsProcessingIndex(indexOid));
/* Set index rebuild count */
- pgstat_progress_update_param(PROGRESS_CLUSTER_INDEX_REBUILD_COUNT,
+ pgstat_progress_update_param(PROGRESS_REPACK_INDEX_REBUILD_COUNT,
i);
i++;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index a4d2cfdcaf..b8209b2acd 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1262,6 +1262,33 @@ CREATE VIEW pg_stat_progress_cluster AS
FROM pg_stat_get_progress_info('CLUSTER') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
+CREATE VIEW pg_stat_progress_repack AS
+ SELECT
+ S.pid AS pid,
+ S.datid AS datid,
+ D.datname AS datname,
+ S.relid AS relid,
+ CASE S.param1 WHEN 1 THEN 'REPACK'
+ END AS command,
+ CASE S.param2 WHEN 0 THEN 'initializing'
+ WHEN 1 THEN 'seq scanning heap'
+ WHEN 2 THEN 'index scanning heap'
+ WHEN 3 THEN 'sorting tuples'
+ WHEN 4 THEN 'writing new heap'
+ WHEN 5 THEN 'swapping relation files'
+ WHEN 6 THEN 'rebuilding index'
+ WHEN 7 THEN 'performing final cleanup'
+ END AS phase,
+ CAST(S.param3 AS oid) AS repack_index_relid,
+ S.param4 AS heap_tuples_scanned,
+ S.param5 AS heap_tuples_written,
+ S.param6 AS heap_blks_total,
+ S.param7 AS heap_blks_scanned,
+ S.param8 AS index_rebuild_count
+ FROM pg_stat_get_progress_info('REPACK') AS S
+ LEFT JOIN pg_database D ON S.datid = D.oid;
+
+
CREATE VIEW pg_stat_progress_create_index AS
SELECT
S.pid AS pid, S.datid AS datid, D.datname AS datname,
diff --git a/src/backend/commands/cluster.c b/src/backend/commands/cluster.c
index 99193f5c88..d0f2588a97 100644
--- a/src/backend/commands/cluster.c
+++ b/src/backend/commands/cluster.c
@@ -46,6 +46,7 @@
#include "storage/lmgr.h"
#include "storage/predicate.h"
#include "utils/acl.h"
+#include "utils/formatting.h"
#include "utils/fmgroids.h"
#include "utils/guc.h"
#include "utils/inval.h"
@@ -67,17 +68,33 @@ typedef struct
Oid indexOid;
} RelToCluster;
-
-static void cluster_multiple_rels(List *rtcs, ClusterParams *params);
-static void rebuild_relation(Relation OldHeap, Relation index, bool verbose);
+/*
+ * Map the value of ClusterCommand to string.
+ */
+#define CLUSTER_COMMAND_STR(cmd) ((cmd) == CLUSTER_COMMAND_CLUSTER ? \
+ "cluster" : \
+ ((cmd) == CLUSTER_COMMAND_REPACK ? \
+ "repack" : "vacuum"))
+
+static void cluster_multiple_rels(List *rtcs, ClusterParams *params,
+ ClusterCommand cmd);
+static void rebuild_relation(Relation OldHeap, Relation index, bool verbose,
+ ClusterCommand cmd);
static void copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
- bool verbose, bool *pSwapToastByContent,
+ bool verbose, ClusterCommand cmd,
+ bool *pSwapToastByContent,
TransactionId *pFreezeXid, MultiXactId *pCutoffMulti);
static List *get_tables_to_cluster(MemoryContext cluster_context);
+static List *get_tables_to_repack(MemoryContext repack_context);
static List *get_tables_to_cluster_partitioned(MemoryContext cluster_context,
- Oid indexOid);
-static bool cluster_is_permitted_for_relation(Oid relid, Oid userid);
-
+ Oid relid, bool rel_is_index,
+ ClusterCommand cmd);
+static bool cluster_is_permitted_for_relation(Oid relid, Oid userid,
+ ClusterCommand cmd);
+static Relation process_single_relation(RangeVar *relation, char *indexname,
+ ClusterCommand cmd,
+ ClusterParams *params,
+ Oid *indexOid_p);
/*---------------------------------------------------------------------------
* This cluster code allows for clustering multiple tables at once. Because
@@ -133,72 +150,11 @@ cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel)
if (stmt->relation != NULL)
{
- /* This is the single-relation case. */
- Oid tableOid;
-
- /*
- * Find, lock, and check permissions on the table. We obtain
- * AccessExclusiveLock right away to avoid lock-upgrade hazard in the
- * single-transaction case.
- */
- tableOid = RangeVarGetRelidExtended(stmt->relation,
- AccessExclusiveLock,
- 0,
- RangeVarCallbackMaintainsTable,
- NULL);
- rel = table_open(tableOid, NoLock);
-
- /*
- * Reject clustering a remote temp table ... their local buffer
- * manager is not going to cope.
- */
- if (RELATION_IS_OTHER_TEMP(rel))
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("cannot cluster temporary tables of other sessions")));
-
- if (stmt->indexname == NULL)
- {
- ListCell *index;
-
- /* We need to find the index that has indisclustered set. */
- foreach(index, RelationGetIndexList(rel))
- {
- indexOid = lfirst_oid(index);
- if (get_index_isclustered(indexOid))
- break;
- indexOid = InvalidOid;
- }
-
- if (!OidIsValid(indexOid))
- ereport(ERROR,
- (errcode(ERRCODE_UNDEFINED_OBJECT),
- errmsg("there is no previously clustered index for table \"%s\"",
- stmt->relation->relname)));
- }
- else
- {
- /*
- * The index is expected to be in the same namespace as the
- * relation.
- */
- indexOid = get_relname_relid(stmt->indexname,
- rel->rd_rel->relnamespace);
- if (!OidIsValid(indexOid))
- ereport(ERROR,
- (errcode(ERRCODE_UNDEFINED_OBJECT),
- errmsg("index \"%s\" for table \"%s\" does not exist",
- stmt->indexname, stmt->relation->relname)));
- }
-
- /* For non-partitioned tables, do what we came here to do. */
- if (rel->rd_rel->relkind != RELKIND_PARTITIONED_TABLE)
- {
- cluster_rel(rel, indexOid, ¶ms);
- /* cluster_rel closes the relation, but keeps lock */
-
+ rel = process_single_relation(stmt->relation, stmt->indexname,
+ CLUSTER_COMMAND_CLUSTER, ¶ms,
+ &indexOid);
+ if (rel == NULL)
return;
- }
}
/*
@@ -230,8 +186,11 @@ cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel)
if (rel != NULL)
{
Assert(rel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE);
- check_index_is_clusterable(rel, indexOid, AccessShareLock);
- rtcs = get_tables_to_cluster_partitioned(cluster_context, indexOid);
+ check_index_is_clusterable(rel, indexOid, AccessShareLock,
+ CLUSTER_COMMAND_CLUSTER);
+ rtcs = get_tables_to_cluster_partitioned(cluster_context, indexOid,
+ true,
+ CLUSTER_COMMAND_CLUSTER);
/* close relation, releasing lock on parent table */
table_close(rel, AccessExclusiveLock);
@@ -243,7 +202,7 @@ cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel)
}
/* Do the job. */
- cluster_multiple_rels(rtcs, ¶ms);
+ cluster_multiple_rels(rtcs, ¶ms, CLUSTER_COMMAND_CLUSTER);
/* Start a new transaction for the cleanup work. */
StartTransactionCommand();
@@ -260,7 +219,8 @@ cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel)
* return.
*/
static void
-cluster_multiple_rels(List *rtcs, ClusterParams *params)
+cluster_multiple_rels(List *rtcs, ClusterParams *params,
+ ClusterCommand cmd)
{
ListCell *lc;
@@ -283,7 +243,7 @@ cluster_multiple_rels(List *rtcs, ClusterParams *params)
rel = table_open(rtc->tableOid, AccessExclusiveLock);
/* Process this table */
- cluster_rel(rel, rtc->indexOid, params);
+ cluster_rel(rel, rtc->indexOid, params, cmd);
/* cluster_rel closes the relation, but keeps lock */
PopActiveSnapshot();
@@ -306,9 +266,13 @@ cluster_multiple_rels(List *rtcs, ClusterParams *params)
* If indexOid is InvalidOid, the table will be rewritten in physical order
* instead of index order. This is the new implementation of VACUUM FULL,
* and error messages should refer to the operation as VACUUM not CLUSTER.
+ *
+ * 'cmd' indicates which commands is being executed. REPACK should be the only
+ * caller of this function in the future.
*/
void
-cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params)
+cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
+ ClusterCommand cmd)
{
Oid tableOid = RelationGetRelid(OldHeap);
Oid save_userid;
@@ -317,19 +281,33 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params)
bool verbose = ((params->options & CLUOPT_VERBOSE) != 0);
bool recheck = ((params->options & CLUOPT_RECHECK) != 0);
Relation index;
+ const char *cmd_str = CLUSTER_COMMAND_STR(cmd);
Assert(CheckRelationLockedByMe(OldHeap, AccessExclusiveLock, false));
/* Check for user-requested abort. */
CHECK_FOR_INTERRUPTS();
- pgstat_progress_start_command(PROGRESS_COMMAND_CLUSTER, tableOid);
- if (OidIsValid(indexOid))
- pgstat_progress_update_param(PROGRESS_CLUSTER_COMMAND,
+ if (cmd == CLUSTER_COMMAND_REPACK)
+ pgstat_progress_start_command(PROGRESS_COMMAND_REPACK, tableOid);
+ else
+ pgstat_progress_start_command(PROGRESS_COMMAND_CLUSTER, tableOid);
+
+ if (cmd == CLUSTER_COMMAND_REPACK)
+ pgstat_progress_update_param(PROGRESS_REPACK_COMMAND,
+ PROGRESS_REPACK_COMMAND_REPACK);
+ else if (OidIsValid(indexOid))
+ {
+ Assert(cmd == CLUSTER_COMMAND_CLUSTER);
+ pgstat_progress_update_param(PROGRESS_REPACK_COMMAND,
PROGRESS_CLUSTER_COMMAND_CLUSTER);
+ }
else
- pgstat_progress_update_param(PROGRESS_CLUSTER_COMMAND,
+ {
+ Assert(cmd == CLUSTER_COMMAND_VACUUM);
+ pgstat_progress_update_param(PROGRESS_REPACK_COMMAND,
PROGRESS_CLUSTER_COMMAND_VACUUM_FULL);
+ }
/*
* Switch to the table owner's userid, so that any index functions are run
@@ -353,7 +331,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params)
if (recheck)
{
/* Check that the user still has privileges for the relation */
- if (!cluster_is_permitted_for_relation(tableOid, save_userid))
+ if (!cluster_is_permitted_for_relation(tableOid, save_userid, cmd))
{
relation_close(OldHeap, AccessExclusiveLock);
goto out;
@@ -403,39 +381,38 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params)
* would work in most respects, but the index would only get marked as
* indisclustered in the current database, leading to unexpected behavior
* if CLUSTER were later invoked in another database.
+ *
+ * REPACK does not set indisclustered. XXX Not sure I understand the
+ * comment above: how can an attribute be set "only in the current
+ * database"?
*/
- if (OidIsValid(indexOid) && OldHeap->rd_rel->relisshared)
+ if (cmd == CLUSTER_COMMAND_CLUSTER && OldHeap->rd_rel->relisshared)
ereport(ERROR,
(errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("cannot cluster a shared catalog")));
+ errmsg("cannot %s a shared catalog", cmd_str)));
/*
* Don't process temp tables of other backends ... their local buffer
* manager is not going to cope.
*/
if (RELATION_IS_OTHER_TEMP(OldHeap))
- {
- if (OidIsValid(indexOid))
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("cannot cluster temporary tables of other sessions")));
- else
- ereport(ERROR,
- (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("cannot vacuum temporary tables of other sessions")));
- }
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot %s temporary tables of other sessions",
+ cmd_str)));
/*
* Also check for active uses of the relation in the current transaction,
* including open scans and pending AFTER trigger events.
*/
- CheckTableNotInUse(OldHeap, OidIsValid(indexOid) ? "CLUSTER" : "VACUUM");
+ CheckTableNotInUse(OldHeap, asc_toupper(cmd_str, strlen(cmd_str)));
/* Check heap and index are valid to cluster on */
if (OidIsValid(indexOid))
{
/* verify the index is good and lock it */
- check_index_is_clusterable(OldHeap, indexOid, AccessExclusiveLock);
+ check_index_is_clusterable(OldHeap, indexOid, AccessExclusiveLock,
+ cmd);
/* also open it */
index = index_open(indexOid, NoLock);
}
@@ -469,7 +446,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params)
TransferPredicateLocksToHeapRelation(OldHeap);
/* rebuild_relation does all the dirty work */
- rebuild_relation(OldHeap, index, verbose);
+ rebuild_relation(OldHeap, index, verbose, cmd);
/* rebuild_relation closes OldHeap, and index if valid */
out:
@@ -491,9 +468,11 @@ out:
* protection here.
*/
void
-check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode)
+check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode,
+ ClusterCommand cmd)
{
Relation OldIndex;
+ const char *cmd_str = CLUSTER_COMMAND_STR(cmd);
OldIndex = index_open(indexOid, lockmode);
@@ -512,8 +491,8 @@ check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode)
if (!OldIndex->rd_indam->amclusterable)
ereport(ERROR,
(errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("cannot cluster on index \"%s\" because access method does not support clustering",
- RelationGetRelationName(OldIndex))));
+ errmsg("cannot %s on index \"%s\" because access method does not support clustering",
+ cmd_str, RelationGetRelationName(OldIndex))));
/*
* Disallow clustering on incomplete indexes (those that might not index
@@ -524,7 +503,8 @@ check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode)
if (!heap_attisnull(OldIndex->rd_indextuple, Anum_pg_index_indpred, NULL))
ereport(ERROR,
(errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("cannot cluster on partial index \"%s\"",
+ errmsg("cannot %s on partial index \"%s\"",
+ cmd_str,
RelationGetRelationName(OldIndex))));
/*
@@ -538,8 +518,8 @@ check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode)
if (!OldIndex->rd_index->indisvalid)
ereport(ERROR,
(errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
- errmsg("cannot cluster on invalid index \"%s\"",
- RelationGetRelationName(OldIndex))));
+ errmsg("cannot %s on invalid index \"%s\"",
+ cmd_str, RelationGetRelationName(OldIndex))));
/* Drop relcache refcnt on OldIndex, but keep lock */
index_close(OldIndex, NoLock);
@@ -626,7 +606,8 @@ mark_index_clustered(Relation rel, Oid indexOid, bool is_internal)
* On exit, they are closed, but locks on them are not released.
*/
static void
-rebuild_relation(Relation OldHeap, Relation index, bool verbose)
+rebuild_relation(Relation OldHeap, Relation index, bool verbose,
+ ClusterCommand cmd)
{
Oid tableOid = RelationGetRelid(OldHeap);
Oid accessMethod = OldHeap->rd_rel->relam;
@@ -664,7 +645,7 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose)
NewHeap = table_open(OIDNewHeap, NoLock);
/* Copy the heap data into the new table in the desired order */
- copy_table_data(NewHeap, OldHeap, index, verbose,
+ copy_table_data(NewHeap, OldHeap, index, verbose, cmd,
&swap_toast_by_content, &frozenXid, &cutoffMulti);
@@ -829,8 +810,8 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
*/
static void
copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verbose,
- bool *pSwapToastByContent, TransactionId *pFreezeXid,
- MultiXactId *pCutoffMulti)
+ ClusterCommand cmd, bool *pSwapToastByContent,
+ TransactionId *pFreezeXid, MultiXactId *pCutoffMulti)
{
Relation relRelation;
HeapTuple reltup;
@@ -845,6 +826,7 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
tups_recently_dead = 0;
BlockNumber num_pages;
int elevel = verbose ? INFO : DEBUG2;
+ const char *cmd_str = CLUSTER_COMMAND_STR(cmd);
PGRUsage ru0;
char *nspname;
@@ -958,18 +940,21 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
/* Log what we're doing */
if (OldIndex != NULL && !use_sort)
ereport(elevel,
- (errmsg("clustering \"%s.%s\" using index scan on \"%s\"",
+ (errmsg("%sing \"%s.%s\" using index scan on \"%s\"",
+ cmd_str,
nspname,
RelationGetRelationName(OldHeap),
RelationGetRelationName(OldIndex))));
else if (use_sort)
ereport(elevel,
- (errmsg("clustering \"%s.%s\" using sequential scan and sort",
+ (errmsg("%sing \"%s.%s\" using sequential scan and sort",
+ cmd_str,
nspname,
RelationGetRelationName(OldHeap))));
else
ereport(elevel,
- (errmsg("vacuuming \"%s.%s\"",
+ (errmsg("%sing \"%s.%s\"",
+ cmd_str,
nspname,
RelationGetRelationName(OldHeap))));
@@ -1453,8 +1438,8 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
int i;
/* Report that we are now swapping relation files */
- pgstat_progress_update_param(PROGRESS_CLUSTER_PHASE,
- PROGRESS_CLUSTER_PHASE_SWAP_REL_FILES);
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_SWAP_REL_FILES);
/* Zero out possible results from swapped_relation_files */
memset(mapped_tables, 0, sizeof(mapped_tables));
@@ -1504,14 +1489,14 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
reindex_flags |= REINDEX_REL_FORCE_INDEXES_PERMANENT;
/* Report that we are now reindexing relations */
- pgstat_progress_update_param(PROGRESS_CLUSTER_PHASE,
- PROGRESS_CLUSTER_PHASE_REBUILD_INDEX);
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_REBUILD_INDEX);
reindex_relation(NULL, OIDOldHeap, reindex_flags, &reindex_params);
/* Report that we are now doing clean up */
- pgstat_progress_update_param(PROGRESS_CLUSTER_PHASE,
- PROGRESS_CLUSTER_PHASE_FINAL_CLEANUP);
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_FINAL_CLEANUP);
/*
* If the relation being rebuilt is pg_class, swap_relation_files()
@@ -1661,7 +1646,8 @@ get_tables_to_cluster(MemoryContext cluster_context)
index = (Form_pg_index) GETSTRUCT(indexTuple);
- if (!cluster_is_permitted_for_relation(index->indrelid, GetUserId()))
+ if (!cluster_is_permitted_for_relation(index->indrelid, GetUserId(),
+ CLUSTER_COMMAND_CLUSTER))
continue;
/* Use a permanent memory context for the result list */
@@ -1682,14 +1668,67 @@ get_tables_to_cluster(MemoryContext cluster_context)
}
/*
- * Given an index on a partitioned table, return a list of RelToCluster for
+ * Like get_tables_to_cluster(), but do not care about indexes.
+ */
+static List *
+get_tables_to_repack(MemoryContext repack_context)
+{
+ Relation relrelation;
+ TableScanDesc scan;
+ HeapTuple tuple;
+ MemoryContext old_context;
+ List *rtcs = NIL;
+
+ /*
+ * Get all indexes that have indisclustered set and that the current user
+ * has the appropriate privileges for.
+ */
+ relrelation = table_open(RelationRelationId, AccessShareLock);
+ scan = table_beginscan_catalog(relrelation, 0, NULL);
+ while ((tuple = heap_getnext(scan, ForwardScanDirection)) != NULL)
+ {
+ RelToCluster *rtc;
+ Form_pg_class relrelation = (Form_pg_class) GETSTRUCT(tuple);
+ Oid relid = relrelation->oid;
+
+ /* Only interested in relations. */
+ if (get_rel_relkind(relid) != RELKIND_RELATION)
+ continue;
+
+ if (!cluster_is_permitted_for_relation(relid, GetUserId(),
+ CLUSTER_COMMAND_REPACK))
+ continue;
+
+ /* Use a permanent memory context for the result list */
+ old_context = MemoryContextSwitchTo(repack_context);
+
+ rtc = (RelToCluster *) palloc(sizeof(RelToCluster));
+ rtc->tableOid = relid;
+ rtc->indexOid = InvalidOid;
+ rtcs = lappend(rtcs, rtc);
+
+ MemoryContextSwitchTo(old_context);
+ }
+ table_endscan(scan);
+
+ relation_close(relrelation, AccessShareLock);
+
+ return rtcs;
+}
+
+/*
+ * Given a partitioned table or its index, return a list of RelToCluster for
* all the children leaves tables/indexes.
*
* Like expand_vacuum_rel, but here caller must hold AccessExclusiveLock
* on the table containing the index.
+ *
+ * 'rel_is_index' tells whether 'relid' is that of an index (true) or of the
+ * owning relation.
*/
static List *
-get_tables_to_cluster_partitioned(MemoryContext cluster_context, Oid indexOid)
+get_tables_to_cluster_partitioned(MemoryContext cluster_context, Oid relid,
+ bool rel_is_index, ClusterCommand cmd)
{
List *inhoids;
ListCell *lc;
@@ -1697,17 +1736,33 @@ get_tables_to_cluster_partitioned(MemoryContext cluster_context, Oid indexOid)
MemoryContext old_context;
/* Do not lock the children until they're processed */
- inhoids = find_all_inheritors(indexOid, NoLock, NULL);
+ inhoids = find_all_inheritors(relid, NoLock, NULL);
foreach(lc, inhoids)
{
- Oid indexrelid = lfirst_oid(lc);
- Oid relid = IndexGetRelation(indexrelid, false);
+ Oid inhoid = lfirst_oid(lc);
+ Oid inhrelid,
+ inhindid;
RelToCluster *rtc;
- /* consider only leaf indexes */
- if (get_rel_relkind(indexrelid) != RELKIND_INDEX)
- continue;
+ if (rel_is_index)
+ {
+ /* consider only leaf indexes */
+ if (get_rel_relkind(inhoid) != RELKIND_INDEX)
+ continue;
+
+ inhrelid = IndexGetRelation(inhoid, false);
+ inhindid = inhoid;
+ }
+ else
+ {
+ /* consider only leaf relations */
+ if (get_rel_relkind(inhoid) != RELKIND_RELATION)
+ continue;
+
+ inhrelid = inhoid;
+ inhindid = InvalidOid;
+ }
/*
* It's possible that the user does not have privileges to CLUSTER the
@@ -1715,15 +1770,15 @@ get_tables_to_cluster_partitioned(MemoryContext cluster_context, Oid indexOid)
* table. We skip any partitions which the user is not permitted to
* CLUSTER.
*/
- if (!cluster_is_permitted_for_relation(relid, GetUserId()))
+ if (!cluster_is_permitted_for_relation(inhrelid, GetUserId(), cmd))
continue;
/* Use a permanent memory context for the result list */
old_context = MemoryContextSwitchTo(cluster_context);
rtc = (RelToCluster *) palloc(sizeof(RelToCluster));
- rtc->tableOid = relid;
- rtc->indexOid = indexrelid;
+ rtc->tableOid = inhrelid;
+ rtc->indexOid = inhindid;
rtcs = lappend(rtcs, rtc);
MemoryContextSwitchTo(old_context);
@@ -1737,13 +1792,192 @@ get_tables_to_cluster_partitioned(MemoryContext cluster_context, Oid indexOid)
* function emits a WARNING.
*/
static bool
-cluster_is_permitted_for_relation(Oid relid, Oid userid)
+cluster_is_permitted_for_relation(Oid relid, Oid userid, ClusterCommand cmd)
{
if (pg_class_aclcheck(relid, userid, ACL_MAINTAIN) == ACLCHECK_OK)
return true;
ereport(WARNING,
- (errmsg("permission denied to cluster \"%s\", skipping it",
+ (errmsg("permission denied to %s \"%s\", skipping it",
+ CLUSTER_COMMAND_STR(cmd),
get_rel_name(relid))));
return false;
}
+
+/*
+ * REPACK is intended to be a replacement of both CLUSTER and VACUUM FULL.
+ */
+void
+repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
+{
+ ListCell *lc;
+ ClusterParams params = {0};
+ bool verbose = false;
+ Relation rel = NULL;
+ Oid indexOid = InvalidOid;
+ MemoryContext repack_context;
+ List *rtcs;
+
+ /* Parse option list */
+ foreach(lc, stmt->params)
+ {
+ DefElem *opt = (DefElem *) lfirst(lc);
+
+ if (strcmp(opt->defname, "verbose") == 0)
+ verbose = defGetBoolean(opt);
+ else
+ ereport(ERROR,
+ (errcode(ERRCODE_SYNTAX_ERROR),
+ errmsg("unrecognized REPACK option \"%s\"",
+ opt->defname),
+ parser_errposition(pstate, opt->location)));
+ }
+
+ params.options = (verbose ? CLUOPT_VERBOSE : 0);
+
+ if (stmt->relation != NULL)
+ {
+ rel = process_single_relation(stmt->relation, stmt->indexname,
+ CLUSTER_COMMAND_REPACK, ¶ms,
+ &indexOid);
+ if (rel == NULL)
+ return;
+ }
+
+ /*
+ * By here, we know we are in a multi-table situation. In order to avoid
+ * holding locks for too long, we want to process each table in its own
+ * transaction. This forces us to disallow running inside a user
+ * transaction block.
+ */
+ PreventInTransactionBlock(isTopLevel, "REPACK");
+
+ /* Also, we need a memory context to hold our list of relations */
+ repack_context = AllocSetContextCreate(PortalContext,
+ "Repack",
+ ALLOCSET_DEFAULT_SIZES);
+
+ params.options |= CLUOPT_RECHECK;
+ if (rel != NULL)
+ {
+ Oid relid;
+ bool rel_is_index;
+
+ Assert(rel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE);
+
+ if (OidIsValid(indexOid))
+ {
+ relid = indexOid;
+ rel_is_index = true;
+ }
+ else
+ {
+ relid = RelationGetRelid(rel);
+ rel_is_index = false;
+ }
+ rtcs = get_tables_to_cluster_partitioned(repack_context, relid,
+ rel_is_index,
+ CLUSTER_COMMAND_REPACK);
+
+ /* close relation, releasing lock on parent table */
+ table_close(rel, AccessExclusiveLock);
+ }
+ else
+ rtcs = get_tables_to_repack(repack_context);
+
+ /* Do the job. */
+ cluster_multiple_rels(rtcs, ¶ms, CLUSTER_COMMAND_REPACK);
+
+ /* Start a new transaction for the cleanup work. */
+ StartTransactionCommand();
+
+ /* Clean up working storage */
+ MemoryContextDelete(repack_context);
+
+}
+
+/*
+ * REPACK a single relation.
+ *
+ * Return NULL if done, relation reference if the caller needs to process it
+ * (because the relation is partitioned).
+ */
+static Relation
+process_single_relation(RangeVar *relation, char *indexname,
+ ClusterCommand cmd, ClusterParams *params,
+ Oid *indexOid_p)
+{
+ Relation rel;
+ Oid indexOid = InvalidOid;
+
+ /* This is the single-relation case. */
+ Oid tableOid;
+
+ /*
+ * Find, lock, and check permissions on the table. We obtain
+ * AccessExclusiveLock right away to avoid lock-upgrade hazard in the
+ * single-transaction case.
+ */
+ tableOid = RangeVarGetRelidExtended(relation,
+ AccessExclusiveLock,
+ 0,
+ RangeVarCallbackMaintainsTable,
+ NULL);
+ rel = table_open(tableOid, NoLock);
+
+ /*
+ * Reject clustering a remote temp table ... their local buffer manager is
+ * not going to cope.
+ */
+ if (RELATION_IS_OTHER_TEMP(rel))
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot %s temporary tables of other sessions",
+ CLUSTER_COMMAND_STR(cmd))));
+
+ if (indexname == NULL && cmd == CLUSTER_COMMAND_CLUSTER)
+ {
+ ListCell *index;
+
+ /* We need to find the index that has indisclustered set. */
+ foreach(index, RelationGetIndexList(rel))
+ {
+ indexOid = lfirst_oid(index);
+ if (get_index_isclustered(indexOid))
+ break;
+ indexOid = InvalidOid;
+ }
+
+ if (!OidIsValid(indexOid))
+ ereport(ERROR,
+ (errcode(ERRCODE_UNDEFINED_OBJECT),
+ errmsg("there is no previously clustered index for table \"%s\"",
+ relation->relname)));
+ }
+ else if (indexname != NULL)
+ {
+ /*
+ * The index is expected to be in the same namespace as the relation.
+ */
+ indexOid = get_relname_relid(indexname,
+ rel->rd_rel->relnamespace);
+ if (!OidIsValid(indexOid))
+ ereport(ERROR,
+ (errcode(ERRCODE_UNDEFINED_OBJECT),
+ errmsg("index \"%s\" for table \"%s\" does not exist",
+ indexname, relation->relname)));
+ }
+
+ *indexOid_p = indexOid;
+
+ /* For non-partitioned tables, do what we came here to do. */
+ if (rel->rd_rel->relkind != RELKIND_PARTITIONED_TABLE)
+ {
+ cluster_rel(rel, indexOid, params, cmd);
+ /* cluster_rel closes the relation, but keeps lock */
+
+ return NULL;
+ }
+
+ return rel;
+}
diff --git a/src/backend/commands/tablecmds.c b/src/backend/commands/tablecmds.c
index ce7d115667..901cb321c3 100644
--- a/src/backend/commands/tablecmds.c
+++ b/src/backend/commands/tablecmds.c
@@ -15510,7 +15510,8 @@ ATExecClusterOn(Relation rel, const char *indexName, LOCKMODE lockmode)
indexName, RelationGetRelationName(rel))));
/* Check index is valid to cluster on */
- check_index_is_clusterable(rel, indexOid, lockmode);
+ check_index_is_clusterable(rel, indexOid, lockmode,
+ CLUSTER_COMMAND_CLUSTER);
/* And do the work */
mark_index_clustered(rel, indexOid, false);
diff --git a/src/backend/commands/vacuum.c b/src/backend/commands/vacuum.c
index 0239d9bae6..59dddcd31f 100644
--- a/src/backend/commands/vacuum.c
+++ b/src/backend/commands/vacuum.c
@@ -2248,7 +2248,8 @@ vacuum_rel(Oid relid, RangeVar *relation, VacuumParams *params,
cluster_params.options |= CLUOPT_VERBOSE;
/* VACUUM FULL is now a variant of CLUSTER; see cluster.c */
- cluster_rel(rel, InvalidOid, &cluster_params);
+ cluster_rel(rel, InvalidOid, &cluster_params,
+ CLUSTER_COMMAND_VACUUM);
/* cluster_rel closes the relation, but keeps lock */
rel = NULL;
diff --git a/src/backend/parser/gram.y b/src/backend/parser/gram.y
index 7d99c9355c..8b4c226495 100644
--- a/src/backend/parser/gram.y
+++ b/src/backend/parser/gram.y
@@ -298,7 +298,7 @@ static Node *makeRecursiveViewSelect(char *relname, List *aliases, Node *query);
GrantStmt GrantRoleStmt ImportForeignSchemaStmt IndexStmt InsertStmt
ListenStmt LoadStmt LockStmt MergeStmt NotifyStmt ExplainableStmt PreparableStmt
CreateFunctionStmt AlterFunctionStmt ReindexStmt RemoveAggrStmt
- RemoveFuncStmt RemoveOperStmt RenameStmt ReturnStmt RevokeStmt RevokeRoleStmt
+ RemoveFuncStmt RemoveOperStmt RenameStmt RepackStmt ReturnStmt RevokeStmt RevokeRoleStmt
RuleActionStmt RuleActionStmtOrEmpty RuleStmt
SecLabelStmt SelectStmt TransactionStmt TransactionStmtLegacy TruncateStmt
UnlistenStmt UpdateStmt VacuumStmt
@@ -381,7 +381,7 @@ static Node *makeRecursiveViewSelect(char *relname, List *aliases, Node *query);
%type <str> copy_file_name
access_method_clause attr_name
table_access_method_clause name cursor_name file_name
- cluster_index_specification
+ cluster_index_specification repack_index_specification
%type <list> func_name handler_name qual_Op qual_all_Op subquery_Op
opt_inline_handler opt_validator validator_clause
@@ -764,7 +764,7 @@ static Node *makeRecursiveViewSelect(char *relname, List *aliases, Node *query);
QUOTE QUOTES
RANGE READ REAL REASSIGN RECURSIVE REF_P REFERENCES REFERENCING
- REFRESH REINDEX RELATIVE_P RELEASE RENAME REPEATABLE REPLACE REPLICA
+ REFRESH REINDEX RELATIVE_P RELEASE RENAME REPACK REPEATABLE REPLACE REPLICA
RESET RESTART RESTRICT RETURN RETURNING RETURNS REVOKE RIGHT ROLE ROLLBACK ROLLUP
ROUTINE ROUTINES ROW ROWS RULE
@@ -1100,6 +1100,7 @@ stmt:
| RemoveFuncStmt
| RemoveOperStmt
| RenameStmt
+ | RepackStmt
| RevokeStmt
| RevokeRoleStmt
| RuleStmt
@@ -11869,6 +11870,60 @@ cluster_index_specification:
| /*EMPTY*/ { $$ = NULL; }
;
+/*****************************************************************************
+ *
+ * QUERY:
+ * REPACK [ (options) ] [ <qualified_name> [ USING INDEX <index_name> ] ]
+ *
+ *****************************************************************************/
+
+RepackStmt:
+ REPACK qualified_name repack_index_specification
+ {
+ RepackStmt *n = makeNode(RepackStmt);
+
+ n->relation = $2;
+ n->indexname = $3;
+ n->params = NIL;
+ $$ = (Node *) n;
+ }
+
+ | REPACK '(' utility_option_list ')' qualified_name repack_index_specification
+ {
+ RepackStmt *n = makeNode(RepackStmt);
+
+ n->relation = $5;
+ n->indexname = $6;
+ n->params = $3;
+ $$ = (Node *) n;
+ }
+
+ | REPACK
+ {
+ RepackStmt *n = makeNode(RepackStmt);
+
+ n->relation = NULL;
+ n->indexname = NULL;
+ n->params = NIL;
+ $$ = (Node *) n;
+ }
+
+ | REPACK '(' utility_option_list ')'
+ {
+ RepackStmt *n = makeNode(RepackStmt);
+
+ n->relation = NULL;
+ n->indexname = NULL;
+ n->params = $3;
+ $$ = (Node *) n;
+ }
+ ;
+
+repack_index_specification:
+ USING INDEX name { $$ = $3; }
+ | /*EMPTY*/ { $$ = NULL; }
+ ;
+
/*****************************************************************************
*
@@ -17909,6 +17964,7 @@ unreserved_keyword:
| RELATIVE_P
| RELEASE
| RENAME
+ | REPACK
| REPEATABLE
| REPLACE
| REPLICA
@@ -18540,6 +18596,7 @@ bare_label_keyword:
| RELATIVE_P
| RELEASE
| RENAME
+ | REPACK
| REPEATABLE
| REPLACE
| REPLICA
diff --git a/src/backend/tcop/utility.c b/src/backend/tcop/utility.c
index 25fe3d5801..bf3ba3c2ae 100644
--- a/src/backend/tcop/utility.c
+++ b/src/backend/tcop/utility.c
@@ -280,6 +280,7 @@ ClassifyUtilityCommandAsReadOnly(Node *parsetree)
case T_ClusterStmt:
case T_ReindexStmt:
case T_VacuumStmt:
+ case T_RepackStmt:
{
/*
* These commands write WAL, so they're not strictly
@@ -862,6 +863,10 @@ standard_ProcessUtility(PlannedStmt *pstmt,
ExecVacuum(pstate, (VacuumStmt *) parsetree, isTopLevel);
break;
+ case T_RepackStmt:
+ repack(pstate, (RepackStmt *) parsetree, isTopLevel);
+ break;
+
case T_ExplainStmt:
ExplainQuery(pstate, (ExplainStmt *) parsetree, params, dest);
break;
@@ -2869,6 +2874,10 @@ CreateCommandTag(Node *parsetree)
tag = CMDTAG_ANALYZE;
break;
+ case T_RepackStmt:
+ tag = CMDTAG_REPACK;
+ break;
+
case T_ExplainStmt:
tag = CMDTAG_EXPLAIN;
break;
diff --git a/src/backend/utils/adt/pgstatfuncs.c b/src/backend/utils/adt/pgstatfuncs.c
index 0ea41299e0..02ac18fca6 100644
--- a/src/backend/utils/adt/pgstatfuncs.c
+++ b/src/backend/utils/adt/pgstatfuncs.c
@@ -268,6 +268,8 @@ pg_stat_get_progress_info(PG_FUNCTION_ARGS)
cmdtype = PROGRESS_COMMAND_ANALYZE;
else if (pg_strcasecmp(cmd, "CLUSTER") == 0)
cmdtype = PROGRESS_COMMAND_CLUSTER;
+ else if (pg_strcasecmp(cmd, "REPACK") == 0)
+ cmdtype = PROGRESS_COMMAND_REPACK;
else if (pg_strcasecmp(cmd, "CREATE INDEX") == 0)
cmdtype = PROGRESS_COMMAND_CREATE_INDEX;
else if (pg_strcasecmp(cmd, "BASEBACKUP") == 0)
diff --git a/src/bin/psql/tab-complete.in.c b/src/bin/psql/tab-complete.in.c
index 8432be641a..72338fffb2 100644
--- a/src/bin/psql/tab-complete.in.c
+++ b/src/bin/psql/tab-complete.in.c
@@ -1223,7 +1223,7 @@ static const char *const sql_commands[] = {
"DELETE FROM", "DISCARD", "DO", "DROP", "END", "EXECUTE", "EXPLAIN",
"FETCH", "GRANT", "IMPORT FOREIGN SCHEMA", "INSERT INTO", "LISTEN", "LOAD", "LOCK",
"MERGE INTO", "MOVE", "NOTIFY", "PREPARE",
- "REASSIGN", "REFRESH MATERIALIZED VIEW", "REINDEX", "RELEASE",
+ "REASSIGN", "REFRESH MATERIALIZED VIEW", "REINDEX", "RELEASE", "REPACK",
"RESET", "REVOKE", "ROLLBACK",
"SAVEPOINT", "SECURITY LABEL", "SELECT", "SET", "SHOW", "START",
"TABLE", "TRUNCATE", "UNLISTEN", "UPDATE", "VACUUM", "VALUES", "WITH",
@@ -4909,6 +4909,35 @@ match_previous_words(int pattern_id,
COMPLETE_WITH_QUERY(Query_for_list_of_tablespaces);
}
+/* REPACK */
+ else if (Matches("REPACK"))
+ COMPLETE_WITH_SCHEMA_QUERY(Query_for_list_of_clusterables);
+ else if (Matches("REPACK", "(*)"))
+ COMPLETE_WITH_SCHEMA_QUERY(Query_for_list_of_clusterables);
+ /* If we have REPACK <sth>, then add "USING INDEX" */
+ else if (Matches("REPACK", MatchAnyExcept("(")))
+ COMPLETE_WITH("USING INDEX");
+ /* If we have REPACK (*) <sth>, then add "USING INDEX" */
+ else if (Matches("REPACK", "(*)", MatchAny))
+ COMPLETE_WITH("USING INDEX");
+ /* If we have REPACK <sth> USING, then add the index as well */
+ else if (Matches("REPACK", MatchAny, "USING", "INDEX"))
+ {
+ set_completion_reference(prev3_wd);
+ COMPLETE_WITH_SCHEMA_QUERY(Query_for_index_of_table);
+ }
+ else if (HeadMatches("REPACK", "(*") &&
+ !HeadMatches("REPACK", "(*)"))
+ {
+ /*
+ * This fires if we're in an unfinished parenthesized option list.
+ * get_previous_words treats a completed parenthesized option list as
+ * one word, so the above test is correct.
+ */
+ if (ends_with(prev_wd, '(') || ends_with(prev_wd, ','))
+ COMPLETE_WITH("VERBOSE");
+ }
+
/* SECURITY LABEL */
else if (Matches("SECURITY"))
COMPLETE_WITH("LABEL");
diff --git a/src/include/commands/cluster.h b/src/include/commands/cluster.h
index 60088a64cb..c2976905e4 100644
--- a/src/include/commands/cluster.h
+++ b/src/include/commands/cluster.h
@@ -31,10 +31,27 @@ typedef struct ClusterParams
bits32 options; /* bitmask of CLUOPT_* */
} ClusterParams;
+/*
+ * cluster.c currently implements three nearly identical commands: CLUSTER,
+ * VACUUM FULL and REPACK. Where needed, use this enumeration to distinguish
+ * which of these commands is being executed.
+ *
+ * Remove this stuff when removing the (now deprecated) CLUSTER and VACUUM
+ * FULL commands.
+ */
+typedef enum ClusterCommand
+{
+ CLUSTER_COMMAND_CLUSTER,
+ CLUSTER_COMMAND_REPACK,
+ CLUSTER_COMMAND_VACUUM
+} ClusterCommand;
+
extern void cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel);
-extern void cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params);
+extern void cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
+ ClusterCommand cmd);
extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid,
- LOCKMODE lockmode);
+ LOCKMODE lockmode,
+ ClusterCommand cmd);
extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal);
extern Oid make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
@@ -48,4 +65,5 @@ extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
MultiXactId cutoffMulti,
char newrelpersistence);
+extern void repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel);
#endif /* CLUSTER_H */
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 7c736e7b03..7644267e14 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -56,24 +56,48 @@
#define PROGRESS_ANALYZE_PHASE_COMPUTE_EXT_STATS 4
#define PROGRESS_ANALYZE_PHASE_FINALIZE_ANALYZE 5
-/* Progress parameters for cluster */
-#define PROGRESS_CLUSTER_COMMAND 0
-#define PROGRESS_CLUSTER_PHASE 1
-#define PROGRESS_CLUSTER_INDEX_RELID 2
-#define PROGRESS_CLUSTER_HEAP_TUPLES_SCANNED 3
-#define PROGRESS_CLUSTER_HEAP_TUPLES_WRITTEN 4
-#define PROGRESS_CLUSTER_TOTAL_HEAP_BLKS 5
-#define PROGRESS_CLUSTER_HEAP_BLKS_SCANNED 6
-#define PROGRESS_CLUSTER_INDEX_REBUILD_COUNT 7
-
-/* Phases of cluster (as advertised via PROGRESS_CLUSTER_PHASE) */
-#define PROGRESS_CLUSTER_PHASE_SEQ_SCAN_HEAP 1
-#define PROGRESS_CLUSTER_PHASE_INDEX_SCAN_HEAP 2
-#define PROGRESS_CLUSTER_PHASE_SORT_TUPLES 3
-#define PROGRESS_CLUSTER_PHASE_WRITE_NEW_HEAP 4
-#define PROGRESS_CLUSTER_PHASE_SWAP_REL_FILES 5
-#define PROGRESS_CLUSTER_PHASE_REBUILD_INDEX 6
-#define PROGRESS_CLUSTER_PHASE_FINAL_CLEANUP 7
+/*
+ * Progress parameters for REPACK.
+ *
+ * Note: Since REPACK shares some code with CLUSTER, (some of) these values
+ * are also used by CLUSTER. (CLUSTER is now deprecated, so it makes no sense
+ * to introduce separate set of constants.)
+ */
+#define PROGRESS_REPACK_COMMAND 0
+#define PROGRESS_REPACK_PHASE 1
+#define PROGRESS_REPACK_INDEX_RELID 2
+#define PROGRESS_REPACK_HEAP_TUPLES_SCANNED 3
+#define PROGRESS_REPACK_HEAP_TUPLES_WRITTEN 4
+#define PROGRESS_REPACK_TOTAL_HEAP_BLKS 5
+#define PROGRESS_REPACK_HEAP_BLKS_SCANNED 6
+#define PROGRESS_REPACK_INDEX_REBUILD_COUNT 7
+
+/*
+ * Phases of repack (as advertised via PROGRESS_REPACK_PHASE).
+ *
+ * Note: Since REPACK shares some code with CLUSTER, (some of) these values
+ * are also used by CLUSTER. (CLUSTER is now deprecated, so it makes no sense
+ * to introduce separate set of constants.)
+ */
+#define PROGRESS_REPACK_PHASE_SEQ_SCAN_HEAP 1
+#define PROGRESS_REPACK_PHASE_INDEX_SCAN_HEAP 2
+#define PROGRESS_REPACK_PHASE_SORT_TUPLES 3
+#define PROGRESS_REPACK_PHASE_WRITE_NEW_HEAP 4
+#define PROGRESS_REPACK_PHASE_SWAP_REL_FILES 5
+#define PROGRESS_REPACK_PHASE_REBUILD_INDEX 6
+#define PROGRESS_REPACK_PHASE_FINAL_CLEANUP 7
+
+/* Commands of PROGRESS_REPACK */
+#define PROGRESS_REPACK_COMMAND_REPACK 1
+
+/*
+ * Progress parameters for cluster.
+ *
+ * Although we need to report REPACK and CLUSTER in separate views, the
+ * parameters and phases of CLUSTER are a subset of those of REPACK. Therefore
+ * we just use the appropriate values defined for REPACK above instead of
+ * defining a separate set of constants here.
+ */
/* Commands of PROGRESS_CLUSTER */
#define PROGRESS_CLUSTER_COMMAND_CLUSTER 1
diff --git a/src/include/nodes/parsenodes.h b/src/include/nodes/parsenodes.h
index 0b208f51bd..03ed0450df 100644
--- a/src/include/nodes/parsenodes.h
+++ b/src/include/nodes/parsenodes.h
@@ -3914,6 +3914,19 @@ typedef struct ClusterStmt
List *params; /* list of DefElem nodes */
} ClusterStmt;
+/* ----------------------
+ * Repack Statement
+ * ----------------------
+ */
+typedef struct RepackStmt
+{
+ NodeTag type;
+ RangeVar *relation; /* relation being repacked */
+ char *indexname; /* order tuples by this index */
+ List *params; /* list of DefElem nodes */
+} RepackStmt;
+
+
/* ----------------------
* Vacuum and Analyze Statements
*
diff --git a/src/include/parser/kwlist.h b/src/include/parser/kwlist.h
index 40cf090ce6..0932d6fce5 100644
--- a/src/include/parser/kwlist.h
+++ b/src/include/parser/kwlist.h
@@ -373,6 +373,7 @@ PG_KEYWORD("reindex", REINDEX, UNRESERVED_KEYWORD, BARE_LABEL)
PG_KEYWORD("relative", RELATIVE_P, UNRESERVED_KEYWORD, BARE_LABEL)
PG_KEYWORD("release", RELEASE, UNRESERVED_KEYWORD, BARE_LABEL)
PG_KEYWORD("rename", RENAME, UNRESERVED_KEYWORD, BARE_LABEL)
+PG_KEYWORD("repack", REPACK, UNRESERVED_KEYWORD, BARE_LABEL)
PG_KEYWORD("repeatable", REPEATABLE, UNRESERVED_KEYWORD, BARE_LABEL)
PG_KEYWORD("replace", REPLACE, UNRESERVED_KEYWORD, BARE_LABEL)
PG_KEYWORD("replica", REPLICA, UNRESERVED_KEYWORD, BARE_LABEL)
diff --git a/src/include/tcop/cmdtaglist.h b/src/include/tcop/cmdtaglist.h
index d250a714d5..cceb312f2b 100644
--- a/src/include/tcop/cmdtaglist.h
+++ b/src/include/tcop/cmdtaglist.h
@@ -196,6 +196,7 @@ PG_CMDTAG(CMDTAG_REASSIGN_OWNED, "REASSIGN OWNED", false, false, false)
PG_CMDTAG(CMDTAG_REFRESH_MATERIALIZED_VIEW, "REFRESH MATERIALIZED VIEW", true, false, false)
PG_CMDTAG(CMDTAG_REINDEX, "REINDEX", true, false, false)
PG_CMDTAG(CMDTAG_RELEASE, "RELEASE", false, false, false)
+PG_CMDTAG(CMDTAG_REPACK, "REPACK", false, false, false)
PG_CMDTAG(CMDTAG_RESET, "RESET", false, false, false)
PG_CMDTAG(CMDTAG_REVOKE, "REVOKE", true, false, false)
PG_CMDTAG(CMDTAG_REVOKE_ROLE, "REVOKE ROLE", false, false, false)
diff --git a/src/include/utils/backend_progress.h b/src/include/utils/backend_progress.h
index dda813ab40..da3d14bb97 100644
--- a/src/include/utils/backend_progress.h
+++ b/src/include/utils/backend_progress.h
@@ -25,6 +25,7 @@ typedef enum ProgressCommandType
PROGRESS_COMMAND_VACUUM,
PROGRESS_COMMAND_ANALYZE,
PROGRESS_COMMAND_CLUSTER,
+ PROGRESS_COMMAND_REPACK,
PROGRESS_COMMAND_CREATE_INDEX,
PROGRESS_COMMAND_BASEBACKUP,
PROGRESS_COMMAND_COPY,
diff --git a/src/test/regress/expected/cluster.out b/src/test/regress/expected/cluster.out
index 4d40a6809a..ed7df29b8e 100644
--- a/src/test/regress/expected/cluster.out
+++ b/src/test/regress/expected/cluster.out
@@ -254,6 +254,120 @@ ORDER BY 1;
clstr_tst_pkey
(3 rows)
+-- REPACK handles individual tables identically to CLUSTER, but it's worth
+-- checking if it handles table hierarchies identically as well.
+REPACK clstr_tst USING INDEX clstr_tst_c;
+-- Verify that inheritance link still works
+INSERT INTO clstr_tst_inh VALUES (0, 100, 'in child table 2');
+SELECT a,b,c,substring(d for 30), length(d) from clstr_tst;
+ a | b | c | substring | length
+----+-----+------------------+--------------------------------+--------
+ 10 | 14 | catorce | |
+ 18 | 5 | cinco | |
+ 9 | 4 | cuatro | |
+ 26 | 19 | diecinueve | |
+ 12 | 18 | dieciocho | |
+ 30 | 16 | dieciseis | |
+ 24 | 17 | diecisiete | |
+ 2 | 10 | diez | |
+ 23 | 12 | doce | |
+ 11 | 2 | dos | |
+ 25 | 9 | nueve | |
+ 31 | 8 | ocho | |
+ 1 | 11 | once | |
+ 28 | 15 | quince | |
+ 32 | 6 | seis | xyzzyxyzzyxyzzyxyzzyxyzzyxyzzy | 500000
+ 29 | 7 | siete | |
+ 15 | 13 | trece | |
+ 22 | 30 | treinta | |
+ 17 | 32 | treinta y dos | |
+ 3 | 31 | treinta y uno | |
+ 5 | 3 | tres | |
+ 20 | 1 | uno | |
+ 6 | 20 | veinte | |
+ 14 | 25 | veinticinco | |
+ 21 | 24 | veinticuatro | |
+ 4 | 22 | veintidos | |
+ 19 | 29 | veintinueve | |
+ 16 | 28 | veintiocho | |
+ 27 | 26 | veintiseis | |
+ 13 | 27 | veintisiete | |
+ 7 | 23 | veintitres | |
+ 8 | 21 | veintiuno | |
+ 0 | 100 | in child table | |
+ 0 | 100 | in child table 2 | |
+(34 rows)
+
+-- Verify that foreign key link still works
+INSERT INTO clstr_tst (b, c) VALUES (1111, 'this should fail');
+ERROR: insert or update on table "clstr_tst" violates foreign key constraint "clstr_tst_con"
+DETAIL: Key (b)=(1111) is not present in table "clstr_tst_s".
+SELECT conname FROM pg_constraint WHERE conrelid = 'clstr_tst'::regclass
+ORDER BY 1;
+ conname
+----------------------
+ clstr_tst_a_not_null
+ clstr_tst_con
+ clstr_tst_pkey
+(3 rows)
+
+-- Yet another code path: REPACK w/o index.
+REPACK clstr_tst USING INDEX clstr_tst_c;
+-- Verify that inheritance link still works
+INSERT INTO clstr_tst_inh VALUES (0, 100, 'in child table 3');
+SELECT a,b,c,substring(d for 30), length(d) from clstr_tst;
+ a | b | c | substring | length
+----+-----+------------------+--------------------------------+--------
+ 10 | 14 | catorce | |
+ 18 | 5 | cinco | |
+ 9 | 4 | cuatro | |
+ 26 | 19 | diecinueve | |
+ 12 | 18 | dieciocho | |
+ 30 | 16 | dieciseis | |
+ 24 | 17 | diecisiete | |
+ 2 | 10 | diez | |
+ 23 | 12 | doce | |
+ 11 | 2 | dos | |
+ 25 | 9 | nueve | |
+ 31 | 8 | ocho | |
+ 1 | 11 | once | |
+ 28 | 15 | quince | |
+ 32 | 6 | seis | xyzzyxyzzyxyzzyxyzzyxyzzyxyzzy | 500000
+ 29 | 7 | siete | |
+ 15 | 13 | trece | |
+ 22 | 30 | treinta | |
+ 17 | 32 | treinta y dos | |
+ 3 | 31 | treinta y uno | |
+ 5 | 3 | tres | |
+ 20 | 1 | uno | |
+ 6 | 20 | veinte | |
+ 14 | 25 | veinticinco | |
+ 21 | 24 | veinticuatro | |
+ 4 | 22 | veintidos | |
+ 19 | 29 | veintinueve | |
+ 16 | 28 | veintiocho | |
+ 27 | 26 | veintiseis | |
+ 13 | 27 | veintisiete | |
+ 7 | 23 | veintitres | |
+ 8 | 21 | veintiuno | |
+ 0 | 100 | in child table | |
+ 0 | 100 | in child table 2 | |
+ 0 | 100 | in child table 3 | |
+(35 rows)
+
+-- Verify that foreign key link still works
+INSERT INTO clstr_tst (b, c) VALUES (1111, 'this should fail');
+ERROR: insert or update on table "clstr_tst" violates foreign key constraint "clstr_tst_con"
+DETAIL: Key (b)=(1111) is not present in table "clstr_tst_s".
+SELECT conname FROM pg_constraint WHERE conrelid = 'clstr_tst'::regclass
+ORDER BY 1;
+ conname
+----------------------
+ clstr_tst_a_not_null
+ clstr_tst_con
+ clstr_tst_pkey
+(3 rows)
+
SELECT relname, relkind,
EXISTS(SELECT 1 FROM pg_class WHERE oid = c.reltoastrelid) AS hastoast
FROM pg_class c WHERE relname LIKE 'clstr_tst%' ORDER BY relname;
@@ -381,6 +495,35 @@ SELECT * FROM clstr_1;
2
(2 rows)
+-- REPACK w/o argument performs no ordering, so we can only check which tables
+-- have the relfilenode changed.
+RESET SESSION AUTHORIZATION;
+CREATE TEMP TABLE relnodes_old AS
+(SELECT relname, relfilenode
+FROM pg_class
+WHERE relname IN ('clstr_1', 'clstr_2', 'clstr_3'));
+SET SESSION AUTHORIZATION regress_clstr_user;
+SET client_min_messages = ERROR; -- order of "skipping" warnings may vary
+REPACK;
+RESET client_min_messages;
+RESET SESSION AUTHORIZATION;
+CREATE TEMP TABLE relnodes_new AS
+(SELECT relname, relfilenode
+FROM pg_class
+WHERE relname IN ('clstr_1', 'clstr_2', 'clstr_3'));
+-- Do the actual comparison. Unlike CLUSTER, clstr_3 should have been
+-- processed because there is nothing like clustering index here.
+SELECT o.relname FROM relnodes_old o
+JOIN relnodes_new n ON o.relname = n.relname
+WHERE o.relfilenode <> n.relfilenode
+ORDER BY o.relname;
+ relname
+---------
+ clstr_1
+ clstr_3
+(2 rows)
+
+SET SESSION AUTHORIZATION regress_clstr_user;
-- Test MVCC-safety of cluster. There isn't much we can do to verify the
-- results with a single backend...
CREATE TABLE clustertest (key int PRIMARY KEY);
@@ -495,6 +638,43 @@ ALTER TABLE clstrpart SET WITHOUT CLUSTER;
ERROR: cannot mark index clustered in partitioned table
ALTER TABLE clstrpart CLUSTER ON clstrpart_idx;
ERROR: cannot mark index clustered in partitioned table
+-- Check that REPACK sets new relfilenodes: it should process exactly the same
+-- tables as CLUSTER did.
+DROP TABLE old_cluster_info;
+DROP TABLE new_cluster_info;
+CREATE TEMP TABLE old_cluster_info AS SELECT relname, level, relfilenode, relkind FROM pg_partition_tree('clstrpart'::regclass) AS tree JOIN pg_class c ON c.oid=tree.relid ;
+REPACK clstrpart USING INDEX clstrpart_idx;
+CREATE TEMP TABLE new_cluster_info AS SELECT relname, level, relfilenode, relkind FROM pg_partition_tree('clstrpart'::regclass) AS tree JOIN pg_class c ON c.oid=tree.relid ;
+SELECT relname, old.level, old.relkind, old.relfilenode = new.relfilenode FROM old_cluster_info AS old JOIN new_cluster_info AS new USING (relname) ORDER BY relname COLLATE "C";
+ relname | level | relkind | ?column?
+-------------+-------+---------+----------
+ clstrpart | 0 | p | t
+ clstrpart1 | 1 | p | t
+ clstrpart11 | 2 | r | f
+ clstrpart12 | 2 | p | t
+ clstrpart2 | 1 | r | f
+ clstrpart3 | 1 | p | t
+ clstrpart33 | 2 | r | f
+(7 rows)
+
+-- And finally the same for REPACK w/o index.
+DROP TABLE old_cluster_info;
+DROP TABLE new_cluster_info;
+CREATE TEMP TABLE old_cluster_info AS SELECT relname, level, relfilenode, relkind FROM pg_partition_tree('clstrpart'::regclass) AS tree JOIN pg_class c ON c.oid=tree.relid ;
+REPACK clstrpart;
+CREATE TEMP TABLE new_cluster_info AS SELECT relname, level, relfilenode, relkind FROM pg_partition_tree('clstrpart'::regclass) AS tree JOIN pg_class c ON c.oid=tree.relid ;
+SELECT relname, old.level, old.relkind, old.relfilenode = new.relfilenode FROM old_cluster_info AS old JOIN new_cluster_info AS new USING (relname) ORDER BY relname COLLATE "C";
+ relname | level | relkind | ?column?
+-------------+-------+---------+----------
+ clstrpart | 0 | p | t
+ clstrpart1 | 1 | p | t
+ clstrpart11 | 2 | r | f
+ clstrpart12 | 2 | p | t
+ clstrpart2 | 1 | r | f
+ clstrpart3 | 1 | p | t
+ clstrpart33 | 2 | r | f
+(7 rows)
+
DROP TABLE clstrpart;
-- Ownership of partitions is checked
CREATE TABLE ptnowner(i int unique) PARTITION BY LIST (i);
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 62f69ac20b..50d87af2fd 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2041,6 +2041,33 @@ pg_stat_progress_create_index| SELECT s.pid,
s.param15 AS partitions_done
FROM (pg_stat_get_progress_info('CREATE INDEX'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
+pg_stat_progress_repack| SELECT s.pid,
+ s.datid,
+ d.datname,
+ s.relid,
+ CASE s.param1
+ WHEN 1 THEN 'REPACK'::text
+ ELSE NULL::text
+ END AS command,
+ CASE s.param2
+ WHEN 0 THEN 'initializing'::text
+ WHEN 1 THEN 'seq scanning heap'::text
+ WHEN 2 THEN 'index scanning heap'::text
+ WHEN 3 THEN 'sorting tuples'::text
+ WHEN 4 THEN 'writing new heap'::text
+ WHEN 5 THEN 'swapping relation files'::text
+ WHEN 6 THEN 'rebuilding index'::text
+ WHEN 7 THEN 'performing final cleanup'::text
+ ELSE NULL::text
+ END AS phase,
+ (s.param3)::oid AS repack_index_relid,
+ s.param4 AS heap_tuples_scanned,
+ s.param5 AS heap_tuples_written,
+ s.param6 AS heap_blks_total,
+ s.param7 AS heap_blks_scanned,
+ s.param8 AS index_rebuild_count
+ FROM (pg_stat_get_progress_info('REPACK'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
+ LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_progress_vacuum| SELECT s.pid,
s.datid,
d.datname,
diff --git a/src/test/regress/sql/cluster.sql b/src/test/regress/sql/cluster.sql
index b7115f8610..e348e26fbf 100644
--- a/src/test/regress/sql/cluster.sql
+++ b/src/test/regress/sql/cluster.sql
@@ -76,6 +76,33 @@ INSERT INTO clstr_tst (b, c) VALUES (1111, 'this should fail');
SELECT conname FROM pg_constraint WHERE conrelid = 'clstr_tst'::regclass
ORDER BY 1;
+-- REPACK handles individual tables identically to CLUSTER, but it's worth
+-- checking if it handles table hierarchies identically as well.
+REPACK clstr_tst USING INDEX clstr_tst_c;
+
+-- Verify that inheritance link still works
+INSERT INTO clstr_tst_inh VALUES (0, 100, 'in child table 2');
+SELECT a,b,c,substring(d for 30), length(d) from clstr_tst;
+
+-- Verify that foreign key link still works
+INSERT INTO clstr_tst (b, c) VALUES (1111, 'this should fail');
+
+SELECT conname FROM pg_constraint WHERE conrelid = 'clstr_tst'::regclass
+ORDER BY 1;
+
+-- Yet another code path: REPACK w/o index.
+REPACK clstr_tst USING INDEX clstr_tst_c;
+
+-- Verify that inheritance link still works
+INSERT INTO clstr_tst_inh VALUES (0, 100, 'in child table 3');
+SELECT a,b,c,substring(d for 30), length(d) from clstr_tst;
+
+-- Verify that foreign key link still works
+INSERT INTO clstr_tst (b, c) VALUES (1111, 'this should fail');
+
+SELECT conname FROM pg_constraint WHERE conrelid = 'clstr_tst'::regclass
+ORDER BY 1;
+
SELECT relname, relkind,
EXISTS(SELECT 1 FROM pg_class WHERE oid = c.reltoastrelid) AS hastoast
@@ -159,6 +186,34 @@ INSERT INTO clstr_1 VALUES (1);
CLUSTER clstr_1;
SELECT * FROM clstr_1;
+-- REPACK w/o argument performs no ordering, so we can only check which tables
+-- have the relfilenode changed.
+RESET SESSION AUTHORIZATION;
+CREATE TEMP TABLE relnodes_old AS
+(SELECT relname, relfilenode
+FROM pg_class
+WHERE relname IN ('clstr_1', 'clstr_2', 'clstr_3'));
+
+SET SESSION AUTHORIZATION regress_clstr_user;
+SET client_min_messages = ERROR; -- order of "skipping" warnings may vary
+REPACK;
+RESET client_min_messages;
+
+RESET SESSION AUTHORIZATION;
+CREATE TEMP TABLE relnodes_new AS
+(SELECT relname, relfilenode
+FROM pg_class
+WHERE relname IN ('clstr_1', 'clstr_2', 'clstr_3'));
+
+-- Do the actual comparison. Unlike CLUSTER, clstr_3 should have been
+-- processed because there is nothing like clustering index here.
+SELECT o.relname FROM relnodes_old o
+JOIN relnodes_new n ON o.relname = n.relname
+WHERE o.relfilenode <> n.relfilenode
+ORDER BY o.relname;
+
+SET SESSION AUTHORIZATION regress_clstr_user;
+
-- Test MVCC-safety of cluster. There isn't much we can do to verify the
-- results with a single backend...
@@ -229,6 +284,24 @@ SELECT relname, old.level, old.relkind, old.relfilenode = new.relfilenode FROM o
CLUSTER clstrpart;
ALTER TABLE clstrpart SET WITHOUT CLUSTER;
ALTER TABLE clstrpart CLUSTER ON clstrpart_idx;
+
+-- Check that REPACK sets new relfilenodes: it should process exactly the same
+-- tables as CLUSTER did.
+DROP TABLE old_cluster_info;
+DROP TABLE new_cluster_info;
+CREATE TEMP TABLE old_cluster_info AS SELECT relname, level, relfilenode, relkind FROM pg_partition_tree('clstrpart'::regclass) AS tree JOIN pg_class c ON c.oid=tree.relid ;
+REPACK clstrpart USING INDEX clstrpart_idx;
+CREATE TEMP TABLE new_cluster_info AS SELECT relname, level, relfilenode, relkind FROM pg_partition_tree('clstrpart'::regclass) AS tree JOIN pg_class c ON c.oid=tree.relid ;
+SELECT relname, old.level, old.relkind, old.relfilenode = new.relfilenode FROM old_cluster_info AS old JOIN new_cluster_info AS new USING (relname) ORDER BY relname COLLATE "C";
+
+-- And finally the same for REPACK w/o index.
+DROP TABLE old_cluster_info;
+DROP TABLE new_cluster_info;
+CREATE TEMP TABLE old_cluster_info AS SELECT relname, level, relfilenode, relkind FROM pg_partition_tree('clstrpart'::regclass) AS tree JOIN pg_class c ON c.oid=tree.relid ;
+REPACK clstrpart;
+CREATE TEMP TABLE new_cluster_info AS SELECT relname, level, relfilenode, relkind FROM pg_partition_tree('clstrpart'::regclass) AS tree JOIN pg_class c ON c.oid=tree.relid ;
+SELECT relname, old.level, old.relkind, old.relfilenode = new.relfilenode FROM old_cluster_info AS old JOIN new_cluster_info AS new USING (relname) ORDER BY relname COLLATE "C";
+
DROP TABLE clstrpart;
-- Ownership of partitions is checked
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index cfbab589d6..098a7a602c 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -412,6 +412,7 @@ ClientCertName
ClientConnectionInfo
ClientData
ClientSocket
+ClusterCommand
ClonePtrType
ClosePortalStmt
ClosePtrType
@@ -2464,6 +2465,7 @@ ReorderBufferTupleCidKey
ReorderBufferUpdateProgressTxnCB
ReorderTuple
RepOriginId
+RepackStmt
ReparameterizeForeignPathByChild_function
ReplaceVarsFromTargetList_context
ReplaceVarsNoMatchOption
--
2.43.5
[text/x-diff] v08-0002-Move-progress-related-fields-from-PgBackendStatus-to.patch (7.8K, ../127361.1740559688@localhost/3-v08-0002-Move-progress-related-fields-from-PgBackendStatus-to.patch)
download | inline diff:
From 7921185eb112e7c1839981e6d5309493c697a1c1 Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 26 Feb 2025 09:17:20 +0100
Subject: [PATCH 2/9] Move progress related fields from PgBackendStatus to
PgBackendProgress.
REPACK CONCURRENTLY will need to save and restore these fields at some
point. This is because plan_cluster_use_sort() has to be called in a
subtransaction (so that it does not leave any additional locks on the table)
and rollback of that subtransaction clears the progress information.
---
src/backend/access/heap/vacuumlazy.c | 2 +-
src/backend/commands/analyze.c | 2 +-
src/backend/utils/activity/backend_progress.c | 18 +++++++++---------
src/backend/utils/activity/backend_status.c | 4 ++--
src/backend/utils/adt/pgstatfuncs.c | 6 +++---
src/include/utils/backend_progress.h | 14 ++++++++++++++
src/include/utils/backend_status.h | 14 ++------------
7 files changed, 32 insertions(+), 28 deletions(-)
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 1af18a78a2..5a0eb1ca80 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -1100,7 +1100,7 @@ heap_vacuum_rel(Relation rel, VacuumParams *params,
* the st_progress_param array.
*/
appendStringInfo(&buf, _("delay time: %.3f ms\n"),
- (double) MyBEEntry->st_progress_param[PROGRESS_VACUUM_DELAY_TIME] / 1000000.0);
+ (double) MyBEEntry->st_progress.param[PROGRESS_VACUUM_DELAY_TIME] / 1000000.0);
}
if (track_io_timing)
{
diff --git a/src/backend/commands/analyze.c b/src/backend/commands/analyze.c
index cd75954951..49fe8c43ce 100644
--- a/src/backend/commands/analyze.c
+++ b/src/backend/commands/analyze.c
@@ -815,7 +815,7 @@ do_analyze_rel(Relation onerel, VacuumParams *params,
* only updated by the calling process.
*/
appendStringInfo(&buf, _("delay time: %.3f ms\n"),
- (double) MyBEEntry->st_progress_param[PROGRESS_ANALYZE_DELAY_TIME] / 1000000.0);
+ (double) MyBEEntry->st_progress.param[PROGRESS_ANALYZE_DELAY_TIME] / 1000000.0);
}
if (track_io_timing)
{
diff --git a/src/backend/utils/activity/backend_progress.c b/src/backend/utils/activity/backend_progress.c
index 99a8c73bf0..eebc968193 100644
--- a/src/backend/utils/activity/backend_progress.c
+++ b/src/backend/utils/activity/backend_progress.c
@@ -32,9 +32,9 @@ pgstat_progress_start_command(ProgressCommandType cmdtype, Oid relid)
return;
PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
- beentry->st_progress_command = cmdtype;
- beentry->st_progress_command_target = relid;
- MemSet(&beentry->st_progress_param, 0, sizeof(beentry->st_progress_param));
+ beentry->st_progress.command = cmdtype;
+ beentry->st_progress.command_target = relid;
+ MemSet(&beentry->st_progress.param, 0, sizeof(beentry->st_progress.param));
PGSTAT_END_WRITE_ACTIVITY(beentry);
}
@@ -55,7 +55,7 @@ pgstat_progress_update_param(int index, int64 val)
return;
PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
- beentry->st_progress_param[index] = val;
+ beentry->st_progress.param[index] = val;
PGSTAT_END_WRITE_ACTIVITY(beentry);
}
@@ -76,7 +76,7 @@ pgstat_progress_incr_param(int index, int64 incr)
return;
PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
- beentry->st_progress_param[index] += incr;
+ beentry->st_progress.param[index] += incr;
PGSTAT_END_WRITE_ACTIVITY(beentry);
}
@@ -133,7 +133,7 @@ pgstat_progress_update_multi_param(int nparam, const int *index,
{
Assert(index[i] >= 0 && index[i] < PGSTAT_NUM_PROGRESS_PARAM);
- beentry->st_progress_param[index[i]] = val[i];
+ beentry->st_progress.param[index[i]] = val[i];
}
PGSTAT_END_WRITE_ACTIVITY(beentry);
@@ -154,11 +154,11 @@ pgstat_progress_end_command(void)
if (!beentry || !pgstat_track_activities)
return;
- if (beentry->st_progress_command == PROGRESS_COMMAND_INVALID)
+ if (beentry->st_progress.command == PROGRESS_COMMAND_INVALID)
return;
PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
- beentry->st_progress_command = PROGRESS_COMMAND_INVALID;
- beentry->st_progress_command_target = InvalidOid;
+ beentry->st_progress.command = PROGRESS_COMMAND_INVALID;
+ beentry->st_progress.command_target = InvalidOid;
PGSTAT_END_WRITE_ACTIVITY(beentry);
}
diff --git a/src/backend/utils/activity/backend_status.c b/src/backend/utils/activity/backend_status.c
index 5f68ef26ad..647bf863f0 100644
--- a/src/backend/utils/activity/backend_status.c
+++ b/src/backend/utils/activity/backend_status.c
@@ -376,8 +376,8 @@ pgstat_bestart(void)
#endif
lbeentry.st_state = STATE_UNDEFINED;
- lbeentry.st_progress_command = PROGRESS_COMMAND_INVALID;
- lbeentry.st_progress_command_target = InvalidOid;
+ lbeentry.st_progress.command = PROGRESS_COMMAND_INVALID;
+ lbeentry.st_progress.command_target = InvalidOid;
lbeentry.st_query_id = UINT64CONST(0);
/*
diff --git a/src/backend/utils/adt/pgstatfuncs.c b/src/backend/utils/adt/pgstatfuncs.c
index 02ac18fca6..6603b2a64b 100644
--- a/src/backend/utils/adt/pgstatfuncs.c
+++ b/src/backend/utils/adt/pgstatfuncs.c
@@ -299,7 +299,7 @@ pg_stat_get_progress_info(PG_FUNCTION_ARGS)
* Report values for only those backends which are running the given
* command.
*/
- if (beentry->st_progress_command != cmdtype)
+ if (beentry->st_progress.command != cmdtype)
continue;
/* Value available to all callers */
@@ -309,9 +309,9 @@ pg_stat_get_progress_info(PG_FUNCTION_ARGS)
/* show rest of the values including relid only to role members */
if (HAS_PGSTAT_PERMISSIONS(beentry->st_userid))
{
- values[2] = ObjectIdGetDatum(beentry->st_progress_command_target);
+ values[2] = ObjectIdGetDatum(beentry->st_progress.command_target);
for (i = 0; i < PGSTAT_NUM_PROGRESS_PARAM; i++)
- values[i + 3] = Int64GetDatum(beentry->st_progress_param[i]);
+ values[i + 3] = Int64GetDatum(beentry->st_progress.param[i]);
}
else
{
diff --git a/src/include/utils/backend_progress.h b/src/include/utils/backend_progress.h
index da3d14bb97..2f1de46d05 100644
--- a/src/include/utils/backend_progress.h
+++ b/src/include/utils/backend_progress.h
@@ -31,8 +31,22 @@ typedef enum ProgressCommandType
PROGRESS_COMMAND_COPY,
} ProgressCommandType;
+
#define PGSTAT_NUM_PROGRESS_PARAM 20
+/*
+ * Any command which wishes can advertise that it is running by setting
+ * command, command_target, and param[]. command_target should be the OID of
+ * the relation which the command targets (we assume there's just one, as this
+ * is meant for utility commands), but the meaning of each element in the
+ * param array is command-specific.
+ */
+typedef struct PgBackendProgress
+{
+ ProgressCommandType command;
+ Oid command_target;
+ int64 param[PGSTAT_NUM_PROGRESS_PARAM];
+} PgBackendProgress;
extern void pgstat_progress_start_command(ProgressCommandType cmdtype,
Oid relid);
diff --git a/src/include/utils/backend_status.h b/src/include/utils/backend_status.h
index d3d4ff6c5c..a73c76a442 100644
--- a/src/include/utils/backend_status.h
+++ b/src/include/utils/backend_status.h
@@ -155,18 +155,8 @@ typedef struct PgBackendStatus
*/
char *st_activity_raw;
- /*
- * Command progress reporting. Any command which wishes can advertise
- * that it is running by setting st_progress_command,
- * st_progress_command_target, and st_progress_param[].
- * st_progress_command_target should be the OID of the relation which the
- * command targets (we assume there's just one, as this is meant for
- * utility commands), but the meaning of each element in the
- * st_progress_param array is command-specific.
- */
- ProgressCommandType st_progress_command;
- Oid st_progress_command_target;
- int64 st_progress_param[PGSTAT_NUM_PROGRESS_PARAM];
+ /* Command progress reporting. */
+ PgBackendProgress st_progress;
/* query identifier, optionally computed using post_parse_analyze_hook */
uint64 st_query_id;
--
2.43.5
[text/x-diff] v08-0003-Move-conversion-of-a-historic-to-MVCC-snapshot-to-a-.patch (5.4K, ../127361.1740559688@localhost/4-v08-0003-Move-conversion-of-a-historic-to-MVCC-snapshot-to-a-.patch)
download | inline diff:
From ced00311269d331e5b91985e08b2330aa8069dbd Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 26 Feb 2025 09:17:20 +0100
Subject: [PATCH 3/9] Move conversion of a "historic" to MVCC snapshot to a
separate function.
The conversion is now handled by SnapBuildMVCCFromHistoric(). REPACK
CONCURRENTLY will also need it.
---
src/backend/replication/logical/snapbuild.c | 51 +++++++++++++++++----
src/backend/utils/time/snapmgr.c | 3 +-
src/include/replication/snapbuild.h | 1 +
src/include/utils/snapmgr.h | 1 +
4 files changed, 45 insertions(+), 11 deletions(-)
diff --git a/src/backend/replication/logical/snapbuild.c b/src/backend/replication/logical/snapbuild.c
index bd0680dcbe..8c83ff6feb 100644
--- a/src/backend/replication/logical/snapbuild.c
+++ b/src/backend/replication/logical/snapbuild.c
@@ -440,10 +440,7 @@ Snapshot
SnapBuildInitialSnapshot(SnapBuild *builder)
{
Snapshot snap;
- TransactionId xid;
TransactionId safeXid;
- TransactionId *newxip;
- int newxcnt = 0;
Assert(XactIsoLevel == XACT_REPEATABLE_READ);
Assert(builder->building_full_snapshot);
@@ -485,6 +482,31 @@ SnapBuildInitialSnapshot(SnapBuild *builder)
MyProc->xmin = snap->xmin;
+ /* Convert the historic snapshot to MVCC snapshot. */
+ return SnapBuildMVCCFromHistoric(snap, true);
+}
+
+/*
+ * Turn a historic MVCC snapshot into an ordinary MVCC snapshot.
+ *
+ * Unlike a regular (non-historic) MVCC snapshot, the xip array of this
+ * snapshot contains not only running main transactions, but also their
+ * subtransactions. This difference does has no impact on XidInMVCCSnapshot().
+ *
+ * Pass true for 'in_place' if you don't care about modifying the source
+ * snapshot. If you need a new instance, and one that was allocated as a
+ * single chunk of memory, pass false.
+ */
+Snapshot
+SnapBuildMVCCFromHistoric(Snapshot snapshot, bool in_place)
+{
+ TransactionId xid;
+ TransactionId *oldxip = snapshot->xip;
+ uint32 oldxcnt = snapshot->xcnt;
+ TransactionId *newxip;
+ int newxcnt = 0;
+ Snapshot result;
+
/* allocate in transaction context */
newxip = (TransactionId *)
palloc(sizeof(TransactionId) * GetMaxSnapshotXidCount());
@@ -495,7 +517,7 @@ SnapBuildInitialSnapshot(SnapBuild *builder)
* classical snapshot by marking all non-committed transactions as
* in-progress. This can be expensive.
*/
- for (xid = snap->xmin; NormalTransactionIdPrecedes(xid, snap->xmax);)
+ for (xid = snapshot->xmin; NormalTransactionIdPrecedes(xid, snapshot->xmax);)
{
void *test;
@@ -503,7 +525,7 @@ SnapBuildInitialSnapshot(SnapBuild *builder)
* Check whether transaction committed using the decoding snapshot
* meaning of ->xip.
*/
- test = bsearch(&xid, snap->xip, snap->xcnt,
+ test = bsearch(&xid, snapshot->xip, snapshot->xcnt,
sizeof(TransactionId), xidComparator);
if (test == NULL)
@@ -520,11 +542,22 @@ SnapBuildInitialSnapshot(SnapBuild *builder)
}
/* adjust remaining snapshot fields as needed */
- snap->snapshot_type = SNAPSHOT_MVCC;
- snap->xcnt = newxcnt;
- snap->xip = newxip;
+ snapshot->xcnt = newxcnt;
+ snapshot->xip = newxip;
+
+ if (in_place)
+ result = snapshot;
+ else
+ {
+ result = CopySnapshot(snapshot);
+
+ /* Restore the original values so the source is intact. */
+ snapshot->xip = oldxip;
+ snapshot->xcnt = oldxcnt;
+ }
+ result->snapshot_type = SNAPSHOT_MVCC;
- return snap;
+ return result;
}
/*
diff --git a/src/backend/utils/time/snapmgr.c b/src/backend/utils/time/snapmgr.c
index 8f1508b1ee..42bded373b 100644
--- a/src/backend/utils/time/snapmgr.c
+++ b/src/backend/utils/time/snapmgr.c
@@ -153,7 +153,6 @@ typedef struct ExportedSnapshot
static List *exportedSnapshots = NIL;
/* Prototypes for local functions */
-static Snapshot CopySnapshot(Snapshot snapshot);
static void UnregisterSnapshotNoOwner(Snapshot snapshot);
static void FreeSnapshot(Snapshot snapshot);
static void SnapshotResetXmin(void);
@@ -532,7 +531,7 @@ SetTransactionSnapshot(Snapshot sourcesnap, VirtualTransactionId *sourcevxid,
* The copy is palloc'd in TopTransactionContext and has initial refcounts set
* to 0. The returned snapshot has the copied flag set.
*/
-static Snapshot
+Snapshot
CopySnapshot(Snapshot snapshot)
{
Snapshot newsnap;
diff --git a/src/include/replication/snapbuild.h b/src/include/replication/snapbuild.h
index 44031dcf6e..6d4d2d1814 100644
--- a/src/include/replication/snapbuild.h
+++ b/src/include/replication/snapbuild.h
@@ -73,6 +73,7 @@ extern void FreeSnapshotBuilder(SnapBuild *builder);
extern void SnapBuildSnapDecRefcount(Snapshot snap);
extern Snapshot SnapBuildInitialSnapshot(SnapBuild *builder);
+extern Snapshot SnapBuildMVCCFromHistoric(Snapshot snapshot, bool in_place);
extern const char *SnapBuildExportSnapshot(SnapBuild *builder);
extern void SnapBuildClearExportedSnapshot(void);
extern void SnapBuildResetExportedSnapshotState(void);
diff --git a/src/include/utils/snapmgr.h b/src/include/utils/snapmgr.h
index d346be7164..147b190210 100644
--- a/src/include/utils/snapmgr.h
+++ b/src/include/utils/snapmgr.h
@@ -60,6 +60,7 @@ extern Snapshot GetTransactionSnapshot(void);
extern Snapshot GetLatestSnapshot(void);
extern void SnapshotSetCommandId(CommandId curcid);
+extern Snapshot CopySnapshot(Snapshot snapshot);
extern Snapshot GetCatalogSnapshot(Oid relid);
extern Snapshot GetNonHistoricCatalogSnapshot(Oid relid);
extern void InvalidateCatalogSnapshot(void);
--
2.43.5
[text/plain] v08-0004-Add-CONCURRENTLY-option-to-REPACK-command.patch (168.4K, ../127361.1740559688@localhost/5-v08-0004-Add-CONCURRENTLY-option-to-REPACK-command.patch)
download | inline diff:
From 8da24b040720e5a30517076825760c7516c42bfa Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 26 Feb 2025 09:17:20 +0100
Subject: [PATCH 4/9] Add CONCURRENTLY option to REPACK command.
The REPACK command copies the relation data into a new file, creates new
indexes and eventually swaps the files. To make sure that the old file does
not change during the copying, the relation is locked in an exclusive mode,
which prevents applications from both reading and writing. (To keep the data
consistent, we'd only need to prevent the applications from writing, but even
reading needs to be blocked before we can swap the files - otherwise some
applications could continue using the old file. Since we cannot get stronger
lock without releasing the weaker one first, we acquire the exclusive lock in
the beginning and keep it till the end of the processing.)
This patch introduces an alternative workflow, which only requires the
exclusive lock when the relation (and index) files are being swapped.
(Supposedly, the swapping should be pretty fast.) On the other hand, when we
copy the data to the new file, we allow applications to read from the relation
and even write into it.
First, we scan the relation using a "historic snapshot", and insert all the
tuples satisfying this snapshot into the new file. Note that, before creating
that snapshot, we need to make sure that all the other backends treat the
relation as a system catalog: in particular, they must log information on new
command IDs (CIDs). We achieve that by adding the relation ID into a shared
hash table and waiting until all the transactions currently writing into the
table (i.e. transactions possibly not aware of the new entry) have finished.
Second, logical decoding is used to capture the data changes done by
applications during the copying (i.e. changes that do not satisfy the historic
snapshot mentioned above), and those are applied to the new file before we
acquire the exclusive lock we need to swap the files. (Of course, more data
changes can take place while we are waiting for the lock - these will be
applied to the new file after we have acquired the lock, before we swap the
files.)
While copying the data into the new file, we hold a lock that prevents
applications from changing the relation tuple descriptor (tuples inserted into
the old file must fit into the new file). However, as we have to release that
lock before getting the exclusive one, it's possible that someone adds or
drops a column, or changes the data type of an existing one. Therefore we have
to check the tuple descriptor before we swap the files. If we find out that
the tuple descriptor changed, ERROR is raised and all the changes are rolled
back. Since a lot of effort can be wasted in such a case, the ALTER TABLE
command also tries to check if REPACK CONCURRENTLY is running on the same
relation, and raises an ERROR if it is.
Like the existing implementation of REPACK, the variant with the CONCURRENTLY
option also requires an extra space for the new relation and index files
(which coexist with the old files for some time). In addition, the
CONCURRENTLY option might introduce a lag in releasing WAL segments for
archiving / recycling. This is due to the decoding of the data changes done by
application concurrently. However, this lag should not be more than a single
WAL segment.
---
doc/src/sgml/monitoring.sgml | 65 +-
doc/src/sgml/ref/repack.sgml | 116 +-
src/Makefile | 1 +
src/backend/access/heap/heapam.c | 8 +-
src/backend/access/heap/heapam_handler.c | 145 +-
src/backend/access/heap/heapam_visibility.c | 30 +-
src/backend/catalog/index.c | 43 +-
src/backend/catalog/system_views.sql | 30 +-
src/backend/commands/cluster.c | 2667 ++++++++++++++++-
src/backend/commands/matview.c | 2 +-
src/backend/commands/tablecmds.c | 11 +
src/backend/commands/vacuum.c | 12 +-
src/backend/meson.build | 1 +
src/backend/parser/gram.y | 17 +-
src/backend/replication/logical/decode.c | 24 +
src/backend/replication/logical/snapbuild.c | 20 +
.../replication/pgoutput_repack/Makefile | 32 +
.../replication/pgoutput_repack/meson.build | 18 +
.../pgoutput_repack/pgoutput_repack.c | 286 ++
src/backend/storage/ipc/ipci.c | 3 +
src/backend/tcop/utility.c | 10 +
src/backend/utils/activity/backend_progress.c | 16 +
.../utils/activity/wait_event_names.txt | 1 +
src/backend/utils/cache/inval.c | 21 +
src/backend/utils/cache/relcache.c | 5 +
src/backend/utils/time/snapmgr.c | 3 +-
src/bin/psql/tab-complete.in.c | 24 +-
src/include/access/heapam.h | 4 +
src/include/access/tableam.h | 10 +
src/include/catalog/index.h | 3 +
src/include/commands/cluster.h | 93 +-
src/include/commands/progress.h | 17 +-
src/include/nodes/parsenodes.h | 1 +
src/include/replication/snapbuild.h | 1 +
src/include/storage/lockdefs.h | 5 +-
src/include/storage/lwlocklist.h | 1 +
src/include/utils/backend_progress.h | 3 +-
src/include/utils/inval.h | 2 +
src/include/utils/rel.h | 7 +-
src/include/utils/snapmgr.h | 2 +
src/test/regress/expected/rules.out | 29 +-
41 files changed, 3536 insertions(+), 253 deletions(-)
create mode 100644 src/backend/replication/pgoutput_repack/Makefile
create mode 100644 src/backend/replication/pgoutput_repack/meson.build
create mode 100644 src/backend/replication/pgoutput_repack/pgoutput_repack.c
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 58e1becf02..8d73c01c55 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -5780,14 +5780,35 @@ FROM pg_stat_get_backend_idset() AS backendid;
<row>
<entry role="catalog_table_entry"><para role="column_definition">
- <structfield>heap_tuples_written</structfield> <type>bigint</type>
+ <structfield>heap_tuples_inserted</structfield> <type>bigint</type>
</para>
<para>
- Number of heap tuples written.
+ Number of heap tuples inserted.
This counter only advances when the phase is
<literal>seq scanning heap</literal>,
- <literal>index scanning heap</literal>
- or <literal>writing new heap</literal>.
+ <literal>index scanning heap</literal>,
+ <literal>writing new heap</literal>
+ or <literal>catch-up</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_tuples_updated</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap tuples updated.
+ This counter only advances when the phase is <literal>catch-up</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_tuples_deleted</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap tuples deleted.
+ This counter only advances when the phase is <literal>catch-up</literal>.
</para></entry>
</row>
@@ -6003,14 +6024,35 @@ FROM pg_stat_get_backend_idset() AS backendid;
<row>
<entry role="catalog_table_entry"><para role="column_definition">
- <structfield>heap_tuples_written</structfield> <type>bigint</type>
+ <structfield>heap_tuples_inserted</structfield> <type>bigint</type>
</para>
<para>
- Number of heap tuples written.
+ Number of heap tuples inserted.
This counter only advances when the phase is
<literal>seq scanning heap</literal>,
- <literal>index scanning heap</literal>
- or <literal>writing new heap</literal>.
+ <literal>index scanning heap</literal>,
+ <literal>writing new heap</literal>
+ or <literal>catch-up</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_tuples_updated</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap tuples updated.
+ This counter only advances when the phase is <literal>catch-up</literal>.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>heap_tuples_deleted</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of heap tuples deleted.
+ This counter only advances when the phase is <literal>catch-up</literal>.
</para></entry>
</row>
@@ -6091,6 +6133,13 @@ FROM pg_stat_get_backend_idset() AS backendid;
<command>REPACK</command> is currently writing the new heap.
</entry>
</row>
+ <row>
+ <entry><literal>catch-up</literal></entry>
+ <entry>
+ <command>REPACK</command> is currently processing the DML commands that
+ other transactions executed during any of the preceding phase.
+ </entry>
+ </row>
<row>
<entry><literal>swapping relation files</literal></entry>
<entry>
diff --git a/doc/src/sgml/ref/repack.sgml b/doc/src/sgml/ref/repack.sgml
index 84f3c3e3f2..9ee640e351 100644
--- a/doc/src/sgml/ref/repack.sgml
+++ b/doc/src/sgml/ref/repack.sgml
@@ -22,6 +22,7 @@ PostgreSQL documentation
<refsynopsisdiv>
<synopsis>
REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <replaceable class="parameter">table_name</replaceable> [ USING INDEX<replaceable class="parameter">index_name</replaceable> ] ]
+REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] CONCURRENTLY <replaceable class="parameter">table_name</replaceable> [ USING INDEX<replaceable class="parameter">index_name</replaceable> ]
<phrase>where <replaceable class="parameter">option</replaceable> can be one of:</phrase>
@@ -48,7 +49,8 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <re
processes every table and materialized view in the current database that
the current user has the <literal>MAINTAIN</literal> privilege on. This
form of <command>REPACK</command> cannot be executed inside a transaction
- block.
+ block. Also, this form is not allowed if
+ the <literal>CONCURRENTLY</literal> option is used.
</para>
<para>
@@ -61,7 +63,8 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <re
When a table is being repacked, an <literal>ACCESS EXCLUSIVE</literal> lock
is acquired on it. This prevents any other database operations (both reads
and writes) from operating on the table until the <command>REPACK</command>
- is finished.
+ is finished. If you want to keep the table accessible during the repacking,
+ consider using the <literal>CONCURRENTLY</literal> option.
</para>
<refsect2 id="sql-repack-notes-on-clustering" xreflabel="Notes on Clustering">
@@ -160,6 +163,115 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] [ <re
</listitem>
</varlistentry>
+ <varlistentry>
+ <term><literal>CONCURRENTLY</literal></term>
+ <listitem>
+ <para>
+ Allow other transactions to use the table while it is being repacked.
+ </para>
+
+ <para>
+ Internally, <command>REPACK</command> copies the contents of the table
+ (ignoring dead tuples) into a new file, sorted by the specified index,
+ and also creates a new file for each index. Then it swaps the old and
+ new files for the table and all the indexes, and deletes the old
+ files. The <literal>ACCESS EXCLUSIVE</literal> lock is needed to make
+ sure that the old files do not change during the processing because the
+ changes would get lost due to the swap.
+ </para>
+
+ <para>
+ With the <literal>CONCURRENTLY</literal> option, the <literal>ACCESS
+ EXCLUSIVE</literal> lock is only acquired to swap the table and index
+ files. The data changes that took place during the creation of the new
+ table and index files are captured using logical decoding
+ (<xref linkend="logicaldecoding"/>) and applied before
+ the <literal>ACCESS EXCLUSIVE</literal> lock is requested. Thus the lock
+ is typically held only for the time needed to swap the files, which
+ should be pretty short.
+ </para>
+
+ <para>
+ Note that <command>REPACK</command> with the
+ the <literal>CONCURRENTLY</literal> option does not try to order the
+ rows inserted into the table after the repacking started. Also
+ note <command>REPACK</command> might fail to complete due to DDL
+ commands executed on the table by other transactions during the
+ repacking.
+ </para>
+
+ <note>
+ <para>
+ In addition to the temporary space requirements explained in
+ <xref linkend="sql-repack-notes-on-resources"/>,
+ the <literal>CONCURRENTLY</literal> option can add to the usage of
+ temporary space a bit more. The reason is that other transactions can
+ perform DML operations which cannot be applied to the new file until
+ <command>REPACK</command> has copied all the tuples from the old
+ file. Thus the tuples inserted into the old file during the copying are
+ also stored in separately in a temporary file, so they can eventually
+ be applied to the new file.
+ </para>
+
+ <para>
+ Furthermore, the data changes performed during the copying are
+ extracted from <link linkend="wal">write-ahead log</link> (WAL), and
+ this extraction (decoding) only takes place when certain amount of WAL
+ has been written. Therefore, WAL removal can be delayed by this
+ threshold. Currently the threshold is equal to the value of
+ the <link linkend="guc-wal-segment-size"><varname>wal_segment_size</varname></link>
+ configuration parameter.
+ </para>
+ </note>
+
+ <para>
+ The <literal>CONCURRENTLY</literal> option cannot be used in the
+ following cases:
+
+ <itemizedlist>
+ <listitem>
+ <para>
+ The table is <literal>UNLOGGED</literal>.
+ </para>
+ </listitem>
+
+ <listitem>
+ <para>
+ The table is partitioned.
+ </para>
+ </listitem>
+
+ <listitem>
+ <para>
+ The table is a system catalog or a <acronym>TOAST</acronym> table.
+ </para>
+ </listitem>
+
+ <listitem>
+ <para>
+ <command>REPACK</command> is executed inside a transaction block.
+ </para>
+ </listitem>
+
+ <listitem>
+ <para>
+ The <link linkend="guc-wal-level"><varname>wal_level</varname></link>
+ configuration parameter is less than <literal>logical</literal>.
+ </para>
+ </listitem>
+
+ <listitem>
+ <para>
+ The <link linkend="guc-max-replication-slots"><varname>max_replication_slots</varname></link>
+ configuration parameter does not allow for creation of an additional
+ replication slot.
+ </para>
+ </listitem>
+ </itemizedlist>
+ </para>
+ </listitem>
+ </varlistentry>
+
<varlistentry>
<term><literal>VERBOSE</literal></term>
<listitem>
diff --git a/src/Makefile b/src/Makefile
index 2f31a2f20a..b18c9a14ff 100644
--- a/src/Makefile
+++ b/src/Makefile
@@ -23,6 +23,7 @@ SUBDIRS = \
interfaces \
backend/replication/libpqwalreceiver \
backend/replication/pgoutput \
+ backend/replication/pgoutput_repack \
fe_utils \
bin \
pl \
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index fa7935a0ed..cb856a74ee 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -2093,8 +2093,14 @@ heap_insert(Relation relation, HeapTuple tup, CommandId cid,
/*
* If this is a catalog, we need to transmit combo CIDs to properly
* decode, so log that as well.
+ *
+ * For the main heap (as opposed to TOAST), we only receive
+ * HEAP_INSERT_NO_LOGICAL when doing REPACK CONCURRENTLY, in which
+ * case the visibility information does not change. Therefore, there's
+ * no need to update the decoding snapshot.
*/
- if (RelationIsAccessibleInLogicalDecoding(relation))
+ if ((options & HEAP_INSERT_NO_LOGICAL) == 0 &&
+ RelationIsAccessibleInLogicalDecoding(relation))
log_heap_new_cid(relation, heaptup);
/*
diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index 5c3cab8bc2..b2bfd05dc9 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -33,6 +33,7 @@
#include "catalog/index.h"
#include "catalog/storage.h"
#include "catalog/storage_xlog.h"
+#include "commands/cluster.h"
#include "commands/progress.h"
#include "executor/executor.h"
#include "miscadmin.h"
@@ -53,6 +54,9 @@ static void reform_and_rewrite_tuple(HeapTuple tuple,
static bool SampleHeapTupleVisible(TableScanDesc scan, Buffer buffer,
HeapTuple tuple,
OffsetNumber tupoffset);
+static HeapTuple accept_tuple_for_concurrent_copy(HeapTuple tuple,
+ Snapshot snapshot,
+ Buffer buffer);
static BlockNumber heapam_scan_get_blocks_done(HeapScanDesc hscan);
@@ -681,6 +685,8 @@ static void
heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
Relation OldIndex, bool use_sort,
TransactionId OldestXmin,
+ Snapshot snapshot,
+ LogicalDecodingContext *decoding_ctx,
TransactionId *xid_cutoff,
MultiXactId *multi_cutoff,
double *num_tuples,
@@ -701,6 +707,8 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
bool *isnull;
BufferHeapTupleTableSlot *hslot;
BlockNumber prev_cblock = InvalidBlockNumber;
+ bool concurrent = snapshot != NULL;
+ XLogRecPtr end_of_wal_prev = GetFlushRecPtr(NULL);
/* Remember if it's a system catalog */
is_system_catalog = IsSystemRelation(OldHeap);
@@ -779,8 +787,10 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
for (;;)
{
HeapTuple tuple;
+ bool tuple_copied = false;
Buffer buf;
bool isdead;
+ HTSV_Result vis;
CHECK_FOR_INTERRUPTS();
@@ -835,7 +845,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
LockBuffer(buf, BUFFER_LOCK_SHARE);
- switch (HeapTupleSatisfiesVacuum(tuple, OldestXmin, buf))
+ switch ((vis = HeapTupleSatisfiesVacuum(tuple, OldestXmin, buf)))
{
case HEAPTUPLE_DEAD:
/* Definitely dead */
@@ -851,14 +861,15 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
case HEAPTUPLE_INSERT_IN_PROGRESS:
/*
- * Since we hold exclusive lock on the relation, normally the
- * only way to see this is if it was inserted earlier in our
- * own transaction. However, it can happen in system
+ * As long as we hold exclusive lock on the relation, normally
+ * the only way to see this is if it was inserted earlier in
+ * our own transaction. However, it can happen in system
* catalogs, since we tend to release write lock before commit
- * there. Give a warning if neither case applies; but in any
- * case we had better copy it.
+ * there. Also, there's no exclusive lock during concurrent
+ * processing. Give a warning if neither case applies; but in
+ * any case we had better copy it.
*/
- if (!is_system_catalog &&
+ if (!is_system_catalog && !concurrent &&
!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetXmin(tuple->t_data)))
elog(WARNING, "concurrent insert in progress within table \"%s\"",
RelationGetRelationName(OldHeap));
@@ -870,7 +881,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
/*
* Similar situation to INSERT_IN_PROGRESS case.
*/
- if (!is_system_catalog &&
+ if (!is_system_catalog && !concurrent &&
!TransactionIdIsCurrentTransactionId(HeapTupleHeaderGetUpdateXid(tuple->t_data)))
elog(WARNING, "concurrent delete in progress within table \"%s\"",
RelationGetRelationName(OldHeap));
@@ -884,8 +895,6 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
break;
}
- LockBuffer(buf, BUFFER_LOCK_UNLOCK);
-
if (isdead)
{
*tups_vacuumed += 1;
@@ -896,9 +905,47 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
*tups_vacuumed += 1;
*tups_recently_dead -= 1;
}
+
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
continue;
}
+ if (concurrent)
+ {
+ /*
+ * Ignore concurrent changes now, they'll be processed later via
+ * logical decoding.
+ *
+ * INSERT_IN_PROGRESS is rejected right away because our snapshot
+ * should represent a point in time which should precede (or be
+ * equal to) the state of transactions as it was when the
+ * "SatisfiesVacuum" test was performed. Thus
+ * accept_tuple_for_concurrent_copy() should not consider the
+ * tuple inserted.
+ */
+ if (vis == HEAPTUPLE_INSERT_IN_PROGRESS)
+ tuple = NULL;
+ else
+ tuple = accept_tuple_for_concurrent_copy(tuple, snapshot,
+ buf);
+ /* Tuple not suitable for the new heap? */
+ if (tuple == NULL)
+ {
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+ continue;
+ }
+
+ /* Remember that we have to free the tuple eventually. */
+ tuple_copied = true;
+ }
+
+ /*
+ * In the concurrent case, we have a copy of the tuple, so we don't
+ * worry whether the source tuple will be deleted / updated after we
+ * release the lock.
+ */
+ LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+
*num_tuples += 1;
if (tuplesort != NULL)
{
@@ -915,7 +962,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
{
const int ct_index[] = {
PROGRESS_REPACK_HEAP_TUPLES_SCANNED,
- PROGRESS_REPACK_HEAP_TUPLES_WRITTEN
+ PROGRESS_REPACK_HEAP_TUPLES_INSERTED
};
int64 ct_val[2];
@@ -930,6 +977,33 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
ct_val[1] = *num_tuples;
pgstat_progress_update_multi_param(2, ct_index, ct_val);
}
+ if (tuple_copied)
+ heap_freetuple(tuple);
+
+ /*
+ * Process the WAL produced by the load, as well as by other
+ * transactions, so that the replication slot can advance and WAL does
+ * not pile up. Use wal_segment_size as a threshold so that we do not
+ * introduce the decoding overhead too often.
+ *
+ * Of course, we must not apply the changes until the initial load has
+ * completed.
+ *
+ * Note that our insertions into the new table should not be decoded
+ * as we (intentionally) do not write the logical decoding specific
+ * information to WAL.
+ */
+ if (concurrent)
+ {
+ XLogRecPtr end_of_wal;
+
+ end_of_wal = GetFlushRecPtr(NULL);
+ if ((end_of_wal - end_of_wal_prev) > wal_segment_size)
+ {
+ repack_decode_concurrent_changes(decoding_ctx, end_of_wal);
+ end_of_wal_prev = end_of_wal;
+ }
+ }
}
if (indexScan != NULL)
@@ -973,7 +1047,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
values, isnull,
rwstate);
/* Report n_tuples */
- pgstat_progress_update_param(PROGRESS_REPACK_HEAP_TUPLES_WRITTEN,
+ pgstat_progress_update_param(PROGRESS_REPACK_HEAP_TUPLES_INSERTED,
n_tuples);
}
@@ -2626,6 +2700,53 @@ SampleHeapTupleVisible(TableScanDesc scan, Buffer buffer,
}
}
+/*
+ * Return copy of 'tuple' if it has been inserted according to 'snapshot', or
+ * NULL if the insertion took place in the future. If the tuple is already
+ * marked as deleted or updated by a transaction that 'snapshot' still
+ * considers running, clear the deletion / update XID in the header of the
+ * copied tuple. This way the returned tuple is suitable for insertion into
+ * the new heap.
+ */
+static HeapTuple
+accept_tuple_for_concurrent_copy(HeapTuple tuple, Snapshot snapshot,
+ Buffer buffer)
+{
+ HeapTuple result;
+
+ Assert(snapshot->snapshot_type == SNAPSHOT_MVCC);
+
+ /*
+ * First, check if the tuple insertion is visible by our snapshot.
+ */
+ if (!HeapTupleMVCCInserted(tuple, snapshot, buffer))
+ return NULL;
+
+ result = heap_copytuple(tuple);
+
+ /*
+ * If the tuple was deleted / updated but our snapshot still sees it, we
+ * need to keep it. In that case, clear the information that indicates the
+ * deletion / update. Otherwise the tuple chain would stay incomplete (as
+ * we will reject the new tuple above), and the delete / update would fail
+ * if executed later during logical decoding.
+ */
+ if (TransactionIdIsNormal(HeapTupleHeaderGetRawXmax(result->t_data)) &&
+ HeapTupleMVCCNotDeleted(result, snapshot, buffer))
+ {
+ /* TODO More work needed here?*/
+ result->t_data->t_infomask |= HEAP_XMAX_INVALID;
+ HeapTupleHeaderSetXmax(result->t_data, 0);
+ }
+
+ /*
+ * Accept the tuple even if our snapshot considers it deleted - older
+ * snapshots can still see the tuple, while the decoded transactions
+ * should not try to update / delete it again.
+ */
+ return result;
+}
+
/* ------------------------------------------------------------------------
* Definition of the heap table access method.
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index e146605bd5..d9be93aadc 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -955,16 +955,31 @@ HeapTupleSatisfiesDirty(HeapTuple htup, Snapshot snapshot,
* did TransactionIdIsInProgress in each call --- to no avail, as long as the
* inserting/deleting transaction was still running --- which was more cycles
* and more contention on ProcArrayLock.
+ *
+ * The checks are split into two functions, HeapTupleMVCCInserted() and
+ * HeapTupleMVCCNotDeleted(), because they are also useful separately.
*/
static bool
HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
Buffer buffer)
{
- HeapTupleHeader tuple = htup->t_data;
-
Assert(ItemPointerIsValid(&htup->t_self));
Assert(htup->t_tableOid != InvalidOid);
+ return HeapTupleMVCCInserted(htup, snapshot, buffer) &&
+ HeapTupleMVCCNotDeleted(htup, snapshot, buffer);
+}
+
+/*
+ * HeapTupleMVCCInserted
+ * True iff heap tuple was successfully inserted for the given MVCC
+ * snapshot.
+ */
+bool
+HeapTupleMVCCInserted(HeapTuple htup, Snapshot snapshot, Buffer buffer)
+{
+ HeapTupleHeader tuple = htup->t_data;
+
if (!HeapTupleHeaderXminCommitted(tuple))
{
if (HeapTupleHeaderXminInvalid(tuple))
@@ -1073,6 +1088,17 @@ HeapTupleSatisfiesMVCC(HeapTuple htup, Snapshot snapshot,
}
/* by here, the inserting transaction has committed */
+ return true;
+}
+
+/*
+ * HeapTupleMVCCNotDeleted
+ * True iff heap tuple was not deleted for the given MVCC snapshot.
+ */
+bool
+HeapTupleMVCCNotDeleted(HeapTuple htup, Snapshot snapshot, Buffer buffer)
+{
+ HeapTupleHeader tuple = htup->t_data;
if (tuple->t_infomask & HEAP_XMAX_INVALID) /* xid invalid or aborted */
return true;
diff --git a/src/backend/catalog/index.c b/src/backend/catalog/index.c
index c84f67059a..39b121c0b8 100644
--- a/src/backend/catalog/index.c
+++ b/src/backend/catalog/index.c
@@ -1417,22 +1417,7 @@ index_concurrently_create_copy(Relation heapRelation, Oid oldIndexId,
for (int i = 0; i < newInfo->ii_NumIndexAttrs; i++)
opclassOptions[i] = get_attoptions(oldIndexId, i + 1);
- /* Extract statistic targets for each attribute */
- stattargets = palloc0_array(NullableDatum, newInfo->ii_NumIndexAttrs);
- for (int i = 0; i < newInfo->ii_NumIndexAttrs; i++)
- {
- HeapTuple tp;
- Datum dat;
-
- tp = SearchSysCache2(ATTNUM, ObjectIdGetDatum(oldIndexId), Int16GetDatum(i + 1));
- if (!HeapTupleIsValid(tp))
- elog(ERROR, "cache lookup failed for attribute %d of relation %u",
- i + 1, oldIndexId);
- dat = SysCacheGetAttr(ATTNUM, tp, Anum_pg_attribute_attstattarget, &isnull);
- ReleaseSysCache(tp);
- stattargets[i].value = dat;
- stattargets[i].isnull = isnull;
- }
+ stattargets = get_index_stattargets(oldIndexId, newInfo);
/*
* Now create the new index.
@@ -1471,6 +1456,32 @@ index_concurrently_create_copy(Relation heapRelation, Oid oldIndexId,
return newIndexId;
}
+NullableDatum *
+get_index_stattargets(Oid indexid, IndexInfo *indInfo)
+{
+ NullableDatum *stattargets;
+
+ /* Extract statistic targets for each attribute */
+ stattargets = palloc0_array(NullableDatum, indInfo->ii_NumIndexAttrs);
+ for (int i = 0; i < indInfo->ii_NumIndexAttrs; i++)
+ {
+ HeapTuple tp;
+ Datum dat;
+ bool isnull;
+
+ tp = SearchSysCache2(ATTNUM, ObjectIdGetDatum(indexid), Int16GetDatum(i + 1));
+ if (!HeapTupleIsValid(tp))
+ elog(ERROR, "cache lookup failed for attribute %d of relation %u",
+ i + 1, indexid);
+ dat = SysCacheGetAttr(ATTNUM, tp, Anum_pg_attribute_attstattarget, &isnull);
+ ReleaseSysCache(tp);
+ stattargets[i].value = dat;
+ stattargets[i].isnull = isnull;
+ }
+
+ return stattargets;
+}
+
/*
* index_concurrently_build
*
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index b8209b2acd..c301d83d9b 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1249,16 +1249,17 @@ CREATE VIEW pg_stat_progress_cluster AS
WHEN 2 THEN 'index scanning heap'
WHEN 3 THEN 'sorting tuples'
WHEN 4 THEN 'writing new heap'
- WHEN 5 THEN 'swapping relation files'
- WHEN 6 THEN 'rebuilding index'
- WHEN 7 THEN 'performing final cleanup'
+ -- 5 is 'catch-up', but that should not appear here.
+ WHEN 6 THEN 'swapping relation files'
+ WHEN 7 THEN 'rebuilding index'
+ WHEN 8 THEN 'performing final cleanup'
END AS phase,
CAST(S.param3 AS oid) AS cluster_index_relid,
S.param4 AS heap_tuples_scanned,
S.param5 AS heap_tuples_written,
- S.param6 AS heap_blks_total,
- S.param7 AS heap_blks_scanned,
- S.param8 AS index_rebuild_count
+ S.param8 AS heap_blks_total,
+ S.param9 AS heap_blks_scanned,
+ S.param10 AS index_rebuild_count
FROM pg_stat_get_progress_info('CLUSTER') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
@@ -1275,16 +1276,19 @@ CREATE VIEW pg_stat_progress_repack AS
WHEN 2 THEN 'index scanning heap'
WHEN 3 THEN 'sorting tuples'
WHEN 4 THEN 'writing new heap'
- WHEN 5 THEN 'swapping relation files'
- WHEN 6 THEN 'rebuilding index'
- WHEN 7 THEN 'performing final cleanup'
+ WHEN 5 THEN 'catch-up'
+ WHEN 6 THEN 'swapping relation files'
+ WHEN 7 THEN 'rebuilding index'
+ WHEN 8 THEN 'performing final cleanup'
END AS phase,
CAST(S.param3 AS oid) AS repack_index_relid,
S.param4 AS heap_tuples_scanned,
- S.param5 AS heap_tuples_written,
- S.param6 AS heap_blks_total,
- S.param7 AS heap_blks_scanned,
- S.param8 AS index_rebuild_count
+ S.param5 AS heap_tuples_inserted,
+ S.param6 AS heap_tuples_updated,
+ S.param7 AS heap_tuples_deleted,
+ S.param8 AS heap_blks_total,
+ S.param9 AS heap_blks_scanned,
+ S.param10 AS index_rebuild_count
FROM pg_stat_get_progress_info('REPACK') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
diff --git a/src/backend/commands/cluster.c b/src/backend/commands/cluster.c
index d0f2588a97..592ff6041b 100644
--- a/src/backend/commands/cluster.c
+++ b/src/backend/commands/cluster.c
@@ -25,6 +25,10 @@
#include "access/toast_internals.h"
#include "access/transam.h"
#include "access/xact.h"
+#include "access/xlog.h"
+#include "access/xlog_internal.h"
+#include "access/xloginsert.h"
+#include "access/xlogutils.h"
#include "catalog/catalog.h"
#include "catalog/dependency.h"
#include "catalog/heap.h"
@@ -32,6 +36,7 @@
#include "catalog/namespace.h"
#include "catalog/objectaccess.h"
#include "catalog/pg_am.h"
+#include "catalog/pg_control.h"
#include "catalog/pg_inherits.h"
#include "catalog/toasting.h"
#include "commands/cluster.h"
@@ -39,10 +44,15 @@
#include "commands/progress.h"
#include "commands/tablecmds.h"
#include "commands/vacuum.h"
+#include "executor/executor.h"
#include "miscadmin.h"
#include "optimizer/optimizer.h"
#include "pgstat.h"
+#include "replication/decode.h"
+#include "replication/logical.h"
+#include "replication/snapbuild.h"
#include "storage/bufmgr.h"
+#include "storage/ipc.h"
#include "storage/lmgr.h"
#include "storage/predicate.h"
#include "utils/acl.h"
@@ -76,14 +86,96 @@ typedef struct
((cmd) == CLUSTER_COMMAND_REPACK ? \
"repack" : "vacuum"))
+/*
+ * The following definitions are used for concurrent processing.
+ */
+
+/*
+ * OID of the table being repacked by this backend.
+ */
+static Oid repacked_rel = InvalidOid;
+/* The same for its TOAST relation. */
+static Oid repacked_rel_toast = InvalidOid;
+
+/*
+ * The locators are used to avoid logical decoding of data that we do not need
+ * for our table.
+ */
+RelFileLocator repacked_rel_locator = {.relNumber = InvalidOid};
+RelFileLocator repacked_rel_toast_locator = {.relNumber = InvalidOid};
+
+#define REPACK_CONCURRENT_IN_PROGRESS_MSG \
+ "relation \"%s\" is already being processed by REPACK CONCURRENTLY"
+
+/*
+ * Everything we need to call ExecInsertIndexTuples().
+ */
+typedef struct IndexInsertState
+{
+ ResultRelInfo *rri;
+ EState *estate;
+ ExprContext *econtext;
+
+ Relation ident_index;
+} IndexInsertState;
+
+/*
+ * Catalog information to check if another backend changed the relation in
+ * such a way that makes CLUSTE CONCURRENTLY unable to continue. Such changes
+ * are possible because cluster_rel() has to release its lock on the relation
+ * in order to acquire AccessExclusiveLock that it needs to swap the relation
+ * files.
+ *
+ * The most obvious problem is that the tuple descriptor has changed, since
+ * then the tuples we try to insert into the new storage are not guaranteed to
+ * fit into the storage.
+ *
+ * Another problem is relfilenode changed by another backend. It's not
+ * necessarily a correctness issue (e.g. when the other backend ran
+ * cluster_rel()), but it's safer for us to terminate the table processing in
+ * such cases. However, this information is also needs to be checked during
+ * logical decoding, so we store it in global variables repacked_rel_locator
+ * and repacked_rel_toast_locator above.
+ *
+ * Where possible, commands which might change the relation in an incompatible
+ * way should check if REPACK CONCURRENTLY is running, before they start to do
+ * the actual changes (see is_concurrent_repack_in_progress()). Anything else
+ * must be caught by check_catalog_changes(), which uses this structure.
+ */
+typedef struct CatalogState
+{
+ /* Tuple descriptor of the relation. */
+ TupleDesc tupdesc;
+
+ /* The number of indexes tracked. */
+ int ninds;
+ /* The index OIDs. */
+ Oid *ind_oids;
+ /* The index tuple descriptors. */
+ TupleDesc *ind_tupdescs;
+
+ /* The following are copies of the corresponding fields of pg_class. */
+ char relpersistence;
+ char replident;
+
+ /* rd_replidindex */
+ Oid replidindex;
+} CatalogState;
+
+/* The WAL segment being decoded. */
+static XLogSegNo repack_current_segment = 0;
+
static void cluster_multiple_rels(List *rtcs, ClusterParams *params,
- ClusterCommand cmd);
+ ClusterCommand cmd, LOCKMODE lockmode,
+ bool isTopLevel);
static void rebuild_relation(Relation OldHeap, Relation index, bool verbose,
- ClusterCommand cmd);
+ ClusterCommand cmd, bool concurrent);
static void copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
+ Snapshot snapshot, LogicalDecodingContext *decoding_ctx,
bool verbose, ClusterCommand cmd,
bool *pSwapToastByContent,
- TransactionId *pFreezeXid, MultiXactId *pCutoffMulti);
+ TransactionId *pFreezeXid,
+ MultiXactId *pCutoffMulti);
static List *get_tables_to_cluster(MemoryContext cluster_context);
static List *get_tables_to_repack(MemoryContext repack_context);
static List *get_tables_to_cluster_partitioned(MemoryContext cluster_context,
@@ -91,8 +183,91 @@ static List *get_tables_to_cluster_partitioned(MemoryContext cluster_context,
ClusterCommand cmd);
static bool cluster_is_permitted_for_relation(Oid relid, Oid userid,
ClusterCommand cmd);
+static void begin_concurrent_repack(Relation *rel_p, Relation *index_p,
+ bool *entered_p);
+static void end_concurrent_repack(bool error);
+static void cluster_before_shmem_exit_callback(int code, Datum arg);
+static CatalogState *get_catalog_state(Relation rel);
+static void free_catalog_state(CatalogState *state);
+static void check_catalog_changes(Relation rel, CatalogState *cat_state);
+static LogicalDecodingContext *setup_logical_decoding(Oid relid,
+ const char *slotname,
+ TupleDesc tupdesc);
+static HeapTuple get_changed_tuple(char *change);
+static void apply_concurrent_changes(RepackDecodingState *dstate,
+ Relation rel, ScanKey key, int nkeys,
+ IndexInsertState *iistate);
+static void apply_concurrent_insert(Relation rel, ConcurrentChange *change,
+ HeapTuple tup, IndexInsertState *iistate,
+ TupleTableSlot *index_slot);
+static void apply_concurrent_update(Relation rel, HeapTuple tup,
+ HeapTuple tup_target,
+ ConcurrentChange *change,
+ IndexInsertState *iistate,
+ TupleTableSlot *index_slot);
+static void apply_concurrent_delete(Relation rel, HeapTuple tup_target,
+ ConcurrentChange *change);
+static HeapTuple find_target_tuple(Relation rel, ScanKey key, int nkeys,
+ HeapTuple tup_key,
+ IndexInsertState *iistate,
+ TupleTableSlot *ident_slot,
+ IndexScanDesc *scan_p);
+static void process_concurrent_changes(LogicalDecodingContext *ctx,
+ XLogRecPtr end_of_wal,
+ Relation rel_dst,
+ Relation rel_src,
+ ScanKey ident_key,
+ int ident_key_nentries,
+ IndexInsertState *iistate);
+static IndexInsertState *get_index_insert_state(Relation relation,
+ Oid ident_index_id);
+static ScanKey build_identity_key(Oid ident_idx_oid, Relation rel_src,
+ int *nentries);
+static void free_index_insert_state(IndexInsertState *iistate);
+static void cleanup_logical_decoding(LogicalDecodingContext *ctx);
+static void rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
+ Relation cl_index,
+ CatalogState *cat_state,
+ LogicalDecodingContext *ctx,
+ bool swap_toast_by_content,
+ TransactionId frozenXid,
+ MultiXactId cutoffMulti);
+static List *build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes);
+
+/*
+ * Use this API when relation needs to be unlocked, closed and re-opened. If
+ * the relation got dropped while being unlocked, raise ERROR that mentions
+ * the relation name rather than OID.
+ */
+typedef struct RelReopenInfo
+{
+ /*
+ * The relation to be closed. Pointer to the value is stored here so that
+ * the user gets his reference updated automatically on re-opening.
+ *
+ * When calling unlock_and_close_relations(), 'relid' can be passed
+ * instead of 'rel_p' when the caller only needs to gather information for
+ * subsequent opening.
+ */
+ Relation *rel_p;
+ Oid relid;
+
+ char relkind;
+ LOCKMODE lockmode_orig; /* The existing lock mode */
+ LOCKMODE lockmode_new; /* The lock mode after the relation is
+ * re-opened */
+
+ char *relname; /* Relation name, initialized automatically. */
+} RelReopenInfo;
+
+static void init_rel_reopen_info(RelReopenInfo *rri, Relation *rel_p,
+ Oid relid, LOCKMODE lockmode_orig,
+ LOCKMODE lockmode_new);
+static void unlock_and_close_relations(RelReopenInfo *rels, int nrel);
+static void reopen_relations(RelReopenInfo *rels, int nrel);
static Relation process_single_relation(RangeVar *relation, char *indexname,
- ClusterCommand cmd,
+ ClusterCommand cmd, LOCKMODE lockmode,
+ bool isTopLevel,
ClusterParams *params,
Oid *indexOid_p);
@@ -151,8 +326,9 @@ cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel)
if (stmt->relation != NULL)
{
rel = process_single_relation(stmt->relation, stmt->indexname,
- CLUSTER_COMMAND_CLUSTER, ¶ms,
- &indexOid);
+ CLUSTER_COMMAND_CLUSTER,
+ AccessExclusiveLock, isTopLevel,
+ ¶ms, &indexOid);
if (rel == NULL)
return;
}
@@ -202,7 +378,8 @@ cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel)
}
/* Do the job. */
- cluster_multiple_rels(rtcs, ¶ms, CLUSTER_COMMAND_CLUSTER);
+ cluster_multiple_rels(rtcs, ¶ms, CLUSTER_COMMAND_CLUSTER,
+ AccessExclusiveLock, isTopLevel);
/* Start a new transaction for the cleanup work. */
StartTransactionCommand();
@@ -219,8 +396,8 @@ cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel)
* return.
*/
static void
-cluster_multiple_rels(List *rtcs, ClusterParams *params,
- ClusterCommand cmd)
+cluster_multiple_rels(List *rtcs, ClusterParams *params, ClusterCommand cmd,
+ LOCKMODE lockmode, bool isTopLevel)
{
ListCell *lc;
@@ -240,10 +417,10 @@ cluster_multiple_rels(List *rtcs, ClusterParams *params,
/* functions in indexes may want a snapshot set */
PushActiveSnapshot(GetTransactionSnapshot());
- rel = table_open(rtc->tableOid, AccessExclusiveLock);
+ rel = table_open(rtc->tableOid, lockmode);
/* Process this table */
- cluster_rel(rel, rtc->indexOid, params, cmd);
+ cluster_rel(rel, rtc->indexOid, params, cmd, isTopLevel);
/* cluster_rel closes the relation, but keeps lock */
PopActiveSnapshot();
@@ -267,12 +444,18 @@ cluster_multiple_rels(List *rtcs, ClusterParams *params,
* instead of index order. This is the new implementation of VACUUM FULL,
* and error messages should refer to the operation as VACUUM not CLUSTER.
*
+ * Note that, in the concurrent case, the function releases the lock at some
+ * point, in order to get AccessExclusiveLock for the final steps (i.e. to
+ * swap the relation files). To make things simpler, the caller should expect
+ * OldHeap to be closed on return, regardless CLUOPT_CONCURRENT. (The
+ * AccessExclusiveLock is kept till the end of the transaction.)
+ *
* 'cmd' indicates which commands is being executed. REPACK should be the only
* caller of this function in the future.
*/
void
cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
- ClusterCommand cmd)
+ ClusterCommand cmd, bool isTopLevel)
{
Oid tableOid = RelationGetRelid(OldHeap);
Oid save_userid;
@@ -282,8 +465,53 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
bool recheck = ((params->options & CLUOPT_RECHECK) != 0);
Relation index;
const char *cmd_str = CLUSTER_COMMAND_STR(cmd);
+ bool concurrent = ((params->options & CLUOPT_CONCURRENT) != 0);
+ LOCKMODE lmode;
+ bool entered, success;
+
+ /*
+ * Check that the correct lock is held. The lock mode is
+ * AccessExclusiveLock for normal processing and ShareUpdateExclusiveLock
+ * for concurrent processing (so that SELECT, INSERT, UPDATE and DELETE
+ * commands work, but cluster_rel() cannot be called concurrently for the
+ * same relation).
+ */
+ lmode = !concurrent ? AccessExclusiveLock : ShareUpdateExclusiveLock;
+
+ /*
+ * Skip the relation if it's being processed concurrently. In such a case,
+ * we cannot rely on a lock because the other backend needs to release it
+ * temporarily at some point.
+ *
+ * This check should not take place until we have a lock that prevents
+ * another backend from starting VREPACK CONCURRENTLY after our check.
+ */
+ Assert(CheckRelationLockedByMe(OldHeap, lmode, false));
+ if (is_concurrent_repack_in_progress(tableOid))
+ {
+ ereport(NOTICE,
+ (errmsg(REPACK_CONCURRENT_IN_PROGRESS_MSG,
+ RelationGetRelationName(OldHeap))));
+ table_close(OldHeap, lmode);
+ return;
+ }
+
+ /* There are specific requirements on concurrent processing. */
+ if (concurrent)
+ {
+ /*
+ * Make sure we have no XID assigned, otherwise call of
+ * setup_logical_decoding() can cause a deadlock.
+ *
+ * The existence of transaction block actually does not imply that XID
+ * was already assigned, but it very likely is. We might want to check
+ * the result of GetCurrentTransactionIdIfAny() instead, but that
+ * would be less clear from user's perspective.
+ */
+ PreventInTransactionBlock(isTopLevel, "REPACK CONCURRENTLY");
- Assert(CheckRelationLockedByMe(OldHeap, AccessExclusiveLock, false));
+ can_repack_concurrently(OldHeap);
+ }
/* Check for user-requested abort. */
CHECK_FOR_INTERRUPTS();
@@ -333,7 +561,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
/* Check that the user still has privileges for the relation */
if (!cluster_is_permitted_for_relation(tableOid, save_userid, cmd))
{
- relation_close(OldHeap, AccessExclusiveLock);
+ relation_close(OldHeap, lmode);
goto out;
}
@@ -348,7 +576,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
*/
if (RELATION_IS_OTHER_TEMP(OldHeap))
{
- relation_close(OldHeap, AccessExclusiveLock);
+ relation_close(OldHeap, lmode);
goto out;
}
@@ -359,7 +587,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
*/
if (!SearchSysCacheExists1(RELOID, ObjectIdGetDatum(indexOid)))
{
- relation_close(OldHeap, AccessExclusiveLock);
+ relation_close(OldHeap, lmode);
goto out;
}
@@ -370,7 +598,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
if ((params->options & CLUOPT_RECHECK_ISCLUSTERED) != 0 &&
!get_index_isclustered(indexOid))
{
- relation_close(OldHeap, AccessExclusiveLock);
+ relation_close(OldHeap, lmode);
goto out;
}
}
@@ -390,6 +618,11 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
ereport(ERROR,
(errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
errmsg("cannot %s a shared catalog", cmd_str)));
+ /*
+ * The CONCURRENTLY case should have been rejected earlier because it does
+ * not support system catalogs.
+ */
+ Assert(!(OldHeap->rd_rel->relisshared && concurrent));
/*
* Don't process temp tables of other backends ... their local buffer
@@ -411,8 +644,7 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
if (OidIsValid(indexOid))
{
/* verify the index is good and lock it */
- check_index_is_clusterable(OldHeap, indexOid, AccessExclusiveLock,
- cmd);
+ check_index_is_clusterable(OldHeap, indexOid, lmode, cmd);
/* also open it */
index = index_open(indexOid, NoLock);
}
@@ -429,7 +661,8 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
if (OldHeap->rd_rel->relkind == RELKIND_MATVIEW &&
!RelationIsPopulated(OldHeap))
{
- relation_close(OldHeap, AccessExclusiveLock);
+ index_close(index, lmode);
+ relation_close(OldHeap, lmode);
goto out;
}
@@ -442,11 +675,42 @@ cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
* invalid, because we move tuples around. Promote them to relation
* locks. Predicate locks on indexes will be promoted when they are
* reindexed.
+ *
+ * During concurrent processing, the heap as well as its indexes stay in
+ * operation, so we postpone this step until they are locked using
+ * AccessExclusiveLock near the end of the processing.
*/
- TransferPredicateLocksToHeapRelation(OldHeap);
+ if (!concurrent)
+ TransferPredicateLocksToHeapRelation(OldHeap);
/* rebuild_relation does all the dirty work */
- rebuild_relation(OldHeap, index, verbose, cmd);
+ entered = false;
+ success = false;
+ PG_TRY();
+ {
+ /*
+ * For concurrent processing, make sure other transactions treat this
+ * table as if it was a system / user catalog, and WAL the relevant
+ * additional information. ERROR is raised if another backend is
+ * processing the same table.
+ */
+ if (concurrent)
+ {
+ Relation *index_p = index ? &index : NULL;
+
+ begin_concurrent_repack(&OldHeap, index_p, &entered);
+ }
+
+ rebuild_relation(OldHeap, index, verbose, cmd, concurrent);
+ success = true;
+ }
+ PG_FINALLY();
+ {
+ if (concurrent && entered)
+ end_concurrent_repack(!success);
+ }
+ PG_END_TRY();
+
/* rebuild_relation closes OldHeap, and index if valid */
out:
@@ -595,19 +859,86 @@ mark_index_clustered(Relation rel, Oid indexOid, bool is_internal)
table_close(pg_index, RowExclusiveLock);
}
+/*
+ * Check if the CONCURRENTLY option is legal for the relation.
+ */
+void
+can_repack_concurrently(Relation rel)
+{
+ char relpersistence, replident;
+ Oid ident_idx;
+
+ /* Data changes in system relations are not logically decoded. */
+ if (IsCatalogRelation(rel))
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot repack relation \"%s\"",
+ RelationGetRelationName(rel)),
+ errhint("REPACK CONCURRENTLY is not supported for catalog relations.")));
+
+ /*
+ * reorderbuffer.c does not seem to handle processing of TOAST relation
+ * alone.
+ */
+ if (IsToastRelation(rel))
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg("cannot repack relation \"%s\"",
+ RelationGetRelationName(rel)),
+ errhint("REPACK (CONCURRENTLY) is not supported for TOAST relations, unless the main relation is repacked too.")));
+
+ relpersistence = rel->rd_rel->relpersistence;
+ if (relpersistence != RELPERSISTENCE_PERMANENT)
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("cannot repack relation \"%s\"",
+ RelationGetRelationName(rel)),
+ errhint("REPACK (CONCURRENTLY) is only allowed for permanent relations.")));
+
+ /* With NOTHING, WAL does not contain the old tuple. */
+ replident = rel->rd_rel->relreplident;
+ if (replident == REPLICA_IDENTITY_NOTHING)
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("cannot repack relation \"%s\"",
+ RelationGetRelationName(rel)),
+ errhint("Relation \"%s\" has insufficient replication identity.",
+ RelationGetRelationName(rel))));
+
+ /*
+ * Identity index is not set if the replica identity is FULL, but PK might
+ * exist in such a case.
+ */
+ ident_idx = RelationGetReplicaIndex(rel);
+ if (!OidIsValid(ident_idx) && OidIsValid(rel->rd_pkindex))
+ ident_idx = rel->rd_pkindex;
+ if (!OidIsValid(ident_idx))
+ ereport(ERROR,
+ (errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+ errmsg("cannot process relation \"%s\"",
+ RelationGetRelationName(rel)),
+ (errhint("Relation \"%s\" has no identity index.",
+ RelationGetRelationName(rel)))));
+}
+
/*
* rebuild_relation: rebuild an existing relation in index or physical order
*
- * OldHeap: table to rebuild.
+ * OldHeap: table to rebuild. See cluster_rel() for comments on the required
+ * lock strength.
+ *
* index: index to cluster by, or NULL to rewrite in physical order.
*
- * On entry, heap and index (if one is given) must be open, and
- * AccessExclusiveLock held on them.
- * On exit, they are closed, but locks on them are not released.
+ * On entry, heap and index (if one is given) must be open, and the
+ * appropriate lock held on them (AccessExclusiveLock for exclusive processing
+ * and ShareUpdateExclusiveLock for concurrent processing)..
+ *
+ * On exit, they are closed, but still locked with AccessExclusiveLock (The
+ * function handles the lock upgrade if 'concurrent' is true.)
*/
static void
rebuild_relation(Relation OldHeap, Relation index, bool verbose,
- ClusterCommand cmd)
+ ClusterCommand cmd, bool concurrent)
{
Oid tableOid = RelationGetRelid(OldHeap);
Oid accessMethod = OldHeap->rd_rel->relam;
@@ -615,13 +946,81 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
Oid OIDNewHeap;
Relation NewHeap;
char relpersistence;
- bool is_system_catalog;
bool swap_toast_by_content;
TransactionId frozenXid;
MultiXactId cutoffMulti;
+ NameData slotname;
+ LogicalDecodingContext *ctx = NULL;
+ Snapshot snapshot = NULL;
+ CatalogState *cat_state = NULL;
+ LOCKMODE lmode;
+
+ lmode = !concurrent ? AccessExclusiveLock : ShareUpdateExclusiveLock;
+
+ Assert(CheckRelationLockedByMe(OldHeap, lmode, false) &&
+ (index == NULL || CheckRelationLockedByMe(index, lmode, false)));
+
+ if (concurrent)
+ {
+ TupleDesc tupdesc;
+ RelReopenInfo rri[2];
+ int nrel;
+
+ /*
+ * REPACK CONCURRENTLY is not allowed in a transaction block, so this
+ * should never fire.
+ */
+ Assert(GetTopTransactionIdIfAny() == InvalidTransactionId);
- Assert(CheckRelationLockedByMe(OldHeap, AccessExclusiveLock, false) &&
- (index == NULL || CheckRelationLockedByMe(index, AccessExclusiveLock, false)));
+ /*
+ * A single backend should not execute multiple REPACK commands at a
+ * time, so use PID to make the slot unique.
+ */
+ snprintf(NameStr(slotname), NAMEDATALEN, "repack_%d", MyProcPid);
+
+ /*
+ * Gather catalog information so that we can check later if the old
+ * relation has not changed while unlocked.
+ *
+ * Since this function also checks if the relation can be processed,
+ * it's important to call it before we spend notable amount of time to
+ * setup the logical decoding. Not sure though if it's necessary to do
+ * it even earlier.
+ */
+ cat_state = get_catalog_state(OldHeap);
+
+ tupdesc = CreateTupleDescCopy(RelationGetDescr(OldHeap));
+
+ /*
+ * Unlock the relation (and possibly the clustering index) to avoid
+ * deadlock because setup_logical_decoding() will wait for all the
+ * running transactions (with XID assigned) to finish. Some of those
+ * transactions might be waiting for a lock on our relation.
+ */
+ nrel = 0;
+ init_rel_reopen_info(&rri[nrel++], &OldHeap, InvalidOid,
+ ShareUpdateExclusiveLock,
+ ShareUpdateExclusiveLock);
+ if (index)
+ init_rel_reopen_info(&rri[nrel++], &index, InvalidOid,
+ ShareUpdateExclusiveLock,
+ ShareUpdateExclusiveLock);
+ unlock_and_close_relations(rri, nrel);
+
+ /* Prepare to capture the concurrent data changes. */
+ ctx = setup_logical_decoding(tableOid, NameStr(slotname), tupdesc);
+
+ /* Lock the table (and index) again. */
+ reopen_relations(rri, nrel);
+
+ /*
+ * Check if a 'tupdesc' could have changed while the relation was
+ * unlocked.
+ */
+ check_catalog_changes(OldHeap, cat_state);
+
+ snapshot = SnapBuildInitialSnapshotForRepack(ctx->snapshot_builder);
+ }
if (index)
/* Mark the correct index as clustered */
@@ -629,7 +1028,6 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
/* Remember info about rel before closing OldHeap */
relpersistence = OldHeap->rd_rel->relpersistence;
- is_system_catalog = IsSystemRelation(OldHeap);
/*
* Create the transient table that will receive the re-ordered data.
@@ -645,30 +1043,51 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
NewHeap = table_open(OIDNewHeap, NoLock);
/* Copy the heap data into the new table in the desired order */
- copy_table_data(NewHeap, OldHeap, index, verbose, cmd,
- &swap_toast_by_content, &frozenXid, &cutoffMulti);
+ copy_table_data(NewHeap, OldHeap, index, snapshot, ctx, verbose,
+ cmd, &swap_toast_by_content, &frozenXid, &cutoffMulti);
+ if (concurrent)
+ {
+ rebuild_relation_finish_concurrent(NewHeap, OldHeap, index,
+ cat_state, ctx,
+ swap_toast_by_content,
+ frozenXid, cutoffMulti);
+
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_FINAL_CLEANUP);
+
+ /* Done with decoding. */
+ FreeSnapshot(snapshot);
+ free_catalog_state(cat_state);
+ cleanup_logical_decoding(ctx);
+ ReplicationSlotRelease();
+ ReplicationSlotDrop(NameStr(slotname), false);
+ }
+ else
+ {
+ bool is_system_catalog = IsSystemRelation(OldHeap);
- /* Close relcache entries, but keep lock until transaction commit */
- table_close(OldHeap, NoLock);
- if (index)
- index_close(index, NoLock);
+ /* Close relcache entries, but keep lock until transaction commit */
+ table_close(OldHeap, NoLock);
+ if (index)
+ index_close(index, NoLock);
- /*
- * Close the new relation so it can be dropped as soon as the storage is
- * swapped. The relation is not visible to others, so no need to unlock it
- * explicitly.
- */
- table_close(NewHeap, NoLock);
+ /*
+ * Close the new relation so it can be dropped as soon as the storage
+ * is swapped. The relation is not visible to others, so no need to
+ * unlock it explicitly.
+ */
+ table_close(NewHeap, NoLock);
- /*
- * Swap the physical files of the target and transient tables, then
- * rebuild the target's indexes and throw away the transient table.
- */
- finish_heap_swap(tableOid, OIDNewHeap, is_system_catalog,
- swap_toast_by_content, false, true,
- frozenXid, cutoffMulti,
- relpersistence);
+ /*
+ * Swap the physical files of the target and transient tables, then
+ * rebuild the target's indexes and throw away the transient table.
+ */
+ finish_heap_swap(tableOid, OIDNewHeap, is_system_catalog,
+ swap_toast_by_content, false, true, true,
+ frozenXid, cutoffMulti,
+ relpersistence);
+ }
}
@@ -803,14 +1222,18 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
/*
* Do the physical copying of table data.
*
+ * 'snapshot' and 'decoding_ctx': see table_relation_copy_for_cluster(). Pass
+ * iff concurrent processing is required.
+ *
* There are three output parameters:
* *pSwapToastByContent is set true if toast tables must be swapped by content.
* *pFreezeXid receives the TransactionId used as freeze cutoff point.
* *pCutoffMulti receives the MultiXactId used as a cutoff point.
*/
static void
-copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verbose,
- ClusterCommand cmd, bool *pSwapToastByContent,
+copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
+ Snapshot snapshot, LogicalDecodingContext *decoding_ctx,
+ bool verbose, ClusterCommand cmd, bool *pSwapToastByContent,
TransactionId *pFreezeXid, MultiXactId *pCutoffMulti)
{
Relation relRelation;
@@ -829,6 +1252,7 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
const char *cmd_str = CLUSTER_COMMAND_STR(cmd);
PGRUsage ru0;
char *nspname;
+ bool concurrent = snapshot != NULL;
pg_rusage_init(&ru0);
@@ -855,8 +1279,12 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
*
* We don't need to open the toast relation here, just lock it. The lock
* will be held till end of transaction.
+ *
+ * In the REPACK CONCURRENTLY case, the lock does not help because we need
+ * to release it temporarily at some point. Instead, we expect VACUUM /
+ * CLUSTER to skip tables which are present in RepackedRelsHash.
*/
- if (OldHeap->rd_rel->reltoastrelid)
+ if (OldHeap->rd_rel->reltoastrelid && !concurrent)
LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
/*
@@ -932,8 +1360,48 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
* provided, else plain seqscan.
*/
if (OldIndex != NULL && OldIndex->rd_rel->relam == BTREE_AM_OID)
+ {
+ ResourceOwner oldowner = NULL;
+ ResourceOwner resowner = NULL;
+
+ /*
+ * In the CONCURRENT case, use a dedicated resource owner so we don't
+ * leave any additional locks behind us that we cannot release easily.
+ */
+ if (concurrent)
+ {
+ Assert(CheckRelationLockedByMe(OldHeap, ShareUpdateExclusiveLock,
+ false));
+ Assert(CheckRelationLockedByMe(OldIndex, ShareUpdateExclusiveLock,
+ false));
+
+ resowner = ResourceOwnerCreate(CurrentResourceOwner,
+ "plan_cluster_use_sort");
+ oldowner = CurrentResourceOwner;
+ CurrentResourceOwner = resowner;
+ }
+
use_sort = plan_cluster_use_sort(RelationGetRelid(OldHeap),
RelationGetRelid(OldIndex));
+
+ if (concurrent)
+ {
+ CurrentResourceOwner = oldowner;
+
+ /*
+ * We are primarily concerned about locks, but if the planner
+ * happened to allocate any other resources, we should release
+ * them too because we're going to delete the whole resowner.
+ */
+ ResourceOwnerRelease(resowner, RESOURCE_RELEASE_BEFORE_LOCKS,
+ false, false);
+ ResourceOwnerRelease(resowner, RESOURCE_RELEASE_LOCKS,
+ false, false);
+ ResourceOwnerRelease(resowner, RESOURCE_RELEASE_AFTER_LOCKS,
+ false, false);
+ ResourceOwnerDelete(resowner);
+ }
+ }
else
use_sort = false;
@@ -965,7 +1433,9 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
* values (e.g. because the AM doesn't use freezing).
*/
table_relation_copy_for_cluster(OldHeap, NewHeap, OldIndex, use_sort,
- cutoffs.OldestXmin, &cutoffs.FreezeLimit,
+ cutoffs.OldestXmin, snapshot,
+ decoding_ctx,
+ &cutoffs.FreezeLimit,
&cutoffs.MultiXactCutoff,
&num_tuples, &tups_vacuumed,
&tups_recently_dead);
@@ -974,7 +1444,11 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex, bool verb
*pFreezeXid = cutoffs.FreezeLimit;
*pCutoffMulti = cutoffs.MultiXactCutoff;
- /* Reset rd_toastoid just to be tidy --- it shouldn't be looked at again */
+ /*
+ * Reset rd_toastoid just to be tidy --- it shouldn't be looked at
+ * again. In the CONCURRENTLY case, we need to set it again before
+ * applying the concurrent changes.
+ */
NewHeap->rd_toastoid = InvalidOid;
num_pages = RelationGetNumberOfBlocks(NewHeap);
@@ -1427,14 +1901,13 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
bool swap_toast_by_content,
bool check_constraints,
bool is_internal,
+ bool reindex,
TransactionId frozenXid,
MultiXactId cutoffMulti,
char newrelpersistence)
{
ObjectAddress object;
Oid mapped_tables[4];
- int reindex_flags;
- ReindexParams reindex_params = {0};
int i;
/* Report that we are now swapping relation files */
@@ -1460,39 +1933,46 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
if (is_system_catalog)
CacheInvalidateCatalog(OIDOldHeap);
- /*
- * Rebuild each index on the relation (but not the toast table, which is
- * all-new at this point). It is important to do this before the DROP
- * step because if we are processing a system catalog that will be used
- * during DROP, we want to have its indexes available. There is no
- * advantage to the other order anyway because this is all transactional,
- * so no chance to reclaim disk space before commit. We do not need a
- * final CommandCounterIncrement() because reindex_relation does it.
- *
- * Note: because index_build is called via reindex_relation, it will never
- * set indcheckxmin true for the indexes. This is OK even though in some
- * sense we are building new indexes rather than rebuilding existing ones,
- * because the new heap won't contain any HOT chains at all, let alone
- * broken ones, so it can't be necessary to set indcheckxmin.
- */
- reindex_flags = REINDEX_REL_SUPPRESS_INDEX_USE;
- if (check_constraints)
- reindex_flags |= REINDEX_REL_CHECK_CONSTRAINTS;
+ if (reindex)
+ {
+ int reindex_flags;
+ ReindexParams reindex_params = {0};
- /*
- * Ensure that the indexes have the same persistence as the parent
- * relation.
- */
- if (newrelpersistence == RELPERSISTENCE_UNLOGGED)
- reindex_flags |= REINDEX_REL_FORCE_INDEXES_UNLOGGED;
- else if (newrelpersistence == RELPERSISTENCE_PERMANENT)
- reindex_flags |= REINDEX_REL_FORCE_INDEXES_PERMANENT;
+ /*
+ * Rebuild each index on the relation (but not the toast table, which
+ * is all-new at this point). It is important to do this before the
+ * DROP step because if we are processing a system catalog that will
+ * be used during DROP, we want to have its indexes available. There
+ * is no advantage to the other order anyway because this is all
+ * transactional, so no chance to reclaim disk space before commit.
+ * We do not need a final CommandCounterIncrement() because
+ * reindex_relation does it.
+ *
+ * Note: because index_build is called via reindex_relation, it will never
+ * set indcheckxmin true for the indexes. This is OK even though in some
+ * sense we are building new indexes rather than rebuilding existing ones,
+ * because the new heap won't contain any HOT chains at all, let alone
+ * broken ones, so it can't be necessary to set indcheckxmin.
+ */
+ reindex_flags = REINDEX_REL_SUPPRESS_INDEX_USE;
+ if (check_constraints)
+ reindex_flags |= REINDEX_REL_CHECK_CONSTRAINTS;
- /* Report that we are now reindexing relations */
- pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
- PROGRESS_REPACK_PHASE_REBUILD_INDEX);
+ /*
+ * Ensure that the indexes have the same persistence as the parent
+ * relation.
+ */
+ if (newrelpersistence == RELPERSISTENCE_UNLOGGED)
+ reindex_flags |= REINDEX_REL_FORCE_INDEXES_UNLOGGED;
+ else if (newrelpersistence == RELPERSISTENCE_PERMANENT)
+ reindex_flags |= REINDEX_REL_FORCE_INDEXES_PERMANENT;
+
+ /* Report that we are now reindexing relations */
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_REBUILD_INDEX);
- reindex_relation(NULL, OIDOldHeap, reindex_flags, &reindex_params);
+ reindex_relation(NULL, OIDOldHeap, reindex_flags, &reindex_params);
+ }
/* Report that we are now doing clean up */
pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
@@ -1804,89 +2284,1975 @@ cluster_is_permitted_for_relation(Oid relid, Oid userid, ClusterCommand cmd)
return false;
}
+#define REPL_PLUGIN_NAME "pgoutput_repack"
+
/*
- * REPACK is intended to be a replacement of both CLUSTER and VACUUM FULL.
+ * Each relation being processed by REPACK CONCURRENTLY must be in the
+ * repackedRels hashtable.
*/
+typedef struct RepackedRel
+{
+ Oid relid;
+ Oid dbid;
+} RepackedRel;
+
+static HTAB *RepackedRelsHash = NULL;
+
+/* Maximum number of entries in the hashtable. */
+static int maxRepackedRels = 0;
+
+Size
+RepackShmemSize(void)
+{
+ /*
+ * A replication slot is needed for the processing, so use this GUC to
+ * allocate memory for the hashtable.
+ */
+ maxRepackedRels = max_replication_slots;
+
+ return hash_estimate_size(maxRepackedRels, sizeof(RepackedRel));
+}
+
void
-repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
+RepackShmemInit(void)
{
- ListCell *lc;
- ClusterParams params = {0};
- bool verbose = false;
- Relation rel = NULL;
- Oid indexOid = InvalidOid;
- MemoryContext repack_context;
- List *rtcs;
+ HASHCTL info;
- /* Parse option list */
- foreach(lc, stmt->params)
- {
- DefElem *opt = (DefElem *) lfirst(lc);
+ info.keysize = sizeof(RepackedRel);
+ info.entrysize = info.keysize;
- if (strcmp(opt->defname, "verbose") == 0)
- verbose = defGetBoolean(opt);
- else
- ereport(ERROR,
- (errcode(ERRCODE_SYNTAX_ERROR),
- errmsg("unrecognized REPACK option \"%s\"",
- opt->defname),
- parser_errposition(pstate, opt->location)));
+ RepackedRelsHash = ShmemInitHash("Repacked Relations",
+ maxRepackedRels,
+ maxRepackedRels,
+ &info,
+ HASH_ELEM | HASH_BLOBS);
+}
+
+/*
+ * Call this function before REPACK CONCURRENTLY starts to setup logical
+ * decoding. It makes sure that other users of the table put enough
+ * information into WAL.
+ *
+ * The point is that on various places we expect that the table we're
+ * processing is treated like a system catalog. For example, we need to be
+ * able to scan it using a "historic snapshot" anytime during the processing
+ * (as opposed to scanning only at the start point of the decoding, logical
+ * replication does during initial table synchronization), in order to apply
+ * concurrent UPDATE / DELETE commands.
+ *
+ * Since we need to close and reopen the relation here, the 'rel_p' and
+ * 'index_p' arguments are in/out.
+ *
+ * 'enter_p' receives a bool value telling whether relation OID was entered
+ * into the hashtable or not.
+ */
+static void
+begin_concurrent_repack(Relation *rel_p, Relation *index_p,
+ bool *entered_p)
+{
+ Relation rel = *rel_p;
+ Oid relid, toastrelid;
+ RepackedRel key, *entry;
+ bool found;
+ RelReopenInfo rri[2];
+ int nrel;
+ static bool before_shmem_exit_callback_setup = false;
+
+ relid = RelationGetRelid(rel);
+
+ /*
+ * Make sure that we do not leave an entry in RepackedRelsHash if exiting
+ * due to FATAL.
+ */
+ if (!before_shmem_exit_callback_setup)
+ {
+ before_shmem_exit(cluster_before_shmem_exit_callback, 0);
+ before_shmem_exit_callback_setup = true;
}
- params.options = (verbose ? CLUOPT_VERBOSE : 0);
+ memset(&key, 0, sizeof(key));
+ key.relid = relid;
+ key.dbid = MyDatabaseId;
- if (stmt->relation != NULL)
+ *entered_p = false;
+ LWLockAcquire(RepackedRelsLock, LW_EXCLUSIVE);
+ entry = (RepackedRel *)
+ hash_search(RepackedRelsHash, &key, HASH_ENTER_NULL, &found);
+ if (found)
{
- rel = process_single_relation(stmt->relation, stmt->indexname,
- CLUSTER_COMMAND_REPACK, ¶ms,
- &indexOid);
- if (rel == NULL)
- return;
+ /*
+ * Since REPACK CONCURRENTLY takes ShareRowExclusiveLock, a conflict
+ * should occur much earlier. However that lock may be released
+ * temporarily, see below. Anyway, we should complain whatever the
+ * reason of the conflict might be.
+ */
+ ereport(ERROR,
+ (errmsg(REPACK_CONCURRENT_IN_PROGRESS_MSG,
+ RelationGetRelationName(rel))));
}
+ if (entry == NULL)
+ ereport(ERROR,
+ (errmsg("too many requests for REPACK CONCURRENTLY at a time")),
+ (errhint("Please consider increasing the \"max_replication_slots\" configuration parameter.")));
/*
- * By here, we know we are in a multi-table situation. In order to avoid
- * holding locks for too long, we want to process each table in its own
- * transaction. This forces us to disallow running inside a user
- * transaction block.
+ * Even if the insertion of TOAST relid should fail below, the caller has
+ * to do cleanup.
*/
- PreventInTransactionBlock(isTopLevel, "REPACK");
+ *entered_p = true;
- /* Also, we need a memory context to hold our list of relations */
- repack_context = AllocSetContextCreate(PortalContext,
- "Repack",
- ALLOCSET_DEFAULT_SIZES);
+ /*
+ * Enable the callback to remove the entry in case of exit. We should not
+ * do this earlier, otherwise an attempt to insert already existing entry
+ * could make us remove that entry (inserted by another backend) during
+ * ERROR handling.
+ */
+ Assert(!OidIsValid(repacked_rel));
+ repacked_rel = relid;
- params.options |= CLUOPT_RECHECK;
- if (rel != NULL)
+ /*
+ * TOAST relation is not accessed using historic snapshot, but we enter it
+ * here to protect it from being VACUUMed by another backend. (Lock does
+ * not help in the CONCURRENTLY case because cannot hold it continuously
+ * till the end of the transaction.) See the comments on locking TOAST
+ * relation in copy_table_data().
+ */
+ toastrelid = rel->rd_rel->reltoastrelid;
+ if (OidIsValid(toastrelid))
{
- Oid relid;
- bool rel_is_index;
+ key.relid = toastrelid;
+ entry = (RepackedRel *)
+ hash_search(RepackedRelsHash, &key, HASH_ENTER_NULL, &found);
+ if (found)
+ /*
+ * If we could enter the main fork the TOAST should succeed
+ * too. Nevertheless, check.
+ */
+ ereport(ERROR,
+ (errmsg("TOAST relation of \"%s\" is already being processed by REPACK CONCURRENTLY",
+ RelationGetRelationName(rel))));
+ if (entry == NULL)
+ ereport(ERROR,
+ (errmsg("too many requests for REPACK CONCURRENTLY at a time")),
+ (errhint("Please consider increasing the \"max_replication_slots\" configuration parameter.")));
- Assert(rel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE);
+ Assert(!OidIsValid(repacked_rel_toast));
+ repacked_rel_toast = toastrelid;
+ }
+ LWLockRelease(RepackedRelsLock);
- if (OidIsValid(indexOid))
- {
- relid = indexOid;
- rel_is_index = true;
- }
- else
- {
- relid = RelationGetRelid(rel);
- rel_is_index = false;
- }
- rtcs = get_tables_to_cluster_partitioned(repack_context, relid,
- rel_is_index,
- CLUSTER_COMMAND_REPACK);
+ /*
+ * Make sure that other backends are aware of the new hash entry.
+ *
+ * Besides sending the invalidation message, we need to force re-opening
+ * of the relation, which includes the actual invalidation (and thus
+ * checking of our hashtable on the next access).
+ */
+ CacheInvalidateRelcacheImmediate(rel);
+ /*
+ * Since the hashtable only needs to be checked by write transactions,
+ * lock the relation in a mode that conflicts with any DML command. (The
+ * reading transactions are supposed to close the relation before opening
+ * it with higher lock.) Once we have the relation (and its index) locked,
+ * we unlock it immediately and then re-lock using the original mode.
+ */
+ nrel = 0;
+ init_rel_reopen_info(&rri[nrel++], rel_p, InvalidOid,
+ ShareUpdateExclusiveLock, ShareLock);
+ if (index_p)
+ {
+ /*
+ * Another transaction might want to open both the relation and the
+ * index. If it already has the relation lock and is waiting for the
+ * index lock, we should release the index lock, otherwise our request
+ * for ShareLock on the relation can end up in a deadlock.
+ */
+ init_rel_reopen_info(&rri[nrel++], index_p, InvalidOid,
+ ShareUpdateExclusiveLock, ShareLock);
+ }
+ unlock_and_close_relations(rri, nrel);
+ /*
+ * XXX It's not strictly necessary to lock the index here, but it's
+ * probably not worth teaching the "reopen API" about this special case.
+ */
+ reopen_relations(rri, nrel);
+
+ /* Switch back to the original lock. */
+ nrel = 0;
+ init_rel_reopen_info(&rri[nrel++], rel_p, InvalidOid,
+ ShareLock, ShareUpdateExclusiveLock);
+ if (index_p)
+ init_rel_reopen_info(&rri[nrel++], index_p, InvalidOid,
+ ShareLock, ShareUpdateExclusiveLock);
+ unlock_and_close_relations(rri, nrel);
+ reopen_relations(rri, nrel);
+ /* Make sure the reopened relcache entry is used, not the old one. */
+ rel = *rel_p;
+
+ /* Avoid logical decoding of other relations by this backend. */
+ repacked_rel_locator = rel->rd_locator;
+ if (OidIsValid(toastrelid))
+ {
+ Relation toastrel;
+
+ /* Avoid logical decoding of other TOAST relations. */
+ toastrel = table_open(toastrelid, AccessShareLock);
+ repacked_rel_toast_locator = toastrel->rd_locator;
+ table_close(toastrel, AccessShareLock);
+ }
+}
+
+/*
+ * Call this when done with REPACK CONCURRENTLY.
+ *
+ * 'error' tells whether the function is being called in order to handle
+ * error.
+ */
+static void
+end_concurrent_repack(bool error)
+{
+ RepackedRel key;
+ RepackedRel *entry = NULL, *entry_toast = NULL;
+ Oid relid = repacked_rel;
+ Oid toastrelid = repacked_rel_toast;
+
+ /* Remove the relation from the hash if we managed to insert one. */
+ if (OidIsValid(repacked_rel))
+ {
+ memset(&key, 0, sizeof(key));
+ key.relid = repacked_rel;
+ key.dbid = MyDatabaseId;
+ LWLockAcquire(RepackedRelsLock, LW_EXCLUSIVE);
+ entry = hash_search(RepackedRelsHash, &key, HASH_REMOVE, NULL);
+
+ /*
+ * By clearing this variable we also disable
+ * cluster_before_shmem_exit_callback().
+ */
+ repacked_rel = InvalidOid;
+ }
+
+ /* Remove the TOAST relation if there is one. */
+ if (OidIsValid(repacked_rel_toast))
+ {
+ key.relid = repacked_rel_toast;
+ entry_toast = hash_search(RepackedRelsHash, &key, HASH_REMOVE,
+ NULL);
+
+ repacked_rel_toast = InvalidOid;
+ }
+ LWLockRelease(RepackedRelsLock);
+
+ /* Restore normal function of logical decoding. */
+ repacked_rel_locator.relNumber = InvalidOid;
+ repacked_rel_toast_locator.relNumber = InvalidOid;
+
+ /*
+ * On normal completion (!error), we should not really fail to remove the
+ * entry. But if it wasn't there for any reason, raise ERROR to make sure
+ * the transaction is aborted: if other transactions, while changing the
+ * contents of the relation, didn't know that REPACK CONCURRENTLY was in
+ * progress, they could have missed to WAL enough information, and thus we
+ * could have produced an inconsistent table contents.
+ *
+ * On the other hand, if we are already handling an error, there's no
+ * reason to worry about inconsistent contents of the new storage because
+ * the transaction is going to be rolled back anyway. Furthermore, by
+ * raising ERROR here we'd shadow the original error.
+ */
+ if (!error)
+ {
+ char *relname;
+
+ if (OidIsValid(relid) && entry == NULL)
+ {
+ relname = get_rel_name(relid);
+ if (!relname)
+ ereport(ERROR,
+ (errmsg("cache lookup failed for relation %u",
+ relid)));
+
+ ereport(ERROR,
+ (errmsg("relation \"%s\" not found among repacked relations",
+ relname)));
+ }
+
+ /*
+ * Likewise, the TOAST relation should not have disappeared.
+ */
+ if (OidIsValid(toastrelid) && entry_toast == NULL)
+ {
+ relname = get_rel_name(key.relid);
+ if (!relname)
+ ereport(ERROR,
+ (errmsg("cache lookup failed for relation %u",
+ key.relid)));
+
+ ereport(ERROR,
+ (errmsg("relation \"%s\" not found among repacked relations",
+ relname)));
+ }
+ }
+
+ /*
+ * Note: unlike begin_concurrent_repack(), here we do not lock/unlock the
+ * relation: 1) On normal completion, the caller is already holding
+ * AccessExclusiveLock (till the end of the transaction), 2) on ERROR /
+ * FATAL, we try to do the cleanup asap, but the worst case is that other
+ * backends will write unnecessary information to WAL until they close the
+ * relation.
+ */
+}
+
+/*
+ * A wrapper to call end_concurrent_repack() as a before_shmem_exit callback.
+ */
+static void
+cluster_before_shmem_exit_callback(int code, Datum arg)
+{
+ if (OidIsValid(repacked_rel) || OidIsValid(repacked_rel_toast))
+ end_concurrent_repack(true);
+}
+
+/*
+ * Check if relation is currently being processed by REPACK CONCURRENTLY.
+ */
+bool
+is_concurrent_repack_in_progress(Oid relid)
+{
+ RepackedRel key, *entry;
+
+ memset(&key, 0, sizeof(key));
+ key.relid = relid;
+ key.dbid = MyDatabaseId;
+
+ LWLockAcquire(RepackedRelsLock, LW_SHARED);
+ entry = (RepackedRel *)
+ hash_search(RepackedRelsHash, &key, HASH_FIND, NULL);
+ LWLockRelease(RepackedRelsLock);
+
+ return entry != NULL;
+}
+
+/*
+ * Check if REPACK CONCURRENTLY is already running for given relation, and if
+ * so, raise ERROR. The problem is that cluster_rel() needs to release its
+ * lock on the relation temporarily at some point, so our lock alone does not
+ * help. Commands that might break what cluster_rel() is doing should call
+ * this function first.
+ *
+ * Return without checking if lockmode allows for race conditions which would
+ * make the result meaningless. In that case, cluster_rel() itself should
+ * throw ERROR if the relation was changed by us in an incompatible
+ * way. However, if it managed to do most of its work by then, a lot of CPU
+ * time might be wasted.
+ */
+void
+check_for_concurrent_repack(Oid relid, LOCKMODE lockmode)
+{
+ /*
+ * If the caller does not have a lock that conflicts with
+ * ShareUpdateExclusiveLock, the check makes little sense because REPACK
+ * CONCURRENTLY can start anytime after the check.
+ */
+ if (lockmode < ShareUpdateExclusiveLock)
+ return;
+
+ /*
+ * The caller has a lock which conflicts with REPACK CONCURRENTLY, so if
+ * that's not running now, it cannot start until the caller's transaction
+ * has completed.
+ */
+ if (is_concurrent_repack_in_progress(relid))
+ ereport(ERROR,
+ (errcode(ERRCODE_FEATURE_NOT_SUPPORTED),
+ errmsg(REPACK_CONCURRENT_IN_PROGRESS_MSG,
+ get_rel_name(relid))));
+
+}
+
+/*
+ * Check if relation is eligible for REPACK CONCURRENTLY and retrieve the
+ * catalog state to be passed later to check_catalog_changes.
+ *
+ * Caller is supposed to hold (at least) ShareUpdateExclusiveLock on the
+ * relation.
+ */
+static CatalogState *
+get_catalog_state(Relation rel)
+{
+ CatalogState *result = palloc_object(CatalogState);
+ List *ind_oids;
+ ListCell *lc;
+ int ninds, i;
+ char relpersistence = rel->rd_rel->relpersistence;
+ char replident = rel->rd_rel->relreplident;
+ Oid ident_idx = RelationGetReplicaIndex(rel);
+ TupleDesc td_src = RelationGetDescr(rel);
+
+ /*
+ * While gathering the catalog information, check if there is a reason not
+ * to proceed.
+ *
+ * This function was already called, but the relation was unlocked since
+ * (see begin_concurrent_repack()). check_catalog_changes() should catch
+ * any "disruptive" changes in the future.
+ */
+ can_repack_concurrently(rel);
+
+ /* No index should be dropped while we are checking it. */
+ Assert(CheckRelationLockedByMe(rel, ShareUpdateExclusiveLock, true));
+
+ ind_oids = RelationGetIndexList(rel);
+ result->ninds = ninds = list_length(ind_oids);
+ result->ind_oids = palloc_array(Oid, ninds);
+ result->ind_tupdescs = palloc_array(TupleDesc, ninds);
+ i = 0;
+ foreach(lc, ind_oids)
+ {
+ Oid ind_oid = lfirst_oid(lc);
+ Relation index;
+ TupleDesc td_ind_src, td_ind_dst;
+
+ /*
+ * Weaker lock should be o.k. for the index, but this one should not
+ * break anything either.
+ */
+ index = index_open(ind_oid, ShareUpdateExclusiveLock);
+
+ result->ind_oids[i] = RelationGetRelid(index);
+ td_ind_src = RelationGetDescr(index);
+ td_ind_dst = palloc(TupleDescSize(td_ind_src));
+ TupleDescCopy(td_ind_dst, td_ind_src);
+ result->ind_tupdescs[i] = td_ind_dst;
+ i++;
+
+ index_close(index, ShareUpdateExclusiveLock);
+ }
+
+ /* Fill-in the relation info. */
+ result->tupdesc = palloc(TupleDescSize(td_src));
+ TupleDescCopy(result->tupdesc, td_src);
+ result->relpersistence = relpersistence;
+ result->replident = replident;
+ result->replidindex = ident_idx;
+
+ return result;
+}
+
+static void
+free_catalog_state(CatalogState *state)
+{
+ /* We are only interested in indexes. */
+ if (state->ninds == 0)
+ return;
+
+ for (int i = 0; i < state->ninds; i++)
+ FreeTupleDesc(state->ind_tupdescs[i]);
+
+ FreeTupleDesc(state->tupdesc);
+ pfree(state->ind_oids);
+ pfree(state->ind_tupdescs);
+ pfree(state);
+}
+
+/*
+ * Raise ERROR if 'rel' changed in a way that does not allow further
+ * processing of REPACK CONCURRENTLY.
+ *
+ * Besides the relation's tuple descriptor, it's important to check indexes:
+ * concurrent change of index definition (can it happen in other way than
+ * dropping and re-creating the index, accidentally with the same OID?) can be
+ * a problem because we may already have the new index built. If an index was
+ * created or dropped concurrently, we'd fail to swap the index storage. In
+ * any case, we prefer to check the indexes early to get an explicit error
+ * message about the mismatch. Furthermore, the earlier we detect the change,
+ * the fewer CPU cycles we waste.
+ *
+ * Note that we do not check constraints because the transaction which changed
+ * them must have ensured that the existing tuples satisfy the new
+ * constraints. If any DML commands were necessary for that, we will simply
+ * decode them from WAL and apply them to the new storage.
+ *
+ * Caller is supposed to hold (at least) ShareUpdateExclusiveLock on the
+ * relation.
+ */
+static void
+check_catalog_changes(Relation rel, CatalogState *cat_state)
+{
+ Oid reltoastrelid = rel->rd_rel->reltoastrelid;
+ List *ind_oids;
+ ListCell *lc;
+ LOCKMODE lockmode;
+ Oid ident_idx;
+ TupleDesc td, td_cp;
+
+ /* First, check the relation info. */
+
+ /* TOAST is not easy to change, but check. */
+ if (reltoastrelid != repacked_rel_toast)
+ ereport(ERROR,
+ errmsg("TOAST relation of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel)));
+
+ /*
+ * Likewise, check_for_concurrent_repack() should prevent others from
+ * changing the relation file concurrently, but it's our responsibility to
+ * avoid data loss. (The original locators are stored outside cat_state,
+ * but the check belongs to this function.)
+ */
+ if (!RelFileLocatorEquals(rel->rd_locator, repacked_rel_locator))
+ ereport(ERROR,
+ (errmsg("file of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel))));
+ if (OidIsValid(reltoastrelid))
+ {
+ Relation toastrel;
+
+ toastrel = table_open(reltoastrelid, AccessShareLock);
+ if (!RelFileLocatorEquals(toastrel->rd_locator,
+ repacked_rel_toast_locator))
+ ereport(ERROR,
+ (errmsg("file of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(toastrel))));
+ table_close(toastrel, AccessShareLock);
+ }
+
+ if (rel->rd_rel->relpersistence != cat_state->relpersistence)
+ ereport(ERROR,
+ errmsg("persistence of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel)));
+
+ if (cat_state->replident != rel->rd_rel->relreplident)
+ ereport(ERROR,
+ errmsg("replica identity of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel)));
+
+ ident_idx = RelationGetReplicaIndex(rel);
+ if (ident_idx == InvalidOid && rel->rd_pkindex != InvalidOid)
+ ident_idx = rel->rd_pkindex;
+ if (cat_state->replidindex != ident_idx)
+ ereport(ERROR,
+ errmsg("identity index of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel)));
+
+ /*
+ * As cat_state contains a copy (which has the constraint info cleared),
+ * create a temporary copy for the comparison.
+ */
+ td = RelationGetDescr(rel);
+ td_cp = palloc(TupleDescSize(td));
+ TupleDescCopy(td_cp, td);
+ if (!equalTupleDescs(cat_state->tupdesc, td_cp))
+ ereport(ERROR,
+ errmsg("definition of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel)));
+ FreeTupleDesc(td_cp);
+
+ /* Now we are only interested in indexes. */
+ if (cat_state->ninds == 0)
+ return;
+
+ /* No index should be dropped while we are checking the relation. */
+ lockmode = ShareUpdateExclusiveLock;
+ Assert(CheckRelationLockedByMe(rel, lockmode, true));
+
+ ind_oids = RelationGetIndexList(rel);
+ if (list_length(ind_oids) != cat_state->ninds)
+ goto failed_index;
+
+ foreach(lc, ind_oids)
+ {
+ Oid ind_oid = lfirst_oid(lc);
+ int i;
+ TupleDesc tupdesc;
+ Relation index;
+
+ /* Find the index in cat_state. */
+ for (i = 0; i < cat_state->ninds; i++)
+ {
+ if (cat_state->ind_oids[i] == ind_oid)
+ break;
+ }
+ /*
+ * OID not found, i.e. the index was replaced by another one. XXX
+ * Should we yet try to find if an index having the desired tuple
+ * descriptor exists? Or should we always look for the tuple
+ * descriptor and not use OIDs at all?
+ */
+ if (i == cat_state->ninds)
+ goto failed_index;
+
+ /* Check the tuple descriptor. */
+ index = try_index_open(ind_oid, lockmode);
+ if (index == NULL)
+ goto failed_index;
+ tupdesc = RelationGetDescr(index);
+ if (!equalTupleDescs(cat_state->ind_tupdescs[i], tupdesc))
+ goto failed_index;
+ index_close(index, lockmode);
+ }
+
+ return;
+
+failed_index:
+ ereport(ERROR,
+ (errmsg("index(es) of relation \"%s\" changed by another transaction",
+ RelationGetRelationName(rel))));
+}
+
+/*
+ * This function is much like pg_create_logical_replication_slot() except that
+ * the new slot is neither released (if anyone else could read changes from
+ * our slot, we could miss changes other backends do while we copy the
+ * existing data into temporary table), nor persisted (it's easier to handle
+ * crash by restarting all the work from scratch).
+ */
+static LogicalDecodingContext *
+setup_logical_decoding(Oid relid, const char *slotname, TupleDesc tupdesc)
+{
+ LogicalDecodingContext *ctx;
+ RepackDecodingState *dstate;
+
+ /*
+ * Check if we can use logical decoding.
+ */
+ CheckSlotPermissions();
+ CheckLogicalDecodingRequirements();
+
+ /* RS_TEMPORARY so that the slot gets cleaned up on ERROR. */
+ ReplicationSlotCreate(slotname, true, RS_TEMPORARY, false, false, false);
+
+ /*
+ * Neither prepare_write nor do_write callback nor update_progress is
+ * useful for us.
+ *
+ * Regarding the value of need_full_snapshot, we pass false because the
+ * table we are processing is present in RepackedRelsHash and therefore,
+ * regarding logical decoding, treated like a catalog.
+ */
+ ctx = CreateInitDecodingContext(REPL_PLUGIN_NAME,
+ NIL,
+ false,
+ InvalidXLogRecPtr,
+ XL_ROUTINE(.page_read = read_local_xlog_page,
+ .segment_open = wal_segment_open,
+ .segment_close = wal_segment_close),
+ NULL, NULL, NULL);
+
+ /*
+ * We don't have control on setting fast_forward, so at least check it.
+ */
+ Assert(!ctx->fast_forward);
+
+ DecodingContextFindStartpoint(ctx);
+
+ /* Some WAL records should have been read. */
+ Assert(ctx->reader->EndRecPtr != InvalidXLogRecPtr);
+
+ XLByteToSeg(ctx->reader->EndRecPtr, repack_current_segment,
+ wal_segment_size);
+
+ /*
+ * Setup structures to store decoded changes.
+ */
+ dstate = palloc0(sizeof(RepackDecodingState));
+ dstate->relid = relid;
+ dstate->tstore = tuplestore_begin_heap(false, false,
+ maintenance_work_mem);
+ dstate->tupdesc = tupdesc;
+
+ /* Initialize the descriptor to store the changes ... */
+ dstate->tupdesc_change = CreateTemplateTupleDesc(1);
+
+ TupleDescInitEntry(dstate->tupdesc_change, 1, NULL, BYTEAOID, -1, 0);
+ /* ... as well as the corresponding slot. */
+ dstate->tsslot = MakeSingleTupleTableSlot(dstate->tupdesc_change,
+ &TTSOpsMinimalTuple);
+
+ dstate->resowner = ResourceOwnerCreate(CurrentResourceOwner,
+ "logical decoding");
+
+ ctx->output_writer_private = dstate;
+ return ctx;
+}
+
+/*
+ * Retrieve tuple from ConcurrentChange structure.
+ *
+ * The input data starts with the structure but it might not be appropriately
+ * aligned.
+ */
+static HeapTuple
+get_changed_tuple(char *change)
+{
+ HeapTupleData tup_data;
+ HeapTuple result;
+ char *src;
+
+ /*
+ * Ensure alignment before accessing the fields. (This is why we can't use
+ * heap_copytuple() instead of this function.)
+ */
+ src = change + offsetof(ConcurrentChange, tup_data);
+ memcpy(&tup_data, src, sizeof(HeapTupleData));
+
+ result = (HeapTuple) palloc(HEAPTUPLESIZE + tup_data.t_len);
+ memcpy(result, &tup_data, sizeof(HeapTupleData));
+ result->t_data = (HeapTupleHeader) ((char *) result + HEAPTUPLESIZE);
+ src = change + SizeOfConcurrentChange;
+ memcpy(result->t_data, src, result->t_len);
+
+ return result;
+}
+
+/*
+ * Decode logical changes from the WAL sequence up to end_of_wal.
+ */
+void
+repack_decode_concurrent_changes(LogicalDecodingContext *ctx,
+ XLogRecPtr end_of_wal)
+{
+ RepackDecodingState *dstate;
+ ResourceOwner resowner_old;
+ PgBackendProgress progress;
+
+ /*
+ * Invalidate the "present" cache before moving to "(recent) history".
+ */
+ InvalidateSystemCaches();
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+ resowner_old = CurrentResourceOwner;
+ CurrentResourceOwner = dstate->resowner;
+
+ /*
+ * reorderbuffer.c uses internal subtransaction, whose abort ends the
+ * command progress reporting. Save the status here so we can restore when
+ * done with the decoding.
+ */
+ memcpy(&progress, &MyBEEntry->st_progress, sizeof(PgBackendProgress));
+
+ PG_TRY();
+ {
+ while (ctx->reader->EndRecPtr < end_of_wal)
+ {
+ XLogRecord *record;
+ XLogSegNo segno_new;
+ char *errm = NULL;
+ XLogRecPtr end_lsn;
+
+ record = XLogReadRecord(ctx->reader, &errm);
+ if (errm)
+ elog(ERROR, "%s", errm);
+
+ if (record != NULL)
+ LogicalDecodingProcessRecord(ctx, ctx->reader);
+
+ /*
+ * If WAL segment boundary has been crossed, inform the decoding
+ * system that the catalog_xmin can advance. (We can confirm more
+ * often, but a filling a single WAL segment should not take much
+ * time.)
+ */
+ end_lsn = ctx->reader->EndRecPtr;
+ XLByteToSeg(end_lsn, segno_new, wal_segment_size);
+ if (segno_new != repack_current_segment)
+ {
+ LogicalConfirmReceivedLocation(end_lsn);
+ elog(DEBUG1, "REPACK: confirmed receive location %X/%X",
+ (uint32) (end_lsn >> 32), (uint32) end_lsn);
+ repack_current_segment = segno_new;
+ }
+
+ CHECK_FOR_INTERRUPTS();
+ }
+ InvalidateSystemCaches();
+ CurrentResourceOwner = resowner_old;
+ }
+ PG_CATCH();
+ {
+ InvalidateSystemCaches();
+ CurrentResourceOwner = resowner_old;
+ PG_RE_THROW();
+ }
+ PG_END_TRY();
+
+ /* Restore the progress reporting status. */
+ pgstat_progress_restore_state(&progress);
+}
+
+/*
+ * Apply changes that happened during the initial load.
+ *
+ * Scan key is passed by caller, so it does not have to be constructed
+ * multiple times. Key entries have all fields initialized, except for
+ * sk_argument.
+ */
+static void
+apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
+ ScanKey key, int nkeys, IndexInsertState *iistate)
+{
+ TupleTableSlot *index_slot, *ident_slot;
+ HeapTuple tup_old = NULL;
+
+ if (dstate->nchanges == 0)
+ return;
+
+ /* TupleTableSlot is needed to pass the tuple to ExecInsertIndexTuples(). */
+ index_slot = MakeSingleTupleTableSlot(dstate->tupdesc, &TTSOpsHeapTuple);
+ iistate->econtext->ecxt_scantuple = index_slot;
+
+ /* A slot to fetch tuples from identity index. */
+ ident_slot = table_slot_create(rel, NULL);
+
+ while (tuplestore_gettupleslot(dstate->tstore, true, false,
+ dstate->tsslot))
+ {
+ bool shouldFree;
+ HeapTuple tup_change,
+ tup,
+ tup_exist;
+ char *change_raw, *src;
+ ConcurrentChange change;
+ bool isnull[1];
+ Datum values[1];
+
+ CHECK_FOR_INTERRUPTS();
+
+ /* Get the change from the single-column tuple. */
+ tup_change = ExecFetchSlotHeapTuple(dstate->tsslot, false, &shouldFree);
+ heap_deform_tuple(tup_change, dstate->tupdesc_change, values, isnull);
+ Assert(!isnull[0]);
+
+ /* Make sure we access aligned data. */
+ change_raw = (char *) DatumGetByteaP(values[0]);
+ src = (char *) VARDATA(change_raw);
+ memcpy(&change, src, SizeOfConcurrentChange);
+
+ /* TRUNCATE change contains no tuple, so process it separately. */
+ if (change.kind == CHANGE_TRUNCATE)
+ {
+ /*
+ * All the things that ExecuteTruncateGuts() does (such as firing
+ * triggers or handling the DROP_CASCADE behavior) should have
+ * taken place on the source relation. Thus we only do the actual
+ * truncation of the new relation (and its indexes).
+ */
+ heap_truncate_one_rel(rel);
+
+ pfree(tup_change);
+ continue;
+ }
+
+ /*
+ * Extract the tuple from the change. The tuple is copied here because
+ * it might be assigned to 'tup_old', in which case it needs to
+ * survive into the next iteration.
+ */
+ tup = get_changed_tuple(src);
+
+ if (change.kind == CHANGE_UPDATE_OLD)
+ {
+ Assert(tup_old == NULL);
+ tup_old = tup;
+ }
+ else if (change.kind == CHANGE_INSERT)
+ {
+ Assert(tup_old == NULL);
+
+ apply_concurrent_insert(rel, &change, tup, iistate, index_slot);
+
+ pfree(tup);
+ }
+ else if (change.kind == CHANGE_UPDATE_NEW ||
+ change.kind == CHANGE_DELETE)
+ {
+ IndexScanDesc ind_scan = NULL;
+ HeapTuple tup_key;
+
+ if (change.kind == CHANGE_UPDATE_NEW)
+ {
+ tup_key = tup_old != NULL ? tup_old : tup;
+ }
+ else
+ {
+ Assert(tup_old == NULL);
+ tup_key = tup;
+ }
+
+ /*
+ * Find the tuple to be updated or deleted.
+ */
+ tup_exist = find_target_tuple(rel, key, nkeys, tup_key,
+ iistate, ident_slot, &ind_scan);
+ if (tup_exist == NULL)
+ elog(ERROR, "Failed to find target tuple");
+
+ if (change.kind == CHANGE_UPDATE_NEW)
+ apply_concurrent_update(rel, tup, tup_exist, &change, iistate,
+ index_slot);
+ else
+ apply_concurrent_delete(rel, tup_exist, &change);
+
+ if (tup_old != NULL)
+ {
+ pfree(tup_old);
+ tup_old = NULL;
+ }
+
+ pfree(tup);
+ index_endscan(ind_scan);
+ }
+ else
+ elog(ERROR, "Unrecognized kind of change: %d", change.kind);
+
+ /* If there's any change, make it visible to the next iteration. */
+ if (change.kind != CHANGE_UPDATE_OLD)
+ {
+ CommandCounterIncrement();
+ UpdateActiveSnapshotCommandId();
+ }
+
+ /* TTSOpsMinimalTuple has .get_heap_tuple==NULL. */
+ Assert(shouldFree);
+ pfree(tup_change);
+ }
+
+ tuplestore_clear(dstate->tstore);
+ dstate->nchanges = 0;
+
+ /* Cleanup. */
+ ExecDropSingleTupleTableSlot(index_slot);
+ ExecDropSingleTupleTableSlot(ident_slot);
+}
+
+static void
+apply_concurrent_insert(Relation rel, ConcurrentChange *change, HeapTuple tup,
+ IndexInsertState *iistate, TupleTableSlot *index_slot)
+{
+ List *recheck;
+
+
+ heap_insert(rel, tup, GetCurrentCommandId(true), HEAP_INSERT_NO_LOGICAL, NULL);
+
+ /*
+ * Update indexes.
+ *
+ * In case functions in the index need the active snapshot and caller
+ * hasn't set one.
+ */
+ ExecStoreHeapTuple(tup, index_slot, false);
+ recheck = ExecInsertIndexTuples(iistate->rri,
+ index_slot,
+ iistate->estate,
+ false, /* update */
+ false, /* noDupErr */
+ NULL, /* specConflict */
+ NIL, /* arbiterIndexes */
+ false /* onlySummarizing */
+ );
+
+ /*
+ * If recheck is required, it must have been preformed on the source
+ * relation by now. (All the logical changes we process here are already
+ * committed.)
+ */
+ list_free(recheck);
+
+ pgstat_progress_incr_param(PROGRESS_REPACK_HEAP_TUPLES_INSERTED, 1);
+}
+
+static void
+apply_concurrent_update(Relation rel, HeapTuple tup, HeapTuple tup_target,
+ ConcurrentChange *change, IndexInsertState *iistate,
+ TupleTableSlot *index_slot)
+{
+ List *recheck;
+ TU_UpdateIndexes update_indexes;
+
+ /*
+ * Write the new tuple into the new heap. ('tup' gets the TID assigned
+ * here.)
+ */
+ simple_heap_update(rel, &tup_target->t_self, tup, &update_indexes);
+
+ ExecStoreHeapTuple(tup, index_slot, false);
+
+ if (update_indexes != TU_None)
+ {
+ recheck = ExecInsertIndexTuples(iistate->rri,
+ index_slot,
+ iistate->estate,
+ true, /* update */
+ false, /* noDupErr */
+ NULL, /* specConflict */
+ NIL, /* arbiterIndexes */
+ /* onlySummarizing */
+ update_indexes == TU_Summarizing);
+ list_free(recheck);
+ }
+
+ pgstat_progress_incr_param(PROGRESS_REPACK_HEAP_TUPLES_UPDATED, 1);
+}
+
+static void
+apply_concurrent_delete(Relation rel, HeapTuple tup_target,
+ ConcurrentChange *change)
+{
+ simple_heap_delete(rel, &tup_target->t_self);
+
+ pgstat_progress_incr_param(PROGRESS_REPACK_HEAP_TUPLES_DELETED, 1);
+}
+
+/*
+ * Find the tuple to be updated or deleted.
+ *
+ * 'key' is a pre-initialized scan key, into which the function will put the
+ * key values.
+ *
+ * 'tup_key' is a tuple containing the key values for the scan.
+ *
+ * On exit,'*scan_p' contains the scan descriptor used. The caller must close
+ * it when he no longer needs the tuple returned.
+ */
+static HeapTuple
+find_target_tuple(Relation rel, ScanKey key, int nkeys, HeapTuple tup_key,
+ IndexInsertState *iistate,
+ TupleTableSlot *ident_slot, IndexScanDesc *scan_p)
+{
+ IndexScanDesc scan;
+ Form_pg_index ident_form;
+ int2vector *ident_indkey;
+ HeapTuple result = NULL;
+
+ scan = index_beginscan(rel, iistate->ident_index, GetActiveSnapshot(),
+ nkeys, 0);
+ *scan_p = scan;
+ index_rescan(scan, key, nkeys, NULL, 0);
+
+ /* Info needed to retrieve key values from heap tuple. */
+ ident_form = iistate->ident_index->rd_index;
+ ident_indkey = &ident_form->indkey;
+
+ /* Use the incoming tuple to finalize the scan key. */
+ for (int i = 0; i < scan->numberOfKeys; i++)
+ {
+ ScanKey entry;
+ bool isnull;
+ int16 attno_heap;
+
+ entry = &scan->keyData[i];
+ attno_heap = ident_indkey->values[i];
+ entry->sk_argument = heap_getattr(tup_key,
+ attno_heap,
+ rel->rd_att,
+ &isnull);
+ Assert(!isnull);
+ }
+ if (index_getnext_slot(scan, ForwardScanDirection, ident_slot))
+ {
+ bool shouldFree;
+
+ result = ExecFetchSlotHeapTuple(ident_slot, false, &shouldFree);
+ /* TTSOpsBufferHeapTuple has .get_heap_tuple != NULL. */
+ Assert(!shouldFree);
+ }
+
+ return result;
+}
+
+/*
+ * Decode and apply concurrent changes.
+ *
+ * Pass rel_src iff its reltoastrelid is needed.
+ */
+static void
+process_concurrent_changes(LogicalDecodingContext *ctx, XLogRecPtr end_of_wal,
+ Relation rel_dst, Relation rel_src, ScanKey ident_key,
+ int ident_key_nentries, IndexInsertState *iistate)
+{
+ RepackDecodingState *dstate;
+
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_CATCH_UP);
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+
+ repack_decode_concurrent_changes(ctx, end_of_wal);
+
+ if (dstate->nchanges == 0)
+ return;
+
+ PG_TRY();
+ {
+ /*
+ * Make sure that TOAST values can eventually be accessed via the old
+ * relation - see comment in copy_table_data().
+ */
+ if (rel_src)
+ rel_dst->rd_toastoid = rel_src->rd_rel->reltoastrelid;
+
+ apply_concurrent_changes(dstate, rel_dst, ident_key,
+ ident_key_nentries, iistate);
+ }
+ PG_FINALLY();
+ {
+ if (rel_src)
+ rel_dst->rd_toastoid = InvalidOid;
+ }
+ PG_END_TRY();
+}
+
+static IndexInsertState *
+get_index_insert_state(Relation relation, Oid ident_index_id)
+{
+ EState *estate;
+ int i;
+ IndexInsertState *result;
+
+ result = (IndexInsertState *) palloc0(sizeof(IndexInsertState));
+ estate = CreateExecutorState();
+ result->econtext = GetPerTupleExprContext(estate);
+
+ result->rri = (ResultRelInfo *) palloc(sizeof(ResultRelInfo));
+ InitResultRelInfo(result->rri, relation, 0, 0, 0);
+ ExecOpenIndices(result->rri, false);
+
+ /*
+ * Find the relcache entry of the identity index so that we spend no extra
+ * effort to open / close it.
+ */
+ for (i = 0; i < result->rri->ri_NumIndices; i++)
+ {
+ Relation ind_rel;
+
+ ind_rel = result->rri->ri_IndexRelationDescs[i];
+ if (ind_rel->rd_id == ident_index_id)
+ result->ident_index = ind_rel;
+ }
+ if (result->ident_index == NULL)
+ elog(ERROR, "Failed to open identity index");
+
+ /* Only initialize fields needed by ExecInsertIndexTuples(). */
+ result->estate = estate;
+
+ return result;
+}
+
+/*
+ * Build scan key to process logical changes.
+ */
+static ScanKey
+build_identity_key(Oid ident_idx_oid, Relation rel_src, int *nentries)
+{
+ Relation ident_idx_rel;
+ Form_pg_index ident_idx;
+ int n,
+ i;
+ ScanKey result;
+
+ Assert(OidIsValid(ident_idx_oid));
+ ident_idx_rel = index_open(ident_idx_oid, AccessShareLock);
+ ident_idx = ident_idx_rel->rd_index;
+ n = ident_idx->indnatts;
+ result = (ScanKey) palloc(sizeof(ScanKeyData) * n);
+ for (i = 0; i < n; i++)
+ {
+ ScanKey entry;
+ int16 relattno;
+ Form_pg_attribute att;
+ Oid opfamily,
+ opcintype,
+ opno,
+ opcode;
+
+ entry = &result[i];
+ relattno = ident_idx->indkey.values[i];
+ if (relattno >= 1)
+ {
+ TupleDesc desc;
+
+ desc = rel_src->rd_att;
+ att = TupleDescAttr(desc, relattno - 1);
+ }
+ else
+ elog(ERROR, "Unexpected attribute number %d in index", relattno);
+
+ opfamily = ident_idx_rel->rd_opfamily[i];
+ opcintype = ident_idx_rel->rd_opcintype[i];
+ opno = get_opfamily_member(opfamily, opcintype, opcintype,
+ BTEqualStrategyNumber);
+
+ if (!OidIsValid(opno))
+ elog(ERROR, "Failed to find = operator for type %u", opcintype);
+
+ opcode = get_opcode(opno);
+ if (!OidIsValid(opcode))
+ elog(ERROR, "Failed to find = operator for operator %u", opno);
+
+ /* Initialize everything but argument. */
+ ScanKeyInit(entry,
+ i + 1,
+ BTEqualStrategyNumber, opcode,
+ (Datum) NULL);
+ entry->sk_collation = att->attcollation;
+ }
+ index_close(ident_idx_rel, AccessShareLock);
+
+ *nentries = n;
+ return result;
+}
+
+static void
+free_index_insert_state(IndexInsertState *iistate)
+{
+ ExecCloseIndices(iistate->rri);
+ FreeExecutorState(iistate->estate);
+ pfree(iistate->rri);
+ pfree(iistate);
+}
+
+static void
+cleanup_logical_decoding(LogicalDecodingContext *ctx)
+{
+ RepackDecodingState *dstate;
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+
+ ExecDropSingleTupleTableSlot(dstate->tsslot);
+ FreeTupleDesc(dstate->tupdesc_change);
+ FreeTupleDesc(dstate->tupdesc);
+ tuplestore_end(dstate->tstore);
+
+ FreeDecodingContext(ctx);
+}
+
+/*
+ * The final steps of rebuild_relation() for concurrent processing.
+ *
+ * On entry, NewHeap is locked in AccessExclusiveLock mode. OldHeap and its
+ * clustering index (if one is passed) are still locked in a mode that allows
+ * concurrent data changes. On exit, both tables and their indexes are closed,
+ * but locked in AccessExclusiveLock mode.
+ */
+static void
+rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
+ Relation cl_index,
+ CatalogState *cat_state,
+ LogicalDecodingContext *ctx,
+ bool swap_toast_by_content,
+ TransactionId frozenXid,
+ MultiXactId cutoffMulti)
+{
+ LOCKMODE lockmode_old PG_USED_FOR_ASSERTS_ONLY;
+ List *ind_oids_new;
+ Oid old_table_oid = RelationGetRelid(OldHeap);
+ Oid new_table_oid = RelationGetRelid(NewHeap);
+ List *ind_oids_old = RelationGetIndexList(OldHeap);
+ ListCell *lc, *lc2;
+ char relpersistence;
+ bool is_system_catalog;
+ Oid ident_idx_old, ident_idx_new;
+ IndexInsertState *iistate;
+ ScanKey ident_key;
+ int ident_key_nentries;
+ XLogRecPtr wal_insert_ptr, end_of_wal;
+ char dummy_rec_data = '\0';
+ RelReopenInfo *rri = NULL;
+ int nrel;
+ Relation *ind_refs_all, *ind_refs_p;
+
+ /* Like in cluster_rel(). */
+ lockmode_old = ShareUpdateExclusiveLock;
+ Assert(CheckRelationLockedByMe(OldHeap, lockmode_old, false));
+ Assert(cl_index == NULL ||
+ CheckRelationLockedByMe(cl_index, lockmode_old, false));
+ /* This is expected from the caller. */
+ Assert(CheckRelationLockedByMe(NewHeap, AccessExclusiveLock, false));
+
+ ident_idx_old = RelationGetReplicaIndex(OldHeap);
+
+ /*
+ * Unlike the exclusive case, we build new indexes for the new relation
+ * rather than swapping the storage and reindexing the old relation. The
+ * point is that the index build can take some time, so we do it before we
+ * get AccessExclusiveLock on the old heap and therefore we cannot swap
+ * the heap storage yet.
+ *
+ * index_create() will lock the new indexes using AccessExclusiveLock
+ * creation - no need to change that.
+ */
+ ind_oids_new = build_new_indexes(NewHeap, OldHeap, ind_oids_old);
+
+ /*
+ * Processing shouldn't start w/o valid identity index.
+ */
+ Assert(OidIsValid(ident_idx_old));
+
+ /* Find "identity index" on the new relation. */
+ ident_idx_new = InvalidOid;
+ forboth(lc, ind_oids_old, lc2, ind_oids_new)
+ {
+ Oid ind_old = lfirst_oid(lc);
+ Oid ind_new = lfirst_oid(lc2);
+
+ if (ident_idx_old == ind_old)
+ {
+ ident_idx_new = ind_new;
+ break;
+ }
+ }
+ if (!OidIsValid(ident_idx_new))
+ /*
+ * Should not happen, given our lock on the old relation.
+ */
+ ereport(ERROR,
+ (errmsg("Identity index missing on the new relation")));
+
+ /* Executor state to update indexes. */
+ iistate = get_index_insert_state(NewHeap, ident_idx_new);
+
+ /*
+ * Build scan key that we'll use to look for rows to be updated / deleted
+ * during logical decoding.
+ */
+ ident_key = build_identity_key(ident_idx_new, OldHeap, &ident_key_nentries);
+
+ /*
+ * Flush all WAL records inserted so far (possibly except for the last
+ * incomplete page, see GetInsertRecPtr), to minimize the amount of data
+ * we need to flush while holding exclusive lock on the source table.
+ */
+ wal_insert_ptr = GetInsertRecPtr();
+ XLogFlush(wal_insert_ptr);
+ end_of_wal = GetFlushRecPtr(NULL);
+
+ /*
+ * Apply concurrent changes first time, to minimize the time we need to
+ * hold AccessExclusiveLock. (Quite some amount of WAL could have been
+ * written during the data copying and index creation.)
+ */
+ process_concurrent_changes(ctx, end_of_wal, NewHeap,
+ swap_toast_by_content ? OldHeap : NULL,
+ ident_key, ident_key_nentries, iistate);
+
+ /*
+ * Release the locks that allowed concurrent data changes, in order to
+ * acquire the AccessExclusiveLock.
+ */
+ nrel = 0;
+ /*
+ * We unlock the old relation (and its clustering index), but then we will
+ * lock the relation and *all* its indexes because we want to swap their
+ * storage.
+ *
+ * (NewHeap is already locked, as well as its indexes.)
+ */
+ rri = palloc_array(RelReopenInfo, 1 + list_length(ind_oids_old));
+ init_rel_reopen_info(&rri[nrel++], &OldHeap, InvalidOid,
+ ShareUpdateExclusiveLock, AccessExclusiveLock);
+ /* References to the re-opened indexes will be stored in this array. */
+ ind_refs_all = palloc_array(Relation, list_length(ind_oids_old));
+ ind_refs_p = ind_refs_all;
+ /* The clustering index is a special case. */
+ if (cl_index)
+ {
+ *ind_refs_p = cl_index;
+ init_rel_reopen_info(&rri[nrel], ind_refs_p, InvalidOid,
+ ShareUpdateExclusiveLock, AccessExclusiveLock);
+ nrel++;
+ ind_refs_p++;
+ }
+ /*
+ * Initialize also the entries for the other indexes (currently unlocked)
+ * because we will have to lock them.
+ */
+ foreach(lc, ind_oids_old)
+ {
+ Oid ind_oid;
+
+ ind_oid = lfirst_oid(lc);
+ /* Clustering index is already in the array, or there is none. */
+ if (cl_index && RelationGetRelid(cl_index) == ind_oid)
+ continue;
+
+ Assert(nrel < (1 + list_length(ind_oids_old)));
+
+ *ind_refs_p = NULL;
+ init_rel_reopen_info(&rri[nrel],
+ /*
+ * In this special case we do not have the
+ * relcache reference, use OID instead.
+ */
+ ind_refs_p,
+ ind_oid,
+ NoLock, /* Nothing to unlock. */
+ AccessExclusiveLock);
+
+ nrel++;
+ ind_refs_p++;
+ }
+ /* Perform the actual unlocking and re-locking. */
+ unlock_and_close_relations(rri, nrel);
+ reopen_relations(rri, nrel);
+
+ /*
+ * In addition, lock the OldHeap's TOAST relation that we skipped for the
+ * CONCURRENTLY option in copy_table_data(). This lock will be needed to
+ * swap the relation files.
+ */
+ if (OidIsValid(OldHeap->rd_rel->reltoastrelid))
+ LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
+
+ /*
+ * Check if the new indexes match the old ones, i.e. no changes occurred
+ * while OldHeap was unlocked.
+ *
+ * XXX It's probably not necessary to check the relation tuple descriptor
+ * here because the logical decoding was already active when we released
+ * the lock, and thus the corresponding data changes won't be lost.
+ * However processing of those changes might take a lot of time.
+ */
+ check_catalog_changes(OldHeap, cat_state);
+
+ /*
+ * Tuples and pages of the old heap will be gone, but the heap will stay.
+ */
+ TransferPredicateLocksToHeapRelation(OldHeap);
+ /* The same for indexes. */
+ for (int i = 0; i < (nrel - 1); i++)
+ {
+ Relation index = ind_refs_all[i];
+
+ TransferPredicateLocksToHeapRelation(index);
+
+ /*
+ * References to indexes on the old relation are not needed anymore,
+ * however locks stay till the end of the transaction.
+ */
+ index_close(index, NoLock);
+ }
+ pfree(ind_refs_all);
+
+ /*
+ * Flush anything we see in WAL, to make sure that all changes committed
+ * while we were waiting for the exclusive lock are available for
+ * decoding. This should not be necessary if all backends had
+ * synchronous_commit set, but we can't rely on this setting.
+ *
+ * Unfortunately, GetInsertRecPtr() may lag behind the actual insert
+ * position, and GetLastImportantRecPtr() points at the start of the last
+ * record rather than at the end. Thus the simplest way to determine the
+ * insert position is to insert a dummy record and use its LSN.
+ *
+ * XXX Consider using GetLastImportantRecPtr() and adding the size of the
+ * last record (plus the total size of all the page headers the record
+ * spans)?
+ */
+ XLogBeginInsert();
+ XLogRegisterData(&dummy_rec_data, 1);
+ wal_insert_ptr = XLogInsert(RM_XLOG_ID, XLOG_NOOP);
+ XLogFlush(wal_insert_ptr);
+ end_of_wal = GetFlushRecPtr(NULL);
+
+ /* Apply the concurrent changes again. */
+ process_concurrent_changes(ctx, end_of_wal, NewHeap,
+ swap_toast_by_content ? OldHeap : NULL,
+ ident_key, ident_key_nentries, iistate);
+
+ /* Remember info about rel before closing OldHeap */
+ relpersistence = OldHeap->rd_rel->relpersistence;
+ is_system_catalog = IsSystemRelation(OldHeap);
+
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_SWAP_REL_FILES);
+
+ forboth(lc, ind_oids_old, lc2, ind_oids_new)
+ {
+ Oid ind_old = lfirst_oid(lc);
+ Oid ind_new = lfirst_oid(lc2);
+ Oid mapped_tables[4];
+
+ /* Zero out possible results from swapped_relation_files */
+ memset(mapped_tables, 0, sizeof(mapped_tables));
+
+ swap_relation_files(ind_old, ind_new,
+ (old_table_oid == RelationRelationId),
+ swap_toast_by_content,
+ true,
+ InvalidTransactionId,
+ InvalidMultiXactId,
+ mapped_tables);
+
+#ifdef USE_ASSERT_CHECKING
+ /*
+ * Concurrent processing is not supported for system relations, so
+ * there should be no mapped tables.
+ */
+ for (int i = 0; i < 4; i++)
+ Assert(mapped_tables[i] == 0);
+#endif
+ }
+
+ /* The new indexes must be visible for deletion. */
+ CommandCounterIncrement();
+
+ /* Close the old heap but keep lock until transaction commit. */
+ table_close(OldHeap, NoLock);
+ /* Close the new heap. (We didn't have to open its indexes). */
+ table_close(NewHeap, NoLock);
+
+ /* Cleanup what we don't need anymore. (And close the identity index.) */
+ pfree(ident_key);
+ free_index_insert_state(iistate);
+
+ /*
+ * Swap the relations and their TOAST relations and TOAST indexes. This
+ * also drops the new relation and its indexes.
+ *
+ * (System catalogs are currently not supported.)
+ */
+ Assert(!is_system_catalog);
+ finish_heap_swap(old_table_oid, new_table_oid,
+ is_system_catalog,
+ swap_toast_by_content,
+ false, true, false,
+ frozenXid, cutoffMulti,
+ relpersistence);
+
+ pfree(rri);
+}
+
+/*
+ * Build indexes on NewHeap according to those on OldHeap.
+ *
+ * OldIndexes is the list of index OIDs on OldHeap.
+ *
+ * A list of OIDs of the corresponding indexes created on NewHeap is
+ * returned. The order of items does match, so we can use these arrays to swap
+ * index storage.
+ */
+static List *
+build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes)
+{
+ StringInfo ind_name;
+ ListCell *lc;
+ List *result = NIL;
+
+ pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+ PROGRESS_REPACK_PHASE_REBUILD_INDEX);
+
+ ind_name = makeStringInfo();
+
+ foreach(lc, OldIndexes)
+ {
+ Oid ind_oid,
+ ind_oid_new,
+ tbsp_oid;
+ Relation ind;
+ IndexInfo *ind_info;
+ int i,
+ heap_col_id;
+ List *colnames;
+ int16 indnatts;
+ Oid *collations,
+ *opclasses;
+ HeapTuple tup;
+ bool isnull;
+ Datum d;
+ oidvector *oidvec;
+ int2vector *int2vec;
+ size_t oid_arr_size;
+ size_t int2_arr_size;
+ int16 *indoptions;
+ text *reloptions = NULL;
+ bits16 flags;
+ Datum *opclassOptions;
+ NullableDatum *stattargets;
+
+ ind_oid = lfirst_oid(lc);
+ ind = index_open(ind_oid, AccessShareLock);
+ ind_info = BuildIndexInfo(ind);
+
+ tbsp_oid = ind->rd_rel->reltablespace;
+ /*
+ * Index name really doesn't matter, we'll eventually use only their
+ * storage. Just make them unique within the table.
+ */
+ resetStringInfo(ind_name);
+ appendStringInfo(ind_name, "ind_%d",
+ list_cell_number(OldIndexes, lc));
+
+ flags = 0;
+ if (ind->rd_index->indisprimary)
+ flags |= INDEX_CREATE_IS_PRIMARY;
+
+ colnames = NIL;
+ indnatts = ind->rd_index->indnatts;
+ oid_arr_size = sizeof(Oid) * indnatts;
+ int2_arr_size = sizeof(int16) * indnatts;
+
+ collations = (Oid *) palloc(oid_arr_size);
+ for (i = 0; i < indnatts; i++)
+ {
+ char *colname;
+
+ heap_col_id = ind->rd_index->indkey.values[i];
+ if (heap_col_id > 0)
+ {
+ Form_pg_attribute att;
+
+ /* Normal attribute. */
+ att = TupleDescAttr(OldHeap->rd_att, heap_col_id - 1);
+ colname = pstrdup(NameStr(att->attname));
+ collations[i] = att->attcollation;
+ }
+ else if (heap_col_id == 0)
+ {
+ HeapTuple tuple;
+ Form_pg_attribute att;
+
+ /*
+ * Expression column is not present in relcache. What we need
+ * here is an attribute of the *index* relation.
+ */
+ tuple = SearchSysCache2(ATTNUM,
+ ObjectIdGetDatum(ind_oid),
+ Int16GetDatum(i + 1));
+ if (!HeapTupleIsValid(tuple))
+ elog(ERROR,
+ "cache lookup failed for attribute %d of relation %u",
+ i + 1, ind_oid);
+ att = (Form_pg_attribute) GETSTRUCT(tuple);
+ colname = pstrdup(NameStr(att->attname));
+ collations[i] = att->attcollation;
+ ReleaseSysCache(tuple);
+ }
+ else
+ elog(ERROR, "Unexpected column number: %d",
+ heap_col_id);
+
+ colnames = lappend(colnames, colname);
+ }
+
+ /*
+ * Special effort needed for variable length attributes of
+ * Form_pg_index.
+ */
+ tup = SearchSysCache1(INDEXRELID, ObjectIdGetDatum(ind_oid));
+ if (!HeapTupleIsValid(tup))
+ elog(ERROR, "cache lookup failed for index %u", ind_oid);
+ d = SysCacheGetAttr(INDEXRELID, tup, Anum_pg_index_indclass, &isnull);
+ Assert(!isnull);
+ oidvec = (oidvector *) DatumGetPointer(d);
+ opclasses = (Oid *) palloc(oid_arr_size);
+ memcpy(opclasses, oidvec->values, oid_arr_size);
+
+ d = SysCacheGetAttr(INDEXRELID, tup, Anum_pg_index_indoption,
+ &isnull);
+ Assert(!isnull);
+ int2vec = (int2vector *) DatumGetPointer(d);
+ indoptions = (int16 *) palloc(int2_arr_size);
+ memcpy(indoptions, int2vec->values, int2_arr_size);
+ ReleaseSysCache(tup);
+
+ tup = SearchSysCache1(RELOID, ObjectIdGetDatum(ind_oid));
+ if (!HeapTupleIsValid(tup))
+ elog(ERROR, "cache lookup failed for index relation %u", ind_oid);
+ d = SysCacheGetAttr(RELOID, tup, Anum_pg_class_reloptions, &isnull);
+ reloptions = !isnull ? DatumGetTextPCopy(d) : NULL;
+ ReleaseSysCache(tup);
+
+ opclassOptions = palloc0(sizeof(Datum) * ind_info->ii_NumIndexAttrs);
+ for (i = 0; i < ind_info->ii_NumIndexAttrs; i++)
+ opclassOptions[i] = get_attoptions(ind_oid, i + 1);
+
+ stattargets = get_index_stattargets(ind_oid, ind_info);
+
+ /*
+ * Neither parentIndexRelid nor parentConstraintId needs to be passed
+ * since the new catalog entries (pg_constraint, pg_inherits) would
+ * eventually be dropped. Therefore there's no need to record valid
+ * dependency on parents.
+ */
+ ind_oid_new = index_create(NewHeap,
+ ind_name->data,
+ InvalidOid,
+ InvalidOid, /* parentIndexRelid */
+ InvalidOid, /* parentConstraintId */
+ InvalidOid,
+ ind_info,
+ colnames,
+ ind->rd_rel->relam,
+ tbsp_oid,
+ collations,
+ opclasses,
+ opclassOptions,
+ indoptions,
+ stattargets,
+ PointerGetDatum(reloptions),
+ flags, /* flags */
+ 0, /* constr_flags */
+ false, /* allow_system_table_mods */
+ false, /* is_internal */
+ NULL /* constraintId */
+ );
+ result = lappend_oid(result, ind_oid_new);
+
+ index_close(ind, AccessShareLock);
+ list_free_deep(colnames);
+ pfree(collations);
+ pfree(opclasses);
+ pfree(indoptions);
+ if (reloptions)
+ pfree(reloptions);
+ }
+
+ return result;
+}
+
+static void
+init_rel_reopen_info(RelReopenInfo *rri, Relation *rel_p, Oid relid,
+ LOCKMODE lockmode_orig, LOCKMODE lockmode_new)
+{
+ rri->rel_p = rel_p;
+ rri->relid = relid;
+ rri->lockmode_orig = lockmode_orig;
+ rri->lockmode_new = lockmode_new;
+}
+
+/*
+ * Unlock and close relations specified by items of the 'rels' array. 'nrels'
+ * is the number of items.
+ *
+ * Information needed to (re)open the relations (or to issue meaningful ERROR)
+ * is added to the array items.
+ */
+static void
+unlock_and_close_relations(RelReopenInfo *rels, int nrel)
+{
+ int i;
+ RelReopenInfo *rri;
+
+ /*
+ * First, retrieve the information that we will need for re-opening.
+ *
+ * We could close (and unlock) each relation as soon as we have gathered
+ * the related information, but then we would have to be careful not to
+ * unlock the table until we have the info on all its indexes. (Once we
+ * unlock the table, any index can be dropped, and thus we can fail to get
+ * the name we want to report if re-opening fails.) It seem simpler to
+ * separate the work into two iterations.
+ */
+ for (i = 0; i < nrel; i++)
+ {
+ Relation rel;
+
+ rri = &rels[i];
+ rel = *rri->rel_p;
+
+ if (rel)
+ {
+ Assert(CheckRelationLockedByMe(rel, rri->lockmode_orig, false));
+ Assert(!OidIsValid(rri->relid));
+
+ rri->relid = RelationGetRelid(rel);
+ rri->relkind = rel->rd_rel->relkind;
+ rri->relname = pstrdup(RelationGetRelationName(rel));
+ }
+ else
+ {
+ Assert(OidIsValid(rri->relid));
+
+ rri->relname = get_rel_name(rri->relid);
+ rri->relkind = get_rel_relkind(rri->relid);
+ }
+ }
+
+ /* Second, close the relations. */
+ for (i = 0; i < nrel; i++)
+ {
+ Relation rel;
+
+ rri = &rels[i];
+ rel = *rri->rel_p;
+
+ /* Close the relation if the caller passed one. */
+ if (rel)
+ {
+ if (rri->relkind == RELKIND_RELATION)
+ table_close(rel, rri->lockmode_orig);
+ else
+ {
+ Assert(rri->relkind == RELKIND_INDEX);
+
+ index_close(rel, rri->lockmode_orig);
+ }
+ }
+ }
+}
+
+/*
+ * Re-open the relations closed previously by unlock_and_close_relations().
+ */
+static void
+reopen_relations(RelReopenInfo *rels, int nrel)
+{
+ for (int i = 0; i < nrel; i++)
+ {
+ RelReopenInfo *rri = &rels[i];
+ Relation rel;
+
+ if (rri->relkind == RELKIND_RELATION)
+ {
+ rel = try_table_open(rri->relid, rri->lockmode_new);
+ }
+ else
+ {
+ Assert(rri->relkind == RELKIND_INDEX);
+
+ rel = try_index_open(rri->relid, rri->lockmode_new);
+ }
+
+ if (rel == NULL)
+ {
+ const char *kind_str;
+
+ kind_str = (rri->relkind == RELKIND_RELATION) ? "table" : "index";
+ ereport(ERROR,
+ (errmsg("could not open \%s \"%s\"", kind_str,
+ rri->relname),
+ errhint("The %s could have been dropped by another transaction.",
+ kind_str)));
+ }
+ *rri->rel_p = rel;
+
+ pfree(rri->relname);
+ }
+}
+
+/*
+ * REPACK is intended to be a replacement of both CLUSTER and VACUUM FULL.
+ */
+void
+repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
+{
+ ListCell *lc;
+ ClusterParams params = {0};
+ bool verbose = false;
+ Relation rel = NULL;
+ Oid indexOid = InvalidOid;
+ MemoryContext repack_context;
+ List *rtcs;
+ LOCKMODE lockmode;
+
+ /* Parse option list */
+ foreach(lc, stmt->params)
+ {
+ DefElem *opt = (DefElem *) lfirst(lc);
+
+ if (strcmp(opt->defname, "verbose") == 0)
+ verbose = defGetBoolean(opt);
+ else
+ ereport(ERROR,
+ (errcode(ERRCODE_SYNTAX_ERROR),
+ errmsg("unrecognized REPACK option \"%s\"",
+ opt->defname),
+ parser_errposition(pstate, opt->location)));
+ }
+
+ params.options =
+ (verbose ? CLUOPT_VERBOSE : 0) |
+ (stmt->concurrent ? CLUOPT_CONCURRENT : 0);
+
+ /*
+ * Determine the lock mode expected by cluster_rel().
+ *
+ * In the exclusive case, we obtain AccessExclusiveLock right away to
+ * avoid lock-upgrade hazard in the single-transaction case. In the
+ * CONCURRENTLY case, the AccessExclusiveLock will only be used at the end
+ * of processing, supposedly for very short time. Until then, we'll have
+ * to unlock the relation temporarily, so there's no lock-upgrade hazard.
+ */
+ lockmode = (params.options & CLUOPT_CONCURRENT) == 0 ?
+ AccessExclusiveLock : ShareUpdateExclusiveLock;
+
+ if (stmt->relation != NULL)
+ {
+ rel = process_single_relation(stmt->relation, stmt->indexname,
+ CLUSTER_COMMAND_REPACK, lockmode,
+ isTopLevel, ¶ms, &indexOid);
+ if (rel == NULL)
+ return;
+ }
+
+ /*
+ * By here, we know we are in a multi-table situation.
+ *
+ * Concurrent processing is currently considered rather special (e.g. in
+ * terms of resources consumed) so it is not performed in bulk.
+ */
+ if (params.options & CLUOPT_CONCURRENT)
+ {
+ if (rel != NULL)
+ {
+ Assert(rel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE);
+ ereport(ERROR,
+ (errmsg("REPACK CONCURRENTLY not supported for partitioned tables"),
+ errhint("Consider running the command for individual partitions.")));
+ }
+ else
+ ereport(ERROR,
+ (errmsg("REPACK CONCURRENTLY requires explicit table name")));
+ }
+
+ /*
+ * In order to avoid holding locks for too long, we want to process each
+ * table in its own transaction. This forces us to disallow running
+ * inside a user transaction block.
+ */
+ PreventInTransactionBlock(isTopLevel, "REPACK");
+
+ /* Also, we need a memory context to hold our list of relations */
+ repack_context = AllocSetContextCreate(PortalContext,
+ "Repack",
+ ALLOCSET_DEFAULT_SIZES);
+
+ params.options |= CLUOPT_RECHECK;
+ if (rel != NULL)
+ {
+ Oid relid;
+ bool rel_is_index;
+
+ Assert(rel->rd_rel->relkind == RELKIND_PARTITIONED_TABLE);
+ /* See the ereport() above. */
+ Assert((params.options & CLUOPT_CONCURRENT) == 0);
+
+ if (OidIsValid(indexOid))
+ {
+ relid = indexOid;
+ rel_is_index = true;
+ }
+ else
+ {
+ relid = RelationGetRelid(rel);
+ rel_is_index = false;
+ }
+ rtcs = get_tables_to_cluster_partitioned(repack_context, relid,
+ rel_is_index,
+ CLUSTER_COMMAND_REPACK);
/* close relation, releasing lock on parent table */
- table_close(rel, AccessExclusiveLock);
+ table_close(rel, lockmode);
}
else
rtcs = get_tables_to_repack(repack_context);
/* Do the job. */
- cluster_multiple_rels(rtcs, ¶ms, CLUSTER_COMMAND_REPACK);
+ cluster_multiple_rels(rtcs, ¶ms, CLUSTER_COMMAND_REPACK, lockmode,
+ isTopLevel);
+
/* Start a new transaction for the cleanup work. */
StartTransactionCommand();
@@ -1904,7 +4270,8 @@ repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
*/
static Relation
process_single_relation(RangeVar *relation, char *indexname,
- ClusterCommand cmd, ClusterParams *params,
+ ClusterCommand cmd, LOCKMODE lockmode,
+ bool isTopLevel, ClusterParams *params,
Oid *indexOid_p)
{
Relation rel;
@@ -1914,12 +4281,10 @@ process_single_relation(RangeVar *relation, char *indexname,
Oid tableOid;
/*
- * Find, lock, and check permissions on the table. We obtain
- * AccessExclusiveLock right away to avoid lock-upgrade hazard in the
- * single-transaction case.
+ * Find, lock, and check permissions on the table.
*/
tableOid = RangeVarGetRelidExtended(relation,
- AccessExclusiveLock,
+ lockmode,
0,
RangeVarCallbackMaintainsTable,
NULL);
@@ -1973,7 +4338,7 @@ process_single_relation(RangeVar *relation, char *indexname,
/* For non-partitioned tables, do what we came here to do. */
if (rel->rd_rel->relkind != RELKIND_PARTITIONED_TABLE)
{
- cluster_rel(rel, indexOid, params, cmd);
+ cluster_rel(rel, indexOid, params, cmd, isTopLevel);
/* cluster_rel closes the relation, but keeps lock */
return NULL;
diff --git a/src/backend/commands/matview.c b/src/backend/commands/matview.c
index 0bfbc5ca6d..eae34fbe6c 100644
--- a/src/backend/commands/matview.c
+++ b/src/backend/commands/matview.c
@@ -906,7 +906,7 @@ refresh_by_match_merge(Oid matviewOid, Oid tempOid, Oid relowner,
static void
refresh_by_heap_swap(Oid matviewOid, Oid OIDNewHeap, char relpersistence)
{
- finish_heap_swap(matviewOid, OIDNewHeap, false, false, true, true,
+ finish_heap_swap(matviewOid, OIDNewHeap, false, false, true, true, true,
RecentXmin, ReadNextMultiXactId(), relpersistence);
}
diff --git a/src/backend/commands/tablecmds.c b/src/backend/commands/tablecmds.c
index 901cb321c3..364f2b6a81 100644
--- a/src/backend/commands/tablecmds.c
+++ b/src/backend/commands/tablecmds.c
@@ -4527,6 +4527,16 @@ AlterTableInternal(Oid relid, List *cmds, bool recurse)
rel = relation_open(relid, lockmode);
+ /*
+ * If lockmode allows, check if REPACK CONCURRENTLY is in progress. If
+ * lockmode is too weak, cluster_rel() should detect incompatible DDLs
+ * executed by us.
+ *
+ * XXX We might skip the changes for DDLs which do not change the tuple
+ * descriptor.
+ */
+ check_for_concurrent_repack(relid, lockmode);
+
EventTriggerAlterTableRelid(relid);
ATController(NULL, rel, cmds, recurse, lockmode, NULL);
@@ -5960,6 +5970,7 @@ ATRewriteTables(AlterTableStmt *parsetree, List **wqueue, LOCKMODE lockmode,
finish_heap_swap(tab->relid, OIDNewHeap,
false, false, true,
!OidIsValid(tab->newTableSpace),
+ true,
RecentXmin,
ReadNextMultiXactId(),
persistence);
diff --git a/src/backend/commands/vacuum.c b/src/backend/commands/vacuum.c
index 59dddcd31f..30e1bb5719 100644
--- a/src/backend/commands/vacuum.c
+++ b/src/backend/commands/vacuum.c
@@ -123,7 +123,7 @@ static void vac_truncate_clog(TransactionId frozenXID,
TransactionId lastSaneFrozenXid,
MultiXactId lastSaneMinMulti);
static bool vacuum_rel(Oid relid, RangeVar *relation, VacuumParams *params,
- BufferAccessStrategy bstrategy);
+ BufferAccessStrategy bstrategy, bool isTopLevel);
static double compute_parallel_delay(void);
static VacOptValue get_vacoptval_from_boolean(DefElem *def);
static bool vac_tid_reaped(ItemPointer itemptr, void *state);
@@ -633,7 +633,8 @@ vacuum(List *relations, VacuumParams *params, BufferAccessStrategy bstrategy,
if (params->options & VACOPT_VACUUM)
{
- if (!vacuum_rel(vrel->oid, vrel->relation, params, bstrategy))
+ if (!vacuum_rel(vrel->oid, vrel->relation, params, bstrategy,
+ isTopLevel))
continue;
}
@@ -1989,7 +1990,7 @@ vac_truncate_clog(TransactionId frozenXID,
*/
static bool
vacuum_rel(Oid relid, RangeVar *relation, VacuumParams *params,
- BufferAccessStrategy bstrategy)
+ BufferAccessStrategy bstrategy, bool isTopLevel)
{
LOCKMODE lmode;
Relation rel;
@@ -2249,7 +2250,7 @@ vacuum_rel(Oid relid, RangeVar *relation, VacuumParams *params,
/* VACUUM FULL is now a variant of CLUSTER; see cluster.c */
cluster_rel(rel, InvalidOid, &cluster_params,
- CLUSTER_COMMAND_VACUUM);
+ CLUSTER_COMMAND_VACUUM, isTopLevel);
/* cluster_rel closes the relation, but keeps lock */
rel = NULL;
@@ -2295,7 +2296,8 @@ vacuum_rel(Oid relid, RangeVar *relation, VacuumParams *params,
toast_vacuum_params.options |= VACOPT_PROCESS_MAIN;
toast_vacuum_params.toast_parent = relid;
- vacuum_rel(toast_relid, NULL, &toast_vacuum_params, bstrategy);
+ vacuum_rel(toast_relid, NULL, &toast_vacuum_params, bstrategy,
+ isTopLevel);
}
/*
diff --git a/src/backend/meson.build b/src/backend/meson.build
index 2b0db21480..50aa385a58 100644
--- a/src/backend/meson.build
+++ b/src/backend/meson.build
@@ -194,5 +194,6 @@ pg_test_mod_args = pg_mod_args + {
subdir('jit/llvm')
subdir('replication/libpqwalreceiver')
subdir('replication/pgoutput')
+subdir('replication/pgoutput_repack')
subdir('snowball')
subdir('utils/mb/conversion_procs')
diff --git a/src/backend/parser/gram.y b/src/backend/parser/gram.y
index 8b4c226495..5c937b7db8 100644
--- a/src/backend/parser/gram.y
+++ b/src/backend/parser/gram.y
@@ -11874,27 +11874,30 @@ cluster_index_specification:
*
* QUERY:
* REPACK [ (options) ] [ <qualified_name> [ USING INDEX <index_name> ] ]
+ * REPACK [ (options) ] CONCURRENTLY <qualified_name> [ USING INDEX <index_name> ]
*
*****************************************************************************/
RepackStmt:
- REPACK qualified_name repack_index_specification
+ REPACK opt_concurrently qualified_name repack_index_specification
{
RepackStmt *n = makeNode(RepackStmt);
- n->relation = $2;
- n->indexname = $3;
+ n->concurrent = $2;
+ n->relation = $3;
+ n->indexname = $4;
n->params = NIL;
$$ = (Node *) n;
}
- | REPACK '(' utility_option_list ')' qualified_name repack_index_specification
+ | REPACK '(' utility_option_list ')' opt_concurrently qualified_name repack_index_specification
{
RepackStmt *n = makeNode(RepackStmt);
- n->relation = $5;
- n->indexname = $6;
n->params = $3;
+ n->concurrent = $5;
+ n->relation = $6;
+ n->indexname = $7;
$$ = (Node *) n;
}
@@ -11905,6 +11908,7 @@ RepackStmt:
n->relation = NULL;
n->indexname = NULL;
n->params = NIL;
+ n->concurrent = false;
$$ = (Node *) n;
}
@@ -11915,6 +11919,7 @@ RepackStmt:
n->relation = NULL;
n->indexname = NULL;
n->params = $3;
+ n->concurrent = false;
$$ = (Node *) n;
}
;
diff --git a/src/backend/replication/logical/decode.c b/src/backend/replication/logical/decode.c
index 24d88f368d..a6df190747 100644
--- a/src/backend/replication/logical/decode.c
+++ b/src/backend/replication/logical/decode.c
@@ -33,6 +33,7 @@
#include "access/xlogreader.h"
#include "access/xlogrecord.h"
#include "catalog/pg_control.h"
+#include "commands/cluster.h"
#include "replication/decode.h"
#include "replication/logical.h"
#include "replication/message.h"
@@ -467,6 +468,29 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
TransactionId xid = XLogRecGetXid(buf->record);
SnapBuild *builder = ctx->snapshot_builder;
+ /*
+ * Check if REPACK CONCURRENTLY is being performed by this backend. If so,
+ * only decode data changes of the table that it is processing, and the
+ * changes of its TOAST relation.
+ *
+ * (TOAST locator should not be set unless the main is.)
+ */
+ Assert(!OidIsValid(repacked_rel_toast_locator.relNumber) ||
+ OidIsValid(repacked_rel_locator.relNumber));
+
+ if (OidIsValid(repacked_rel_locator.relNumber))
+ {
+ XLogReaderState *r = buf->record;
+ RelFileLocator locator;
+
+ /* Not all records contain the block. */
+ if (XLogRecGetBlockTagExtended(r, 0, &locator, NULL, NULL, NULL) &&
+ !RelFileLocatorEquals(locator, repacked_rel_locator) &&
+ (!OidIsValid(repacked_rel_toast_locator.relNumber) ||
+ !RelFileLocatorEquals(locator, repacked_rel_toast_locator)))
+ return;
+ }
+
ReorderBufferProcessXid(ctx->reorder, xid, buf->origptr);
/*
diff --git a/src/backend/replication/logical/snapbuild.c b/src/backend/replication/logical/snapbuild.c
index 8c83ff6feb..c54a1277cc 100644
--- a/src/backend/replication/logical/snapbuild.c
+++ b/src/backend/replication/logical/snapbuild.c
@@ -486,6 +486,26 @@ SnapBuildInitialSnapshot(SnapBuild *builder)
return SnapBuildMVCCFromHistoric(snap, true);
}
+/*
+ * Build an MVCC snapshot for the initial data load performed by REPACK
+ * CONCURRENTLY command.
+ *
+ * The snapshot will only be used to scan one particular relation, which is
+ * treated like a catalog (therefore ->building_full_snapshot is not
+ * important), and the caller should already have a replication slot setup (so
+ * we do not set MyProc->xmin). XXX Do we yet need to add some restrictions?
+ */
+Snapshot
+SnapBuildInitialSnapshotForRepack(SnapBuild *builder)
+{
+ Snapshot snap;
+
+ Assert(builder->state == SNAPBUILD_CONSISTENT);
+
+ snap = SnapBuildBuildSnapshot(builder);
+ return SnapBuildMVCCFromHistoric(snap, false);
+}
+
/*
* Turn a historic MVCC snapshot into an ordinary MVCC snapshot.
*
diff --git a/src/backend/replication/pgoutput_repack/Makefile b/src/backend/replication/pgoutput_repack/Makefile
new file mode 100644
index 0000000000..4efeb713b7
--- /dev/null
+++ b/src/backend/replication/pgoutput_repack/Makefile
@@ -0,0 +1,32 @@
+#-------------------------------------------------------------------------
+#
+# Makefile--
+# Makefile for src/backend/replication/pgoutput_repack
+#
+# IDENTIFICATION
+# src/backend/replication/pgoutput_repack
+#
+#-------------------------------------------------------------------------
+
+subdir = src/backend/replication/pgoutput_repack
+top_builddir = ../../../..
+include $(top_builddir)/src/Makefile.global
+
+OBJS = \
+ $(WIN32RES) \
+ pgoutput_repack.o
+PGFILEDESC = "pgoutput_repack - logical replication output plugin for REPACK command"
+NAME = pgoutput_repack
+
+all: all-shared-lib
+
+include $(top_srcdir)/src/Makefile.shlib
+
+install: all installdirs install-lib
+
+installdirs: installdirs-lib
+
+uninstall: uninstall-lib
+
+clean distclean: clean-lib
+ rm -f $(OBJS)
diff --git a/src/backend/replication/pgoutput_repack/meson.build b/src/backend/replication/pgoutput_repack/meson.build
new file mode 100644
index 0000000000..133e865a4a
--- /dev/null
+++ b/src/backend/replication/pgoutput_repack/meson.build
@@ -0,0 +1,18 @@
+# Copyright (c) 2022-2024, PostgreSQL Global Development Group
+
+pgoutput_repack_sources = files(
+ 'pgoutput_repack.c',
+)
+
+if host_system == 'windows'
+ pgoutput_repack_sources += rc_lib_gen.process(win32ver_rc, extra_args: [
+ '--NAME', 'pgoutput_repack',
+ '--FILEDESC', 'pgoutput_repack - logical replication output plugin for REPACK command',])
+endif
+
+pgoutput_repack = shared_module('pgoutput_repack',
+ pgoutput_repack_sources,
+ kwargs: pg_mod_args,
+)
+
+backend_targets += pgoutput_repack
diff --git a/src/backend/replication/pgoutput_repack/pgoutput_repack.c b/src/backend/replication/pgoutput_repack/pgoutput_repack.c
new file mode 100644
index 0000000000..1ef9b3cbfd
--- /dev/null
+++ b/src/backend/replication/pgoutput_repack/pgoutput_repack.c
@@ -0,0 +1,286 @@
+/*-------------------------------------------------------------------------
+ *
+ * pgoutput_cluster.c
+ * Logical Replication output plugin for REPACK command
+ *
+ * Copyright (c) 2012-2024, PostgreSQL Global Development Group
+ *
+ * IDENTIFICATION
+ * src/backend/replication/pgoutput_cluster/pgoutput_cluster.c
+ *
+ *-------------------------------------------------------------------------
+ */
+#include "postgres.h"
+
+#include "access/heaptoast.h"
+#include "commands/cluster.h"
+#include "replication/snapbuild.h"
+
+PG_MODULE_MAGIC;
+
+static void plugin_startup(LogicalDecodingContext *ctx,
+ OutputPluginOptions *opt, bool is_init);
+static void plugin_shutdown(LogicalDecodingContext *ctx);
+static void plugin_begin_txn(LogicalDecodingContext *ctx,
+ ReorderBufferTXN *txn);
+static void plugin_commit_txn(LogicalDecodingContext *ctx,
+ ReorderBufferTXN *txn, XLogRecPtr commit_lsn);
+static void plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
+ Relation rel, ReorderBufferChange *change);
+static void plugin_truncate(struct LogicalDecodingContext *ctx,
+ ReorderBufferTXN *txn, int nrelations,
+ Relation relations[],
+ ReorderBufferChange *change);
+static void store_change(LogicalDecodingContext *ctx,
+ ConcurrentChangeKind kind, HeapTuple tuple);
+
+void
+_PG_output_plugin_init(OutputPluginCallbacks *cb)
+{
+ AssertVariableIsOfType(&_PG_output_plugin_init, LogicalOutputPluginInit);
+
+ cb->startup_cb = plugin_startup;
+ cb->begin_cb = plugin_begin_txn;
+ cb->change_cb = plugin_change;
+ cb->truncate_cb = plugin_truncate;
+ cb->commit_cb = plugin_commit_txn;
+ cb->shutdown_cb = plugin_shutdown;
+}
+
+
+/* initialize this plugin */
+static void
+plugin_startup(LogicalDecodingContext *ctx, OutputPluginOptions *opt,
+ bool is_init)
+{
+ ctx->output_plugin_private = NULL;
+
+ /* Probably unnecessary, as we don't use the SQL interface ... */
+ opt->output_type = OUTPUT_PLUGIN_BINARY_OUTPUT;
+
+ if (ctx->output_plugin_options != NIL)
+ {
+ ereport(ERROR,
+ (errcode(ERRCODE_INVALID_PARAMETER_VALUE),
+ errmsg("This plugin does not expect any options")));
+ }
+}
+
+static void
+plugin_shutdown(LogicalDecodingContext *ctx)
+{
+}
+
+/*
+ * As we don't release the slot during processing of particular table, there's
+ * no room for SQL interface, even for debugging purposes. Therefore we need
+ * neither OutputPluginPrepareWrite() nor OutputPluginWrite() in the plugin
+ * callbacks. (Although we might want to write custom callbacks, this API
+ * seems to be unnecessarily generic for our purposes.)
+ */
+
+/* BEGIN callback */
+static void
+plugin_begin_txn(LogicalDecodingContext *ctx, ReorderBufferTXN *txn)
+{
+}
+
+/* COMMIT callback */
+static void
+plugin_commit_txn(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
+ XLogRecPtr commit_lsn)
+{
+}
+
+/*
+ * Callback for individual changed tuples
+ */
+static void
+plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
+ Relation relation, ReorderBufferChange *change)
+{
+ RepackDecodingState *dstate;
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+
+ /* Only interested in one particular relation. */
+ if (relation->rd_id != dstate->relid)
+ return;
+
+ /* Decode entry depending on its type */
+ switch (change->action)
+ {
+ case REORDER_BUFFER_CHANGE_INSERT:
+ {
+ HeapTuple newtuple;
+
+ newtuple = change->data.tp.newtuple != NULL ?
+ change->data.tp.newtuple : NULL;
+
+ /*
+ * Identity checks in the main function should have made this
+ * impossible.
+ */
+ if (newtuple == NULL)
+ elog(ERROR, "Incomplete insert info.");
+
+ store_change(ctx, CHANGE_INSERT, newtuple);
+ }
+ break;
+ case REORDER_BUFFER_CHANGE_UPDATE:
+ {
+ HeapTuple oldtuple,
+ newtuple;
+
+ oldtuple = change->data.tp.oldtuple != NULL ?
+ change->data.tp.oldtuple : NULL;
+ newtuple = change->data.tp.newtuple != NULL ?
+ change->data.tp.newtuple : NULL;
+
+ if (newtuple == NULL)
+ elog(ERROR, "Incomplete update info.");
+
+ if (oldtuple != NULL)
+ store_change(ctx, CHANGE_UPDATE_OLD, oldtuple);
+
+ store_change(ctx, CHANGE_UPDATE_NEW, newtuple);
+ }
+ break;
+ case REORDER_BUFFER_CHANGE_DELETE:
+ {
+ HeapTuple oldtuple;
+
+ oldtuple = change->data.tp.oldtuple ?
+ change->data.tp.oldtuple : NULL;
+
+ if (oldtuple == NULL)
+ elog(ERROR, "Incomplete delete info.");
+
+ store_change(ctx, CHANGE_DELETE, oldtuple);
+ }
+ break;
+ default:
+ /* Should not come here */
+ Assert(false);
+ break;
+ }
+}
+
+static void
+plugin_truncate(struct LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
+ int nrelations, Relation relations[],
+ ReorderBufferChange *change)
+{
+ RepackDecodingState *dstate;
+ int i;
+ Relation relation = NULL;
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+
+ /* Find the relation we are processing. */
+ for (i = 0; i < nrelations; i++)
+ {
+ relation = relations[i];
+
+ if (RelationGetRelid(relation) == dstate->relid)
+ break;
+ }
+
+ /* Is this truncation of another relation? */
+ if (i == nrelations)
+ return;
+
+ store_change(ctx, CHANGE_TRUNCATE, NULL);
+}
+
+/* Store concurrent data change. */
+static void
+store_change(LogicalDecodingContext *ctx, ConcurrentChangeKind kind,
+ HeapTuple tuple)
+{
+ RepackDecodingState *dstate;
+ char *change_raw;
+ ConcurrentChange change;
+ bool flattened = false;
+ Size size;
+ Datum values[1];
+ bool isnull[1];
+ char *dst, *dst_start;
+
+ dstate = (RepackDecodingState *) ctx->output_writer_private;
+
+ size = MAXALIGN(VARHDRSZ) + SizeOfConcurrentChange;
+
+ if (tuple)
+ {
+ /*
+ * ReorderBufferCommit() stores the TOAST chunks in its private memory
+ * context and frees them after having called
+ * apply_change(). Therefore we need flat copy (including TOAST) that
+ * we eventually copy into the memory context which is available to
+ * decode_concurrent_changes().
+ */
+ if (HeapTupleHasExternal(tuple))
+ {
+ /*
+ * toast_flatten_tuple_to_datum() might be more convenient but we
+ * don't want the decompression it does.
+ */
+ tuple = toast_flatten_tuple(tuple, dstate->tupdesc);
+ flattened = true;
+ }
+
+ size += tuple->t_len;
+ }
+
+ /* XXX Isn't there any function / macro to do this? */
+ if (size >= 0x3FFFFFFF)
+ elog(ERROR, "Change is too big.");
+
+ /* Construct the change. */
+ change_raw = (char *) palloc0(size);
+ SET_VARSIZE(change_raw, size);
+ /*
+ * Since the varlena alignment might not be sufficient for the structure,
+ * set the fields in a local instance and remember where it should
+ * eventually be copied.
+ */
+ change.kind = kind;
+ dst_start = (char *) VARDATA(change_raw);
+
+ /* No other information is needed for TRUNCATE. */
+ if (change.kind == CHANGE_TRUNCATE)
+ {
+ memcpy(dst_start, &change, SizeOfConcurrentChange);
+ goto store;
+ }
+
+ /*
+ * Copy the tuple.
+ *
+ * CAUTION: change->tup_data.t_data must be fixed on retrieval!
+ */
+ memcpy(&change.tup_data, tuple, sizeof(HeapTupleData));
+ dst = dst_start + SizeOfConcurrentChange;
+ memcpy(dst, tuple->t_data, tuple->t_len);
+
+ /* The data has been copied. */
+ if (flattened)
+ pfree(tuple);
+
+store:
+ /* Copy the structure so it can be stored. */
+ memcpy(dst_start, &change, SizeOfConcurrentChange);
+
+ /* Store as tuple of 1 bytea column. */
+ values[0] = PointerGetDatum(change_raw);
+ isnull[0] = false;
+ tuplestore_putvalues(dstate->tstore, dstate->tupdesc_change,
+ values, isnull);
+
+ /* Accounting. */
+ dstate->nchanges++;
+
+ /* Cleanup. */
+ pfree(change_raw);
+}
diff --git a/src/backend/storage/ipc/ipci.c b/src/backend/storage/ipc/ipci.c
index 174eed7036..07e477d279 100644
--- a/src/backend/storage/ipc/ipci.c
+++ b/src/backend/storage/ipc/ipci.c
@@ -25,6 +25,7 @@
#include "access/xlogprefetcher.h"
#include "access/xlogrecovery.h"
#include "commands/async.h"
+#include "commands/cluster.h"
#include "miscadmin.h"
#include "pgstat.h"
#include "postmaster/autovacuum.h"
@@ -148,6 +149,7 @@ CalculateShmemSize(int *num_semaphores)
size = add_size(size, WaitEventCustomShmemSize());
size = add_size(size, InjectionPointShmemSize());
size = add_size(size, SlotSyncShmemSize());
+ size = add_size(size, RepackShmemSize());
/* include additional requested shmem from preload libraries */
size = add_size(size, total_addin_request);
@@ -340,6 +342,7 @@ CreateOrAttachShmemStructs(void)
StatsShmemInit();
WaitEventCustomShmemInit();
InjectionPointShmemInit();
+ RepackShmemInit();
}
/*
diff --git a/src/backend/tcop/utility.c b/src/backend/tcop/utility.c
index bf3ba3c2ae..4ee4c47487 100644
--- a/src/backend/tcop/utility.c
+++ b/src/backend/tcop/utility.c
@@ -1307,6 +1307,16 @@ ProcessUtilitySlow(ParseState *pstate,
lockmode = AlterTableGetLockLevel(atstmt->cmds);
relid = AlterTableLookupRelation(atstmt, lockmode);
+ /*
+ * If lockmode allows, check if REPACK CONCURRENT is in
+ * progress. If lockmode is too weak, cluster_rel() should
+ * detect incompatible DDLs executed by us.
+ *
+ * XXX We might skip the changes for DDLs which do not
+ * change the tuple descriptor.
+ */
+ check_for_concurrent_repack(relid, lockmode);
+
if (OidIsValid(relid))
{
AlterTableUtilityContext atcontext;
diff --git a/src/backend/utils/activity/backend_progress.c b/src/backend/utils/activity/backend_progress.c
index eebc968193..e2c84baba9 100644
--- a/src/backend/utils/activity/backend_progress.c
+++ b/src/backend/utils/activity/backend_progress.c
@@ -162,3 +162,19 @@ pgstat_progress_end_command(void)
beentry->st_progress.command_target = InvalidOid;
PGSTAT_END_WRITE_ACTIVITY(beentry);
}
+
+void
+pgstat_progress_restore_state(PgBackendProgress *backup)
+{
+ volatile PgBackendStatus *beentry = MyBEEntry;
+
+ if (!beentry || !pgstat_track_activities)
+ return;
+
+ PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
+ beentry->st_progress.command = backup->command;
+ beentry->st_progress.command_target = backup->command_target;
+ memcpy(MyBEEntry->st_progress.param, backup->param,
+ sizeof(beentry->st_progress.param));
+ PGSTAT_END_WRITE_ACTIVITY(beentry);
+}
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index e199f07162..5a0097d53b 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -346,6 +346,7 @@ WALSummarizer "Waiting to read or update WAL summarization state."
DSMRegistry "Waiting to read or update the dynamic shared memory registry."
InjectionPoint "Waiting to read or update information related to injection points."
SerialControl "Waiting to read or update shared <filename>pg_serial</filename> state."
+RepackedRels "Waiting to read or update information on tables being repacked concurrently."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/backend/utils/cache/inval.c b/src/backend/utils/cache/inval.c
index 700ccb6df9..cb92ddb1e3 100644
--- a/src/backend/utils/cache/inval.c
+++ b/src/backend/utils/cache/inval.c
@@ -1569,6 +1569,27 @@ CacheInvalidateRelcache(Relation relation)
databaseId, relationId);
}
+/*
+ * CacheInvalidateRelcacheImmediate
+ * Send invalidation message for the specified relation's relcache entry.
+ *
+ * Currently this is used in REPACK CONCURRENTLY, to make sure that other
+ * backends are aware that the command is being executed for the relation.
+ */
+void
+CacheInvalidateRelcacheImmediate(Relation relation)
+{
+ SharedInvalidationMessage msg;
+
+ msg.rc.id = SHAREDINVALRELCACHE_ID;
+ msg.rc.dbId = MyDatabaseId;
+ msg.rc.relId = RelationGetRelid(relation);
+ /* check AddCatcacheInvalidationMessage() for an explanation */
+ VALGRIND_MAKE_MEM_DEFINED(&msg, sizeof(msg));
+
+ SendSharedInvalidMessages(&msg, 1);
+}
+
/*
* CacheInvalidateRelcacheAll
* Register invalidation of the whole relcache at the end of command.
diff --git a/src/backend/utils/cache/relcache.c b/src/backend/utils/cache/relcache.c
index 398114373e..1273149178 100644
--- a/src/backend/utils/cache/relcache.c
+++ b/src/backend/utils/cache/relcache.c
@@ -64,6 +64,7 @@
#include "catalog/pg_type.h"
#include "catalog/schemapg.h"
#include "catalog/storage.h"
+#include "commands/cluster.h"
#include "commands/policy.h"
#include "commands/publicationcmds.h"
#include "commands/trigger.h"
@@ -1249,6 +1250,10 @@ retry:
/* make sure relation is marked as having no open file yet */
relation->rd_smgr = NULL;
+ /* Is REPACK CONCURRENTLY in progress? */
+ relation->rd_repack_concurrent =
+ is_concurrent_repack_in_progress(targetRelId);
+
/*
* now we can free the memory allocated for pg_class_tuple
*/
diff --git a/src/backend/utils/time/snapmgr.c b/src/backend/utils/time/snapmgr.c
index 42bded373b..103d1249bb 100644
--- a/src/backend/utils/time/snapmgr.c
+++ b/src/backend/utils/time/snapmgr.c
@@ -154,7 +154,6 @@ static List *exportedSnapshots = NIL;
/* Prototypes for local functions */
static void UnregisterSnapshotNoOwner(Snapshot snapshot);
-static void FreeSnapshot(Snapshot snapshot);
static void SnapshotResetXmin(void);
/* ResourceOwner callbacks to track snapshot references */
@@ -587,7 +586,7 @@ CopySnapshot(Snapshot snapshot)
* FreeSnapshot
* Free the memory associated with a snapshot.
*/
-static void
+void
FreeSnapshot(Snapshot snapshot)
{
Assert(snapshot->regd_count == 0);
diff --git a/src/bin/psql/tab-complete.in.c b/src/bin/psql/tab-complete.in.c
index 72338fffb2..9aefd06481 100644
--- a/src/bin/psql/tab-complete.in.c
+++ b/src/bin/psql/tab-complete.in.c
@@ -4910,18 +4910,26 @@ match_previous_words(int pattern_id,
}
/* REPACK */
- else if (Matches("REPACK"))
+ else if (Matches("REPACK") || Matches("REPACK", "(*)"))
+ COMPLETE_WITH_SCHEMA_QUERY_PLUS(Query_for_list_of_clusterables,
+ "CONCURRENTLY");
+ else if (Matches("REPACK", "CONCURRENTLY"))
COMPLETE_WITH_SCHEMA_QUERY(Query_for_list_of_clusterables);
- else if (Matches("REPACK", "(*)"))
+ else if (Matches("REPACK", "(*)", "CONCURRENTLY"))
COMPLETE_WITH_SCHEMA_QUERY(Query_for_list_of_clusterables);
- /* If we have REPACK <sth>, then add "USING INDEX" */
- else if (Matches("REPACK", MatchAnyExcept("(")))
+ /* If we have REPACK [ CONCURRENTLY ] <sth>, then add "USING INDEX" */
+ else if (Matches("REPACK", MatchAnyExcept("(|CONCURRENTLY")) ||
+ Matches("REPACK", "CONCURRENTLY", MatchAnyExcept("(")))
COMPLETE_WITH("USING INDEX");
- /* If we have REPACK (*) <sth>, then add "USING INDEX" */
- else if (Matches("REPACK", "(*)", MatchAny))
+ /* If we have REPACK (*) [ CONCURRENTLY ] <sth>, then add "USING INDEX" */
+ else if (Matches("REPACK", "(*)", MatchAnyExcept("CONCURRENTLY")) ||
+ Matches("REPACK", "(*)", "CONCURRENTLY", MatchAnyExcept("(")))
COMPLETE_WITH("USING INDEX");
- /* If we have REPACK <sth> USING, then add the index as well */
- else if (Matches("REPACK", MatchAny, "USING", "INDEX"))
+ /*
+ * Complete ... [ (*) ] [ CONCURRENTLY ] <sth> USING INDEX, with a list of
+ * indexes for <sth>.
+ */
+ else if (TailMatches(MatchAnyExcept("(|CONCURRENTLY"), "USING", "INDEX"))
{
set_completion_reference(prev3_wd);
COMPLETE_WITH_SCHEMA_QUERY(Query_for_index_of_table);
diff --git a/src/include/access/heapam.h b/src/include/access/heapam.h
index 1640d9c32f..bdeb2f8354 100644
--- a/src/include/access/heapam.h
+++ b/src/include/access/heapam.h
@@ -421,6 +421,10 @@ extern HTSV_Result HeapTupleSatisfiesVacuumHorizon(HeapTuple htup, Buffer buffer
TransactionId *dead_after);
extern void HeapTupleSetHintBits(HeapTupleHeader tuple, Buffer buffer,
uint16 infomask, TransactionId xid);
+extern bool HeapTupleMVCCInserted(HeapTuple htup, Snapshot snapshot,
+ Buffer buffer);
+extern bool HeapTupleMVCCNotDeleted(HeapTuple htup, Snapshot snapshot,
+ Buffer buffer);
extern bool HeapTupleHeaderIsOnlyLocked(HeapTupleHeader tuple);
extern bool HeapTupleIsSurelyDead(HeapTuple htup,
struct GlobalVisState *vistest);
diff --git a/src/include/access/tableam.h b/src/include/access/tableam.h
index 131c050c15..aa3190986a 100644
--- a/src/include/access/tableam.h
+++ b/src/include/access/tableam.h
@@ -21,6 +21,7 @@
#include "access/sdir.h"
#include "access/xact.h"
#include "executor/tuptable.h"
+#include "replication/logical.h"
#include "storage/read_stream.h"
#include "utils/rel.h"
#include "utils/snapshot.h"
@@ -630,6 +631,8 @@ typedef struct TableAmRoutine
Relation OldIndex,
bool use_sort,
TransactionId OldestXmin,
+ Snapshot snapshot,
+ LogicalDecodingContext *decoding_ctx,
TransactionId *xid_cutoff,
MultiXactId *multi_cutoff,
double *num_tuples,
@@ -1673,6 +1676,10 @@ table_relation_copy_data(Relation rel, const RelFileLocator *newrlocator)
* not needed for the relation's AM
* - *xid_cutoff - ditto
* - *multi_cutoff - ditto
+ * - snapshot - if != NULL, ignore data changes done by transactions that this
+ * (MVCC) snapshot considers still in-progress or in the future.
+ * - decoding_ctx - logical decoding context, to capture concurrent data
+ * changes.
*
* Output parameters:
* - *xid_cutoff - rel's new relfrozenxid value, may be invalid
@@ -1685,6 +1692,8 @@ table_relation_copy_for_cluster(Relation OldTable, Relation NewTable,
Relation OldIndex,
bool use_sort,
TransactionId OldestXmin,
+ Snapshot snapshot,
+ LogicalDecodingContext *decoding_ctx,
TransactionId *xid_cutoff,
MultiXactId *multi_cutoff,
double *num_tuples,
@@ -1693,6 +1702,7 @@ table_relation_copy_for_cluster(Relation OldTable, Relation NewTable,
{
OldTable->rd_tableam->relation_copy_for_cluster(OldTable, NewTable, OldIndex,
use_sort, OldestXmin,
+ snapshot, decoding_ctx,
xid_cutoff, multi_cutoff,
num_tuples, tups_vacuumed,
tups_recently_dead);
diff --git a/src/include/catalog/index.h b/src/include/catalog/index.h
index 4daa8bef5e..66431cc19e 100644
--- a/src/include/catalog/index.h
+++ b/src/include/catalog/index.h
@@ -100,6 +100,9 @@ extern Oid index_concurrently_create_copy(Relation heapRelation,
Oid tablespaceOid,
const char *newName);
+extern NullableDatum *get_index_stattargets(Oid indexid,
+ IndexInfo *indInfo);
+
extern void index_concurrently_build(Oid heapRelationId,
Oid indexRelationId);
diff --git a/src/include/commands/cluster.h b/src/include/commands/cluster.h
index c2976905e4..6fb5f5509c 100644
--- a/src/include/commands/cluster.h
+++ b/src/include/commands/cluster.h
@@ -13,10 +13,15 @@
#ifndef CLUSTER_H
#define CLUSTER_H
+#include "nodes/execnodes.h"
#include "nodes/parsenodes.h"
#include "parser/parse_node.h"
+#include "replication/logical.h"
#include "storage/lock.h"
+#include "storage/relfilelocator.h"
#include "utils/relcache.h"
+#include "utils/resowner.h"
+#include "utils/tuplestore.h"
/* flag bits for ClusterParams->options */
@@ -24,6 +29,7 @@
#define CLUOPT_RECHECK 0x02 /* recheck relation state */
#define CLUOPT_RECHECK_ISCLUSTERED 0x04 /* recheck relation state for
* indisclustered */
+#define CLUOPT_CONCURRENT 0x08 /* allow concurrent data changes */
/* options for CLUSTER */
typedef struct ClusterParams
@@ -46,14 +52,91 @@ typedef enum ClusterCommand
CLUSTER_COMMAND_VACUUM
} ClusterCommand;
+/*
+ * The following definitions are used by REPACK CONCURRENTLY.
+ */
+
+extern RelFileLocator repacked_rel_locator;
+extern RelFileLocator repacked_rel_toast_locator;
+
+typedef enum
+{
+ CHANGE_INSERT,
+ CHANGE_UPDATE_OLD,
+ CHANGE_UPDATE_NEW,
+ CHANGE_DELETE,
+ CHANGE_TRUNCATE
+} ConcurrentChangeKind;
+
+typedef struct ConcurrentChange
+{
+ /* See the enum above. */
+ ConcurrentChangeKind kind;
+
+ /*
+ * The actual tuple.
+ *
+ * The tuple data follows the ConcurrentChange structure. Before use make
+ * sure the tuple is correctly aligned (ConcurrentChange can be stored as
+ * bytea) and that tuple->t_data is fixed.
+ */
+ HeapTupleData tup_data;
+} ConcurrentChange;
+
+#define SizeOfConcurrentChange (offsetof(ConcurrentChange, tup_data) + \
+ sizeof(HeapTupleData))
+
+/*
+ * Logical decoding state.
+ *
+ * Here we store the data changes that we decode from WAL while the table
+ * contents is being copied to a new storage. Also the necessary metadata
+ * needed to apply these changes to the table is stored here.
+ */
+typedef struct RepackDecodingState
+{
+ /* The relation whose changes we're decoding. */
+ Oid relid;
+
+ /*
+ * Decoded changes are stored here. Although we try to avoid excessive
+ * batches, it can happen that the changes need to be stored to disk. The
+ * tuplestore does this transparently.
+ */
+ Tuplestorestate *tstore;
+
+ /* The current number of changes in tstore. */
+ double nchanges;
+
+ /*
+ * Descriptor to store the ConcurrentChange structure serialized (bytea).
+ * We can't store the tuple directly because tuplestore only supports
+ * minimum tuple and we may need to transfer OID system column from the
+ * output plugin. Also we need to transfer the change kind, so it's better
+ * to put everything in the structure than to use 2 tuplestores "in
+ * parallel".
+ */
+ TupleDesc tupdesc_change;
+
+ /* Tuple descriptor needed to update indexes. */
+ TupleDesc tupdesc;
+
+ /* Slot to retrieve data from tstore. */
+ TupleTableSlot *tsslot;
+
+ ResourceOwner resowner;
+} RepackDecodingState;
+
extern void cluster(ParseState *pstate, ClusterStmt *stmt, bool isTopLevel);
extern void cluster_rel(Relation OldHeap, Oid indexOid, ClusterParams *params,
- ClusterCommand cmd);
+ ClusterCommand cmd, bool isTopLevel);
extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid,
LOCKMODE lockmode,
ClusterCommand cmd);
extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal);
-
+extern void can_repack_concurrently(Relation rel);
+extern void repack_decode_concurrent_changes(LogicalDecodingContext *ctx,
+ XLogRecPtr end_of_wal);
extern Oid make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
char relpersistence, LOCKMODE lockmode);
extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
@@ -61,9 +144,15 @@ extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
bool swap_toast_by_content,
bool check_constraints,
bool is_internal,
+ bool reindex,
TransactionId frozenXid,
MultiXactId cutoffMulti,
char newrelpersistence);
+extern Size RepackShmemSize(void);
+extern void RepackShmemInit(void);
+extern bool is_concurrent_repack_in_progress(Oid relid);
+extern void check_for_concurrent_repack(Oid relid, LOCKMODE lockmode);
+
extern void repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel);
#endif /* CLUSTER_H */
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 7644267e14..6b1b1a4c1a 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -67,10 +67,12 @@
#define PROGRESS_REPACK_PHASE 1
#define PROGRESS_REPACK_INDEX_RELID 2
#define PROGRESS_REPACK_HEAP_TUPLES_SCANNED 3
-#define PROGRESS_REPACK_HEAP_TUPLES_WRITTEN 4
-#define PROGRESS_REPACK_TOTAL_HEAP_BLKS 5
-#define PROGRESS_REPACK_HEAP_BLKS_SCANNED 6
-#define PROGRESS_REPACK_INDEX_REBUILD_COUNT 7
+#define PROGRESS_REPACK_HEAP_TUPLES_INSERTED 4
+#define PROGRESS_REPACK_HEAP_TUPLES_UPDATED 5
+#define PROGRESS_REPACK_HEAP_TUPLES_DELETED 6
+#define PROGRESS_REPACK_TOTAL_HEAP_BLKS 7
+#define PROGRESS_REPACK_HEAP_BLKS_SCANNED 8
+#define PROGRESS_REPACK_INDEX_REBUILD_COUNT 9
/*
* Phases of repack (as advertised via PROGRESS_REPACK_PHASE).
@@ -83,9 +85,10 @@
#define PROGRESS_REPACK_PHASE_INDEX_SCAN_HEAP 2
#define PROGRESS_REPACK_PHASE_SORT_TUPLES 3
#define PROGRESS_REPACK_PHASE_WRITE_NEW_HEAP 4
-#define PROGRESS_REPACK_PHASE_SWAP_REL_FILES 5
-#define PROGRESS_REPACK_PHASE_REBUILD_INDEX 6
-#define PROGRESS_REPACK_PHASE_FINAL_CLEANUP 7
+#define PROGRESS_REPACK_PHASE_CATCH_UP 5
+#define PROGRESS_REPACK_PHASE_SWAP_REL_FILES 6
+#define PROGRESS_REPACK_PHASE_REBUILD_INDEX 8
+#define PROGRESS_REPACK_PHASE_FINAL_CLEANUP 8
/* Commands of PROGRESS_REPACK */
#define PROGRESS_REPACK_COMMAND_REPACK 1
diff --git a/src/include/nodes/parsenodes.h b/src/include/nodes/parsenodes.h
index 03ed0450df..d6053cab9f 100644
--- a/src/include/nodes/parsenodes.h
+++ b/src/include/nodes/parsenodes.h
@@ -3924,6 +3924,7 @@ typedef struct RepackStmt
RangeVar *relation; /* relation being repacked */
char *indexname; /* order tuples by this index */
List *params; /* list of DefElem nodes */
+ bool concurrent; /* allow concurrent access? */
} RepackStmt;
diff --git a/src/include/replication/snapbuild.h b/src/include/replication/snapbuild.h
index 6d4d2d1814..802fc4b082 100644
--- a/src/include/replication/snapbuild.h
+++ b/src/include/replication/snapbuild.h
@@ -73,6 +73,7 @@ extern void FreeSnapshotBuilder(SnapBuild *builder);
extern void SnapBuildSnapDecRefcount(Snapshot snap);
extern Snapshot SnapBuildInitialSnapshot(SnapBuild *builder);
+extern Snapshot SnapBuildInitialSnapshotForRepack(SnapBuild *builder);
extern Snapshot SnapBuildMVCCFromHistoric(Snapshot snapshot, bool in_place);
extern const char *SnapBuildExportSnapshot(SnapBuild *builder);
extern void SnapBuildClearExportedSnapshot(void);
diff --git a/src/include/storage/lockdefs.h b/src/include/storage/lockdefs.h
index 7f3ba0352f..b0d81b736d 100644
--- a/src/include/storage/lockdefs.h
+++ b/src/include/storage/lockdefs.h
@@ -36,8 +36,9 @@ typedef int LOCKMODE;
#define AccessShareLock 1 /* SELECT */
#define RowShareLock 2 /* SELECT FOR UPDATE/FOR SHARE */
#define RowExclusiveLock 3 /* INSERT, UPDATE, DELETE */
-#define ShareUpdateExclusiveLock 4 /* VACUUM (non-FULL), ANALYZE, CREATE
- * INDEX CONCURRENTLY */
+#define ShareUpdateExclusiveLock 4 /* VACUUM (non-exclusive), ANALYZE, CREATE
+ * INDEX CONCURRENTLY, REPACK
+ * CONCURRENTLY */
#define ShareLock 5 /* CREATE INDEX (WITHOUT CONCURRENTLY) */
#define ShareRowExclusiveLock 6 /* like EXCLUSIVE MODE, but allows ROW
* SHARE */
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index cf56545238..f07973b458 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -83,3 +83,4 @@ PG_LWLOCK(49, WALSummarizer)
PG_LWLOCK(50, DSMRegistry)
PG_LWLOCK(51, InjectionPoint)
PG_LWLOCK(52, SerialControl)
+PG_LWLOCK(54, RepackedRels)
diff --git a/src/include/utils/backend_progress.h b/src/include/utils/backend_progress.h
index 2f1de46d05..cea341276b 100644
--- a/src/include/utils/backend_progress.h
+++ b/src/include/utils/backend_progress.h
@@ -36,7 +36,7 @@ typedef enum ProgressCommandType
/*
* Any command which wishes can advertise that it is running by setting
- * command, command_target, and param[]. command_target should be the OID of
+ * ommand, command_target, and param[]. command_target should be the OID of
* the relation which the command targets (we assume there's just one, as this
* is meant for utility commands), but the meaning of each element in the
* param array is command-specific.
@@ -56,6 +56,7 @@ extern void pgstat_progress_parallel_incr_param(int index, int64 incr);
extern void pgstat_progress_update_multi_param(int nparam, const int *index,
const int64 *val);
extern void pgstat_progress_end_command(void);
+extern void pgstat_progress_restore_state(PgBackendProgress *backup);
#endif /* BACKEND_PROGRESS_H */
diff --git a/src/include/utils/inval.h b/src/include/utils/inval.h
index 40658ba2ff..6b2faed672 100644
--- a/src/include/utils/inval.h
+++ b/src/include/utils/inval.h
@@ -49,6 +49,8 @@ extern void CacheInvalidateCatalog(Oid catalogId);
extern void CacheInvalidateRelcache(Relation relation);
+extern void CacheInvalidateRelcacheImmediate(Relation relation);
+
extern void CacheInvalidateRelcacheAll(void);
extern void CacheInvalidateRelcacheByTuple(HeapTuple classTuple);
diff --git a/src/include/utils/rel.h b/src/include/utils/rel.h
index db3e504c3d..741b29226d 100644
--- a/src/include/utils/rel.h
+++ b/src/include/utils/rel.h
@@ -253,6 +253,9 @@ typedef struct RelationData
bool pgstat_enabled; /* should relation stats be counted */
/* use "struct" here to avoid needing to include pgstat.h: */
struct PgStat_TableStatus *pgstat_info; /* statistics collection area */
+
+ /* Is REPACK CONCURRENTLY being performed on this relation? */
+ bool rd_repack_concurrent;
} RelationData;
@@ -691,7 +694,9 @@ RelationCloseSmgr(Relation relation)
#define RelationIsAccessibleInLogicalDecoding(relation) \
(XLogLogicalInfoActive() && \
RelationNeedsWAL(relation) && \
- (IsCatalogRelation(relation) || RelationIsUsedAsCatalogTable(relation)))
+ (IsCatalogRelation(relation) || \
+ RelationIsUsedAsCatalogTable(relation) || \
+ (relation)->rd_repack_concurrent))
/*
* RelationIsLogicallyLogged
diff --git a/src/include/utils/snapmgr.h b/src/include/utils/snapmgr.h
index 147b190210..5eeabdc6c4 100644
--- a/src/include/utils/snapmgr.h
+++ b/src/include/utils/snapmgr.h
@@ -61,6 +61,8 @@ extern Snapshot GetLatestSnapshot(void);
extern void SnapshotSetCommandId(CommandId curcid);
extern Snapshot CopySnapshot(Snapshot snapshot);
+extern void FreeSnapshot(Snapshot snapshot);
+
extern Snapshot GetCatalogSnapshot(Oid relid);
extern Snapshot GetNonHistoricCatalogSnapshot(Oid relid);
extern void InvalidateCatalogSnapshot(void);
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 50d87af2fd..587c0c85b0 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -1969,17 +1969,17 @@ pg_stat_progress_cluster| SELECT s.pid,
WHEN 2 THEN 'index scanning heap'::text
WHEN 3 THEN 'sorting tuples'::text
WHEN 4 THEN 'writing new heap'::text
- WHEN 5 THEN 'swapping relation files'::text
- WHEN 6 THEN 'rebuilding index'::text
- WHEN 7 THEN 'performing final cleanup'::text
+ WHEN 6 THEN 'swapping relation files'::text
+ WHEN 7 THEN 'rebuilding index'::text
+ WHEN 8 THEN 'performing final cleanup'::text
ELSE NULL::text
END AS phase,
(s.param3)::oid AS cluster_index_relid,
s.param4 AS heap_tuples_scanned,
s.param5 AS heap_tuples_written,
- s.param6 AS heap_blks_total,
- s.param7 AS heap_blks_scanned,
- s.param8 AS index_rebuild_count
+ s.param8 AS heap_blks_total,
+ s.param9 AS heap_blks_scanned,
+ s.param10 AS index_rebuild_count
FROM (pg_stat_get_progress_info('CLUSTER'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_progress_copy| SELECT s.pid,
@@ -2055,17 +2055,20 @@ pg_stat_progress_repack| SELECT s.pid,
WHEN 2 THEN 'index scanning heap'::text
WHEN 3 THEN 'sorting tuples'::text
WHEN 4 THEN 'writing new heap'::text
- WHEN 5 THEN 'swapping relation files'::text
- WHEN 6 THEN 'rebuilding index'::text
- WHEN 7 THEN 'performing final cleanup'::text
+ WHEN 5 THEN 'catch-up'::text
+ WHEN 6 THEN 'swapping relation files'::text
+ WHEN 7 THEN 'rebuilding index'::text
+ WHEN 8 THEN 'performing final cleanup'::text
ELSE NULL::text
END AS phase,
(s.param3)::oid AS repack_index_relid,
s.param4 AS heap_tuples_scanned,
- s.param5 AS heap_tuples_written,
- s.param6 AS heap_blks_total,
- s.param7 AS heap_blks_scanned,
- s.param8 AS index_rebuild_count
+ s.param5 AS heap_tuples_inserted,
+ s.param6 AS heap_tuples_updated,
+ s.param7 AS heap_tuples_deleted,
+ s.param8 AS heap_blks_total,
+ s.param9 AS heap_blks_scanned,
+ s.param10 AS index_rebuild_count
FROM (pg_stat_get_progress_info('REPACK'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_progress_vacuum| SELECT s.pid,
--
2.43.5
[text/x-diff] v08-0005-Preserve-visibility-information-of-the-concurrent-da.patch (39.6K, ../127361.1740559688@localhost/6-v08-0005-Preserve-visibility-information-of-the-concurrent-da.patch)
download | inline diff:
From d69d88cade2d164ca019c3fd5eea8e8ff2a6eea3 Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 26 Feb 2025 09:17:20 +0100
Subject: [PATCH 5/9] Preserve visibility information of the concurrent data
changes.
As explained in the commit message of the preceding patch of the series, the
data changes done by applications while REPACK CONCURRENTLY is copying the
table contents to a new file are decoded from WAL and eventually also applied
to the new file. To reduce the complexity a little bit, the preceding patch
uses the current transaction (i.e. transaction opened by the REPACK command)
to execute those INSERT, UPDATE and DELETE commands.
However, REPACK is not expected to change visibility of tuples. Therefore,
this patch fixes the handling of the "concurrent data changes". Now the tuples
written into the new table storage have the same XID and command ID (CID) as
they had in the old storage.
Related change we do here is that the data changes (INSERT, UPDATE, DELETE) we
"replay" on the new storage are not logically decoded. First, the logical
decoding subsystem does not expect that already committed transaction is
decoded again. Second, repeated decoding would be just wasted effort.
---
src/backend/access/common/toast_internals.c | 3 +-
src/backend/access/heap/heapam.c | 73 ++++++++----
src/backend/access/heap/heapam_handler.c | 14 ++-
src/backend/access/transam/xact.c | 52 ++++++++
src/backend/commands/cluster.c | 111 ++++++++++++++++--
src/backend/replication/logical/decode.c | 76 ++++++++++--
src/backend/replication/logical/snapbuild.c | 22 ++--
.../pgoutput_repack/pgoutput_repack.c | 68 +++++++++--
src/include/access/heapam.h | 15 ++-
src/include/access/heapam_xlog.h | 2 +
src/include/access/xact.h | 2 +
src/include/commands/cluster.h | 18 +++
src/include/utils/snapshot.h | 3 +
13 files changed, 389 insertions(+), 70 deletions(-)
diff --git a/src/backend/access/common/toast_internals.c b/src/backend/access/common/toast_internals.c
index 7d8be8346c..75d889ec72 100644
--- a/src/backend/access/common/toast_internals.c
+++ b/src/backend/access/common/toast_internals.c
@@ -320,7 +320,8 @@ toast_save_datum(Relation rel, Datum value,
memcpy(VARDATA(&chunk_data), data_p, chunk_size);
toasttup = heap_form_tuple(toasttupDesc, t_values, t_isnull);
- heap_insert(toastrel, toasttup, mycid, options, NULL);
+ heap_insert(toastrel, toasttup, GetCurrentTransactionId(), mycid,
+ options, NULL);
/*
* Create the index entry. We cheat a little here by not using
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index cb856a74ee..66d21e3c9f 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -60,7 +60,8 @@ static HeapTuple heap_prepare_insert(Relation relation, HeapTuple tup,
static XLogRecPtr log_heap_update(Relation reln, Buffer oldbuf,
Buffer newbuf, HeapTuple oldtup,
HeapTuple newtup, HeapTuple old_key_tuple,
- bool all_visible_cleared, bool new_all_visible_cleared);
+ bool all_visible_cleared, bool new_all_visible_cleared,
+ bool wal_logical);
#ifdef USE_ASSERT_CHECKING
static void check_lock_if_inplace_updateable_rel(Relation relation,
ItemPointer otid,
@@ -1989,7 +1990,7 @@ ReleaseBulkInsertStatePin(BulkInsertState bistate)
/*
* heap_insert - insert tuple into a heap
*
- * The new tuple is stamped with current transaction ID and the specified
+ * The new tuple is stamped with specified transaction ID and the specified
* command ID.
*
* See table_tuple_insert for comments about most of the input flags, except
@@ -2005,15 +2006,16 @@ ReleaseBulkInsertStatePin(BulkInsertState bistate)
* reflected into *tup.
*/
void
-heap_insert(Relation relation, HeapTuple tup, CommandId cid,
- int options, BulkInsertState bistate)
+heap_insert(Relation relation, HeapTuple tup, TransactionId xid,
+ CommandId cid, int options, BulkInsertState bistate)
{
- TransactionId xid = GetCurrentTransactionId();
HeapTuple heaptup;
Buffer buffer;
Buffer vmbuffer = InvalidBuffer;
bool all_visible_cleared = false;
+ Assert(TransactionIdIsValid(xid));
+
/* Cheap, simplistic check that the tuple matches the rel's rowtype. */
Assert(HeapTupleHeaderGetNatts(tup->t_data) <=
RelationGetNumberOfAttributes(relation));
@@ -2644,7 +2646,8 @@ heap_multi_insert(Relation relation, TupleTableSlot **slots, int ntuples,
void
simple_heap_insert(Relation relation, HeapTuple tup)
{
- heap_insert(relation, tup, GetCurrentCommandId(true), 0, NULL);
+ heap_insert(relation, tup, GetCurrentTransactionId(),
+ GetCurrentCommandId(true), 0, NULL);
}
/*
@@ -2701,11 +2704,11 @@ xmax_infomask_changed(uint16 new_infomask, uint16 old_infomask)
*/
TM_Result
heap_delete(Relation relation, ItemPointer tid,
- CommandId cid, Snapshot crosscheck, bool wait,
- TM_FailureData *tmfd, bool changingPart)
+ TransactionId xid, CommandId cid, Snapshot crosscheck, bool wait,
+ TM_FailureData *tmfd, bool changingPart,
+ bool wal_logical)
{
TM_Result result;
- TransactionId xid = GetCurrentTransactionId();
ItemId lp;
HeapTupleData tp;
Page page;
@@ -2722,6 +2725,7 @@ heap_delete(Relation relation, ItemPointer tid,
bool old_key_copied = false;
Assert(ItemPointerIsValid(tid));
+ Assert(TransactionIdIsValid(xid));
/*
* Forbid this during a parallel operation, lest it allocate a combo CID.
@@ -2947,7 +2951,8 @@ l1:
* Compute replica identity tuple before entering the critical section so
* we don't PANIC upon a memory allocation failure.
*/
- old_key_tuple = ExtractReplicaIdentity(relation, &tp, true, &old_key_copied);
+ old_key_tuple = wal_logical ?
+ ExtractReplicaIdentity(relation, &tp, true, &old_key_copied) : NULL;
/*
* If this is the first possibly-multixact-able operation in the current
@@ -3015,8 +3020,12 @@ l1:
/*
* For logical decode we need combo CIDs to properly decode the
* catalog
+ *
+ * Like in heap_insert(), visibility is unchanged when called from
+ * VACUUM FULL / CLUSTER.
*/
- if (RelationIsAccessibleInLogicalDecoding(relation))
+ if (wal_logical &&
+ RelationIsAccessibleInLogicalDecoding(relation))
log_heap_new_cid(relation, &tp);
xlrec.flags = 0;
@@ -3037,6 +3046,15 @@ l1:
xlrec.flags |= XLH_DELETE_CONTAINS_OLD_KEY;
}
+ /*
+ * Unlike UPDATE, DELETE is decoded even if there is no old key, so it
+ * does not help to clear both XLH_DELETE_CONTAINS_OLD_TUPLE and
+ * XLH_DELETE_CONTAINS_OLD_KEY. Thus we need an extra flag. TODO
+ * Consider not decoding tuples w/o the old tuple/key instead.
+ */
+ if (!wal_logical)
+ xlrec.flags |= XLH_DELETE_NO_LOGICAL;
+
XLogBeginInsert();
XLogRegisterData(&xlrec, SizeOfHeapDelete);
@@ -3126,10 +3144,11 @@ simple_heap_delete(Relation relation, ItemPointer tid)
TM_Result result;
TM_FailureData tmfd;
- result = heap_delete(relation, tid,
+ result = heap_delete(relation, tid, GetCurrentTransactionId(),
GetCurrentCommandId(true), InvalidSnapshot,
true /* wait for commit */ ,
- &tmfd, false /* changingPart */ );
+ &tmfd, false, /* changingPart */
+ true /* wal_logical */);
switch (result)
{
case TM_SelfModified:
@@ -3168,12 +3187,11 @@ simple_heap_delete(Relation relation, ItemPointer tid)
*/
TM_Result
heap_update(Relation relation, ItemPointer otid, HeapTuple newtup,
- CommandId cid, Snapshot crosscheck, bool wait,
- TM_FailureData *tmfd, LockTupleMode *lockmode,
- TU_UpdateIndexes *update_indexes)
+ TransactionId xid, CommandId cid, Snapshot crosscheck,
+ bool wait, TM_FailureData *tmfd, LockTupleMode *lockmode,
+ TU_UpdateIndexes *update_indexes, bool wal_logical)
{
TM_Result result;
- TransactionId xid = GetCurrentTransactionId();
Bitmapset *hot_attrs;
Bitmapset *sum_attrs;
Bitmapset *key_attrs;
@@ -3213,6 +3231,7 @@ heap_update(Relation relation, ItemPointer otid, HeapTuple newtup,
infomask2_new_tuple;
Assert(ItemPointerIsValid(otid));
+ Assert(TransactionIdIsValid(xid));
/* Cheap, simplistic check that the tuple matches the rel's rowtype. */
Assert(HeapTupleHeaderGetNatts(newtup->t_data) <=
@@ -4050,8 +4069,12 @@ l2:
/*
* For logical decoding we need combo CIDs to properly decode the
* catalog.
+ *
+ * Like in heap_insert(), visibility is unchanged when called from
+ * VACUUM FULL / CLUSTER.
*/
- if (RelationIsAccessibleInLogicalDecoding(relation))
+ if (wal_logical &&
+ RelationIsAccessibleInLogicalDecoding(relation))
{
log_heap_new_cid(relation, &oldtup);
log_heap_new_cid(relation, heaptup);
@@ -4061,7 +4084,8 @@ l2:
newbuf, &oldtup, heaptup,
old_key_tuple,
all_visible_cleared,
- all_visible_cleared_new);
+ all_visible_cleared_new,
+ wal_logical);
if (newbuf != buffer)
{
PageSetLSN(BufferGetPage(newbuf), recptr);
@@ -4416,10 +4440,10 @@ simple_heap_update(Relation relation, ItemPointer otid, HeapTuple tup,
TM_FailureData tmfd;
LockTupleMode lockmode;
- result = heap_update(relation, otid, tup,
+ result = heap_update(relation, otid, tup, GetCurrentTransactionId(),
GetCurrentCommandId(true), InvalidSnapshot,
true /* wait for commit */ ,
- &tmfd, &lockmode, update_indexes);
+ &tmfd, &lockmode, update_indexes, true);
switch (result)
{
case TM_SelfModified:
@@ -8750,7 +8774,8 @@ static XLogRecPtr
log_heap_update(Relation reln, Buffer oldbuf,
Buffer newbuf, HeapTuple oldtup, HeapTuple newtup,
HeapTuple old_key_tuple,
- bool all_visible_cleared, bool new_all_visible_cleared)
+ bool all_visible_cleared, bool new_all_visible_cleared,
+ bool wal_logical)
{
xl_heap_update xlrec;
xl_heap_header xlhdr;
@@ -8761,10 +8786,12 @@ log_heap_update(Relation reln, Buffer oldbuf,
suffixlen = 0;
XLogRecPtr recptr;
Page page = BufferGetPage(newbuf);
- bool need_tuple_data = RelationIsLogicallyLogged(reln);
+ bool need_tuple_data;
bool init;
int bufflags;
+ need_tuple_data = RelationIsLogicallyLogged(reln) && wal_logical;
+
/* Caller should not call me on a non-WAL-logged relation */
Assert(RelationNeedsWAL(reln));
diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index b2bfd05dc9..5876a96e79 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -252,7 +252,8 @@ heapam_tuple_insert(Relation relation, TupleTableSlot *slot, CommandId cid,
tuple->t_tableOid = slot->tts_tableOid;
/* Perform the insertion, and copy the resulting ItemPointer */
- heap_insert(relation, tuple, cid, options, bistate);
+ heap_insert(relation, tuple, GetCurrentTransactionId(), cid, options,
+ bistate);
ItemPointerCopy(&tuple->t_self, &slot->tts_tid);
if (shouldFree)
@@ -275,7 +276,8 @@ heapam_tuple_insert_speculative(Relation relation, TupleTableSlot *slot,
options |= HEAP_INSERT_SPECULATIVE;
/* Perform the insertion, and copy the resulting ItemPointer */
- heap_insert(relation, tuple, cid, options, bistate);
+ heap_insert(relation, tuple, GetCurrentTransactionId(), cid, options,
+ bistate);
ItemPointerCopy(&tuple->t_self, &slot->tts_tid);
if (shouldFree)
@@ -309,7 +311,8 @@ heapam_tuple_delete(Relation relation, ItemPointer tid, CommandId cid,
* the storage itself is cleaning the dead tuples by itself, it is the
* time to call the index tuple deletion also.
*/
- return heap_delete(relation, tid, cid, crosscheck, wait, tmfd, changingPart);
+ return heap_delete(relation, tid, GetCurrentTransactionId(), cid,
+ crosscheck, wait, tmfd, changingPart, true);
}
@@ -327,8 +330,9 @@ heapam_tuple_update(Relation relation, ItemPointer otid, TupleTableSlot *slot,
slot->tts_tableOid = RelationGetRelid(relation);
tuple->t_tableOid = slot->tts_tableOid;
- result = heap_update(relation, otid, tuple, cid, crosscheck, wait,
- tmfd, lockmode, update_indexes);
+ result = heap_update(relation, otid, tuple, GetCurrentTransactionId(),
+ cid, crosscheck, wait,
+ tmfd, lockmode, update_indexes, true);
ItemPointerCopy(&tuple->t_self, &slot->tts_tid);
/*
diff --git a/src/backend/access/transam/xact.c b/src/backend/access/transam/xact.c
index 1b4f21a88d..0bfd329847 100644
--- a/src/backend/access/transam/xact.c
+++ b/src/backend/access/transam/xact.c
@@ -125,6 +125,18 @@ static FullTransactionId XactTopFullTransactionId = {InvalidTransactionId};
static int nParallelCurrentXids = 0;
static TransactionId *ParallelCurrentXids;
+/*
+ * Another case that requires TransactionIdIsCurrentTransactionId() to behave
+ * specially is when REPACK CONCURRENTLY is processing data changes made in
+ * the old storage of a table by other transactions. When applying the changes
+ * to the new storage, the backend executing the CLUSTER command needs to act
+ * on behalf on those other transactions. The transactions responsible for the
+ * changes in the old storage are stored in this array, sorted by
+ * xidComparator.
+ */
+static int nRepackCurrentXids = 0;
+static TransactionId *RepackCurrentXids = NULL;
+
/*
* Miscellaneous flag bits to record events which occur on the top level
* transaction. These flags are only persisted in MyXactFlags and are intended
@@ -971,6 +983,8 @@ TransactionIdIsCurrentTransactionId(TransactionId xid)
int low,
high;
+ Assert(nRepackCurrentXids == 0);
+
low = 0;
high = nParallelCurrentXids - 1;
while (low <= high)
@@ -990,6 +1004,21 @@ TransactionIdIsCurrentTransactionId(TransactionId xid)
return false;
}
+ /*
+ * When executing CLUSTER CONCURRENTLY, the array of current transactions
+ * is given.
+ */
+ if (nRepackCurrentXids > 0)
+ {
+ Assert(nParallelCurrentXids == 0);
+
+ return bsearch(&xid,
+ RepackCurrentXids,
+ nRepackCurrentXids,
+ sizeof(TransactionId),
+ xidComparator) != NULL;
+ }
+
/*
* We will return true for the Xid of the current subtransaction, any of
* its subcommitted children, any of its parents, or any of their
@@ -5628,6 +5657,29 @@ EndParallelWorkerTransaction(void)
CurrentTransactionState->blockState = TBLOCK_DEFAULT;
}
+/*
+ * SetRepackCurrentXids
+ * Set the XID array that TransactionIdIsCurrentTransactionId() should
+ * use.
+ */
+void
+SetRepackCurrentXids(TransactionId *xip, int xcnt)
+{
+ RepackCurrentXids = xip;
+ nRepackCurrentXids = xcnt;
+}
+
+/*
+ * ResetRepackCurrentXids
+ * Undo the effect of SetRepackCurrentXids().
+ */
+void
+ResetRepackCurrentXids(void)
+{
+ RepackCurrentXids = NULL;
+ nRepackCurrentXids = 0;
+}
+
/*
* ShowTransactionState
* Debug support
diff --git a/src/backend/commands/cluster.c b/src/backend/commands/cluster.c
index 592ff6041b..b336d760f2 100644
--- a/src/backend/commands/cluster.c
+++ b/src/backend/commands/cluster.c
@@ -209,6 +209,7 @@ static void apply_concurrent_delete(Relation rel, HeapTuple tup_target,
ConcurrentChange *change);
static HeapTuple find_target_tuple(Relation rel, ScanKey key, int nkeys,
HeapTuple tup_key,
+ Snapshot snapshot,
IndexInsertState *iistate,
TupleTableSlot *ident_slot,
IndexScanDesc *scan_p);
@@ -2960,6 +2961,9 @@ setup_logical_decoding(Oid relid, const char *slotname, TupleDesc tupdesc)
dstate->relid = relid;
dstate->tstore = tuplestore_begin_heap(false, false,
maintenance_work_mem);
+#ifdef USE_ASSERT_CHECKING
+ dstate->last_change_xid = InvalidTransactionId;
+#endif
dstate->tupdesc = tupdesc;
/* Initialize the descriptor to store the changes ... */
@@ -3115,6 +3119,7 @@ apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
tup_exist;
char *change_raw, *src;
ConcurrentChange change;
+ Snapshot snapshot;
bool isnull[1];
Datum values[1];
@@ -3183,8 +3188,30 @@ apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
/*
* Find the tuple to be updated or deleted.
+ *
+ * As the table being CLUSTERed concurrently is considered an
+ * "user catalog", new CID is WAL-logged and decoded. And since we
+ * use the same XID that the original DMLs did, the snapshot used
+ * for the logical decoding (by now converted to a non-historic
+ * MVCC snapshot) should see the tuples inserted previously into
+ * the new heap and/or updated there.
+ */
+ snapshot = change.snapshot;
+
+ /*
+ * Set what should be considered current transaction (and
+ * subtransactions) during visibility check.
+ *
+ * Note that this snapshot was created from a historic snapshot
+ * using SnapBuildMVCCFromHistoric(), which does not touch
+ * 'subxip'. Thus, unlike in a regular MVCC snapshot, the array
+ * only contains the transactions whose data changes we are
+ * applying, and its subtransactions. That's exactly what we need
+ * to check if particular xact is a "current transaction:".
*/
- tup_exist = find_target_tuple(rel, key, nkeys, tup_key,
+ SetRepackCurrentXids(snapshot->subxip, snapshot->subxcnt);
+
+ tup_exist = find_target_tuple(rel, key, nkeys, tup_key, snapshot,
iistate, ident_slot, &ind_scan);
if (tup_exist == NULL)
elog(ERROR, "Failed to find target tuple");
@@ -3195,6 +3222,8 @@ apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
else
apply_concurrent_delete(rel, tup_exist, &change);
+ ResetRepackCurrentXids();
+
if (tup_old != NULL)
{
pfree(tup_old);
@@ -3207,11 +3236,14 @@ apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
else
elog(ERROR, "Unrecognized kind of change: %d", change.kind);
- /* If there's any change, make it visible to the next iteration. */
- if (change.kind != CHANGE_UPDATE_OLD)
+ /* Free the snapshot if this is the last change that needed it. */
+ Assert(change.snapshot->active_count > 0);
+ change.snapshot->active_count--;
+ if (change.snapshot->active_count == 0)
{
- CommandCounterIncrement();
- UpdateActiveSnapshotCommandId();
+ if (change.snapshot == dstate->snapshot)
+ dstate->snapshot = NULL;
+ FreeSnapshot(change.snapshot);
}
/* TTSOpsMinimalTuple has .get_heap_tuple==NULL. */
@@ -3231,10 +3263,30 @@ static void
apply_concurrent_insert(Relation rel, ConcurrentChange *change, HeapTuple tup,
IndexInsertState *iistate, TupleTableSlot *index_slot)
{
+ Snapshot snapshot = change->snapshot;
List *recheck;
+ /*
+ * For INSERT, the visibility information is not important, but we use the
+ * snapshot to get CID. Index functions might need the whole snapshot
+ * anyway.
+ */
+ SetRepackCurrentXids(snapshot->subxip, snapshot->subxcnt);
- heap_insert(rel, tup, GetCurrentCommandId(true), HEAP_INSERT_NO_LOGICAL, NULL);
+ /*
+ * Write the tuple into the new heap.
+ *
+ * The snapshot is the one we used to decode the insert (though converted
+ * to "non-historic" MVCC snapshot), i.e. the snapshot's curcid is the
+ * tuple CID incremented by one (due to the "new CID" WAL record that got
+ * written along with the INSERT record). Thus if we want to use the
+ * original CID, we need to subtract 1 from curcid.
+ */
+ Assert(snapshot->curcid != InvalidCommandId &&
+ snapshot->curcid > FirstCommandId);
+
+ heap_insert(rel, tup, change->xid, snapshot->curcid - 1,
+ HEAP_INSERT_NO_LOGICAL, NULL);
/*
* Update indexes.
@@ -3242,6 +3294,7 @@ apply_concurrent_insert(Relation rel, ConcurrentChange *change, HeapTuple tup,
* In case functions in the index need the active snapshot and caller
* hasn't set one.
*/
+ PushActiveSnapshot(snapshot);
ExecStoreHeapTuple(tup, index_slot, false);
recheck = ExecInsertIndexTuples(iistate->rri,
index_slot,
@@ -3252,6 +3305,8 @@ apply_concurrent_insert(Relation rel, ConcurrentChange *change, HeapTuple tup,
NIL, /* arbiterIndexes */
false /* onlySummarizing */
);
+ PopActiveSnapshot();
+ ResetRepackCurrentXids();
/*
* If recheck is required, it must have been preformed on the source
@@ -3269,18 +3324,36 @@ apply_concurrent_update(Relation rel, HeapTuple tup, HeapTuple tup_target,
TupleTableSlot *index_slot)
{
List *recheck;
+ LockTupleMode lockmode;
TU_UpdateIndexes update_indexes;
+ TM_Result res;
+ Snapshot snapshot = change->snapshot;
+ TM_FailureData tmfd;
/*
* Write the new tuple into the new heap. ('tup' gets the TID assigned
* here.)
+ *
+ * Regarding CID, see the comment in apply_concurrent_insert().
*/
- simple_heap_update(rel, &tup_target->t_self, tup, &update_indexes);
+ Assert(snapshot->curcid != InvalidCommandId &&
+ snapshot->curcid > FirstCommandId);
+
+ res = heap_update(rel, &tup_target->t_self, tup,
+ change->xid, snapshot->curcid - 1,
+ InvalidSnapshot,
+ false, /* no wait - only we are doing changes */
+ &tmfd, &lockmode, &update_indexes,
+ /* wal_logical */
+ false);
+ if (res != TM_Ok)
+ ereport(ERROR, (errmsg("failed to apply concurrent UPDATE")));
ExecStoreHeapTuple(tup, index_slot, false);
if (update_indexes != TU_None)
{
+ PushActiveSnapshot(snapshot);
recheck = ExecInsertIndexTuples(iistate->rri,
index_slot,
iistate->estate,
@@ -3290,6 +3363,7 @@ apply_concurrent_update(Relation rel, HeapTuple tup, HeapTuple tup_target,
NIL, /* arbiterIndexes */
/* onlySummarizing */
update_indexes == TU_Summarizing);
+ PopActiveSnapshot();
list_free(recheck);
}
@@ -3300,7 +3374,22 @@ static void
apply_concurrent_delete(Relation rel, HeapTuple tup_target,
ConcurrentChange *change)
{
- simple_heap_delete(rel, &tup_target->t_self);
+ TM_Result res;
+ TM_FailureData tmfd;
+ Snapshot snapshot = change->snapshot;
+
+ /* Regarding CID, see the comment in apply_concurrent_insert(). */
+ Assert(snapshot->curcid != InvalidCommandId &&
+ snapshot->curcid > FirstCommandId);
+
+ res = heap_delete(rel, &tup_target->t_self, change->xid,
+ snapshot->curcid - 1, InvalidSnapshot, false,
+ &tmfd, false,
+ /* wal_logical */
+ false);
+
+ if (res != TM_Ok)
+ ereport(ERROR, (errmsg("failed to apply concurrent DELETE")));
pgstat_progress_incr_param(PROGRESS_REPACK_HEAP_TUPLES_DELETED, 1);
}
@@ -3318,7 +3407,7 @@ apply_concurrent_delete(Relation rel, HeapTuple tup_target,
*/
static HeapTuple
find_target_tuple(Relation rel, ScanKey key, int nkeys, HeapTuple tup_key,
- IndexInsertState *iistate,
+ Snapshot snapshot, IndexInsertState *iistate,
TupleTableSlot *ident_slot, IndexScanDesc *scan_p)
{
IndexScanDesc scan;
@@ -3326,7 +3415,7 @@ find_target_tuple(Relation rel, ScanKey key, int nkeys, HeapTuple tup_key,
int2vector *ident_indkey;
HeapTuple result = NULL;
- scan = index_beginscan(rel, iistate->ident_index, GetActiveSnapshot(),
+ scan = index_beginscan(rel, iistate->ident_index, snapshot,
nkeys, 0);
*scan_p = scan;
index_rescan(scan, key, nkeys, NULL, 0);
@@ -3398,6 +3487,8 @@ process_concurrent_changes(LogicalDecodingContext *ctx, XLogRecPtr end_of_wal,
}
PG_FINALLY();
{
+ ResetRepackCurrentXids();
+
if (rel_src)
rel_dst->rd_toastoid = InvalidOid;
}
diff --git a/src/backend/replication/logical/decode.c b/src/backend/replication/logical/decode.c
index a6df190747..55abda75d1 100644
--- a/src/backend/replication/logical/decode.c
+++ b/src/backend/replication/logical/decode.c
@@ -469,9 +469,18 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
SnapBuild *builder = ctx->snapshot_builder;
/*
- * Check if REPACK CONCURRENTLY is being performed by this backend. If so,
- * only decode data changes of the table that it is processing, and the
- * changes of its TOAST relation.
+ * If the change is not intended for logical decoding, do not even
+ * establish transaction for it. This is particularly important if the
+ * record was generated by CLUSTER CONCURRENTLY because this command uses
+ * the original XID when doing changes in the new storage. The decoding
+ * subsystem probably does not expect to see the same transaction multiple
+ * times.
+ */
+
+ /*
+ * First, check if REPACK CONCURRENTLY is being performed by this
+ * backend. If so, only decode data changes of the table that it is
+ * processing, and the changes of its TOAST relation.
*
* (TOAST locator should not be set unless the main is.)
*/
@@ -491,6 +500,60 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
return;
}
+ /*
+ * Second, skip records which do not contain sufficient information for
+ * the decoding.
+ *
+ * The backend executing CLUSTER CONCURRENTLY should not return here
+ * because the records which passed the checks above should contain be
+ * eligible for decoding. However, CLUSTER CONCURRENTLY generates WAL when
+ * writing data into the new table, which should not be decoded by the
+ * other backends. This is where the other backends skip them.
+ */
+ switch (info)
+ {
+ case XLOG_HEAP_INSERT:
+ {
+ xl_heap_insert *rec;
+
+ rec = (xl_heap_insert *) XLogRecGetData(buf->record);
+ /*
+ * (Besides insertion into the main heap by CLUSTER CONCURRENTLY,
+ * this does happen when raw_heap_insert marks the TOAST record as
+ * HEAP_INSERT_NO_LOGICAL).
+ */
+ if ((rec->flags & XLH_INSERT_CONTAINS_NEW_TUPLE) == 0)
+ return;
+
+ break;
+ }
+
+ case XLOG_HEAP_HOT_UPDATE:
+ case XLOG_HEAP_UPDATE:
+ {
+ xl_heap_update *rec;
+
+ rec = (xl_heap_update *) XLogRecGetData(buf->record);
+ if ((rec->flags &
+ (XLH_UPDATE_CONTAINS_NEW_TUPLE |
+ XLH_UPDATE_CONTAINS_OLD_TUPLE |
+ XLH_UPDATE_CONTAINS_OLD_KEY)) == 0)
+ return;
+
+ break;
+ }
+
+ case XLOG_HEAP_DELETE:
+ {
+ xl_heap_delete *rec;
+
+ rec = (xl_heap_delete *) XLogRecGetData(buf->record);
+ if (rec->flags & XLH_DELETE_NO_LOGICAL)
+ return;
+ break;
+ }
+ }
+
ReorderBufferProcessXid(ctx->reorder, xid, buf->origptr);
/*
@@ -923,13 +986,6 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
xlrec = (xl_heap_insert *) XLogRecGetData(r);
- /*
- * Ignore insert records without new tuples (this does happen when
- * raw_heap_insert marks the TOAST record as HEAP_INSERT_NO_LOGICAL).
- */
- if (!(xlrec->flags & XLH_INSERT_CONTAINS_NEW_TUPLE))
- return;
-
/* only interested in our database */
XLogRecGetBlockTag(r, 0, &target_locator, NULL, NULL);
if (target_locator.dbOid != ctx->slot->data.database)
diff --git a/src/backend/replication/logical/snapbuild.c b/src/backend/replication/logical/snapbuild.c
index c54a1277cc..554fe83f4b 100644
--- a/src/backend/replication/logical/snapbuild.c
+++ b/src/backend/replication/logical/snapbuild.c
@@ -155,7 +155,7 @@ static bool ExportInProgress = false;
static void SnapBuildPurgeOlderTxn(SnapBuild *builder);
/* snapshot building/manipulation/distribution functions */
-static Snapshot SnapBuildBuildSnapshot(SnapBuild *builder);
+static Snapshot SnapBuildBuildSnapshot(SnapBuild *builder, XLogRecPtr lsn);
static void SnapBuildFreeSnapshot(Snapshot snap);
@@ -352,12 +352,17 @@ SnapBuildSnapDecRefcount(Snapshot snap)
* Build a new snapshot, based on currently committed catalog-modifying
* transactions.
*
+ * 'lsn' is the location of the commit record (of a catalog-changing
+ * transaction) that triggered creation of the snapshot. Pass
+ * InvalidXLogRecPtr for the transaction base snapshot or if it the user of
+ * the snapshot should not need the LSN.
+ *
* In-progress transactions with catalog access are *not* allowed to modify
* these snapshots; they have to copy them and fill in appropriate ->curcid
* and ->subxip/subxcnt values.
*/
static Snapshot
-SnapBuildBuildSnapshot(SnapBuild *builder)
+SnapBuildBuildSnapshot(SnapBuild *builder, XLogRecPtr lsn)
{
Snapshot snapshot;
Size ssize;
@@ -425,6 +430,7 @@ SnapBuildBuildSnapshot(SnapBuild *builder)
snapshot->active_count = 0;
snapshot->regd_count = 0;
snapshot->snapXactCompletionCount = 0;
+ snapshot->lsn = lsn;
return snapshot;
}
@@ -461,7 +467,7 @@ SnapBuildInitialSnapshot(SnapBuild *builder)
if (TransactionIdIsValid(MyProc->xmin))
elog(ERROR, "cannot build an initial slot snapshot when MyProc->xmin already is valid");
- snap = SnapBuildBuildSnapshot(builder);
+ snap = SnapBuildBuildSnapshot(builder, InvalidXLogRecPtr);
/*
* We know that snap->xmin is alive, enforced by the logical xmin
@@ -502,7 +508,7 @@ SnapBuildInitialSnapshotForRepack(SnapBuild *builder)
Assert(builder->state == SNAPBUILD_CONSISTENT);
- snap = SnapBuildBuildSnapshot(builder);
+ snap = SnapBuildBuildSnapshot(builder, InvalidXLogRecPtr);
return SnapBuildMVCCFromHistoric(snap, false);
}
@@ -636,7 +642,7 @@ SnapBuildGetOrBuildSnapshot(SnapBuild *builder)
/* only build a new snapshot if we don't have a prebuilt one */
if (builder->snapshot == NULL)
{
- builder->snapshot = SnapBuildBuildSnapshot(builder);
+ builder->snapshot = SnapBuildBuildSnapshot(builder, InvalidXLogRecPtr);
/* increase refcount for the snapshot builder */
SnapBuildSnapIncRefcount(builder->snapshot);
}
@@ -716,7 +722,7 @@ SnapBuildProcessChange(SnapBuild *builder, TransactionId xid, XLogRecPtr lsn)
/* only build a new snapshot if we don't have a prebuilt one */
if (builder->snapshot == NULL)
{
- builder->snapshot = SnapBuildBuildSnapshot(builder);
+ builder->snapshot = SnapBuildBuildSnapshot(builder, lsn);
/* increase refcount for the snapshot builder */
SnapBuildSnapIncRefcount(builder->snapshot);
}
@@ -1085,7 +1091,7 @@ SnapBuildCommitTxn(SnapBuild *builder, XLogRecPtr lsn, TransactionId xid,
if (builder->snapshot)
SnapBuildSnapDecRefcount(builder->snapshot);
- builder->snapshot = SnapBuildBuildSnapshot(builder);
+ builder->snapshot = SnapBuildBuildSnapshot(builder, lsn);
/* we might need to execute invalidations, add snapshot */
if (!ReorderBufferXidHasBaseSnapshot(builder->reorder, xid))
@@ -1910,7 +1916,7 @@ SnapBuildRestore(SnapBuild *builder, XLogRecPtr lsn)
{
SnapBuildSnapDecRefcount(builder->snapshot);
}
- builder->snapshot = SnapBuildBuildSnapshot(builder);
+ builder->snapshot = SnapBuildBuildSnapshot(builder, InvalidXLogRecPtr);
SnapBuildSnapIncRefcount(builder->snapshot);
ReorderBufferSetRestartPoint(builder->reorder, lsn);
diff --git a/src/backend/replication/pgoutput_repack/pgoutput_repack.c b/src/backend/replication/pgoutput_repack/pgoutput_repack.c
index 1ef9b3cbfd..d42d93a8b6 100644
--- a/src/backend/replication/pgoutput_repack/pgoutput_repack.c
+++ b/src/backend/replication/pgoutput_repack/pgoutput_repack.c
@@ -32,7 +32,8 @@ static void plugin_truncate(struct LogicalDecodingContext *ctx,
Relation relations[],
ReorderBufferChange *change);
static void store_change(LogicalDecodingContext *ctx,
- ConcurrentChangeKind kind, HeapTuple tuple);
+ ConcurrentChangeKind kind, HeapTuple tuple,
+ TransactionId xid);
void
_PG_output_plugin_init(OutputPluginCallbacks *cb)
@@ -100,6 +101,7 @@ plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
Relation relation, ReorderBufferChange *change)
{
RepackDecodingState *dstate;
+ Snapshot snapshot;
dstate = (RepackDecodingState *) ctx->output_writer_private;
@@ -107,6 +109,48 @@ plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
if (relation->rd_id != dstate->relid)
return;
+ /*
+ * Catalog snapshot is fine because the table we are processing is
+ * temporarily considered a user catalog table.
+ */
+ snapshot = GetCatalogSnapshot(InvalidOid);
+ Assert(snapshot->snapshot_type == SNAPSHOT_HISTORIC_MVCC);
+ Assert(!snapshot->suboverflowed);
+
+ /*
+ * This should not happen, but if we don't have enough information to
+ * apply a new snapshot, the consequences would be bad. Thus prefer ERROR
+ * to Assert().
+ */
+ if (XLogRecPtrIsInvalid(snapshot->lsn))
+ ereport(ERROR, (errmsg("snapshot has invalid LSN")));
+
+ /*
+ * reorderbuffer.c changes the catalog snapshot as soon as it sees a new
+ * CID or a commit record of a catalog-changing transaction.
+ */
+ if (dstate->snapshot == NULL || snapshot->lsn != dstate->snapshot_lsn ||
+ snapshot->curcid != dstate->snapshot->curcid)
+ {
+ /* CID should not go backwards. */
+ Assert(dstate->snapshot == NULL ||
+ snapshot->curcid >= dstate->snapshot->curcid ||
+ change->txn->xid != dstate->last_change_xid);
+
+ /*
+ * XXX Is it a problem that the copy is created in
+ * TopTransactionContext?
+ *
+ * XXX Wouldn't it be o.k. for SnapBuildMVCCFromHistoric() to set xcnt
+ * to 0 instead of converting xip in this case? The point is that
+ * transactions which are still in progress from the perspective of
+ * reorderbuffer.c could not be replayed yet, so we do not need to
+ * examine their XIDs.
+ */
+ dstate->snapshot = SnapBuildMVCCFromHistoric(snapshot, false);
+ dstate->snapshot_lsn = snapshot->lsn;
+ }
+
/* Decode entry depending on its type */
switch (change->action)
{
@@ -124,7 +168,7 @@ plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
if (newtuple == NULL)
elog(ERROR, "Incomplete insert info.");
- store_change(ctx, CHANGE_INSERT, newtuple);
+ store_change(ctx, CHANGE_INSERT, newtuple, change->txn->xid);
}
break;
case REORDER_BUFFER_CHANGE_UPDATE:
@@ -141,9 +185,11 @@ plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
elog(ERROR, "Incomplete update info.");
if (oldtuple != NULL)
- store_change(ctx, CHANGE_UPDATE_OLD, oldtuple);
+ store_change(ctx, CHANGE_UPDATE_OLD, oldtuple,
+ change->txn->xid);
- store_change(ctx, CHANGE_UPDATE_NEW, newtuple);
+ store_change(ctx, CHANGE_UPDATE_NEW, newtuple,
+ change->txn->xid);
}
break;
case REORDER_BUFFER_CHANGE_DELETE:
@@ -156,7 +202,7 @@ plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
if (oldtuple == NULL)
elog(ERROR, "Incomplete delete info.");
- store_change(ctx, CHANGE_DELETE, oldtuple);
+ store_change(ctx, CHANGE_DELETE, oldtuple, change->txn->xid);
}
break;
default:
@@ -190,13 +236,13 @@ plugin_truncate(struct LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
if (i == nrelations)
return;
- store_change(ctx, CHANGE_TRUNCATE, NULL);
+ store_change(ctx, CHANGE_TRUNCATE, NULL, InvalidTransactionId);
}
/* Store concurrent data change. */
static void
store_change(LogicalDecodingContext *ctx, ConcurrentChangeKind kind,
- HeapTuple tuple)
+ HeapTuple tuple, TransactionId xid)
{
RepackDecodingState *dstate;
char *change_raw;
@@ -264,6 +310,11 @@ store_change(LogicalDecodingContext *ctx, ConcurrentChangeKind kind,
dst = dst_start + SizeOfConcurrentChange;
memcpy(dst, tuple->t_data, tuple->t_len);
+ /* Initialize the other fields. */
+ change.xid = xid;
+ change.snapshot = dstate->snapshot;
+ dstate->snapshot->active_count++;
+
/* The data has been copied. */
if (flattened)
pfree(tuple);
@@ -277,6 +328,9 @@ store:
isnull[0] = false;
tuplestore_putvalues(dstate->tstore, dstate->tupdesc_change,
values, isnull);
+#ifdef USE_ASSERT_CHECKING
+ dstate->last_change_xid = xid;
+#endif
/* Accounting. */
dstate->nchanges++;
diff --git a/src/include/access/heapam.h b/src/include/access/heapam.h
index bdeb2f8354..b0c6f1d916 100644
--- a/src/include/access/heapam.h
+++ b/src/include/access/heapam.h
@@ -325,21 +325,24 @@ extern BulkInsertState GetBulkInsertState(void);
extern void FreeBulkInsertState(BulkInsertState);
extern void ReleaseBulkInsertStatePin(BulkInsertState bistate);
-extern void heap_insert(Relation relation, HeapTuple tup, CommandId cid,
- int options, BulkInsertState bistate);
+extern void heap_insert(Relation relation, HeapTuple tup, TransactionId xid,
+ CommandId cid, int options, BulkInsertState bistate);
extern void heap_multi_insert(Relation relation, struct TupleTableSlot **slots,
int ntuples, CommandId cid, int options,
BulkInsertState bistate);
extern TM_Result heap_delete(Relation relation, ItemPointer tid,
- CommandId cid, Snapshot crosscheck, bool wait,
- struct TM_FailureData *tmfd, bool changingPart);
+ TransactionId xid, CommandId cid,
+ Snapshot crosscheck, bool wait,
+ struct TM_FailureData *tmfd, bool changingPart,
+ bool wal_logical);
extern void heap_finish_speculative(Relation relation, ItemPointer tid);
extern void heap_abort_speculative(Relation relation, ItemPointer tid);
extern TM_Result heap_update(Relation relation, ItemPointer otid,
- HeapTuple newtup,
+ HeapTuple newtup, TransactionId xid,
CommandId cid, Snapshot crosscheck, bool wait,
struct TM_FailureData *tmfd, LockTupleMode *lockmode,
- TU_UpdateIndexes *update_indexes);
+ TU_UpdateIndexes *update_indexes,
+ bool wal_logical);
extern TM_Result heap_lock_tuple(Relation relation, HeapTuple tuple,
CommandId cid, LockTupleMode mode, LockWaitPolicy wait_policy,
bool follow_updates,
diff --git a/src/include/access/heapam_xlog.h b/src/include/access/heapam_xlog.h
index 277df6b3cf..8d4af07f84 100644
--- a/src/include/access/heapam_xlog.h
+++ b/src/include/access/heapam_xlog.h
@@ -104,6 +104,8 @@
#define XLH_DELETE_CONTAINS_OLD_KEY (1<<2)
#define XLH_DELETE_IS_SUPER (1<<3)
#define XLH_DELETE_IS_PARTITION_MOVE (1<<4)
+/* See heap_delete() */
+#define XLH_DELETE_NO_LOGICAL (1<<5)
/* convenience macro for checking whether any form of old tuple was logged */
#define XLH_DELETE_CONTAINS_OLD \
diff --git a/src/include/access/xact.h b/src/include/access/xact.h
index b2bc10ee04..fbb66d559b 100644
--- a/src/include/access/xact.h
+++ b/src/include/access/xact.h
@@ -482,6 +482,8 @@ extern Size EstimateTransactionStateSpace(void);
extern void SerializeTransactionState(Size maxsize, char *start_address);
extern void StartParallelWorkerTransaction(char *tstatespace);
extern void EndParallelWorkerTransaction(void);
+extern void SetRepackCurrentXids(TransactionId *xip, int xcnt);
+extern void ResetRepackCurrentXids(void);
extern bool IsTransactionBlock(void);
extern bool IsTransactionOrTransactionBlock(void);
extern char TransactionBlockStatusCode(void);
diff --git a/src/include/commands/cluster.h b/src/include/commands/cluster.h
index 6fb5f5509c..ef3cb55751 100644
--- a/src/include/commands/cluster.h
+++ b/src/include/commands/cluster.h
@@ -73,6 +73,14 @@ typedef struct ConcurrentChange
/* See the enum above. */
ConcurrentChangeKind kind;
+ /* Transaction that changes the data. */
+ TransactionId xid;
+
+ /*
+ * Historic catalog snapshot that was used to decode this change.
+ */
+ Snapshot snapshot;
+
/*
* The actual tuple.
*
@@ -104,6 +112,8 @@ typedef struct RepackDecodingState
* tuplestore does this transparently.
*/
Tuplestorestate *tstore;
+ /* XID of the last change added to tstore. */
+ TransactionId last_change_xid PG_USED_FOR_ASSERTS_ONLY;
/* The current number of changes in tstore. */
double nchanges;
@@ -124,6 +134,14 @@ typedef struct RepackDecodingState
/* Slot to retrieve data from tstore. */
TupleTableSlot *tsslot;
+ /*
+ * Historic catalog snapshot that was used to decode the most recent
+ * change.
+ */
+ Snapshot snapshot;
+ /* LSN of the record */
+ XLogRecPtr snapshot_lsn;
+
ResourceOwner resowner;
} RepackDecodingState;
diff --git a/src/include/utils/snapshot.h b/src/include/utils/snapshot.h
index 0e546ec149..014f27db7d 100644
--- a/src/include/utils/snapshot.h
+++ b/src/include/utils/snapshot.h
@@ -13,6 +13,7 @@
#ifndef SNAPSHOT_H
#define SNAPSHOT_H
+#include "access/xlogdefs.h"
#include "lib/pairingheap.h"
@@ -201,6 +202,8 @@ typedef struct SnapshotData
uint32 regd_count; /* refcount on RegisteredSnapshots */
pairingheap_node ph_node; /* link in the RegisteredSnapshots heap */
+ XLogRecPtr lsn; /* position in the WAL stream when taken */
+
/*
* The transaction completion count at the time GetSnapshotData() built
* this snapshot. Allows to avoid re-computing static snapshots when no
--
2.43.5
[text/x-diff] v08-0006-Add-regression-tests.patch (10.6K, ../127361.1740559688@localhost/7-v08-0006-Add-regression-tests.patch)
download | inline diff:
From 2d84095db2a68290e0e462ed7e84759eeed3568f Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 26 Feb 2025 09:17:20 +0100
Subject: [PATCH 6/9] Add regression tests.
As this patch series adds the CONCURRENTLY option to the REPACK command, it's
appropriate to test that the "concurrent data changes" (i.e. changes done by
application while we are copying the table contents to the new storage) are
processed correctly.
Injection points are used to stop the data copying at some point. While the
backend in charge of the copying is waiting on the injection point, another
backend runs some INSERT, UPDATE and DELETE commands on the table. Then we
wake up the first backend and let the REPACK CONCURRENTLY command
finish. Finally we check that all the "concurrent data changes" are present in
the table and that they contain the correct visibility information.
---
src/backend/commands/cluster.c | 7 +
src/test/modules/injection_points/Makefile | 3 +-
.../injection_points/expected/repack.out | 113 ++++++++++++++
.../modules/injection_points/logical.conf | 1 +
src/test/modules/injection_points/meson.build | 4 +
.../injection_points/specs/repack.spec | 140 ++++++++++++++++++
6 files changed, 267 insertions(+), 1 deletion(-)
create mode 100644 src/test/modules/injection_points/expected/repack.out
create mode 100644 src/test/modules/injection_points/logical.conf
create mode 100644 src/test/modules/injection_points/specs/repack.spec
diff --git a/src/backend/commands/cluster.c b/src/backend/commands/cluster.c
index b336d760f2..1ff0fcd1d9 100644
--- a/src/backend/commands/cluster.c
+++ b/src/backend/commands/cluster.c
@@ -59,6 +59,7 @@
#include "utils/formatting.h"
#include "utils/fmgroids.h"
#include "utils/guc.h"
+#include "utils/injection_point.h"
#include "utils/inval.h"
#include "utils/lsyscache.h"
#include "utils/memutils.h"
@@ -3710,6 +3711,12 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
*/
ident_key = build_identity_key(ident_idx_new, OldHeap, &ident_key_nentries);
+ /*
+ * During testing, wait for another backend to perform concurrent data
+ * changes which we will process below.
+ */
+ INJECTION_POINT("repack-concurrently-before-lock");
+
/*
* Flush all WAL records inserted so far (possibly except for the last
* incomplete page, see GetInsertRecPtr), to minimize the amount of data
diff --git a/src/test/modules/injection_points/Makefile b/src/test/modules/injection_points/Makefile
index e680991f8d..405d0811b4 100644
--- a/src/test/modules/injection_points/Makefile
+++ b/src/test/modules/injection_points/Makefile
@@ -14,7 +14,8 @@ PGFILEDESC = "injection_points - facility for injection points"
REGRESS = injection_points hashagg reindex_conc
REGRESS_OPTS = --dlpath=$(top_builddir)/src/test/regress
-ISOLATION = basic inplace syscache-update-pruned
+ISOLATION = basic inplace syscache-update-pruned repack
+ISOLATION_OPTS = --temp-config $(top_srcdir)/src/test/modules/injection_points/logical.conf
TAP_TESTS = 1
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
new file mode 100644
index 0000000000..49a736ed61
--- /dev/null
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -0,0 +1,113 @@
+Parsed test spec with 2 sessions
+
+starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
+injection_points_attach
+-----------------------
+
+(1 row)
+
+step wait_before_lock:
+ REPACK CONCURRENTLY repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step change_existing:
+ UPDATE repack_test SET i=10 where i=1;
+ UPDATE repack_test SET j=20 where i=2;
+ UPDATE repack_test SET i=30 where i=3;
+ UPDATE repack_test SET i=40 where i=30;
+ DELETE FROM repack_test WHERE i=4;
+
+step change_new:
+ INSERT INTO repack_test(i, j) VALUES (5, 5), (6, 6), (7, 7), (8, 8);
+ UPDATE repack_test SET i=50 where i=5;
+ UPDATE repack_test SET j=60 where i=6;
+ DELETE FROM repack_test WHERE i=7;
+
+step change_subxact1:
+ BEGIN;
+ INSERT INTO repack_test(i, j) VALUES (100, 100);
+ SAVEPOINT s1;
+ UPDATE repack_test SET i=101 where i=100;
+ SAVEPOINT s2;
+ UPDATE repack_test SET i=102 where i=101;
+ COMMIT;
+
+step change_subxact2:
+ BEGIN;
+ SAVEPOINT s1;
+ INSERT INTO repack_test(i, j) VALUES (110, 110);
+ ROLLBACK TO SAVEPOINT s1;
+ INSERT INTO repack_test(i, j) VALUES (110, 111);
+ COMMIT;
+
+step check2:
+ INSERT INTO relfilenodes(node)
+ SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+ SELECT i, j FROM repack_test ORDER BY i, j;
+
+ INSERT INTO data_s2(_xmin, _cmin, i, j)
+ SELECT xmin, cmin, i, j FROM repack_test;
+
+ i| j
+---+---
+ 2| 20
+ 6| 60
+ 8| 8
+ 10| 1
+ 40| 3
+ 50| 5
+102|100
+110|111
+(8 rows)
+
+step wakeup_before_lock:
+ SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1:
+ INSERT INTO relfilenodes(node)
+ SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+ SELECT count(DISTINCT node) FROM relfilenodes;
+
+ SELECT i, j FROM repack_test ORDER BY i, j;
+
+ INSERT INTO data_s1(_xmin, _cmin, i, j)
+ SELECT xmin, cmin, i, j FROM repack_test;
+
+ SELECT count(*)
+ FROM data_s1 d1 FULL JOIN data_s2 d2 USING (_xmin, _cmin, i, j)
+ WHERE d1.i ISNULL OR d2.i ISNULL;
+
+count
+-----
+ 2
+(1 row)
+
+ i| j
+---+---
+ 2| 20
+ 6| 60
+ 8| 8
+ 10| 1
+ 40| 3
+ 50| 5
+102|100
+110|111
+(8 rows)
+
+count
+-----
+ 0
+(1 row)
+
+injection_points_detach
+-----------------------
+
+(1 row)
+
diff --git a/src/test/modules/injection_points/logical.conf b/src/test/modules/injection_points/logical.conf
new file mode 100644
index 0000000000..c8f264bc6c
--- /dev/null
+++ b/src/test/modules/injection_points/logical.conf
@@ -0,0 +1 @@
+wal_level = logical
\ No newline at end of file
diff --git a/src/test/modules/injection_points/meson.build b/src/test/modules/injection_points/meson.build
index d61149712f..0e3c47ba99 100644
--- a/src/test/modules/injection_points/meson.build
+++ b/src/test/modules/injection_points/meson.build
@@ -46,9 +46,13 @@ tests += {
'specs': [
'basic',
'inplace',
+ 'repack',
'syscache-update-pruned',
],
'runningcheck': false, # see syscache-update-pruned
+ # 'repack' requires wal_level = 'logical'.
+ 'regress_args': ['--temp-config', files('logical.conf')],
+
},
'tap': {
'env': {
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
new file mode 100644
index 0000000000..5aa8983f98
--- /dev/null
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -0,0 +1,140 @@
+# Prefix the system columns with underscore as they are not allowed as column
+# names.
+setup
+{
+ CREATE EXTENSION injection_points;
+
+ CREATE TABLE repack_test(i int PRIMARY KEY, j int);
+ INSERT INTO repack_test(i, j) VALUES (1, 1), (2, 2), (3, 3), (4, 4);
+
+ CREATE TABLE relfilenodes(node oid);
+
+ CREATE TABLE data_s1(_xmin xid, _cmin cid, i int, j int);
+ CREATE TABLE data_s2(_xmin xid, _cmin cid, i int, j int);
+}
+
+teardown
+{
+ DROP TABLE repack_test;
+ DROP EXTENSION injection_points;
+
+ DROP TABLE relfilenodes;
+ DROP TABLE data_s1;
+ DROP TABLE data_s2;
+}
+
+session s1
+setup
+{
+ SELECT injection_points_set_local();
+ SELECT injection_points_attach('repack-concurrently-before-lock', 'wait');
+}
+# Perform the initial load and wait for s2 to do some data changes.
+step wait_before_lock
+{
+ REPACK CONCURRENTLY repack_test USING INDEX repack_test_pkey;
+}
+# Check the table from the perspective of s1.
+#
+# Besides the contents, we also check that relfilenode has changed.
+#
+# xmin and cmin columns are used to check that we do not change tuple
+# visibility information. Since we do not expect xmin to stay unchanged across
+# test runs, it cannot appear in the output text. Instead, have each session
+# write the contents into a table and use FULL JOIN to check if the outputs
+# are identical.
+step check1
+{
+ INSERT INTO relfilenodes(node)
+ SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+ SELECT count(DISTINCT node) FROM relfilenodes;
+
+ SELECT i, j FROM repack_test ORDER BY i, j;
+
+ INSERT INTO data_s1(_xmin, _cmin, i, j)
+ SELECT xmin, cmin, i, j FROM repack_test;
+
+ SELECT count(*)
+ FROM data_s1 d1 FULL JOIN data_s2 d2 USING (_xmin, _cmin, i, j)
+ WHERE d1.i ISNULL OR d2.i ISNULL;
+}
+teardown
+{
+ SELECT injection_points_detach('repack-concurrently-before-lock');
+}
+
+session s2
+# Change the existing data. UPDATE changes both key and non-key columns. Also
+# update one row twice to test whether tuple version generated by this session
+# can be found.
+step change_existing
+{
+ UPDATE repack_test SET i=10 where i=1;
+ UPDATE repack_test SET j=20 where i=2;
+ UPDATE repack_test SET i=30 where i=3;
+ UPDATE repack_test SET i=40 where i=30;
+ DELETE FROM repack_test WHERE i=4;
+}
+# Insert new rows and UPDATE / DELETE some of them. Again, update both key and
+# non-key column.
+step change_new
+{
+ INSERT INTO repack_test(i, j) VALUES (5, 5), (6, 6), (7, 7), (8, 8);
+ UPDATE repack_test SET i=50 where i=5;
+ UPDATE repack_test SET j=60 where i=6;
+ DELETE FROM repack_test WHERE i=7;
+}
+
+# When applying concurrent data changes, we should see the effects of an
+# in-progress subtransaction.
+step change_subxact1
+{
+ BEGIN;
+ INSERT INTO repack_test(i, j) VALUES (100, 100);
+ SAVEPOINT s1;
+ UPDATE repack_test SET i=101 where i=100;
+ SAVEPOINT s2;
+ UPDATE repack_test SET i=102 where i=101;
+ COMMIT;
+}
+
+# When applying concurrent data changes, we should not see the effects of a
+# rolled back subtransaction.
+step change_subxact2
+{
+ BEGIN;
+ SAVEPOINT s1;
+ INSERT INTO repack_test(i, j) VALUES (110, 110);
+ ROLLBACK TO SAVEPOINT s1;
+ INSERT INTO repack_test(i, j) VALUES (110, 111);
+ COMMIT;
+}
+
+# Check the table from the perspective of s2.
+step check2
+{
+ INSERT INTO relfilenodes(node)
+ SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+ SELECT i, j FROM repack_test ORDER BY i, j;
+
+ INSERT INTO data_s2(_xmin, _cmin, i, j)
+ SELECT xmin, cmin, i, j FROM repack_test;
+}
+step wakeup_before_lock
+{
+ SELECT injection_points_wakeup('repack-concurrently-before-lock');
+}
+
+# Test if data changes introduced while one session is performing REPACK
+# CONCURRENTLY find their way into the table.
+permutation
+ wait_before_lock
+ change_existing
+ change_new
+ change_subxact1
+ change_subxact2
+ check2
+ wakeup_before_lock
+ check1
--
2.43.5
[text/x-diff] v08-0007-Introduce-repack_max_xlock_time-configuration-variab.patch (20.3K, ../127361.1740559688@localhost/8-v08-0007-Introduce-repack_max_xlock_time-configuration-variab.patch)
download | inline diff:
From 25026937685b8c6b146f95efdd4dfedffc1716e0 Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 26 Feb 2025 09:17:20 +0100
Subject: [PATCH 7/9] Introduce repack_max_xlock_time configuration variable.
When executing REPACK CONCURRENTLY, we need the AccessExclusiveLock to swap
the relation files and that should require pretty short time. However, on a
busy system, other backends might change non-negligible amount of data in the
table while we are waiting for the lock. Since these changes must be applied
to the new storage before the swap, the time we eventually hold the lock might
become non-negligible too.
If the user is worried about this situation, he can set repack_max_xlock_time
to the maximum time for which the exclusive lock may be held. If this amount
of time is not sufficient to complete the REPACK CONCURRENTLY command, ERROR
is raised and the command is canceled.
---
doc/src/sgml/config.sgml | 31 ++++
doc/src/sgml/ref/repack.sgml | 9 +-
src/backend/access/heap/heapam_handler.c | 3 +-
src/backend/commands/cluster.c | 133 +++++++++++++++---
src/backend/utils/misc/guc_tables.c | 14 ++
src/backend/utils/misc/postgresql.conf.sample | 1 +
src/include/commands/cluster.h | 5 +-
.../injection_points/expected/repack.out | 74 +++++++++-
.../injection_points/specs/repack.spec | 42 ++++++
9 files changed, 292 insertions(+), 20 deletions(-)
diff --git a/doc/src/sgml/config.sgml b/doc/src/sgml/config.sgml
index e55700f35b..758cf6849d 100644
--- a/doc/src/sgml/config.sgml
+++ b/doc/src/sgml/config.sgml
@@ -10888,6 +10888,37 @@ dynamic_library_path = 'C:\tools\postgresql;H:\my_project\lib;$libdir'
</listitem>
</varlistentry>
+ <varlistentry id="guc-repack-max-xclock-time" xreflabel="repack_max_xlock_time">
+ <term><varname>repack_max_xlock_time</varname> (<type>integer</type>)
+ <indexterm>
+ <primary><varname>repack_max_xlock_time</varname> configuration parameter</primary>
+ </indexterm>
+ </term>
+ <listitem>
+ <para>
+ This is the maximum amount of time to hold an exclusive lock on a
+ table by <command>REPACK</command> with
+ the <literal>CONCURRENTLY</literal> option. Typically, these commands
+ should not need the lock for longer time
+ than <command>TRUNCATE</command> does. However, additional time might
+ be needed if the system is too busy. (See <xref linkend="sql-repack"/>
+ for explanation how the <literal>CONCURRENTLY</literal> option works.)
+ </para>
+
+ <para>
+ If you want to restrict the lock time, set this variable to the
+ highest acceptable value. If it appears during the processing that
+ additional time is needed to release the lock, the command will be
+ cancelled.
+ </para>
+
+ <para>
+ The default value is 0, which means that the lock is not released
+ until the concurrent data changes are processed.
+ </para>
+ </listitem>
+ </varlistentry>
+
</variablelist>
</sect1>
diff --git a/doc/src/sgml/ref/repack.sgml b/doc/src/sgml/ref/repack.sgml
index 9ee640e351..0c250689d1 100644
--- a/doc/src/sgml/ref/repack.sgml
+++ b/doc/src/sgml/ref/repack.sgml
@@ -188,7 +188,14 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] CONCU
(<xref linkend="logicaldecoding"/>) and applied before
the <literal>ACCESS EXCLUSIVE</literal> lock is requested. Thus the lock
is typically held only for the time needed to swap the files, which
- should be pretty short.
+ should be pretty short. However, the time might still be noticeable if
+ too many data changes have been done to the table while
+ <command>REPACK</command> was waiting for the lock: those changes must
+ be processed just before the files are swapped, while the
+ <literal>ACCESS EXCLUSIVE</literal> lock is being held. If you are
+ worried about this situation, set
+ the <link linkend="guc-repack-max-xclock-time"><varname>repack_max_xlock_time</varname></link>
+ configuration parameter to a value that your applications can tolerate.
</para>
<para>
diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index 5876a96e79..beec45b18e 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -1004,7 +1004,8 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
end_of_wal = GetFlushRecPtr(NULL);
if ((end_of_wal - end_of_wal_prev) > wal_segment_size)
{
- repack_decode_concurrent_changes(decoding_ctx, end_of_wal);
+ repack_decode_concurrent_changes(decoding_ctx, end_of_wal,
+ NULL);
end_of_wal_prev = end_of_wal;
}
}
diff --git a/src/backend/commands/cluster.c b/src/backend/commands/cluster.c
index 1ff0fcd1d9..a5790d77b5 100644
--- a/src/backend/commands/cluster.c
+++ b/src/backend/commands/cluster.c
@@ -17,6 +17,8 @@
*/
#include "postgres.h"
+#include <sys/time.h>
+
#include "access/amapi.h"
#include "access/heapam.h"
#include "access/multixact.h"
@@ -108,6 +110,15 @@ RelFileLocator repacked_rel_toast_locator = {.relNumber = InvalidOid};
#define REPACK_CONCURRENT_IN_PROGRESS_MSG \
"relation \"%s\" is already being processed by REPACK CONCURRENTLY"
+/*
+ * The maximum time to hold AccessExclusiveLock during the final
+ * processing. Note that only the execution time of
+ * process_concurrent_changes() is included here. The very last steps like
+ * swap_relation_files() shouldn't get blocked and it'd be wrong to consider
+ * them a reason to abort otherwise completed processing.
+ */
+int repack_max_xlock_time = 0;
+
/*
* Everything we need to call ExecInsertIndexTuples().
*/
@@ -197,7 +208,8 @@ static LogicalDecodingContext *setup_logical_decoding(Oid relid,
static HeapTuple get_changed_tuple(char *change);
static void apply_concurrent_changes(RepackDecodingState *dstate,
Relation rel, ScanKey key, int nkeys,
- IndexInsertState *iistate);
+ IndexInsertState *iistate,
+ struct timeval *must_complete);
static void apply_concurrent_insert(Relation rel, ConcurrentChange *change,
HeapTuple tup, IndexInsertState *iistate,
TupleTableSlot *index_slot);
@@ -214,13 +226,15 @@ static HeapTuple find_target_tuple(Relation rel, ScanKey key, int nkeys,
IndexInsertState *iistate,
TupleTableSlot *ident_slot,
IndexScanDesc *scan_p);
-static void process_concurrent_changes(LogicalDecodingContext *ctx,
+static bool process_concurrent_changes(LogicalDecodingContext *ctx,
XLogRecPtr end_of_wal,
Relation rel_dst,
Relation rel_src,
ScanKey ident_key,
int ident_key_nentries,
- IndexInsertState *iistate);
+ IndexInsertState *iistate,
+ struct timeval *must_complete);
+static bool processing_time_elapsed(struct timeval *must_complete);
static IndexInsertState *get_index_insert_state(Relation relation,
Oid ident_index_id);
static ScanKey build_identity_key(Oid ident_idx_oid, Relation rel_src,
@@ -3016,7 +3030,8 @@ get_changed_tuple(char *change)
*/
void
repack_decode_concurrent_changes(LogicalDecodingContext *ctx,
- XLogRecPtr end_of_wal)
+ XLogRecPtr end_of_wal,
+ struct timeval *must_complete)
{
RepackDecodingState *dstate;
ResourceOwner resowner_old;
@@ -3054,6 +3069,9 @@ repack_decode_concurrent_changes(LogicalDecodingContext *ctx,
if (record != NULL)
LogicalDecodingProcessRecord(ctx, ctx->reader);
+ if (processing_time_elapsed(must_complete))
+ break;
+
/*
* If WAL segment boundary has been crossed, inform the decoding
* system that the catalog_xmin can advance. (We can confirm more
@@ -3096,7 +3114,8 @@ repack_decode_concurrent_changes(LogicalDecodingContext *ctx,
*/
static void
apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
- ScanKey key, int nkeys, IndexInsertState *iistate)
+ ScanKey key, int nkeys, IndexInsertState *iistate,
+ struct timeval *must_complete)
{
TupleTableSlot *index_slot, *ident_slot;
HeapTuple tup_old = NULL;
@@ -3126,6 +3145,9 @@ apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
CHECK_FOR_INTERRUPTS();
+ Assert(dstate->nchanges > 0);
+ dstate->nchanges--;
+
/* Get the change from the single-column tuple. */
tup_change = ExecFetchSlotHeapTuple(dstate->tsslot, false, &shouldFree);
heap_deform_tuple(tup_change, dstate->tupdesc_change, values, isnull);
@@ -3250,10 +3272,22 @@ apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
/* TTSOpsMinimalTuple has .get_heap_tuple==NULL. */
Assert(shouldFree);
pfree(tup_change);
+
+ /*
+ * If there is a limit on the time of completion, check it
+ * now. However, make sure the loop does not break if tup_old was set
+ * in the previous iteration. In such a case we could not resume the
+ * processing in the next call.
+ */
+ if (must_complete && tup_old == NULL &&
+ processing_time_elapsed(must_complete))
+ /* The next call will process the remaining changes. */
+ break;
}
- tuplestore_clear(dstate->tstore);
- dstate->nchanges = 0;
+ /* If we could not apply all the changes, the next call will do. */
+ if (dstate->nchanges == 0)
+ tuplestore_clear(dstate->tstore);
/* Cleanup. */
ExecDropSingleTupleTableSlot(index_slot);
@@ -3456,11 +3490,15 @@ find_target_tuple(Relation rel, ScanKey key, int nkeys, HeapTuple tup_key,
* Decode and apply concurrent changes.
*
* Pass rel_src iff its reltoastrelid is needed.
+ *
+ * Returns true if must_complete is NULL or if managed to complete by the time
+ * *must_complete indicates.
*/
-static void
+static bool
process_concurrent_changes(LogicalDecodingContext *ctx, XLogRecPtr end_of_wal,
Relation rel_dst, Relation rel_src, ScanKey ident_key,
- int ident_key_nentries, IndexInsertState *iistate)
+ int ident_key_nentries, IndexInsertState *iistate,
+ struct timeval *must_complete)
{
RepackDecodingState *dstate;
@@ -3469,10 +3507,19 @@ process_concurrent_changes(LogicalDecodingContext *ctx, XLogRecPtr end_of_wal,
dstate = (RepackDecodingState *) ctx->output_writer_private;
- repack_decode_concurrent_changes(ctx, end_of_wal);
+ repack_decode_concurrent_changes(ctx, end_of_wal, must_complete);
+
+ if (processing_time_elapsed(must_complete))
+ /* Caller is responsible for applying the changes. */
+ return false;
+ /*
+ * *must_complete not reached, so there are really no changes. (It's
+ * possible to see no changes just because not enough time was left for
+ * the decoding.)
+ */
if (dstate->nchanges == 0)
- return;
+ return true;
PG_TRY();
{
@@ -3484,7 +3531,7 @@ process_concurrent_changes(LogicalDecodingContext *ctx, XLogRecPtr end_of_wal,
rel_dst->rd_toastoid = rel_src->rd_rel->reltoastrelid;
apply_concurrent_changes(dstate, rel_dst, ident_key,
- ident_key_nentries, iistate);
+ ident_key_nentries, iistate, must_complete);
}
PG_FINALLY();
{
@@ -3494,6 +3541,28 @@ process_concurrent_changes(LogicalDecodingContext *ctx, XLogRecPtr end_of_wal,
rel_dst->rd_toastoid = InvalidOid;
}
PG_END_TRY();
+
+ /*
+ * apply_concurrent_changes() does check the processing time, so if some
+ * changes are left, we ran out of time.
+ */
+ return dstate->nchanges == 0;
+}
+
+/*
+ * Check if the current time is beyond *must_complete.
+ */
+static bool
+processing_time_elapsed(struct timeval *must_complete)
+{
+ struct timeval now;
+
+ if (must_complete == NULL)
+ return false;
+
+ gettimeofday(&now, NULL);
+
+ return timercmp(&now, must_complete, >);
}
static IndexInsertState *
@@ -3654,6 +3723,8 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
RelReopenInfo *rri = NULL;
int nrel;
Relation *ind_refs_all, *ind_refs_p;
+ struct timeval t_end;
+ struct timeval *t_end_ptr = NULL;
/* Like in cluster_rel(). */
lockmode_old = ShareUpdateExclusiveLock;
@@ -3733,7 +3804,8 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
*/
process_concurrent_changes(ctx, end_of_wal, NewHeap,
swap_toast_by_content ? OldHeap : NULL,
- ident_key, ident_key_nentries, iistate);
+ ident_key, ident_key_nentries, iistate,
+ NULL);
/*
* Release the locks that allowed concurrent data changes, in order to
@@ -3855,9 +3927,38 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
end_of_wal = GetFlushRecPtr(NULL);
/* Apply the concurrent changes again. */
- process_concurrent_changes(ctx, end_of_wal, NewHeap,
- swap_toast_by_content ? OldHeap : NULL,
- ident_key, ident_key_nentries, iistate);
+ /*
+ * This time we have the exclusive lock on the table, so make sure that
+ * repack_max_xlock_time is not exceeded.
+ */
+ if (repack_max_xlock_time > 0)
+ {
+ int64 usec;
+ struct timeval t_start;
+
+ gettimeofday(&t_start, NULL);
+ /* Add the whole seconds. */
+ t_end.tv_sec = t_start.tv_sec + repack_max_xlock_time / 1000;
+ /* Add the rest, expressed in microseconds. */
+ usec = t_start.tv_usec + 1000 * (repack_max_xlock_time % 1000);
+ /* The number of microseconds could have overflown. */
+ t_end.tv_sec += usec / USECS_PER_SEC;
+ t_end.tv_usec = usec % USECS_PER_SEC;
+ t_end_ptr = &t_end;
+ }
+ /*
+ * During testing, stop here to simulate excessive processing time.
+ */
+ INJECTION_POINT("repack-concurrently-after-lock");
+
+ if (!process_concurrent_changes(ctx, end_of_wal, NewHeap,
+ swap_toast_by_content ? OldHeap : NULL,
+ ident_key, ident_key_nentries, iistate,
+ t_end_ptr))
+ ereport(ERROR,
+ (errmsg("could not process concurrent data changes in time"),
+ errhint("Please consider adjusting \"repack_max_xlock_time\".")));
+
/* Remember info about rel before closing OldHeap */
relpersistence = OldHeap->rd_rel->relpersistence;
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index ad25cbb39c..a1a19f2cbd 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -39,6 +39,7 @@
#include "catalog/namespace.h"
#include "catalog/storage.h"
#include "commands/async.h"
+#include "commands/cluster.h"
#include "commands/event_trigger.h"
#include "commands/tablespace.h"
#include "commands/trigger.h"
@@ -2814,6 +2815,19 @@ struct config_int ConfigureNamesInt[] =
1600000000, 0, 2100000000,
NULL, NULL, NULL
},
+ {
+ {"repack_max_xlock_time", PGC_USERSET, LOCK_MANAGEMENT,
+ gettext_noop("Maximum time for REPACK CONCURRENTLY to keep table locked."),
+ gettext_noop(
+ "The table is locked in exclusive mode during the final stage of processing. "
+ "If the lock time exceeds this value, error is raised and the lock is "
+ "released. Set to zero if you don't care how long the lock can be held."),
+ GUC_UNIT_MS
+ },
+ &repack_max_xlock_time,
+ 0, 0, INT_MAX,
+ NULL, NULL, NULL
+ },
/*
* See also CheckRequiredParameterValues() if this parameter changes
diff --git a/src/backend/utils/misc/postgresql.conf.sample b/src/backend/utils/misc/postgresql.conf.sample
index 5362ff8051..b59f8aae7e 100644
--- a/src/backend/utils/misc/postgresql.conf.sample
+++ b/src/backend/utils/misc/postgresql.conf.sample
@@ -744,6 +744,7 @@ autovacuum_worker_slots = 16 # autovacuum worker slots to allocate
#lock_timeout = 0 # in milliseconds, 0 is disabled
#idle_in_transaction_session_timeout = 0 # in milliseconds, 0 is disabled
#idle_session_timeout = 0 # in milliseconds, 0 is disabled
+#repack_max_xlock_time = 0
#bytea_output = 'hex' # hex, escape
#xmlbinary = 'base64'
#xmloption = 'content'
diff --git a/src/include/commands/cluster.h b/src/include/commands/cluster.h
index ef3cb55751..f5600bf4f6 100644
--- a/src/include/commands/cluster.h
+++ b/src/include/commands/cluster.h
@@ -59,6 +59,8 @@ typedef enum ClusterCommand
extern RelFileLocator repacked_rel_locator;
extern RelFileLocator repacked_rel_toast_locator;
+extern PGDLLIMPORT int repack_max_xlock_time;
+
typedef enum
{
CHANGE_INSERT,
@@ -154,7 +156,8 @@ extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid,
extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal);
extern void can_repack_concurrently(Relation rel);
extern void repack_decode_concurrent_changes(LogicalDecodingContext *ctx,
- XLogRecPtr end_of_wal);
+ XLogRecPtr end_of_wal,
+ struct timeval *must_complete);
extern Oid make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
char relpersistence, LOCKMODE lockmode);
extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index 49a736ed61..f2728d9422 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 4 sessions
starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
injection_points_attach
@@ -111,3 +111,75 @@ injection_points_detach
(1 row)
+injection_points_detach
+-----------------------
+
+(1 row)
+
+
+starting permutation: wait_after_lock wakeup_after_lock
+injection_points_attach
+-----------------------
+
+(1 row)
+
+step wait_after_lock:
+ REPACK CONCURRENTLY repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step wakeup_after_lock:
+ SELECT injection_points_wakeup('repack-concurrently-after-lock');
+
+injection_points_wakeup
+-----------------------
+
+(1 row)
+
+step wait_after_lock: <... completed>
+injection_points_detach
+-----------------------
+
+(1 row)
+
+injection_points_detach
+-----------------------
+
+(1 row)
+
+
+starting permutation: wait_after_lock after_lock_delay wakeup_after_lock
+injection_points_attach
+-----------------------
+
+(1 row)
+
+step wait_after_lock:
+ REPACK CONCURRENTLY repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step after_lock_delay:
+ SELECT pg_sleep(1.5);
+
+pg_sleep
+--------
+
+(1 row)
+
+step wakeup_after_lock:
+ SELECT injection_points_wakeup('repack-concurrently-after-lock');
+
+injection_points_wakeup
+-----------------------
+
+(1 row)
+
+step wait_after_lock: <... completed>
+ERROR: could not process concurrent data changes in time
+injection_points_detach
+-----------------------
+
+(1 row)
+
+injection_points_detach
+-----------------------
+
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index 5aa8983f98..0f45f9d254 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -127,6 +127,34 @@ step wakeup_before_lock
SELECT injection_points_wakeup('repack-concurrently-before-lock');
}
+session s3
+setup
+{
+ SELECT injection_points_set_local();
+ SELECT injection_points_attach('repack-concurrently-after-lock', 'wait');
+ SET repack_max_xlock_time TO '1s';
+}
+# Perform the initial load, lock the table in exclusive mode and wait. s4 will
+# cancel the waiting.
+step wait_after_lock
+{
+ REPACK CONCURRENTLY repack_test USING INDEX repack_test_pkey;
+}
+teardown
+{
+ SELECT injection_points_detach('repack-concurrently-after-lock');
+}
+
+session s4
+step wakeup_after_lock
+{
+ SELECT injection_points_wakeup('repack-concurrently-after-lock');
+}
+step after_lock_delay
+{
+ SELECT pg_sleep(1.5);
+}
+
# Test if data changes introduced while one session is performing REPACK
# CONCURRENTLY find their way into the table.
permutation
@@ -138,3 +166,17 @@ permutation
check2
wakeup_before_lock
check1
+
+# Test the repack_max_xlock_time configuration variable.
+#
+# First, cancel waiting on the injection point immediately. That way, REPACK
+# should complete.
+permutation
+ wait_after_lock
+ wakeup_after_lock
+# Second, cancel the waiting with a delay that violates
+# repack_max_xlock_time.
+permutation
+ wait_after_lock
+ after_lock_delay
+ wakeup_after_lock
--
2.43.5
[text/x-diff] v08-0008-Enable-logical-decoding-transiently-only-for-REPACK-.patch (24.2K, ../127361.1740559688@localhost/9-v08-0008-Enable-logical-decoding-transiently-only-for-REPACK-.patch)
download | inline diff:
From c5adedcec21f4523243af2c0e5c0d63ef2e54b23 Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 26 Feb 2025 09:17:21 +0100
Subject: [PATCH 8/9] Enable logical decoding transiently, only for REPACK
CONCURRENTLY.
As REPACK CONCURRENTLY uses logical decoding, it requires wal_level to be set
to 'logical', while 'replica' is the default value. If logical replication is
not used, users will probably be reluctant to set the GUC to 'logical' because
it can affect server performance (by writing additional information to WAL)
and because it cannot be changed to 'logical' only for the time REPACK
CONCURRENTLY is running: change of this GUC requires server restart to take
effect.
This patch teaches postgres backend to recognize whether it should consider
wal_level='logical' "locally" for particular transaction, even if the
wal_level GUC is actually set to 'replica'. Also it ensures that the logical
decoding specific information is added to WAL only for the tables which are
currently being processed by REPACK CONCURRENTLY.
If the logical decoding is enabled this way, only temporary replication slots
should be created. The problem of permanent slot is that it is restored during
server restart, and the restore fails if wal_level is not "globally"
'logical'.
There is an independent work in progres to enable logical decoding transiently
[1]. ISTM that this is too "heavyweight" solution for our problem. And I think
that these two approaches are not mutually exclusive: once [1] is committed,
we only need to adjust the XLogLogicalInfoActive() macro.
[1] https://www.postgresql.org/message-id/CAD21AoCVLeLYq09pQPaWs%2BJwdni5FuJ8v2jgq-u9_uFbcp6UbA%40mail.gmail.com
---
src/backend/access/transam/parallel.c | 8 ++
src/backend/access/transam/xact.c | 106 ++++++++++++++---
src/backend/access/transam/xlog.c | 1 +
src/backend/commands/cluster.c | 107 ++++++++++++++----
src/backend/replication/logical/logical.c | 9 +-
src/backend/storage/ipc/standby.c | 4 +-
src/include/access/xlog.h | 15 ++-
src/include/commands/cluster.h | 1 +
src/include/utils/rel.h | 6 +-
src/test/modules/injection_points/Makefile | 1 -
.../modules/injection_points/logical.conf | 1 -
src/test/modules/injection_points/meson.build | 3 -
12 files changed, 216 insertions(+), 46 deletions(-)
delete mode 100644 src/test/modules/injection_points/logical.conf
diff --git a/src/backend/access/transam/parallel.c b/src/backend/access/transam/parallel.c
index 4ab5df9213..4e6e92c4db 100644
--- a/src/backend/access/transam/parallel.c
+++ b/src/backend/access/transam/parallel.c
@@ -97,6 +97,7 @@ typedef struct FixedParallelState
TimestampTz xact_ts;
TimestampTz stmt_ts;
SerializableXactHandle serializable_xact_handle;
+ int wal_level_transient;
/* Mutex protects remaining fields. */
slock_t mutex;
@@ -351,6 +352,7 @@ InitializeParallelDSM(ParallelContext *pcxt)
fps->xact_ts = GetCurrentTransactionStartTimestamp();
fps->stmt_ts = GetCurrentStatementStartTimestamp();
fps->serializable_xact_handle = ShareSerializableXact();
+ fps->wal_level_transient = wal_level_transient;
SpinLockInit(&fps->mutex);
fps->last_xlog_end = 0;
shm_toc_insert(pcxt->toc, PARALLEL_KEY_FIXED, fps);
@@ -1546,6 +1548,12 @@ ParallelWorkerMain(Datum main_arg)
/* Attach to the leader's serializable transaction, if SERIALIZABLE. */
AttachSerializableXact(fps->serializable_xact_handle);
+ /*
+ * Restore the information whether this worker should behave as if
+ * wal_level was WAL_LEVEL_LOGICAL..
+ */
+ wal_level_transient = fps->wal_level_transient;
+
/*
* We've initialized all of our state now; nothing should change
* hereafter.
diff --git a/src/backend/access/transam/xact.c b/src/backend/access/transam/xact.c
index 0bfd329847..0e4f234440 100644
--- a/src/backend/access/transam/xact.c
+++ b/src/backend/access/transam/xact.c
@@ -36,6 +36,7 @@
#include "catalog/pg_enum.h"
#include "catalog/storage.h"
#include "commands/async.h"
+#include "commands/cluster.h"
#include "commands/tablecmds.h"
#include "commands/trigger.h"
#include "common/pg_prng.h"
@@ -137,6 +138,12 @@ static TransactionId *ParallelCurrentXids;
static int nRepackCurrentXids = 0;
static TransactionId *RepackCurrentXids = NULL;
+/*
+ * Have we determined the value of wal_level_transient for the current
+ * transaction?
+ */
+static bool wal_level_transient_checked = false;
+
/*
* Miscellaneous flag bits to record events which occur on the top level
* transaction. These flags are only persisted in MyXactFlags and are intended
@@ -648,6 +655,7 @@ AssignTransactionId(TransactionState s)
bool isSubXact = (s->parent != NULL);
ResourceOwner currentOwner;
bool log_unknown_top = false;
+ bool set_wal_level_transient = false;
/* Assert that caller didn't screw up */
Assert(!FullTransactionIdIsValid(s->fullTransactionId));
@@ -662,6 +670,32 @@ AssignTransactionId(TransactionState s)
(errcode(ERRCODE_INVALID_TRANSACTION_STATE),
errmsg("cannot assign transaction IDs during a parallel operation")));
+ /*
+ * The first call (i.e. the first write) in the transaction tree
+ * determines whether the whole transaction assumes logical decoding or
+ * not.
+ */
+ if (!wal_level_transient_checked)
+ {
+ Assert(wal_level_transient == WAL_LEVEL_MINIMAL);
+
+ /*
+ * Do not repeat the check when calling this function for parent
+ * transactions.
+ */
+ wal_level_transient_checked = true;
+
+ /*
+ * Remember that the actual check is needed. We cannot do it until the
+ * top-level transaction has its XID assigned, see comments below.
+ *
+ * There is no use case for overriding MINIMAL, and LOGICAL cannot be
+ * overridden as such.
+ */
+ if (wal_level == WAL_LEVEL_REPLICA)
+ set_wal_level_transient = true;
+ }
+
/*
* Ensure parent(s) have XIDs, so that a child always has an XID later
* than its parent. Mustn't recurse here, or we might get a stack
@@ -691,20 +725,6 @@ AssignTransactionId(TransactionState s)
pfree(parents);
}
- /*
- * When wal_level=logical, guarantee that a subtransaction's xid can only
- * be seen in the WAL stream if its toplevel xid has been logged before.
- * If necessary we log an xact_assignment record with fewer than
- * PGPROC_MAX_CACHED_SUBXIDS. Note that it is fine if didLogXid isn't set
- * for a transaction even though it appears in a WAL record, we just might
- * superfluously log something. That can happen when an xid is included
- * somewhere inside a wal record, but not in XLogRecord->xl_xid, like in
- * xl_standby_locks.
- */
- if (isSubXact && XLogLogicalInfoActive() &&
- !TopTransactionStateData.didLogXid)
- log_unknown_top = true;
-
/*
* Generate a new FullTransactionId and record its xid in PGPROC and
* pg_subtrans.
@@ -729,6 +749,54 @@ AssignTransactionId(TransactionState s)
if (!isSubXact)
RegisterPredicateLockingXid(XidFromFullTransactionId(s->fullTransactionId));
+ /*
+ * Check if this transaction should consider wal_level=logical.
+ *
+ * Sometimes we need to turn on the logical decoding transiently although
+ * wal_level=WAL_LEVEL_REPLICA. Currently we do so when at least one table
+ * is being clustered concurrently, i.e. when we should assume that
+ * changes done by this transaction will be decoded. In such a case we
+ * adjust the value of XLogLogicalInfoActive() by setting
+ * wal_level_transient to LOGICAL.
+ *
+ * It's important not to do this check until the XID of the top-level
+ * transaction is in ProcGlobal: if the decoding becomes mandatory right
+ * after the check, our transaction will fail to write the necessary
+ * information to WAL. However, if the top-level transaction is already in
+ * ProcGlobal, its XID is guaranteed to appear in the xl_running_xacts
+ * record and therefore the snapshot builder will not try to decode the
+ * transaction (because it assumes it could have missed the initial part
+ * of the transaction).
+ *
+ * On the other hand, if the decoding became mandatory between the actual
+ * XID assignment and now, the transaction will WAL the decoding specific
+ * information unnecessarily. Let's assume that such race conditions do
+ * not happen too often.
+ */
+ if (set_wal_level_transient)
+ {
+ /*
+ * Check for the operation that enables the logical decoding
+ * transiently.
+ */
+ if (is_concurrent_repack_in_progress(InvalidOid))
+ wal_level_transient = WAL_LEVEL_LOGICAL;
+ }
+
+ /*
+ * When wal_level=logical, guarantee that a subtransaction's xid can only
+ * be seen in the WAL stream if its toplevel xid has been logged before.
+ * If necessary we log an xact_assignment record with fewer than
+ * PGPROC_MAX_CACHED_SUBXIDS. Note that it is fine if didLogXid isn't set
+ * for a transaction even though it appears in a WAL record, we just might
+ * superfluously log something. That can happen when an xid is included
+ * somewhere inside a wal record, but not in XLogRecord->xl_xid, like in
+ * xl_standby_locks.
+ */
+ if (isSubXact && XLogLogicalInfoActive() &&
+ !TopTransactionStateData.didLogXid)
+ log_unknown_top = true;
+
/*
* Acquire lock on the transaction XID. (We assume this cannot block.) We
* have to ensure that the lock is assigned to the transaction's own
@@ -2243,6 +2311,16 @@ StartTransaction(void)
if (TransactionTimeout > 0)
enable_timeout_after(TRANSACTION_TIMEOUT, TransactionTimeout);
+ /*
+ * wal_level_transient can override wal_level for individual transactions,
+ * which effectively enables logical decoding for them. At the moment we
+ * don't know if this transaction will write any data changes to be
+ * decoded. Should it do, AssignTransactionId() will check if the decoding
+ * needs to be considered.
+ */
+ wal_level_transient = WAL_LEVEL_MINIMAL;
+ wal_level_transient_checked = false;
+
ShowTransactionState("StartTransaction");
}
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 799fc739e1..4a8a43a6da 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -129,6 +129,7 @@ bool wal_recycle = true;
bool log_checkpoints = true;
int wal_sync_method = DEFAULT_WAL_SYNC_METHOD;
int wal_level = WAL_LEVEL_REPLICA;
+int wal_level_transient = WAL_LEVEL_MINIMAL;
int CommitDelay = 0; /* precommit delay in microseconds */
int CommitSiblings = 5; /* # concurrent xacts needed to sleep */
int wal_retrieve_retry_interval = 5000;
diff --git a/src/backend/commands/cluster.c b/src/backend/commands/cluster.c
index a5790d77b5..910ff9fa91 100644
--- a/src/backend/commands/cluster.c
+++ b/src/backend/commands/cluster.c
@@ -1298,7 +1298,7 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
*
* In the REPACK CONCURRENTLY case, the lock does not help because we need
* to release it temporarily at some point. Instead, we expect VACUUM /
- * CLUSTER to skip tables which are present in RepackedRelsHash.
+ * CLUSTER to skip tables which are present in repackedRels->hashtable.
*/
if (OldHeap->rd_rel->reltoastrelid && !concurrent)
LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
@@ -2312,7 +2312,16 @@ typedef struct RepackedRel
Oid dbid;
} RepackedRel;
-static HTAB *RepackedRelsHash = NULL;
+typedef struct RepackedRels
+{
+ /* Hashtable of RepackedRel elements. */
+ HTAB *hashtable;
+
+ /* The number of elements in the hashtable.. */
+ pg_atomic_uint32 nrels;
+} RepackedRels;
+
+static RepackedRels *repackedRels = NULL;
/* Maximum number of entries in the hashtable. */
static int maxRepackedRels = 0;
@@ -2320,28 +2329,44 @@ static int maxRepackedRels = 0;
Size
RepackShmemSize(void)
{
+ Size result;
+
+ result = sizeof(RepackedRels);
+
/*
* A replication slot is needed for the processing, so use this GUC to
* allocate memory for the hashtable.
*/
maxRepackedRels = max_replication_slots;
- return hash_estimate_size(maxRepackedRels, sizeof(RepackedRel));
+ result += hash_estimate_size(maxRepackedRels, sizeof(RepackedRel));
+ return result;
}
void
RepackShmemInit(void)
{
+ bool found;
HASHCTL info;
+ repackedRels = ShmemInitStruct("Repacked Relations",
+ sizeof(RepackedRels),
+ &found);
+ if (!IsUnderPostmaster)
+ {
+ Assert(!found);
+ pg_atomic_init_u32(&repackedRels->nrels, 0);
+ }
+ else
+ Assert(found);
+
info.keysize = sizeof(RepackedRel);
info.entrysize = info.keysize;
-
- RepackedRelsHash = ShmemInitHash("Repacked Relations",
- maxRepackedRels,
- maxRepackedRels,
- &info,
- HASH_ELEM | HASH_BLOBS);
+ repackedRels->hashtable = ShmemInitHash("Repacked Relations Hash",
+ maxRepackedRels,
+ maxRepackedRels,
+ &info,
+ HASH_ELEM | HASH_BLOBS);
}
/*
@@ -2373,12 +2398,13 @@ begin_concurrent_repack(Relation *rel_p, Relation *index_p,
RelReopenInfo rri[2];
int nrel;
static bool before_shmem_exit_callback_setup = false;
+ uint32 nrels PG_USED_FOR_ASSERTS_ONLY;
relid = RelationGetRelid(rel);
/*
- * Make sure that we do not leave an entry in RepackedRelsHash if exiting
- * due to FATAL.
+ * Make sure that we do not leave an entry in repackedRels->Hashtable if
+ * exiting due to FATAL.
*/
if (!before_shmem_exit_callback_setup)
{
@@ -2393,7 +2419,7 @@ begin_concurrent_repack(Relation *rel_p, Relation *index_p,
*entered_p = false;
LWLockAcquire(RepackedRelsLock, LW_EXCLUSIVE);
entry = (RepackedRel *)
- hash_search(RepackedRelsHash, &key, HASH_ENTER_NULL, &found);
+ hash_search(repackedRels->hashtable, &key, HASH_ENTER_NULL, &found);
if (found)
{
/*
@@ -2411,6 +2437,10 @@ begin_concurrent_repack(Relation *rel_p, Relation *index_p,
(errmsg("too many requests for REPACK CONCURRENTLY at a time")),
(errhint("Please consider increasing the \"max_replication_slots\" configuration parameter.")));
+ /* Increment the number of relations. */
+ nrels = pg_atomic_fetch_add_u32(&repackedRels->nrels, 1);
+ Assert(nrels < maxRepackedRels);
+
/*
* Even if the insertion of TOAST relid should fail below, the caller has
* to do cleanup.
@@ -2438,7 +2468,8 @@ begin_concurrent_repack(Relation *rel_p, Relation *index_p,
{
key.relid = toastrelid;
entry = (RepackedRel *)
- hash_search(RepackedRelsHash, &key, HASH_ENTER_NULL, &found);
+ hash_search(repackedRels->hashtable, &key, HASH_ENTER_NULL,
+ &found);
if (found)
/*
* If we could enter the main fork the TOAST should succeed
@@ -2452,6 +2483,10 @@ begin_concurrent_repack(Relation *rel_p, Relation *index_p,
(errmsg("too many requests for REPACK CONCURRENTLY at a time")),
(errhint("Please consider increasing the \"max_replication_slots\" configuration parameter.")));
+ /* Increment the number of relations. */
+ nrels = pg_atomic_fetch_add_u32(&repackedRels->nrels, 1);
+ Assert(nrels < maxRepackedRels);
+
Assert(!OidIsValid(repacked_rel_toast));
repacked_rel_toast = toastrelid;
}
@@ -2531,6 +2566,7 @@ end_concurrent_repack(bool error)
RepackedRel *entry = NULL, *entry_toast = NULL;
Oid relid = repacked_rel;
Oid toastrelid = repacked_rel_toast;
+ uint32 nrels PG_USED_FOR_ASSERTS_ONLY;
/* Remove the relation from the hash if we managed to insert one. */
if (OidIsValid(repacked_rel))
@@ -2539,23 +2575,32 @@ end_concurrent_repack(bool error)
key.relid = repacked_rel;
key.dbid = MyDatabaseId;
LWLockAcquire(RepackedRelsLock, LW_EXCLUSIVE);
- entry = hash_search(RepackedRelsHash, &key, HASH_REMOVE, NULL);
+ entry = hash_search(repackedRels->hashtable, &key, HASH_REMOVE,
+ NULL);
/*
* By clearing this variable we also disable
* cluster_before_shmem_exit_callback().
*/
repacked_rel = InvalidOid;
+
+ /* Decrement the number of relations. */
+ nrels = pg_atomic_fetch_sub_u32(&repackedRels->nrels, 1);
+ Assert(nrels > 0);
}
/* Remove the TOAST relation if there is one. */
if (OidIsValid(repacked_rel_toast))
{
key.relid = repacked_rel_toast;
- entry_toast = hash_search(RepackedRelsHash, &key, HASH_REMOVE,
+ entry_toast = hash_search(repackedRels->hashtable, &key, HASH_REMOVE,
NULL);
repacked_rel_toast = InvalidOid;
+
+ /* Decrement the number of relations. */
+ nrels = pg_atomic_fetch_sub_u32(&repackedRels->nrels, 1);
+ Assert(nrels > 0);
}
LWLockRelease(RepackedRelsLock);
@@ -2621,7 +2666,7 @@ end_concurrent_repack(bool error)
}
/*
- * A wrapper to call end_concurrent_repack() as a before_shmem_exit callback.
+ * A wrapper to call end_concurrent_cluster() as a before_shmem_exit callback.
*/
static void
cluster_before_shmem_exit_callback(int code, Datum arg)
@@ -2632,24 +2677,48 @@ cluster_before_shmem_exit_callback(int code, Datum arg)
/*
* Check if relation is currently being processed by REPACK CONCURRENTLY.
+ *
+ * If relid is InvalidOid, check if any relation is being processed.
*/
bool
is_concurrent_repack_in_progress(Oid relid)
{
RepackedRel key, *entry;
+ /*
+ * If the caller is interested whether any relation is being repacked,
+ * just use the counter.
+ */
+ if (!OidIsValid(relid))
+ {
+ if (pg_atomic_read_u32(&repackedRels->nrels) > 0)
+ return true;
+ else
+ return false;
+ }
+
+ /* For particular relation we need to search in the hashtable. */
memset(&key, 0, sizeof(key));
key.relid = relid;
key.dbid = MyDatabaseId;
LWLockAcquire(RepackedRelsLock, LW_SHARED);
entry = (RepackedRel *)
- hash_search(RepackedRelsHash, &key, HASH_FIND, NULL);
+ hash_search(repackedRels->hashtable, &key, HASH_FIND, NULL);
LWLockRelease(RepackedRelsLock);
return entry != NULL;
}
+/*
+ * Is this backend performing REPACK CONCURRENTLY?
+ */
+bool
+is_concurrent_repack_run_by_me(void)
+{
+ return OidIsValid(repacked_rel);
+}
+
/*
* Check if REPACK CONCURRENTLY is already running for given relation, and if
* so, raise ERROR. The problem is that cluster_rel() needs to release its
@@ -2944,8 +3013,8 @@ setup_logical_decoding(Oid relid, const char *slotname, TupleDesc tupdesc)
* useful for us.
*
* Regarding the value of need_full_snapshot, we pass false because the
- * table we are processing is present in RepackedRelsHash and therefore,
- * regarding logical decoding, treated like a catalog.
+ * table we are processing is present in repackedRels->hashtable and
+ * therefore, regarding logical decoding, treated like a catalog.
*/
ctx = CreateInitDecodingContext(REPL_PLUGIN_NAME,
NIL,
diff --git a/src/backend/replication/logical/logical.c b/src/backend/replication/logical/logical.c
index 8ea846bfc3..e5790d3fe8 100644
--- a/src/backend/replication/logical/logical.c
+++ b/src/backend/replication/logical/logical.c
@@ -30,6 +30,7 @@
#include "access/xact.h"
#include "access/xlogutils.h"
+#include "commands/cluster.h"
#include "fmgr.h"
#include "miscadmin.h"
#include "pgstat.h"
@@ -112,10 +113,12 @@ CheckLogicalDecodingRequirements(void)
/*
* NB: Adding a new requirement likely means that RestoreSlotFromDisk()
- * needs the same check.
+ * needs the same check. (Except that only temporary slots should be
+ * created for REPACK CONCURRENTLY, which effectively raises wal_level to
+ * LOGICAL.)
*/
-
- if (wal_level < WAL_LEVEL_LOGICAL)
+ if ((wal_level < WAL_LEVEL_LOGICAL && !is_concurrent_repack_run_by_me())
+ || wal_level < WAL_LEVEL_REPLICA)
ereport(ERROR,
(errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
errmsg("logical decoding requires \"wal_level\" >= \"logical\"")));
diff --git a/src/backend/storage/ipc/standby.c b/src/backend/storage/ipc/standby.c
index 5acb4508f8..413bcc1add 100644
--- a/src/backend/storage/ipc/standby.c
+++ b/src/backend/storage/ipc/standby.c
@@ -1313,13 +1313,13 @@ LogStandbySnapshot(void)
* record. Fortunately this routine isn't executed frequently, and it's
* only a shared lock.
*/
- if (wal_level < WAL_LEVEL_LOGICAL)
+ if (!XLogLogicalInfoActive())
LWLockRelease(ProcArrayLock);
recptr = LogCurrentRunningXacts(running);
/* Release lock if we kept it longer ... */
- if (wal_level >= WAL_LEVEL_LOGICAL)
+ if (XLogLogicalInfoActive())
LWLockRelease(ProcArrayLock);
/* GetRunningTransactionData() acquired XidGenLock, we must release it */
diff --git a/src/include/access/xlog.h b/src/include/access/xlog.h
index d313099c02..a325bb1d16 100644
--- a/src/include/access/xlog.h
+++ b/src/include/access/xlog.h
@@ -95,6 +95,12 @@ typedef enum RecoveryState
extern PGDLLIMPORT int wal_level;
+/*
+ * wal_level_transient overrides wal_level if logical decoding needs to be
+ * enabled transiently.
+ */
+extern PGDLLIMPORT int wal_level_transient;
+
/* Is WAL archiving enabled (always or only while server is running normally)? */
#define XLogArchivingActive() \
(AssertMacro(XLogArchiveMode == ARCHIVE_MODE_OFF || wal_level >= WAL_LEVEL_REPLICA), XLogArchiveMode > ARCHIVE_MODE_OFF)
@@ -122,8 +128,13 @@ extern PGDLLIMPORT int wal_level;
/* Do we need to WAL-log information required only for Hot Standby and logical replication? */
#define XLogStandbyInfoActive() (wal_level >= WAL_LEVEL_REPLICA)
-/* Do we need to WAL-log information required only for logical replication? */
-#define XLogLogicalInfoActive() (wal_level >= WAL_LEVEL_LOGICAL)
+/*
+ * Do we need to WAL-log information required only for logical replication?
+ *
+ * wal_level_transient overrides wal_level if logical decoding needs to be
+ * active transiently.
+ */
+#define XLogLogicalInfoActive() (Max(wal_level, wal_level_transient) == WAL_LEVEL_LOGICAL)
#ifdef WAL_DEBUG
extern PGDLLIMPORT bool XLOG_DEBUG;
diff --git a/src/include/commands/cluster.h b/src/include/commands/cluster.h
index f5600bf4f6..3ed3066b36 100644
--- a/src/include/commands/cluster.h
+++ b/src/include/commands/cluster.h
@@ -173,6 +173,7 @@ extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
extern Size RepackShmemSize(void);
extern void RepackShmemInit(void);
extern bool is_concurrent_repack_in_progress(Oid relid);
+extern bool is_concurrent_repack_run_by_me(void);
extern void check_for_concurrent_repack(Oid relid, LOCKMODE lockmode);
extern void repack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel);
diff --git a/src/include/utils/rel.h b/src/include/utils/rel.h
index 741b29226d..81f5348a4f 100644
--- a/src/include/utils/rel.h
+++ b/src/include/utils/rel.h
@@ -709,12 +709,16 @@ RelationCloseSmgr(Relation relation)
* it would complicate decoding slightly for little gain). Note that we *do*
* log information for user defined catalog tables since they presumably are
* interesting to the user...
+ *
+ * If particular relations require that, the logical decoding can be active
+ * even if wal_level is REPLICA. Do not log other relations in that case.
*/
#define RelationIsLogicallyLogged(relation) \
(XLogLogicalInfoActive() && \
RelationNeedsWAL(relation) && \
(relation)->rd_rel->relkind != RELKIND_FOREIGN_TABLE && \
- !IsCatalogRelation(relation))
+ !IsCatalogRelation(relation) && \
+ (wal_level == WAL_LEVEL_LOGICAL || (relation)->rd_repack_concurrent))
/* routines in utils/cache/relcache.c */
extern void RelationIncrementReferenceCount(Relation rel);
diff --git a/src/test/modules/injection_points/Makefile b/src/test/modules/injection_points/Makefile
index 405d0811b4..4f6c0ca3a8 100644
--- a/src/test/modules/injection_points/Makefile
+++ b/src/test/modules/injection_points/Makefile
@@ -15,7 +15,6 @@ REGRESS = injection_points hashagg reindex_conc
REGRESS_OPTS = --dlpath=$(top_builddir)/src/test/regress
ISOLATION = basic inplace syscache-update-pruned repack
-ISOLATION_OPTS = --temp-config $(top_srcdir)/src/test/modules/injection_points/logical.conf
TAP_TESTS = 1
diff --git a/src/test/modules/injection_points/logical.conf b/src/test/modules/injection_points/logical.conf
deleted file mode 100644
index c8f264bc6c..0000000000
--- a/src/test/modules/injection_points/logical.conf
+++ /dev/null
@@ -1 +0,0 @@
-wal_level = logical
\ No newline at end of file
diff --git a/src/test/modules/injection_points/meson.build b/src/test/modules/injection_points/meson.build
index 0e3c47ba99..716e5619aa 100644
--- a/src/test/modules/injection_points/meson.build
+++ b/src/test/modules/injection_points/meson.build
@@ -50,9 +50,6 @@ tests += {
'syscache-update-pruned',
],
'runningcheck': false, # see syscache-update-pruned
- # 'repack' requires wal_level = 'logical'.
- 'regress_args': ['--temp-config', files('logical.conf')],
-
},
'tap': {
'env': {
--
2.43.5
[text/x-diff] v08-0009-Call-logical_rewrite_heap_tuple-when-applying-concur.patch (26.1K, ../127361.1740559688@localhost/10-v08-0009-Call-logical_rewrite_heap_tuple-when-applying-concur.patch)
download | inline diff:
From 36fa8657637b9d1738d03405807a2ae1799e3637 Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 26 Feb 2025 09:17:21 +0100
Subject: [PATCH 9/9] Call logical_rewrite_heap_tuple() when applying
concurrent data changes.
This was implemented for the sake of completeness, but I think it's currently
not needed. Possible use cases could be:
1. REPACK CONCURRENTLY can process system catalogs.
System catalogs are scanned using a historic snapshot during logical decoding,
and the "combo CIDs" information is needed for that. Since "combo CID" is
associated with the "file locator" and that locator is changed by REPACK, this
command must record the information on individual tuples being moved from the
old file to the new one. This is what logical_rewrite_heap_tuple() does.
However, the logical decoding subsystem currently does not support decoding of
data changes in the system catalog. Therefore, the CONCURRENTLY option cannot
be used for system catalogs.
2. REPACK CONCURRENTLY is processing a relation, but once it has released all
the locks (in order to get the exclusive lock), another backend runs REPACK
CONCURRENTLY on the same table. Since the relation is treated as a system
catalog while these commands are processing it (so it can be scanned using a
historic snapshot during the "initial load"), it is important that the 2nd
backend does not break decoding of the "combo CIDs" performed by the 1st
backend.
However, it's not practical to let multiple backends run REPACK CONCURRENTLY
on the same relation, so we forbid that.
---
src/backend/access/heap/heapam_handler.c | 2 +-
src/backend/access/heap/rewriteheap.c | 65 ++++++-----
src/backend/commands/cluster.c | 110 +++++++++++++++---
src/backend/replication/logical/decode.c | 41 ++++++-
.../pgoutput_repack/pgoutput_repack.c | 21 ++--
src/include/access/rewriteheap.h | 5 +-
src/include/commands/cluster.h | 3 +
src/include/replication/reorderbuffer.h | 7 ++
8 files changed, 194 insertions(+), 60 deletions(-)
diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index beec45b18e..f8528b3acd 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -730,7 +730,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
/* Initialize the rewrite operation */
rwstate = begin_heap_rewrite(OldHeap, NewHeap, OldestXmin, *xid_cutoff,
- *multi_cutoff);
+ *multi_cutoff, true);
/* Set up sorting if wanted */
diff --git a/src/backend/access/heap/rewriteheap.c b/src/backend/access/heap/rewriteheap.c
index e6d2b5fced..94b603423d 100644
--- a/src/backend/access/heap/rewriteheap.c
+++ b/src/backend/access/heap/rewriteheap.c
@@ -214,10 +214,8 @@ static void raw_heap_insert(RewriteState state, HeapTuple tup);
/* internal logical remapping prototypes */
static void logical_begin_heap_rewrite(RewriteState state);
-static void logical_rewrite_heap_tuple(RewriteState state, ItemPointerData old_tid, HeapTuple new_tuple);
static void logical_end_heap_rewrite(RewriteState state);
-
/*
* Begin a rewrite of a table
*
@@ -226,18 +224,19 @@ static void logical_end_heap_rewrite(RewriteState state);
* oldest_xmin xid used by the caller to determine which tuples are dead
* freeze_xid xid before which tuples will be frozen
* cutoff_multi multixact before which multis will be removed
+ * tid_chains need to maintain TID chains?
*
* Returns an opaque RewriteState, allocated in current memory context,
* to be used in subsequent calls to the other functions.
*/
RewriteState
begin_heap_rewrite(Relation old_heap, Relation new_heap, TransactionId oldest_xmin,
- TransactionId freeze_xid, MultiXactId cutoff_multi)
+ TransactionId freeze_xid, MultiXactId cutoff_multi,
+ bool tid_chains)
{
RewriteState state;
MemoryContext rw_cxt;
MemoryContext old_cxt;
- HASHCTL hash_ctl;
/*
* To ease cleanup, make a separate context that will contain the
@@ -262,29 +261,34 @@ begin_heap_rewrite(Relation old_heap, Relation new_heap, TransactionId oldest_xm
state->rs_cxt = rw_cxt;
state->rs_bulkstate = smgr_bulk_start_rel(new_heap, MAIN_FORKNUM);
- /* Initialize hash tables used to track update chains */
- hash_ctl.keysize = sizeof(TidHashKey);
- hash_ctl.entrysize = sizeof(UnresolvedTupData);
- hash_ctl.hcxt = state->rs_cxt;
-
- state->rs_unresolved_tups =
- hash_create("Rewrite / Unresolved ctids",
- 128, /* arbitrary initial size */
- &hash_ctl,
- HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
-
- hash_ctl.entrysize = sizeof(OldToNewMappingData);
+ if (tid_chains)
+ {
+ HASHCTL hash_ctl;
+
+ /* Initialize hash tables used to track update chains */
+ hash_ctl.keysize = sizeof(TidHashKey);
+ hash_ctl.entrysize = sizeof(UnresolvedTupData);
+ hash_ctl.hcxt = state->rs_cxt;
+
+ state->rs_unresolved_tups =
+ hash_create("Rewrite / Unresolved ctids",
+ 128, /* arbitrary initial size */
+ &hash_ctl,
+ HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
+
+ hash_ctl.entrysize = sizeof(OldToNewMappingData);
+
+ state->rs_old_new_tid_map =
+ hash_create("Rewrite / Old to new tid map",
+ 128, /* arbitrary initial size */
+ &hash_ctl,
+ HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
+ }
- state->rs_old_new_tid_map =
- hash_create("Rewrite / Old to new tid map",
- 128, /* arbitrary initial size */
- &hash_ctl,
- HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
+ logical_begin_heap_rewrite(state);
MemoryContextSwitchTo(old_cxt);
- logical_begin_heap_rewrite(state);
-
return state;
}
@@ -303,12 +307,15 @@ end_heap_rewrite(RewriteState state)
* Write any remaining tuples in the UnresolvedTups table. If we have any
* left, they should in fact be dead, but let's err on the safe side.
*/
- hash_seq_init(&seq_status, state->rs_unresolved_tups);
-
- while ((unresolved = hash_seq_search(&seq_status)) != NULL)
+ if (state->rs_unresolved_tups)
{
- ItemPointerSetInvalid(&unresolved->tuple->t_data->t_ctid);
- raw_heap_insert(state, unresolved->tuple);
+ hash_seq_init(&seq_status, state->rs_unresolved_tups);
+
+ while ((unresolved = hash_seq_search(&seq_status)) != NULL)
+ {
+ ItemPointerSetInvalid(&unresolved->tuple->t_data->t_ctid);
+ raw_heap_insert(state, unresolved->tuple);
+ }
}
/* Write the last page, if any */
@@ -995,7 +1002,7 @@ logical_rewrite_log_mapping(RewriteState state, TransactionId xid,
* Perform logical remapping for a tuple that's mapped from old_tid to
* new_tuple->t_self by rewrite_heap_tuple() if necessary for the tuple.
*/
-static void
+void
logical_rewrite_heap_tuple(RewriteState state, ItemPointerData old_tid,
HeapTuple new_tuple)
{
diff --git a/src/backend/commands/cluster.c b/src/backend/commands/cluster.c
index 910ff9fa91..562d778f62 100644
--- a/src/backend/commands/cluster.c
+++ b/src/backend/commands/cluster.c
@@ -23,6 +23,7 @@
#include "access/heapam.h"
#include "access/multixact.h"
#include "access/relscan.h"
+#include "access/rewriteheap.h"
#include "access/tableam.h"
#include "access/toast_internals.h"
#include "access/transam.h"
@@ -209,17 +210,21 @@ static HeapTuple get_changed_tuple(char *change);
static void apply_concurrent_changes(RepackDecodingState *dstate,
Relation rel, ScanKey key, int nkeys,
IndexInsertState *iistate,
- struct timeval *must_complete);
+ struct timeval *must_complete,
+ RewriteState rwstate);
static void apply_concurrent_insert(Relation rel, ConcurrentChange *change,
HeapTuple tup, IndexInsertState *iistate,
- TupleTableSlot *index_slot);
+ TupleTableSlot *index_slot,
+ RewriteState rwstate);
static void apply_concurrent_update(Relation rel, HeapTuple tup,
HeapTuple tup_target,
ConcurrentChange *change,
IndexInsertState *iistate,
- TupleTableSlot *index_slot);
+ TupleTableSlot *index_slot,
+ RewriteState rwstate);
static void apply_concurrent_delete(Relation rel, HeapTuple tup_target,
- ConcurrentChange *change);
+ ConcurrentChange *change,
+ RewriteState rwstate);
static HeapTuple find_target_tuple(Relation rel, ScanKey key, int nkeys,
HeapTuple tup_key,
Snapshot snapshot,
@@ -233,7 +238,8 @@ static bool process_concurrent_changes(LogicalDecodingContext *ctx,
ScanKey ident_key,
int ident_key_nentries,
IndexInsertState *iistate,
- struct timeval *must_complete);
+ struct timeval *must_complete,
+ RewriteState rwstate);
static bool processing_time_elapsed(struct timeval *must_complete);
static IndexInsertState *get_index_insert_state(Relation relation,
Oid ident_index_id);
@@ -3184,7 +3190,7 @@ repack_decode_concurrent_changes(LogicalDecodingContext *ctx,
static void
apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
ScanKey key, int nkeys, IndexInsertState *iistate,
- struct timeval *must_complete)
+ struct timeval *must_complete, RewriteState rwstate)
{
TupleTableSlot *index_slot, *ident_slot;
HeapTuple tup_old = NULL;
@@ -3258,7 +3264,8 @@ apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
{
Assert(tup_old == NULL);
- apply_concurrent_insert(rel, &change, tup, iistate, index_slot);
+ apply_concurrent_insert(rel, &change, tup, iistate, index_slot,
+ rwstate);
pfree(tup);
}
@@ -3266,7 +3273,7 @@ apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
change.kind == CHANGE_DELETE)
{
IndexScanDesc ind_scan = NULL;
- HeapTuple tup_key;
+ HeapTuple tup_key, tup_exist_cp;
if (change.kind == CHANGE_UPDATE_NEW)
{
@@ -3308,11 +3315,23 @@ apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
if (tup_exist == NULL)
elog(ERROR, "Failed to find target tuple");
+ /*
+ * Update the mapping for xmax of the old version.
+ *
+ * Use a copy ('tup_exist' can point to shared buffer) with xmin
+ * invalid because mapping of that should have been written on
+ * insertion.
+ */
+ tup_exist_cp = heap_copytuple(tup_exist);
+ HeapTupleHeaderSetXmin(tup_exist_cp->t_data, InvalidTransactionId);
+ logical_rewrite_heap_tuple(rwstate, change.old_tid, tup_exist_cp);
+ pfree(tup_exist_cp);
+
if (change.kind == CHANGE_UPDATE_NEW)
apply_concurrent_update(rel, tup, tup_exist, &change, iistate,
- index_slot);
+ index_slot, rwstate);
else
- apply_concurrent_delete(rel, tup_exist, &change);
+ apply_concurrent_delete(rel, tup_exist, &change, rwstate);
ResetRepackCurrentXids();
@@ -3365,9 +3384,12 @@ apply_concurrent_changes(RepackDecodingState *dstate, Relation rel,
static void
apply_concurrent_insert(Relation rel, ConcurrentChange *change, HeapTuple tup,
- IndexInsertState *iistate, TupleTableSlot *index_slot)
+ IndexInsertState *iistate, TupleTableSlot *index_slot,
+ RewriteState rwstate)
{
+ HeapTupleHeader tup_hdr = tup->t_data;
Snapshot snapshot = change->snapshot;
+ ItemPointerData old_tid;
List *recheck;
/*
@@ -3377,6 +3399,9 @@ apply_concurrent_insert(Relation rel, ConcurrentChange *change, HeapTuple tup,
*/
SetRepackCurrentXids(snapshot->subxip, snapshot->subxcnt);
+ /* Remember location in the old heap. */
+ ItemPointerCopy(&tup_hdr->t_ctid, &old_tid);
+
/*
* Write the tuple into the new heap.
*
@@ -3392,6 +3417,14 @@ apply_concurrent_insert(Relation rel, ConcurrentChange *change, HeapTuple tup,
heap_insert(rel, tup, change->xid, snapshot->curcid - 1,
HEAP_INSERT_NO_LOGICAL, NULL);
+ /*
+ * Update the mapping for xmin. (xmax should be invalid). This is needed
+ * because, during the processing, the table is considered an "user
+ * catalog".
+ */
+ Assert(!TransactionIdIsValid(HeapTupleHeaderGetRawXmax(tup->t_data)));
+ logical_rewrite_heap_tuple(rwstate, old_tid, tup);
+
/*
* Update indexes.
*
@@ -3425,15 +3458,22 @@ apply_concurrent_insert(Relation rel, ConcurrentChange *change, HeapTuple tup,
static void
apply_concurrent_update(Relation rel, HeapTuple tup, HeapTuple tup_target,
ConcurrentChange *change, IndexInsertState *iistate,
- TupleTableSlot *index_slot)
+ TupleTableSlot *index_slot, RewriteState rwstate)
{
List *recheck;
LockTupleMode lockmode;
TU_UpdateIndexes update_indexes;
+ ItemPointerData tid_new_old_heap, tid_old_new_heap;
TM_Result res;
Snapshot snapshot = change->snapshot;
TM_FailureData tmfd;
+ /* Location of the new tuple in the old heap. */
+ ItemPointerCopy(&tup->t_data->t_ctid, &tid_new_old_heap);
+
+ /* Location of the existing tuple in the new heap. */
+ ItemPointerCopy(&tup_target->t_self, &tid_old_new_heap);
+
/*
* Write the new tuple into the new heap. ('tup' gets the TID assigned
* here.)
@@ -3443,7 +3483,7 @@ apply_concurrent_update(Relation rel, HeapTuple tup, HeapTuple tup_target,
Assert(snapshot->curcid != InvalidCommandId &&
snapshot->curcid > FirstCommandId);
- res = heap_update(rel, &tup_target->t_self, tup,
+ res = heap_update(rel, &tid_old_new_heap, tup,
change->xid, snapshot->curcid - 1,
InvalidSnapshot,
false, /* no wait - only we are doing changes */
@@ -3453,6 +3493,10 @@ apply_concurrent_update(Relation rel, HeapTuple tup, HeapTuple tup_target,
if (res != TM_Ok)
ereport(ERROR, (errmsg("failed to apply concurrent UPDATE")));
+ /* Update the mapping for xmin of the new version. */
+ Assert(!TransactionIdIsValid(HeapTupleHeaderGetRawXmax(tup->t_data)));
+ logical_rewrite_heap_tuple(rwstate, tid_new_old_heap, tup);
+
ExecStoreHeapTuple(tup, index_slot, false);
if (update_indexes != TU_None)
@@ -3476,8 +3520,9 @@ apply_concurrent_update(Relation rel, HeapTuple tup, HeapTuple tup_target,
static void
apply_concurrent_delete(Relation rel, HeapTuple tup_target,
- ConcurrentChange *change)
+ ConcurrentChange *change, RewriteState rwstate)
{
+ ItemPointerData tid_old_new_heap;
TM_Result res;
TM_FailureData tmfd;
Snapshot snapshot = change->snapshot;
@@ -3486,7 +3531,10 @@ apply_concurrent_delete(Relation rel, HeapTuple tup_target,
Assert(snapshot->curcid != InvalidCommandId &&
snapshot->curcid > FirstCommandId);
- res = heap_delete(rel, &tup_target->t_self, change->xid,
+ /* Location of the existing tuple in the new heap. */
+ ItemPointerCopy(&tup_target->t_self, &tid_old_new_heap);
+
+ res = heap_delete(rel, &tid_old_new_heap, change->xid,
snapshot->curcid - 1, InvalidSnapshot, false,
&tmfd, false,
/* wal_logical */
@@ -3567,7 +3615,8 @@ static bool
process_concurrent_changes(LogicalDecodingContext *ctx, XLogRecPtr end_of_wal,
Relation rel_dst, Relation rel_src, ScanKey ident_key,
int ident_key_nentries, IndexInsertState *iistate,
- struct timeval *must_complete)
+ struct timeval *must_complete,
+ RewriteState rwstate)
{
RepackDecodingState *dstate;
@@ -3600,7 +3649,8 @@ process_concurrent_changes(LogicalDecodingContext *ctx, XLogRecPtr end_of_wal,
rel_dst->rd_toastoid = rel_src->rd_rel->reltoastrelid;
apply_concurrent_changes(dstate, rel_dst, ident_key,
- ident_key_nentries, iistate, must_complete);
+ ident_key_nentries, iistate, must_complete,
+ rwstate);
}
PG_FINALLY();
{
@@ -3785,6 +3835,7 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
bool is_system_catalog;
Oid ident_idx_old, ident_idx_new;
IndexInsertState *iistate;
+ RewriteState rwstate;
ScanKey ident_key;
int ident_key_nentries;
XLogRecPtr wal_insert_ptr, end_of_wal;
@@ -3870,11 +3921,26 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
* Apply concurrent changes first time, to minimize the time we need to
* hold AccessExclusiveLock. (Quite some amount of WAL could have been
* written during the data copying and index creation.)
+ *
+ * Now we are processing individual tuples, so pass false for
+ * 'tid_chains'. Since rwstate is now only needed for
+ * logical_begin_heap_rewrite(), none of the transaction IDs needs to be
+ * valid.
*/
+ rwstate = begin_heap_rewrite(OldHeap, NewHeap,
+ InvalidTransactionId,
+ InvalidTransactionId,
+ InvalidTransactionId,
+ false);
process_concurrent_changes(ctx, end_of_wal, NewHeap,
swap_toast_by_content ? OldHeap : NULL,
ident_key, ident_key_nentries, iistate,
- NULL);
+ NULL, rwstate);
+ /*
+ * OldHeap will be closed, so we need to initialize rwstate again for the
+ * next call of process_concurrent_changes().
+ */
+ end_heap_rewrite(rwstate);
/*
* Release the locks that allowed concurrent data changes, in order to
@@ -3996,6 +4062,11 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
end_of_wal = GetFlushRecPtr(NULL);
/* Apply the concurrent changes again. */
+ rwstate = begin_heap_rewrite(OldHeap, NewHeap,
+ InvalidTransactionId,
+ InvalidTransactionId,
+ InvalidTransactionId,
+ false);
/*
* This time we have the exclusive lock on the table, so make sure that
* repack_max_xlock_time is not exceeded.
@@ -4023,11 +4094,12 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
if (!process_concurrent_changes(ctx, end_of_wal, NewHeap,
swap_toast_by_content ? OldHeap : NULL,
ident_key, ident_key_nentries, iistate,
- t_end_ptr))
+ t_end_ptr, rwstate))
ereport(ERROR,
(errmsg("could not process concurrent data changes in time"),
errhint("Please consider adjusting \"repack_max_xlock_time\".")));
+ end_heap_rewrite(rwstate);
/* Remember info about rel before closing OldHeap */
relpersistence = OldHeap->rd_rel->relpersistence;
diff --git a/src/backend/replication/logical/decode.c b/src/backend/replication/logical/decode.c
index 55abda75d1..973867c58a 100644
--- a/src/backend/replication/logical/decode.c
+++ b/src/backend/replication/logical/decode.c
@@ -983,11 +983,13 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
xl_heap_insert *xlrec;
ReorderBufferChange *change;
RelFileLocator target_locator;
+ BlockNumber blknum;
+ HeapTupleHeader tuphdr;
xlrec = (xl_heap_insert *) XLogRecGetData(r);
/* only interested in our database */
- XLogRecGetBlockTag(r, 0, &target_locator, NULL, NULL);
+ XLogRecGetBlockTag(r, 0, &target_locator, NULL, &blknum);
if (target_locator.dbOid != ctx->slot->data.database)
return;
@@ -1012,6 +1014,13 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
DecodeXLogTuple(tupledata, datalen, change->data.tp.newtuple);
+ /*
+ * CTID is needed for logical_rewrite_heap_tuple(), when doing REPACK
+ * CONCURRENTLY.
+ */
+ tuphdr = change->data.tp.newtuple->t_data;
+ ItemPointerSet(&tuphdr->t_ctid, blknum, xlrec->offnum);
+
change->data.tp.clear_toast_afterwards = true;
ReorderBufferQueueChange(ctx->reorder, XLogRecGetXid(r), buf->origptr,
@@ -1033,11 +1042,14 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
ReorderBufferChange *change;
char *data;
RelFileLocator target_locator;
+ BlockNumber old_blknum, new_blknum;
xlrec = (xl_heap_update *) XLogRecGetData(r);
+ /* Retrieve blknum, so that we can compose CTID below. */
+ XLogRecGetBlockTag(r, 0, &target_locator, NULL, &new_blknum);
+
/* only interested in our database */
- XLogRecGetBlockTag(r, 0, &target_locator, NULL, NULL);
if (target_locator.dbOid != ctx->slot->data.database)
return;
@@ -1054,6 +1066,7 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
{
Size datalen;
Size tuplelen;
+ HeapTupleHeader tuphdr;
data = XLogRecGetBlockData(r, 0, &datalen);
@@ -1063,6 +1076,13 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
ReorderBufferGetTupleBuf(ctx->reorder, tuplelen);
DecodeXLogTuple(data, datalen, change->data.tp.newtuple);
+
+ /*
+ * CTID is needed for logical_rewrite_heap_tuple(), when doing REPACK
+ * CONCURRENTLY.
+ */
+ tuphdr = change->data.tp.newtuple->t_data;
+ ItemPointerSet(&tuphdr->t_ctid, new_blknum, xlrec->new_offnum);
}
if (xlrec->flags & XLH_UPDATE_CONTAINS_OLD)
@@ -1081,6 +1101,14 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
DecodeXLogTuple(data, datalen, change->data.tp.oldtuple);
}
+ /*
+ * Remember the old tuple CTID, for the sake of
+ * logical_rewrite_heap_tuple().
+ */
+ if (!XLogRecGetBlockTagExtended(r, 1, NULL, NULL, &old_blknum, NULL))
+ old_blknum = new_blknum;
+ ItemPointerSet(&change->data.tp.old_tid, old_blknum, xlrec->old_offnum);
+
change->data.tp.clear_toast_afterwards = true;
ReorderBufferQueueChange(ctx->reorder, XLogRecGetXid(r), buf->origptr,
@@ -1099,11 +1127,12 @@ DecodeDelete(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
xl_heap_delete *xlrec;
ReorderBufferChange *change;
RelFileLocator target_locator;
+ BlockNumber blknum;
xlrec = (xl_heap_delete *) XLogRecGetData(r);
/* only interested in our database */
- XLogRecGetBlockTag(r, 0, &target_locator, NULL, NULL);
+ XLogRecGetBlockTag(r, 0, &target_locator, NULL, &blknum);
if (target_locator.dbOid != ctx->slot->data.database)
return;
@@ -1135,6 +1164,12 @@ DecodeDelete(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
DecodeXLogTuple((char *) xlrec + SizeOfHeapDelete,
datalen, change->data.tp.oldtuple);
+
+ /*
+ * CTID is needed for logical_rewrite_heap_tuple(), when doing REPACK
+ * CONCURRENTLY.
+ */
+ ItemPointerSet(&change->data.tp.old_tid, blknum, xlrec->offnum);
}
change->data.tp.clear_toast_afterwards = true;
diff --git a/src/backend/replication/pgoutput_repack/pgoutput_repack.c b/src/backend/replication/pgoutput_repack/pgoutput_repack.c
index d42d93a8b6..71b010c351 100644
--- a/src/backend/replication/pgoutput_repack/pgoutput_repack.c
+++ b/src/backend/replication/pgoutput_repack/pgoutput_repack.c
@@ -33,7 +33,7 @@ static void plugin_truncate(struct LogicalDecodingContext *ctx,
ReorderBufferChange *change);
static void store_change(LogicalDecodingContext *ctx,
ConcurrentChangeKind kind, HeapTuple tuple,
- TransactionId xid);
+ TransactionId xid, ItemPointer old_tid);
void
_PG_output_plugin_init(OutputPluginCallbacks *cb)
@@ -168,7 +168,8 @@ plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
if (newtuple == NULL)
elog(ERROR, "Incomplete insert info.");
- store_change(ctx, CHANGE_INSERT, newtuple, change->txn->xid);
+ store_change(ctx, CHANGE_INSERT, newtuple, change->txn->xid,
+ NULL);
}
break;
case REORDER_BUFFER_CHANGE_UPDATE:
@@ -186,10 +187,10 @@ plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
if (oldtuple != NULL)
store_change(ctx, CHANGE_UPDATE_OLD, oldtuple,
- change->txn->xid);
+ change->txn->xid, NULL);
store_change(ctx, CHANGE_UPDATE_NEW, newtuple,
- change->txn->xid);
+ change->txn->xid, &change->data.tp.old_tid);
}
break;
case REORDER_BUFFER_CHANGE_DELETE:
@@ -202,7 +203,8 @@ plugin_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
if (oldtuple == NULL)
elog(ERROR, "Incomplete delete info.");
- store_change(ctx, CHANGE_DELETE, oldtuple, change->txn->xid);
+ store_change(ctx, CHANGE_DELETE, oldtuple, change->txn->xid,
+ &change->data.tp.old_tid);
}
break;
default:
@@ -236,13 +238,13 @@ plugin_truncate(struct LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
if (i == nrelations)
return;
- store_change(ctx, CHANGE_TRUNCATE, NULL, InvalidTransactionId);
+ store_change(ctx, CHANGE_TRUNCATE, NULL, InvalidTransactionId, NULL);
}
/* Store concurrent data change. */
static void
store_change(LogicalDecodingContext *ctx, ConcurrentChangeKind kind,
- HeapTuple tuple, TransactionId xid)
+ HeapTuple tuple, TransactionId xid, ItemPointer old_tid)
{
RepackDecodingState *dstate;
char *change_raw;
@@ -315,6 +317,11 @@ store_change(LogicalDecodingContext *ctx, ConcurrentChangeKind kind,
change.snapshot = dstate->snapshot;
dstate->snapshot->active_count++;
+ if (old_tid)
+ ItemPointerCopy(old_tid, &change.old_tid);
+ else
+ ItemPointerSetInvalid(&change.old_tid);
+
/* The data has been copied. */
if (flattened)
pfree(tuple);
diff --git a/src/include/access/rewriteheap.h b/src/include/access/rewriteheap.h
index 99c3f362ad..eebda35c7c 100644
--- a/src/include/access/rewriteheap.h
+++ b/src/include/access/rewriteheap.h
@@ -23,11 +23,14 @@ typedef struct RewriteStateData *RewriteState;
extern RewriteState begin_heap_rewrite(Relation old_heap, Relation new_heap,
TransactionId oldest_xmin, TransactionId freeze_xid,
- MultiXactId cutoff_multi);
+ MultiXactId cutoff_multi, bool tid_chains);
extern void end_heap_rewrite(RewriteState state);
extern void rewrite_heap_tuple(RewriteState state, HeapTuple old_tuple,
HeapTuple new_tuple);
extern bool rewrite_heap_dead_tuple(RewriteState state, HeapTuple old_tuple);
+extern void logical_rewrite_heap_tuple(RewriteState state,
+ ItemPointerData old_tid,
+ HeapTuple new_tuple);
/*
* On-Disk data format for an individual logical rewrite mapping.
diff --git a/src/include/commands/cluster.h b/src/include/commands/cluster.h
index 3ed3066b36..db029c62cf 100644
--- a/src/include/commands/cluster.h
+++ b/src/include/commands/cluster.h
@@ -78,6 +78,9 @@ typedef struct ConcurrentChange
/* Transaction that changes the data. */
TransactionId xid;
+ /* For UPDATE / DELETE, the location of the old tuple version. */
+ ItemPointerData old_tid;
+
/*
* Historic catalog snapshot that was used to decode this change.
*/
diff --git a/src/include/replication/reorderbuffer.h b/src/include/replication/reorderbuffer.h
index 517a8e3634..d0b1b48ef0 100644
--- a/src/include/replication/reorderbuffer.h
+++ b/src/include/replication/reorderbuffer.h
@@ -104,6 +104,13 @@ typedef struct ReorderBufferChange
HeapTuple oldtuple;
/* valid for INSERT || UPDATE */
HeapTuple newtuple;
+
+ /*
+ * REPACK CONCURRENTLY needs the old TID, even if the old tuple
+ * itself is not WAL-logged (i.e. when the identity key does not
+ * change).
+ */
+ ItemPointerData old_tid;
} tp;
/*
--
2.43.5
view thread (88+ messages) latest in thread
Message-ID: <127361.1740559688@localhost>
Permalink: ../127361.1740559688@localhost/
Also on: postgresql.org/message-id/127361.1740559688@localhost
reply
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-hackers@postgresql.org
Cc: ah@cybertec.at, marcos@f10.com.br, alvherre@alvh.no-ip.org, mbanck@gmx.net, zhjwpku@gmail.com, reshkekirill@gmail.com, pavel.stehule@gmail.com, michael@paquier.xyz
Subject: Re: why there is not VACUUM FULL CONCURRENTLY?
In-Reply-To: <127361.1740559688@localhost>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox