agora inbox for pgsql-hackers@postgresql.org  
help / color / mirror / Atom feed
From: Antonin Houska <ah@cybertec.at>
To: pgsql-hackers@lists.postgresql.org
Subject: Re: REPACK enhancements
Date: Wed, 15 Jul 2026 10:47:28 +0200
Message-ID: <108776.1784105248@localhost> (raw)
In-Reply-To: <109367.1781614382@localhost>
References: <109367.1781614382@localhost>

Antonin Houska <ah@cybertec.at> wrote:

> This patch set is for the next development cycle. It tries to relax some
> limitations of the REPACK (CONCURRENTLY) command.

This is the patch set rebased. (It also fixes race conditions in
repack_snapshots.spec that I haven't encountered before).

-- 
Antonin Houska
Web: https://www.cybertec-postgresql.com
From 56e791da5b1d17d0798bfc71da7b4ef95836e8e6 Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 15 Jul 2026 10:37:12 +0200
Subject: [PATCH 4/8] Use multiple snapshots to copy the data.

REPACK (CONCURRENTLY) does not prevent applications from using the table that
is being processed, however it can prevent the xmin horizon from advancing and
thus restrict VACUUM for the whole database. This patch adds the ability to
use particular snapshot only for certain range of pages. Each time that range
is processed, a new snapshot is built, which supposedly has its xmin higher
than the previous snapshot.

Note that we still use the same XID throughout the REPACK execution, and that
also prevents the xmin horizon from advancing. The following part of this
series will fix the problem.

To use multiple snapshots, the data copying works as follows:

  1. Have the logical decoding system build a snapshot S0 for range R0 at
     LSN0. This snapshot sees all the data changes whose commit records have
     LSN < LSN0.

  2. Copy the pages in that range to the new relation. The changes not visible
     to the snapshot (because their transactions are still running from the
     POV of the snapshot) will appear in the output of the logical decoding
     system as soon as their commit records are decoded.

  3. Perform logical decoding of all changes we find in WAL for the table
     we're repacking, but only apply those that affect the range R0 in the old
     relation. (Naturally, we cannot apply ones that belong to other pages
     because it's impossible to UPDATE / DELETE a row in the new relation if
     it hasn't been copied yet.) Then consider LSN1 to be the position of the
     end of the last WAL record decoded.

  4. Build a new snapshot S1 at position LSN1, i.e. one that sees all the data
     whose commit records are at WAL positions < LSN1. Use this snapshot to
     copy the range of pages R1.

  5. Perform logical decoding like in step 3, but out of this next set, only
     apply changes belonging to ranges R0 *and* R1 in the old table.

  6. etc

Special attention needs to be paid to UPDATES that span page ranges. For
example, if the old tuple is in range R0, but the new tuple is in R1, and R1
hasn't been copied yet, we only DELETE the old version from the new
relation. The new version will be handled during processing of range R1. The
snapshot S1 will be based on WAL position following that UPDATE, so it'll see
the new tuple if its transaction's commit record is at WAL position lower than
the position where we built the snapshot. On the other hand, if the commit
record appears at higher position than the that of the snapshot, the insertion
of the new tuple will be decoded and replayed later (after the copying of
range R1 has completed).

Likewise, if the old tuple is in range R1 (not yet copied) but the new tuple
is in R0, we only perform INSERT on the new relation. The deletion of the old
version will either be visible to the snapshot S1 (i.e. the snapshot won't see
the old version), or replayed later.

Due to these cross-range UPDATEs, we must apply the changes pertaining to
given range before processing of the next range starts. Specifically, if
UPDATE becomes DELETE for specific range, that DELETE must be replayed soon
enough so that we don't see both old and new tuple when building the identity
index. The problem is that if the UPDATE does not change the identity key,
we'd end up with duplicate key values.

Even if the USING INDEX clause is specified, a sequential scan is used to
retrieve the tuples from the old relation: the approach described above
requires that the tuples are in CTID order. For sorting we use a regular table
("auxiliary table"), on which we create the clustering index and scan it. The
scan output is inserted into the new relation. Tuplesort is not appropriate
here because it has no identity index, so it's not possible to apply the
decoded changes to it "eagerly", as explained above.

A new GUC repack_snapshot_after can be used to set the number of pages per
snapshot. It's currently classified as DEVELOPER_OPTIONS and may be replaced
by a constant after enough evaluation is done.
---
 src/backend/access/heap/heapam_handler.c      | 238 ++++-
 src/backend/commands/repack.c                 | 876 +++++++++++++-----
 src/backend/commands/repack_worker.c          |  97 +-
 src/backend/replication/logical/decode.c      |  83 +-
 src/backend/replication/logical/logical.c     |  30 +-
 .../replication/logical/reorderbuffer.c       |  40 +
 src/backend/replication/logical/snapbuild.c   |  30 +-
 src/backend/replication/pgrepack/pgrepack.c   |  24 +-
 src/backend/utils/misc/guc_parameters.dat     |  10 +
 src/backend/utils/misc/guc_tables.c           |   1 +
 src/include/access/tableam.h                  |  14 +-
 src/include/commands/repack.h                 |  64 +-
 src/include/commands/repack_internal.h        |  13 +-
 src/include/replication/logical.h             |   2 +-
 src/include/replication/reorderbuffer.h       |   7 +
 src/test/modules/injection_points/Makefile    |   1 +
 .../expected/repack_snapshots.out             | 401 ++++++++
 src/test/modules/injection_points/meson.build |   1 +
 .../specs/repack_snapshots.spec               | 272 ++++++
 19 files changed, 1906 insertions(+), 298 deletions(-)
 create mode 100644 src/test/modules/injection_points/expected/repack_snapshots.out
 create mode 100644 src/test/modules/injection_points/specs/repack_snapshots.spec

diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index 97f850ebf73..1096d9d4dc9 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -43,12 +43,16 @@
 #include "storage/lmgr.h"
 #include "storage/lock.h"
 #include "storage/predicate.h"
+#include "storage/proc.h"
 #include "storage/procarray.h"
 #include "storage/smgr.h"
 #include "utils/builtins.h"
+#include "utils/injection_point.h"
 #include "utils/rel.h"
 #include "utils/tuplesort.h"
 
+static Snapshot finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
+									 BlockNumber start, BlockNumber *end_p);
 static void reform_and_rewrite_tuple(TupleTableSlot *src, TupleTableSlot *reform,
 									 RewriteState rwstate);
 
@@ -583,15 +587,14 @@ static void
 heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 								 Relation OldIndex, bool use_sort,
 								 TransactionId OldestXmin,
-								 Snapshot snapshot,
 								 TransactionId *xid_cutoff,
 								 MultiXactId *multi_cutoff,
 								 double *num_tuples,
 								 double *tups_vacuumed,
-								 double *tups_recently_dead)
+								 double *tups_recently_dead,
+								 void *tableam_data)
 {
 	RewriteState rwstate;
-	BulkInsertState bistate;
 	IndexScanDesc indexScan;
 	TableScanDesc tableScan;
 	HeapScanDesc heapScan;
@@ -602,7 +605,11 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	TupleTableSlot *reform_slot;
 	BufferHeapTupleTableSlot *hslot;
 	BlockNumber prev_cblock = InvalidBlockNumber;
-	bool		concurrent = snapshot != NULL;
+	ChangeContext *chgcxt = (ChangeContext *) tableam_data;
+	bool		concurrent = chgcxt != NULL;
+	Snapshot	snapshot = NULL;
+	BlockNumber range_start = InvalidBlockNumber;
+	BlockNumber range_end = InvalidBlockNumber;
 
 	/* Remember if it's a system catalog */
 	is_system_catalog = IsSystemRelation(OldHeap);
@@ -623,14 +630,11 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	else
 		rwstate = NULL;
 
-	/* In concurrent mode, prepare for bulk-insert operation. */
-	if (concurrent)
-		bistate = GetBulkInsertState();
-	else
-		bistate = NULL;
-
-	/* Set up sorting if wanted */
-	if (use_sort)
+	/*
+	 * Set up sorting if wanted. CONCURRENTLY sorts the tuple w/o tuplesort,
+	 * see below.
+	 */
+	if (use_sort && !concurrent)
 		tuplesort = tuplesort_begin_cluster(oldTupDesc, OldIndex,
 											maintenance_work_mem,
 											NULL, TUPLESORT_NONE);
@@ -642,8 +646,11 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	 * that still need to be copied, we scan with SnapshotAny and use
 	 * HeapTupleSatisfiesVacuum for the visibility test.
 	 *
-	 * In the CONCURRENTLY case, we do regular MVCC visibility tests, using
-	 * the snapshot passed by the caller.
+	 * In the CONCURRENTLY case, we do regular MVCC visibility tests. The
+	 * snapshot changes several times during the scan so that we do not block
+	 * the progress of the xmin horizon for VACUUM too much.  Index scan
+	 * should not be used because it returns tuples in random order, which
+	 * makes it impossible to split the scan into block ranges.
 	 */
 	if (OldIndex != NULL && !use_sort)
 	{
@@ -653,6 +660,8 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		};
 		int64		ci_val[2];
 
+		Assert(!concurrent);
+
 		/* Set phase and OIDOldIndex to columns */
 		ci_val[0] = PROGRESS_REPACK_PHASE_INDEX_SCAN_HEAP;
 		ci_val[1] = RelationGetRelid(OldIndex);
@@ -660,10 +669,8 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 		tableScan = NULL;
 		heapScan = NULL;
-		indexScan = index_beginscan(OldHeap, OldIndex,
-									snapshot ? snapshot : SnapshotAny,
-									NULL, 0, 0,
-									SO_NONE);
+		indexScan = index_beginscan(OldHeap, OldIndex, SnapshotAny, NULL, 0,
+									0, SO_NONE);
 		index_rescan(indexScan, NULL, 0, NULL, 0);
 	}
 	else
@@ -672,16 +679,28 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
 									 PROGRESS_REPACK_PHASE_SEQ_SCAN_HEAP);
 
-		tableScan = table_beginscan(OldHeap,
-									snapshot ? snapshot : SnapshotAny,
-									0, (ScanKey) NULL,
+		tableScan = table_beginscan(OldHeap, SnapshotAny, 0, (ScanKey) NULL,
 									SO_NONE);
 		heapScan = (HeapScanDesc) tableScan;
+
+		/*
+		 * In CONCURRENTLY mode we scan the table by ranges of blocks and the
+		 * algorithm below expects forward direction. (No other direction
+		 * should be set here regardless concurrently anyway.)
+		 */
+		Assert(heapScan->rs_dir == ForwardScanDirection || !concurrent);
 		indexScan = NULL;
 
 		/* Set total heap blocks */
 		pgstat_progress_update_param(PROGRESS_REPACK_TOTAL_HEAP_BLKS,
 									 heapScan->rs_nblocks);
+
+		/* Setup the first range. */
+		if (concurrent)
+		{
+			range_start = heapScan->rs_startblock;
+			range_end = range_start + repack_pages_per_snapshot;
+		}
 	}
 
 	slot = table_slot_create(OldHeap, NULL);
@@ -689,6 +708,30 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	reform_slot = MakeSingleTupleTableSlot(RelationGetDescr(OldHeap),
 										   &TTSOpsVirtual);
 
+	if (concurrent)
+	{
+		/*
+		 * Do not block the progress of xmin horizons.
+		 */
+		PopActiveSnapshot();
+		InvalidateCatalogSnapshot();
+
+		/*
+		 * As there is no snapshot, our xmin should be invalid now.
+		 *
+		 * XXX xid can still be valid. The next patches in the series fix
+		 * that.
+		 */
+		Assert(!TransactionIdIsValid(MyProc->xmin));
+
+		/*
+		 * Wait until the worker has the initial snapshot and retrieve it.
+		 */
+		snapshot = repack_get_snapshot(chgcxt);
+
+		PushActiveSnapshot(snapshot);
+	}
+
 	/*
 	 * Scan through the OldHeap, either in OldIndex order or sequentially;
 	 * copy each tuple into the NewHeap, or transiently to the tuplesort
@@ -705,6 +748,9 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 		if (indexScan != NULL)
 		{
+			/* See above. */
+			Assert(!concurrent);
+
 			if (!index_getnext_slot(indexScan, ForwardScanDirection, slot))
 				break;
 
@@ -726,6 +772,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 				 */
 				pgstat_progress_update_param(PROGRESS_REPACK_HEAP_BLKS_SCANNED,
 											 heapScan->rs_nblocks);
+
 				break;
 			}
 
@@ -842,10 +889,39 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 				continue;
 			}
 		}
+		else
+		{
+			BlockNumber blkno;
+			bool		visible;
+
+			/*
+			 * With CONCURRENTLY, we use each snapshot only for certain range
+			 * of pages, so that VACUUM does not get blocked for too long. So
+			 * first check if the tuple falls into the current range.
+			 */
+			blkno = BufferGetBlockNumber(buf);
+
+			Assert(BlockNumberIsValid(range_end));
+
+			/* End of the current range or wraparound? */
+			if (blkno >= range_end || blkno < range_start)
+				snapshot = finalize_block_range(chgcxt, blkno, range_start,
+												&range_end);
+
+			/* Finally check the tuple visibility. */
+			LockBuffer(buf, BUFFER_LOCK_SHARE);
+			visible = HeapTupleSatisfiesVisibility(tuple, snapshot, buf);
+			LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+
+			if (!visible)
+				continue;
+		}
 
 		*num_tuples += 1;
 		if (tuplesort != NULL)
 		{
+			Assert(!concurrent);
+
 			tuplesort_putheaptuple(tuplesort, tuple);
 
 			/*
@@ -866,7 +942,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 			if (!concurrent)
 				reform_and_rewrite_tuple(slot, reform_slot, rwstate);
 			else
-				heap_insert_for_repack(NewHeap, slot, reform_slot, bistate);
+				heap_insert_for_repack(chgcxt, slot, reform_slot);
 
 			/*
 			 * In indexscan mode and also VACUUM FULL, report increase in
@@ -878,6 +954,28 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		}
 	}
 
+	if (concurrent)
+	{
+		XLogRecPtr	end_of_wal;
+
+		/*
+		 * Process the changes belonging to the last range.
+		 */
+		end_of_wal = GetFlushRecPtr(NULL);
+		repack_process_concurrent_changes(chgcxt, end_of_wal,
+										  InvalidBlockNumber,
+										  InvalidBlockNumber,
+										  false, false);
+
+		/*
+		 * There was an active transaction snapshot on entry, so push one
+		 * before return.
+		 */
+		PopActiveSnapshot();
+		PushActiveSnapshot(GetTransactionSnapshot());
+
+	}
+
 	if (indexScan != NULL)
 		index_endscan(indexScan);
 	if (tableScan != NULL)
@@ -923,10 +1021,13 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 			ExecStoreHeapTuple(tuple, slot, false);
 
 			n_tuples += 1;
-			if (!concurrent)
-				reform_and_rewrite_tuple(slot, reform_slot, rwstate);
-			else
-				heap_insert_for_repack(NewHeap, slot, reform_slot, bistate);
+
+			/*
+			 * The CONCURRENTLY mode uses auxiliary tables rather than
+			 * tuplesort.
+			 */
+			Assert(!concurrent);
+			reform_and_rewrite_tuple(slot, reform_slot, rwstate);
 
 			/* Report n_tuples */
 			pgstat_progress_update_param(PROGRESS_REPACK_HEAP_TUPLES_INSERTED,
@@ -942,8 +1043,89 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	/* Write out any remaining tuples, and fsync if needed */
 	if (rwstate)
 		end_heap_rewrite(rwstate);
-	if (bistate)
-		FreeBulkInsertState(bistate);
+}
+
+/*
+ * Finalize processing of the current block range.
+ *
+ * 'cur' is the current block, 'start' is the first block of the current
+ * range.
+ *
+ * '*end_p': on entry, the first block beyond the current range, on exit, the
+ * first block beyond the new range.
+ *
+ * Return the snapshot for the scan of the new range.
+ */
+static Snapshot
+finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
+					 BlockNumber start, BlockNumber *end_p)
+{
+	BlockNumber end = *end_p;
+	XLogRecPtr	end_of_wal;
+	Snapshot	snapshot;
+
+	/*
+	 * Wait here when testing how snapshot is changed at page boundary.
+	 */
+	INJECTION_POINT("repack-concurrently-new-range", NULL);
+
+	/*
+	 * Decode all the concurrent data changes committed so far before
+	 * requesting the next snapshot - these changes are applicable on top of
+	 * the current snapshot. Since we only copied part of the table so far,
+	 * only changes applicable to that part can be applied.
+	 *
+	 * It's important to apply the changes before we start copying the next
+	 * range of blocks. Without that, in case of concurrent UPDATE, we could
+	 * end up with both old and new tuple present in the new table: the old
+	 * still visible in the current range and the new already visible in the
+	 * following range (for which we'll use more recent snapshot). Thus it'd
+	 * be non-trivial to apply the UPDATE later. By replaying it now, we get
+	 * rid of the old tuple in the current range.
+	 */
+	end_of_wal = GetFlushRecPtr(NULL);
+	repack_process_concurrent_changes(chgcxt, end_of_wal, start, end, true,
+									  false);
+
+	/*
+	 * A new snapshot will be pushed below. Note that it's important to not do
+	 * this earlier, because - while processing the concurrent data changes -
+	 * we might have needed to fetch TOASTed values from the old relation -
+	 * see the UPDATE-to-INSERT conversion in apply_concurrent_changes(). As
+	 * this snapshot protects the data copied from VACUUM, it should also
+	 * protect the TOAST values referenced by the consequent UPDATE
+	 * statements.
+	 */
+	PopActiveSnapshot();
+	InvalidateCatalogSnapshot();
+
+	/* See above. */
+	Assert(!TransactionIdIsValid(MyProc->xmin));
+
+	/*
+	 * XXX It might be worth Assert(CatalogSnapshot == NULL) here, however
+	 * that symbol is not external.
+	 */
+
+	/*
+	 * Compute the end of the new range by aligning 'cur' to a multiple of
+	 * range boundary. This accounts for the possibility that some block
+	 * numbers could have been skipped (due to pages being empty) or that the
+	 * block number could have wrapped around.
+	 */
+	end = cur + repack_pages_per_snapshot - (cur % repack_pages_per_snapshot);
+	*end_p = end;
+
+	/*
+	 * Get the snapshot for the next range - it should have been built at the
+	 * position right after the last change decoded. Data present in the next
+	 * range of blocks will either be visible to the snapshot or appear in the
+	 * next batch of decoded changes.
+	 */
+	snapshot = repack_get_snapshot(chgcxt);
+	PushActiveSnapshot(snapshot);
+
+	return snapshot;
 }
 
 /*
diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 392332b4b2b..0bf19d07db5 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -33,6 +33,7 @@
 #include "postgres.h"
 
 #include "access/amapi.h"
+#include "access/detoast.h"
 #include "access/heapam.h"
 #include "access/multixact.h"
 #include "access/relscan.h"
@@ -95,6 +96,12 @@ typedef struct
 	Oid			indexOid;
 } RelToCluster;
 
+/*
+ * When REPACK (CONCURRENTLY) copies data to the new heap, a new snapshot is
+ * built after processing this many pages. XXX Tune the value.
+ */
+int			repack_pages_per_snapshot = 1024;
+
 /*
  * Backend-local information to control the decoding worker.
  */
@@ -128,11 +135,11 @@ static void check_concurrent_repack_requirements(Relation rel,
 static void rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 							 Oid ident_idx);
 static void copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
-							Snapshot snapshot,
 							bool verbose,
 							bool *pSwapToastByContent,
 							TransactionId *pFreezeXid,
-							MultiXactId *pCutoffMulti);
+							MultiXactId *pCutoffMulti,
+							ChangeContext *chgcxt);
 static List *get_tables_to_repack(RepackCommand cmd, bool usingindex,
 								  MemoryContext permcxt);
 static List *get_tables_to_repack_partitioned(RepackCommand cmd,
@@ -141,23 +148,26 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd,
 static bool repack_is_permitted_for_relation(RepackCommand cmd,
 											 Oid relid, Oid userid);
 
-static void apply_concurrent_changes(ChangeContext *chgcxt);
+static void apply_concurrent_changes(ChangeContext *chgcxt,
+									 BlockNumber range_start,
+									 BlockNumber range_end);
 static void apply_concurrent_insert(RepackDest *dest, TupleTableSlot *slot);
 static void apply_concurrent_update(RepackDest *dest,
 									TupleTableSlot *spilled_tuple,
 									TupleTableSlot *ondisk_tuple);
 static void apply_concurrent_delete(Relation rel, TupleTableSlot *slot);
 static void restore_tuple(BufFile *file, Relation relation,
-						  TupleTableSlot *slot);
+						  TupleTableSlot *slot, BlockNumber *block_nr_p,
+						  BlockNumber *old_block_nr_p);
 static void adjust_toast_pointers(Relation relation, TupleTableSlot *dest,
 								  TupleTableSlot *src);
+static bool is_block_in_range(BlockNumber blknum, BlockNumber start,
+							  BlockNumber end);
 static bool find_target_tuple(RepackDest *dest, TupleTableSlot *locator,
 							  TupleTableSlot *retrieved);
-static bool identity_key_equal(RepackDest *dest, TupleTableSlot *locator,
+static bool identity_key_equal(RepackDest *dest,
+							   TupleTableSlot *locator,
 							   TupleTableSlot *candidate);
-static void process_concurrent_changes(XLogRecPtr end_of_wal,
-									   ChangeContext *chgcxt,
-									   bool done);
 static void initialize_change_context(ChangeContext *chgcxt,
 									  Relation relation,
 									  Oid ident_index_id);
@@ -168,8 +178,12 @@ static void release_change_dest(RepackDest *dest);
 static void rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 											   Oid identIdx,
 											   TransactionId frozenXid,
-											   MultiXactId cutoffMulti);
+											   MultiXactId cutoffMulti,
+											   ChangeContext *chgcxt);
+static void process_auxiliary_table(ChangeContext *chgcxt, Relation OldHeap,
+									Oid identIdx);
 static List *build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes);
+static Oid	build_new_index(Relation NewHeap, Relation OldHeap, Oid oldindex);
 static void copy_index_constraints(Relation old_index, Oid new_index_id,
 								   Oid new_heap_id);
 static void copy_attribute_defaults(Oid old_heap_oid, Oid new_heap_oid);
@@ -183,7 +197,6 @@ static Oid	determine_clustered_index(Relation rel, bool usingindex,
 static void start_repack_decoding_worker(Oid relid);
 static void stop_repack_decoding_worker(void);
 static void stop_repack_decoding_worker_cb(int code, Datum arg);
-static Snapshot get_initial_snapshot(DecodingWorker *worker);
 
 static void ProcessRepackMessage(StringInfo msg);
 static const char *RepackCommandAsString(RepackCommand cmd);
@@ -960,6 +973,14 @@ check_concurrent_repack_requirements(Relation rel, Oid *ident_idx_p)
 						RelationGetRelationName(rel)));
 	}
 
+	/*
+	 * In the CONCURRENTLY mode we don't want to use the same snapshot
+	 * throughout the whole processing, as it could block the progress of xmin
+	 * horizon. Assert should be ok as we already disallow transaction block
+	 * in the CONCURRENTLY case.
+	 */
+	Assert(!IsolationUsesXactSnapshot());
+
 	*ident_idx_p = ident_idx;
 }
 
@@ -996,7 +1017,7 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	TransactionId frozenXid;
 	MultiXactId cutoffMulti;
 	bool		concurrent = OidIsValid(ident_idx);
-	Snapshot	snapshot = NULL;
+	ChangeContext *chgcxt = NULL;
 #if USE_ASSERT_CHECKING
 	LOCKMODE	lmode;
 
@@ -1033,13 +1054,6 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 		 * REPACK CONCURRENTLY.
 		 */
 		start_repack_decoding_worker(tableOid);
-
-		/*
-		 * Wait until the worker has the initial snapshot and retrieve it.
-		 */
-		snapshot = get_initial_snapshot(decoding_worker);
-
-		PushActiveSnapshot(snapshot);
 	}
 
 	/* for CLUSTER or REPACK USING INDEX, mark the index as the one to use */
@@ -1062,26 +1076,115 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	Assert(CheckRelationOidLockedByMe(OIDNewHeap, AccessExclusiveLock, false));
 	NewHeap = table_open(OIDNewHeap, NoLock);
 
-	/*
-	 * In concurrent mode, create a copy of the attribute defaults on the temp
-	 * table, which the executor needs when replaying concurrent data changes.
-	 */
 	if (concurrent)
+	{
+		bool		need_aux_rel;
+
+		/*
+		 * Auxiliary table is needed for clustering in the CONCURRENTLY mode,
+		 * see comments in ChangeContext. FIXME Non-btree indexes are allowed
+		 * historically, but in general, these can hardly define any useful
+		 * order. We ignore them here.
+		 */
+		need_aux_rel = index != NULL && index->rd_rel->relam == BTREE_AM_OID;
+
+		/* Gather information to apply concurrent changes. */
+		chgcxt = palloc0_object(ChangeContext);
+
+		/*
+		 * Create a copy of the attribute defaults on the temp table, which
+		 * the executor needs when replaying concurrent data changes.
+		 */
 		copy_attribute_defaults(tableOid, OIDNewHeap);
 
+		if (!need_aux_rel)
+		{
+			Oid			ident_idx_new;
+
+			/*
+			 * Create the identity index. We will need it during data copying
+			 * so that we can apply the data changes at the appropriate time -
+			 * see comments around the call of
+			 * repack_process_concurrent_changes() with block range specified.
+			 *
+			 * XXX NewHeap is empty - should we pass INDEX_CREATE_SKIP_BUILD?
+			 */
+			ident_idx_new = build_new_index(NewHeap, OldHeap, ident_idx);
+
+			initialize_change_context(chgcxt, NewHeap, ident_idx_new);
+		}
+		else
+		{
+			Oid			aux_oid;
+			Relation	aux_rel;
+			Oid			aux_ident_idx;
+
+			/*
+			 * As the concurrent data changes will be applied to the auxiliary
+			 * heap, the new heap does not need the identity index yet. We'll
+			 * build it after having copied the data from the auxiliary heap.
+			 * (Bulk insert should be more efficient.)
+			 */
+			initialize_change_context(chgcxt, NewHeap, InvalidOid);
+
+			/*
+			 * Like above, but only temporary - no other backend should need
+			 * it.
+			 */
+			aux_oid = make_new_heap(tableOid, tableSpace, accessMethod,
+									RELPERSISTENCE_TEMP, NoLock);
+			Assert(CheckRelationOidLockedByMe(aux_oid, AccessExclusiveLock,
+											  false));
+			aux_rel = table_open(aux_oid, NoLock);
+
+
+			/*
+			 * Copy the attribute defaults as we did for the new heap above -
+			 * the concurrent changes also need to be applied to the auxiliary
+			 * table.
+			 */
+			copy_attribute_defaults(tableOid, aux_oid);
+
+			/*
+			 * The same for identity index. (The additional
+			 * ShareUpdateExclusiveLock on ident_idx is not a problem, it'll
+			 * be released at the end of transaction.)
+			 */
+			aux_ident_idx = build_new_index(aux_rel, OldHeap, ident_idx);
+
+			/*
+			 * Make the relation ready for use.
+			 */
+			chgcxt->cc_dest_aux = palloc0_object(RepackDest);
+			initialize_change_dest(chgcxt->cc_dest_aux, aux_rel,
+								   aux_ident_idx);
+
+			/*
+			 * Set OID of the old relation's clustering index if it's
+			 * different from the identity index. Otherwise set InvalidOid to
+			 * indicate that the identity index should be used for clustering.
+			 */
+			if (RelationGetRelid(index) != ident_idx)
+				chgcxt->cc_clustering_index = RelationGetRelid(index);
+			else
+				chgcxt->cc_clustering_index = InvalidOid;
+		}
+	}
+
 	/* Copy the heap data into the new table in the desired order */
-	copy_table_data(NewHeap, OldHeap, index, snapshot, verbose,
-					&swap_toast_by_content, &frozenXid, &cutoffMulti);
+	copy_table_data(NewHeap, OldHeap, index, verbose,
+					&swap_toast_by_content, &frozenXid, &cutoffMulti,
+					chgcxt);
 
 	/* The historic snapshot won't be needed anymore. */
-	if (snapshot)
+	if (concurrent)
 	{
-		PopActiveSnapshot();
+		/*
+		 * Make sure the active snapshot can see the data copied, so the rows
+		 * can be updated / deleted.
+		 */
 		UpdateActiveSnapshotCommandId();
-	}
 
-	if (concurrent)
-	{
 		Assert(!swap_toast_by_content);
 
 		/*
@@ -1092,10 +1195,16 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 			index_close(index, NoLock);
 
 		rebuild_relation_finish_concurrent(NewHeap, OldHeap, ident_idx,
-										   frozenXid, cutoffMulti);
+										   frozenXid, cutoffMulti, chgcxt);
 
 		pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
 									 PROGRESS_REPACK_PHASE_FINAL_CLEANUP);
+
+		/*
+		 * REPACK (CONCURRENTLY) launches separate transaction(s) so it
+		 * shouldn't rely on the current portal to pop the active snapshot.
+		 */
+		PopActiveSnapshot();
 	}
 	else
 	{
@@ -1258,10 +1367,9 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
  * Insert tuple when processing REPACK CONCURRENTLY.
  *
  * rewriteheap.c is not used in the CONCURRENTLY case because it'd be
- * difficult to do the same in the catch-up phase (as the logical
- * decoding does not provide us with sufficient visibility
- * information). Thus we must use heap_insert() both during the
- * catch-up and here.
+ * difficult to do the same in the catch-up phase (as the logical decoding
+ * does not provide us with sufficient visibility information). Thus we must
+ * use heap_insert() both during the catch-up and here.
  *
  * 'reform' is a slot to use for tuple "reforming", typically to get set
  * values of dropped columns to NULL.
@@ -1269,20 +1377,27 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
  * We pass the NO_LOGICAL flag to heap_insert() in order to skip logical
  * decoding: as soon as REPACK CONCURRENTLY swaps the relation files, it drops
  * this relation, so no logical replication subscription should need the data.
- *
- * BulkInsertState is used because many tuples are inserted in the typical
- * case.
  */
 void
-heap_insert_for_repack(Relation rel, TupleTableSlot *src,
-					   TupleTableSlot *reform, BulkInsertStateData *bistate)
+heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
+					   TupleTableSlot *reform)
 {
 	HeapTuple	tuple;
 	bool		shouldFree;
 	TupleTableSlot *slot;
+	RepackDest *dest;
+
+	/*
+	 * Use the current auxiliary table as output if one is active, otherwise
+	 * insert the tuple into the actual destination table.
+	 */
+	if (chgcxt->cc_dest_aux)
+		dest = chgcxt->cc_dest_aux;
+	else
+		dest = &chgcxt->cc_dest;
 
 	tuple = ExecFetchSlotHeapTuple(src, false, &shouldFree);
-	if (tuple_needs_reform(tuple, src->tts_tupleDescriptor))
+	if (reform != NULL && tuple_needs_reform(tuple, src->tts_tupleDescriptor))
 	{
 		clear_dropped_attributes(tuple, reform);
 		slot = reform;
@@ -1297,8 +1412,16 @@ heap_insert_for_repack(Relation rel, TupleTableSlot *src,
 	if (shouldFree)
 		heap_freetuple(tuple);
 
-	table_tuple_insert(rel, slot, GetCurrentCommandId(true),
-					   TABLE_INSERT_NO_LOGICAL, bistate);
+	table_tuple_insert(dest->rel, slot, GetCurrentCommandId(true),
+					   TABLE_INSERT_NO_LOGICAL, dest->bistate);
+
+	/*
+	 * Insert the tuple into the identity index. initialize_change_context()
+	 * may skip opening of indexes if the identity index is not needed
+	 * immediately.
+	 */
+	if (dest->rri)
+		ExecInsertIndexTuples(dest->rri, dest->estate, 0, slot, NIL, NULL);
 }
 
 bool
@@ -1348,9 +1471,6 @@ clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
 /*
  * Do the physical copying of table data.
  *
- * 'snapshot' and 'decoding_ctx': see table_relation_copy_for_cluster(). Pass
- * iff concurrent processing is required.
- *
  * There are three output parameters:
  * *pSwapToastByContent is set true if toast tables must be swapped by content.
  * *pFreezeXid receives the TransactionId used as freeze cutoff point.
@@ -1358,8 +1478,9 @@ clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
  */
 static void
 copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
-				Snapshot snapshot, bool verbose, bool *pSwapToastByContent,
-				TransactionId *pFreezeXid, MultiXactId *pCutoffMulti)
+				bool verbose, bool *pSwapToastByContent,
+				TransactionId *pFreezeXid, MultiXactId *pCutoffMulti,
+				ChangeContext *chgcxt)
 {
 	Relation	relRelation;
 	HeapTuple	reltup;
@@ -1376,7 +1497,7 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 	int			elevel = verbose ? INFO : DEBUG2;
 	PGRUsage	ru0;
 	char	   *nspname;
-	bool		concurrent = snapshot != NULL;
+	bool		concurrent = chgcxt != NULL;
 	LOCKMODE	lmode;
 
 	lmode = RepackLockLevel(concurrent);
@@ -1480,18 +1601,28 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 			cutoffs.MultiXactCutoff = relminmxid;
 	}
 
-	/*
-	 * Decide whether to use an indexscan or seqscan-and-optional-sort to scan
-	 * the OldHeap.  We know how to use a sort to duplicate the ordering of a
-	 * btree index, and will use seqscan-and-sort for that case if the planner
-	 * tells us it's cheaper.  Otherwise, always indexscan if an index is
-	 * provided, else plain seqscan.
-	 */
-	if (OldIndex != NULL && OldIndex->rd_rel->relam == BTREE_AM_OID)
-		use_sort = plan_cluster_use_sort(RelationGetRelid(OldHeap),
-										 RelationGetRelid(OldIndex));
+	if (!concurrent)
+	{
+		/*
+		 * Decide whether to use an indexscan or seqscan-and-optional-sort to
+		 * scan the OldHeap.  We know how to use a sort to duplicate the
+		 * ordering of a btree index, and will use seqscan-and-sort for that
+		 * case if the planner tells us it's cheaper.  Otherwise, always
+		 * indexscan if an index is provided, else plain seqscan.
+		 */
+		if (OldIndex != NULL && OldIndex->rd_rel->relam == BTREE_AM_OID)
+			use_sort = plan_cluster_use_sort(RelationGetRelid(OldHeap),
+											 RelationGetRelid(OldIndex));
+		else
+			use_sort = false;
+	}
 	else
-		use_sort = false;
+	{
+		/*
+		 * To use multiple snapshots, we need to read the table sequentially.
+		 */
+		use_sort = true;
+	}
 
 	/* Log what we're doing */
 	if (OldIndex != NULL && !use_sort)
@@ -1518,11 +1649,11 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 	 * values (e.g. because the AM doesn't use freezing).
 	 */
 	table_relation_copy_for_cluster(OldHeap, NewHeap, OldIndex, use_sort,
-									cutoffs.OldestXmin, snapshot,
+									cutoffs.OldestXmin,
 									&cutoffs.FreezeLimit,
 									&cutoffs.MultiXactCutoff,
 									&num_tuples, &tups_vacuumed,
-									&tups_recently_dead);
+									&tups_recently_dead, chgcxt);
 
 	/* return selected values to caller, get set as relfrozenxid/minmxid */
 	*pFreezeXid = cutoffs.FreezeLimit;
@@ -2429,6 +2560,8 @@ repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid)
  * instead return the opened and locked relcache entry, so that caller can
  * process the partitions using the multiple-table handling code.  In this
  * case, if an index name is given, it's up to the caller to resolve it.
+ *
+ * A new transaction is started in either case.
  */
 static Relation
 process_single_relation(RepackStmt *stmt, LOCKMODE lockmode, bool isTopLevel,
@@ -2441,6 +2574,31 @@ process_single_relation(RepackStmt *stmt, LOCKMODE lockmode, bool isTopLevel,
 	Assert(stmt->command == REPACK_COMMAND_CLUSTER ||
 		   stmt->command == REPACK_COMMAND_REPACK);
 
+	if (params->options & CLUOPT_CONCURRENT)
+	{
+		/*
+		 * Since REPACK (CONCURRENTLY) pops the active snapshot during the
+		 * processing (it creates and pushes snapshots on its own), and since
+		 * that snapshot can be referenced by the current portal, we need to
+		 * make sure that the portal has no dangling pointer to the snapshot.
+		 * Starting a new transaction seems to be the simplest way.
+		 *
+		 * XXX The following patches in the series make this unnecessary, as
+		 * they start new transactions for other reasons elsewhere.
+		 */
+		PopActiveSnapshot();
+		CommitTransactionCommand();
+
+		/* Start a new transaction. */
+		StartTransactionCommand();
+
+		/*
+		 * Functions in indexes may want a snapshot set. Note that the portal
+		 * is not aware of this one, so the caller needs to pop it explicitly.
+		 */
+		PushActiveSnapshot(GetTransactionSnapshot());
+	}
+
 	/*
 	 * Make sure ANALYZE is specified if a column list is present.
 	 */
@@ -2580,10 +2738,11 @@ RepackCommandAsString(RepackCommand cmd)
 }
 
 /*
- * Apply all the changes provided by decoding worker.
+ * Apply data changes that affect pages in given range.
  */
 static void
-apply_concurrent_changes(ChangeContext *chgcxt)
+apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
+						 BlockNumber range_end)
 {
 	ConcurrentChangeKind kind = '\0';
 	RepackDest *dest;
@@ -2592,18 +2751,31 @@ apply_concurrent_changes(ChangeContext *chgcxt)
 	TupleTableSlot *old_update_tuple;
 	TupleTableSlot *ondisk_tuple;
 	bool		have_old_tuple = false;
+	bool		check_range;
 	MemoryContext oldcxt;
 	DecodingWorkerShared *shared;
 	char		fname[MAXPGPATH];
 	BufFile    *file;
 
-	dest = &chgcxt->cc_dest;
+	/*
+	 * Use the auxiliary table if one exists, otherwise the "final"
+	 * destination table.
+	 */
+	dest = chgcxt->cc_dest_aux ? chgcxt->cc_dest_aux : &chgcxt->cc_dest;
 	rel = dest->rel;
 
+	/*
+	 * Range needs to be checked if the bounds are specified. Expect either
+	 * both or none.
+	 */
+	Assert(BlockNumberIsValid(range_start) == BlockNumberIsValid(range_end));
+	check_range = BlockNumberIsValid(range_start);
+
 	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
 
 	/* Open the file containing the changes. */
-	DecodingWorkerFileName(fname, shared->relid, chgcxt->cc_file_seq);
+	DecodingWorkerFileName(fname, shared->relid, chgcxt->cc_file_seq_changes,
+						   false);
 	file = BufFileOpenFileSet(&shared->sfs.fs, fname, O_RDONLY, false);
 
 	spilled_tuple = MakeSingleTupleTableSlot(RelationGetDescr(rel),
@@ -2619,6 +2791,9 @@ apply_concurrent_changes(ChangeContext *chgcxt)
 	{
 		size_t		nread;
 		ConcurrentChangeKind prevkind = kind;
+		BlockNumber block,
+					old_block;
+		BlockNumber *old_block_p;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -2633,7 +2808,7 @@ apply_concurrent_changes(ChangeContext *chgcxt)
 		 */
 		if (kind == CHANGE_UPDATE_OLD)
 		{
-			restore_tuple(file, rel, old_update_tuple);
+			restore_tuple(file, rel, old_update_tuple, NULL, NULL);
 			have_old_tuple = true;
 			continue;
 		}
@@ -2657,22 +2832,39 @@ apply_concurrent_changes(ChangeContext *chgcxt)
 
 		/*
 		 * Now restore the tuple into the slot and execute the change.
+		 *
+		 * old_block is only stored with UPDATE_NEW.
 		 */
-		restore_tuple(file, rel, spilled_tuple);
+		old_block_p = kind == CHANGE_UPDATE_NEW ? &old_block : NULL;
+		restore_tuple(file, rel, spilled_tuple, &block, old_block_p);
 
 		if (kind == CHANGE_INSERT)
 		{
-			apply_concurrent_insert(dest, spilled_tuple);
+			/*
+			 * Only insert the tuple if it fits into the current range (or if
+			 * range does not matter).
+			 */
+			if (!check_range ||
+				is_block_in_range(block, range_start, range_end))
+				apply_concurrent_insert(dest, spilled_tuple);
 		}
 		else if (kind == CHANGE_DELETE)
 		{
-			bool		found;
+			/*
+			 * Only delete the tuple if it fits into the current range (or if
+			 * range does not matter).
+			 */
+			if (!check_range ||
+				is_block_in_range(block, range_start, range_end))
+			{
+				bool		found;
 
-			/* Find the tuple to be deleted */
-			found = find_target_tuple(dest, spilled_tuple, ondisk_tuple);
-			if (!found)
-				elog(ERROR, "could not find target tuple");
-			apply_concurrent_delete(rel, ondisk_tuple);
+				/* Find the tuple to be deleted */
+				found = find_target_tuple(dest, spilled_tuple, ondisk_tuple);
+				if (!found)
+					elog(ERROR, "could not find target tuple");
+				apply_concurrent_delete(rel, ondisk_tuple);
+			}
 		}
 		else if (kind == CHANGE_UPDATE_NEW)
 		{
@@ -2684,21 +2876,71 @@ apply_concurrent_changes(ChangeContext *chgcxt)
 			else
 				key = spilled_tuple;
 
-			/* Find the tuple to be updated or deleted. */
-			found = find_target_tuple(dest, key, ondisk_tuple);
-			if (!found)
-				elog(ERROR, "could not find target tuple");
-
 			/*
-			 * If 'tup' contains TOAST pointers, they point to the old
-			 * relation's toast. Copy the corresponding TOAST pointers for the
-			 * new relation from the existing tuple. (The fact that we
-			 * received a TOAST pointer here implies that the attribute hasn't
-			 * changed.)
+			 * Perform normal update if both old and new version are in the
+			 * current range.
 			 */
-			adjust_toast_pointers(rel, spilled_tuple, ondisk_tuple);
+			if (!check_range ||
+				(is_block_in_range(old_block, range_start, range_end) &&
+				 is_block_in_range(block, range_start, range_end)))
+			{
+				/* Find the tuple to be updated or deleted. */
+				found = find_target_tuple(dest, key, ondisk_tuple);
+				if (!found)
+					elog(ERROR, "could not find target tuple");
 
-			apply_concurrent_update(dest, spilled_tuple, ondisk_tuple);
+				/*
+				 * If 'spilled_tuple' contains TOAST pointers, they point to
+				 * the old relation's toast. Copy the corresponding TOAST
+				 * pointers for the new relation from the existing tuple. (The
+				 * fact that we received a TOAST pointer here implies that the
+				 * attribute hasn't changed.)
+				 */
+				adjust_toast_pointers(rel, spilled_tuple, ondisk_tuple);
+
+				apply_concurrent_update(dest, spilled_tuple, ondisk_tuple);
+			}
+			else
+			{
+				Assert(check_range);
+
+				if (is_block_in_range(block, range_start, range_end))
+				{
+					/*
+					 * The old key is in another range, so only insert the new
+					 * one into the current range. The old version should not
+					 * be visible to the snapshot that we'll use to copy the
+					 * other range.
+					 *
+					 * Unlike UPDATE, there's no old tuple to copy the TOAST
+					 * pointers from. Therefore pass NULL for the source
+					 * tuple, to enforce detoasting of the TOAST pointers in
+					 * 'spilled_tuple'.
+					 */
+					adjust_toast_pointers(rel, spilled_tuple, NULL);
+
+					apply_concurrent_insert(dest, spilled_tuple);
+				}
+				else if (is_block_in_range(old_block, range_start, range_end))
+				{
+					found = find_target_tuple(dest, key, ondisk_tuple);
+					if (!found)
+						elog(ERROR, "could not find target tuple");
+
+					/*
+					 * The new key is in another range, so only delete the old
+					 * one from the current range. The new version should be
+					 * visible to the snapshot that we'll use to copy the
+					 * other range.
+					 */
+					apply_concurrent_delete(rel, ondisk_tuple);
+				}
+
+				/*
+				 * Otherwise, both tuple versions belong to another range, so
+				 * there's nothing to do here.
+				 */
+			}
 
 			ExecClearTuple(old_update_tuple);
 			have_old_tuple = false;
@@ -2820,7 +3062,8 @@ apply_concurrent_delete(Relation rel, TupleTableSlot *slot)
  * smaller than MaxAllocSize but the whole tuple is bigger.
  */
 static void
-restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot)
+restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot,
+			  BlockNumber *block_nr_p, BlockNumber *old_block_nr_p)
 {
 	uint32		t_len;
 	HeapTuple	tup;
@@ -2832,7 +3075,6 @@ restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot)
 	tup->t_data = (HeapTupleHeader) ((char *) tup + HEAPTUPLESIZE);
 	BufFileReadExact(file, tup->t_data, t_len);
 	tup->t_len = t_len;
-	ItemPointerSetInvalid(&tup->t_self);
 	tup->t_tableOid = RelationGetRelid(relation);
 
 	/*
@@ -2841,6 +3083,12 @@ restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot)
 	 */
 	ExecForceStoreHeapTuple(tup, slot, false);
 
+	/* Handle TID separate because not all tuple slots care about it. */
+	if (block_nr_p)
+		*block_nr_p = ItemPointerGetBlockNumber(&tup->t_data->t_ctid);
+	if (old_block_nr_p)
+		BufFileReadExact(file, old_block_nr_p, sizeof(BlockNumber));
+
 	/*
 	 * Next, read any attributes we stored separately into the tts_values
 	 * array elements expecting them, if any.  This matches
@@ -2890,10 +3138,12 @@ restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot)
 
 /*
  * Adjust 'dest' replacing any EXTERNAL_ONDISK toast pointers with the
- * corresponding ones from 'src'.
+ * corresponding ones from 'src'. If 'src' is NULL, replace the toast pointer
+ * with the actual value.
  */
 static void
-adjust_toast_pointers(Relation relation, TupleTableSlot *dest, TupleTableSlot *src)
+adjust_toast_pointers(Relation relation, TupleTableSlot *dest,
+					  TupleTableSlot *src)
 {
 	TupleDesc	desc = dest->tts_tupleDescriptor;
 
@@ -2914,9 +3164,44 @@ adjust_toast_pointers(Relation relation, TupleTableSlot *dest, TupleTableSlot *s
 		varlena_dst = (varlena *) DatumGetPointer(dest->tts_values[i]);
 		if (!VARATT_IS_EXTERNAL_ONDISK(varlena_dst))
 			continue;
-		slot_getsomeattrs(src, i + 1);
 
-		dest->tts_values[i] = src->tts_values[i];
+		/*
+		 * Ideally we just copy the value, but if there is no source tuple, we
+		 * need to detoast the value.
+		 */
+		if (src)
+		{
+			slot_getsomeattrs(src, i + 1);
+			dest->tts_values[i] = src->tts_values[i];
+		}
+		else
+		{
+			varlena    *detoasted;
+
+			detoasted = detoast_external_attr(varlena_dst);
+			dest->tts_values[i] = PointerGetDatum(detoasted);
+		}
+	}
+}
+
+/*
+ * Check if tuple originates from given range of blocks that have already been
+ * copied.
+ */
+static bool
+is_block_in_range(BlockNumber blknum, BlockNumber start, BlockNumber end)
+{
+	Assert(BlockNumberIsValid(start) && BlockNumberIsValid(end));
+	Assert(BlockNumberIsValid(blknum));
+
+	if (start < end)
+		return blknum >= start && blknum < end;
+	else
+	{
+		/* Has the scan position wrapped around? */
+		Assert(start > end);
+
+		return blknum >= start || blknum < end;
 	}
 }
 
@@ -3012,75 +3297,19 @@ identity_key_equal(RepackDest *dest, TupleTableSlot *locator,
 }
 
 /*
- * Decode and apply concurrent changes, up to (and including) the record whose
- * LSN is 'end_of_wal'.
- *
- * XXX the names "process_concurrent_changes" and "apply_concurrent_changes"
- * are far too similar to each other.
- */
-static void
-process_concurrent_changes(XLogRecPtr end_of_wal, ChangeContext *chgcxt, bool done)
-{
-	DecodingWorkerShared *shared;
-	char		fname[MAXPGPATH];
-	BufFile    *file;
-
-	pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
-								 PROGRESS_REPACK_PHASE_CATCH_UP);
-
-	/* Ask the worker for the file. */
-	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
-	SpinLockAcquire(&shared->mutex);
-	shared->lsn_upto = end_of_wal;
-	shared->done = done;
-	SpinLockRelease(&shared->mutex);
-
-	/*
-	 * The worker needs to finish processing of the current WAL record. Even
-	 * if it's idle, it'll need to close the output file. Thus we're likely to
-	 * wait, so prepare for sleep.
-	 */
-	ConditionVariablePrepareToSleep(&shared->cv);
-	for (;;)
-	{
-		int			last_exported;
-
-		SpinLockAcquire(&shared->mutex);
-		last_exported = shared->last_exported;
-		SpinLockRelease(&shared->mutex);
-
-		/*
-		 * Has the worker exported the file we are waiting for?
-		 */
-		if (last_exported == chgcxt->cc_file_seq)
-			break;
-
-		ConditionVariableSleep(&shared->cv, WAIT_EVENT_REPACK_WORKER_EXPORT);
-	}
-	ConditionVariableCancelSleep();
-
-	/* Open the file. */
-	DecodingWorkerFileName(fname, shared->relid, chgcxt->cc_file_seq);
-	file = BufFileOpenFileSet(&shared->sfs.fs, fname, O_RDONLY, false);
-	apply_concurrent_changes(chgcxt);
-
-	BufFileClose(file);
-
-	/* Get ready for the next file. */
-	chgcxt->cc_file_seq++;
-}
-
-/*
- * Initialize the ChangeContext struct for the given relation, with
- * the given index as identity index.
+ * Initialize the ChangeContext struct for the given relation.
  */
 static void
-initialize_change_context(ChangeContext *chgcxt,
-						  Relation relation, Oid ident_index_id)
+initialize_change_context(ChangeContext *chgcxt, Relation relation,
+						  Oid ident_index_id)
 {
 	initialize_change_dest(&chgcxt->cc_dest, relation, ident_index_id);
 
-	chgcxt->cc_file_seq = WORKER_FILE_SNAPSHOT + 1;
+	chgcxt->cc_file_seq_snapshot = 0;
+	chgcxt->cc_file_seq_changes = 0;
+
+	chgcxt->cc_dest_aux = NULL;
+	chgcxt->cc_clustering_index = InvalidOid;
 }
 
 /*
@@ -3090,6 +3319,8 @@ static void
 release_change_context(ChangeContext *chgcxt)
 {
 	release_change_dest(&chgcxt->cc_dest);
+	if (chgcxt->cc_dest_aux)
+		release_change_dest(chgcxt->cc_dest_aux);
 }
 
 /*
@@ -3104,6 +3335,10 @@ initialize_change_dest(RepackDest *dest, Relation relation,
 	dest->rel = relation;
 	dest->bistate = GetBulkInsertState();
 
+	/* If there's no identity index yet, there should be no indexes at all. */
+	if (!OidIsValid(ident_index_id))
+		return;
+
 	/* Only initialize fields needed by ExecInsertIndexTuples(). */
 	dest->estate = CreateExecutorState();
 
@@ -3241,6 +3476,11 @@ static void
 release_change_dest(RepackDest *dest)
 {
 	FreeBulkInsertState(dest->bistate);
+
+	/* It's possible that no indexes were opened during initialization. */
+	if (dest->rri == NULL)
+		return;
+
 	ExecCloseIndices(dest->rri);
 	FreeExecutorState(dest->estate);
 	/* XXX are these pfrees necessary? */
@@ -3259,7 +3499,8 @@ release_change_dest(RepackDest *dest)
 static void
 rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 								   Oid identIdx, TransactionId frozenXid,
-								   MultiXactId cutoffMulti)
+								   MultiXactId cutoffMulti,
+								   ChangeContext *chgcxt)
 {
 	List	   *ind_oids_new;
 	Oid			old_table_oid = RelationGetRelid(OldHeap);
@@ -3269,14 +3510,17 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 			   *lc2;
 	char		relpersistence;
 	bool		is_system_catalog;
-	Oid			ident_idx_new;
 	XLogRecPtr	end_of_wal;
 	List	   *indexrels;
-	ChangeContext chgcxt;
+	List	   *inds_tmp = NIL;
 
 	Assert(CheckRelationLockedByMe(OldHeap, ShareUpdateExclusiveLock, false));
 	Assert(CheckRelationLockedByMe(NewHeap, AccessExclusiveLock, false));
 
+	/* If we have the auxiliary table, this is the moment we should use it. */
+	if (chgcxt->cc_dest_aux)
+		process_auxiliary_table(chgcxt, OldHeap, identIdx);
+
 	/*
 	 * Unlike the exclusive case, we build new indexes for the new relation
 	 * rather than swapping the storage and reindexing the old relation. The
@@ -3292,32 +3536,26 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	 * might not be enough for commands like ALTER INDEX ... SET ... (Those
 	 * are not necessarily dangerous, but can make user confused if the
 	 * changes they do get lost due to REPACK.)
+	 *
+	 * As the identity index already had to be built, skip it here. XXX
+	 * Consider if the retail inserts during data copying (in the case w/o
+	 * auxiliary table) can be a problem in terms of index layout. Shouldn't
+	 * we drop the identity index and build it using bulk insert too?
 	 */
+	foreach_oid(ind_oid, ind_oids_old)
+	{
+		if (ind_oid != identIdx)
+			inds_tmp = lappend_oid(inds_tmp, ind_oid);
+	}
+	ind_oids_old = inds_tmp;
 	ind_oids_new = build_new_indexes(NewHeap, OldHeap, ind_oids_old);
 
 	/*
-	 * The identity index in the new relation appears in the same relative
-	 * position as the corresponding index in the old relation.  Find it.
+	 * The identity index will be involved in the following processing.
 	 */
-	ident_idx_new = InvalidOid;
-	foreach_oid(ind_old, ind_oids_old)
-	{
-		if (identIdx == ind_old)
-		{
-			int			pos = foreach_current_index(ind_old);
-
-			if (list_length(ind_oids_new) <= pos)
-				elog(ERROR, "list of new indexes too short");
-			ident_idx_new = list_nth_oid(ind_oids_new, pos);
-			break;
-		}
-	}
-	if (!OidIsValid(ident_idx_new))
-		elog(ERROR, "could not find index matching \"%s\" at the new relation",
-			 get_rel_name(identIdx));
-
-	/* Gather information to apply concurrent changes. */
-	initialize_change_context(&chgcxt, NewHeap, ident_idx_new);
+	ind_oids_old = lappend_oid(ind_oids_old, identIdx);
+	ind_oids_new = lappend_oid(ind_oids_new,
+							   RelationGetRelid(chgcxt->cc_dest.ident_index));
 
 	/*
 	 * During testing, wait for another backend to perform concurrent data
@@ -3334,11 +3572,13 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	end_of_wal = GetFlushRecPtr(NULL);
 
 	/*
-	 * Apply concurrent changes first time, to minimize the time we need to
-	 * hold AccessExclusiveLock. (Quite some amount of WAL could have been
+	 * Decode and apply concurrent changes again, to minimize the time we need
+	 * to hold AccessExclusiveLock. (Quite some amount of WAL could have been
 	 * written during the data copying and index creation.)
 	 */
-	process_concurrent_changes(end_of_wal, &chgcxt, false);
+	repack_process_concurrent_changes(chgcxt, end_of_wal,
+									  InvalidBlockNumber, InvalidBlockNumber,
+									  false, false);
 
 	/*
 	 * Acquire AccessExclusiveLock on the table, its TOAST relation (if there
@@ -3392,10 +3632,12 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	end_of_wal = GetFlushRecPtr(NULL);
 
 	/*
-	 * Apply the concurrent changes again. Indicate that the decoding worker
-	 * won't be needed anymore.
+	 * Decode and apply the concurrent changes again. Indicate that the
+	 * decoding worker won't be needed anymore.
 	 */
-	process_concurrent_changes(end_of_wal, &chgcxt, true);
+	repack_process_concurrent_changes(chgcxt, end_of_wal,
+									  InvalidBlockNumber, InvalidBlockNumber,
+									  false, true);
 
 	/* Remember info about rel before closing OldHeap */
 	relpersistence = OldHeap->rd_rel->relpersistence;
@@ -3443,7 +3685,7 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	table_close(NewHeap, NoLock);
 
 	/* Cleanup what we don't need anymore. (And close the identity index.) */
-	release_change_context(&chgcxt);
+	release_change_context(chgcxt);
 
 	/*
 	 * Swap the relations and their TOAST relations and TOAST indexes. This
@@ -3462,6 +3704,103 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 					 relpersistence);
 }
 
+/*
+ * Copy the contents of the auxiliary table to the new table in the desired
+ * order, then drop the auxiliary table.
+ */
+static void
+process_auxiliary_table(ChangeContext *chgcxt, Relation OldHeap, Oid identIdx)
+{
+	RepackDest *dest = chgcxt->cc_dest_aux;
+	Oid			ident_idx_new;
+	Relation	clustering_index;
+	IndexScanDesc scan;
+	TupleTableSlot *slot;
+	Oid			aux_oid;
+	ObjectAddress object;
+	Relation	rel;
+
+	/*
+	 * First, make sure the clustering index exists.
+	 */
+	if (OidIsValid(chgcxt->cc_clustering_index))
+	{
+		Oid			cl_ind_oid;
+
+		/*
+		 * Create it according to the clustering index on the old relation.
+		 */
+		cl_ind_oid = build_new_index(dest->rel, OldHeap,
+									 chgcxt->cc_clustering_index);
+		clustering_index = index_open(cl_ind_oid, NoLock);
+	}
+	else
+	{
+		/* The identity index is also the clustering index. */
+		clustering_index = dest->ident_index;
+	}
+
+	/*
+	 * Now do the copying. Before starting, clear ->cc_dest_aux so that
+	 * insertions go to the final table, rather than the auxiliary one.
+	 */
+	chgcxt->cc_dest_aux = NULL;
+	slot = table_slot_create(dest->rel, NULL);
+
+	/*
+	 * Note: the current active snapshot blocks the progress of xmin
+	 * horizon(s). The next patches in the series should fix this by using a
+	 * new kind of snapshot (which we can use here because there are no
+	 * transaction aborts in the auxiliary table).
+	 */
+	scan = index_beginscan(dest->rel, clustering_index, GetActiveSnapshot(),
+						   NULL, 0, 0, SO_NONE);
+	index_rescan(scan, NULL, 0, NULL, 0);
+	for (;;)
+	{
+		CHECK_FOR_INTERRUPTS();
+
+		if (!index_getnext_slot(scan, ForwardScanDirection, slot))
+			break;
+
+		/*
+		 * Reforming should have been performed during insertions into the
+		 * auxiliary table.
+		 */
+		heap_insert_for_repack(chgcxt, slot, NULL);
+	}
+	index_endscan(scan);
+	ExecDropSingleTupleTableSlot(slot);
+
+	/*
+	 * Close the relation, its identity index and clustering index if we had
+	 * to open it above. Lock will be released on commit.
+	 */
+	aux_oid = RelationGetRelid(dest->rel);
+	table_close(dest->rel, NoLock);
+	if (OidIsValid(chgcxt->cc_clustering_index))
+		index_close(clustering_index, NoLock);
+	/* Here we close the other indexes. */
+	release_change_dest(dest);
+
+	/* Drop the auxiliary table. */
+	object.classId = RelationRelationId;
+	object.objectId = aux_oid;
+	object.objectSubId = 0;
+	performDeletion(&object, DROP_RESTRICT, PERFORM_DELETION_INTERNAL);
+
+	/* Build the identity index on the new relation. */
+	ident_idx_new = build_new_index(chgcxt->cc_dest.rel, OldHeap, identIdx);
+
+	/*
+	 * Make the new heap ready to use the index for future replaying of
+	 * concurrent changes.
+	 */
+	rel = chgcxt->cc_dest.rel;
+	release_change_dest(&chgcxt->cc_dest);
+	initialize_change_dest(&chgcxt->cc_dest, rel, ident_idx_new);
+}
+
 /*
  * Build indexes on NewHeap according to those on OldHeap.
  *
@@ -3477,34 +3816,48 @@ build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes)
 {
 	List	   *result = NIL;
 
-	pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
-								 PROGRESS_REPACK_PHASE_REBUILD_INDEX);
-
 	foreach_oid(oldindex, OldIndexes)
 	{
 		Oid			newindex;
-		char	   *newName;
-		Relation	ind;
-
-		ind = index_open(oldindex, ShareUpdateExclusiveLock);
-
-		newName = ChooseRelationName(get_rel_name(oldindex),
-									 NULL,
-									 "repacknew",
-									 get_rel_namespace(ind->rd_index->indrelid),
-									 false);
-		newindex = index_create_copy(NewHeap, INDEX_CREATE_SUPPRESS_PROGRESS,
-									 oldindex, ind->rd_rel->reltablespace,
-									 newName);
-		copy_index_constraints(ind, newindex, RelationGetRelid(NewHeap));
-		result = lappend_oid(result, newindex);
 
-		index_close(ind, NoLock);
+		newindex = build_new_index(NewHeap, OldHeap, oldindex);
+		result = lappend_oid(result, newindex);
 	}
 
 	return result;
 }
 
+/*
+ * Subroutine of build_new_indexes().
+ */
+static Oid
+build_new_index(Relation NewHeap, Relation OldHeap, Oid oldindex)
+{
+	Oid			newindex;
+	char	   *newName;
+	Relation	ind;
+
+	pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+								 PROGRESS_REPACK_PHASE_REBUILD_INDEX);
+
+	ind = index_open(oldindex, ShareUpdateExclusiveLock);
+
+	newName = ChooseRelationName(get_rel_name(oldindex),
+								 NULL,
+								 "repacknew",
+								 get_rel_namespace(ind->rd_index->indrelid),
+								 false);
+	/* Functions in indexes may want a snapshot set. */
+	Assert(ActiveSnapshotSet());
+	newindex = index_create_copy(NewHeap, INDEX_CREATE_SUPPRESS_PROGRESS,
+								 oldindex, ind->rd_rel->reltablespace,
+								 newName);
+	copy_index_constraints(ind, newindex, RelationGetRelid(NewHeap));
+	index_close(ind, NoLock);
+
+	return newindex;
+}
+
 /*
  * Create a transient copy of a constraint -- supported by a transient
  * copy of the index that supports the original constraint.
@@ -3693,10 +4046,13 @@ start_repack_decoding_worker(Oid relid)
 
 	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
 	shared->initialized = false;
+	/* Snapshot is the first thing we need from the worker. */
+	shared->snapshot_requested = true;
 	shared->lsn_upto = InvalidXLogRecPtr;
 	shared->done = false;
 	SharedFileSetInit(&shared->sfs, decoding_worker->seg);
-	shared->last_exported = -1;
+	shared->last_exported_snapshot = -1;
+	shared->last_exported_changes = -1;
 	SpinLockInit(&shared->mutex);
 	shared->dbid = MyDatabaseId;
 
@@ -3818,10 +4174,10 @@ stop_repack_decoding_worker_cb(int code, Datum arg)
 }
 
 /*
- * Get the initial snapshot from the decoding worker.
+ * Get snapshot from the decoding worker.
  */
-static Snapshot
-get_initial_snapshot(DecodingWorker *worker)
+Snapshot
+repack_get_snapshot(ChangeContext *chgcxt)
 {
 	DecodingWorkerShared *shared;
 	char		fname[MAXPGPATH];
@@ -3830,12 +4186,13 @@ get_initial_snapshot(DecodingWorker *worker)
 	char	   *snap_space;
 	Snapshot	snapshot;
 
-	shared = (DecodingWorkerShared *) dsm_segment_address(worker->seg);
+	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
 
 	/*
-	 * The worker needs to initialize the logical decoding, which usually
-	 * takes some time. Therefore it makes sense to prepare for the sleep
-	 * first.
+	 * For the first snapshot request, the worker needs to initialize the
+	 * logical decoding, which usually takes some time. Therefore it makes
+	 * sense to prepare for the sleep first. Does it make sense to skip the
+	 * preparation on the next requests?
 	 */
 	ConditionVariablePrepareToSleep(&shared->cv);
 	for (;;)
@@ -3843,13 +4200,13 @@ get_initial_snapshot(DecodingWorker *worker)
 		int			last_exported;
 
 		SpinLockAcquire(&shared->mutex);
-		last_exported = shared->last_exported;
+		last_exported = shared->last_exported_snapshot;
 		SpinLockRelease(&shared->mutex);
 
 		/*
 		 * Has the worker exported the file we are waiting for?
 		 */
-		if (last_exported == WORKER_FILE_SNAPSHOT)
+		if (last_exported == chgcxt->cc_file_seq_snapshot)
 			break;
 
 		ConditionVariableSleep(&shared->cv, WAIT_EVENT_REPACK_WORKER_EXPORT);
@@ -3857,20 +4214,98 @@ get_initial_snapshot(DecodingWorker *worker)
 	ConditionVariableCancelSleep();
 
 	/* Read the snapshot from a file. */
-	DecodingWorkerFileName(fname, shared->relid, WORKER_FILE_SNAPSHOT);
+	DecodingWorkerFileName(fname, shared->relid, chgcxt->cc_file_seq_snapshot,
+						   true);
 	file = BufFileOpenFileSet(&shared->sfs.fs, fname, O_RDONLY, false);
 	BufFileReadExact(file, &snap_size, sizeof(snap_size));
 	snap_space = (char *) palloc(snap_size);
 	BufFileReadExact(file, snap_space, snap_size);
 	BufFileClose(file);
 
+#ifdef USE_ASSERT_CHECKING
+	SpinLockAcquire(&shared->mutex);
+	Assert(!shared->snapshot_requested);
+	shared->snapshot_requested = false;
+	SpinLockRelease(&shared->mutex);
+#endif
+
 	/* Restore it. */
 	snapshot = RestoreSnapshot(snap_space);
 	pfree(snap_space);
 
+	/* Get ready for the next snapshot. */
+	chgcxt->cc_file_seq_snapshot++;
+
 	return snapshot;
 }
 
+/*
+ * Get concurrent changes, up to (and including) the record whose LSN is
+ * 'end_of_wal', from the decoding worker, and apply them to the new table. If
+ * block range is specified, only apply changes related to that range.
+ *
+ * If 'request_snapshot' is true, the snapshot built at LSN following the last
+ * data change needs to be exported too.
+ */
+void
+repack_process_concurrent_changes(ChangeContext *chgcxt,
+								  XLogRecPtr end_of_wal,
+								  BlockNumber range_start,
+								  BlockNumber range_end,
+								  bool request_snapshot, bool done)
+{
+	DecodingWorkerShared *shared;
+
+	pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+								 PROGRESS_REPACK_PHASE_CATCH_UP);
+
+	/* Ask the worker for the file. */
+	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
+	SpinLockAcquire(&shared->mutex);
+	shared->lsn_upto = end_of_wal;
+	Assert(!shared->snapshot_requested);
+	shared->snapshot_requested = request_snapshot;
+	shared->done = done;
+	SpinLockRelease(&shared->mutex);
+
+	/*
+	 * The worker needs to finish processing of the current WAL record. Even
+	 * if it's idle, it'll need to close the output file. Thus we're likely to
+	 * wait, so prepare for sleep.
+	 */
+	ConditionVariablePrepareToSleep(&shared->cv);
+	for (;;)
+	{
+		int			last_exported;
+
+		SpinLockAcquire(&shared->mutex);
+		last_exported = shared->last_exported_changes;
+		SpinLockRelease(&shared->mutex);
+
+		/*
+		 * Has the worker exported the file we are waiting for?
+		 */
+		if (last_exported == chgcxt->cc_file_seq_changes)
+			break;
+
+		ConditionVariableSleep(&shared->cv, WAIT_EVENT_REPACK_WORKER_EXPORT);
+	}
+	ConditionVariableCancelSleep();
+
+#ifdef USE_ASSERT_CHECKING
+	/* No file is exported until the worker exports the next one. */
+	SpinLockAcquire(&shared->mutex);
+	Assert(XLogRecPtrIsInvalid(shared->lsn_upto));
+	SpinLockRelease(&shared->mutex);
+#endif
+
+	/* Apply the changes to the new table. */
+	apply_concurrent_changes(chgcxt, range_start, range_end);
+
+	/* Get ready for the next set of changes. */
+	chgcxt->cc_file_seq_changes++;
+}
+
 /*
  * Generate worker's file name into 'fname', which must be of size MAXPGPATH.
  * If relations of the same 'relid' happen to be processed at the same time,
@@ -3878,10 +4313,13 @@ get_initial_snapshot(DecodingWorker *worker)
  * be involved.
  */
 void
-DecodingWorkerFileName(char *fname, Oid relid, uint32 seq)
+DecodingWorkerFileName(char *fname, Oid relid, uint32 seq, bool snapshot)
 {
 	/* The PID is already present in the fileset name, so we needn't add it */
-	snprintf(fname, MAXPGPATH, "%u-%u", relid, seq);
+	if (!snapshot)
+		snprintf(fname, MAXPGPATH, "%u-%u", relid, seq);
+	else
+		snprintf(fname, MAXPGPATH, "%u-%u-snapshot", relid, seq);
 }
 
 /*
diff --git a/src/backend/commands/repack_worker.c b/src/backend/commands/repack_worker.c
index db9ff057cc6..461d60ec0ca 100644
--- a/src/backend/commands/repack_worker.c
+++ b/src/backend/commands/repack_worker.c
@@ -33,8 +33,7 @@
 static void RepackWorkerShutdown(int code, Datum arg);
 static LogicalDecodingContext *repack_setup_logical_decoding(Oid relid);
 static void repack_cleanup_logical_decoding(LogicalDecodingContext *ctx);
-static void export_initial_snapshot(Snapshot snapshot,
-									DecodingWorkerShared *shared);
+static void export_snapshot(Snapshot snapshot, DecodingWorkerShared *shared);
 static bool decode_concurrent_changes(LogicalDecodingContext *ctx,
 									  DecodingWorkerShared *shared);
 
@@ -65,6 +64,8 @@ RepackWorkerMain(Datum main_arg)
 	shm_mq_handle *mqh;
 	LogicalDecodingContext *decoding_ctx;
 	SharedFileSet *sfs;
+	RepackDecodingState *dstate;
+	MemoryContext oldcxt;
 	Snapshot	snapshot;
 
 	am_repack_worker = true;
@@ -118,7 +119,9 @@ RepackWorkerMain(Datum main_arg)
 	 * anything in the shared memory until we have serialized the snapshot.
 	 */
 	SpinLockAcquire(&shared->mutex);
-	Assert(!XLogRecPtrIsValid(shared->lsn_upto));
+	/* Initially we're expected to provide a snapshot and only that. */
+	Assert(shared->snapshot_requested &&
+		   XLogRecPtrIsInvalid(shared->lsn_upto));
 	sfs = &shared->sfs;
 	SpinLockRelease(&shared->mutex);
 
@@ -139,9 +142,25 @@ RepackWorkerMain(Datum main_arg)
 	XactIsoLevel = XACT_REPEATABLE_READ;
 	XactReadOnly = true;
 
-	/* Build the initial snapshot and export it. */
+	/*
+	 * Build the initial snapshot and export it.
+	 *
+	 * Since there is no API to free the "external snapshot", and since such
+	 * snapshot is not guaranteed to be flat (i.e. pfree() is not appropriate)
+	 * the easiest way to clean it up is to use a separate memory context for
+	 * it.
+	 */
+	dstate = (RepackDecodingState *) decoding_ctx->output_writer_private;
+	MemoryContextReset(dstate->snapshot_cxt);
+	oldcxt = MemoryContextSwitchTo(dstate->snapshot_cxt);
 	snapshot = SnapBuildInitialSnapshot(decoding_ctx->snapshot_builder);
-	export_initial_snapshot(snapshot, shared);
+	MemoryContextSwitchTo(oldcxt);
+	export_snapshot(snapshot, shared);
+
+	/*
+	 * Adjust the replication slot's xmin so that VACUUM can do more work.
+	 */
+	LogicalIncreaseXminForSlot(InvalidXLogRecPtr, snapshot->xmin, false);
 
 	/*
 	 * Only historic snapshots should be used now. Do not let us restrict the
@@ -307,7 +326,7 @@ repack_cleanup_logical_decoding(LogicalDecodingContext *ctx)
  * Make snapshot available to the backend that launched the decoding worker.
  */
 static void
-export_initial_snapshot(Snapshot snapshot, DecodingWorkerShared *shared)
+export_snapshot(Snapshot snapshot, DecodingWorkerShared *shared)
 {
 	char		fname[MAXPGPATH];
 	BufFile    *file;
@@ -318,7 +337,9 @@ export_initial_snapshot(Snapshot snapshot, DecodingWorkerShared *shared)
 	snap_space = (char *) palloc(snap_size);
 	SerializeSnapshot(snapshot, snap_space);
 
-	DecodingWorkerFileName(fname, shared->relid, shared->last_exported + 1);
+	DecodingWorkerFileName(fname, shared->relid,
+						   shared->last_exported_snapshot + 1,
+						   true);
 	file = BufFileCreateFileSet(&shared->sfs.fs, fname);
 	/* To make restoration easier, write the snapshot size first. */
 	BufFileWrite(file, &snap_size, sizeof(snap_size));
@@ -328,7 +349,8 @@ export_initial_snapshot(Snapshot snapshot, DecodingWorkerShared *shared)
 
 	/* Increase the counter to tell the backend that the file is available. */
 	SpinLockAcquire(&shared->mutex);
-	shared->last_exported++;
+	shared->last_exported_snapshot++;
+	shared->snapshot_requested = false;
 	SpinLockRelease(&shared->mutex);
 	ConditionVariableSignal(&shared->cv);
 }
@@ -343,6 +365,7 @@ decode_concurrent_changes(LogicalDecodingContext *ctx,
 						  DecodingWorkerShared *shared)
 {
 	RepackDecodingState *dstate;
+	bool		snapshot_requested;
 	XLogRecPtr	lsn_upto;
 	bool		done;
 	char		fname[MAXPGPATH];
@@ -350,11 +373,14 @@ decode_concurrent_changes(LogicalDecodingContext *ctx,
 	dstate = (RepackDecodingState *) ctx->output_writer_private;
 
 	/* Open the output file. */
-	DecodingWorkerFileName(fname, shared->relid, shared->last_exported + 1);
+	DecodingWorkerFileName(fname, shared->relid,
+						   shared->last_exported_changes + 1,
+						   false);
 	dstate->file = BufFileCreateFileSet(&shared->sfs.fs, fname);
 
 	SpinLockAcquire(&shared->mutex);
 	lsn_upto = shared->lsn_upto;
+	snapshot_requested = shared->snapshot_requested;
 	done = shared->done;
 	SpinLockRelease(&shared->mutex);
 
@@ -437,6 +463,7 @@ decode_concurrent_changes(LogicalDecodingContext *ctx,
 		{
 			SpinLockAcquire(&shared->mutex);
 			lsn_upto = shared->lsn_upto;
+			snapshot_requested = shared->snapshot_requested;
 			/* 'done' should be set at the same time as 'lsn_upto' */
 			done = shared->done;
 			SpinLockRelease(&shared->mutex);
@@ -483,9 +510,59 @@ decode_concurrent_changes(LogicalDecodingContext *ctx,
 	 */
 	BufFileClose(dstate->file);
 	dstate->file = NULL;
+
+	/*
+	 * Before publishing the data changes, export the snapshot too if
+	 * requested. Publishing both at once makes sense because both are needed
+	 * at the same time, and it's simpler.
+	 */
+	if (snapshot_requested)
+	{
+		Snapshot	snapshot;
+		MemoryContext oldcxt;
+
+		/* See comments about memory context in RepackWorkerMain(). */
+		MemoryContextReset(dstate->snapshot_cxt);
+		oldcxt = MemoryContextSwitchTo(dstate->snapshot_cxt);
+
+		/*
+		 * SnapBuildInitialSnapshot() assumes invalid XID, so set it. We do
+		 * not use the snapshot, so it's ok.
+		 */
+		MyProc->xmin = InvalidTransactionId;
+		snapshot = SnapBuildInitialSnapshot(ctx->snapshot_builder);
+		MemoryContextSwitchTo(oldcxt);
+		export_snapshot(snapshot, shared);
+
+		/*
+		 * Adjust the replication slot's xmin so that VACUUM can do more work.
+		 */
+		LogicalIncreaseXminForSlot(InvalidXLogRecPtr, snapshot->xmin, false);
+	}
+	else
+	{
+		/*
+		 * If data changes were requested but no following snapshot, we don't
+		 * care about xmin horizon because the heap copying should be done by
+		 * now.
+		 */
+		LogicalIncreaseXminForSlot(InvalidXLogRecPtr, InvalidTransactionId,
+								   false);
+
+	}
+
+	/*
+	 * Make sure the xmin of our slot is taken into account when computing new
+	 * VACUUM horizons.
+	 */
+	ReplicationSlotsComputeRequiredXmin(false);
+
+	/*
+	 * Now increase the counter(s) to announce that the output is available.
+	 */
 	SpinLockAcquire(&shared->mutex);
+	shared->last_exported_changes++;
 	shared->lsn_upto = InvalidXLogRecPtr;
-	shared->last_exported++;
 	SpinLockRelease(&shared->mutex);
 	ConditionVariableSignal(&shared->cv);
 
diff --git a/src/backend/replication/logical/decode.c b/src/backend/replication/logical/decode.c
index c944be4ac83..c3722b5c623 100644
--- a/src/backend/replication/logical/decode.c
+++ b/src/backend/replication/logical/decode.c
@@ -921,6 +921,7 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	xl_heap_insert *xlrec;
 	ReorderBufferChange *change;
 	RelFileLocator target_locator;
+	BlockNumber blknum;
 
 	xlrec = (xl_heap_insert *) XLogRecGetData(r);
 
@@ -932,7 +933,7 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		return;
 
 	/* only interested in our database */
-	XLogRecGetBlockTag(r, 0, &target_locator, NULL, NULL);
+	XLogRecGetBlockTag(r, 0, &target_locator, NULL, &blknum);
 	if (target_locator.dbOid != ctx->slot->data.database)
 		return;
 
@@ -947,7 +948,8 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		change->action = REORDER_BUFFER_CHANGE_INTERNAL_SPEC_INSERT;
 	change->origin_id = XLogRecGetOrigin(r);
 
-	memcpy(&change->data.tp.rlocator, &target_locator, sizeof(RelFileLocator));
+	memcpy(&change->data.tp.rlocator, &target_locator,
+		   sizeof(RelFileLocator));
 
 	tupledata = XLogRecGetBlockData(r, 0, &datalen);
 	tuplelen = datalen - SizeOfHeapHeader;
@@ -957,6 +959,20 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 
 	DecodeXLogTuple(tupledata, datalen, change->data.tp.newtuple);
 
+	/*
+	 * REPACK (CONCURRENTLY) needs block number to check if the corresponding
+	 * part of the table was already copied.  XXX Should we only do this if
+	 * AmRepackWorker()? It might save a few cycles, but not sure it's good to
+	 * leave the fields unset in other cases.
+	 */
+	{
+		HeapTupleHeader header;
+
+		header = change->data.tp.newtuple->t_data;
+		/* offnum is not really needed, but let's set valid pointer. */
+		ItemPointerSet(&header->t_ctid, blknum, xlrec->offnum);
+	}
+
 	change->data.tp.clear_toast_afterwards = true;
 
 	ReorderBufferQueueChange(ctx->reorder, XLogRecGetXid(r), buf->origptr,
@@ -978,11 +994,13 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	ReorderBufferChange *change;
 	char	   *data;
 	RelFileLocator target_locator;
+	BlockNumber new_blknum,
+				old_blknum;
 
 	xlrec = (xl_heap_update *) XLogRecGetData(r);
 
 	/* only interested in our database */
-	XLogRecGetBlockTag(r, 0, &target_locator, NULL, NULL);
+	XLogRecGetBlockTag(r, 0, &target_locator, NULL, &new_blknum);
 	if (target_locator.dbOid != ctx->slot->data.database)
 		return;
 
@@ -990,6 +1008,11 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	if (FilterByOrigin(ctx, XLogRecGetOrigin(r)))
 		return;
 
+	if (XLogRecHasBlockRef(r, 1))
+		XLogRecGetBlockTag(r, 1, NULL, NULL, &old_blknum);
+	else
+		old_blknum = new_blknum;
+
 	change = ReorderBufferAllocChange(ctx->reorder);
 	change->action = REORDER_BUFFER_CHANGE_UPDATE;
 	change->origin_id = XLogRecGetOrigin(r);
@@ -1008,6 +1031,20 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 			ReorderBufferAllocTupleBuf(ctx->reorder, tuplelen);
 
 		DecodeXLogTuple(data, datalen, change->data.tp.newtuple);
+
+		/*
+		 * REPACK (CONCURRENTLY) needs block numbers to check if the
+		 * corresponding part of the table was already copied. XXX Do this
+		 * only if AmRepackWorker()?
+		 */
+		{
+			HeapTupleHeader header;
+
+			header = change->data.tp.newtuple->t_data;
+			/* offnum is not really needed, but let's set valid pointer. */
+			ItemPointerSet(&header->t_ctid, new_blknum, xlrec->new_offnum);
+			change->data.tp.old_blknum = old_blknum;
+		}
 	}
 
 	if (xlrec->flags & XLH_UPDATE_CONTAINS_OLD)
@@ -1044,6 +1081,7 @@ DecodeDelete(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	xl_heap_delete *xlrec;
 	ReorderBufferChange *change;
 	RelFileLocator target_locator;
+	BlockNumber blknum;
 
 	xlrec = (xl_heap_delete *) XLogRecGetData(r);
 
@@ -1057,7 +1095,7 @@ DecodeDelete(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		return;
 
 	/* only interested in our database */
-	XLogRecGetBlockTag(r, 0, &target_locator, NULL, NULL);
+	XLogRecGetBlockTag(r, 0, &target_locator, NULL, &blknum);
 	if (target_locator.dbOid != ctx->slot->data.database)
 		return;
 
@@ -1089,6 +1127,19 @@ DecodeDelete(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 
 		DecodeXLogTuple((char *) xlrec + SizeOfHeapDelete,
 						datalen, change->data.tp.oldtuple);
+
+		/*
+		 * REPACK (CONCURRENTLY) needs block number to check if the
+		 * corresponding part of the table was already copied. XXX Do this
+		 * only if AmRepackWorker()?
+		 */
+		{
+			HeapTupleHeader header;
+
+			header = change->data.tp.oldtuple->t_data;
+			/* offnum is not really needed, but let's set valid pointer. */
+			ItemPointerSet(&header->t_ctid, blknum, xlrec->offnum);
+		}
 	}
 
 	change->data.tp.clear_toast_afterwards = true;
@@ -1148,8 +1199,11 @@ DecodeMultiInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	char	   *tupledata;
 	Size		tuplelen;
 	RelFileLocator rlocator;
+	BlockNumber blknum;
+	bool		isinit;
 
 	xlrec = (xl_heap_multi_insert *) XLogRecGetData(r);
+	isinit = (XLogRecGetInfo(r) & XLOG_HEAP_INIT_PAGE) != 0;
 
 	/*
 	 * Ignore insert records without new tuples.  This happens when a
@@ -1159,7 +1213,7 @@ DecodeMultiInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		return;
 
 	/* only interested in our database */
-	XLogRecGetBlockTag(r, 0, &rlocator, NULL, NULL);
+	XLogRecGetBlockTag(r, 0, &rlocator, NULL, &blknum);
 	if (rlocator.dbOid != ctx->slot->data.database)
 		return;
 
@@ -1227,6 +1281,25 @@ DecodeMultiInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		else
 			change->data.tp.clear_toast_afterwards = false;
 
+		/*
+		 * REPACK (CONCURRENTLY) needs block number to check if the
+		 * corresponding part of the table was already copied.
+		 */
+		if (AmRepackWorker())
+		{
+			OffsetNumber offnum;
+
+			/*
+			 * offnum is not really needed, but let's set valid pointer. (It
+			 * will be invalid anyway if the page was initially empty.)
+			 */
+			if (isinit)
+				offnum = FirstOffsetNumber + i;
+			else
+				offnum = xlrec->offsets[i];
+			ItemPointerSet(&header->t_ctid, blknum, offnum);
+		}
+
 		ReorderBufferQueueChange(ctx->reorder, XLogRecGetXid(r),
 								 buf->origptr, change, false);
 
diff --git a/src/backend/replication/logical/logical.c b/src/backend/replication/logical/logical.c
index 3541fc793e4..fdcaa5036b4 100644
--- a/src/backend/replication/logical/logical.c
+++ b/src/backend/replication/logical/logical.c
@@ -1659,14 +1659,17 @@ update_progress_txn_cb_wrapper(ReorderBuffer *cache, ReorderBufferTXN *txn,
 
 /*
  * Set the required catalog xmin horizon for historic snapshots in the current
- * replication slot.
+ * replication slot if catalog is TRUE, or xmin if catalog is FALSE.
  *
  * Note that in the most cases, we won't be able to immediately use the xmin
  * to increase the xmin horizon: we need to wait till the client has confirmed
- * receiving current_lsn with LogicalConfirmReceivedLocation().
+ * receiving current_lsn with LogicalConfirmReceivedLocation(). However,
+ * catalog=FALSE is only allowed for temporary replication slots, so the
+ * horizon is applied immediately.
  */
 void
-LogicalIncreaseXminForSlot(XLogRecPtr current_lsn, TransactionId xmin)
+LogicalIncreaseXminForSlot(XLogRecPtr current_lsn, TransactionId xmin,
+						   bool catalog)
 {
 	bool		updated_xmin = false;
 	ReplicationSlot *slot;
@@ -1677,6 +1680,27 @@ LogicalIncreaseXminForSlot(XLogRecPtr current_lsn, TransactionId xmin)
 	Assert(slot != NULL);
 
 	SpinLockAcquire(&slot->mutex);
+	if (!catalog)
+	{
+		/*
+		 * The non-catalog horizon can only advance in temporary slots, so
+		 * update it in the shared memory immediately (w/o requiring prior
+		 * saving to disk).
+		 */
+		Assert(slot->data.persistency == RS_TEMPORARY);
+
+		/*
+		 * The horizon must not go backwards, however it's o.k. to become
+		 * invalid.
+		 */
+		Assert(!TransactionIdIsValid(slot->effective_xmin) ||
+			   !TransactionIdIsValid(xmin) ||
+			   TransactionIdFollowsOrEquals(xmin, slot->effective_xmin));
+
+		slot->effective_xmin = xmin;
+		SpinLockRelease(&slot->mutex);
+		return;
+	}
 
 	/*
 	 * don't overwrite if we already have a newer xmin. This can happen if we
diff --git a/src/backend/replication/logical/reorderbuffer.c b/src/backend/replication/logical/reorderbuffer.c
index d06d0d8c9be..cae2b099e69 100644
--- a/src/backend/replication/logical/reorderbuffer.c
+++ b/src/backend/replication/logical/reorderbuffer.c
@@ -3727,6 +3727,40 @@ ReorderBufferXidHasCatalogChanges(ReorderBuffer *rb, TransactionId xid)
 	return rbtxn_has_catalog_changes(txn);
 }
 
+/*
+ * Check if a transaction (or its subtransaction) contains a heap change.
+ */
+bool
+ReorderBufferXidHasHeapChanges(ReorderBuffer *rb, TransactionId xid)
+{
+	ReorderBufferTXN *txn;
+	dlist_iter	iter;
+
+	txn = ReorderBufferTXNByXid(rb, xid, false, NULL, InvalidXLogRecPtr,
+								false);
+	if (txn == NULL)
+		return false;
+
+	dlist_foreach(iter, &txn->changes)
+	{
+		ReorderBufferChange *change;
+
+		change = dlist_container(ReorderBufferChange, node, iter.cur);
+
+		switch (change->action)
+		{
+			case REORDER_BUFFER_CHANGE_INSERT:
+			case REORDER_BUFFER_CHANGE_UPDATE:
+			case REORDER_BUFFER_CHANGE_DELETE:
+				return true;
+			default:
+				break;
+		}
+	}
+
+	return false;
+}
+
 /*
  * ReorderBufferXidHasBaseSnapshot
  *		Have we already set the base snapshot for the given txn/subtxn?
@@ -5224,6 +5258,12 @@ ReorderBufferToastReplace(ReorderBuffer *rb, ReorderBufferTXN *txn,
 	Assert(newtup->t_len <= MaxHeapTupleSize);
 	Assert(newtup->t_data == (HeapTupleHeader) ((char *) newtup + HEAPTUPLESIZE));
 
+	/*
+	 * Preserve TID - REPACK relies on it when dealing with block ranges. XXX
+	 * Shouldn't we add a new field to ReorderBufferChange instead?
+	 */
+	tmphtup->t_data->t_ctid = newtup->t_data->t_ctid;
+
 	memcpy(newtup->t_data, tmphtup->t_data, tmphtup->t_len);
 	newtup->t_len = tmphtup->t_len;
 
diff --git a/src/backend/replication/logical/snapbuild.c b/src/backend/replication/logical/snapbuild.c
index b1e37ef6792..dd7d584e15b 100644
--- a/src/backend/replication/logical/snapbuild.c
+++ b/src/backend/replication/logical/snapbuild.c
@@ -128,6 +128,7 @@
 #include "access/heapam_xlog.h"
 #include "access/transam.h"
 #include "access/xact.h"
+#include "commands/repack.h"
 #include "common/file_utils.h"
 #include "miscadmin.h"
 #include "pgstat.h"
@@ -983,6 +984,13 @@ SnapBuildCommitTxn(SnapBuild *builder, XLogRecPtr lsn, TransactionId xid,
 		}
 	}
 
+	/*
+	 * REPACK decoding worker may need timetravel anytime. It takes
+	 * responsibility for tracking transaction commits, see below.
+	 */
+	else if (AmRepackWorker())
+		needs_timetravel = true;
+
 	for (nxact = 0; nxact < nsubxacts; nxact++)
 	{
 		TransactionId subxid = subxacts[nxact];
@@ -990,8 +998,12 @@ SnapBuildCommitTxn(SnapBuild *builder, XLogRecPtr lsn, TransactionId xid,
 		/*
 		 * Add subtransaction to base snapshot if catalog modifying, we don't
 		 * distinguish to toplevel transactions there.
+		 *
+		 * See comments on REPACK worker below.
 		 */
-		if (SnapBuildXidHasCatalogChanges(builder, subxid, xinfo))
+		if (SnapBuildXidHasCatalogChanges(builder, subxid, xinfo) ||
+			(AmRepackWorker() &&
+			 ReorderBufferXidHasHeapChanges(builder->reorder, xid)))
 		{
 			sub_needs_timetravel = true;
 			needs_snapshot = true;
@@ -1019,8 +1031,18 @@ SnapBuildCommitTxn(SnapBuild *builder, XLogRecPtr lsn, TransactionId xid,
 		}
 	}
 
-	/* if top-level modified catalog, it'll need a snapshot */
-	if (SnapBuildXidHasCatalogChanges(builder, xid, xinfo))
+	/*
+	 * If top-level modified catalog, it'll need a snapshot.
+	 *
+	 * If we're decoding changes on behalf of REPACK (CONCURRENTLY), only
+	 * changes of the relation being processed are decoded - see
+	 * heap_decode(). Thus any heap change we find here must belong to that
+	 * relation. Add the transaction so that we can keep building snapshots to
+	 * scan that relation.
+	 */
+	if (SnapBuildXidHasCatalogChanges(builder, xid, xinfo) ||
+		(AmRepackWorker() &&
+		 ReorderBufferXidHasHeapChanges(builder->reorder, xid)))
 	{
 		elog(DEBUG2, "found top level transaction %u, with catalog changes",
 			 xid);
@@ -1188,7 +1210,7 @@ SnapBuildProcessRunningXacts(SnapBuild *builder, XLogRecPtr lsn, xl_running_xact
 		xmin = running->oldestRunningXid;
 	elog(DEBUG3, "xmin: %u, xmax: %u, oldest running: %u, oldest xmin: %u",
 		 builder->xmin, builder->xmax, running->oldestRunningXid, xmin);
-	LogicalIncreaseXminForSlot(lsn, xmin);
+	LogicalIncreaseXminForSlot(lsn, xmin, true);
 
 	/*
 	 * Also tell the slot where we can restart decoding from. We don't want to
diff --git a/src/backend/replication/pgrepack/pgrepack.c b/src/backend/replication/pgrepack/pgrepack.c
index 5c5095bde4e..239796b3335 100644
--- a/src/backend/replication/pgrepack/pgrepack.c
+++ b/src/backend/replication/pgrepack/pgrepack.c
@@ -33,7 +33,8 @@ static void repack_commit_txn(LogicalDecodingContext *ctx,
 static void repack_process_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
 								  Relation relation, ReorderBufferChange *change);
 static void repack_store_change(LogicalDecodingContext *ctx, Relation relation,
-								ConcurrentChangeKind kind, HeapTuple tuple);
+								ConcurrentChangeKind kind, HeapTuple tuple,
+								BlockNumber old_blknum);
 
 void
 _PG_output_plugin_init(OutputPluginCallbacks *cb)
@@ -67,6 +68,9 @@ repack_startup(LogicalDecodingContext *ctx, OutputPluginOptions *opt,
 	dstate->change_cxt = AllocSetContextCreate(ctx->context,
 											   "REPACK - change",
 											   ALLOCSET_DEFAULT_SIZES);
+	dstate->snapshot_cxt = AllocSetContextCreate(ctx->context,
+												 "REPACK - snapshot",
+												 ALLOCSET_DEFAULT_SIZES);
 	/* repack_setup_logical_decoding fills in the rest */
 	ctx->output_writer_private = dstate;
 
@@ -136,7 +140,8 @@ repack_process_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
 				if (newtuple == NULL)
 					elog(ERROR, "incomplete insert info");
 
-				repack_store_change(ctx, relation, CHANGE_INSERT, newtuple);
+				repack_store_change(ctx, relation, CHANGE_INSERT, newtuple,
+									InvalidBlockNumber);
 			}
 			break;
 		case REORDER_BUFFER_CHANGE_UPDATE:
@@ -151,9 +156,11 @@ repack_process_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
 					elog(ERROR, "incomplete update info");
 
 				if (oldtuple != NULL)
-					repack_store_change(ctx, relation, CHANGE_UPDATE_OLD, oldtuple);
+					repack_store_change(ctx, relation, CHANGE_UPDATE_OLD, oldtuple,
+										InvalidBlockNumber);
 
-				repack_store_change(ctx, relation, CHANGE_UPDATE_NEW, newtuple);
+				repack_store_change(ctx, relation, CHANGE_UPDATE_NEW, newtuple,
+									change->data.tp.old_blknum);
 			}
 			break;
 		case REORDER_BUFFER_CHANGE_DELETE:
@@ -165,7 +172,8 @@ repack_process_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
 				if (oldtuple == NULL)
 					elog(ERROR, "incomplete delete info");
 
-				repack_store_change(ctx, relation, CHANGE_DELETE, oldtuple);
+				repack_store_change(ctx, relation, CHANGE_DELETE, oldtuple,
+									InvalidBlockNumber);
 			}
 			break;
 		default:
@@ -192,7 +200,8 @@ repack_process_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
  */
 static void
 repack_store_change(LogicalDecodingContext *ctx, Relation relation,
-					ConcurrentChangeKind kind, HeapTuple tuple)
+					ConcurrentChangeKind kind, HeapTuple tuple,
+					BlockNumber old_blknum)
 {
 	RepackDecodingState *dstate;
 	MemoryContext oldcxt;
@@ -288,6 +297,9 @@ repack_store_change(LogicalDecodingContext *ctx, Relation relation,
 	 */
 	BufFileWrite(file, &tuple->t_len, sizeof(tuple->t_len));
 	BufFileWrite(file, tuple->t_data, tuple->t_len);
+	/* If old_blknum is specified, write it too. */
+	if (old_blknum != InvalidBlockNumber)
+		BufFileWrite(file, &old_blknum, sizeof(old_blknum));
 
 	/* Then, write the number of external attributes we found. */
 	natt_ext = list_length(attrs_ext);
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index d421cdbde76..ca2783107b1 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2544,6 +2544,16 @@
   boot_val => 'true',
 },
 
+# TODO Tune boot_val, 1024 is probably too low.
+{ name => 'repack_snapshot_after', type => 'int', context => 'PGC_USERSET', group => 'DEVELOPER_OPTIONS',
+  short_desc => 'Number of pages REPACK (CONCURRENTLY) can read using a single snapshot.',
+  flags => 'GUC_UNIT_BLOCKS | GUC_NOT_IN_SAMPLE',
+  variable => 'repack_pages_per_snapshot',
+  boot_val => '1024',
+  min => '1',
+  max => 'INT_MAX',
+}
+
 { name => 'reserved_connections', type => 'int', context => 'PGC_POSTMASTER', group => 'CONN_AUTH_SETTINGS',
   short_desc => 'Sets the number of connection slots reserved for roles with privileges of pg_use_reserved_connections.',
   variable => 'ReservedConnections',
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 90aa374b3ec..6d8106d9445 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -44,6 +44,7 @@
 #include "commands/async.h"
 #include "commands/extension.h"
 #include "commands/event_trigger.h"
+#include "commands/repack.h"
 #include "commands/tablespace.h"
 #include "commands/trigger.h"
 #include "commands/user.h"
diff --git a/src/include/access/tableam.h b/src/include/access/tableam.h
index f2c36696bca..132248c5d43 100644
--- a/src/include/access/tableam.h
+++ b/src/include/access/tableam.h
@@ -666,12 +666,12 @@ typedef struct TableAmRoutine
 											  Relation OldIndex,
 											  bool use_sort,
 											  TransactionId OldestXmin,
-											  Snapshot snapshot,
 											  TransactionId *xid_cutoff,
 											  MultiXactId *multi_cutoff,
 											  double *num_tuples,
 											  double *tups_vacuumed,
-											  double *tups_recently_dead);
+											  double *tups_recently_dead,
+											  void *tableam_data);
 
 	/*
 	 * React to VACUUM command on the relation. The VACUUM can be triggered by
@@ -1733,8 +1733,6 @@ table_relation_copy_data(Relation rel, const RelFileLocator *newrlocator)
  *   not needed for the relation's AM
  * - *xid_cutoff - ditto
  * - *multi_cutoff - ditto
- * - snapshot - if != NULL, ignore data changes done by transactions that this
- *	 (MVCC) snapshot considers still in-progress or in the future.
  *
  * Output parameters:
  * - *xid_cutoff - rel's new relfrozenxid value, may be invalid
@@ -1747,19 +1745,19 @@ table_relation_copy_for_cluster(Relation OldTable, Relation NewTable,
 								Relation OldIndex,
 								bool use_sort,
 								TransactionId OldestXmin,
-								Snapshot snapshot,
 								TransactionId *xid_cutoff,
 								MultiXactId *multi_cutoff,
 								double *num_tuples,
 								double *tups_vacuumed,
-								double *tups_recently_dead)
+								double *tups_recently_dead,
+								void *tableam_data)
 {
 	OldTable->rd_tableam->relation_copy_for_cluster(OldTable, NewTable, OldIndex,
 													use_sort, OldestXmin,
-													snapshot,
 													xid_cutoff, multi_cutoff,
 													num_tuples, tups_vacuumed,
-													tups_recently_dead);
+													tups_recently_dead,
+													tableam_data);
 }
 
 /*
diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h
index 07f887e99f6..8af73f8c81f 100644
--- a/src/include/commands/repack.h
+++ b/src/include/commands/repack.h
@@ -17,11 +17,14 @@
 
 #include "access/hio.h"
 #include "access/skey.h"
+#include "access/xlogdefs.h"
 #include "nodes/execnodes.h"
 #include "nodes/parsenodes.h"
 #include "parser/parse_node.h"
+#include "storage/block.h"
 #include "storage/lockdefs.h"
 #include "utils/relcache.h"
+#include "utils/snapshot.h"
 
 
 /* flag bits for ClusterParams->options */
@@ -66,13 +69,14 @@ typedef struct RepackDest
 
 	/* The latest column we need to deform to have the tuple identity */
 	AttrNumber	last_key_attno;
-} RepackDest;
 
-/*
- * The first file exported by the decoding worker must contain a snapshot, the
- * following ones contain the data changes.
- */
-#define WORKER_FILE_SNAPSHOT	0
+	/*
+	 * Range of blocks in the old table the contents of this table comes from.
+	 * Note that range_end is the first block of the next range.
+	 */
+	BlockNumber range_start;
+	BlockNumber range_end;
+} RepackDest;
 
 /*
  * Information needed to apply concurrent data changes.
@@ -84,10 +88,42 @@ typedef struct ChangeContext
 	/* The destination table. */
 	RepackDest	cc_dest;
 
-	/* Sequential number of the file containing the changes. */
-	int			cc_file_seq;
+	/* Sequential number of the file containing snapshot. */
+	int			cc_file_seq_snapshot;
+	/* Sequential number of the file containing data changes. */
+	int			cc_file_seq_changes;
+
+	/*
+	 * Auxiliary table to store ordered tuples temporarily.
+	 *
+	 * When the new relation needs to be clustered, we use this table instead
+	 * of tuplesort. The problem with a tuplesort is that data changes need to
+	 * be applied at range boundary (see heapam_relation_copy_for_cluster()
+	 * for more information), however it's not possible to look-up and change
+	 * tuples in tuplestore.
+	 *
+	 * Once the contents of the REPACKed table has been copied into the
+	 * auxiliary table, we build the clustering index (unless it's the same as
+	 * the identity index) and scan it to get the tuple in the desired order.
+	 * XXX Is it worth putting the contents into a tuplestore and sorting it?
+	 * Not sure, it'd require disk space for one more copy and the copying
+	 * itself is not free.
+	 *
+	 * TODO 1) make the tables unlogged, 2) if REPACK locks the TOAST relation
+	 * too (not sure it does) try to preserve TOAST pointers, instead of
+	 * storing them to TOAST relations of these tables, 3) Check that the
+	 * tables are dropped on transaction abort.
+	 */
+	RepackDest *cc_dest_aux;
+
+	/*
+	 * The index that defines ordering of the old table.
+	 */
+	Oid			cc_clustering_index;
 } ChangeContext;
 
+extern PGDLLIMPORT int repack_pages_per_snapshot;
+
 extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel);
 
 extern void cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid,
@@ -98,9 +134,8 @@ extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal);
 
 extern Oid	make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
 						  char relpersistence, LOCKMODE lockmode);
-extern void heap_insert_for_repack(Relation rel, TupleTableSlot *src,
-								   TupleTableSlot *reform,
-								   BulkInsertStateData *bistate);
+extern void heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
+								   TupleTableSlot *reform);
 extern bool tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc);
 extern void clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform);
 extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
@@ -112,7 +147,12 @@ extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 							 TransactionId frozenXid,
 							 MultiXactId cutoffMulti,
 							 char newrelpersistence);
-
+extern Snapshot repack_get_snapshot(ChangeContext *chgcxt);
+extern void repack_process_concurrent_changes(ChangeContext *chgcxt,
+											  XLogRecPtr end_of_wal,
+											  BlockNumber range_start,
+											  BlockNumber range_end,
+											  bool request_snapshot, bool done);
 extern void HandleRepackMessageInterrupt(void);
 extern void ProcessRepackMessages(void);
 
diff --git a/src/include/commands/repack_internal.h b/src/include/commands/repack_internal.h
index 42111aa4ae3..8003e999864 100644
--- a/src/include/commands/repack_internal.h
+++ b/src/include/commands/repack_internal.h
@@ -44,6 +44,8 @@ typedef struct RepackDecodingState
 
 	/* Per-change memory context. */
 	MemoryContext change_cxt;
+	/* Per-snapshot memory context. */
+	MemoryContext snapshot_cxt;
 
 	/* A tuple slot used to pass tuples back and forth */
 	TupleTableSlot *slot;
@@ -67,6 +69,9 @@ typedef struct DecodingWorkerShared
 	/* Is the decoding initialized? */
 	bool		initialized;
 
+	/* Set to request a snapshot. */
+	bool		snapshot_requested;
+
 	/*
 	 * Once the worker has reached this LSN, it should close the current
 	 * output file and either create a new one or exit, according to the field
@@ -74,6 +79,8 @@ typedef struct DecodingWorkerShared
 	 * the WAL available and keep checking this field. It is ok if the worker
 	 * had already decoded records whose LSN is >= lsn_upto before this field
 	 * has been set.
+	 *
+	 * Set a valid LSN to request data changes.
 	 */
 	XLogRecPtr	lsn_upto;
 
@@ -84,7 +91,8 @@ typedef struct DecodingWorkerShared
 	SharedFileSet sfs;
 
 	/* Number of the last file exported by the worker. */
-	int			last_exported;
+	int			last_exported_snapshot;
+	int			last_exported_changes;
 
 	/* Synchronize access to the fields above. */
 	slock_t		mutex;
@@ -116,7 +124,8 @@ typedef struct DecodingWorkerShared
 	char		error_queue[FLEXIBLE_ARRAY_MEMBER];
 } DecodingWorkerShared;
 
-extern void DecodingWorkerFileName(char *fname, Oid relid, uint32 seq);
+extern void DecodingWorkerFileName(char *fname, Oid relid, uint32 seq,
+								   bool snapshot);
 
 
 #endif							/* REPACK_INTERNAL_H */
diff --git a/src/include/replication/logical.h b/src/include/replication/logical.h
index 6e0b7628001..37315a424e0 100644
--- a/src/include/replication/logical.h
+++ b/src/include/replication/logical.h
@@ -138,7 +138,7 @@ extern bool DecodingContextReady(LogicalDecodingContext *ctx);
 extern void FreeDecodingContext(LogicalDecodingContext *ctx);
 
 extern void LogicalIncreaseXminForSlot(XLogRecPtr current_lsn,
-									   TransactionId xmin);
+									   TransactionId xmin, bool catalog);
 extern void LogicalIncreaseRestartDecodingForSlot(XLogRecPtr current_lsn,
 												  XLogRecPtr restart_lsn);
 extern void LogicalConfirmReceivedLocation(XLogRecPtr lsn);
diff --git a/src/include/replication/reorderbuffer.h b/src/include/replication/reorderbuffer.h
index ff825e4b7b2..cdefc4808df 100644
--- a/src/include/replication/reorderbuffer.h
+++ b/src/include/replication/reorderbuffer.h
@@ -104,6 +104,12 @@ typedef struct ReorderBufferChange
 			HeapTuple	oldtuple;
 			/* valid for INSERT || UPDATE */
 			HeapTuple	newtuple;
+
+			/*
+			 * valid for UPDATE - this is the physical location of the old
+			 * tuple version, valid even if 'oldtuple' is NULL.
+			 */
+			BlockNumber old_blknum;
 		}			tp;
 
 		/*
@@ -763,6 +769,7 @@ extern void ReorderBufferProcessXid(ReorderBuffer *rb, TransactionId xid, XLogRe
 
 extern void ReorderBufferXidSetCatalogChanges(ReorderBuffer *rb, TransactionId xid, XLogRecPtr lsn);
 extern bool ReorderBufferXidHasCatalogChanges(ReorderBuffer *rb, TransactionId xid);
+extern bool ReorderBufferXidHasHeapChanges(ReorderBuffer *rb, TransactionId xid);
 extern bool ReorderBufferXidHasBaseSnapshot(ReorderBuffer *rb, TransactionId xid);
 
 extern bool ReorderBufferRememberPrepareInfo(ReorderBuffer *rb, TransactionId xid,
diff --git a/src/test/modules/injection_points/Makefile b/src/test/modules/injection_points/Makefile
index c01d2fb095c..9c942599c49 100644
--- a/src/test/modules/injection_points/Makefile
+++ b/src/test/modules/injection_points/Makefile
@@ -15,6 +15,7 @@ REGRESS_OPTS = --dlpath=$(top_builddir)/src/test/regress
 ISOLATION = basic \
 	    inplace \
 	    repack \
+	    repack_snapshots \
 	    repack_temporal \
 	    repack_temporal_multirange \
 	    repack_toast \
diff --git a/src/test/modules/injection_points/expected/repack_snapshots.out b/src/test/modules/injection_points/expected/repack_snapshots.out
new file mode 100644
index 00000000000..98bac9ea882
--- /dev/null
+++ b/src/test/modules/injection_points/expected/repack_snapshots.out
@@ -0,0 +1,401 @@
+Parsed test spec with 2 sessions
+
+starting permutation: load repack change_new_beyond change_old_beyond check2 wakeup_new_range check1
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step load: 
+	SELECT load(1);
+
+load
+----
+    
+(1 row)
+
+step repack: 
+	REPACK (CONCURRENTLY) repack_test;
+ <waiting ...>
+step change_new_beyond: 
+	INSERT INTO repack_test
+	SELECT max(i) + 1, gen_external()
+	FROM repack_test
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT min(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        0|        1
+(1 row)
+
+step change_old_beyond: 
+	DELETE FROM repack_test
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        1|        1
+(1 row)
+
+step check2: 
+	INSERT INTO data_s2(i, j)
+	SELECT i, j FROM repack_test;
+
+step wakeup_new_range: 
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step repack: <... completed>
+step check1: 
+	INSERT INTO data_s1(i, j)
+	SELECT i, j FROM repack_test;
+
+	SELECT count(*) > 0 FROM repack_test;
+
+	SELECT count(*)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	WHERE d1.i ISNULL OR d2.i ISNULL;
+
+?column?
+--------
+t       
+(1 row)
+
+count
+-----
+    0
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: load2 load2_vacuum repack change_old_beyond2 check2 wakeup_new_range wait_new_range wakeup_new_range check1
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step load2: 
+	SELECT load(2);
+
+	DELETE FROM repack_test WHERE i < 100;
+
+load
+----
+    
+(1 row)
+
+step load2_vacuum: 
+	VACUUM repack_test;
+
+step repack: 
+	REPACK (CONCURRENTLY) repack_test;
+ <waiting ...>
+step change_old_beyond2: 
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test WHERE tid_block(ctid) = 1)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block|tid_block
+---------+---------
+        1|        0
+(1 row)
+
+step check2: 
+	INSERT INTO data_s2(i, j)
+	SELECT i, j FROM repack_test;
+
+step wakeup_new_range: 
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_new_range: 
+	SELECT wait_for_blocks_scanned('repack_test'::regclass, 3);
+
+wait_for_blocks_scanned
+-----------------------
+                       
+(1 row)
+
+step wakeup_new_range: 
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step repack: <... completed>
+step check1: 
+	INSERT INTO data_s1(i, j)
+	SELECT i, j FROM repack_test;
+
+	SELECT count(*) > 0 FROM repack_test;
+
+	SELECT count(*)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	WHERE d1.i ISNULL OR d2.i ISNULL;
+
+?column?
+--------
+t       
+(1 row)
+
+count
+-----
+    0
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: load repack_pkey change_new_beyond change_old_beyond check2 wakeup_new_range check1 check1_order_asc
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step load: 
+	SELECT load(1);
+
+load
+----
+    
+(1 row)
+
+step repack_pkey: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step change_new_beyond: 
+	INSERT INTO repack_test
+	SELECT max(i) + 1, gen_external()
+	FROM repack_test
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT min(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        0|        1
+(1 row)
+
+step change_old_beyond: 
+	DELETE FROM repack_test
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        1|        1
+(1 row)
+
+step check2: 
+	INSERT INTO data_s2(i, j)
+	SELECT i, j FROM repack_test;
+
+step wakeup_new_range: 
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step repack_pkey: <... completed>
+step check1: 
+	INSERT INTO data_s1(i, j)
+	SELECT i, j FROM repack_test;
+
+	SELECT count(*) > 0 FROM repack_test;
+
+	SELECT count(*)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	WHERE d1.i ISNULL OR d2.i ISNULL;
+
+?column?
+--------
+t       
+(1 row)
+
+count
+-----
+    0
+(1 row)
+
+step check1_order_asc: 
+	SELECT i FROM repack_test LIMIT 10;
+
+ i
+--
+ 2
+ 3
+ 4
+ 5
+ 6
+ 7
+ 8
+ 9
+10
+11
+(10 rows)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: load repack_other_index change_new_beyond change_old_beyond check2 wakeup_new_range check1 check1_order_desc
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step load: 
+	SELECT load(1);
+
+load
+----
+    
+(1 row)
+
+step repack_other_index: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_i_idx;
+ <waiting ...>
+step change_new_beyond: 
+	INSERT INTO repack_test
+	SELECT max(i) + 1, gen_external()
+	FROM repack_test
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT min(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        0|        1
+(1 row)
+
+step change_old_beyond: 
+	DELETE FROM repack_test
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        1|        1
+(1 row)
+
+step check2: 
+	INSERT INTO data_s2(i, j)
+	SELECT i, j FROM repack_test;
+
+step wakeup_new_range: 
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step repack_other_index: <... completed>
+step check1: 
+	INSERT INTO data_s1(i, j)
+	SELECT i, j FROM repack_test;
+
+	SELECT count(*) > 0 FROM repack_test;
+
+	SELECT count(*)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	WHERE d1.i ISNULL OR d2.i ISNULL;
+
+?column?
+--------
+t       
+(1 row)
+
+count
+-----
+    0
+(1 row)
+
+step check1_order_desc: 
+	WITH tmp(diff) as (
+		SELECT i - lag(i, 1, 10000) OVER (ORDER BY ctid)
+		FROM repack_test
+		LIMIT 10)
+	SELECT * FROM tmp WHERE diff > -1;
+
+diff
+----
+(0 rows)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/meson.build b/src/test/modules/injection_points/meson.build
index 59dba1cb023..d432a6b8f76 100644
--- a/src/test/modules/injection_points/meson.build
+++ b/src/test/modules/injection_points/meson.build
@@ -46,6 +46,7 @@ tests += {
       'basic',
       'inplace',
       'repack',
+      'repack_snapshots',
       'repack_temporal',
       'repack_temporal_multirange',
       'repack_toast',
diff --git a/src/test/modules/injection_points/specs/repack_snapshots.spec b/src/test/modules/injection_points/specs/repack_snapshots.spec
new file mode 100644
index 00000000000..23b478fd6f4
--- /dev/null
+++ b/src/test/modules/injection_points/specs/repack_snapshots.spec
@@ -0,0 +1,272 @@
+# REPACK (CONCURRENTLY) - use one snapshot per block range.
+setup
+{
+	CREATE EXTENSION injection_points;
+
+	CREATE TABLE repack_test(i int PRIMARY KEY, j text);
+	CREATE INDEX ON repack_test(i DESC);
+	CREATE TABLE relfilenodes(node oid);
+
+	CREATE TABLE data_s1(i int, j text);
+	CREATE TABLE data_s2(i int, j text);
+
+	-- Keep inserting tuples into repack_test until several tuples need to
+        -- be inserted into block number last_block.
+	CREATE FUNCTION load(last_block int)
+	RETURNS void
+	LANGUAGE 'plpgsql'
+	AS $$
+	    DECLARE
+		cnt	int;
+	    BEGIN
+		INSERT INTO repack_test VALUES (1, gen_external());
+
+		LOOP
+		    WITH max(m) AS (SELECT max(i) FROM repack_test)
+		    INSERT INTO repack_test(i, j)
+		    SELECT m + x, gen_external()
+		    FROM generate_series(1, 100) s(x), max;
+
+		    SELECT count(*)
+		    FROM repack_test WHERE tid_block(ctid) = last_block
+		    INTO cnt;
+
+		    IF cnt >= 10 THEN
+			EXIT;
+		    END IF;
+		END LOOP;
+	    END;
+	$$;
+
+	-- Generate a string of random characters that is not likely to be
+	-- compressed, but is big enough to be stored externally.
+	CREATE FUNCTION gen_external()
+	RETURNS text
+	LANGUAGE sql as $$
+		SELECT string_agg(chr(65 + trunc(25 * random())::int), '')
+		FROM generate_series(1, 2048) s(x);
+	$$;
+
+	-- Wait until the number of blocks involved in the scan by REPACK has
+	-- reached _blocks.
+	CREATE FUNCTION wait_for_blocks_scanned(_relid oid, _blocks int)
+	RETURNS void
+	LANGUAGE 'plpgsql'
+	AS $$
+		DECLARE
+			cnt	int;
+		BEGIN
+			LOOP
+				PERFORM pg_stat_clear_snapshot();
+
+				SELECT heap_blks_scanned
+				FROM pg_stat_progress_repack
+				WHERE relid = _relid
+				INTO cnt;
+
+				IF cnt >= _blocks THEN
+					EXIT;
+				END IF;
+
+				PERFORM pg_sleep(0.1);
+			END LOOP;
+		END;
+	$$;
+}
+
+teardown
+{
+	DROP TABLE repack_test;
+	DROP EXTENSION injection_points;
+
+	DROP TABLE relfilenodes;
+	DROP TABLE data_s1;
+	DROP TABLE data_s2;
+
+	DROP FUNCTION load(int);
+	DROP FUNCTION gen_external();
+	DROP FUNCTION wait_for_blocks_scanned(oid, int);
+}
+
+session s1
+setup
+{
+	SET repack_snapshot_after = 1;
+
+	SELECT injection_points_set_local();
+	SELECT injection_points_attach('repack-concurrently-new-range', 'wait');
+}
+# The most practical way to test the corner cases is to set range size to 1
+# block. To initialize, insert new tuples until we have several tuples in the
+# 2nd block.
+step load
+{
+	SELECT load(1);
+}
+# Start the initial load and wait when the first range has been completed.
+step repack
+{
+	REPACK (CONCURRENTLY) repack_test;
+}
+# The same, but with clustering.
+step repack_pkey
+{
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+}
+# Clustering by other than the identity index.
+step repack_other_index
+{
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_i_idx;
+}
+# Check the table from the perspective of s1.
+step check1
+{
+	INSERT INTO data_s1(i, j)
+	SELECT i, j FROM repack_test;
+
+	SELECT count(*) > 0 FROM repack_test;
+
+	SELECT count(*)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	WHERE d1.i ISNULL OR d2.i ISNULL;
+}
+# Check the ordering where appropriate. We don't know the exact number of
+# rows, so just check a sample. (The replayed concurrent changes are not
+# ordered, but those shouldn't fit into the first 10 rows.)
+step check1_order_asc
+{
+	SELECT i FROM repack_test LIMIT 10;
+}
+# Due to the special way of loading the data (see the load() function above)
+# we don't know the maximum value. To make the test output deterministic,
+# check for cases where the current row is not lower than the previous row.
+step check1_order_desc
+{
+	WITH tmp(diff) as (
+		SELECT i - lag(i, 1, 10000) OVER (ORDER BY ctid)
+		FROM repack_test
+		LIMIT 10)
+	SELECT * FROM tmp WHERE diff > -1;
+}
+teardown
+{
+	SELECT injection_points_detach('repack-concurrently-new-range');
+}
+
+session s2
+# Test processing of changes such that the new tuple is beyond the current
+# range. Specifically for UPDATE, the old tuple should be in the current
+# range. So when applying it, we have to convert it to DELETE.
+step change_new_beyond
+{
+	INSERT INTO repack_test
+	SELECT max(i) + 1, gen_external()
+	FROM repack_test
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT min(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+}
+# Test processing of changes such that the old tuple is beyond the current
+# range. The UPDATE puts also the new tuple beyond the current range.
+step change_old_beyond
+{
+	DELETE FROM repack_test
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+}
+# Arrange for UPDATE to put the new tuples into block 0. In particular, fill
+# the block 1 and delete some tuples from block 1.
+step load2
+{
+	SELECT load(2);
+
+	DELETE FROM repack_test WHERE i < 100;
+}
+# Logically this belongs to the previous step, however VACUUM cannot run
+# inside a transaction block.
+step load2_vacuum
+{
+	VACUUM repack_test;
+}
+# This UPDATE should put the new tuple into block 0 (the current range). So
+# when replaying it, we have to convert it to INSERT. That includes fetching
+# the old tuple's TOAST from the TOAST table because the old tuple is not
+# available during the replay.
+step change_old_beyond2
+{
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test WHERE tid_block(ctid) = 1)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+}
+# Check the table from the perspective of s4.
+step check2
+{
+	INSERT INTO data_s2(i, j)
+	SELECT i, j FROM repack_test;
+}
+step wakeup_new_range
+{
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+}
+# This step is used to make sure that REPACK is waiting on an injection point
+# before finalizing the 2nd block. However, the coding is such that the
+# counter is incremented at the top of the loop (see
+# heapam_relation_copy_for_cluster()), so we wait until it's at least 3.
+step wait_new_range
+{
+	SELECT wait_for_blocks_scanned('repack_test'::regclass, 3);
+}
+
+# Test if snapshots are used correctly to scan block ranges.
+permutation
+	load
+	repack
+	change_new_beyond
+	change_old_beyond
+	check2
+	wakeup_new_range
+	check1
+
+# Special attention is needed to update tuple in block 1 so that the new tuple
+# appears in block 0. The preparation includes VACUUM, which in turn cannot
+# proceed while REPACK is in progress. That's why we need a separate
+# permutation. Note that two wake-ups are needed as we have two range
+# boundaries now. However we need to wait in between to make sure that the
+# second waiting started before we try to wake it up.
+permutation
+	load2
+	load2_vacuum
+	repack
+	change_old_beyond2
+	check2
+	wakeup_new_range
+	wait_new_range
+	wakeup_new_range
+	check1
+
+# The first permutation with identity index as the clustering index.
+permutation
+	load
+	repack_pkey
+	change_new_beyond
+	change_old_beyond
+	check2
+	wakeup_new_range
+	check1
+	check1_order_asc
+# The first permutation with another clustering index.
+permutation
+	load
+	repack_other_index
+	change_new_beyond
+	change_old_beyond
+	check2
+	wakeup_new_range
+	check1
+	check1_order_desc
-- 
2.52.0
From ee135ee4c57d978594d483af7fcb3997c3e5094b Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 15 Jul 2026 10:37:12 +0200
Subject: [PATCH 8/8] Make REPACK (CONCURRENTLY) MVCC-safe.

In the initial implementation of REPACK (CONCURRENTLY), tuples in the new
table are marked with XID of the transaction the command executes in. This is
simpler to implement, but can cause surprising behavior: on specific
conditions, the table contents may become invisible for transactions that
should see it. (Both ALTER TABLE and TRUNCATE commands have this problem.)
Besides that, as REPACK (CONCURRENTLY) needs to have XID assigned (to mark the
new tuples) while it copies the data (which can take long time), it can
restrict the progress of the xmin horizons for VACUUM quite a bit.

This patch teaches REPACK (CONCURRENTLY) to transfer the visibility
information from the old table to the new one. Thus there's no need to have
XID assigned during the copying, and therefore VACUUM (of other tables) is not
restricted anymore.

A new type of snapshot is used to look for the existing tuples in the new
table, in order to apply the concurrent data changes (i.e. changes performed
by other transactions during the initial copying). This snapshot assumes that
the new table does not contain any tuples created by aborted
transactions. This is simpler and more efficient than building historic MVCC
snapshots for the look-ups.

One problem that needs special attention is tuple freezing. Besides page
pruning code, freezing is currently implemented only in rewriteheap.c. Thus it
seems simpler to teach REPACK (CONCURRENTLY) to copy the data using
rewriteheap.c rather than implementing freezing of individual tuples
independent from page pruning. (The latter would include WAL logging of the
individual frozen tuples, whereas rewriteheap.c logs the whole page at once.)
Disadvantage of this approach is that we only freeze tuples during the initial
copying, while the data changes decoded from WAL are replayed without
freezing.

This patch also uses rewriteheap.c (in particular raw_heap_insert()) to insert
TOAST tuples. That seems to be the easiest way to avoid getting XID assigned
just to insert TOAST. However, due to not using a new XID, even TOAST needs to
be frozen now - again, we use the existing code in rewriteheap.c.
---
 doc/src/sgml/mvcc.sgml                        |  12 +-
 doc/src/sgml/ref/repack.sgml                  |   9 -
 src/backend/access/common/toast_internals.c   |  45 +-
 src/backend/access/heap/heapam.c              |  84 +++-
 src/backend/access/heap/heapam_handler.c      | 265 +++++++++--
 src/backend/access/heap/heapam_visibility.c   |  62 +++
 src/backend/access/heap/heaptoast.c           |  23 +-
 src/backend/access/heap/rewriteheap.c         | 263 ++++++++---
 src/backend/access/table/tableam.c            |   1 +
 src/backend/access/table/toast_helper.c       |  11 +-
 src/backend/access/transam/xloginsert.c       |  31 +-
 src/backend/access/transam/xlogrecovery.c     |   7 +-
 src/backend/catalog/indexing.c                |   2 +-
 src/backend/commands/matview.c                |   3 +-
 src/backend/commands/repack.c                 | 433 ++++++++++++------
 src/backend/commands/tablecmds.c              |   1 +
 src/backend/executor/nodeModifyTable.c        |   1 +
 src/backend/replication/logical/decode.c      | 110 ++++-
 .../replication/logical/reorderbuffer.c       |   3 +
 .../utils/activity/wait_event_names.txt       |   1 +
 src/include/access/heapam.h                   |   4 +-
 src/include/access/heaptoast.h                |   7 +-
 src/include/access/rewriteheap.h              |  13 +-
 src/include/access/tableam.h                  |  28 +-
 src/include/access/toast_helper.h             |  14 +-
 src/include/access/toast_internals.h          |   9 +-
 src/include/access/xlog_internal.h            |   2 +-
 src/include/access/xloginsert.h               |   1 +
 src/include/access/xlogrecord.h               |   8 +
 src/include/commands/repack.h                 |  17 +-
 src/include/utils/snapmgr.h                   |  12 +
 src/include/utils/snapshot.h                  |  27 +-
 .../injection_points/expected/repack.out      |  10 +-
 .../injection_points/specs/repack.spec        |  23 +-
 34 files changed, 1200 insertions(+), 342 deletions(-)

diff --git a/doc/src/sgml/mvcc.sgml b/doc/src/sgml/mvcc.sgml
index 9cb52302f23..8a6bb165ecb 100644
--- a/doc/src/sgml/mvcc.sgml
+++ b/doc/src/sgml/mvcc.sgml
@@ -1883,17 +1883,15 @@ SELECT pg_advisory_lock(q.id) FROM
    <title>Caveats</title>
 
    <para>
-    Some commands, currently only <link linkend="sql-truncate"><command>TRUNCATE</command></link>, the
-    table-rewriting forms of <link linkend="sql-altertable"><command>ALTER
-    TABLE</command></link> and <command>REPACK</command> with
-    the <literal>CONCURRENTLY</literal> option, are not
+    Some DDL commands, currently only <link linkend="sql-truncate"><command>TRUNCATE</command></link> and the
+    table-rewriting forms of <link linkend="sql-altertable"><command>ALTER TABLE</command></link>, are not
     MVCC-safe.  This means that after the truncation or rewrite commits, the
     table will appear empty to concurrent transactions, if they are using a
-    snapshot taken before the command committed.  This will only be an
+    snapshot taken before the DDL command committed.  This will only be an
     issue for a transaction that did not access the table in question
-    before the command started &mdash; any transaction that has done so
+    before the DDL command started &mdash; any transaction that has done so
     would hold at least an <literal>ACCESS SHARE</literal> table lock,
-    which would block the truncating or rewriting command until that transaction completes.
+    which would block the DDL command until that transaction completes.
     So these commands will not cause any apparent inconsistency in the
     table contents for successive queries on the target table, but they
     could cause visible inconsistency between the contents of the target
diff --git a/doc/src/sgml/ref/repack.sgml b/doc/src/sgml/ref/repack.sgml
index 0cb72b6b289..acbd92ef79d 100644
--- a/doc/src/sgml/ref/repack.sgml
+++ b/doc/src/sgml/ref/repack.sgml
@@ -300,15 +300,6 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] USING
        </listitem>
       </itemizedlist>
      </para>
-
-     <warning>
-      <para>
-       <command>REPACK</command> with the <literal>CONCURRENTLY</literal>
-       option is not MVCC-safe, see <xref linkend="mvcc-caveats"/> for
-       details.
-      </para>
-     </warning>
-
     </listitem>
    </varlistentry>
 
diff --git a/src/backend/access/common/toast_internals.c b/src/backend/access/common/toast_internals.c
index 77d42e7ed65..ba80fed0ca9 100644
--- a/src/backend/access/common/toast_internals.c
+++ b/src/backend/access/common/toast_internals.c
@@ -112,12 +112,15 @@ toast_compress_datum(Datum value, char cmethod)
  * rel: the main relation we're working with (not the toast rel!)
  * value: datum to be pushed to toast storage
  * oldexternal: if not NULL, toast pointer previously representing the datum
+ * rwstate: state needed for "raw insert".
+ * tup_main: the tuple whose attribute we're saving
  * options: options to be passed to heap_insert() for toast rows
  * ----------
  */
 Datum
 toast_save_datum(Relation rel, Datum value,
-				 varlena *oldexternal, uint32 options)
+				 varlena *oldexternal, RewriteState rwstate,
+				 HeapTuple tup_main, uint32 options)
 {
 	Relation	toastrel;
 	Relation   *toastidxs;
@@ -311,7 +314,40 @@ toast_save_datum(Relation rel, Datum value,
 
 		toasttup = heap_form_tuple(toasttupDesc, t_values, t_isnull);
 
-		heap_insert(toastrel, toasttup, mycid, options, NULL);
+		if (rwstate == NULL)
+		{
+			/*
+			 * With REUSE_XID we'd need regular freezing, however that is
+			 * currently implemented only as a part of VACUUM or in
+			 * rewriteheap.c. On the other hand, tuples containing a new XID
+			 * can be marked as frozen in special cases - see the current uses
+			 * of TABLE_INSERT_FROZEN.
+			 */
+			Assert((options & HEAP_INSERT_FROZEN) == 0 ||
+				   (options & TABLE_REUSE_XID) == 0);
+
+			/*
+			 * If an existing XID should be used, the entire visibility info
+			 * of the TOAST tuple should be equal to that of corresponding
+			 * tuple in the main table.
+			 */
+			if (options & TABLE_REUSE_XID)
+				rewrite_copy_visibility_info(toasttup, tup_main);
+
+			heap_insert(toastrel, toasttup, mycid, options, NULL);
+		}
+		else
+		{
+			/*
+			 * During heap rewrite, XID is always reused and the tuple is
+			 * always frozen - we do not expect the user to tell us what to
+			 * do.
+			 */
+			Assert((options & HEAP_INSERT_FROZEN) == 0 &&
+				   (options & TABLE_REUSE_XID) == 0);
+
+			rewrite_heap_tuple_no_chains(rwstate, tup_main, toasttup, true);
+		}
 
 		/*
 		 * Create the index entry.  We cheat a little here by not using
@@ -373,7 +409,8 @@ toast_save_datum(Relation rel, Datum value,
  * ----------
  */
 void
-toast_delete_datum(Relation rel, Datum value, bool is_speculative)
+toast_delete_datum(Relation rel, Datum value, bool is_speculative,
+				   TransactionId xid)
 {
 	varlena    *attr = (varlena *) DatumGetPointer(value);
 	varatt_external toast_pointer;
@@ -425,7 +462,7 @@ toast_delete_datum(Relation rel, Datum value, bool is_speculative)
 		if (is_speculative)
 			heap_abort_speculative(toastrel, &toasttup->t_self);
 		else
-			simple_heap_delete(toastrel, &toasttup->t_self);
+			simple_heap_delete(toastrel, &toasttup->t_self, xid);
 	}
 
 	/*
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 9cdc221675b..35012434d27 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -63,7 +63,7 @@ static XLogRecPtr log_heap_update(Relation reln, Buffer oldbuf,
 								  Buffer newbuf, HeapTuple oldtup,
 								  HeapTuple newtup, HeapTuple old_key_tuple,
 								  bool all_visible_cleared, bool new_all_visible_cleared,
-								  bool walLogical);
+								  bool walLogical, TransactionId xid);
 #ifdef USE_ASSERT_CHECKING
 static void check_lock_if_inplace_updateable_rel(Relation relation,
 												 const ItemPointerData *otid,
@@ -2004,7 +2004,7 @@ void
 heap_insert(Relation relation, HeapTuple tup, CommandId cid,
 			uint32 options, BulkInsertState bistate)
 {
-	TransactionId xid = GetCurrentTransactionId();
+	TransactionId xid;
 	HeapTuple	heaptup;
 	Buffer		buffer;
 	Page		page;
@@ -2017,6 +2017,13 @@ heap_insert(Relation relation, HeapTuple tup, CommandId cid,
 
 	AssertHasSnapshotForToast(relation);
 
+	/* The caller might need to preserve the existing xmin. */
+	if ((options & TABLE_REUSE_XID) == 0)
+		xid = GetCurrentTransactionId();
+	else
+		xid = HeapTupleHeaderGetXmin(tup->t_data);
+	Assert(TransactionIdIsValid(xid));
+
 	/*
 	 * Fill in tuple header fields and toast the tuple if necessary.
 	 *
@@ -2159,6 +2166,13 @@ heap_insert(Relation relation, HeapTuple tup, CommandId cid,
 		/* filtering by origin on a row level is much more efficient */
 		XLogSetRecordFlags(XLOG_INCLUDE_ORIGIN);
 
+		/*
+		 * Even if we don't have XID assigned, valid XID is necessary for
+		 * recovery and streaming replication to work.
+		 */
+		if (options & TABLE_REUSE_XID)
+			XLogSetRecordXid(xid);
+
 		recptr = XLogInsert(RM_HEAP_ID, info);
 
 		PageSetLSN(page, recptr);
@@ -2236,7 +2250,7 @@ heap_prepare_insert(Relation relation, HeapTuple tup, TransactionId xid,
 		return tup;
 	}
 	else if (HeapTupleHasExternal(tup) || tup->t_len > TOAST_TUPLE_THRESHOLD)
-		return heap_toast_insert_or_update(relation, tup, NULL, options);
+		return heap_toast_insert_or_update(relation, tup, NULL, NULL, options);
 	else
 		return tup;
 }
@@ -2299,6 +2313,8 @@ heap_multi_insert(Relation relation, TupleTableSlot **slots, int ntuples,
 
 	/* currently not needed (thus unsupported) for heap_multi_insert() */
 	Assert(!(options & HEAP_INSERT_NO_LOGICAL));
+	/* Likewise. */
+	Assert(!(options & TABLE_REUSE_XID));
 
 	AssertHasSnapshotForToast(relation);
 
@@ -2715,11 +2731,11 @@ xmax_infomask_changed(uint16 new_infomask, uint16 old_infomask)
  */
 TM_Result
 heap_delete(Relation relation, const ItemPointerData *tid,
-			CommandId cid, uint32 options, Snapshot crosscheck,
+			TransactionId xid, CommandId cid, uint32 options,
+			Snapshot crosscheck,
 			bool wait, TM_FailureData *tmfd)
 {
 	TM_Result	result;
-	TransactionId xid = GetCurrentTransactionId();
 	ItemId		lp;
 	HeapTupleData tp;
 	Page		page;
@@ -2736,11 +2752,14 @@ heap_delete(Relation relation, const ItemPointerData *tid,
 	bool		all_visible_cleared = false;
 	HeapTuple	old_key_tuple = NULL;	/* replica identity of the tuple */
 	bool		old_key_copied = false;
+	bool		override_xid = false;
 
 	Assert(ItemPointerIsValid(tid));
 
 	AssertHasSnapshotForToast(relation);
 
+	Assert((options & TABLE_REUSE_XID) == 0);
+
 	/*
 	 * Forbid this during a parallel operation, lest it allocate a combo CID.
 	 * Other workers might need that combo CID for visibility checks, and we
@@ -2751,6 +2770,12 @@ heap_delete(Relation relation, const ItemPointerData *tid,
 				(errcode(ERRCODE_INVALID_TRANSACTION_STATE),
 				 errmsg("cannot delete tuples during a parallel operation")));
 
+	/* Caller can override the xid. */
+	if (!TransactionIdIsValid(xid))
+		xid = GetCurrentTransactionId();
+	else
+		override_xid = true;
+
 	block = ItemPointerGetBlockNumber(tid);
 	buffer = ReadBuffer(relation, block);
 	page = BufferGetPage(buffer);
@@ -3089,6 +3114,9 @@ l1:
 
 		/* filtering by origin on a row level is much more efficient */
 		XLogSetRecordFlags(XLOG_INCLUDE_ORIGIN);
+		/* See heap_insert() for details. */
+		if (override_xid)
+			XLogSetRecordXid(xid);
 
 		recptr = XLogInsert(RM_HEAP_ID, XLOG_HEAP_DELETE);
 
@@ -3115,7 +3143,7 @@ l1:
 		Assert(!HeapTupleHasExternal(&tp));
 	}
 	else if (HeapTupleHasExternal(&tp))
-		heap_toast_delete(relation, &tp, false);
+		heap_toast_delete(relation, &tp, false, xid);
 
 	/*
 	 * Mark tuple for invalidation from system caches at next command
@@ -3148,14 +3176,18 @@ l1:
  * the target tuple are not expected (for example, because we have a lock
  * on the relation associated with the tuple).  Any failure is reported
  * via ereport().
+ *
+ * XXX Add simple_heap_delete_xid() so that the signature of this function can
+ * stay intact?
  */
 void
-simple_heap_delete(Relation relation, const ItemPointerData *tid)
+simple_heap_delete(Relation relation, const ItemPointerData *tid,
+				   TransactionId xid)
 {
 	TM_Result	result;
 	TM_FailureData tmfd;
 
-	result = heap_delete(relation, tid,
+	result = heap_delete(relation, tid, xid,
 						 GetCurrentCommandId(true),
 						 0,
 						 InvalidSnapshot,
@@ -3204,7 +3236,7 @@ heap_update(Relation relation, const ItemPointerData *otid, HeapTuple newtup,
 			TU_UpdateIndexes *update_indexes)
 {
 	TM_Result	result;
-	TransactionId xid = GetCurrentTransactionId();
+	TransactionId xid;
 	Bitmapset  *hot_attrs;
 	Bitmapset  *sum_attrs;
 	Bitmapset  *key_attrs;
@@ -3267,6 +3299,13 @@ heap_update(Relation relation, const ItemPointerData *otid, HeapTuple newtup,
 	check_lock_if_inplace_updateable_rel(relation, otid, newtup);
 #endif
 
+	/* The caller might need to preserve the existing xmin. */
+	if ((options & TABLE_REUSE_XID) == 0)
+		xid = GetCurrentTransactionId();
+	else
+		xid = HeapTupleHeaderGetXmin(newtup->t_data);
+	Assert(TransactionIdIsValid(xid));
+
 	/*
 	 * Fetch the list of attributes to be checked for various operations.
 	 *
@@ -3871,7 +3910,8 @@ l2:
 		if (need_toast)
 		{
 			/* Note we always use WAL and FSM during updates */
-			heaptup = heap_toast_insert_or_update(relation, newtup, &oldtup, 0);
+			heaptup = heap_toast_insert_or_update(relation, newtup, &oldtup,
+												  NULL, options);
 			newtupsize = MAXALIGN(heaptup->t_len);
 		}
 		else
@@ -4099,7 +4139,9 @@ l2:
 								 old_key_tuple,
 								 all_visible_cleared,
 								 all_visible_cleared_new,
-								 walLogical);
+								 walLogical,
+								 options & TABLE_REUSE_XID ? xid : InvalidTransactionId);
+
 		if (newbuf != buffer)
 		{
 			PageSetLSN(newpage, recptr);
@@ -5297,7 +5339,9 @@ compute_new_xmax_infomask(TransactionId xmax, uint16 old_infomask,
 	uint16		new_infomask,
 				new_infomask2;
 
-	Assert(TransactionIdIsCurrentTransactionId(add_to_xmax));
+	/* REPACK (CONCURRENTLY) might not have XID assigned. */
+	Assert(TransactionIdIsCurrentTransactionId(add_to_xmax) ||
+		   !TransactionIdIsValid(GetTopTransactionIdIfAny()));
 
 l5:
 	new_infomask = 0;
@@ -6269,7 +6313,9 @@ heap_abort_speculative(Relation relation, const ItemPointerData *tid)
 	if (HeapTupleHasExternal(&tp))
 	{
 		Assert(!IsToastRelation(relation));
-		heap_toast_delete(relation, &tp, true);
+		/* XID overriding is not needed for speculative abort. */
+		Assert(TransactionIdIsValid(GetCurrentTransactionIdIfAny()));
+		heap_toast_delete(relation, &tp, true, GetCurrentTransactionId());
 	}
 
 	/*
@@ -6695,7 +6741,8 @@ FreezeMultiXactId(MultiXactId multi, uint16 t_infomask,
 		pagefrz->freeze_required = true;
 		return InvalidTransactionId;
 	}
-	else if (MultiXactIdPrecedes(multi, cutoffs->relminmxid))
+	else if (MultiXactIdIsValid(cutoffs->relminmxid) &&
+			 MultiXactIdPrecedes(multi, cutoffs->relminmxid))
 		ereport(ERROR,
 				(errcode(ERRCODE_DATA_CORRUPTED),
 				 errmsg_internal("found multixact %u from before relminmxid %u",
@@ -7199,7 +7246,8 @@ heap_prepare_freeze_tuple(HeapTupleHeader tuple,
 	else if (TransactionIdIsNormal(xid))
 	{
 		/* Raw xmax is normal XID */
-		if (TransactionIdPrecedes(xid, cutoffs->relfrozenxid))
+		if (TransactionIdIsValid(cutoffs->relfrozenxid) &&
+			TransactionIdPrecedes(xid, cutoffs->relfrozenxid))
 			ereport(ERROR,
 					(errcode(ERRCODE_DATA_CORRUPTED),
 					 errmsg_internal("found xmax %u from before relfrozenxid %u",
@@ -8776,7 +8824,7 @@ log_heap_update(Relation reln, Buffer oldbuf,
 				Buffer newbuf, HeapTuple oldtup, HeapTuple newtup,
 				HeapTuple old_key_tuple,
 				bool all_visible_cleared, bool new_all_visible_cleared,
-				bool walLogical)
+				bool walLogical, TransactionId xid)
 {
 	xl_heap_update xlrec;
 	xl_heap_header xlhdr;
@@ -8983,6 +9031,10 @@ log_heap_update(Relation reln, Buffer oldbuf,
 	/* filtering by origin on a row level is much more efficient */
 	XLogSetRecordFlags(XLOG_INCLUDE_ORIGIN);
 
+	/* See heap_insert() for details. */
+	if (TransactionIdIsValid(xid))
+		XLogSetRecordXid(xid);
+
 	recptr = XLogInsert(RM_HEAP_ID, info);
 
 	return recptr;
diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index 1096d9d4dc9..1eec39e6ebb 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -50,11 +50,24 @@
 #include "utils/injection_point.h"
 #include "utils/rel.h"
 #include "utils/tuplesort.h"
-
-static Snapshot finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
-									 BlockNumber start, BlockNumber *end_p);
+#include "utils/wait_event_types.h"
+
+static void update_identity_index(ChangeContext *chgcxt,
+								  BlockNumber range_start,
+								  BlockNumber range_end);
+static Snapshot finalize_block_range(Relation rel_old, Relation rel_dst,
+									 TransactionId oldest_xmin,
+									 TransactionId xid_cutoff,
+									 MultiXactId multi_cutoff,
+									 ChangeContext *chgcxt, BlockNumber cur,
+									 BlockNumber start, BlockNumber *end_p,
+									 RewriteState *rwstate_p,
+									 BlockNumber *range_start_dst_p);
 static void reform_and_rewrite_tuple(TupleTableSlot *src, TupleTableSlot *reform,
 									 RewriteState rwstate);
+static void heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
+								   TupleTableSlot *reform,
+								   RewriteState rwstate);
 
 static bool SampleHeapTupleVisible(TableScanDesc scan, Buffer buffer,
 								   HeapTuple tuple,
@@ -199,7 +212,8 @@ heapam_tuple_complete_speculative(Relation relation, TupleTableSlot *slot,
 }
 
 static TM_Result
-heapam_tuple_delete(Relation relation, ItemPointer tid, CommandId cid,
+heapam_tuple_delete(Relation relation, ItemPointer tid, TransactionId xid,
+					CommandId cid,
 					uint32 options, Snapshot snapshot, Snapshot crosscheck,
 					bool wait, TM_FailureData *tmfd)
 {
@@ -208,7 +222,7 @@ heapam_tuple_delete(Relation relation, ItemPointer tid, CommandId cid,
 	 * the storage itself is cleaning the dead tuples by itself, it is the
 	 * time to call the index tuple deletion also.
 	 */
-	return heap_delete(relation, tid, cid, options, crosscheck, wait,
+	return heap_delete(relation, tid, xid, cid, options, crosscheck, wait,
 					   tmfd);
 }
 
@@ -610,25 +624,30 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	Snapshot	snapshot = NULL;
 	BlockNumber range_start = InvalidBlockNumber;
 	BlockNumber range_end = InvalidBlockNumber;
+	BlockNumber range_start_new = InvalidBlockNumber;
+	BlockNumber range_end_new = InvalidBlockNumber;
+	Relation	rel_dst;
 
 	/* Remember if it's a system catalog */
 	is_system_catalog = IsSystemRelation(OldHeap);
 
+	/* Determine the destination for the tuples. */
+	if (!concurrent || chgcxt->cc_dest_aux == NULL)
+		rel_dst = NewHeap;
+	else
+		rel_dst = chgcxt->cc_dest_aux->rel;
+
 	/*
 	 * Valid smgr_targblock implies something already wrote to the relation.
 	 * This may be harmless, but this function hasn't planned for it.
 	 */
-	Assert(RelationGetTargetBlock(NewHeap) == InvalidBlockNumber);
+	Assert(RelationGetTargetBlock(rel_dst) == InvalidBlockNumber);
 
 	/*
-	 * In non-concurrent mode, initialize the rewrite operation.  This is not
-	 * needed in concurrent mode.
+	 * Initialize the rewrite operation.
 	 */
-	if (!concurrent)
-		rwstate = begin_heap_rewrite(OldHeap, NewHeap, OldestXmin,
-									 *xid_cutoff, *multi_cutoff);
-	else
-		rwstate = NULL;
+	rwstate = begin_heap_rewrite(OldHeap, rel_dst, OldestXmin, *xid_cutoff,
+								 *multi_cutoff, concurrent);
 
 	/*
 	 * Set up sorting if wanted. CONCURRENTLY sorts the tuple w/o tuplesort,
@@ -700,6 +719,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		{
 			range_start = heapScan->rs_startblock;
 			range_end = range_start + repack_pages_per_snapshot;
+			range_start_new = RelationGetNumberOfBlocks(rel_dst);
 		}
 	}
 
@@ -717,24 +737,40 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		InvalidateCatalogSnapshot();
 
 		/*
-		 * As there is no snapshot, our xmin should be invalid now.
-		 *
-		 * XXX xid can still be valid. The next patches in the series fix
-		 * that.
+		 * As there is no snapshot, our xmin should be invalid now. xid should
+		 * be invalid too because our transactions didn't have to do any
+		 * writes yet.
 		 */
 		Assert(!TransactionIdIsValid(MyProc->xmin));
+		Assert(!TransactionIdIsValid(MyProc->xid));
 
 		/*
 		 * Wait until the worker has the initial snapshot and retrieve it.
 		 */
 		snapshot = repack_get_snapshot(chgcxt);
+		chgcxt->cc_last_snapshot_xmin = snapshot->xmin;
+
+		/*
+		 * Since we currently do not freeze tuples when replaying data
+		 * changes, the replayed transactions should not precede the new value
+		 * of relfrozenxid. (Only transactions having XID >= snapshot->xmin
+		 * will be replayed.)
+		 *
+		 * Regarding multixacts, the corresponding cutoff is probably not
+		 * trivial to determine, however there should not be any multixacts in
+		 * the new relation at all: we do not replay tuple locking records an
+		 * that's ok because the tuple locks should no longer exist at the
+		 * moment we acquire AccessExclusiveLock on the old relation.
+		 */
+		if (TransactionIdFollows(*xid_cutoff, snapshot->xmin))
+			*xid_cutoff = snapshot->xmin;
 
 		PushActiveSnapshot(snapshot);
 	}
 
 	/*
 	 * Scan through the OldHeap, either in OldIndex order or sequentially;
-	 * copy each tuple into the NewHeap, or transiently to the tuplesort
+	 * copy each tuple into the rel_dst, or transiently to the tuplesort
 	 * module.  Note that we don't bother sorting dead tuples (they won't get
 	 * to the new table anyway).
 	 */
@@ -905,8 +941,12 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 			/* End of the current range or wraparound? */
 			if (blkno >= range_end || blkno < range_start)
-				snapshot = finalize_block_range(chgcxt, blkno, range_start,
-												&range_end);
+				snapshot = finalize_block_range(OldHeap, rel_dst,
+												OldestXmin, *xid_cutoff,
+												*multi_cutoff,
+												chgcxt, blkno, range_start,
+												&range_end, &rwstate,
+												&range_start_new);
 
 			/* Finally check the tuple visibility. */
 			LockBuffer(buf, BUFFER_LOCK_SHARE);
@@ -942,7 +982,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 			if (!concurrent)
 				reform_and_rewrite_tuple(slot, reform_slot, rwstate);
 			else
-				heap_insert_for_repack(chgcxt, slot, reform_slot);
+				heap_insert_for_repack(chgcxt, slot, reform_slot, rwstate);
 
 			/*
 			 * In indexscan mode and also VACUUM FULL, report increase in
@@ -958,6 +998,13 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	{
 		XLogRecPtr	end_of_wal;
 
+		/* Write out any remaining tuples, and fsync if needed */
+		end_heap_rewrite(rwstate);
+
+		/* Finalize the last range in the new relation. */
+		range_end_new = RelationGetNumberOfBlocks(rel_dst);
+		update_identity_index(chgcxt, range_start_new, range_end_new);
+
 		/*
 		 * Process the changes belonging to the last range.
 		 */
@@ -969,11 +1016,16 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 		/*
 		 * There was an active transaction snapshot on entry, so push one
-		 * before return.
+		 * before return. While there is no active snapshot, also invalidate
+		 * the catalog snapshot so that the xmin horizons for VACUUM can
+		 * advance.
 		 */
 		PopActiveSnapshot();
+		InvalidateCatalogSnapshot();
+		Assert(!TransactionIdIsValid(MyProc->xmin));
+		Assert(!TransactionIdIsValid(MyProc->xid));
+		Assert(!HaveRegisteredOrActiveSnapshot());
 		PushActiveSnapshot(GetTransactionSnapshot());
-
 	}
 
 	if (indexScan != NULL)
@@ -1040,9 +1092,70 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 	ExecDropSingleTupleTableSlot(reform_slot);
 
-	/* Write out any remaining tuples, and fsync if needed */
-	if (rwstate)
+	if (!concurrent)
+	{
+		/*
+		 * In the CONCURRENTLY case, we had to do away with the rewrite state
+		 * earlier.
+		 */
 		end_heap_rewrite(rwstate);
+	}
+}
+
+/*
+ * Scan range of the new / auxiliary relation and insert all the tuples into
+ * the identity index. range_end is the first block of the next range.
+ */
+static void
+update_identity_index(ChangeContext *chgcxt, BlockNumber range_start,
+					  BlockNumber range_end)
+{
+	RepackDest *dest;
+	BlockNumber first_block,
+				last_block;
+	ItemPointerData mintid,
+				maxtid;
+	TableScanDesc scan;
+	TupleTableSlot *slot;
+
+	/*
+	 * If an auxiliary table exists, it's the one we're copying the data into.
+	 */
+	if (chgcxt->cc_dest_aux)
+		dest = chgcxt->cc_dest_aux;
+	else
+		dest = &chgcxt->cc_dest;
+
+	first_block = range_start;
+
+	if (range_end > range_start)
+		last_block = range_end - 1;
+	else
+	{
+		Assert(range_start == range_end);
+
+		last_block = range_end;
+	}
+
+	ItemPointerSet(&mintid, first_block, FirstOffsetNumber);
+	ItemPointerSet(&maxtid, last_block, MaxOffsetNumber);
+
+	/* XXX flags? */
+	scan = table_beginscan_tidrange(dest->rel, SnapshotAny, &mintid, &maxtid,
+									0);
+	slot = table_slot_create(dest->rel, NULL);
+	while (table_scan_getnextslot(scan, ForwardScanDirection, slot))
+		ExecInsertIndexTuples(dest->rri, dest->estate, 0, slot, NIL, NULL);
+	ExecDropSingleTupleTableSlot(slot);
+	table_endscan(scan);
+
+	/*
+	 * Index insertion could have accessed catalog when checking constraints,
+	 * so make sure that we no longer block the progress of xmin horizons for
+	 * VACUUM. (It's ok to have a snapshot throughout range scan, so there's
+	 * no point in doing this invalidation more often.)
+	 */
+	InvalidateCatalogSnapshot();
 }
 
 /*
@@ -1055,20 +1168,47 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
  * first block beyond the new range.
  *
  * Return the snapshot for the scan of the new range.
+ *
+ * TODO Consider a structure to accommodate (most of) the arguments.
  */
 static Snapshot
-finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
-					 BlockNumber start, BlockNumber *end_p)
+finalize_block_range(Relation rel_old, Relation rel_dst,
+					 TransactionId oldest_xmin,
+					 TransactionId xid_cutoff,
+					 MultiXactId multi_cutoff,
+					 ChangeContext *chgcxt, BlockNumber cur,
+					 BlockNumber start, BlockNumber *end_p,
+					 RewriteState *rwstate_p,
+					 BlockNumber *range_start_dst_p)
 {
 	BlockNumber end = *end_p;
+	RewriteState rwstate = *rwstate_p;
 	XLogRecPtr	end_of_wal;
 	Snapshot	snapshot;
+	BlockNumber range_end_dst;
 
 	/*
 	 * Wait here when testing how snapshot is changed at page boundary.
 	 */
 	INJECTION_POINT("repack-concurrently-new-range", NULL);
 
+	/*
+	 * Make sure the data is flushed to file before changes can be applied to
+	 * it. (Nothing of it should be in shared buffers so far.)
+	 */
+	end_heap_rewrite(rwstate);
+
+	/*
+	 * Now that the data can be accessed via shared buffers, scan it and add
+	 * each tuple to the identity index - this is necessary to update the
+	 * concurrent data changes below. We could not do that earlier because
+	 * even insertion into index might need to fetch heap tuples, in order to
+	 * check unique or exclusion constraints.
+	 */
+	range_end_dst = RelationGetNumberOfBlocks(rel_dst);
+	update_identity_index(chgcxt, *range_start_dst_p, range_end_dst);
+	*range_start_dst_p = range_end_dst;
+
 	/*
 	 * Decode all the concurrent data changes committed so far before
 	 * requesting the next snapshot - these changes are applicable on top of
@@ -1101,6 +1241,7 @@ finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
 
 	/* See above. */
 	Assert(!TransactionIdIsValid(MyProc->xmin));
+	Assert(!TransactionIdIsValid(MyProc->xid));
 
 	/*
 	 * XXX It might be worth Assert(CatalogSnapshot == NULL) here, however
@@ -1123,8 +1264,20 @@ finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
 	 * next batch of decoded changes.
 	 */
 	snapshot = repack_get_snapshot(chgcxt);
+
+	Assert(TransactionIdFollowsOrEquals(snapshot->xmin,
+										chgcxt->cc_last_snapshot_xmin));
+	chgcxt->cc_last_snapshot_xmin = snapshot->xmin;
+
 	PushActiveSnapshot(snapshot);
 
+	/*
+	 * Prepare for bulk insert of the next set of tuples. We rely on it to
+	 * start on a new page, even if the last existing page is not full.
+	 */
+	*rwstate_p = begin_heap_rewrite(rel_old, rel_dst, oldest_xmin, xid_cutoff,
+									multi_cutoff, true);
+
 	return snapshot;
 }
 
@@ -2554,6 +2707,64 @@ reform_and_rewrite_tuple(TupleTableSlot *src, TupleTableSlot *reform,
 		heap_freetuple(newtuple);
 }
 
+/*
+ * Insert tuple when processing REPACK CONCURRENTLY.
+ *
+ * 'reform' is a slot to use for tuple "reforming", typically to get set
+ * values of dropped columns to NULL.
+ *
+ * We pass the NO_LOGICAL flag to heap_insert() in order to skip logical
+ * decoding: as soon as REPACK CONCURRENTLY swaps the relation files, it drops
+ * this relation, so no logical replication subscription should need the data.
+ */
+static void
+heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
+					   TupleTableSlot *reform, RewriteState rwstate)
+{
+	HeapTuple	tuple,
+				new_tuple;
+	TransactionId xid;
+	bool		shouldFree,
+				shouldFreeNew;
+	TupleTableSlot *slot;
+	bool		freeze;
+
+	if (chgcxt->cc_dest_aux)
+	{
+		/* Will freeze when copying data to the new table. */
+		freeze = false;
+	}
+	else
+		freeze = true;
+
+	Assert(TTS_IS_BUFFERTUPLE(src));
+	tuple = ExecFetchSlotHeapTuple(src, false, &shouldFree);
+	xid = HeapTupleHeaderGetXmin(tuple->t_data);
+	if (reform != NULL && tuple_needs_reform(tuple, src->tts_tupleDescriptor))
+	{
+		clear_dropped_attributes(tuple, reform);
+		slot = reform;
+	}
+	else
+		slot = src;
+
+	/* Make sure we have a copy of the tuple, and set its XID. */
+	new_tuple = ExecFetchSlotHeapTuple(slot, true, &shouldFreeNew);
+	HeapTupleHeaderSetXmin(new_tuple->t_data, xid);
+
+	/* Perform the insertion. */
+	rewrite_heap_tuple_no_chains(rwstate, tuple, new_tuple, freeze);
+
+	/* The insertion shouldn't have caused XID assignment. */
+	Assert(!TransactionIdIsValid(GetCurrentTransactionIdIfAny()));
+
+	/* Cleanup. */
+	if (shouldFree)
+		heap_freetuple(tuple);
+	if (shouldFreeNew)
+		heap_freetuple(new_tuple);
+}
+
 /*
  * Check visibility of the tuple.
  */
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 361b76e5065..f37e4848667 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -1363,6 +1363,65 @@ HeapTupleSatisfiesNonVacuumable(HeapTuple htup, Snapshot snapshot,
 	return res != HEAPTUPLE_DEAD;
 }
 
+/*
+ * HeapTupleSatisfiesNewHeap
+ *		Consider all transactions committed or current.
+ *
+ * See SNAPSHOT_NEW_HEAP's definition for the intended behaviour.
+ *
+ */
+static bool
+HeapTupleSatisfiesNewHeap(HeapTuple htup, Snapshot snapshot, Buffer buffer)
+{
+	HeapTupleHeader tuple = htup->t_data;
+
+	Assert(ItemPointerIsValid(&htup->t_self));
+	Assert(htup->t_tableOid != InvalidOid);
+
+	/* xmin should always be there. */
+	Assert(TransactionIdIsValid(HeapTupleHeaderGetXmin(tuple)));
+
+	/*
+	 * No one should have the chance to set XMIN_INVALID until REPACK has
+	 * finished, and we do not need it.
+	 */
+	Assert(!HeapTupleHeaderXminInvalid(tuple));
+
+	/*
+	 * Unlike that, XMIN_COMMITTED might have been set earlier by REPACK
+	 * itself, but we don't need it here.
+	 */
+
+	/*
+	 * Set XMIN_COMMITTED to make the next checks (by any snapshot) faster.
+	 *
+	 * TODO Set the flag in the initial load and in
+	 * apply_concurrent_changes().
+	 */
+	SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
+				HeapTupleHeaderGetRawXmax(tuple));
+
+	/* Inserted tuples have XMAX_INVALID set. */
+	if (tuple->t_infomask & HEAP_XMAX_INVALID)
+		return true;
+
+	/*
+	 * REPACK could have set XMAX_COMMITTED during UPDATE or DELETE, or below.
+	 */
+	if (tuple->t_infomask & HEAP_XMAX_COMMITTED)
+		return false;
+
+	if (!TransactionIdIsValid(HeapTupleHeaderGetRawXmax(tuple)))
+		return true;
+
+	/*
+	 * Set XMAX_COMMITTED to make the next checks faster.
+	 */
+	SetHintBits(tuple, buffer, HEAP_XMAX_COMMITTED,
+				HeapTupleHeaderGetRawXmax(tuple));
+
+	return false;
+}
 
 /*
  * HeapTupleIsSurelyDead
@@ -1747,6 +1806,9 @@ HeapTupleSatisfiesVisibility(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 			return HeapTupleSatisfiesHistoricMVCC(htup, snapshot, buffer);
 		case SNAPSHOT_NON_VACUUMABLE:
 			return HeapTupleSatisfiesNonVacuumable(htup, snapshot, buffer);
+		case SNAPSHOT_NEW_HEAP:
+			return HeapTupleSatisfiesNewHeap(htup, snapshot, buffer);
+
 	}
 
 	return false;				/* keep compiler quiet */
diff --git a/src/backend/access/heap/heaptoast.c b/src/backend/access/heap/heaptoast.c
index 03f885a25b0..4a6a674b80e 100644
--- a/src/backend/access/heap/heaptoast.c
+++ b/src/backend/access/heap/heaptoast.c
@@ -40,7 +40,8 @@
  * ----------
  */
 void
-heap_toast_delete(Relation rel, HeapTuple oldtup, bool is_speculative)
+heap_toast_delete(Relation rel, HeapTuple oldtup, bool is_speculative,
+				  TransactionId xid)
 {
 	TupleDesc	tupleDesc;
 	Datum		toast_values[MaxHeapAttributeNumber];
@@ -70,7 +71,8 @@ heap_toast_delete(Relation rel, HeapTuple oldtup, bool is_speculative)
 	heap_deform_tuple(oldtup, tupleDesc, toast_values, toast_isnull);
 
 	/* Do the real work. */
-	toast_delete_external(rel, toast_values, toast_isnull, is_speculative);
+	toast_delete_external(rel, toast_values, toast_isnull, is_speculative,
+						  xid);
 }
 
 
@@ -84,6 +86,7 @@ heap_toast_delete(Relation rel, HeapTuple oldtup, bool is_speculative)
  *	newtup: the candidate new tuple to be inserted
  *	oldtup: the old row version for UPDATE, or NULL for INSERT
  *	options: options to be passed to heap_insert() for toast rows
+ *	rwstate: if valid, use raw_heap_insert()
  * Result:
  *	either newtup if no toasting is needed, or a palloc'd modified tuple
  *	that is what should actually get stored
@@ -94,7 +97,7 @@ heap_toast_delete(Relation rel, HeapTuple oldtup, bool is_speculative)
  */
 HeapTuple
 heap_toast_insert_or_update(Relation rel, HeapTuple newtup, HeapTuple oldtup,
-							uint32 options)
+							RewriteState rwstate, uint32 options)
 {
 	HeapTuple	result_tuple;
 	TupleDesc	tupleDesc;
@@ -109,6 +112,7 @@ heap_toast_insert_or_update(Relation rel, HeapTuple newtup, HeapTuple oldtup,
 	Datum		toast_oldvalues[MaxHeapAttributeNumber];
 	ToastAttrInfo toast_attr[MaxHeapAttributeNumber];
 	ToastTupleContext ttc;
+	TransactionId xid;
 
 	/*
 	 * Ignore the INSERT_SPECULATIVE option. Speculative insertions/super
@@ -156,6 +160,13 @@ heap_toast_insert_or_update(Relation rel, HeapTuple newtup, HeapTuple oldtup,
 	ttc.ttc_attr = toast_attr;
 	toast_tuple_init(&ttc);
 
+	/*
+	 * raw_heap_insert() may be needed for insertion into the TOAST table. In
+	 * that case, visibility information will be retrieved from 'newtup'.
+	 */
+	ttc.ttc_rwstate = rwstate;
+	ttc.ttc_tup_main = newtup;
+
 	/* ----------
 	 * Compress and/or save external until data fits into target length
 	 *
@@ -330,7 +341,11 @@ heap_toast_insert_or_update(Relation rel, HeapTuple newtup, HeapTuple oldtup,
 	else
 		result_tuple = newtup;
 
-	toast_tuple_cleanup(&ttc);
+	if (options & TABLE_REUSE_XID)
+		xid = HeapTupleHeaderGetXmin(newtup->t_data);
+	else
+		xid = InvalidTransactionId; /* The current transaction. */
+	toast_tuple_cleanup(&ttc, xid);
 
 	return result_tuple;
 }
diff --git a/src/backend/access/heap/rewriteheap.c b/src/backend/access/heap/rewriteheap.c
index 0ccd392c3dc..809dfba1b55 100644
--- a/src/backend/access/heap/rewriteheap.c
+++ b/src/backend/access/heap/rewriteheap.c
@@ -107,6 +107,7 @@
 #include "access/heapam.h"
 #include "access/heapam_xlog.h"
 #include "access/heaptoast.h"
+#include "access/multixact.h"
 #include "access/rewriteheap.h"
 #include "access/transam.h"
 #include "access/xact.h"
@@ -119,6 +120,7 @@
 #include "storage/bufmgr.h"
 #include "storage/bulk_write.h"
 #include "storage/fd.h"
+#include "storage/lmgr.h"
 #include "storage/procarray.h"
 #include "utils/memutils.h"
 #include "utils/rel.h"
@@ -151,6 +153,12 @@ typedef struct RewriteStateData
 	HTAB	   *rs_old_new_tid_map; /* unmatched B tuples */
 	HTAB	   *rs_logical_mappings;	/* logical remapping files */
 	uint32		rs_num_rewrite_mappings;	/* # in memory mappings */
+
+	/*
+	 * If this is initialized, raw_heap_insert() is also used for TOAST
+	 * relation.
+	 */
+	struct RewriteStateData *toast;
 } RewriteStateData;
 
 /*
@@ -211,6 +219,11 @@ typedef struct RewriteMappingDataEntry
 
 
 /* prototypes for internal functions */
+static RewriteState begin_heap_rewrite_common(Relation old_heap,
+											  Relation new_heap,
+											  TransactionId oldest_xmin,
+											  TransactionId freeze_xid,
+											  MultiXactId cutoff_multi);
 static void raw_heap_insert(RewriteState state, HeapTuple tup);
 
 /* internal logical remapping prototypes */
@@ -227,18 +240,19 @@ static void logical_end_heap_rewrite(RewriteState state);
  * oldest_xmin	xid used by the caller to determine which tuples are dead
  * freeze_xid	xid before which tuples will be frozen
  * cutoff_multi	multixact before which multis will be removed
+ * no_chains	only raw insert (and freezing), do not care of HOT chains
  *
  * Returns an opaque RewriteState, allocated in current memory context,
  * to be used in subsequent calls to the other functions.
  */
 RewriteState
 begin_heap_rewrite(Relation old_heap, Relation new_heap, TransactionId oldest_xmin,
-				   TransactionId freeze_xid, MultiXactId cutoff_multi)
+				   TransactionId freeze_xid, MultiXactId cutoff_multi,
+				   bool no_chains)
 {
-	RewriteState state;
 	MemoryContext rw_cxt;
 	MemoryContext old_cxt;
-	HASHCTL		hash_ctl;
+	RewriteState state;
 
 	/*
 	 * To ease cleanup, make a separate context that will contain the
@@ -249,9 +263,99 @@ begin_heap_rewrite(Relation old_heap, Relation new_heap, TransactionId oldest_xm
 								   ALLOCSET_DEFAULT_SIZES);
 	old_cxt = MemoryContextSwitchTo(rw_cxt);
 
+	state = begin_heap_rewrite_common(old_heap, new_heap, oldest_xmin,
+									  freeze_xid, cutoff_multi);
+	state->rs_cxt = rw_cxt;
+
+	if (!no_chains)
+	{
+		HASHCTL		hash_ctl;
+
+		/* Initialize hash tables used to track update chains */
+		hash_ctl.keysize = sizeof(TidHashKey);
+		hash_ctl.entrysize = sizeof(UnresolvedTupData);
+		hash_ctl.hcxt = state->rs_cxt;
+
+		state->rs_unresolved_tups =
+			hash_create("Rewrite / Unresolved ctids",
+						128,	/* arbitrary initial size */
+						&hash_ctl,
+						HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
+
+		hash_ctl.entrysize = sizeof(OldToNewMappingData);
+
+		state->rs_old_new_tid_map =
+			hash_create("Rewrite / Old to new tid map",
+						128,	/* arbitrary initial size */
+						&hash_ctl,
+						HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
+
+		logical_begin_heap_rewrite(state);
+	}
+	else
+	{
+		Oid			toastid;
+
+		/*
+		 * The current user of this mode, REPACK (CONCURRENTLY), does not want
+		 * XID assigned at all, so "raw insert" is also used for TOAST.
+		 *
+		 * XXX "raw insert" would be helpful for REPACK w/o CONCURRENTLY too,
+		 * as it writes the whole pages to WAL. The problem here is that it
+		 * might check the existence of chunk OIDs in the new TOAST relation
+		 * (see the hacks with rd_toastoid in toast_save_datum()), and that
+		 * doesn't work while the rewrite is still in progress: the relation
+		 * pages are not guaranteed to be flushed to disk (and read into
+		 * shared buffers) before the end of the rewrite.
+		 */
+		toastid = new_heap->rd_rel->reltoastrelid;
+		if (OidIsValid(toastid))
+		{
+			Oid			toastid_old;
+			Relation	toast_rel;
+			Relation	toast_rel_old = NULL;
+
+			/* New relation's TOAST should already be locked. */
+			Assert(CheckRelationOidLockedByMe(toastid, AccessExclusiveLock,
+											  false));
+			toast_rel = table_open(toastid, NoLock);
+
+			toastid_old = old_heap ? old_heap->rd_rel->reltoastrelid :
+				InvalidTransactionId;
+			if (OidIsValid(toastid_old))
+			{
+				/*
+				 * Currently we do not lock the old relation's TOAST. Use the
+				 * same lock mode we use for the parent relation.
+				 */
+				toast_rel_old = table_open(toastid_old,
+										   ShareUpdateExclusiveLock);
+			}
+
+			/* Create the state for TOAST insertions. */
+			state->toast = begin_heap_rewrite_common(toast_rel_old,
+													 toast_rel,
+													 oldest_xmin,
+													 freeze_xid,
+													 cutoff_multi);
+		}
+	}
+
+	MemoryContextSwitchTo(old_cxt);
+
+	return state;
+}
+
+static RewriteState
+begin_heap_rewrite_common(Relation old_heap, Relation new_heap,
+						  TransactionId oldest_xmin,
+						  TransactionId freeze_xid, MultiXactId cutoff_multi)
+
+{
+	RewriteState state;
+
 	/* Create and fill in the state struct */
 	state = palloc0_object(RewriteStateData);
-
 	state->rs_old_rel = old_heap;
 	state->rs_new_rel = new_heap;
 	state->rs_buffer = NULL;
@@ -260,32 +364,8 @@ begin_heap_rewrite(Relation old_heap, Relation new_heap, TransactionId oldest_xm
 	state->rs_oldest_xmin = oldest_xmin;
 	state->rs_freeze_xid = freeze_xid;
 	state->rs_cutoff_multi = cutoff_multi;
-	state->rs_cxt = rw_cxt;
 	state->rs_bulkstate = smgr_bulk_start_rel(new_heap, MAIN_FORKNUM);
 
-	/* Initialize hash tables used to track update chains */
-	hash_ctl.keysize = sizeof(TidHashKey);
-	hash_ctl.entrysize = sizeof(UnresolvedTupData);
-	hash_ctl.hcxt = state->rs_cxt;
-
-	state->rs_unresolved_tups =
-		hash_create("Rewrite / Unresolved ctids",
-					128,		/* arbitrary initial size */
-					&hash_ctl,
-					HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
-
-	hash_ctl.entrysize = sizeof(OldToNewMappingData);
-
-	state->rs_old_new_tid_map =
-		hash_create("Rewrite / Old to new tid map",
-					128,		/* arbitrary initial size */
-					&hash_ctl,
-					HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
-
-	MemoryContextSwitchTo(old_cxt);
-
-	logical_begin_heap_rewrite(state);
-
 	return state;
 }
 
@@ -300,16 +380,20 @@ end_heap_rewrite(RewriteState state)
 	HASH_SEQ_STATUS seq_status;
 	UnresolvedTup unresolved;
 
-	/*
-	 * Write any remaining tuples in the UnresolvedTups table. If we have any
-	 * left, they should in fact be dead, but let's err on the safe side.
-	 */
-	hash_seq_init(&seq_status, state->rs_unresolved_tups);
-
-	while ((unresolved = hash_seq_search(&seq_status)) != NULL)
+	if (state->rs_unresolved_tups)
 	{
-		ItemPointerSetInvalid(&unresolved->tuple->t_data->t_ctid);
-		raw_heap_insert(state, unresolved->tuple);
+		/*
+		 * Write any remaining tuples in the UnresolvedTups table. If we have
+		 * any left, they should in fact be dead, but let's err on the safe
+		 * side.
+		 */
+		hash_seq_init(&seq_status, state->rs_unresolved_tups);
+
+		while ((unresolved = hash_seq_search(&seq_status)) != NULL)
+		{
+			ItemPointerSetInvalid(&unresolved->tuple->t_data->t_ctid);
+			raw_heap_insert(state, unresolved->tuple);
+		}
 	}
 
 	/* Write the last page, if any */
@@ -318,10 +402,29 @@ end_heap_rewrite(RewriteState state)
 		smgr_bulk_write(state->rs_bulkstate, state->rs_blockno, state->rs_buffer, true);
 		state->rs_buffer = NULL;
 	}
-
 	smgr_bulk_finish(state->rs_bulkstate);
 
-	logical_end_heap_rewrite(state);
+	/* The same for TOAST */
+	if (state->toast)
+	{
+		RewriteState toast = state->toast;
+
+		if (toast->rs_buffer)
+		{
+			smgr_bulk_write(toast->rs_bulkstate, toast->rs_blockno,
+							toast->rs_buffer, true);
+			toast->rs_buffer = NULL;
+		}
+		smgr_bulk_finish(toast->rs_bulkstate);
+
+		/* Close relation(s) opened by begin_heap_rewrite(). */
+		table_close(toast->rs_new_rel, NoLock);
+		if (toast->rs_old_rel)
+			table_close(toast->rs_old_rel, ShareUpdateExclusiveLock);
+	}
+
+	if (state->rs_logical_rewrite)
+		logical_end_heap_rewrite(state);
 
 	/* Deleting the context frees everything */
 	MemoryContextDelete(state->rs_cxt);
@@ -350,30 +453,13 @@ rewrite_heap_tuple(RewriteState state,
 
 	old_cxt = MemoryContextSwitchTo(state->rs_cxt);
 
-	/*
-	 * Copy the original tuple's visibility information into new_tuple.
-	 *
-	 * XXX we might later need to copy some t_infomask2 bits, too? Right now,
-	 * we intentionally clear the HOT status bits.
-	 */
-	memcpy(&new_tuple->t_data->t_choice.t_heap,
-		   &old_tuple->t_data->t_choice.t_heap,
-		   sizeof(HeapTupleFields));
-
-	new_tuple->t_data->t_infomask &= ~HEAP_XACT_MASK;
-	new_tuple->t_data->t_infomask2 &= ~HEAP2_XACT_MASK;
-	new_tuple->t_data->t_infomask |=
-		old_tuple->t_data->t_infomask & HEAP_XACT_MASK;
+	rewrite_copy_visibility_info(new_tuple, old_tuple);
 
 	/*
 	 * While we have our hands on the tuple, we may as well freeze any
 	 * eligible xmin or xmax, so that future VACUUM effort can be saved.
 	 */
-	heap_freeze_tuple(new_tuple->t_data,
-					  state->rs_old_rel->rd_rel->relfrozenxid,
-					  state->rs_old_rel->rd_rel->relminmxid,
-					  state->rs_freeze_xid,
-					  state->rs_cutoff_multi);
+	rewrite_freeze_tuple(state, new_tuple);
 
 	/*
 	 * Invalid ctid means that ctid should point to the tuple itself. We'll
@@ -534,6 +620,24 @@ rewrite_heap_tuple(RewriteState state,
 	MemoryContextSwitchTo(old_cxt);
 }
 
+/*
+ * Like rewrite_heap_tuple(), but do not care about hot chains. The user
+ * should have used the appropriate snapshot to pick at most one tuple of the
+ * chain - this is typical for REPACK (CONCURRENTLY).
+ */
+void
+rewrite_heap_tuple_no_chains(RewriteState state, HeapTuple old_tuple,
+							 HeapTuple new_tuple, bool freeze)
+{
+	if (new_tuple != old_tuple)
+		rewrite_copy_visibility_info(new_tuple, old_tuple);
+
+	if (freeze)
+		rewrite_freeze_tuple(state, new_tuple);
+
+	raw_heap_insert(state, new_tuple);
+}
+
 /*
  * Register a dead tuple with an ongoing rewrite. Dead tuples are not
  * copied to the new table, but we still make note of them so that we
@@ -628,7 +732,7 @@ raw_heap_insert(RewriteState state, HeapTuple tup)
 		options |= HEAP_INSERT_NO_LOGICAL;
 
 		heaptup = heap_toast_insert_or_update(state->rs_new_rel, tup, NULL,
-											  options);
+											  state->toast, options);
 	}
 	else
 		heaptup = tup;
@@ -704,6 +808,47 @@ raw_heap_insert(RewriteState state, HeapTuple tup)
 		heap_freetuple(heaptup);
 }
 
+/*
+ * Freeze tuple. 'old_tuple' provides the initial visibility information.
+ */
+void
+rewrite_freeze_tuple(RewriteState state, HeapTuple tuple)
+{
+	TransactionId relfrozenxid = InvalidTransactionId;
+	MultiXactId relminmxid = InvalidMultiXactId;
+
+	/* The old relation may be missing if dealing with TOAST. */
+	if (state->rs_old_rel)
+	{
+		relfrozenxid = state->rs_old_rel->rd_rel->relfrozenxid;
+		relminmxid = state->rs_old_rel->rd_rel->relminmxid;
+	}
+
+	heap_freeze_tuple(tuple->t_data, relfrozenxid, relminmxid,
+					  state->rs_freeze_xid,
+					  state->rs_cutoff_multi);
+}
+
+/*
+ * Copy the old tuple's visibility information into the new tuple.
+ */
+void
+rewrite_copy_visibility_info(HeapTuple new_tuple, HeapTuple old_tuple)
+{
+	/*
+	 * XXX we might later need to copy some t_infomask2 bits, too? Right now,
+	 * we intentionally clear the HOT status bits.
+	 */
+	memcpy(&new_tuple->t_data->t_choice.t_heap,
+		   &old_tuple->t_data->t_choice.t_heap,
+		   sizeof(HeapTupleFields));
+
+	new_tuple->t_data->t_infomask &= ~HEAP_XACT_MASK;
+	new_tuple->t_data->t_infomask2 &= ~HEAP2_XACT_MASK;
+	new_tuple->t_data->t_infomask |=
+		old_tuple->t_data->t_infomask & HEAP_XACT_MASK;
+}
+
 /* ------------------------------------------------------------------------
  * Logical rewrite support
  *
diff --git a/src/backend/access/table/tableam.c b/src/backend/access/table/tableam.c
index 68ff0966f1c..43b853d949e 100644
--- a/src/backend/access/table/tableam.c
+++ b/src/backend/access/table/tableam.c
@@ -319,6 +319,7 @@ simple_table_tuple_delete(Relation rel, ItemPointer tid, Snapshot snapshot)
 	TM_FailureData tmfd;
 
 	result = table_tuple_delete(rel, tid,
+								InvalidTransactionId,
 								GetCurrentCommandId(true),
 								0, snapshot, InvalidSnapshot,
 								true /* wait for commit */ ,
diff --git a/src/backend/access/table/toast_helper.c b/src/backend/access/table/toast_helper.c
index 2f2022d9951..e0c3af410d4 100644
--- a/src/backend/access/table/toast_helper.c
+++ b/src/backend/access/table/toast_helper.c
@@ -261,7 +261,7 @@ toast_tuple_externalize(ToastTupleContext *ttc, int attribute, uint32 options)
 
 	attr->tai_colflags |= TOASTCOL_IGNORE;
 	*value = toast_save_datum(ttc->ttc_rel, old_value, attr->tai_oldexternal,
-							  options);
+							  ttc->ttc_rwstate, ttc->ttc_tup_main, options);
 	if ((attr->tai_colflags & TOASTCOL_NEEDS_FREE) != 0)
 		pfree(DatumGetPointer(old_value));
 	attr->tai_colflags |= TOASTCOL_NEEDS_FREE;
@@ -272,7 +272,7 @@ toast_tuple_externalize(ToastTupleContext *ttc, int attribute, uint32 options)
  * Perform appropriate cleanup after one tuple has been subjected to TOAST.
  */
 void
-toast_tuple_cleanup(ToastTupleContext *ttc)
+toast_tuple_cleanup(ToastTupleContext *ttc, TransactionId xid)
 {
 	TupleDesc	tupleDesc = ttc->ttc_rel->rd_att;
 	int			numAttrs = tupleDesc->natts;
@@ -305,7 +305,8 @@ toast_tuple_cleanup(ToastTupleContext *ttc)
 			ToastAttrInfo *attr = &ttc->ttc_attr[i];
 
 			if ((attr->tai_colflags & TOASTCOL_NEEDS_DELETE_OLD) != 0)
-				toast_delete_datum(ttc->ttc_rel, ttc->ttc_oldvalues[i], false);
+				toast_delete_datum(ttc->ttc_rel, ttc->ttc_oldvalues[i], false,
+								   xid);
 		}
 	}
 }
@@ -316,7 +317,7 @@ toast_tuple_cleanup(ToastTupleContext *ttc)
  */
 void
 toast_delete_external(Relation rel, const Datum *values, const bool *isnull,
-					  bool is_speculative)
+					  bool is_speculative, TransactionId xid)
 {
 	TupleDesc	tupleDesc = rel->rd_att;
 	int			numAttrs = tupleDesc->natts;
@@ -331,7 +332,7 @@ toast_delete_external(Relation rel, const Datum *values, const bool *isnull,
 			if (isnull[i])
 				continue;
 			else if (VARATT_IS_EXTERNAL_ONDISK(DatumGetPointer(value)))
-				toast_delete_datum(rel, value, is_speculative);
+				toast_delete_datum(rel, value, is_speculative, xid);
 		}
 	}
 }
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index f2e10b82b7d..ed20a35f7d6 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -105,6 +105,9 @@ static uint64 mainrdata_len;	/* total # of bytes in chain */
 /* flags for the in-progress insertion */
 static uint8 curinsert_flags = 0;
 
+/* XID to override the XID of the current transaction. */
+static TransactionId curinsert_xid = InvalidTransactionId;
+
 /*
  * These are used to hold the record header while constructing a record.
  * 'hdr_scratch' is not a plain variable, but is palloc'd at initialization,
@@ -235,6 +238,7 @@ XLogResetInsertion(void)
 	mainrdata_len = 0;
 	mainrdata_last = (XLogRecData *) &mainrdata_head;
 	curinsert_flags = 0;
+	curinsert_xid = InvalidTransactionId;
 	begininsert_called = false;
 }
 
@@ -467,6 +471,18 @@ XLogSetRecordFlags(uint8 flags)
 	curinsert_flags |= flags;
 }
 
+/*
+ * Set XID status flags for the upcoming WAL record.
+ *
+ * Useful when creating WAL records on behalf of another transaction.
+ */
+void
+XLogSetRecordXid(TransactionId xid)
+{
+	Assert(begininsert_called);
+	curinsert_xid = xid;
+}
+
 /*
  * Insert an XLOG record having the specified RMID and info bytes, with the
  * body of the record being the data and buffer references registered earlier
@@ -928,6 +944,12 @@ XLogRecordAssemble(RmgrId rmid, uint8 info,
 	{
 		TransactionId xid = GetTopTransactionIdIfAny();
 
+		/*
+		 * On curinsert_xid: if it's set, it's only for recovery and streaming
+		 * replication to work. On the other hand, the record shouldn't be
+		 * logically decoded, so we don't care if the toplevel XID is invalid.
+		 */
+
 		/* Set the flag that the top xid is included in the WAL */
 		*topxid_included = true;
 
@@ -1000,7 +1022,14 @@ XLogRecordAssemble(RmgrId rmid, uint8 info,
 	 * once we know where in the WAL the record will be inserted. The CRC does
 	 * not include the record header yet.
 	 */
-	rechdr->xl_xid = GetCurrentTransactionIdIfAny();
+	if (!TransactionIdIsValid(curinsert_xid))
+		rechdr->xl_xid = GetCurrentTransactionIdIfAny();
+	else
+	{
+		/* The overriding XID should be handled specially. */
+		rechdr->xl_xid = curinsert_xid;
+		info |= XLR_XID_REPLAYED;
+	}
 	rechdr->xl_tot_len = (uint32) total_len;
 	rechdr->xl_info = info;
 	rechdr->xl_rmid = rmid;
diff --git a/src/backend/access/transam/xlogrecovery.c b/src/backend/access/transam/xlogrecovery.c
index c0ae4d3f63f..cae7318506d 100644
--- a/src/backend/access/transam/xlogrecovery.c
+++ b/src/backend/access/transam/xlogrecovery.c
@@ -1949,10 +1949,13 @@ ApplyWalRecord(XLogReaderState *xlogreader, XLogRecord *record, TimeLineID *repl
 	SpinLockRelease(&XLogRecoveryCtl->info_lck);
 
 	/*
-	 * If we are attempting to enter Hot Standby mode, process XIDs we see
+	 * If we are attempting to enter Hot Standby mode, process XIDs we see.
+	 *
+	 * "replayed" changes should not get into the array again.
 	 */
 	if (standbyState >= STANDBY_INITIALIZED &&
-		TransactionIdIsValid(record->xl_xid))
+		TransactionIdIsValid(record->xl_xid) &&
+		(record->xl_info & XLR_XID_REPLAYED) == 0)
 		RecordKnownAssignedTransactionIds(record->xl_xid);
 
 	/*
diff --git a/src/backend/catalog/indexing.c b/src/backend/catalog/indexing.c
index fd7d2ec0e3a..e09e9fcd8ac 100644
--- a/src/backend/catalog/indexing.c
+++ b/src/backend/catalog/indexing.c
@@ -364,5 +364,5 @@ CatalogTupleUpdateWithInfo(Relation heapRel, const ItemPointerData *otid, HeapTu
 void
 CatalogTupleDelete(Relation heapRel, const ItemPointerData *tid)
 {
-	simple_heap_delete(heapRel, tid);
+	simple_heap_delete(heapRel, tid, InvalidTransactionId);
 }
diff --git a/src/backend/commands/matview.c b/src/backend/commands/matview.c
index 9d490da5f81..372cf57e9f4 100644
--- a/src/backend/commands/matview.c
+++ b/src/backend/commands/matview.c
@@ -894,7 +894,8 @@ refresh_by_heap_swap(Oid matviewOid, Oid OIDNewHeap, char relpersistence)
 {
 	finish_heap_swap(matviewOid, OIDNewHeap, false, false, true, true,
 					 true,		/* reindex */
-					 RecentXmin, ReadNextMultiXactId(), relpersistence);
+					 RecentXmin, ReadNextMultiXactId(), false,
+					 relpersistence);
 }
 
 /*
diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 7390fd303f6..32f15683b6e 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -73,6 +73,7 @@
 #include "storage/lmgr.h"
 #include "storage/predicate.h"
 #include "storage/proc.h"
+#include "storage/procarray.h"
 #include "utils/acl.h"
 #include "utils/fmgroids.h"
 #include "utils/guc.h"
@@ -127,6 +128,7 @@ typedef struct ChangeContexBackup
 	int			file_seq_snapshot;
 	int			file_seq_changes;
 	Oid			clustering_index;
+	TransactionId	last_snapshot_xmin;
 } ChangeContextBackup;
 
 /*
@@ -187,7 +189,11 @@ static void copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldInde
 							bool *pSwapToastByContent,
 							TransactionId *pFreezeXid,
 							MultiXactId *pCutoffMulti,
+							double *p_num_tuples,
 							ChangeContext *chgcxt);
+static void copy_table_data_update_stats(Relation OldHeap, Relation NewHeap,
+										 BlockNumber num_pages,
+										 double num_tuples);
 static void update_relation_cutoffs(Oid relid, TransactionId frozenXid,
 									MultiXactId cutoffMulti);
 static List *get_tables_to_repack(RepackCommand cmd, bool usingindex,
@@ -201,14 +207,21 @@ static bool repack_is_permitted_for_relation(RepackCommand cmd,
 static void apply_concurrent_changes(ChangeContext *chgcxt,
 									 BlockNumber range_start,
 									 BlockNumber range_end);
-static void apply_concurrent_insert(RepackDest *dest, TupleTableSlot *slot);
+static void apply_concurrent_insert(RepackDest *dest,
+									TupleTableSlot *spill_tuple,
+									TupleTableSlot *new_tuple,
+									TransactionId xid);
 static void apply_concurrent_update(RepackDest *dest,
 									TupleTableSlot *spilled_tuple,
-									TupleTableSlot *ondisk_tuple);
-static void apply_concurrent_delete(Relation rel, TupleTableSlot *slot);
+									TupleTableSlot *ondisk_tuple,
+									TupleTableSlot *new_tuple,
+									TransactionId xid);
+static void apply_concurrent_delete(Relation rel, TupleTableSlot *slot,
+									TransactionId xid);
 static void restore_tuple(BufFile *file, Relation relation,
 						  TupleTableSlot *slot, BlockNumber *block_nr_p,
-						  BlockNumber *old_block_nr_p);
+						  BlockNumber *old_block_nr_p,
+						  TransactionId *xid_p);
 static void adjust_toast_pointers(Relation relation, TupleTableSlot *dest,
 								  TupleTableSlot *src);
 static bool is_block_in_range(BlockNumber blknum, BlockNumber start,
@@ -238,11 +251,14 @@ static void rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHea
 											   Oid identIdx,
 											   TransactionId frozenXid,
 											   MultiXactId cutoffMulti,
+											   double num_tuples,
 											   ChangeContext *chgcxt);
 static ChangeContext *process_auxiliary_table(ChangeContext *chgcxt,
 											  Relation *pOldHeap,
 											  Relation *pNewHeap,
-											  Oid identIdx);
+											  Oid identIdx,
+											  TransactionId freeze_xid,
+											  MultiXactId cutoff_multi);
 static List *build_new_indexes(List *OldIndexes, Relation *p_old,
 							   Relation *p_new,
 							   ChangeContext **p_chgcxt);
@@ -1073,6 +1089,7 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	TransactionId frozenXid;
 	MultiXactId cutoffMulti;
 	bool		concurrent = OidIsValid(ident_idx);
+	double		num_tuples = 0;
 	IndexBuildSecurity ibsec;
 	ChangeContext *chgcxt = NULL;
 #if USE_ASSERT_CHECKING
@@ -1162,17 +1179,10 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	/* Copy the heap data into the new table in the desired order */
 	copy_table_data(NewHeap, OldHeap, index, verbose,
 					&swap_toast_by_content, &frozenXid, &cutoffMulti,
-					chgcxt);
+					&num_tuples, chgcxt);
 
-	/* The historic snapshot won't be needed anymore. */
 	if (concurrent)
 	{
-		/*
-		 * Make sure the active snapshot can see the data copied, so the rows
-		 * can be updated / deleted.
-		 */
-		UpdateActiveSnapshotCommandId();
-
 		Assert(!swap_toast_by_content);
 
 		/*
@@ -1183,7 +1193,8 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 			index_close(index, NoLock);
 
 		rebuild_relation_finish_concurrent(NewHeap, OldHeap, ident_idx,
-										   frozenXid, cutoffMulti, chgcxt);
+										   frozenXid, cutoffMulti, num_tuples,
+										   chgcxt);
 
 		pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
 									 PROGRESS_REPACK_PHASE_FINAL_CLEANUP);
@@ -1213,11 +1224,15 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 		/*
 		 * Swap the physical files of the target and transient tables, then
 		 * rebuild the target's indexes and throw away the transient table.
+		 *
+		 * If swap_toast_by_content is false, we don't need to update the
+		 * cutoffs because the TOAST relation is new.
 		 */
 		finish_heap_swap(tableOid, OIDNewHeap, is_system_catalog,
 						 swap_toast_by_content, false, true,
 						 true,	/* reindex */
 						 frozenXid, cutoffMulti,
+						 swap_toast_by_content, /* update_toast_cutoffs */
 						 relpersistence);
 	}
 
@@ -1682,67 +1697,6 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
 	return OIDNewHeap;
 }
 
-/*
- * Insert tuple when processing REPACK CONCURRENTLY.
- *
- * rewriteheap.c is not used in the CONCURRENTLY case because it'd be
- * difficult to do the same in the catch-up phase (as the logical decoding
- * does not provide us with sufficient visibility information). Thus we must
- * use heap_insert() both during the catch-up and here.
- *
- * 'reform' is a slot to use for tuple "reforming", typically to get set
- * values of dropped columns to NULL.
- *
- * We pass the NO_LOGICAL flag to heap_insert() in order to skip logical
- * decoding: as soon as REPACK CONCURRENTLY swaps the relation files, it drops
- * this relation, so no logical replication subscription should need the data.
- */
-void
-heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
-					   TupleTableSlot *reform)
-{
-	HeapTuple	tuple;
-	bool		shouldFree;
-	TupleTableSlot *slot;
-	RepackDest *dest;
-
-	/*
-	 * Use the current auxiliary table as output if one is active, otherwise
-	 * insert the tuple into the actual destination table.
-	 */
-	if (chgcxt->cc_dest_aux)
-		dest = chgcxt->cc_dest_aux;
-	else
-		dest = &chgcxt->cc_dest;
-
-	tuple = ExecFetchSlotHeapTuple(src, false, &shouldFree);
-	if (reform != NULL && tuple_needs_reform(tuple, src->tts_tupleDescriptor))
-	{
-		clear_dropped_attributes(tuple, reform);
-		slot = reform;
-	}
-	else
-		slot = src;
-
-	/*
-	 * clear_dropped_attributes() should have deformed the tuple, so nothing
-	 * should depend on it now.
-	 */
-	if (shouldFree)
-		heap_freetuple(tuple);
-
-	table_tuple_insert(dest->rel, slot, GetCurrentCommandId(true),
-					   TABLE_INSERT_NO_LOGICAL, dest->bistate);
-
-	/*
-	 * Insert the tuple into the identity index. initialize_change_context()
-	 * may skip opening of indexes if the identity index is not needed
-	 * immediately.
-	 */
-	if (dest->rri)
-		ExecInsertIndexTuples(dest->rri, dest->estate, 0, slot, NIL, NULL);
-}
-
 bool
 tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc)
 {
@@ -1776,14 +1730,22 @@ clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
 {
 	TupleDesc	tupDesc = reform->tts_tupleDescriptor;
 
-	/* Assuming 'reform' is virtual, this deforms the tuple. */
-	Assert(TTS_IS_VIRTUAL(reform));
+	Assert(TTS_IS_VIRTUAL(reform) || TTS_IS_HEAPTUPLE(reform));
 	ExecForceStoreHeapTuple(tuple, reform, false);
 
 	for (int i = 0; i < tupDesc->natts; i++)
 	{
 		if (TupleDescCompactAttr(tupDesc, i)->attisdropped)
+		{
+			/*
+			 * If 'reform' is virtual, all the attributes are already
+			 * deformed. XXX Should we use the virtual slot at all? .
+			 */
+			if (TTS_IS_HEAPTUPLE(reform))
+				slot_getsomeattrs(reform, i + 1);
+
 			reform->tts_isnull[i] = true;
+		}
 	}
 }
 
@@ -1794,16 +1756,14 @@ clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
  * *pSwapToastByContent is set true if toast tables must be swapped by content.
  * *pFreezeXid receives the TransactionId used as freeze cutoff point.
  * *pCutoffMulti receives the MultiXactId used as a cutoff point.
+ * *p_num_tuples receives the number of tuples copied.
  */
 static void
 copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 				bool verbose, bool *pSwapToastByContent,
 				TransactionId *pFreezeXid, MultiXactId *pCutoffMulti,
-				ChangeContext *chgcxt)
+				double *p_num_tuples, ChangeContext *chgcxt)
 {
-	Relation	relRelation;
-	HeapTuple	reltup;
-	Form_pg_class relform;
 	TupleDesc	oldTupDesc PG_USED_FOR_ASSERTS_ONLY;
 	TupleDesc	newTupDesc PG_USED_FOR_ASSERTS_ONLY;
 	VacuumParams params;
@@ -1812,7 +1772,6 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 	double		num_tuples = 0,
 				tups_vacuumed = 0,
 				tups_recently_dead = 0;
-	BlockNumber num_pages;
 	int			elevel = verbose ? INFO : DEBUG2;
 	PGRUsage	ru0;
 	char	   *nspname;
@@ -1985,8 +1944,6 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 	 */
 	NewHeap->rd_toastoid = InvalidOid;
 
-	num_pages = RelationGetNumberOfBlocks(NewHeap);
-
 	/* Log what we did */
 	ereport(elevel,
 			(errmsg("\"%s.%s\": found %.0f removable, %.0f nonremovable row versions in %u pages",
@@ -1999,6 +1956,35 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 					   tups_recently_dead,
 					   pg_rusage_show(&ru0))));
 
+	/*
+	 * Update pg_class fields. In the CONCURRENTLY case we do it later because
+	 * 1) the catalog update triggers XID assignment, 2) the work is split
+	 * into several transactions, so the catalog update should take place in
+	 * the last one.
+	 */
+	if (!concurrent)
+	{
+		BlockNumber num_pages;
+
+		num_pages = RelationGetNumberOfBlocks(NewHeap);
+
+		copy_table_data_update_stats(OldHeap, NewHeap, num_pages, num_tuples);
+	}
+
+	*p_num_tuples = num_tuples;
+}
+
+/*
+ * Sub-routine of copy_table_data(), to update pg_class.
+ */
+static void
+copy_table_data_update_stats(Relation OldHeap, Relation NewHeap,
+							 BlockNumber num_pages, double num_tuples)
+{
+	Relation	relRelation;
+	HeapTuple	reltup;
+	Form_pg_class relform;
+
 	/* Update pg_class to reflect the correct values of pages and tuples. */
 	relRelation = table_open(RelationRelationId, RowExclusiveLock);
 
@@ -2446,6 +2432,9 @@ update_relation_cutoffs(Oid relid, TransactionId frozenXid,
 /*
  * Remove the transient table that was built by make_new_heap, and finish
  * cleaning up (including rebuilding all indexes on the old heap).
+ *
+ * 'update_toast_cutoffs' tells whether relfrozenxid and relminmxid of the
+ * TOAST relation should be updated too.
  */
 void
 finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
@@ -2456,6 +2445,7 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 				 bool reindex,
 				 TransactionId frozenXid,
 				 MultiXactId cutoffMulti,
+				 bool update_toast_cutoffs,
 				 char newrelpersistence)
 {
 	ObjectAddress object;
@@ -2463,6 +2453,17 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 	Oid			oid_old_toastid;
 	int			i;
 
+	/*
+	 * In the swap-toast-by-content case, we always need to update the
+	 * cutoffs. In the swap-toast-links case, we usually assume we don't need
+	 * to change the toast table's relfrozenxid: the new version of the toast
+	 * table should already have relfrozenxid set to RecentXmin, which is good
+	 * enough. However, there's a special case - REPACK (CONCURRENTLY) - which
+	 * still needs to update the cutoffs - see the related call for more info.
+	 */
+	Assert((swap_toast_by_content && update_toast_cutoffs) ||
+		   !swap_toast_by_content);
+
 	/* Report that we are now swapping relation files */
 	pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
 								 PROGRESS_REPACK_PHASE_SWAP_REL_FILES);
@@ -2540,12 +2541,8 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 	 */
 	CommandCounterIncrement();
 	update_relation_cutoffs(OIDOldHeap, frozenXid, cutoffMulti);
-
-	/*
-	 * The same for TOAST, if needed. In the swap-toast-links case, the new
-	 * the toast table should already have relfrozenxid set to RecentXmin.
-	 */
-	if (OidIsValid(oid_old_toastid) && swap_toast_by_content)
+	/* The same for TOAST, if requested. */
+	if (OidIsValid(oid_old_toastid) && update_toast_cutoffs)
 		update_relation_cutoffs(oid_old_toastid, frozenXid, cutoffMulti);
 
 	/* Destroy new heap with old filenumber */
@@ -3041,12 +3038,14 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 	TupleTableSlot *spilled_tuple;
 	TupleTableSlot *old_update_tuple;
 	TupleTableSlot *ondisk_tuple;
+	TupleTableSlot *new_tuple;
 	bool		have_old_tuple = false;
 	bool		check_range;
 	MemoryContext oldcxt;
 	DecodingWorkerShared *shared;
 	char		fname[MAXPGPATH];
 	BufFile    *file;
+	SnapshotData SnapshotNewHeap;
 
 	/*
 	 * Use the auxiliary table if one exists, otherwise the "final"
@@ -3075,16 +3074,25 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 											table_slot_callbacks(rel));
 	old_update_tuple = MakeSingleTupleTableSlot(RelationGetDescr(rel),
 												&TTSOpsVirtual);
+	new_tuple = MakeSingleTupleTableSlot(RelationGetDescr(rel),
+										 &TTSOpsHeapTuple);
 
 	oldcxt = MemoryContextSwitchTo(GetPerTupleMemoryContext(dest->estate));
 
+	/*
+	 * Finding tuples to UPDATE / DELETE is exactly the purpose of
+	 * SNAPSHOT_NEW_HEAP.
+	 */
+	InitNewHeapSnapshot(SnapshotNewHeap);
+	PushActiveSnapshot(&SnapshotNewHeap);
+
 	while (true)
 	{
 		size_t		nread;
-		ConcurrentChangeKind prevkind = kind;
 		BlockNumber block,
 					old_block;
 		BlockNumber *old_block_p;
+		TransactionId xid;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -3099,35 +3107,18 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 		 */
 		if (kind == CHANGE_UPDATE_OLD)
 		{
-			restore_tuple(file, rel, old_update_tuple, NULL, NULL);
+			restore_tuple(file, rel, old_update_tuple, NULL, NULL, NULL);
 			have_old_tuple = true;
 			continue;
 		}
 
-		/*
-		 * Just before an UPDATE or DELETE, we must update the command
-		 * counter, because the change could refer to a tuple that we have
-		 * just inserted; and before an INSERT, we have to do this also if the
-		 * previous command was either update or delete.
-		 *
-		 * With this approach we don't spend so many CCIs for long strings of
-		 * only INSERTs, which can't affect one another.
-		 */
-		if (kind == CHANGE_UPDATE_NEW || kind == CHANGE_DELETE ||
-			(kind == CHANGE_INSERT && (prevkind == CHANGE_UPDATE_NEW ||
-									   prevkind == CHANGE_DELETE)))
-		{
-			CommandCounterIncrement();
-			UpdateActiveSnapshotCommandId();
-		}
-
 		/*
 		 * Now restore the tuple into the slot and execute the change.
 		 *
 		 * old_block is only stored with UPDATE_NEW.
 		 */
 		old_block_p = kind == CHANGE_UPDATE_NEW ? &old_block : NULL;
-		restore_tuple(file, rel, spilled_tuple, &block, old_block_p);
+		restore_tuple(file, rel, spilled_tuple, &block, old_block_p, &xid);
 
 		if (kind == CHANGE_INSERT)
 		{
@@ -3137,7 +3128,7 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 			 */
 			if (!check_range ||
 				is_block_in_range(block, range_start, range_end))
-				apply_concurrent_insert(dest, spilled_tuple);
+				apply_concurrent_insert(dest, spilled_tuple, new_tuple, xid);
 		}
 		else if (kind == CHANGE_DELETE)
 		{
@@ -3154,7 +3145,7 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 				found = find_target_tuple(dest, spilled_tuple, ondisk_tuple);
 				if (!found)
 					elog(ERROR, "could not find target tuple");
-				apply_concurrent_delete(rel, ondisk_tuple);
+				apply_concurrent_delete(rel, ondisk_tuple, xid);
 			}
 		}
 		else if (kind == CHANGE_UPDATE_NEW)
@@ -3189,7 +3180,8 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 				 */
 				adjust_toast_pointers(rel, spilled_tuple, ondisk_tuple);
 
-				apply_concurrent_update(dest, spilled_tuple, ondisk_tuple);
+				apply_concurrent_update(dest, spilled_tuple, ondisk_tuple,
+										new_tuple, xid);
 			}
 			else
 			{
@@ -3210,7 +3202,8 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 					 */
 					adjust_toast_pointers(rel, spilled_tuple, NULL);
 
-					apply_concurrent_insert(dest, spilled_tuple);
+					apply_concurrent_insert(dest, spilled_tuple, new_tuple,
+											xid);
 				}
 				else if (is_block_in_range(old_block, range_start, range_end))
 				{
@@ -3224,7 +3217,7 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 					 * visible to the snapshot that we'll use to copy the
 					 * other range.
 					 */
-					apply_concurrent_delete(rel, ondisk_tuple);
+					apply_concurrent_delete(rel, ondisk_tuple, xid);
 				}
 
 				/*
@@ -3241,11 +3234,13 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 
 		ResetPerTupleExprContext(dest->estate);
 	}
+	PopActiveSnapshot();
 
 	/* Cleanup. */
 	ExecDropSingleTupleTableSlot(spilled_tuple);
 	ExecDropSingleTupleTableSlot(ondisk_tuple);
 	ExecDropSingleTupleTableSlot(old_update_tuple);
+	ExecDropSingleTupleTableSlot(new_tuple);
 
 	MemoryContextSwitchTo(oldcxt);
 
@@ -3254,44 +3249,85 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 
 /*
  * Apply an insert from the spill of concurrent changes to the new copy of the
- * table.
+ * table. 'new_tuple' is the source for table AM.
  */
 static void
-apply_concurrent_insert(RepackDest *dest, TupleTableSlot *slot)
+apply_concurrent_insert(RepackDest *dest, TupleTableSlot *spill_tuple,
+						TupleTableSlot *new_tuple, TransactionId xid)
 {
-	/* Put the tuple in the table, but make sure it won't be decoded */
-	table_tuple_insert(dest->rel, slot, GetCurrentCommandId(true),
-					   TABLE_INSERT_NO_LOGICAL, NULL);
+	HeapTuple	tup;
+	bool		shouldFree;
+
+	/* Copy the contents to a slot that preserves the header fields. */
+	Assert(TTS_IS_HEAPTUPLE(new_tuple));
+	ExecCopySlot(new_tuple, spill_tuple);
+
+	/* Get pointer to the contained tuple (not a copy). */
+	tup = ExecFetchSlotHeapTuple(new_tuple, false, &shouldFree);
+	Assert(!shouldFree);
+
+	/* Set the XID. */
+	HeapTupleHeaderSetXmin(tup->t_data, xid);
+
+	/*
+	 * Put the tuple in the table, but make sure it won't be decoded. At the
+	 * same time, request that the XID we set above is used, instead of
+	 * generating a new one.
+	 *
+	 * FirstCommandId is ok in the new table because the transaction that
+	 * inserted the tuple has already committed, and no other transaction
+	 * should ever need the CID.
+	 */
+	table_tuple_insert(dest->rel, new_tuple, FirstCommandId,
+					   TABLE_INSERT_NO_LOGICAL | TABLE_REUSE_XID,
+					   NULL);
 
 	/* Update indexes with this new tuple. */
 	ExecInsertIndexTuples(dest->rri,
 						  dest->estate,
 						  0,
-						  slot,
+						  new_tuple,
 						  NIL, NULL);
 	pgstat_progress_incr_param(PROGRESS_REPACK_HEAP_TUPLES_INSERTED, 1);
 }
 
 /*
  * Apply an update from the spill of concurrent changes to the new copy of the
- * table.
+ * table. 'new_tuple' is the source for table AM.
  */
 static void
 apply_concurrent_update(RepackDest *dest, TupleTableSlot *spilled_tuple,
-						TupleTableSlot *ondisk_tuple)
+						TupleTableSlot *ondisk_tuple,
+						TupleTableSlot *new_tuple, TransactionId xid)
 {
+	HeapTuple	tup;
+	bool		shouldFree;
 	Relation	rel = dest->rel;
 	LockTupleMode lockmode;
 	TM_FailureData tmfd;
 	TU_UpdateIndexes update_indexes;
 	TM_Result	res;
 
+	/* Copy the contents to a slot that preserves the header fields. */
+	Assert(TTS_IS_HEAPTUPLE(new_tuple));
+	ExecCopySlot(new_tuple, spilled_tuple);
+
+	/* Get pointer to the contained tuple (not a copy). */
+	tup = ExecFetchSlotHeapTuple(new_tuple, false, &shouldFree);
+	Assert(!shouldFree);
+
+	/* Set the XID. */
+	HeapTupleHeaderSetXmin(tup->t_data, xid);
+
 	/*
 	 * Carry out the update, skipping logical decoding for it.
+	 *
+	 * See comments in apply_concurrent_insert() to understand why
+	 * FirstCommandId is ok in the new table.
 	 */
-	res = table_tuple_update(rel, &(ondisk_tuple->tts_tid), spilled_tuple,
-							 GetCurrentCommandId(true),
-							 TABLE_UPDATE_NO_LOGICAL,
+	res = table_tuple_update(rel, &(ondisk_tuple->tts_tid), new_tuple,
+							 FirstCommandId,
+							 TABLE_UPDATE_NO_LOGICAL | TABLE_REUSE_XID,
 							 InvalidSnapshot,
 							 InvalidSnapshot,
 							 false,
@@ -3311,7 +3347,7 @@ apply_concurrent_update(RepackDest *dest, TupleTableSlot *spilled_tuple,
 		ExecInsertIndexTuples(dest->rri,
 							  dest->estate,
 							  flags,
-							  spilled_tuple,
+							  new_tuple,
 							  NIL, NULL);
 	}
 
@@ -3319,16 +3355,22 @@ apply_concurrent_update(RepackDest *dest, TupleTableSlot *spilled_tuple,
 }
 
 static void
-apply_concurrent_delete(Relation rel, TupleTableSlot *slot)
+apply_concurrent_delete(Relation rel, TupleTableSlot *slot, TransactionId xid)
 {
 	TM_Result	res;
 	TM_FailureData tmfd;
 
 	/*
 	 * Delete tuple from the new heap, skipping logical decoding for it.
+	 *
+	 * See comments in heap_insert_for_repack() to understand why
+	 * FirstCommandId is ok in the new table.
+	 *
+	 * See comments in apply_concurrent_insert() to understand why
+	 * FirstCommandId is ok in the new table.
 	 */
 	res = table_tuple_delete(rel, &(slot->tts_tid),
-							 GetCurrentCommandId(true),
+							 xid, FirstCommandId,
 							 TABLE_DELETE_NO_LOGICAL,
 							 InvalidSnapshot, InvalidSnapshot,
 							 false,
@@ -3354,7 +3396,8 @@ apply_concurrent_delete(Relation rel, TupleTableSlot *slot)
  */
 static void
 restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot,
-			  BlockNumber *block_nr_p, BlockNumber *old_block_nr_p)
+			  BlockNumber *block_nr_p, BlockNumber *old_block_nr_p,
+			  TransactionId *xid_p)
 {
 	uint32		t_len;
 	HeapTuple	tup;
@@ -3377,6 +3420,8 @@ restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot,
 	/* Handle TID separate because not all tuple slots care about it. */
 	if (block_nr_p)
 		*block_nr_p = ItemPointerGetBlockNumber(&tup->t_data->t_ctid);
+	if (xid_p)
+		*xid_p = HeapTupleHeaderGetXmin(tup->t_data);
 	if (old_block_nr_p)
 		BufFileReadExact(file, old_block_nr_p, sizeof(BlockNumber));
 
@@ -3601,6 +3646,7 @@ initialize_change_context(ChangeContext *chgcxt, Relation relation,
 
 	chgcxt->cc_dest_aux = NULL;
 	chgcxt->cc_clustering_index = InvalidOid;
+	chgcxt->cc_last_snapshot_xmin = InvalidTransactionId;
 }
 
 /*
@@ -3816,6 +3862,7 @@ backup_change_context(ChangeContext *chgcxt, ChangeContextBackup *backup)
 	backup->file_seq_snapshot = chgcxt->cc_file_seq_snapshot;
 	backup->file_seq_changes = chgcxt->cc_file_seq_changes;
 	backup->clustering_index = chgcxt->cc_clustering_index;
+	backup->last_snapshot_xmin = chgcxt->cc_last_snapshot_xmin;
 }
 
 /*
@@ -3849,6 +3896,7 @@ reinitialize_change_context(ChangeContextBackup *backup)
 	chgcxt->cc_file_seq_snapshot = backup->file_seq_snapshot;
 	chgcxt->cc_file_seq_changes = backup->file_seq_changes;
 	chgcxt->cc_clustering_index = backup->clustering_index;
+	chgcxt->cc_last_snapshot_xmin = backup->last_snapshot_xmin;
 
 	return chgcxt;
 }
@@ -3961,6 +4009,7 @@ static void
 rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 								   Oid identIdx, TransactionId frozenXid,
 								   MultiXactId cutoffMulti,
+								   double num_tuples,
 								   ChangeContext *chgcxt)
 {
 	List	   *ind_oids_new;
@@ -3975,6 +4024,7 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	MemoryContext oldcxt;
 	List	   *indexrels;
 	List	   *inds_tmp = NIL;
+	BlockNumber num_pages;
 
 	Assert(CheckRelationLockedByMe(OldHeap, ShareUpdateExclusiveLock, false));
 	Assert(CheckRelationLockedByMe(NewHeap, AccessExclusiveLock, false));
@@ -3986,7 +4036,8 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	 * cache entries updated.
 	 */
 	if (chgcxt->cc_dest_aux)
-		chgcxt = process_auxiliary_table(chgcxt, &OldHeap, &NewHeap, identIdx);
+		chgcxt = process_auxiliary_table(chgcxt, &OldHeap, &NewHeap, identIdx,
+										 frozenXid, cutoffMulti);
 
 	/*
 	 * Unlike the exclusive case, we build new indexes for the new relation
@@ -4028,6 +4079,43 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	ind_oids_new = lappend_oid(ind_oids_new,
 							   RelationGetRelid(chgcxt->cc_dest.ident_index));
 
+	/*
+	 * Since we haven't copied "recently dead" tuples into the new heap, we
+	 * must not finish the processing until they are considered dead by all
+	 * backends.
+	 *
+	 * In particular, the VACUUM xmin horizon for the table must be at least
+	 * xmin of the last snapshot that we used to copy the data.  That means
+	 * even the least recently deleted tuples we omitted from the copying
+	 * (because we considered them dead) must be considered dead by anyone.
+	 *
+	 * Note: Although some time should have elapsed since the data copying
+	 * stage (at least the time to build the indexes), we might get stuck here
+	 * due to another backend running REPACK because its snapshot does not
+	 * allow the xmin horizon to advance for some time.
+	 *
+	 * TODO Consider this when determining the value of
+	 * repack_pages_per_snapshot (currently GUC, in the future preferably a
+	 * constant). Is this worth an additional phase in progress reporting?
+	 */
+	Assert(TransactionIdIsValid(chgcxt->cc_last_snapshot_xmin));
+	while (true)
+	{
+		TransactionId oldest_xmin;
+
+		oldest_xmin = GetOldestNonRemovableTransactionId(OldHeap);
+		if (TransactionIdFollowsOrEquals(oldest_xmin,
+										 chgcxt->cc_last_snapshot_xmin))
+			break;
+
+		/* Wait before the next check. */
+		(void) WaitLatch(MyLatch,
+						 WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+						 1000L,
+						 WAIT_EVENT_REPACK_MVCC_SAFETY);
+		ResetLatch(MyLatch);
+	}
+
 	/*
 	 * During testing, wait for another backend to perform concurrent data
 	 * changes which we will process below.
@@ -4149,6 +4237,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	/* The new indexes must be visible for deletion. */
 	CommandCounterIncrement();
 
+	/* Update the numbers of pages and tuples in pg_class. */
+	num_pages = RelationGetNumberOfBlocks(NewHeap);
+	copy_table_data_update_stats(OldHeap, NewHeap, num_pages, num_tuples);
+
 	/* Close the old heap but keep lock until transaction commit. */
 	table_close(OldHeap, NoLock);
 	/* Close the new heap. (We didn't have to open its indexes). */
@@ -4161,6 +4253,12 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	 * Swap the relations and their TOAST relations and TOAST indexes. This
 	 * also drops the new relation and its indexes.
 	 *
+	 * update_toast_cutoffs is true because REPACK (CONCURRENTLY) does not
+	 * freeze tuples decoded from WAL, and because RecentXmin is not affected
+	 * by logical decoding. Thus if we accepted relfrozenxid of the new TOAST
+	 * relation (derived from RecentXmin), it could incorrectly tell that we
+	 * froze more recent XID's than we actually did.
+	 *
 	 * (System catalogs are currently not supported.)
 	 */
 	Assert(!is_system_catalog);
@@ -4171,6 +4269,7 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 					 true,
 					 false,		/* reindex */
 					 frozenXid, cutoffMulti,
+					 true,		/* update_toast_cutoffs */
 					 relpersistence);
 }
 
@@ -4185,7 +4284,8 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
  */
 static ChangeContext *
 process_auxiliary_table(ChangeContext *chgcxt, Relation *pOldHeap,
-						Relation *pNewHeap, Oid identIdx)
+						Relation *pNewHeap, Oid identIdx,
+						TransactionId freeze_xid, MultiXactId cutoff_multi)
 {
 	RepackDest *dest = chgcxt->cc_dest_aux;
 	Oid			ident_idx_new;
@@ -4195,6 +4295,8 @@ process_auxiliary_table(ChangeContext *chgcxt, Relation *pOldHeap,
 	Oid			aux_oid;
 	ObjectAddress object;
 	Relation	rel;
+	SnapshotData SnapshotNewHeap;
+	RewriteState rwstate;
 
 	/*
 	 * First, make sure the clustering index exists.
@@ -4224,37 +4326,81 @@ process_auxiliary_table(ChangeContext *chgcxt, Relation *pOldHeap,
 		clustering_index = dest->ident_index;
 	}
 
+	/* Now do the copying. */
+	slot = table_slot_create(dest->rel, NULL);
+
 	/*
-	 * Now do the copying. Before starting, clear ->cc_dest_aux so that
-	 * insertions go to the final table, rather than the auxiliary one.
+	 * No point in specifying the auxiliary relation as the old one: we
+	 * haven't frozen tuples when inserting them (one freezing is enough, see
+	 * below), so the tuples do not satisfy the relfrozenxid / relminmxid
+	 * limits, and thus the following freezing would fail.
 	 */
-	chgcxt->cc_dest_aux = NULL;
-	slot = table_slot_create(dest->rel, NULL);
+	rwstate = begin_heap_rewrite(NULL, chgcxt->cc_dest.rel,
+	/* oldest_xmin only needed for rewriting */
+								 InvalidTransactionId,
+								 freeze_xid, cutoff_multi,
+								 true);
 
 	/*
-	 * Note: the current active snapshot blocks the progress of xmin
-	 * horizon(s). The next patches in the series should fix this by using a
-	 * new kind of snapshot (which we can use here because there are no
-	 * transaction aborts in the auxiliary table).
+	 * Scan of the auxiliary table can take long time, but the SnapshotNewHeap
+	 * snapshot can be used here (because there should be no aborted
+	 * insertions in the table), so the scan should not affect the xmin
+	 * horizons.
 	 */
+	PopActiveSnapshot();
+	InitNewHeapSnapshot(SnapshotNewHeap);
+	PushActiveSnapshot(&SnapshotNewHeap);
 	scan = index_beginscan(dest->rel, clustering_index, GetActiveSnapshot(),
 						   NULL, 0, 0, SO_NONE);
 	index_rescan(scan, NULL, 0, NULL, 0);
 	for (;;)
 	{
+		HeapTuple	tuple;
+		bool		shouldFree;
+
 		CHECK_FOR_INTERRUPTS();
 
 		if (!index_getnext_slot(scan, ForwardScanDirection, slot))
 			break;
 
+		/* Make sure we have a writable copy of the tuple. */
+		tuple = ExecFetchSlotHeapTuple(slot, true, &shouldFree);
+
 		/*
-		 * Reforming should have been performed during insertions into the
-		 * auxiliary table.
+		 * This kind of slot maintains the tuple header. We don't need to copy
+		 * the contents into a slot of other kind because reforming was
+		 * performed when populating the auxiliary table.
+		 */
+		Assert(TTS_IS_BUFFERTUPLE(slot));
+		Assert(TransactionIdIsValid(HeapTupleHeaderGetXmin(tuple->t_data)));
+
+		/*
+		 * Insert the tuple into the new relation, and freeze it while doing
+		 * so.
+		 *
+		 * Since our copy is already writable, the tuple can be passed for
+		 * both old and new tuple.
 		 */
-		heap_insert_for_repack(chgcxt, slot, NULL);
+		rewrite_heap_tuple_no_chains(rwstate, tuple, tuple, true);
+
+		if (shouldFree)
+			pfree(tuple);
 	}
 	index_endscan(scan);
+	PopActiveSnapshot();
+	InvalidateCatalogSnapshot();
+
+	/*
+	 * We should not be restricting the progress of xmin horizons at the
+	 * moment.
+	 */
+	Assert(!TransactionIdIsValid(MyProc->xmin));
+	Assert(!TransactionIdIsValid(MyProc->xid));
+	Assert(!HaveRegisteredOrActiveSnapshot());
+
+	PushActiveSnapshot(GetTransactionSnapshot());
 	ExecDropSingleTupleTableSlot(slot);
+	end_heap_rewrite(rwstate);
 
 	/*
 	 * Close the relation, its identity index and clustering index if we had
@@ -4266,6 +4412,7 @@ process_auxiliary_table(ChangeContext *chgcxt, Relation *pOldHeap,
 		index_close(clustering_index, NoLock);
 	/* Here we close the other indexes. */
 	release_change_dest(dest);
+	chgcxt->cc_dest_aux = NULL;
 
 	/* Drop the auxiliary table. */
 	object.classId = RelationRelationId;
diff --git a/src/backend/commands/tablecmds.c b/src/backend/commands/tablecmds.c
index 427fe733153..5c684235e22 100644
--- a/src/backend/commands/tablecmds.c
+++ b/src/backend/commands/tablecmds.c
@@ -6122,6 +6122,7 @@ ATRewriteTables(AlterTableStmt *parsetree, List **wqueue, LOCKMODE lockmode,
 							 true,	/* reindex */
 							 RecentXmin,
 							 ReadNextMultiXactId(),
+							 false, /* update_toast_cutoffs */
 							 persistence);
 
 			InvokeObjectPostAlterHook(RelationRelationId, tab->relid, 0);
diff --git a/src/backend/executor/nodeModifyTable.c b/src/backend/executor/nodeModifyTable.c
index b9781eb3b95..265e80b765b 100644
--- a/src/backend/executor/nodeModifyTable.c
+++ b/src/backend/executor/nodeModifyTable.c
@@ -1765,6 +1765,7 @@ ExecDeleteAct(ModifyTableContext *context, ResultRelInfo *resultRelInfo,
 		options |= TABLE_DELETE_CHANGING_PARTITION;
 
 	return table_tuple_delete(resultRelInfo->ri_RelationDesc, tupleid,
+							  InvalidTransactionId,
 							  estate->es_output_cid,
 							  options,
 							  estate->es_snapshot,
diff --git a/src/backend/replication/logical/decode.c b/src/backend/replication/logical/decode.c
index c3722b5c623..e3c94f83875 100644
--- a/src/backend/replication/logical/decode.c
+++ b/src/backend/replication/logical/decode.c
@@ -425,6 +425,18 @@ heap2_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	TransactionId xid = XLogRecGetXid(buf->record);
 	SnapBuild  *builder = ctx->snapshot_builder;
 
+	/*
+	 * XLOG_HEAP2_MULTI_INSERT is not replayed.
+	 */
+	Assert((XLogRecGetInfo(buf->record) & XLR_XID_REPLAYED) == 0);
+
+	/* See heap_decode(). */
+	if (change_useless_for_repack(buf))
+	{
+		Assert(!ctx->fast_forward);
+		return;
+	}
+
 	ReorderBufferProcessXid(ctx->reorder, xid, buf->origptr);
 
 	/*
@@ -442,8 +454,7 @@ heap2_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	{
 		case XLOG_HEAP2_MULTI_INSERT:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeMultiInsert(ctx, buf);
 			break;
 		case XLOG_HEAP2_NEW_CID:
@@ -488,6 +499,50 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	TransactionId xid = XLogRecGetXid(buf->record);
 	SnapBuild  *builder = ctx->snapshot_builder;
 
+	/*
+	 * REPACK decoding should only decode changes of the relation being
+	 * processed. Decoding changes of other tables would not only introduce
+	 * performance overhead, it would also make us process already committed
+	 * (and replayed) XIDs again - that's probably not expected by the logical
+	 * decoding system.
+	 *
+	 * The problem is that during the replay, REPACK can (and does) avoid WAL
+	 * logging the extra information needed for logical decoding, however this
+	 * is checked in DecodeInsert(), DecodeUpdate(), etc., which is too late.
+	 * Moreover, these function do not distinguish whether the transaction is
+	 * being decoded the first time, or if the WAL records originate from the
+	 * replay phase of REPACK.
+	 *
+	 * Unlike the fast-forward case (see comments below), REPACK does not need
+	 * the base snapshot for the transaction until it receives a change that
+	 * really needs to be decoded. Thus it's ok to skip
+	 * SnapBuildProcessChange().
+	 *
+	 * (With fast-forward, we must not omit the setup of the transaction base
+	 * snapshot because the changes skipped by fast-forward initially may need
+	 * to be decoded after restart. Thus the base snapshot may be needed after
+	 * the restart too. If we didn't create the snapshot in the fast-forward
+	 * mode, the snapshot builder's xmin would advance too eagerly, so the
+	 * same snapshot wouldn't work after restart.)
+	 *
+	 * First, filter out WAL records generated by REPACK (CONCURRENTLY)
+	 * replaying the data changes of other transactions - these transactions
+	 * have already been decoded, so no backend / worker should decode them
+	 * again.
+	 */
+	if (XLogRecGetInfo(buf->record) & XLR_XID_REPLAYED)
+		return;
+
+	/*
+	 * Now let REPACK decoding worker filter out changes of tables other than
+	 * the one whose REPACKing it's involved in.
+	 */
+	if (change_useless_for_repack(buf))
+	{
+		Assert(!ctx->fast_forward);
+		return;
+	}
+
 	ReorderBufferProcessXid(ctx->reorder, xid, buf->origptr);
 
 	/*
@@ -505,8 +560,7 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	{
 		case XLOG_HEAP_INSERT:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeInsert(ctx, buf);
 			break;
 
@@ -518,22 +572,19 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		case XLOG_HEAP_HOT_UPDATE:
 		case XLOG_HEAP_UPDATE:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeUpdate(ctx, buf);
 			break;
 
 		case XLOG_HEAP_DELETE:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeDelete(ctx, buf);
 			break;
 
 		case XLOG_HEAP_TRUNCATE:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeTruncate(ctx, buf);
 			break;
 
@@ -549,8 +600,7 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 
 		case XLOG_HEAP_CONFIRM:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeSpecConfirm(ctx, buf);
 			break;
 
@@ -960,15 +1010,17 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	DecodeXLogTuple(tupledata, datalen, change->data.tp.newtuple);
 
 	/*
-	 * REPACK (CONCURRENTLY) needs block number to check if the corresponding
-	 * part of the table was already copied.  XXX Should we only do this if
-	 * AmRepackWorker()? It might save a few cycles, but not sure it's good to
-	 * leave the fields unset in other cases.
+	 * REPACK (CONCURRENTLY) needs xmin to preserve visibility information and
+	 * block number to check if the corresponding part of the table was
+	 * already copied.  XXX Should we only do this if AmRepackWorker()? It
+	 * might save a few cycles, but not sure it's good to leave the fields
+	 * unset in other cases.
 	 */
 	{
 		HeapTupleHeader header;
 
 		header = change->data.tp.newtuple->t_data;
+		HeapTupleHeaderSetXmin(header, XLogRecGetXid(r));
 		/* offnum is not really needed, but let's set valid pointer. */
 		ItemPointerSet(&header->t_ctid, blknum, xlrec->offnum);
 	}
@@ -1033,14 +1085,15 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		DecodeXLogTuple(data, datalen, change->data.tp.newtuple);
 
 		/*
-		 * REPACK (CONCURRENTLY) needs block numbers to check if the
-		 * corresponding part of the table was already copied. XXX Do this
-		 * only if AmRepackWorker()?
+		 * REPACK (CONCURRENTLY) needs xmin to preserve visibility information
+		 * and block numbers to check if the corresponding part of the table
+		 * was already copied. XXX Do this only if AmRepackWorker()?
 		 */
 		{
 			HeapTupleHeader header;
 
 			header = change->data.tp.newtuple->t_data;
+			HeapTupleHeaderSetXmin(header, XLogRecGetXid(r));
 			/* offnum is not really needed, but let's set valid pointer. */
 			ItemPointerSet(&header->t_ctid, new_blknum, xlrec->new_offnum);
 			change->data.tp.old_blknum = old_blknum;
@@ -1129,14 +1182,20 @@ DecodeDelete(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 						datalen, change->data.tp.oldtuple);
 
 		/*
-		 * REPACK (CONCURRENTLY) needs block number to check if the
-		 * corresponding part of the table was already copied. XXX Do this
-		 * only if AmRepackWorker()?
+		 * REPACK (CONCURRENTLY) needs xmax to preserve visibility information
+		 * and block number to check if the corresponding part of the table
+		 * was already copied. XXX Do this only if AmRepackWorker()?
 		 */
 		{
 			HeapTupleHeader header;
 
 			header = change->data.tp.oldtuple->t_data;
+
+			/*
+			 * xmax makes more sense here, but we don't want restore_tuple()
+			 * to pay attention to the change kind, so use xmin here as well.
+			 */
+			HeapTupleHeaderSetXmin(header, XLogRecGetXid(r));
 			/* offnum is not really needed, but let's set valid pointer. */
 			ItemPointerSet(&header->t_ctid, blknum, xlrec->offnum);
 		}
@@ -1282,13 +1341,16 @@ DecodeMultiInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 			change->data.tp.clear_toast_afterwards = false;
 
 		/*
-		 * REPACK (CONCURRENTLY) needs block number to check if the
-		 * corresponding part of the table was already copied.
+		 * REPACK (CONCURRENTLY) needs xmin to preserve visibility information
+		 * and block number to check if the corresponding part of the table
+		 * was already copied.
 		 */
 		if (AmRepackWorker())
 		{
 			OffsetNumber offnum;
 
+			HeapTupleHeaderSetXmin(header, XLogRecGetXid(r));
+
 			/*
 			 * offnum is not really needed, but let's set valid pointer. (It
 			 * will be invalid anyway if the page was initially empty.)
diff --git a/src/backend/replication/logical/reorderbuffer.c b/src/backend/replication/logical/reorderbuffer.c
index cae2b099e69..626e9400724 100644
--- a/src/backend/replication/logical/reorderbuffer.c
+++ b/src/backend/replication/logical/reorderbuffer.c
@@ -5263,6 +5263,9 @@ ReorderBufferToastReplace(ReorderBuffer *rb, ReorderBufferTXN *txn,
 	 * Shouldn't we add a new field to ReorderBufferChange instead?
 	 */
 	tmphtup->t_data->t_ctid = newtup->t_data->t_ctid;
+	/* Likewise, preserve XID - REPACK needs it to be MVCC-safe. */
+	HeapTupleHeaderSetXmin(tmphtup->t_data,
+						   HeapTupleHeaderGetXmin(newtup->t_data));
 
 	memcpy(newtup->t_data, tmphtup->t_data, tmphtup->t_len);
 	newtup->t_len = tmphtup->t_len;
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 1016502d042..30fcd952132 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -184,6 +184,7 @@ PG_SLEEP	"Waiting due to a call to <function>pg_sleep</function> or a sibling fu
 RECOVERY_APPLY_DELAY	"Waiting to apply WAL during recovery because of a delay setting."
 RECOVERY_RETRIEVE_RETRY_INTERVAL	"Waiting during recovery when WAL data is not available from any source (<filename>pg_wal</filename>, archive or stream)."
 REGISTER_SYNC_REQUEST	"Waiting while sending synchronization requests to the checkpointer, because the request queue is full."
+REPACK_MVCC_SAFETY	"Waiting until not copied tuples are considered dead."
 SPIN_DELAY	"Waiting while acquiring a contended spinlock."
 VACUUM_DELAY	"Waiting in a cost-based vacuum delay point."
 VACUUM_TRUNCATE	"Waiting to acquire an exclusive lock to truncate off any empty pages at the end of a table vacuumed."
diff --git a/src/include/access/heapam.h b/src/include/access/heapam.h
index 5176478c295..bc6eadcb978 100644
--- a/src/include/access/heapam.h
+++ b/src/include/access/heapam.h
@@ -380,6 +380,7 @@ extern void heap_multi_insert(Relation relation, TupleTableSlot **slots,
 							  int ntuples, CommandId cid, uint32 options,
 							  BulkInsertState bistate);
 extern TM_Result heap_delete(Relation relation, const ItemPointerData *tid,
+							 TransactionId xid,
 							 CommandId cid, uint32 options, Snapshot crosscheck,
 							 bool wait, TM_FailureData *tmfd);
 extern void heap_finish_speculative(Relation relation, const ItemPointerData *tid);
@@ -422,7 +423,8 @@ extern bool heap_tuple_should_freeze(HeapTupleHeader tuple,
 extern bool heap_tuple_needs_eventual_freeze(HeapTupleHeader tuple);
 
 extern void simple_heap_insert(Relation relation, HeapTuple tup);
-extern void simple_heap_delete(Relation relation, const ItemPointerData *tid);
+extern void simple_heap_delete(Relation relation, const ItemPointerData *tid,
+							   TransactionId xid);
 extern void simple_heap_update(Relation relation, const ItemPointerData *otid,
 							   HeapTuple tup, TU_UpdateIndexes *update_indexes);
 
diff --git a/src/include/access/heaptoast.h b/src/include/access/heaptoast.h
index 631cb1836b9..36d1d65c131 100644
--- a/src/include/access/heaptoast.h
+++ b/src/include/access/heaptoast.h
@@ -14,6 +14,7 @@
 #define HEAPTOAST_H
 
 #include "access/htup_details.h"
+#include "access/rewriteheap.h"
 #include "storage/lockdefs.h"
 #include "utils/relcache.h"
 
@@ -95,7 +96,9 @@
  * ----------
  */
 extern HeapTuple heap_toast_insert_or_update(Relation rel, HeapTuple newtup,
-											 HeapTuple oldtup, uint32 options);
+											 HeapTuple oldtup,
+											 RewriteState rwstate,
+											 uint32 options);
 
 /* ----------
  * heap_toast_delete -
@@ -104,7 +107,7 @@ extern HeapTuple heap_toast_insert_or_update(Relation rel, HeapTuple newtup,
  * ----------
  */
 extern void heap_toast_delete(Relation rel, HeapTuple oldtup,
-							  bool is_speculative);
+							  bool is_speculative, TransactionId xid);
 
 /* ----------
  * toast_flatten_tuple -
diff --git a/src/include/access/rewriteheap.h b/src/include/access/rewriteheap.h
index 6ccf7b45c04..80d928e0e26 100644
--- a/src/include/access/rewriteheap.h
+++ b/src/include/access/rewriteheap.h
@@ -22,12 +22,21 @@
 typedef struct RewriteStateData *RewriteState;
 
 extern RewriteState begin_heap_rewrite(Relation old_heap, Relation new_heap,
-									   TransactionId oldest_xmin, TransactionId freeze_xid,
-									   MultiXactId cutoff_multi);
+									   TransactionId oldest_xmin,
+									   TransactionId freeze_xid,
+									   MultiXactId cutoff_multi,
+									   bool no_chains);
 extern void end_heap_rewrite(RewriteState state);
 extern void rewrite_heap_tuple(RewriteState state, HeapTuple old_tuple,
 							   HeapTuple new_tuple);
+extern void rewrite_heap_tuple_no_chains(RewriteState state,
+										 HeapTuple old_tuple,
+										 HeapTuple new_tuple,
+										 bool freeze);
 extern bool rewrite_heap_dead_tuple(RewriteState state, HeapTuple old_tuple);
+extern void rewrite_freeze_tuple(RewriteState state, HeapTuple tuple);
+extern void rewrite_copy_visibility_info(HeapTuple new_tuple,
+										 HeapTuple old_tuple);
 
 /*
  * On-Disk data format for an individual logical rewrite mapping.
diff --git a/src/include/access/tableam.h b/src/include/access/tableam.h
index 132248c5d43..1af01235c8b 100644
--- a/src/include/access/tableam.h
+++ b/src/include/access/tableam.h
@@ -292,6 +292,12 @@ typedef struct TM_IndexDeleteOp
 /* "options" flag bits for table_tuple_update */
 #define TABLE_UPDATE_NO_LOGICAL					(1 << 0)
 
+/*
+ * For INSERT or UPDATE, use XID (xmin) contained the in new tuple rather than
+ * the XID of the current transaction.
+ */
+#define TABLE_REUSE_XID							(1 << 31)
+
 /* flag bits for table_tuple_lock */
 /* Follow tuples whose update is in progress if lock modes don't conflict  */
 #define TUPLE_LOCK_FLAG_LOCK_UPDATE_IN_PROGRESS	(1 << 0)
@@ -568,6 +574,7 @@ typedef struct TableAmRoutine
 	/* see table_tuple_delete() for reference about parameters */
 	TM_Result	(*tuple_delete) (Relation rel,
 								 ItemPointer tid,
+								 TransactionId xid,
 								 CommandId cid,
 								 uint32 options,
 								 Snapshot snapshot,
@@ -1458,6 +1465,14 @@ static inline void
 table_tuple_insert(Relation rel, TupleTableSlot *slot, CommandId cid,
 				   uint32 options, BulkInsertStateData *bistate)
 {
+	/*
+	 * TABLE_REUSE_XID restricts the slot type because not all slots preserve
+	 * the visibility information. XXX Isn't this a reason to pass the xid as
+	 * an argument?
+	 */
+	Assert((options & TABLE_REUSE_XID) == 0 || TTS_IS_HEAPTUPLE(slot) ||
+		   TTS_IS_BUFFERTUPLE(slot));
+
 	rel->rd_tableam->tuple_insert(rel, slot, cid, options,
 								  bistate);
 }
@@ -1479,6 +1494,9 @@ table_tuple_insert_speculative(Relation rel, TupleTableSlot *slot,
 							   BulkInsertStateData *bistate,
 							   uint32 specToken)
 {
+	/* TABLE_REUSE_XID is currently not needed here. */
+	Assert((options & TABLE_REUSE_XID) == 0);
+
 	rel->rd_tableam->tuple_insert_speculative(rel, slot, cid, options,
 											  bistate, specToken);
 }
@@ -1526,6 +1544,7 @@ table_multi_insert(Relation rel, TupleTableSlot **slots, int nslots,
  * Input parameters:
  *	rel - table to be modified (caller must hold suitable lock)
  *	tid - TID of tuple to be deleted
+ *	xid - XID to use or InvalidTransactionId for the current transaction
  *	cid - delete command ID (used for visibility test, and stored into
  *		cmax if successful)
  *	options - bitmask of options.  Supported values:
@@ -1546,11 +1565,12 @@ table_multi_insert(Relation rel, TupleTableSlot **slots, int nslots,
  * TM_FailureData for additional info.
  */
 static inline TM_Result
-table_tuple_delete(Relation rel, ItemPointer tid, CommandId cid,
+table_tuple_delete(Relation rel, ItemPointer tid, TransactionId xid,
+				   CommandId cid,
 				   uint32 options, Snapshot snapshot, Snapshot crosscheck,
 				   bool wait, TM_FailureData *tmfd)
 {
-	return rel->rd_tableam->tuple_delete(rel, tid, cid, options,
+	return rel->rd_tableam->tuple_delete(rel, tid, xid, cid, options,
 										 snapshot, crosscheck,
 										 wait, tmfd);
 }
@@ -1601,6 +1621,10 @@ table_tuple_update(Relation rel, ItemPointer otid, TupleTableSlot *slot,
 				   bool wait, TM_FailureData *tmfd, LockTupleMode *lockmode,
 				   TU_UpdateIndexes *update_indexes)
 {
+	/* See table_tuple_insert(). */
+	Assert((options & TABLE_REUSE_XID) == 0 || TTS_IS_HEAPTUPLE(slot) ||
+		   TTS_IS_BUFFERTUPLE(slot));
+
 	return rel->rd_tableam->tuple_update(rel, otid, slot,
 										 cid, options, snapshot, crosscheck,
 										 wait, tmfd,
diff --git a/src/include/access/toast_helper.h b/src/include/access/toast_helper.h
index 2ec92397f26..f86f88c774e 100644
--- a/src/include/access/toast_helper.h
+++ b/src/include/access/toast_helper.h
@@ -14,6 +14,7 @@
 #ifndef TOAST_HELPER_H
 #define TOAST_HELPER_H
 
+#include "access/rewriteheap.h"
 #include "utils/rel.h"
 
 /*
@@ -60,6 +61,15 @@ typedef struct
 	 */
 	uint8		ttc_flags;
 	ToastAttrInfo *ttc_attr;
+
+	/*
+	 * These fields are needed when the option HEAP_INSERT_REUSE_XID_FREEZE
+	 * was passed to heap_toast_insert_or_update(). We could actually use
+	 * normal insert, but that would require WAL support of
+	 * heap_freeze_tuple().
+	 */
+	RewriteState ttc_rwstate;
+	HeapTuple	ttc_tup_main;
 } ToastTupleContext;
 
 /*
@@ -108,9 +118,9 @@ extern int	toast_tuple_find_biggest_attribute(ToastTupleContext *ttc,
 extern void toast_tuple_try_compression(ToastTupleContext *ttc, int attribute);
 extern void toast_tuple_externalize(ToastTupleContext *ttc, int attribute,
 									uint32 options);
-extern void toast_tuple_cleanup(ToastTupleContext *ttc);
+extern void toast_tuple_cleanup(ToastTupleContext *ttc, TransactionId xid);
 
 extern void toast_delete_external(Relation rel, const Datum *values, const bool *isnull,
-								  bool is_speculative);
+								  bool is_speculative, TransactionId xid);
 
 #endif
diff --git a/src/include/access/toast_internals.h b/src/include/access/toast_internals.h
index bf45889a642..62da4a18d63 100644
--- a/src/include/access/toast_internals.h
+++ b/src/include/access/toast_internals.h
@@ -12,6 +12,7 @@
 #ifndef TOAST_INTERNALS_H
 #define TOAST_INTERNALS_H
 
+#include "access/rewriteheap.h"
 #include "access/toast_compression.h"
 #include "storage/lockdefs.h"
 #include "utils/relcache.h"
@@ -48,9 +49,13 @@ typedef struct toast_compress_header
 extern Datum toast_compress_datum(Datum value, char cmethod);
 extern Oid	toast_get_valid_index(Oid toastoid, LOCKMODE lock);
 
-extern void toast_delete_datum(Relation rel, Datum value, bool is_speculative);
+extern void toast_delete_datum(Relation rel, Datum value, bool is_speculative,
+							   TransactionId xid);
 extern Datum toast_save_datum(Relation rel, Datum value,
-							  varlena *oldexternal, uint32 options);
+							  varlena *oldexternal,
+							  RewriteState rwstate,
+							  HeapTuple tup_main,
+							  uint32 options);
 
 extern int	toast_open_indexes(Relation toastrel,
 							   LOCKMODE lock,
diff --git a/src/include/access/xlog_internal.h b/src/include/access/xlog_internal.h
index 55663e6f4af..be718993401 100644
--- a/src/include/access/xlog_internal.h
+++ b/src/include/access/xlog_internal.h
@@ -32,7 +32,7 @@
 /*
  * Each page of XLOG file has a header like this:
  */
-#define XLOG_PAGE_MAGIC 0xD120	/* can be used as WAL version indicator */
+#define XLOG_PAGE_MAGIC 0xD121	/* can be used as WAL version indicator */
 
 typedef struct XLogPageHeaderData
 {
diff --git a/src/include/access/xloginsert.h b/src/include/access/xloginsert.h
index 91dfbd5627f..a0cea5fde01 100644
--- a/src/include/access/xloginsert.h
+++ b/src/include/access/xloginsert.h
@@ -43,6 +43,7 @@
 /* prototypes for public functions in xloginsert.c: */
 extern void XLogBeginInsert(void);
 extern void XLogSetRecordFlags(uint8 flags);
+extern void XLogSetRecordXid(TransactionId xid);
 extern XLogRecPtr XLogInsert(RmgrId rmid, uint8 info);
 extern XLogRecPtr XLogSimpleInsertInt64(RmgrId rmid, uint8 info, int64 value);
 extern void XLogEnsureRecordSpace(int max_block_id, int ndatas);
diff --git a/src/include/access/xlogrecord.h b/src/include/access/xlogrecord.h
index e8999d3fe91..bf3c1965e05 100644
--- a/src/include/access/xlogrecord.h
+++ b/src/include/access/xlogrecord.h
@@ -90,6 +90,14 @@ typedef struct XLogRecord
  */
 #define XLR_CHECK_CONSISTENCY	0x02
 
+/*
+ * The record contains a data change that was already committed and now is
+ * being applied to a new relation due to rewriting. The original XID is
+ * needed to keep the rewriting MVCC-safe, however the transaction should be
+ * ignored by logical decoding and it should not get into KnownAssignedXids.
+ */
+#define XLR_XID_REPLAYED		0x04
+
 /*
  * Header info for block data appended to an XLOG record.
  *
diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h
index ad8125790b0..9e2f3e491c6 100644
--- a/src/include/commands/repack.h
+++ b/src/include/commands/repack.h
@@ -16,6 +16,7 @@
 #include <signal.h>
 
 #include "access/hio.h"
+#include "access/rewriteheap.h"
 #include "access/skey.h"
 #include "access/xlogdefs.h"
 #include "catalog/index.h"
@@ -110,10 +111,10 @@ typedef struct ChangeContext
 	 * Not sure, it'd require disk space for one more copy and the copying
 	 * itself is not free.
 	 *
-	 * TODO 1) make the tables unlogged, 2) if REPACK locks the TOAST relation
-	 * too (not sure it does) try to preserve TOAST pointers, instead of
-	 * storing them to TOAST relations of these tables, 3) Check that the
-	 * tables are dropped on transaction abort.
+	 * XXX REPACK currently does not lock the old TOAST relation. If it did,
+	 * we could perhaps copy TOAST pointers from the old relation to the
+	 * auxiliary relation, so that the auxiliary relation would not need its
+	 * own TOAST relation.
 	 */
 	RepackDest *cc_dest_aux;
 
@@ -127,6 +128,11 @@ typedef struct ChangeContext
 	 * functions. This is needed when starting a new transaction.
 	 */
 	IndexBuildSecurity cc_ind_build_sec;
+
+	/*
+	 * xmin of the last snapshot used to copy data.
+	 */
+	TransactionId cc_last_snapshot_xmin;
 } ChangeContext;
 
 extern PGDLLIMPORT int repack_pages_per_snapshot;
@@ -142,8 +148,6 @@ extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal);
 extern Oid	make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
 						  char relpersistence, LOCKMODE lockmode,
 						  bool auxiliary);
-extern void heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
-								   TupleTableSlot *reform);
 extern bool tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc);
 extern void clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform);
 extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
@@ -154,6 +158,7 @@ extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 							 bool reindex,
 							 TransactionId frozenXid,
 							 MultiXactId cutoffMulti,
+							 bool update_toast_cutoffs,
 							 char newrelpersistence);
 extern Snapshot repack_get_snapshot(ChangeContext *chgcxt);
 extern void repack_process_concurrent_changes(ChangeContext *chgcxt,
diff --git a/src/include/utils/snapmgr.h b/src/include/utils/snapmgr.h
index 1c550096393..b0c40df7934 100644
--- a/src/include/utils/snapmgr.h
+++ b/src/include/utils/snapmgr.h
@@ -51,6 +51,18 @@ extern PGDLLIMPORT SnapshotData SnapshotToastData;
 	((snapshotdata).snapshot_type = SNAPSHOT_NON_VACUUMABLE, \
 	 (snapshotdata).vistest = (vistestp))
 
+/*
+ * NewHeap snapshot needs to be used as the active snapshot at some point, so
+ * initialize the fields related to PushActiveSnapshot().
+ */
+#define InitNewHeapSnapshot(snapshotdata)  \
+	((snapshotdata).snapshot_type = SNAPSHOT_NEW_HEAP, \
+	 (snapshotdata).regd_count = 0, \
+	 (snapshotdata).active_count = 0, \
+	 (snapshotdata).copied = false, \
+	 (snapshotdata).xcnt = 0, \
+	 (snapshotdata).subxcnt = 0)
+
 /*
  * Is the snapshot implemented as an MVCC snapshot (i.e. it uses
  * SNAPSHOT_MVCC)? If so, there will be at most one visible tuple in a chain
diff --git a/src/include/utils/snapshot.h b/src/include/utils/snapshot.h
index 9766aabcad4..f91cb43a5a2 100644
--- a/src/include/utils/snapshot.h
+++ b/src/include/utils/snapshot.h
@@ -112,6 +112,29 @@ typedef enum SnapshotType
 	 * horizon to use.
 	 */
 	SNAPSHOT_NON_VACUUMABLE,
+
+	/*
+	 * The effects of all transactions are visible. Unlike SNAPSHOT_DIRTY,
+	 * aborted (sub)transactions are not expected.
+	 *
+	 * This is specific to applying data changes to the new heap by the REPACK
+	 * command - that replays applies changes done in the old heap by
+	 * transactions that have already committed. No other transactions can
+	 * access the new heap while this snapshot is in use.
+	 *
+	 * Therefore, whenever a transaction being applied looks for a tuple to
+	 * update or delete, it can assume that the insertion of any candidate
+	 * tuple was already committed - otherwise the inserting transaction
+	 * wouldn't have been applied.
+	 *
+	 * By considering effects of all transactions visible we also ensure that
+	 * a transaction can update / delete tuples that it inserted itself.
+	 *
+	 * TODO Consider better name. Would SNAPSHOT_BOOTSTRAP be confusing?
+	 * During cluster bootstrap we also consider all changes committed
+	 * immediately.
+	 */
+	SNAPSHOT_NEW_HEAP,
 } SnapshotType;
 
 typedef struct SnapshotData *Snapshot;
@@ -127,8 +150,8 @@ typedef struct SnapshotData *Snapshot;
  * * Historic MVCC snapshots used during logical decoding
  * * snapshots passed to HeapTupleSatisfiesDirty()
  * * snapshots passed to HeapTupleSatisfiesNonVacuumable()
- * * snapshots used for SatisfiesAny, Toast, Self where no members are
- *	 accessed.
+ * * snapshots used for SatisfiesAny, Toast, Self, NewHeap where no members
+ *	 are accessed.
  *
  * TODO: It's probably a good idea to split this struct using a NodeTag
  * similar to how parser and executor nodes are handled, with one type for
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..9919c93fe7f 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -45,8 +45,8 @@ step check2:
 
 	SELECT i, j FROM repack_test ORDER BY i, j;
 
-	INSERT INTO data_s2(i, j)
-	SELECT i, j FROM repack_test;
+	INSERT INTO data_s2(i, j, _xmin)
+	SELECT i, j, xmin FROM repack_test;
 
   i|  j
 ---+---
@@ -77,11 +77,11 @@ step check1:
 
 	SELECT i, j FROM repack_test ORDER BY i, j;
 
-	INSERT INTO data_s1(i, j)
-	SELECT i, j FROM repack_test;
+	INSERT INTO data_s1(i, j, _xmin)
+	SELECT i, j, xmin FROM repack_test;
 
 	SELECT count(*)
-	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j, _xmin)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 
 count
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index 7896d1456ad..3e15f5db31f 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -9,8 +9,8 @@ setup
 
 	CREATE TABLE relfilenodes(node oid);
 
-	CREATE TABLE data_s1(i int, j int);
-	CREATE TABLE data_s2(i int, j int);
+	CREATE TABLE data_s1(i int, j int, _xmin xid);
+	CREATE TABLE data_s2(i int, j int, _xmin xid);
 }
 
 teardown
@@ -39,7 +39,8 @@ step wait_before_lock
 # Besides the contents, we also check that relfilenode has changed.
 
 # Have each session write the contents into a table and use FULL JOIN to check
-# if the outputs are identical.
+# if the outputs are identical. xmin is included in order to check the MVCC
+# safety.
 step check1
 {
 	INSERT INTO relfilenodes(node)
@@ -49,11 +50,11 @@ step check1
 
 	SELECT i, j FROM repack_test ORDER BY i, j;
 
-	INSERT INTO data_s1(i, j)
-	SELECT i, j FROM repack_test;
+	INSERT INTO data_s1(i, j, _xmin)
+	SELECT i, j, xmin FROM repack_test;
 
 	SELECT count(*)
-	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j, _xmin)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
 teardown
@@ -85,10 +86,6 @@ step change_new
 
 # When applying concurrent data changes, we should see the effects of an
 # in-progress subtransaction.
-#
-# XXX Not sure this test is useful now - it was designed for the patch that
-# preserves tuple visibility and which therefore modifies
-# TransactionIdIsCurrentTransactionId().
 step change_subxact1
 {
 	BEGIN;
@@ -102,8 +99,6 @@ step change_subxact1
 
 # When applying concurrent data changes, we should not see the effects of a
 # rolled back subtransaction.
-#
-# XXX Is this test useful? See above.
 step change_subxact2
 {
 	BEGIN;
@@ -122,8 +117,8 @@ step check2
 
 	SELECT i, j FROM repack_test ORDER BY i, j;
 
-	INSERT INTO data_s2(i, j)
-	SELECT i, j FROM repack_test;
+	INSERT INTO data_s2(i, j, _xmin)
+	SELECT i, j, xmin FROM repack_test;
 }
 step wakeup_before_lock
 {
-- 
2.52.0

Attachments:

  [text/x-diff] v02-0001-Use-tuple-slot-to-pass-tuples-for-rewriting.patch (11.5K, ../108776.1784105248@localhost/2-v02-0001-Use-tuple-slot-to-pass-tuples-for-rewriting.patch)
  download | inline diff:
From d3b84b47e0ff5a0d9bc9b5ffc949b74610bdd84c Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 15 Jul 2026 10:37:11 +0200
Subject: [PATCH 1/8] Use tuple slot to pass tuples for rewriting.

This patch tries to adopt the preferable way of handling tuples, i.e. pass the
containing tuple slot rather than the actual tuple. The motivation is that
heap_insert_for_repack() will need a slot in the near future, in order to call
ExecInsertIndexTuples(). Also, in order to set values of dropped attributes to
to NULL, it seems more compact to pass a tuple slot as a workspace than two
arrays (one for values and one for nulls).
---
 src/backend/access/heap/heapam_handler.c | 199 +++++++++++++----------
 1 file changed, 110 insertions(+), 89 deletions(-)

diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index bf87430cf01..4db8a16f691 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -48,14 +48,13 @@
 #include "utils/rel.h"
 #include "utils/tuplesort.h"
 
-static void reform_and_rewrite_tuple(HeapTuple tuple,
-									 Relation OldHeap, Relation NewHeap,
-									 Datum *values, bool *isnull, RewriteState rwstate);
-static void heap_insert_for_repack(HeapTuple tuple, Relation OldHeap,
-								   Relation NewHeap, Datum *values, bool *isnull,
-								   BulkInsertState bistate);
-static HeapTuple reform_tuple(HeapTuple tuple, Relation OldHeap,
-							  Relation NewHeap, Datum *values, bool *isnull);
+static void reform_and_rewrite_tuple(TupleTableSlot *src, TupleTableSlot *reform,
+									 RewriteState rwstate);
+static void heap_insert_for_repack(Relation rel, TupleTableSlot *src,
+								   TupleTableSlot *reform,
+								   BulkInsertStateData *bistate);
+static bool tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc);
+static void clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform);
 
 static bool SampleHeapTupleVisible(TableScanDesc scan, Buffer buffer,
 								   HeapTuple tuple,
@@ -603,11 +602,8 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	bool		is_system_catalog;
 	Tuplesortstate *tuplesort;
 	TupleDesc	oldTupDesc = RelationGetDescr(OldHeap);
-	TupleDesc	newTupDesc = RelationGetDescr(NewHeap);
 	TupleTableSlot *slot;
-	int			natts;
-	Datum	   *values;
-	bool	   *isnull;
+	TupleTableSlot *reform_slot;
 	BufferHeapTupleTableSlot *hslot;
 	BlockNumber prev_cblock = InvalidBlockNumber;
 	bool		concurrent = snapshot != NULL;
@@ -621,11 +617,6 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	 */
 	Assert(RelationGetTargetBlock(NewHeap) == InvalidBlockNumber);
 
-	/* Preallocate values/isnull arrays */
-	natts = newTupDesc->natts;
-	values = palloc_array(Datum, natts);
-	isnull = palloc_array(bool, natts);
-
 	/*
 	 * In non-concurrent mode, initialize the rewrite operation.  This is not
 	 * needed in concurrent mode.
@@ -699,6 +690,8 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 	slot = table_slot_create(OldHeap, NULL);
 	hslot = (BufferHeapTupleTableSlot *) slot;
+	reform_slot = MakeSingleTupleTableSlot(RelationGetDescr(OldHeap),
+										   &TTSOpsVirtual);
 
 	/*
 	 * Scan through the OldHeap, either in OldIndex order or sequentially;
@@ -875,11 +868,9 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 			int64		ct_val[2];
 
 			if (!concurrent)
-				reform_and_rewrite_tuple(tuple, OldHeap, NewHeap,
-										 values, isnull, rwstate);
+				reform_and_rewrite_tuple(slot, reform_slot, rwstate);
 			else
-				heap_insert_for_repack(tuple, OldHeap, NewHeap,
-									   values, isnull, bistate);
+				heap_insert_for_repack(NewHeap, slot, reform_slot, bistate);
 
 			/*
 			 * In indexscan mode and also VACUUM FULL, report increase in
@@ -895,8 +886,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		index_endscan(indexScan);
 	if (tableScan != NULL)
 		table_endscan(tableScan);
-	if (slot)
-		ExecDropSingleTupleTableSlot(slot);
+	ExecDropSingleTupleTableSlot(slot);
 
 	/*
 	 * In scan-and-sort mode, complete the sort, then read out all live tuples
@@ -916,6 +906,8 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
 									 PROGRESS_REPACK_PHASE_WRITE_NEW_HEAP);
 
+		slot = MakeSingleTupleTableSlot(RelationGetDescr(OldHeap),
+										&TTSOpsHeapTuple);
 		for (;;)
 		{
 			HeapTuple	tuple;
@@ -926,33 +918,36 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 			if (tuple == NULL)
 				break;
 
+			/*
+			 * XXX Ideally we should use tuplesort_gettupleslot() above, but
+			 * it retrieves minimal tuples and tuplesort_puttupleslot() cannot
+			 * get the tuple descriptor from tuplesort created by
+			 * tuplesort_begin_cluster().
+			 */
+			ExecStoreHeapTuple(tuple, slot, false);
+
 			n_tuples += 1;
 			if (!concurrent)
-				reform_and_rewrite_tuple(tuple,
-										 OldHeap, NewHeap,
-										 values, isnull,
-										 rwstate);
+				reform_and_rewrite_tuple(slot, reform_slot, rwstate);
 			else
-				heap_insert_for_repack(tuple, OldHeap, NewHeap,
-									   values, isnull, bistate);
+				heap_insert_for_repack(NewHeap, slot, reform_slot, bistate);
 
 			/* Report n_tuples */
 			pgstat_progress_update_param(PROGRESS_REPACK_HEAP_TUPLES_INSERTED,
 										 n_tuples);
 		}
 
+		ExecDropSingleTupleTableSlot(slot);
 		tuplesort_end(tuplesort);
 	}
 
+	ExecDropSingleTupleTableSlot(reform_slot);
+
 	/* Write out any remaining tuples, and fsync if needed */
 	if (rwstate)
 		end_heap_rewrite(rwstate);
 	if (bistate)
 		FreeBulkInsertState(bistate);
-
-	/* Clean up */
-	pfree(values);
-	pfree(isnull);
 }
 
 /*
@@ -2340,21 +2335,45 @@ heapam_scan_sample_next_tuple(TableScanDesc scan, SampleScanState *scanstate,
  * currently only known to happen as an after-effect of ALTER TABLE
  * SET WITHOUT OIDS.
  *
- * So, we must reconstruct the tuple from component Datums.
+ * So, we must reconstruct the tuple from component Datums. 'reform' slot
+ * is a workspace for this reconstruction.
  */
 static void
-reform_and_rewrite_tuple(HeapTuple tuple,
-						 Relation OldHeap, Relation NewHeap,
-						 Datum *values, bool *isnull, RewriteState rwstate)
+reform_and_rewrite_tuple(TupleTableSlot *src, TupleTableSlot *reform,
+						 RewriteState rwstate)
 {
-	HeapTuple	newtuple;
+	HeapTuple	tuple,
+				newtuple;
+	bool		shouldFree,
+				shouldFreeNew;
 
-	newtuple = reform_tuple(tuple, OldHeap, NewHeap, values, isnull);
+	/*
+	 * The old tuple will not be modified, so do not request materialization.
+	 * (A copy can be created for specific slot type though, e.g.
+	 * TTSOpsMinimalTuple.)
+	 */
+	tuple = ExecFetchSlotHeapTuple(src, false, &shouldFree);
+	if (tuple_needs_reform(tuple, src->tts_tupleDescriptor))
+	{
+		clear_dropped_attributes(tuple, reform);
+
+		/* No need to materialize, copy will be created anyway. */
+		Assert(TTS_IS_VIRTUAL(reform));
+		newtuple = ExecFetchSlotHeapTuple(reform, false, &shouldFreeNew);
+	}
+	else
+	{
+		newtuple = heap_copytuple(tuple);
+		shouldFreeNew = true;
+	}
 
 	/* The heap rewrite module does the rest */
 	rewrite_heap_tuple(rwstate, tuple, newtuple);
 
-	heap_freetuple(newtuple);
+	if (shouldFree)
+		heap_freetuple(tuple);
+	if (shouldFreeNew)
+		heap_freetuple(newtuple);
 }
 
 /*
@@ -2366,6 +2385,9 @@ reform_and_rewrite_tuple(HeapTuple tuple,
  * information). Thus we must use heap_insert() both during the
  * catch-up and here.
  *
+ * 'reform' is a slot to use for tuple "reforming", typically to get set
+ * values of dropped columns to NULL.
+ *
  * We pass the NO_LOGICAL flag to heap_insert() in order to skip logical
  * decoding: as soon as REPACK CONCURRENTLY swaps the relation files, it drops
  * this relation, so no logical replication subscription should need the data.
@@ -2374,76 +2396,75 @@ reform_and_rewrite_tuple(HeapTuple tuple,
  * case.
  */
 static void
-heap_insert_for_repack(HeapTuple tuple, Relation OldHeap, Relation NewHeap,
-					   Datum *values, bool *isnull, BulkInsertState bistate)
+heap_insert_for_repack(Relation rel, TupleTableSlot *src,
+					   TupleTableSlot *reform, BulkInsertStateData *bistate)
 {
-	HeapTuple	newtuple;
+	HeapTuple	tuple;
+	bool		shouldFree;
+	TupleTableSlot *slot;
 
-	newtuple = reform_tuple(tuple, OldHeap, NewHeap, values, isnull);
+	tuple = ExecFetchSlotHeapTuple(src, false, &shouldFree);
+	if (tuple_needs_reform(tuple, src->tts_tupleDescriptor))
+	{
+		clear_dropped_attributes(tuple, reform);
+		slot = reform;
+	}
+	else
+		slot = src;
 
-	heap_insert(NewHeap, newtuple, GetCurrentCommandId(true),
-				HEAP_INSERT_NO_LOGICAL, bistate);
+	/*
+	 * clear_dropped_attributes() should have deformed the tuple, so nothing
+	 * should depend on it now.
+	 */
+	if (shouldFree)
+		heap_freetuple(tuple);
 
-	heap_freetuple(newtuple);
+	table_tuple_insert(rel, slot, GetCurrentCommandId(true),
+					   TABLE_INSERT_NO_LOGICAL, bistate);
 }
 
-/*
- * Subroutine for reform_and_rewrite_tuple and heap_insert_for_repack.
- *
- * Deform the given tuple, set values of dropped columns to NULL, and fill in
- * any values from attmissingval; then form a new tuple and return it.  If no
- * attributes need to be changed, a copy of the original tuple is returned.
- * Caller is responsible for freeing the returned tuple.
- *
- * XXX this coding assumes that both relations have the same tupledesc.
- */
-static HeapTuple
-reform_tuple(HeapTuple tuple, Relation OldHeap, Relation NewHeap,
-			 Datum *values, bool *isnull)
+static bool
+tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc)
 {
-	TupleDesc	oldTupDesc = RelationGetDescr(OldHeap);
-	TupleDesc	newTupDesc = RelationGetDescr(NewHeap);
-	bool		needs_reform = false;
-
 	/*
 	 * A short tuple might require values from attmissing val, so activate the
 	 * coding unconditionally in that case.  The value might legitimally be
 	 * NULL otherwise, so this is slightly wasteful, but it probably beats
 	 * having to test each attribute for presence of attmissingval each time.
 	 */
-	if (HeapTupleHeaderGetNatts(tuple->t_data) < newTupDesc->natts)
-		needs_reform = true;
+	if (HeapTupleHeaderGetNatts(tuple->t_data) < tupDesc->natts)
+		return true;
 
-	/*
-	 * If the column has been dropped but a value is still present, we can
-	 * optimize storage now by getting rid of it.
-	 */
-	if (!needs_reform)
+	/* Does it have dropped attributes? */
+	for (int i = 0; i < tupDesc->natts; i++)
 	{
-		for (int i = 0; i < newTupDesc->natts; i++)
-		{
-			if (TupleDescCompactAttr(newTupDesc, i)->attisdropped &&
-				!heap_attisnull(tuple, i + 1, newTupDesc))
-			{
-				needs_reform = true;
-				break;
-			}
-		}
+		if (TupleDescCompactAttr(tupDesc, i)->attisdropped &&
+			!heap_attisnull(tuple, i + 1, tupDesc))
+			return true;
 	}
 
-	/* Skip work if no changes are needed */
-	if (!needs_reform)
-		return heap_copytuple(tuple);
+	return false;
+}
 
-	heap_deform_tuple(tuple, oldTupDesc, values, isnull);
+/*
+ * Subroutine for reform_and_rewrite_tuple and heap_insert_for_repack.
+ *
+ * Set values of dropped columns to NULL,
+ */
+static void
+clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
+{
+	TupleDesc	tupDesc = reform->tts_tupleDescriptor;
+
+	/* Assuming 'reform' is virtual, this deforms the tuple. */
+	Assert(TTS_IS_VIRTUAL(reform));
+	ExecForceStoreHeapTuple(tuple, reform, false);
 
-	for (int i = 0; i < newTupDesc->natts; i++)
+	for (int i = 0; i < tupDesc->natts; i++)
 	{
-		if (TupleDescCompactAttr(newTupDesc, i)->attisdropped)
-			isnull[i] = true;
+		if (TupleDescCompactAttr(tupDesc, i)->attisdropped)
+			reform->tts_isnull[i] = true;
 	}
-
-	return heap_form_tuple(newTupDesc, values, isnull);
 }
 
 /*
-- 
2.52.0

  [text/x-diff] v02-0002-Move-functions-to-repack.c.patch (8.4K, ../108776.1784105248@localhost/3-v02-0002-Move-functions-to-repack.c.patch)
  download | inline diff:
From e7d0670c2bf49cea1772e4fd2bb0fbfad0bf7adc Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 15 Jul 2026 10:37:11 +0200
Subject: [PATCH 2/8] Move functions to repack.c.

In the next patches, the functions will be called from other modules. Use a
separate diff for the move so that the following diffs are a bit easier to
read.
---
 src/backend/access/heap/heapam_handler.c | 97 +-----------------------
 src/backend/commands/repack.c            | 91 ++++++++++++++++++++++
 src/include/commands/repack.h            |  7 ++
 3 files changed, 99 insertions(+), 96 deletions(-)

diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index 4db8a16f691..97f850ebf73 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -34,6 +34,7 @@
 #include "catalog/storage.h"
 #include "catalog/storage_xlog.h"
 #include "commands/progress.h"
+#include "commands/repack.h"
 #include "executor/executor.h"
 #include "miscadmin.h"
 #include "pgstat.h"
@@ -50,11 +51,6 @@
 
 static void reform_and_rewrite_tuple(TupleTableSlot *src, TupleTableSlot *reform,
 									 RewriteState rwstate);
-static void heap_insert_for_repack(Relation rel, TupleTableSlot *src,
-								   TupleTableSlot *reform,
-								   BulkInsertStateData *bistate);
-static bool tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc);
-static void clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform);
 
 static bool SampleHeapTupleVisible(TableScanDesc scan, Buffer buffer,
 								   HeapTuple tuple,
@@ -2376,97 +2372,6 @@ reform_and_rewrite_tuple(TupleTableSlot *src, TupleTableSlot *reform,
 		heap_freetuple(newtuple);
 }
 
-/*
- * Insert tuple when processing REPACK CONCURRENTLY.
- *
- * rewriteheap.c is not used in the CONCURRENTLY case because it'd be
- * difficult to do the same in the catch-up phase (as the logical
- * decoding does not provide us with sufficient visibility
- * information). Thus we must use heap_insert() both during the
- * catch-up and here.
- *
- * 'reform' is a slot to use for tuple "reforming", typically to get set
- * values of dropped columns to NULL.
- *
- * We pass the NO_LOGICAL flag to heap_insert() in order to skip logical
- * decoding: as soon as REPACK CONCURRENTLY swaps the relation files, it drops
- * this relation, so no logical replication subscription should need the data.
- *
- * BulkInsertState is used because many tuples are inserted in the typical
- * case.
- */
-static void
-heap_insert_for_repack(Relation rel, TupleTableSlot *src,
-					   TupleTableSlot *reform, BulkInsertStateData *bistate)
-{
-	HeapTuple	tuple;
-	bool		shouldFree;
-	TupleTableSlot *slot;
-
-	tuple = ExecFetchSlotHeapTuple(src, false, &shouldFree);
-	if (tuple_needs_reform(tuple, src->tts_tupleDescriptor))
-	{
-		clear_dropped_attributes(tuple, reform);
-		slot = reform;
-	}
-	else
-		slot = src;
-
-	/*
-	 * clear_dropped_attributes() should have deformed the tuple, so nothing
-	 * should depend on it now.
-	 */
-	if (shouldFree)
-		heap_freetuple(tuple);
-
-	table_tuple_insert(rel, slot, GetCurrentCommandId(true),
-					   TABLE_INSERT_NO_LOGICAL, bistate);
-}
-
-static bool
-tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc)
-{
-	/*
-	 * A short tuple might require values from attmissing val, so activate the
-	 * coding unconditionally in that case.  The value might legitimally be
-	 * NULL otherwise, so this is slightly wasteful, but it probably beats
-	 * having to test each attribute for presence of attmissingval each time.
-	 */
-	if (HeapTupleHeaderGetNatts(tuple->t_data) < tupDesc->natts)
-		return true;
-
-	/* Does it have dropped attributes? */
-	for (int i = 0; i < tupDesc->natts; i++)
-	{
-		if (TupleDescCompactAttr(tupDesc, i)->attisdropped &&
-			!heap_attisnull(tuple, i + 1, tupDesc))
-			return true;
-	}
-
-	return false;
-}
-
-/*
- * Subroutine for reform_and_rewrite_tuple and heap_insert_for_repack.
- *
- * Set values of dropped columns to NULL,
- */
-static void
-clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
-{
-	TupleDesc	tupDesc = reform->tts_tupleDescriptor;
-
-	/* Assuming 'reform' is virtual, this deforms the tuple. */
-	Assert(TTS_IS_VIRTUAL(reform));
-	ExecForceStoreHeapTuple(tuple, reform, false);
-
-	for (int i = 0; i < tupDesc->natts; i++)
-	{
-		if (TupleDescCompactAttr(tupDesc, i)->attisdropped)
-			reform->tts_isnull[i] = true;
-	}
-}
-
 /*
  * Check visibility of the tuple.
  */
diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 02883fe34a4..19927c15bf3 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -1288,6 +1288,97 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
 	return OIDNewHeap;
 }
 
+/*
+ * Insert tuple when processing REPACK CONCURRENTLY.
+ *
+ * rewriteheap.c is not used in the CONCURRENTLY case because it'd be
+ * difficult to do the same in the catch-up phase (as the logical
+ * decoding does not provide us with sufficient visibility
+ * information). Thus we must use heap_insert() both during the
+ * catch-up and here.
+ *
+ * 'reform' is a slot to use for tuple "reforming", typically to get set
+ * values of dropped columns to NULL.
+ *
+ * We pass the NO_LOGICAL flag to heap_insert() in order to skip logical
+ * decoding: as soon as REPACK CONCURRENTLY swaps the relation files, it drops
+ * this relation, so no logical replication subscription should need the data.
+ *
+ * BulkInsertState is used because many tuples are inserted in the typical
+ * case.
+ */
+void
+heap_insert_for_repack(Relation rel, TupleTableSlot *src,
+					   TupleTableSlot *reform, BulkInsertStateData *bistate)
+{
+	HeapTuple	tuple;
+	bool		shouldFree;
+	TupleTableSlot *slot;
+
+	tuple = ExecFetchSlotHeapTuple(src, false, &shouldFree);
+	if (tuple_needs_reform(tuple, src->tts_tupleDescriptor))
+	{
+		clear_dropped_attributes(tuple, reform);
+		slot = reform;
+	}
+	else
+		slot = src;
+
+	/*
+	 * clear_dropped_attributes() should have deformed the tuple, so nothing
+	 * should depend on it now.
+	 */
+	if (shouldFree)
+		heap_freetuple(tuple);
+
+	table_tuple_insert(rel, slot, GetCurrentCommandId(true),
+					   TABLE_INSERT_NO_LOGICAL, bistate);
+}
+
+bool
+tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc)
+{
+	/*
+	 * A short tuple might require values from attmissing val, so activate the
+	 * coding unconditionally in that case.  The value might legitimally be
+	 * NULL otherwise, so this is slightly wasteful, but it probably beats
+	 * having to test each attribute for presence of attmissingval each time.
+	 */
+	if (HeapTupleHeaderGetNatts(tuple->t_data) < tupDesc->natts)
+		return true;
+
+	/* Does it have dropped attributes? */
+	for (int i = 0; i < tupDesc->natts; i++)
+	{
+		if (TupleDescCompactAttr(tupDesc, i)->attisdropped &&
+			!heap_attisnull(tuple, i + 1, tupDesc))
+			return true;
+	}
+
+	return false;
+}
+
+/*
+ * Subroutine for reform_and_rewrite_tuple and heap_insert_for_repack.
+ *
+ * Set values of dropped columns to NULL,
+ */
+void
+clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
+{
+	TupleDesc	tupDesc = reform->tts_tupleDescriptor;
+
+	/* Assuming 'reform' is virtual, this deforms the tuple. */
+	Assert(TTS_IS_VIRTUAL(reform));
+	ExecForceStoreHeapTuple(tuple, reform, false);
+
+	for (int i = 0; i < tupDesc->natts; i++)
+	{
+		if (TupleDescCompactAttr(tupDesc, i)->attisdropped)
+			reform->tts_isnull[i] = true;
+	}
+}
+
 /*
  * Do the physical copying of table data.
  *
diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h
index 45e5440a311..27105c10591 100644
--- a/src/include/commands/repack.h
+++ b/src/include/commands/repack.h
@@ -15,6 +15,8 @@
 
 #include <signal.h>
 
+#include "access/hio.h"
+#include "nodes/execnodes.h"
 #include "nodes/parsenodes.h"
 #include "parser/parse_node.h"
 #include "storage/lockdefs.h"
@@ -48,6 +50,11 @@ extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal);
 
 extern Oid	make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
 						  char relpersistence, LOCKMODE lockmode);
+extern void heap_insert_for_repack(Relation rel, TupleTableSlot *src,
+								   TupleTableSlot *reform,
+								   BulkInsertStateData *bistate);
+extern bool tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc);
+extern void clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform);
 extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 							 bool is_system_catalog,
 							 bool swap_toast_by_content,
-- 
2.52.0

  [text/x-diff] v02-0003-Introduce-RepackDest-structure.patch (18.8K, ../108776.1784105248@localhost/4-v02-0003-Introduce-RepackDest-structure.patch)
  download | inline diff:
From f13c7cb3d8780f2f2ecb8b1c8b8d8808fb71a95f Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 15 Jul 2026 10:37:12 +0200
Subject: [PATCH 3/8] Introduce RepackDest structure.

This is for ChangeContext to handle insertions into two relations: besides the
new relation (whose file will eventually be used by the REPACKed relation), an
"auxiliary relation" is needed sometimes, in order to get the data sorted. The
concept is introduced and explained later in the patch series.
---
 src/backend/commands/repack.c    | 223 ++++++++++++++++---------------
 src/include/commands/repack.h    |  48 +++++++
 src/tools/pgindent/typedefs.list |   1 +
 3 files changed, 161 insertions(+), 111 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 19927c15bf3..392332b4b2b 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -95,40 +95,6 @@ typedef struct
 	Oid			indexOid;
 } RelToCluster;
 
-/*
- * The first file exported by the decoding worker must contain a snapshot, the
- * following ones contain the data changes.
- */
-#define WORKER_FILE_SNAPSHOT	0
-
-/*
- * Information needed to apply concurrent data changes.
- */
-typedef struct ChangeContext
-{
-	/* The relation the changes are applied to. */
-	Relation	cc_rel;
-
-	/* Needed to update indexes of cc_rel. */
-	ResultRelInfo *cc_rri;
-	EState	   *cc_estate;
-
-	/*
-	 * Existing tuples to UPDATE and DELETE are located via this index. We
-	 * keep the scankey in partially initialized state to avoid repeated work.
-	 * sk_argument is completed on the fly.
-	 */
-	Relation	cc_ident_index;
-	ScanKey		cc_ident_key;
-	int			cc_ident_key_nentries;
-
-	/* The latest column we need to deform to have the tuple identity */
-	AttrNumber	cc_last_key_attno;
-
-	/* Sequential number of the file containing the changes. */
-	int			cc_file_seq;
-} ChangeContext;
-
 /*
  * Backend-local information to control the decoding worker.
  */
@@ -175,22 +141,19 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd,
 static bool repack_is_permitted_for_relation(RepackCommand cmd,
 											 Oid relid, Oid userid);
 
-static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt);
-static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot,
-									ChangeContext *chgcxt);
-static void apply_concurrent_update(Relation rel, TupleTableSlot *spilled_tuple,
-									TupleTableSlot *ondisk_tuple,
-									ChangeContext *chgcxt);
+static void apply_concurrent_changes(ChangeContext *chgcxt);
+static void apply_concurrent_insert(RepackDest *dest, TupleTableSlot *slot);
+static void apply_concurrent_update(RepackDest *dest,
+									TupleTableSlot *spilled_tuple,
+									TupleTableSlot *ondisk_tuple);
 static void apply_concurrent_delete(Relation rel, TupleTableSlot *slot);
 static void restore_tuple(BufFile *file, Relation relation,
 						  TupleTableSlot *slot);
 static void adjust_toast_pointers(Relation relation, TupleTableSlot *dest,
 								  TupleTableSlot *src);
-static bool find_target_tuple(Relation rel, ChangeContext *chgcxt,
-							  TupleTableSlot *locator,
+static bool find_target_tuple(RepackDest *dest, TupleTableSlot *locator,
 							  TupleTableSlot *retrieved);
-static bool identity_key_equal(ChangeContext *chgcxt,
-							   TupleTableSlot *locator,
+static bool identity_key_equal(RepackDest *dest, TupleTableSlot *locator,
 							   TupleTableSlot *candidate);
 static void process_concurrent_changes(XLogRecPtr end_of_wal,
 									   ChangeContext *chgcxt,
@@ -199,6 +162,9 @@ static void initialize_change_context(ChangeContext *chgcxt,
 									  Relation relation,
 									  Oid ident_index_id);
 static void release_change_context(ChangeContext *chgcxt);
+static void initialize_change_dest(RepackDest *dest, Relation relation,
+								   Oid ident_index_id);
+static void release_change_dest(RepackDest *dest);
 static void rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 											   Oid identIdx,
 											   TransactionId frozenXid,
@@ -2614,18 +2580,31 @@ RepackCommandAsString(RepackCommand cmd)
 }
 
 /*
- * Apply all the changes stored in 'file'.
+ * Apply all the changes provided by decoding worker.
  */
 static void
-apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt)
+apply_concurrent_changes(ChangeContext *chgcxt)
 {
 	ConcurrentChangeKind kind = '\0';
-	Relation	rel = chgcxt->cc_rel;
+	RepackDest *dest;
+	Relation	rel;
 	TupleTableSlot *spilled_tuple;
 	TupleTableSlot *old_update_tuple;
 	TupleTableSlot *ondisk_tuple;
 	bool		have_old_tuple = false;
 	MemoryContext oldcxt;
+	DecodingWorkerShared *shared;
+	char		fname[MAXPGPATH];
+	BufFile    *file;
+
+	dest = &chgcxt->cc_dest;
+	rel = dest->rel;
+
+	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
+
+	/* Open the file containing the changes. */
+	DecodingWorkerFileName(fname, shared->relid, chgcxt->cc_file_seq);
+	file = BufFileOpenFileSet(&shared->sfs.fs, fname, O_RDONLY, false);
 
 	spilled_tuple = MakeSingleTupleTableSlot(RelationGetDescr(rel),
 											 &TTSOpsVirtual);
@@ -2634,7 +2613,7 @@ apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt)
 	old_update_tuple = MakeSingleTupleTableSlot(RelationGetDescr(rel),
 												&TTSOpsVirtual);
 
-	oldcxt = MemoryContextSwitchTo(GetPerTupleMemoryContext(chgcxt->cc_estate));
+	oldcxt = MemoryContextSwitchTo(GetPerTupleMemoryContext(dest->estate));
 
 	while (true)
 	{
@@ -2683,14 +2662,14 @@ apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt)
 
 		if (kind == CHANGE_INSERT)
 		{
-			apply_concurrent_insert(rel, spilled_tuple, chgcxt);
+			apply_concurrent_insert(dest, spilled_tuple);
 		}
 		else if (kind == CHANGE_DELETE)
 		{
 			bool		found;
 
 			/* Find the tuple to be deleted */
-			found = find_target_tuple(rel, chgcxt, spilled_tuple, ondisk_tuple);
+			found = find_target_tuple(dest, spilled_tuple, ondisk_tuple);
 			if (!found)
 				elog(ERROR, "could not find target tuple");
 			apply_concurrent_delete(rel, ondisk_tuple);
@@ -2706,7 +2685,7 @@ apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt)
 				key = spilled_tuple;
 
 			/* Find the tuple to be updated or deleted. */
-			found = find_target_tuple(rel, chgcxt, key, ondisk_tuple);
+			found = find_target_tuple(dest, key, ondisk_tuple);
 			if (!found)
 				elog(ERROR, "could not find target tuple");
 
@@ -2719,7 +2698,7 @@ apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt)
 			 */
 			adjust_toast_pointers(rel, spilled_tuple, ondisk_tuple);
 
-			apply_concurrent_update(rel, spilled_tuple, ondisk_tuple, chgcxt);
+			apply_concurrent_update(dest, spilled_tuple, ondisk_tuple);
 
 			ExecClearTuple(old_update_tuple);
 			have_old_tuple = false;
@@ -2727,7 +2706,7 @@ apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt)
 		else
 			elog(ERROR, "unrecognized kind of change: %d", kind);
 
-		ResetPerTupleExprContext(chgcxt->cc_estate);
+		ResetPerTupleExprContext(dest->estate);
 	}
 
 	/* Cleanup. */
@@ -2736,6 +2715,8 @@ apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt)
 	ExecDropSingleTupleTableSlot(old_update_tuple);
 
 	MemoryContextSwitchTo(oldcxt);
+
+	BufFileClose(file);
 }
 
 /*
@@ -2743,16 +2724,15 @@ apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt)
  * table.
  */
 static void
-apply_concurrent_insert(Relation rel, TupleTableSlot *slot,
-						ChangeContext *chgcxt)
+apply_concurrent_insert(RepackDest *dest, TupleTableSlot *slot)
 {
 	/* Put the tuple in the table, but make sure it won't be decoded */
-	table_tuple_insert(rel, slot, GetCurrentCommandId(true),
+	table_tuple_insert(dest->rel, slot, GetCurrentCommandId(true),
 					   TABLE_INSERT_NO_LOGICAL, NULL);
 
 	/* Update indexes with this new tuple. */
-	ExecInsertIndexTuples(chgcxt->cc_rri,
-						  chgcxt->cc_estate,
+	ExecInsertIndexTuples(dest->rri,
+						  dest->estate,
 						  0,
 						  slot,
 						  NIL, NULL);
@@ -2764,10 +2744,10 @@ apply_concurrent_insert(Relation rel, TupleTableSlot *slot,
  * table.
  */
 static void
-apply_concurrent_update(Relation rel, TupleTableSlot *spilled_tuple,
-						TupleTableSlot *ondisk_tuple,
-						ChangeContext *chgcxt)
+apply_concurrent_update(RepackDest *dest, TupleTableSlot *spilled_tuple,
+						TupleTableSlot *ondisk_tuple)
 {
+	Relation	rel = dest->rel;
 	LockTupleMode lockmode;
 	TM_FailureData tmfd;
 	TU_UpdateIndexes update_indexes;
@@ -2795,8 +2775,8 @@ apply_concurrent_update(Relation rel, TupleTableSlot *spilled_tuple,
 
 		if (update_indexes == TU_Summarizing)
 			flags |= EIIT_ONLY_SUMMARIZING;
-		ExecInsertIndexTuples(chgcxt->cc_rri,
-							  chgcxt->cc_estate,
+		ExecInsertIndexTuples(dest->rri,
+							  dest->estate,
 							  flags,
 							  spilled_tuple,
 							  NIL, NULL);
@@ -2948,10 +2928,11 @@ adjust_toast_pointers(Relation relation, TupleTableSlot *dest, TupleTableSlot *s
  * not found, return false.
  */
 static bool
-find_target_tuple(Relation rel, ChangeContext *chgcxt, TupleTableSlot *locator,
+find_target_tuple(RepackDest *dest, TupleTableSlot *locator,
 				  TupleTableSlot *retrieved)
 {
-	Form_pg_index idx = chgcxt->cc_ident_index->rd_index;
+	Relation	rel = dest->rel;
+	Form_pg_index idx = dest->ident_index->rd_index;
 	IndexScanDesc scan;
 	bool		retval = false;
 
@@ -2962,9 +2943,9 @@ find_target_tuple(Relation rel, ChangeContext *chgcxt, TupleTableSlot *locator,
 	 *
 	 * Use the incoming tuple to finalize the scan key.
 	 */
-	for (int i = 0; i < chgcxt->cc_ident_key_nentries; i++)
+	for (int i = 0; i < dest->ident_key_nentries; i++)
 	{
-		ScanKey		entry = &chgcxt->cc_ident_key[i];
+		ScanKey		entry = &dest->ident_key[i];
 		AttrNumber	attno = idx->indkey.values[i];
 
 		entry->sk_argument = locator->tts_values[attno - 1];
@@ -2972,13 +2953,13 @@ find_target_tuple(Relation rel, ChangeContext *chgcxt, TupleTableSlot *locator,
 	}
 
 	/* XXX no instrumentation for now */
-	scan = index_beginscan(rel, chgcxt->cc_ident_index, GetActiveSnapshot(),
-						   NULL, chgcxt->cc_ident_key_nentries, 0, 0);
-	index_rescan(scan, chgcxt->cc_ident_key, chgcxt->cc_ident_key_nentries, NULL, 0);
+	scan = index_beginscan(rel, dest->ident_index, GetActiveSnapshot(),
+						   NULL, dest->ident_key_nentries, 0, 0);
+	index_rescan(scan, dest->ident_key, dest->ident_key_nentries, NULL, 0);
 	while (index_getnext_slot(scan, ForwardScanDirection, retrieved))
 	{
 		/* Be wary of temporal constraints */
-		if (scan->xs_recheck && !identity_key_equal(chgcxt, locator, retrieved))
+		if (scan->xs_recheck && !identity_key_equal(dest, locator, retrieved))
 		{
 			CHECK_FOR_INTERRUPTS();
 			continue;
@@ -3001,16 +2982,16 @@ find_target_tuple(Relation rel, ChangeContext *chgcxt, TupleTableSlot *locator,
  * used for temporal constraints.
  */
 static bool
-identity_key_equal(ChangeContext *chgcxt, TupleTableSlot *locator,
+identity_key_equal(RepackDest *dest, TupleTableSlot *locator,
 				   TupleTableSlot *candidate)
 {
-	slot_getsomeattrs(locator, chgcxt->cc_last_key_attno);
-	slot_getsomeattrs(candidate, chgcxt->cc_last_key_attno);
+	slot_getsomeattrs(locator, dest->last_key_attno);
+	slot_getsomeattrs(candidate, dest->last_key_attno);
 
-	for (int i = 0; i < chgcxt->cc_ident_key_nentries; i++)
+	for (int i = 0; i < dest->ident_key_nentries; i++)
 	{
-		ScanKey		entry = &chgcxt->cc_ident_key[i];
-		AttrNumber	attno = chgcxt->cc_ident_index->rd_index->indkey.values[i];
+		ScanKey		entry = &dest->ident_key[i];
+		AttrNumber	attno = dest->ident_index->rd_index->indkey.values[i];
 
 		Assert(attno > 0);
 
@@ -3081,7 +3062,7 @@ process_concurrent_changes(XLogRecPtr end_of_wal, ChangeContext *chgcxt, bool do
 	/* Open the file. */
 	DecodingWorkerFileName(fname, shared->relid, chgcxt->cc_file_seq);
 	file = BufFileOpenFileSet(&shared->sfs.fs, fname, O_RDONLY, false);
-	apply_concurrent_changes(file, chgcxt);
+	apply_concurrent_changes(chgcxt);
 
 	BufFileClose(file);
 
@@ -3097,10 +3078,34 @@ static void
 initialize_change_context(ChangeContext *chgcxt,
 						  Relation relation, Oid ident_index_id)
 {
-	chgcxt->cc_rel = relation;
+	initialize_change_dest(&chgcxt->cc_dest, relation, ident_index_id);
+
+	chgcxt->cc_file_seq = WORKER_FILE_SNAPSHOT + 1;
+}
+
+/*
+ * Free up resources taken by a ChangeContext.
+ */
+static void
+release_change_context(ChangeContext *chgcxt)
+{
+	release_change_dest(&chgcxt->cc_dest);
+}
+
+/*
+ * Initialize the RepackDest struct for the given relation, with the given
+ * index as identity index. InvalidOid can be specified to only make the
+ * relation ready for insertions.
+ */
+static void
+initialize_change_dest(RepackDest *dest, Relation relation,
+					   Oid ident_index_id)
+{
+	dest->rel = relation;
+	dest->bistate = GetBulkInsertState();
 
 	/* Only initialize fields needed by ExecInsertIndexTuples(). */
-	chgcxt->cc_estate = CreateExecutorState();
+	dest->estate = CreateExecutorState();
 
 	/*
 	 * Set up a range table for the executor, containing our repacked table as
@@ -3149,41 +3154,41 @@ initialize_change_context(ChangeContext *chgcxt,
 		perminfo->updatedCols = updatedCols;
 
 		/* finally we can initialize the range table proper */
-		ExecInitRangeTable(chgcxt->cc_estate, list_make1(rte), perminfos,
+		ExecInitRangeTable(dest->estate, list_make1(rte), perminfos,
 						   bms_make_singleton(1));
 	}
 
 	/* Set up our ResultRelInfo to use for index updates */
-	chgcxt->cc_rri = makeNode(ResultRelInfo);
-	InitResultRelInfo(chgcxt->cc_rri, relation, 1, NULL, 0);
-	ExecOpenIndices(chgcxt->cc_rri, false);
+	dest->rri = makeNode(ResultRelInfo);
+	InitResultRelInfo(dest->rri, relation, 1, NULL, 0);
+	ExecOpenIndices(dest->rri, false);
 
 	/*
 	 * The table's relcache entry already has the relcache entry for the
 	 * identity index; find that.
 	 */
-	chgcxt->cc_ident_index = NULL;
-	for (int i = 0; i < chgcxt->cc_rri->ri_NumIndices; i++)
+	dest->ident_index = NULL;
+	for (int i = 0; i < dest->rri->ri_NumIndices; i++)
 	{
 		Relation	ind_rel;
 
-		ind_rel = chgcxt->cc_rri->ri_IndexRelationDescs[i];
+		ind_rel = dest->rri->ri_IndexRelationDescs[i];
 		if (ind_rel->rd_id == ident_index_id)
 		{
-			chgcxt->cc_ident_index = ind_rel;
+			dest->ident_index = ind_rel;
 			break;
 		}
 	}
-	if (chgcxt->cc_ident_index == NULL)
+	if (dest->ident_index == NULL)
 		elog(ERROR, "could not find identity index");
 
 	/* Set up for scanning said identity index */
 	{
 		Form_pg_index indexForm;
 
-		indexForm = chgcxt->cc_ident_index->rd_index;
-		chgcxt->cc_ident_key_nentries = indexForm->indnkeyatts;
-		chgcxt->cc_ident_key = (ScanKey) palloc_array(ScanKeyData, indexForm->indnkeyatts);
+		indexForm = dest->ident_index->rd_index;
+		dest->ident_key_nentries = indexForm->indnkeyatts;
+		dest->ident_key = (ScanKey) palloc_array(ScanKeyData, indexForm->indnkeyatts);
 		for (int i = 0; i < indexForm->indnkeyatts; i++)
 		{
 			ScanKey		entry;
@@ -3193,12 +3198,12 @@ initialize_change_context(ChangeContext *chgcxt,
 						opcode;
 			StrategyNumber eq_strategy;
 
-			entry = &chgcxt->cc_ident_key[i];
+			entry = &dest->ident_key[i];
 
-			opfamily = chgcxt->cc_ident_index->rd_opfamily[i];
-			opcintype = chgcxt->cc_ident_index->rd_opcintype[i];
+			opfamily = dest->ident_index->rd_opfamily[i];
+			opcintype = dest->ident_index->rd_opcintype[i];
 			eq_strategy = IndexAmTranslateCompareType(COMPARE_EQ,
-													  chgcxt->cc_ident_index->rd_rel->relam,
+													  dest->ident_index->rd_rel->relam,
 													  opfamily, false);
 			if (eq_strategy == InvalidStrategy)
 				elog(ERROR, "could not find equality strategy for index operator family %u for type %u",
@@ -3217,34 +3222,30 @@ initialize_change_context(ChangeContext *chgcxt,
 						i + 1,
 						eq_strategy, opcode,
 						(Datum) 0);
-			entry->sk_collation = chgcxt->cc_ident_index->rd_indcollation[i];
+			entry->sk_collation = dest->ident_index->rd_indcollation[i];
 		}
 	}
 
 	/* Determine the last column we must deform to read the identity */
-	chgcxt->cc_last_key_attno = InvalidAttrNumber;
-	for (int i = 0; i < chgcxt->cc_ident_key_nentries; i++)
+	dest->last_key_attno = InvalidAttrNumber;
+	for (int i = 0; i < dest->ident_key_nentries; i++)
 	{
-		AttrNumber	attno = chgcxt->cc_ident_index->rd_index->indkey.values[i];
+		AttrNumber	attno = dest->ident_index->rd_index->indkey.values[i];
 
 		Assert(attno > 0);
-		chgcxt->cc_last_key_attno = Max(chgcxt->cc_last_key_attno, attno);
+		dest->last_key_attno = Max(dest->last_key_attno, attno);
 	}
-
-	chgcxt->cc_file_seq = WORKER_FILE_SNAPSHOT + 1;
 }
 
-/*
- * Free up resources taken by a ChangeContext.
- */
 static void
-release_change_context(ChangeContext *chgcxt)
+release_change_dest(RepackDest *dest)
 {
-	ExecCloseIndices(chgcxt->cc_rri);
-	FreeExecutorState(chgcxt->cc_estate);
+	FreeBulkInsertState(dest->bistate);
+	ExecCloseIndices(dest->rri);
+	FreeExecutorState(dest->estate);
 	/* XXX are these pfrees necessary? */
-	pfree(chgcxt->cc_rri);
-	pfree(chgcxt->cc_ident_key);
+	pfree(dest->rri);
+	pfree(dest->ident_key);
 }
 
 /*
diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h
index 27105c10591..07f887e99f6 100644
--- a/src/include/commands/repack.h
+++ b/src/include/commands/repack.h
@@ -16,6 +16,7 @@
 #include <signal.h>
 
 #include "access/hio.h"
+#include "access/skey.h"
 #include "nodes/execnodes.h"
 #include "nodes/parsenodes.h"
 #include "parser/parse_node.h"
@@ -39,6 +40,53 @@ typedef struct ClusterParams
 
 extern PGDLLIMPORT volatile sig_atomic_t RepackMessagePending;
 
+/*
+ * Table to apply concurrent data changes to.
+ */
+typedef struct RepackDest
+{
+	/* The relation the changes are applied to. */
+	Relation	rel;
+
+	BulkInsertStateData *bistate;
+
+	/* Needed to update indexes of cc_rel. */
+	ResultRelInfo *rri;
+	EState	   *estate;
+
+	/*
+	 * Existing tuples to UPDATE and DELETE are located via this index. We
+	 * keep the scankey in partially initialized state to avoid repeated work.
+	 * sk_argument is completed on the fly.
+	 */
+	Relation	ident_index;
+	ScanKey		ident_key;
+
+	int			ident_key_nentries;
+
+	/* The latest column we need to deform to have the tuple identity */
+	AttrNumber	last_key_attno;
+} RepackDest;
+
+/*
+ * The first file exported by the decoding worker must contain a snapshot, the
+ * following ones contain the data changes.
+ */
+#define WORKER_FILE_SNAPSHOT	0
+
+/*
+ * Information needed to apply concurrent data changes.
+ *
+ * XXX Now that it's in *.h file, rename to RepackChangeContext?
+ */
+typedef struct ChangeContext
+{
+	/* The destination table. */
+	RepackDest	cc_dest;
+
+	/* Sequential number of the file containing the changes. */
+	int			cc_file_seq;
+} ChangeContext;
 
 extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel);
 
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 56c1f997f88..25ac7079099 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -2664,6 +2664,7 @@ ReorderBufferTupleCidKey
 ReorderBufferUpdateProgressTxnCB
 ReorderTuple
 RepackCommand
+RepackDest
 RepackDecodingState
 RepackStmt
 ReparameterizeForeignPathByChild_function
-- 
2.52.0

  [text/plain] v02-0004-Use-multiple-snapshots-to-copy-the-data.patch (111.8K, ../108776.1784105248@localhost/5-v02-0004-Use-multiple-snapshots-to-copy-the-data.patch)
  download | inline diff:
From 56e791da5b1d17d0798bfc71da7b4ef95836e8e6 Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 15 Jul 2026 10:37:12 +0200
Subject: [PATCH 4/8] Use multiple snapshots to copy the data.

REPACK (CONCURRENTLY) does not prevent applications from using the table that
is being processed, however it can prevent the xmin horizon from advancing and
thus restrict VACUUM for the whole database. This patch adds the ability to
use particular snapshot only for certain range of pages. Each time that range
is processed, a new snapshot is built, which supposedly has its xmin higher
than the previous snapshot.

Note that we still use the same XID throughout the REPACK execution, and that
also prevents the xmin horizon from advancing. The following part of this
series will fix the problem.

To use multiple snapshots, the data copying works as follows:

  1. Have the logical decoding system build a snapshot S0 for range R0 at
     LSN0. This snapshot sees all the data changes whose commit records have
     LSN < LSN0.

  2. Copy the pages in that range to the new relation. The changes not visible
     to the snapshot (because their transactions are still running from the
     POV of the snapshot) will appear in the output of the logical decoding
     system as soon as their commit records are decoded.

  3. Perform logical decoding of all changes we find in WAL for the table
     we're repacking, but only apply those that affect the range R0 in the old
     relation. (Naturally, we cannot apply ones that belong to other pages
     because it's impossible to UPDATE / DELETE a row in the new relation if
     it hasn't been copied yet.) Then consider LSN1 to be the position of the
     end of the last WAL record decoded.

  4. Build a new snapshot S1 at position LSN1, i.e. one that sees all the data
     whose commit records are at WAL positions < LSN1. Use this snapshot to
     copy the range of pages R1.

  5. Perform logical decoding like in step 3, but out of this next set, only
     apply changes belonging to ranges R0 *and* R1 in the old table.

  6. etc

Special attention needs to be paid to UPDATES that span page ranges. For
example, if the old tuple is in range R0, but the new tuple is in R1, and R1
hasn't been copied yet, we only DELETE the old version from the new
relation. The new version will be handled during processing of range R1. The
snapshot S1 will be based on WAL position following that UPDATE, so it'll see
the new tuple if its transaction's commit record is at WAL position lower than
the position where we built the snapshot. On the other hand, if the commit
record appears at higher position than the that of the snapshot, the insertion
of the new tuple will be decoded and replayed later (after the copying of
range R1 has completed).

Likewise, if the old tuple is in range R1 (not yet copied) but the new tuple
is in R0, we only perform INSERT on the new relation. The deletion of the old
version will either be visible to the snapshot S1 (i.e. the snapshot won't see
the old version), or replayed later.

Due to these cross-range UPDATEs, we must apply the changes pertaining to
given range before processing of the next range starts. Specifically, if
UPDATE becomes DELETE for specific range, that DELETE must be replayed soon
enough so that we don't see both old and new tuple when building the identity
index. The problem is that if the UPDATE does not change the identity key,
we'd end up with duplicate key values.

Even if the USING INDEX clause is specified, a sequential scan is used to
retrieve the tuples from the old relation: the approach described above
requires that the tuples are in CTID order. For sorting we use a regular table
("auxiliary table"), on which we create the clustering index and scan it. The
scan output is inserted into the new relation. Tuplesort is not appropriate
here because it has no identity index, so it's not possible to apply the
decoded changes to it "eagerly", as explained above.

A new GUC repack_snapshot_after can be used to set the number of pages per
snapshot. It's currently classified as DEVELOPER_OPTIONS and may be replaced
by a constant after enough evaluation is done.
---
 src/backend/access/heap/heapam_handler.c      | 238 ++++-
 src/backend/commands/repack.c                 | 876 +++++++++++++-----
 src/backend/commands/repack_worker.c          |  97 +-
 src/backend/replication/logical/decode.c      |  83 +-
 src/backend/replication/logical/logical.c     |  30 +-
 .../replication/logical/reorderbuffer.c       |  40 +
 src/backend/replication/logical/snapbuild.c   |  30 +-
 src/backend/replication/pgrepack/pgrepack.c   |  24 +-
 src/backend/utils/misc/guc_parameters.dat     |  10 +
 src/backend/utils/misc/guc_tables.c           |   1 +
 src/include/access/tableam.h                  |  14 +-
 src/include/commands/repack.h                 |  64 +-
 src/include/commands/repack_internal.h        |  13 +-
 src/include/replication/logical.h             |   2 +-
 src/include/replication/reorderbuffer.h       |   7 +
 src/test/modules/injection_points/Makefile    |   1 +
 .../expected/repack_snapshots.out             | 401 ++++++++
 src/test/modules/injection_points/meson.build |   1 +
 .../specs/repack_snapshots.spec               | 272 ++++++
 19 files changed, 1906 insertions(+), 298 deletions(-)
 create mode 100644 src/test/modules/injection_points/expected/repack_snapshots.out
 create mode 100644 src/test/modules/injection_points/specs/repack_snapshots.spec

diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index 97f850ebf73..1096d9d4dc9 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -43,12 +43,16 @@
 #include "storage/lmgr.h"
 #include "storage/lock.h"
 #include "storage/predicate.h"
+#include "storage/proc.h"
 #include "storage/procarray.h"
 #include "storage/smgr.h"
 #include "utils/builtins.h"
+#include "utils/injection_point.h"
 #include "utils/rel.h"
 #include "utils/tuplesort.h"
 
+static Snapshot finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
+									 BlockNumber start, BlockNumber *end_p);
 static void reform_and_rewrite_tuple(TupleTableSlot *src, TupleTableSlot *reform,
 									 RewriteState rwstate);
 
@@ -583,15 +587,14 @@ static void
 heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 								 Relation OldIndex, bool use_sort,
 								 TransactionId OldestXmin,
-								 Snapshot snapshot,
 								 TransactionId *xid_cutoff,
 								 MultiXactId *multi_cutoff,
 								 double *num_tuples,
 								 double *tups_vacuumed,
-								 double *tups_recently_dead)
+								 double *tups_recently_dead,
+								 void *tableam_data)
 {
 	RewriteState rwstate;
-	BulkInsertState bistate;
 	IndexScanDesc indexScan;
 	TableScanDesc tableScan;
 	HeapScanDesc heapScan;
@@ -602,7 +605,11 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	TupleTableSlot *reform_slot;
 	BufferHeapTupleTableSlot *hslot;
 	BlockNumber prev_cblock = InvalidBlockNumber;
-	bool		concurrent = snapshot != NULL;
+	ChangeContext *chgcxt = (ChangeContext *) tableam_data;
+	bool		concurrent = chgcxt != NULL;
+	Snapshot	snapshot = NULL;
+	BlockNumber range_start = InvalidBlockNumber;
+	BlockNumber range_end = InvalidBlockNumber;
 
 	/* Remember if it's a system catalog */
 	is_system_catalog = IsSystemRelation(OldHeap);
@@ -623,14 +630,11 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	else
 		rwstate = NULL;
 
-	/* In concurrent mode, prepare for bulk-insert operation. */
-	if (concurrent)
-		bistate = GetBulkInsertState();
-	else
-		bistate = NULL;
-
-	/* Set up sorting if wanted */
-	if (use_sort)
+	/*
+	 * Set up sorting if wanted. CONCURRENTLY sorts the tuple w/o tuplesort,
+	 * see below.
+	 */
+	if (use_sort && !concurrent)
 		tuplesort = tuplesort_begin_cluster(oldTupDesc, OldIndex,
 											maintenance_work_mem,
 											NULL, TUPLESORT_NONE);
@@ -642,8 +646,11 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	 * that still need to be copied, we scan with SnapshotAny and use
 	 * HeapTupleSatisfiesVacuum for the visibility test.
 	 *
-	 * In the CONCURRENTLY case, we do regular MVCC visibility tests, using
-	 * the snapshot passed by the caller.
+	 * In the CONCURRENTLY case, we do regular MVCC visibility tests. The
+	 * snapshot changes several times during the scan so that we do not block
+	 * the progress of the xmin horizon for VACUUM too much.  Index scan
+	 * should not be used because it returns tuples in random order, which
+	 * makes it impossible to split the scan into block ranges.
 	 */
 	if (OldIndex != NULL && !use_sort)
 	{
@@ -653,6 +660,8 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		};
 		int64		ci_val[2];
 
+		Assert(!concurrent);
+
 		/* Set phase and OIDOldIndex to columns */
 		ci_val[0] = PROGRESS_REPACK_PHASE_INDEX_SCAN_HEAP;
 		ci_val[1] = RelationGetRelid(OldIndex);
@@ -660,10 +669,8 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 		tableScan = NULL;
 		heapScan = NULL;
-		indexScan = index_beginscan(OldHeap, OldIndex,
-									snapshot ? snapshot : SnapshotAny,
-									NULL, 0, 0,
-									SO_NONE);
+		indexScan = index_beginscan(OldHeap, OldIndex, SnapshotAny, NULL, 0,
+									0, SO_NONE);
 		index_rescan(indexScan, NULL, 0, NULL, 0);
 	}
 	else
@@ -672,16 +679,28 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
 									 PROGRESS_REPACK_PHASE_SEQ_SCAN_HEAP);
 
-		tableScan = table_beginscan(OldHeap,
-									snapshot ? snapshot : SnapshotAny,
-									0, (ScanKey) NULL,
+		tableScan = table_beginscan(OldHeap, SnapshotAny, 0, (ScanKey) NULL,
 									SO_NONE);
 		heapScan = (HeapScanDesc) tableScan;
+
+		/*
+		 * In CONCURRENTLY mode we scan the table by ranges of blocks and the
+		 * algorithm below expects forward direction. (No other direction
+		 * should be set here regardless concurrently anyway.)
+		 */
+		Assert(heapScan->rs_dir == ForwardScanDirection || !concurrent);
 		indexScan = NULL;
 
 		/* Set total heap blocks */
 		pgstat_progress_update_param(PROGRESS_REPACK_TOTAL_HEAP_BLKS,
 									 heapScan->rs_nblocks);
+
+		/* Setup the first range. */
+		if (concurrent)
+		{
+			range_start = heapScan->rs_startblock;
+			range_end = range_start + repack_pages_per_snapshot;
+		}
 	}
 
 	slot = table_slot_create(OldHeap, NULL);
@@ -689,6 +708,30 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	reform_slot = MakeSingleTupleTableSlot(RelationGetDescr(OldHeap),
 										   &TTSOpsVirtual);
 
+	if (concurrent)
+	{
+		/*
+		 * Do not block the progress of xmin horizons.
+		 */
+		PopActiveSnapshot();
+		InvalidateCatalogSnapshot();
+
+		/*
+		 * As there is no snapshot, our xmin should be invalid now.
+		 *
+		 * XXX xid can still be valid. The next patches in the series fix
+		 * that.
+		 */
+		Assert(!TransactionIdIsValid(MyProc->xmin));
+
+		/*
+		 * Wait until the worker has the initial snapshot and retrieve it.
+		 */
+		snapshot = repack_get_snapshot(chgcxt);
+
+		PushActiveSnapshot(snapshot);
+	}
+
 	/*
 	 * Scan through the OldHeap, either in OldIndex order or sequentially;
 	 * copy each tuple into the NewHeap, or transiently to the tuplesort
@@ -705,6 +748,9 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 		if (indexScan != NULL)
 		{
+			/* See above. */
+			Assert(!concurrent);
+
 			if (!index_getnext_slot(indexScan, ForwardScanDirection, slot))
 				break;
 
@@ -726,6 +772,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 				 */
 				pgstat_progress_update_param(PROGRESS_REPACK_HEAP_BLKS_SCANNED,
 											 heapScan->rs_nblocks);
+
 				break;
 			}
 
@@ -842,10 +889,39 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 				continue;
 			}
 		}
+		else
+		{
+			BlockNumber blkno;
+			bool		visible;
+
+			/*
+			 * With CONCURRENTLY, we use each snapshot only for certain range
+			 * of pages, so that VACUUM does not get blocked for too long. So
+			 * first check if the tuple falls into the current range.
+			 */
+			blkno = BufferGetBlockNumber(buf);
+
+			Assert(BlockNumberIsValid(range_end));
+
+			/* End of the current range or wraparound? */
+			if (blkno >= range_end || blkno < range_start)
+				snapshot = finalize_block_range(chgcxt, blkno, range_start,
+												&range_end);
+
+			/* Finally check the tuple visibility. */
+			LockBuffer(buf, BUFFER_LOCK_SHARE);
+			visible = HeapTupleSatisfiesVisibility(tuple, snapshot, buf);
+			LockBuffer(buf, BUFFER_LOCK_UNLOCK);
+
+			if (!visible)
+				continue;
+		}
 
 		*num_tuples += 1;
 		if (tuplesort != NULL)
 		{
+			Assert(!concurrent);
+
 			tuplesort_putheaptuple(tuplesort, tuple);
 
 			/*
@@ -866,7 +942,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 			if (!concurrent)
 				reform_and_rewrite_tuple(slot, reform_slot, rwstate);
 			else
-				heap_insert_for_repack(NewHeap, slot, reform_slot, bistate);
+				heap_insert_for_repack(chgcxt, slot, reform_slot);
 
 			/*
 			 * In indexscan mode and also VACUUM FULL, report increase in
@@ -878,6 +954,28 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		}
 	}
 
+	if (concurrent)
+	{
+		XLogRecPtr	end_of_wal;
+
+		/*
+		 * Process the changes belonging to the last range.
+		 */
+		end_of_wal = GetFlushRecPtr(NULL);
+		repack_process_concurrent_changes(chgcxt, end_of_wal,
+										  InvalidBlockNumber,
+										  InvalidBlockNumber,
+										  false, false);
+
+		/*
+		 * There was an active transaction snapshot on entry, so push one
+		 * before return.
+		 */
+		PopActiveSnapshot();
+		PushActiveSnapshot(GetTransactionSnapshot());
+
+	}
+
 	if (indexScan != NULL)
 		index_endscan(indexScan);
 	if (tableScan != NULL)
@@ -923,10 +1021,13 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 			ExecStoreHeapTuple(tuple, slot, false);
 
 			n_tuples += 1;
-			if (!concurrent)
-				reform_and_rewrite_tuple(slot, reform_slot, rwstate);
-			else
-				heap_insert_for_repack(NewHeap, slot, reform_slot, bistate);
+
+			/*
+			 * The CONCURRENTLY mode uses auxiliary tables rather than
+			 * tuplesort.
+			 */
+			Assert(!concurrent);
+			reform_and_rewrite_tuple(slot, reform_slot, rwstate);
 
 			/* Report n_tuples */
 			pgstat_progress_update_param(PROGRESS_REPACK_HEAP_TUPLES_INSERTED,
@@ -942,8 +1043,89 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	/* Write out any remaining tuples, and fsync if needed */
 	if (rwstate)
 		end_heap_rewrite(rwstate);
-	if (bistate)
-		FreeBulkInsertState(bistate);
+}
+
+/*
+ * Finalize processing of the current block range.
+ *
+ * 'cur' is the current block, 'start' is the first block of the current
+ * range.
+ *
+ * '*end_p': on entry, the first block beyond the current range, on exit, the
+ * first block beyond the new range.
+ *
+ * Return the snapshot for the scan of the new range.
+ */
+static Snapshot
+finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
+					 BlockNumber start, BlockNumber *end_p)
+{
+	BlockNumber end = *end_p;
+	XLogRecPtr	end_of_wal;
+	Snapshot	snapshot;
+
+	/*
+	 * Wait here when testing how snapshot is changed at page boundary.
+	 */
+	INJECTION_POINT("repack-concurrently-new-range", NULL);
+
+	/*
+	 * Decode all the concurrent data changes committed so far before
+	 * requesting the next snapshot - these changes are applicable on top of
+	 * the current snapshot. Since we only copied part of the table so far,
+	 * only changes applicable to that part can be applied.
+	 *
+	 * It's important to apply the changes before we start copying the next
+	 * range of blocks. Without that, in case of concurrent UPDATE, we could
+	 * end up with both old and new tuple present in the new table: the old
+	 * still visible in the current range and the new already visible in the
+	 * following range (for which we'll use more recent snapshot). Thus it'd
+	 * be non-trivial to apply the UPDATE later. By replaying it now, we get
+	 * rid of the old tuple in the current range.
+	 */
+	end_of_wal = GetFlushRecPtr(NULL);
+	repack_process_concurrent_changes(chgcxt, end_of_wal, start, end, true,
+									  false);
+
+	/*
+	 * A new snapshot will be pushed below. Note that it's important to not do
+	 * this earlier, because - while processing the concurrent data changes -
+	 * we might have needed to fetch TOASTed values from the old relation -
+	 * see the UPDATE-to-INSERT conversion in apply_concurrent_changes(). As
+	 * this snapshot protects the data copied from VACUUM, it should also
+	 * protect the TOAST values referenced by the consequent UPDATE
+	 * statements.
+	 */
+	PopActiveSnapshot();
+	InvalidateCatalogSnapshot();
+
+	/* See above. */
+	Assert(!TransactionIdIsValid(MyProc->xmin));
+
+	/*
+	 * XXX It might be worth Assert(CatalogSnapshot == NULL) here, however
+	 * that symbol is not external.
+	 */
+
+	/*
+	 * Compute the end of the new range by aligning 'cur' to a multiple of
+	 * range boundary. This accounts for the possibility that some block
+	 * numbers could have been skipped (due to pages being empty) or that the
+	 * block number could have wrapped around.
+	 */
+	end = cur + repack_pages_per_snapshot - (cur % repack_pages_per_snapshot);
+	*end_p = end;
+
+	/*
+	 * Get the snapshot for the next range - it should have been built at the
+	 * position right after the last change decoded. Data present in the next
+	 * range of blocks will either be visible to the snapshot or appear in the
+	 * next batch of decoded changes.
+	 */
+	snapshot = repack_get_snapshot(chgcxt);
+	PushActiveSnapshot(snapshot);
+
+	return snapshot;
 }
 
 /*
diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 392332b4b2b..0bf19d07db5 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -33,6 +33,7 @@
 #include "postgres.h"
 
 #include "access/amapi.h"
+#include "access/detoast.h"
 #include "access/heapam.h"
 #include "access/multixact.h"
 #include "access/relscan.h"
@@ -95,6 +96,12 @@ typedef struct
 	Oid			indexOid;
 } RelToCluster;
 
+/*
+ * When REPACK (CONCURRENTLY) copies data to the new heap, a new snapshot is
+ * built after processing this many pages. XXX Tune the value.
+ */
+int			repack_pages_per_snapshot = 1024;
+
 /*
  * Backend-local information to control the decoding worker.
  */
@@ -128,11 +135,11 @@ static void check_concurrent_repack_requirements(Relation rel,
 static void rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 							 Oid ident_idx);
 static void copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
-							Snapshot snapshot,
 							bool verbose,
 							bool *pSwapToastByContent,
 							TransactionId *pFreezeXid,
-							MultiXactId *pCutoffMulti);
+							MultiXactId *pCutoffMulti,
+							ChangeContext *chgcxt);
 static List *get_tables_to_repack(RepackCommand cmd, bool usingindex,
 								  MemoryContext permcxt);
 static List *get_tables_to_repack_partitioned(RepackCommand cmd,
@@ -141,23 +148,26 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd,
 static bool repack_is_permitted_for_relation(RepackCommand cmd,
 											 Oid relid, Oid userid);
 
-static void apply_concurrent_changes(ChangeContext *chgcxt);
+static void apply_concurrent_changes(ChangeContext *chgcxt,
+									 BlockNumber range_start,
+									 BlockNumber range_end);
 static void apply_concurrent_insert(RepackDest *dest, TupleTableSlot *slot);
 static void apply_concurrent_update(RepackDest *dest,
 									TupleTableSlot *spilled_tuple,
 									TupleTableSlot *ondisk_tuple);
 static void apply_concurrent_delete(Relation rel, TupleTableSlot *slot);
 static void restore_tuple(BufFile *file, Relation relation,
-						  TupleTableSlot *slot);
+						  TupleTableSlot *slot, BlockNumber *block_nr_p,
+						  BlockNumber *old_block_nr_p);
 static void adjust_toast_pointers(Relation relation, TupleTableSlot *dest,
 								  TupleTableSlot *src);
+static bool is_block_in_range(BlockNumber blknum, BlockNumber start,
+							  BlockNumber end);
 static bool find_target_tuple(RepackDest *dest, TupleTableSlot *locator,
 							  TupleTableSlot *retrieved);
-static bool identity_key_equal(RepackDest *dest, TupleTableSlot *locator,
+static bool identity_key_equal(RepackDest *dest,
+							   TupleTableSlot *locator,
 							   TupleTableSlot *candidate);
-static void process_concurrent_changes(XLogRecPtr end_of_wal,
-									   ChangeContext *chgcxt,
-									   bool done);
 static void initialize_change_context(ChangeContext *chgcxt,
 									  Relation relation,
 									  Oid ident_index_id);
@@ -168,8 +178,12 @@ static void release_change_dest(RepackDest *dest);
 static void rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 											   Oid identIdx,
 											   TransactionId frozenXid,
-											   MultiXactId cutoffMulti);
+											   MultiXactId cutoffMulti,
+											   ChangeContext *chgcxt);
+static void process_auxiliary_table(ChangeContext *chgcxt, Relation OldHeap,
+									Oid identIdx);
 static List *build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes);
+static Oid	build_new_index(Relation NewHeap, Relation OldHeap, Oid oldindex);
 static void copy_index_constraints(Relation old_index, Oid new_index_id,
 								   Oid new_heap_id);
 static void copy_attribute_defaults(Oid old_heap_oid, Oid new_heap_oid);
@@ -183,7 +197,6 @@ static Oid	determine_clustered_index(Relation rel, bool usingindex,
 static void start_repack_decoding_worker(Oid relid);
 static void stop_repack_decoding_worker(void);
 static void stop_repack_decoding_worker_cb(int code, Datum arg);
-static Snapshot get_initial_snapshot(DecodingWorker *worker);
 
 static void ProcessRepackMessage(StringInfo msg);
 static const char *RepackCommandAsString(RepackCommand cmd);
@@ -960,6 +973,14 @@ check_concurrent_repack_requirements(Relation rel, Oid *ident_idx_p)
 						RelationGetRelationName(rel)));
 	}
 
+	/*
+	 * In the CONCURRENTLY mode we don't want to use the same snapshot
+	 * throughout the whole processing, as it could block the progress of xmin
+	 * horizon. Assert should be ok as we already disallow transaction block
+	 * in the CONCURRENTLY case.
+	 */
+	Assert(!IsolationUsesXactSnapshot());
+
 	*ident_idx_p = ident_idx;
 }
 
@@ -996,7 +1017,7 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	TransactionId frozenXid;
 	MultiXactId cutoffMulti;
 	bool		concurrent = OidIsValid(ident_idx);
-	Snapshot	snapshot = NULL;
+	ChangeContext *chgcxt = NULL;
 #if USE_ASSERT_CHECKING
 	LOCKMODE	lmode;
 
@@ -1033,13 +1054,6 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 		 * REPACK CONCURRENTLY.
 		 */
 		start_repack_decoding_worker(tableOid);
-
-		/*
-		 * Wait until the worker has the initial snapshot and retrieve it.
-		 */
-		snapshot = get_initial_snapshot(decoding_worker);
-
-		PushActiveSnapshot(snapshot);
 	}
 
 	/* for CLUSTER or REPACK USING INDEX, mark the index as the one to use */
@@ -1062,26 +1076,115 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	Assert(CheckRelationOidLockedByMe(OIDNewHeap, AccessExclusiveLock, false));
 	NewHeap = table_open(OIDNewHeap, NoLock);
 
-	/*
-	 * In concurrent mode, create a copy of the attribute defaults on the temp
-	 * table, which the executor needs when replaying concurrent data changes.
-	 */
 	if (concurrent)
+	{
+		bool		need_aux_rel;
+
+		/*
+		 * Auxiliary table is needed for clustering in the CONCURRENTLY mode,
+		 * see comments in ChangeContext. FIXME Non-btree indexes are allowed
+		 * historically, but in general, these can hardly define any useful
+		 * order. We ignore them here.
+		 */
+		need_aux_rel = index != NULL && index->rd_rel->relam == BTREE_AM_OID;
+
+		/* Gather information to apply concurrent changes. */
+		chgcxt = palloc0_object(ChangeContext);
+
+		/*
+		 * Create a copy of the attribute defaults on the temp table, which
+		 * the executor needs when replaying concurrent data changes.
+		 */
 		copy_attribute_defaults(tableOid, OIDNewHeap);
 
+		if (!need_aux_rel)
+		{
+			Oid			ident_idx_new;
+
+			/*
+			 * Create the identity index. We will need it during data copying
+			 * so that we can apply the data changes at the appropriate time -
+			 * see comments around the call of
+			 * repack_process_concurrent_changes() with block range specified.
+			 *
+			 * XXX NewHeap is empty - should we pass INDEX_CREATE_SKIP_BUILD?
+			 */
+			ident_idx_new = build_new_index(NewHeap, OldHeap, ident_idx);
+
+			initialize_change_context(chgcxt, NewHeap, ident_idx_new);
+		}
+		else
+		{
+			Oid			aux_oid;
+			Relation	aux_rel;
+			Oid			aux_ident_idx;
+
+			/*
+			 * As the concurrent data changes will be applied to the auxiliary
+			 * heap, the new heap does not need the identity index yet. We'll
+			 * build it after having copied the data from the auxiliary heap.
+			 * (Bulk insert should be more efficient.)
+			 */
+			initialize_change_context(chgcxt, NewHeap, InvalidOid);
+
+			/*
+			 * Like above, but only temporary - no other backend should need
+			 * it.
+			 */
+			aux_oid = make_new_heap(tableOid, tableSpace, accessMethod,
+									RELPERSISTENCE_TEMP, NoLock);
+			Assert(CheckRelationOidLockedByMe(aux_oid, AccessExclusiveLock,
+											  false));
+			aux_rel = table_open(aux_oid, NoLock);
+
+
+			/*
+			 * Copy the attribute defaults as we did for the new heap above -
+			 * the concurrent changes also need to be applied to the auxiliary
+			 * table.
+			 */
+			copy_attribute_defaults(tableOid, aux_oid);
+
+			/*
+			 * The same for identity index. (The additional
+			 * ShareUpdateExclusiveLock on ident_idx is not a problem, it'll
+			 * be released at the end of transaction.)
+			 */
+			aux_ident_idx = build_new_index(aux_rel, OldHeap, ident_idx);
+
+			/*
+			 * Make the relation ready for use.
+			 */
+			chgcxt->cc_dest_aux = palloc0_object(RepackDest);
+			initialize_change_dest(chgcxt->cc_dest_aux, aux_rel,
+								   aux_ident_idx);
+
+			/*
+			 * Set OID of the old relation's clustering index if it's
+			 * different from the identity index. Otherwise set InvalidOid to
+			 * indicate that the identity index should be used for clustering.
+			 */
+			if (RelationGetRelid(index) != ident_idx)
+				chgcxt->cc_clustering_index = RelationGetRelid(index);
+			else
+				chgcxt->cc_clustering_index = InvalidOid;
+		}
+	}
+
 	/* Copy the heap data into the new table in the desired order */
-	copy_table_data(NewHeap, OldHeap, index, snapshot, verbose,
-					&swap_toast_by_content, &frozenXid, &cutoffMulti);
+	copy_table_data(NewHeap, OldHeap, index, verbose,
+					&swap_toast_by_content, &frozenXid, &cutoffMulti,
+					chgcxt);
 
 	/* The historic snapshot won't be needed anymore. */
-	if (snapshot)
+	if (concurrent)
 	{
-		PopActiveSnapshot();
+		/*
+		 * Make sure the active snapshot can see the data copied, so the rows
+		 * can be updated / deleted.
+		 */
 		UpdateActiveSnapshotCommandId();
-	}
 
-	if (concurrent)
-	{
 		Assert(!swap_toast_by_content);
 
 		/*
@@ -1092,10 +1195,16 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 			index_close(index, NoLock);
 
 		rebuild_relation_finish_concurrent(NewHeap, OldHeap, ident_idx,
-										   frozenXid, cutoffMulti);
+										   frozenXid, cutoffMulti, chgcxt);
 
 		pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
 									 PROGRESS_REPACK_PHASE_FINAL_CLEANUP);
+
+		/*
+		 * REPACK (CONCURRENTLY) launches separate transaction(s) so it
+		 * shouldn't rely on the current portal to pop the active snapshot.
+		 */
+		PopActiveSnapshot();
 	}
 	else
 	{
@@ -1258,10 +1367,9 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
  * Insert tuple when processing REPACK CONCURRENTLY.
  *
  * rewriteheap.c is not used in the CONCURRENTLY case because it'd be
- * difficult to do the same in the catch-up phase (as the logical
- * decoding does not provide us with sufficient visibility
- * information). Thus we must use heap_insert() both during the
- * catch-up and here.
+ * difficult to do the same in the catch-up phase (as the logical decoding
+ * does not provide us with sufficient visibility information). Thus we must
+ * use heap_insert() both during the catch-up and here.
  *
  * 'reform' is a slot to use for tuple "reforming", typically to get set
  * values of dropped columns to NULL.
@@ -1269,20 +1377,27 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
  * We pass the NO_LOGICAL flag to heap_insert() in order to skip logical
  * decoding: as soon as REPACK CONCURRENTLY swaps the relation files, it drops
  * this relation, so no logical replication subscription should need the data.
- *
- * BulkInsertState is used because many tuples are inserted in the typical
- * case.
  */
 void
-heap_insert_for_repack(Relation rel, TupleTableSlot *src,
-					   TupleTableSlot *reform, BulkInsertStateData *bistate)
+heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
+					   TupleTableSlot *reform)
 {
 	HeapTuple	tuple;
 	bool		shouldFree;
 	TupleTableSlot *slot;
+	RepackDest *dest;
+
+	/*
+	 * Use the current auxiliary table as output if one is active, otherwise
+	 * insert the tuple into the actual destination table.
+	 */
+	if (chgcxt->cc_dest_aux)
+		dest = chgcxt->cc_dest_aux;
+	else
+		dest = &chgcxt->cc_dest;
 
 	tuple = ExecFetchSlotHeapTuple(src, false, &shouldFree);
-	if (tuple_needs_reform(tuple, src->tts_tupleDescriptor))
+	if (reform != NULL && tuple_needs_reform(tuple, src->tts_tupleDescriptor))
 	{
 		clear_dropped_attributes(tuple, reform);
 		slot = reform;
@@ -1297,8 +1412,16 @@ heap_insert_for_repack(Relation rel, TupleTableSlot *src,
 	if (shouldFree)
 		heap_freetuple(tuple);
 
-	table_tuple_insert(rel, slot, GetCurrentCommandId(true),
-					   TABLE_INSERT_NO_LOGICAL, bistate);
+	table_tuple_insert(dest->rel, slot, GetCurrentCommandId(true),
+					   TABLE_INSERT_NO_LOGICAL, dest->bistate);
+
+	/*
+	 * Insert the tuple into the identity index. initialize_change_context()
+	 * may skip opening of indexes if the identity index is not needed
+	 * immediately.
+	 */
+	if (dest->rri)
+		ExecInsertIndexTuples(dest->rri, dest->estate, 0, slot, NIL, NULL);
 }
 
 bool
@@ -1348,9 +1471,6 @@ clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
 /*
  * Do the physical copying of table data.
  *
- * 'snapshot' and 'decoding_ctx': see table_relation_copy_for_cluster(). Pass
- * iff concurrent processing is required.
- *
  * There are three output parameters:
  * *pSwapToastByContent is set true if toast tables must be swapped by content.
  * *pFreezeXid receives the TransactionId used as freeze cutoff point.
@@ -1358,8 +1478,9 @@ clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
  */
 static void
 copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
-				Snapshot snapshot, bool verbose, bool *pSwapToastByContent,
-				TransactionId *pFreezeXid, MultiXactId *pCutoffMulti)
+				bool verbose, bool *pSwapToastByContent,
+				TransactionId *pFreezeXid, MultiXactId *pCutoffMulti,
+				ChangeContext *chgcxt)
 {
 	Relation	relRelation;
 	HeapTuple	reltup;
@@ -1376,7 +1497,7 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 	int			elevel = verbose ? INFO : DEBUG2;
 	PGRUsage	ru0;
 	char	   *nspname;
-	bool		concurrent = snapshot != NULL;
+	bool		concurrent = chgcxt != NULL;
 	LOCKMODE	lmode;
 
 	lmode = RepackLockLevel(concurrent);
@@ -1480,18 +1601,28 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 			cutoffs.MultiXactCutoff = relminmxid;
 	}
 
-	/*
-	 * Decide whether to use an indexscan or seqscan-and-optional-sort to scan
-	 * the OldHeap.  We know how to use a sort to duplicate the ordering of a
-	 * btree index, and will use seqscan-and-sort for that case if the planner
-	 * tells us it's cheaper.  Otherwise, always indexscan if an index is
-	 * provided, else plain seqscan.
-	 */
-	if (OldIndex != NULL && OldIndex->rd_rel->relam == BTREE_AM_OID)
-		use_sort = plan_cluster_use_sort(RelationGetRelid(OldHeap),
-										 RelationGetRelid(OldIndex));
+	if (!concurrent)
+	{
+		/*
+		 * Decide whether to use an indexscan or seqscan-and-optional-sort to
+		 * scan the OldHeap.  We know how to use a sort to duplicate the
+		 * ordering of a btree index, and will use seqscan-and-sort for that
+		 * case if the planner tells us it's cheaper.  Otherwise, always
+		 * indexscan if an index is provided, else plain seqscan.
+		 */
+		if (OldIndex != NULL && OldIndex->rd_rel->relam == BTREE_AM_OID)
+			use_sort = plan_cluster_use_sort(RelationGetRelid(OldHeap),
+											 RelationGetRelid(OldIndex));
+		else
+			use_sort = false;
+	}
 	else
-		use_sort = false;
+	{
+		/*
+		 * To use multiple snapshots, we need to read the table sequentially.
+		 */
+		use_sort = true;
+	}
 
 	/* Log what we're doing */
 	if (OldIndex != NULL && !use_sort)
@@ -1518,11 +1649,11 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 	 * values (e.g. because the AM doesn't use freezing).
 	 */
 	table_relation_copy_for_cluster(OldHeap, NewHeap, OldIndex, use_sort,
-									cutoffs.OldestXmin, snapshot,
+									cutoffs.OldestXmin,
 									&cutoffs.FreezeLimit,
 									&cutoffs.MultiXactCutoff,
 									&num_tuples, &tups_vacuumed,
-									&tups_recently_dead);
+									&tups_recently_dead, chgcxt);
 
 	/* return selected values to caller, get set as relfrozenxid/minmxid */
 	*pFreezeXid = cutoffs.FreezeLimit;
@@ -2429,6 +2560,8 @@ repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid)
  * instead return the opened and locked relcache entry, so that caller can
  * process the partitions using the multiple-table handling code.  In this
  * case, if an index name is given, it's up to the caller to resolve it.
+ *
+ * A new transaction is started in either case.
  */
 static Relation
 process_single_relation(RepackStmt *stmt, LOCKMODE lockmode, bool isTopLevel,
@@ -2441,6 +2574,31 @@ process_single_relation(RepackStmt *stmt, LOCKMODE lockmode, bool isTopLevel,
 	Assert(stmt->command == REPACK_COMMAND_CLUSTER ||
 		   stmt->command == REPACK_COMMAND_REPACK);
 
+	if (params->options & CLUOPT_CONCURRENT)
+	{
+		/*
+		 * Since REPACK (CONCURRENTLY) pops the active snapshot during the
+		 * processing (it creates and pushes snapshots on its own), and since
+		 * that snapshot can be referenced by the current portal, we need to
+		 * make sure that the portal has no dangling pointer to the snapshot.
+		 * Starting a new transaction seems to be the simplest way.
+		 *
+		 * XXX The following patches in the series make this unnecessary, as
+		 * they start new transactions for other reasons elsewhere.
+		 */
+		PopActiveSnapshot();
+		CommitTransactionCommand();
+
+		/* Start a new transaction. */
+		StartTransactionCommand();
+
+		/*
+		 * Functions in indexes may want a snapshot set. Note that the portal
+		 * is not aware of this one, so the caller needs to pop it explicitly.
+		 */
+		PushActiveSnapshot(GetTransactionSnapshot());
+	}
+
 	/*
 	 * Make sure ANALYZE is specified if a column list is present.
 	 */
@@ -2580,10 +2738,11 @@ RepackCommandAsString(RepackCommand cmd)
 }
 
 /*
- * Apply all the changes provided by decoding worker.
+ * Apply data changes that affect pages in given range.
  */
 static void
-apply_concurrent_changes(ChangeContext *chgcxt)
+apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
+						 BlockNumber range_end)
 {
 	ConcurrentChangeKind kind = '\0';
 	RepackDest *dest;
@@ -2592,18 +2751,31 @@ apply_concurrent_changes(ChangeContext *chgcxt)
 	TupleTableSlot *old_update_tuple;
 	TupleTableSlot *ondisk_tuple;
 	bool		have_old_tuple = false;
+	bool		check_range;
 	MemoryContext oldcxt;
 	DecodingWorkerShared *shared;
 	char		fname[MAXPGPATH];
 	BufFile    *file;
 
-	dest = &chgcxt->cc_dest;
+	/*
+	 * Use the auxiliary table if one exists, otherwise the "final"
+	 * destination table.
+	 */
+	dest = chgcxt->cc_dest_aux ? chgcxt->cc_dest_aux : &chgcxt->cc_dest;
 	rel = dest->rel;
 
+	/*
+	 * Range needs to be checked if the bounds are specified. Expect either
+	 * both or none.
+	 */
+	Assert(BlockNumberIsValid(range_start) == BlockNumberIsValid(range_end));
+	check_range = BlockNumberIsValid(range_start);
+
 	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
 
 	/* Open the file containing the changes. */
-	DecodingWorkerFileName(fname, shared->relid, chgcxt->cc_file_seq);
+	DecodingWorkerFileName(fname, shared->relid, chgcxt->cc_file_seq_changes,
+						   false);
 	file = BufFileOpenFileSet(&shared->sfs.fs, fname, O_RDONLY, false);
 
 	spilled_tuple = MakeSingleTupleTableSlot(RelationGetDescr(rel),
@@ -2619,6 +2791,9 @@ apply_concurrent_changes(ChangeContext *chgcxt)
 	{
 		size_t		nread;
 		ConcurrentChangeKind prevkind = kind;
+		BlockNumber block,
+					old_block;
+		BlockNumber *old_block_p;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -2633,7 +2808,7 @@ apply_concurrent_changes(ChangeContext *chgcxt)
 		 */
 		if (kind == CHANGE_UPDATE_OLD)
 		{
-			restore_tuple(file, rel, old_update_tuple);
+			restore_tuple(file, rel, old_update_tuple, NULL, NULL);
 			have_old_tuple = true;
 			continue;
 		}
@@ -2657,22 +2832,39 @@ apply_concurrent_changes(ChangeContext *chgcxt)
 
 		/*
 		 * Now restore the tuple into the slot and execute the change.
+		 *
+		 * old_block is only stored with UPDATE_NEW.
 		 */
-		restore_tuple(file, rel, spilled_tuple);
+		old_block_p = kind == CHANGE_UPDATE_NEW ? &old_block : NULL;
+		restore_tuple(file, rel, spilled_tuple, &block, old_block_p);
 
 		if (kind == CHANGE_INSERT)
 		{
-			apply_concurrent_insert(dest, spilled_tuple);
+			/*
+			 * Only insert the tuple if it fits into the current range (or if
+			 * range does not matter).
+			 */
+			if (!check_range ||
+				is_block_in_range(block, range_start, range_end))
+				apply_concurrent_insert(dest, spilled_tuple);
 		}
 		else if (kind == CHANGE_DELETE)
 		{
-			bool		found;
+			/*
+			 * Only delete the tuple if it fits into the current range (or if
+			 * range does not matter).
+			 */
+			if (!check_range ||
+				is_block_in_range(block, range_start, range_end))
+			{
+				bool		found;
 
-			/* Find the tuple to be deleted */
-			found = find_target_tuple(dest, spilled_tuple, ondisk_tuple);
-			if (!found)
-				elog(ERROR, "could not find target tuple");
-			apply_concurrent_delete(rel, ondisk_tuple);
+				/* Find the tuple to be deleted */
+				found = find_target_tuple(dest, spilled_tuple, ondisk_tuple);
+				if (!found)
+					elog(ERROR, "could not find target tuple");
+				apply_concurrent_delete(rel, ondisk_tuple);
+			}
 		}
 		else if (kind == CHANGE_UPDATE_NEW)
 		{
@@ -2684,21 +2876,71 @@ apply_concurrent_changes(ChangeContext *chgcxt)
 			else
 				key = spilled_tuple;
 
-			/* Find the tuple to be updated or deleted. */
-			found = find_target_tuple(dest, key, ondisk_tuple);
-			if (!found)
-				elog(ERROR, "could not find target tuple");
-
 			/*
-			 * If 'tup' contains TOAST pointers, they point to the old
-			 * relation's toast. Copy the corresponding TOAST pointers for the
-			 * new relation from the existing tuple. (The fact that we
-			 * received a TOAST pointer here implies that the attribute hasn't
-			 * changed.)
+			 * Perform normal update if both old and new version are in the
+			 * current range.
 			 */
-			adjust_toast_pointers(rel, spilled_tuple, ondisk_tuple);
+			if (!check_range ||
+				(is_block_in_range(old_block, range_start, range_end) &&
+				 is_block_in_range(block, range_start, range_end)))
+			{
+				/* Find the tuple to be updated or deleted. */
+				found = find_target_tuple(dest, key, ondisk_tuple);
+				if (!found)
+					elog(ERROR, "could not find target tuple");
 
-			apply_concurrent_update(dest, spilled_tuple, ondisk_tuple);
+				/*
+				 * If 'spilled_tuple' contains TOAST pointers, they point to
+				 * the old relation's toast. Copy the corresponding TOAST
+				 * pointers for the new relation from the existing tuple. (The
+				 * fact that we received a TOAST pointer here implies that the
+				 * attribute hasn't changed.)
+				 */
+				adjust_toast_pointers(rel, spilled_tuple, ondisk_tuple);
+
+				apply_concurrent_update(dest, spilled_tuple, ondisk_tuple);
+			}
+			else
+			{
+				Assert(check_range);
+
+				if (is_block_in_range(block, range_start, range_end))
+				{
+					/*
+					 * The old key is in another range, so only insert the new
+					 * one into the current range. The old version should not
+					 * be visible to the snapshot that we'll use to copy the
+					 * other range.
+					 *
+					 * Unlike UPDATE, there's no old tuple to copy the TOAST
+					 * pointers from. Therefore pass NULL for the source
+					 * tuple, to enforce detoasting of the TOAST pointers in
+					 * 'spilled_tuple'.
+					 */
+					adjust_toast_pointers(rel, spilled_tuple, NULL);
+
+					apply_concurrent_insert(dest, spilled_tuple);
+				}
+				else if (is_block_in_range(old_block, range_start, range_end))
+				{
+					found = find_target_tuple(dest, key, ondisk_tuple);
+					if (!found)
+						elog(ERROR, "could not find target tuple");
+
+					/*
+					 * The new key is in another range, so only delete the old
+					 * one from the current range. The new version should be
+					 * visible to the snapshot that we'll use to copy the
+					 * other range.
+					 */
+					apply_concurrent_delete(rel, ondisk_tuple);
+				}
+
+				/*
+				 * Otherwise, both tuple versions belong to another range, so
+				 * there's nothing to do here.
+				 */
+			}
 
 			ExecClearTuple(old_update_tuple);
 			have_old_tuple = false;
@@ -2820,7 +3062,8 @@ apply_concurrent_delete(Relation rel, TupleTableSlot *slot)
  * smaller than MaxAllocSize but the whole tuple is bigger.
  */
 static void
-restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot)
+restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot,
+			  BlockNumber *block_nr_p, BlockNumber *old_block_nr_p)
 {
 	uint32		t_len;
 	HeapTuple	tup;
@@ -2832,7 +3075,6 @@ restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot)
 	tup->t_data = (HeapTupleHeader) ((char *) tup + HEAPTUPLESIZE);
 	BufFileReadExact(file, tup->t_data, t_len);
 	tup->t_len = t_len;
-	ItemPointerSetInvalid(&tup->t_self);
 	tup->t_tableOid = RelationGetRelid(relation);
 
 	/*
@@ -2841,6 +3083,12 @@ restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot)
 	 */
 	ExecForceStoreHeapTuple(tup, slot, false);
 
+	/* Handle TID separate because not all tuple slots care about it. */
+	if (block_nr_p)
+		*block_nr_p = ItemPointerGetBlockNumber(&tup->t_data->t_ctid);
+	if (old_block_nr_p)
+		BufFileReadExact(file, old_block_nr_p, sizeof(BlockNumber));
+
 	/*
 	 * Next, read any attributes we stored separately into the tts_values
 	 * array elements expecting them, if any.  This matches
@@ -2890,10 +3138,12 @@ restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot)
 
 /*
  * Adjust 'dest' replacing any EXTERNAL_ONDISK toast pointers with the
- * corresponding ones from 'src'.
+ * corresponding ones from 'src'. If 'src' is NULL, replace the toast pointer
+ * with the actual value.
  */
 static void
-adjust_toast_pointers(Relation relation, TupleTableSlot *dest, TupleTableSlot *src)
+adjust_toast_pointers(Relation relation, TupleTableSlot *dest,
+					  TupleTableSlot *src)
 {
 	TupleDesc	desc = dest->tts_tupleDescriptor;
 
@@ -2914,9 +3164,44 @@ adjust_toast_pointers(Relation relation, TupleTableSlot *dest, TupleTableSlot *s
 		varlena_dst = (varlena *) DatumGetPointer(dest->tts_values[i]);
 		if (!VARATT_IS_EXTERNAL_ONDISK(varlena_dst))
 			continue;
-		slot_getsomeattrs(src, i + 1);
 
-		dest->tts_values[i] = src->tts_values[i];
+		/*
+		 * Ideally we just copy the value, but if there is no source tuple, we
+		 * need to detoast the value.
+		 */
+		if (src)
+		{
+			slot_getsomeattrs(src, i + 1);
+			dest->tts_values[i] = src->tts_values[i];
+		}
+		else
+		{
+			varlena    *detoasted;
+
+			detoasted = detoast_external_attr(varlena_dst);
+			dest->tts_values[i] = PointerGetDatum(detoasted);
+		}
+	}
+}
+
+/*
+ * Check if tuple originates from given range of blocks that have already been
+ * copied.
+ */
+static bool
+is_block_in_range(BlockNumber blknum, BlockNumber start, BlockNumber end)
+{
+	Assert(BlockNumberIsValid(start) && BlockNumberIsValid(end));
+	Assert(BlockNumberIsValid(blknum));
+
+	if (start < end)
+		return blknum >= start && blknum < end;
+	else
+	{
+		/* Has the scan position wrapped around? */
+		Assert(start > end);
+
+		return blknum >= start || blknum < end;
 	}
 }
 
@@ -3012,75 +3297,19 @@ identity_key_equal(RepackDest *dest, TupleTableSlot *locator,
 }
 
 /*
- * Decode and apply concurrent changes, up to (and including) the record whose
- * LSN is 'end_of_wal'.
- *
- * XXX the names "process_concurrent_changes" and "apply_concurrent_changes"
- * are far too similar to each other.
- */
-static void
-process_concurrent_changes(XLogRecPtr end_of_wal, ChangeContext *chgcxt, bool done)
-{
-	DecodingWorkerShared *shared;
-	char		fname[MAXPGPATH];
-	BufFile    *file;
-
-	pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
-								 PROGRESS_REPACK_PHASE_CATCH_UP);
-
-	/* Ask the worker for the file. */
-	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
-	SpinLockAcquire(&shared->mutex);
-	shared->lsn_upto = end_of_wal;
-	shared->done = done;
-	SpinLockRelease(&shared->mutex);
-
-	/*
-	 * The worker needs to finish processing of the current WAL record. Even
-	 * if it's idle, it'll need to close the output file. Thus we're likely to
-	 * wait, so prepare for sleep.
-	 */
-	ConditionVariablePrepareToSleep(&shared->cv);
-	for (;;)
-	{
-		int			last_exported;
-
-		SpinLockAcquire(&shared->mutex);
-		last_exported = shared->last_exported;
-		SpinLockRelease(&shared->mutex);
-
-		/*
-		 * Has the worker exported the file we are waiting for?
-		 */
-		if (last_exported == chgcxt->cc_file_seq)
-			break;
-
-		ConditionVariableSleep(&shared->cv, WAIT_EVENT_REPACK_WORKER_EXPORT);
-	}
-	ConditionVariableCancelSleep();
-
-	/* Open the file. */
-	DecodingWorkerFileName(fname, shared->relid, chgcxt->cc_file_seq);
-	file = BufFileOpenFileSet(&shared->sfs.fs, fname, O_RDONLY, false);
-	apply_concurrent_changes(chgcxt);
-
-	BufFileClose(file);
-
-	/* Get ready for the next file. */
-	chgcxt->cc_file_seq++;
-}
-
-/*
- * Initialize the ChangeContext struct for the given relation, with
- * the given index as identity index.
+ * Initialize the ChangeContext struct for the given relation.
  */
 static void
-initialize_change_context(ChangeContext *chgcxt,
-						  Relation relation, Oid ident_index_id)
+initialize_change_context(ChangeContext *chgcxt, Relation relation,
+						  Oid ident_index_id)
 {
 	initialize_change_dest(&chgcxt->cc_dest, relation, ident_index_id);
 
-	chgcxt->cc_file_seq = WORKER_FILE_SNAPSHOT + 1;
+	chgcxt->cc_file_seq_snapshot = 0;
+	chgcxt->cc_file_seq_changes = 0;
+
+	chgcxt->cc_dest_aux = NULL;
+	chgcxt->cc_clustering_index = InvalidOid;
 }
 
 /*
@@ -3090,6 +3319,8 @@ static void
 release_change_context(ChangeContext *chgcxt)
 {
 	release_change_dest(&chgcxt->cc_dest);
+	if (chgcxt->cc_dest_aux)
+		release_change_dest(chgcxt->cc_dest_aux);
 }
 
 /*
@@ -3104,6 +3335,10 @@ initialize_change_dest(RepackDest *dest, Relation relation,
 	dest->rel = relation;
 	dest->bistate = GetBulkInsertState();
 
+	/* If there's no identity index yet, there should be no indexes at all. */
+	if (!OidIsValid(ident_index_id))
+		return;
+
 	/* Only initialize fields needed by ExecInsertIndexTuples(). */
 	dest->estate = CreateExecutorState();
 
@@ -3241,6 +3476,11 @@ static void
 release_change_dest(RepackDest *dest)
 {
 	FreeBulkInsertState(dest->bistate);
+
+	/* It's possible that no indexes were opened during initialization. */
+	if (dest->rri == NULL)
+		return;
+
 	ExecCloseIndices(dest->rri);
 	FreeExecutorState(dest->estate);
 	/* XXX are these pfrees necessary? */
@@ -3259,7 +3499,8 @@ release_change_dest(RepackDest *dest)
 static void
 rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 								   Oid identIdx, TransactionId frozenXid,
-								   MultiXactId cutoffMulti)
+								   MultiXactId cutoffMulti,
+								   ChangeContext *chgcxt)
 {
 	List	   *ind_oids_new;
 	Oid			old_table_oid = RelationGetRelid(OldHeap);
@@ -3269,14 +3510,17 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 			   *lc2;
 	char		relpersistence;
 	bool		is_system_catalog;
-	Oid			ident_idx_new;
 	XLogRecPtr	end_of_wal;
 	List	   *indexrels;
-	ChangeContext chgcxt;
+	List	   *inds_tmp = NIL;
 
 	Assert(CheckRelationLockedByMe(OldHeap, ShareUpdateExclusiveLock, false));
 	Assert(CheckRelationLockedByMe(NewHeap, AccessExclusiveLock, false));
 
+	/* If we have the auxiliary table, this is the moment we should use it. */
+	if (chgcxt->cc_dest_aux)
+		process_auxiliary_table(chgcxt, OldHeap, identIdx);
+
 	/*
 	 * Unlike the exclusive case, we build new indexes for the new relation
 	 * rather than swapping the storage and reindexing the old relation. The
@@ -3292,32 +3536,26 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	 * might not be enough for commands like ALTER INDEX ... SET ... (Those
 	 * are not necessarily dangerous, but can make user confused if the
 	 * changes they do get lost due to REPACK.)
+	 *
+	 * As the identity index already had to be built, skip it here. XXX
+	 * Consider if the retail inserts during data copying (in the case w/o
+	 * auxiliary table) can be a problem in terms of index layout. Shouldn't
+	 * we drop the identity index and build it using bulk insert too?
 	 */
+	foreach_oid(ind_oid, ind_oids_old)
+	{
+		if (ind_oid != identIdx)
+			inds_tmp = lappend_oid(inds_tmp, ind_oid);
+	}
+	ind_oids_old = inds_tmp;
 	ind_oids_new = build_new_indexes(NewHeap, OldHeap, ind_oids_old);
 
 	/*
-	 * The identity index in the new relation appears in the same relative
-	 * position as the corresponding index in the old relation.  Find it.
+	 * The identity index will be involved in the following processing.
 	 */
-	ident_idx_new = InvalidOid;
-	foreach_oid(ind_old, ind_oids_old)
-	{
-		if (identIdx == ind_old)
-		{
-			int			pos = foreach_current_index(ind_old);
-
-			if (list_length(ind_oids_new) <= pos)
-				elog(ERROR, "list of new indexes too short");
-			ident_idx_new = list_nth_oid(ind_oids_new, pos);
-			break;
-		}
-	}
-	if (!OidIsValid(ident_idx_new))
-		elog(ERROR, "could not find index matching \"%s\" at the new relation",
-			 get_rel_name(identIdx));
-
-	/* Gather information to apply concurrent changes. */
-	initialize_change_context(&chgcxt, NewHeap, ident_idx_new);
+	ind_oids_old = lappend_oid(ind_oids_old, identIdx);
+	ind_oids_new = lappend_oid(ind_oids_new,
+							   RelationGetRelid(chgcxt->cc_dest.ident_index));
 
 	/*
 	 * During testing, wait for another backend to perform concurrent data
@@ -3334,11 +3572,13 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	end_of_wal = GetFlushRecPtr(NULL);
 
 	/*
-	 * Apply concurrent changes first time, to minimize the time we need to
-	 * hold AccessExclusiveLock. (Quite some amount of WAL could have been
+	 * Decode and apply concurrent changes again, to minimize the time we need
+	 * to hold AccessExclusiveLock. (Quite some amount of WAL could have been
 	 * written during the data copying and index creation.)
 	 */
-	process_concurrent_changes(end_of_wal, &chgcxt, false);
+	repack_process_concurrent_changes(chgcxt, end_of_wal,
+									  InvalidBlockNumber, InvalidBlockNumber,
+									  false, false);
 
 	/*
 	 * Acquire AccessExclusiveLock on the table, its TOAST relation (if there
@@ -3392,10 +3632,12 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	end_of_wal = GetFlushRecPtr(NULL);
 
 	/*
-	 * Apply the concurrent changes again. Indicate that the decoding worker
-	 * won't be needed anymore.
+	 * Decode and apply the concurrent changes again. Indicate that the
+	 * decoding worker won't be needed anymore.
 	 */
-	process_concurrent_changes(end_of_wal, &chgcxt, true);
+	repack_process_concurrent_changes(chgcxt, end_of_wal,
+									  InvalidBlockNumber, InvalidBlockNumber,
+									  false, true);
 
 	/* Remember info about rel before closing OldHeap */
 	relpersistence = OldHeap->rd_rel->relpersistence;
@@ -3443,7 +3685,7 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	table_close(NewHeap, NoLock);
 
 	/* Cleanup what we don't need anymore. (And close the identity index.) */
-	release_change_context(&chgcxt);
+	release_change_context(chgcxt);
 
 	/*
 	 * Swap the relations and their TOAST relations and TOAST indexes. This
@@ -3462,6 +3704,103 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 					 relpersistence);
 }
 
+/*
+ * Copy the contents of the auxiliary table to the new table in the desired
+ * order, then drop the auxiliary table.
+ */
+static void
+process_auxiliary_table(ChangeContext *chgcxt, Relation OldHeap, Oid identIdx)
+{
+	RepackDest *dest = chgcxt->cc_dest_aux;
+	Oid			ident_idx_new;
+	Relation	clustering_index;
+	IndexScanDesc scan;
+	TupleTableSlot *slot;
+	Oid			aux_oid;
+	ObjectAddress object;
+	Relation	rel;
+
+	/*
+	 * First, make sure the clustering index exists.
+	 */
+	if (OidIsValid(chgcxt->cc_clustering_index))
+	{
+		Oid			cl_ind_oid;
+
+		/*
+		 * Create it according to the clustering index on the old relation.
+		 */
+		cl_ind_oid = build_new_index(dest->rel, OldHeap,
+									 chgcxt->cc_clustering_index);
+		clustering_index = index_open(cl_ind_oid, NoLock);
+	}
+	else
+	{
+		/* The identity index is also the clustering index. */
+		clustering_index = dest->ident_index;
+	}
+
+	/*
+	 * Now do the copying. Before starting, clear ->cc_dest_aux so that
+	 * insertions go to the final table, rather than the auxiliary one.
+	 */
+	chgcxt->cc_dest_aux = NULL;
+	slot = table_slot_create(dest->rel, NULL);
+
+	/*
+	 * Note: the current active snapshot blocks the progress of xmin
+	 * horizon(s). The next patches in the series should fix this by using a
+	 * new kind of snapshot (which we can use here because there are no
+	 * transaction aborts in the auxiliary table).
+	 */
+	scan = index_beginscan(dest->rel, clustering_index, GetActiveSnapshot(),
+						   NULL, 0, 0, SO_NONE);
+	index_rescan(scan, NULL, 0, NULL, 0);
+	for (;;)
+	{
+		CHECK_FOR_INTERRUPTS();
+
+		if (!index_getnext_slot(scan, ForwardScanDirection, slot))
+			break;
+
+		/*
+		 * Reforming should have been performed during insertions into the
+		 * auxiliary table.
+		 */
+		heap_insert_for_repack(chgcxt, slot, NULL);
+	}
+	index_endscan(scan);
+	ExecDropSingleTupleTableSlot(slot);
+
+	/*
+	 * Close the relation, its identity index and clustering index if we had
+	 * to open it above. Lock will be released on commit.
+	 */
+	aux_oid = RelationGetRelid(dest->rel);
+	table_close(dest->rel, NoLock);
+	if (OidIsValid(chgcxt->cc_clustering_index))
+		index_close(clustering_index, NoLock);
+	/* Here we close the other indexes. */
+	release_change_dest(dest);
+
+	/* Drop the auxiliary table. */
+	object.classId = RelationRelationId;
+	object.objectId = aux_oid;
+	object.objectSubId = 0;
+	performDeletion(&object, DROP_RESTRICT, PERFORM_DELETION_INTERNAL);
+
+	/* Build the identity index on the new relation. */
+	ident_idx_new = build_new_index(chgcxt->cc_dest.rel, OldHeap, identIdx);
+
+	/*
+	 * Make the new heap ready to use the index for future replaying of
+	 * concurrent changes.
+	 */
+	rel = chgcxt->cc_dest.rel;
+	release_change_dest(&chgcxt->cc_dest);
+	initialize_change_dest(&chgcxt->cc_dest, rel, ident_idx_new);
+}
+
 /*
  * Build indexes on NewHeap according to those on OldHeap.
  *
@@ -3477,34 +3816,48 @@ build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes)
 {
 	List	   *result = NIL;
 
-	pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
-								 PROGRESS_REPACK_PHASE_REBUILD_INDEX);
-
 	foreach_oid(oldindex, OldIndexes)
 	{
 		Oid			newindex;
-		char	   *newName;
-		Relation	ind;
-
-		ind = index_open(oldindex, ShareUpdateExclusiveLock);
-
-		newName = ChooseRelationName(get_rel_name(oldindex),
-									 NULL,
-									 "repacknew",
-									 get_rel_namespace(ind->rd_index->indrelid),
-									 false);
-		newindex = index_create_copy(NewHeap, INDEX_CREATE_SUPPRESS_PROGRESS,
-									 oldindex, ind->rd_rel->reltablespace,
-									 newName);
-		copy_index_constraints(ind, newindex, RelationGetRelid(NewHeap));
-		result = lappend_oid(result, newindex);
 
-		index_close(ind, NoLock);
+		newindex = build_new_index(NewHeap, OldHeap, oldindex);
+		result = lappend_oid(result, newindex);
 	}
 
 	return result;
 }
 
+/*
+ * Subroutine of build_new_indexes().
+ */
+static Oid
+build_new_index(Relation NewHeap, Relation OldHeap, Oid oldindex)
+{
+	Oid			newindex;
+	char	   *newName;
+	Relation	ind;
+
+	pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+								 PROGRESS_REPACK_PHASE_REBUILD_INDEX);
+
+	ind = index_open(oldindex, ShareUpdateExclusiveLock);
+
+	newName = ChooseRelationName(get_rel_name(oldindex),
+								 NULL,
+								 "repacknew",
+								 get_rel_namespace(ind->rd_index->indrelid),
+								 false);
+	/* Functions in indexes may want a snapshot set. */
+	Assert(ActiveSnapshotSet());
+	newindex = index_create_copy(NewHeap, INDEX_CREATE_SUPPRESS_PROGRESS,
+								 oldindex, ind->rd_rel->reltablespace,
+								 newName);
+	copy_index_constraints(ind, newindex, RelationGetRelid(NewHeap));
+	index_close(ind, NoLock);
+
+	return newindex;
+}
+
 /*
  * Create a transient copy of a constraint -- supported by a transient
  * copy of the index that supports the original constraint.
@@ -3693,10 +4046,13 @@ start_repack_decoding_worker(Oid relid)
 
 	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
 	shared->initialized = false;
+	/* Snapshot is the first thing we need from the worker. */
+	shared->snapshot_requested = true;
 	shared->lsn_upto = InvalidXLogRecPtr;
 	shared->done = false;
 	SharedFileSetInit(&shared->sfs, decoding_worker->seg);
-	shared->last_exported = -1;
+	shared->last_exported_snapshot = -1;
+	shared->last_exported_changes = -1;
 	SpinLockInit(&shared->mutex);
 	shared->dbid = MyDatabaseId;
 
@@ -3818,10 +4174,10 @@ stop_repack_decoding_worker_cb(int code, Datum arg)
 }
 
 /*
- * Get the initial snapshot from the decoding worker.
+ * Get snapshot from the decoding worker.
  */
-static Snapshot
-get_initial_snapshot(DecodingWorker *worker)
+Snapshot
+repack_get_snapshot(ChangeContext *chgcxt)
 {
 	DecodingWorkerShared *shared;
 	char		fname[MAXPGPATH];
@@ -3830,12 +4186,13 @@ get_initial_snapshot(DecodingWorker *worker)
 	char	   *snap_space;
 	Snapshot	snapshot;
 
-	shared = (DecodingWorkerShared *) dsm_segment_address(worker->seg);
+	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
 
 	/*
-	 * The worker needs to initialize the logical decoding, which usually
-	 * takes some time. Therefore it makes sense to prepare for the sleep
-	 * first.
+	 * For the first snapshot request, the worker needs to initialize the
+	 * logical decoding, which usually takes some time. Therefore it makes
+	 * sense to prepare for the sleep first. Does it make sense to skip the
+	 * preparation on the next requests?
 	 */
 	ConditionVariablePrepareToSleep(&shared->cv);
 	for (;;)
@@ -3843,13 +4200,13 @@ get_initial_snapshot(DecodingWorker *worker)
 		int			last_exported;
 
 		SpinLockAcquire(&shared->mutex);
-		last_exported = shared->last_exported;
+		last_exported = shared->last_exported_snapshot;
 		SpinLockRelease(&shared->mutex);
 
 		/*
 		 * Has the worker exported the file we are waiting for?
 		 */
-		if (last_exported == WORKER_FILE_SNAPSHOT)
+		if (last_exported == chgcxt->cc_file_seq_snapshot)
 			break;
 
 		ConditionVariableSleep(&shared->cv, WAIT_EVENT_REPACK_WORKER_EXPORT);
@@ -3857,20 +4214,98 @@ get_initial_snapshot(DecodingWorker *worker)
 	ConditionVariableCancelSleep();
 
 	/* Read the snapshot from a file. */
-	DecodingWorkerFileName(fname, shared->relid, WORKER_FILE_SNAPSHOT);
+	DecodingWorkerFileName(fname, shared->relid, chgcxt->cc_file_seq_snapshot,
+						   true);
 	file = BufFileOpenFileSet(&shared->sfs.fs, fname, O_RDONLY, false);
 	BufFileReadExact(file, &snap_size, sizeof(snap_size));
 	snap_space = (char *) palloc(snap_size);
 	BufFileReadExact(file, snap_space, snap_size);
 	BufFileClose(file);
 
+#ifdef USE_ASSERT_CHECKING
+	SpinLockAcquire(&shared->mutex);
+	Assert(!shared->snapshot_requested);
+	shared->snapshot_requested = false;
+	SpinLockRelease(&shared->mutex);
+#endif
+
 	/* Restore it. */
 	snapshot = RestoreSnapshot(snap_space);
 	pfree(snap_space);
 
+	/* Get ready for the next snapshot. */
+	chgcxt->cc_file_seq_snapshot++;
+
 	return snapshot;
 }
 
+/*
+ * Get concurrent changes, up to (and including) the record whose LSN is
+ * 'end_of_wal', from the decoding worker, and apply them to the new table. If
+ * block range is specified, only apply changes related to that range.
+ *
+ * If 'request_snapshot' is true, the snapshot built at LSN following the last
+ * data change needs to be exported too.
+ */
+void
+repack_process_concurrent_changes(ChangeContext *chgcxt,
+								  XLogRecPtr end_of_wal,
+								  BlockNumber range_start,
+								  BlockNumber range_end,
+								  bool request_snapshot, bool done)
+{
+	DecodingWorkerShared *shared;
+
+	pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
+								 PROGRESS_REPACK_PHASE_CATCH_UP);
+
+	/* Ask the worker for the file. */
+	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
+	SpinLockAcquire(&shared->mutex);
+	shared->lsn_upto = end_of_wal;
+	Assert(!shared->snapshot_requested);
+	shared->snapshot_requested = request_snapshot;
+	shared->done = done;
+	SpinLockRelease(&shared->mutex);
+
+	/*
+	 * The worker needs to finish processing of the current WAL record. Even
+	 * if it's idle, it'll need to close the output file. Thus we're likely to
+	 * wait, so prepare for sleep.
+	 */
+	ConditionVariablePrepareToSleep(&shared->cv);
+	for (;;)
+	{
+		int			last_exported;
+
+		SpinLockAcquire(&shared->mutex);
+		last_exported = shared->last_exported_changes;
+		SpinLockRelease(&shared->mutex);
+
+		/*
+		 * Has the worker exported the file we are waiting for?
+		 */
+		if (last_exported == chgcxt->cc_file_seq_changes)
+			break;
+
+		ConditionVariableSleep(&shared->cv, WAIT_EVENT_REPACK_WORKER_EXPORT);
+	}
+	ConditionVariableCancelSleep();
+
+#ifdef USE_ASSERT_CHECKING
+	/* No file is exported until the worker exports the next one. */
+	SpinLockAcquire(&shared->mutex);
+	Assert(XLogRecPtrIsInvalid(shared->lsn_upto));
+	SpinLockRelease(&shared->mutex);
+#endif
+
+	/* Apply the changes to the new table. */
+	apply_concurrent_changes(chgcxt, range_start, range_end);
+
+	/* Get ready for the next set of changes. */
+	chgcxt->cc_file_seq_changes++;
+}
+
 /*
  * Generate worker's file name into 'fname', which must be of size MAXPGPATH.
  * If relations of the same 'relid' happen to be processed at the same time,
@@ -3878,10 +4313,13 @@ get_initial_snapshot(DecodingWorker *worker)
  * be involved.
  */
 void
-DecodingWorkerFileName(char *fname, Oid relid, uint32 seq)
+DecodingWorkerFileName(char *fname, Oid relid, uint32 seq, bool snapshot)
 {
 	/* The PID is already present in the fileset name, so we needn't add it */
-	snprintf(fname, MAXPGPATH, "%u-%u", relid, seq);
+	if (!snapshot)
+		snprintf(fname, MAXPGPATH, "%u-%u", relid, seq);
+	else
+		snprintf(fname, MAXPGPATH, "%u-%u-snapshot", relid, seq);
 }
 
 /*
diff --git a/src/backend/commands/repack_worker.c b/src/backend/commands/repack_worker.c
index db9ff057cc6..461d60ec0ca 100644
--- a/src/backend/commands/repack_worker.c
+++ b/src/backend/commands/repack_worker.c
@@ -33,8 +33,7 @@
 static void RepackWorkerShutdown(int code, Datum arg);
 static LogicalDecodingContext *repack_setup_logical_decoding(Oid relid);
 static void repack_cleanup_logical_decoding(LogicalDecodingContext *ctx);
-static void export_initial_snapshot(Snapshot snapshot,
-									DecodingWorkerShared *shared);
+static void export_snapshot(Snapshot snapshot, DecodingWorkerShared *shared);
 static bool decode_concurrent_changes(LogicalDecodingContext *ctx,
 									  DecodingWorkerShared *shared);
 
@@ -65,6 +64,8 @@ RepackWorkerMain(Datum main_arg)
 	shm_mq_handle *mqh;
 	LogicalDecodingContext *decoding_ctx;
 	SharedFileSet *sfs;
+	RepackDecodingState *dstate;
+	MemoryContext oldcxt;
 	Snapshot	snapshot;
 
 	am_repack_worker = true;
@@ -118,7 +119,9 @@ RepackWorkerMain(Datum main_arg)
 	 * anything in the shared memory until we have serialized the snapshot.
 	 */
 	SpinLockAcquire(&shared->mutex);
-	Assert(!XLogRecPtrIsValid(shared->lsn_upto));
+	/* Initially we're expected to provide a snapshot and only that. */
+	Assert(shared->snapshot_requested &&
+		   XLogRecPtrIsInvalid(shared->lsn_upto));
 	sfs = &shared->sfs;
 	SpinLockRelease(&shared->mutex);
 
@@ -139,9 +142,25 @@ RepackWorkerMain(Datum main_arg)
 	XactIsoLevel = XACT_REPEATABLE_READ;
 	XactReadOnly = true;
 
-	/* Build the initial snapshot and export it. */
+	/*
+	 * Build the initial snapshot and export it.
+	 *
+	 * Since there is no API to free the "external snapshot", and since such
+	 * snapshot is not guaranteed to be flat (i.e. pfree() is not appropriate)
+	 * the easiest way to clean it up is to use a separate memory context for
+	 * it.
+	 */
+	dstate = (RepackDecodingState *) decoding_ctx->output_writer_private;
+	MemoryContextReset(dstate->snapshot_cxt);
+	oldcxt = MemoryContextSwitchTo(dstate->snapshot_cxt);
 	snapshot = SnapBuildInitialSnapshot(decoding_ctx->snapshot_builder);
-	export_initial_snapshot(snapshot, shared);
+	MemoryContextSwitchTo(oldcxt);
+	export_snapshot(snapshot, shared);
+
+	/*
+	 * Adjust the replication slot's xmin so that VACUUM can do more work.
+	 */
+	LogicalIncreaseXminForSlot(InvalidXLogRecPtr, snapshot->xmin, false);
 
 	/*
 	 * Only historic snapshots should be used now. Do not let us restrict the
@@ -307,7 +326,7 @@ repack_cleanup_logical_decoding(LogicalDecodingContext *ctx)
  * Make snapshot available to the backend that launched the decoding worker.
  */
 static void
-export_initial_snapshot(Snapshot snapshot, DecodingWorkerShared *shared)
+export_snapshot(Snapshot snapshot, DecodingWorkerShared *shared)
 {
 	char		fname[MAXPGPATH];
 	BufFile    *file;
@@ -318,7 +337,9 @@ export_initial_snapshot(Snapshot snapshot, DecodingWorkerShared *shared)
 	snap_space = (char *) palloc(snap_size);
 	SerializeSnapshot(snapshot, snap_space);
 
-	DecodingWorkerFileName(fname, shared->relid, shared->last_exported + 1);
+	DecodingWorkerFileName(fname, shared->relid,
+						   shared->last_exported_snapshot + 1,
+						   true);
 	file = BufFileCreateFileSet(&shared->sfs.fs, fname);
 	/* To make restoration easier, write the snapshot size first. */
 	BufFileWrite(file, &snap_size, sizeof(snap_size));
@@ -328,7 +349,8 @@ export_initial_snapshot(Snapshot snapshot, DecodingWorkerShared *shared)
 
 	/* Increase the counter to tell the backend that the file is available. */
 	SpinLockAcquire(&shared->mutex);
-	shared->last_exported++;
+	shared->last_exported_snapshot++;
+	shared->snapshot_requested = false;
 	SpinLockRelease(&shared->mutex);
 	ConditionVariableSignal(&shared->cv);
 }
@@ -343,6 +365,7 @@ decode_concurrent_changes(LogicalDecodingContext *ctx,
 						  DecodingWorkerShared *shared)
 {
 	RepackDecodingState *dstate;
+	bool		snapshot_requested;
 	XLogRecPtr	lsn_upto;
 	bool		done;
 	char		fname[MAXPGPATH];
@@ -350,11 +373,14 @@ decode_concurrent_changes(LogicalDecodingContext *ctx,
 	dstate = (RepackDecodingState *) ctx->output_writer_private;
 
 	/* Open the output file. */
-	DecodingWorkerFileName(fname, shared->relid, shared->last_exported + 1);
+	DecodingWorkerFileName(fname, shared->relid,
+						   shared->last_exported_changes + 1,
+						   false);
 	dstate->file = BufFileCreateFileSet(&shared->sfs.fs, fname);
 
 	SpinLockAcquire(&shared->mutex);
 	lsn_upto = shared->lsn_upto;
+	snapshot_requested = shared->snapshot_requested;
 	done = shared->done;
 	SpinLockRelease(&shared->mutex);
 
@@ -437,6 +463,7 @@ decode_concurrent_changes(LogicalDecodingContext *ctx,
 		{
 			SpinLockAcquire(&shared->mutex);
 			lsn_upto = shared->lsn_upto;
+			snapshot_requested = shared->snapshot_requested;
 			/* 'done' should be set at the same time as 'lsn_upto' */
 			done = shared->done;
 			SpinLockRelease(&shared->mutex);
@@ -483,9 +510,59 @@ decode_concurrent_changes(LogicalDecodingContext *ctx,
 	 */
 	BufFileClose(dstate->file);
 	dstate->file = NULL;
+
+	/*
+	 * Before publishing the data changes, export the snapshot too if
+	 * requested. Publishing both at once makes sense because both are needed
+	 * at the same time, and it's simpler.
+	 */
+	if (snapshot_requested)
+	{
+		Snapshot	snapshot;
+		MemoryContext oldcxt;
+
+		/* See comments about memory context in RepackWorkerMain(). */
+		MemoryContextReset(dstate->snapshot_cxt);
+		oldcxt = MemoryContextSwitchTo(dstate->snapshot_cxt);
+
+		/*
+		 * SnapBuildInitialSnapshot() assumes invalid XID, so set it. We do
+		 * not use the snapshot, so it's ok.
+		 */
+		MyProc->xmin = InvalidTransactionId;
+		snapshot = SnapBuildInitialSnapshot(ctx->snapshot_builder);
+		MemoryContextSwitchTo(oldcxt);
+		export_snapshot(snapshot, shared);
+
+		/*
+		 * Adjust the replication slot's xmin so that VACUUM can do more work.
+		 */
+		LogicalIncreaseXminForSlot(InvalidXLogRecPtr, snapshot->xmin, false);
+	}
+	else
+	{
+		/*
+		 * If data changes were requested but no following snapshot, we don't
+		 * care about xmin horizon because the heap copying should be done by
+		 * now.
+		 */
+		LogicalIncreaseXminForSlot(InvalidXLogRecPtr, InvalidTransactionId,
+								   false);
+
+	}
+
+	/*
+	 * Make sure the xmin of our slot is taken into account when computing new
+	 * VACUUM horizons.
+	 */
+	ReplicationSlotsComputeRequiredXmin(false);
+
+	/*
+	 * Now increase the counter(s) to announce that the output is available.
+	 */
 	SpinLockAcquire(&shared->mutex);
+	shared->last_exported_changes++;
 	shared->lsn_upto = InvalidXLogRecPtr;
-	shared->last_exported++;
 	SpinLockRelease(&shared->mutex);
 	ConditionVariableSignal(&shared->cv);
 
diff --git a/src/backend/replication/logical/decode.c b/src/backend/replication/logical/decode.c
index c944be4ac83..c3722b5c623 100644
--- a/src/backend/replication/logical/decode.c
+++ b/src/backend/replication/logical/decode.c
@@ -921,6 +921,7 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	xl_heap_insert *xlrec;
 	ReorderBufferChange *change;
 	RelFileLocator target_locator;
+	BlockNumber blknum;
 
 	xlrec = (xl_heap_insert *) XLogRecGetData(r);
 
@@ -932,7 +933,7 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		return;
 
 	/* only interested in our database */
-	XLogRecGetBlockTag(r, 0, &target_locator, NULL, NULL);
+	XLogRecGetBlockTag(r, 0, &target_locator, NULL, &blknum);
 	if (target_locator.dbOid != ctx->slot->data.database)
 		return;
 
@@ -947,7 +948,8 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		change->action = REORDER_BUFFER_CHANGE_INTERNAL_SPEC_INSERT;
 	change->origin_id = XLogRecGetOrigin(r);
 
-	memcpy(&change->data.tp.rlocator, &target_locator, sizeof(RelFileLocator));
+	memcpy(&change->data.tp.rlocator, &target_locator,
+		   sizeof(RelFileLocator));
 
 	tupledata = XLogRecGetBlockData(r, 0, &datalen);
 	tuplelen = datalen - SizeOfHeapHeader;
@@ -957,6 +959,20 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 
 	DecodeXLogTuple(tupledata, datalen, change->data.tp.newtuple);
 
+	/*
+	 * REPACK (CONCURRENTLY) needs block number to check if the corresponding
+	 * part of the table was already copied.  XXX Should we only do this if
+	 * AmRepackWorker()? It might save a few cycles, but not sure it's good to
+	 * leave the fields unset in other cases.
+	 */
+	{
+		HeapTupleHeader header;
+
+		header = change->data.tp.newtuple->t_data;
+		/* offnum is not really needed, but let's set valid pointer. */
+		ItemPointerSet(&header->t_ctid, blknum, xlrec->offnum);
+	}
+
 	change->data.tp.clear_toast_afterwards = true;
 
 	ReorderBufferQueueChange(ctx->reorder, XLogRecGetXid(r), buf->origptr,
@@ -978,11 +994,13 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	ReorderBufferChange *change;
 	char	   *data;
 	RelFileLocator target_locator;
+	BlockNumber new_blknum,
+				old_blknum;
 
 	xlrec = (xl_heap_update *) XLogRecGetData(r);
 
 	/* only interested in our database */
-	XLogRecGetBlockTag(r, 0, &target_locator, NULL, NULL);
+	XLogRecGetBlockTag(r, 0, &target_locator, NULL, &new_blknum);
 	if (target_locator.dbOid != ctx->slot->data.database)
 		return;
 
@@ -990,6 +1008,11 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	if (FilterByOrigin(ctx, XLogRecGetOrigin(r)))
 		return;
 
+	if (XLogRecHasBlockRef(r, 1))
+		XLogRecGetBlockTag(r, 1, NULL, NULL, &old_blknum);
+	else
+		old_blknum = new_blknum;
+
 	change = ReorderBufferAllocChange(ctx->reorder);
 	change->action = REORDER_BUFFER_CHANGE_UPDATE;
 	change->origin_id = XLogRecGetOrigin(r);
@@ -1008,6 +1031,20 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 			ReorderBufferAllocTupleBuf(ctx->reorder, tuplelen);
 
 		DecodeXLogTuple(data, datalen, change->data.tp.newtuple);
+
+		/*
+		 * REPACK (CONCURRENTLY) needs block numbers to check if the
+		 * corresponding part of the table was already copied. XXX Do this
+		 * only if AmRepackWorker()?
+		 */
+		{
+			HeapTupleHeader header;
+
+			header = change->data.tp.newtuple->t_data;
+			/* offnum is not really needed, but let's set valid pointer. */
+			ItemPointerSet(&header->t_ctid, new_blknum, xlrec->new_offnum);
+			change->data.tp.old_blknum = old_blknum;
+		}
 	}
 
 	if (xlrec->flags & XLH_UPDATE_CONTAINS_OLD)
@@ -1044,6 +1081,7 @@ DecodeDelete(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	xl_heap_delete *xlrec;
 	ReorderBufferChange *change;
 	RelFileLocator target_locator;
+	BlockNumber blknum;
 
 	xlrec = (xl_heap_delete *) XLogRecGetData(r);
 
@@ -1057,7 +1095,7 @@ DecodeDelete(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		return;
 
 	/* only interested in our database */
-	XLogRecGetBlockTag(r, 0, &target_locator, NULL, NULL);
+	XLogRecGetBlockTag(r, 0, &target_locator, NULL, &blknum);
 	if (target_locator.dbOid != ctx->slot->data.database)
 		return;
 
@@ -1089,6 +1127,19 @@ DecodeDelete(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 
 		DecodeXLogTuple((char *) xlrec + SizeOfHeapDelete,
 						datalen, change->data.tp.oldtuple);
+
+		/*
+		 * REPACK (CONCURRENTLY) needs block number to check if the
+		 * corresponding part of the table was already copied. XXX Do this
+		 * only if AmRepackWorker()?
+		 */
+		{
+			HeapTupleHeader header;
+
+			header = change->data.tp.oldtuple->t_data;
+			/* offnum is not really needed, but let's set valid pointer. */
+			ItemPointerSet(&header->t_ctid, blknum, xlrec->offnum);
+		}
 	}
 
 	change->data.tp.clear_toast_afterwards = true;
@@ -1148,8 +1199,11 @@ DecodeMultiInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	char	   *tupledata;
 	Size		tuplelen;
 	RelFileLocator rlocator;
+	BlockNumber blknum;
+	bool		isinit;
 
 	xlrec = (xl_heap_multi_insert *) XLogRecGetData(r);
+	isinit = (XLogRecGetInfo(r) & XLOG_HEAP_INIT_PAGE) != 0;
 
 	/*
 	 * Ignore insert records without new tuples.  This happens when a
@@ -1159,7 +1213,7 @@ DecodeMultiInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		return;
 
 	/* only interested in our database */
-	XLogRecGetBlockTag(r, 0, &rlocator, NULL, NULL);
+	XLogRecGetBlockTag(r, 0, &rlocator, NULL, &blknum);
 	if (rlocator.dbOid != ctx->slot->data.database)
 		return;
 
@@ -1227,6 +1281,25 @@ DecodeMultiInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		else
 			change->data.tp.clear_toast_afterwards = false;
 
+		/*
+		 * REPACK (CONCURRENTLY) needs block number to check if the
+		 * corresponding part of the table was already copied.
+		 */
+		if (AmRepackWorker())
+		{
+			OffsetNumber offnum;
+
+			/*
+			 * offnum is not really needed, but let's set valid pointer. (It
+			 * will be invalid anyway if the page was initially empty.)
+			 */
+			if (isinit)
+				offnum = FirstOffsetNumber + i;
+			else
+				offnum = xlrec->offsets[i];
+			ItemPointerSet(&header->t_ctid, blknum, offnum);
+		}
+
 		ReorderBufferQueueChange(ctx->reorder, XLogRecGetXid(r),
 								 buf->origptr, change, false);
 
diff --git a/src/backend/replication/logical/logical.c b/src/backend/replication/logical/logical.c
index 3541fc793e4..fdcaa5036b4 100644
--- a/src/backend/replication/logical/logical.c
+++ b/src/backend/replication/logical/logical.c
@@ -1659,14 +1659,17 @@ update_progress_txn_cb_wrapper(ReorderBuffer *cache, ReorderBufferTXN *txn,
 
 /*
  * Set the required catalog xmin horizon for historic snapshots in the current
- * replication slot.
+ * replication slot if catalog is TRUE, or xmin if catalog is FALSE.
  *
  * Note that in the most cases, we won't be able to immediately use the xmin
  * to increase the xmin horizon: we need to wait till the client has confirmed
- * receiving current_lsn with LogicalConfirmReceivedLocation().
+ * receiving current_lsn with LogicalConfirmReceivedLocation(). However,
+ * catalog=FALSE is only allowed for temporary replication slots, so the
+ * horizon is applied immediately.
  */
 void
-LogicalIncreaseXminForSlot(XLogRecPtr current_lsn, TransactionId xmin)
+LogicalIncreaseXminForSlot(XLogRecPtr current_lsn, TransactionId xmin,
+						   bool catalog)
 {
 	bool		updated_xmin = false;
 	ReplicationSlot *slot;
@@ -1677,6 +1680,27 @@ LogicalIncreaseXminForSlot(XLogRecPtr current_lsn, TransactionId xmin)
 	Assert(slot != NULL);
 
 	SpinLockAcquire(&slot->mutex);
+	if (!catalog)
+	{
+		/*
+		 * The non-catalog horizon can only advance in temporary slots, so
+		 * update it in the shared memory immediately (w/o requiring prior
+		 * saving to disk).
+		 */
+		Assert(slot->data.persistency == RS_TEMPORARY);
+
+		/*
+		 * The horizon must not go backwards, however it's o.k. to become
+		 * invalid.
+		 */
+		Assert(!TransactionIdIsValid(slot->effective_xmin) ||
+			   !TransactionIdIsValid(xmin) ||
+			   TransactionIdFollowsOrEquals(xmin, slot->effective_xmin));
+
+		slot->effective_xmin = xmin;
+		SpinLockRelease(&slot->mutex);
+		return;
+	}
 
 	/*
 	 * don't overwrite if we already have a newer xmin. This can happen if we
diff --git a/src/backend/replication/logical/reorderbuffer.c b/src/backend/replication/logical/reorderbuffer.c
index d06d0d8c9be..cae2b099e69 100644
--- a/src/backend/replication/logical/reorderbuffer.c
+++ b/src/backend/replication/logical/reorderbuffer.c
@@ -3727,6 +3727,40 @@ ReorderBufferXidHasCatalogChanges(ReorderBuffer *rb, TransactionId xid)
 	return rbtxn_has_catalog_changes(txn);
 }
 
+/*
+ * Check if a transaction (or its subtransaction) contains a heap change.
+ */
+bool
+ReorderBufferXidHasHeapChanges(ReorderBuffer *rb, TransactionId xid)
+{
+	ReorderBufferTXN *txn;
+	dlist_iter	iter;
+
+	txn = ReorderBufferTXNByXid(rb, xid, false, NULL, InvalidXLogRecPtr,
+								false);
+	if (txn == NULL)
+		return false;
+
+	dlist_foreach(iter, &txn->changes)
+	{
+		ReorderBufferChange *change;
+
+		change = dlist_container(ReorderBufferChange, node, iter.cur);
+
+		switch (change->action)
+		{
+			case REORDER_BUFFER_CHANGE_INSERT:
+			case REORDER_BUFFER_CHANGE_UPDATE:
+			case REORDER_BUFFER_CHANGE_DELETE:
+				return true;
+			default:
+				break;
+		}
+	}
+
+	return false;
+}
+
 /*
  * ReorderBufferXidHasBaseSnapshot
  *		Have we already set the base snapshot for the given txn/subtxn?
@@ -5224,6 +5258,12 @@ ReorderBufferToastReplace(ReorderBuffer *rb, ReorderBufferTXN *txn,
 	Assert(newtup->t_len <= MaxHeapTupleSize);
 	Assert(newtup->t_data == (HeapTupleHeader) ((char *) newtup + HEAPTUPLESIZE));
 
+	/*
+	 * Preserve TID - REPACK relies on it when dealing with block ranges. XXX
+	 * Shouldn't we add a new field to ReorderBufferChange instead?
+	 */
+	tmphtup->t_data->t_ctid = newtup->t_data->t_ctid;
+
 	memcpy(newtup->t_data, tmphtup->t_data, tmphtup->t_len);
 	newtup->t_len = tmphtup->t_len;
 
diff --git a/src/backend/replication/logical/snapbuild.c b/src/backend/replication/logical/snapbuild.c
index b1e37ef6792..dd7d584e15b 100644
--- a/src/backend/replication/logical/snapbuild.c
+++ b/src/backend/replication/logical/snapbuild.c
@@ -128,6 +128,7 @@
 #include "access/heapam_xlog.h"
 #include "access/transam.h"
 #include "access/xact.h"
+#include "commands/repack.h"
 #include "common/file_utils.h"
 #include "miscadmin.h"
 #include "pgstat.h"
@@ -983,6 +984,13 @@ SnapBuildCommitTxn(SnapBuild *builder, XLogRecPtr lsn, TransactionId xid,
 		}
 	}
 
+	/*
+	 * REPACK decoding worker may need timetravel anytime. It takes
+	 * responsibility for tracking transaction commits, see below.
+	 */
+	else if (AmRepackWorker())
+		needs_timetravel = true;
+
 	for (nxact = 0; nxact < nsubxacts; nxact++)
 	{
 		TransactionId subxid = subxacts[nxact];
@@ -990,8 +998,12 @@ SnapBuildCommitTxn(SnapBuild *builder, XLogRecPtr lsn, TransactionId xid,
 		/*
 		 * Add subtransaction to base snapshot if catalog modifying, we don't
 		 * distinguish to toplevel transactions there.
+		 *
+		 * See comments on REPACK worker below.
 		 */
-		if (SnapBuildXidHasCatalogChanges(builder, subxid, xinfo))
+		if (SnapBuildXidHasCatalogChanges(builder, subxid, xinfo) ||
+			(AmRepackWorker() &&
+			 ReorderBufferXidHasHeapChanges(builder->reorder, xid)))
 		{
 			sub_needs_timetravel = true;
 			needs_snapshot = true;
@@ -1019,8 +1031,18 @@ SnapBuildCommitTxn(SnapBuild *builder, XLogRecPtr lsn, TransactionId xid,
 		}
 	}
 
-	/* if top-level modified catalog, it'll need a snapshot */
-	if (SnapBuildXidHasCatalogChanges(builder, xid, xinfo))
+	/*
+	 * If top-level modified catalog, it'll need a snapshot.
+	 *
+	 * If we're decoding changes on behalf of REPACK (CONCURRENTLY), only
+	 * changes of the relation being processed are decoded - see
+	 * heap_decode(). Thus any heap change we find here must belong to that
+	 * relation. Add the transaction so that we can keep building snapshots to
+	 * scan that relation.
+	 */
+	if (SnapBuildXidHasCatalogChanges(builder, xid, xinfo) ||
+		(AmRepackWorker() &&
+		 ReorderBufferXidHasHeapChanges(builder->reorder, xid)))
 	{
 		elog(DEBUG2, "found top level transaction %u, with catalog changes",
 			 xid);
@@ -1188,7 +1210,7 @@ SnapBuildProcessRunningXacts(SnapBuild *builder, XLogRecPtr lsn, xl_running_xact
 		xmin = running->oldestRunningXid;
 	elog(DEBUG3, "xmin: %u, xmax: %u, oldest running: %u, oldest xmin: %u",
 		 builder->xmin, builder->xmax, running->oldestRunningXid, xmin);
-	LogicalIncreaseXminForSlot(lsn, xmin);
+	LogicalIncreaseXminForSlot(lsn, xmin, true);
 
 	/*
 	 * Also tell the slot where we can restart decoding from. We don't want to
diff --git a/src/backend/replication/pgrepack/pgrepack.c b/src/backend/replication/pgrepack/pgrepack.c
index 5c5095bde4e..239796b3335 100644
--- a/src/backend/replication/pgrepack/pgrepack.c
+++ b/src/backend/replication/pgrepack/pgrepack.c
@@ -33,7 +33,8 @@ static void repack_commit_txn(LogicalDecodingContext *ctx,
 static void repack_process_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
 								  Relation relation, ReorderBufferChange *change);
 static void repack_store_change(LogicalDecodingContext *ctx, Relation relation,
-								ConcurrentChangeKind kind, HeapTuple tuple);
+								ConcurrentChangeKind kind, HeapTuple tuple,
+								BlockNumber old_blknum);
 
 void
 _PG_output_plugin_init(OutputPluginCallbacks *cb)
@@ -67,6 +68,9 @@ repack_startup(LogicalDecodingContext *ctx, OutputPluginOptions *opt,
 	dstate->change_cxt = AllocSetContextCreate(ctx->context,
 											   "REPACK - change",
 											   ALLOCSET_DEFAULT_SIZES);
+	dstate->snapshot_cxt = AllocSetContextCreate(ctx->context,
+												 "REPACK - snapshot",
+												 ALLOCSET_DEFAULT_SIZES);
 	/* repack_setup_logical_decoding fills in the rest */
 	ctx->output_writer_private = dstate;
 
@@ -136,7 +140,8 @@ repack_process_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
 				if (newtuple == NULL)
 					elog(ERROR, "incomplete insert info");
 
-				repack_store_change(ctx, relation, CHANGE_INSERT, newtuple);
+				repack_store_change(ctx, relation, CHANGE_INSERT, newtuple,
+									InvalidBlockNumber);
 			}
 			break;
 		case REORDER_BUFFER_CHANGE_UPDATE:
@@ -151,9 +156,11 @@ repack_process_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
 					elog(ERROR, "incomplete update info");
 
 				if (oldtuple != NULL)
-					repack_store_change(ctx, relation, CHANGE_UPDATE_OLD, oldtuple);
+					repack_store_change(ctx, relation, CHANGE_UPDATE_OLD, oldtuple,
+										InvalidBlockNumber);
 
-				repack_store_change(ctx, relation, CHANGE_UPDATE_NEW, newtuple);
+				repack_store_change(ctx, relation, CHANGE_UPDATE_NEW, newtuple,
+									change->data.tp.old_blknum);
 			}
 			break;
 		case REORDER_BUFFER_CHANGE_DELETE:
@@ -165,7 +172,8 @@ repack_process_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
 				if (oldtuple == NULL)
 					elog(ERROR, "incomplete delete info");
 
-				repack_store_change(ctx, relation, CHANGE_DELETE, oldtuple);
+				repack_store_change(ctx, relation, CHANGE_DELETE, oldtuple,
+									InvalidBlockNumber);
 			}
 			break;
 		default:
@@ -192,7 +200,8 @@ repack_process_change(LogicalDecodingContext *ctx, ReorderBufferTXN *txn,
  */
 static void
 repack_store_change(LogicalDecodingContext *ctx, Relation relation,
-					ConcurrentChangeKind kind, HeapTuple tuple)
+					ConcurrentChangeKind kind, HeapTuple tuple,
+					BlockNumber old_blknum)
 {
 	RepackDecodingState *dstate;
 	MemoryContext oldcxt;
@@ -288,6 +297,9 @@ repack_store_change(LogicalDecodingContext *ctx, Relation relation,
 	 */
 	BufFileWrite(file, &tuple->t_len, sizeof(tuple->t_len));
 	BufFileWrite(file, tuple->t_data, tuple->t_len);
+	/* If old_blknum is specified, write it too. */
+	if (old_blknum != InvalidBlockNumber)
+		BufFileWrite(file, &old_blknum, sizeof(old_blknum));
 
 	/* Then, write the number of external attributes we found. */
 	natt_ext = list_length(attrs_ext);
diff --git a/src/backend/utils/misc/guc_parameters.dat b/src/backend/utils/misc/guc_parameters.dat
index d421cdbde76..ca2783107b1 100644
--- a/src/backend/utils/misc/guc_parameters.dat
+++ b/src/backend/utils/misc/guc_parameters.dat
@@ -2544,6 +2544,16 @@
   boot_val => 'true',
 },
 
+# TODO Tune boot_val, 1024 is probably too low.
+{ name => 'repack_snapshot_after', type => 'int', context => 'PGC_USERSET', group => 'DEVELOPER_OPTIONS',
+  short_desc => 'Number of pages REPACK (CONCURRENTLY) can read using a single snapshot.',
+  flags => 'GUC_UNIT_BLOCKS | GUC_NOT_IN_SAMPLE',
+  variable => 'repack_pages_per_snapshot',
+  boot_val => '1024',
+  min => '1',
+  max => 'INT_MAX',
+}
+
 { name => 'reserved_connections', type => 'int', context => 'PGC_POSTMASTER', group => 'CONN_AUTH_SETTINGS',
   short_desc => 'Sets the number of connection slots reserved for roles with privileges of pg_use_reserved_connections.',
   variable => 'ReservedConnections',
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index 90aa374b3ec..6d8106d9445 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -44,6 +44,7 @@
 #include "commands/async.h"
 #include "commands/extension.h"
 #include "commands/event_trigger.h"
+#include "commands/repack.h"
 #include "commands/tablespace.h"
 #include "commands/trigger.h"
 #include "commands/user.h"
diff --git a/src/include/access/tableam.h b/src/include/access/tableam.h
index f2c36696bca..132248c5d43 100644
--- a/src/include/access/tableam.h
+++ b/src/include/access/tableam.h
@@ -666,12 +666,12 @@ typedef struct TableAmRoutine
 											  Relation OldIndex,
 											  bool use_sort,
 											  TransactionId OldestXmin,
-											  Snapshot snapshot,
 											  TransactionId *xid_cutoff,
 											  MultiXactId *multi_cutoff,
 											  double *num_tuples,
 											  double *tups_vacuumed,
-											  double *tups_recently_dead);
+											  double *tups_recently_dead,
+											  void *tableam_data);
 
 	/*
 	 * React to VACUUM command on the relation. The VACUUM can be triggered by
@@ -1733,8 +1733,6 @@ table_relation_copy_data(Relation rel, const RelFileLocator *newrlocator)
  *   not needed for the relation's AM
  * - *xid_cutoff - ditto
  * - *multi_cutoff - ditto
- * - snapshot - if != NULL, ignore data changes done by transactions that this
- *	 (MVCC) snapshot considers still in-progress or in the future.
  *
  * Output parameters:
  * - *xid_cutoff - rel's new relfrozenxid value, may be invalid
@@ -1747,19 +1745,19 @@ table_relation_copy_for_cluster(Relation OldTable, Relation NewTable,
 								Relation OldIndex,
 								bool use_sort,
 								TransactionId OldestXmin,
-								Snapshot snapshot,
 								TransactionId *xid_cutoff,
 								MultiXactId *multi_cutoff,
 								double *num_tuples,
 								double *tups_vacuumed,
-								double *tups_recently_dead)
+								double *tups_recently_dead,
+								void *tableam_data)
 {
 	OldTable->rd_tableam->relation_copy_for_cluster(OldTable, NewTable, OldIndex,
 													use_sort, OldestXmin,
-													snapshot,
 													xid_cutoff, multi_cutoff,
 													num_tuples, tups_vacuumed,
-													tups_recently_dead);
+													tups_recently_dead,
+													tableam_data);
 }
 
 /*
diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h
index 07f887e99f6..8af73f8c81f 100644
--- a/src/include/commands/repack.h
+++ b/src/include/commands/repack.h
@@ -17,11 +17,14 @@
 
 #include "access/hio.h"
 #include "access/skey.h"
+#include "access/xlogdefs.h"
 #include "nodes/execnodes.h"
 #include "nodes/parsenodes.h"
 #include "parser/parse_node.h"
+#include "storage/block.h"
 #include "storage/lockdefs.h"
 #include "utils/relcache.h"
+#include "utils/snapshot.h"
 
 
 /* flag bits for ClusterParams->options */
@@ -66,13 +69,14 @@ typedef struct RepackDest
 
 	/* The latest column we need to deform to have the tuple identity */
 	AttrNumber	last_key_attno;
-} RepackDest;
 
-/*
- * The first file exported by the decoding worker must contain a snapshot, the
- * following ones contain the data changes.
- */
-#define WORKER_FILE_SNAPSHOT	0
+	/*
+	 * Range of blocks in the old table the contents of this table comes from.
+	 * Note that range_end is the first block of the next range.
+	 */
+	BlockNumber range_start;
+	BlockNumber range_end;
+} RepackDest;
 
 /*
  * Information needed to apply concurrent data changes.
@@ -84,10 +88,42 @@ typedef struct ChangeContext
 	/* The destination table. */
 	RepackDest	cc_dest;
 
-	/* Sequential number of the file containing the changes. */
-	int			cc_file_seq;
+	/* Sequential number of the file containing snapshot. */
+	int			cc_file_seq_snapshot;
+	/* Sequential number of the file containing data changes. */
+	int			cc_file_seq_changes;
+
+	/*
+	 * Auxiliary table to store ordered tuples temporarily.
+	 *
+	 * When the new relation needs to be clustered, we use this table instead
+	 * of tuplesort. The problem with a tuplesort is that data changes need to
+	 * be applied at range boundary (see heapam_relation_copy_for_cluster()
+	 * for more information), however it's not possible to look-up and change
+	 * tuples in tuplestore.
+	 *
+	 * Once the contents of the REPACKed table has been copied into the
+	 * auxiliary table, we build the clustering index (unless it's the same as
+	 * the identity index) and scan it to get the tuple in the desired order.
+	 * XXX Is it worth putting the contents into a tuplestore and sorting it?
+	 * Not sure, it'd require disk space for one more copy and the copying
+	 * itself is not free.
+	 *
+	 * TODO 1) make the tables unlogged, 2) if REPACK locks the TOAST relation
+	 * too (not sure it does) try to preserve TOAST pointers, instead of
+	 * storing them to TOAST relations of these tables, 3) Check that the
+	 * tables are dropped on transaction abort.
+	 */
+	RepackDest *cc_dest_aux;
+
+	/*
+	 * The index that defines ordering of the old table.
+	 */
+	Oid			cc_clustering_index;
 } ChangeContext;
 
+extern PGDLLIMPORT int repack_pages_per_snapshot;
+
 extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel);
 
 extern void cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid,
@@ -98,9 +134,8 @@ extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal);
 
 extern Oid	make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
 						  char relpersistence, LOCKMODE lockmode);
-extern void heap_insert_for_repack(Relation rel, TupleTableSlot *src,
-								   TupleTableSlot *reform,
-								   BulkInsertStateData *bistate);
+extern void heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
+								   TupleTableSlot *reform);
 extern bool tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc);
 extern void clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform);
 extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
@@ -112,7 +147,12 @@ extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 							 TransactionId frozenXid,
 							 MultiXactId cutoffMulti,
 							 char newrelpersistence);
-
+extern Snapshot repack_get_snapshot(ChangeContext *chgcxt);
+extern void repack_process_concurrent_changes(ChangeContext *chgcxt,
+											  XLogRecPtr end_of_wal,
+											  BlockNumber range_start,
+											  BlockNumber range_end,
+											  bool request_snapshot, bool done);
 extern void HandleRepackMessageInterrupt(void);
 extern void ProcessRepackMessages(void);
 
diff --git a/src/include/commands/repack_internal.h b/src/include/commands/repack_internal.h
index 42111aa4ae3..8003e999864 100644
--- a/src/include/commands/repack_internal.h
+++ b/src/include/commands/repack_internal.h
@@ -44,6 +44,8 @@ typedef struct RepackDecodingState
 
 	/* Per-change memory context. */
 	MemoryContext change_cxt;
+	/* Per-snapshot memory context. */
+	MemoryContext snapshot_cxt;
 
 	/* A tuple slot used to pass tuples back and forth */
 	TupleTableSlot *slot;
@@ -67,6 +69,9 @@ typedef struct DecodingWorkerShared
 	/* Is the decoding initialized? */
 	bool		initialized;
 
+	/* Set to request a snapshot. */
+	bool		snapshot_requested;
+
 	/*
 	 * Once the worker has reached this LSN, it should close the current
 	 * output file and either create a new one or exit, according to the field
@@ -74,6 +79,8 @@ typedef struct DecodingWorkerShared
 	 * the WAL available and keep checking this field. It is ok if the worker
 	 * had already decoded records whose LSN is >= lsn_upto before this field
 	 * has been set.
+	 *
+	 * Set a valid LSN to request data changes.
 	 */
 	XLogRecPtr	lsn_upto;
 
@@ -84,7 +91,8 @@ typedef struct DecodingWorkerShared
 	SharedFileSet sfs;
 
 	/* Number of the last file exported by the worker. */
-	int			last_exported;
+	int			last_exported_snapshot;
+	int			last_exported_changes;
 
 	/* Synchronize access to the fields above. */
 	slock_t		mutex;
@@ -116,7 +124,8 @@ typedef struct DecodingWorkerShared
 	char		error_queue[FLEXIBLE_ARRAY_MEMBER];
 } DecodingWorkerShared;
 
-extern void DecodingWorkerFileName(char *fname, Oid relid, uint32 seq);
+extern void DecodingWorkerFileName(char *fname, Oid relid, uint32 seq,
+								   bool snapshot);
 
 
 #endif							/* REPACK_INTERNAL_H */
diff --git a/src/include/replication/logical.h b/src/include/replication/logical.h
index 6e0b7628001..37315a424e0 100644
--- a/src/include/replication/logical.h
+++ b/src/include/replication/logical.h
@@ -138,7 +138,7 @@ extern bool DecodingContextReady(LogicalDecodingContext *ctx);
 extern void FreeDecodingContext(LogicalDecodingContext *ctx);
 
 extern void LogicalIncreaseXminForSlot(XLogRecPtr current_lsn,
-									   TransactionId xmin);
+									   TransactionId xmin, bool catalog);
 extern void LogicalIncreaseRestartDecodingForSlot(XLogRecPtr current_lsn,
 												  XLogRecPtr restart_lsn);
 extern void LogicalConfirmReceivedLocation(XLogRecPtr lsn);
diff --git a/src/include/replication/reorderbuffer.h b/src/include/replication/reorderbuffer.h
index ff825e4b7b2..cdefc4808df 100644
--- a/src/include/replication/reorderbuffer.h
+++ b/src/include/replication/reorderbuffer.h
@@ -104,6 +104,12 @@ typedef struct ReorderBufferChange
 			HeapTuple	oldtuple;
 			/* valid for INSERT || UPDATE */
 			HeapTuple	newtuple;
+
+			/*
+			 * valid for UPDATE - this is the physical location of the old
+			 * tuple version, valid even if 'oldtuple' is NULL.
+			 */
+			BlockNumber old_blknum;
 		}			tp;
 
 		/*
@@ -763,6 +769,7 @@ extern void ReorderBufferProcessXid(ReorderBuffer *rb, TransactionId xid, XLogRe
 
 extern void ReorderBufferXidSetCatalogChanges(ReorderBuffer *rb, TransactionId xid, XLogRecPtr lsn);
 extern bool ReorderBufferXidHasCatalogChanges(ReorderBuffer *rb, TransactionId xid);
+extern bool ReorderBufferXidHasHeapChanges(ReorderBuffer *rb, TransactionId xid);
 extern bool ReorderBufferXidHasBaseSnapshot(ReorderBuffer *rb, TransactionId xid);
 
 extern bool ReorderBufferRememberPrepareInfo(ReorderBuffer *rb, TransactionId xid,
diff --git a/src/test/modules/injection_points/Makefile b/src/test/modules/injection_points/Makefile
index c01d2fb095c..9c942599c49 100644
--- a/src/test/modules/injection_points/Makefile
+++ b/src/test/modules/injection_points/Makefile
@@ -15,6 +15,7 @@ REGRESS_OPTS = --dlpath=$(top_builddir)/src/test/regress
 ISOLATION = basic \
 	    inplace \
 	    repack \
+	    repack_snapshots \
 	    repack_temporal \
 	    repack_temporal_multirange \
 	    repack_toast \
diff --git a/src/test/modules/injection_points/expected/repack_snapshots.out b/src/test/modules/injection_points/expected/repack_snapshots.out
new file mode 100644
index 00000000000..98bac9ea882
--- /dev/null
+++ b/src/test/modules/injection_points/expected/repack_snapshots.out
@@ -0,0 +1,401 @@
+Parsed test spec with 2 sessions
+
+starting permutation: load repack change_new_beyond change_old_beyond check2 wakeup_new_range check1
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step load: 
+	SELECT load(1);
+
+load
+----
+    
+(1 row)
+
+step repack: 
+	REPACK (CONCURRENTLY) repack_test;
+ <waiting ...>
+step change_new_beyond: 
+	INSERT INTO repack_test
+	SELECT max(i) + 1, gen_external()
+	FROM repack_test
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT min(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        0|        1
+(1 row)
+
+step change_old_beyond: 
+	DELETE FROM repack_test
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        1|        1
+(1 row)
+
+step check2: 
+	INSERT INTO data_s2(i, j)
+	SELECT i, j FROM repack_test;
+
+step wakeup_new_range: 
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step repack: <... completed>
+step check1: 
+	INSERT INTO data_s1(i, j)
+	SELECT i, j FROM repack_test;
+
+	SELECT count(*) > 0 FROM repack_test;
+
+	SELECT count(*)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	WHERE d1.i ISNULL OR d2.i ISNULL;
+
+?column?
+--------
+t       
+(1 row)
+
+count
+-----
+    0
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: load2 load2_vacuum repack change_old_beyond2 check2 wakeup_new_range wait_new_range wakeup_new_range check1
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step load2: 
+	SELECT load(2);
+
+	DELETE FROM repack_test WHERE i < 100;
+
+load
+----
+    
+(1 row)
+
+step load2_vacuum: 
+	VACUUM repack_test;
+
+step repack: 
+	REPACK (CONCURRENTLY) repack_test;
+ <waiting ...>
+step change_old_beyond2: 
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test WHERE tid_block(ctid) = 1)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block|tid_block
+---------+---------
+        1|        0
+(1 row)
+
+step check2: 
+	INSERT INTO data_s2(i, j)
+	SELECT i, j FROM repack_test;
+
+step wakeup_new_range: 
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_new_range: 
+	SELECT wait_for_blocks_scanned('repack_test'::regclass, 3);
+
+wait_for_blocks_scanned
+-----------------------
+                       
+(1 row)
+
+step wakeup_new_range: 
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step repack: <... completed>
+step check1: 
+	INSERT INTO data_s1(i, j)
+	SELECT i, j FROM repack_test;
+
+	SELECT count(*) > 0 FROM repack_test;
+
+	SELECT count(*)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	WHERE d1.i ISNULL OR d2.i ISNULL;
+
+?column?
+--------
+t       
+(1 row)
+
+count
+-----
+    0
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: load repack_pkey change_new_beyond change_old_beyond check2 wakeup_new_range check1 check1_order_asc
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step load: 
+	SELECT load(1);
+
+load
+----
+    
+(1 row)
+
+step repack_pkey: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step change_new_beyond: 
+	INSERT INTO repack_test
+	SELECT max(i) + 1, gen_external()
+	FROM repack_test
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT min(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        0|        1
+(1 row)
+
+step change_old_beyond: 
+	DELETE FROM repack_test
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        1|        1
+(1 row)
+
+step check2: 
+	INSERT INTO data_s2(i, j)
+	SELECT i, j FROM repack_test;
+
+step wakeup_new_range: 
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step repack_pkey: <... completed>
+step check1: 
+	INSERT INTO data_s1(i, j)
+	SELECT i, j FROM repack_test;
+
+	SELECT count(*) > 0 FROM repack_test;
+
+	SELECT count(*)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	WHERE d1.i ISNULL OR d2.i ISNULL;
+
+?column?
+--------
+t       
+(1 row)
+
+count
+-----
+    0
+(1 row)
+
+step check1_order_asc: 
+	SELECT i FROM repack_test LIMIT 10;
+
+ i
+--
+ 2
+ 3
+ 4
+ 5
+ 6
+ 7
+ 8
+ 9
+10
+11
+(10 rows)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: load repack_other_index change_new_beyond change_old_beyond check2 wakeup_new_range check1 check1_order_desc
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step load: 
+	SELECT load(1);
+
+load
+----
+    
+(1 row)
+
+step repack_other_index: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_i_idx;
+ <waiting ...>
+step change_new_beyond: 
+	INSERT INTO repack_test
+	SELECT max(i) + 1, gen_external()
+	FROM repack_test
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT min(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        0|        1
+(1 row)
+
+step change_old_beyond: 
+	DELETE FROM repack_test
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+
+tid_block
+---------
+        1
+(1 row)
+
+tid_block|tid_block
+---------+---------
+        1|        1
+(1 row)
+
+step check2: 
+	INSERT INTO data_s2(i, j)
+	SELECT i, j FROM repack_test;
+
+step wakeup_new_range: 
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step repack_other_index: <... completed>
+step check1: 
+	INSERT INTO data_s1(i, j)
+	SELECT i, j FROM repack_test;
+
+	SELECT count(*) > 0 FROM repack_test;
+
+	SELECT count(*)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	WHERE d1.i ISNULL OR d2.i ISNULL;
+
+?column?
+--------
+t       
+(1 row)
+
+count
+-----
+    0
+(1 row)
+
+step check1_order_desc: 
+	WITH tmp(diff) as (
+		SELECT i - lag(i, 1, 10000) OVER (ORDER BY ctid)
+		FROM repack_test
+		LIMIT 10)
+	SELECT * FROM tmp WHERE diff > -1;
+
+diff
+----
+(0 rows)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/meson.build b/src/test/modules/injection_points/meson.build
index 59dba1cb023..d432a6b8f76 100644
--- a/src/test/modules/injection_points/meson.build
+++ b/src/test/modules/injection_points/meson.build
@@ -46,6 +46,7 @@ tests += {
       'basic',
       'inplace',
       'repack',
+      'repack_snapshots',
       'repack_temporal',
       'repack_temporal_multirange',
       'repack_toast',
diff --git a/src/test/modules/injection_points/specs/repack_snapshots.spec b/src/test/modules/injection_points/specs/repack_snapshots.spec
new file mode 100644
index 00000000000..23b478fd6f4
--- /dev/null
+++ b/src/test/modules/injection_points/specs/repack_snapshots.spec
@@ -0,0 +1,272 @@
+# REPACK (CONCURRENTLY) - use one snapshot per block range.
+setup
+{
+	CREATE EXTENSION injection_points;
+
+	CREATE TABLE repack_test(i int PRIMARY KEY, j text);
+	CREATE INDEX ON repack_test(i DESC);
+	CREATE TABLE relfilenodes(node oid);
+
+	CREATE TABLE data_s1(i int, j text);
+	CREATE TABLE data_s2(i int, j text);
+
+	-- Keep inserting tuples into repack_test until several tuples need to
+        -- be inserted into block number last_block.
+	CREATE FUNCTION load(last_block int)
+	RETURNS void
+	LANGUAGE 'plpgsql'
+	AS $$
+	    DECLARE
+		cnt	int;
+	    BEGIN
+		INSERT INTO repack_test VALUES (1, gen_external());
+
+		LOOP
+		    WITH max(m) AS (SELECT max(i) FROM repack_test)
+		    INSERT INTO repack_test(i, j)
+		    SELECT m + x, gen_external()
+		    FROM generate_series(1, 100) s(x), max;
+
+		    SELECT count(*)
+		    FROM repack_test WHERE tid_block(ctid) = last_block
+		    INTO cnt;
+
+		    IF cnt >= 10 THEN
+			EXIT;
+		    END IF;
+		END LOOP;
+	    END;
+	$$;
+
+	-- Generate a string of random characters that is not likely to be
+	-- compressed, but is big enough to be stored externally.
+	CREATE FUNCTION gen_external()
+	RETURNS text
+	LANGUAGE sql as $$
+		SELECT string_agg(chr(65 + trunc(25 * random())::int), '')
+		FROM generate_series(1, 2048) s(x);
+	$$;
+
+	-- Wait until the number of blocks involved in the scan by REPACK has
+	-- reached _blocks.
+	CREATE FUNCTION wait_for_blocks_scanned(_relid oid, _blocks int)
+	RETURNS void
+	LANGUAGE 'plpgsql'
+	AS $$
+		DECLARE
+			cnt	int;
+		BEGIN
+			LOOP
+				PERFORM pg_stat_clear_snapshot();
+
+				SELECT heap_blks_scanned
+				FROM pg_stat_progress_repack
+				WHERE relid = _relid
+				INTO cnt;
+
+				IF cnt >= _blocks THEN
+					EXIT;
+				END IF;
+
+				PERFORM pg_sleep(0.1);
+			END LOOP;
+		END;
+	$$;
+}
+
+teardown
+{
+	DROP TABLE repack_test;
+	DROP EXTENSION injection_points;
+
+	DROP TABLE relfilenodes;
+	DROP TABLE data_s1;
+	DROP TABLE data_s2;
+
+	DROP FUNCTION load(int);
+	DROP FUNCTION gen_external();
+	DROP FUNCTION wait_for_blocks_scanned(oid, int);
+}
+
+session s1
+setup
+{
+	SET repack_snapshot_after = 1;
+
+	SELECT injection_points_set_local();
+	SELECT injection_points_attach('repack-concurrently-new-range', 'wait');
+}
+# The most practical way to test the corner cases is to set range size to 1
+# block. To initialize, insert new tuples until we have several tuples in the
+# 2nd block.
+step load
+{
+	SELECT load(1);
+}
+# Start the initial load and wait when the first range has been completed.
+step repack
+{
+	REPACK (CONCURRENTLY) repack_test;
+}
+# The same, but with clustering.
+step repack_pkey
+{
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+}
+# Clustering by other than the identity index.
+step repack_other_index
+{
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_i_idx;
+}
+# Check the table from the perspective of s1.
+step check1
+{
+	INSERT INTO data_s1(i, j)
+	SELECT i, j FROM repack_test;
+
+	SELECT count(*) > 0 FROM repack_test;
+
+	SELECT count(*)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	WHERE d1.i ISNULL OR d2.i ISNULL;
+}
+# Check the ordering where appropriate. We don't know the exact number of
+# rows, so just check a sample. (The replayed concurrent changes are not
+# ordered, but those shouldn't fit into the first 10 rows.)
+step check1_order_asc
+{
+	SELECT i FROM repack_test LIMIT 10;
+}
+# Due to the special way of loading the data (see the load() function above)
+# we don't know the maximum value. To make the test output deterministic,
+# check for cases where the current row is not lower than the previous row.
+step check1_order_desc
+{
+	WITH tmp(diff) as (
+		SELECT i - lag(i, 1, 10000) OVER (ORDER BY ctid)
+		FROM repack_test
+		LIMIT 10)
+	SELECT * FROM tmp WHERE diff > -1;
+}
+teardown
+{
+	SELECT injection_points_detach('repack-concurrently-new-range');
+}
+
+session s2
+# Test processing of changes such that the new tuple is beyond the current
+# range. Specifically for UPDATE, the old tuple should be in the current
+# range. So when applying it, we have to convert it to DELETE.
+step change_new_beyond
+{
+	INSERT INTO repack_test
+	SELECT max(i) + 1, gen_external()
+	FROM repack_test
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT min(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+}
+# Test processing of changes such that the old tuple is beyond the current
+# range. The UPDATE puts also the new tuple beyond the current range.
+step change_old_beyond
+{
+	DELETE FROM repack_test
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(ctid);
+
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+}
+# Arrange for UPDATE to put the new tuples into block 0. In particular, fill
+# the block 1 and delete some tuples from block 1.
+step load2
+{
+	SELECT load(2);
+
+	DELETE FROM repack_test WHERE i < 100;
+}
+# Logically this belongs to the previous step, however VACUUM cannot run
+# inside a transaction block.
+step load2_vacuum
+{
+	VACUUM repack_test;
+}
+# This UPDATE should put the new tuple into block 0 (the current range). So
+# when replaying it, we have to convert it to INSERT. That includes fetching
+# the old tuple's TOAST from the TOAST table because the old tuple is not
+# available during the replay.
+step change_old_beyond2
+{
+	UPDATE repack_test SET j = gen_external()
+	WHERE ctid=(SELECT max(ctid) FROM repack_test WHERE tid_block(ctid) = 1)
+	RETURNING tid_block(OLD.ctid), tid_block(NEW.ctid);
+}
+# Check the table from the perspective of s4.
+step check2
+{
+	INSERT INTO data_s2(i, j)
+	SELECT i, j FROM repack_test;
+}
+step wakeup_new_range
+{
+	SELECT injection_points_wakeup('repack-concurrently-new-range');
+}
+# This step is used to make sure that REPACK is waiting on an injection point
+# before finalizing the 2nd block. However, the coding is such that the
+# counter is incremented at the top of the loop (see
+# heapam_relation_copy_for_cluster()), so we wait until it's at least 3.
+step wait_new_range
+{
+	SELECT wait_for_blocks_scanned('repack_test'::regclass, 3);
+}
+
+# Test if snapshots are used correctly to scan block ranges.
+permutation
+	load
+	repack
+	change_new_beyond
+	change_old_beyond
+	check2
+	wakeup_new_range
+	check1
+
+# Special attention is needed to update tuple in block 1 so that the new tuple
+# appears in block 0. The preparation includes VACUUM, which in turn cannot
+# proceed while REPACK is in progress. That's why we need a separate
+# permutation. Note that two wake-ups are needed as we have two range
+# boundaries now. However we need to wait in between to make sure that the
+# second waiting started before we try to wake it up.
+permutation
+	load2
+	load2_vacuum
+	repack
+	change_old_beyond2
+	check2
+	wakeup_new_range
+	wait_new_range
+	wakeup_new_range
+	check1
+
+# The first permutation with identity index as the clustering index.
+permutation
+	load
+	repack_pkey
+	change_new_beyond
+	change_old_beyond
+	check2
+	wakeup_new_range
+	check1
+	check1_order_asc
+# The first permutation with another clustering index.
+permutation
+	load
+	repack_other_index
+	change_new_beyond
+	change_old_beyond
+	check2
+	wakeup_new_range
+	check1
+	check1_order_desc
-- 
2.52.0

  [text/x-diff] v02-0005-Simplify-the-way-restrictions-are-imposed-on-index-f.patch (22.8K, ../108776.1784105248@localhost/6-v02-0005-Simplify-the-way-restrictions-are-imposed-on-index-f.patch)
  download | inline diff:
From 2a2d091f00911b2cb4fd449aff6d4ecda321894b Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 15 Jul 2026 10:37:12 +0200
Subject: [PATCH 5/8] Simplify the way restrictions are imposed on index
 functions.

Whenever we expect possible execution of index functions, we need to make sure
that they execute with the appropriate privileges. Also, the core should not
see (and use) values of GUC parameters introduced by the index functions.

This patch introduces functions enable_index_build_security() and
disable_index_build_security() which make the security measures less
verbose. It's needed for the upcoming enhancements of REPACK (CONCURRENTLY),
but looks like useful refactoring anyway.
---
 src/backend/access/brin/brin.c   | 32 ++++----------
 src/backend/catalog/index.c      | 72 ++++++++++++++++++--------------
 src/backend/commands/analyze.c   | 21 +++-------
 src/backend/commands/indexcmds.c | 49 +++++++---------------
 src/backend/commands/repack.c    | 24 +++--------
 src/backend/commands/tablecmds.c | 24 +++--------
 src/backend/commands/vacuum.c    | 23 +++-------
 src/include/catalog/index.h      | 10 +++++
 src/tools/pgindent/typedefs.list |  1 +
 9 files changed, 95 insertions(+), 161 deletions(-)

diff --git a/src/backend/access/brin/brin.c b/src/backend/access/brin/brin.c
index bdb30752e09..2359bffa518 100644
--- a/src/backend/access/brin/brin.c
+++ b/src/backend/access/brin/brin.c
@@ -1391,10 +1391,8 @@ brin_summarize_range(PG_FUNCTION_ARGS)
 	Oid			heapoid;
 	Relation	indexRel;
 	Relation	heapRel;
-	Oid			save_userid;
-	int			save_sec_context;
-	int			save_nestlevel;
 	double		numSummarized = 0;
+	IndexBuildSecurity ibsec;
 
 	if (RecoveryInProgress())
 		ereport(ERROR,
@@ -1420,27 +1418,12 @@ brin_summarize_range(PG_FUNCTION_ARGS)
 		heapRel = table_open(heapoid, ShareUpdateExclusiveLock);
 
 		/*
-		 * Autovacuum calls us.  For its benefit, switch to the table owner's
-		 * userid, so that any index functions are run as that user.  Also
-		 * lock down security-restricted operations and arrange to make GUC
-		 * variable changes local to this command.  This is harmless, albeit
-		 * unnecessary, when called from SQL, because we fail shortly if the
-		 * user does not own the index.
+		 * Prevent index functions from doing what they are not supposed to.
 		 */
-		GetUserIdAndSecContext(&save_userid, &save_sec_context);
-		SetUserIdAndSecContext(heapRel->rd_rel->relowner,
-							   save_sec_context | SECURITY_RESTRICTED_OPERATION);
-		save_nestlevel = NewGUCNestLevel();
-		RestrictSearchPath();
+		enable_index_build_security(heapRel->rd_rel->relowner, &ibsec);
 	}
 	else
-	{
 		heapRel = NULL;
-		/* Set these just to suppress "uninitialized variable" warnings */
-		save_userid = InvalidOid;
-		save_sec_context = -1;
-		save_nestlevel = -1;
-	}
 
 	indexRel = index_open(indexoid, ShareUpdateExclusiveLock);
 
@@ -1453,7 +1436,8 @@ brin_summarize_range(PG_FUNCTION_ARGS)
 						RelationGetRelationName(indexRel))));
 
 	/* User must own the index (comparable to privileges needed for VACUUM) */
-	if (heapRel != NULL && !object_ownercheck(RelationRelationId, indexoid, save_userid))
+	if (heapRel != NULL && !object_ownercheck(RelationRelationId, indexoid,
+											  ibsec.userid))
 		aclcheck_error(ACLCHECK_NOT_OWNER, OBJECT_INDEX,
 					   RelationGetRelationName(indexRel));
 
@@ -1477,11 +1461,9 @@ brin_summarize_range(PG_FUNCTION_ARGS)
 				 errmsg("index \"%s\" is not valid",
 						RelationGetRelationName(indexRel))));
 
-	/* Roll back any GUC changes executed by index functions */
-	AtEOXact_GUC(false, save_nestlevel);
+	/* Relax the restrictions imposed above. */
+	disable_index_build_security(&ibsec);
 
-	/* Restore userid and security context */
-	SetUserIdAndSecContext(save_userid, save_sec_context);
 
 	index_close(indexRel, ShareUpdateExclusiveLock);
 	table_close(heapRel, ShareUpdateExclusiveLock);
diff --git a/src/backend/catalog/index.c b/src/backend/catalog/index.c
index 81bba4beac7..8eabe232d1e 100644
--- a/src/backend/catalog/index.c
+++ b/src/backend/catalog/index.c
@@ -1504,11 +1504,9 @@ index_concurrently_build(Oid heapRelationId,
 						 Oid indexRelationId)
 {
 	Relation	heapRel;
-	Oid			save_userid;
-	int			save_sec_context;
-	int			save_nestlevel;
 	Relation	indexRelation;
 	IndexInfo  *indexInfo;
+	IndexBuildSecurity ibsec;
 
 	/* This had better make sure that a snapshot is active */
 	Assert(ActiveSnapshotSet());
@@ -1517,15 +1515,9 @@ index_concurrently_build(Oid heapRelationId,
 	heapRel = table_open(heapRelationId, ShareUpdateExclusiveLock);
 
 	/*
-	 * Switch to the table owner's userid, so that any index functions are run
-	 * as that user.  Also lock down security-restricted operations and
-	 * arrange to make GUC variable changes local to this command.
+	 * Prevent index functions from doing what they are not supposed to.
 	 */
-	GetUserIdAndSecContext(&save_userid, &save_sec_context);
-	SetUserIdAndSecContext(heapRel->rd_rel->relowner,
-						   save_sec_context | SECURITY_RESTRICTED_OPERATION);
-	save_nestlevel = NewGUCNestLevel();
-	RestrictSearchPath();
+	enable_index_build_security(heapRel->rd_rel->relowner, &ibsec);
 
 	indexRelation = index_open(indexRelationId, RowExclusiveLock);
 
@@ -1542,11 +1534,8 @@ index_concurrently_build(Oid heapRelationId,
 	/* Now build the index */
 	index_build(heapRel, indexRelation, indexInfo, false, true, true);
 
-	/* Roll back any GUC changes executed by index functions */
-	AtEOXact_GUC(false, save_nestlevel);
-
-	/* Restore userid and security context */
-	SetUserIdAndSecContext(save_userid, save_sec_context);
+	/* Relax the restrictions imposed above. */
+	disable_index_build_security(&ibsec);
 
 	/* Close both the relations, but keep the locks */
 	table_close(heapRel, NoLock);
@@ -3375,9 +3364,7 @@ validate_index(Oid heapId, Oid indexId, Snapshot snapshot)
 	IndexInfo  *indexInfo;
 	IndexVacuumInfo ivinfo;
 	ValidateIndexState state;
-	Oid			save_userid;
-	int			save_sec_context;
-	int			save_nestlevel;
+	IndexBuildSecurity ibsec;
 
 	{
 		const int	progress_index[] = {
@@ -3399,15 +3386,9 @@ validate_index(Oid heapId, Oid indexId, Snapshot snapshot)
 	heapRelation = table_open(heapId, ShareUpdateExclusiveLock);
 
 	/*
-	 * Switch to the table owner's userid, so that any index functions are run
-	 * as that user.  Also lock down security-restricted operations and
-	 * arrange to make GUC variable changes local to this command.
+	 * Prevent index functions from doing what they are not supposed to.
 	 */
-	GetUserIdAndSecContext(&save_userid, &save_sec_context);
-	SetUserIdAndSecContext(heapRelation->rd_rel->relowner,
-						   save_sec_context | SECURITY_RESTRICTED_OPERATION);
-	save_nestlevel = NewGUCNestLevel();
-	RestrictSearchPath();
+	enable_index_build_security(heapRelation->rd_rel->relowner, &ibsec);
 
 	indexRelation = index_open(indexId, RowExclusiveLock);
 
@@ -3486,11 +3467,8 @@ validate_index(Oid heapId, Oid indexId, Snapshot snapshot)
 		 "validate_index found %.0f heap tuples, %.0f index tuples; inserted %.0f missing tuples",
 		 state.htups, state.itups, state.tups_inserted);
 
-	/* Roll back any GUC changes executed by index functions */
-	AtEOXact_GUC(false, save_nestlevel);
-
-	/* Restore userid and security context */
-	SetUserIdAndSecContext(save_userid, save_sec_context);
+	/* Relax the restrictions imposed above. */
+	disable_index_build_security(&ibsec);
 
 	/* Close rels, but keep locks */
 	index_close(indexRelation, NoLock);
@@ -4115,6 +4093,36 @@ reindex_relation(const ReindexStmt *stmt, Oid relid, int flags,
 	return result;
 }
 
+/*
+ * Before building an index, witch to the table owner's userid, so that any
+ * index functions are run as that user. Also lock down security-restricted
+ * operations and arrange to make GUC variable changes local to this command.
+ *
+ * Information needed later by disable_index_build_security() is stored in
+ * *sec.
+ */
+void
+enable_index_build_security(Oid userid, IndexBuildSecurity *sec)
+{
+	GetUserIdAndSecContext(&sec->userid, &sec->sec_context);
+	SetUserIdAndSecContext(userid,
+						   sec->sec_context | SECURITY_RESTRICTED_OPERATION);
+	sec->nestlevel = NewGUCNestLevel();
+	RestrictSearchPath();
+}
+
+/*
+ * Undo what enable_index_build_security() did.
+ */
+void
+disable_index_build_security(IndexBuildSecurity *sec)
+{
+	/* Roll back any GUC changes executed by index functions */
+	AtEOXact_GUC(false, sec->nestlevel);
+
+	/* Restore userid and security context */
+	SetUserIdAndSecContext(sec->userid, sec->sec_context);
+}
 
 /* ----------------------------------------------------------------
  *		System index reindexing support
diff --git a/src/backend/commands/analyze.c b/src/backend/commands/analyze.c
index f66e80b757c..dc2ad77ef9a 100644
--- a/src/backend/commands/analyze.c
+++ b/src/backend/commands/analyze.c
@@ -328,14 +328,12 @@ do_analyze_rel(Relation onerel, const VacuumParams *params,
 	PGRUsage	ru0;
 	TimestampTz starttime = 0;
 	MemoryContext caller_context;
-	Oid			save_userid;
-	int			save_sec_context;
-	int			save_nestlevel;
 	WalUsage	startwalusage = pgWalUsage;
 	BufferUsage startbufferusage = pgBufferUsage;
 	BufferUsage bufferusage;
 	PgStat_Counter startreadtime = 0;
 	PgStat_Counter startwritetime = 0;
+	IndexBuildSecurity ibsec;
 
 	verbose = (params->options & VACOPT_VERBOSE) != 0;
 	instrument = (verbose || (AmAutoVacuumWorkerProcess() &&
@@ -361,15 +359,9 @@ do_analyze_rel(Relation onerel, const VacuumParams *params,
 	caller_context = MemoryContextSwitchTo(anl_context);
 
 	/*
-	 * Switch to the table owner's userid, so that any index functions are run
-	 * as that user.  Also lock down security-restricted operations and
-	 * arrange to make GUC variable changes local to this command.
+	 * Prevent index functions from doing what they are not supposed to.
 	 */
-	GetUserIdAndSecContext(&save_userid, &save_sec_context);
-	SetUserIdAndSecContext(onerel->rd_rel->relowner,
-						   save_sec_context | SECURITY_RESTRICTED_OPERATION);
-	save_nestlevel = NewGUCNestLevel();
-	RestrictSearchPath();
+	enable_index_build_security(onerel->rd_rel->relowner, &ibsec);
 
 	/*
 	 * When verbose or autovacuum logging is used, initialize a resource usage
@@ -858,11 +850,8 @@ do_analyze_rel(Relation onerel, const VacuumParams *params,
 		}
 	}
 
-	/* Roll back any GUC changes executed by index functions */
-	AtEOXact_GUC(false, save_nestlevel);
-
-	/* Restore userid and security context */
-	SetUserIdAndSecContext(save_userid, save_sec_context);
+	/* Relax the restrictions imposed above. */
+	disable_index_build_security(&ibsec);
 
 	/* Restore current context and release memory */
 	MemoryContextSwitchTo(caller_context);
diff --git a/src/backend/commands/indexcmds.c b/src/backend/commands/indexcmds.c
index 713bb5d10f1..5bc28c94139 100644
--- a/src/backend/commands/indexcmds.c
+++ b/src/backend/commands/indexcmds.c
@@ -699,6 +699,8 @@ DefineIndex(ParseState *pstate,
 	 * Switch to the table owner's userid, so that any index functions are run
 	 * as that user.  Also lock down security-restricted operations.  We
 	 * already arranged to make GUC variable changes local to this command.
+	 *
+	 * XXX Use enable_index_build_security()?
 	 */
 	GetUserIdAndSecContext(&root_save_userid, &root_save_sec_context);
 	SetUserIdAndSecContext(rel->rd_rel->relowner,
@@ -1386,22 +1388,16 @@ DefineIndex(ParseState *pstate,
 			{
 				Oid			childRelid = part_oids[i];
 				Relation	childrel;
-				Oid			child_save_userid;
-				int			child_save_sec_context;
-				int			child_save_nestlevel;
 				List	   *childidxs;
 				ListCell   *cell;
 				AttrMap    *attmap;
 				bool		found = false;
+				IndexBuildSecurity child_ibsec;
 
 				childrel = table_open(childRelid, lockmode);
 
-				GetUserIdAndSecContext(&child_save_userid,
-									   &child_save_sec_context);
-				SetUserIdAndSecContext(childrel->rd_rel->relowner,
-									   child_save_sec_context | SECURITY_RESTRICTED_OPERATION);
-				child_save_nestlevel = NewGUCNestLevel();
-				RestrictSearchPath();
+				enable_index_build_security(childrel->rd_rel->relowner,
+											&child_ibsec);
 
 				/*
 				 * Don't try to create indexes on foreign tables, though. Skip
@@ -1418,9 +1414,7 @@ DefineIndex(ParseState *pstate,
 								 errdetail("Table \"%s\" contains partitions that are foreign tables.",
 										   RelationGetRelationName(rel))));
 
-					AtEOXact_GUC(false, child_save_nestlevel);
-					SetUserIdAndSecContext(child_save_userid,
-										   child_save_sec_context);
+					disable_index_build_security(&child_ibsec);
 					table_close(childrel, lockmode);
 					continue;
 				}
@@ -1505,9 +1499,7 @@ DefineIndex(ParseState *pstate,
 				}
 
 				list_free(childidxs);
-				AtEOXact_GUC(false, child_save_nestlevel);
-				SetUserIdAndSecContext(child_save_userid,
-									   child_save_sec_context);
+				disable_index_build_security(&child_ibsec);
 				table_close(childrel, NoLock);
 
 				/*
@@ -1534,7 +1526,7 @@ DefineIndex(ParseState *pstate,
 					 * Recurse as the starting user ID.  Callee will use that
 					 * for permission checks, then switch again.
 					 */
-					Assert(GetUserId() == child_save_userid);
+					Assert(GetUserId() == child_ibsec.userid);
 					SetUserIdAndSecContext(root_save_userid,
 										   root_save_sec_context);
 					childAddr =
@@ -1547,8 +1539,8 @@ DefineIndex(ParseState *pstate,
 									is_alter_table, check_rights,
 									check_not_in_use,
 									skip_build, quiet);
-					SetUserIdAndSecContext(child_save_userid,
-										   child_save_sec_context);
+					SetUserIdAndSecContext(child_ibsec.userid,
+										   child_ibsec.sec_context);
 
 					/*
 					 * Check if the index just created is valid or not, as it
@@ -4051,27 +4043,19 @@ ReindexRelationConcurrently(const ReindexStmt *stmt, Oid relationOid, const Rein
 		Oid			newIndexId;
 		Relation	indexRel;
 		Relation	heapRel;
-		Oid			save_userid;
-		int			save_sec_context;
-		int			save_nestlevel;
 		Relation	newIndexRel;
 		LockRelId  *lockrelid;
 		Oid			tablespaceid;
+		IndexBuildSecurity ibsec;
 
 		indexRel = index_open(idx->indexId, ShareUpdateExclusiveLock);
 		heapRel = table_open(indexRel->rd_index->indrelid,
 							 ShareUpdateExclusiveLock);
 
 		/*
-		 * Switch to the table owner's userid, so that any index functions are
-		 * run as that user.  Also lock down security-restricted operations
-		 * and arrange to make GUC variable changes local to this command.
+		 * Prevent index functions from doing what they are not supposed to.
 		 */
-		GetUserIdAndSecContext(&save_userid, &save_sec_context);
-		SetUserIdAndSecContext(heapRel->rd_rel->relowner,
-							   save_sec_context | SECURITY_RESTRICTED_OPERATION);
-		save_nestlevel = NewGUCNestLevel();
-		RestrictSearchPath();
+		enable_index_build_security(heapRel->rd_rel->relowner, &ibsec);
 
 		/* determine safety of this index for set_indexsafe_procflags */
 		idx->safe = (RelationGetIndexExpressions(indexRel) == NIL &&
@@ -4159,11 +4143,8 @@ ReindexRelationConcurrently(const ReindexStmt *stmt, Oid relationOid, const Rein
 		index_close(indexRel, NoLock);
 		index_close(newIndexRel, NoLock);
 
-		/* Roll back any GUC changes executed by index functions */
-		AtEOXact_GUC(false, save_nestlevel);
-
-		/* Restore userid and security context */
-		SetUserIdAndSecContext(save_userid, save_sec_context);
+		/* Relax the restrictions imposed above. */
+		disable_index_build_security(&ibsec);
 
 		table_close(heapRel, NoLock);
 
diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 0bf19d07db5..1ad453a39a0 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -514,13 +514,11 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid,
 	Oid			tableOid = RelationGetRelid(OldHeap);
 	Relation	index;
 	LOCKMODE	lmode;
-	Oid			save_userid;
-	int			save_sec_context;
-	int			save_nestlevel;
 	bool		verbose = ((params->options & CLUOPT_VERBOSE) != 0);
 	bool		recheck = ((params->options & CLUOPT_RECHECK) != 0);
 	bool		concurrent = ((params->options & CLUOPT_CONCURRENT) != 0);
 	Oid			ident_idx = InvalidOid;
+	IndexBuildSecurity ibsec;
 
 	/* Determine the lock mode to use. */
 	lmode = RepackLockLevel(concurrent);
@@ -539,15 +537,9 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid,
 	pgstat_progress_update_param(PROGRESS_REPACK_COMMAND, cmd);
 
 	/*
-	 * Switch to the table owner's userid, so that any index functions are run
-	 * as that user.  Also lock down security-restricted operations and
-	 * arrange to make GUC variable changes local to this command.
+	 * Prevent index functions from doing what they are not supposed to.
 	 */
-	GetUserIdAndSecContext(&save_userid, &save_sec_context);
-	SetUserIdAndSecContext(OldHeap->rd_rel->relowner,
-						   save_sec_context | SECURITY_RESTRICTED_OPERATION);
-	save_nestlevel = NewGUCNestLevel();
-	RestrictSearchPath();
+	enable_index_build_security(OldHeap->rd_rel->relowner, &ibsec);
 
 	/*
 	 * Recheck that the relation is still what it was when we started.
@@ -557,7 +549,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid,
 	 * not-previously-clustered index.
 	 */
 	if (recheck &&
-		!cluster_rel_recheck(cmd, OldHeap, indexOid, save_userid,
+		!cluster_rel_recheck(cmd, OldHeap, indexOid, GetUserId(),
 							 lmode, params->options))
 		goto out;
 
@@ -681,11 +673,8 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid,
 		rebuild_relation(OldHeap, index, verbose, ident_idx);
 
 out:
-	/* Roll back any GUC changes executed by index functions */
-	AtEOXact_GUC(false, save_nestlevel);
-
-	/* Restore userid and security context */
-	SetUserIdAndSecContext(save_userid, save_sec_context);
+	/* Relax the restrictions imposed above. */
+	disable_index_build_security(&ibsec);
 
 	pgstat_progress_end_command();
 }
@@ -1234,7 +1223,6 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	}
 }
 
-
 /*
  * Create the transient table that will be filled with new data during
  * CLUSTER, ALTER TABLE, and similar operations.  The transient table
diff --git a/src/backend/commands/tablecmds.c b/src/backend/commands/tablecmds.c
index cb93c3e935a..d691b317011 100644
--- a/src/backend/commands/tablecmds.c
+++ b/src/backend/commands/tablecmds.c
@@ -23762,9 +23762,7 @@ ATExecMergePartitions(List **wqueue, AlteredTableInfo *tab, Relation rel,
 	Oid			defaultPartOid;
 	Oid			existingRelid;
 	Oid			ownerId = InvalidOid;
-	Oid			save_userid;
-	int			save_sec_context;
-	int			save_nestlevel;
+	IndexBuildSecurity ibsec;
 
 	/*
 	 * Check ownership of merged partitions - partitions with different owners
@@ -23896,18 +23894,9 @@ ATExecMergePartitions(List **wqueue, AlteredTableInfo *tab, Relation rel,
 	newPartRel = createPartitionTable(wqueue, cmd->name, rel, ownerId);
 
 	/*
-	 * Switch to the table owner's userid, so that any index functions are run
-	 * as that user.  Also, lockdown security-restricted operations and
-	 * arrange to make GUC variable changes local to this command.
-	 *
-	 * Need to do it after determining the namespace in the
-	 * createPartitionTable() call.
+	 * Prevent index functions from doing what they are not supposed to.
 	 */
-	GetUserIdAndSecContext(&save_userid, &save_sec_context);
-	SetUserIdAndSecContext(ownerId,
-						   save_sec_context | SECURITY_RESTRICTED_OPERATION);
-	save_nestlevel = NewGUCNestLevel();
-	RestrictSearchPath();
+	enable_index_build_security(ownerId, &ibsec);
 
 	/* Copy data from merged partitions to the new partition. */
 	MergePartitionsMoveRows(wqueue, mergingPartitions, newPartRel);
@@ -23944,11 +23933,8 @@ ATExecMergePartitions(List **wqueue, AlteredTableInfo *tab, Relation rel,
 	/* Keep the lock until commit. */
 	table_close(newPartRel, NoLock);
 
-	/* Roll back any GUC changes executed by index functions. */
-	AtEOXact_GUC(false, save_nestlevel);
-
-	/* Restore the userid and security context. */
-	SetUserIdAndSecContext(save_userid, save_sec_context);
+	/* Relax the restrictions imposed above. */
+	disable_index_build_security(&ibsec);
 }
 
 /*
diff --git a/src/backend/commands/vacuum.c b/src/backend/commands/vacuum.c
index 38539a6fd3d..5d1cbc382fa 100644
--- a/src/backend/commands/vacuum.c
+++ b/src/backend/commands/vacuum.c
@@ -34,6 +34,7 @@
 #include "access/tableam.h"
 #include "access/transam.h"
 #include "access/xact.h"
+#include "catalog/index.h"
 #include "catalog/namespace.h"
 #include "catalog/pg_database.h"
 #include "catalog/pg_inherits.h"
@@ -2017,10 +2018,8 @@ vacuum_rel(Oid relid, RangeVar *relation, VacuumParams params,
 	LockRelId	lockrelid;
 	Oid			priv_relid;
 	Oid			toast_relid;
-	Oid			save_userid;
-	int			save_sec_context;
-	int			save_nestlevel;
 	VacuumParams toast_vacuum_params;
+	IndexBuildSecurity ibsec;
 
 	/*
 	 * This function scribbles on the parameters, so make a copy early to
@@ -2270,16 +2269,9 @@ vacuum_rel(Oid relid, RangeVar *relation, VacuumParams params,
 		toast_relid = InvalidOid;
 
 	/*
-	 * Switch to the table owner's userid, so that any index functions are run
-	 * as that user.  Also lock down security-restricted operations and
-	 * arrange to make GUC variable changes local to this command. (This is
-	 * unnecessary, but harmless, for lazy VACUUM.)
+	 * Prevent index functions from doing what they are not supposed to.
 	 */
-	GetUserIdAndSecContext(&save_userid, &save_sec_context);
-	SetUserIdAndSecContext(rel->rd_rel->relowner,
-						   save_sec_context | SECURITY_RESTRICTED_OPERATION);
-	save_nestlevel = NewGUCNestLevel();
-	RestrictSearchPath();
+	enable_index_build_security(rel->rd_rel->relowner, &ibsec);
 
 	/*
 	 * If PROCESS_MAIN is set (the default), it's time to vacuum the main
@@ -2310,11 +2302,8 @@ vacuum_rel(Oid relid, RangeVar *relation, VacuumParams params,
 			table_relation_vacuum(rel, &params, bstrategy);
 	}
 
-	/* Roll back any GUC changes executed by index functions */
-	AtEOXact_GUC(false, save_nestlevel);
-
-	/* Restore userid and security context */
-	SetUserIdAndSecContext(save_userid, save_sec_context);
+	/* Relax the restrictions imposed above. */
+	disable_index_build_security(&ibsec);
 
 	/* all done with this class, but hold lock until commit */
 	if (rel)
diff --git a/src/include/catalog/index.h b/src/include/catalog/index.h
index 9aee8226347..dd9ae8119e5 100644
--- a/src/include/catalog/index.h
+++ b/src/include/catalog/index.h
@@ -172,6 +172,16 @@ extern void reindex_index(const ReindexStmt *stmt, Oid indexId,
 extern bool reindex_relation(const ReindexStmt *stmt, Oid relid, int flags,
 							 const ReindexParams *params);
 
+typedef struct IndexBuildSecurity
+{
+	Oid			userid;
+	int			sec_context;
+	int			nestlevel;
+} IndexBuildSecurity;
+
+extern void enable_index_build_security(Oid userid, IndexBuildSecurity *sec);
+extern void disable_index_build_security(IndexBuildSecurity *sec);
+
 extern bool ReindexIsProcessingHeap(Oid heapOid);
 extern bool ReindexIsProcessingIndex(Oid indexOid);
 
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 25ac7079099..0ec534d15d7 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1327,6 +1327,7 @@ IndexAttachInfo
 IndexAttrBitmapKind
 IndexBuildCallback
 IndexBuildResult
+IndexBuildSecurity
 IndexBulkDeleteCallback
 IndexBulkDeleteResult
 IndexClause
-- 
2.52.0

  [text/x-diff] v02-0006-Use-separate-transactions-for-catalog-changes.patch (43.2K, ../108776.1784105248@localhost/7-v02-0006-Use-separate-transactions-for-catalog-changes.patch)
  download | inline diff:
From 3569755de1cc6c885091aa9f49a715b7fdec1f38 Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 15 Jul 2026 10:37:12 +0200
Subject: [PATCH 6/8] Use separate transactions for catalog changes.

This is another part of the effort to eliminate the impact of REPACK
(CONCURRENTLY) on VACUUM xmin horizon. While one of the previous patches
avoids using the same snapshot for long time, here we try to avoid using the
same XID throughout the REPACK execution. The idea is that we only get XID
assigned for a catalog change, and as soon as it's done, we start a new
transaction. We still need XID to load the data into the new table, but this
will be fixed by the next patch in the series.

A side effect of committing the catalog changes early is that the new table is
visible to other transactions while REPACK is still running. The other
transactions are not able to access it though since we keep it locked in
AccessExclusiveMode all the time. However, by committing the transaction that
created the table we'd lose the lock. Therefore we use a session lock until
the table gets locked in the following transaction.

Another problem is that, once the creation of the new table got committed, the
table will not be dropped automatically if REPACK ends up with ERROR (or
crash). We resolve this by creating a dependency in the system catalog. When
REPACK starts, it first uses the dependency information to find out if some
table depends on the table that should be processed, and if it does, it checks
if its pg_class(relrewrite) attribute points to the table to be processed. In
such a case it drops the new table because it evidently belongs to the
previous (failed) run.

Besides creation of the new table, separate transaction is also used to each
index,
---
 src/backend/commands/matview.c   |   2 +-
 src/backend/commands/repack.c    | 880 +++++++++++++++++++++++++------
 src/backend/commands/tablecmds.c |   2 +-
 src/include/commands/repack.h    |  10 +-
 src/tools/pgindent/typedefs.list |   2 +
 5 files changed, 743 insertions(+), 153 deletions(-)

diff --git a/src/backend/commands/matview.c b/src/backend/commands/matview.c
index f7d8007f796..9d490da5f81 100644
--- a/src/backend/commands/matview.c
+++ b/src/backend/commands/matview.c
@@ -317,7 +317,7 @@ RefreshMatViewByOid(Oid matviewOid, bool is_create, bool skipData,
 	 */
 	OIDNewHeap = make_new_heap(matviewOid, tableSpace,
 							   matviewRel->rd_rel->relam,
-							   relpersistence, ExclusiveLock);
+							   relpersistence, ExclusiveLock, false);
 	Assert(CheckRelationOidLockedByMe(OIDNewHeap, AccessExclusiveLock, false));
 
 	/* Generate the data, if wanted. */
diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 1ad453a39a0..89695335127 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -51,6 +51,7 @@
 #include "catalog/pg_am.h"
 #include "catalog/pg_attrdef.h"
 #include "catalog/pg_constraint.h"
+#include "catalog/pg_depend.h"
 #include "catalog/pg_inherits.h"
 #include "catalog/toasting.h"
 #include "commands/defrem.h"
@@ -96,6 +97,38 @@ typedef struct
 	Oid			indexOid;
 } RelToCluster;
 
+/*
+ * Information needed to close and re-open relation on transaction boundary.
+ */
+typedef struct RelReopenInfo
+{
+	Oid			relid;
+	Relation   *p_rel;
+	bool		is_new;
+} RelReopenInfo;
+
+/*
+ * Information needed to release ans re-initialize ChangeContext on
+ * transaction boundary.
+ */
+typedef struct ChangeContexBackup
+{
+	/* The new relation. */
+	Relation	rel;
+	RelReopenInfo rel_ri;
+	Oid			ident_index;
+
+	/* Auxiliary relation. */
+	Relation	rel_aux;
+	RelReopenInfo rel_aux_ri;
+	Oid			ident_index_aux;
+
+	/* Common fields */
+	int			file_seq_snapshot;
+	int			file_seq_changes;
+	Oid			clustering_index;
+} ChangeContextBackup;
+
 /*
  * When REPACK (CONCURRENTLY) copies data to the new heap, a new snapshot is
  * built after processing this many pages. XXX Tune the value.
@@ -120,6 +153,13 @@ typedef struct DecodingWorker
 /* Pointer to currently running decoding worker. */
 static DecodingWorker *decoding_worker = NULL;
 
+/*
+ * In the CONCURRENTLY mode, we need multiple transactions to process a single
+ * table. Information that needs to survive commit should be stored in this
+ * context.
+ */
+static MemoryContext repack_cxt = NULL;
+
 /*
  * Is there a message sent by a repack worker that the backend needs to
  * receive?
@@ -134,6 +174,14 @@ static void check_concurrent_repack_requirements(Relation rel,
 												 Oid *ident_idx_p);
 static void rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 							 Oid ident_idx);
+static ChangeContext *make_new_heap_for_repack(Oid tablespace,
+											   Oid access_method,
+											   char relpersistence,
+											   Oid ident_idx,
+											   Relation *p_old_heap,
+											   Relation *p_clustering_index,
+											   IndexBuildSecurity *ibsec);
+static List *find_new_heaps(Oid relid_old);
 static void copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 							bool verbose,
 							bool *pSwapToastByContent,
@@ -175,15 +223,30 @@ static void release_change_context(ChangeContext *chgcxt);
 static void initialize_change_dest(RepackDest *dest, Relation relation,
 								   Oid ident_index_id);
 static void release_change_dest(RepackDest *dest);
+static void backup_change_context(ChangeContext *chgcxt,
+								  ChangeContextBackup *backup);
+static ChangeContext *reinitialize_change_context(ChangeContextBackup *backup);
+static void enable_session_locks(Relation rel, LOCKMODE lmode, bool is_new);
+static void disable_session_locks(Relation rel, LOCKMODE lmode, bool is_new);
+static void prepare_relation_for_reopening(Relation *p_rel,
+										   RelReopenInfo *backup,
+										   bool is_new);
+static void reopen_relation(RelReopenInfo *backup);
 static void rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 											   Oid identIdx,
 											   TransactionId frozenXid,
 											   MultiXactId cutoffMulti,
 											   ChangeContext *chgcxt);
-static void process_auxiliary_table(ChangeContext *chgcxt, Relation OldHeap,
-									Oid identIdx);
-static List *build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes);
+static ChangeContext *process_auxiliary_table(ChangeContext *chgcxt,
+											  Relation *pOldHeap,
+											  Relation *pNewHeap,
+											  Oid identIdx);
+static List *build_new_indexes(List *OldIndexes, Relation *p_old,
+							   Relation *p_new,
+							   ChangeContext **p_chgcxt);
 static Oid	build_new_index(Relation NewHeap, Relation OldHeap, Oid oldindex);
+static ChangeContext *start_new_transaction(ChangeContext *chgcxt,
+											Relation *p_old, Relation *p_new);
 static void copy_index_constraints(Relation old_index, Oid new_index_id,
 								   Oid new_heap_id);
 static void copy_attribute_defaults(Oid old_heap_oid, Oid new_heap_oid);
@@ -518,7 +581,6 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid,
 	bool		recheck = ((params->options & CLUOPT_RECHECK) != 0);
 	bool		concurrent = ((params->options & CLUOPT_CONCURRENT) != 0);
 	Oid			ident_idx = InvalidOid;
-	IndexBuildSecurity ibsec;
 
 	/* Determine the lock mode to use. */
 	lmode = RepackLockLevel(concurrent);
@@ -536,11 +598,6 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid,
 	pgstat_progress_start_command(PROGRESS_COMMAND_REPACK, tableOid);
 	pgstat_progress_update_param(PROGRESS_REPACK_COMMAND, cmd);
 
-	/*
-	 * Prevent index functions from doing what they are not supposed to.
-	 */
-	enable_index_build_security(OldHeap->rd_rel->relowner, &ibsec);
-
 	/*
 	 * Recheck that the relation is still what it was when we started.
 	 *
@@ -662,6 +719,17 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid,
 	 */
 	if (concurrent)
 	{
+		/*
+		 * In the CONCURRENTLY mode, new transactions may be started. Make
+		 * sure we can preserve information across transaction boundaries.
+		 */
+		if (repack_cxt == NULL)
+			repack_cxt = AllocSetContextCreate(TopMemoryContext,
+											   "Repack - Top",
+											   ALLOCSET_DEFAULT_SIZES);
+		else
+			MemoryContextReset(repack_cxt);
+
 		PG_ENSURE_ERROR_CLEANUP(stop_repack_decoding_worker_cb, 0);
 		{
 			rebuild_relation(OldHeap, index, verbose, ident_idx);
@@ -673,9 +741,6 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid,
 		rebuild_relation(OldHeap, index, verbose, ident_idx);
 
 out:
-	/* Relax the restrictions imposed above. */
-	disable_index_build_security(&ibsec);
-
 	pgstat_progress_end_command();
 }
 
@@ -999,13 +1064,14 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	Oid			tableOid = RelationGetRelid(OldHeap);
 	Oid			accessMethod = OldHeap->rd_rel->relam;
 	Oid			tableSpace = OldHeap->rd_rel->reltablespace;
-	Oid			OIDNewHeap;
-	Relation	NewHeap;
+	Oid			OIDNewHeap = InvalidOid;
+	Relation	NewHeap = NULL;
 	char		relpersistence;
 	bool		swap_toast_by_content;
 	TransactionId frozenXid;
 	MultiXactId cutoffMulti;
 	bool		concurrent = OidIsValid(ident_idx);
+	IndexBuildSecurity ibsec;
 	ChangeContext *chgcxt = NULL;
 #if USE_ASSERT_CHECKING
 	LOCKMODE	lmode;
@@ -1016,8 +1082,15 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	Assert(index == NULL || CheckRelationLockedByMe(index, lmode, false));
 #endif
 
+	/*
+	 * Prevent index functions from doing what they are not supposed to.
+	 */
+	enable_index_build_security(OldHeap->rd_rel->relowner, &ibsec);
+
 	if (concurrent)
 	{
+		MemoryContext oldcxt;
+
 		/*
 		 * The worker needs to be member of the locking group we're the leader
 		 * of. We ought to become the leader before the worker starts. The
@@ -1041,8 +1114,12 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 		 * Not sure this risk is worth unlocking/locking the table (and its
 		 * clustering index) and checking again if it's still eligible for
 		 * REPACK CONCURRENTLY.
+		 *
+		 * Use a transaction context that survives transaction commit(s).
 		 */
+		oldcxt = MemoryContextSwitchTo(repack_cxt);
 		start_repack_decoding_worker(tableOid);
+		MemoryContextSwitchTo(oldcxt);
 	}
 
 	/* for CLUSTER or REPACK USING INDEX, mark the index as the one to use */
@@ -1052,112 +1129,32 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	/* Remember info about rel before closing OldHeap */
 	relpersistence = OldHeap->rd_rel->relpersistence;
 
-	/*
-	 * Create the transient table that will receive the re-ordered data.
-	 *
-	 * OldHeap is already locked, so no need to lock it again.  make_new_heap
-	 * obtains AccessExclusiveLock on the new heap and its toast table.
-	 */
-	OIDNewHeap = make_new_heap(tableOid, tableSpace,
-							   accessMethod,
-							   relpersistence,
-							   NoLock);
-	Assert(CheckRelationOidLockedByMe(OIDNewHeap, AccessExclusiveLock, false));
-	NewHeap = table_open(OIDNewHeap, NoLock);
-
 	if (concurrent)
 	{
-		bool		need_aux_rel;
-
 		/*
-		 * Auxiliary table is needed for clustering in the CONCURRENTLY mode,
-		 * see comments in ChangeContext. FIXME Non-btree indexes are allowed
-		 * historically, but in general, these can hardly define any useful
-		 * order. We ignore them here.
+		 * Create the transient table that will receive the re-ordered data,
+		 * and possibly the auxiliary table for the ordered data.
 		 */
-		need_aux_rel = index != NULL && index->rd_rel->relam == BTREE_AM_OID;
-
-		/* Gather information to apply concurrent changes. */
-		chgcxt = palloc0_object(ChangeContext);
+		chgcxt = make_new_heap_for_repack(tableSpace, accessMethod,
+										  relpersistence, ident_idx,
+										  &OldHeap, &index, &ibsec);
+		NewHeap = chgcxt->cc_dest.rel;
+	}
+	else
+	{
+		Assert(CheckRelationLockedByMe(OldHeap, AccessExclusiveLock, false));
 
 		/*
-		 * Create a copy of the attribute defaults on the temp table, which
-		 * the executor needs when replaying concurrent data changes.
+		 * Nothing special here. No locking as the caller should already have
+		 * a lock on the old relation, and the new one will be locked by
+		 * make_new_heap() anyway.
 		 */
-		copy_attribute_defaults(tableOid, OIDNewHeap);
-
-		if (!need_aux_rel)
-		{
-			Oid			ident_idx_new;
-
-			/*
-			 * Create the identity index. We will need it during data copying
-			 * so that we can apply the data changes at the appropriate time -
-			 * see comments around the call of
-			 * repack_process_concurrent_changes() with block range specified.
-			 *
-			 * XXX NewHeap is empty - should we pass INDEX_CREATE_SKIP_BUILD?
-			 */
-			ident_idx_new = build_new_index(NewHeap, OldHeap, ident_idx);
-
-			initialize_change_context(chgcxt, NewHeap, ident_idx_new);
-		}
-		else
-		{
-			Oid			aux_oid;
-			Relation	aux_rel;
-			Oid			aux_ident_idx;
-
-			/*
-			 * As the concurrent data changes will be applied to the auxiliary
-			 * heap, the new heap does not need the identity index yet. We'll
-			 * build it after having copied the data from the auxiliary heap.
-			 * (Bulk insert should be more efficient.)
-			 */
-			initialize_change_context(chgcxt, NewHeap, InvalidOid);
-
-			/*
-			 * Like above, but only temporary - no other backend should need
-			 * it.
-			 */
-			aux_oid = make_new_heap(tableOid, tableSpace, accessMethod,
-									RELPERSISTENCE_TEMP, NoLock);
-			Assert(CheckRelationOidLockedByMe(aux_oid, AccessExclusiveLock,
-											  false));
-			aux_rel = table_open(aux_oid, NoLock);
-
-
-			/*
-			 * Copy the attribute defaults as we did for the new heap above -
-			 * the concurrent changes also need to be applied to the auxiliary
-			 * table.
-			 */
-			copy_attribute_defaults(tableOid, aux_oid);
-
-			/*
-			 * The same for identity index. (The additional
-			 * ShareUpdateExclusiveLock on ident_idx is not a problem, it'll
-			 * be released at the end of transaction.)
-			 */
-			aux_ident_idx = build_new_index(aux_rel, OldHeap, ident_idx);
-
-			/*
-			 * Make the relation ready for use.
-			 */
-			chgcxt->cc_dest_aux = palloc0_object(RepackDest);
-			initialize_change_dest(chgcxt->cc_dest_aux, aux_rel,
-								   aux_ident_idx);
-
-			/*
-			 * Set OID of the old relation's clustering index if it's
-			 * different from the identity index. Otherwise set InvalidOid to
-			 * indicate that the identity index should be used for clustering.
-			 */
-			if (RelationGetRelid(index) != ident_idx)
-				chgcxt->cc_clustering_index = RelationGetRelid(index);
-			else
-				chgcxt->cc_clustering_index = InvalidOid;
-		}
+		OIDNewHeap = make_new_heap(tableOid, tableSpace, accessMethod,
+								   relpersistence, NoLock, false);
+		/* No additional lock needed. */
+		Assert(CheckRelationOidLockedByMe(OIDNewHeap, AccessExclusiveLock,
+										  false));
+		NewHeap = table_open(OIDNewHeap, NoLock);
 	}
 
 	/* Copy the heap data into the new table in the desired order */
@@ -1221,6 +1218,335 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 						 frozenXid, cutoffMulti,
 						 relpersistence);
 	}
+
+	/* Relax the restrictions imposed above. */
+	disable_index_build_security(&ibsec);
+}
+
+/*
+ * Wrapper around make_new_heap() to handle specifics of REPACK
+ * (CONCURRENTLY).
+ *
+ * 'ident_idx' is the identity index, to handle replaying of concurrent data
+ * changes to the new heap.
+ *
+ * The function starts a new transaction. Therefore, both old heap and
+ * clustering index are passed in the form of pointer so the caller has the
+ * correct entries in the new transaction. New heap is returned in the
+ * ChangeContext object. On return, the new heap will be locked in
+ * AccessExclusiveLock mode.
+ */
+static ChangeContext *
+make_new_heap_for_repack(Oid tablespace, Oid access_method,
+						 char relpersistence, Oid ident_idx,
+						 Relation *p_old_heap, Relation *p_clustering_index,
+						 IndexBuildSecurity *ibsec)
+{
+	Relation	old_heap = *p_old_heap;
+	Relation	clustering_index = *p_clustering_index;
+	List	   *existing = NIL;
+	Relation	new_heap;
+	Oid			old_heap_oid = RelationGetRelid(old_heap);
+	bool		need_aux_rel;
+	Oid			new_heap_oid;
+	ObjectAddress old_obj,
+				new_obj;
+	Oid			ident_idx_new;
+	Oid			aux_oid = InvalidOid;
+	Oid			aux_ident_idx = InvalidOid;
+	Oid			clustering_index_oid = InvalidOid;
+	Relation	aux_rel = NULL;
+	ChangeContext *chgcxt;
+
+	/*
+	 * In the CONCURRENTLY case, we do the catalog changes in a separate
+	 * transaction so that we do not have XID assigned while copying the data.
+	 * (That XID would block the progress of the xmin horizon for VACUUM.)
+	 * Therefore, the new relation will be created in a separate transaction
+	 * and thus be visible to other transactions, although our lock should not
+	 * let them access it. However, if previous run of REPACK ended with
+	 * ERROR, such relation may still be there.
+	 *
+	 * Besides that, an auxiliary relation (for tuple sorting) might exist.
+	 * Thus we expect a list of zero, one or two items.
+	 *
+	 * Since the relations should depend on the old one (see below), we use
+	 * the pg_depend catalog to find it.
+	 */
+	existing = find_new_heaps(old_heap_oid);
+	Assert(list_length(existing) <= 2);
+	foreach_oid(existing_oid, existing)
+	{
+		/* There appears to be one, so simply drop it. */
+		new_obj.classId = RelationRelationId;
+		new_obj.objectId = existing_oid;
+		new_obj.objectSubId = InvalidOid;
+
+		/*
+		 * User is not supposed to create any dependencies on the relation, so
+		 * CASCADE.
+		 */
+		performDeletion(&new_obj, DROP_CASCADE, PERFORM_DELETION_INTERNAL);
+	}
+
+	/*
+	 * The old heap should already be locked by the caller. As for the new
+	 * one, lock is needed because the commit below will make it visible to
+	 * other transactions. make_new_heap() should lock it.
+	 */
+	Assert(CheckRelationLockedByMe(old_heap, ShareUpdateExclusiveLock,
+								   false));
+	new_heap_oid = make_new_heap(old_heap_oid, tablespace, access_method,
+								 relpersistence, NoLock, false);
+	Assert(CheckRelationOidLockedByMe(new_heap_oid, AccessExclusiveLock,
+									  false));
+
+	new_heap = table_open(new_heap_oid, NoLock);
+
+	/*
+	 * Create a copy of the attribute defaults on the temp table, which the
+	 * executor needs when replaying concurrent data changes.
+	 */
+	copy_attribute_defaults(old_heap_oid, new_heap_oid);
+
+	/*
+	 * Auxiliary table is needed for clustering in the CONCURRENTLY mode, see
+	 * comments in ChangeContext. FIXME Non-btree indexes are allowed
+	 * historically, but in general, these can hardly define any useful order.
+	 * We ignore them here.
+	 */
+	need_aux_rel = clustering_index != NULL &&
+		clustering_index->rd_rel->relam == BTREE_AM_OID;
+	if (!need_aux_rel)
+	{
+		/*
+		 * Create the identity index. We will need it during data copying so
+		 * that we can apply the data changes at the appropriate time - see
+		 * comments around the call of repack_process_concurrent_changes()
+		 * with block range specified.
+		 *
+		 * XXX NewHeap is empty - should we pass INDEX_CREATE_SKIP_BUILD?
+		 */
+		ident_idx_new = build_new_index(new_heap, old_heap, ident_idx);
+	}
+	else
+	{
+		/*
+		 * As the concurrent data changes will be applied to the auxiliary
+		 * heap, the new heap does not need the identity index yet. We'll
+		 * build it after having copied the data from the auxiliary heap.
+		 * (Bulk insert should be more efficient.)
+		 */
+		ident_idx_new = InvalidOid;
+
+		/*
+		 * Create the auxiliary table - like the new heap above, but only
+		 * unlogged - neither crash recovery nor replication is needed.
+		 */
+		aux_oid = make_new_heap(old_heap_oid, tablespace, access_method,
+								RELPERSISTENCE_UNLOGGED, NoLock, true);
+		Assert(CheckRelationOidLockedByMe(aux_oid, AccessExclusiveLock,
+										  false));
+
+		/*
+		 * The same for identity index. (The additional
+		 * ShareUpdateExclusiveLock on ident_idx is not a problem, it'll be
+		 * released at the end of transaction.)
+		 */
+		aux_rel = table_open(aux_oid, NoLock);
+		aux_ident_idx = build_new_index(aux_rel, old_heap, ident_idx);
+
+		clustering_index_oid = RelationGetRelid(clustering_index);
+
+		/*
+		 * Copy the attribute defaults as we did for the new heap above - the
+		 * concurrent changes also need to be applied to the auxiliary table.
+		 */
+		copy_attribute_defaults(old_heap_oid, aux_oid);
+	}
+
+	/*
+	 * Create dependencies on the old heap - it ensures that dropping the old
+	 * heap automatically cascades to the new ones.
+	 */
+	old_obj.classId = RelationRelationId;
+	old_obj.objectId = old_heap_oid;
+	old_obj.objectSubId = InvalidOid;
+	new_obj.classId = RelationRelationId;
+	new_obj.objectId = new_heap_oid;
+	new_obj.objectSubId = InvalidOid;
+	recordDependencyOn(&new_obj, &old_obj, DEPENDENCY_AUTO);
+
+	if (OidIsValid(aux_oid))
+	{
+		new_obj.objectId = aux_oid;
+		recordDependencyOn(&new_obj, &old_obj, DEPENDENCY_AUTO);
+	}
+
+	/*
+	 * Commit will release the locks, so make sure no one can alter or drop
+	 * the relations until we're done.
+	 */
+	enable_session_locks(old_heap, ShareUpdateExclusiveLock, false);
+	/* NoLock for simplicity, commit will release the existing lock anyway. */
+	table_close(old_heap, NoLock);
+
+	/* Likewise, acquire session lock for the new heap. */
+	enable_session_locks(new_heap, AccessExclusiveLock, true);
+	table_close(new_heap, NoLock);
+
+	/*
+	 * The same for the auxiliary relation and clustering index if ordering is
+	 * required.
+	 */
+	if (need_aux_rel)
+	{
+		enable_session_locks(aux_rel, AccessExclusiveLock, true);
+		table_close(aux_rel, NoLock);
+
+		/* is_new does not matter for index. */
+		enable_session_locks(clustering_index, ShareUpdateExclusiveLock,
+							 false);
+		index_close(clustering_index, NoLock);
+	}
+
+	/*
+	 * Even if the clustering index is not appropriate for sorting, we ought
+	 * to close it before committing the transaction.
+	 */
+	else if (clustering_index)
+		index_close(clustering_index, NoLock);
+
+	/*
+	 * Commit the current transaction and start a new one.
+	 */
+	disable_index_build_security(ibsec);
+	PopActiveSnapshot();
+	CommitTransactionCommand();
+	StartTransactionCommand();
+	PushActiveSnapshot(GetTransactionSnapshot());
+	enable_index_build_security(ibsec->userid, ibsec);
+
+	/*
+	 * Open the new heap in the new transaction and disable the session
+	 * lock(s).
+	 */
+	new_heap = table_open(new_heap_oid, NoLock);
+	disable_session_locks(new_heap, AccessExclusiveLock, true);
+
+	/*
+	 * Allocate and initialize memory for the output information. We shouldn't
+	 * have done it in the memory context of the previous transaction.
+	 */
+	chgcxt = palloc0_object(ChangeContext);
+	initialize_change_context(chgcxt, new_heap, ident_idx_new);
+	chgcxt->cc_ind_build_sec = *ibsec;
+	if (need_aux_rel)
+	{
+		/*
+		 * Make the auxiliary relation ready for use.
+		 */
+		chgcxt->cc_dest_aux = palloc0_object(RepackDest);
+		aux_rel = table_open(aux_oid, NoLock);
+		disable_session_locks(aux_rel, AccessExclusiveLock, true);
+		initialize_change_dest(chgcxt->cc_dest_aux, aux_rel,
+							   aux_ident_idx);
+
+		/*
+		 * Set OID of the old relation's clustering index if it's different
+		 * from the identity index. Otherwise set InvalidOid to indicate that
+		 * the identity index should be used for clustering.
+		 */
+		if (clustering_index_oid != ident_idx)
+			chgcxt->cc_clustering_index = clustering_index_oid;
+		else
+			chgcxt->cc_clustering_index = InvalidOid;
+
+		/* Re-open the clustering index. */
+		clustering_index = index_open(clustering_index_oid, NoLock);
+		/* is_new does not matter for index. */
+		disable_session_locks(clustering_index, ShareUpdateExclusiveLock,
+							  false);
+	}
+	else
+		clustering_index = NULL;
+
+	/*
+	 * Re-open the old heap.
+	 */
+	old_heap = table_open(old_heap_oid, NoLock);
+	disable_session_locks(old_heap, ShareUpdateExclusiveLock, false);
+
+	*p_old_heap = old_heap;
+	*p_clustering_index = clustering_index;
+
+	return chgcxt;
+}
+
+/*
+ * Check if new heap, and possibly also an auxiliary heap, already exists for
+ * given old heap - typically due to failed REPACK (CONCURRENTLY). Return OIDs
+ * of in a list or NIL if there is none.
+ */
+static List *
+find_new_heaps(Oid relid_old)
+{
+	Relation	depRel;
+	ScanKeyData key[3];
+	SysScanDesc scan;
+	HeapTuple	tup;
+	List	   *result = NIL;
+
+	ScanKeyInit(&key[0],
+				Anum_pg_depend_refclassid,
+				BTEqualStrategyNumber, F_OIDEQ,
+				RelationRelationId);
+	ScanKeyInit(&key[1],
+				Anum_pg_depend_refobjid,
+				BTEqualStrategyNumber, F_OIDEQ,
+				ObjectIdGetDatum(relid_old));
+	ScanKeyInit(&key[2],
+				Anum_pg_depend_refobjsubid,
+				BTEqualStrategyNumber, F_OIDEQ,
+				ObjectIdGetDatum(InvalidOid));
+
+	depRel = table_open(DependRelationId, AccessShareLock);
+	scan = systable_beginscan(depRel, DependReferenceIndexId, true,
+							  NULL, 3, key);
+	while (HeapTupleIsValid(tup = systable_getnext(scan)))
+	{
+		Form_pg_depend depform;
+
+		/*
+		 * There should be AUTO dependency from another table, which is the
+		 * new or the auxiliary one. If we found any other, just ignore them.
+		 */
+		depform = (Form_pg_depend) GETSTRUCT(tup);
+		if (depform->classid == RelationRelationId &&
+			depform->deptype == DEPENDENCY_AUTO)
+		{
+			Relation	rel;
+			Form_pg_class classForm;
+			Oid			matched = InvalidOid;
+
+			rel = relation_open(depform->objid, AccessShareLock);
+
+			/*
+			 * The new relation's relrewrite should point to the old relation.
+			 */
+			classForm = rel->rd_rel;
+			if (classForm->relrewrite == relid_old)
+				matched = depform->objid;
+			table_close(rel, AccessShareLock);
+
+			if (OidIsValid(matched))
+				result = lappend_oid(result, matched);
+		}
+	}
+	systable_endscan(scan);
+	table_close(depRel, AccessShareLock);
+
+	return result;
 }
 
 /*
@@ -1235,7 +1561,7 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
  */
 Oid
 make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
-			  char relpersistence, LOCKMODE lockmode)
+			  char relpersistence, LOCKMODE lockmode, bool auxiliary)
 {
 	TupleDesc	OldHeapDesc;
 	char		NewHeapName[NAMEDATALEN];
@@ -1246,6 +1572,7 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
 	Datum		reloptions;
 	bool		isNull;
 	Oid			namespaceid;
+	const char *suffix;
 
 	OldHeap = table_open(OIDOldHeap, lockmode);
 	OldHeapDesc = RelationGetDescr(OldHeap);
@@ -1285,7 +1612,9 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
 	 * mapped.  This simplifies swap_relation_files, and is absolutely
 	 * necessary for rebuilding pg_class, for reasons explained there.
 	 */
-	snprintf(NewHeapName, sizeof(NewHeapName), "pg_temp_%u", OIDOldHeap);
+	suffix = auxiliary ? "_aux" : "";
+	snprintf(NewHeapName, sizeof(NewHeapName), "pg_temp_%u%s", OIDOldHeap,
+			 suffix);
 
 	OIDNewHeap = heap_create_with_catalog(NewHeapName,
 										  namespaceid,
@@ -2562,31 +2891,6 @@ process_single_relation(RepackStmt *stmt, LOCKMODE lockmode, bool isTopLevel,
 	Assert(stmt->command == REPACK_COMMAND_CLUSTER ||
 		   stmt->command == REPACK_COMMAND_REPACK);
 
-	if (params->options & CLUOPT_CONCURRENT)
-	{
-		/*
-		 * Since REPACK (CONCURRENTLY) pops the active snapshot during the
-		 * processing (it creates and pushes snapshots on its own), and since
-		 * that snapshot can be referenced by the current portal, we need to
-		 * make sure that the portal has no dangling pointer to the snapshot.
-		 * Starting a new transaction seems to be the simplest way.
-		 *
-		 * XXX The following patches in the series make this unnecessary, as
-		 * they start new transactions for other reasons elsewhere.
-		 */
-		PopActiveSnapshot();
-		CommitTransactionCommand();
-
-		/* Start a new transaction. */
-		StartTransactionCommand();
-
-		/*
-		 * Functions in indexes may want a snapshot set. Note that the portal
-		 * is not aware of this one, so the caller needs to pop it explicitly.
-		 */
-		PushActiveSnapshot(GetTransactionSnapshot());
-	}
-
 	/*
 	 * Make sure ANALYZE is specified if a column list is present.
 	 */
@@ -3476,6 +3780,176 @@ release_change_dest(RepackDest *dest)
 	pfree(dest->ident_key);
 }
 
+/*
+ * Copy information needed to re-initialize 'chgcxt' to 'backup'.
+ */
+static void
+backup_change_context(ChangeContext *chgcxt, ChangeContextBackup *backup)
+{
+	RepackDest *dest;
+
+	memset(backup, 0, sizeof(ChangeContextBackup));
+
+	/* Store information on the new relation. */
+	dest = &chgcxt->cc_dest;
+	backup->rel = dest->rel;
+	prepare_relation_for_reopening(&backup->rel, &backup->rel_ri, true);
+	backup->ident_index = dest->ident_index ?
+		RelationGetRelid(dest->ident_index) : InvalidOid;
+
+	/* Store information on the auxiliary relation if one exists. */
+	if (chgcxt->cc_dest_aux)
+	{
+		dest = chgcxt->cc_dest_aux;
+		backup->rel_aux = dest->rel;
+		prepare_relation_for_reopening(&backup->rel_aux, &backup->rel_aux_ri,
+									   true);
+		backup->ident_index_aux = dest->ident_index ?
+			RelationGetRelid(dest->ident_index) : InvalidOid;
+	}
+	else
+	{
+		/* There should be no reopening. */
+		Assert(backup->rel_aux_ri.relid == InvalidOid);
+	}
+
+	/* Backup the common fields. */
+	backup->file_seq_snapshot = chgcxt->cc_file_seq_snapshot;
+	backup->file_seq_changes = chgcxt->cc_file_seq_changes;
+	backup->clustering_index = chgcxt->cc_clustering_index;
+}
+
+/*
+ * Create and initialize an instance of ChangeContext, based on previously
+ * existing instance. It includes re-opening of relations and indexes.
+ */
+static ChangeContext *
+reinitialize_change_context(ChangeContextBackup *backup)
+{
+	ChangeContext *chgcxt;
+	RelReopenInfo *rri;
+
+	chgcxt = palloc0_object(ChangeContext);
+
+	/* Re-initialize the new relation part. */
+	rri = &backup->rel_ri;
+	reopen_relation(rri);
+	initialize_change_context(chgcxt, backup->rel, backup->ident_index);
+
+	/* The same for the auxiliary relation, if it exists. */
+	rri = &backup->rel_aux_ri;
+	if (OidIsValid(rri->relid))
+	{
+		reopen_relation(rri);
+		chgcxt->cc_dest_aux = palloc0_object(RepackDest);
+		initialize_change_dest(chgcxt->cc_dest_aux, backup->rel_aux,
+							   backup->ident_index_aux);
+	}
+
+	/* Restore the common fields. */
+	chgcxt->cc_file_seq_snapshot = backup->file_seq_snapshot;
+	chgcxt->cc_file_seq_changes = backup->file_seq_changes;
+	chgcxt->cc_clustering_index = backup->clustering_index;
+
+	return chgcxt;
+}
+
+/*
+ * Lock the relation and its TOAST relation using a session lock.
+ */
+static void
+enable_session_locks(Relation rel, LOCKMODE lmode, bool is_new)
+{
+	LockRelId	lock;
+	Oid			toastid;
+
+	Assert(CheckRelationLockedByMe(rel, lmode, false));
+	lock = rel->rd_lockInfo.lockRelId;
+	LockRelationIdForSession(&lock, lmode);
+
+	/*
+	 * The same for TOAST relation. XXX Is it ok that we do not lock TOAST
+	 * relation of the old relation until we start swapping the files? (The
+	 * new relation's TOAST is locked on creation, so we keep it that way.)
+	 */
+	toastid = rel->rd_rel->reltoastrelid;
+	if (is_new && OidIsValid(toastid))
+	{
+		Assert(CheckRelationOidLockedByMe(toastid, lmode, false));
+		lock.relId = toastid;
+		LockRelationIdForSession(&lock, lmode);
+	}
+}
+
+/*
+ * Disable session locks acquired earlier by enable_session_locks(), but first
+ * acquire normal (transactional) lock(s).
+ */
+static void
+disable_session_locks(Relation rel, LOCKMODE lmode, bool is_new)
+{
+	LockRelId	lock;
+	Oid			toastid;
+
+	LockRelation(rel, lmode);
+	lock = rel->rd_lockInfo.lockRelId;
+	UnlockRelationIdForSession(&lock, lmode);
+	Assert(CheckRelationLockedByMe(rel, lmode, false));
+
+	/* The same for TOAST relation. See enable_session_locks(). */
+	toastid = rel->rd_rel->reltoastrelid;
+	if (is_new && OidIsValid(toastid))
+	{
+		lock.relId = toastid;
+		LockRelationId(&lock, lmode);
+		UnlockRelationIdForSession(&lock, lmode);
+		Assert(CheckRelationOidLockedByMe(toastid, lmode, false));
+	}
+}
+
+/*
+ * Remember OID of a relation, lock it using a session lock and close it. and a
+ *
+ * The lock mode is always ShareUpdateExclusiveLock, as this is related to
+ * REPACK (CONCURRENTLY).
+ */
+static void
+prepare_relation_for_reopening(Relation *p_rel, RelReopenInfo *backup,
+							   bool is_new)
+{
+	Relation	rel = *p_rel;
+	LOCKMODE	lmode;
+
+	lmode = is_new ? AccessExclusiveLock : ShareUpdateExclusiveLock;
+
+	/* Remember what we need to be able to re-open the relation later. */
+	backup->relid = RelationGetRelid(rel);
+	backup->p_rel = p_rel;
+	backup->is_new = is_new;
+
+	/*
+	 * Close it, but make sure the lock is not lost at transaction boundary.
+	 */
+	enable_session_locks(rel, lmode, is_new);
+	relation_close(rel, NoLock);
+}
+
+/*
+ * Open the relation, acquire transactional lock on it and release the session
+ * lock acquired by prepare_relation_for_reopening().
+ */
+static void
+reopen_relation(RelReopenInfo *backup)
+{
+	Relation	rel;
+	LOCKMODE	lmode;
+
+	lmode = backup->is_new ? AccessExclusiveLock : ShareUpdateExclusiveLock;
+	rel = relation_open(backup->relid, lmode);
+	disable_session_locks(rel, lmode, backup->is_new);
+	*backup->p_rel = rel;
+}
+
 /*
  * The final steps of rebuild_relation() for concurrent processing.
  *
@@ -3493,21 +3967,27 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	List	   *ind_oids_new;
 	Oid			old_table_oid = RelationGetRelid(OldHeap);
 	Oid			new_table_oid = RelationGetRelid(NewHeap);
-	List	   *ind_oids_old = RelationGetIndexList(OldHeap);
+	List	   *ind_oids_old;
 	ListCell   *lc,
 			   *lc2;
 	char		relpersistence;
 	bool		is_system_catalog;
 	XLogRecPtr	end_of_wal;
+	MemoryContext oldcxt;
 	List	   *indexrels;
 	List	   *inds_tmp = NIL;
 
 	Assert(CheckRelationLockedByMe(OldHeap, ShareUpdateExclusiveLock, false));
 	Assert(CheckRelationLockedByMe(NewHeap, AccessExclusiveLock, false));
 
-	/* If we have the auxiliary table, this is the moment we should use it. */
+	/*
+	 * If we have the auxiliary table, this is the moment we should use it.
+	 * Note that the function can start new transaction(s). In that case it
+	 * re-opens the new and old heap, so pass pointers to get the relation
+	 * cache entries updated.
+	 */
 	if (chgcxt->cc_dest_aux)
-		process_auxiliary_table(chgcxt, OldHeap, identIdx);
+		chgcxt = process_auxiliary_table(chgcxt, &OldHeap, &NewHeap, identIdx);
 
 	/*
 	 * Unlike the exclusive case, we build new indexes for the new relation
@@ -3530,13 +4010,17 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	 * auxiliary table) can be a problem in terms of index layout. Shouldn't
 	 * we drop the identity index and build it using bulk insert too?
 	 */
+	oldcxt = MemoryContextSwitchTo(repack_cxt);
+	ind_oids_old = RelationGetIndexList(OldHeap);
 	foreach_oid(ind_oid, ind_oids_old)
 	{
 		if (ind_oid != identIdx)
 			inds_tmp = lappend_oid(inds_tmp, ind_oid);
 	}
+	MemoryContextSwitchTo(oldcxt);
 	ind_oids_old = inds_tmp;
-	ind_oids_new = build_new_indexes(NewHeap, OldHeap, ind_oids_old);
+	ind_oids_new = build_new_indexes(ind_oids_old, &OldHeap, &NewHeap,
+									 &chgcxt);
 
 	/*
 	 * The identity index will be involved in the following processing.
@@ -3695,9 +4179,15 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 /*
  * Copy the contents of the auxiliary table to the new table in the desired
  * order, then drop the auxiliary table.
+ *
+ * All the opened relations may need to be closed and re-opened because the
+ * function may start a new transaction. That's why a pointer to old and new
+ * heap is passed. 'chgcxt' is re-created in that case, so user should use the
+ * returned value.
  */
-static void
-process_auxiliary_table(ChangeContext *chgcxt, Relation OldHeap, Oid identIdx)
+static ChangeContext *
+process_auxiliary_table(ChangeContext *chgcxt, Relation *pOldHeap,
+						Relation *pNewHeap, Oid identIdx)
 {
 	RepackDest *dest = chgcxt->cc_dest_aux;
 	Oid			ident_idx_new;
@@ -3718,8 +4208,16 @@ process_auxiliary_table(ChangeContext *chgcxt, Relation OldHeap, Oid identIdx)
 		/*
 		 * Create it according to the clustering index on the old relation.
 		 */
-		cl_ind_oid = build_new_index(dest->rel, OldHeap,
+		cl_ind_oid = build_new_index(dest->rel, *pOldHeap,
 									 chgcxt->cc_clustering_index);
+
+		/*
+		 * The index build generated a new XID, so commit the transaction and
+		 * start a new one.
+		 */
+		chgcxt = start_new_transaction(chgcxt, pOldHeap, pNewHeap);
+		dest = chgcxt->cc_dest_aux;
+
 		clustering_index = index_open(cl_ind_oid, NoLock);
 	}
 	else
@@ -3778,7 +4276,13 @@ process_auxiliary_table(ChangeContext *chgcxt, Relation OldHeap, Oid identIdx)
 	performDeletion(&object, DROP_RESTRICT, PERFORM_DELETION_INTERNAL);
 
 	/* Build the identity index on the new relation. */
-	ident_idx_new = build_new_index(chgcxt->cc_dest.rel, OldHeap, identIdx);
+	ident_idx_new = build_new_index(chgcxt->cc_dest.rel, *pOldHeap, identIdx);
+
+	/*
+	 * The index build generated a new XID, so commit the transaction and
+	 * start a new one.
+	 */
+	chgcxt = start_new_transaction(chgcxt, pOldHeap, pNewHeap);
 
 	/*
 	 * Make the new heap ready to use the index for future replaying of
@@ -3787,6 +4291,8 @@ process_auxiliary_table(ChangeContext *chgcxt, Relation OldHeap, Oid identIdx)
 	rel = chgcxt->cc_dest.rel;
 	release_change_dest(&chgcxt->cc_dest);
 	initialize_change_dest(&chgcxt->cc_dest, rel, ident_idx_new);
+
+	return chgcxt;
 }
 
 /*
@@ -3798,20 +4304,47 @@ process_auxiliary_table(ChangeContext *chgcxt, Relation OldHeap, Oid identIdx)
  * A list of OIDs of the corresponding indexes created on NewHeap is
  * returned. The order of items does match, so we can use these arrays to swap
  * index storage.
+ *
+ * A separate transaction is used for each index. Therefore caller needs to
+ * pass pointers to the relcache entries so that we can provide him with new
+ * entries after reopening the relations. Likewise, pointer to ChangeContext
+ * is used to return the re-created instance.
  */
 static List *
-build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes)
+build_new_indexes(List *OldIndexes, Relation *p_old, Relation *p_new,
+				  ChangeContext **p_chgcxt)
 {
+	ChangeContext *chgcxt = *p_chgcxt;
 	List	   *result = NIL;
 
 	foreach_oid(oldindex, OldIndexes)
 	{
+		Relation	rel_old = *p_old;
+		Relation	rel_new = *p_new;
 		Oid			newindex;
+		MemoryContext oldcxt;
+
+		newindex = build_new_index(rel_new, rel_old, oldindex);
 
-		newindex = build_new_index(NewHeap, OldHeap, oldindex);
+		/*
+		 * Allocate the list in repack_cxt so it survives the current
+		 * transaction.
+		 */
+		oldcxt = MemoryContextSwitchTo(repack_cxt);
 		result = lappend_oid(result, newindex);
+		MemoryContextSwitchTo(oldcxt);
+
+		/*
+		 * Use one transaction per index, to limit the impact on xmin
+		 * horizons. XXX Does it make sense to even first commit the catalog
+		 * changes and then do the build in a new transaction (which should
+		 * not set MyProc->xmin until the build is complete)? Not sure.
+		 */
+		chgcxt = start_new_transaction(chgcxt, p_old, p_new);
 	}
 
+	*p_chgcxt = chgcxt;
+
 	return result;
 }
 
@@ -3835,17 +4368,62 @@ build_new_index(Relation NewHeap, Relation OldHeap, Oid oldindex)
 								 "repacknew",
 								 get_rel_namespace(ind->rd_index->indrelid),
 								 false);
+
 	/* Functions in indexes may want a snapshot set. */
 	Assert(ActiveSnapshotSet());
 	newindex = index_create_copy(NewHeap, INDEX_CREATE_SUPPRESS_PROGRESS,
 								 oldindex, ind->rd_rel->reltablespace,
 								 newName);
+
 	copy_index_constraints(ind, newindex, RelationGetRelid(NewHeap));
 	index_close(ind, NoLock);
 
 	return newindex;
 }
 
+/*
+ * Start a new transaction, but make sure that the caller still has valid
+ * relcache entries as well as valid ChangeContext. The problem is that
+ * relcache references should be closed before transaction commit, so we need
+ * them re-opened in the new transaction. Also, commit releases all locks
+ * acquired in the transaction, however we don't want to lose our locks. We
+ * work it around by using session locks temporarily.
+ */
+static ChangeContext *
+start_new_transaction(ChangeContext *chgcxt, Relation *p_old, Relation *p_new)
+{
+	ChangeContextBackup chgcxt_backup;
+	RelReopenInfo rri_old;
+	MemoryContext oldcxt;
+	IndexBuildSecurity ibsec = chgcxt->cc_ind_build_sec;
+
+	/* This also closes the new relation. */
+	backup_change_context(chgcxt, &chgcxt_backup);
+	release_change_context(chgcxt);
+	prepare_relation_for_reopening(p_old, &rri_old, false);
+
+	disable_index_build_security(&ibsec);
+	PopActiveSnapshot();
+	CommitTransactionCommand();
+	StartTransactionCommand();
+	PushActiveSnapshot(GetTransactionSnapshot());
+	enable_index_build_security(ibsec.userid, &ibsec);
+
+	/* Create 'chgcxt' using re-opened relations. */
+	oldcxt = MemoryContextSwitchTo(repack_cxt);
+	/* This also re-opens the new relation. */
+	chgcxt = reinitialize_change_context(&chgcxt_backup);
+	chgcxt->cc_ind_build_sec = ibsec;
+	MemoryContextSwitchTo(oldcxt);
+
+	reopen_relation(&rri_old);
+
+	/* Update the callers relcache entry for the new relation. */
+	*p_new = chgcxt->cc_dest.rel;
+
+	return chgcxt;
+}
+
 /*
  * Create a transient copy of a constraint -- supported by a transient
  * copy of the index that supports the original constraint.
@@ -4031,6 +4609,8 @@ start_repack_decoding_worker(Oid relid)
 	size = BUFFERALIGN(offsetof(DecodingWorkerShared, error_queue)) +
 		BUFFERALIGN(REPACK_ERROR_QUEUE_SIZE);
 	decoding_worker->seg = dsm_create(size, 0);
+	/* The segment should not be detached at transaction boundary. */
+	dsm_pin_mapping(decoding_worker->seg);
 
 	shared = (DecodingWorkerShared *) dsm_segment_address(decoding_worker->seg);
 	shared->initialized = false;
diff --git a/src/backend/commands/tablecmds.c b/src/backend/commands/tablecmds.c
index d691b317011..427fe733153 100644
--- a/src/backend/commands/tablecmds.c
+++ b/src/backend/commands/tablecmds.c
@@ -6099,7 +6099,7 @@ ATRewriteTables(AlterTableStmt *parsetree, List **wqueue, LOCKMODE lockmode,
 			 * unlogged anyway.
 			 */
 			OIDNewHeap = make_new_heap(tab->relid, NewTableSpace, NewAccessMethod,
-									   persistence, lockmode);
+									   persistence, lockmode, false);
 
 			/*
 			 * Copy the heap data into the new table with the desired
diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h
index 8af73f8c81f..ad8125790b0 100644
--- a/src/include/commands/repack.h
+++ b/src/include/commands/repack.h
@@ -18,6 +18,7 @@
 #include "access/hio.h"
 #include "access/skey.h"
 #include "access/xlogdefs.h"
+#include "catalog/index.h"
 #include "nodes/execnodes.h"
 #include "nodes/parsenodes.h"
 #include "parser/parse_node.h"
@@ -120,6 +121,12 @@ typedef struct ChangeContext
 	 * The index that defines ordering of the old table.
 	 */
 	Oid			cc_clustering_index;
+
+	/*
+	 * Information needed to disable / enable security restrictions on index
+	 * functions. This is needed when starting a new transaction.
+	 */
+	IndexBuildSecurity cc_ind_build_sec;
 } ChangeContext;
 
 extern PGDLLIMPORT int repack_pages_per_snapshot;
@@ -133,7 +140,8 @@ extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid,
 extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal);
 
 extern Oid	make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
-						  char relpersistence, LOCKMODE lockmode);
+						  char relpersistence, LOCKMODE lockmode,
+						  bool auxiliary);
 extern void heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
 								   TupleTableSlot *reform);
 extern bool tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc);
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 0ec534d15d7..5a7e10cee9c 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -431,6 +431,7 @@ CatalogId
 CatalogIdMapEntry
 CatalogIndexState
 ChangeContext
+ChangeContextBackup
 ChangeVarNodes_callback
 ChangeVarNodes_context
 ChannelName
@@ -2611,6 +2612,7 @@ RelMapping
 RelOptInfo
 RelOptKind
 RelPathStr
+RelReopenInfo
 RelStatsInfo
 RelSyncCallbackFunction
 RelToCheck
-- 
2.52.0

  [text/x-diff] v02-0007-Decouple-updating-of-freezing-information-from-swap_.patch (8.1K, ../108776.1784105248@localhost/8-v02-0007-Decouple-updating-of-freezing-information-from-swap_.patch)
  download | inline diff:
From 79d5c8edb56448f51168238d7d48ee20ac52500d Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 15 Jul 2026 10:37:12 +0200
Subject: [PATCH 7/8] Decouple updating of freezing information from
 swap_relation_files().

So far, the relfrozenxid and relminmxid attributes of pg_class used to be
updated by swap_relation_files(), however, not every call of the function
should do that. On the contrary, new feature of REPACK (CONCURRENTLY) will
need to update this information in one case that would be tricky to handle
using swap_relation_files().
---
 src/backend/commands/repack.c | 114 +++++++++++++++++-----------------
 1 file changed, 56 insertions(+), 58 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 89695335127..7390fd303f6 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -188,6 +188,8 @@ static void copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldInde
 							TransactionId *pFreezeXid,
 							MultiXactId *pCutoffMulti,
 							ChangeContext *chgcxt);
+static void update_relation_cutoffs(Oid relid, TransactionId frozenXid,
+									MultiXactId cutoffMulti);
 static List *get_tables_to_repack(RepackCommand cmd, bool usingindex,
 								  MemoryContext permcxt);
 static List *get_tables_to_repack_partitioned(RepackCommand cmd,
@@ -2039,13 +2041,6 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
  * links) while the latter is the only way to handle cases in which a toast
  * table is added or removed altogether.
  *
- * Additionally, the first relation is marked with relfrozenxid set to
- * frozenXid.  It seems a bit ugly to have this here, but the caller would
- * have to do it anyway, so having it here saves a heap_update.  Note: in
- * the swap-toast-links case, we assume we don't need to change the toast
- * table's relfrozenxid: the new version of the toast table should already
- * have relfrozenxid set to RecentXmin, which is good enough.
- *
  * Lastly, if r2 and its toast table and toast index (if any) are mapped,
  * their OIDs are emitted into mapped_tables[].  This is hacky but beats
  * having to look the information up again later in finish_heap_swap.
@@ -2054,9 +2049,7 @@ static void
 swap_relation_files(Oid r1, Oid r2, bool target_is_pg_class,
 					bool swap_toast_by_content,
 					bool is_internal,
-					TransactionId frozenXid,
-					MultiXactId cutoffMulti,
-					Oid *mapped_tables)
+					Oid *mapped_tables, Oid *r1_toastid)
 {
 	Relation	relRelation;
 	HeapTuple	reltup1,
@@ -2088,6 +2081,9 @@ swap_relation_files(Oid r1, Oid r2, bool target_is_pg_class,
 	relam1 = relform1->relam;
 	relam2 = relform2->relam;
 
+	if (r1_toastid)
+		*r1_toastid = relform1->reltoastrelid;
+
 	if (RelFileNumberIsValid(relfilenumber1) &&
 		RelFileNumberIsValid(relfilenumber2))
 	{
@@ -2203,15 +2199,6 @@ swap_relation_files(Oid r1, Oid r2, bool target_is_pg_class,
 	 * and then fail to commit the pg_class update.
 	 */
 
-	/* set rel1's frozen Xid and minimum MultiXid */
-	if (relform1->relkind != RELKIND_INDEX)
-	{
-		Assert(!TransactionIdIsValid(frozenXid) ||
-			   TransactionIdIsNormal(frozenXid));
-		relform1->relfrozenxid = frozenXid;
-		relform1->relminmxid = cutoffMulti;
-	}
-
 	/* swap size statistics too, since new rel has freshly-updated stats */
 	{
 		int32		swap_pages;
@@ -2313,9 +2300,8 @@ swap_relation_files(Oid r1, Oid r2, bool target_is_pg_class,
 									target_is_pg_class,
 									swap_toast_by_content,
 									is_internal,
-									frozenXid,
-									cutoffMulti,
-									mapped_tables);
+									mapped_tables,
+									NULL);
 			}
 			else
 			{
@@ -2416,9 +2402,8 @@ swap_relation_files(Oid r1, Oid r2, bool target_is_pg_class,
 							target_is_pg_class,
 							swap_toast_by_content,
 							is_internal,
-							InvalidTransactionId,
-							InvalidMultiXactId,
-							mapped_tables);
+							mapped_tables,
+							NULL);
 	}
 
 	/* Clean up. */
@@ -2428,6 +2413,36 @@ swap_relation_files(Oid r1, Oid r2, bool target_is_pg_class,
 	table_close(relRelation, RowExclusiveLock);
 }
 
+/*
+ * Update relfrozenxid and relminmxid attributes of pg_class entry, specified
+ * by OID.
+ */
+static void
+update_relation_cutoffs(Oid relid, TransactionId frozenXid,
+						MultiXactId cutoffMulti)
+{
+	Relation	relRelation;
+	HeapTuple	reltup;
+	Form_pg_class relform;
+	CatalogIndexState indstate;
+
+	relRelation = table_open(RelationRelationId, RowExclusiveLock);
+	reltup = SearchSysCacheCopy1(RELOID, ObjectIdGetDatum(relid));
+	if (!HeapTupleIsValid(reltup))
+		elog(ERROR, "cache lookup failed for relation %u", relid);
+	relform = (Form_pg_class) GETSTRUCT(reltup);
+	relform->relfrozenxid = frozenXid;
+	relform->relminmxid = cutoffMulti;
+
+	indstate = CatalogOpenIndexes(relRelation);
+	CatalogTupleUpdateWithInfo(relRelation, &reltup->t_self, reltup,
+							   indstate);
+	CatalogCloseIndexes(indstate);
+
+	heap_freetuple(reltup);
+	table_close(relRelation, RowExclusiveLock);
+}
+
 /*
  * Remove the transient table that was built by make_new_heap, and finish
  * cleaning up (including rebuilding all indexes on the old heap).
@@ -2445,6 +2460,7 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 {
 	ObjectAddress object;
 	Oid			mapped_tables[4];
+	Oid			oid_old_toastid;
 	int			i;
 
 	/* Report that we are now swapping relation files */
@@ -2461,7 +2477,7 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 	swap_relation_files(OIDOldHeap, OIDNewHeap,
 						(OIDOldHeap == RelationRelationId),
 						swap_toast_by_content, is_internal,
-						frozenXid, cutoffMulti, mapped_tables);
+						mapped_tables, &oid_old_toastid);
 
 	/*
 	 * If it's a system catalog, queue a sinval message to flush all catcaches
@@ -2517,37 +2533,20 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 								 PROGRESS_REPACK_PHASE_FINAL_CLEANUP);
 
 	/*
-	 * If the relation being rebuilt is pg_class, swap_relation_files()
-	 * couldn't update pg_class's own pg_class entry (check comments in
-	 * swap_relation_files()), thus relfrozenxid was not updated. That's
-	 * annoying because a potential reason for doing a VACUUM FULL is a
-	 * imminent or actual anti-wraparound shutdown.  So, now that we can
-	 * access the new relation using its indices, update relfrozenxid.
-	 * pg_class doesn't have a toast relation, so we don't need to update the
-	 * corresponding toast relation. Not that there's little point moving all
-	 * relfrozenxid updates here since swap_relation_files() needs to write to
-	 * pg_class for non-mapped relations anyway.
+	 * Update relfrozenxid and relminmxid. CCI is needed to see the most
+	 * recent pg_class tuple, created by swap_relation_files(). However that
+	 * CCI makes us use the new relation file. Therefore we cannot do this
+	 * before indexes have been rebuilt too.
 	 */
-	if (OIDOldHeap == RelationRelationId)
-	{
-		Relation	relRelation;
-		HeapTuple	reltup;
-		Form_pg_class relform;
-
-		relRelation = table_open(RelationRelationId, RowExclusiveLock);
-
-		reltup = SearchSysCacheCopy1(RELOID, ObjectIdGetDatum(OIDOldHeap));
-		if (!HeapTupleIsValid(reltup))
-			elog(ERROR, "cache lookup failed for relation %u", OIDOldHeap);
-		relform = (Form_pg_class) GETSTRUCT(reltup);
-
-		relform->relfrozenxid = frozenXid;
-		relform->relminmxid = cutoffMulti;
-
-		CatalogTupleUpdate(relRelation, &reltup->t_self, reltup);
+	CommandCounterIncrement();
+	update_relation_cutoffs(OIDOldHeap, frozenXid, cutoffMulti);
 
-		table_close(relRelation, RowExclusiveLock);
-	}
+	/*
+	 * The same for TOAST, if needed. In the swap-toast-links case, the new
+	 * the toast table should already have relfrozenxid set to RecentXmin.
+	 */
+	if (OidIsValid(oid_old_toastid) && swap_toast_by_content)
+		update_relation_cutoffs(oid_old_toastid, frozenXid, cutoffMulti);
 
 	/* Destroy new heap with old filenumber */
 	object.classId = RelationRelationId;
@@ -4133,9 +4132,8 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 							(old_table_oid == RelationRelationId),
 							false,	/* swap_toast_by_content */
 							true,
-							InvalidTransactionId,
-							InvalidMultiXactId,
-							mapped_tables);
+							mapped_tables,
+							NULL);
 
 #ifdef USE_ASSERT_CHECKING
 
-- 
2.52.0

  [text/plain] v02-0008-Make-REPACK-CONCURRENTLY-MVCC-safe.patch (115.9K, ../108776.1784105248@localhost/9-v02-0008-Make-REPACK-CONCURRENTLY-MVCC-safe.patch)
  download | inline diff:
From ee135ee4c57d978594d483af7fcb3997c3e5094b Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Wed, 15 Jul 2026 10:37:12 +0200
Subject: [PATCH 8/8] Make REPACK (CONCURRENTLY) MVCC-safe.

In the initial implementation of REPACK (CONCURRENTLY), tuples in the new
table are marked with XID of the transaction the command executes in. This is
simpler to implement, but can cause surprising behavior: on specific
conditions, the table contents may become invisible for transactions that
should see it. (Both ALTER TABLE and TRUNCATE commands have this problem.)
Besides that, as REPACK (CONCURRENTLY) needs to have XID assigned (to mark the
new tuples) while it copies the data (which can take long time), it can
restrict the progress of the xmin horizons for VACUUM quite a bit.

This patch teaches REPACK (CONCURRENTLY) to transfer the visibility
information from the old table to the new one. Thus there's no need to have
XID assigned during the copying, and therefore VACUUM (of other tables) is not
restricted anymore.

A new type of snapshot is used to look for the existing tuples in the new
table, in order to apply the concurrent data changes (i.e. changes performed
by other transactions during the initial copying). This snapshot assumes that
the new table does not contain any tuples created by aborted
transactions. This is simpler and more efficient than building historic MVCC
snapshots for the look-ups.

One problem that needs special attention is tuple freezing. Besides page
pruning code, freezing is currently implemented only in rewriteheap.c. Thus it
seems simpler to teach REPACK (CONCURRENTLY) to copy the data using
rewriteheap.c rather than implementing freezing of individual tuples
independent from page pruning. (The latter would include WAL logging of the
individual frozen tuples, whereas rewriteheap.c logs the whole page at once.)
Disadvantage of this approach is that we only freeze tuples during the initial
copying, while the data changes decoded from WAL are replayed without
freezing.

This patch also uses rewriteheap.c (in particular raw_heap_insert()) to insert
TOAST tuples. That seems to be the easiest way to avoid getting XID assigned
just to insert TOAST. However, due to not using a new XID, even TOAST needs to
be frozen now - again, we use the existing code in rewriteheap.c.
---
 doc/src/sgml/mvcc.sgml                        |  12 +-
 doc/src/sgml/ref/repack.sgml                  |   9 -
 src/backend/access/common/toast_internals.c   |  45 +-
 src/backend/access/heap/heapam.c              |  84 +++-
 src/backend/access/heap/heapam_handler.c      | 265 +++++++++--
 src/backend/access/heap/heapam_visibility.c   |  62 +++
 src/backend/access/heap/heaptoast.c           |  23 +-
 src/backend/access/heap/rewriteheap.c         | 263 ++++++++---
 src/backend/access/table/tableam.c            |   1 +
 src/backend/access/table/toast_helper.c       |  11 +-
 src/backend/access/transam/xloginsert.c       |  31 +-
 src/backend/access/transam/xlogrecovery.c     |   7 +-
 src/backend/catalog/indexing.c                |   2 +-
 src/backend/commands/matview.c                |   3 +-
 src/backend/commands/repack.c                 | 433 ++++++++++++------
 src/backend/commands/tablecmds.c              |   1 +
 src/backend/executor/nodeModifyTable.c        |   1 +
 src/backend/replication/logical/decode.c      | 110 ++++-
 .../replication/logical/reorderbuffer.c       |   3 +
 .../utils/activity/wait_event_names.txt       |   1 +
 src/include/access/heapam.h                   |   4 +-
 src/include/access/heaptoast.h                |   7 +-
 src/include/access/rewriteheap.h              |  13 +-
 src/include/access/tableam.h                  |  28 +-
 src/include/access/toast_helper.h             |  14 +-
 src/include/access/toast_internals.h          |   9 +-
 src/include/access/xlog_internal.h            |   2 +-
 src/include/access/xloginsert.h               |   1 +
 src/include/access/xlogrecord.h               |   8 +
 src/include/commands/repack.h                 |  17 +-
 src/include/utils/snapmgr.h                   |  12 +
 src/include/utils/snapshot.h                  |  27 +-
 .../injection_points/expected/repack.out      |  10 +-
 .../injection_points/specs/repack.spec        |  23 +-
 34 files changed, 1200 insertions(+), 342 deletions(-)

diff --git a/doc/src/sgml/mvcc.sgml b/doc/src/sgml/mvcc.sgml
index 9cb52302f23..8a6bb165ecb 100644
--- a/doc/src/sgml/mvcc.sgml
+++ b/doc/src/sgml/mvcc.sgml
@@ -1883,17 +1883,15 @@ SELECT pg_advisory_lock(q.id) FROM
    <title>Caveats</title>
 
    <para>
-    Some commands, currently only <link linkend="sql-truncate"><command>TRUNCATE</command></link>, the
-    table-rewriting forms of <link linkend="sql-altertable"><command>ALTER
-    TABLE</command></link> and <command>REPACK</command> with
-    the <literal>CONCURRENTLY</literal> option, are not
+    Some DDL commands, currently only <link linkend="sql-truncate"><command>TRUNCATE</command></link> and the
+    table-rewriting forms of <link linkend="sql-altertable"><command>ALTER TABLE</command></link>, are not
     MVCC-safe.  This means that after the truncation or rewrite commits, the
     table will appear empty to concurrent transactions, if they are using a
-    snapshot taken before the command committed.  This will only be an
+    snapshot taken before the DDL command committed.  This will only be an
     issue for a transaction that did not access the table in question
-    before the command started &mdash; any transaction that has done so
+    before the DDL command started &mdash; any transaction that has done so
     would hold at least an <literal>ACCESS SHARE</literal> table lock,
-    which would block the truncating or rewriting command until that transaction completes.
+    which would block the DDL command until that transaction completes.
     So these commands will not cause any apparent inconsistency in the
     table contents for successive queries on the target table, but they
     could cause visible inconsistency between the contents of the target
diff --git a/doc/src/sgml/ref/repack.sgml b/doc/src/sgml/ref/repack.sgml
index 0cb72b6b289..acbd92ef79d 100644
--- a/doc/src/sgml/ref/repack.sgml
+++ b/doc/src/sgml/ref/repack.sgml
@@ -300,15 +300,6 @@ REPACK [ ( <replaceable class="parameter">option</replaceable> [, ...] ) ] USING
        </listitem>
       </itemizedlist>
      </para>
-
-     <warning>
-      <para>
-       <command>REPACK</command> with the <literal>CONCURRENTLY</literal>
-       option is not MVCC-safe, see <xref linkend="mvcc-caveats"/> for
-       details.
-      </para>
-     </warning>
-
     </listitem>
    </varlistentry>
 
diff --git a/src/backend/access/common/toast_internals.c b/src/backend/access/common/toast_internals.c
index 77d42e7ed65..ba80fed0ca9 100644
--- a/src/backend/access/common/toast_internals.c
+++ b/src/backend/access/common/toast_internals.c
@@ -112,12 +112,15 @@ toast_compress_datum(Datum value, char cmethod)
  * rel: the main relation we're working with (not the toast rel!)
  * value: datum to be pushed to toast storage
  * oldexternal: if not NULL, toast pointer previously representing the datum
+ * rwstate: state needed for "raw insert".
+ * tup_main: the tuple whose attribute we're saving
  * options: options to be passed to heap_insert() for toast rows
  * ----------
  */
 Datum
 toast_save_datum(Relation rel, Datum value,
-				 varlena *oldexternal, uint32 options)
+				 varlena *oldexternal, RewriteState rwstate,
+				 HeapTuple tup_main, uint32 options)
 {
 	Relation	toastrel;
 	Relation   *toastidxs;
@@ -311,7 +314,40 @@ toast_save_datum(Relation rel, Datum value,
 
 		toasttup = heap_form_tuple(toasttupDesc, t_values, t_isnull);
 
-		heap_insert(toastrel, toasttup, mycid, options, NULL);
+		if (rwstate == NULL)
+		{
+			/*
+			 * With REUSE_XID we'd need regular freezing, however that is
+			 * currently implemented only as a part of VACUUM or in
+			 * rewriteheap.c. On the other hand, tuples containing a new XID
+			 * can be marked as frozen in special cases - see the current uses
+			 * of TABLE_INSERT_FROZEN.
+			 */
+			Assert((options & HEAP_INSERT_FROZEN) == 0 ||
+				   (options & TABLE_REUSE_XID) == 0);
+
+			/*
+			 * If an existing XID should be used, the entire visibility info
+			 * of the TOAST tuple should be equal to that of corresponding
+			 * tuple in the main table.
+			 */
+			if (options & TABLE_REUSE_XID)
+				rewrite_copy_visibility_info(toasttup, tup_main);
+
+			heap_insert(toastrel, toasttup, mycid, options, NULL);
+		}
+		else
+		{
+			/*
+			 * During heap rewrite, XID is always reused and the tuple is
+			 * always frozen - we do not expect the user to tell us what to
+			 * do.
+			 */
+			Assert((options & HEAP_INSERT_FROZEN) == 0 &&
+				   (options & TABLE_REUSE_XID) == 0);
+
+			rewrite_heap_tuple_no_chains(rwstate, tup_main, toasttup, true);
+		}
 
 		/*
 		 * Create the index entry.  We cheat a little here by not using
@@ -373,7 +409,8 @@ toast_save_datum(Relation rel, Datum value,
  * ----------
  */
 void
-toast_delete_datum(Relation rel, Datum value, bool is_speculative)
+toast_delete_datum(Relation rel, Datum value, bool is_speculative,
+				   TransactionId xid)
 {
 	varlena    *attr = (varlena *) DatumGetPointer(value);
 	varatt_external toast_pointer;
@@ -425,7 +462,7 @@ toast_delete_datum(Relation rel, Datum value, bool is_speculative)
 		if (is_speculative)
 			heap_abort_speculative(toastrel, &toasttup->t_self);
 		else
-			simple_heap_delete(toastrel, &toasttup->t_self);
+			simple_heap_delete(toastrel, &toasttup->t_self, xid);
 	}
 
 	/*
diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 9cdc221675b..35012434d27 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -63,7 +63,7 @@ static XLogRecPtr log_heap_update(Relation reln, Buffer oldbuf,
 								  Buffer newbuf, HeapTuple oldtup,
 								  HeapTuple newtup, HeapTuple old_key_tuple,
 								  bool all_visible_cleared, bool new_all_visible_cleared,
-								  bool walLogical);
+								  bool walLogical, TransactionId xid);
 #ifdef USE_ASSERT_CHECKING
 static void check_lock_if_inplace_updateable_rel(Relation relation,
 												 const ItemPointerData *otid,
@@ -2004,7 +2004,7 @@ void
 heap_insert(Relation relation, HeapTuple tup, CommandId cid,
 			uint32 options, BulkInsertState bistate)
 {
-	TransactionId xid = GetCurrentTransactionId();
+	TransactionId xid;
 	HeapTuple	heaptup;
 	Buffer		buffer;
 	Page		page;
@@ -2017,6 +2017,13 @@ heap_insert(Relation relation, HeapTuple tup, CommandId cid,
 
 	AssertHasSnapshotForToast(relation);
 
+	/* The caller might need to preserve the existing xmin. */
+	if ((options & TABLE_REUSE_XID) == 0)
+		xid = GetCurrentTransactionId();
+	else
+		xid = HeapTupleHeaderGetXmin(tup->t_data);
+	Assert(TransactionIdIsValid(xid));
+
 	/*
 	 * Fill in tuple header fields and toast the tuple if necessary.
 	 *
@@ -2159,6 +2166,13 @@ heap_insert(Relation relation, HeapTuple tup, CommandId cid,
 		/* filtering by origin on a row level is much more efficient */
 		XLogSetRecordFlags(XLOG_INCLUDE_ORIGIN);
 
+		/*
+		 * Even if we don't have XID assigned, valid XID is necessary for
+		 * recovery and streaming replication to work.
+		 */
+		if (options & TABLE_REUSE_XID)
+			XLogSetRecordXid(xid);
+
 		recptr = XLogInsert(RM_HEAP_ID, info);
 
 		PageSetLSN(page, recptr);
@@ -2236,7 +2250,7 @@ heap_prepare_insert(Relation relation, HeapTuple tup, TransactionId xid,
 		return tup;
 	}
 	else if (HeapTupleHasExternal(tup) || tup->t_len > TOAST_TUPLE_THRESHOLD)
-		return heap_toast_insert_or_update(relation, tup, NULL, options);
+		return heap_toast_insert_or_update(relation, tup, NULL, NULL, options);
 	else
 		return tup;
 }
@@ -2299,6 +2313,8 @@ heap_multi_insert(Relation relation, TupleTableSlot **slots, int ntuples,
 
 	/* currently not needed (thus unsupported) for heap_multi_insert() */
 	Assert(!(options & HEAP_INSERT_NO_LOGICAL));
+	/* Likewise. */
+	Assert(!(options & TABLE_REUSE_XID));
 
 	AssertHasSnapshotForToast(relation);
 
@@ -2715,11 +2731,11 @@ xmax_infomask_changed(uint16 new_infomask, uint16 old_infomask)
  */
 TM_Result
 heap_delete(Relation relation, const ItemPointerData *tid,
-			CommandId cid, uint32 options, Snapshot crosscheck,
+			TransactionId xid, CommandId cid, uint32 options,
+			Snapshot crosscheck,
 			bool wait, TM_FailureData *tmfd)
 {
 	TM_Result	result;
-	TransactionId xid = GetCurrentTransactionId();
 	ItemId		lp;
 	HeapTupleData tp;
 	Page		page;
@@ -2736,11 +2752,14 @@ heap_delete(Relation relation, const ItemPointerData *tid,
 	bool		all_visible_cleared = false;
 	HeapTuple	old_key_tuple = NULL;	/* replica identity of the tuple */
 	bool		old_key_copied = false;
+	bool		override_xid = false;
 
 	Assert(ItemPointerIsValid(tid));
 
 	AssertHasSnapshotForToast(relation);
 
+	Assert((options & TABLE_REUSE_XID) == 0);
+
 	/*
 	 * Forbid this during a parallel operation, lest it allocate a combo CID.
 	 * Other workers might need that combo CID for visibility checks, and we
@@ -2751,6 +2770,12 @@ heap_delete(Relation relation, const ItemPointerData *tid,
 				(errcode(ERRCODE_INVALID_TRANSACTION_STATE),
 				 errmsg("cannot delete tuples during a parallel operation")));
 
+	/* Caller can override the xid. */
+	if (!TransactionIdIsValid(xid))
+		xid = GetCurrentTransactionId();
+	else
+		override_xid = true;
+
 	block = ItemPointerGetBlockNumber(tid);
 	buffer = ReadBuffer(relation, block);
 	page = BufferGetPage(buffer);
@@ -3089,6 +3114,9 @@ l1:
 
 		/* filtering by origin on a row level is much more efficient */
 		XLogSetRecordFlags(XLOG_INCLUDE_ORIGIN);
+		/* See heap_insert() for details. */
+		if (override_xid)
+			XLogSetRecordXid(xid);
 
 		recptr = XLogInsert(RM_HEAP_ID, XLOG_HEAP_DELETE);
 
@@ -3115,7 +3143,7 @@ l1:
 		Assert(!HeapTupleHasExternal(&tp));
 	}
 	else if (HeapTupleHasExternal(&tp))
-		heap_toast_delete(relation, &tp, false);
+		heap_toast_delete(relation, &tp, false, xid);
 
 	/*
 	 * Mark tuple for invalidation from system caches at next command
@@ -3148,14 +3176,18 @@ l1:
  * the target tuple are not expected (for example, because we have a lock
  * on the relation associated with the tuple).  Any failure is reported
  * via ereport().
+ *
+ * XXX Add simple_heap_delete_xid() so that the signature of this function can
+ * stay intact?
  */
 void
-simple_heap_delete(Relation relation, const ItemPointerData *tid)
+simple_heap_delete(Relation relation, const ItemPointerData *tid,
+				   TransactionId xid)
 {
 	TM_Result	result;
 	TM_FailureData tmfd;
 
-	result = heap_delete(relation, tid,
+	result = heap_delete(relation, tid, xid,
 						 GetCurrentCommandId(true),
 						 0,
 						 InvalidSnapshot,
@@ -3204,7 +3236,7 @@ heap_update(Relation relation, const ItemPointerData *otid, HeapTuple newtup,
 			TU_UpdateIndexes *update_indexes)
 {
 	TM_Result	result;
-	TransactionId xid = GetCurrentTransactionId();
+	TransactionId xid;
 	Bitmapset  *hot_attrs;
 	Bitmapset  *sum_attrs;
 	Bitmapset  *key_attrs;
@@ -3267,6 +3299,13 @@ heap_update(Relation relation, const ItemPointerData *otid, HeapTuple newtup,
 	check_lock_if_inplace_updateable_rel(relation, otid, newtup);
 #endif
 
+	/* The caller might need to preserve the existing xmin. */
+	if ((options & TABLE_REUSE_XID) == 0)
+		xid = GetCurrentTransactionId();
+	else
+		xid = HeapTupleHeaderGetXmin(newtup->t_data);
+	Assert(TransactionIdIsValid(xid));
+
 	/*
 	 * Fetch the list of attributes to be checked for various operations.
 	 *
@@ -3871,7 +3910,8 @@ l2:
 		if (need_toast)
 		{
 			/* Note we always use WAL and FSM during updates */
-			heaptup = heap_toast_insert_or_update(relation, newtup, &oldtup, 0);
+			heaptup = heap_toast_insert_or_update(relation, newtup, &oldtup,
+												  NULL, options);
 			newtupsize = MAXALIGN(heaptup->t_len);
 		}
 		else
@@ -4099,7 +4139,9 @@ l2:
 								 old_key_tuple,
 								 all_visible_cleared,
 								 all_visible_cleared_new,
-								 walLogical);
+								 walLogical,
+								 options & TABLE_REUSE_XID ? xid : InvalidTransactionId);
+
 		if (newbuf != buffer)
 		{
 			PageSetLSN(newpage, recptr);
@@ -5297,7 +5339,9 @@ compute_new_xmax_infomask(TransactionId xmax, uint16 old_infomask,
 	uint16		new_infomask,
 				new_infomask2;
 
-	Assert(TransactionIdIsCurrentTransactionId(add_to_xmax));
+	/* REPACK (CONCURRENTLY) might not have XID assigned. */
+	Assert(TransactionIdIsCurrentTransactionId(add_to_xmax) ||
+		   !TransactionIdIsValid(GetTopTransactionIdIfAny()));
 
 l5:
 	new_infomask = 0;
@@ -6269,7 +6313,9 @@ heap_abort_speculative(Relation relation, const ItemPointerData *tid)
 	if (HeapTupleHasExternal(&tp))
 	{
 		Assert(!IsToastRelation(relation));
-		heap_toast_delete(relation, &tp, true);
+		/* XID overriding is not needed for speculative abort. */
+		Assert(TransactionIdIsValid(GetCurrentTransactionIdIfAny()));
+		heap_toast_delete(relation, &tp, true, GetCurrentTransactionId());
 	}
 
 	/*
@@ -6695,7 +6741,8 @@ FreezeMultiXactId(MultiXactId multi, uint16 t_infomask,
 		pagefrz->freeze_required = true;
 		return InvalidTransactionId;
 	}
-	else if (MultiXactIdPrecedes(multi, cutoffs->relminmxid))
+	else if (MultiXactIdIsValid(cutoffs->relminmxid) &&
+			 MultiXactIdPrecedes(multi, cutoffs->relminmxid))
 		ereport(ERROR,
 				(errcode(ERRCODE_DATA_CORRUPTED),
 				 errmsg_internal("found multixact %u from before relminmxid %u",
@@ -7199,7 +7246,8 @@ heap_prepare_freeze_tuple(HeapTupleHeader tuple,
 	else if (TransactionIdIsNormal(xid))
 	{
 		/* Raw xmax is normal XID */
-		if (TransactionIdPrecedes(xid, cutoffs->relfrozenxid))
+		if (TransactionIdIsValid(cutoffs->relfrozenxid) &&
+			TransactionIdPrecedes(xid, cutoffs->relfrozenxid))
 			ereport(ERROR,
 					(errcode(ERRCODE_DATA_CORRUPTED),
 					 errmsg_internal("found xmax %u from before relfrozenxid %u",
@@ -8776,7 +8824,7 @@ log_heap_update(Relation reln, Buffer oldbuf,
 				Buffer newbuf, HeapTuple oldtup, HeapTuple newtup,
 				HeapTuple old_key_tuple,
 				bool all_visible_cleared, bool new_all_visible_cleared,
-				bool walLogical)
+				bool walLogical, TransactionId xid)
 {
 	xl_heap_update xlrec;
 	xl_heap_header xlhdr;
@@ -8983,6 +9031,10 @@ log_heap_update(Relation reln, Buffer oldbuf,
 	/* filtering by origin on a row level is much more efficient */
 	XLogSetRecordFlags(XLOG_INCLUDE_ORIGIN);
 
+	/* See heap_insert() for details. */
+	if (TransactionIdIsValid(xid))
+		XLogSetRecordXid(xid);
+
 	recptr = XLogInsert(RM_HEAP_ID, info);
 
 	return recptr;
diff --git a/src/backend/access/heap/heapam_handler.c b/src/backend/access/heap/heapam_handler.c
index 1096d9d4dc9..1eec39e6ebb 100644
--- a/src/backend/access/heap/heapam_handler.c
+++ b/src/backend/access/heap/heapam_handler.c
@@ -50,11 +50,24 @@
 #include "utils/injection_point.h"
 #include "utils/rel.h"
 #include "utils/tuplesort.h"
-
-static Snapshot finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
-									 BlockNumber start, BlockNumber *end_p);
+#include "utils/wait_event_types.h"
+
+static void update_identity_index(ChangeContext *chgcxt,
+								  BlockNumber range_start,
+								  BlockNumber range_end);
+static Snapshot finalize_block_range(Relation rel_old, Relation rel_dst,
+									 TransactionId oldest_xmin,
+									 TransactionId xid_cutoff,
+									 MultiXactId multi_cutoff,
+									 ChangeContext *chgcxt, BlockNumber cur,
+									 BlockNumber start, BlockNumber *end_p,
+									 RewriteState *rwstate_p,
+									 BlockNumber *range_start_dst_p);
 static void reform_and_rewrite_tuple(TupleTableSlot *src, TupleTableSlot *reform,
 									 RewriteState rwstate);
+static void heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
+								   TupleTableSlot *reform,
+								   RewriteState rwstate);
 
 static bool SampleHeapTupleVisible(TableScanDesc scan, Buffer buffer,
 								   HeapTuple tuple,
@@ -199,7 +212,8 @@ heapam_tuple_complete_speculative(Relation relation, TupleTableSlot *slot,
 }
 
 static TM_Result
-heapam_tuple_delete(Relation relation, ItemPointer tid, CommandId cid,
+heapam_tuple_delete(Relation relation, ItemPointer tid, TransactionId xid,
+					CommandId cid,
 					uint32 options, Snapshot snapshot, Snapshot crosscheck,
 					bool wait, TM_FailureData *tmfd)
 {
@@ -208,7 +222,7 @@ heapam_tuple_delete(Relation relation, ItemPointer tid, CommandId cid,
 	 * the storage itself is cleaning the dead tuples by itself, it is the
 	 * time to call the index tuple deletion also.
 	 */
-	return heap_delete(relation, tid, cid, options, crosscheck, wait,
+	return heap_delete(relation, tid, xid, cid, options, crosscheck, wait,
 					   tmfd);
 }
 
@@ -610,25 +624,30 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	Snapshot	snapshot = NULL;
 	BlockNumber range_start = InvalidBlockNumber;
 	BlockNumber range_end = InvalidBlockNumber;
+	BlockNumber range_start_new = InvalidBlockNumber;
+	BlockNumber range_end_new = InvalidBlockNumber;
+	Relation	rel_dst;
 
 	/* Remember if it's a system catalog */
 	is_system_catalog = IsSystemRelation(OldHeap);
 
+	/* Determine the destination for the tuples. */
+	if (!concurrent || chgcxt->cc_dest_aux == NULL)
+		rel_dst = NewHeap;
+	else
+		rel_dst = chgcxt->cc_dest_aux->rel;
+
 	/*
 	 * Valid smgr_targblock implies something already wrote to the relation.
 	 * This may be harmless, but this function hasn't planned for it.
 	 */
-	Assert(RelationGetTargetBlock(NewHeap) == InvalidBlockNumber);
+	Assert(RelationGetTargetBlock(rel_dst) == InvalidBlockNumber);
 
 	/*
-	 * In non-concurrent mode, initialize the rewrite operation.  This is not
-	 * needed in concurrent mode.
+	 * Initialize the rewrite operation.
 	 */
-	if (!concurrent)
-		rwstate = begin_heap_rewrite(OldHeap, NewHeap, OldestXmin,
-									 *xid_cutoff, *multi_cutoff);
-	else
-		rwstate = NULL;
+	rwstate = begin_heap_rewrite(OldHeap, rel_dst, OldestXmin, *xid_cutoff,
+								 *multi_cutoff, concurrent);
 
 	/*
 	 * Set up sorting if wanted. CONCURRENTLY sorts the tuple w/o tuplesort,
@@ -700,6 +719,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		{
 			range_start = heapScan->rs_startblock;
 			range_end = range_start + repack_pages_per_snapshot;
+			range_start_new = RelationGetNumberOfBlocks(rel_dst);
 		}
 	}
 
@@ -717,24 +737,40 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 		InvalidateCatalogSnapshot();
 
 		/*
-		 * As there is no snapshot, our xmin should be invalid now.
-		 *
-		 * XXX xid can still be valid. The next patches in the series fix
-		 * that.
+		 * As there is no snapshot, our xmin should be invalid now. xid should
+		 * be invalid too because our transactions didn't have to do any
+		 * writes yet.
 		 */
 		Assert(!TransactionIdIsValid(MyProc->xmin));
+		Assert(!TransactionIdIsValid(MyProc->xid));
 
 		/*
 		 * Wait until the worker has the initial snapshot and retrieve it.
 		 */
 		snapshot = repack_get_snapshot(chgcxt);
+		chgcxt->cc_last_snapshot_xmin = snapshot->xmin;
+
+		/*
+		 * Since we currently do not freeze tuples when replaying data
+		 * changes, the replayed transactions should not precede the new value
+		 * of relfrozenxid. (Only transactions having XID >= snapshot->xmin
+		 * will be replayed.)
+		 *
+		 * Regarding multixacts, the corresponding cutoff is probably not
+		 * trivial to determine, however there should not be any multixacts in
+		 * the new relation at all: we do not replay tuple locking records an
+		 * that's ok because the tuple locks should no longer exist at the
+		 * moment we acquire AccessExclusiveLock on the old relation.
+		 */
+		if (TransactionIdFollows(*xid_cutoff, snapshot->xmin))
+			*xid_cutoff = snapshot->xmin;
 
 		PushActiveSnapshot(snapshot);
 	}
 
 	/*
 	 * Scan through the OldHeap, either in OldIndex order or sequentially;
-	 * copy each tuple into the NewHeap, or transiently to the tuplesort
+	 * copy each tuple into the rel_dst, or transiently to the tuplesort
 	 * module.  Note that we don't bother sorting dead tuples (they won't get
 	 * to the new table anyway).
 	 */
@@ -905,8 +941,12 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 			/* End of the current range or wraparound? */
 			if (blkno >= range_end || blkno < range_start)
-				snapshot = finalize_block_range(chgcxt, blkno, range_start,
-												&range_end);
+				snapshot = finalize_block_range(OldHeap, rel_dst,
+												OldestXmin, *xid_cutoff,
+												*multi_cutoff,
+												chgcxt, blkno, range_start,
+												&range_end, &rwstate,
+												&range_start_new);
 
 			/* Finally check the tuple visibility. */
 			LockBuffer(buf, BUFFER_LOCK_SHARE);
@@ -942,7 +982,7 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 			if (!concurrent)
 				reform_and_rewrite_tuple(slot, reform_slot, rwstate);
 			else
-				heap_insert_for_repack(chgcxt, slot, reform_slot);
+				heap_insert_for_repack(chgcxt, slot, reform_slot, rwstate);
 
 			/*
 			 * In indexscan mode and also VACUUM FULL, report increase in
@@ -958,6 +998,13 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 	{
 		XLogRecPtr	end_of_wal;
 
+		/* Write out any remaining tuples, and fsync if needed */
+		end_heap_rewrite(rwstate);
+
+		/* Finalize the last range in the new relation. */
+		range_end_new = RelationGetNumberOfBlocks(rel_dst);
+		update_identity_index(chgcxt, range_start_new, range_end_new);
+
 		/*
 		 * Process the changes belonging to the last range.
 		 */
@@ -969,11 +1016,16 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 		/*
 		 * There was an active transaction snapshot on entry, so push one
-		 * before return.
+		 * before return. While there is no active snapshot, also invalidate
+		 * the catalog snapshot so that the xmin horizons for VACUUM can
+		 * advance.
 		 */
 		PopActiveSnapshot();
+		InvalidateCatalogSnapshot();
+		Assert(!TransactionIdIsValid(MyProc->xmin));
+		Assert(!TransactionIdIsValid(MyProc->xid));
+		Assert(!HaveRegisteredOrActiveSnapshot());
 		PushActiveSnapshot(GetTransactionSnapshot());
-
 	}
 
 	if (indexScan != NULL)
@@ -1040,9 +1092,70 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
 
 	ExecDropSingleTupleTableSlot(reform_slot);
 
-	/* Write out any remaining tuples, and fsync if needed */
-	if (rwstate)
+	if (!concurrent)
+	{
+		/*
+		 * In the CONCURRENTLY case, we had to do away with the rewrite state
+		 * earlier.
+		 */
 		end_heap_rewrite(rwstate);
+	}
+}
+
+/*
+ * Scan range of the new / auxiliary relation and insert all the tuples into
+ * the identity index. range_end is the first block of the next range.
+ */
+static void
+update_identity_index(ChangeContext *chgcxt, BlockNumber range_start,
+					  BlockNumber range_end)
+{
+	RepackDest *dest;
+	BlockNumber first_block,
+				last_block;
+	ItemPointerData mintid,
+				maxtid;
+	TableScanDesc scan;
+	TupleTableSlot *slot;
+
+	/*
+	 * If an auxiliary table exists, it's the one we're copying the data into.
+	 */
+	if (chgcxt->cc_dest_aux)
+		dest = chgcxt->cc_dest_aux;
+	else
+		dest = &chgcxt->cc_dest;
+
+	first_block = range_start;
+
+	if (range_end > range_start)
+		last_block = range_end - 1;
+	else
+	{
+		Assert(range_start == range_end);
+
+		last_block = range_end;
+	}
+
+	ItemPointerSet(&mintid, first_block, FirstOffsetNumber);
+	ItemPointerSet(&maxtid, last_block, MaxOffsetNumber);
+
+	/* XXX flags? */
+	scan = table_beginscan_tidrange(dest->rel, SnapshotAny, &mintid, &maxtid,
+									0);
+	slot = table_slot_create(dest->rel, NULL);
+	while (table_scan_getnextslot(scan, ForwardScanDirection, slot))
+		ExecInsertIndexTuples(dest->rri, dest->estate, 0, slot, NIL, NULL);
+	ExecDropSingleTupleTableSlot(slot);
+	table_endscan(scan);
+
+	/*
+	 * Index insertion could have accessed catalog when checking constraints,
+	 * so make sure that we no longer block the progress of xmin horizons for
+	 * VACUUM. (It's ok to have a snapshot throughout range scan, so there's
+	 * no point in doing this invalidation more often.)
+	 */
+	InvalidateCatalogSnapshot();
 }
 
 /*
@@ -1055,20 +1168,47 @@ heapam_relation_copy_for_cluster(Relation OldHeap, Relation NewHeap,
  * first block beyond the new range.
  *
  * Return the snapshot for the scan of the new range.
+ *
+ * TODO Consider a structure to accommodate (most of) the arguments.
  */
 static Snapshot
-finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
-					 BlockNumber start, BlockNumber *end_p)
+finalize_block_range(Relation rel_old, Relation rel_dst,
+					 TransactionId oldest_xmin,
+					 TransactionId xid_cutoff,
+					 MultiXactId multi_cutoff,
+					 ChangeContext *chgcxt, BlockNumber cur,
+					 BlockNumber start, BlockNumber *end_p,
+					 RewriteState *rwstate_p,
+					 BlockNumber *range_start_dst_p)
 {
 	BlockNumber end = *end_p;
+	RewriteState rwstate = *rwstate_p;
 	XLogRecPtr	end_of_wal;
 	Snapshot	snapshot;
+	BlockNumber range_end_dst;
 
 	/*
 	 * Wait here when testing how snapshot is changed at page boundary.
 	 */
 	INJECTION_POINT("repack-concurrently-new-range", NULL);
 
+	/*
+	 * Make sure the data is flushed to file before changes can be applied to
+	 * it. (Nothing of it should be in shared buffers so far.)
+	 */
+	end_heap_rewrite(rwstate);
+
+	/*
+	 * Now that the data can be accessed via shared buffers, scan it and add
+	 * each tuple to the identity index - this is necessary to update the
+	 * concurrent data changes below. We could not do that earlier because
+	 * even insertion into index might need to fetch heap tuples, in order to
+	 * check unique or exclusion constraints.
+	 */
+	range_end_dst = RelationGetNumberOfBlocks(rel_dst);
+	update_identity_index(chgcxt, *range_start_dst_p, range_end_dst);
+	*range_start_dst_p = range_end_dst;
+
 	/*
 	 * Decode all the concurrent data changes committed so far before
 	 * requesting the next snapshot - these changes are applicable on top of
@@ -1101,6 +1241,7 @@ finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
 
 	/* See above. */
 	Assert(!TransactionIdIsValid(MyProc->xmin));
+	Assert(!TransactionIdIsValid(MyProc->xid));
 
 	/*
 	 * XXX It might be worth Assert(CatalogSnapshot == NULL) here, however
@@ -1123,8 +1264,20 @@ finalize_block_range(ChangeContext *chgcxt, BlockNumber cur,
 	 * next batch of decoded changes.
 	 */
 	snapshot = repack_get_snapshot(chgcxt);
+
+	Assert(TransactionIdFollowsOrEquals(snapshot->xmin,
+										chgcxt->cc_last_snapshot_xmin));
+	chgcxt->cc_last_snapshot_xmin = snapshot->xmin;
+
 	PushActiveSnapshot(snapshot);
 
+	/*
+	 * Prepare for bulk insert of the next set of tuples. We rely on it to
+	 * start on a new page, even if the last existing page is not full.
+	 */
+	*rwstate_p = begin_heap_rewrite(rel_old, rel_dst, oldest_xmin, xid_cutoff,
+									multi_cutoff, true);
+
 	return snapshot;
 }
 
@@ -2554,6 +2707,64 @@ reform_and_rewrite_tuple(TupleTableSlot *src, TupleTableSlot *reform,
 		heap_freetuple(newtuple);
 }
 
+/*
+ * Insert tuple when processing REPACK CONCURRENTLY.
+ *
+ * 'reform' is a slot to use for tuple "reforming", typically to get set
+ * values of dropped columns to NULL.
+ *
+ * We pass the NO_LOGICAL flag to heap_insert() in order to skip logical
+ * decoding: as soon as REPACK CONCURRENTLY swaps the relation files, it drops
+ * this relation, so no logical replication subscription should need the data.
+ */
+static void
+heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
+					   TupleTableSlot *reform, RewriteState rwstate)
+{
+	HeapTuple	tuple,
+				new_tuple;
+	TransactionId xid;
+	bool		shouldFree,
+				shouldFreeNew;
+	TupleTableSlot *slot;
+	bool		freeze;
+
+	if (chgcxt->cc_dest_aux)
+	{
+		/* Will freeze when copying data to the new table. */
+		freeze = false;
+	}
+	else
+		freeze = true;
+
+	Assert(TTS_IS_BUFFERTUPLE(src));
+	tuple = ExecFetchSlotHeapTuple(src, false, &shouldFree);
+	xid = HeapTupleHeaderGetXmin(tuple->t_data);
+	if (reform != NULL && tuple_needs_reform(tuple, src->tts_tupleDescriptor))
+	{
+		clear_dropped_attributes(tuple, reform);
+		slot = reform;
+	}
+	else
+		slot = src;
+
+	/* Make sure we have a copy of the tuple, and set its XID. */
+	new_tuple = ExecFetchSlotHeapTuple(slot, true, &shouldFreeNew);
+	HeapTupleHeaderSetXmin(new_tuple->t_data, xid);
+
+	/* Perform the insertion. */
+	rewrite_heap_tuple_no_chains(rwstate, tuple, new_tuple, freeze);
+
+	/* The insertion shouldn't have caused XID assignment. */
+	Assert(!TransactionIdIsValid(GetCurrentTransactionIdIfAny()));
+
+	/* Cleanup. */
+	if (shouldFree)
+		heap_freetuple(tuple);
+	if (shouldFreeNew)
+		heap_freetuple(new_tuple);
+}
+
 /*
  * Check visibility of the tuple.
  */
diff --git a/src/backend/access/heap/heapam_visibility.c b/src/backend/access/heap/heapam_visibility.c
index 361b76e5065..f37e4848667 100644
--- a/src/backend/access/heap/heapam_visibility.c
+++ b/src/backend/access/heap/heapam_visibility.c
@@ -1363,6 +1363,65 @@ HeapTupleSatisfiesNonVacuumable(HeapTuple htup, Snapshot snapshot,
 	return res != HEAPTUPLE_DEAD;
 }
 
+/*
+ * HeapTupleSatisfiesNewHeap
+ *		Consider all transactions committed or current.
+ *
+ * See SNAPSHOT_NEW_HEAP's definition for the intended behaviour.
+ *
+ */
+static bool
+HeapTupleSatisfiesNewHeap(HeapTuple htup, Snapshot snapshot, Buffer buffer)
+{
+	HeapTupleHeader tuple = htup->t_data;
+
+	Assert(ItemPointerIsValid(&htup->t_self));
+	Assert(htup->t_tableOid != InvalidOid);
+
+	/* xmin should always be there. */
+	Assert(TransactionIdIsValid(HeapTupleHeaderGetXmin(tuple)));
+
+	/*
+	 * No one should have the chance to set XMIN_INVALID until REPACK has
+	 * finished, and we do not need it.
+	 */
+	Assert(!HeapTupleHeaderXminInvalid(tuple));
+
+	/*
+	 * Unlike that, XMIN_COMMITTED might have been set earlier by REPACK
+	 * itself, but we don't need it here.
+	 */
+
+	/*
+	 * Set XMIN_COMMITTED to make the next checks (by any snapshot) faster.
+	 *
+	 * TODO Set the flag in the initial load and in
+	 * apply_concurrent_changes().
+	 */
+	SetHintBits(tuple, buffer, HEAP_XMIN_COMMITTED,
+				HeapTupleHeaderGetRawXmax(tuple));
+
+	/* Inserted tuples have XMAX_INVALID set. */
+	if (tuple->t_infomask & HEAP_XMAX_INVALID)
+		return true;
+
+	/*
+	 * REPACK could have set XMAX_COMMITTED during UPDATE or DELETE, or below.
+	 */
+	if (tuple->t_infomask & HEAP_XMAX_COMMITTED)
+		return false;
+
+	if (!TransactionIdIsValid(HeapTupleHeaderGetRawXmax(tuple)))
+		return true;
+
+	/*
+	 * Set XMAX_COMMITTED to make the next checks faster.
+	 */
+	SetHintBits(tuple, buffer, HEAP_XMAX_COMMITTED,
+				HeapTupleHeaderGetRawXmax(tuple));
+
+	return false;
+}
 
 /*
  * HeapTupleIsSurelyDead
@@ -1747,6 +1806,9 @@ HeapTupleSatisfiesVisibility(HeapTuple htup, Snapshot snapshot, Buffer buffer)
 			return HeapTupleSatisfiesHistoricMVCC(htup, snapshot, buffer);
 		case SNAPSHOT_NON_VACUUMABLE:
 			return HeapTupleSatisfiesNonVacuumable(htup, snapshot, buffer);
+		case SNAPSHOT_NEW_HEAP:
+			return HeapTupleSatisfiesNewHeap(htup, snapshot, buffer);
+
 	}
 
 	return false;				/* keep compiler quiet */
diff --git a/src/backend/access/heap/heaptoast.c b/src/backend/access/heap/heaptoast.c
index 03f885a25b0..4a6a674b80e 100644
--- a/src/backend/access/heap/heaptoast.c
+++ b/src/backend/access/heap/heaptoast.c
@@ -40,7 +40,8 @@
  * ----------
  */
 void
-heap_toast_delete(Relation rel, HeapTuple oldtup, bool is_speculative)
+heap_toast_delete(Relation rel, HeapTuple oldtup, bool is_speculative,
+				  TransactionId xid)
 {
 	TupleDesc	tupleDesc;
 	Datum		toast_values[MaxHeapAttributeNumber];
@@ -70,7 +71,8 @@ heap_toast_delete(Relation rel, HeapTuple oldtup, bool is_speculative)
 	heap_deform_tuple(oldtup, tupleDesc, toast_values, toast_isnull);
 
 	/* Do the real work. */
-	toast_delete_external(rel, toast_values, toast_isnull, is_speculative);
+	toast_delete_external(rel, toast_values, toast_isnull, is_speculative,
+						  xid);
 }
 
 
@@ -84,6 +86,7 @@ heap_toast_delete(Relation rel, HeapTuple oldtup, bool is_speculative)
  *	newtup: the candidate new tuple to be inserted
  *	oldtup: the old row version for UPDATE, or NULL for INSERT
  *	options: options to be passed to heap_insert() for toast rows
+ *	rwstate: if valid, use raw_heap_insert()
  * Result:
  *	either newtup if no toasting is needed, or a palloc'd modified tuple
  *	that is what should actually get stored
@@ -94,7 +97,7 @@ heap_toast_delete(Relation rel, HeapTuple oldtup, bool is_speculative)
  */
 HeapTuple
 heap_toast_insert_or_update(Relation rel, HeapTuple newtup, HeapTuple oldtup,
-							uint32 options)
+							RewriteState rwstate, uint32 options)
 {
 	HeapTuple	result_tuple;
 	TupleDesc	tupleDesc;
@@ -109,6 +112,7 @@ heap_toast_insert_or_update(Relation rel, HeapTuple newtup, HeapTuple oldtup,
 	Datum		toast_oldvalues[MaxHeapAttributeNumber];
 	ToastAttrInfo toast_attr[MaxHeapAttributeNumber];
 	ToastTupleContext ttc;
+	TransactionId xid;
 
 	/*
 	 * Ignore the INSERT_SPECULATIVE option. Speculative insertions/super
@@ -156,6 +160,13 @@ heap_toast_insert_or_update(Relation rel, HeapTuple newtup, HeapTuple oldtup,
 	ttc.ttc_attr = toast_attr;
 	toast_tuple_init(&ttc);
 
+	/*
+	 * raw_heap_insert() may be needed for insertion into the TOAST table. In
+	 * that case, visibility information will be retrieved from 'newtup'.
+	 */
+	ttc.ttc_rwstate = rwstate;
+	ttc.ttc_tup_main = newtup;
+
 	/* ----------
 	 * Compress and/or save external until data fits into target length
 	 *
@@ -330,7 +341,11 @@ heap_toast_insert_or_update(Relation rel, HeapTuple newtup, HeapTuple oldtup,
 	else
 		result_tuple = newtup;
 
-	toast_tuple_cleanup(&ttc);
+	if (options & TABLE_REUSE_XID)
+		xid = HeapTupleHeaderGetXmin(newtup->t_data);
+	else
+		xid = InvalidTransactionId; /* The current transaction. */
+	toast_tuple_cleanup(&ttc, xid);
 
 	return result_tuple;
 }
diff --git a/src/backend/access/heap/rewriteheap.c b/src/backend/access/heap/rewriteheap.c
index 0ccd392c3dc..809dfba1b55 100644
--- a/src/backend/access/heap/rewriteheap.c
+++ b/src/backend/access/heap/rewriteheap.c
@@ -107,6 +107,7 @@
 #include "access/heapam.h"
 #include "access/heapam_xlog.h"
 #include "access/heaptoast.h"
+#include "access/multixact.h"
 #include "access/rewriteheap.h"
 #include "access/transam.h"
 #include "access/xact.h"
@@ -119,6 +120,7 @@
 #include "storage/bufmgr.h"
 #include "storage/bulk_write.h"
 #include "storage/fd.h"
+#include "storage/lmgr.h"
 #include "storage/procarray.h"
 #include "utils/memutils.h"
 #include "utils/rel.h"
@@ -151,6 +153,12 @@ typedef struct RewriteStateData
 	HTAB	   *rs_old_new_tid_map; /* unmatched B tuples */
 	HTAB	   *rs_logical_mappings;	/* logical remapping files */
 	uint32		rs_num_rewrite_mappings;	/* # in memory mappings */
+
+	/*
+	 * If this is initialized, raw_heap_insert() is also used for TOAST
+	 * relation.
+	 */
+	struct RewriteStateData *toast;
 } RewriteStateData;
 
 /*
@@ -211,6 +219,11 @@ typedef struct RewriteMappingDataEntry
 
 
 /* prototypes for internal functions */
+static RewriteState begin_heap_rewrite_common(Relation old_heap,
+											  Relation new_heap,
+											  TransactionId oldest_xmin,
+											  TransactionId freeze_xid,
+											  MultiXactId cutoff_multi);
 static void raw_heap_insert(RewriteState state, HeapTuple tup);
 
 /* internal logical remapping prototypes */
@@ -227,18 +240,19 @@ static void logical_end_heap_rewrite(RewriteState state);
  * oldest_xmin	xid used by the caller to determine which tuples are dead
  * freeze_xid	xid before which tuples will be frozen
  * cutoff_multi	multixact before which multis will be removed
+ * no_chains	only raw insert (and freezing), do not care of HOT chains
  *
  * Returns an opaque RewriteState, allocated in current memory context,
  * to be used in subsequent calls to the other functions.
  */
 RewriteState
 begin_heap_rewrite(Relation old_heap, Relation new_heap, TransactionId oldest_xmin,
-				   TransactionId freeze_xid, MultiXactId cutoff_multi)
+				   TransactionId freeze_xid, MultiXactId cutoff_multi,
+				   bool no_chains)
 {
-	RewriteState state;
 	MemoryContext rw_cxt;
 	MemoryContext old_cxt;
-	HASHCTL		hash_ctl;
+	RewriteState state;
 
 	/*
 	 * To ease cleanup, make a separate context that will contain the
@@ -249,9 +263,99 @@ begin_heap_rewrite(Relation old_heap, Relation new_heap, TransactionId oldest_xm
 								   ALLOCSET_DEFAULT_SIZES);
 	old_cxt = MemoryContextSwitchTo(rw_cxt);
 
+	state = begin_heap_rewrite_common(old_heap, new_heap, oldest_xmin,
+									  freeze_xid, cutoff_multi);
+	state->rs_cxt = rw_cxt;
+
+	if (!no_chains)
+	{
+		HASHCTL		hash_ctl;
+
+		/* Initialize hash tables used to track update chains */
+		hash_ctl.keysize = sizeof(TidHashKey);
+		hash_ctl.entrysize = sizeof(UnresolvedTupData);
+		hash_ctl.hcxt = state->rs_cxt;
+
+		state->rs_unresolved_tups =
+			hash_create("Rewrite / Unresolved ctids",
+						128,	/* arbitrary initial size */
+						&hash_ctl,
+						HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
+
+		hash_ctl.entrysize = sizeof(OldToNewMappingData);
+
+		state->rs_old_new_tid_map =
+			hash_create("Rewrite / Old to new tid map",
+						128,	/* arbitrary initial size */
+						&hash_ctl,
+						HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
+
+		logical_begin_heap_rewrite(state);
+	}
+	else
+	{
+		Oid			toastid;
+
+		/*
+		 * The current user of this mode, REPACK (CONCURRENTLY), does not want
+		 * XID assigned at all, so "raw insert" is also used for TOAST.
+		 *
+		 * XXX "raw insert" would be helpful for REPACK w/o CONCURRENTLY too,
+		 * as it writes the whole pages to WAL. The problem here is that it
+		 * might check the existence of chunk OIDs in the new TOAST relation
+		 * (see the hacks with rd_toastoid in toast_save_datum()), and that
+		 * doesn't work while the rewrite is still in progress: the relation
+		 * pages are not guaranteed to be flushed to disk (and read into
+		 * shared buffers) before the end of the rewrite.
+		 */
+		toastid = new_heap->rd_rel->reltoastrelid;
+		if (OidIsValid(toastid))
+		{
+			Oid			toastid_old;
+			Relation	toast_rel;
+			Relation	toast_rel_old = NULL;
+
+			/* New relation's TOAST should already be locked. */
+			Assert(CheckRelationOidLockedByMe(toastid, AccessExclusiveLock,
+											  false));
+			toast_rel = table_open(toastid, NoLock);
+
+			toastid_old = old_heap ? old_heap->rd_rel->reltoastrelid :
+				InvalidTransactionId;
+			if (OidIsValid(toastid_old))
+			{
+				/*
+				 * Currently we do not lock the old relation's TOAST. Use the
+				 * same lock mode we use for the parent relation.
+				 */
+				toast_rel_old = table_open(toastid_old,
+										   ShareUpdateExclusiveLock);
+			}
+
+			/* Create the state for TOAST insertions. */
+			state->toast = begin_heap_rewrite_common(toast_rel_old,
+													 toast_rel,
+													 oldest_xmin,
+													 freeze_xid,
+													 cutoff_multi);
+		}
+	}
+
+	MemoryContextSwitchTo(old_cxt);
+
+	return state;
+}
+
+static RewriteState
+begin_heap_rewrite_common(Relation old_heap, Relation new_heap,
+						  TransactionId oldest_xmin,
+						  TransactionId freeze_xid, MultiXactId cutoff_multi)
+
+{
+	RewriteState state;
+
 	/* Create and fill in the state struct */
 	state = palloc0_object(RewriteStateData);
-
 	state->rs_old_rel = old_heap;
 	state->rs_new_rel = new_heap;
 	state->rs_buffer = NULL;
@@ -260,32 +364,8 @@ begin_heap_rewrite(Relation old_heap, Relation new_heap, TransactionId oldest_xm
 	state->rs_oldest_xmin = oldest_xmin;
 	state->rs_freeze_xid = freeze_xid;
 	state->rs_cutoff_multi = cutoff_multi;
-	state->rs_cxt = rw_cxt;
 	state->rs_bulkstate = smgr_bulk_start_rel(new_heap, MAIN_FORKNUM);
 
-	/* Initialize hash tables used to track update chains */
-	hash_ctl.keysize = sizeof(TidHashKey);
-	hash_ctl.entrysize = sizeof(UnresolvedTupData);
-	hash_ctl.hcxt = state->rs_cxt;
-
-	state->rs_unresolved_tups =
-		hash_create("Rewrite / Unresolved ctids",
-					128,		/* arbitrary initial size */
-					&hash_ctl,
-					HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
-
-	hash_ctl.entrysize = sizeof(OldToNewMappingData);
-
-	state->rs_old_new_tid_map =
-		hash_create("Rewrite / Old to new tid map",
-					128,		/* arbitrary initial size */
-					&hash_ctl,
-					HASH_ELEM | HASH_BLOBS | HASH_CONTEXT);
-
-	MemoryContextSwitchTo(old_cxt);
-
-	logical_begin_heap_rewrite(state);
-
 	return state;
 }
 
@@ -300,16 +380,20 @@ end_heap_rewrite(RewriteState state)
 	HASH_SEQ_STATUS seq_status;
 	UnresolvedTup unresolved;
 
-	/*
-	 * Write any remaining tuples in the UnresolvedTups table. If we have any
-	 * left, they should in fact be dead, but let's err on the safe side.
-	 */
-	hash_seq_init(&seq_status, state->rs_unresolved_tups);
-
-	while ((unresolved = hash_seq_search(&seq_status)) != NULL)
+	if (state->rs_unresolved_tups)
 	{
-		ItemPointerSetInvalid(&unresolved->tuple->t_data->t_ctid);
-		raw_heap_insert(state, unresolved->tuple);
+		/*
+		 * Write any remaining tuples in the UnresolvedTups table. If we have
+		 * any left, they should in fact be dead, but let's err on the safe
+		 * side.
+		 */
+		hash_seq_init(&seq_status, state->rs_unresolved_tups);
+
+		while ((unresolved = hash_seq_search(&seq_status)) != NULL)
+		{
+			ItemPointerSetInvalid(&unresolved->tuple->t_data->t_ctid);
+			raw_heap_insert(state, unresolved->tuple);
+		}
 	}
 
 	/* Write the last page, if any */
@@ -318,10 +402,29 @@ end_heap_rewrite(RewriteState state)
 		smgr_bulk_write(state->rs_bulkstate, state->rs_blockno, state->rs_buffer, true);
 		state->rs_buffer = NULL;
 	}
-
 	smgr_bulk_finish(state->rs_bulkstate);
 
-	logical_end_heap_rewrite(state);
+	/* The same for TOAST */
+	if (state->toast)
+	{
+		RewriteState toast = state->toast;
+
+		if (toast->rs_buffer)
+		{
+			smgr_bulk_write(toast->rs_bulkstate, toast->rs_blockno,
+							toast->rs_buffer, true);
+			toast->rs_buffer = NULL;
+		}
+		smgr_bulk_finish(toast->rs_bulkstate);
+
+		/* Close relation(s) opened by begin_heap_rewrite(). */
+		table_close(toast->rs_new_rel, NoLock);
+		if (toast->rs_old_rel)
+			table_close(toast->rs_old_rel, ShareUpdateExclusiveLock);
+	}
+
+	if (state->rs_logical_rewrite)
+		logical_end_heap_rewrite(state);
 
 	/* Deleting the context frees everything */
 	MemoryContextDelete(state->rs_cxt);
@@ -350,30 +453,13 @@ rewrite_heap_tuple(RewriteState state,
 
 	old_cxt = MemoryContextSwitchTo(state->rs_cxt);
 
-	/*
-	 * Copy the original tuple's visibility information into new_tuple.
-	 *
-	 * XXX we might later need to copy some t_infomask2 bits, too? Right now,
-	 * we intentionally clear the HOT status bits.
-	 */
-	memcpy(&new_tuple->t_data->t_choice.t_heap,
-		   &old_tuple->t_data->t_choice.t_heap,
-		   sizeof(HeapTupleFields));
-
-	new_tuple->t_data->t_infomask &= ~HEAP_XACT_MASK;
-	new_tuple->t_data->t_infomask2 &= ~HEAP2_XACT_MASK;
-	new_tuple->t_data->t_infomask |=
-		old_tuple->t_data->t_infomask & HEAP_XACT_MASK;
+	rewrite_copy_visibility_info(new_tuple, old_tuple);
 
 	/*
 	 * While we have our hands on the tuple, we may as well freeze any
 	 * eligible xmin or xmax, so that future VACUUM effort can be saved.
 	 */
-	heap_freeze_tuple(new_tuple->t_data,
-					  state->rs_old_rel->rd_rel->relfrozenxid,
-					  state->rs_old_rel->rd_rel->relminmxid,
-					  state->rs_freeze_xid,
-					  state->rs_cutoff_multi);
+	rewrite_freeze_tuple(state, new_tuple);
 
 	/*
 	 * Invalid ctid means that ctid should point to the tuple itself. We'll
@@ -534,6 +620,24 @@ rewrite_heap_tuple(RewriteState state,
 	MemoryContextSwitchTo(old_cxt);
 }
 
+/*
+ * Like rewrite_heap_tuple(), but do not care about hot chains. The user
+ * should have used the appropriate snapshot to pick at most one tuple of the
+ * chain - this is typical for REPACK (CONCURRENTLY).
+ */
+void
+rewrite_heap_tuple_no_chains(RewriteState state, HeapTuple old_tuple,
+							 HeapTuple new_tuple, bool freeze)
+{
+	if (new_tuple != old_tuple)
+		rewrite_copy_visibility_info(new_tuple, old_tuple);
+
+	if (freeze)
+		rewrite_freeze_tuple(state, new_tuple);
+
+	raw_heap_insert(state, new_tuple);
+}
+
 /*
  * Register a dead tuple with an ongoing rewrite. Dead tuples are not
  * copied to the new table, but we still make note of them so that we
@@ -628,7 +732,7 @@ raw_heap_insert(RewriteState state, HeapTuple tup)
 		options |= HEAP_INSERT_NO_LOGICAL;
 
 		heaptup = heap_toast_insert_or_update(state->rs_new_rel, tup, NULL,
-											  options);
+											  state->toast, options);
 	}
 	else
 		heaptup = tup;
@@ -704,6 +808,47 @@ raw_heap_insert(RewriteState state, HeapTuple tup)
 		heap_freetuple(heaptup);
 }
 
+/*
+ * Freeze tuple. 'old_tuple' provides the initial visibility information.
+ */
+void
+rewrite_freeze_tuple(RewriteState state, HeapTuple tuple)
+{
+	TransactionId relfrozenxid = InvalidTransactionId;
+	MultiXactId relminmxid = InvalidMultiXactId;
+
+	/* The old relation may be missing if dealing with TOAST. */
+	if (state->rs_old_rel)
+	{
+		relfrozenxid = state->rs_old_rel->rd_rel->relfrozenxid;
+		relminmxid = state->rs_old_rel->rd_rel->relminmxid;
+	}
+
+	heap_freeze_tuple(tuple->t_data, relfrozenxid, relminmxid,
+					  state->rs_freeze_xid,
+					  state->rs_cutoff_multi);
+}
+
+/*
+ * Copy the old tuple's visibility information into the new tuple.
+ */
+void
+rewrite_copy_visibility_info(HeapTuple new_tuple, HeapTuple old_tuple)
+{
+	/*
+	 * XXX we might later need to copy some t_infomask2 bits, too? Right now,
+	 * we intentionally clear the HOT status bits.
+	 */
+	memcpy(&new_tuple->t_data->t_choice.t_heap,
+		   &old_tuple->t_data->t_choice.t_heap,
+		   sizeof(HeapTupleFields));
+
+	new_tuple->t_data->t_infomask &= ~HEAP_XACT_MASK;
+	new_tuple->t_data->t_infomask2 &= ~HEAP2_XACT_MASK;
+	new_tuple->t_data->t_infomask |=
+		old_tuple->t_data->t_infomask & HEAP_XACT_MASK;
+}
+
 /* ------------------------------------------------------------------------
  * Logical rewrite support
  *
diff --git a/src/backend/access/table/tableam.c b/src/backend/access/table/tableam.c
index 68ff0966f1c..43b853d949e 100644
--- a/src/backend/access/table/tableam.c
+++ b/src/backend/access/table/tableam.c
@@ -319,6 +319,7 @@ simple_table_tuple_delete(Relation rel, ItemPointer tid, Snapshot snapshot)
 	TM_FailureData tmfd;
 
 	result = table_tuple_delete(rel, tid,
+								InvalidTransactionId,
 								GetCurrentCommandId(true),
 								0, snapshot, InvalidSnapshot,
 								true /* wait for commit */ ,
diff --git a/src/backend/access/table/toast_helper.c b/src/backend/access/table/toast_helper.c
index 2f2022d9951..e0c3af410d4 100644
--- a/src/backend/access/table/toast_helper.c
+++ b/src/backend/access/table/toast_helper.c
@@ -261,7 +261,7 @@ toast_tuple_externalize(ToastTupleContext *ttc, int attribute, uint32 options)
 
 	attr->tai_colflags |= TOASTCOL_IGNORE;
 	*value = toast_save_datum(ttc->ttc_rel, old_value, attr->tai_oldexternal,
-							  options);
+							  ttc->ttc_rwstate, ttc->ttc_tup_main, options);
 	if ((attr->tai_colflags & TOASTCOL_NEEDS_FREE) != 0)
 		pfree(DatumGetPointer(old_value));
 	attr->tai_colflags |= TOASTCOL_NEEDS_FREE;
@@ -272,7 +272,7 @@ toast_tuple_externalize(ToastTupleContext *ttc, int attribute, uint32 options)
  * Perform appropriate cleanup after one tuple has been subjected to TOAST.
  */
 void
-toast_tuple_cleanup(ToastTupleContext *ttc)
+toast_tuple_cleanup(ToastTupleContext *ttc, TransactionId xid)
 {
 	TupleDesc	tupleDesc = ttc->ttc_rel->rd_att;
 	int			numAttrs = tupleDesc->natts;
@@ -305,7 +305,8 @@ toast_tuple_cleanup(ToastTupleContext *ttc)
 			ToastAttrInfo *attr = &ttc->ttc_attr[i];
 
 			if ((attr->tai_colflags & TOASTCOL_NEEDS_DELETE_OLD) != 0)
-				toast_delete_datum(ttc->ttc_rel, ttc->ttc_oldvalues[i], false);
+				toast_delete_datum(ttc->ttc_rel, ttc->ttc_oldvalues[i], false,
+								   xid);
 		}
 	}
 }
@@ -316,7 +317,7 @@ toast_tuple_cleanup(ToastTupleContext *ttc)
  */
 void
 toast_delete_external(Relation rel, const Datum *values, const bool *isnull,
-					  bool is_speculative)
+					  bool is_speculative, TransactionId xid)
 {
 	TupleDesc	tupleDesc = rel->rd_att;
 	int			numAttrs = tupleDesc->natts;
@@ -331,7 +332,7 @@ toast_delete_external(Relation rel, const Datum *values, const bool *isnull,
 			if (isnull[i])
 				continue;
 			else if (VARATT_IS_EXTERNAL_ONDISK(DatumGetPointer(value)))
-				toast_delete_datum(rel, value, is_speculative);
+				toast_delete_datum(rel, value, is_speculative, xid);
 		}
 	}
 }
diff --git a/src/backend/access/transam/xloginsert.c b/src/backend/access/transam/xloginsert.c
index f2e10b82b7d..ed20a35f7d6 100644
--- a/src/backend/access/transam/xloginsert.c
+++ b/src/backend/access/transam/xloginsert.c
@@ -105,6 +105,9 @@ static uint64 mainrdata_len;	/* total # of bytes in chain */
 /* flags for the in-progress insertion */
 static uint8 curinsert_flags = 0;
 
+/* XID to override the XID of the current transaction. */
+static TransactionId curinsert_xid = InvalidTransactionId;
+
 /*
  * These are used to hold the record header while constructing a record.
  * 'hdr_scratch' is not a plain variable, but is palloc'd at initialization,
@@ -235,6 +238,7 @@ XLogResetInsertion(void)
 	mainrdata_len = 0;
 	mainrdata_last = (XLogRecData *) &mainrdata_head;
 	curinsert_flags = 0;
+	curinsert_xid = InvalidTransactionId;
 	begininsert_called = false;
 }
 
@@ -467,6 +471,18 @@ XLogSetRecordFlags(uint8 flags)
 	curinsert_flags |= flags;
 }
 
+/*
+ * Set XID status flags for the upcoming WAL record.
+ *
+ * Useful when creating WAL records on behalf of another transaction.
+ */
+void
+XLogSetRecordXid(TransactionId xid)
+{
+	Assert(begininsert_called);
+	curinsert_xid = xid;
+}
+
 /*
  * Insert an XLOG record having the specified RMID and info bytes, with the
  * body of the record being the data and buffer references registered earlier
@@ -928,6 +944,12 @@ XLogRecordAssemble(RmgrId rmid, uint8 info,
 	{
 		TransactionId xid = GetTopTransactionIdIfAny();
 
+		/*
+		 * On curinsert_xid: if it's set, it's only for recovery and streaming
+		 * replication to work. On the other hand, the record shouldn't be
+		 * logically decoded, so we don't care if the toplevel XID is invalid.
+		 */
+
 		/* Set the flag that the top xid is included in the WAL */
 		*topxid_included = true;
 
@@ -1000,7 +1022,14 @@ XLogRecordAssemble(RmgrId rmid, uint8 info,
 	 * once we know where in the WAL the record will be inserted. The CRC does
 	 * not include the record header yet.
 	 */
-	rechdr->xl_xid = GetCurrentTransactionIdIfAny();
+	if (!TransactionIdIsValid(curinsert_xid))
+		rechdr->xl_xid = GetCurrentTransactionIdIfAny();
+	else
+	{
+		/* The overriding XID should be handled specially. */
+		rechdr->xl_xid = curinsert_xid;
+		info |= XLR_XID_REPLAYED;
+	}
 	rechdr->xl_tot_len = (uint32) total_len;
 	rechdr->xl_info = info;
 	rechdr->xl_rmid = rmid;
diff --git a/src/backend/access/transam/xlogrecovery.c b/src/backend/access/transam/xlogrecovery.c
index c0ae4d3f63f..cae7318506d 100644
--- a/src/backend/access/transam/xlogrecovery.c
+++ b/src/backend/access/transam/xlogrecovery.c
@@ -1949,10 +1949,13 @@ ApplyWalRecord(XLogReaderState *xlogreader, XLogRecord *record, TimeLineID *repl
 	SpinLockRelease(&XLogRecoveryCtl->info_lck);
 
 	/*
-	 * If we are attempting to enter Hot Standby mode, process XIDs we see
+	 * If we are attempting to enter Hot Standby mode, process XIDs we see.
+	 *
+	 * "replayed" changes should not get into the array again.
 	 */
 	if (standbyState >= STANDBY_INITIALIZED &&
-		TransactionIdIsValid(record->xl_xid))
+		TransactionIdIsValid(record->xl_xid) &&
+		(record->xl_info & XLR_XID_REPLAYED) == 0)
 		RecordKnownAssignedTransactionIds(record->xl_xid);
 
 	/*
diff --git a/src/backend/catalog/indexing.c b/src/backend/catalog/indexing.c
index fd7d2ec0e3a..e09e9fcd8ac 100644
--- a/src/backend/catalog/indexing.c
+++ b/src/backend/catalog/indexing.c
@@ -364,5 +364,5 @@ CatalogTupleUpdateWithInfo(Relation heapRel, const ItemPointerData *otid, HeapTu
 void
 CatalogTupleDelete(Relation heapRel, const ItemPointerData *tid)
 {
-	simple_heap_delete(heapRel, tid);
+	simple_heap_delete(heapRel, tid, InvalidTransactionId);
 }
diff --git a/src/backend/commands/matview.c b/src/backend/commands/matview.c
index 9d490da5f81..372cf57e9f4 100644
--- a/src/backend/commands/matview.c
+++ b/src/backend/commands/matview.c
@@ -894,7 +894,8 @@ refresh_by_heap_swap(Oid matviewOid, Oid OIDNewHeap, char relpersistence)
 {
 	finish_heap_swap(matviewOid, OIDNewHeap, false, false, true, true,
 					 true,		/* reindex */
-					 RecentXmin, ReadNextMultiXactId(), relpersistence);
+					 RecentXmin, ReadNextMultiXactId(), false,
+					 relpersistence);
 }
 
 /*
diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 7390fd303f6..32f15683b6e 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -73,6 +73,7 @@
 #include "storage/lmgr.h"
 #include "storage/predicate.h"
 #include "storage/proc.h"
+#include "storage/procarray.h"
 #include "utils/acl.h"
 #include "utils/fmgroids.h"
 #include "utils/guc.h"
@@ -127,6 +128,7 @@ typedef struct ChangeContexBackup
 	int			file_seq_snapshot;
 	int			file_seq_changes;
 	Oid			clustering_index;
+	TransactionId	last_snapshot_xmin;
 } ChangeContextBackup;
 
 /*
@@ -187,7 +189,11 @@ static void copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldInde
 							bool *pSwapToastByContent,
 							TransactionId *pFreezeXid,
 							MultiXactId *pCutoffMulti,
+							double *p_num_tuples,
 							ChangeContext *chgcxt);
+static void copy_table_data_update_stats(Relation OldHeap, Relation NewHeap,
+										 BlockNumber num_pages,
+										 double num_tuples);
 static void update_relation_cutoffs(Oid relid, TransactionId frozenXid,
 									MultiXactId cutoffMulti);
 static List *get_tables_to_repack(RepackCommand cmd, bool usingindex,
@@ -201,14 +207,21 @@ static bool repack_is_permitted_for_relation(RepackCommand cmd,
 static void apply_concurrent_changes(ChangeContext *chgcxt,
 									 BlockNumber range_start,
 									 BlockNumber range_end);
-static void apply_concurrent_insert(RepackDest *dest, TupleTableSlot *slot);
+static void apply_concurrent_insert(RepackDest *dest,
+									TupleTableSlot *spill_tuple,
+									TupleTableSlot *new_tuple,
+									TransactionId xid);
 static void apply_concurrent_update(RepackDest *dest,
 									TupleTableSlot *spilled_tuple,
-									TupleTableSlot *ondisk_tuple);
-static void apply_concurrent_delete(Relation rel, TupleTableSlot *slot);
+									TupleTableSlot *ondisk_tuple,
+									TupleTableSlot *new_tuple,
+									TransactionId xid);
+static void apply_concurrent_delete(Relation rel, TupleTableSlot *slot,
+									TransactionId xid);
 static void restore_tuple(BufFile *file, Relation relation,
 						  TupleTableSlot *slot, BlockNumber *block_nr_p,
-						  BlockNumber *old_block_nr_p);
+						  BlockNumber *old_block_nr_p,
+						  TransactionId *xid_p);
 static void adjust_toast_pointers(Relation relation, TupleTableSlot *dest,
 								  TupleTableSlot *src);
 static bool is_block_in_range(BlockNumber blknum, BlockNumber start,
@@ -238,11 +251,14 @@ static void rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHea
 											   Oid identIdx,
 											   TransactionId frozenXid,
 											   MultiXactId cutoffMulti,
+											   double num_tuples,
 											   ChangeContext *chgcxt);
 static ChangeContext *process_auxiliary_table(ChangeContext *chgcxt,
 											  Relation *pOldHeap,
 											  Relation *pNewHeap,
-											  Oid identIdx);
+											  Oid identIdx,
+											  TransactionId freeze_xid,
+											  MultiXactId cutoff_multi);
 static List *build_new_indexes(List *OldIndexes, Relation *p_old,
 							   Relation *p_new,
 							   ChangeContext **p_chgcxt);
@@ -1073,6 +1089,7 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	TransactionId frozenXid;
 	MultiXactId cutoffMulti;
 	bool		concurrent = OidIsValid(ident_idx);
+	double		num_tuples = 0;
 	IndexBuildSecurity ibsec;
 	ChangeContext *chgcxt = NULL;
 #if USE_ASSERT_CHECKING
@@ -1162,17 +1179,10 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 	/* Copy the heap data into the new table in the desired order */
 	copy_table_data(NewHeap, OldHeap, index, verbose,
 					&swap_toast_by_content, &frozenXid, &cutoffMulti,
-					chgcxt);
+					&num_tuples, chgcxt);
 
-	/* The historic snapshot won't be needed anymore. */
 	if (concurrent)
 	{
-		/*
-		 * Make sure the active snapshot can see the data copied, so the rows
-		 * can be updated / deleted.
-		 */
-		UpdateActiveSnapshotCommandId();
-
 		Assert(!swap_toast_by_content);
 
 		/*
@@ -1183,7 +1193,8 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 			index_close(index, NoLock);
 
 		rebuild_relation_finish_concurrent(NewHeap, OldHeap, ident_idx,
-										   frozenXid, cutoffMulti, chgcxt);
+										   frozenXid, cutoffMulti, num_tuples,
+										   chgcxt);
 
 		pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
 									 PROGRESS_REPACK_PHASE_FINAL_CLEANUP);
@@ -1213,11 +1224,15 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose,
 		/*
 		 * Swap the physical files of the target and transient tables, then
 		 * rebuild the target's indexes and throw away the transient table.
+		 *
+		 * If swap_toast_by_content is false, we don't need to update the
+		 * cutoffs because the TOAST relation is new.
 		 */
 		finish_heap_swap(tableOid, OIDNewHeap, is_system_catalog,
 						 swap_toast_by_content, false, true,
 						 true,	/* reindex */
 						 frozenXid, cutoffMulti,
+						 swap_toast_by_content, /* update_toast_cutoffs */
 						 relpersistence);
 	}
 
@@ -1682,67 +1697,6 @@ make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
 	return OIDNewHeap;
 }
 
-/*
- * Insert tuple when processing REPACK CONCURRENTLY.
- *
- * rewriteheap.c is not used in the CONCURRENTLY case because it'd be
- * difficult to do the same in the catch-up phase (as the logical decoding
- * does not provide us with sufficient visibility information). Thus we must
- * use heap_insert() both during the catch-up and here.
- *
- * 'reform' is a slot to use for tuple "reforming", typically to get set
- * values of dropped columns to NULL.
- *
- * We pass the NO_LOGICAL flag to heap_insert() in order to skip logical
- * decoding: as soon as REPACK CONCURRENTLY swaps the relation files, it drops
- * this relation, so no logical replication subscription should need the data.
- */
-void
-heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
-					   TupleTableSlot *reform)
-{
-	HeapTuple	tuple;
-	bool		shouldFree;
-	TupleTableSlot *slot;
-	RepackDest *dest;
-
-	/*
-	 * Use the current auxiliary table as output if one is active, otherwise
-	 * insert the tuple into the actual destination table.
-	 */
-	if (chgcxt->cc_dest_aux)
-		dest = chgcxt->cc_dest_aux;
-	else
-		dest = &chgcxt->cc_dest;
-
-	tuple = ExecFetchSlotHeapTuple(src, false, &shouldFree);
-	if (reform != NULL && tuple_needs_reform(tuple, src->tts_tupleDescriptor))
-	{
-		clear_dropped_attributes(tuple, reform);
-		slot = reform;
-	}
-	else
-		slot = src;
-
-	/*
-	 * clear_dropped_attributes() should have deformed the tuple, so nothing
-	 * should depend on it now.
-	 */
-	if (shouldFree)
-		heap_freetuple(tuple);
-
-	table_tuple_insert(dest->rel, slot, GetCurrentCommandId(true),
-					   TABLE_INSERT_NO_LOGICAL, dest->bistate);
-
-	/*
-	 * Insert the tuple into the identity index. initialize_change_context()
-	 * may skip opening of indexes if the identity index is not needed
-	 * immediately.
-	 */
-	if (dest->rri)
-		ExecInsertIndexTuples(dest->rri, dest->estate, 0, slot, NIL, NULL);
-}
-
 bool
 tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc)
 {
@@ -1776,14 +1730,22 @@ clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
 {
 	TupleDesc	tupDesc = reform->tts_tupleDescriptor;
 
-	/* Assuming 'reform' is virtual, this deforms the tuple. */
-	Assert(TTS_IS_VIRTUAL(reform));
+	Assert(TTS_IS_VIRTUAL(reform) || TTS_IS_HEAPTUPLE(reform));
 	ExecForceStoreHeapTuple(tuple, reform, false);
 
 	for (int i = 0; i < tupDesc->natts; i++)
 	{
 		if (TupleDescCompactAttr(tupDesc, i)->attisdropped)
+		{
+			/*
+			 * If 'reform' is virtual, all the attributes are already
+			 * deformed. XXX Should we use the virtual slot at all? .
+			 */
+			if (TTS_IS_HEAPTUPLE(reform))
+				slot_getsomeattrs(reform, i + 1);
+
 			reform->tts_isnull[i] = true;
+		}
 	}
 }
 
@@ -1794,16 +1756,14 @@ clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform)
  * *pSwapToastByContent is set true if toast tables must be swapped by content.
  * *pFreezeXid receives the TransactionId used as freeze cutoff point.
  * *pCutoffMulti receives the MultiXactId used as a cutoff point.
+ * *p_num_tuples receives the number of tuples copied.
  */
 static void
 copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 				bool verbose, bool *pSwapToastByContent,
 				TransactionId *pFreezeXid, MultiXactId *pCutoffMulti,
-				ChangeContext *chgcxt)
+				double *p_num_tuples, ChangeContext *chgcxt)
 {
-	Relation	relRelation;
-	HeapTuple	reltup;
-	Form_pg_class relform;
 	TupleDesc	oldTupDesc PG_USED_FOR_ASSERTS_ONLY;
 	TupleDesc	newTupDesc PG_USED_FOR_ASSERTS_ONLY;
 	VacuumParams params;
@@ -1812,7 +1772,6 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 	double		num_tuples = 0,
 				tups_vacuumed = 0,
 				tups_recently_dead = 0;
-	BlockNumber num_pages;
 	int			elevel = verbose ? INFO : DEBUG2;
 	PGRUsage	ru0;
 	char	   *nspname;
@@ -1985,8 +1944,6 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 	 */
 	NewHeap->rd_toastoid = InvalidOid;
 
-	num_pages = RelationGetNumberOfBlocks(NewHeap);
-
 	/* Log what we did */
 	ereport(elevel,
 			(errmsg("\"%s.%s\": found %.0f removable, %.0f nonremovable row versions in %u pages",
@@ -1999,6 +1956,35 @@ copy_table_data(Relation NewHeap, Relation OldHeap, Relation OldIndex,
 					   tups_recently_dead,
 					   pg_rusage_show(&ru0))));
 
+	/*
+	 * Update pg_class fields. In the CONCURRENTLY case we do it later because
+	 * 1) the catalog update triggers XID assignment, 2) the work is split
+	 * into several transactions, so the catalog update should take place in
+	 * the last one.
+	 */
+	if (!concurrent)
+	{
+		BlockNumber num_pages;
+
+		num_pages = RelationGetNumberOfBlocks(NewHeap);
+
+		copy_table_data_update_stats(OldHeap, NewHeap, num_pages, num_tuples);
+	}
+
+	*p_num_tuples = num_tuples;
+}
+
+/*
+ * Sub-routine of copy_table_data(), to update pg_class.
+ */
+static void
+copy_table_data_update_stats(Relation OldHeap, Relation NewHeap,
+							 BlockNumber num_pages, double num_tuples)
+{
+	Relation	relRelation;
+	HeapTuple	reltup;
+	Form_pg_class relform;
+
 	/* Update pg_class to reflect the correct values of pages and tuples. */
 	relRelation = table_open(RelationRelationId, RowExclusiveLock);
 
@@ -2446,6 +2432,9 @@ update_relation_cutoffs(Oid relid, TransactionId frozenXid,
 /*
  * Remove the transient table that was built by make_new_heap, and finish
  * cleaning up (including rebuilding all indexes on the old heap).
+ *
+ * 'update_toast_cutoffs' tells whether relfrozenxid and relminmxid of the
+ * TOAST relation should be updated too.
  */
 void
 finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
@@ -2456,6 +2445,7 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 				 bool reindex,
 				 TransactionId frozenXid,
 				 MultiXactId cutoffMulti,
+				 bool update_toast_cutoffs,
 				 char newrelpersistence)
 {
 	ObjectAddress object;
@@ -2463,6 +2453,17 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 	Oid			oid_old_toastid;
 	int			i;
 
+	/*
+	 * In the swap-toast-by-content case, we always need to update the
+	 * cutoffs. In the swap-toast-links case, we usually assume we don't need
+	 * to change the toast table's relfrozenxid: the new version of the toast
+	 * table should already have relfrozenxid set to RecentXmin, which is good
+	 * enough. However, there's a special case - REPACK (CONCURRENTLY) - which
+	 * still needs to update the cutoffs - see the related call for more info.
+	 */
+	Assert((swap_toast_by_content && update_toast_cutoffs) ||
+		   !swap_toast_by_content);
+
 	/* Report that we are now swapping relation files */
 	pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
 								 PROGRESS_REPACK_PHASE_SWAP_REL_FILES);
@@ -2540,12 +2541,8 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 	 */
 	CommandCounterIncrement();
 	update_relation_cutoffs(OIDOldHeap, frozenXid, cutoffMulti);
-
-	/*
-	 * The same for TOAST, if needed. In the swap-toast-links case, the new
-	 * the toast table should already have relfrozenxid set to RecentXmin.
-	 */
-	if (OidIsValid(oid_old_toastid) && swap_toast_by_content)
+	/* The same for TOAST, if requested. */
+	if (OidIsValid(oid_old_toastid) && update_toast_cutoffs)
 		update_relation_cutoffs(oid_old_toastid, frozenXid, cutoffMulti);
 
 	/* Destroy new heap with old filenumber */
@@ -3041,12 +3038,14 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 	TupleTableSlot *spilled_tuple;
 	TupleTableSlot *old_update_tuple;
 	TupleTableSlot *ondisk_tuple;
+	TupleTableSlot *new_tuple;
 	bool		have_old_tuple = false;
 	bool		check_range;
 	MemoryContext oldcxt;
 	DecodingWorkerShared *shared;
 	char		fname[MAXPGPATH];
 	BufFile    *file;
+	SnapshotData SnapshotNewHeap;
 
 	/*
 	 * Use the auxiliary table if one exists, otherwise the "final"
@@ -3075,16 +3074,25 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 											table_slot_callbacks(rel));
 	old_update_tuple = MakeSingleTupleTableSlot(RelationGetDescr(rel),
 												&TTSOpsVirtual);
+	new_tuple = MakeSingleTupleTableSlot(RelationGetDescr(rel),
+										 &TTSOpsHeapTuple);
 
 	oldcxt = MemoryContextSwitchTo(GetPerTupleMemoryContext(dest->estate));
 
+	/*
+	 * Finding tuples to UPDATE / DELETE is exactly the purpose of
+	 * SNAPSHOT_NEW_HEAP.
+	 */
+	InitNewHeapSnapshot(SnapshotNewHeap);
+	PushActiveSnapshot(&SnapshotNewHeap);
+
 	while (true)
 	{
 		size_t		nread;
-		ConcurrentChangeKind prevkind = kind;
 		BlockNumber block,
 					old_block;
 		BlockNumber *old_block_p;
+		TransactionId xid;
 
 		CHECK_FOR_INTERRUPTS();
 
@@ -3099,35 +3107,18 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 		 */
 		if (kind == CHANGE_UPDATE_OLD)
 		{
-			restore_tuple(file, rel, old_update_tuple, NULL, NULL);
+			restore_tuple(file, rel, old_update_tuple, NULL, NULL, NULL);
 			have_old_tuple = true;
 			continue;
 		}
 
-		/*
-		 * Just before an UPDATE or DELETE, we must update the command
-		 * counter, because the change could refer to a tuple that we have
-		 * just inserted; and before an INSERT, we have to do this also if the
-		 * previous command was either update or delete.
-		 *
-		 * With this approach we don't spend so many CCIs for long strings of
-		 * only INSERTs, which can't affect one another.
-		 */
-		if (kind == CHANGE_UPDATE_NEW || kind == CHANGE_DELETE ||
-			(kind == CHANGE_INSERT && (prevkind == CHANGE_UPDATE_NEW ||
-									   prevkind == CHANGE_DELETE)))
-		{
-			CommandCounterIncrement();
-			UpdateActiveSnapshotCommandId();
-		}
-
 		/*
 		 * Now restore the tuple into the slot and execute the change.
 		 *
 		 * old_block is only stored with UPDATE_NEW.
 		 */
 		old_block_p = kind == CHANGE_UPDATE_NEW ? &old_block : NULL;
-		restore_tuple(file, rel, spilled_tuple, &block, old_block_p);
+		restore_tuple(file, rel, spilled_tuple, &block, old_block_p, &xid);
 
 		if (kind == CHANGE_INSERT)
 		{
@@ -3137,7 +3128,7 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 			 */
 			if (!check_range ||
 				is_block_in_range(block, range_start, range_end))
-				apply_concurrent_insert(dest, spilled_tuple);
+				apply_concurrent_insert(dest, spilled_tuple, new_tuple, xid);
 		}
 		else if (kind == CHANGE_DELETE)
 		{
@@ -3154,7 +3145,7 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 				found = find_target_tuple(dest, spilled_tuple, ondisk_tuple);
 				if (!found)
 					elog(ERROR, "could not find target tuple");
-				apply_concurrent_delete(rel, ondisk_tuple);
+				apply_concurrent_delete(rel, ondisk_tuple, xid);
 			}
 		}
 		else if (kind == CHANGE_UPDATE_NEW)
@@ -3189,7 +3180,8 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 				 */
 				adjust_toast_pointers(rel, spilled_tuple, ondisk_tuple);
 
-				apply_concurrent_update(dest, spilled_tuple, ondisk_tuple);
+				apply_concurrent_update(dest, spilled_tuple, ondisk_tuple,
+										new_tuple, xid);
 			}
 			else
 			{
@@ -3210,7 +3202,8 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 					 */
 					adjust_toast_pointers(rel, spilled_tuple, NULL);
 
-					apply_concurrent_insert(dest, spilled_tuple);
+					apply_concurrent_insert(dest, spilled_tuple, new_tuple,
+											xid);
 				}
 				else if (is_block_in_range(old_block, range_start, range_end))
 				{
@@ -3224,7 +3217,7 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 					 * visible to the snapshot that we'll use to copy the
 					 * other range.
 					 */
-					apply_concurrent_delete(rel, ondisk_tuple);
+					apply_concurrent_delete(rel, ondisk_tuple, xid);
 				}
 
 				/*
@@ -3241,11 +3234,13 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 
 		ResetPerTupleExprContext(dest->estate);
 	}
+	PopActiveSnapshot();
 
 	/* Cleanup. */
 	ExecDropSingleTupleTableSlot(spilled_tuple);
 	ExecDropSingleTupleTableSlot(ondisk_tuple);
 	ExecDropSingleTupleTableSlot(old_update_tuple);
+	ExecDropSingleTupleTableSlot(new_tuple);
 
 	MemoryContextSwitchTo(oldcxt);
 
@@ -3254,44 +3249,85 @@ apply_concurrent_changes(ChangeContext *chgcxt, BlockNumber range_start,
 
 /*
  * Apply an insert from the spill of concurrent changes to the new copy of the
- * table.
+ * table. 'new_tuple' is the source for table AM.
  */
 static void
-apply_concurrent_insert(RepackDest *dest, TupleTableSlot *slot)
+apply_concurrent_insert(RepackDest *dest, TupleTableSlot *spill_tuple,
+						TupleTableSlot *new_tuple, TransactionId xid)
 {
-	/* Put the tuple in the table, but make sure it won't be decoded */
-	table_tuple_insert(dest->rel, slot, GetCurrentCommandId(true),
-					   TABLE_INSERT_NO_LOGICAL, NULL);
+	HeapTuple	tup;
+	bool		shouldFree;
+
+	/* Copy the contents to a slot that preserves the header fields. */
+	Assert(TTS_IS_HEAPTUPLE(new_tuple));
+	ExecCopySlot(new_tuple, spill_tuple);
+
+	/* Get pointer to the contained tuple (not a copy). */
+	tup = ExecFetchSlotHeapTuple(new_tuple, false, &shouldFree);
+	Assert(!shouldFree);
+
+	/* Set the XID. */
+	HeapTupleHeaderSetXmin(tup->t_data, xid);
+
+	/*
+	 * Put the tuple in the table, but make sure it won't be decoded. At the
+	 * same time, request that the XID we set above is used, instead of
+	 * generating a new one.
+	 *
+	 * FirstCommandId is ok in the new table because the transaction that
+	 * inserted the tuple has already committed, and no other transaction
+	 * should ever need the CID.
+	 */
+	table_tuple_insert(dest->rel, new_tuple, FirstCommandId,
+					   TABLE_INSERT_NO_LOGICAL | TABLE_REUSE_XID,
+					   NULL);
 
 	/* Update indexes with this new tuple. */
 	ExecInsertIndexTuples(dest->rri,
 						  dest->estate,
 						  0,
-						  slot,
+						  new_tuple,
 						  NIL, NULL);
 	pgstat_progress_incr_param(PROGRESS_REPACK_HEAP_TUPLES_INSERTED, 1);
 }
 
 /*
  * Apply an update from the spill of concurrent changes to the new copy of the
- * table.
+ * table. 'new_tuple' is the source for table AM.
  */
 static void
 apply_concurrent_update(RepackDest *dest, TupleTableSlot *spilled_tuple,
-						TupleTableSlot *ondisk_tuple)
+						TupleTableSlot *ondisk_tuple,
+						TupleTableSlot *new_tuple, TransactionId xid)
 {
+	HeapTuple	tup;
+	bool		shouldFree;
 	Relation	rel = dest->rel;
 	LockTupleMode lockmode;
 	TM_FailureData tmfd;
 	TU_UpdateIndexes update_indexes;
 	TM_Result	res;
 
+	/* Copy the contents to a slot that preserves the header fields. */
+	Assert(TTS_IS_HEAPTUPLE(new_tuple));
+	ExecCopySlot(new_tuple, spilled_tuple);
+
+	/* Get pointer to the contained tuple (not a copy). */
+	tup = ExecFetchSlotHeapTuple(new_tuple, false, &shouldFree);
+	Assert(!shouldFree);
+
+	/* Set the XID. */
+	HeapTupleHeaderSetXmin(tup->t_data, xid);
+
 	/*
 	 * Carry out the update, skipping logical decoding for it.
+	 *
+	 * See comments in apply_concurrent_insert() to understand why
+	 * FirstCommandId is ok in the new table.
 	 */
-	res = table_tuple_update(rel, &(ondisk_tuple->tts_tid), spilled_tuple,
-							 GetCurrentCommandId(true),
-							 TABLE_UPDATE_NO_LOGICAL,
+	res = table_tuple_update(rel, &(ondisk_tuple->tts_tid), new_tuple,
+							 FirstCommandId,
+							 TABLE_UPDATE_NO_LOGICAL | TABLE_REUSE_XID,
 							 InvalidSnapshot,
 							 InvalidSnapshot,
 							 false,
@@ -3311,7 +3347,7 @@ apply_concurrent_update(RepackDest *dest, TupleTableSlot *spilled_tuple,
 		ExecInsertIndexTuples(dest->rri,
 							  dest->estate,
 							  flags,
-							  spilled_tuple,
+							  new_tuple,
 							  NIL, NULL);
 	}
 
@@ -3319,16 +3355,22 @@ apply_concurrent_update(RepackDest *dest, TupleTableSlot *spilled_tuple,
 }
 
 static void
-apply_concurrent_delete(Relation rel, TupleTableSlot *slot)
+apply_concurrent_delete(Relation rel, TupleTableSlot *slot, TransactionId xid)
 {
 	TM_Result	res;
 	TM_FailureData tmfd;
 
 	/*
 	 * Delete tuple from the new heap, skipping logical decoding for it.
+	 *
+	 * See comments in heap_insert_for_repack() to understand why
+	 * FirstCommandId is ok in the new table.
+	 *
+	 * See comments in apply_concurrent_insert() to understand why
+	 * FirstCommandId is ok in the new table.
 	 */
 	res = table_tuple_delete(rel, &(slot->tts_tid),
-							 GetCurrentCommandId(true),
+							 xid, FirstCommandId,
 							 TABLE_DELETE_NO_LOGICAL,
 							 InvalidSnapshot, InvalidSnapshot,
 							 false,
@@ -3354,7 +3396,8 @@ apply_concurrent_delete(Relation rel, TupleTableSlot *slot)
  */
 static void
 restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot,
-			  BlockNumber *block_nr_p, BlockNumber *old_block_nr_p)
+			  BlockNumber *block_nr_p, BlockNumber *old_block_nr_p,
+			  TransactionId *xid_p)
 {
 	uint32		t_len;
 	HeapTuple	tup;
@@ -3377,6 +3420,8 @@ restore_tuple(BufFile *file, Relation relation, TupleTableSlot *slot,
 	/* Handle TID separate because not all tuple slots care about it. */
 	if (block_nr_p)
 		*block_nr_p = ItemPointerGetBlockNumber(&tup->t_data->t_ctid);
+	if (xid_p)
+		*xid_p = HeapTupleHeaderGetXmin(tup->t_data);
 	if (old_block_nr_p)
 		BufFileReadExact(file, old_block_nr_p, sizeof(BlockNumber));
 
@@ -3601,6 +3646,7 @@ initialize_change_context(ChangeContext *chgcxt, Relation relation,
 
 	chgcxt->cc_dest_aux = NULL;
 	chgcxt->cc_clustering_index = InvalidOid;
+	chgcxt->cc_last_snapshot_xmin = InvalidTransactionId;
 }
 
 /*
@@ -3816,6 +3862,7 @@ backup_change_context(ChangeContext *chgcxt, ChangeContextBackup *backup)
 	backup->file_seq_snapshot = chgcxt->cc_file_seq_snapshot;
 	backup->file_seq_changes = chgcxt->cc_file_seq_changes;
 	backup->clustering_index = chgcxt->cc_clustering_index;
+	backup->last_snapshot_xmin = chgcxt->cc_last_snapshot_xmin;
 }
 
 /*
@@ -3849,6 +3896,7 @@ reinitialize_change_context(ChangeContextBackup *backup)
 	chgcxt->cc_file_seq_snapshot = backup->file_seq_snapshot;
 	chgcxt->cc_file_seq_changes = backup->file_seq_changes;
 	chgcxt->cc_clustering_index = backup->clustering_index;
+	chgcxt->cc_last_snapshot_xmin = backup->last_snapshot_xmin;
 
 	return chgcxt;
 }
@@ -3961,6 +4009,7 @@ static void
 rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 								   Oid identIdx, TransactionId frozenXid,
 								   MultiXactId cutoffMulti,
+								   double num_tuples,
 								   ChangeContext *chgcxt)
 {
 	List	   *ind_oids_new;
@@ -3975,6 +4024,7 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	MemoryContext oldcxt;
 	List	   *indexrels;
 	List	   *inds_tmp = NIL;
+	BlockNumber num_pages;
 
 	Assert(CheckRelationLockedByMe(OldHeap, ShareUpdateExclusiveLock, false));
 	Assert(CheckRelationLockedByMe(NewHeap, AccessExclusiveLock, false));
@@ -3986,7 +4036,8 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	 * cache entries updated.
 	 */
 	if (chgcxt->cc_dest_aux)
-		chgcxt = process_auxiliary_table(chgcxt, &OldHeap, &NewHeap, identIdx);
+		chgcxt = process_auxiliary_table(chgcxt, &OldHeap, &NewHeap, identIdx,
+										 frozenXid, cutoffMulti);
 
 	/*
 	 * Unlike the exclusive case, we build new indexes for the new relation
@@ -4028,6 +4079,43 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	ind_oids_new = lappend_oid(ind_oids_new,
 							   RelationGetRelid(chgcxt->cc_dest.ident_index));
 
+	/*
+	 * Since we haven't copied "recently dead" tuples into the new heap, we
+	 * must not finish the processing until they are considered dead by all
+	 * backends.
+	 *
+	 * In particular, the VACUUM xmin horizon for the table must be at least
+	 * xmin of the last snapshot that we used to copy the data.  That means
+	 * even the least recently deleted tuples we omitted from the copying
+	 * (because we considered them dead) must be considered dead by anyone.
+	 *
+	 * Note: Although some time should have elapsed since the data copying
+	 * stage (at least the time to build the indexes), we might get stuck here
+	 * due to another backend running REPACK because its snapshot does not
+	 * allow the xmin horizon to advance for some time.
+	 *
+	 * TODO Consider this when determining the value of
+	 * repack_pages_per_snapshot (currently GUC, in the future preferably a
+	 * constant). Is this worth an additional phase in progress reporting?
+	 */
+	Assert(TransactionIdIsValid(chgcxt->cc_last_snapshot_xmin));
+	while (true)
+	{
+		TransactionId oldest_xmin;
+
+		oldest_xmin = GetOldestNonRemovableTransactionId(OldHeap);
+		if (TransactionIdFollowsOrEquals(oldest_xmin,
+										 chgcxt->cc_last_snapshot_xmin))
+			break;
+
+		/* Wait before the next check. */
+		(void) WaitLatch(MyLatch,
+						 WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+						 1000L,
+						 WAIT_EVENT_REPACK_MVCC_SAFETY);
+		ResetLatch(MyLatch);
+	}
+
 	/*
 	 * During testing, wait for another backend to perform concurrent data
 	 * changes which we will process below.
@@ -4149,6 +4237,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	/* The new indexes must be visible for deletion. */
 	CommandCounterIncrement();
 
+	/* Update the numbers of pages and tuples in pg_class. */
+	num_pages = RelationGetNumberOfBlocks(NewHeap);
+	copy_table_data_update_stats(OldHeap, NewHeap, num_pages, num_tuples);
+
 	/* Close the old heap but keep lock until transaction commit. */
 	table_close(OldHeap, NoLock);
 	/* Close the new heap. (We didn't have to open its indexes). */
@@ -4161,6 +4253,12 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 	 * Swap the relations and their TOAST relations and TOAST indexes. This
 	 * also drops the new relation and its indexes.
 	 *
+	 * update_toast_cutoffs is true because REPACK (CONCURRENTLY) does not
+	 * freeze tuples decoded from WAL, and because RecentXmin is not affected
+	 * by logical decoding. Thus if we accepted relfrozenxid of the new TOAST
+	 * relation (derived from RecentXmin), it could incorrectly tell that we
+	 * froze more recent XID's than we actually did.
+	 *
 	 * (System catalogs are currently not supported.)
 	 */
 	Assert(!is_system_catalog);
@@ -4171,6 +4269,7 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 					 true,
 					 false,		/* reindex */
 					 frozenXid, cutoffMulti,
+					 true,		/* update_toast_cutoffs */
 					 relpersistence);
 }
 
@@ -4185,7 +4284,8 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
  */
 static ChangeContext *
 process_auxiliary_table(ChangeContext *chgcxt, Relation *pOldHeap,
-						Relation *pNewHeap, Oid identIdx)
+						Relation *pNewHeap, Oid identIdx,
+						TransactionId freeze_xid, MultiXactId cutoff_multi)
 {
 	RepackDest *dest = chgcxt->cc_dest_aux;
 	Oid			ident_idx_new;
@@ -4195,6 +4295,8 @@ process_auxiliary_table(ChangeContext *chgcxt, Relation *pOldHeap,
 	Oid			aux_oid;
 	ObjectAddress object;
 	Relation	rel;
+	SnapshotData SnapshotNewHeap;
+	RewriteState rwstate;
 
 	/*
 	 * First, make sure the clustering index exists.
@@ -4224,37 +4326,81 @@ process_auxiliary_table(ChangeContext *chgcxt, Relation *pOldHeap,
 		clustering_index = dest->ident_index;
 	}
 
+	/* Now do the copying. */
+	slot = table_slot_create(dest->rel, NULL);
+
 	/*
-	 * Now do the copying. Before starting, clear ->cc_dest_aux so that
-	 * insertions go to the final table, rather than the auxiliary one.
+	 * No point in specifying the auxiliary relation as the old one: we
+	 * haven't frozen tuples when inserting them (one freezing is enough, see
+	 * below), so the tuples do not satisfy the relfrozenxid / relminmxid
+	 * limits, and thus the following freezing would fail.
 	 */
-	chgcxt->cc_dest_aux = NULL;
-	slot = table_slot_create(dest->rel, NULL);
+	rwstate = begin_heap_rewrite(NULL, chgcxt->cc_dest.rel,
+	/* oldest_xmin only needed for rewriting */
+								 InvalidTransactionId,
+								 freeze_xid, cutoff_multi,
+								 true);
 
 	/*
-	 * Note: the current active snapshot blocks the progress of xmin
-	 * horizon(s). The next patches in the series should fix this by using a
-	 * new kind of snapshot (which we can use here because there are no
-	 * transaction aborts in the auxiliary table).
+	 * Scan of the auxiliary table can take long time, but the SnapshotNewHeap
+	 * snapshot can be used here (because there should be no aborted
+	 * insertions in the table), so the scan should not affect the xmin
+	 * horizons.
 	 */
+	PopActiveSnapshot();
+	InitNewHeapSnapshot(SnapshotNewHeap);
+	PushActiveSnapshot(&SnapshotNewHeap);
 	scan = index_beginscan(dest->rel, clustering_index, GetActiveSnapshot(),
 						   NULL, 0, 0, SO_NONE);
 	index_rescan(scan, NULL, 0, NULL, 0);
 	for (;;)
 	{
+		HeapTuple	tuple;
+		bool		shouldFree;
+
 		CHECK_FOR_INTERRUPTS();
 
 		if (!index_getnext_slot(scan, ForwardScanDirection, slot))
 			break;
 
+		/* Make sure we have a writable copy of the tuple. */
+		tuple = ExecFetchSlotHeapTuple(slot, true, &shouldFree);
+
 		/*
-		 * Reforming should have been performed during insertions into the
-		 * auxiliary table.
+		 * This kind of slot maintains the tuple header. We don't need to copy
+		 * the contents into a slot of other kind because reforming was
+		 * performed when populating the auxiliary table.
+		 */
+		Assert(TTS_IS_BUFFERTUPLE(slot));
+		Assert(TransactionIdIsValid(HeapTupleHeaderGetXmin(tuple->t_data)));
+
+		/*
+		 * Insert the tuple into the new relation, and freeze it while doing
+		 * so.
+		 *
+		 * Since our copy is already writable, the tuple can be passed for
+		 * both old and new tuple.
 		 */
-		heap_insert_for_repack(chgcxt, slot, NULL);
+		rewrite_heap_tuple_no_chains(rwstate, tuple, tuple, true);
+
+		if (shouldFree)
+			pfree(tuple);
 	}
 	index_endscan(scan);
+	PopActiveSnapshot();
+	InvalidateCatalogSnapshot();
+
+	/*
+	 * We should not be restricting the progress of xmin horizons at the
+	 * moment.
+	 */
+	Assert(!TransactionIdIsValid(MyProc->xmin));
+	Assert(!TransactionIdIsValid(MyProc->xid));
+	Assert(!HaveRegisteredOrActiveSnapshot());
+
+	PushActiveSnapshot(GetTransactionSnapshot());
 	ExecDropSingleTupleTableSlot(slot);
+	end_heap_rewrite(rwstate);
 
 	/*
 	 * Close the relation, its identity index and clustering index if we had
@@ -4266,6 +4412,7 @@ process_auxiliary_table(ChangeContext *chgcxt, Relation *pOldHeap,
 		index_close(clustering_index, NoLock);
 	/* Here we close the other indexes. */
 	release_change_dest(dest);
+	chgcxt->cc_dest_aux = NULL;
 
 	/* Drop the auxiliary table. */
 	object.classId = RelationRelationId;
diff --git a/src/backend/commands/tablecmds.c b/src/backend/commands/tablecmds.c
index 427fe733153..5c684235e22 100644
--- a/src/backend/commands/tablecmds.c
+++ b/src/backend/commands/tablecmds.c
@@ -6122,6 +6122,7 @@ ATRewriteTables(AlterTableStmt *parsetree, List **wqueue, LOCKMODE lockmode,
 							 true,	/* reindex */
 							 RecentXmin,
 							 ReadNextMultiXactId(),
+							 false, /* update_toast_cutoffs */
 							 persistence);
 
 			InvokeObjectPostAlterHook(RelationRelationId, tab->relid, 0);
diff --git a/src/backend/executor/nodeModifyTable.c b/src/backend/executor/nodeModifyTable.c
index b9781eb3b95..265e80b765b 100644
--- a/src/backend/executor/nodeModifyTable.c
+++ b/src/backend/executor/nodeModifyTable.c
@@ -1765,6 +1765,7 @@ ExecDeleteAct(ModifyTableContext *context, ResultRelInfo *resultRelInfo,
 		options |= TABLE_DELETE_CHANGING_PARTITION;
 
 	return table_tuple_delete(resultRelInfo->ri_RelationDesc, tupleid,
+							  InvalidTransactionId,
 							  estate->es_output_cid,
 							  options,
 							  estate->es_snapshot,
diff --git a/src/backend/replication/logical/decode.c b/src/backend/replication/logical/decode.c
index c3722b5c623..e3c94f83875 100644
--- a/src/backend/replication/logical/decode.c
+++ b/src/backend/replication/logical/decode.c
@@ -425,6 +425,18 @@ heap2_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	TransactionId xid = XLogRecGetXid(buf->record);
 	SnapBuild  *builder = ctx->snapshot_builder;
 
+	/*
+	 * XLOG_HEAP2_MULTI_INSERT is not replayed.
+	 */
+	Assert((XLogRecGetInfo(buf->record) & XLR_XID_REPLAYED) == 0);
+
+	/* See heap_decode(). */
+	if (change_useless_for_repack(buf))
+	{
+		Assert(!ctx->fast_forward);
+		return;
+	}
+
 	ReorderBufferProcessXid(ctx->reorder, xid, buf->origptr);
 
 	/*
@@ -442,8 +454,7 @@ heap2_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	{
 		case XLOG_HEAP2_MULTI_INSERT:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeMultiInsert(ctx, buf);
 			break;
 		case XLOG_HEAP2_NEW_CID:
@@ -488,6 +499,50 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	TransactionId xid = XLogRecGetXid(buf->record);
 	SnapBuild  *builder = ctx->snapshot_builder;
 
+	/*
+	 * REPACK decoding should only decode changes of the relation being
+	 * processed. Decoding changes of other tables would not only introduce
+	 * performance overhead, it would also make us process already committed
+	 * (and replayed) XIDs again - that's probably not expected by the logical
+	 * decoding system.
+	 *
+	 * The problem is that during the replay, REPACK can (and does) avoid WAL
+	 * logging the extra information needed for logical decoding, however this
+	 * is checked in DecodeInsert(), DecodeUpdate(), etc., which is too late.
+	 * Moreover, these function do not distinguish whether the transaction is
+	 * being decoded the first time, or if the WAL records originate from the
+	 * replay phase of REPACK.
+	 *
+	 * Unlike the fast-forward case (see comments below), REPACK does not need
+	 * the base snapshot for the transaction until it receives a change that
+	 * really needs to be decoded. Thus it's ok to skip
+	 * SnapBuildProcessChange().
+	 *
+	 * (With fast-forward, we must not omit the setup of the transaction base
+	 * snapshot because the changes skipped by fast-forward initially may need
+	 * to be decoded after restart. Thus the base snapshot may be needed after
+	 * the restart too. If we didn't create the snapshot in the fast-forward
+	 * mode, the snapshot builder's xmin would advance too eagerly, so the
+	 * same snapshot wouldn't work after restart.)
+	 *
+	 * First, filter out WAL records generated by REPACK (CONCURRENTLY)
+	 * replaying the data changes of other transactions - these transactions
+	 * have already been decoded, so no backend / worker should decode them
+	 * again.
+	 */
+	if (XLogRecGetInfo(buf->record) & XLR_XID_REPLAYED)
+		return;
+
+	/*
+	 * Now let REPACK decoding worker filter out changes of tables other than
+	 * the one whose REPACKing it's involved in.
+	 */
+	if (change_useless_for_repack(buf))
+	{
+		Assert(!ctx->fast_forward);
+		return;
+	}
+
 	ReorderBufferProcessXid(ctx->reorder, xid, buf->origptr);
 
 	/*
@@ -505,8 +560,7 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	{
 		case XLOG_HEAP_INSERT:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeInsert(ctx, buf);
 			break;
 
@@ -518,22 +572,19 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		case XLOG_HEAP_HOT_UPDATE:
 		case XLOG_HEAP_UPDATE:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeUpdate(ctx, buf);
 			break;
 
 		case XLOG_HEAP_DELETE:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeDelete(ctx, buf);
 			break;
 
 		case XLOG_HEAP_TRUNCATE:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeTruncate(ctx, buf);
 			break;
 
@@ -549,8 +600,7 @@ heap_decode(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 
 		case XLOG_HEAP_CONFIRM:
 			if (SnapBuildProcessChange(builder, xid, buf->origptr) &&
-				!ctx->fast_forward &&
-				!change_useless_for_repack(buf))
+				!ctx->fast_forward)
 				DecodeSpecConfirm(ctx, buf);
 			break;
 
@@ -960,15 +1010,17 @@ DecodeInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 	DecodeXLogTuple(tupledata, datalen, change->data.tp.newtuple);
 
 	/*
-	 * REPACK (CONCURRENTLY) needs block number to check if the corresponding
-	 * part of the table was already copied.  XXX Should we only do this if
-	 * AmRepackWorker()? It might save a few cycles, but not sure it's good to
-	 * leave the fields unset in other cases.
+	 * REPACK (CONCURRENTLY) needs xmin to preserve visibility information and
+	 * block number to check if the corresponding part of the table was
+	 * already copied.  XXX Should we only do this if AmRepackWorker()? It
+	 * might save a few cycles, but not sure it's good to leave the fields
+	 * unset in other cases.
 	 */
 	{
 		HeapTupleHeader header;
 
 		header = change->data.tp.newtuple->t_data;
+		HeapTupleHeaderSetXmin(header, XLogRecGetXid(r));
 		/* offnum is not really needed, but let's set valid pointer. */
 		ItemPointerSet(&header->t_ctid, blknum, xlrec->offnum);
 	}
@@ -1033,14 +1085,15 @@ DecodeUpdate(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 		DecodeXLogTuple(data, datalen, change->data.tp.newtuple);
 
 		/*
-		 * REPACK (CONCURRENTLY) needs block numbers to check if the
-		 * corresponding part of the table was already copied. XXX Do this
-		 * only if AmRepackWorker()?
+		 * REPACK (CONCURRENTLY) needs xmin to preserve visibility information
+		 * and block numbers to check if the corresponding part of the table
+		 * was already copied. XXX Do this only if AmRepackWorker()?
 		 */
 		{
 			HeapTupleHeader header;
 
 			header = change->data.tp.newtuple->t_data;
+			HeapTupleHeaderSetXmin(header, XLogRecGetXid(r));
 			/* offnum is not really needed, but let's set valid pointer. */
 			ItemPointerSet(&header->t_ctid, new_blknum, xlrec->new_offnum);
 			change->data.tp.old_blknum = old_blknum;
@@ -1129,14 +1182,20 @@ DecodeDelete(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 						datalen, change->data.tp.oldtuple);
 
 		/*
-		 * REPACK (CONCURRENTLY) needs block number to check if the
-		 * corresponding part of the table was already copied. XXX Do this
-		 * only if AmRepackWorker()?
+		 * REPACK (CONCURRENTLY) needs xmax to preserve visibility information
+		 * and block number to check if the corresponding part of the table
+		 * was already copied. XXX Do this only if AmRepackWorker()?
 		 */
 		{
 			HeapTupleHeader header;
 
 			header = change->data.tp.oldtuple->t_data;
+
+			/*
+			 * xmax makes more sense here, but we don't want restore_tuple()
+			 * to pay attention to the change kind, so use xmin here as well.
+			 */
+			HeapTupleHeaderSetXmin(header, XLogRecGetXid(r));
 			/* offnum is not really needed, but let's set valid pointer. */
 			ItemPointerSet(&header->t_ctid, blknum, xlrec->offnum);
 		}
@@ -1282,13 +1341,16 @@ DecodeMultiInsert(LogicalDecodingContext *ctx, XLogRecordBuffer *buf)
 			change->data.tp.clear_toast_afterwards = false;
 
 		/*
-		 * REPACK (CONCURRENTLY) needs block number to check if the
-		 * corresponding part of the table was already copied.
+		 * REPACK (CONCURRENTLY) needs xmin to preserve visibility information
+		 * and block number to check if the corresponding part of the table
+		 * was already copied.
 		 */
 		if (AmRepackWorker())
 		{
 			OffsetNumber offnum;
 
+			HeapTupleHeaderSetXmin(header, XLogRecGetXid(r));
+
 			/*
 			 * offnum is not really needed, but let's set valid pointer. (It
 			 * will be invalid anyway if the page was initially empty.)
diff --git a/src/backend/replication/logical/reorderbuffer.c b/src/backend/replication/logical/reorderbuffer.c
index cae2b099e69..626e9400724 100644
--- a/src/backend/replication/logical/reorderbuffer.c
+++ b/src/backend/replication/logical/reorderbuffer.c
@@ -5263,6 +5263,9 @@ ReorderBufferToastReplace(ReorderBuffer *rb, ReorderBufferTXN *txn,
 	 * Shouldn't we add a new field to ReorderBufferChange instead?
 	 */
 	tmphtup->t_data->t_ctid = newtup->t_data->t_ctid;
+	/* Likewise, preserve XID - REPACK needs it to be MVCC-safe. */
+	HeapTupleHeaderSetXmin(tmphtup->t_data,
+						   HeapTupleHeaderGetXmin(newtup->t_data));
 
 	memcpy(newtup->t_data, tmphtup->t_data, tmphtup->t_len);
 	newtup->t_len = tmphtup->t_len;
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 1016502d042..30fcd952132 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -184,6 +184,7 @@ PG_SLEEP	"Waiting due to a call to <function>pg_sleep</function> or a sibling fu
 RECOVERY_APPLY_DELAY	"Waiting to apply WAL during recovery because of a delay setting."
 RECOVERY_RETRIEVE_RETRY_INTERVAL	"Waiting during recovery when WAL data is not available from any source (<filename>pg_wal</filename>, archive or stream)."
 REGISTER_SYNC_REQUEST	"Waiting while sending synchronization requests to the checkpointer, because the request queue is full."
+REPACK_MVCC_SAFETY	"Waiting until not copied tuples are considered dead."
 SPIN_DELAY	"Waiting while acquiring a contended spinlock."
 VACUUM_DELAY	"Waiting in a cost-based vacuum delay point."
 VACUUM_TRUNCATE	"Waiting to acquire an exclusive lock to truncate off any empty pages at the end of a table vacuumed."
diff --git a/src/include/access/heapam.h b/src/include/access/heapam.h
index 5176478c295..bc6eadcb978 100644
--- a/src/include/access/heapam.h
+++ b/src/include/access/heapam.h
@@ -380,6 +380,7 @@ extern void heap_multi_insert(Relation relation, TupleTableSlot **slots,
 							  int ntuples, CommandId cid, uint32 options,
 							  BulkInsertState bistate);
 extern TM_Result heap_delete(Relation relation, const ItemPointerData *tid,
+							 TransactionId xid,
 							 CommandId cid, uint32 options, Snapshot crosscheck,
 							 bool wait, TM_FailureData *tmfd);
 extern void heap_finish_speculative(Relation relation, const ItemPointerData *tid);
@@ -422,7 +423,8 @@ extern bool heap_tuple_should_freeze(HeapTupleHeader tuple,
 extern bool heap_tuple_needs_eventual_freeze(HeapTupleHeader tuple);
 
 extern void simple_heap_insert(Relation relation, HeapTuple tup);
-extern void simple_heap_delete(Relation relation, const ItemPointerData *tid);
+extern void simple_heap_delete(Relation relation, const ItemPointerData *tid,
+							   TransactionId xid);
 extern void simple_heap_update(Relation relation, const ItemPointerData *otid,
 							   HeapTuple tup, TU_UpdateIndexes *update_indexes);
 
diff --git a/src/include/access/heaptoast.h b/src/include/access/heaptoast.h
index 631cb1836b9..36d1d65c131 100644
--- a/src/include/access/heaptoast.h
+++ b/src/include/access/heaptoast.h
@@ -14,6 +14,7 @@
 #define HEAPTOAST_H
 
 #include "access/htup_details.h"
+#include "access/rewriteheap.h"
 #include "storage/lockdefs.h"
 #include "utils/relcache.h"
 
@@ -95,7 +96,9 @@
  * ----------
  */
 extern HeapTuple heap_toast_insert_or_update(Relation rel, HeapTuple newtup,
-											 HeapTuple oldtup, uint32 options);
+											 HeapTuple oldtup,
+											 RewriteState rwstate,
+											 uint32 options);
 
 /* ----------
  * heap_toast_delete -
@@ -104,7 +107,7 @@ extern HeapTuple heap_toast_insert_or_update(Relation rel, HeapTuple newtup,
  * ----------
  */
 extern void heap_toast_delete(Relation rel, HeapTuple oldtup,
-							  bool is_speculative);
+							  bool is_speculative, TransactionId xid);
 
 /* ----------
  * toast_flatten_tuple -
diff --git a/src/include/access/rewriteheap.h b/src/include/access/rewriteheap.h
index 6ccf7b45c04..80d928e0e26 100644
--- a/src/include/access/rewriteheap.h
+++ b/src/include/access/rewriteheap.h
@@ -22,12 +22,21 @@
 typedef struct RewriteStateData *RewriteState;
 
 extern RewriteState begin_heap_rewrite(Relation old_heap, Relation new_heap,
-									   TransactionId oldest_xmin, TransactionId freeze_xid,
-									   MultiXactId cutoff_multi);
+									   TransactionId oldest_xmin,
+									   TransactionId freeze_xid,
+									   MultiXactId cutoff_multi,
+									   bool no_chains);
 extern void end_heap_rewrite(RewriteState state);
 extern void rewrite_heap_tuple(RewriteState state, HeapTuple old_tuple,
 							   HeapTuple new_tuple);
+extern void rewrite_heap_tuple_no_chains(RewriteState state,
+										 HeapTuple old_tuple,
+										 HeapTuple new_tuple,
+										 bool freeze);
 extern bool rewrite_heap_dead_tuple(RewriteState state, HeapTuple old_tuple);
+extern void rewrite_freeze_tuple(RewriteState state, HeapTuple tuple);
+extern void rewrite_copy_visibility_info(HeapTuple new_tuple,
+										 HeapTuple old_tuple);
 
 /*
  * On-Disk data format for an individual logical rewrite mapping.
diff --git a/src/include/access/tableam.h b/src/include/access/tableam.h
index 132248c5d43..1af01235c8b 100644
--- a/src/include/access/tableam.h
+++ b/src/include/access/tableam.h
@@ -292,6 +292,12 @@ typedef struct TM_IndexDeleteOp
 /* "options" flag bits for table_tuple_update */
 #define TABLE_UPDATE_NO_LOGICAL					(1 << 0)
 
+/*
+ * For INSERT or UPDATE, use XID (xmin) contained the in new tuple rather than
+ * the XID of the current transaction.
+ */
+#define TABLE_REUSE_XID							(1 << 31)
+
 /* flag bits for table_tuple_lock */
 /* Follow tuples whose update is in progress if lock modes don't conflict  */
 #define TUPLE_LOCK_FLAG_LOCK_UPDATE_IN_PROGRESS	(1 << 0)
@@ -568,6 +574,7 @@ typedef struct TableAmRoutine
 	/* see table_tuple_delete() for reference about parameters */
 	TM_Result	(*tuple_delete) (Relation rel,
 								 ItemPointer tid,
+								 TransactionId xid,
 								 CommandId cid,
 								 uint32 options,
 								 Snapshot snapshot,
@@ -1458,6 +1465,14 @@ static inline void
 table_tuple_insert(Relation rel, TupleTableSlot *slot, CommandId cid,
 				   uint32 options, BulkInsertStateData *bistate)
 {
+	/*
+	 * TABLE_REUSE_XID restricts the slot type because not all slots preserve
+	 * the visibility information. XXX Isn't this a reason to pass the xid as
+	 * an argument?
+	 */
+	Assert((options & TABLE_REUSE_XID) == 0 || TTS_IS_HEAPTUPLE(slot) ||
+		   TTS_IS_BUFFERTUPLE(slot));
+
 	rel->rd_tableam->tuple_insert(rel, slot, cid, options,
 								  bistate);
 }
@@ -1479,6 +1494,9 @@ table_tuple_insert_speculative(Relation rel, TupleTableSlot *slot,
 							   BulkInsertStateData *bistate,
 							   uint32 specToken)
 {
+	/* TABLE_REUSE_XID is currently not needed here. */
+	Assert((options & TABLE_REUSE_XID) == 0);
+
 	rel->rd_tableam->tuple_insert_speculative(rel, slot, cid, options,
 											  bistate, specToken);
 }
@@ -1526,6 +1544,7 @@ table_multi_insert(Relation rel, TupleTableSlot **slots, int nslots,
  * Input parameters:
  *	rel - table to be modified (caller must hold suitable lock)
  *	tid - TID of tuple to be deleted
+ *	xid - XID to use or InvalidTransactionId for the current transaction
  *	cid - delete command ID (used for visibility test, and stored into
  *		cmax if successful)
  *	options - bitmask of options.  Supported values:
@@ -1546,11 +1565,12 @@ table_multi_insert(Relation rel, TupleTableSlot **slots, int nslots,
  * TM_FailureData for additional info.
  */
 static inline TM_Result
-table_tuple_delete(Relation rel, ItemPointer tid, CommandId cid,
+table_tuple_delete(Relation rel, ItemPointer tid, TransactionId xid,
+				   CommandId cid,
 				   uint32 options, Snapshot snapshot, Snapshot crosscheck,
 				   bool wait, TM_FailureData *tmfd)
 {
-	return rel->rd_tableam->tuple_delete(rel, tid, cid, options,
+	return rel->rd_tableam->tuple_delete(rel, tid, xid, cid, options,
 										 snapshot, crosscheck,
 										 wait, tmfd);
 }
@@ -1601,6 +1621,10 @@ table_tuple_update(Relation rel, ItemPointer otid, TupleTableSlot *slot,
 				   bool wait, TM_FailureData *tmfd, LockTupleMode *lockmode,
 				   TU_UpdateIndexes *update_indexes)
 {
+	/* See table_tuple_insert(). */
+	Assert((options & TABLE_REUSE_XID) == 0 || TTS_IS_HEAPTUPLE(slot) ||
+		   TTS_IS_BUFFERTUPLE(slot));
+
 	return rel->rd_tableam->tuple_update(rel, otid, slot,
 										 cid, options, snapshot, crosscheck,
 										 wait, tmfd,
diff --git a/src/include/access/toast_helper.h b/src/include/access/toast_helper.h
index 2ec92397f26..f86f88c774e 100644
--- a/src/include/access/toast_helper.h
+++ b/src/include/access/toast_helper.h
@@ -14,6 +14,7 @@
 #ifndef TOAST_HELPER_H
 #define TOAST_HELPER_H
 
+#include "access/rewriteheap.h"
 #include "utils/rel.h"
 
 /*
@@ -60,6 +61,15 @@ typedef struct
 	 */
 	uint8		ttc_flags;
 	ToastAttrInfo *ttc_attr;
+
+	/*
+	 * These fields are needed when the option HEAP_INSERT_REUSE_XID_FREEZE
+	 * was passed to heap_toast_insert_or_update(). We could actually use
+	 * normal insert, but that would require WAL support of
+	 * heap_freeze_tuple().
+	 */
+	RewriteState ttc_rwstate;
+	HeapTuple	ttc_tup_main;
 } ToastTupleContext;
 
 /*
@@ -108,9 +118,9 @@ extern int	toast_tuple_find_biggest_attribute(ToastTupleContext *ttc,
 extern void toast_tuple_try_compression(ToastTupleContext *ttc, int attribute);
 extern void toast_tuple_externalize(ToastTupleContext *ttc, int attribute,
 									uint32 options);
-extern void toast_tuple_cleanup(ToastTupleContext *ttc);
+extern void toast_tuple_cleanup(ToastTupleContext *ttc, TransactionId xid);
 
 extern void toast_delete_external(Relation rel, const Datum *values, const bool *isnull,
-								  bool is_speculative);
+								  bool is_speculative, TransactionId xid);
 
 #endif
diff --git a/src/include/access/toast_internals.h b/src/include/access/toast_internals.h
index bf45889a642..62da4a18d63 100644
--- a/src/include/access/toast_internals.h
+++ b/src/include/access/toast_internals.h
@@ -12,6 +12,7 @@
 #ifndef TOAST_INTERNALS_H
 #define TOAST_INTERNALS_H
 
+#include "access/rewriteheap.h"
 #include "access/toast_compression.h"
 #include "storage/lockdefs.h"
 #include "utils/relcache.h"
@@ -48,9 +49,13 @@ typedef struct toast_compress_header
 extern Datum toast_compress_datum(Datum value, char cmethod);
 extern Oid	toast_get_valid_index(Oid toastoid, LOCKMODE lock);
 
-extern void toast_delete_datum(Relation rel, Datum value, bool is_speculative);
+extern void toast_delete_datum(Relation rel, Datum value, bool is_speculative,
+							   TransactionId xid);
 extern Datum toast_save_datum(Relation rel, Datum value,
-							  varlena *oldexternal, uint32 options);
+							  varlena *oldexternal,
+							  RewriteState rwstate,
+							  HeapTuple tup_main,
+							  uint32 options);
 
 extern int	toast_open_indexes(Relation toastrel,
 							   LOCKMODE lock,
diff --git a/src/include/access/xlog_internal.h b/src/include/access/xlog_internal.h
index 55663e6f4af..be718993401 100644
--- a/src/include/access/xlog_internal.h
+++ b/src/include/access/xlog_internal.h
@@ -32,7 +32,7 @@
 /*
  * Each page of XLOG file has a header like this:
  */
-#define XLOG_PAGE_MAGIC 0xD120	/* can be used as WAL version indicator */
+#define XLOG_PAGE_MAGIC 0xD121	/* can be used as WAL version indicator */
 
 typedef struct XLogPageHeaderData
 {
diff --git a/src/include/access/xloginsert.h b/src/include/access/xloginsert.h
index 91dfbd5627f..a0cea5fde01 100644
--- a/src/include/access/xloginsert.h
+++ b/src/include/access/xloginsert.h
@@ -43,6 +43,7 @@
 /* prototypes for public functions in xloginsert.c: */
 extern void XLogBeginInsert(void);
 extern void XLogSetRecordFlags(uint8 flags);
+extern void XLogSetRecordXid(TransactionId xid);
 extern XLogRecPtr XLogInsert(RmgrId rmid, uint8 info);
 extern XLogRecPtr XLogSimpleInsertInt64(RmgrId rmid, uint8 info, int64 value);
 extern void XLogEnsureRecordSpace(int max_block_id, int ndatas);
diff --git a/src/include/access/xlogrecord.h b/src/include/access/xlogrecord.h
index e8999d3fe91..bf3c1965e05 100644
--- a/src/include/access/xlogrecord.h
+++ b/src/include/access/xlogrecord.h
@@ -90,6 +90,14 @@ typedef struct XLogRecord
  */
 #define XLR_CHECK_CONSISTENCY	0x02
 
+/*
+ * The record contains a data change that was already committed and now is
+ * being applied to a new relation due to rewriting. The original XID is
+ * needed to keep the rewriting MVCC-safe, however the transaction should be
+ * ignored by logical decoding and it should not get into KnownAssignedXids.
+ */
+#define XLR_XID_REPLAYED		0x04
+
 /*
  * Header info for block data appended to an XLOG record.
  *
diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h
index ad8125790b0..9e2f3e491c6 100644
--- a/src/include/commands/repack.h
+++ b/src/include/commands/repack.h
@@ -16,6 +16,7 @@
 #include <signal.h>
 
 #include "access/hio.h"
+#include "access/rewriteheap.h"
 #include "access/skey.h"
 #include "access/xlogdefs.h"
 #include "catalog/index.h"
@@ -110,10 +111,10 @@ typedef struct ChangeContext
 	 * Not sure, it'd require disk space for one more copy and the copying
 	 * itself is not free.
 	 *
-	 * TODO 1) make the tables unlogged, 2) if REPACK locks the TOAST relation
-	 * too (not sure it does) try to preserve TOAST pointers, instead of
-	 * storing them to TOAST relations of these tables, 3) Check that the
-	 * tables are dropped on transaction abort.
+	 * XXX REPACK currently does not lock the old TOAST relation. If it did,
+	 * we could perhaps copy TOAST pointers from the old relation to the
+	 * auxiliary relation, so that the auxiliary relation would not need its
+	 * own TOAST relation.
 	 */
 	RepackDest *cc_dest_aux;
 
@@ -127,6 +128,11 @@ typedef struct ChangeContext
 	 * functions. This is needed when starting a new transaction.
 	 */
 	IndexBuildSecurity cc_ind_build_sec;
+
+	/*
+	 * xmin of the last snapshot used to copy data.
+	 */
+	TransactionId cc_last_snapshot_xmin;
 } ChangeContext;
 
 extern PGDLLIMPORT int repack_pages_per_snapshot;
@@ -142,8 +148,6 @@ extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal);
 extern Oid	make_new_heap(Oid OIDOldHeap, Oid NewTableSpace, Oid NewAccessMethod,
 						  char relpersistence, LOCKMODE lockmode,
 						  bool auxiliary);
-extern void heap_insert_for_repack(ChangeContext *chgcxt, TupleTableSlot *src,
-								   TupleTableSlot *reform);
 extern bool tuple_needs_reform(HeapTuple tuple, TupleDesc tupDesc);
 extern void clear_dropped_attributes(HeapTuple tuple, TupleTableSlot *reform);
 extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
@@ -154,6 +158,7 @@ extern void finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
 							 bool reindex,
 							 TransactionId frozenXid,
 							 MultiXactId cutoffMulti,
+							 bool update_toast_cutoffs,
 							 char newrelpersistence);
 extern Snapshot repack_get_snapshot(ChangeContext *chgcxt);
 extern void repack_process_concurrent_changes(ChangeContext *chgcxt,
diff --git a/src/include/utils/snapmgr.h b/src/include/utils/snapmgr.h
index 1c550096393..b0c40df7934 100644
--- a/src/include/utils/snapmgr.h
+++ b/src/include/utils/snapmgr.h
@@ -51,6 +51,18 @@ extern PGDLLIMPORT SnapshotData SnapshotToastData;
 	((snapshotdata).snapshot_type = SNAPSHOT_NON_VACUUMABLE, \
 	 (snapshotdata).vistest = (vistestp))
 
+/*
+ * NewHeap snapshot needs to be used as the active snapshot at some point, so
+ * initialize the fields related to PushActiveSnapshot().
+ */
+#define InitNewHeapSnapshot(snapshotdata)  \
+	((snapshotdata).snapshot_type = SNAPSHOT_NEW_HEAP, \
+	 (snapshotdata).regd_count = 0, \
+	 (snapshotdata).active_count = 0, \
+	 (snapshotdata).copied = false, \
+	 (snapshotdata).xcnt = 0, \
+	 (snapshotdata).subxcnt = 0)
+
 /*
  * Is the snapshot implemented as an MVCC snapshot (i.e. it uses
  * SNAPSHOT_MVCC)? If so, there will be at most one visible tuple in a chain
diff --git a/src/include/utils/snapshot.h b/src/include/utils/snapshot.h
index 9766aabcad4..f91cb43a5a2 100644
--- a/src/include/utils/snapshot.h
+++ b/src/include/utils/snapshot.h
@@ -112,6 +112,29 @@ typedef enum SnapshotType
 	 * horizon to use.
 	 */
 	SNAPSHOT_NON_VACUUMABLE,
+
+	/*
+	 * The effects of all transactions are visible. Unlike SNAPSHOT_DIRTY,
+	 * aborted (sub)transactions are not expected.
+	 *
+	 * This is specific to applying data changes to the new heap by the REPACK
+	 * command - that replays applies changes done in the old heap by
+	 * transactions that have already committed. No other transactions can
+	 * access the new heap while this snapshot is in use.
+	 *
+	 * Therefore, whenever a transaction being applied looks for a tuple to
+	 * update or delete, it can assume that the insertion of any candidate
+	 * tuple was already committed - otherwise the inserting transaction
+	 * wouldn't have been applied.
+	 *
+	 * By considering effects of all transactions visible we also ensure that
+	 * a transaction can update / delete tuples that it inserted itself.
+	 *
+	 * TODO Consider better name. Would SNAPSHOT_BOOTSTRAP be confusing?
+	 * During cluster bootstrap we also consider all changes committed
+	 * immediately.
+	 */
+	SNAPSHOT_NEW_HEAP,
 } SnapshotType;
 
 typedef struct SnapshotData *Snapshot;
@@ -127,8 +150,8 @@ typedef struct SnapshotData *Snapshot;
  * * Historic MVCC snapshots used during logical decoding
  * * snapshots passed to HeapTupleSatisfiesDirty()
  * * snapshots passed to HeapTupleSatisfiesNonVacuumable()
- * * snapshots used for SatisfiesAny, Toast, Self where no members are
- *	 accessed.
+ * * snapshots used for SatisfiesAny, Toast, Self, NewHeap where no members
+ *	 are accessed.
  *
  * TODO: It's probably a good idea to split this struct using a NodeTag
  * similar to how parser and executor nodes are handled, with one type for
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..9919c93fe7f 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -45,8 +45,8 @@ step check2:
 
 	SELECT i, j FROM repack_test ORDER BY i, j;
 
-	INSERT INTO data_s2(i, j)
-	SELECT i, j FROM repack_test;
+	INSERT INTO data_s2(i, j, _xmin)
+	SELECT i, j, xmin FROM repack_test;
 
   i|  j
 ---+---
@@ -77,11 +77,11 @@ step check1:
 
 	SELECT i, j FROM repack_test ORDER BY i, j;
 
-	INSERT INTO data_s1(i, j)
-	SELECT i, j FROM repack_test;
+	INSERT INTO data_s1(i, j, _xmin)
+	SELECT i, j, xmin FROM repack_test;
 
 	SELECT count(*)
-	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j, _xmin)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 
 count
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index 7896d1456ad..3e15f5db31f 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -9,8 +9,8 @@ setup
 
 	CREATE TABLE relfilenodes(node oid);
 
-	CREATE TABLE data_s1(i int, j int);
-	CREATE TABLE data_s2(i int, j int);
+	CREATE TABLE data_s1(i int, j int, _xmin xid);
+	CREATE TABLE data_s2(i int, j int, _xmin xid);
 }
 
 teardown
@@ -39,7 +39,8 @@ step wait_before_lock
 # Besides the contents, we also check that relfilenode has changed.
 
 # Have each session write the contents into a table and use FULL JOIN to check
-# if the outputs are identical.
+# if the outputs are identical. xmin is included in order to check the MVCC
+# safety.
 step check1
 {
 	INSERT INTO relfilenodes(node)
@@ -49,11 +50,11 @@ step check1
 
 	SELECT i, j FROM repack_test ORDER BY i, j;
 
-	INSERT INTO data_s1(i, j)
-	SELECT i, j FROM repack_test;
+	INSERT INTO data_s1(i, j, _xmin)
+	SELECT i, j, xmin FROM repack_test;
 
 	SELECT count(*)
-	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
+	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j, _xmin)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
 teardown
@@ -85,10 +86,6 @@ step change_new
 
 # When applying concurrent data changes, we should see the effects of an
 # in-progress subtransaction.
-#
-# XXX Not sure this test is useful now - it was designed for the patch that
-# preserves tuple visibility and which therefore modifies
-# TransactionIdIsCurrentTransactionId().
 step change_subxact1
 {
 	BEGIN;
@@ -102,8 +99,6 @@ step change_subxact1
 
 # When applying concurrent data changes, we should not see the effects of a
 # rolled back subtransaction.
-#
-# XXX Is this test useful? See above.
 step change_subxact2
 {
 	BEGIN;
@@ -122,8 +117,8 @@ step check2
 
 	SELECT i, j FROM repack_test ORDER BY i, j;
 
-	INSERT INTO data_s2(i, j)
-	SELECT i, j FROM repack_test;
+	INSERT INTO data_s2(i, j, _xmin)
+	SELECT i, j, xmin FROM repack_test;
 }
 step wakeup_before_lock
 {
-- 
2.52.0

view thread (14+ messages)  latest in thread

Message-ID: <108776.1784105248@localhost>
Permalink:  ../108776.1784105248@localhost/
Also on:    postgresql.org/message-id/108776.1784105248@localhost

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: ah@cybertec.at, pgsql-hackers@lists.postgresql.org
  Subject: Re: REPACK enhancements
  In-Reply-To: <108776.1784105248@localhost>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox