Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1peHby-0004hb-Bs for pgsql-bugs@arkaria.postgresql.org; Mon, 20 Mar 2023 15:39:50 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.92) (envelope-from ) id 1peHbx-0006gY-30 for pgsql-bugs@arkaria.postgresql.org; Mon, 20 Mar 2023 15:39:49 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1peHbw-0006gP-R0 for pgsql-bugs@lists.postgresql.org; Mon, 20 Mar 2023 15:39:48 +0000 Received: from sss.pgh.pa.us ([66.207.139.130]) by makus.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1peHbu-0003Db-Iw for pgsql-bugs@lists.postgresql.org; Mon, 20 Mar 2023 15:39:47 +0000 Received: from sss1.sss.pgh.pa.us (localhost [127.0.0.1]) by sss.pgh.pa.us (8.15.2/8.15.2) with ESMTP id 32KFdixF941276; Mon, 20 Mar 2023 11:39:44 -0400 From: Tom Lane To: Francisco Reinolds cc: pgsql-bugs@lists.postgresql.org Subject: Re: BUG #17846: pg_dump doesn't properly dump with paused WAL replay In-reply-to: References: <17846-1a0e5ce976f4c01a@postgresql.org> <3777456.1678979457@sss.pgh.pa.us> Comments: In-reply-to Francisco Reinolds message dated "Mon, 20 Mar 2023 14:31:20 +0100" MIME-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-ID: <941274.1679326784.1@sss.pgh.pa.us> Content-Transfer-Encoding: 8bit Date: Mon, 20 Mar 2023 11:39:44 -0400 Message-ID: <941275.1679326784@sss.pgh.pa.us> List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk [ please keep the mailing list cc'd ] Francisco Reinolds writes: > On 16-03-2023 16:10, Tom Lane wrote: >> I really have no idea what's going on there, but can you show the exact >> pg_dump command(s) being issued? I'm particularly curious whether you >> are using parallel dump. The same for the failing pg_restore. > Of course: > - pg_dump: pg_dump --port 5432 --host localhost --verbose > --format=directory --jobs=8 --file= --dbname= > - pg_restore: pg_restore --exit-on-error --cluster 13/ > --dbname= --port --format=directory --jobs=8 > --use-list=/tmp/tmpsote5wvm --clean --if-exists Hmm, so the fact that the dump is being done in parallel is very likely relevant. Perhaps parallelism on the restore is also relevant, not sure. Can you try running each of those steps not-parallel to see if the problem goes away? I'm also slightly troubled by the --use-list option, and am wondering if faulty creation of the restore list could be a contributing factor. The error looks like missing data row(s) not missing schema objects; but perhaps if the problematic table(s) are partitioned then one could lead to the other? Could we see the DDL definition for the problematic table(s)? >> Also, are all the moving parts (primary server, secondary server, >> pg_dump, pg_restore) exactly the same PG version? > So, the version of both the primary and the secondary servers match, 13.8, > but the server of the instance where we run the backup verifications does > not, it's currently sitting at 13.6 Hmm. With some unsupported assumptions about your schema, I could believe that some of the 13.9 bug fixes are relevant, particularly * Fix construction of per-partition foreign key constraints while doing ALTER TABLE ATTACH PARTITION (Jehan-Guillaume de Rorthais, Álvaro Herrera) Previously, incorrect or duplicate constraints could be constructed for the newly-added partition. regards, tom lane