Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1pcpFv-00036K-Vb for pgsql-bugs@arkaria.postgresql.org; Thu, 16 Mar 2023 15:11:03 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.92) (envelope-from ) id 1pcpFu-00086d-Js for pgsql-bugs@arkaria.postgresql.org; Thu, 16 Mar 2023 15:11:02 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1pcpFu-00085b-Be for pgsql-bugs@lists.postgresql.org; Thu, 16 Mar 2023 15:11:02 +0000 Received: from sss.pgh.pa.us ([66.207.139.130]) by magus.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1pcpFr-0000rq-RW for pgsql-bugs@lists.postgresql.org; Thu, 16 Mar 2023 15:11:01 +0000 Received: from sss1.sss.pgh.pa.us (localhost [127.0.0.1]) by sss.pgh.pa.us (8.15.2/8.15.2) with ESMTP id 32GFAvXt3777457; Thu, 16 Mar 2023 11:10:57 -0400 From: Tom Lane To: francisco.reinolds@channable.com cc: pgsql-bugs@lists.postgresql.org Subject: Re: BUG #17846: pg_dump doesn't properly dump with paused WAL replay In-reply-to: <17846-1a0e5ce976f4c01a@postgresql.org> References: <17846-1a0e5ce976f4c01a@postgresql.org> Comments: In-reply-to PG Bug reporting form message dated "Thu, 16 Mar 2023 11:01:42 -0000" MIME-Version: 1.0 Content-Type: text/plain; charset="us-ascii" Content-ID: <3777455.1678979457.1@sss.pgh.pa.us> Date: Thu, 16 Mar 2023 11:10:57 -0400 Message-ID: <3777456.1678979457@sss.pgh.pa.us> List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk PG Bug reporting form writes: > For backups, we use pg_dump to perform a full database dump. Before we start > a backup, we pause the WAL replay on the secondary, unpausing it after it is > concluded. This was done since we previously encountered problems with > pg_dump failing when an AccessExclusiveLock was held on a table that pg_dump > was going to dump. > For some time we faced no problems with this setup, but starting some months > ago, we started witnessing sporadic failures when we attempted to restore > the dumps of one of our databases, to verify the dump's integrity. These > restore failures would occur due to a key not being present in a table: I really have no idea what's going on there, but can you show the exact pg_dump command(s) being issued? I'm particularly curious whether you are using parallel dump. The same for the failing pg_restore. Also, are all the moving parts (primary server, secondary server, pg_dump, pg_restore) exactly the same PG version? > We have managed, with some help from the Postgres IRC channel (special > thanks to user nickb), to work around the problem. The solution was to begin > a transaction, and extract a snapshot that'd be passed as a pg_dump > argument, and only then pause WAL replay. From our understanding, pg_dump > should already implicitly pick a suitable point to start the dump but it > apparently is not the case, hence the bug report. It's the other way around: the replay mechanism should not damage any data that's visible to an open snapshot. So I agree this smells like a bug, but we don't have enough info here to reproduce it. regards, tom lane