pg.ddx.io  pgsql-bugs@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Tom Lane <tgl@sss.pgh.pa.us>
To: francisco.reinolds@channable.com
Cc: pgsql-bugs@lists.postgresql.org
Subject: Re: BUG #17846: pg_dump doesn't properly dump with paused WAL replay
Date: Thu, 16 Mar 2023 11:10:57 -0400
Message-ID: <3777456.1678979457@sss.pgh.pa.us> (raw)
In-Reply-To: <17846-1a0e5ce976f4c01a@postgresql.org>
References: <17846-1a0e5ce976f4c01a@postgresql.org>

PG Bug reporting form <noreply@postgresql.org> writes:
> For backups, we use pg_dump to perform a full database dump. Before we start
> a backup, we pause the WAL replay on the secondary, unpausing it after it is
> concluded. This was done since we previously encountered problems with
> pg_dump failing when an AccessExclusiveLock was held on a table that pg_dump
> was going to dump.

> For some time we faced no problems with this setup, but starting some months
> ago, we started witnessing sporadic failures when we attempted to restore
> the dumps of one of our databases, to verify the dump's integrity. These
> restore failures would occur due to a key not being present in a table:

I really have no idea what's going on there, but can you show the exact
pg_dump command(s) being issued?  I'm particularly curious whether you
are using parallel dump.  The same for the failing pg_restore.

Also, are all the moving parts (primary server, secondary server,
pg_dump, pg_restore) exactly the same PG version?

> We have managed, with some help from the Postgres IRC channel (special
> thanks to user nickb), to work around the problem. The solution was to begin
> a transaction, and extract a snapshot that'd be passed as a pg_dump
> argument, and only then pause WAL replay. From our understanding, pg_dump
> should already implicitly pick a suitable point to start the dump but it
> apparently is not the case, hence the bug report.

It's the other way around: the replay mechanism should not damage
any data that's visible to an open snapshot.  So I agree this smells
like a bug, but we don't have enough info here to reproduce it.

			regards, tom lane





view thread (4+ messages)  latest in thread

Message-ID: <3777456.1678979457@sss.pgh.pa.us>
Permalink:  ../3777456.1678979457@sss.pgh.pa.us/
Also on:    postgresql.org/message-id/3777456.1678979457@sss.pgh.pa.us

 · 

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-bugs@postgresql.org
  Cc: tgl@sss.pgh.pa.us, francisco.reinolds@channable.com, pgsql-bugs@lists.postgresql.org
  Subject: Re: BUG #17846: pg_dump doesn't properly dump with paused WAL replay
  In-Reply-To: <3777456.1678979457@sss.pgh.pa.us>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox