From: Arseny Sher <a.sher@postgrespro.ru>
To: Dan Katz <dkatz@joor.com>
Cc: Andres Freund <andres@anarazel.de>
Cc: Alvaro Herrera <alvherre@2ndquadrant.com>
Cc: Hsu\, John <hsuchen@amazon.com>
Cc: pgsql-bugs\@lists.postgresql.org <pgsql-bugs@lists.postgresql.org>
Subject: Re: ERROR: subtransaction logged without previous top-level txn record
Date: Fri, 31 Jan 2020 00:22:46 +0300
Message-ID: <87ftfwwsex.fsf@ars-thinkpad> (raw)
In-Reply-To: <CAJiDyLG0K0zr5v5Cpbu4xXP7zEDww2aX8vMejxHEJwFfeNP4qw@mail.gmail.com>
References: <AB5978B2-1772-4FEE-A245-74C91704ECB0@amazon.com>
<87ftjifoql.fsf@ars-thinkpad>
<20191024213157.7pm6niybfxgpvmgg@alap3.anarazel.de>
<87eez1fh48.fsf@ars-thinkpad>
<87k185ba4x.fsf@ars-thinkpad>
<CAJiDyLFh4Fx0VMdf0Emwxsq03RfswT2v_cwR4FqoBvksPG1yvw@mail.gmail.com>
<87o8w7vv8h.fsf@ars-thinkpad>
<CAJiDyLG0K0zr5v5Cpbu4xXP7zEDww2aX8vMejxHEJwFfeNP4qw@mail.gmail.com>
Hi,
Dan Katz <dkatz@joor.com> writes:
> Arseny,
>
> I was hoping you could give me some insights about how this bug might
> appear with multiple replications slots. For example if I have two
> replication slots would you expect both slots to see the same error, even
> if they were started, consumed or the LSN was confirmed-flushed at
> different times?
Well, to encounter this you must happen to interrupt decoding session
(e.g. shutdown server) when restart_lsn (LSN since WAL will be read next
time) is at unfortunate position, as described in
https://www.postgresql.org/message-id/87ftjifoql.fsf%40ars-thinkpad
Generally each slot has its own restart_lsn, so if one decoding session
stucked on this issue, another one won't necessarily fail at the same
time. However, restart_lsn can be advanced only to certain points,
mainly xl_running_xacts records, which is logged every 15 seconds. So if
all consumers acknowledge changes fast enough, it is quite likely that
during shutdown restart_lsn will be the same for all slots -- which
means either all of them will stuck on further decoding or all of them
won't. If not, different slots might have different restart_lsn and
probably won't fail at the same time; but encountering this issue even
once suggests that your workload makes possibility of such problematic
restart_lsn perceptible (i.e. many subtransactions). And each
restart_lsn probably has approximately the same chance to be 'bad'
(provided the workload is even).
We need a committer familiar with this code to look here...
--
Arseny Sher
Postgres Professional: http://www.postgrespro.com
The Russian Postgres Company
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-hackers@postgresql.org
Cc: a.sher@postgrespro.ru, dkatz@joor.com, andres@anarazel.de, alvherre@2ndquadrant.com, hsuchen@amazon.com, pgsql-bugs@lists.postgresql.org
Subject: Re: ERROR: subtransaction logged without previous top-level txn record
In-Reply-To: <87ftfwwsex.fsf@ars-thinkpad>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox