pg.ddx.io  pgsql-hackers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Arseny Sher <a.sher@postgrespro.ru>
To: Dan Katz <dkatz@joor.com>
Cc: Andres Freund <andres@anarazel.de>
Cc: Alvaro Herrera <alvherre@2ndquadrant.com>
Cc: Hsu\, John <hsuchen@amazon.com>
Cc: pgsql-bugs\@lists.postgresql.org <pgsql-bugs@lists.postgresql.org>
Subject: Re: ERROR: subtransaction logged without previous top-level txn record
Date: Fri, 31 Jan 2020 00:22:46 +0300
Message-ID: <87ftfwwsex.fsf@ars-thinkpad> (raw)
In-Reply-To: <CAJiDyLG0K0zr5v5Cpbu4xXP7zEDww2aX8vMejxHEJwFfeNP4qw@mail.gmail.com>
References: <AB5978B2-1772-4FEE-A245-74C91704ECB0@amazon.com>
	<87ftjifoql.fsf@ars-thinkpad>
	<20191024213157.7pm6niybfxgpvmgg@alap3.anarazel.de>
	<87eez1fh48.fsf@ars-thinkpad>
	<87k185ba4x.fsf@ars-thinkpad>
	<CAJiDyLFh4Fx0VMdf0Emwxsq03RfswT2v_cwR4FqoBvksPG1yvw@mail.gmail.com>
	<87o8w7vv8h.fsf@ars-thinkpad>
	<CAJiDyLG0K0zr5v5Cpbu4xXP7zEDww2aX8vMejxHEJwFfeNP4qw@mail.gmail.com>

Hi,

Dan Katz <dkatz@joor.com> writes:

> Arseny,
>
> I was hoping you could give me some insights about how this bug might
> appear with multiple replications slots. For example if I have two
> replication slots would you expect both slots to see the same error, even
> if they were started, consumed or the LSN was confirmed-flushed at
> different times?

Well, to encounter this you must happen to interrupt decoding session
(e.g. shutdown server) when restart_lsn (LSN since WAL will be read next
time) is at unfortunate position, as described in
https://www.postgresql.org/message-id/87ftjifoql.fsf%40ars-thinkpad

Generally each slot has its own restart_lsn, so if one decoding session
stucked on this issue, another one won't necessarily fail at the same
time. However, restart_lsn can be advanced only to certain points,
mainly xl_running_xacts records, which is logged every 15 seconds. So if
all consumers acknowledge changes fast enough, it is quite likely that
during shutdown restart_lsn will be the same for all slots -- which
means either all of them will stuck on further decoding or all of them
won't. If not, different slots might have different restart_lsn and
probably won't fail at the same time; but encountering this issue even
once suggests that your workload makes possibility of such problematic
restart_lsn perceptible (i.e. many subtransactions). And each
restart_lsn probably has approximately the same chance to be 'bad'
(provided the workload is even).


We need a committer familiar with this code to look here...


--
Arseny Sher
Postgres Professional: http://www.postgrespro.com
The Russian Postgres Company





view thread (40+ messages)  latest in thread

Message-ID: <87ftfwwsex.fsf@ars-thinkpad>
Permalink:  ../87ftfwwsex.fsf@ars-thinkpad/
Also on:    postgresql.org/message-id/87ftfwwsex.fsf@ars-thinkpad

 · 

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: a.sher@postgrespro.ru, dkatz@joor.com, andres@anarazel.de, alvherre@2ndquadrant.com, hsuchen@amazon.com, pgsql-bugs@lists.postgresql.org
  Subject: Re: ERROR: subtransaction logged without previous top-level txn record
  In-Reply-To: <87ftfwwsex.fsf@ars-thinkpad>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox