pg.ddx.io  pgsql-hackers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Arseny Sher <a.sher@postgrespro.ru>
To: Amit Kapila <amit.kapila16@gmail.com>
Cc: Andres Freund <andres@anarazel.de>
Cc: Hsu\, John <hsuchen@amazon.com>
Cc: pgsql-bugs\@lists.postgresql.org <pgsql-bugs@lists.postgresql.org>
Subject: Re: ERROR: subtransaction logged without previous top-level txn record
Date: Mon, 03 Feb 2020 12:20:30 +0300
Message-ID: <8736bs81sx.fsf@ars-thinkpad> (raw)
In-Reply-To: <CAA4eK1L=MDbmGu5-+BmY7Svc07jr+ZabiH8C_qo3RSc8pgUpDQ@mail.gmail.com>
References: <AB5978B2-1772-4FEE-A245-74C91704ECB0@amazon.com>
	<87ftjifoql.fsf@ars-thinkpad>
	<20191024213157.7pm6niybfxgpvmgg@alap3.anarazel.de>
	<87eez1fh48.fsf@ars-thinkpad>
	<CAA4eK1L=MDbmGu5-+BmY7Svc07jr+ZabiH8C_qo3RSc8pgUpDQ@mail.gmail.com>


Amit Kapila <amit.kapila16@gmail.com> writes:

>> I don't see a bug here. At least in reproduced scenario I see false
>> alert, as explained above: transaction with skipped xl_xact_assignment
>> won't be streamed as it finishes before confirmed_flush_lsn.
>>
>
> Does this guarantee come from the fact that we need to wait for such a
> transaction before reaching a consistent snapshot state?  If not, can
> you explain a bit more what makes you say so?

Right, see FULL_SNAPSHOT -> SNAPBUILD_CONSISTENT transition -- it exists
exactly for this purpose: once we have good snapshot, we need to wait
for all running xacts to finish to see all xacts we are promising to
stream in full. This ensures <restart_lsn, confirmed_flush_lsn> pair is
good (reading WAL since the former is enough to see all xacts committing
after the latter in full) initially, and slot advancement arrangements
ensure it stays good forever (see
LogicalIncreaseRestartDecodingForSlot).

Well, almost. This is true as long initial snapshot construction process
goes the long way of building the snapshot by itself. If it happens to
pick up from disk ready snapshot pickled there by another decoding
session, it fast path'es to SNAPBUILD_CONSISTENT, which is technically a
bug as described in
https://www.postgresql.org/message-id/87ftjifoql.fsf%40ars-thinkpad

In theory, this bug could indeed lead to 'subtransaction logged without
previous top-level txn record' error. In practice, I think its
possibility is disappearingly small -- process of slot creation must be
intervened in a very short gap by another decoder who serializes its
snapshot (see the exact sequence of steps in the mail above). What is
much more probable (doesn't involve new slot creation and relatively
easily reproducible without sleeps) is false alert triggered by unlucky
position of restart_lsn.


Surely we still must fix it. I just mean
  - People definitely encountered false alert, not this bug
    (at least because nobody said this was immediately after slot
    creation).
  - I've no bright ideas how to relax the check to make it proper
    without additional complications and I'm pretty sure this is
    impossible (again, see above for details), so I'd remove it.


--
Arseny Sher
Postgres Professional: http://www.postgrespro.com
The Russian Postgres Company





view thread (40+ messages)  latest in thread

Message-ID: <8736bs81sx.fsf@ars-thinkpad>
Permalink:  ../8736bs81sx.fsf@ars-thinkpad/
Also on:    postgresql.org/message-id/8736bs81sx.fsf@ars-thinkpad

 · 

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: a.sher@postgrespro.ru, amit.kapila16@gmail.com, andres@anarazel.de, hsuchen@amazon.com, pgsql-bugs@lists.postgresql.org
  Subject: Re: ERROR: subtransaction logged without previous top-level txn record
  In-Reply-To: <8736bs81sx.fsf@ars-thinkpad>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox