From: Justin Pryzby <pryzby@telsasoft.com>
To: Amit Kapila <amit.kapila16@gmail.com>
Cc: Masahiko Sawada <masahiko.sawada@2ndquadrant.com>
Cc: Alvaro Herrera <alvherre@2ndquadrant.com>
Cc: Andres Freund <andres@anarazel.de>
Cc: Michael Paquier <michael@paquier.xyz>
Cc: pgsql-hackers <pgsql-hackers@postgresql.org>
Subject: Re: error context for vacuum to include block number
Date: Fri, 27 Mar 2020 01:16:01 -0500
Message-ID: <20200327061601.GE20103@telsasoft.com> (raw)
In-Reply-To: <20200327044424.GD20103@telsasoft.com>
References: <20200325124155.GU21443@telsasoft.com>
<CAA4eK1+=ToXurzjOcPyQ4j6Vz2nG2C-cUdg=yZoUpN+g1k5m6w@mail.gmail.com>
<20200326044115.GB28385@telsasoft.com>
<CAA4eK1LBtp+qpsz8uR-oj2h6jzhLEVo9TEdm46anKM-n=8nS0g@mail.gmail.com>
<CAA4eK1L33gvQ1z-CHctzTWng2HfmFn0cDJif_mpBOvG47Mymow@mail.gmail.com>
<CA+fd4k61uPzeAWT1EZOwPWq7DkhgVTrC_wmUNCzVW+MbP9Wjfg@mail.gmail.com>
<20200326150457.GB17431@telsasoft.com>
<20200326221752.GR17431@telsasoft.com>
<CAA4eK1+b5-aJcbpvqCFB7Qi0zLKTZNJmA+Js644M2YeiEyTZtw@mail.gmail.com>
<20200327044424.GD20103@telsasoft.com>
On Thu, Mar 26, 2020 at 11:44:24PM -0500, Justin Pryzby wrote:
> On Fri, Mar 27, 2020 at 09:49:29AM +0530, Amit Kapila wrote:
> > On Fri, Mar 27, 2020 at 3:47 AM Justin Pryzby <pryzby@telsasoft.com> wrote:
> > >
> > > > Hm, I was just wondering what happens if an error happens *during*
> > > > update_vacuum_error_cbarg(). It seems like if we set
> > > > errcbarg->phase=VACUUM_INDEX before setting errcbarg->indname=indname, then an
> > > > error would cause a crash.
> > > >
> >
> > Can't that be avoided if you check if cbarg->indname is non-null in
> > vacuum_error_callback as we are already doing for
> > VACUUM_ERRCB_PHASE_TRUNCATE?
> >
> > > > And if we pfree and set indname before phase, it'd
> > > > be a problem when going from an index phase to non-index phase.
> >
> > How is it possible that we move to the non-index phase without
> > clearing indname as we always revert back the old phase information?
>
> The crash scenario I'm trying to avoid would be like statement_timeout or other
> asynchronous event occurring between two non-atomic operations.
>
> I said that there's an issue no matter what order we set indname/phase;
> If we wrote:
> |cbarg->indname = indname;
> |cbarg->phase = phase;
> ..and hit a timeout (or similar) between setting indname=NULL but before
> setting phase=VACUUM_INDEX, then we can crash due to null pointer.
>
> But if we write:
> |cbarg->phase = phase;
> |if (cbarg->indname) {pfree(cbarg->indname);}
> |cbarg->indname = indname ? pstrdup(indname) : NULL;
> ..then we can still crash if we timeout between freeing cbarg->indname and
> setting it to null, due to acccessing a pfreed allocation.
If "phase" is updated before "indname", I'm able to induce a synthetic crash
like this:
+if (errinfo->phase==VACUUM_ERRCB_PHASE_VACUUM_INDEX && errinfo->indname==NULL)
+{
+kill(getpid(), SIGINT);
+pg_sleep(1); // that's needed since signals are delivered asynchronously
+}
And another crash if we do this after pfree but before setting indname.
+if (errinfo->phase==VACUUM_ERRCB_PHASE_VACUUM_INDEX && errinfo->indname!=NULL)
+{
+kill(getpid(), SIGINT);
+pg_sleep(1);
+}
I'm not sure if those are possible outside of "induced" errors. Maybe the
function is essentially atomic due to no CHECK_FOR_INTERRUPTS or similar?
--
Justin
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-hackers@postgresql.org
Cc: pryzby@telsasoft.com, amit.kapila16@gmail.com, masahiko.sawada@2ndquadrant.com, alvherre@2ndquadrant.com, andres@anarazel.de, michael@paquier.xyz
Subject: Re: error context for vacuum to include block number
In-Reply-To: <20200327061601.GE20103@telsasoft.com>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox