Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_CBC_SHA1:256) (Exim 4.89) (envelope-from ) id 1hnBQo-0001OJ-Ex for pgsql-hackers@arkaria.postgresql.org; Tue, 16 Jul 2019 00:34:58 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.89) (envelope-from ) id 1hnBQn-0006Yw-3g for pgsql-hackers@arkaria.postgresql.org; Tue, 16 Jul 2019 00:34:57 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_CBC_SHA1:256) (Exim 4.89) (envelope-from ) id 1hnBQm-0006Yp-R3 for pgsql-hackers@lists.postgresql.org; Tue, 16 Jul 2019 00:34:56 +0000 Received: from mail-wr1-x444.google.com ([2a00:1450:4864:20::444]) by magus.postgresql.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_CBC_SHA1:256) (Exim 4.89) (envelope-from ) id 1hnBQj-0002pu-VN for pgsql-hackers@postgresql.org; Tue, 16 Jul 2019 00:34:56 +0000 Received: by mail-wr1-x444.google.com with SMTP id g17so18951915wrr.5 for ; Mon, 15 Jul 2019 17:34:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=2ndquadrant-com.20150623.gappssmtp.com; s=20150623; h=date:from:to:cc:subject:message-id:references:mime-version :content-disposition:in-reply-to:user-agent; bh=Xw//Yiunnilch80N7uSf3+vKVcwLHQHkfDyX/QxhFrU=; b=CMyxFPy+3VlX7w1gA14Uwuq3MzJQ7glAY9wXNNSZu4ZLkytsHly3sXe937+GMEWh9X wc3HFHwmWEIRHaMqGxZ5r54CKpHznuCBwA+IRiyafCQRuLuKwaxdVKYzX+DO9iHU+vgj 6Mv5grl4eh6Ej2XB2Ky2nabYBDo5szgQuX4L+YyNB2FguEJ7OFieYiH5i86KqY0LhVLx mwe5KrXmmbtVeS0rlOIclFV8V6xQ110KX85jTeqFYy9nNrhy9V0N+2cPvIL7IRlDmIqF T4Ow7DWmEGL2Axrh8pKg0PPpgiNBtvEv7i4TkXcyrrTqIEg8qsU/9U5Fpz0Kvn/4Ta4p FNZA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:date:from:to:cc:subject:message-id:references :mime-version:content-disposition:in-reply-to:user-agent; bh=Xw//Yiunnilch80N7uSf3+vKVcwLHQHkfDyX/QxhFrU=; b=m34G3QGXyJfdAhPu+eifyYQBAPDo5tb835SUTDsg6vJaVWvgN46iLAiZ93/cf7jHFJ Ti9sp2TtXd+T+e6ZtYonAAOM6NkNp6irYTsDb0pkMhLQv+WPU1Yf25hLCgquol/0lxGK Dzdybdbe13fELpZH8zJ4h+QkeSgPqbQFKIc2WMPPqJ/VxzWVKVWL20qAdkFMv/wcoKYa yJNWLIdZDgtqhXennx5VizLCGStKwsWcYsY/zBXa7P/Y6ikgrb9l7Flqa9zLthBI8Dfu Qgb/PoPnsPp5g2XldQqPPf5285QKTjnFpUNo4MWFZnv+2KxDqMorC00uHR6qvi0lzhW8 5rkQ== X-Gm-Message-State: APjAAAXkTpMVjOdWN+0jNQvW9JEYBTHoS2MKhzAvEh/qJkAsDcofoUec eXVvf8w0PNJFGiVEiQuQ1f87jw== X-Google-Smtp-Source: APXvYqyNGa/jj11iDeGfEvm73e0TJ12p0NC/thMp706gVq7V6a/+fR9oDE1pFzTXbEu5jjKOcmvxZw== X-Received: by 2002:adf:c508:: with SMTP id q8mr31080804wrf.148.1563237291228; Mon, 15 Jul 2019 17:34:51 -0700 (PDT) Received: from localhost (ip-86-49-253-160.net.upcbroadband.cz. [86.49.253.160]) by smtp.gmail.com with ESMTPSA id p6sm17650549wrq.97.2019.07.15.17.34.50 (version=TLS1_3 cipher=AEAD-AES256-GCM-SHA384 bits=256/256); Mon, 15 Jul 2019 17:34:50 -0700 (PDT) Date: Tue, 16 Jul 2019 02:34:49 +0200 From: Tomas Vondra To: Jerry Sievers Cc: pgsql-hackers@postgresql.org Subject: Re: SegFault on 9.6.14 Message-ID: <20190716003449.fjegtxinhrqubysu@development> References: <87ims2amh6.fsf@jsievers.enova.com> <20190716001529.uvpfngvxtlu5h2fr@development> <87blxuakv4.fsf@jsievers.enova.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii; format=flowed Content-Disposition: inline In-Reply-To: <87blxuakv4.fsf@jsievers.enova.com> User-Agent: NeoMutt/20180716-1444-295967 List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Precedence: bulk On Mon, Jul 15, 2019 at 07:22:55PM -0500, Jerry Sievers wrote: >Tomas Vondra writes: > >> On Mon, Jul 15, 2019 at 06:48:05PM -0500, Jerry Sievers wrote: >> >>>Greetings Hackers. >>> >>>We have a reproduceable case of $subject that issues a backtrace such as >>>seen below. >>> >>>The query that I'd prefer to sanitize before sending is <30 lines of at >>>a glance, not terribly complex logic. >>> >>>It nonetheless dies hard after a few seconds of running and as expected, >>>results in an automatic all-backend restart. >>> >>>Please advise on how to proceed. Thanks! >>> >>>bt >>>#0 initscan (scan=scan@entry=0x55d7a7daa0b0, key=0x0, keep_startblock=keep_startblock@entry=1 '\001') >>> at /build/postgresql-9.6-5O8OLM/postgresql-9.6-9.6.14/build/../src/backend/access/heap/heapam.c:233 >>>#1 0x000055d7a72fa8d0 in heap_rescan (scan=0x55d7a7daa0b0, key=key@entry=0x0) at /build/postgresql-9.6-5O8OLM/postgresql-9.6-9.6.14/build/../src/backend/access/heap/heapam.c:1529 >>>#2 0x000055d7a7451fef in ExecReScanSeqScan (node=node@entry=0x55d7a7d85100) at /build/postgresql-9.6-5O8OLM/postgresql-9.6-9.6.14/build/../src/backend/executor/nodeSeqscan.c:280 >>>#3 0x000055d7a742d36e in ExecReScan (node=0x55d7a7d85100) at /build/postgresql-9.6-5O8OLM/postgresql-9.6-9.6.14/build/../src/backend/executor/execAmi.c:158 >>>#4 0x000055d7a7445d38 in ExecReScanGather (node=node@entry=0x55d7a7d84d30) at /build/postgresql-9.6-5O8OLM/postgresql-9.6-9.6.14/build/../src/backend/executor/nodeGather.c:475 >>>#5 0x000055d7a742d255 in ExecReScan (node=0x55d7a7d84d30) at /build/postgresql-9.6-5O8OLM/postgresql-9.6-9.6.14/build/../src/backend/executor/execAmi.c:166 >>>#6 0x000055d7a7448673 in ExecReScanHashJoin (node=node@entry=0x55d7a7d84110) at /build/postgresql-9.6-5O8OLM/postgresql-9.6-9.6.14/build/../src/backend/executor/nodeHashjoin.c:1019 >>>#7 0x000055d7a742d29e in ExecReScan (node=node@entry=0x55d7a7d84110) at /build/postgresql-9.6-5O8OLM/postgresql-9.6-9.6.14/build/../src/backend/executor/execAmi.c:226 >>> >>> >> >> Hmmm, that means it's crashing here: >> >> if (scan->rs_parallel != NULL) >> scan->rs_nblocks = scan->rs_parallel->phs_nblocks; <--- here >> else >> scan->rs_nblocks = RelationGetNumberOfBlocks(scan->rs_rd); >> >> But clearly, scan is valid (otherwise it'd crash on the if condition), >> and scan->rs_parallel must me non-NULL. Which probably means the pointer >> is (no longer) valid. >> >> Could it be that the rs_parallel DSM disappears on rescan, or something >> like that? > >No clue but something I just tried was to disable parallelism by setting >max_parallel_workers_per_gather to 0 and however the query has not >finished after a few minutes, there is no crash. > That might be a hint my rough analysis was somewhat correct. The question is whether the non-parallel plan does the same thing. Maybe it picks a plan that does not require rescans, or something like that. >Please advise. > It would be useful to see (a) exacution plan of the query, (b) full backtrace and (c) a bit of context for the place where it crashed. Something like (in gdb): bt full list p *scan regards -- Tomas Vondra http://www.2ndQuadrant.com PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services