Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.94.2) (envelope-from ) id 1s6R3Z-00BuYF-04 for pgsql-hackers@arkaria.postgresql.org; Mon, 13 May 2024 08:29:13 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.94.2) (envelope-from ) id 1s6R3W-00GA2B-Qe for pgsql-hackers@arkaria.postgresql.org; Mon, 13 May 2024 08:29:11 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.94.2) (envelope-from ) id 1s6R3V-00GA23-ST for pgsql-hackers@lists.postgresql.org; Mon, 13 May 2024 08:29:10 +0000 Received: from m16.mail.163.com ([220.197.31.4]) by makus.postgresql.org with esmtp (Exim 4.94.2) (envelope-from ) id 1s6R3Q-000m11-0M for pgsql-hackers@lists.postgresql.org; Mon, 13 May 2024 08:29:07 +0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:Subject:Date:Message-ID:MIME-Version: Content-Type; bh=Y1vAOA9se4wNN+oCX8DNisd8vDGm/udH8B4OpZP+fB4=; b=c3x2skJFakqWkbKVeyd/ik/aLpixcPQgLggZoYebZpqLeGeE53OdTw6Nz3ZXbH KNIKAsjYCIP7TO6kDsiM52b3ozBzgI/OInYBfCoA0at6dv62iG+EIlE9TDryzhLC I+jOhC9IHtSE3sBLauOgp5lL+fndH5MuiQJATR18a7GhM= Received: from 8235eee8a2a0 (unknown [140.205.118.142]) by gzga-smtp-mta-g2-1 (Coremail) with SMTP id _____wDn7+ZEz0FmSoZJEQ--.41013S3; Mon, 13 May 2024 16:28:53 +0800 (CST) References: <6ab4003f-a8b8-4d75-a67f-f25ad98582dc@enterprisedb.com> <87pltvmgdm.fsf@163.com> <3b721981-6fa3-4698-a9b6-70b2d8e8fa3b@enterprisedb.com> User-agent: mu4e 1.10.7; emacs 29.1 From: Andy Fan To: Tomas Vondra Cc: pgsql-hackers@lists.postgresql.org Subject: Re: Parallel CREATE INDEX for GIN indexes Date: Mon, 13 May 2024 16:19:43 +0800 In-reply-to: <3b721981-6fa3-4698-a9b6-70b2d8e8fa3b@enterprisedb.com> Message-ID: <87y18ektdn.fsf@163.com> MIME-Version: 1.0 Content-Type: text/plain X-CM-TRANSID: _____wDn7+ZEz0FmSoZJEQ--.41013S3 X-Coremail-Antispam: 1Uf129KBjvJXoWxXw4fZrykWF1kZrWUGr1DGFg_yoW5WrykpF Waqa4akF4DGrW8ZF12vw40yFyjyw4rtFWUArn5CrWUC398Zrs29ryDKrW29F1kurn7CF1Y vF4jgwn8Cw4qvaDanT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0ztg4kNUUUUU= X-Originating-IP: [140.205.118.142] X-CM-SenderInfo: x2klx3xlid0iqsrtqiywtou0bp/xtbBZw3dU2V4HFeEbwABsY List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk Tomas Vondra writes: >>> 7) v20240502-0007-Detect-wrap-around-in-parallel-callback.patch >>> >>> There's one more efficiency problem - the parallel scans are required to >>> be synchronized, i.e. the scan may start half-way through the table, and >>> then wrap around. Which however means the TID list will have a very wide >>> range of TID values, essentially the min and max of for the key. >> >> I have two questions here and both of them are generall gin index questions >> rather than the patch here. >> >> 1. What does the "wrap around" mean in the "the scan may start half-way >> through the table, and then wrap around". Searching "wrap" in >> gin/README gets nothing. >> > > The "wrap around" is about the scan used to read data from the table > when building the index. A "sync scan" may start e.g. at TID (1000,0) > and read till the end of the table, and then wraps and returns the > remaining part at the beginning of the table for blocks 0-999. > > This means the callback would not see a monotonically increasing > sequence of TIDs. > > Which is why the serial build disables sync scans, allowing simply > appending values to the sorted list, and even with regular flushes of > data into the index we can simply append data to the posting lists. Thanks for the hints, I know the sync strategy comes from syncscan.c now. >>> Without 0006 this would cause frequent failures of the index build, with >>> the error I already mentioned: >>> >>> ERROR: could not split GIN page; all old items didn't fit >> 2. I can't understand the below error. >> >>> ERROR: could not split GIN page; all old items didn't fit > if (!append || ItemPointerCompare(&maxOldItem, &remaining) >= 0) > elog(ERROR, "could not split GIN page; all old items didn't fit"); > > It can fail simply because of the !append part. Got it, Thanks! >> If we split the blocks among worker 1-block by 1-block, we will have a >> serious issue like here. If we can have N-block by N-block, and N-block >> is somehow fill the work_mem which makes the dedicated temp file, we >> can make things much better, can we? > I don't understand the question. The blocks are distributed to workers > by the parallel table scan, and it certainly does not do that block by > block. But even it it did, that's not a problem for this code. OK, I get ParallelBlockTableScanWorkerData.phsw_chunk_size is designed for this. > The problem is that if the scan wraps around, then one of the TID lists > for a given worker will have the min TID and max TID, so it will overlap > with every other TID list for the same key in that worker. And when the > worker does the merging, this list will force a "full" merge sort for > all TID lists (for that key), which is very expensive. OK. Thanks for all the answers, they are pretty instructive! -- Best Regards Andy Fan