Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1lE7VI-0003gD-8m for pgsql-hackers@arkaria.postgresql.org; Mon, 22 Feb 2021 09:27:44 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.92) (envelope-from ) id 1lE7VH-0001Qe-7h for pgsql-hackers@arkaria.postgresql.org; Mon, 22 Feb 2021 09:27:43 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1lE7VG-0001Q4-Se for pgsql-hackers@lists.postgresql.org; Mon, 22 Feb 2021 09:27:43 +0000 Received: from oss.nttdata.com ([49.212.34.109]) by makus.postgresql.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1lE7VD-0008NB-TD for pgsql-hackers@postgresql.org; Mon, 22 Feb 2021 09:27:41 +0000 Received: from hnk.local (p2710230-ipbf1703funabasi.chiba.ocn.ne.jp [123.216.93.230]) by oss.nttdata.com (Postfix) with ESMTPSA id 9C7636276D; Mon, 22 Feb 2021 18:27:36 +0900 (JST) X-Virus-Status: Clean X-Virus-Scanned: clamav-milter 0.102.3 at oss.nttdata.com Subject: Re: adding wait_start column to pg_locks To: torikoshia Cc: Ian Lawrence Barwick , Robert Haas , Justin Pryzby , pgsql-hackers References: <8002219d-b999-12fa-2327-49afe40e5fbb@oss.nttdata.com> <88c451c7-2729-4329-fbae-6ed7664adea1@oss.nttdata.com> <3db333a9-31f7-a225-8038-eeb835ba8f37@oss.nttdata.com> <990e5d4b-075f-101f-aec5-bb44f9b30550@oss.nttdata.com> <40dfaa75-1058-e811-1f7c-4cf7203a3068@oss.nttdata.com> <64102638-7fc3-1941-12f1-e139e3463c03@oss.nttdata.com> <067d5597e5de82fe47e5716991ab65c4@oss.nttdata.com> <337d5311-6113-15e0-bd85-90b47b55f5ae@oss.nttdata.com> <655f1a5b-4fa8-53ba-c3e2-8cee31587d37@oss.nttdata.com> <1df88660-6f08-cc6e-b7e2-f85296a2bdab@oss.nttdata.com> <62fe634bce475fb7e38e6fb3a2fff124@oss.nttdata.com> From: Fujii Masao Message-ID: <0630dd64-a97f-28d1-5c5d-640cbfc293cc@oss.nttdata.com> Date: Mon, 22 Feb 2021 18:27:36 +0900 User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.14; rv:78.0) Gecko/20100101 Thunderbird/78.7.1 MIME-Version: 1.0 In-Reply-To: <62fe634bce475fb7e38e6fb3a2fff124@oss.nttdata.com> Content-Type: text/plain; charset=utf-8; format=flowed Content-Language: en-US Content-Transfer-Encoding: 8bit List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Precedence: bulk On 2021/02/18 16:26, torikoshia wrote: > On 2021-02-16 16:59, Fujii Masao wrote: >> On 2021/02/15 15:17, Fujii Masao wrote: >>> >>> >>> On 2021/02/10 10:43, Fujii Masao wrote: >>>> >>>> >>>> On 2021/02/09 23:31, torikoshia wrote: >>>>> On 2021-02-09 22:54, Fujii Masao wrote: >>>>>> On 2021/02/09 19:11, Fujii Masao wrote: >>>>>>> >>>>>>> >>>>>>> On 2021/02/09 18:13, Fujii Masao wrote: >>>>>>>> >>>>>>>> >>>>>>>> On 2021/02/09 17:48, torikoshia wrote: >>>>>>>>> On 2021-02-05 18:49, Fujii Masao wrote: >>>>>>>>>> On 2021/02/05 0:03, torikoshia wrote: >>>>>>>>>>> On 2021-02-03 11:23, Fujii Masao wrote: >>>>>>>>>>>>> 64-bit fetches are not atomic on some platforms. So spinlock is necessary when updating "waitStart" without holding the partition lock? Also GetLockStatusData() needs spinlock when reading "waitStart"? >>>>>>>>>>>> >>>>>>>>>>>> Also it might be worth thinking to use 64-bit atomic operations like >>>>>>>>>>>> pg_atomic_read_u64(), for that. >>>>>>>>>>> >>>>>>>>>>> Thanks for your suggestion and advice! >>>>>>>>>>> >>>>>>>>>>> In the attached patch I used pg_atomic_read_u64() and pg_atomic_write_u64(). >>>>>>>>>>> >>>>>>>>>>> waitStart is TimestampTz i.e., int64, but it seems pg_atomic_read_xxx and pg_atomic_write_xxx only supports unsigned int, so I cast the type. >>>>>>>>>>> >>>>>>>>>>> I may be using these functions not correctly, so if something is wrong, I would appreciate any comments. >>>>>>>>>>> >>>>>>>>>>> >>>>>>>>>>> About the documentation, since your suggestion seems better than v6, I used it as is. >>>>>>>>>> >>>>>>>>>> Thanks for updating the patch! >>>>>>>>>> >>>>>>>>>> +    if (pg_atomic_read_u64(&MyProc->waitStart) == 0) >>>>>>>>>> +        pg_atomic_write_u64(&MyProc->waitStart, >>>>>>>>>> + pg_atomic_read_u64((pg_atomic_uint64 *) &now)); >>>>>>>>>> >>>>>>>>>> pg_atomic_read_u64() is really necessary? I think that >>>>>>>>>> "pg_atomic_write_u64(&MyProc->waitStart, now)" is enough. >>>>>>>>>> >>>>>>>>>> +        deadlockStart = get_timeout_start_time(DEADLOCK_TIMEOUT); >>>>>>>>>> +        pg_atomic_write_u64(&MyProc->waitStart, >>>>>>>>>> +                    pg_atomic_read_u64((pg_atomic_uint64 *) &deadlockStart)); >>>>>>>>>> >>>>>>>>>> Same as above. >>>>>>>>>> >>>>>>>>>> +        /* >>>>>>>>>> +         * Record waitStart reusing the deadlock timeout timer. >>>>>>>>>> +         * >>>>>>>>>> +         * It would be ideal this can be synchronously done with updating >>>>>>>>>> +         * lock information. Howerver, since it gives performance impacts >>>>>>>>>> +         * to hold partitionLock longer time, we do it here asynchronously. >>>>>>>>>> +         */ >>>>>>>>>> >>>>>>>>>> IMO it's better to comment why we reuse the deadlock timeout timer. >>>>>>>>>> >>>>>>>>>>      proc->waitStatus = waitStatus; >>>>>>>>>> +    pg_atomic_init_u64(&MyProc->waitStart, 0); >>>>>>>>>> >>>>>>>>>> pg_atomic_write_u64() should be used instead? Because waitStart can be >>>>>>>>>> accessed concurrently there. >>>>>>>>>> >>>>>>>>>> I updated the patch and addressed the above review comments. Patch attached. >>>>>>>>>> Barring any objection, I will commit this version. >>>>>>>>> >>>>>>>>> Thanks for modifying the patch! >>>>>>>>> I agree with your comments. >>>>>>>>> >>>>>>>>> BTW, I ran pgbench several times before and after applying >>>>>>>>> this patch. >>>>>>>>> >>>>>>>>> The environment is virtual machine(CentOS 8), so this is >>>>>>>>> just for reference, but there were no significant difference >>>>>>>>> in latency or tps(both are below 1%). >>>>>>>> >>>>>>>> Thanks for the test! I pushed the patch. >>>>>>> >>>>>>> But I reverted the patch because buildfarm members rorqual and >>>>>>> prion don't like the patch. I'm trying to investigate the cause >>>>>>> of this failures. >>>>>>> >>>>>>> https://buildfarm.postgresql.org/cgi-bin/show_log.pl?nm=rorqual&dt=2021-02-09%2009%3A20%3A10 >>>>>> >>>>>> -    relation     | locktype |        mode >>>>>> ------------------+----------+--------------------- >>>>>> - test_prepared_1 | relation | RowExclusiveLock >>>>>> - test_prepared_1 | relation | AccessExclusiveLock >>>>>> -(2 rows) >>>>>> - >>>>>> +ERROR:  invalid spinlock number: 0 >>>>>> >>>>>> "rorqual" reported that the above error happened in the server built with >>>>>> --disable-atomics --disable-spinlocks when reading pg_locks after >>>>>> the transaction was prepared. The cause of this issue is that "waitStart" >>>>>> atomic variable in the dummy proc created at the end of prepare transaction >>>>>> was not initialized. I updated the patch so that pg_atomic_init_u64() is >>>>>> called for the "waitStart" in the dummy proc for prepared transaction. >>>>>> Patch attached. I confirmed that the patched server built with >>>>>> --disable-atomics --disable-spinlocks passed all the regression tests. >>>>> >>>>> Thanks for fixing the bug, I also tested v9.patch configured with >>>>> --disable-atomics --disable-spinlocks on my environment and confirmed >>>>> that all tests have passed. >>>> >>>> Thanks for the test! >>>> >>>> I found another bug in the patch. InitProcess() initializes "waitStart", >>>> but previously InitAuxiliaryProcess() did not. This could cause "invalid >>>> spinlock number" error when reading pg_locks in the standby server. >>>> I fixed that. Attached is the updated version of the patch. >>> >>> I pushed this version. Thanks! >> >> While reading the patch again, I found two minor things. >> >> 1. As discussed in another thread [1], the atomic variable "waitStart" should >>   be initialized at the postmaster startup rather than the startup of each >>   child process. I changed "waitStart" so that it's initialized in >>   InitProcGlobal() and also reset to 0 by using pg_atomic_write_u64() in >>   InitProcess() and InitAuxiliaryProcess(). >> >> 2. Thanks to the above change, InitProcGlobal() initializes "waitStart" >>   even in PGPROC entries for prepare transactions. But those entries are >>   zeroed in MarkAsPreparingGuts(), so "waitStart" needs to be initialized >>   again. Currently TwoPhaseGetDummyProc() initializes "waitStart" in the >>   PGPROC entry for prepare transaction. But it's better to do that in >>   MarkAsPreparingGuts() instead because that function initializes other >>   PGPROC variables. So I moved that initialization code from >>   TwoPhaseGetDummyProc() to MarkAsPreparingGuts(). >> >> Patch attached. Thought? > > Thanks for updating the patch! > > It seems to me that the modification is right. > I ran some regression tests but didn't find problems. Thanks for the review and test! I pushed the patch. Regards, -- Fujii Masao Advanced Computing Technology Center Research and Development Headquarters NTT DATA CORPORATION