Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1ja9PE-0007kI-DO for pgsql-hackers@arkaria.postgresql.org; Sun, 17 May 2020 02:52:00 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.92) (envelope-from ) id 1ja9PB-00019c-Pc for pgsql-hackers@arkaria.postgresql.org; Sun, 17 May 2020 02:51:57 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1ja9PB-000177-ID for pgsql-hackers@lists.postgresql.org; Sun, 17 May 2020 02:51:57 +0000 Received: from mail-qv1-xf42.google.com ([2607:f8b0:4864:20::f42]) by makus.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.92) (envelope-from ) id 1ja9P8-0000bd-Rp for pgsql-hackers@lists.postgresql.org; Sun, 17 May 2020 02:51:56 +0000 Received: by mail-qv1-xf42.google.com with SMTP id g20so3101791qvb.9 for ; Sat, 16 May 2020 19:51:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=2ndquadrant-com.20150623.gappssmtp.com; s=20150623; h=date:from:to:cc:subject:message-id:mime-version:content-disposition :content-transfer-encoding:in-reply-to:user-agent; bh=tExEowZGz2VnHrUtaOGLklykp0hjOokmTO7ZSGVWpow=; b=XAK+q/Jlq5d0xxRBABMdaQDrTRjRdVOt6MgPwa8Elvs+Doi3kb1vJ/pmuquJwC1sqm XsQiv5mDT739AyW8VERDaWt/oEr4tyG2AtjrhOQbWP7F4mbdVincmcq1yy78VWS+F+rZ Uv7DAhTXmi0mqv7CkZUVbMVL6504yN/NGjtJAPpu4KKJm++yH9YfD3UseT4DtRQZbNqL T0zlpDdkhR+WgOucK7Un4MGyhNiGQI0/ER65iTQEy3do1bzwA0eF2XgJfXXZIW1BwZgB MHgeWw/TII/hEeN9HVcay7WZvY88HsP9QtmLBl7/IxlAN/AUY78SWYz0TAQmZojoY3sz TM+g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:date:from:to:cc:subject:message-id:mime-version :content-disposition:content-transfer-encoding:in-reply-to :user-agent; bh=tExEowZGz2VnHrUtaOGLklykp0hjOokmTO7ZSGVWpow=; b=MAK05I7IoSlcitWAUh/9IQMglhsxvnh1Ol9wv+t1sEPPS1Ea3xq/kccfWs7ou2koEz qEowxuDam1evqNcBMrCJGsxGiHULTNR1immroHDvkL+I20whsfM6kWNrZpTqyIv4mf32 FyoeXsoExCL/KXmQtSz7TJRvu72/C0Y+wQPLxB9VKv9rX1aLqU7ireTagmD/01unx2Su bOjoCJcB15bnX5wy51jv3dvCurfJ8tgHxHO/YQ1dcMEQhg/U9xIDAirBW3M+lEKiYkaG DxYm5Tsm9peM8b+f+tv/iUaRcwoON8u5EXgiaNWHkXKUX5S2iQEoZiJydEtF2VyAsRcU pY7A== X-Gm-Message-State: AOAM530XVx/1t7xhb5Fo/AwflyjGIksQ+VpO8cquo7QGRT2hG5nPQ/fi kEhBaZR+8/4G1P8dru73AodVSw== X-Google-Smtp-Source: ABdhPJygPCD2pYQmMi6cj8clsYlcZrIrgjYnLVbmIyA8+17/F6CX9uWnMtCuuTvY+Iu96M0+tfFyJw== X-Received: by 2002:a0c:f409:: with SMTP id h9mr10426457qvl.187.1589683913853; Sat, 16 May 2020 19:51:53 -0700 (PDT) Received: from nimloth.alvh.no-ip.org ([190.95.18.252]) by smtp.gmail.com with ESMTPSA id q10sm6133383qtk.54.2020.05.16.19.51.52 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 16 May 2020 19:51:53 -0700 (PDT) Received: by nimloth.alvh.no-ip.org (Postfix, from userid 1000) id AA1A8300703; Sat, 16 May 2020 22:51:50 -0400 (-04) Date: Sat, 16 May 2020 22:51:50 -0400 From: Alvaro Herrera To: Andres Freund Cc: Kyotaro Horiguchi , jgdr@dalibo.com, michael@paquier.xyz, sawada.mshk@gmail.com, peter.eisentraut@2ndquadrant.com, pgsql-hackers@lists.postgresql.org, thomas.munro@enterprisedb.com, sk@zsrv.org, michael.paquier@gmail.com Subject: Re: [HACKERS] Restricting maximum keep segments by repslots Message-ID: <20200517025150.GA12478@alvherre.pgsql> MIME-Version: 1.0 Content-Type: text/plain; charset=iso-8859-1 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <20200517010005.jzaf2245w4rrgs2o@alap3.anarazel.de> User-Agent: Mutt/1.10.1 (2018-07-13) List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Precedence: bulk On 2020-May-16, Andres Freund wrote: > I, independent of this patch, added a few additional paths in which > checkpointer's latch is reset, and I found a few shutdowns in regression > tests to be extremely slow / timing out. The reason for that is that > the only check for interrupts is at the top of the loop. So if > checkpointer gets SIGUSR2 we don't see ShutdownRequestPending until we > decide to do a checkpoint for other reasons. Ah, yeah, this seems a genuine bug. > I also suspect that it could have harmful consequences to not do a > AbsorbSyncRequests() if something "ate" the set latch. I traced through this when looking over the previous fix, and given that checkpoint execution itself calls AbsorbSyncRequests frequently, I don't think this one qualifies as a bug. > I don't think it's reasonable to expect this much code between a > ResetLatch and WaitLatch to never reset a latch. So I think we need to > make the coding more robust in face of that. Without having to duplicate > the top and the bottom of the loop. That makes sense to me. > One way to do that would be to WaitLatch() call to much earlier, and > only do a WaitLatch() if do_checkpoint is false. Roughly like in the > attached. Hm. I'd do "WaitLatch() / continue" in the "!do_checkpoint" block, and put the checpkoint code not in the else block; seems easier to read to me. While we're here, can we change CreateCheckPoint to return true so that we can do ckpt_performed = do_restartpoint ? CreateRestartPoint(flags) : CreateCheckPoint(flags); instead of the mess we have there now? (Also add a comment that CreateCheckPoint must not return false, to avoid messing with the schedule) -- Álvaro Herrera https://www.2ndQuadrant.com/ PostgreSQL Development, 24x7 Support, Remote DBA, Training & Services