Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.94.2) (envelope-from ) id 1rxurd-00D5ai-At for pgsql-hackers@arkaria.postgresql.org; Fri, 19 Apr 2024 20:29:41 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.94.2) (envelope-from ) id 1rxurb-005Lav-Ks for pgsql-hackers@arkaria.postgresql.org; Fri, 19 Apr 2024 20:29:39 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.94.2) (envelope-from ) id 1rxurb-005Lan-BL for pgsql-hackers@lists.postgresql.org; Fri, 19 Apr 2024 20:29:39 +0000 Received: from mail-io1-xd2e.google.com ([2607:f8b0:4864:20::d2e]) by makus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 (Exim 4.94.2) (envelope-from ) id 1rxurX-003fTf-PN for pgsql-hackers@postgresql.org; Fri, 19 Apr 2024 20:29:37 +0000 Received: by mail-io1-xd2e.google.com with SMTP id ca18e2360f4ac-7d6c4c97875so98932739f.3 for ; Fri, 19 Apr 2024 13:29:35 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1713558574; x=1714163374; darn=postgresql.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=RAO2IkNoYxPFoW/O0vFqotAvn5eXRtpa3qrysadwUco=; b=U8dL2P41kTCFMaW/E4O1u0u6z+G01u/nFhvdRC3bqVNTro8Gq5i83OMSKw3DEQpaXV 0JDvqLz6jmq+y7EQsLnwjB56XGT+DtfXb2l5zoypzdzeGWSqvT+sxHUZxXCIWhhym55L gWGAaXpbewgL6M33pULwLeqR/77wLADOPol9Bjf/mJaPkOKGJ8Ig+232YbpkGMK3t216 FPGe9jAcOI1oCyi5hOUUYh7qtyAMGIuHGb2IHTgfZCsjwkefddgnbjtYIXGJ440C65Sq REm8I2kuw3zTv61B4xq1rf4ko6+rIIa+aMRk8MSFYnFlzd0Dn0/1QG7PKSv2NS0HaP5q 3Kgw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1713558574; x=1714163374; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=RAO2IkNoYxPFoW/O0vFqotAvn5eXRtpa3qrysadwUco=; b=jL55aqrVSnjuZWgp68A32FpzI6DE9Qto0J6y4lWTaj9e0lWKWNi8OTURvawyfRdeQt EzRhvBLYjr+kHVCLzu03IJ2JIcLqfIwc1tZyy8IqEIfSbxnXTTbFmuBoBbhfpDGF7Rf7 2QBPmgcbogoHYF4fLZ8jlI7Mw++fVMgnOAoZAH8Va/5pKxLTwH1o90UdNxjTOGlHKcnO IHR2+iOZEibKwMrQGY0/se4Yjn5O1PNcbbO11Njp7F6nkphVCOdNtLecAx/gKdZJsinb NcvwI3cC+26RxvdKlyRRTUZ+BEF4n46tRx0oaT0D3ThbabMc3oppbyhLTmQ9YUocAOEw 1+cw== X-Forwarded-Encrypted: i=1; AJvYcCX7DM02qW4U9bnN+n024rzl39sFGKBwf6Hz9jDxoihIeAoG3dpoIp1ngWkzef/F432tm6pWjb67OSt6k5xu61dLlWSx+DqPk0NQQuQJ X-Gm-Message-State: AOJu0YzFIrgn43pvU3FOQIiYKSkTRsis6vdMvjx5xDoo6CoYiP1yfXd+ Nb+8fU+k1d4Pa6vBa6mYxSyQNgU+R2wdijiSS7Rb+TP5hBnTaqg8 X-Google-Smtp-Source: AGHT+IE8TnOez3iqz5ZqR98RM7/d2lnQNKn1fVblAFGX0cA710UH6rS7UfAzYrJs9bM9tMUaibddoQ== X-Received: by 2002:a6b:7c4a:0:b0:7da:522b:f377 with SMTP id b10-20020a6b7c4a000000b007da522bf377mr2648316ioq.15.1713558574447; Fri, 19 Apr 2024 13:29:34 -0700 (PDT) Received: from nathanxps13 (162-195-168-172.lightspeed.stlsmo.sbcglobal.net. [162.195.168.172]) by smtp.gmail.com with ESMTPSA id w12-20020a056638138c00b00484c1fbd7c0sm899404jad.121.2024.04.19.13.29.33 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 19 Apr 2024 13:29:33 -0700 (PDT) Date: Fri, 19 Apr 2024 15:29:31 -0500 From: Nathan Bossart To: Robert Haas Cc: "Imseih (AWS), Sami" , "pgsql-hackers@postgresql.org" Subject: Re: allow changing autovacuum_max_workers without restarting Message-ID: <20240419202931.GA57652@nathanxps13> References: <20240411192423.GB2005410@nathanxps13> <20240412184000.GA2416270@nathanxps13> <9CD7B92C-3330-4E79-A84E-95E99B1FD926@amazon.com> <20240413194450.GA2537802@nathanxps13> <25FBADC0-52E7-4DE6-B043-68E5F77D4CB1@amazon.com> <20240414144058.GA2819458@nathanxps13> <37F50121-2DC5-4B7D-AB8F-FC2A0CD827A5@amazon.com> <20240419154322.GA3988554@nathanxps13> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk On Fri, Apr 19, 2024 at 02:42:13PM -0400, Robert Haas wrote: > I think this could help a bunch of users, but I'd still like to > complain, not so much with the desire to kill this patch as with the > desire to broaden the conversation. I think I subconsciously hoped this would spark a bigger discussion... > Now, before this patch, there is a fairly good reason for that, which > is that we need to reserve shared memory resources for each autovacuum > worker that might potentially run, and the system can't know how much > shared memory you'd like to reserve for that purpose. But if that were > the only problem, then this patch would probably just be proposing to > crank up the default value of that parameter rather than introducing a > second one. I bet Nathan isn't proposing that because his intuition is > that it will work out badly, and I think he's right. I bet that > cranking up the number of allowed workers will often result in running > more workers than we really should. One possible negative consequence > is that we'll end up with multiple processes fighting over the disk in > a situation where they should just take turns. I suspect there are > also ways that we can be harmed - in broadly similar fashion - by cost > balancing. Even if we were content to bump up the default value of autovacuum_max_workers and tell folks to just mess with the cost settings, there are still probably many cases where bumping up the number of workers further would be necessary. If you have a zillion tables, turning cost-based vacuuming off completely may be insufficient to keep up, at which point your options become limited. It can be difficult to tell whether you might end up in this situation over time as your workload evolves. In any case, it's not clear to me that bumping up the default value of autovacuum_max_workers would do more good than harm. I get the idea that the default of 3 is sufficient for a lot of clusters, so there'd really be little upside to changing it AFAICT. (I guess this proves your point about my intuition.) > So I feel like what this proposal reveals is that we know that our > algorithm for ramping up the number of running workers doesn't really > work. And maybe that's just a consequence of the general problem that > we have no global information about how much vacuuming work there is > to be done at any given time, and therefore we cannot take any kind of > sensible guess about whether 1 more worker will help or hurt. Or, > maybe there's some way to do better than what we do today without a > big rewrite. I'm not sure. I don't think this patch should be burdened > with solving the general problem here. But I do think the general > problem is worth some discussion. I certainly don't want to hold up $SUBJECT for a larger rewrite of autovacuum scheduling, but I also don't want to shy away from a larger rewrite if it's an idea whose time has come. I'm looking forward to hearing your ideas in your pgconf.dev talk. -- Nathan Bossart Amazon Web Services: https://aws.amazon.com