Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1myDnt-0008Al-Us for pgsql-hackers@arkaria.postgresql.org; Fri, 17 Dec 2021 14:01:46 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.92) (envelope-from ) id 1myDnr-0006z3-2x for pgsql-hackers@arkaria.postgresql.org; Fri, 17 Dec 2021 14:01:43 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1myDnq-0006w2-OR for pgsql-hackers@lists.postgresql.org; Fri, 17 Dec 2021 14:01:42 +0000 Received: from mail-wr1-x435.google.com ([2a00:1450:4864:20::435]) by makus.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.92) (envelope-from ) id 1myDnj-0007UW-KT for pgsql-hackers@lists.postgresql.org; Fri, 17 Dec 2021 14:01:41 +0000 Received: by mail-wr1-x435.google.com with SMTP id i22so4176240wrb.13 for ; Fri, 17 Dec 2021 06:01:35 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=aiven.io; s=google; h=from:to:cc:subject:date:message-id:organization:in-reply-to :references:mime-version:content-transfer-encoding; bh=/BhUIHCVywFjth109rRdwGGLpzznW73+Vr0XFg5xVVU=; b=D8HfGMU+Whefv8d4UCJOSQICQs+k+k+O9KR++ioiHNxi2OvP0OtSstBfC7Z9t1iVi4 ZnquLvAKNd49zFgWWKRsMyHrAVtzX1l7z7ffhVqkJwzZxhKFHbHmxIzHDvsXcOIqjIew 4/TzvZQ6kkCoK1JrxtkIMSx3cHiZZaRztbTVE= X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:from:to:cc:subject:date:message-id:organization :in-reply-to:references:mime-version:content-transfer-encoding; bh=/BhUIHCVywFjth109rRdwGGLpzznW73+Vr0XFg5xVVU=; b=mWJ33w/k7vlxlahBJ2x3SMu9/ntnUIw8FBH7q+Kvx7tf1atn7Wuc99p/4R79lZ0XLg mSnHBtAtEfcywK7I5NCvfRMycD9vdLZyX5/deJpddzkfIFDBszCTbUOfOuG7EsVNVmTT SdHm/VkAf7u6xZy3aJ23Lo6YW0VZD6DlSg1xhXWEUufiTL+rCN2XblcWjNO2Wyx8Xvce dFlxYSlYdkuONN2XLd5DBYDFQONyf794FgEzlTDRuRVVyg23+IDVFlPyeAkGtQ1vTmcK 9jy8vJ1pt2EhY4qTZm+17Nbny9oJ5AtXrzPdSG/LkLttfXCQ1pSzRz420mBlDP4XJI00 q+nA== X-Gm-Message-State: AOAM53337AiACC2HapYrJcDYZvOjqDbzotO8+zO4/eQEL3HRhBcf4qNt P+/h9wQJmUgBBPQlPtsAiRD5Mg== X-Google-Smtp-Source: ABdhPJw4cKt3Db+ftgGUvUsCsC2OriBA37IwTLaFslltoU09pBjVzIkRH0sSMvEgT8cgWlo5mlCqoA== X-Received: by 2002:a05:6000:18ad:: with SMTP id b13mr2698625wri.195.1639749694058; Fri, 17 Dec 2021 06:01:34 -0800 (PST) Received: from aivenronan.localnet (static-176-158-121-96.ftth.abo.bbox.fr. [176.158.121.96]) by smtp.gmail.com with ESMTPSA id g3sm4815648wrd.112.2021.12.17.06.01.33 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 17 Dec 2021 06:01:33 -0800 (PST) From: Ronan Dunklau To: David Rowley , pgsql-hackers@lists.postgresql.org, Tomas Vondra Cc: Andres Freund , Tomas Vondra Subject: Re: Use generation context to speed up tuplesorts Date: Fri, 17 Dec 2021 15:00:25 +0100 Message-ID: <4776839.iZASKD2KPV@aivenronan> Organization: aiven In-Reply-To: References: <7285172.GXAFRqVoOG@aivenronan> MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset="iso-8859-1" List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk Le vendredi 17 d=E9cembre 2021, 14:39:10 CET Tomas Vondra a =E9crit : > I wasn't really suggesting to investigate those other allocators in this > patch - it seems like a task requiring a pretty significant amount of > work/time. My point was that we should make it reasonably easy to add > tweaks for those other environments, if someone is interested enough to > do the legwork. >=20 > >> 2) In fact, I wonder if different glibc versions behave differently? > >> Hopefully it's not changing that much, though. Ditto kernel versions, > >> but the mmap/sbrk interface is likely more stable. We can test this. > >=20 > > That could be tested, yes. As a matter of fact, a commit removing the > > upper > > limit for MALLOC_MMAP_THRESHOLD has just been committed yesterday to > > glibc, > > which means we can service much bigger allocation without mmap. >=20 > Yeah, I noticed that commit too. Most systems stick to one glibc > version, so it'll take time to reach most systems. Let's continue with > just one glibc version and then maybe test other versions. Yes, I also need to figure out how to detect we're using glibc as I'm not v= ery=20 familiar with configure.=20 >=20 > >> 3) If we bump the thresholds, won't that work against reusing the > >> memory? I mean, if we free a whole block (from any allocator we have), > >> glibc might return it to kernel, depending on mmap threshold value. It= 's > >> not guaranteed, but increasing the malloc thresholds will make that ev= en > >> less likely. So we might just as well increase the minimum block size, > >> with about the same effect, no? > >=20 > > It is my understanding that malloc will try to compact memory by moving= it > > around. So the memory should be actually be released to the kernel at s= ome > > point. In the meantime, malloc can reuse it for our next invocation (wh= ich > > can be in a different memory context on our side). > >=20 > > If we increase the minimum block size, this is memory we will actually > >=20 > > reserve, and it will not protect us against the ramping-up behaviour: > > - the first allocation of a big block may be over mmap_threshold, and > > serviced>=20 > > by an expensive mmap > >=20 > > - when it's free, the threshold is doubled > > - next invocation is serviced by an sbrk call > > - freeing it will be above the trim threshold, and it will be returned. > >=20 > > After several "big" allocations, the thresholds will raise to their > > maximum > > values (well, it used to, I need to check what happens with that latest > > patch of glibc...) > >=20 > > This will typically happen several times as malloc doubles the threshold > > each time. This is probably the reason quadrupling the block sizes was > > more effective. >=20 > Hmmm, OK. Can we we benchmark the case with large initial block size, at > least for comparison? The benchmark I called "fixed" was with a fixed block size of=20 ALLOCSET_DEFAULT_MAXSIZE (first proposed patch) and showed roughly the same= =20 performance profile as the growing blocks + malloc tuning. But if I understand correctly, you implemented the growing blocks logic af= ter=20 concerns about wasting memory with a constant large block size. If we tune= =20 malloc, that memory would not be wasted if we don't alloc it, just not=20 released as eagerly when it's allocated.=20 Or do you want a benchmark with an even bigger initial block size ? With th= e=20 growing blocks patch with a large initial size ? I can run either, I just w= ant=20 to understand what is interesting to you. One thing that would be interesting would be to trace the total amount of=20 memory allocated in the different cases. This is something I will need to d= o=20 anyway when I propose that patch;=20 Best regards, =2D-=20 Ronan Dunklau