Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1mv49s-00025n-08 for pgsql-hackers@arkaria.postgresql.org; Wed, 08 Dec 2021 21:07:24 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.92) (envelope-from ) id 1mv49p-00071i-DU for pgsql-hackers@arkaria.postgresql.org; Wed, 08 Dec 2021 21:07:21 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1mv49o-00071X-Ol for pgsql-hackers@lists.postgresql.org; Wed, 08 Dec 2021 21:07:21 +0000 Received: from mail-ed1-x534.google.com ([2a00:1450:4864:20::534]) by makus.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.92) (envelope-from ) id 1mv49l-0008T3-NX for pgsql-hackers@lists.postgresql.org; Wed, 08 Dec 2021 21:07:19 +0000 Received: by mail-ed1-x534.google.com with SMTP id l25so12627472eda.11 for ; Wed, 08 Dec 2021 13:07:17 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=enterprisedb.com; s=google; h=message-id:date:mime-version:user-agent:subject:content-language:to :cc:references:from:in-reply-to:content-transfer-encoding; bh=2jifa+5kfqO6FRzV9+GRZxQGLoMqOgb9AlN8vmX35h4=; b=PFBsJIMhYi3zKmwa6mx19N12KiOLc56A25xHlK1/e5FYXTcVLfjLY/yjDOLtaB+k+d 0IFaTPpAivAQlrDYnHmAP+JJYFzzNj19l65hfDR5d7DaKcPCfl9k4G268ic4/HVVOgcv DNfcu3VwEoB94tR6BhhKF4j3+nEOFmOzMomjyj2NP/ShV4xx/FXHzkbqFxz3Enxj/b+q KenGkE1/QXfbXCmETsIBAzF2kZuZgMkJa1DR1NVIywqMqUmuTB8v2INId/Neerya0qTi TuT60cfnDG1HgkAvrEB98F5jQXIiSOohz6KfRBmdM6Nrjo1GUDFi2EfsXXuLh6xD9o13 ObRA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:message-id:date:mime-version:user-agent:subject :content-language:to:cc:references:from:in-reply-to :content-transfer-encoding; bh=2jifa+5kfqO6FRzV9+GRZxQGLoMqOgb9AlN8vmX35h4=; b=deITUE4g+i+nnlfJIrNowYEfPDOKZ4neWIeyiGzkHZZw3MreEpC/cMa1PbsW/WnYgM a2ut9IL/zcm8Mjn4UAOQkOOsQb+xjEZIDDh0vKY28n2r9Q2FtVTs7gjg3kXQRI1djrz+ PKGEmbV1ZORG4zAHOe7muvumk7sD+817/cMdBLGWZp0/qVtApNsYVh/4v3DLPUTblDUP 31AR7MJZXxP+pOta8yh/61MP44pPl+6uSm55gQBoM6pe8I4ZHkan1eXE8dzcIzXtSRhm auEJdQX2gBbZEBTChN8ZBUwHHer0tIXrm3H0OK/bNF5ozWN+QIaZGOxVWYfA4q7fVavB bP1g== X-Gm-Message-State: AOAM532ApLEwPF6Vnj26Z6fBtQfg4j4ZTKQdF8fAV1Slcr/Aseb4PmpA p1aVYBZag4V8l2rX+9KtT+Bd6v7JBLMj7tzlmBMIAqZwtlh8Z5MnEC7NM962cgE6xGTySZtP085 tauhIW27T+6ztDdhdI4D7DRjJjCOqScbOCgnhM6AFhgQsC5rm4GN7CYA5R21ahnuzQ8XznecJun XEsPx06lsHUiGEUH+bRcEM7SITVU5Nw1OrsdgSLb25Xz1TN1wqZglu X-Google-Smtp-Source: ABdhPJy+at90eIAmoo5RnOP/uib5O9LjCnc4i/dnJU5tMKELZq/fjlRdqqcxtPqNI4/0ijqNO8/bVw== X-Received: by 2002:a05:6402:42:: with SMTP id f2mr23293705edu.204.1638997635881; Wed, 08 Dec 2021 13:07:15 -0800 (PST) Received: from [10.137.0.19] (ip-86-49-251-241.net.upcbroadband.cz. [86.49.251.241]) by smtp.gmail.com with ESMTPSA id og38sm1901020ejc.5.2021.12.08.13.07.14 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Wed, 08 Dec 2021 13:07:15 -0800 (PST) Message-ID: <567d1ea7-bfeb-bd97-1a7f-b13d2770258c@enterprisedb.com> Date: Wed, 8 Dec 2021 22:07:12 +0100 MIME-Version: 1.0 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:91.0) Gecko/20100101 Thunderbird/91.2.0 Subject: Re: Use generation context to speed up tuplesorts Content-Language: en-US To: Ronan Dunklau , David Rowley , pgsql-hackers@lists.postgresql.org Cc: Andres Freund , Tomas Vondra References: <36829a8a-63d0-c428-89d5-07e49561973e@enterprisedb.com> <8046109.NyiUUSuA9g@aivenronan> From: Tomas Vondra In-Reply-To: <8046109.NyiUUSuA9g@aivenronan> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-CLOUD-SEC-AV-Info: enterprisedb,google_mail,monitor X-CLOUD-SEC-AV-Sent: true X-Gm-Spam: 0 X-Gm-Phishy: 0 List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk On 12/8/21 16:51, Ronan Dunklau wrote: > Le jeudi 9 septembre 2021, 15:37:59 CET Tomas Vondra a écrit : >> And now comes the funny part - if I run it in the same backend as the >> "full" benchmark, I get roughly the same results: >> >> block_size | chunk_size | mem_allocated | alloc_ms | free_ms >> ------------+------------+---------------+----------+--------- >> 32768 | 512 | 806256640 | 37159 | 76669 >> >> but if I reconnect and run it in the new backend, I get this: >> >> block_size | chunk_size | mem_allocated | alloc_ms | free_ms >> ------------+------------+---------------+----------+--------- >> 32768 | 512 | 806158336 | 233909 | 100785 >> (1 row) >> >> It does not matter if I wait a bit before running the query, if I run it >> repeatedly, etc. The machine is not doing anything else, the CPU is set >> to use "performance" governor, etc. > > I've reproduced the behaviour you mention. > I also noticed asm_exc_page_fault showing up in the perf report in that case. > > Running an strace on it shows that in one case, we have a lot of brk calls, > while when we run in the same process as the previous tests, we don't. > > My suspicion is that the previous workload makes glibc malloc change it's > trim_threshold and possibly other dynamic options, which leads to constantly > moving the brk pointer in one case and not the other. > > Running your fifo test with absurd malloc options shows that indeed that might > be the case (I needed to change several, because changing one disable the > dynamic adjustment for every single one of them, and malloc would fall back to > using mmap and freeing it on each iteration): > > mallopt(M_TOP_PAD, 1024 * 1024 * 1024); > mallopt(M_TRIM_THRESHOLD, 256 * 1024 * 1024); > mallopt(M_MMAP_THRESHOLD, 4*1024*1024*sizeof(long)); > > I get the following results for your self contained test. I ran the query > twice, in each case, seeing the importance of the first allocation and the > subsequent ones: > > With default malloc options: > > block_size | chunk_size | mem_allocated | alloc_ms | free_ms > ------------+------------+---------------+----------+--------- > 32768 | 512 | 795836416 | 300156 | 207557 > > block_size | chunk_size | mem_allocated | alloc_ms | free_ms > ------------+------------+---------------+----------+--------- > 32768 | 512 | 795836416 | 211942 | 77207 > > > With the oversized values above: > > block_size | chunk_size | mem_allocated | alloc_ms | free_ms > ------------+------------+---------------+----------+--------- > 32768 | 512 | 795836416 | 219000 | 36223 > > > block_size | chunk_size | mem_allocated | alloc_ms | free_ms > ------------+------------+---------------+----------+--------- > 32768 | 512 | 795836416 | 75761 | 78082 > (1 row) > > I can't tell how representative your benchmark extension would be of real life > allocation / free patterns, but there is probably something we can improve > here. > Thanks for looking at this. I think those allocation / free patterns are fairly extreme, and there probably are no workloads doing exactly this. The idea is the actual workloads are likely some combination of these extreme cases. > I'll try to see if I can understand more precisely what is happening. > Thanks, that'd be helpful. Maybe we can learn something about tuning malloc parameters to get significantly better performance. regards -- Tomas Vondra EnterpriseDB: http://www.enterprisedb.com The Enterprise PostgreSQL Company