Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1w7bYD-005WFL-1J for pgsql-hackers@arkaria.postgresql.org; Tue, 31 Mar 2026 16:02:46 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.96) (envelope-from ) id 1w7bYA-00B8WR-2z for pgsql-hackers@arkaria.postgresql.org; Tue, 31 Mar 2026 16:02:43 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1w7bYA-00B8WH-1a for pgsql-hackers@lists.postgresql.org; Tue, 31 Mar 2026 16:02:43 +0000 Received: from mail.postgrespro.ru ([93.174.132.70]) by makus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.98.2) (envelope-from ) id 1w7bY7-00000001zZS-3KlJ for pgsql-hackers@postgresql.org; Tue, 31 Mar 2026 16:02:41 +0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=postgrespro.ru; s=mx2023; t=1774972956; bh=6NRxYRZ6ogrBEVk4Jz5yHhPWgM1cq0vxgxaWkSUQyVg=; h=Message-ID:Date:User-Agent:Subject:To:Cc:References:From: In-Reply-To:From; b=Ko7p/QP8+hNBb80NMG/r/rOxjso7JYn9j9wuuIoPU6w9XpzjG+zL6vZMbvMY/9Kpr MNKBWhv1UF+7uVB66OMX8uKTh4Fk+eSBPVN00RmgXfZww2yoN5zYp/1kxMUrTQieS7 9G25dpPrZoFeZNVF7SR3bymoOVflIjYgOF+y1X+5dxmyRFI4UBopNEbsewsntrLBFw rXCJf/gZVyimHJkfB7smlA3RbgBaEL/Go45Pfw2QJKWmI35FCPfIvfIWV5hYQ3gR7e QCJZcEiJ5uE9Wv3yR+hdAbMprMIEfg6Mp5bxonk5NEVIYlSBnxDr1e6hjTrKtxfDnI FtmAe/44tOsCQ== Received: from [192.168.3.127] (broadband-5-228-113-160.ip.moscow.rt.ru [5.228.113.160]) (using TLSv1.3 with cipher TLS_AES_128_GCM_SHA256 (128/128 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (Client did not present a certificate) (Authenticated sender: y.sokolov@postgrespro.ru) by mail.postgrespro.ru (Postfix/465) with ESMTPSA id B306F5FFD4; Tue, 31 Mar 2026 19:02:35 +0300 (MSK) Message-ID: <5bf667f3-5270-4b19-a08f-0facbecdff68@postgrespro.ru> Date: Tue, 31 Mar 2026 19:02:33 +0300 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: Buffer locking is special (hints, checksums, AIO writes) To: Andres Freund , Melanie Plageman Cc: Noah Misch , Heikki Linnakangas , Kirill Reshke , Matthias van de Meent , pgsql-hackers@postgresql.org, Thomas Munro , Robert Haas , Michael Paquier References: <5ubipyssiju5twkb7zgqwdr7q2vhpkpmuelxfpanetlk6ofnop@hvxb4g2amb2d> <68e89de8-5f6c-4eaf-a800-e16a5e487667@iki.fi> <20260215195239.ce.noahmisch@microsoft.com> Content-Language: en-US From: Yura Sokolov In-Reply-To: Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-KSMG-AntiPhishing: NotDetected, bases: 2026/03/31 15:37:00 X-KSMG-AntiSpam-Interceptor-Info: not scanned X-KSMG-AntiSpam-Status: not scanned, disabled by settings X-KSMG-AntiVirus: Kaspersky Secure Mail Gateway, version 2.1.0.7854, bases: 2026/03/31 13:30:00 #28358221 X-KSMG-AntiVirus-Status: NotDetected, skipped X-KSMG-LinksScanning: not scanned, disabled by settings X-KSMG-Message-Action: skipped X-KSMG-Rule-ID: 1 List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk 27.03.2026 23:00, Andres Freund wrote: > Hi, > > On 2026-03-25 18:35:55 -0400, Andres Freund wrote: >> Running it through valgrind and then will work on reading through one more >> time and pushing them. > > And done. > > Phew, this project took way longer than I'd though it'd take. In addition to bug with BM_IO_ERROR [1] , I found race condition in PinBuffer in this lines of code: if (unlikely(skip_if_not_valid && !(old_buf_state & BM_VALID))) return false; /* * We're not allowed to increase the refcount while the buffer * header spinlock is held. Wait for the lock to be released. */ if (old_buf_state & BM_LOCKED) old_buf_state = WaitBufHdrUnlocked(buf); While we waited for buffer header for being unlocked, it may become invalid, isn't it? Therefore, check related to skip_if_not_valid have to happen after waiting. .... Another question: previously we had to wait for buffer for being unlocked because UnlockBufHdr wrote to buf->state unconditionally, therefore our pin increment could be lost. Now UnlockBufHdr and UnlockBufHdrExt does proper atomic operations and preserves concurrent changes. Are we still need to wait? Most of time PinBuffer is called protected by BufTable's partition LWLock, therefore buffer may not be changed in dramatic way. But call in ReadRecentBuffer is the exception. It is not protected by partition lock and have to make additional checks. That is why you introduced skip_if_not_valid. Does optimization of ReadRecentBuffer pays for WaitBufHdrUnlocked? [1] https://www.postgresql.org/message-id/57a8ea70-32a8-4596-bb68-5d2990b83380%40postgrespro.ru -- regards Yura Sokolov aka funny-falcon