Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_CBC_SHA1:256) (Exim 4.89) (envelope-from ) id 1hVOes-0002oX-Mt for pgsql-hackers@arkaria.postgresql.org; Mon, 27 May 2019 23:03:59 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.89) (envelope-from ) id 1hVOdr-0000gJ-NT for pgsql-hackers@arkaria.postgresql.org; Mon, 27 May 2019 23:02:55 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_CBC_SHA1:256) (Exim 4.89) (envelope-from ) id 1hVOdr-0000fh-7o for pgsql-hackers@lists.postgresql.org; Mon, 27 May 2019 23:02:55 +0000 Received: from mail-pg1-x542.google.com ([2607:f8b0:4864:20::542]) by magus.postgresql.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_CBC_SHA1:256) (Exim 4.89) (envelope-from ) id 1hVOdh-0002fc-M7 for pgsql-hackers@postgresql.org; Mon, 27 May 2019 23:02:52 +0000 Received: by mail-pg1-x542.google.com with SMTP id w34so5074409pga.12 for ; Mon, 27 May 2019 16:02:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=leadboat.com; s=google; h=date:from:to:cc:subject:message-id:references:mime-version :content-disposition:in-reply-to:user-agent; bh=2BVR/ozKfTEZs2i8XhHz4d1v/YQMg06hMv/LIWVDTOo=; b=eMGLhXDu6V6uunUuzYhcKkZ6Q3U4SNHx7c4aPJbGN186q6+4jZJljQ68cyAuKSEgO9 Ebc8JC7NRwSJW4O3kSBNDDhG5XWIl/FAhvmzIWMEx76rHlCPPcxqhDyr2Vn5mddla2bJ /kPZL3sQo8ohXkNSn5Gem9y0rbvoVqG4xvLvI= X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:date:from:to:cc:subject:message-id:references :mime-version:content-disposition:in-reply-to:user-agent; bh=2BVR/ozKfTEZs2i8XhHz4d1v/YQMg06hMv/LIWVDTOo=; b=nAvMXWg6pUw8odzC0JR797eaKeFc/ZUNq8+S484FY1WDiHFDc2SJYHnbyW56Fz45uI bswm2xSWRfUhY9ezJytYawhrzZZ9IA69cGq4dLQ20T/zmhoEa7iEDEKp3jJS12I3vV8v YNimRxlK6BD8FhT9w4rODfVDxUj7+/IuLxBpq29GaWgSBOS+gw9FiTnsVkSG4vnJusb0 jBwhZsZSy+cdVaOiI6dWR3ue5Bq2OyKhVa9NAlmifNnlthyOwRv7PXN9mlodFBdPx92a VAXKQeTpGmobevMT8P0qDdt1v/V/j/VyTFbGFP0UBewpAD02+DBCaIRoL/OS2M1JtHGG mAIw== X-Gm-Message-State: APjAAAWbJb/HaYAF8LVxQyaJt/PpwTDfCnE8KJxsIlmv6pMrBCtPuKUK gOPTSZ2nLaA3CBYWrnfXZlODfw== X-Google-Smtp-Source: APXvYqzeuEReOSKA2oQJqXLFu05ohm5ocSwQ8mgucGBt/VvJahL64bZ5CPGx4VrpE9S4JbkF+bC4Vg== X-Received: by 2002:a63:4f1c:: with SMTP id d28mr27984930pgb.353.1558998162220; Mon, 27 May 2019 16:02:42 -0700 (PDT) Received: from gust.leadboat.com ([208.54.39.214]) by smtp.gmail.com with ESMTPSA id e14sm12033927pff.60.2019.05.27.16.02.38 (version=TLS1_2 cipher=ECDHE-RSA-AES128-GCM-SHA256 bits=128/128); Mon, 27 May 2019 16:02:41 -0700 (PDT) Date: Mon, 27 May 2019 19:02:25 -0400 From: Noah Misch To: Kyotaro HORIGUCHI Cc: pgsql-hackers@postgresql.org, 9erthalion6@gmail.com, andrew.dunstan@2ndquadrant.com, hlinnaka@iki.fi, robertmhaas@gmail.com, michael@paquier.xyz Subject: Re: [HACKERS] WAL logging problem in 9.4.3? Message-ID: <20190527230225.GA59385@gust.leadboat.com> References: <20190520.155430.215084510.horiguchi.kyotaro@lab.ntt.co.jp> <20190521.212948.34357392.horiguchi.kyotaro@lab.ntt.co.jp> <20190525023332.GE1624191@rfd.leadboat.com> <20190527.140826.258215605.horiguchi.kyotaro@lab.ntt.co.jp> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20190527.140826.258215605.horiguchi.kyotaro@lab.ntt.co.jp> User-Agent: Mutt/1.5.24 (2015-08-30) List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Precedence: bulk On Mon, May 27, 2019 at 02:08:26PM +0900, Kyotaro HORIGUCHI wrote: > At Fri, 24 May 2019 19:33:32 -0700, Noah Misch wrote in <20190525023332.GE1624191@rfd.leadboat.com> > > On Mon, May 20, 2019 at 03:54:30PM +0900, Kyotaro HORIGUCHI wrote: > > > Following this direction, the attached PoC works *at least for* > > > the wal_optimization TAP tests, but doing pending flush not in > > > smgr but in relcache. > > > > This task, syncing files created in the current transaction, is not the kind > > of task normally assigned to a cache. We already have a module, storage.c, > > that maintains state about files created in the current transaction. Why did > > you use relcache instead of storage.c? > > The reason was at-commit sync needs buffer flush beforehand. But > FlushRelationBufferWithoutRelCache() in v11 can do > that. storage.c is reasonable as the place. Okay. I do want this to work in 9.5 and later, but I'm not aware of a reason relcache.c would be a better code location in older branches. Unless you think of a reason to prefer relcache.c, please use storage.c. > > On Tue, May 21, 2019 at 09:29:48PM +0900, Kyotaro HORIGUCHI wrote: > > > This is a tidier version of the patch. > > > > > - Move the substantial work to table/index AMs. > > > > > > Each AM can decide whether to support WAL skip or not. > > > Currently heap and nbtree support it. > > > > Why would an AM find it important to disable WAL skip? > > The reason is currently it's AM's responsibility to decide > whether to skip WAL or not. I see. Skipping the sync would be a mere optimization; no AM would require it for correctness. An AM might want RelationNeedsWAL() to keep returning true despite the sync happening, perhaps because it persists data somewhere other than the forks of pg_class.relfilenode. Since the index and table APIs already assume one relfilenode captures all persistent data, I'm not seeing a use case for an AM overriding this behavior. Let's take away the AM's responsibility for this decision, making the system simpler. A future patch could let AM code decide, if someone find a real-world use case for AM-specific logic around when to skip WAL.