pg.ddx.io  pgsql-hackers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Alvaro Herrera <alvherre@alvh.no-ip.org>
To: Robert Haas <robertmhaas@gmail.com>
Cc: Bossart, Nathan <bossartn@amazon.com>
Cc: Dipesh Pandit <dipesh.pandit@gmail.com>
Cc: Kyotaro Horiguchi <horikyota.ntt@gmail.com>
Cc: Jeevan Ladhe <jeevan.ladhe@enterprisedb.com>
Cc: Stephen Frost <sfrost@snowman.net>
Cc: Andres Freund <andres@anarazel.de>
Cc: Hannu Krosing <hannuk@google.com>
Cc: pgsql-hackers@postgresql.org <pgsql-hackers@postgresql.org>
Subject: Re: .ready and .done files considered harmful
Date: Mon, 20 Sep 2021 17:42:26 -0300
Message-ID: <202109202042.uziogfq245yw@alvherre.pgsql> (raw)
In-Reply-To: <CA+Tgmoa+Z8eudDLOn-+EFMAts3uMKZE=LRNiS6VbTqGUrwOgzA@mail.gmail.com>

On 2021-Sep-20, Robert Haas wrote:

> I was thinking that this might increase the number of directory scans
> by a pretty large amount when we repeatedly catch up, then 1 new file
> gets added, then we catch up, etc.

I was going to say that perhaps we can avoid repeated scans by having a
bitmap of future files that were found by a scan; so if we need to do
one scan, we keep track of the presence of the next (say) 64 files in
our timeline, and then we only have to do another scan when we need to
archive a file that wasn't present the last time we scanned.  However:

> But I guess your thought process is that such directory scans, even if
> they happen many times per second, can't really be that expensive,
> since the directory can't have much in it. Which seems like a fair
> point. I wonder if there are any situations in which there's not much
> to archive but the archive_status directory still contains tons of
> files.

(If we take this stance, which seems reasonable to me, then we don't
need to optimize.)  But perhaps we should complain if we find extraneous
files in archive_status -- Then it'd be on the users' heads not to leave
tons of files that would slow down the scan.

-- 
Álvaro Herrera           39°49'30"S 73°17'W  —  https://www.EnterpriseDB.com/
Maybe there's lots of data loss but the records of data loss are also lost.
(Lincoln Yeoh)





view thread (118+ messages)  latest in thread

Message-ID: <202109202042.uziogfq245yw@alvherre.pgsql>
Permalink:  ../202109202042.uziogfq245yw@alvherre.pgsql/
Also on:    postgresql.org/message-id/202109202042.uziogfq245yw@alvherre.pgsql

 · 

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: alvherre@alvh.no-ip.org, robertmhaas@gmail.com, bossartn@amazon.com, dipesh.pandit@gmail.com, horikyota.ntt@gmail.com, jeevan.ladhe@enterprisedb.com, sfrost@snowman.net, andres@anarazel.de, hannuk@google.com
  Subject: Re: .ready and .done files considered harmful
  In-Reply-To: <202109202042.uziogfq245yw@alvherre.pgsql>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox