Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1oVgZF-0000v9-W9 for pgsql-hackers@arkaria.postgresql.org; Tue, 06 Sep 2022 21:57:14 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.92) (envelope-from ) id 1oVgZD-0008Cz-Og for pgsql-hackers@arkaria.postgresql.org; Tue, 06 Sep 2022 21:57:11 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1oVgZD-0008Cp-FK for pgsql-hackers@lists.postgresql.org; Tue, 06 Sep 2022 21:57:11 +0000 Received: from mail-pg1-x535.google.com ([2607:f8b0:4864:20::535]) by makus.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.92) (envelope-from ) id 1oVgZA-0004AG-OF for pgsql-hackers@lists.postgresql.org; Tue, 06 Sep 2022 21:57:10 +0000 Received: by mail-pg1-x535.google.com with SMTP id t65so871806pgt.2 for ; Tue, 06 Sep 2022 14:57:08 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20210112; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date; bh=mW4w17WdcGWgJyCgHAZSG+3LLoHq4gr4zSS0iucROL8=; b=RAISUHjTbMu3r0wl+N9rHvPZCTvniDAA7z5Drg8r6Pn8QNcGq/GQxxJda1fORWzg5H aqCojSmucHXI5Cbxsb7XsrZx4jj+lyVx+pSwohW9P8QrVBtAgJfA08JcTGW0eD3CksvC VSn11D+FvS7vlMdmqJRT2HPBIeBlXyzVt3QU5ZTXmWg3K89LK0TmvjO68lq3AID1mGKo iZnxoI1fhjUXaK0/jzIxF9A1X+IXgiKcBh9oolw3RJSvWP7qzsqGpzB4PGa0ZROznEFg lvbv/6sp3svLlU7Tu3towcqWmL80C+90Fv8c2OuN3OXaa+iV0O7NApVDaagI5ip/3JV7 8fFA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date; bh=mW4w17WdcGWgJyCgHAZSG+3LLoHq4gr4zSS0iucROL8=; b=ydQXQXZGe57CVcp7fpOvuvQK9j5N6bIYKEjktAqQMDspauLGd0hChmsTmCcUjIAHG4 SX5RW0O50YORBKCVW+TAxSre61YQ1dO1mm6ijlaLvVOyja8W5ZdYwsUZs87NXj66ZC8L X1pNKb6M0cjNaOrSUhCNQHjvarwVDg3dMKeqJz2d5Be2KSuIoh4QBhNSgdQOolCUGprH BRsNdXbj6emXBRds6aPtQHQd/jal5GG/A37d0UHbKekeNiplqIvdWZ87YdbmDY/vxyqk GEkageD2e3FWFrqq9jMJ7ZbJPrsjyfj2+1VRYUkEckdfwe95rZYw9HHhFHFsZhLd09Cd nYpg== X-Gm-Message-State: ACgBeo3am7tTbn9oYxz1X3MGsXsT6RxpV4dZJMsNIZOk8uu8dIwqaZK9 IppzfOAFatmBChsci9XPZzs= X-Google-Smtp-Source: AA6agR6RS0sTWN4AJZBcrSPv0A4262A0+JC4NBg6UDxQZ5qZTFDf2g/PzEGPRo3XEZuw1AV/1HDbuA== X-Received: by 2002:a65:6749:0:b0:434:1f8b:cb97 with SMTP id c9-20020a656749000000b004341f8bcb97mr614710pgu.360.1662501427752; Tue, 06 Sep 2022 14:57:07 -0700 (PDT) Received: from nathanxps13 ([50.47.162.83]) by smtp.gmail.com with ESMTPSA id a6-20020a170902900600b001769206a766sm7304928plp.307.2022.09.06.14.57.06 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 06 Sep 2022 14:57:06 -0700 (PDT) Date: Tue, 6 Sep 2022 14:57:04 -0700 From: Nathan Bossart To: Bharath Rupireddy Cc: Cary Huang , PostgreSQL Hackers , SATYANARAYANA NARLAPURAM Subject: Re: Switching XLog source from archive to streaming when primary available Message-ID: <20220906215704.GA2084086@nathanxps13> References: <165610083853.2517.17504309233817317001.pgcf@coridan.postgresql.org> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk + + wal_source_switch_interval configuration parameter + I don't want to bikeshed on the name too much, but I do think we need something more descriptive. I'm thinking of something like streaming_replication_attempt_interval or streaming_replication_retry_interval. + Specifies how long the standby server should wait before switching WAL + source from WAL archive to primary (streaming replication). This can + happen either during the standby initial recovery or after a previous + failed attempt to stream WAL from the primary. I'm not sure what the second sentence means. In general, I think the explanation in your commit message is much clearer: The standby makes an attempt to read WAL from primary after wal_retrieve_retry_interval milliseconds reading from archive. + If this value is specified without units, it is taken as milliseconds. + The default value is 5 seconds. A setting of 0 + disables the feature. 5 seconds seems low. I would expect the default to be 1-5 minutes. I think it's important to strike a balance between interrupting archive recovery to attempt streaming replication and letting archive recovery make progress. + * Try reading WAL from primary after every wal_source_switch_interval + * milliseconds, when state machine is in XLOG_FROM_ARCHIVE state. If + * successful, the state machine moves to XLOG_FROM_STREAM state, otherwise + * it falls back to XLOG_FROM_ARCHIVE state. It's not clear to me how this is expected to interact with the pg_wal phase of standby recovery. As the docs note [0], standby servers loop through archive recovery, recovery from pg_wal, and streaming replication. Does this cause the pg_wal phase to be skipped (i.e., the standby goes straight from archive recovery to streaming replication)? I wonder if it'd be better for this mechanism to simply move the standby to the pg_wal phase so that the usual ordering isn't changed. + if (!first_time && + TimestampDifferenceExceeds(last_switch_time, curr_time, + wal_source_switch_interval)) Shouldn't this also check that wal_source_switch_interval is not set to 0? + elog(DEBUG2, + "trying to switch WAL source to %s after fetching WAL from %s for %d milliseconds", + xlogSourceNames[XLOG_FROM_STREAM], + xlogSourceNames[currentSource], + wal_source_switch_interval); + + last_switch_time = curr_time; Shouldn't the last_switch_time be set when the state machine first enters XLOG_FROM_ARCHIVE? IIUC this logic is currently counting time spent elsewhere (e.g., XLOG_FROM_STREAM) when determining whether to force a source switch. This would mean that a standby that has spent a lot of time in streaming replication before failing would flip to XLOG_FROM_ARCHIVE, immediately flip back to XLOG_FROM_STREAM, and then likely flip back to XLOG_FROM_ARCHIVE when it failed again. Given the standby will wait for wal_retrieve_retry_interval before going back to XLOG_FROM_ARCHIVE, it seems like we could end up rapidly looping between sources. Perhaps I am misunderstanding how this is meant to work. + { + {"wal_source_switch_interval", PGC_SIGHUP, REPLICATION_STANDBY, + gettext_noop("Sets the time to wait before switching WAL source from archive to primary"), + gettext_noop("0 turns this feature off."), + GUC_UNIT_MS + }, + &wal_source_switch_interval, + 5000, 0, INT_MAX, + NULL, NULL, NULL + }, I wonder if the lower bound should be higher to avoid switching unnecessarily rapidly between WAL sources. I see that WaitForWALToBecomeAvailable() ensures that standbys do not switch from XLOG_FROM_STREAM to XLOG_FROM_ARCHIVE more often than once per wal_retrieve_retry_interval. Perhaps wal_retrieve_retry_interval should be the lower bound for this GUC, too. Or maybe WaitForWALToBecomeAvailable() should make sure that the standby makes at least once attempt to restore the file from archive before switching to streaming replication. [0] https://www.postgresql.org/docs/current/warm-standby.html#STANDBY-SERVER-OPERATION -- Nathan Bossart Amazon Web Services: https://aws.amazon.com