agora inbox for pgsql-admin@postgresql.org  
help / color / mirror / Atom feed
Replication is stuck
5+ messages / 3 participants
[nested] [flat]

* Replication is stuck
@ 2024-06-23 12:01  Murthy Nunna <mnunna@fnal.gov>
  0 siblings, 1 reply; 5+ messages in thread

From: Murthy Nunna @ 2024-06-23 12:01 UTC (permalink / raw)
  To: pgsql-admin

I am running pg14.4. I use WAL replication in a stand-by server which is 7-days behind primary (recovery_min_apply_delay = 7d)

My replication is stuck. It looks like it is repeatedly applying same WAL file. The next WAL file(s) are very much there.

I restarted cluster but it didn't fix the issue.

I appreciate any help you can provide before I rebuild the stand-by. I am trying to find the root cause. If 0000000100013D94000000FF is corrupted how can we tell?

2024-06-23 06:54:57 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:02 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:07 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:12 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:17 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:22 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:27 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:32 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:37 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:42 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive


There are no missing WALs:

ls -ltr 0000000100013D95000000* |more
-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000000
-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000001
-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000002
-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000003
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000004
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000005
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000006
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000007
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000008
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000009
-rw------- 1 postgres postgres 16777216 Jun 14 19:41 0000000100013D950000000A
-rw------- 1 postgres postgres 16777216 Jun 14 19:41 0000000100013D950000000B

^ permalink  raw  reply  [nested|flat] 5+ messages in thread

* Re: Replication is stuck
@ 2024-06-23 12:15  Ninad Shah <ninad.shah@percona.com>
  parent: Murthy Nunna <mnunna@fnal.gov>
  0 siblings, 1 reply; 5+ messages in thread

From: Ninad Shah @ 2024-06-23 12:15 UTC (permalink / raw)
  To: Murthy Nunna <mnunna@fnal.gov>; +Cc: pgsql-admin

Hi Murthy,

Would you please generate a pg_waldump of
0000000100013D94000000FF, 0000000100013D94000000FE
and 0000000100013D9500000000?

Thanks,

--

<https://www.percona.com/;

Ninad Shah
PostgreSQL DBA I, Managed Services

e: ninad.shah@percona.com

 w: www.percona.com

Databases Run Better With Percona


On Sun, Jun 23, 2024 at 5:32 PM Murthy Nunna <mnunna@fnal.gov> wrote:

> I am running pg14.4. I use WAL replication in a stand-by server which is
> 7-days behind primary (recovery_min_apply_delay = 7d)
>
>
>
> My replication is stuck. It looks like it is repeatedly applying same WAL
> file. The next WAL file(s) are very much there.
>
>
>
> I restarted cluster but it didn’t fix the issue.
>
>
>
> I appreciate any help you can provide before I rebuild the stand-by. I am
> trying to find the root cause. If 0000000100013D94000000FF is corrupted how
> can we tell?
>
>
>
> 2024-06-23 06:54:57 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:02 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:07 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:12 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:17 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:22 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:27 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:32 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:37 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:42 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
>
>
>
>
> There are no missing WALs:
>
>
>
> ls -ltr 0000000100013D95000000* |more
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:39
> 0000000100013D9500000000
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:39
> 0000000100013D9500000001
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:39
> 0000000100013D9500000002
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:39
> 0000000100013D9500000003
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000004
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000005
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000006
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000007
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000008
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000009
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:41
> 0000000100013D950000000A
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:41
> 0000000100013D950000000B
>
>
>
>
>
>
>

^ permalink  raw  reply  [nested|flat] 5+ messages in thread

* RE: Replication is stuck
@ 2024-06-23 12:34  Murthy Nunna <mnunna@fnal.gov>
  parent: Ninad Shah <ninad.shah@percona.com>
  0 siblings, 1 reply; 5+ messages in thread

From: Murthy Nunna @ 2024-06-23 12:34 UTC (permalink / raw)
  To: Ninad Shah <ninad.shah@percona.com>; +Cc: pgsql-admin

Thanks, Ninad. Looks like there is some error in 0000000100013D94000000FF. Any way to tell if this is logical corruption or physical corruption. In other words if this is file system corruption or of postgres generated corrupted file?

pg_waldump -q 0000000100013D94000000FE
[no errors]

pg_waldump -q 0000000100013D94000000FF
pg_waldump: fatal: error in WAL record at 13D94/FFBFFF48: invalid magic number 0000 in log segment 0000000100013D94000000FF, offset 12582912

pg_waldump -q 0000000100013D9500000000
[no errors]


From: Ninad Shah <ninad.shah@percona.com>
Sent: Sunday, June 23, 2024 7:16 AM
To: Murthy Nunna <mnunna@fnal.gov>
Cc: pgsql-admin@postgresql.org
Subject: Re: Replication is stuck


[EXTERNAL] – This message is from an external sender
Hi Murthy,

Would you please generate a pg_waldump of 0000000100013D94000000FF, 0000000100013D94000000FE and 0000000100013D9500000000?

Thanks,
--

[https://lh4.googleusercontent.com/zZlmF0li-Mz9lyOHHJmHhlsc-kXsksD5zSd8_0s-C-tRJ2EE5bHh_Ens51AYkdKyVi...]<https://urldefense.proofpoint.com/v2/url?u=https-3A__www.percona.com_&d=DwMFaQ&c=gRgGjJ3BkIs...;

Ninad Shah
PostgreSQL DBA I, Managed Services

e: ninad.shah@percona.com<mailto:ninad.shah@percona.com>

 w: www.percona.com<https://urldefense.proofpoint.com/v2/url?u=http-3A__www.percona.com_&d=DwMFaQ&c=gRgGjJ3BkIsb...;

Databases Run Better With Percona


On Sun, Jun 23, 2024 at 5:32 PM Murthy Nunna <mnunna@fnal.gov<mailto:mnunna@fnal.gov>> wrote:
I am running pg14.4. I use WAL replication in a stand-by server which is 7-days behind primary (recovery_min_apply_delay = 7d)

My replication is stuck. It looks like it is repeatedly applying same WAL file. The next WAL file(s) are very much there.

I restarted cluster but it didn’t fix the issue.

I appreciate any help you can provide before I rebuild the stand-by. I am trying to find the root cause. If 0000000100013D94000000FF is corrupted how can we tell?

2024-06-23 06:54:57 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:02 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:07 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:12 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:17 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:22 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:27 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:32 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:37 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive
2024-06-23 06:55:42 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive


There are no missing WALs:

ls -ltr 0000000100013D95000000* |more
-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000000
-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000001
-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000002
-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000003
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000004
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000005
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000006
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000007
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000008
-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000009
-rw------- 1 postgres postgres 16777216 Jun 14 19:41 0000000100013D950000000A
-rw------- 1 postgres postgres 16777216 Jun 14 19:41 0000000100013D950000000B





^ permalink  raw  reply  [nested|flat] 5+ messages in thread

* Re: Replication is stuck
@ 2024-06-23 12:38  Ninad Shah <ninad.shah@percona.com>
  parent: Murthy Nunna <mnunna@fnal.gov>
  0 siblings, 1 reply; 5+ messages in thread

From: Ninad Shah @ 2024-06-23 12:38 UTC (permalink / raw)
  To: Murthy Nunna <mnunna@fnal.gov>; +Cc: pgsql-admin

Your WAL file is corrupted. It's not possible to restore.


Thanks,

--

<https://www.percona.com/;

Ninad Shah
PostgreSQL DBA I, Managed Services

e: ninad.shah@percona.com

 w: www.percona.com

Databases Run Better With Percona


On Sun, Jun 23, 2024 at 6:04 PM Murthy Nunna <mnunna@fnal.gov> wrote:

> Thanks, Ninad. Looks like there is some error in 0000000100013D94000000FF.
> Any way to tell if this is logical corruption or physical corruption. In
> other words if this is file system corruption or of postgres generated
> corrupted file?
>
>
>
> pg_waldump -q 0000000100013D94000000FE
>
> [no errors]
>
>
>
> pg_waldump -q 0000000100013D94000000FF
>
> pg_waldump: fatal: error in WAL record at 13D94/FFBFFF48: invalid magic
> number 0000 in log segment 0000000100013D94000000FF, offset 12582912
>
>
>
> pg_waldump -q 0000000100013D9500000000
>
> [no errors]
>
>
>
>
>
> *From:* Ninad Shah <ninad.shah@percona.com>
> *Sent:* Sunday, June 23, 2024 7:16 AM
> *To:* Murthy Nunna <mnunna@fnal.gov>
> *Cc:* pgsql-admin@postgresql.org
> *Subject:* Re: Replication is stuck
>
>
>
> [EXTERNAL] – This message is from an external sender
>
> Hi Murthy,
>
>
>
> Would you please generate a pg_waldump of
> 0000000100013D94000000FF, 0000000100013D94000000FE
> and 0000000100013D9500000000?
>
>
> Thanks,
>
> --
>
>
> <https://url.avanan.click/v2/___https://urldefense.proofpoint.com/v2/url?u=https-3A__www.percona.com_...;
>
> *Ninad Shah*
> PostgreSQL DBA I, *Managed Services*
>
> *e:* ninad.shah@percona.com
>
>  *w:* www.percona.com
> <https://url.avanan.click/v2/___https://urldefense.proofpoint.com/v2/url?u=http-3A__www.percona.com_&...;
>
> *Databases Run Better With Percona*
>
>
>
>
>
> On Sun, Jun 23, 2024 at 5:32 PM Murthy Nunna <mnunna@fnal.gov> wrote:
>
> I am running pg14.4. I use WAL replication in a stand-by server which is
> 7-days behind primary (recovery_min_apply_delay = 7d)
>
>
>
> My replication is stuck. It looks like it is repeatedly applying same WAL
> file. The next WAL file(s) are very much there.
>
>
>
> I restarted cluster but it didn’t fix the issue.
>
>
>
> I appreciate any help you can provide before I rebuild the stand-by. I am
> trying to find the root cause. If 0000000100013D94000000FF is corrupted how
> can we tell?
>
>
>
> 2024-06-23 06:54:57 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:02 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:07 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:12 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:17 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:22 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:27 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:32 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:37 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
> 2024-06-23 06:55:42 CDT []LOG:  restored log file
> "0000000100013D94000000FF" from archive
>
>
>
>
>
> There are no missing WALs:
>
>
>
> ls -ltr 0000000100013D95000000* |more
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:39
> 0000000100013D9500000000
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:39
> 0000000100013D9500000001
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:39
> 0000000100013D9500000002
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:39
> 0000000100013D9500000003
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000004
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000005
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000006
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000007
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000008
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:40
> 0000000100013D9500000009
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:41
> 0000000100013D950000000A
>
> -rw------- 1 postgres postgres 16777216 Jun 14 19:41
> 0000000100013D950000000B
>
>
>
>
>
>
>
>

^ permalink  raw  reply  [nested|flat] 5+ messages in thread

* RE: Replication is stuck
@ 2024-06-23 12:40  lennam@incisivetechgroup.com
  parent: Ninad Shah <ninad.shah@percona.com>
  0 siblings, 0 replies; 5+ messages in thread

From: lennam@incisivetechgroup.com @ 2024-06-23 12:40 UTC (permalink / raw)
  To: 'Ninad Shah' <ninad.shah@percona.com>; 'Murthy Nunna' <mnunna@fnal.gov>; +Cc: pgsql-admin

If WAL lag is  more than 7 days , rebuild the the replication only solution 

 

From: Ninad Shah <ninad.shah@percona.com> 
Sent: Sunday, June 23, 2024 8:38 AM
To: Murthy Nunna <mnunna@fnal.gov>
Cc: pgsql-admin@postgresql.org
Subject: Re: Replication is stuck

 

Your WAL file is corrupted. It's not possible to restore.

 




Thanks,

--


 <https://www.percona.com/; 

Ninad Shah
PostgreSQL DBA I, Managed Services

e: ninad.shah@percona.com <mailto:ninad.shah@percona.com> 

 w:  <http://www.percona.com/; www.percona.com


Databases Run Better With Percona

	

 

 

On Sun, Jun 23, 2024 at 6:04 PM Murthy Nunna <mnunna@fnal.gov <mailto:mnunna@fnal.gov> > wrote:

Thanks, Ninad. Looks like there is some error in 0000000100013D94000000FF. Any way to tell if this is logical corruption or physical corruption. In other words if this is file system corruption or of postgres generated corrupted file?

 

pg_waldump -q 0000000100013D94000000FE

[no errors]

 

pg_waldump -q 0000000100013D94000000FF

pg_waldump: fatal: error in WAL record at 13D94/FFBFFF48: invalid magic number 0000 in log segment 0000000100013D94000000FF, offset 12582912

 

pg_waldump -q 0000000100013D9500000000

[no errors]

 

 

From: Ninad Shah <ninad.shah@percona.com <mailto:ninad.shah@percona.com> > 
Sent: Sunday, June 23, 2024 7:16 AM
To: Murthy Nunna <mnunna@fnal.gov <mailto:mnunna@fnal.gov> >
Cc: pgsql-admin@postgresql.org <mailto:pgsql-admin@postgresql.org> 
Subject: Re: Replication is stuck

 

[EXTERNAL] – This message is from an external sender

Hi Murthy, 

 

Would you please generate a pg_waldump of 0000000100013D94000000FF, 0000000100013D94000000FE and 0000000100013D9500000000?




Thanks,

--


 <https://url.avanan.click/v2/___https:/urldefense.proofpoint.com/v2/url?u=https-3A__www.percona.com_&...; 

Ninad Shah
PostgreSQL DBA I, Managed Services

e: ninad.shah@percona.com <mailto:ninad.shah@percona.com> 

 w:  <https://url.avanan.click/v2/___https:/urldefense.proofpoint.com/v2/url?u=http-3A__www.percona.com_&a...; www.percona.com


Databases Run Better With Percona

	

 

 

On Sun, Jun 23, 2024 at 5:32 PM Murthy Nunna <mnunna@fnal.gov <mailto:mnunna@fnal.gov> > wrote:

I am running pg14.4. I use WAL replication in a stand-by server which is 7-days behind primary (recovery_min_apply_delay = 7d)

 

My replication is stuck. It looks like it is repeatedly applying same WAL file. The next WAL file(s) are very much there.

 

I restarted cluster but it didn’t fix the issue.

 

I appreciate any help you can provide before I rebuild the stand-by. I am trying to find the root cause. If 0000000100013D94000000FF is corrupted how can we tell?

 

2024-06-23 06:54:57 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive

2024-06-23 06:55:02 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive

2024-06-23 06:55:07 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive

2024-06-23 06:55:12 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive

2024-06-23 06:55:17 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive

2024-06-23 06:55:22 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive

2024-06-23 06:55:27 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive

2024-06-23 06:55:32 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive

2024-06-23 06:55:37 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive

2024-06-23 06:55:42 CDT []LOG:  restored log file "0000000100013D94000000FF" from archive

 

 

There are no missing WALs:

 

ls -ltr 0000000100013D95000000* |more

-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000000

-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000001

-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000002

-rw------- 1 postgres postgres 16777216 Jun 14 19:39 0000000100013D9500000003

-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000004

-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000005

-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000006

-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000007

-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000008

-rw------- 1 postgres postgres 16777216 Jun 14 19:40 0000000100013D9500000009

-rw------- 1 postgres postgres 16777216 Jun 14 19:41 0000000100013D950000000A

-rw------- 1 postgres postgres 16777216 Jun 14 19:41 0000000100013D950000000B

 

 

 

^ permalink  raw  reply  [nested|flat] 5+ messages in thread


end of thread, other threads:[~2024-06-23 12:40 UTC | newest]

Thread overview: 5+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2024-06-23 12:01 Replication is stuck Murthy Nunna <mnunna@fnal.gov>
2024-06-23 12:15 ` Ninad Shah <ninad.shah@percona.com>
2024-06-23 12:34   ` Murthy Nunna <mnunna@fnal.gov>
2024-06-23 12:38     ` Ninad Shah <ninad.shah@percona.com>
2024-06-23 12:40       ` lennam@incisivetechgroup.com

This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox