pg.ddx.io  pgsql-admin@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Joe Conway <mail@joeconway.com>
To: Joseph Hammerman <joe.hammerman@datadoghq.com>
Cc: Tom Lane <tgl@sss.pgh.pa.us>
Cc: pgsql-admin@lists.postgresql.org
Subject: Re: Getting out ahead of OOM
Date: Wed, 12 Mar 2025 22:35:17 -0400
Message-ID: <e52f00fc-ffe3-4745-8082-b492f44c6693@joeconway.com> (raw)
In-Reply-To: <CAHs7QM_P-wqprxTmD5sjF9ShX2kLWGdR9gatPBtO+FfLqfWr6Q@mail.gmail.com>
References: <CAHs7QM_8SRfybxg_Pq3orGw5anMkV7OzKYfxDXSBQVWzcXf3Bw@mail.gmail.com>
	<1117736.1741375563@sss.pgh.pa.us>
	<efcca885-a954-43e2-99b8-8b993678c72c@joeconway.com>
	<CAHs7QM_P-wqprxTmD5sjF9ShX2kLWGdR9gatPBtO+FfLqfWr6Q@mail.gmail.com>

On 3/12/25 18:21, Joseph Hammerman wrote:
> Joe, can you expand on your recommendation to use cgroup-v2? We're 
> trying to collect our complete rationale for our request to our internal 
> team that is tasked with rolling out this configuration change.

cgroup-v2 has a much better measure of memory pressure (see PSI[1]), 
better ability to reclaim memory pages[4], safer delegation, and other 
advantages. When I last looked the kube support for it was still brand 
new, but it appears to be well supported now [2][3]. In particular, this 
statement from [3] is important:

   Memory QoS uses memory.high to throttle workload approaching its
   memory limit, ensuring that the system is not overwhelmed by
   instantaneous memory allocation.

With cgroup-v1 a kube memory limit would set memory.limit and usage of 
the pod (sum across all processes in the pod cgroup) was tracked with 
memory.usage_in_bytes. Whenever the latter exceeds the former, the OOM 
killer will whack the process in the pod cgroup with the highest 
oom_score, irrespective of how much free memory may be available at the 
host level.

With cgroup-v2 it appears that kube uses memory.high[5], which is more 
of a throttle/soft limit. In cgroup-v2 there is also a new memory.max[6] 
which is essentially the same as what memory.limit was in v1. Exceeding 
memory.max would invoke the OOM killer, but since kubernetes limits the 
pod memory with memory.high, the OOM killer should be avoided.

Note that I cannot claim a bunch of hands on experience with this 
(cgroup-v2 with kubernetes), so please do your own testing and YMMV, etc.

[1] https://docs.kernel.org/accounting/psi.html#psi
[2] https://kubernetes.io/docs/concepts/architecture/cgroups/
[3] 
https://kubernetes.io/docs/concepts/workloads/pods/pod-qos/#memory-qos-with-cgroup-v2
[4] 
https://docs.kernel.org/admin-guide/cgroup-v2.html#:~:text=memory.reclaim
[5] https://docs.kernel.org/admin-guide/cgroup-v2.html#:~:text=memory.high
[6] https://docs.kernel.org/admin-guide/cgroup-v2.html#:~:text=memory.max


-- 
Joe Conway
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com





view thread (6+ messages)

Message-ID: <e52f00fc-ffe3-4745-8082-b492f44c6693@joeconway.com>
Permalink:  ../e52f00fc-ffe3-4745-8082-b492f44c6693@joeconway.com/
Also on:    postgresql.org/message-id/e52f00fc-ffe3-4745-8082-b492f44c6693@joeconway.com

 · 

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-admin@postgresql.org
  Cc: mail@joeconway.com, joe.hammerman@datadoghq.com, tgl@sss.pgh.pa.us, pgsql-admin@lists.postgresql.org
  Subject: Re: Getting out ahead of OOM
  In-Reply-To: <e52f00fc-ffe3-4745-8082-b492f44c6693@joeconway.com>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox