pg.ddx.io pgsql-admin@postgresql.org mailing list archive
help / color / mirror / Atom feed From: Joe Conway <mail@joeconway.com>
To: Joseph Hammerman <joe.hammerman@datadoghq.com>
Cc: Tom Lane <tgl@sss.pgh.pa.us>
Cc: pgsql-admin@lists.postgresql.org
Subject: Re: Getting out ahead of OOM
Date: Wed, 12 Mar 2025 22:35:17 -0400
Message-ID: <e52f00fc-ffe3-4745-8082-b492f44c6693@joeconway.com> (raw )
In-Reply-To: <CAHs7QM_P-wqprxTmD5sjF9ShX2kLWGdR9gatPBtO+FfLqfWr6Q@mail.gmail.com >
References: <CAHs7QM_8SRfybxg_Pq3orGw5anMkV7OzKYfxDXSBQVWzcXf3Bw@mail.gmail.com >
<1117736.1741375563@sss.pgh.pa.us >
<efcca885-a954-43e2-99b8-8b993678c72c@joeconway.com >
<CAHs7QM_P-wqprxTmD5sjF9ShX2kLWGdR9gatPBtO+FfLqfWr6Q@mail.gmail.com >
On 3/12/25 18:21, Joseph Hammerman wrote:
> Joe, can you expand on your recommendation to use cgroup-v2? We're
> trying to collect our complete rationale for our request to our internal
> team that is tasked with rolling out this configuration change.
cgroup-v2 has a much better measure of memory pressure (see PSI[1]),
better ability to reclaim memory pages[4], safer delegation, and other
advantages. When I last looked the kube support for it was still brand
new, but it appears to be well supported now [2][3]. In particular, this
statement from [3] is important:
Memory QoS uses memory.high to throttle workload approaching its
memory limit, ensuring that the system is not overwhelmed by
instantaneous memory allocation.
With cgroup-v1 a kube memory limit would set memory.limit and usage of
the pod (sum across all processes in the pod cgroup) was tracked with
memory.usage_in_bytes. Whenever the latter exceeds the former, the OOM
killer will whack the process in the pod cgroup with the highest
oom_score, irrespective of how much free memory may be available at the
host level.
With cgroup-v2 it appears that kube uses memory.high[5], which is more
of a throttle/soft limit. In cgroup-v2 there is also a new memory.max[6]
which is essentially the same as what memory.limit was in v1. Exceeding
memory.max would invoke the OOM killer, but since kubernetes limits the
pod memory with memory.high, the OOM killer should be avoided.
Note that I cannot claim a bunch of hands on experience with this
(cgroup-v2 with kubernetes), so please do your own testing and YMMV, etc.
[1] https://docs.kernel.org/accounting/psi.html#psi
[2] https://kubernetes.io/docs/concepts/architecture/cgroups/
[3]
https://kubernetes.io/docs/concepts/workloads/pods/pod-qos/#memory-qos-with-cgroup-v2
[4]
https://docs.kernel.org/admin-guide/cgroup-v2.html#:~:text=memory.reclaim
[5] https://docs.kernel.org/admin-guide/cgroup-v2.html#:~:text=memory.high
[6] https://docs.kernel.org/admin-guide/cgroup-v2.html#:~:text=memory.max
--
Joe Conway
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
view thread (6+ messages)
Message-ID: <e52f00fc-ffe3-4745-8082-b492f44c6693@joeconway.com>
Permalink: ../e52f00fc-ffe3-4745-8082-b492f44c6693@joeconway.com/
Also on: postgresql.org/message-id/e52f00fc-ffe3-4745-8082-b492f44c6693@joeconway.com
copy link · copy postgr.es
reply Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-admin@postgresql.org
Cc: mail@joeconway.com, joe.hammerman@datadoghq.com, tgl@sss.pgh.pa.us, pgsql-admin@lists.postgresql.org
Subject: Re: Getting out ahead of OOM
In-Reply-To: <e52f00fc-ffe3-4745-8082-b492f44c6693@joeconway.com>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox