Memory QoS has graduated to Beta in Kubernetes v1.37 and is now enabled by default. On Linux nodes running cgroup v2, the feature uses the memory controller to give the kernel better guidance on how to treat container memory. It was first introduced as Alpha in v1.22, and expanded in v1.36 with tiered memory reservation.
This post covers what changed in v1.37, what the Beta promotion means for cluster operators, and how to configure the feature.
The MemoryQoS feature gate is now Beta in v1.37.
This means every v1.37 kubelet has the feature gate turned on without any
configuration change. Turning on the feature by default is safe because the
default kubelet configuration does not enable memory throttling or memory
reservation. No memory.high, memory.min, or memory.low values are
written to cgroups unless you explicitly configure them.
You can opt into specific behaviors through kubelet configuration fields:
memoryThrottlingFactor (for example, 0.9) to enable memory.high throttling on Burstable and BestEffort containers. The default is null, which means no throttling.memoryReservationPolicy to TieredReservation to enable tiered memory protection via memory.min and memory.low. The default is None, which means no memory reservation.memoryThrottlingFactor changed to nullIn earlier Alpha releases, memoryThrottlingFactor defaulted to 0.9, which
meant enabling the feature gate caused the kubelet to set
memory.high on containers. In v1.37, the default is null, so the kubelet
does not set memory.high unless you configure a value.
This change was made because, with the feature gate now on by default, an
automatic memory.high could throttle workloads that were previously running
without throttling. Making it null ensures that upgrading to v1.37 does not
change runtime behavior for existing clusters.
If your kubelet configuration file already contains an explicit
memoryThrottlingFactor value, that value is preserved during the upgrade and
throttling continues to work as before. If your configuration file does not
include memoryThrottlingFactor, the kubelet uses the new null default and
stops setting memory.high. To keep throttling in that case, add
memoryThrottlingFactor explicitly:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
memoryThrottlingFactor: 0.9
For full details on configuring Memory QoS, see Memory QoS with cgroup v2, Configuring memory reservation, and System requirements
Set memoryThrottlingFactor to a value between 0 and 1. The kubelet uses this
factor to calculate memory.high for Burstable and BestEffort containers. See
Memory throttling
for how memory.high is calculated for each QoS class.
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
memoryThrottlingFactor: 0.9
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
memoryThrottlingFactor: 0.9
memoryReservationPolicy: TieredReservation
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
memoryReservationPolicy: TieredReservation
To disable the feature after upgrading, set the feature gate to false
and ensure a compatible kubelet configuration. The kubelet rejects the configuration if
memoryThrottlingFactor is set to anything other than the former default of 0.9, or if
memoryReservationPolicy is TieredReservation, so remove or adjust those fields if you set them.
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
featureGates:
MemoryQoS: false
When the feature gate is off, or memoryReservationPolicy is not TieredReservation, the
kubelet resets stale protection at startup on cgroup v2 nodes: memory.min=0 and memory.low=0
on the root kubepods cgroup, and memory.low=0 on the Burstable QoS cgroup.
For containers, stale memory.high values are reset to max on reconciliation paths such as restart or resize.
memoryReservationPolicy applies to every pod on the node. With TieredReservation, every Guaranteed pod gets memory.min and every Burstable pod gets memory.low; there is no way to opt individual pods in or out. A node that mixes workloads needing hard reservation with workloads that should stay reclaimable has to choose one policy for all of them.
Hard reservation also covers everything charged to the container's cgroup, including page cache, so a pod that reads large files can hold memory the kernel would otherwise reclaim to serve its neighbors.
SIG Node is tracking both in kubernetes/kubernetes#140246. If this affects you, that issue is the best place to describe your workload.
The next milestone for Memory QoS is graduation to GA. Feedback from Beta users will shape any remaining adjustments before that step. If you run into issues, please file bugs at kubernetes/kubernetes.
This feature is driven by SIG Node. If you are interested in contributing or have feedback, you can reach out through: