fix(k8s): bring the over-reserved memory limits down to measured use

Six pods reserved far more memory than they have ever touched. uptime-kuma held
a 3Gi limit against 469M of measured p95, metube 2Gi against 72M, convertx
1.5Gi against 85M, netbird-server 1Gi against 97M, searxng 700Mi against 134M
and bentopdf 700Mi against 4M. Every one of them is a ceiling the scheduler
counts against the node while the memory sits unused.

Requests move down with the limits but never below the measured p95, so none of
these becomes an eviction candidate as a side effect of being right-sized. The
limits keep between 2.2x and 11.6x over the observed max, which is the figure
that decides whether a pod gets OOM-killed during a burst.

Net effect across the six: requests -557M, limits -4.6Gi, all of it ceiling that
was never in use. This is the first change that actually gives memory back.

CPU limits are left exactly as they were. They were not part of the sizing pass,
they are not being hit on a node sitting at 5% CPU, and removing them is a
separate decision from moving memory.

Verified: each limit is above the container's own observed max and each request
is above its p95, and 16/16 local gates pass.
This commit is contained in:
forust committed 2026-09-28 10:20:24 +02:00
1 parent 16aaeb60c1
commit 2a4f215546
6 files changed
+17 -11

No files matched your search

+3 -2
View File
@@ -31,12 +31,13 @@ spec:
name: bentopdf
ports:
- containerPort: 8080
# p95 4M, max 11M over 7 days. Was 50Mi/700Mi.
resources:
requests:
memory: "50Mi"
memory: "32Mi"
cpu: "50m"
ephemeral-storage: "100Mi"
limits:
memory: "700Mi"
memory: "128Mi"
cpu: "700m"
ephemeral-storage: "5Gi"