Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.
Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.
Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.
prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.
Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.
CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.
Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
Six pods reserved far more memory than they have ever touched. uptime-kuma held
a 3Gi limit against 469M of measured p95, metube 2Gi against 72M, convertx
1.5Gi against 85M, netbird-server 1Gi against 97M, searxng 700Mi against 134M
and bentopdf 700Mi against 4M. Every one of them is a ceiling the scheduler
counts against the node while the memory sits unused.
Requests move down with the limits but never below the measured p95, so none of
these becomes an eviction candidate as a side effect of being right-sized. The
limits keep between 2.2x and 11.6x over the observed max, which is the figure
that decides whether a pod gets OOM-killed during a burst.
Net effect across the six: requests -557M, limits -4.6Gi, all of it ceiling that
was never in use. This is the first change that actually gives memory back.
CPU limits are left exactly as they were. They were not part of the sizing pass,
they are not being hit on a node sitting at 5% CPU, and removing them is a
separate decision from moving memory.
Verified: each limit is above the container's own observed max and each request
is above its p95, and 16/16 local gates pass.