Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.
Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.
Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.
prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.
Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.
CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.
Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in 1505b63 therefore could not recover from a bad image.
Pin the digest at apply time. A digest is not knowable when a manifest is
written, so render_pinned resolves it on the way into the cluster and the
digest is never committed. Git keeps a readable `:prod`, Renovate keeps
seeing exactly the manifests it saw before, and the previous revision of
each workload now holds the digest that was actually serving, so undo
restores those exact bytes.
imagePullPolicy is dropped from the manifests rather than set to
IfNotPresent: a reference that is not `:latest` already defaults to it, and
that is what the Kubernetes docs ask for alongside a digest.
An unresolvable image is fatal instead of a warning, because carrying on
would quietly apply a mutable tag again.
restart_stale_images keeps its comparison but is no longer how a rebuild
reaches the cluster -- the pinned template rolls out on its own now. What
is left is a drift check for hand-run `kubectl set image`, so it matches
the container by repository: a pod's status now reports `repo@sha256:...`
while the manifest still says `:prod`.
The build job stops pushing `:latest` altogether, which removes the tag
that a dev branch could otherwise move under a prod deploy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>