fix(k8s): size the remaining workloads against measured use

Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.

Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.

Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.

prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.

Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.

CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.

Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
This commit is contained in:
forust committed 2026-09-28 10:38:19 +02:00
1 parent 2ad4fa1b82
commit f9e4623ade
19 files changed
+42 -42

No files matched your search

+1 -1
View File
@@ -73,7 +73,7 @@ spec:
memory: "1.5Gi"
cpu: "300m"
requests:
memory: "500Mi"
memory: "1Gi"
cpu: "50m"
ports:
- containerPort: 3000
+3 -3
View File
@@ -52,7 +52,7 @@ spec:
- containerPort: 9000
resources:
requests:
memory: "700Mi"
memory: "768Mi"
cpu: "300m"
limits:
memory: "1.5Gi"
@@ -86,8 +86,8 @@ spec:
name: authentik-secrets
resources:
requests:
memory: "512Mi"
memory: "320Mi"
cpu: "300m"
limits:
memory: "1Gi"
memory: "768Mi"
cpu: "700m"
+2 -2
View File
@@ -22,10 +22,10 @@ spec:
imagePullPolicy: Always
resources:
requests:
memory: "20Mi"
memory: "32Mi"
cpu: "30m"
limits:
memory: "64Mi"
memory: "128Mi"
cpu: "50m"
envFrom:
- secretRef:
+2 -2
View File
@@ -30,8 +30,8 @@ spec:
key: TUNNEL_TOKEN
resources:
requests:
memory: "32Mi"
memory: "128Mi"
cpu: "30m"
limits:
memory: "128Mi"
memory: "256Mi"
cpu: "200m"
+2 -2
View File
@@ -28,10 +28,10 @@ spec:
resources:
requests:
cpu: 25m
memory: 64Mi
memory: 32Mi
limits:
cpu: 250m
memory: 256Mi
memory: 128Mi
readinessProbe:
exec:
command: ["redis-cli", "ping"]
+2 -2
View File
@@ -38,10 +38,10 @@ spec:
resources:
requests:
cpu: 25m
memory: 96Mi
memory: 32Mi
limits:
cpu: 250m
memory: 256Mi
memory: 128Mi
readinessProbe:
exec:
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
+2 -2
View File
@@ -67,7 +67,7 @@ spec:
resources:
requests:
cpu: "50m"
memory: "128Mi"
memory: "192Mi"
limits:
cpu: "600m"
memory: "512Mi"
memory: "384Mi"
+2 -2
View File
@@ -47,10 +47,10 @@ spec:
mountPath: /data
resources:
requests:
memory: "512Mi"
memory: "320Mi"
cpu: "300m"
limits:
memory: "1.5Gi"
memory: "1Gi"
cpu: "1300m"
volumes:
- name: gitea-data
+2 -2
View File
@@ -57,10 +57,10 @@ spec:
resources:
requests:
cpu: "50m"
memory: "64Mi"
memory: "32Mi"
limits:
cpu: "200m"
memory: "256Mi"
memory: "128Mi"
volumes:
- name: glance-config
configMap:
+4 -4
View File
@@ -39,10 +39,10 @@ spec:
failureThreshold: 3
resources:
requests:
memory: "10Mi"
memory: "32Mi"
cpu: "20m"
limits:
memory: "100Mi"
memory: "128Mi"
cpu: "50m"
---
apiVersion: v1
@@ -86,8 +86,8 @@ spec:
failureThreshold: 3
resources:
requests:
memory: "10Mi"
memory: "32Mi"
cpu: "20m"
limits:
memory: "100Mi"
memory: "128Mi"
cpu: "50m"
+4 -4
View File
@@ -57,10 +57,10 @@ singleBinary:
storageClass: local-path-retain
resources:
requests:
memory: "512Mi"
memory: "256Mi"
cpu: "200m"
limits:
memory: "2Gi"
memory: "1Gi"
cpu: "1000m"
# Zeroed: unused in SingleBinary mode (chart validation requires it).
@@ -75,10 +75,10 @@ gateway:
replicas: 1
resources:
requests:
memory: "64Mi"
memory: "32Mi"
cpu: "50m"
limits:
memory: "256Mi"
memory: "128Mi"
cpu: "300m"
monitoring:
+2 -2
View File
@@ -163,10 +163,10 @@ spec:
failureThreshold: 5
resources:
requests:
memory: "64Mi"
memory: "32Mi"
cpu: "50m"
limits:
memory: "256Mi"
memory: "128Mi"
cpu: "300m"
---
apiVersion: v1
+2 -2
View File
@@ -97,7 +97,7 @@ spec:
resources:
requests:
cpu: "100m"
memory: "512Mi"
memory: "1Gi"
limits:
cpu: "2"
memory: "2Gi"
@@ -164,7 +164,7 @@ spec:
memory: "256Mi"
limits:
cpu: "1"
memory: "1Gi"
memory: "512Mi"
volumes:
- name: netbox-config
configMap:
+2 -2
View File
@@ -51,8 +51,8 @@ spec:
key: NETRONOME__DB_PASSWORD
resources:
requests:
memory: "100Mi"
memory: "64Mi"
cpu: "100m"
limits:
memory: "512Mi"
memory: "256Mi"
cpu: "500m"
+2 -2
View File
@@ -67,10 +67,10 @@ prometheus:
storage: 40Gi
resources:
requests:
memory: "700Mi"
memory: "768Mi"
cpu: 200m
limits:
memory: "2Gi"
memory: "2560Mi"
alertmanager:
alertmanagerSpec:
configSecret: alertmanager-config
+2 -2
View File
@@ -66,10 +66,10 @@ spec:
periodSeconds: 30
resources:
requests:
memory: "128Mi"
memory: "160Mi"
cpu: "100m"
limits:
memory: "512Mi"
memory: "384Mi"
cpu: "500m"
volumes:
- name: rackpeek-config
+2 -2
View File
@@ -38,10 +38,10 @@ spec:
mountPath: /app/data
resources:
requests:
memory: "128Mi"
memory: "160Mi"
cpu: "100m"
limits:
memory: "256Mi"
memory: "384Mi"
cpu: "300m"
livenessProbe:
httpGet:
+2 -2
View File
@@ -36,10 +36,10 @@ spec:
resources:
requests:
cpu: "100m"
memory: "128Mi"
memory: "80Mi"
limits:
cpu: "500m"
memory: "512Mi"
memory: "256Mi"
volumeMounts:
- name: vaultwarden-data
mountPath: /data
+2 -2
View File
@@ -51,10 +51,10 @@ spec:
mountPath: /etc/x-ui
resources:
requests:
memory: "128Mi"
memory: "192Mi"
cpu: "100m"
limits:
memory: "1Gi"
memory: "512Mi"
cpu: "1000m"
volumes:
- name: x-ui-db