Files
homelab/netbird/k8s/netbird.yaml
T
forust f9e4623ade fix(k8s): size the remaining workloads against measured use
Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.

Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.

Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.

prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.

Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.

CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.

Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
2026-09-28 10:38:19 +02:00

183 lines
4.4 KiB
YAML

apiVersion: v1
kind: Service
metadata:
name: netbird-server-service
namespace: netbird
spec:
selector:
app: netbird-server
ports:
- port: 80
name: http
targetPort: 80
protocol: TCP
- port: 3478
name: stun
targetPort: 3478
protocol: UDP
---
apiVersion: v1
kind: Service
metadata:
name: netbird-dashboard-service
namespace: netbird
spec:
selector:
app: netbird-dashboard
ports:
- port: 80
name: http
targetPort: 80
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: netbird-server-deployment
namespace: netbird
spec:
replicas: 1
selector:
matchLabels:
app: netbird-server
template:
metadata:
labels:
app: netbird-server
spec:
containers:
- name: netbird-server
image: netbirdio/netbird-server:0.79.0
command: ["/bin/sh", "/opt/netbird/entrypoint.sh", "--config", "/run/netbird/config.yaml"]
envFrom:
- configMapRef:
name: netbird-config
ports:
- containerPort: 80
name: http
protocol: TCP
- containerPort: 3478
name: stun
protocol: UDP
volumeMounts:
- name: netbird-data
mountPath: /var/lib/netbird
- name: netbird-files
mountPath: /opt/netbird
readOnly: true
- name: netbird-secrets
mountPath: /run/secrets/relay_auth_secret
subPath: relay_auth_secret
readOnly: true
- name: netbird-secrets
mountPath: /run/secrets/datastore_encryption_key
subPath: datastore_encryption_key
readOnly: true
- name: netbird-run
mountPath: /run/netbird
readinessProbe:
tcpSocket:
port: 80
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 5
livenessProbe:
tcpSocket:
port: 80
initialDelaySeconds: 60
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 5
# p95 97M, max 102M over 7 days. Was 256Mi/1Gi.
resources:
requests:
memory: "128Mi"
cpu: "250m"
limits:
memory: "384Mi"
cpu: "1000m"
volumes:
- name: netbird-data
persistentVolumeClaim:
claimName: netbird-pvc
- name: netbird-files
configMap:
name: netbird-config
defaultMode: 0755
items:
- key: config.template.yaml
path: config.template.yaml
- key: entrypoint.sh
path: entrypoint.sh
- name: netbird-secrets
secret:
secretName: netbird-secrets
items:
- key: relay_auth_secret
path: relay_auth_secret
- key: datastore_encryption_key
path: datastore_encryption_key
- name: netbird-run
emptyDir:
medium: Memory
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: netbird-dashboard-deployment
namespace: netbird
spec:
replicas: 1
selector:
matchLabels:
app: netbird-dashboard
template:
metadata:
labels:
app: netbird-dashboard
spec:
containers:
- name: dashboard
image: netbirdio/dashboard:v2.93.0
envFrom:
- configMapRef:
name: netbird-config
ports:
- containerPort: 80
name: http
readinessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 15
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 5
livenessProbe:
httpGet:
path: /
port: 80
initialDelaySeconds: 30
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 5
resources:
requests:
memory: "32Mi"
cpu: "50m"
limits:
memory: "128Mi"
cpu: "300m"
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: netbird-pvc
namespace: netbird
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 2Gi