Commit Graph
12 Commits
Author SHA1 Message Date
forust a6af69dca0 fix(k8s): Recreate singletons and trim requests for scheduler headroom
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-compose (push) Successful in 3s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 13s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 18s
ci / build (push) Successful in 1m30s
RollingUpdate with default maxSurge needs a spare pod the single node does not have (99% CPU requested), so multi-workload restarts end Pending and verify times out. Recreate on all replicas:1 Deployments (immich-server and bentopdf keep RollingUpdate at replicas 2). Also trims CPU/memory requests toward measured use (adguard, authentik, gitea, netbox, uptime-kuma, netbird-server) and gives traefik requests/limits so it is no longer BestEffort.
2026-09-28 22:36:46 +02:00
forust f9e4623ade fix(k8s): size the remaining workloads against measured use
Finishes the sizing pass over every workload the deploy actually manages. Each
request is at or above the container's p95 over the last seven days, so nothing
is sized below what it is known to use, and each limit is between 1.6x and 5x
the observed max, which is the figure that decides whether a burst gets an
OOMKill.

Some of these go up, and that is the point. adguard was holding 975M against a
500Mi request and netbox 962M against 512Mi, so both sat permanently above
their own request and were standing eviction candidates on a node that has
about 300M of headroom. Raising a request costs scheduler room; leaving it low
costs the pod its place in the queue when the node gets tight.

Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the
authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis,
glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the
loki gateway -- each reserved 4x to 16x more than they have ever touched.

prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it
compacts its TSDB in place and that is a burst worth budgeting for rather than
throttling.

Two of these limits are close enough to the observed max to be worth watching
rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so
the ceiling is a date, not a margin. That was true before this change too; the
pod sizing does not fix it and the cache needs bounding.

CPU limits are untouched throughout. Leaving postgres alone as well: it sits in
an uncommitted file that belongs to other work in progress.

Verified: every request is at or above p95 and every limit above the observed
max across all 74 containers, and 16/16 local gates pass.
2026-09-28 10:38:19 +02:00
forustandCopilot 23ed72826a chore(images): pin service image updates
Replace floating service images with reviewable tags or digests.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-14 09:16:01 +02:00
forust 861d89d36a ci(deploy): split runtime by k8s/active marker
services marked k8s/active are applied via kubectl; the rest via docker
compose. inactive services with k8s/ keep only routing manifests
(external Services, EndpointSlices, Ingresses) to reach docker backends.
headscale/nextcloud routing moved to k8s/routing/.

validations: compose config --quiet + kubectl apply --dry-run=client.
namespace manifests applied first. pull_policy:build stacks get
build+push before up so the registry image stays fresh.
2026-09-06 20:39:12 +02:00
forust f424d91405 Tighten lint workflow scope
lint / prettier (push) Successful in 8s
lint / ruff (push) Failing after 5s
lint / yamllint (push) Successful in 6s
lint / hadolint (push) Failing after 4s
validate / yaml (push) Successful in 6s
validate / k8s (push) Successful in 4s
2026-06-21 21:51:20 +02:00
forust 2de6131ba7 chore: apply prettier formatting across compose files and configs 2026-06-19 12:03:11 +02:00
forust 1ea669220b chore(k8): adjusted system resources requests and limits based on manual monitoring 2026-06-10 19:37:16 +02:00
forust 7a3708f70c lint: yaml spaces and tabs 2026-06-09 12:59:36 +02:00
forust f3d04e935c k8s: cloudflare-ddns deployment via kubernetes
Deploy to Server / deploy (push) Has been cancelled
2026-06-07 21:04:03 +02:00
forust d988b0f7db refactor: change Cloudlfare DDNS config to env-based
Deploy to Server / deploy (push) Failing after 14m7s
2026-05-08 17:38:41 +02:00
forust 45ce789f58 chore: compose cleanup:
- Remove TZ envs
- +- unified compose structure
- Minify where possible
- Remove <service>.internal routers
- Remove service specifications where possible
Affected services:
- adguardhome
- authentik
- cfddns
- checkmk
- dockmon
- downtify
- gitea
- glance
- headscale
- homepages
- metube
- nextcloud
- penpot
- portainer
- termix
- traefik
- uptime-kuma

TODO: Move data from directory to volumes
2026-01-19 23:33:50 +01:00
forust 2d05a1911c feat: cloudflare IP sync (ddns) 2025-11-30 03:28:41 +01:00