main
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
939fad4a23 |
chore(renovate,crowdsec): group minor/patch updates and whitelist VPS
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 6s
ci / lint-prettier (push) Successful in 11s
ci / lint-ruff (push) Successful in 5s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 6s
renovate-ci / validate-renovate (push) Successful in 10s
ci / build (push) Successful in 17s
Renovate: group all minor updates and all patch updates into reviewable PRs. Crowdsec: whitelist static VPS IP. Remove stale cloudflared/k8s/active marker. |
||
|
|
12d5cc6202 |
Merge pull request 'chore(deps): update docker.gitea.com/gitea docker tag to v28' (#74) from renovate/docker.gitea.com-gitea-28.x into main
ci / lint-compose (push) Successful in 8s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / lint-prettier (push) Successful in 13s
ci / lint-ruff (push) Successful in 5s
ci / validate (push) Successful in 6s
renovate-ci / validate-renovate (push) Successful in 11s
ci / build (push) Successful in 16s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/74 |
||
|
|
a6b84d86b4 |
Merge pull request 'chore(deps): update gitea/gitea docker tag to v28' (#75) from renovate/gitea-gitea-28.x into main
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 8s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/75 |
||
|
|
f3a5cedc80 |
Merge pull request 'chore(deps): update renovate/renovate docker tag to v44.132.2' (#69) from renovate/renovate-self-update into main
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 5s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 14s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 6s
ci / build (push) Successful in 19s
renovate-ci / validate-renovate (push) Successful in 1m10s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/69 |
||
|
|
e208ccae62 |
Merge pull request 'chore(deps): update ghcr.io/alexta69/metube docker tag to v2026.09.29' (#70) from renovate/container-patch-updates into main
renovate-ci / validate-renovate (push) Successful in 8s
ci / lint-compose (push) Successful in 8s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 6s
ci / lint-prettier (push) Successful in 14s
ci / lint-ruff (push) Successful in 5s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 5s
ci / build (push) Successful in 18s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/70 |
||
|
|
d3892d2ed5 |
Merge pull request 'chore(deps): update ghcr.io/henriquesebastiao/downtify docker tag to v3.4.0' (#71) from renovate/ghcr.io-henriquesebastiao-downtify-3.x into main
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 8s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/71 Reviewed-by: forust <vzlomdsisma@gmail.com> |
||
|
|
7164b48275 |
Merge pull request 'chore(deps): update ghcr.io/lukegus/termix docker tag to v2.9.0' (#72) from renovate/ghcr.io-lukegus-termix-2.x into main
renovate-ci / validate-renovate (push) Successful in 8s
ci / lint-compose (push) Successful in 8s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 14s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 8s
ci / lint-dockerfiles (push) Successful in 4s
ci / validate (push) Successful in 6s
ci / build (push) Successful in 17s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/72 |
||
|
|
956596e6aa |
Merge pull request 'chore(deps): update docker.n8n.io/n8nio/n8n docker tag to v2.42.2' (#68) from renovate/docker.n8n.io-n8nio-n8n-2.x into main
renovate-ci / validate-renovate (push) Successful in 8s
ci / lint-compose (push) Successful in 7s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 13s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 5s
ci / build (push) Successful in 17s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/68 |
||
|
|
65709e005b |
Merge pull request 'chore(deps): update netbirdio/netbird-server docker tag to v0.80.0' (#73) from renovate/netbirdio-netbird-server-0.x into main
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 12s
ci / lint-ruff (push) Successful in 5s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 7s
renovate-ci / validate-renovate (push) Successful in 8s
ci / build (push) Successful in 18s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/73 |
||
|
|
78e6363fe2 |
chore(gitea): add git.forust.xyz route
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 5s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 12s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 4s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (push) Successful in 8s
ci / build (push) Successful in 17s
|
||
|
|
22c2f1e108 |
extract dtek_notif subtree to its own repo
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 10s
|
||
|
|
64ba4ce9f1 |
Revert "chore(prometheus): drop dead grafana routes and cert"
This reverts commit
|
||
|
|
7a81b4ea8b | feat(adguard): add netbird sidecar | ||
|
|
0859479c0f |
feat(ingress): replace traefik crowdsec plugin with firewall bouncer
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 12s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (push) Successful in 7s
ci / build (push) Failing after 14m22s
Move L3 enforcement to the host firewall-bouncer (systemd, nftables): drop the Traefik plugin, its secrets volume and the crowdsec Middleware, remove bouncer refs from all IngressRoutes. Disable the http-generic-bf scenario (403-burst bans hurt legit automation under L3 enforcement). Add a Gateway API PoC for homepages prod and CrowdSec PrometheusRule alerts. |
||
|
|
872f64b887 |
fix(ci): silence intentional SC2029 in ssh-run.sh
ci / lint-compose (push) Successful in 11s
ci / lint-actionlint (push) Successful in 7s
ci / lint-shellcheck (push) Successful in 9s
ci / lint-prettier (push) Successful in 16s
ci / lint-ruff (push) Successful in 7s
ci / lint-yaml (push) Successful in 12s
ci / lint-dockerfiles (push) Successful in 8s
ci / validate (push) Successful in 8s
renovate-ci / validate-renovate (push) Successful in 22s
ci / build (push) Successful in 38s
|
||
|
|
a90fb19ed4 |
fix(traefik): size probes for HDD stalls
ci / lint-compose (push) Successful in 12s
ci / lint-actionlint (push) Successful in 9s
ci / lint-shellcheck (push) Failing after 25s
ci / lint-prettier (push) Successful in 19s
ci / lint-ruff (push) Successful in 9s
ci / lint-yaml (push) Successful in 12s
ci / lint-dockerfiles (push) Successful in 8s
ci / validate (push) Successful in 9s
ci / build (push) Skipped
renovate-ci / validate-renovate (push) Successful in 12s
Single replica is the whole ingress; liveness kills at 2s timeouts took every public service down in a loop. Same stall-sized budgets as postgres/metallb. Applied live via helm (pinned 41.5.0, values from git). |
||
|
|
6f4cd03f4b |
feat(deploy): serialize apply stages and wait for calm node
apply-k8s and apply-compose share a workstation flock so host docker churn never overlaps cluster churn. Each helm upgrade and the apply loop wait up to 10m for load <28 first, so a deploy never piles onto an already-hot node (the load-40/netbird-death/pending-helm spiral). |
||
|
|
e56662194b |
fix(ci): log in to registry for manifest-only main pushes
ci / lint-compose (push) Successful in 11s
ci / lint-actionlint (push) Successful in 6s
ci / lint-shellcheck (push) Successful in 9s
ci / lint-prettier (push) Successful in 15s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 13s
ci / lint-dockerfiles (push) Successful in 8s
ci / validate (push) Successful in 9s
renovate-ci / validate-renovate (push) Successful in 10s
ci / build (push) Successful in 27s
The pin step writes manifest PUTs on every main push, but login was gated on services != ''. Manifest-only pushes skipped login and pushed anonymously (401); this only worked before via a stale persistent login on the old runner. |
||
|
|
4856348a6e |
chore(searxng,glance): disable services
renovate-ci / validate-renovate (push) Successful in 15s
ci / lint-compose (push) Successful in 13s
ci / lint-actionlint (push) Successful in 7s
ci / lint-shellcheck (push) Successful in 11s
ci / lint-prettier (push) Successful in 18s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 13s
ci / lint-dockerfiles (push) Successful in 8s
ci / validate (push) Successful in 8s
ci / build (push) Failing after 19s
Workloads, services, routes and certs removed from the cluster. Configs, manifests and PVCs kept for easy re-enable via k8s/active. |
||
|
|
49d0da1dd0 |
chore(termix): disable service
renovate-ci / validate-renovate (push) Successful in 11s
ci / lint-compose (push) Successful in 17s
ci / lint-actionlint (push) Successful in 8s
ci / lint-shellcheck (push) Successful in 10s
ci / lint-prettier (push) Successful in 18s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 13s
ci / lint-dockerfiles (push) Successful in 8s
ci / validate (push) Successful in 12s
ci / build (push) Failing after 17s
Deployment was stuck in ImagePullBackOff and only generated log churn. Workload, service, routes and certs removed from the cluster; PVC, config and manifests kept so it can be re-enabled by restoring k8s/active. |
||
|
|
6bc938436f |
Merge pull request 'chore(deps): update grafana/grafana docker tag to v13.2.3' (#67) from renovate/grafana-monorepo into main
ci / lint-compose (push) Successful in 11s
ci / lint-actionlint (push) Successful in 7s
ci / lint-shellcheck (push) Successful in 10s
ci / lint-prettier (push) Successful in 18s
ci / lint-ruff (push) Successful in 8s
ci / lint-yaml (push) Successful in 12s
ci / lint-dockerfiles (push) Successful in 7s
renovate-ci / validate-renovate (push) Successful in 12s
ci / validate (push) Successful in 8s
ci / build (push) Failing after 2m13s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/67 |
||
|
|
5dcad7eb38 |
fix(ci): resolve uv from BIN_DIR in install-ci-tools
ci / lint-compose (push) Successful in 14s
ci / lint-actionlint (push) Successful in 10s
ci / lint-shellcheck (push) Successful in 11s
ci / lint-prettier (push) Successful in 17s
ci / lint-ruff (push) Successful in 9s
ci / lint-yaml (push) Successful in 12s
ci / lint-dockerfiles (push) Successful in 9s
ci / validate (push) Successful in 9s
renovate-ci / validate-renovate (push) Successful in 1m34s
ci / build (push) Failing after 3m3s
Callers prepend BIN_DIR to PATH only after the script exits, so the bare uv invocation in install_uv_tool died with 127 on clean runners. Export BIN_DIR to PATH inside the script and invoke the just-installed binary by absolute path. |
||
|
|
d0872bc918 |
feat(deploy): autodeploy off unless AUTODEPLOY=true
renovate-ci / validate-renovate (push) Successful in 2m1s
ci / lint-shellcheck (push) Successful in 24s
ci / lint-prettier (push) Successful in 6s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
ci / lint-compose (push) Successful in 16s
ci / lint-actionlint (push) Successful in 8s
ci / build (push) Failing after 18s
Pushes no longer reach the cluster by default; set the AUTODEPLOY repo variable to 'true' to re-enable, or dispatch manually. Skips propagate through the existing needs chain. |
||
|
|
91dd749a29 |
feat(deploy): AUTODEPLOY kill-switch via repo variable
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 3s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 15s
ci / build (push) Canceled after 0s
Set AUTODEPLOY=false under Settings -> Actions -> Variables and pushes stop deploying with no commit; unset means on. Manual Run workflow always bypasses the switch. The existing needs/skipped chaining propagates the skip through verify and smoke untouched. |
||
|
|
b6e1dc0362 |
fix(immich): size postgres probes for HDD stalls
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 17s
ci / build (push) Successful in 13s
Postmaster was SIGKILLed in a loop: 70s fsync stalls on the loaded rotational disk outlasted the 5-minute startup budget and the 60s liveness tolerance, and every kill bought another full WAL replay. Startup budget 15min, liveness 5x60s. Already applied live with kubectl; this keeps git in sync. |
||
|
|
a2af854867 |
fix(alerts): measure traefik latency per route via exported_service
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 11s
ci / build (push) Successful in 6s
The service label the rule grouped by is the scrape target name, not the backend: with honorLabels=false the per-route value traefik emits is renamed to exported_service, so the old rule measured one global aggregate and both exclusions matched nothing. Group the latency and 5xx alerts by exported_service with exclusions in traefik's real namespace-routename-hash format. Already applied live with kubectl; this commit keeps git in sync. |
||
|
|
314ec0cda7 |
chore(prometheus): drop dead grafana routes and cert
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 19s
ci / build (push) Successful in 6s
No grafana is deployed from this stack, so grafana-prod/grafana-local only ever served 503s. Live objects already deleted directly; this keeps git from resurrecting them on the next deploy. |
||
|
|
e45ae10204 |
feat(immich): migrate postgres 14 to 16 with dump/restore
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 5s
ci / lint-ruff (push) Successful in 3s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 3s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 1m7s
ci / build (push) Successful in 11s
Reverts the v16 revert: data migrated via pg_dumpall from PG14 into a fresh PG16 data directory (same extension versions vectorchord 0.4.3/pgvectors 0.2.0). Verified 1944 assets, 2 users, 10 albums on 16.10. Note: the restore OOMed once at the 2Gi limit during index builds and recovered via WAL replay; worth watching under PG16 load. |
||
|
|
357ee26111 |
Revert "Merge pull request 'chore(deps): update ghcr.io/immich-app/postgres docker tag to v16' (#65) from renovate/ghcr.io-immich-app-postgres-16.x into main"
ci / lint-compose (push) Successful in 5s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 34s
ci / build (push) Successful in 6s
This reverts commit |
||
|
|
436fd1ecae |
chore: extract userbot subtree to its own repo
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 17s
ci / build (push) Successful in 8s
userbot/ (bot plus panel) now lives at /home/forust/userbot as a clone of forust/userbot instead of a subtree in homelab. Cleans up the pipeline references that only existed for it: scan-deps, test-backend and test-frontend jobs, the userbot build matrix entries, the deploy panel hook, and the userbot-only pyrightconfig. Running cluster workloads are untouched; homelab just stops building and testing upstream's code. |
||
|
|
a564dd8b67 |
fix(alerts): exclude netbird streams from traefik latency alert
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 12s
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 8s
ci / build (push) Successful in 10s
SignalExchange ConnectStream holds 60s gRPC streams by design; at night they exceed 5% of samples and pin P95 to the 5.0s bucket ceiling, flapping the alert. Also fixes the stale xui exclusion pattern, which matched no real service label. |
||
|
|
c2c90c490d |
Merge pull request 'chore(deps): update ghcr.io/autobrr/netronome docker tag to v0.15.0' (#62) from renovate/ghcr.io-autobrr-netronome-0.x into main
ci / lint-compose (push) Successful in 2s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 16s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/62 |
||
|
|
3ecc12300a |
Merge pull request 'chore(deps): update ghcr.io/alexta69/metube docker tag to v2026.09.28' (#61) from renovate/container-patch-updates into main
renovate-ci / validate-renovate (push) Successful in 13s
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/61 |
||
|
|
2a5d690e34 |
Merge pull request 'chore(deps): update netbirdio/dashboard docker tag to v2.94.0' (#63) from renovate/netbirdio-dashboard-2.x into main
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 19s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/63 |
||
|
|
a8c4e9bcbe |
Merge pull request 'chore(deps): update renovate/renovate docker tag to v44.117.0' (#64) from renovate/renovate-self-update into main
renovate-ci / validate-renovate (push) Successful in 16s
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
ci / build (push) Successful in 9s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/64 |
||
|
|
0c9743e241 |
Merge pull request 'chore(deps): update ghcr.io/immich-app/postgres docker tag to v16' (#65) from renovate/ghcr.io-immich-app-postgres-16.x into main
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 14s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 27s
ci / build (push) Successful in 8s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/65 |
||
|
|
2cb06debc5 |
fix(deploy): retry registry lookups with a timeout
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 7s
A single blink of the registry failed render_pinned for the whole file and redded the apply stage. registry_digest now retries 3 times under a 25s timeout with a warning per attempt; empty still means unresolvable and callers report it by name as before. |
||
|
|
a6af69dca0 |
fix(k8s): Recreate singletons and trim requests for scheduler headroom
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-compose (push) Successful in 3s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 13s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 18s
ci / build (push) Successful in 1m30s
RollingUpdate with default maxSurge needs a spare pod the single node does not have (99% CPU requested), so multi-workload restarts end Pending and verify times out. Recreate on all replicas:1 Deployments (immich-server and bentopdf keep RollingUpdate at replicas 2). Also trims CPU/memory requests toward measured use (adguard, authentik, gitea, netbox, uptime-kuma, netbird-server) and gives traefik requests/limits so it is no longer BestEffort. |
||
|
|
76f39da90c |
fix(ci): skip heavy jobs on renovate branches, automerge digest and patch
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Canceled after 0s
Renovate branches only carry version/digest bumps, so scan-deps, test-backend, test-frontend and build just burn runner time on the box that also serves prod. Static checks and validate still run. Digest and patch updates automerge (playwright, helm and major rules below still override to no-automerge). Also replaces deprecated helm --atomic with --wait --rollback-on-failure. |
||
|
|
86730ff0c4 |
Merge pull request 'chore(deps): update ghcr.io/c4illin/convertx docker tag to v0.19.0' (#55) from renovate/ghcr.io-c4illin-convertx-0.x into main
renovate-ci / validate-renovate (push) Successful in 6s
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 14s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
ci / build (push) Successful in 10s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/55 |
||
|
|
deeaefa695 |
Merge pull request 'chore(deps): update ghcr.io/henriquesebastiao/downtify docker tag to v3.2.0' (#60) from renovate/ghcr.io-henriquesebastiao-downtify-3.x into main
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 18s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/60 |
||
|
|
742d78944b |
Merge pull request 'chore(deps): update ghcr.io/alexta69/metube docker tag to v2026.09.27' (#54) from renovate/container-patch-updates into main
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 23s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/54 |
||
|
|
33c54ac830 |
Merge pull request 'chore(deps): update docker.io/valkey/valkey:9 docker digest to 418652c' (#59) from renovate/docker.io-valkey-valkey-9 into main
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-ruff (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 11s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/59 |
||
|
|
b4f76373bb |
fix(gitea): serve issue search from postgres instead of reindexing on boot
renovate-ci / validate-renovate (push) Successful in 5s
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 9s
ci / test-frontend (push) Successful in 13s
ci / validate (push) Successful in 2s
ci / build (push) Successful in 7s
cron.rebuild_issue_indexer runs at start, so every gitea pod restart reindexed the whole issue index and read ~7MB/s off the rotational disk for an hour. |
||
|
|
74adf38d63 |
feat(alerts): cover OOM kills, restart loops and evictions
The OOMKilled container behind the immich crash loop was invisible: PodCrashLooping only fires once kubelet has already given up and started the backoff. |
||
|
|
f5b2f89f38 |
style(immich): match repo prettier quoting in compose file
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 9s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 16s
ci / test-backend (push) Successful in 9s
ci / test-frontend (push) Successful in 13s
ci / validate (push) Successful in 3s
ci / build (push) Successful in 9s
|
||
|
|
8e63284240 |
feat(immich): add self-hosted photo backup with dedicated postgres
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Failing after 4s
ci / lint-ruff (push) Successful in 3s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 20s
ci / test-backend (push) Successful in 8s
ci / validate (push) Successful in 3s
ci / build (push) Skipped
renovate-ci / validate-renovate (push) Successful in 11s
ci / test-frontend (push) Successful in 14s
Server x2, machine learning, valkey and VectorChord postgres on local storage, Traefik routes for external and internal access. |
||
|
|
1d81410cd8 |
fix(traefik): persist plugin storage on a PVC
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 3s
ci / lint-shellcheck (push) Successful in 4s
ci / lint-prettier (push) Successful in 5s
ci / lint-ruff (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 18s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 4s
renovate-ci / validate-renovate (push) Successful in 1m28s
ci / build (push) Successful in 37s
Mount traefik-plugins PVC at /plugins-storage instead of the chart default emptyDir, so the crowdsec-bouncer download survives node reboots. Without this Traefik starts before the network is ready, the download from plugins.traefik.io times out, plugins get disabled and every route behind the middleware returns 404/503 until a manual restart. |
||
|
|
2f891a5d31 |
fix(postgres): give probes room on an I/O-bound single node
renovate-ci / validate-renovate (push) Canceled after 21s
ci / lint-compose (push) Successful in 8s
ci / lint-actionlint (push) Successful in 3s
ci / lint-shellcheck (push) Successful in 4s
ci / lint-prettier (push) Successful in 4s
ci / lint-ruff (push) Successful in 10s
ci / lint-yaml (push) Successful in 4s
ci / lint-dockerfiles (push) Successful in 4s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 9s
ci / test-frontend (push) Successful in 21s
ci / validate (push) Successful in 5s
ci / build (push) Successful in 22s
pg_isready with the 1s default times out under I/O stall and kubelet kills a healthy postgres mid-recovery; each kill restarts a multi-minute fsync from zero and loops forever. readiness/liveness timeout 5s, liveness threshold 5, startup budget 15min. |
||
|
|
cda0022d81 |
fix(renovate): reap finished job pods with ttlSecondsAfterFinished
ci / lint-compose (push) Successful in 5s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 4s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 12s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (push) Successful in 9m17s
ci / build (push) Failing after 28m20s
History limits never delete manual 'create job --from' runs, so Failed pods accumulated for a week. Keep a day for debugging, reap the rest. |
||
|
|
dde6b1c743 |
fix(deploy): recover helm releases from pending-* and skip helm-owned rollbacks
An --atomic upgrade whose own rollback never finishes leaves the release in pending-*, blocking every future run until a human rolls back (loki rev 18/21). Recover automatically before and after each upgrade, and fail loud when recovery does not land on deployed. Also skip helm-managed workloads in rollback_workloads: rollout undo there would step back to the revision --atomic just escaped. |
||
|
|
7c4843c88c |
fix(loki): unblock gateway rollout on a single node
Chart default is required podAntiAffinity on hostname plus RollingUpdate 25%/25%, which is maxUnavailable=0 at replicas=1: the new pod stays Unschedulable while the old one lives, and the old one never leaves while the new one is not Ready. Null the affinity (an empty map deep-merges with the default and keeps the rule) and set maxSurge/maxUnavailable to 1. |
||
|
|
f9e4623ade |
fix(k8s): size the remaining workloads against measured use
Finishes the sizing pass over every workload the deploy actually manages. Each request is at or above the container's p95 over the last seven days, so nothing is sized below what it is known to use, and each limit is between 1.6x and 5x the observed max, which is the figure that decides whether a burst gets an OOMKill. Some of these go up, and that is the point. adguard was holding 975M against a 500Mi request and netbox 962M against 512Mi, so both sat permanently above their own request and were standing eviction candidates on a node that has about 300M of headroom. Raising a request costs scheduler room; leaving it low costs the pod its place in the queue when the node gets tight. Others come down. loki ran with a 2Gi limit on 224M, gitea 1.5Gi on 305M, the authentik worker 1Gi on 305M, and a tail of single-purpose pods -- redis, glance, the two homepages, cfddns, session-keeper, the netbird dashboard, the loki gateway -- each reserved 4x to 16x more than they have ever touched. prometheus gets the opposite treatment: 768Mi/2560Mi, above its p95, because it compacts its TSDB in place and that is a burst worth budgeting for rather than throttling. Two of these limits are close enough to the observed max to be worth watching rather than trusting: adguard at 1.5x, and its DNS cache grows monotonically, so the ceiling is a date, not a margin. That was true before this change too; the pod sizing does not fix it and the cache needs bounding. CPU limits are untouched throughout. Leaving postgres alone as well: it sits in an uncommitted file that belongs to other work in progress. Verified: every request is at or above p95 and every limit above the observed max across all 74 containers, and 16/16 local gates pass. |
||
|
|
2ad4fa1b82 |
chore(portainer): stop deploying a container manager nothing routes to
Portainer had been running for 111 days with a 512Mi request and a 2Gi limit against 50M of measured use, on a node that is short of memory. It is a UI over the Docker socket; nothing in the repo or the cluster depends on it. The marker goes, not the manifests. `portainer/k8s/active` is what puts these files in the deploy's manifest set, so without it the next push leaves the namespace alone and the manifests stay on disk for a one-command return. This also matters for the smoke stage: that host list is built from the active directories, so `portainer.forust.xyz` leaves it and the new router check does not go looking for a route to a service we just retired. In the cluster the Deployment, the Service and both IngressRoutes are deleted. The routes go first: leaving an IngressRoute behind a deleted Service keeps a Traefik router pointing at nothing, which answers 502 while looking perfectly healthy to the stage that just started checking for routers. Deliberately kept, so this is reversible rather than destructive: the namespace, the 2Gi `portainer-data-pvc` and both Certificates stay. Re-enabling is `git checkout HEAD~1 -- portainer/k8s/active` plus an apply, and no Let's Encrypt quota is spent reissuing the production certificate. `glance` still links to `portainer.forust.xyz` and that tile will now be a 404. Left alone on purpose. The Cloudflare record is manual and cfddns only ever creates records, so `portainer.forust.xyz` keeps resolving until it is removed in the dashboard, same as `dockmon.forust.xyz`. Verified: no Traefik router matches portainer any more, all 19 hosts left in the smoke list still have a router, and 16/16 local gates pass. |
||
|
|
2a4f215546 |
fix(k8s): bring the over-reserved memory limits down to measured use
Six pods reserved far more memory than they have ever touched. uptime-kuma held a 3Gi limit against 469M of measured p95, metube 2Gi against 72M, convertx 1.5Gi against 85M, netbird-server 1Gi against 97M, searxng 700Mi against 134M and bentopdf 700Mi against 4M. Every one of them is a ceiling the scheduler counts against the node while the memory sits unused. Requests move down with the limits but never below the measured p95, so none of these becomes an eviction candidate as a side effect of being right-sized. The limits keep between 2.2x and 11.6x over the observed max, which is the figure that decides whether a pod gets OOM-killed during a burst. Net effect across the six: requests -557M, limits -4.6Gi, all of it ceiling that was never in use. This is the first change that actually gives memory back. CPU limits are left exactly as they were. They were not part of the sizing pass, they are not being hit on a node sitting at 5% CPU, and removing them is a separate decision from moving memory. Verified: each limit is above the container's own observed max and each request is above its p95, and 16/16 local gates pass. |
||
|
|
16aaeb60c1 |
fix(k8s): set requests and limits on the pods that shipped with neither
Eighteen containers had no memory limit at all, so nothing on the node could bound them. Three of the values files even claimed to set resources: Helm does not complain about a key it does not recognise, so the block sat there looking like a limit while the pod ran unbounded. alloy is the one that mattered. The chart reads `alloy.resources`; the file had `controller.resources`, so the DaemonSet that tails every pod log on the node shipped with nothing at all. `kubeStateMetrics` is the same trap in a different shape -- that is the condition key, the values live under `kube-state-metrics` -- and `configReloader` in the alloy chart sits at the top level rather than under `alloy`. Each one is verified by rendering the chart and reading the resources back off the containers, because a values key that is ignored looks exactly like one that works. reloader turned out to be set and still wrong: 64Mi request against a measured p95 of 73M, so the pod ran permanently above its own request and stayed a standing eviction candidate. That is the pod that restarts every other pod, so it is the last one that should be evicted. Raised to 96Mi. Requests are set at p95 throughout, grafana, playwright and alloy included. Left at the values first proposed they would have sat below their own p95 and queued for eviction ahead of everything smaller. CPU limits are deliberately absent: the node is I/O bound at 5% CPU, and CFS throttling would turn disk wait into runnable-throttled, which is the failure mode that took the node down. The prometheus and alertmanager configReloader sidecars are left open: chart 86.2.3 does not template the key, so reaching those two containers needs a postRenderer. Verified: all four charts render with the resources landing on the intended containers, and 16/16 local gates pass. |
||
|
|
a5409edbf2 |
fix(deploy): fail the smoke stage when Traefik has no route for a host
ci / lint-compose (push) Successful in 5s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
renovate-ci / validate-renovate (push) Successful in 1m15s
ci / test-frontend (push) Successful in 14s
ci / validate (push) Successful in 4s
ci / test-backend (push) Failing after 13m28s
ci / build (push) Skipped
The smoke stage treats any HTTP response as proof the service is serving, which is right -- a 302 to a login or a 404 from a path the app does not serve still means the chain is intact. But a 404 is not evidence of that on its own: a router Traefik refused to build answers with exactly the same 404 and nothing behind it. That is not hypothetical. The crowdsec bouncer is a plugin, and when Traefik cannot fetch it at startup it disables the plugin without failing, then drops every router whose chain referenced it. Sixteen routes answered 404 and the stage printed `ok` for all sixteen, because a dropped router and an unserved path are indistinguishable from outside. The Kubernetes objects cannot tell us either: the IngressRoute is still sitting there looking healthy, the router Traefik built from it is simply not there. So ask Traefik. api.insecure is already on for the internal entrypoint and the router list says which hosts it matches right now. Every probed host has to appear in that list. HTTP routers only -- the TCP ones match on a HostSNI wildcard and the UDP ones carry no rule at all, both selected by entrypoint and port, so neither can answer the question. A router mid-rollout is legitimately absent for a moment, so the list is re-read twice over 20s; a plugin that failed to load stays absent and waiting cannot rescue it. An unreadable router list fails the stage rather than skipping the check, since a check that cannot run is not a passing check. Verified against the live cluster: all 23 public routes have a router and the stage passes. With gitea, grafana and uptime removed from that list the probes still answer and the stage fails on exactly those three. |
||
|
|
c70d2db3a1 |
ci: deploy the image the commit built, not whatever the tag points at
Every service tracked the mutable `:prod` tag, so a deploy applied whatever that tag happened to name at the time rather than the commit it was deploying. A rollback had no way to state what it was rolling back to, and two deploys of one commit could land different images. CI now publishes an immutable `sha-<commit12>` tag beside `:prod` on main, and re-tags it for every image a push did not rebuild. That re-tag copies the manifest list, so no layer moves. The deploy resolves the immutable tag to a digest and pins the workload to it, and only falls back to the moving tag when the immutable one cannot be resolved -- which it says out loud, because that fallback is the deploy quietly ceasing to be reproducible from its own commit. The image list comes out of the tree with git grep rather than being written out a second time, so adding a service no longer means keeping two lists in step. build also gains the three jobs it was skipping -- scan-deps, test-backend, test-frontend -- so a change that breaks them cannot be tagged at all. The two run blocks where a mid-loop failure was survivable now run under set -euo pipefail: the build loop and the service detector both carried on past an error and could report a green build having produced nothing. The registry password moves from run: substitution into an env: block. A quote, a backtick or a $(...) in the password is parsed as shell before the command ever runs, and a login that failed that way looked exactly like a build that failed. The apply and verify timeouts stay at 45 and 30 minutes. The comments now record the arithmetic that says so rather than leaving the numbers to be raised on the next scare: three no-op helm upgrades run 3-5 minutes, one broken release is a single 10 minute rollback because the loop aborts on the first failure, and the apply loop itself is about a minute. That is roughly 15 minutes of work against a 45 minute budget. verify is 32 workloads at 8 wide -- four waves of 300 seconds, 20 minutes -- which leaves room for two serial rollbacks, and only becomes derivable at 45 once rollback_workloads is parallelised. |
||
|
|
24dd82e801 |
fix(crowdsec): make the bouncer trust the mobile range independently
The parser-stage whitelist already covers 84.245.64.0/18, so an event from the phone is dropped before it reaches a bucket and no decision is ever created for it - confirmed against 72h of traefik access logs, where the phone shows up as 84.245.120.147, inside that /18. But that left the bouncer's own ClientTrustedIPs without the range, so the guarantee rested on a single config. If the parser whitelist ever stops matching, a ban would be created and then served against the phone, which is the one thing that must not happen: the address belongs to a carrier, so it comes back to us by rotation and a 4h ban is not survivable from the device. ClientTrustedIPs bypasses the bouncer and the decision cache entirely, so repeating the range there holds even if a decision exists for any reason. All nine parser-stage ranges are now mirrored in the bouncer, and the bouncer has no range the parser stage does not know about. The file header now records that this middleware must be applied together with a traefik restart. Applying it alone wedges the plugin: in stream mode handleStreamTicker runs over package-level globals that no reconfiguration stops, so every route referencing the middleware answers 404 with 'invalid middleware crowdsec-crowdsec-bouncer@kubernetescrd' until the pod is replaced. Re-applying the prior config does not recover it and the config is not the cause - NewChecker is a plain net.ParseCIDR and cannot fail on a valid range. That cost 21 routes down before the restart requirement was found; recovery is a pod replace, ~35s. Verified live: middleware applied, traefik restarted, 17 of 20 hosts serving (the three exceptions are unchanged and unrelated - searxng returns its own 429, checkmk is down with 503, and one host is local-only), zero invalid-middleware errors, and the LAPI still shows /v1/decisions/stream polls at the 15s interval. |
||
|
|
11e92fdf4e |
fix(deploy): bound ssh hangs and retry the stage on transport loss
A connection that died silently used to hang until the job timeout, and the stage was never re-run. One flaky TCP session cost a whole 45-minute apply, and the symptom - a job that stops mid-output with no error - is what made the last few deploy failures expensive to read. ServerAliveInterval/CountMax cap how long a dead peer goes unnoticed at ~60s, ConnectTimeout caps setup. Only exit 255 - ssh's own transport failures - is retried, up to three attempts with a growing gap. A stage that fails on its own merits exits with the remote's status, so a real failure surfaces its own log immediately instead of being repeated three times over 45 minutes. The stages are declarative applies, so re-running one that had already committed is harmless. The stage environment now goes through `env` as separate argv entries rather than one interpolated string, so nothing in REPO, DEPLOY_SHA or DEPLOY_SNAPSHOT_DIR is re-split by the remote shell. Verified against a stubbed ssh: clean run attempts once, a single transport failure recovers on attempt 2 and exits 0, three failures give up preserving 255, and a stage failing with 1 or 7 attempts once and passes the code through unchanged. Also records why USERBOT_IMAGE stays on the prod tag: render_pinned rewrites only plain `image:` lines, and this ref is what the panel injects into the per-instance Deployments it creates, so those instances track the tag rather than the panel's own resolved digest. The two panel-created instances currently in the cluster are digest-pinned, so the panel does accept one either way; the tag is the choice, not a limitation. |
||
|
|
b1f98fc148 |
Revert "fix(compose): stop pointing at the tag the build dropped" for dtek_notif
dtek-notif is being picked up again, so leave its compose alone. It is also the one image not rebuilt since the build dropped :latest, so :prod does not exist for it yet - unlike the other six, which resolve. Keeping this file out of the change also keeps dtek_notif/* out of the build's changed-service detector, so pushing does not build an image nobody asked for. |
||
|
|
9b91b5847e |
fix(compose): stop pointing at the tag the build dropped
These five stacks still asked for :latest, but the build stopped pushing it - on main it only pushes main and prod, on dev only dev. Every one of these images therefore resolved only because the registry still had a stale :latest from before that change, and the next time one of them was actually built the reference would have dangled. Named for dtek-notif: of the seven, only that one still resolved at :prod, because it is the only image not rebuilt since the build dropped :latest - and it is the only one of these five stacks the deploy does not manage (no `active` marker, and its file is docker-compose.yaml, which the COMPOSE_STACKS glob does not even match). Touching dtek_notif/docker-compose.yaml matches dtek_notif/* in the build's changed-service detector, so pushing this builds it and publishes :prod for it too. Verified the other six resolve at :prod in the registry. |
||
|
|
892790822d |
fix(crowdsec): stop the 403 loop that banned our own VPS and runner
Chasing why apply-k8s kept dying mid-run turned up a self-inflicted
ban loop. 585 of 586 LePresidente/http-generic-403-bf alerts in the LAPI
came from 193.181.211.79 - our own VPS - POSTing
/management.ManagementService/GetServerKey, i.e. the NetBird client's
own management call. netbird-server had never seen a single one of them,
so the 403 was not NetBird's: the bouncer was rejecting the request
before it got there. A banned peer keeps retrying, each retry is another
403, and the scenario turns five 403s in ten seconds into a 4h ban, so
the loop kept re-arming the ban it was serving. The hourly janitor step
that deleted those decisions hourly was masking all of it.
* netbird/k8s/ingress.yaml - drop the bouncer from the mesh API routes
(gRPC-gateway management, signal, relay, /api, /oauth2). A ban there
locks a peer out of the network it needs to reach anything else, and
those endpoints authenticate by NetBird token, not by a login form.
The dashboard keeps the bouncer; it is a real login surface.
netbird-local was already exempt, so this makes prod match.
* crowdsec-middleware.yaml - CrowdsecMode stream instead of live. live
blocked on GET /v1/decisions per request, so a burst saturated the
LAPI and the plugin 403'd IPs that were never banned. v1.3.3 ignores
UpdateMaxFailure in live, so fail-open is only reachable in stream;
-1 now means an unreachable LAPI degrades to "no protection" rather
than "every site 403". 15s poll instead of the 60s default, because
the runner shares one public IP with the house.
Also corrects the key name: HTTPTimeoutSeconds, not
CrowdsecLapiTimeout, which never existed and was being silently
dropped, leaving the 10s default. Back at 10, not the 2s
|
||
|
|
f7cd75d65e |
fix(crowdsec): stop the bouncer from 403-ing our own deploys
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 32s
ci / lint-ruff (push) Successful in 0s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / build (push) Successful in 1s
Every bouncer-protected request blocks on a synchronous GET /v1/decisions against the LAPI, and the plugin fails CLOSED with 403 when that lookup exceeds its timeout. The LAPI was capped at 400m/500Mi on a single replica, so idle lookups measured 1.3-7.4s and a deploy burst pushed them past the fork's implicit 10s default: gitea answered 403 for ten seconds straight, and containerd turned those 403s on gcr.forust.xyz into ErrImagePull/ImagePullBackOff on freshly rolled pods. Three changes, plus the 403 feedback loop that made it sticky: * gitea/k8s/ingress.yaml - drop the bouncer from the /v2 registry route. A deploy fires hundreds of parallel authenticated OCI requests (runner Action API, manifest inspect per own image, containerd pulls, smoke probes) and scanners gain nothing from a registry that already does its own token auth. The web UI route keeps the bouncer. * crowdsec-values.yaml - LAPI to 1500m/1Gi. Replicas stay at 1 on purpose: LAPI is stateful (BoltDB plus credentials on two RWO PVCs), so a second replica would corrupt the decision store. * crowdsec-middleware.yaml - CrowdsecLapiTimeout: 2s instead of the implicit 10s, so a slow LAPI costs a fast 403 rather than a 10s hang. LePresidente/http-generic-403-bf then banned us for our own 403s: five POSTs answered 403 within 10s earn a 4h ban, and the hairpin-NAT address 192.168.88.1 that the Gitea Actions runner presents to Traefik is not covered by the home-dynamic-IP whitelist. That scenario cannot be dropped per-scenario - it is baked into the hub item crowdsecurity/http-generic-bf v0.9, and disabling the whole base-http-scenarios collection would cost ~40 useful detections. So the janitor now deletes its decisions hourly and a new forust/lan whitelist postoverflow covers 192.168.88.0/24. First janitor run removed 165 decisions; none of the remaining ones are local. Measured after: gitea 200 in 25-148ms (was 403 at 10001ms), a 60-way parallel burst all 200 with a 422ms max, gcr /v2 back to its 401 auth challenge, LAPI at 60m CPU with no throttling. |
||
|
|
6a9a460769 |
fix(deploy): let a pinning failure explain itself
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 6s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 26s
ci / build (push) Successful in 1s
registry_digest was written to return an empty string for a ref the registry does not have, so render_pinned could print "cannot resolve <ref>" and stop. It could not do that. Every caller runs under set -euo pipefail, pipefail reports the rightmost non-zero stage, and the failed docker manifest inspect made the assignment itself fail, which set -e turns into an immediate exit. render_pinned therefore died silently on the first unresolvable ref: nothing on stderr, nothing on stdout, exit 1. The apply loop piped that empty stream into kubectl, so the whole deploy stopped with "error: no objects passed to apply" - kubectl guessing at a cause, with the actual reason nowhere in the log. The missing message is the reason the |
||
|
|
ac0f845da6 |
fix(deploy): stop the panel from pointing at a tag the build dropped
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 11s
renovate-ci / validate-renovate (push) Successful in 12s
ci / validate (push) Successful in 3s
ci / build (push) Successful in 37s
USERBOT_IMAGE named userbot:latest, but the build stopped pushing latest when images became :prod. render_pinned resolves every gcr.forust.xyz reference it finds in a manifest, not just the ones it rewrites, so that env var made the whole apply abort with "cannot resolve userbot:latest; applying nothing". Naming it :prod also makes the build job rebuild userbot, which is what restores forust/userbot:prod - the tag is listed in the registry but its manifest 404s, so the image: field in userbots.yaml had nothing to resolve either. Found by resolving every digest apply-k8s will need before spending a run on it: 6 of 8 resolved, and both failures were userbot. |
||
|
|
f54589a05c |
fix(deploy): put the snapshot where the deploy user can write it
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / test-backend (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 7s
renovate-ci / validate-renovate (push) Successful in 29s
ci / build (push) Successful in 1s
The first deploy to actually run died on its very first action, and the error the other job reported was only the consequence. DEPLOY_SNAPSHOT_DIR defaulted to /var/backups/homelab-deploy. The deploy is unprivileged, and this Arch host has no /var/backups at all, so snapshot_dir's mkdir -p had to create it under root-owned /var and got Permission denied. It refused to go on, which is exactly what the guard is for, so no workload was touched - but the verify job then found no pointer and could only say to go look by hand. Defaulting to the deploy user's own XDG state directory fixes it with no root and no setup step, and keeps the guard: an unwritable snapshot dir still stops the deploy before the first apply. ssh-run.sh now forwards DEPLOY_SNAPSHOT_DIR too, so the path is overridable without editing the library. Verified on the workstation as the unprivileged user: pointer published, commit recorded, 71 workload generations and three helm releases captured, and the stale-pointer refusal still works. |
||
|
|
f49d91b63d |
Update deploy-lib.sh
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 0s
|
||
|
|
aafa74b70a |
chore: give local verification tooling a gitignored home
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 10s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 3s
renovate-ci / validate-renovate (push) Successful in 25s
ci / build (push) Successful in 1s
The pinned CI tools and the pre-push verification script were living in a temp directory that did not survive. tmp/ is now ignored, so ruff (which respects .gitignore by default) and the git ls-files globs the lint jobs use both skip it without needing to be told. |
||
|
|
4f74fe1778 |
ci: stop inheriting the runner's python and node
The runner executes jobs on the host rather than in a container, so a
workflow that says 'python3' or 'npm' is really saying 'whatever this
machine happens to have today'. Both of the test jobs added in
|
||
|
|
a5d384a4d8 |
feat(deploy): probe every active service after a deploy, rollouts included
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Failing after 13s
ci / test-backend (push) Failing after 10s
ci / test-frontend (push) Failing after 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 59s
ci / build (push) Successful in 1m50s
verify-k8s watches rollouts, which reports that pods converged. It cannot tell
a converged pod from a serving one. A Service selector pointing at a port
nothing listens on, a 500 from the app itself, a Traefik route that stopped
matching, a pod that OOMKilled early enough to still count as Available for the
duration of the check -- all of those are green at the rollout level and broken
for whoever opens the URL.
So ask what users ask. A smoke stage probes the public route of every active
service and fails on a transport error, a 5xx, or a 000, which curl reports
when it exits cleanly and nothing replied. Everything else passes, including 4xx:
a 404 from a path the service does not serve and a 302 to a login both prove
Traefik matched the host, the Service resolved to a pod and the pod answered,
which is the whole claim being tested.
An empty host list is an error, not a pass. Zero names means the extraction
broke, and reporting a clean deploy off a broken grep is the failure mode this
job exists to catch.
It runs on always() and after verify-k8s rather than before it, because a
rollback is when a route most needs re-checking. It only skips when verify-k8s
did, which is when nothing was deployed at all.
Two things worth writing down, because both were wrong on the first pass:
Stripping comments before reading the routes is not optional. naio and xui are
still in the tree commented out, and a plain grep picks both up and then reports
two services as unreachable when nobody ever deployed them. The apex
forust.xyz also needs a filter that admits it, so /\.forust\.xyz$/ quietly
dropped the site root.
And the 5xx test was written as ${code%%[0-9]*} != 5, which is empty for every
three-digit code, so a 500 was reported as ok. A case glob on the leading digit
is what actually works.
23 routes answer today, in 1.6s. The 5xx branch is the one part no live service
here exercises, so it was checked by running the block over 200 through 599 and
000 rather than against a real response.
|
||
|
|
af9a22fea9 |
fix(converters): spell workstation right in the compose router label
The local bentopdf router matched Host(`pdf.wokstation.internal`), a hostname that resolves to nothing. converters/k8s next to it has always had the correct spelling, so the service is reachable in the path that actually deploys and this one is dormant -- converters/ has no active marker, so select_manifests skips the stack. It would only bite whoever switches the service back to Compose and then wonders why one of the three internal routes 404s. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
baedea504d |
ci: scan dependencies for new advisories and stop handing out a write token
Two things, both about not finding out late. No workflow declared `permissions`, so all eighteen jobs across the four workflows ran on a token with the default full repository scope. Every one of them only checks out code, and deploy reaches the cluster over SSH with the deploy key, and Renovate writes through its own bot PAT rather than the Actions token. So `contents: read` is all any of them needed. The panel image ships 15 known advisories and nothing was looking. Add a scan-deps job that fails on anything new, and record the eight current ones by ID in the workflow. It is a list rather than a baseline count so that the diff that accepts an advisory says so in words, and it lives in our workflow instead of the package manifest so a subtree sync from forust/userbot cannot quietly widen the exemption. Both halves were checked to fail on a regression, not just to pass today: removing one --ignore-vuln turns the Python step red, and dropping --audit-level to moderate turns the npm one red on the devalue advisory. npm audits production dependencies only. All seven findings in the full tree are build- or test-time: the esbuild advisory needs a vite dev server exposed to the internet, and nanoid's infinite loop needs a custom generator called with size 0, which postcss does not do. None are in the 91 kB bundle the panel serves, so failing on them would be noise that trains people to ignore the job. The starlette entries are the reason the job is not "fail on everything": fastapi 0.115.12 pins starlette<0.47.0 and the last four fixes need 0.49.1 through 1.3.1, so clearing them is a jump to fastapi 0.141.x and is upstream's call, not a drive-by. Four of the seven are reachable in principle, which the comment on the job sets out. The panel answers only on userbot.workstation.internal with no public route, which is what keeps those four from being an internet-facing DoS. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
30995ee009 |
ci: pin the last four linters instead of trusting the runner
prettier, ruff, yamllint and hadolint were the only CI tools still called bare, straight off whatever the runner happened to have installed. Pin them in tool-versions.env like the other three and install them the same way, so the versions Renovate moves are the versions CI runs. Each pinned version equals what is already on the runner, so this changes what CI does not at all today. It changes what CI does on a rebuilt runner: the pinned one gets installed over the drift. The four need four different mechanisms, which is why this is not one pattern: hadolint a bare binary per platform, like actionlint ruff, yamllint PyPI wheels, unpacked by uv prettier an npm tarball, unpacked by tar prettier is the awkward one. Its entry point requires ../package.json relative to its own real path, so copying the single file out -- which is what every other installer here does -- yields a module-not-found at the first run. It keeps its package directory in a versioned one next to a relative symlink, and the tarball ships bin/ without the exec bit, so that needs chmod too. hadolint's release names one platform uname-style and the other Go-style (x86_64 but arm64), which 404s on the first architecture if you assume otherwise. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
3a05d86e3e |
ci: run svelte-check, which was already a dependency with no script
svelte-check sits in devDependencies at ^4.7.3 and nothing in the repository ever invoked it, so the type errors it reports had no path to a human. Point a script at it and run it in the frontend job, and it is clean: 0 errors, 0 warnings. It shares the one `npm ci` with the test step. A second install would have doubled the slowest part of the job to learn exactly the same thing. The lockfile is untouched, because scripts are not part of what it pins. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c00a4724f5 |
ci: actually run the test suites that exist in the tree
The panel ships 25 pytest tests and 2 vitest tests. Nothing executed them:
there was no job, no local dev loop, and nothing that would have noticed when
one of them rotted. They pass, and they are 8 seconds of work, which is the
argument for having them.
Both jobs mirror how the image is built rather than how a developer would run
them by hand: `npm ci` because that is what the Dockerfile does, so the tree
under test is the tree that ships, and requirements-dev.txt through uv, which
is now pinned like the other CI tools.
The backend job runs `python -m pytest`, not bare `pytest`. The tests import
`app.*` relative to the backend directory, and only the `-m` form puts the
working directory on sys.path.
ruff format --check joins ruff check in the lint job. It needed
|
||
|
|
284e19ef88 |
ci: pin uv, the tool that builds the pytest venv
The panel backend has 25 pytest tests that no workflow has ever run. Making them run needs a throwaway virtualenv, and uv is what builds it in seconds against the pinned requirements-dev.txt. Installing it through install-ci-tools.sh rather than assuming it is on the runner keeps the version in one place, where the other three tools already live, and where the Renovate regex manager can move it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0ae0df7473 |
style: format the last 10 files that ruff format disagreed with
ruff.toml has declared `quote-style = "single"` and line-length 120 since the lint job landed, and 118 of 128 files follow it. The panel backend and the netbox configuration were written in black/prettier style instead, so a `ruff format --check` would have failed on them from the start. Bring them onto the style the repository already declares, which is what makes the check adoptable at all. Formatting only: apart from quote style the diff is multi-line expressions joined where they fit inside 120 columns. Both suites still pass afterwards (25 pytest, and ruff check is clean). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
fddd82704f |
ci: keep pull requests away from the production admission webhooks
`kubectl apply --dry-run=server` persists nothing, but it does execute the admission webhooks of the real API server. The validate job runs on pull_request with no branch guard, so anyone able to open a PR could run arbitrary manifest content through cert-manager and Traefik in production. Limit the step to pushes to main. A pull request loses nothing by it: only main is ever deployed, and this job has to complete successfully before the deploy workflow is allowed to start, so a bad CRD is still caught before anything reaches the cluster -- on the push instead of on the PR. The skip is announced rather than silent, so a missing server-side pass does not read as a pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0691536f28 |
fix(deploy): roll out our images by digest instead of a moving tag
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in
|
||
|
|
2b9e34ba4a |
fix(k8s): add readiness probes so a bad image cannot look healthy
Six of eight workloads had no readinessProbe, so a pod turned Ready the moment its process started. The verify job relies on `rollout status`, so it passed for images that crash-looped or served errors, which left the rollback safety net inert. Each probe targets the path the service is actually reached on: - homepages: / (verified 200) - error-pages: /404.html, the path Traefik's errorPages middleware requests. / returns 403 by design and would never pass. - webinar-checker: /health (verified 200). /metrics also answers, but it is a Prometheus endpoint, not a readiness signal. The two userbot deployments stay without probes: they expose no port and no session file, and the panel reaches Telegram through its own client. A truthful signal there needs a health endpoint in the app itself. |
||
|
|
30d2b83efe |
chore(reloader): manage the reloader release from the repository
Reloader has been running since 23 September and is what makes the reloader.stakater.com/auto annotation on a pod template do anything, but the repository only held a namespace. It was a release someone installed by hand, so it was invisible to review, invisible to Renovate, and one reinstall away from being silently dropped. Declaring it in HELM_RELEASES pins the chart version somewhere Renovate can update it, and the active marker means the namespace is applied before the upgrade instead of only existing as a side effect of the original install. The values file sets nothing the running release does not already do, apart from resource requests and limits, which the chart leaves empty. |
||
|
|
db7bccfd89 |
fix(deploy): restart workloads whose image tag moved past what they run
Our manifests pin images to `:latest`, so a rebuild leaves the pod template byte-identical. kubectl apply sees no change, creates no ReplicaSet and pulls nothing, and the cluster keeps serving the previous build. imagePullPolicy: Always does not help, because it only decides whether a pod that *is* starting pulls, and no pod ever starts. All eight workloads that consume an image from our own registry were affected. Three of them had been running code from 23 September, and the single hardcoded `rollout restart deployment/userbot-panel` covered one of the eight. Restarting everything unconditionally was not the answer either: that bounces healthy services on every deploy, error-pages included, and the brief window where nothing answers is exactly what error-pages exists to prevent. So compare what each workload actually runs against what the tag resolves to now, and restart only the ones that differ. When the tag still points at the running digest nothing happens, so a redeploy that changed no image is a no-op. Scope is the repository, deliberately. Five more workloads run our images but have no manifest here, and they are applied out of band. Walking the manifests rather than the cluster means this can never reach them. The digest is resolved for the node architecture. A multi-arch tag also carries `unknown/unknown` entries for the build attestation, and a pod's imageID is always the per-platform digest, so comparing the wrong entry would mark everything stale forever. Once a restart happens it bumps the generation, which is what makes the change visible to changed_workloads and therefore watchable and revertible by the verify stage. |
||
|
|
1505b638ce |
fix(deploy): verify and roll back in a separate job
verify_workloads ended on `[ -s "$failed_file" ]`, which is the opposite of what its own contract says. A non-empty file means something failed, so the function returned success exactly when a workload never came up, and failure when everything was fine. Every rollback was therefore skipped, and every deploy that changed anything ended red with an empty failure list and a bogus "Rolled back successfully". Worse, the check only ever ran at the end of stage_apply_k8s, inside the same process as the apply. A job killed by timeout-minutes, cancelled by a new push, or cut off by a dropped SSH connection never reached it, which is precisely when a rollback matters. The three helm upgrades alone can consume the whole 30-minute job budget, so that path was reachable. Verification now lives in its own job, gated on always(), so it runs whatever happened to the apply. The apply stage publishes its pre-apply snapshot through DEPLOY_SNAPSHOT_DIR/current before touching anything, and the verify stage picks it up from there. A snapshot whose recorded commit does not match the deploy is refused rather than trusted, so a stale pointer from an earlier run cannot make the rollback revert the wrong workloads. An unwritable snapshot directory now fails the deploy up front instead of silently continuing without a way back. cancel-in-progress becomes false for the same reason: cancelling a run kills the apply job and takes the verify job with it, which is the failure this change exists to prevent. Both applies are idempotent, so queueing costs little. The SSH key moves to a per-run directory removed on exit, and the deploy is pinned to the exact commit CI validated. |
||
|
|
7ce727bc8a |
chore(renovate): move config under renovate/ and validate it in CI
The config lived in renovate.json at the repo root while everything else Renovate-related sat under renovate/, and renovate/config.js was a second, unused source of truth. Both are gone: renovate/renovate.json is now the only config file. Because the CronJob in the cluster cannot read the repository, its ConfigMap carries an inlined copy of the config. That copy is generated, and sync-renovate-configmap.sh --check now fails the build when it drifts from the source file. The workflows also stop carrying a copy of the renovate/renovate image tag. They read it from renovate/k8s/cronjob.yaml, so the version validated in CI is the version that actually runs in the cluster. ci.yaml validates the config with renovate-config-validator, checks the generated ConfigMap, and kubeconforms the CronJob's own manifests. |
||
|
|
f22793e32e |
ci: lint workflows and shell scripts, validate k8s against the API server
Adds three lint jobs (actionlint, shellcheck, compose) and a server-side dry-run of the active manifests. Previously the only k8s check was kubeconform, which has no schemas for CRDs, so every IngressRoute, Certificate, PrometheusRule and Middleware was silently skipped. The server-side pass needs the live API server because that is the only place the real CRD schemas and the cert-manager / Traefik admission webhooks exist. It is scoped to services carrying a k8s/active marker, since dry-run needs the target namespace to exist. userbot/ is excluded from shellcheck: it is a git subtree, and linting upstream's scripts would let a routine subtree pull turn the deploy gate red on code we do not own. kubeconform, shellcheck and actionlint are now installed from pinned versions in tool-versions.env rather than picked up from the runner's PATH. The Compose helper is shared with the deploy workflow so both check the same file set the same way. |
||
|
|
4a8d4feea0 |
fix(netbox): run probes with curl directly, not python -c
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 20s
ci / build (push) Successful in 1s
deploy / preflight (push) Successful in 2s
deploy / validate (push) Successful in 1m53s
deploy / apply-k8s (push) Successful in 1m47s
deploy / apply-compose (push) Successful in 11s
python -c 'exec /usr/bin/curl ...' is shell syntax, not Python, so every probe raised SyntaxError and the pod never became ready. Verified curl against /login/ returns HTTP 200. |
||
|
|
1d9a85bef9 |
fix(edu-master): stop 20722d false alert on zeroed last_success gauge
ci / validate (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m52s
deploy / apply-k8s (push) Successful in 1m48s
deploy / apply-compose (push) Successful in 13s
checker.py initialises last_success to 0, so right after a pod restart `time() - last_success` equals the current epoch. The rule compared that against 300, went firing instantly, and humanizeDuration rendered the raw epoch delta as ~20722d. The last_run > 0 guard did not help because a run happens long before the first success. Guard the duration rule on last_success > 0, keeping the duration expression on the left of `and` so $value stays the real gap, and add a separate WebinarCheckerNeverSucceeded rule for the zeroed-gauge case so a checker that has never succeeded is still caught. |
||
|
|
b625d30568 |
ci(deploy): cancel superseded deploys on new push
deploy / preflight (push) Successful in 1s
deploy / validate (push) Successful in 1m53s
deploy / apply-k8s (push) Successful in 1m48s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 6s
ci / validate (push) Successful in 1s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 1s
deploy / apply-compose (push) Successful in 12s
Queued deploy runs were deploying origin/main at start anyway (preflight reset), so waiting runs duplicated the newest deploy instead of their own commit. Cancel them. |
||
|
|
d018a441af |
fix(deploy): move netbox from compose to k8s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 17s
deploy / preflight (push) Successful in 2s
ci / build (push) Successful in 1s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
deploy / validate (push) Canceled after 4s
compose.yaml is a local-only stand on 127.0.0.1:8000 per README; the live service runs in-cluster. The root active marker made apply-compose fail on gitignored .env vars. |
||
|
|
ecb254017d |
fix(searxng): use existing image tag 2026.9.25-12f8b6515
ci / validate (push) Successful in 2s
ci / build (push) Successful in 1s
deploy / apply-k8s (push) Successful in 2m23s
deploy / apply-compose (push) Failing after 8s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 18s
deploy / validate (push) Successful in 1m46s
2026.09.13-d4ce87c23 was never published (upstream tags month without leading zero); rollout stuck in ImagePullBackOff. Verified replacement tag exists on Docker Hub. |
||
|
|
41f18ea993 |
fix(deploy): move netbird from compose to k8s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 19s
ci / build (push) Successful in 2s
deploy / validate (push) Successful in 1m45s
deploy / apply-k8s (push) Successful in 1m52s
deploy / apply-compose (push) Failing after 24s
netbird runs in-cluster; the root active marker made apply-compose pick up netbird/compose.yaml and fail on gitignored .env vars. Drop the compose marker and enable k8s/active instead. |
||
|
|
91c344fe2c |
fix(monitoring): disable control-plane alerts and scrapes on k0s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 3s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 24s
deploy / validate (push) Successful in 1m42s
deploy / apply-k8s (push) Successful in 1m45s
ci / build (push) Successful in 2s
deploy / apply-compose (push) Failing after 9s
k0s runs kube-controller-manager, kube-scheduler and etcd inside its own process rather than as pods, so the chart's Services never get endpoints and the targets stay permanently absent. Drop the matching ServiceMonitors and their Down/HighCommitDurations rules; kube-proxy and kubelet do get endpoints on k0s and stay enabled. |
||
|
|
cab6ef2102 |
Merge branch 'feat/netbird'
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 2s
deploy / validate (push) Successful in 1m43s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 1s
deploy / apply-k8s (push) Successful in 3m15s
deploy / apply-compose (push) Failing after 41s
|
||
|
|
a2ff9515a3 |
fix(deploy): validate compose without workstation secrets
renovate-ci / validate-renovate (push) Skipped
deploy / validate (push) Skipped
ci / lint-prettier (push) Successful in 2s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
docker compose config required real values for gitignored .env files and secrets, so validate always failed on stacks with :? guards (netbird, netbox). Validate structure only via --no-interpolate, --no-env-resolution and --no-path-resolution, keeping normalization and consistency checks. |
||
|
|
62d39ee4f1 |
Merge pull request 'chore(deps): update netbirdio/dashboard docker tag to v2.93.0' (#53) from renovate/netbirdio-dashboard-2.x into main
ci / lint-dockerfiles (push) Successful in 1s
deploy / preflight (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 9s
deploy / validate (push) Failing after 2s
deploy / apply-k8s (push) Skipped
deploy / apply-compose (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / validate (push) Successful in 1s
ci / build (push) Successful in 1s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/53 |
||
|
|
f4df6d4437 |
Merge pull request 'chore(deps): update container patch updates' (#38) from renovate/container-patch-updates into main
renovate-ci / validate-renovate (push) Skipped
deploy / preflight (push) Successful in 3s
deploy / validate (push) Failing after 2s
deploy / apply-k8s (push) Skipped
deploy / apply-compose (push) Skipped
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / lint-prettier (pull_request) Successful in 3s
ci / lint-ruff (pull_request) Successful in 1s
ci / lint-yaml (pull_request) Successful in 2s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
ci / build (push) Skipped
ci / lint-dockerfiles (pull_request) Successful in 1s
renovate-ci / validate-renovate (pull_request) Successful in 21s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/38 |
||
|
|
b423119632 |
Merge pull request 'chore(deps): update renovate/renovate docker tag to v44.115.9' (#47) from renovate/renovate-renovate-44.x into main
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 3s
ci / build (push) Canceled after 0s
deploy / preflight (push) Successful in 3s
deploy / validate (push) Canceled after 0s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 1m30s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/47 |
||
|
|
6ad33d75e6 |
Merge pull request 'chore(deps): update prom/prometheus docker tag to v3.15.0' (#48) from renovate/prom-prometheus-3.x into main
ci / lint-prettier (push) Successful in 3s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
deploy / preflight (push) Successful in 2s
ci / build (push) Canceled after 0s
deploy / validate (push) Canceled after 0s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 24s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/48 |
||
|
|
b7a1835adb |
Merge pull request 'feat(netbird): add tailscale-alternative' (#50) from feat/netbird into main
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 1s
deploy / preflight (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / build (push) Canceled after 0s
deploy / validate (push) Canceled after 0s
deploy / apply-k8s (push) Canceled after 0s
deploy / apply-compose (push) Canceled after 0s
renovate-ci / validate-renovate (push) Successful in 21s
Reviewed-on: https://gitea.forust.xyz/forust/homelab/pulls/50 |