f196099491d307347663058cedc73f7ec6afdb88
18
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0859479c0f |
feat(ingress): replace traefik crowdsec plugin with firewall bouncer
ci / lint-compose (push) Successful in 9s
ci / lint-actionlint (push) Successful in 4s
ci / lint-shellcheck (push) Successful in 7s
ci / lint-prettier (push) Successful in 12s
ci / lint-ruff (push) Successful in 6s
ci / lint-yaml (push) Successful in 9s
ci / lint-dockerfiles (push) Successful in 5s
ci / validate (push) Successful in 5s
renovate-ci / validate-renovate (push) Successful in 7s
ci / build (push) Failing after 14m22s
Move L3 enforcement to the host firewall-bouncer (systemd, nftables): drop the Traefik plugin, its secrets volume and the crowdsec Middleware, remove bouncer refs from all IngressRoutes. Disable the http-generic-bf scenario (403-burst bans hurt legit automation under L3 enforcement). Add a Gateway API PoC for homepages prod and CrowdSec PrometheusRule alerts. |
||
|
|
6f4cd03f4b |
feat(deploy): serialize apply stages and wait for calm node
apply-k8s and apply-compose share a workstation flock so host docker churn never overlaps cluster churn. Each helm upgrade and the apply loop wait up to 10m for load <28 first, so a deploy never piles onto an already-hot node (the load-40/netbird-death/pending-helm spiral). |
||
|
|
436fd1ecae |
chore: extract userbot subtree to its own repo
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 2s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 17s
ci / build (push) Successful in 8s
userbot/ (bot plus panel) now lives at /home/forust/userbot as a clone of forust/userbot instead of a subtree in homelab. Cleans up the pipeline references that only existed for it: scan-deps, test-backend and test-frontend jobs, the userbot build matrix entries, the deploy panel hook, and the userbot-only pyrightconfig. Running cluster workloads are untouched; homelab just stops building and testing upstream's code. |
||
|
|
2cb06debc5 |
fix(deploy): retry registry lookups with a timeout
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / scan-deps (push) Successful in 19s
ci / test-backend (push) Successful in 8s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 14s
ci / build (push) Successful in 7s
A single blink of the registry failed render_pinned for the whole file and redded the apply stage. registry_digest now retries 3 times under a 25s timeout with a warning per attempt; empty still means unresolvable and callers report it by name as before. |
||
|
|
76f39da90c |
fix(ci): skip heavy jobs on renovate branches, automerge digest and patch
ci / lint-compose (push) Canceled after 0s
ci / lint-actionlint (push) Canceled after 0s
ci / lint-shellcheck (push) Canceled after 0s
ci / lint-prettier (push) Canceled after 0s
ci / lint-ruff (push) Canceled after 0s
ci / lint-yaml (push) Canceled after 0s
ci / lint-dockerfiles (push) Canceled after 0s
ci / scan-deps (push) Canceled after 0s
ci / test-backend (push) Canceled after 0s
ci / test-frontend (push) Canceled after 0s
ci / validate (push) Canceled after 0s
ci / build (push) Canceled after 0s
renovate-ci / validate-renovate (push) Canceled after 0s
Renovate branches only carry version/digest bumps, so scan-deps, test-backend, test-frontend and build just burn runner time on the box that also serves prod. Static checks and validate still run. Digest and patch updates automerge (playwright, helm and major rules below still override to no-automerge). Also replaces deprecated helm --atomic with --wait --rollback-on-failure. |
||
|
|
dde6b1c743 |
fix(deploy): recover helm releases from pending-* and skip helm-owned rollbacks
An --atomic upgrade whose own rollback never finishes leaves the release in pending-*, blocking every future run until a human rolls back (loki rev 18/21). Recover automatically before and after each upgrade, and fail loud when recovery does not land on deployed. Also skip helm-managed workloads in rollback_workloads: rollout undo there would step back to the revision --atomic just escaped. |
||
|
|
a5409edbf2 |
fix(deploy): fail the smoke stage when Traefik has no route for a host
ci / lint-compose (push) Successful in 5s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 3s
ci / lint-prettier (push) Successful in 2s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
renovate-ci / validate-renovate (push) Successful in 1m15s
ci / test-frontend (push) Successful in 14s
ci / validate (push) Successful in 4s
ci / test-backend (push) Failing after 13m28s
ci / build (push) Skipped
The smoke stage treats any HTTP response as proof the service is serving, which is right -- a 302 to a login or a 404 from a path the app does not serve still means the chain is intact. But a 404 is not evidence of that on its own: a router Traefik refused to build answers with exactly the same 404 and nothing behind it. That is not hypothetical. The crowdsec bouncer is a plugin, and when Traefik cannot fetch it at startup it disables the plugin without failing, then drops every router whose chain referenced it. Sixteen routes answered 404 and the stage printed `ok` for all sixteen, because a dropped router and an unserved path are indistinguishable from outside. The Kubernetes objects cannot tell us either: the IngressRoute is still sitting there looking healthy, the router Traefik built from it is simply not there. So ask Traefik. api.insecure is already on for the internal entrypoint and the router list says which hosts it matches right now. Every probed host has to appear in that list. HTTP routers only -- the TCP ones match on a HostSNI wildcard and the UDP ones carry no rule at all, both selected by entrypoint and port, so neither can answer the question. A router mid-rollout is legitimately absent for a moment, so the list is re-read twice over 20s; a plugin that failed to load stays absent and waiting cannot rescue it. An unreadable router list fails the stage rather than skipping the check, since a check that cannot run is not a passing check. Verified against the live cluster: all 23 public routes have a router and the stage passes. With gitea, grafana and uptime removed from that list the probes still answer and the stage fails on exactly those three. |
||
|
|
c70d2db3a1 |
ci: deploy the image the commit built, not whatever the tag points at
Every service tracked the mutable `:prod` tag, so a deploy applied whatever that tag happened to name at the time rather than the commit it was deploying. A rollback had no way to state what it was rolling back to, and two deploys of one commit could land different images. CI now publishes an immutable `sha-<commit12>` tag beside `:prod` on main, and re-tags it for every image a push did not rebuild. That re-tag copies the manifest list, so no layer moves. The deploy resolves the immutable tag to a digest and pins the workload to it, and only falls back to the moving tag when the immutable one cannot be resolved -- which it says out loud, because that fallback is the deploy quietly ceasing to be reproducible from its own commit. The image list comes out of the tree with git grep rather than being written out a second time, so adding a service no longer means keeping two lists in step. build also gains the three jobs it was skipping -- scan-deps, test-backend, test-frontend -- so a change that breaks them cannot be tagged at all. The two run blocks where a mid-loop failure was survivable now run under set -euo pipefail: the build loop and the service detector both carried on past an error and could report a green build having produced nothing. The registry password moves from run: substitution into an env: block. A quote, a backtick or a $(...) in the password is parsed as shell before the command ever runs, and a login that failed that way looked exactly like a build that failed. The apply and verify timeouts stay at 45 and 30 minutes. The comments now record the arithmetic that says so rather than leaving the numbers to be raised on the next scare: three no-op helm upgrades run 3-5 minutes, one broken release is a single 10 minute rollback because the loop aborts on the first failure, and the apply loop itself is about a minute. That is roughly 15 minutes of work against a 45 minute budget. verify is 32 workloads at 8 wide -- four waves of 300 seconds, 20 minutes -- which leaves room for two serial rollbacks, and only becomes derivable at 45 once rollback_workloads is parallelised. |
||
|
|
6a9a460769 |
fix(deploy): let a pinning failure explain itself
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 15s
ci / test-backend (push) Successful in 6s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 26s
ci / build (push) Successful in 1s
registry_digest was written to return an empty string for a ref the registry does not have, so render_pinned could print "cannot resolve <ref>" and stop. It could not do that. Every caller runs under set -euo pipefail, pipefail reports the rightmost non-zero stage, and the failed docker manifest inspect made the assignment itself fail, which set -e turns into an immediate exit. render_pinned therefore died silently on the first unresolvable ref: nothing on stderr, nothing on stdout, exit 1. The apply loop piped that empty stream into kubectl, so the whole deploy stopped with "error: no objects passed to apply" - kubectl guessing at a cause, with the actual reason nowhere in the log. The missing message is the reason the |
||
|
|
f54589a05c |
fix(deploy): put the snapshot where the deploy user can write it
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / test-backend (push) Successful in 7s
ci / lint-dockerfiles (push) Successful in 3s
ci / scan-deps (push) Successful in 15s
ci / test-frontend (push) Successful in 11s
ci / validate (push) Successful in 7s
renovate-ci / validate-renovate (push) Successful in 29s
ci / build (push) Successful in 1s
The first deploy to actually run died on its very first action, and the error the other job reported was only the consequence. DEPLOY_SNAPSHOT_DIR defaulted to /var/backups/homelab-deploy. The deploy is unprivileged, and this Arch host has no /var/backups at all, so snapshot_dir's mkdir -p had to create it under root-owned /var and got Permission denied. It refused to go on, which is exactly what the guard is for, so no workload was touched - but the verify job then found no pointer and could only say to go look by hand. Defaulting to the deploy user's own XDG state directory fixes it with no root and no setup step, and keeps the guard: an unwritable snapshot dir still stops the deploy before the first apply. ssh-run.sh now forwards DEPLOY_SNAPSHOT_DIR too, so the path is overridable without editing the library. Verified on the workstation as the unprivileged user: pointer published, commit recorded, 71 workload generations and three helm releases captured, and the stale-pointer refusal still works. |
||
|
|
f49d91b63d |
Update deploy-lib.sh
ci / lint-compose (push) Successful in 3s
ci / lint-actionlint (push) Successful in 2s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 3s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Successful in 17s
ci / test-backend (push) Successful in 7s
ci / test-frontend (push) Successful in 10s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 13s
ci / build (push) Successful in 0s
|
||
|
|
a5d384a4d8 |
feat(deploy): probe every active service after a deploy, rollouts included
ci / lint-compose (push) Successful in 4s
ci / lint-actionlint (push) Successful in 1s
ci / lint-shellcheck (push) Successful in 2s
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 2s
ci / scan-deps (push) Failing after 13s
ci / test-backend (push) Failing after 10s
ci / test-frontend (push) Failing after 1s
ci / validate (push) Successful in 2s
renovate-ci / validate-renovate (push) Successful in 59s
ci / build (push) Successful in 1m50s
verify-k8s watches rollouts, which reports that pods converged. It cannot tell
a converged pod from a serving one. A Service selector pointing at a port
nothing listens on, a 500 from the app itself, a Traefik route that stopped
matching, a pod that OOMKilled early enough to still count as Available for the
duration of the check -- all of those are green at the rollout level and broken
for whoever opens the URL.
So ask what users ask. A smoke stage probes the public route of every active
service and fails on a transport error, a 5xx, or a 000, which curl reports
when it exits cleanly and nothing replied. Everything else passes, including 4xx:
a 404 from a path the service does not serve and a 302 to a login both prove
Traefik matched the host, the Service resolved to a pod and the pod answered,
which is the whole claim being tested.
An empty host list is an error, not a pass. Zero names means the extraction
broke, and reporting a clean deploy off a broken grep is the failure mode this
job exists to catch.
It runs on always() and after verify-k8s rather than before it, because a
rollback is when a route most needs re-checking. It only skips when verify-k8s
did, which is when nothing was deployed at all.
Two things worth writing down, because both were wrong on the first pass:
Stripping comments before reading the routes is not optional. naio and xui are
still in the tree commented out, and a plain grep picks both up and then reports
two services as unreachable when nobody ever deployed them. The apex
forust.xyz also needs a filter that admits it, so /\.forust\.xyz$/ quietly
dropped the site root.
And the 5xx test was written as ${code%%[0-9]*} != 5, which is empty for every
three-digit code, so a 500 was reported as ok. A case glob on the leading digit
is what actually works.
23 routes answer today, in 1.6s. The 5xx branch is the one part no live service
here exercises, so it was checked by running the block over 200 through 599 and
000 rather than against a real response.
|
||
|
|
0691536f28 |
fix(deploy): roll out our images by digest instead of a moving tag
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in
|
||
|
|
30d2b83efe |
chore(reloader): manage the reloader release from the repository
Reloader has been running since 23 September and is what makes the reloader.stakater.com/auto annotation on a pod template do anything, but the repository only held a namespace. It was a release someone installed by hand, so it was invisible to review, invisible to Renovate, and one reinstall away from being silently dropped. Declaring it in HELM_RELEASES pins the chart version somewhere Renovate can update it, and the active marker means the namespace is applied before the upgrade instead of only existing as a side effect of the original install. The values file sets nothing the running release does not already do, apart from resource requests and limits, which the chart leaves empty. |
||
|
|
db7bccfd89 |
fix(deploy): restart workloads whose image tag moved past what they run
Our manifests pin images to `:latest`, so a rebuild leaves the pod template byte-identical. kubectl apply sees no change, creates no ReplicaSet and pulls nothing, and the cluster keeps serving the previous build. imagePullPolicy: Always does not help, because it only decides whether a pod that *is* starting pulls, and no pod ever starts. All eight workloads that consume an image from our own registry were affected. Three of them had been running code from 23 September, and the single hardcoded `rollout restart deployment/userbot-panel` covered one of the eight. Restarting everything unconditionally was not the answer either: that bounces healthy services on every deploy, error-pages included, and the brief window where nothing answers is exactly what error-pages exists to prevent. So compare what each workload actually runs against what the tag resolves to now, and restart only the ones that differ. When the tag still points at the running digest nothing happens, so a redeploy that changed no image is a no-op. Scope is the repository, deliberately. Five more workloads run our images but have no manifest here, and they are applied out of band. Walking the manifests rather than the cluster means this can never reach them. The digest is resolved for the node architecture. A multi-arch tag also carries `unknown/unknown` entries for the build attestation, and a pod's imageID is always the per-platform digest, so comparing the wrong entry would mark everything stale forever. Once a restart happens it bumps the generation, which is what makes the change visible to changed_workloads and therefore watchable and revertible by the verify stage. |
||
|
|
1505b638ce |
fix(deploy): verify and roll back in a separate job
verify_workloads ended on `[ -s "$failed_file" ]`, which is the opposite of what its own contract says. A non-empty file means something failed, so the function returned success exactly when a workload never came up, and failure when everything was fine. Every rollback was therefore skipped, and every deploy that changed anything ended red with an empty failure list and a bogus "Rolled back successfully". Worse, the check only ever ran at the end of stage_apply_k8s, inside the same process as the apply. A job killed by timeout-minutes, cancelled by a new push, or cut off by a dropped SSH connection never reached it, which is precisely when a rollback matters. The three helm upgrades alone can consume the whole 30-minute job budget, so that path was reachable. Verification now lives in its own job, gated on always(), so it runs whatever happened to the apply. The apply stage publishes its pre-apply snapshot through DEPLOY_SNAPSHOT_DIR/current before touching anything, and the verify stage picks it up from there. A snapshot whose recorded commit does not match the deploy is refused rather than trusted, so a stale pointer from an earlier run cannot make the rollback revert the wrong workloads. An unwritable snapshot directory now fails the deploy up front instead of silently continuing without a way back. cancel-in-progress becomes false for the same reason: cancelling a run kills the apply job and takes the verify job with it, which is the failure this change exists to prevent. Both applies are idempotent, so queueing costs little. The SSH key moves to a per-run directory removed on exit, and the deploy is pinned to the exact commit CI validated. |
||
|
|
a2ff9515a3 |
fix(deploy): validate compose without workstation secrets
renovate-ci / validate-renovate (push) Skipped
deploy / validate (push) Skipped
ci / lint-prettier (push) Successful in 2s
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
docker compose config required real values for gitignored .env files and secrets, so validate always failed on stacks with :? guards (netbird, netbox). Validate structure only via --no-interpolate, --no-env-resolution and --no-path-resolution, keeping normalization and consistency checks. |
||
|
|
648b354951 |
refactor(deploy): marker-driven selection (k8s/active, root active); enable headscale/nextcloud hybrid, disable dockmon/kener/downtify/n8n
ci / lint-prettier (push) Successful in 3s
ci / lint-ruff (push) Successful in 1s
ci / validate (push) Successful in 2s
deploy / preflight (push) Successful in 1s
renovate-ci / validate-renovate (push) Successful in 8s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / build (push) Successful in 1s
deploy / validate (push) Successful in 1m40s
deploy / apply-k8s (push) Successful in 1m41s
deploy / apply-compose (push) Successful in 13s
|