Set AUTODEPLOY=false under Settings -> Actions -> Variables and pushes stop deploying with no commit; unset means on. Manual Run workflow always bypasses the switch. The existing needs/skipped chaining propagates the skip through verify and smoke untouched.
userbot/ (bot plus panel) now lives at /home/forust/userbot as a clone of forust/userbot instead of a subtree in homelab. Cleans up the pipeline references that only existed for it: scan-deps, test-backend and test-frontend jobs, the userbot build matrix entries, the deploy panel hook, and the userbot-only pyrightconfig. Running cluster workloads are untouched; homelab just stops building and testing upstream's code.
A single blink of the registry failed render_pinned for the whole file and redded the apply stage. registry_digest now retries 3 times under a 25s timeout with a warning per attempt; empty still means unresolvable and callers report it by name as before.
Renovate branches only carry version/digest bumps, so scan-deps, test-backend, test-frontend and build just burn runner time on the box that also serves prod. Static checks and validate still run. Digest and patch updates automerge (playwright, helm and major rules below still override to no-automerge). Also replaces deprecated helm --atomic with --wait --rollback-on-failure.
An --atomic upgrade whose own rollback never finishes leaves the release in pending-*, blocking every future run until a human rolls back (loki rev 18/21). Recover automatically before and after each upgrade, and fail loud when recovery does not land on deployed. Also skip helm-managed workloads in rollback_workloads: rollout undo there would step back to the revision --atomic just escaped.
The smoke stage treats any HTTP response as proof the service is serving,
which is right -- a 302 to a login or a 404 from a path the app does not
serve still means the chain is intact. But a 404 is not evidence of that on
its own: a router Traefik refused to build answers with exactly the same
404 and nothing behind it.
That is not hypothetical. The crowdsec bouncer is a plugin, and when Traefik
cannot fetch it at startup it disables the plugin without failing, then drops
every router whose chain referenced it. Sixteen routes answered 404 and the
stage printed `ok` for all sixteen, because a dropped router and an unserved
path are indistinguishable from outside.
The Kubernetes objects cannot tell us either: the IngressRoute is still
sitting there looking healthy, the router Traefik built from it is simply not
there. So ask Traefik. api.insecure is already on for the internal entrypoint
and the router list says which hosts it matches right now.
Every probed host has to appear in that list. HTTP routers only -- the TCP
ones match on a HostSNI wildcard and the UDP ones carry no rule at all, both
selected by entrypoint and port, so neither can answer the question. A router
mid-rollout is legitimately absent for a moment, so the list is re-read twice
over 20s; a plugin that failed to load stays absent and waiting cannot rescue
it. An unreadable router list fails the stage rather than skipping the check,
since a check that cannot run is not a passing check.
Verified against the live cluster: all 23 public routes have a router and the
stage passes. With gitea, grafana and uptime removed from that list the
probes still answer and the stage fails on exactly those three.
Every service tracked the mutable `:prod` tag, so a deploy applied whatever
that tag happened to name at the time rather than the commit it was
deploying. A rollback had no way to state what it was rolling back to, and
two deploys of one commit could land different images.
CI now publishes an immutable `sha-<commit12>` tag beside `:prod` on main,
and re-tags it for every image a push did not rebuild. That re-tag copies
the manifest list, so no layer moves. The deploy resolves the immutable tag
to a digest and pins the workload to it, and only falls back to the moving
tag when the immutable one cannot be resolved -- which it says out loud,
because that fallback is the deploy quietly ceasing to be reproducible from
its own commit.
The image list comes out of the tree with git grep rather than being written
out a second time, so adding a service no longer means keeping two lists in
step.
build also gains the three jobs it was skipping -- scan-deps, test-backend,
test-frontend -- so a change that breaks them cannot be tagged at all. The
two run blocks where a mid-loop failure was survivable now run under
set -euo pipefail: the build loop and the service detector both carried on
past an error and could report a green build having produced nothing.
The registry password moves from run: substitution into an env: block. A
quote, a backtick or a $(...) in the password is parsed as shell before the
command ever runs, and a login that failed that way looked exactly like a
build that failed.
The apply and verify timeouts stay at 45 and 30 minutes. The comments now
record the arithmetic that says so rather than leaving the numbers to be
raised on the next scare: three no-op helm upgrades run 3-5 minutes, one
broken release is a single 10 minute rollback because the loop aborts on
the first failure, and the apply loop itself is about a minute. That is
roughly 15 minutes of work against a 45 minute budget. verify is 32
workloads at 8 wide -- four waves of 300 seconds, 20 minutes -- which
leaves room for two serial rollbacks, and only becomes derivable at 45 once
rollback_workloads is parallelised.
A connection that died silently used to hang until the job timeout, and
the stage was never re-run. One flaky TCP session cost a whole
45-minute apply, and the symptom - a job that stops mid-output with no
error - is what made the last few deploy failures expensive to read.
ServerAliveInterval/CountMax cap how long a dead peer goes unnoticed at
~60s, ConnectTimeout caps setup. Only exit 255 - ssh's own transport
failures - is retried, up to three attempts with a growing gap. A stage
that fails on its own merits exits with the remote's status, so a real
failure surfaces its own log immediately instead of being repeated three
times over 45 minutes. The stages are declarative applies, so re-running
one that had already committed is harmless.
The stage environment now goes through `env` as separate argv entries
rather than one interpolated string, so nothing in REPO, DEPLOY_SHA or
DEPLOY_SNAPSHOT_DIR is re-split by the remote shell.
Verified against a stubbed ssh: clean run attempts once, a single
transport failure recovers on attempt 2 and exits 0, three failures give
up preserving 255, and a stage failing with 1 or 7 attempts once and
passes the code through unchanged.
Also records why USERBOT_IMAGE stays on the prod tag: render_pinned
rewrites only plain `image:` lines, and this ref is what the panel
injects into the per-instance Deployments it creates, so those instances
track the tag rather than the panel's own resolved digest. The two
panel-created instances currently in the cluster are digest-pinned, so
the panel does accept one either way; the tag is the choice, not a
limitation.
registry_digest was written to return an empty string for a ref the registry
does not have, so render_pinned could print "cannot resolve <ref>" and stop.
It could not do that. Every caller runs under set -euo pipefail, pipefail
reports the rightmost non-zero stage, and the failed docker manifest inspect
made the assignment itself fail, which set -e turns into an immediate exit.
render_pinned therefore died silently on the first unresolvable ref: nothing on
stderr, nothing on stdout, exit 1. The apply loop piped that empty stream into
kubectl, so the whole deploy stopped with "error: no objects passed to apply" -
kubectl guessing at a cause, with the actual reason nowhere in the log. The
missing message is the reason the f54589a run looked like a network death.
Reproduced against the old file with a docker stub that always fails: identical
to the ac0f845 log. The || true makes the empty string reachable, and the apply
loop now names the file that failed instead of letting kubectl speak.
The first deploy to actually run died on its very first action, and the
error the other job reported was only the consequence.
DEPLOY_SNAPSHOT_DIR defaulted to /var/backups/homelab-deploy. The deploy
is unprivileged, and this Arch host has no /var/backups at all, so
snapshot_dir's mkdir -p had to create it under root-owned /var and got
Permission denied. It refused to go on, which is exactly what the guard
is for, so no workload was touched - but the verify job then found no
pointer and could only say to go look by hand.
Defaulting to the deploy user's own XDG state directory fixes it with no
root and no setup step, and keeps the guard: an unwritable snapshot dir
still stops the deploy before the first apply.
ssh-run.sh now forwards DEPLOY_SNAPSHOT_DIR too, so the path is
overridable without editing the library. Verified on the workstation as
the unprivileged user: pointer published, commit recorded, 71 workload
generations and three helm releases captured, and the stale-pointer
refusal still works.
The runner executes jobs on the host rather than in a container, so a
workflow that says 'python3' or 'npm' is really saying 'whatever this
machine happens to have today'. Both of the test jobs added in c00a472
were red on the first run for exactly that reason.
uv venv with no --python takes the first interpreter it finds, which is
the host's 3.14 here. pyrogram's sync.py calls the bare
asyncio.get_event_loop() that 3.14 no longer auto-creates, so three
tests died at collection. Pinned to 3.13, which is both what uv will
fetch when the host has none and what python:3.13-slim actually builds.
npm was missing outright, and turned up an hour later as npm 12 on node
26 - the same push, minutes apart. Neither is the panel image's
node:22-alpine, so node is now installed from the official tarball the
way the other tools are, at the image's major. npm is checked by running
it rather than by looking it up, so a name that resolves to something
broken reads as not installed.
Renovate keeps NODE_VERSION in step with the Dockerfile's node: tag, and
the two are one grouped dependency: CI that tests on a different major
than it builds on is a gate that can pass over a real break.
verify-k8s watches rollouts, which reports that pods converged. It cannot tell
a converged pod from a serving one. A Service selector pointing at a port
nothing listens on, a 500 from the app itself, a Traefik route that stopped
matching, a pod that OOMKilled early enough to still count as Available for the
duration of the check -- all of those are green at the rollout level and broken
for whoever opens the URL.
So ask what users ask. A smoke stage probes the public route of every active
service and fails on a transport error, a 5xx, or a 000, which curl reports
when it exits cleanly and nothing replied. Everything else passes, including 4xx:
a 404 from a path the service does not serve and a 302 to a login both prove
Traefik matched the host, the Service resolved to a pod and the pod answered,
which is the whole claim being tested.
An empty host list is an error, not a pass. Zero names means the extraction
broke, and reporting a clean deploy off a broken grep is the failure mode this
job exists to catch.
It runs on always() and after verify-k8s rather than before it, because a
rollback is when a route most needs re-checking. It only skips when verify-k8s
did, which is when nothing was deployed at all.
Two things worth writing down, because both were wrong on the first pass:
Stripping comments before reading the routes is not optional. naio and xui are
still in the tree commented out, and a plain grep picks both up and then reports
two services as unreachable when nobody ever deployed them. The apex
forust.xyz also needs a filter that admits it, so /\.forust\.xyz$/ quietly
dropped the site root.
And the 5xx test was written as ${code%%[0-9]*} != 5, which is empty for every
three-digit code, so a 500 was reported as ok. A case glob on the leading digit
is what actually works.
23 routes answer today, in 1.6s. The 5xx branch is the one part no live service
here exercises, so it was checked by running the block over 200 through 599 and
000 rather than against a real response.
Two things, both about not finding out late.
No workflow declared `permissions`, so all eighteen jobs across the four
workflows ran on a token with the default full repository scope. Every one of
them only checks out code, and deploy reaches the cluster over SSH with the
deploy key, and Renovate writes through its own bot PAT rather than the
Actions token. So `contents: read` is all any of them needed.
The panel image ships 15 known advisories and nothing was looking. Add a
scan-deps job that fails on anything new, and record the eight current ones by
ID in the workflow. It is a list rather than a baseline count so that the diff
that accepts an advisory says so in words, and it lives in our workflow
instead of the package manifest so a subtree sync from forust/userbot cannot
quietly widen the exemption.
Both halves were checked to fail on a regression, not just to pass today:
removing one --ignore-vuln turns the Python step red, and dropping
--audit-level to moderate turns the npm one red on the devalue advisory.
npm audits production dependencies only. All seven findings in the full tree
are build- or test-time: the esbuild advisory needs a vite dev server exposed
to the internet, and nanoid's infinite loop needs a custom generator called
with size 0, which postcss does not do. None are in the 91 kB bundle the panel
serves, so failing on them would be noise that trains people to ignore the
job.
The starlette entries are the reason the job is not "fail on everything":
fastapi 0.115.12 pins starlette<0.47.0 and the last four fixes need 0.49.1
through 1.3.1, so clearing them is a jump to fastapi 0.141.x and is upstream's
call, not a drive-by. Four of the seven are reachable in principle, which the
comment on the job sets out. The panel answers only on
userbot.workstation.internal with no public route, which is what keeps those
four from being an internet-facing DoS.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
prettier, ruff, yamllint and hadolint were the only CI tools still called bare,
straight off whatever the runner happened to have installed. Pin them in
tool-versions.env like the other three and install them the same way, so the
versions Renovate moves are the versions CI runs.
Each pinned version equals what is already on the runner, so this changes what
CI does not at all today. It changes what CI does on a rebuilt runner: the
pinned one gets installed over the drift.
The four need four different mechanisms, which is why this is not one pattern:
hadolint a bare binary per platform, like actionlint
ruff,
yamllint PyPI wheels, unpacked by uv
prettier an npm tarball, unpacked by tar
prettier is the awkward one. Its entry point requires ../package.json relative
to its own real path, so copying the single file out -- which is what every
other installer here does -- yields a module-not-found at the first run. It
keeps its package directory in a versioned one next to a relative symlink, and
the tarball ships bin/ without the exec bit, so that needs chmod too.
hadolint's release names one platform uname-style and the other Go-style
(x86_64 but arm64), which 404s on the first architecture if you assume
otherwise.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
svelte-check sits in devDependencies at ^4.7.3 and nothing in the repository
ever invoked it, so the type errors it reports had no path to a human. Point a
script at it and run it in the frontend job, and it is clean: 0 errors, 0
warnings.
It shares the one `npm ci` with the test step. A second install would have
doubled the slowest part of the job to learn exactly the same thing.
The lockfile is untouched, because scripts are not part of what it pins.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The panel ships 25 pytest tests and 2 vitest tests. Nothing executed them:
there was no job, no local dev loop, and nothing that would have noticed when
one of them rotted. They pass, and they are 8 seconds of work, which is the
argument for having them.
Both jobs mirror how the image is built rather than how a developer would run
them by hand: `npm ci` because that is what the Dockerfile does, so the tree
under test is the tree that ships, and requirements-dev.txt through uv, which
is now pinned like the other CI tools.
The backend job runs `python -m pytest`, not bare `pytest`. The tests import
`app.*` relative to the backend directory, and only the `-m` form puts the
working directory on sys.path.
ruff format --check joins ruff check in the lint job. It needed 0ae0df7 to be
addable, since ten files disagreed with the style ruff.toml has always
declared.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The panel backend has 25 pytest tests that no workflow has ever run. Making
them run needs a throwaway virtualenv, and uv is what builds it in seconds
against the pinned requirements-dev.txt.
Installing it through install-ci-tools.sh rather than assuming it is on the
runner keeps the version in one place, where the other three tools already
live, and where the Renovate regex manager can move it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`kubectl apply --dry-run=server` persists nothing, but it does execute the
admission webhooks of the real API server. The validate job runs on
pull_request with no branch guard, so anyone able to open a PR could run
arbitrary manifest content through cert-manager and Traefik in production.
Limit the step to pushes to main. A pull request loses nothing by it: only
main is ever deployed, and this job has to complete successfully before the
deploy workflow is allowed to start, so a bad CRD is still caught before
anything reaches the cluster -- on the push instead of on the PR.
The skip is announced rather than silent, so a missing server-side pass does
not read as a pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`kubectl rollout undo` restores the previous ReplicaSet's pod template
verbatim. While that template names a tag, the rollback does not roll back
the image: the tag has already moved, so the reverted pod pulls the very
build that just failed and the cluster stays broken. The safety net added
in 1505b63 therefore could not recover from a bad image.
Pin the digest at apply time. A digest is not knowable when a manifest is
written, so render_pinned resolves it on the way into the cluster and the
digest is never committed. Git keeps a readable `:prod`, Renovate keeps
seeing exactly the manifests it saw before, and the previous revision of
each workload now holds the digest that was actually serving, so undo
restores those exact bytes.
imagePullPolicy is dropped from the manifests rather than set to
IfNotPresent: a reference that is not `:latest` already defaults to it, and
that is what the Kubernetes docs ask for alongside a digest.
An unresolvable image is fatal instead of a warning, because carrying on
would quietly apply a mutable tag again.
restart_stale_images keeps its comparison but is no longer how a rebuild
reaches the cluster -- the pinned template rolls out on its own now. What
is left is a drift check for hand-run `kubectl set image`, so it matches
the container by repository: a pod's status now reports `repo@sha256:...`
while the manifest still says `:prod`.
The build job stops pushing `:latest` altogether, which removes the tag
that a dev branch could otherwise move under a prod deploy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reloader has been running since 23 September and is what makes the
reloader.stakater.com/auto annotation on a pod template do anything, but the
repository only held a namespace. It was a release someone installed by hand,
so it was invisible to review, invisible to Renovate, and one reinstall away from
being silently dropped.
Declaring it in HELM_RELEASES pins the chart version somewhere Renovate can
update it, and the active marker means the namespace is applied before the
upgrade instead of only existing as a side effect of the original install.
The values file sets nothing the running release does not already do, apart from
resource requests and limits, which the chart leaves empty.
Our manifests pin images to `:latest`, so a rebuild leaves the pod template
byte-identical. kubectl apply sees no change, creates no ReplicaSet and pulls
nothing, and the cluster keeps serving the previous build. imagePullPolicy:
Always does not help, because it only decides whether a pod that *is* starting
pulls, and no pod ever starts.
All eight workloads that consume an image from our own registry were affected.
Three of them had been running code from 23 September, and the single hardcoded
`rollout restart deployment/userbot-panel` covered one of the eight.
Restarting everything unconditionally was not the answer either: that bounces
healthy services on every deploy, error-pages included, and the brief window
where nothing answers is exactly what error-pages exists to prevent. So compare
what each workload actually runs against what the tag resolves to now, and
restart only the ones that differ. When the tag still points at the running
digest nothing happens, so a redeploy that changed no image is a no-op.
Scope is the repository, deliberately. Five more workloads run our images but
have no manifest here, and they are applied out of band. Walking the manifests
rather than the cluster means this can never reach them.
The digest is resolved for the node architecture. A multi-arch tag also carries
`unknown/unknown` entries for the build attestation, and a pod's imageID is
always the per-platform digest, so comparing the wrong entry would mark
everything stale forever.
Once a restart happens it bumps the generation, which is what makes the change
visible to changed_workloads and therefore watchable and revertible by the
verify stage.
verify_workloads ended on `[ -s "$failed_file" ]`, which is the opposite
of what its own contract says. A non-empty file means something failed, so
the function returned success exactly when a workload never came up, and
failure when everything was fine. Every rollback was therefore skipped,
and every deploy that changed anything ended red with an empty failure
list and a bogus "Rolled back successfully".
Worse, the check only ever ran at the end of stage_apply_k8s, inside the
same process as the apply. A job killed by timeout-minutes, cancelled by
a new push, or cut off by a dropped SSH connection never reached it, which
is precisely when a rollback matters. The three helm upgrades alone can
consume the whole 30-minute job budget, so that path was reachable.
Verification now lives in its own job, gated on always(), so it runs
whatever happened to the apply. The apply stage publishes its pre-apply
snapshot through DEPLOY_SNAPSHOT_DIR/current before touching anything,
and the verify stage picks it up from there. A snapshot whose recorded
commit does not match the deploy is refused rather than trusted, so a
stale pointer from an earlier run cannot make the rollback revert the
wrong workloads. An unwritable snapshot directory now fails the deploy up
front instead of silently continuing without a way back.
cancel-in-progress becomes false for the same reason: cancelling a run
kills the apply job and takes the verify job with it, which is the failure
this change exists to prevent. Both applies are idempotent, so queueing
costs little. The SSH key moves to a per-run directory removed on exit,
and the deploy is pinned to the exact commit CI validated.
The config lived in renovate.json at the repo root while everything else
Renovate-related sat under renovate/, and renovate/config.js was a second,
unused source of truth. Both are gone: renovate/renovate.json is now the
only config file.
Because the CronJob in the cluster cannot read the repository, its
ConfigMap carries an inlined copy of the config. That copy is generated,
and sync-renovate-configmap.sh --check now fails the build when it drifts
from the source file.
The workflows also stop carrying a copy of the renovate/renovate image
tag. They read it from renovate/k8s/cronjob.yaml, so the version validated
in CI is the version that actually runs in the cluster.
ci.yaml validates the config with renovate-config-validator, checks the
generated ConfigMap, and kubeconforms the CronJob's own manifests.
Adds three lint jobs (actionlint, shellcheck, compose) and a server-side
dry-run of the active manifests. Previously the only k8s check was
kubeconform, which has no schemas for CRDs, so every IngressRoute,
Certificate, PrometheusRule and Middleware was silently skipped.
The server-side pass needs the live API server because that is the only
place the real CRD schemas and the cert-manager / Traefik admission
webhooks exist. It is scoped to services carrying a k8s/active marker,
since dry-run needs the target namespace to exist. userbot/ is excluded
from shellcheck: it is a git subtree, and linting upstream's scripts would
let a routine subtree pull turn the deploy gate red on code we do not own.
kubeconform, shellcheck and actionlint are now installed from pinned
versions in tool-versions.env rather than picked up from the runner's
PATH. The Compose helper is shared with the deploy workflow so both
check the same file set the same way.
Queued deploy runs were deploying origin/main at start anyway (preflight reset), so waiting runs duplicated the newest deploy instead of their own commit. Cancel them.
docker compose config required real values for gitignored .env files and secrets, so validate always failed on stacks with :? guards (netbird, netbox). Validate structure only via --no-interpolate, --no-env-resolution and --no-path-resolution, keeping normalization and consistency checks.
Deploy workflow uses git-tracked manifests, DISABLED flag and kustomize overlays; add webinar-checker metrics with ServiceMonitor and alerts; upgrade shared postgres to 17 with statuspage DB and probes/resources.
Ruff and yamllint ship in Arch repos, drop their container pulls. Prettier keeps the node container but mounts persistent npm cache. Hadolint and kubeconform stay containerized (AUR-only / hermetic pin).
Drop duplicated .github/workflows (helm blocks ported to .gitea deploy first). Add ci concurrency with cancel on branches, pin checkout to SHA and prettier to 3.9.8, registry layer cache for image builds. renovate-run gains config validation and dry-run input.
renovate-run workflow_dispatch runs pinned renovate via docker on self-hosted runner. Sync helm-values manager into config.js/configmap, bump compose and validator pins to 44.97.2.
- ensure_prerequisites runs on startup, not per-request; kube config
errors surface as 503 PanelError
- serialize provisioning with a lock; drop per-endpoint prereq checks
- guard SPA fallback against path traversal (relative_to)
- add backend tests for auth flow, k8s service, spa routing; ci comment
for legacy userbot deployments
services marked k8s/active are applied via kubectl; the rest via docker
compose. inactive services with k8s/ keep only routing manifests
(external Services, EndpointSlices, Ingresses) to reach docker backends.
headscale/nextcloud routing moved to k8s/routing/.
validations: compose config --quiet + kubectl apply --dry-run=client.
namespace manifests applied first. pull_policy:build stacks get
build+push before up so the registry image stays fresh.