8.9 KiB
Homelab CI/CD
The native Gitea runners run on vps; production runs on workstation.
Main-branch checks and image builds use homelab:host. Pull request and
non-main checks use homelab-pr:host under a separate account without Docker
access. The homelab-pr runner is registered at User scope for forust, so
any repository under that account can schedule jobs that request this label.
Each runner accepts one job at a time; the build waits for every check to pass.
CI and deploy runs also show a summary with
the release SHA, image build or reuse results, deploy mode, selected services,
and image digests. Failed runs keep a summary of completed image builds, stage
results, apply results, and recorded Kubernetes recovery. The final deploy
summary is in the smoke job; earlier jobs show the state observed at that time.
Apply success is separate from health and recovery. Update the installed
workstation controller with setup-workstation.sh when no deploy is running.
No job images or Kubernetes credentials are needed on the VPS. Builds use one
pinned BuildKit helper container. CI and deploy are separate workflows.
Runner installation
Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl,
GNU tar/xz, flock and systemd using the host's package manager. Keep the existing
Gitea runner 3.0.2 binary at /usr/local/bin/gitea-runner.
From this checkout on the VPS:
sudo bash .gitea/runner/setup-runner.sh
The installer reuses /var/lib/gitea-runner/.runner and the existing service.
For a new host, install the same runner binary and register as gitea-runner
using the registration token interactively, label homelab:host, and working
directory /var/lib/gitea-runner; then rerun the installer. Tokens never belong
in this repository or command-line examples.
Pinned tools live in the runner user's ~/.cache/homelab-ci; CI repairs version
drift there. Installations are locked. Buildx uses only the homelab-ci builder,
pushes directly to the registry, and caps retained local cache at 1 GiB with a
2 GiB free-space target. This is not a hard limit on peak build disk usage.
Nothing runs docker system prune, removes unrelated images, or deletes volumes.
Pull request runner
Install the unprivileged host runner on the VPS:
sudo bash .gitea/runner/setup-pr-runner.sh
Get a registration token from the user Actions runner settings. Run the
installer in a terminal. It asks for the token without echoing it, registers the
runner as homelab-pr with label homelab-pr:host, then enables the service.
The work directory is /var/lib/gitea-pr-runner. Confirm that Gitea lists the
runner as User scope before merging the workflow change. An unmatched label can
fall back to the default job image.
Renovate PR validation uses pull_request_target, which reads the workflow from
the base branch. It checks out the PR head only after runner selection and runs
that code on homelab-pr. Keep this workflow read-only and do not add secrets.
The PR runner has a separate home and tool cache. Do not add it to the docker
group or give it access to /var/run/docker.sock. It runs repository code from
pull requests, so keep its registration and permissions separate from the
trusted homelab runner. This separates users and host permissions, but both
runners still share the VPS kernel and network. Use a disposable VM if PRs from
untrusted external authors must be fully isolated.
Workstation setup
As the existing SSH deploy user on workstation:
sudo loginctl enable-linger forust
bash .gitea/runner/setup-workstation.sh
The controller uses /srv/homelab as the persistent configuration tree and makes
a detached source worktree for each SHA. It never resets /srv/homelab, moves
local configuration, renames Compose projects, or changes volume names.
The installer records the current Kubernetes context and cluster UID in
~/.config/homelab-deploy/environment. Check these before installing.
Configure Gitea Actions Variables:
DEPLOY_HOST,DEPLOY_USER,DEPLOY_PORT: the existing VPS-to-workstation SSH endpoint.DEPLOY_KNOWN_HOSTS: workstation's verified SSH host key entry for that endpoint.AUTODEPLOY:falseinitially;trueenables deployment after successful main CI.
Keep DEPLOY_SSH_KEY, REGISTRY_USERNAME and REGISTRY_PASSWORD in Actions
Secrets. Legacy endpoint secrets remain accepted during migration. The Actions
token must have repository read and Actions read access for release downloads.
The deploy user's existing Docker registry authentication remains necessary.
Releases and deployment
CI publishes release-<full SHA> as a Gitea artifact with all three owned image
digests and build input fingerprints. Unchanged images are reused only from a
successful main CI artifact, never from :prod. Expired artifacts cause CI to
rebuild images; they block deployment until CI is rerun.
Run deploy from main with deploy_ref=main or a checked SHA:
full: required for the first baseline; reconcile all active components.changed: compare with the last fully successful production deploy.plan: validate configuration and show selection without changing production resources.refresh_images=true: explicitly refresh mutable third-party Compose tags.
The manual and automatic paths both require successful CI, a successful build
job and the exact SHA's release artifact. PRs cannot publish images or deploy.
Removed resources are reported and require explicit removal; no automatic prune.
Service dependencies are listed in .gitea/deploy-dependencies.json.
A workstation user systemd service holds the deploy lock across validation,
sequential apply, verification and smoke checks. SSH clients only submit/follow:
disconnecting or cancelling the Actions client does not kill production apply.
Retrying the same run ID does not start another apply. ExecStopPost recovers
interrupted runs before the unit finishes. Kubernetes rolls back to captured
revisions; configuration and persistent data are not reverted.
Status and recovery
--retry repeats failed verification and smoke checks, never apply. Recovery
keeps a failed deploy out of the successful baseline, even after rollback.
On workstation (replace the numeric ID with Actions run ID and attempt):
python3 ~/.local/lib/homelab-deploy/controller.py status 123-1
python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry
journalctl --user -u homelab-deploy@123-1
Runs live in ~/.local/state/homelab-deploy/runs. Compose stores resolved configs
with restricted permissions; these may contain credentials and must never be
uploaded as CI artifacts. Stage logs print the exact manual recovery command
using compose-before/<stack>.json, the original project directory and project
name. Compose does not automatically roll back, and Nextcloud AIO's child
containers remain managed by AIO. Preserve its own backups for data recovery.
The controller retains twenty successful/planned runs and preserves failures. Update the workstation dispatcher only when no deploy is running.
Validation and migration rollback
python3 -m unittest discover -s tests -v
bash .gitea/tests/deploy-validation.sh
Test on a separate namespace before the initial production full run. Check a
failed rollout, interrupted SSH and repeated run ID, and verify that an isolated
service change does not upgrade unrelated Helm releases or Compose stacks.
To roll back the migration, disable autodeploy and finish or recover the remote
run first. Restore the runner config/unit from .before-<timestamp> backups,
reload systemd and restart the runner. Restore the prior workflows from Git.
Production data and persistent volumes stay where they were. Do not remove run
state or Compose recovery files until recovery is confirmed.
Compose configuration recovery
Successful deploys save the complete resolved Compose configuration in
~/.local/state/homelab-deploy/compose-configs/. These files can contain secrets.
Keep them private and do not commit or upload them.
The next deploy uses this configuration for its recovery file, including old
commands, environment, mounts, ports, and removed services. The recovery command
uses --remove-orphans to remove services added by the failed deploy. It does
not restore volume data or reverse database migrations.
On the first run after this update, the controller can use the Compose file
from the previous successful run. If that file is absent, it reads the persistent
checkout and checks its service configuration hashes against existing containers.
A mismatch stops preflight. Restore the previous configuration before retrying.
Update the installed controller with bash .gitea/runner/setup-workstation.sh
from the reviewed checkout before using this change.
New namespaces are checked during preflight. Server validation of their resources runs after namespace creation and before application resources are applied. Plan mode does not create namespaces. A failed deferred check can leave an empty namespace; inspect it before removing it.