Homelab CI/CD
The native Gitea runners run on vps; production runs on workstation.
Main-branch checks and image builds use homelab:host. Pull request checks use
homelab-pr in a Docker job container. CI PR checks use pull_request_target,
so Gitea loads the workflow from the trusted base branch. That event then runs
untrusted PR code, so the workflow must select homelab-pr before checkout and
must not expose secrets. The CI validation jobs grant only contents: read and
checkout the explicit PR head SHA with persist-credentials: false. Register
homelab-pr at repository scope so only this repository can schedule its jobs.
Each runner accepts one job at a time; the build waits for every check to pass.
CI and deploy runs also show a summary with
the release SHA, image build or reuse results, deploy mode, selected services,
and image digests. Failed runs keep a summary of completed image builds, stage
results, apply results, and recorded Kubernetes recovery. The final deploy
summary is in the smoke job; earlier jobs show the state observed at that time.
Apply success is separate from health and recovery. Update the installed
workstation controller with setup-workstation.sh when no deploy is running.
No job images or Kubernetes credentials are needed on the VPS. Builds use one
pinned BuildKit helper container. CI and deploy are separate workflows.
Runner installation
Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl,
GNU tar/xz, flock and systemd using the host's package manager. Keep the existing
Gitea runner 3.0.2 binary at /usr/local/bin/gitea-runner.
From this checkout on the VPS:
sudo bash .gitea/runner/setup-runner.sh
The installer reuses /var/lib/gitea-runner/.runner and the existing service.
For a new host, install the same runner binary and register as gitea-runner
using the registration token interactively, label homelab:host, and working
directory /var/lib/gitea-runner; then rerun the installer. Tokens never belong
in this repository or command-line examples.
Pinned tools live in the runner user's ~/.cache/homelab-ci; CI repairs version
drift there. Installations are locked. Buildx uses only the homelab-ci builder,
pushes directly to the registry, and caps retained local cache at 1 GiB with a
2 GiB free-space target. This is not a hard limit on peak build disk usage.
Nothing runs docker system prune, removes unrelated images, or deletes volumes.
Pull request runner
Install the PR container runner on the VPS:
sudo bash .gitea/runner/setup-pr-runner.sh
Create a runner registration token from this repository's Actions runner
settings. Run the installer in a terminal. It asks for the token without
echoing it and registers homelab-pr with label
homelab-pr:docker://docker.gitea.com/runner-images:ubuntu-24.04-v26.09.01. Confirm that
Gitea lists the runner at Repository scope. The service runs as
gitea-pr-runner; systemd grants that service access to the Docker socket with
SupplementaryGroups=docker. Keep the account itself out of the docker
group. The work directory is /var/lib/gitea-pr-runner.
The runner config disables privileged containers, forbids workflow volume mounts, and prevents the Docker socket from being mounted into job and action containers. Do not mount the runner home or its host-side tool cache into a job. The existing host-side cache is retained, but PR job containers cannot read it. An unmatched label can fall back to the default job image; check the registered label before enabling PR checks.
The installer reuses /var/lib/gitea-pr-runner/.runner when it exists. That
file keeps the registration scope assigned by Gitea. To move an existing
User-scoped runner to Repository scope, stop the service, remove the old runner
from Gitea, back up and remove that registration file, then run the installer
with a token created in this repository's Actions runner settings. Confirm the
new scope in Gitea before enabling PR checks.
The CI and Renovate workflows use pull_request_target, which reads the
workflow from the base branch. They select homelab-pr before checking out PR
code. The explicit head SHA and persist-credentials: false are mandatory:
without the latter, checkout can leave the job token in Git configuration.
Keep PR validation read-only and do not add Actions secrets. In the checked-in
workflows, only a push to main or a manual CI run on main can select the
trusted homelab runner. Gitea schedules jobs by matching runs-on labels; the
runner does not restrict jobs by event or branch. Keep Gitea's approval gate for
fork PR workflows enabled. Verify the live Gitea version and approval setting
before relying on this gate; the image tag in the repository does not prove the
version currently running. Before approving a fork workflow run, review all new
and changed workflow files: a PR-defined pull_request workflow can request
the homelab label. Automatic CI and Renovate PR checks use the trusted base
workflow and select only homelab-pr. The release and deploy jobs stay on the
trusted runner.
The runner service can access the host Docker daemon, but job and action containers do not receive its socket or arbitrary host mounts. The runner and job containers still share the VPS kernel and Docker daemon. A container escape can therefore affect the host and other workloads. This is container isolation, not VM isolation; use disposable VMs for PRs that require a separate kernel and Docker daemon.
Workstation setup
As the existing SSH deploy user on workstation:
sudo loginctl enable-linger forust
bash .gitea/runner/setup-workstation.sh
The controller uses /srv/homelab as the persistent configuration tree and makes
a detached source worktree for each SHA. It never resets /srv/homelab, moves
local configuration, renames Compose projects, or changes volume names.
The installer records the current Kubernetes context and cluster UID in
~/.config/homelab-deploy/environment. Check these before installing.
Configure Gitea Actions Variables:
DEPLOY_HOST,DEPLOY_USER,DEPLOY_PORT: the existing VPS-to-workstation SSH endpoint.DEPLOY_KNOWN_HOSTS: workstation's verified SSH host key entry for that endpoint.AUTODEPLOY:falseinitially;trueenables deployment after successful main CI.
Keep DEPLOY_SSH_KEY, REGISTRY_USERNAME and REGISTRY_PASSWORD in Actions
Secrets. Legacy endpoint secrets remain accepted during migration. The Actions
token must have repository read and Actions read access for release downloads.
The deploy user's existing Docker registry authentication remains necessary.
Releases and deployment
CI publishes release-<full SHA> as a Gitea artifact with all three owned image
digests and build input fingerprints. Unchanged images are reused only from a
successful main CI artifact, never from :prod. Expired artifacts cause CI to
rebuild images; they block deployment until CI is rerun.
Run deploy from main with deploy_ref=main or a checked SHA:
full: required for the first baseline; reconcile all active components.changed: compare with the last fully successful production deploy.plan: validate configuration and show selection without changing production resources.refresh_images=true: explicitly refresh mutable third-party Compose tags.
The manual and automatic paths both require successful CI, a successful build
job and the exact SHA's release artifact. PRs cannot publish images or deploy.
Removed resources are reported and require explicit removal; no automatic prune.
Service dependencies are listed in .gitea/deploy-dependencies.json.
A workstation user systemd service holds the deploy lock across validation,
sequential apply, verification and smoke checks. SSH clients only submit/follow:
disconnecting or cancelling the Actions client does not kill production apply.
Retrying the same run ID does not start another apply. ExecStopPost recovers
interrupted runs before the unit finishes. Kubernetes rolls back to captured
revisions; configuration and persistent data are not reverted.
Status and recovery
--retry repeats failed verification and smoke checks, never apply. Recovery
keeps a failed deploy out of the successful baseline, even after rollback.
On workstation (replace the numeric ID with Actions run ID and attempt):
python3 ~/.local/lib/homelab-deploy/controller.py status 123-1
python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry
journalctl --user -u homelab-deploy@123-1
Runs live in ~/.local/state/homelab-deploy/runs. Compose stores resolved configs
with restricted permissions; these may contain credentials and must never be
uploaded as CI artifacts. Stage logs print the exact manual recovery command
using compose-before/<stack>.json, the original project directory and project
name. Compose does not automatically roll back, and Nextcloud AIO's child
containers remain managed by AIO. Preserve its own backups for data recovery.
The controller retains twenty successful/planned runs and preserves failures. Update the workstation dispatcher only when no deploy is running.
Validation and migration rollback
python3 -m unittest discover -s tests -v
bash .gitea/tests/deploy-validation.sh
Test on a separate namespace before the initial production full run. Check a
failed rollout, interrupted SSH and repeated run ID, and verify that an isolated
service change does not upgrade unrelated Helm releases or Compose stacks.
To roll back the migration, disable autodeploy and finish or recover the remote
run first. Restore the runner config/unit from .before-<timestamp> backups,
reload systemd and restart the runner. Restore the prior workflows from Git.
Production data and persistent volumes stay where they were. Do not remove run
state or Compose recovery files until recovery is confirmed.
Compose configuration recovery
Successful deploys save the complete resolved Compose configuration in
~/.local/state/homelab-deploy/compose-configs/. These files can contain secrets.
Keep them private and do not commit or upload them.
The next deploy uses this configuration for its recovery file, including old
commands, environment, mounts, ports, and removed services. The recovery command
uses --remove-orphans to remove services added by the failed deploy. It does
not restore volume data or reverse database migrations.
On the first run after this update, the controller can use the Compose file
from the previous successful run. If that file is absent, it reads the persistent
checkout and checks its service configuration hashes against existing containers.
A mismatch stops preflight. Restore the previous configuration before retrying.
Update the installed controller with bash .gitea/runner/setup-workstation.sh
from the reviewed checkout before using this change.
New namespaces are checked during preflight. Server validation of their resources runs after namespace creation and before application resources are applied. Plan mode does not create namespaces. A failed deferred check can leave an empty namespace; inspect it before removing it.