# Homelab CI/CD The native Gitea runners run on **vps**; production runs on **workstation**. Main-branch checks and image builds use `homelab:host`. Pull request and non-main checks use `homelab-pr:host` under a separate account without Docker access. The `homelab-pr` runner is registered at User scope for `forust`, so any repository under that account can schedule jobs that request this label. Each runner accepts one job at a time; the build waits for every check to pass. CI and deploy runs also show a summary with the release SHA, image build or reuse results, deploy mode, selected services, and image digests. Failed runs keep a summary of completed image builds, stage results, apply results, and recorded Kubernetes recovery. The final deploy summary is in the smoke job; earlier jobs show the state observed at that time. Apply success is separate from health and recovery. Update the installed workstation controller with `setup-workstation.sh` when no deploy is running. No job images or Kubernetes credentials are needed on the VPS. Builds use one pinned BuildKit helper container. CI and deploy are separate workflows. ## Runner installation Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl, GNU tar/xz, flock and systemd using the host's package manager. Keep the existing Gitea runner 3.0.2 binary at `/usr/local/bin/gitea-runner`. From this checkout on the VPS: ```sh sudo bash .gitea/runner/setup-runner.sh ``` The installer reuses `/var/lib/gitea-runner/.runner` and the existing service. For a new host, install the same runner binary and register as `gitea-runner` using the registration token interactively, label `homelab:host`, and working directory `/var/lib/gitea-runner`; then rerun the installer. Tokens never belong in this repository or command-line examples. Pinned tools live in the runner user's `~/.cache/homelab-ci`; CI repairs version drift there. Installations are locked. Buildx uses only the `homelab-ci` builder, pushes directly to the registry, and caps retained local cache at 1 GiB with a 2 GiB free-space target. This is not a hard limit on peak build disk usage. Nothing runs `docker system prune`, removes unrelated images, or deletes volumes. ### Pull request runner Install the unprivileged host runner on the VPS: ```sh sudo bash .gitea/runner/setup-pr-runner.sh ``` Get a registration token from the user Actions runner settings. Run the installer in a terminal. It asks for the token without echoing it, registers the runner as `homelab-pr` with label `homelab-pr:host`, then enables the service. The work directory is `/var/lib/gitea-pr-runner`. Confirm that Gitea lists the runner as User scope before merging the workflow change. An unmatched label can fall back to the default job image. Renovate PR validation uses `pull_request_target`, which reads the workflow from the base branch. It checks out the PR head only after runner selection and runs that code on `homelab-pr`. Keep this workflow read-only and do not add secrets. The PR runner has a separate home and tool cache. Do not add it to the `docker` group or give it access to `/var/run/docker.sock`. It runs repository code from pull requests, so keep its registration and permissions separate from the trusted `homelab` runner. This separates users and host permissions, but both runners still share the VPS kernel and network. Use a disposable VM if PRs from untrusted external authors must be fully isolated. ## Workstation setup As the existing SSH deploy user on workstation: ```sh sudo loginctl enable-linger forust bash .gitea/runner/setup-workstation.sh ``` The controller uses `/srv/homelab` as the persistent configuration tree and makes a detached source worktree for each SHA. It never resets `/srv/homelab`, moves local configuration, renames Compose projects, or changes volume names. The installer records the current Kubernetes context and cluster UID in `~/.config/homelab-deploy/environment`. Check these before installing. Configure Gitea Actions Variables: - `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`: the existing VPS-to-workstation SSH endpoint. - `DEPLOY_KNOWN_HOSTS`: workstation's verified SSH host key entry for that endpoint. - `AUTODEPLOY`: `false` initially; `true` enables deployment after successful main CI. Keep `DEPLOY_SSH_KEY`, `REGISTRY_USERNAME` and `REGISTRY_PASSWORD` in Actions Secrets. Legacy endpoint secrets remain accepted during migration. The Actions token must have repository read and Actions read access for release downloads. The deploy user's existing Docker registry authentication remains necessary. ## Releases and deployment CI publishes `release-` as a Gitea artifact with all three owned image digests and build input fingerprints. Unchanged images are reused only from a successful main CI artifact, never from `:prod`. Expired artifacts cause CI to rebuild images; they block deployment until CI is rerun. Run deploy from main with `deploy_ref=main` or a checked SHA: - `full`: required for the first baseline; reconcile all active components. - `changed`: compare with the last fully successful production deploy. - `plan`: validate configuration and show selection without changing production resources. - `refresh_images=true`: explicitly refresh mutable third-party Compose tags. The manual and automatic paths both require successful CI, a successful build job and the exact SHA's release artifact. PRs cannot publish images or deploy. Removed resources are reported and require explicit removal; no automatic prune. Service dependencies are listed in `.gitea/deploy-dependencies.json`. A workstation user systemd service holds the deploy lock across validation, sequential apply, verification and smoke checks. SSH clients only submit/follow: disconnecting or cancelling the Actions client does not kill production apply. Retrying the same run ID does not start another apply. `ExecStopPost` recovers interrupted runs before the unit finishes. Kubernetes rolls back to captured revisions; configuration and persistent data are not reverted. ## Status and recovery `--retry` repeats failed verification and smoke checks, never apply. Recovery keeps a failed deploy out of the successful baseline, even after rollback. On workstation (replace the numeric ID with Actions run ID and attempt): ```sh python3 ~/.local/lib/homelab-deploy/controller.py status 123-1 python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry journalctl --user -u homelab-deploy@123-1 ``` Runs live in `~/.local/state/homelab-deploy/runs`. Compose stores resolved configs with restricted permissions; these may contain credentials and must never be uploaded as CI artifacts. Stage logs print the exact manual recovery command using `compose-before/.json`, the original project directory and project name. Compose does not automatically roll back, and Nextcloud AIO's child containers remain managed by AIO. Preserve its own backups for data recovery. The controller retains twenty successful/planned runs and preserves failures. Update the workstation dispatcher only when no deploy is running. ## Validation and migration rollback ```sh python3 -m unittest discover -s tests -v bash .gitea/tests/deploy-validation.sh ``` Test on a separate namespace before the initial production `full` run. Check a failed rollout, interrupted SSH and repeated run ID, and verify that an isolated service change does not upgrade unrelated Helm releases or Compose stacks. To roll back the migration, disable autodeploy and finish or recover the remote run first. Restore the runner config/unit from `.before-` backups, reload systemd and restart the runner. Restore the prior workflows from Git. Production data and persistent volumes stay where they were. Do not remove run state or Compose recovery files until recovery is confirmed. ### Compose configuration recovery Successful deploys save the complete resolved Compose configuration in `~/.local/state/homelab-deploy/compose-configs/`. These files can contain secrets. Keep them private and do not commit or upload them. The next deploy uses this configuration for its recovery file, including old commands, environment, mounts, ports, and removed services. The recovery command uses `--remove-orphans` to remove services added by the failed deploy. It does not restore volume data or reverse database migrations. On the first run after this update, the controller can use the Compose file from the previous successful run. If that file is absent, it reads the persistent checkout and checks its service configuration hashes against existing containers. A mismatch stops preflight. Restore the previous configuration before retrying. Update the installed controller with `bash .gitea/runner/setup-workstation.sh` from the reviewed checkout before using this change. New namespaces are checked during preflight. Server validation of their resources runs after namespace creation and before application resources are applied. Plan mode does not create namespaces. A failed deferred check can leave an empty namespace; inspect it before removing it.