docs: sync service guides with current main
ci / Compose (pull_request) Successful in 27s
ci / Workflows (pull_request) Successful in 14s
ci / Shell (pull_request) Successful in 34s
ci / Python and tests (pull_request) Successful in 19s
ci / YAML (pull_request) Successful in 17s
ci / Dockerfiles (pull_request) Successful in 6s
ci / Formatting (pull_request) Successful in 36s
ci / Kubernetes (pull_request) Successful in 14s
ci / image-plan (pull_request) Skipped
ci / Image (${{ matrix.name }}) (pull_request) Skipped
ci / build (pull_request) Skipped
renovate-ci / validate-renovate (pull_request_target) Successful in 3m13s

This commit is contained in:
forust committed 2026-10-08 21:23:46 +02:00
commit d85bdf5dbd
126 files changed
+4971 -4076

No files matched your search

+50
View File
@@ -0,0 +1,50 @@
# EDU ownership handoff
## Status
The EDU ownership handoff is complete. The homelab repository no longer owns
EDU workloads, images, routes, alerts, or deployment selection. The EDU
repository is the only deployment owner: [forust/edu-master](https://git.forust.xyz/forust/edu-master).
Homelab PRs #99 and #105 are merged. PR #105 removed the EDU subtree and its
build, deploy, rollback, verification, route-probe, and registry references.
It also added the serial image build matrix for the homelab services. This
handoff record is the only remaining EDU-specific file in homelab Git.
The dedicated workstation checkout is `/srv/edu-master`, at release
`4f2b2a0e37dc11ac2c75441a15076c178e219d37`. It contains `k8s/active`; root
`active` is absent. The old untracked `/srv/homelab/edu_master` checkout was
moved outside the homelab repository to
`/srv/edu-master-legacy-archive-20261007/edu_master`. Its private files remain
mode `0600` inside an archive directory with mode `0700`. The homelab deploy
checkout has no EDU marker or tracked EDU application/deployment files.
`AUTODEPLOY=false` remains in place for homelab deployment.
## Release evidence
EDU PR #4 merged after its review and CI checks. Main-push CI run 1652 passed
all validation and both image builds. Deploy run 1653 passed for the exact main
SHA above.
The workstation rollout completed for both Deployments. The deployment
verified `/health` and `/live` with HTTP 200, Redis AUTH, session TTL of 1058
seconds, a delivery backlog of zero, and all nine EDU vmalert rules with
matching expressions and healthy evaluation.
The images now run by digest:
- Session keeper: `sha256:998dea51aa3015fd9cabefb0f53b030157a650c3bef72e02fe84f17d5762613d`
- Webinar checker: `sha256:92f3c1fa2bb7f9b4680a9fc76a5b33dfbea8ef3dd9c6490ebc45876fd4c54461`
Redis StatefulSet was unchanged. PVC `redis-data-pvc` remains bound to PV
`pvc-a4f2a79a-363a-4c12-ae91-92cdfc2a0d2e` with capacity 1 GiB. The existing
runtime Secret and Fernet key were preserved during the handoff. Notification
delivery was verified before closeout, as confirmed by the operator. The
deployment did not record downtime.
The release rollback snapshot is
`/home/forust/.local/state/edu-master-deploy/20261007T180541Z-4f2b2a0e37dc11ac2c75441a15076c178e219d37`.
The handoff data snapshot remains at
`/home/forust/.local/state/edu-master-deploy/handoff-20261007T080838Z`.
Both snapshots are outside Git. Do not restore old Redis data unless recovery
requires it. Never delete or recreate the Redis PVC.
+48 -72
View File
@@ -1,88 +1,64 @@
# Build and deployment workflows
# CI and deployment
Gitea Actions checks this repository, builds its custom images, and deploys
selected services to the workstation. Workflows use the self-hosted runner labels
`linux`, `arch`, and `homelab`; deployment jobs also require `prod`.
Gitea Actions validates changes, builds the repository's custom images, and can
deploy selected services to the workstation. CI and production deployment use
separate workflows. See the [runner and recovery guide](runner/README.md) for
installation, configuration, and operator commands.
## Checks
## CI
`ci.yaml` runs Compose validation, actionlint, ShellCheck, Prettier, Ruff,
yamllint, hadolint, and kubeconform. Tool versions are pinned in
`workflows/tool-versions.env` and installed by `install-ci-tools.sh`.
`workflows/ci.yaml` runs Compose, workflow, shell, formatting, Python and unit
test, YAML, Dockerfile, and Kubernetes checks. Pull requests and non-main refs
use the unprivileged `homelab-pr` runner. Main-branch CI uses `homelab`. Tool
versions are pinned in `workflows/tool-versions.env`.
Compose CI checks structure without resolving local environment files or paths.
On the reviewed main commit it only discovers standard filenames; the
`fix/deploy-validation` branch adds the manual Compose entry points too.
Compose CI checks every committed Compose file without requiring ignored `.env`
files. Kubernetes checks validate known schemas; unknown CRDs are skipped.
Kubeconform validates known resource schemas. Unknown CRDs are skipped. On main,
CI also attempts server-side dry-runs for marked services; these require an
existing namespace and contact the cluster's admission webhooks. A cluster that
is unreachable produces a warning and skips that CI pass. Deploy validation has
its own dry-run stage.
On main, CI plans builds for the three owned images: `error-pages`,
`forust-homepage`, and `xdfnx-homepage`. It builds changed inputs or reuses a
digest from a successful earlier main run. The successful build job publishes a
release artifact for the exact commit SHA. Pull requests do not publish images.
`renovate-ci.yaml` validates Renovate settings and checks that its generated
ConfigMap matches `renovate/renovate.json`.
## Deployment gate
## Image builds
`workflows/deploy.yaml` starts a deployment after successful main CI when the
`AUTODEPLOY` Actions variable is `true`. Manual dispatch uses the same gate: the
requested `main` ref or commit must have successful main CI and its matching
release artifact. A manual dispatch does not bypass validation.
CI builds changed custom images for `errorpages`, both `homepages` variants, and
the two `edu_master` Python services. Main builds publish `main`, `prod`, and a
commit tag. Dev builds publish `dev`. Build jobs wait for the lint and manifest
checks.
The workflow supports these modes:
Kubernetes deployment resolves the lab's own registry images to digests, preferring
commit-specific tags. Third-party image versions remain declared in the manifests.
- `changed`: select active services changed since the last successful deploy.
- `full`: select all active services; use this for the first baseline.
- `plan`: validate and show the selection without applying production resources.
## Deploy selection
`refresh_images=true` explicitly refreshes mutable third-party Compose tags.
`workflows/deploy-lib.sh` owns the stage logic; `ssh-run.sh` invokes it on the
workstation through SSH. Kubernetes selection uses `k8s/active`; Compose selection
uses an `active` file beside a standard `compose.yaml` or `compose.yml`.
Kustomize overlays are supported, although the current tree primarily contains
plain manifests.
## Selection and rollout
Secret files, examples, Helm values, and patch files are excluded from plain
manifest selection. Create local Kubernetes Secrets separately in their target
namespaces. The Helm table lists Prometheus, Loki, Alloy, and Reloader, with each
release controlled by its configured marker. Other charts need separate setup.
The active markers define automatic deployment. `<service>/active` selects a
standard Compose file; `<service>/k8s/active` selects Kubernetes resources. Helm
releases have their own markers in `workflows/deploy-lib.sh`. Service
dependencies are declared in `deploy-dependencies.json`. Removed resources are
reported for manual review; the workflow does not prune them automatically.
## Trigger and required settings
The workstation controller runs the checked source in a per-SHA worktree. It
validates configuration, applies Kubernetes and Compose changes in sequence,
verifies changed Kubernetes workloads, and checks public routes. A durable
systemd service continues the rollout if the Actions SSH client disconnects.
The workflow checks the exact CI release before it submits a deployment.
Automatic deployment follows a successful main CI run when the repository Actions
variable `AUTODEPLOY` is `true`. The manual deploy workflow bypasses that switch
and targets the fetched main branch when no validated commit SHA is provided.
A manual dispatch does not prove that this commit passed CI.
Kubernetes recovery uses captured workload revisions. It does not restore
ConfigMaps, Secrets, database schemas, or persistent data. Compose recovery is
manual and does not restore volume data or reverse migrations. Keep backups for
stateful services. The runner guide documents status, retry, logs, and recovery
commands.
Configure the Actions secrets `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_SSH_KEY`, and,
where needed, `DEPLOY_PORT` and `DEPLOY_PATH`. Registry publishing uses
`REGISTRY_USERNAME` and `REGISTRY_PASSWORD`. The remote user needs access to Git,
Docker, kubectl, Helm, jq, and the state directory used for snapshots.
## Settings
Keep `APPLY_PRUNE` false on the reviewed implementation: its per-file prune loop
is unsafe. `fix/deploy-prune-guard` rejects that option before changes are applied.
Preflight fetches and resets the remote checkout. It refuses when tracked files
have local changes; ignored local env and Secret files stay in place. Do not use
a development checkout with uncommitted tracked changes as the deployment target.
## Stages and recovery
1. Preflight fetches the target commit and checks the remote working tree.
2. Validate selects services, parses Compose, performs Kubernetes dry-runs, and
checks referenced Secrets.
3. Apply Kubernetes records a workload snapshot, upgrades selected Helm releases,
applies resources, and refreshes owned custom images.
4. Apply Compose recreates marked stacks and checks container state.
5. Verify Kubernetes checks changed workloads and attempts rollback for failures.
6. Smoke probes public routes after verification.
The two apply jobs share a remote lock. Workflow concurrency queues deployments
rather than interrupting an older apply. Snapshots live under
`$XDG_STATE_HOME/homelab-deploy`, or `~/.local/state/homelab-deploy` by default.
They contain the pre-apply workload data and commit identifier.
Rollback uses workload revisions. It does not restore ConfigMaps, Secrets,
database schemas, or data. Helm-owned workloads are handled through the Helm
upgrade's rollback path; the generic rollback skips them. Compose has no automatic
rollback. See the [review](../docs/repository-review.md) for remaining recovery
limitations, including SSH retries and serial rollback timing.
Configure `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`, and the verified
`DEPLOY_KNOWN_HOSTS` entry as Actions variables. Keep `DEPLOY_SSH_KEY`,
`REGISTRY_USERNAME`, and `REGISTRY_PASSWORD` in Actions secrets. The workstation
also needs its existing registry authentication. Set `AUTODEPLOY=false` until
automatic production deploys are intended.
+1
View File
@@ -7,4 +7,5 @@ self-hosted-runner:
labels:
- arch
- homelab
- homelab-pr
- prod
+3
View File
@@ -0,0 +1,3 @@
{
"postgres": ["authentik", "gitea", "immich", "n8n", "netbox", "netronome"]
}
+181
View File
@@ -0,0 +1,181 @@
# Homelab CI/CD
The native Gitea runners run on **vps**; production runs on **workstation**.
Main-branch checks and image builds use `homelab:host`. Pull request and
non-main checks use `homelab-pr:host` under a separate account without Docker
access. The `homelab-pr` runner is registered at User scope for `forust`, so
any repository under that account can schedule jobs that request this label.
Each runner accepts one job at a time; the build waits for every check to pass.
CI and deploy runs also show a summary with
the release SHA, image build or reuse results, deploy mode, selected services,
and image digests. Failed runs keep a summary of completed image builds, stage
results, apply results, and recorded Kubernetes recovery. The final deploy
summary is in the smoke job; earlier jobs show the state observed at that time.
Apply success is separate from health and recovery. Update the installed
workstation controller with `setup-workstation.sh` when no deploy is running.
No job images or Kubernetes credentials are needed on the VPS. Builds use one
pinned BuildKit helper container. CI and deploy are separate workflows.
## Runner installation
Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl,
GNU tar/xz, flock and systemd using the host's package manager. Keep the existing
Gitea runner 3.0.2 binary at `/usr/local/bin/gitea-runner`.
From this checkout on the VPS:
```sh
sudo bash .gitea/runner/setup-runner.sh
```
The installer reuses `/var/lib/gitea-runner/.runner` and the existing service.
For a new host, install the same runner binary and register as `gitea-runner`
using the registration token interactively, label `homelab:host`, and working
directory `/var/lib/gitea-runner`; then rerun the installer. Tokens never belong
in this repository or command-line examples.
Pinned tools live in the runner user's `~/.cache/homelab-ci`; CI repairs version
drift there. Installations are locked. Buildx uses only the `homelab-ci` builder,
pushes directly to the registry, and caps retained local cache at 1 GiB with a
2 GiB free-space target. This is not a hard limit on peak build disk usage.
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
### Pull request runner
Install the unprivileged host runner on the VPS:
```sh
sudo bash .gitea/runner/setup-pr-runner.sh
```
Get a registration token from the user Actions runner settings. Run the
installer in a terminal. It asks for the token without echoing it, registers the
runner as `homelab-pr` with label `homelab-pr:host`, then enables the service.
The work directory is `/var/lib/gitea-pr-runner`. Confirm that Gitea lists the
runner as User scope before merging the workflow change. An unmatched label can
fall back to the default job image.
Renovate PR validation uses `pull_request_target`, which reads the workflow from
the base branch. It checks out the PR head only after runner selection and runs
that code on `homelab-pr`. Keep this workflow read-only and do not add secrets.
The PR runner has a separate home and tool cache. Do not add it to the `docker`
group or give it access to `/var/run/docker.sock`. It runs repository code from
pull requests, so keep its registration and permissions separate from the
trusted `homelab` runner. This separates users and host permissions, but both
runners still share the VPS kernel and network. Use a disposable VM if PRs from
untrusted external authors must be fully isolated.
## Workstation setup
As the existing SSH deploy user on workstation:
```sh
sudo loginctl enable-linger forust
bash .gitea/runner/setup-workstation.sh
```
The controller uses `/srv/homelab` as the persistent configuration tree and makes
a detached source worktree for each SHA. It never resets `/srv/homelab`, moves
local configuration, renames Compose projects, or changes volume names.
The installer records the current Kubernetes context and cluster UID in
`~/.config/homelab-deploy/environment`. Check these before installing.
Configure Gitea Actions Variables:
- `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`: the existing VPS-to-workstation SSH endpoint.
- `DEPLOY_KNOWN_HOSTS`: workstation's verified SSH host key entry for that endpoint.
- `AUTODEPLOY`: `false` initially; `true` enables deployment after successful main CI.
Keep `DEPLOY_SSH_KEY`, `REGISTRY_USERNAME` and `REGISTRY_PASSWORD` in Actions
Secrets. Legacy endpoint secrets remain accepted during migration. The Actions
token must have repository read and Actions read access for release downloads.
The deploy user's existing Docker registry authentication remains necessary.
## Releases and deployment
CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image
digests and build input fingerprints. Unchanged images are reused only from a
successful main CI artifact, never from `:prod`. Expired artifacts cause CI to
rebuild images; they block deployment until CI is rerun.
Run deploy from main with `deploy_ref=main` or a checked SHA:
- `full`: required for the first baseline; reconcile all active components.
- `changed`: compare with the last fully successful production deploy.
- `plan`: validate configuration and show selection without changing production resources.
- `refresh_images=true`: explicitly refresh mutable third-party Compose tags.
The manual and automatic paths both require successful CI, a successful build
job and the exact SHA's release artifact. PRs cannot publish images or deploy.
Removed resources are reported and require explicit removal; no automatic prune.
Service dependencies are listed in `.gitea/deploy-dependencies.json`.
A workstation user systemd service holds the deploy lock across validation,
sequential apply, verification and smoke checks. SSH clients only submit/follow:
disconnecting or cancelling the Actions client does not kill production apply.
Retrying the same run ID does not start another apply. `ExecStopPost` recovers
interrupted runs before the unit finishes. Kubernetes rolls back to captured
revisions; configuration and persistent data are not reverted.
## Status and recovery
`--retry` repeats failed verification and smoke checks, never apply. Recovery
keeps a failed deploy out of the successful baseline, even after rollback.
On workstation (replace the numeric ID with Actions run ID and attempt):
```sh
python3 ~/.local/lib/homelab-deploy/controller.py status 123-1
python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry
journalctl --user -u homelab-deploy@123-1
```
Runs live in `~/.local/state/homelab-deploy/runs`. Compose stores resolved configs
with restricted permissions; these may contain credentials and must never be
uploaded as CI artifacts. Stage logs print the exact manual recovery command
using `compose-before/<stack>.json`, the original project directory and project
name. Compose does not automatically roll back, and Nextcloud AIO's child
containers remain managed by AIO. Preserve its own backups for data recovery.
The controller retains twenty successful/planned runs and preserves failures.
Update the workstation dispatcher only when no deploy is running.
## Validation and migration rollback
```sh
python3 -m unittest discover -s tests -v
bash .gitea/tests/deploy-validation.sh
```
Test on a separate namespace before the initial production `full` run. Check a
failed rollout, interrupted SSH and repeated run ID, and verify that an isolated
service change does not upgrade unrelated Helm releases or Compose stacks.
To roll back the migration, disable autodeploy and finish or recover the remote
run first. Restore the runner config/unit from `.before-<timestamp>` backups,
reload systemd and restart the runner. Restore the prior workflows from Git.
Production data and persistent volumes stay where they were. Do not remove run
state or Compose recovery files until recovery is confirmed.
### Compose configuration recovery
Successful deploys save the complete resolved Compose configuration in
`~/.local/state/homelab-deploy/compose-configs/`. These files can contain secrets.
Keep them private and do not commit or upload them.
The next deploy uses this configuration for its recovery file, including old
commands, environment, mounts, ports, and removed services. The recovery command
uses `--remove-orphans` to remove services added by the failed deploy. It does
not restore volume data or reverse database migrations.
On the first run after this update, the controller can use the Compose file
from the previous successful run. If that file is absent, it reads the persistent
checkout and checks its service configuration hashes against existing containers.
A mismatch stops preflight. Restore the previous configuration before retrying.
Update the installed controller with `bash .gitea/runner/setup-workstation.sh`
from the reviewed checkout before using this change.
New namespaces are checked during preflight. Server validation of their resources
runs after namespace creation and before application resources are applied.
Plan mode does not create namespaces. A failed deferred check can leave an empty
namespace; inspect it before removing it.
+11
View File
@@ -0,0 +1,11 @@
[worker.oci]
gc = true
reservedSpace = "256MB"
maxUsedSpace = "1GB"
minFreeSpace = "2GB"
[[worker.oci.gcpolicy]]
reservedSpace = "256MB"
maxUsedSpace = "1GB"
minFreeSpace = "2GB"
all = true
+10
View File
@@ -0,0 +1,10 @@
runner:
file: /var/lib/gitea-runner/.runner
capacity: 1
timeout: 5h
labels:
- homelab:host
cache:
enabled: false
container:
docker_host: unix:///var/run/docker.sock
+18
View File
@@ -0,0 +1,18 @@
[Unit]
Description=Gitea Actions runner
After=network-online.target docker.service
Wants=network-online.target
[Service]
User=gitea-runner
Group=gitea-runner
SupplementaryGroups=docker
WorkingDirectory=/var/lib/gitea-runner
Environment=PATH=/var/lib/gitea-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-runner/config.yaml
Restart=on-failure
RestartSec=5
UMask=0077
[Install]
WantedBy=multi-user.target
+12
View File
@@ -0,0 +1,12 @@
[Unit]
Description=Homelab deploy %i
[Service]
Type=exec
EnvironmentFile=%h/.config/homelab-deploy/environment
ExecStart=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py execute %i
ExecStopPost=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py recover %i
RuntimeMaxSec=5h
TimeoutStopSec=135min
KillMode=control-group
UMask=0077
+8
View File
@@ -0,0 +1,8 @@
runner:
file: /var/lib/gitea-pr-runner/.runner
capacity: 1
timeout: 5h
labels:
- homelab-pr:host
cache:
enabled: false
+27
View File
@@ -0,0 +1,27 @@
[Unit]
Description=Gitea Actions untrusted pull request runner
After=network-online.target
Wants=network-online.target
[Service]
User=gitea-pr-runner
Group=gitea-pr-runner
WorkingDirectory=/var/lib/gitea-pr-runner
Environment=HOME=/var/lib/gitea-pr-runner
Environment=PATH=/var/lib/gitea-pr-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-pr-runner/config.yaml
Restart=on-failure
RestartSec=5
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=full
ProtectHome=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
RestrictSUIDSGID=yes
LockPersonality=yes
UMask=0077
[Install]
WantedBy=multi-user.target
+56
View File
@@ -0,0 +1,56 @@
#!/usr/bin/env bash
# Install a native runner for untrusted PR jobs without Docker access.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
for tool in cp cut date getent id install runuser systemctl useradd; do
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
done
command -v /usr/local/bin/gitea-runner >/dev/null || {
echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2
exit 1
}
id gitea-pr-runner >/dev/null 2>&1 || \
useradd --system --create-home --home-dir /var/lib/gitea-pr-runner --shell /usr/bin/bash gitea-pr-runner
runner_home="$(getent passwd gitea-pr-runner | cut -d: -f6)"
[ "$runner_home" = /var/lib/gitea-pr-runner ] || {
echo 'Unexpected PR runner home; inspect the existing service first' >&2
exit 1
}
case " $(id -nG gitea-pr-runner) " in
*' docker '*)
echo 'The PR runner account must not belong to the docker group' >&2
exit 1
;;
esac
install -d -m 0755 /etc/gitea-pr-runner
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
for existing in /etc/gitea-pr-runner/config.yaml /etc/systemd/system/gitea-pr-runner.service; do
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
done
install -m 0644 "$here/pr-config.yaml" /etc/gitea-pr-runner/config.yaml
install -m 0644 "$here/pr-runner.service" /etc/systemd/system/gitea-pr-runner.service
if [ ! -f /var/lib/gitea-pr-runner/.runner ]; then
read -r -s -p 'Enter the Gitea repository runner registration token: ' runner_token
printf '\n'
[ -n "$runner_token" ] || { echo 'Runner token is required' >&2; exit 1; }
export GITEA_RUNNER_REGISTRATION_TOKEN="$runner_token"
unset runner_token
runuser --preserve-environment -u gitea-pr-runner -- \
/usr/local/bin/gitea-runner register \
--config /etc/gitea-pr-runner/config.yaml \
--instance https://gitea.forust.xyz \
--name homelab-pr \
--labels homelab-pr:host \
--no-interactive
unset GITEA_RUNNER_REGISTRATION_TOKEN
fi
chmod 0600 /var/lib/gitea-pr-runner/.runner
systemctl daemon-reload
systemctl enable --now gitea-pr-runner.service
systemctl restart gitea-pr-runner.service
echo "PR runner ready. Configuration backups: *.before-$stamp"
+48
View File
@@ -0,0 +1,48 @@
#!/usr/bin/env bash
# Native host runner, with pinned user-space tools and no extra CI images.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
for tool in docker curl python3 git tar xz flock runuser systemctl; do
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
done
docker info >/dev/null
docker compose version >/dev/null
docker buildx version >/dev/null
id gitea-runner >/dev/null 2>&1 || useradd --system --create-home --home-dir /var/lib/gitea-runner --shell /usr/bin/bash gitea-runner
# Reuse the established service account and runner registration.
runner_home="$(getent passwd gitea-runner | cut -d: -f6)"
[ "$runner_home" = /var/lib/gitea-runner ] || { echo 'Unexpected runner home; inspect the existing service first' >&2; exit 1; }
runuser -u gitea-runner -- docker info >/dev/null || { echo "The runner user needs access to Docker before setup" >&2; exit 1; }
command -v gitea-runner >/dev/null || { echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2; exit 1; }
mkdir -p /etc/gitea-runner
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
for existing in /etc/gitea-runner/config.yaml /etc/systemd/system/gitea-runner.service; do
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
done
scratch="$(mktemp -d)"
trap 'rm -rf "$scratch"' EXIT
chmod 755 "$scratch"
install -m 0644 "$here/../workflows/install-ci-tools.sh" "$here/../workflows/tool-versions.env" "$scratch/"
runuser -u gitea-runner -- bash "$scratch/install-ci-tools.sh"
install -m 0644 "$here/config.yaml" /etc/gitea-runner/config.yaml
python3 - <<'PYLABELS'
import json
from pathlib import Path
registration = Path('/var/lib/gitea-runner/.runner')
if registration.exists():
labels = json.loads(registration.read_text()).get('labels', [])
labels = [label for label in labels if isinstance(label, str) and label.split(':')[0] != 'homelab']
labels.append('homelab:host')
config = Path('/etc/gitea-runner/config.yaml')
config.write_text(config.read_text().replace(' - homelab:host', '\n'.join(' - ' + json.dumps(label) for label in labels)))
PYLABELS
install -m 0644 "$here/gitea-runner.service" /etc/systemd/system/gitea-runner.service
if [ ! -f /var/lib/gitea-runner/.runner ]; then
echo 'Register once as gitea-runner with homelab:host before starting the service.'
exit 0
fi
systemctl daemon-reload
systemctl enable --now gitea-runner.service
systemctl restart gitea-runner.service
echo "Runner ready. Configuration backups: *.before-$stamp"
+32
View File
@@ -0,0 +1,32 @@
#!/usr/bin/env bash
# Run as the existing deploy user on workstation. Never resets the working tree.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
repo="${HOMELAB_REPO:-/srv/homelab}"
for tool in python3 git kubectl helm docker flock timeout; do
command -v "$tool" >/dev/null || { echo "Install missing dependency: $tool" >&2; exit 1; }
done
[ -d "$repo/.git" ] || { echo "Missing deploy checkout: $repo" >&2; exit 1; }
[[ "$repo" =~ ^/[A-Za-z0-9_./-]+$ ]] || { echo 'Deploy path must be absolute and contain no whitespace' >&2; exit 1; }
if [ "$(loginctl show-user "$USER" -p Linger --value)" != yes ]; then
echo "Run once: sudo loginctl enable-linger $USER" >&2
exit 1
fi
config="${XDG_CONFIG_HOME:-$HOME/.config}/homelab-deploy"
mkdir -p "$config" "$HOME/.local/lib/homelab-deploy" "$HOME/.config/systemd/user"
chmod 700 "$config"
if [ ! -f "$config/environment" ]; then
context="$(kubectl config current-context)"
cluster_uid="$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')"
printf 'HOMELAB_REPO=%s\nKUBE_CONTEXT=%s\nEXPECTED_CLUSTER_UID=%s\n' "$repo" "$context" "$cluster_uid" >"$config/environment"
chmod 600 "$config/environment"
fi
# Do not replace a dispatcher while an existing deploy uses it.
if systemctl --user list-units 'homelab-deploy@*' --state=running --no-legend | grep -q .; then
echo 'An existing deploy is running; wait before updating the controller' >&2
exit 1
fi
install -m 0755 "$here/../workflows/deploy-controller.py" "$HOME/.local/lib/homelab-deploy/controller.py"
install -m 0644 "$here/homelab-deploy@.service" "$HOME/.config/systemd/user/homelab-deploy@.service"
systemctl --user daemon-reload
echo 'Controller ready. Run a checked main SHA in full mode for the initial baseline.'
+156
View File
@@ -0,0 +1,156 @@
#!/usr/bin/env bash
# Local regressions only: kubectl is mocked and Docker is used for config parsing.
set -euo pipefail
repo="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
scratch="$(mktemp -d)"
trap 'rm -rf "$scratch"' EXIT
mkdir -p "$scratch/repo/app" "$scratch/repo/postgres" "$scratch/repo/netbird" "$scratch/repo/renovate"
git -C "$scratch/repo" init -q
for file in app/compose.yaml postgres/shared-compose.yaml netbird/client.compose.yaml renovate/renovate-compose.yaml; do
touch "$scratch/repo/$file"
done
git -C "$scratch/repo" add .
# shellcheck source=../workflows/compose-lint.sh
source "$repo/.gitea/workflows/compose-lint.sh"
actual="$(cd "$scratch/repo" && compose_files)"
expected=$'app/compose.yaml\nnetbird/client.compose.yaml\npostgres/shared-compose.yaml\nrenovate/renovate-compose.yaml'
[ "$actual" = "$expected" ] || { echo 'Compose discovery missed a file' >&2; exit 1; }
cat >"$scratch/compose.yaml" <<'YAML'
services:
example:
image: busybox:1.37.0
environment:
REQUIRED: ${HOMELAB_TEST_REQUIRED:?required for this regression}
YAML
unset HOMELAB_TEST_REQUIRED
if validate_compose_file "$scratch/compose.yaml" >"$scratch/config.log" 2>&1; then
echo 'Full Compose validation accepted a missing variable' >&2
exit 1
fi
grep -q 'required for this regression' "$scratch/config.log"
HOMELAB_TEST_REQUIRED=present validate_compose_file "$scratch/compose.yaml"
cat >"$scratch/resources.json" <<'JSON'
{"kind":"List","items":[
{"kind":"Deployment","metadata":{"namespace":"app"},"spec":{"template":{"spec":{
"containers":[{"envFrom":[{"secretRef":{"name":"credentials"}},{"secretRef":{"name":"optional","optional":true}}],"env":[{"valueFrom":{"secretKeyRef":{"name":"credentials","key":"password"}}}]}],
"initContainers":[{"envFrom":[{"secretRef":{"name":"init"}}]}],
"imagePullSecrets":[{"name":"registry"}],
"volumes":[{"secret":{"secretName":"mounted"}},{"projected":{"sources":[{"secret":{"name":"projected"}},{"secret":{"name":"optional-projected","optional":true}}]}}]
}}}},
{"kind":"CronJob","metadata":{},"spec":{"jobTemplate":{"spec":{"template":{"spec":{"containers":[{"envFrom":[{"secretRef":{"name":"cron"}}]}]}}}}}},
{"kind":"IngressRoute","metadata":{"namespace":"app"},"spec":{"tls":{"secretName":"controller-issued-tls"}}}
]}
JSON
actual="$(jq -r -f "$repo/.gitea/workflows/secret-references.jq" "$scratch/resources.json" | sort)"
expected=$'app credentials\napp init\napp mounted\napp projected\napp registry\ndefault cron'
[ "$actual" = "$expected" ] || { echo "Unexpected Secret references: $actual" >&2; exit 1; }
REPO="$repo"
# shellcheck source=../workflows/deploy-lib.sh
source "$repo/.gitea/workflows/deploy-lib.sh"
K8S_MANIFESTS=("$scratch/resources.json")
KUSTOMIZE_APPS=()
# No live cluster access. Reject credentials in app even if they exist elsewhere.
kubectl() {
case "$1" in
create) cat "$scratch/resources.json" ;;
get)
if [ "$3" = credentials ] && [ "$5" = app ]; then
return 1
fi
return 0
;;
*) echo "Unexpected kubectl invocation: $*" >&2; return 1 ;;
esac
}
if check_referenced_secrets >"$scratch/secrets.log"; then
echo 'Namespace-scoped Secret check accepted a missing Secret' >&2
exit 1
fi
grep -q 'MISSING OR UNREADABLE: app/credentials' "$scratch/secrets.log"
# API/rendering errors must not produce an empty reference list and pass.
kubectl() { return 1; }
if ! skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
echo 'VMAgent preflight did not skip an uninstalled CRD' >&2
exit 1
fi
kubectl() { return 0; }
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
echo 'VMAgent preflight skipped an installed CRD' >&2
exit 1
fi
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/victoria.yaml"; then
echo 'VMAgent preflight skipped an unrelated manifest' >&2
exit 1
fi
kubectl() { return 1; }
if check_referenced_secrets >"$scratch/secrets.log"; then
echo 'Secret check accepted a failed manifest render' >&2
exit 1
fi
# New declared namespaces defer only their own resources during preflight.
render_selected_resources() {
cat <<'JSON'
{"apiVersion":"v1","kind":"List","items":[
{"apiVersion":"v1","kind":"Namespace","metadata":{"name":"new"}},
{"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"new-config","namespace":"new"}},
{"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"existing-config","namespace":"default"}}
]}
JSON
}
kubectl() {
case "$1" in
get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}}]}' ;;
apply) cat >"$scratch/server-input.json" ;;
*) return 1 ;;
esac
}
validate_server_resources true
jq -e '.items | length == 2 and all(.metadata.name != "new-config")' "$scratch/server-input.json" >/dev/null
if validate_server_resources false 2>"$scratch/deferred.log"; then
echo 'Post-namespace validation accepted a missing namespace' >&2
exit 1
fi
kubectl() {
case "$1" in
get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}},{"metadata":{"name":"new"}}]}' ;;
apply) cat >"$scratch/server-input.json" ;;
*) return 1 ;;
esac
}
validate_server_resources false
jq -e '.items | length == 3' "$scratch/server-input.json" >/dev/null
render_selected_resources() {
printf '%s\n' '{"items":[{"kind":"ConfigMap","metadata":{"name":"bad","namespace":"undeclared"}}]}'
}
if validate_server_resources true 2>"$scratch/undeclared.log"; then
echo 'Preflight accepted an undeclared missing namespace' >&2
exit 1
fi
# Count services, not characters in the newline-separated service names.
compose() {
case "$*" in
*'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{},"init":{"restart":"no"}}}' ;;
*'ps --status running --services') printf '%s\n' headscale headplane web ;;
*) return 1 ;;
esac
}
verify_compose_stack example.yaml >"$scratch/compose-count.log"
grep -qF 'all 3 service(s) running' "$scratch/compose-count.log"
compose() {
case "$*" in
*'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{}}}' ;;
*'ps --status running --services') printf '%s\n' headscale headplane ;;
*) return 0 ;;
esac
}
if verify_compose_stack example.yaml >"$scratch/compose-missing.log"; then
echo 'Compose verification accepted a missing service' >&2
exit 1
fi
grep -qF 'NOT RUNNING: web' "$scratch/compose-missing.log"
printf '%s\n' 'Deploy validation regressions passed.'
+357 -386
View File
@@ -1,38 +1,25 @@
name: ci
on:
"on":
push:
branches:
- "**"
pull_request:
workflow_dispatch:
# Every job here is checkout plus local tools. The token needs to read the tree
# and nothing else, and saying so keeps a future step that reaches for the API
# from quietly holding a token that can write to the repository.
- main
pull_request: null
workflow_dispatch: null
permissions:
contents: read
actions: read
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
env:
REGISTRY: gcr.forust.xyz
jobs:
lint-compose:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
compose:
name: Compose
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
# Structure check for every committed Compose file, active or not.
# Interpolation, env-file and bind-mount resolution are all switched off,
# because inactive stacks have no .env here and would only fail on their
# ${VAR:?} guards. Active stacks get the full check with interpolation in
# the deploy workflow, where the real .env files live.
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Validate Compose files
shell: bash
run: |
@@ -61,35 +48,77 @@ jobs:
exit 1
fi
echo "checked ${#files[@]} Compose file(s)"
lint-actionlint:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Compose
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.source.conclusion == 'failure'
&& 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
workflows:
name: Workflows
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint shellcheck)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint Gitea Actions workflows with actionlint
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint)"
export PATH="$tools_dir:$PATH"
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
lint-shellcheck:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Workflows
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
shell:
name: Shell
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck jq)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint shell scripts with ShellCheck
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck)"
export PATH="$tools_dir:$PATH"
mapfile -t scripts < <(
git ls-files '*.sh' ':(glob)**/*.bash'
)
@@ -98,20 +127,42 @@ jobs:
exit 0
fi
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
lint-prettier:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
bash .gitea/tests/deploy-validation.sh
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Shell
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
formatting:
name: Formatting
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Check formatting with Prettier
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
export PATH="$tools_dir:$PATH"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Check formatting with Prettier
shell: bash
run: |
set -euo pipefail
mapfile -t prettier_files < <(
git ls-files \
@@ -125,36 +176,79 @@ jobs:
fi
prettier --check --ignore-unknown "${prettier_files[@]}"
lint-ruff:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Formatting
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
python:
name: Python and tests
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff jq)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint and format-check Python with Ruff
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff)"
export PATH="$tools_dir:$PATH"
ruff check .
ruff format --check .
lint-yaml:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
ruff check . .gitea/workflows
ruff format --check . .gitea/workflows
python3 -m unittest discover -s tests -v
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Python and tests
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
yaml:
name: YAML
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint YAML syntax
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
export PATH="$tools_dir:$PATH"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint YAML syntax
shell: bash
run: |
set -euo pipefail
mapfile -t yaml_files < <(
git ls-files '*.yaml' '*.yml' \
@@ -168,20 +262,41 @@ jobs:
fi
yamllint -c .yamllint "${yaml_files[@]}"
lint-dockerfiles:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 10
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: YAML
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
dockerfiles:
name: Dockerfiles
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Lint Dockerfiles
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
export PATH="$tools_dir:$PATH"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint Dockerfiles
shell: bash
run: |
set -euo pipefail
mapfile -t dockerfiles < <(
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
@@ -193,20 +308,41 @@ jobs:
fi
hadolint -c .hadolint.yaml "${dockerfiles[@]}"
validate:
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 20
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Dockerfiles
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
kubernetes:
name: Kubernetes
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Validate Kubernetes manifests against JSON schemas
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Validate Kubernetes manifests against JSON schemas
shell: bash
run: |
set -euo pipefail
mapfile -t manifests < <(
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
@@ -223,327 +359,162 @@ jobs:
-ignore-missing-schemas \
-summary \
"${manifests[@]}"
# kubeconform has no schemas for CRDs, so every IngressRoute, Certificate,
# PrometheusRule, Middleware, ServersTransport and ServiceMonitor is silently
# skipped above. The live API server knows the real CRD schemas (and runs the
# cert-manager / Traefik admission webhooks), so validate there too.
#
# Only services marked with a k8s/active marker are checked: server-side
# dry-run needs the target namespace to exist, and inactive services are not
# deployed. Services being enabled for the first time are still covered by
# the JSON-schema pass above.
#
# Main pushes only. `--dry-run=server` persists nothing, but it does execute
# the admission webhooks of the production API server, so anyone able to open
# a pull request would be able to run arbitrary manifest content through
# cert-manager and Traefik. A pull request has nothing to gain from it either:
# only main is ever deployed, and this job runs to completion before the
# deploy workflow is allowed to start, so a bad CRD is still caught before
# anything reaches the cluster -- just on the push rather than on the PR.
- name: Note the server-side check is not running here
if: github.event_name == 'pull_request' || github.ref != 'refs/heads/main'
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Kubernetes
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
echo "::notice::Skipping the server-side dry-run. It executes the cert-manager and" \
"Traefik admission webhooks against the production API server, so it is limited" \
"to pushes to main. CRDs are still schema-checked by kubeconform above, and the" \
"server-side pass still runs on main before the deploy."
- name: Validate active manifests against the live API server
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
shell: bash
run: |
set -euo pipefail
if ! kubectl get --raw='/readyz' --request-timeout=10s >/dev/null 2>&1; then
echo "::warning::Cluster unreachable — skipped server-side validation of CRDs (IngressRoute, Certificate, PrometheusRule). Review manifest changes manually."
exit 0
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
mapfile -t k8s_dirs < <(
git ls-files '*.yaml' '*.yml' \
| grep -E '(^|/)k8s/' \
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
| sort -u
)
manifests=()
kustomize_apps=()
for dir in "${k8s_dirs[@]}"; do
if [ ! -f "${dir}/active" ]; then
echo "skip (no k8s/active): ${dir}"
continue
fi
if [ -f "${dir}/overlays/prod/kustomization.yaml" ]; then
kustomize_apps+=("${dir}/overlays/prod")
elif [ -f "${dir}/base/kustomization.yaml" ]; then
kustomize_apps+=("${dir}/base")
else
while IFS= read -r f; do
[ -n "$f" ] && manifests+=("$f")
done < <(
git ls-files "${dir}/*.yaml" "${dir}/*.yml" \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
)
fi
done
echo "server-side dry-run: ${#manifests[@]} manifests, ${#kustomize_apps[@]} kustomize apps"
failed=0
for m in ${manifests[@]+"${manifests[@]}"}; do
if ! out="$(kubectl apply --dry-run=server -f "$m" 2>&1)"; then
failed=1
echo "::error file=${m}::$(printf '%s' "$out" | head -1)"
fi
done
for k in ${kustomize_apps[@]+"${kustomize_apps[@]}"}; do
if ! out="$(kubectl apply -k "$k" --dry-run=server 2>&1)"; then
failed=1
echo "::error file=${k}::$(printf '%s' "$out" | head -1)"
fi
done
if [ "$failed" -ne 0 ]; then
echo "Server-side validation failed. The API server (or an admission webhook) rejected these manifests."
exit 1
fi
echo "server-side dry-run: all active manifests accepted by the API server"
build:
needs:
# The panel's scan-deps/test-backend/test-frontend jobs gated here until
# userbot moved to its own repo; upstream's code is upstream's gate now.
# The rule is unchanged: publishing and passing the checks are the same
# gate, so a commit that fails any of these still cannot move :prod.
[lint-actionlint, lint-shellcheck, lint-compose, lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev') && !startsWith(github.ref_name, 'renovate/')
runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 60
image-plan:
needs: [compose, workflows, shell, formatting, python, yaml, dockerfiles, kubernetes]
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
runs-on: homelab
timeout-minutes: 10
outputs:
services: ${{ steps.services.outputs.services }}
matrix: ${{ steps.plan.outputs.matrix }}
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
id: source
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
with:
fetch-depth: 0
- name: Detect build inputs against successful CI
id: plan
env:
GITEA_TOKEN: ${{ github.token }}
run: python3 .gitea/workflows/release.py prepare --output build-plan.json
- name: Store the image plan
id: artifact
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: build-plan
path: build-plan.json
if-no-files-found: error
retention-days: 30
- name: Detect changed docker-built services
id: services
- name: Write the plan result
if: always()
env:
SUMMARY_CHECK: Image plan
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP: >-
${{ steps.plan.conclusion == 'failure' && 'Build input detection' ||
steps.artifact.conclusion == 'failure' && 'Plan upload' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
set -euo pipefail
base="${{ github.event.before }}"
if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then
base="$(git rev-list --max-parents=0 HEAD)"
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## Image plan\n\nResult: %s\n' "$SUMMARY_RESULT" >>"$GITHUB_STEP_SUMMARY" || true
fi
# A failed diff used to leave changed_files empty, which reads exactly
# like "nothing to build": the job went green having built nothing and
# the tag never moved. The status is checked, not assumed.
if ! changed="$(git diff --name-only "$base" "${GITHUB_SHA}")"; then
echo "::error::cannot diff ${base}..${GITHUB_SHA}"
exit 1
fi
mapfile -t changed_files <<<"$changed"
services=()
add_service() {
local name="$1"
local seen=0
for existing in "${services[@]}"; do
if [ "$existing" = "$name" ]; then
seen=1
break
fi
done
if [ "$seen" -eq 0 ]; then
services+=("$name")
fi
}
for file in "${changed_files[@]}"; do
case "$file" in
errorpages/*)
add_service errorpages
;;
homepages/*)
add_service homepages
;;
edu_master/phpsessid-bot/*|edu_master/webinar-checker/*|edu_master/compose.yaml)
add_service edu_master
;;
esac
done
if [ "${#services[@]}" -eq 0 ]; then
echo "No docker-built services changed."
echo "services=" >> "$GITHUB_OUTPUT"
exit 0
fi
printf '%s\n' "${services[@]}" | tee /tmp/services.txt
echo "services=$(paste -sd, /tmp/services.txt)" >> "$GITHUB_OUTPUT"
- name: Log in to registry
# The pin step below also writes (manifest PUTs), and it runs on every
# main push — including manifest-only ones where services is empty. A
# stale persistent login on the old runner used to mask this; a clean
# runner pushes anonymously and gets 401.
if: steps.services.outputs.services != '' || github.ref_name == 'main'
images:
name: Image (${{ matrix.name }})
needs: [image-plan]
if: needs.image-plan.result == 'success'
runs-on: homelab
timeout-minutes: 60
strategy:
max-parallel: 1
fail-fast: false
matrix: ${{ fromJSON(needs.image-plan.outputs.matrix || '{"include":[{"name":"inactive"}]}') }}
steps:
- name: Checkout repository
id: source
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
- name: Download the checked image plan
id: inputs
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
name: build-plan
- name: Build or reuse this image
id: check
env:
IMAGE_NAME: ${{ matrix.name }}
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: python3 .gitea/workflows/release.py image --image "$IMAGE_NAME" --output image.json
- name: Store the image result
id: artifact
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: image-${{ matrix.name }}
path: image.json
if-no-files-found: error
retention-days: 30
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Image (${{ matrix.name }})
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP: >-
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
# Through env, not by substitution into the script. A secret written
# into a run: block is pasted into the shell source before bash parses
# it, so a password containing a quote, a backtick or $(...) becomes
# code that runs. Masking the value in the log does not prevent that.
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
# Retain the build job name required by the immutable release deployment gate.
build:
needs: [image-plan, images]
runs-on: homelab
timeout-minutes: 15
steps:
- name: Checkout repository
id: source
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
- name: Download all image results
id: inputs
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
path: artifacts
- name: Pin SHA tags and write the complete release
id: check
env:
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: |
set -euo pipefail
printf '%s' "$REGISTRY_PASSWORD" | docker login "${REGISTRY}" \
-u "$REGISTRY_USERNAME" \
--password-stdin
- name: Build and push changed images
if: steps.services.outputs.services != ''
run: >-
python3 .gitea/workflows/release.py finalize
--plan artifacts/build-plan/build-plan.json
- name: Store commit release
id: artifact
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: release-${{ github.sha }}
path: release.json
if-no-files-found: error
retention-days: 30
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Image release and SHA tags
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP: >-
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
# This step was the one run: block in the workflow without it, and it
# is the one that cannot afford it: a docker push that failed partway
# through the loop used to be followed by more pushes, the loop's exit
# status came from the last one, and the job went green with half the
# images missing from the registry.
set -euo pipefail
IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}"
# Tags for this push. The commit-pinned name is the point of this
# step: the deploy resolves it in preference to :prod, so a deploy
# that sat in the queue behind a later push still gets the build of
# the commit CI validated, instead of whatever :prod points at by the
# time it runs. See render_pinned in deploy-lib.sh.
commit_tag=""
if [ "${GITHUB_REF_NAME}" = "main" ]; then
commit_tag="sha-${GITHUB_SHA:0:12}"
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
set_tags() {
tags=()
case "${GITHUB_REF_NAME}" in
main) tags+=("main" "prod") ;;
dev) tags+=("dev") ;;
esac
if [ -n "$commit_tag" ]; then
tags+=("$commit_tag")
fi
}
for service in "${services[@]}"; do
case "$service" in
errorpages)
image="${REGISTRY}/forust/error-pages"
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" errorpages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
;;
homepages)
for variant in forust xdfnx; do
case "$variant" in
forust)
image="${REGISTRY}/forust/forust-homepage"
;;
xdfnx)
image="${REGISTRY}/forust/xdfnx-homepage"
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" -f "homepages/Dockerfile.${variant}" homepages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
edu_master)
for variant in session-keeper webinar-checker; do
case "$variant" in
session-keeper)
context="edu_master/phpsessid-bot"
image="${REGISTRY}/forust/session-keeper"
;;
webinar-checker)
context="edu_master/webinar-checker"
image="${REGISTRY}/forust/webinar-checker"
;;
esac
set_tags
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" "$context"
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
esac
done
# Every image the tree names has to carry the commit-pinned name, not only
# the ones this push rebuilt. A push that touches nothing but manifests
# builds nothing, and its deploy would then find no commit-pinned tag to
# resolve and quietly fall back to the moving :prod - which is the whole
# failure the commit-pinned name exists to remove.
#
# Re-tagging copies the manifest list and transfers no layers, so pinning
# six images that already exist costs six registry writes.
#
# The list is derived from the tree rather than written out here, so an
# image added to a manifest is covered without a second place to update.
- name: Pin the commit name on the images this push did not rebuild
if: github.ref_name == 'main'
shell: bash
run: |
set -euo pipefail
commit_tag="sha-${GITHUB_SHA:0:12}"
mapfile -t repos < <(
git grep -hoE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+' -- '*.yaml' '*.yml' \
| sort -u
)
if [ "${#repos[@]}" -eq 0 ]; then
echo "No own images referenced by the tree."
exit 0
fi
echo "pinning ${#repos[@]} image(s) to $commit_tag"
for repo in "${repos[@]}"; do
if docker buildx imagetools inspect "$repo:$commit_tag" >/dev/null 2>&1; then
echo " already built by this push: ${repo##*/}"
continue
fi
if ! docker buildx imagetools inspect "$repo:prod" >/dev/null 2>&1; then
echo " WARNING: ${repo##*/} has no :prod to pin and no build produced it"
continue
fi
docker buildx imagetools create --tag "$repo:$commit_tag" "$repo:prod"
echo " pinned ${repo##*/}"
done
+1 -2
View File
@@ -21,8 +21,7 @@
# All committed Compose files, including the ones deploy never starts.
compose_files() {
git ls-files \
'*/compose.yaml' '*/compose.yml' 'compose.yaml' 'compose.yml' \
'*/docker-compose.yaml' '*/docker-compose.yml'
'*compose.yaml' '*compose.yml'
}
# Prints the flags that turn `docker compose config` into the general check.
+164
View File
@@ -0,0 +1,164 @@
#!/usr/bin/env python3
"""Resolve Compose images without changing project names or local bind paths."""
import json
import os
import re
import subprocess
import sys
from pathlib import Path
def output(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
def resolve(reference):
if '@sha256:' in reference:
return reference
descriptor = json.loads(
output('docker', 'buildx', 'imagetools', 'inspect', reference, '--format', '{{json .Manifest}}')
)
digest = descriptor['digest']
if not re.fullmatch(r'sha256:[0-9a-f]{64}', digest):
raise ValueError(f'Invalid registry digest for {reference}')
# Strip tag only from the final path segment (registry ports are preserved).
repository = reference.rsplit('/', 1)
repository[-1] = repository[-1].split(':')[0]
return '/'.join(repository) + '@' + digest
def prepare(source_file):
config_repo = Path(os.environ['CONFIG_REPO'])
source_repo = Path(os.environ['REPO'])
directory = Path(os.environ['RUN_DIR'])
relative = source_file.relative_to(source_repo)
project_dir = config_repo / relative.parent
base = ['docker', 'compose', '--project-directory', str(project_dir), '-f', str(source_file)]
config = json.loads(output(*base, 'config', '--format', 'json', cwd=config_repo))
project = config['name']
previous_file = directory / 'previous.json'
previous = json.loads(previous_file.read_text()) if previous_file.exists() else {}
images_file = directory / 'compose-images.json'
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
release = json.loads((directory / 'release.json').read_text())
state = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
baseline = state / 'compose-configs' / f'{relative.parent.name}.json'
if not baseline.exists() and re.fullmatch(r'[0-9]+-[0-9]+', previous.get('run_id', '')):
baseline = state / 'runs' / previous['run_id'] / 'compose' / baseline.name
bootstrap = not baseline.exists()
if not bootstrap:
before = json.loads(baseline.read_text())
else:
# Bootstrap from the persistent configuration, never from the new source.
persistent_file = config_repo / relative
if persistent_file.exists():
before = json.loads(
output(
'docker',
'compose',
'--project-directory',
str(project_dir),
'-f',
str(persistent_file),
'config',
'--format',
'json',
cwd=config_repo,
)
)
elif output('docker', 'ps', '-aq', '--filter', f'label=com.docker.compose.project={project}'):
raise ValueError(f'{project}: no previous Compose configuration; restore it before deploy')
else:
before = {'name': project, 'services': {}}
if before['name'] != project:
raise ValueError('Compose project name changed; manual migration is required')
for service, settings in config['services'].items():
reference = settings.get('image')
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
if not reference or settings.get('build'):
raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
image_repo = reference.split('@')[0].rsplit('/', 1)
image_repo[-1] = image_repo[-1].split(':')[0]
image_repo = '/'.join(image_repo)
# Nextcloud AIO validates the mastercontainer image reference and rejects
# a digest. Keep its configured tag so AIO can start and manage its stack.
if nextcloud_aio_master:
pinned = reference
elif image_repo in release['images']:
pinned = image_repo + '@' + release['images'][image_repo]
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
pinned = locks[reference]
else:
pinned = resolve(reference)
settings['image'] = pinned
locks[reference] = pinned
for service, settings in before['services'].items():
reference = settings['image']
image_repo = reference.split('@')[0].rsplit('/', 1)
image_repo[-1] = image_repo[-1].split(':')[0]
image_repo = '/'.join(image_repo)
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
# Capture what is running, not the current value of its mutable tag.
ids = output(
'docker',
'ps',
'-aq',
'--filter',
f'label=com.docker.compose.project={project}',
'--filter',
f'label=com.docker.compose.service={service}',
).splitlines()
actual = set()
if bootstrap and ids:
expected_hash = output(
'docker',
'compose',
'--project-directory',
str(project_dir),
'-f',
str(persistent_file),
'config',
'--hash',
service,
cwd=config_repo,
).split()[-1]
for container in ids:
running_hash = output(
'docker',
'inspect',
container,
'--format',
'{{ index .Config.Labels "com.docker.compose.config-hash" }}',
)
if running_hash != expected_hash:
raise ValueError(
f'{project}/{service}: persistent config differs from running config; restore the previous config'
)
for container in ids:
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
if len(actual) > 1:
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
# AIO also rejects a digest in its recovery config. Preserve its tag in
# both deploy and recovery files.
if nextcloud_aio_master:
before['services'][service]['image'] = reference
else:
before['services'][service]['image'] = next(iter(actual)) if actual else reference
for name, data in (('compose', config), ('compose-before', before)):
folder = directory / name
folder.mkdir(mode=0o700, exist_ok=True)
destination = folder / f'{relative.parent.name}.json'
destination.write_text(json.dumps(data, indent=2) + '\n')
destination.chmod(0o600)
images_file.write_text(json.dumps(locks, indent=2) + '\n')
print(f'Compose {project}: images pinned; local paths preserved')
print(
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never --remove-orphans'
)
if __name__ == '__main__':
prepare(Path(sys.argv[1]))
+427
View File
@@ -0,0 +1,427 @@
#!/usr/bin/env python3
"""Durable workstation deployment controller. Install with setup-workstation.sh."""
import argparse
import contextlib
import fcntl
import importlib.util
import json
import math
import os
import re
import shutil
import subprocess
import sys
import time
from pathlib import Path
STATE = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
CONFIG_REPO = Path(os.environ.get('HOMELAB_REPO', '/srv/homelab'))
RUN_ID = re.compile(r'[0-9]+-[0-9]+')
def command(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
def atomic_json(path, data):
temporary = path.with_suffix('.tmp')
temporary.write_text(json.dumps(data, indent=2) + '\n')
temporary.chmod(0o600)
temporary.replace(path)
@contextlib.contextmanager
def lock(name):
STATE.mkdir(mode=0o700, parents=True, exist_ok=True)
with (STATE / name).open('a') as stream:
fcntl.flock(stream, fcntl.LOCK_EX)
yield
def load_module(name, path):
spec = importlib.util.spec_from_file_location(name, path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def run_directory(run_id):
if not RUN_ID.fullmatch(run_id):
raise ValueError('Run ID must be numeric workflow-id and attempt')
return STATE / 'runs' / run_id
def start(run_id):
payload = sys.stdin.buffer.read(256 * 1024 + 1)
if len(payload) > 256 * 1024:
raise ValueError('Deploy request exceeds 256 KiB')
request = json.loads(payload)
sha = request['release']['sha']
if not re.fullmatch(r'[0-9a-f]{40}', sha) or request['mode'] not in ('changed', 'full', 'plan'):
raise ValueError('Invalid deploy SHA or mode')
if not isinstance(request['refresh_images'], bool):
raise ValueError('refresh_images must be boolean')
directory = run_directory(run_id)
with lock('prepare.lock'):
if (directory / 'request.json').exists():
if json.loads((directory / 'request.json').read_text()) != request:
raise ValueError('Run ID already belongs to a different request')
else:
directory.mkdir(mode=0o700, parents=True, exist_ok=True)
command('git', '-C', str(CONFIG_REPO), 'fetch', '--quiet', 'origin', 'main')
command('git', '-C', str(CONFIG_REPO), 'merge-base', '--is-ancestor', sha, 'origin/main')
if not (directory / 'source').exists():
command('git', '-C', str(CONFIG_REPO), 'worktree', 'add', '--detach', str(directory / 'source'), sha)
if command('git', '-C', str(directory / 'source'), 'rev-parse', 'HEAD') != sha:
raise ValueError('Prepared source does not match deploy SHA')
release_module = load_module('release', directory / 'source/.gitea/workflows/release.py')
release_module.validate_release(request['release'], sha)
atomic_json(directory / 'release.json', request['release'])
atomic_json(directory / 'request.json', request)
if not (directory / 'status.json').exists():
atomic_json(directory / 'status.json', {'state': 'queued', 'stages': {}})
# Starting an existing active or finished ID is idempotent; never re-apply it.
if json.loads((directory / 'status.json').read_text())['state'] == 'queued':
command('systemctl', '--user', 'start', '--no-block', f'homelab-deploy@{run_id}.service')
print(f'Accepted deploy {run_id} ({sha})')
def environment(directory):
request = json.loads((directory / 'request.json').read_text())
return {
**os.environ,
'REPO': str(directory / 'source'),
'CONFIG_REPO': str(CONFIG_REPO),
'RUN_DIR': str(directory),
'DEPLOY_SHA': request['release']['sha'],
'RELEASE_FILE': str(directory / 'release.json'),
'DEPLOY_PLAN': str(directory / 'plan.json'),
'DEPLOY_SNAPSHOT_DIR': str(directory / 'snapshot'),
'REFRESH_IMAGES': str(request['refresh_images']).lower(),
'ROLLOUT_PARALLELISM': '4',
}
def stage(directory, name, budget):
status = json.loads((directory / 'status.json').read_text())
if name in status['stages'] and status['stages'][name].get('result') in ('success', 'failure'):
return status['stages'][name]['result'] == 'success'
started = time.time()
status['stages'][name] = {'result': 'running', 'started': started}
atomic_json(directory / 'status.json', status)
script = directory / 'source/.gitea/workflows/deploy-stage.sh'
with (directory / f'{name}.log').open('a') as log:
# timeout kills the whole stage process group, including children, before recovery.
result = subprocess.run( # noqa: S603, S607
[
shutil.which('timeout') or '/usr/bin/timeout',
'--signal=TERM',
'--kill-after=30s',
str(budget),
'bash',
str(script),
name,
],
env=environment(directory),
stdout=log,
stderr=subprocess.STDOUT,
check=False,
).returncode
status = json.loads((directory / 'status.json').read_text())
status['stages'][name].update(
result='success' if result == 0 else 'failure', exit_code=result, seconds=round(time.time() - started)
)
atomic_json(directory / 'status.json', status)
return result == 0
def make_plan(directory):
source = directory / 'source'
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
request = json.loads((directory / 'request.json').read_text())
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
# Helm 4 lists every release status by default and removed the --all flag.
helm = json.loads(command('helm', 'list', '-A', '-o', 'json'))
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
if request['refresh_images']:
plan['selected']['compose'] = plan['active']['compose']
atomic_json(directory / 'plan.json', plan)
if previous:
atomic_json(directory / 'previous.json', previous)
# Local config is deliberately separate from the immutable Git source.
return plan
def finish_success(directory, plan):
# Repeating finalization after a crash is safe while holding deploy.lock.
plan['run_id'] = directory.name
path = directory / 'compose-images.json'
previous = directory / 'previous.json'
plan['compose-images'] = (
json.loads(path.read_text())
if path.exists()
else json.loads(previous.read_text()).get('compose-images', {})
if previous.exists()
else {}
)
configs = STATE / 'compose-configs'
configs.mkdir(mode=0o700, exist_ok=True)
for config in (directory / 'compose').glob('*.json'):
atomic_json(configs / config.name, json.loads(config.read_text()))
atomic_json(STATE / 'last-success.json', plan)
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'success'
atomic_json(directory / 'status.json', status)
try:
retain_completed(directory)
except (OSError, subprocess.CalledProcessError) as error:
print(f'Retention deferred: {error}', flush=True)
def recover(directory, retry=False):
status = json.loads((directory / 'status.json').read_text())
if status['state'] in ('success', 'planned'):
return
completed = ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke')
if all(status['stages'].get(name, {}).get('result') == 'success' for name in completed):
finish_success(directory, json.loads((directory / 'plan.json').read_text()))
return
if retry:
for name in ('verify-k8s', 'smoke'):
if status['stages'].get(name, {}).get('result') == 'failure':
del status['stages'][name]
atomic_json(directory / 'status.json', status)
snapshot = directory / 'snapshot/current'
if snapshot.exists():
stage(directory, 'verify-k8s', 7200)
stage(directory, 'smoke', 600)
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'failure'
atomic_json(directory / 'status.json', status)
def execute(run_id):
directory = run_directory(run_id)
with lock('deploy.lock'):
status = json.loads((directory / 'status.json').read_text())
if status['state'] != 'queued':
return
# A crashed predecessor must be recovered before another apply begins.
for other in (STATE / 'runs').iterdir():
if (
other != directory
and (other / 'status.json').exists()
and json.loads((other / 'status.json').read_text())['state'] == 'running'
):
raise ValueError(f'Interrupted deploy {other.name}; run recover first')
status['state'] = 'running'
atomic_json(directory / 'status.json', status)
phase = 'plan'
try:
plan = make_plan(directory)
print(
json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}),
flush=True,
)
phase = 'doctor'
if not stage(directory, 'doctor', 600):
raise RuntimeError('Preflight failed')
phase = 'validate'
if not stage(directory, 'validate', 1200):
raise RuntimeError('Validation failed')
if json.loads((directory / 'request.json').read_text())['mode'] == 'plan':
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'planned'
atomic_json(directory / 'status.json', status)
return
# Budget includes both rollout checks and rollback waves, plus API overhead.
phase = 'Recovery budget'
count = int(
command(
'bash',
str(directory / 'source/.gitea/workflows/deploy-stage.sh'),
'workload-count',
env=environment(directory),
)
)
verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120)
if verify_budget > 7200:
raise ValueError('More than two hours of recovery required; split this deploy')
phase = 'apply-k8s'
k8s_ok = stage(directory, 'apply-k8s', 2700)
phase = 'apply-compose'
compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False
phase = 'verify-k8s'
verify_ok = stage(directory, 'verify-k8s', verify_budget)
phase = 'smoke'
smoke_ok = stage(directory, 'smoke', 600)
if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)):
raise RuntimeError('Deploy failed; inspect stage logs and recovery report')
phase = 'Save the successful baseline'
finish_success(directory, plan)
except Exception as error:
status = json.loads((directory / 'status.json').read_text())
status['failure_stage'] = next(
(name for name, result in status['stages'].items() if result.get('result') == 'failure'), phase
)
atomic_json(directory / 'status.json', status)
with (directory / 'controller.log').open('a') as stream:
stream.write(f'{error}\n')
recover(directory)
raise
def retain_completed(current):
finished = []
for directory in (STATE / 'runs').iterdir():
status_file = directory / 'status.json'
if status_file.exists() and json.loads(status_file.read_text())['state'] in ('success', 'planned'):
finished.append(directory)
for directory in sorted(finished, key=lambda p: p.stat().st_mtime, reverse=True)[20:]:
if directory == current:
continue
command('git', '-C', str(CONFIG_REPO), 'worktree', 'remove', '--force', str(directory / 'source'))
shutil.rmtree(directory)
def follow(run_id, phase):
directory = run_directory(run_id)
groups = {
'apply': ('doctor', 'validate', 'apply-k8s', 'apply-compose'),
'verify': ('verify-k8s',),
'smoke': ('smoke',),
}
names = groups[phase]
offsets = {}
while True:
status = json.loads((directory / 'status.json').read_text())
for name in (*names, 'controller'):
path = directory / f'{name}.log'
if path.exists():
with path.open() as stream:
stream.seek(offsets.get(name, 0))
content = stream.read()
if content:
print(content, end='', flush=True)
offsets[name] = stream.tell()
stages = status['stages']
if all(stages.get(name, {}).get('result') in ('success', 'failure') for name in names):
return all(stages[name]['result'] == 'success' for name in names)
if status['state'] in ('success', 'failure', 'planned'):
return status['state'] in ('success', 'planned')
time.sleep(3)
def summary(run_id):
directory = run_directory(run_id)
request = json.loads((directory / 'request.json').read_text())
release = request['release']
plan_file = directory / 'plan.json'
lines = [
f'## Deploy `{release["sha"]}`',
'',
f'- Mode: `{request["mode"]}`',
f'- Refresh third-party images: `{request["refresh_images"]}`',
]
status = json.loads((directory / 'status.json').read_text())
if status.get('failure_stage'):
lines.append(f'- Failed stage: **{status["failure_stage"]}**')
lines.extend(
[
'',
f'- Observed run state: **{status["state"]}**',
'',
'### Stage results',
'| Stage | Result | Exit code |',
'| --- | --- | --- |',
]
)
for name in ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke'):
stage_result = status['stages'].get(name, {})
lines.append(f'| {name} | {stage_result.get("result", "not started")} | {stage_result.get("exit_code", "—")} |')
lines.extend(['', '### Apply and Helm recovery results'])
events_file = directory / 'apply-events.jsonl'
events = []
if events_file.exists():
for line in events_file.read_text().splitlines():
try:
events.append(json.loads(line))
except json.JSONDecodeError:
lines.append('- An operation record is incomplete. Check the stage log.')
latest = {(event['action'], event['target']): event['result'] for event in events}
lines.extend(f'- `{action}` `{target}`: **{result}**' for (action, target), result in latest.items())
if not latest:
lines.append('- No apply results were recorded.')
lines.append('- A completed apply does not confirm health. See verification and smoke results.')
lines.extend(['', '### Kubernetes recovery'])
pointer = directory / 'snapshot/current'
failed = Path(pointer.read_text().strip()) / 'failed-workloads' if pointer.exists() else None
if failed and failed.exists():
contents = failed.read_text()
counts = dict(re.findall(r'^(ROLLED_BACK|UNRECOVERED)=([0-9]+)$', contents, re.MULTILINE))
if not contents.strip():
lines.append('- No failed workloads were recorded. See the verification result above.')
elif counts:
lines.append(f'- Workloads restored: **{counts.get("ROLLED_BACK", "unknown")}**')
lines.append(f'- Workloads that need manual recovery: **{counts.get("UNRECOVERED", "unknown")}**')
else:
lines.append('- Rollback has no recorded result yet. Check the verification log.')
else:
lines.append('- No workload rollback was recorded. This does not confirm health.')
lines.append('- Compose requires manual recovery. Use the saved command in the apply log.')
if not plan_file.exists():
lines.extend(['', 'Plan was not created. Check the controller log.'])
print('\n'.join(lines))
return
plan = json.loads(plan_file.read_text())
lines.extend(['', '### Selected services'])
count = 0
for kind, services in plan['selected'].items():
for service in services:
lines.append(f'- `{kind}`: `{service}`')
count += 1
if not count:
lines.append('- None')
lines.extend(['', '### Selected Helm releases'])
lines.extend(f'- `{release}`' for release in plan.get('helm', []))
if not plan.get('helm'):
lines.append('- None')
lines.extend(['', '### Images pinned in the checked release'])
lines.extend(f'- `{image}@{digest}`' for image, digest in sorted(release['images'].items()))
lines.extend(['', '### Removed resources requiring manual review'])
lines.extend(f'- `{item}`' for item in plan.get('removed', []))
if not plan.get('removed'):
lines.append('- None')
print('\n'.join(lines))
def main():
os.umask(0o077)
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow', 'summary'))
parser.add_argument('run_id')
parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke'))
parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply')
args = parser.parse_args()
directory = run_directory(args.run_id)
if args.action == 'start':
start(args.run_id)
elif args.action == 'execute':
execute(args.run_id)
elif args.action == 'recover':
with lock('deploy.lock'):
recover(directory, retry=args.retry)
elif args.action == 'status':
print((directory / 'status.json').read_text())
if (directory / 'plan.json').exists():
plan = json.loads((directory / 'plan.json').read_text())
print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2))
elif args.action == 'summary':
summary(args.run_id)
elif not follow(args.run_id, args.phase):
sys.exit(1)
if __name__ == '__main__':
main()
File diff suppressed because it is too large. Load diff
+124
View File
@@ -0,0 +1,124 @@
#!/usr/bin/env python3
"""Calculate selected components against the last fully successful deploy."""
import hashlib
import json
import re
import subprocess
from pathlib import Path
def output(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
def tracked(repo):
return output('git', '-C', str(repo), 'ls-files').splitlines()
def helm_releases(repo):
text = (repo / '.gitea/workflows/deploy-lib.sh').read_text()
return [line.split('|') for line in re.findall(r'^ "([^"\n]+\|[^"\n]+)"$', text, re.MULTILINE)]
def inventory(repo):
files = tracked(repo)
k8s = sorted(
{f.split('/k8s/')[0] for f in files if '/k8s/' in f and (repo / f.split('/k8s/')[0] / 'k8s/active').is_file()}
)
compose = sorted(
{
str(Path(f).parent)
for f in files
if Path(f).name in ('compose.yaml', 'compose.yml') and (repo / Path(f).parent / 'active').is_file()
}
)
return {'k8s': k8s, 'compose': compose}
def file_hash(path):
return hashlib.sha256(path.read_bytes()).hexdigest() if path.is_file() else 'missing'
def make_plan(repo, config_repo, release, previous, mode, live_helm):
active = inventory(repo)
all_services = set(active['k8s'] + active['compose'])
helm_inputs = {}
helm_selected = []
for name, chart, namespace, version, values, marker in helm_releases(repo):
if not (repo / marker).is_file():
continue
value_path = repo / values if (repo / values).is_file() else config_repo / values
if not value_path.is_file():
raise ValueError(f'Missing Helm values: {values}')
stamp = hashlib.sha256(f'{chart}|{version}|{file_hash(value_path)}'.encode()).hexdigest()
helm_inputs[name] = stamp
live = next((h for h in live_helm if h['name'] == name and h['namespace'] == namespace), None)
if (
mode == 'full'
or previous is None
or previous.get('helm_inputs', {}).get(name) != stamp
or live is None
or live.get('status') != 'deployed'
or live.get('chart') != f'{chart.split("/")[-1]}-{version}'
):
helm_selected.append(name)
local_inputs = {}
for service in all_services:
candidates = [config_repo / service / '.env']
if service in active['compose']:
candidates.append(config_repo / '.env')
cfg = config_repo / service / 'config'
if cfg.is_dir():
candidates.extend(
p for p in cfg.rglob('*') if p.is_file() and p.suffix in ('.yaml', '.yml', '.json', '.conf')
)
local_inputs[service] = hashlib.sha256(
'\n'.join(f'{p.relative_to(config_repo)}:{file_hash(p)}' for p in sorted(candidates)).encode()
).hexdigest()
if previous is None:
if mode == 'changed':
raise ValueError('No successful baseline; run deploy in full mode first')
changed = set(all_services)
removed = []
else:
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
changed = {service for service in all_services for path in paths if path.startswith(service + '/')}
if any(path.startswith('.gitea/') for path in paths):
changed |= all_services
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
for file in tracked(repo):
owners = {service for service in all_services if file.startswith(service + '/')}
if not owners or not file.endswith(('.yaml', '.yml')):
continue
text = (repo / file).read_text()
if any(
image in text and previous.get('images', {}).get(image) != digest
for image, digest in release['images'].items()
):
changed |= owners
removed = sorted(
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
- all_services
)
removed += [path for path in paths if '/k8s/' in path and not (repo / path).exists()]
if mode == 'full':
changed = set(all_services)
dependencies = json.loads((repo / '.gitea/deploy-dependencies.json').read_text())
while True:
expanded = changed | {dependent for service in changed for dependent in dependencies.get(service, [])}
if expanded == changed:
break
changed = expanded
return {
'version': 1,
'sha': release['sha'],
'images': release['images'],
'active': active,
'selected': {kind: sorted(set(services) & changed) for kind, services in active.items()},
'helm': helm_selected,
'helm_inputs': helm_inputs,
'local_inputs': local_inputs,
'removed': sorted(set(removed)),
'full_smoke': mode == 'full' or 'traefik' in changed,
}
+10
View File
@@ -0,0 +1,10 @@
#!/usr/bin/env bash
set -euo pipefail
source "${REPO:?}/.gitea/workflows/deploy-lib.sh"
case "${1:?stage required}" in
workload-count)
select_manifests >/dev/null
selected_workload_refs | sort -u | wc -l
;;
*) run_stage "$1" ;;
esac
+100 -157
View File
@@ -1,197 +1,140 @@
name: deploy
on:
# Deploy only what CI already validated. workflow_run is used instead of
# workflow_dispatch so a red lint/validate run can never reach the cluster.
workflow_run:
workflows: [ci]
branches: [main]
types: [completed]
workflow_dispatch:
inputs:
deploy_ref:
description: "Commit already checked by successful main CI (main or SHA)"
default: main
required: true
deploy_mode:
description: "First deploy requires full; plan changes no production resources"
type: choice
options: [changed, full, plan]
default: changed
refresh_images:
description: "Explicitly refresh mutable third-party Compose tags"
type: boolean
default: false
# The deploy jobs read the tree, then reach the cluster over SSH with the
# deploy key. The Actions token itself is not part of that path, so it gets
# read-only contents and no more.
permissions:
contents: read
actions: read
concurrency:
group: deploy-main
# Queue instead of cancelling. Cancelling a run kills the apply job mid-loop and
# takes the verify job down with it, so a superseded deploy would leave the
# cluster half-applied and unchecked — the exact failure the verify job exists
# to catch. kubectl apply and docker compose up are both idempotent, so letting
# the older run finish and then deploying the newer commit costs little.
cancel-in-progress: false
env:
DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
DEPLOY_HOST: ${{ vars.DEPLOY_HOST || secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ vars.DEPLOY_PORT || secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ vars.DEPLOY_USER || secrets.DEPLOY_USER }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
# workflow_run's own GITHUB_SHA points at the branch head, not at the commit the
# finished ci run checked. Pin the exact validated commit instead, so a push
# landing mid-deploy cannot make the workstation deploy something else. Also
# what the verify job checks the snapshot against. Empty for workflow_dispatch,
# which falls back to the current origin/main.
DEPLOY_SHA: ${{ github.event.workflow_run.head_sha }}
DEPLOY_KNOWN_HOSTS: ${{ vars.DEPLOY_KNOWN_HOSTS }}
DEPLOY_RUN_ID: ${{ github.run_id }}-${{ github.run_attempt || 1 }}
DEPLOY_MODE: ${{ inputs.deploy_mode || 'changed' }}
REFRESH_IMAGES: ${{ inputs.refresh_images && 'true' || 'false' }}
jobs:
preflight:
# Autodeploy defaults to OFF: pushes deploy only when the AUTODEPLOY repo
# variable is set to 'true' (Settings -> Actions -> Variables). A manual
# Run workflow always bypasses the switch: dispatching it is the explicit
# intent to deploy.
gate:
if: >-
github.ref == 'refs/heads/main' &&
(vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') &&
(github.event_name != 'workflow_run' ||
(github.event.workflow_run.conclusion == 'success' &&
github.event.workflow_run.head_branch == 'main'))
runs-on: [self-hosted, linux, arch, homelab, prod]
(github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.head_branch == 'main'))
runs-on: homelab
timeout-minutes: 10
outputs:
sha: ${{ steps.release.outputs.sha }}
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Fetch and reset workstation
shell: bash
with:
fetch-depth: 0
- name: Check successful CI and download the exact commit release
id: release
env:
GITEA_TOKEN: ${{ github.token }}
DEPLOY_REF: ${{ inputs.deploy_ref || 'main' }}
EVENT_SHA: ${{ github.event.workflow_run.head_sha }}
run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA"
- name: Submit durable deploy to workstation
run: bash .gitea/workflows/ssh-run.sh start
- name: Write the request result
if: always()
env:
REQUEST_RESULT: ${{ job.status }}
CHECKED_SHA: ${{ steps.release.outputs.sha }}
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh preflight
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## Deploy request\n\n- Result: **%s**\n- Checked commit: %s\n- Mode: %s\n' "$REQUEST_RESULT" "${CHECKED_SHA:-not checked}" "$DEPLOY_MODE" >>"$GITHUB_STEP_SUMMARY"
if [ "$REQUEST_RESULT" != success ]; then
echo 'Open the failed step log. If SSH submission failed, check the remote controller state.' >>"$GITHUB_STEP_SUMMARY"
fi
fi
validate:
needs: [preflight]
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 20
apply:
needs: [gate]
runs-on: homelab
timeout-minutes: 120
steps:
- name: Checkout repository
- name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Dry-run manifests and check Secrets
shell: bash
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow validation and sequential Kubernetes / Compose apply
run: bash .gitea/workflows/ssh-run.sh apply
- name: Write the deploy result
if: always()
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh validate
if [ -f .gitea/workflows/ssh-run.sh ]; then
bash .gitea/workflows/ssh-run.sh summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
apply-k8s:
needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
# Apply only, no verification, so this is just the work itself: snapshot,
# then sequential `helm upgrade --install --wait --rollback-on-failure --timeout 10m`, then the apply loop.
# Verification has its own job and its own budget.
#
# 45 is roughly four times the measured cost of the stage, which is
# deliberately not raised on a theory:
#
# helm, healthy 3 no-op upgrades ~3-5 min
# helm, one release bad rollback-on-failure spends its 10m, ~10-15 min
# then rolls that one back
# apply loop ~40 manifests, 4 of which ~1 min
# resolve an image digest
# restart_stale_images 7.6s to find 8 workloads, ~0.5 min
# 9.8s to resolve their digests
#
# The helm figure is one release, not three: `set -e` aborts
# upgrade_helm_releases on the first failure, so a broken release costs
# 10m and the other two are never attempted. Multiplying 10m by three
# overstates the worst case by 20 minutes.
#
# The 45 minutes this was last raised to 45 were still not enough, and the
# job logs for those runs no longer exist, so what actually consumed the
# budget is not known - the two measurable candidates above account for
# ~15 of it. The unbounded `docker manifest inspect` against the registry's
# known hang mode is now bounded inside registry_digest (25s timeout, 3
# attempts): a dead registry fails each owned image after ~85s instead of
# hanging the stage, and a blinking one is retried instead of failing the
# whole apply file. Still open: make the stage announce which manifest it
# is working on, so a killed run leaves a diagnosable last line.
timeout-minutes: 45
verify:
needs: [gate, apply]
if: always() && needs.gate.result == 'success'
runs-on: homelab
timeout-minutes: 130
steps:
- name: Checkout repository
- name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Apply Kubernetes manifests
shell: bash
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow workload verification and recovery
run: bash .gitea/workflows/ssh-run.sh verify
- name: Write the deploy result
if: always()
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-k8s
if [ -f .gitea/workflows/ssh-run.sh ]; then
bash .gitea/workflows/ssh-run.sh summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
apply-compose:
needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Redeploy docker compose stacks
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-compose
# Watches the workloads this deploy changed and rolls back the ones that never
# became healthy. Runs even when the apply jobs failed, timed out or were
# cancelled — that is the whole point of splitting it out. `always()` is what
# lets it start after a failed dependency; the needs on apply-compose are a
# barrier, so verification begins only once both applies are done.
verify-k8s:
needs: [apply-k8s, apply-compose]
if: >-
always() &&
needs.apply-k8s.result != 'skipped' &&
needs.apply-compose.result != 'skipped'
runs-on: [self-hosted, linux, arch, homelab, prod]
# Not raised, because the arithmetic does not close.
#
# 32 workloads are under management and the wave width is 8, so the verify
# itself is 4 waves of ROLLOUT_TIMEOUT (300s) = 20 minutes worst case, when
# every rollout times out rather than converging. That is already 20 of 30.
#
# The other 10 would have to absorb rollback, and rollback_workloads is a
# serial `while read` loop at 300s per failed workload. 10 minutes buys two.
# Any larger number is buying a bigger multiple of an unbounded term rather
# than covering a known cost: 60 minutes buys eight, and 60 minutes is
# therefore not a bound, it is a guess with two digits.
#
# The number becomes derivable the moment rollback uses the same wave width
# as the verify: 32 failures then cost 4 waves = 20 minutes instead of 160,
# and 45 covers verify plus rollback at full width. That change is to the
# recovery path and is not folded into a timeout edit.
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Verify workloads and roll back on failure
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh verify-k8s
# Asks the public route of every active service whether it is actually
# serving, which the rollout check above structurally cannot: a pod can
# converge and still be crash-looping, or be listening on a port no Service
# points at, or answer 500.
#
# `always()` for the same reason verify-k8s has it, and it runs after that job
# specifically because a rollback is when a route most needs re-checking. The
# needs is a barrier, not a filter: whether verify-k8s passed, failed or was
# cancelled, the probes are what say whether the cluster is serving, and
# suppressing them on a rollback would hide the one run where the answer
# matters most.
smoke:
needs: [verify-k8s]
if: always() && needs.verify-k8s.result != 'skipped'
runs-on: [self-hosted, linux, arch, homelab, prod]
timeout-minutes: 10
needs: [gate, verify]
if: always() && needs.gate.result == 'success'
runs-on: homelab
timeout-minutes: 15
steps:
- name: Checkout repository
- name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Probe the public route of every active service
shell: bash
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow public route checks
run: bash .gitea/workflows/ssh-run.sh smoke
- name: Write the deploy result
if: always()
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh smoke
if [ -f .gitea/workflows/ssh-run.sh ]; then
bash .gitea/workflows/ssh-run.sh summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
+60 -29
View File
@@ -13,9 +13,14 @@ here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=tool-versions.env
. "$here/tool-versions.env"
TOOLS_DIR="${TOOLS_DIR:-${RUNNER_TEMP:-/tmp}/homelab-tools}"
TOOLS_DIR="${TOOLS_DIR:-${XDG_CACHE_HOME:-$HOME/.cache}/homelab-ci}"
BIN_DIR="$TOOLS_DIR/bin"
mkdir -p "$BIN_DIR"
# A runner may accept overlapping workflows even though each workflow is sequential.
exec 9>"$TOOLS_DIR/install.lock"
flock -w 300 9
export UV_TOOL_DIR="$TOOLS_DIR/uv-tools"
export UV_CACHE_DIR="$TOOLS_DIR/uv-cache"
# The just-installed tools must resolve inside this script too: callers only
# prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would
# miss the binary install_uv just placed (exit 127 on a clean runner).
@@ -48,7 +53,7 @@ esac
fetch() {
# fetch <url> <dest>
if command -v curl >/dev/null 2>&1; then
curl -sSLf --retry 3 -o "$2" "$1"
curl -sSLf --connect-timeout 15 --max-time 120 --retry 3 -o "$2" "$1"
elif command -v wget >/dev/null 2>&1; then
wget -q -O "$2" "$1"
else
@@ -88,10 +93,13 @@ installed_version() {
# at_version <command> <expected>
at_version() {
case "$(installed_version "$1")" in
*"$2"*) return 0 ;;
*) return 1 ;;
esac
local version expected="${2#v}"
version="$(installed_version "$1")"
if [[ "$version" =~ (^|[^0-9.])v?([0-9]+(\.[0-9]+)+) ]]; then
[ "${BASH_REMATCH[2]}" = "$expected" ]
else
return 1
fi
}
install_kubeconform() {
@@ -120,6 +128,15 @@ install_shellcheck() {
rm -rf "$tmp"
}
install_jq() {
if at_version jq "${JQ_VERSION}"; then
return 0
fi
fetch "https://github.com/jqlang/jq/releases/download/jq-${JQ_VERSION}/jq-linux-${goarch}" \
"$BIN_DIR/jq"
chmod 0755 "$BIN_DIR/jq"
}
install_uv() {
if at_version uv "${UV_VERSION}"; then
return 0
@@ -167,6 +184,7 @@ install_pip_audit() {
}
install_prettier() {
install_node
if at_version prettier "${PRETTIER_VERSION}"; then
return 0
fi
@@ -227,28 +245,41 @@ install_actionlint() {
rm -rf "$tmp"
}
wanted=("$@")
if [ "${#wanted[@]}" -eq 0 ]; then
wanted=(kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
fi
main() {
wanted=("$@")
if [ "${#wanted[@]}" -eq 0 ]; then
wanted=(node jq kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
fi
for tool in "${wanted[@]}"; do
case "$tool" in
kubeconform) install_kubeconform ;;
shellcheck) install_shellcheck ;;
actionlint) install_actionlint ;;
prettier) install_prettier ;;
ruff) install_ruff ;;
yamllint) install_yamllint ;;
pip-audit) install_pip_audit ;;
hadolint) install_hadolint ;;
node) install_node ;;
uv) install_uv ;;
*)
echo "install-ci-tools: unknown tool: $tool" >&2
exit 1
;;
esac
done
for tool in "${wanted[@]}"; do
case "$tool" in
kubeconform) install_kubeconform ;;
shellcheck) install_shellcheck ;;
jq) install_jq ;;
actionlint) install_actionlint ;;
prettier) install_prettier ;;
ruff) install_ruff ;;
yamllint) install_yamllint ;;
pip-audit) install_pip_audit ;;
hadolint) install_hadolint ;;
node) install_node ;;
uv) install_uv ;;
*)
echo "install-ci-tools: unknown tool: $tool" >&2
exit 1
;;
esac
done
printf '%s\n' "$BIN_DIR"
for old in "$BIN_DIR"/node-* "$BIN_DIR"/prettier-*; do
[ -d "$old" ] || continue
case "$(basename "$old")" in
"node-$NODE_VERSION"|"prettier-$PRETTIER_VERSION") ;;
*) rm -rf "$old" ;;
esac
done
if [ -x "$BIN_DIR/uv" ]; then "$BIN_DIR/uv" cache prune >/dev/null; fi
printf '%s\n' "$BIN_DIR"
}
if [ "${BASH_SOURCE[0]}" = "$0" ]; then main "$@"; fi
+519
View File
@@ -0,0 +1,519 @@
#!/usr/bin/env python3
"""CI release artifacts and the SHA-specific Gitea deployment gate (stdlib only)."""
import argparse
import hashlib
import io
import itertools
import json
import os
import re
import shutil
import subprocess
import sys
import tempfile
import urllib.error
import urllib.parse
import urllib.request
import zipfile
from pathlib import Path
SHA = re.compile(r'[0-9a-f]{40}')
DIGEST = re.compile(r'sha256:[0-9a-f]{64}')
IMAGES = {
'error-pages': ('errorpages', 'errorpages/Dockerfile'),
'forust-homepage': ('homepages', 'homepages/Dockerfile.forust'),
'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'),
}
def command(*args, **kwargs):
"""Arguments are passed directly to the executable, never to a shell."""
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
def validate_release(data, sha=None):
if data.get('version') != 1 or not SHA.fullmatch(data.get('sha', '')):
raise ValueError('Invalid release version or SHA')
if sha is not None and data['sha'] != sha:
raise ValueError('Release SHA does not match the checked CI commit')
expected = {f'gcr.forust.xyz/forust/{name}' for name in IMAGES}
if set(data.get('images', {})) != expected:
raise ValueError('Release must contain all owned images')
if not all(DIGEST.fullmatch(value) for value in data['images'].values()):
raise ValueError('Release has an invalid image digest')
if set(data.get('inputs', {})) != expected or not all(
re.fullmatch(r'[0-9a-f]{64}', value) for value in data['inputs'].values()
):
raise ValueError('Release has invalid build input fingerprints')
return data
class NoRedirect(urllib.request.HTTPRedirectHandler):
def redirect_request(self, _req, _fp, _code, _msg, _headers, _newurl):
return None
class Gitea:
def __init__(self):
self.origin = os.environ['GITHUB_SERVER_URL'].rstrip('/')
if urllib.parse.urlsplit(self.origin).scheme != 'https':
raise ValueError('Gitea API must use HTTPS')
self.repository = os.environ['GITHUB_REPOSITORY']
if not re.fullmatch(r'[\w.-]+/[\w.-]+', self.repository):
raise ValueError('Invalid Gitea repository')
self.token = os.environ['GITEA_TOKEN']
self.base = f'{self.origin}/api/v1/repos/{self.repository}'
def request(self, url, *, archive=False):
if not url.startswith(self.base + '/'):
raise ValueError('Refusing to send the Actions token to another origin')
req = urllib.request.Request(url, headers={'Authorization': f'token {self.token}'}) # noqa: S310 -- HTTPS origin validated above
opener = urllib.request.build_opener(NoRedirect())
try:
response = opener.open(req, timeout=30) # noqa: S310
except urllib.error.HTTPError as error:
if not archive or error.code not in (301, 302, 303, 307, 308):
raise RuntimeError(f'Gitea API returned HTTP {error.code}') from None
target = urllib.parse.urljoin(url, error.headers['Location'])
if urllib.parse.urlsplit(target).scheme != 'https':
raise ValueError('Artifact redirect must use HTTPS') from None
# Signed storage redirects must never receive the Gitea token.
response = urllib.request.urlopen(target, timeout=30) # noqa: S310
with response:
payload = response.read(8 * 1024 * 1024 + 1)
if len(payload) > 8 * 1024 * 1024:
raise ValueError('Gitea response exceeds 8 MiB')
return payload if archive else json.loads(payload)
def pages(self, path, key, **params):
for page in range(1, 101):
query = urllib.parse.urlencode({**params, 'page': page, 'limit': 50})
data = self.request(f'{self.base}/{path}?{query}')
entries = data[key]
yield from entries
if len(entries) < 50:
return
raise RuntimeError('Gitea pagination limit exceeded')
def successful_runs(self, sha=None):
params = {'branch': 'main', 'status': 'success', 'exclude_pull_requests': 'true'}
if sha:
params['head_sha'] = sha
for run in self.pages('actions/workflows/ci.yaml/runs', 'workflow_runs', **params):
if (
run.get('status') == 'completed'
and run.get('conclusion') == 'success'
and run.get('head_branch') == 'main'
and run.get('event') in ('push', 'workflow_dispatch')
and (run.get('repository') or {}).get('full_name') == self.repository
and (run.get('head_repository') or run.get('repository') or {}).get('full_name') == self.repository
and (sha is None or run.get('head_sha') == sha)
):
yield run
def release(self, run):
sha = run['head_sha']
jobs = list(self.pages(f'actions/runs/{run["id"]}/jobs', 'jobs'))
# A green workflow with a skipped build must not authorize a deploy.
if not any(job.get('name') == 'build' and job.get('conclusion') == 'success' for job in jobs):
raise ValueError('CI build job did not succeed')
artifacts = self.request(f'{self.base}/actions/runs/{run["id"]}/artifacts')['artifacts']
matching = [a for a in artifacts if a['name'] == f'release-{sha}' and not a.get('expired')]
if len(matching) != 1:
raise ValueError('CI release artifact is missing, expired or ambiguous; rerun CI')
blob = self.request(f'{self.base}/actions/artifacts/{matching[0]["id"]}/zip', archive=True)
with zipfile.ZipFile(io.BytesIO(blob)) as archive:
files = [entry for entry in archive.infolist() if not entry.is_dir()]
if len(files) != 1 or files[0].filename != 'release.json' or files[0].file_size > 256 * 1024:
raise ValueError('Unexpected release archive contents')
return validate_release(json.loads(archive.read(files[0])), sha)
def fingerprint(context, dockerfile):
tree = command('git', 'ls-tree', '-r', 'HEAD', '--', context, dockerfile, '.gitea/workflows/release.py')
return hashlib.sha256(tree.encode()).hexdigest()
def gate(output, requested_ref, event_sha):
command('git', 'fetch', '--quiet', 'origin', 'main')
if event_sha:
if not SHA.fullmatch(event_sha):
raise ValueError('Invalid workflow_run SHA')
sha = event_sha
else:
if requested_ref == 'main':
requested_ref = 'origin/main'
sha = command('git', 'rev-parse', '--verify', '--end-of-options', f'{requested_ref}^{{commit}}')
if not SHA.fullmatch(sha):
raise ValueError('Invalid deploy SHA')
command('git', 'merge-base', '--is-ancestor', sha, 'origin/main')
api = Gitea()
runs = list(api.successful_runs(sha))
if not runs:
raise ValueError(f'No successful main CI for {sha}; run CI before deploying')
release = api.release(max(runs, key=lambda run: run['id']))
output.write_text(json.dumps(release, indent=2) + '\n')
if os.environ.get('GITHUB_OUTPUT'):
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
stream.write(f'sha={sha}\n')
print(f'CI gate accepted {sha}')
def prepare_images(output):
sha = command('git', 'rev-parse', 'HEAD')
if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
raise ValueError('Build checkout does not match GITHUB_SHA')
api = Gitea()
previous = None
for run in sorted(itertools.islice(api.successful_runs(), 50), key=lambda item: item['id'], reverse=True):
if str(run['id']) == os.environ.get('GITHUB_RUN_ID'):
continue
try:
previous = api.release(run)
break
except ValueError:
# Expired artifacts only cost a rebuild; mutable tags are never a fallback.
continue
targets = []
for name, (context, dockerfile) in IMAGES.items():
image = f'gcr.forust.xyz/forust/{name}'
inputs = fingerprint(context, dockerfile)
old_digest = (previous or {}).get('images', {}).get(image)
targets.append(
{
'name': name,
'image': image,
'context': context,
'dockerfile': dockerfile,
'inputs': inputs,
'reuse_digest': old_digest if (previous or {}).get('inputs', {}).get(image) == inputs else None,
}
)
output.write_text(json.dumps({'sha': sha, 'targets': targets}, indent=2) + '\n')
if os.environ.get('GITHUB_OUTPUT'):
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
stream.write('matrix=' + json.dumps({'include': targets}, separators=(',', ':')) + '\n')
print(f'Prepared {len(targets)} image jobs; {sum(t["reuse_digest"] is None for t in targets)} require builds')
def checked_plan(path):
data = json.loads(path.read_text())
sha = command('git', 'rev-parse', 'HEAD')
if data.get('sha') != sha or sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
raise ValueError('Image plan does not match the checked source commit')
targets = data.get('targets', [])
if sorted(t['name'] for t in targets) != sorted(IMAGES):
raise ValueError('Image plan must contain each owned image once')
for target in targets:
name = target['name']
context, dockerfile = IMAGES[name]
if (target['context'], target['dockerfile'], target['image']) != (
context,
dockerfile,
f'gcr.forust.xyz/forust/{name}',
) or target['inputs'] != fingerprint(context, dockerfile):
raise ValueError('Image plan has invalid build inputs')
if target['reuse_digest'] is not None and not DIGEST.fullmatch(target['reuse_digest']):
raise ValueError('Image plan has an invalid reuse digest')
return data
def build_images(output, report, name, plan):
data = checked_plan(plan)
sha = data['sha']
target = next(t for t in data['targets'] if t['name'] == name)
context, dockerfile = IMAGES[name]
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
builder_config = Path.home() / '.cache/homelab-ci/buildx'
builder_config.mkdir(parents=True, exist_ok=True)
env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)}
try:
report['phase'] = 'Registry login'
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'login',
'gcr.forust.xyz',
'-u',
os.environ['REGISTRY_USERNAME'],
'--password-stdin',
],
input=os.environ['REGISTRY_PASSWORD'],
text=True,
check=True,
env=env,
)
report['phase'] = 'Prepare the builder'
builder = 'homelab-ci'
versions = dict(
re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE)
)
image = versions['BUILDKIT_IMAGE']
signature = builder_config / 'homelab-ci-image'
exists = (
subprocess.run( # noqa: S603
[shutil.which('docker') or '/usr/bin/docker', 'buildx', 'inspect', builder],
capture_output=True,
env=env,
).returncode
== 0
)
if exists and (not signature.exists() or signature.read_text().strip() != image):
command('docker', 'buildx', 'rm', '--keep-state', builder, env=env)
exists = False
if not exists:
command(
'docker',
'buildx',
'create',
'--name',
builder,
'--driver',
'docker-container',
'--driver-opt',
f'image={image}',
'--buildkitd-config',
'.gitea/runner/buildkitd.toml',
env=env,
)
signature.write_text(image + '\n')
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
report['images'] = release['images']
report['phase'] = f'Build or reuse {name}'
report['current'] = name
image = f'gcr.forust.xyz/forust/{name}'
inputs = target['inputs']
old_digest = target['reuse_digest']
exists = False
if old_digest:
exists = (
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'buildx',
'imagetools',
'inspect',
f'{image}@{old_digest}',
],
capture_output=True,
env=env,
timeout=60,
).returncode
== 0
)
if exists:
print(f'Reuse {name}: inputs unchanged')
digest = old_digest
else:
print(f'Build {name}', flush=True)
metadata = Path(docker_config) / 'metadata.json'
command(
'docker',
'buildx',
'build',
'--builder',
builder,
'--platform',
'linux/amd64',
'--provenance=false',
'--cache-from',
f'type=registry,ref={image}:buildcache',
'--cache-to',
f'type=registry,ref={image}:buildcache,mode=max',
'--output',
f'type=image,name={image},push-by-digest=true,name-canonical=true,push=true',
'--metadata-file',
str(metadata),
'--file',
dockerfile,
context,
env=env,
)
digest = json.loads(metadata.read_text())['containerimage.digest']
if not isinstance(digest, str) or not DIGEST.fullmatch(digest):
raise ValueError('Image job returned an invalid digest')
release['images'][image] = digest
release['inputs'][image] = inputs
report['reused' if exists else 'built'].append(name)
output.write_text(json.dumps(release, indent=2) + '\n')
report['current'] = None
report['phase'] = 'Image result file saved'
finally:
# Cleanup errors must neither leak credentials nor mask the original build error.
try:
subprocess.run( # noqa: S603
[
shutil.which('docker') or '/usr/bin/docker',
'buildx',
'prune',
'--builder',
'homelab-ci',
'--force',
'--max-used-space',
'1gb',
],
env=env,
timeout=60,
)
except (OSError, subprocess.TimeoutExpired):
print('CI builder cache cleanup deferred', flush=True)
finally:
shutil.rmtree(docker_config)
def write_summary(lines):
path = os.environ.get('GITHUB_STEP_SUMMARY')
if path:
try:
with Path(path).open('a') as stream:
stream.write('\n'.join(lines) + '\n\n')
except OSError:
print('WARNING: cannot write the job summary')
def check_summary():
lines = [
f'## {os.environ["SUMMARY_CHECK"]}',
'',
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
f'- Result: **{os.environ["SUMMARY_RESULT"]}**',
]
if os.environ.get('SUMMARY_FAILED_STEP'):
lines.append(f'- Failed step: {os.environ["SUMMARY_FAILED_STEP"]}')
if os.environ['SUMMARY_RESULT'] != 'success':
lines.append('- Open the failed step log for the error details.')
write_summary(lines)
def build(output, name, plan):
report = {'phase': 'Check the source commit', 'current': None, 'built': [], 'reused': [], 'images': {}}
result = 'failure'
try:
build_images(output, report, name, plan)
result = 'success'
finally:
lines = [
f'## Image build result `{name}`',
'',
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
'',
f'- Result: **{result}**',
f'- Last stage: {report["phase"]}',
]
if result == 'failure':
lines.append('- This image job failed. The complete release cannot be published. Open the failed step log.')
if result == 'success':
lines.append('- This is one image result. The final build job must publish the complete release.')
if report['current']:
lines.append(f'- Image at the failure: `{report["current"]}`')
for title, key in (('Built', 'built'), ('Reused from successful CI', 'reused')):
lines.extend(['', f'### {title}'])
lines.extend(f'- `{name}`' for name in report[key])
if not report[key]:
lines.append('- None')
lines.extend(['', '### Completed image digests'])
lines.extend(f'- `{image}@{digest}`' for image, digest in report['images'].items())
if not report['images']:
lines.append('- None')
write_summary(lines)
def render(stream, destination):
release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA'])
image_line = re.compile(
r"^(\s*(?:-\s*)?image:\s*)(['\"]?)(gcr\.forust\.xyz/forust/[\w.-]+)(?::[\w.-]+|@sha256:[0-9a-f]{64})\2(\s*(?:#.*)?)$"
)
rendered = []
for line in stream:
match = image_line.fullmatch(line.rstrip('\n'))
if match:
prefix, quote, image, tail = match.groups()
if image not in release['images']:
raise ValueError(f'Owned image missing from checked release: {image}')
line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n'
elif re.match(r'\s*(?:-\s*)?image:', line) and 'gcr.forust.xyz/forust/' in line:
raise ValueError('Unsupported owned image syntax; refusing to apply a mutable tag')
rendered.append(line)
destination.writelines(rendered)
def finalize_images(output, fragments, plan):
data = checked_plan(plan)
sha = data['sha']
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
for name in IMAGES:
fragment = json.loads((fragments / f'image-{name}' / 'image.json').read_text())
image = f'gcr.forust.xyz/forust/{name}'
if fragment.get('sha') != sha or fragment.get('version') != 1 or set(fragment.get('images', {})) != {image}:
raise ValueError('Image job artifact is missing or belongs to another commit')
target = next(t for t in data['targets'] if t['name'] == name)
if fragment.get('inputs') != {image: target['inputs']}:
raise ValueError('Image artifact does not match the build plan')
release['images'].update(fragment['images'])
release['inputs'].update(fragment['inputs'])
validate_release(release, sha)
# Only a complete set of successful image jobs can publish the release tags.
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
env = {**os.environ, 'DOCKER_CONFIG': docker_config}
try:
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'login',
'gcr.forust.xyz',
'-u',
os.environ['REGISTRY_USERNAME'],
'--password-stdin',
],
input=os.environ['REGISTRY_PASSWORD'],
text=True,
check=True,
env=env,
)
for image, digest in release['images'].items():
command(
'docker',
'buildx',
'imagetools',
'create',
'--prefer-index=false',
'--tag',
f'{image}:sha-{sha}',
f'{image}@{digest}',
env=env,
timeout=90,
)
output.write_text(json.dumps(release, indent=2) + '\n')
finally:
shutil.rmtree(docker_config)
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('prepare', 'image', 'finalize', 'gate', 'render', 'check-summary'))
parser.add_argument('--output', type=Path, default=Path('release.json'))
parser.add_argument('--ref', default='main')
parser.add_argument('--event-sha', default='')
parser.add_argument('--image', choices=IMAGES)
parser.add_argument('--plan', type=Path, default=Path('build-plan.json'))
parser.add_argument('--fragments', type=Path, default=Path('artifacts'))
args = parser.parse_args()
if args.action == 'check-summary':
check_summary()
elif args.action == 'render':
render(sys.stdin, sys.stdout)
elif args.action == 'gate':
gate(args.output, args.ref, args.event_sha)
elif args.action == 'prepare':
prepare_images(args.output)
elif args.action == 'image':
if not args.image:
parser.error('--image is required')
build(args.output, args.image, args.plan)
else:
finalize_images(args.output, args.fragments, args.plan)
if __name__ == '__main__':
main()
+41 -17
View File
@@ -1,10 +1,26 @@
name: renovate-ci
on:
pull_request:
# Read the workflow from the trusted base branch. PR code runs only on the
# unprivileged runner selected below.
pull_request_target:
paths:
- "renovate/**"
- ".gitea/workflows/renovate-ci.yaml"
- ".gitea/workflows/sync-renovate-configmap.sh"
- ".gitea/workflows/compose-lint.sh"
- ".gitea/workflows/install-ci-tools.sh"
- ".gitea/workflows/tool-versions.env"
push:
branches:
- main
paths:
- "renovate/**"
- ".gitea/workflows/renovate-ci.yaml"
- ".gitea/workflows/sync-renovate-configmap.sh"
- ".gitea/workflows/compose-lint.sh"
- ".gitea/workflows/install-ci-tools.sh"
- ".gitea/workflows/tool-versions.env"
workflow_dispatch:
permissions:
@@ -12,37 +28,47 @@ permissions:
jobs:
validate-renovate:
runs-on: [self-hosted, linux, arch, homelab]
runs-on: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 20
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ github.event_name == 'pull_request_target' && github.event.pull_request.head.sha || github.sha }}
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag,
# so the same version that runs in the cluster is the one validated here.
- name: Resolve the deployed Renovate image
# renovate/k8s/cronjob.yaml is the single source of truth for the version.
- name: Resolve the deployed Renovate version
id: image
shell: bash
run: |
set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)"
if [ -z "$image" ]; then
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
if [[ ! "$image" =~ ^renovate/renovate:([0-9]+\.[0-9]+\.[0-9]+)$ ]]; then
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
exit 1
fi
echo "using $image"
echo "image=$image" >> "$GITHUB_OUTPUT"
version="${BASH_REMATCH[1]}"
echo "using Renovate $version"
printf 'version=%s\n' "$version" >> "$GITHUB_OUTPUT"
- name: Validate Renovate repository config
- name: Prepare pinned validation tools
shell: bash
run: |
set -euo pipefail
docker run --rm \
-v "$PWD/renovate:/opt/renovate:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
"${{ steps.image.outputs.image }}" \
renovate-config-validator /opt/renovate/renovate.json
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform node)"
echo "$tools_dir" >> "$GITHUB_PATH"
- name: Validate Renovate repository config
shell: bash
env:
RENOVATE_VERSION: ${{ steps.image.outputs.version }}
run: |
set -euo pipefail
npm_cache="$(mktemp -d "${RUNNER_TEMP:-/tmp}/renovate-npm-cache.XXXXXXXX")"
trap 'rm -rf "$npm_cache"' EXIT
NPM_CONFIG_CACHE="$npm_cache" RENOVATE_CONFIG_FILE="$PWD/renovate/renovate.json" \
npm exec --yes --package="renovate@${RENOVATE_VERSION}" -- renovate-config-validator
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
# carries an inlined copy of the config. Fail if it no longer matches.
@@ -56,8 +82,6 @@ jobs:
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
export PATH="$tools_dir:$PATH"
kubeconform \
-strict \
-ignore-missing-schemas \
+12 -6
View File
@@ -32,11 +32,14 @@ concurrency:
jobs:
run-renovate:
runs-on: [self-hosted, linux, arch, homelab]
if: github.ref == 'refs/heads/main'
runs-on: homelab
timeout-minutes: 60
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: refs/heads/main
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
# Reading it here means this workflow validates and runs the exact version
@@ -48,21 +51,23 @@ jobs:
set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)"
if [ -z "$image" ]; then
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
if [[ ! "$image" =~ ^renovate/renovate:[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
exit 1
fi
echo "using $image"
echo "image=$image" >> "$GITHUB_OUTPUT"
printf 'image=%s\n' "$image" >> "$GITHUB_OUTPUT"
- name: Validate Renovate config
shell: bash
env:
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
run: |
set -euo pipefail
docker run --rm \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
"${{ steps.image.outputs.image }}" \
"$RENOVATE_IMAGE" \
renovate-config-validator
- name: Run Renovate
@@ -73,6 +78,7 @@ jobs:
RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
LOG_LEVEL: ${{ inputs.log_level }}
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
run: |
set -euo pipefail
@@ -89,4 +95,4 @@ jobs:
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
-e RENOVATE_BASE_DIR=/tmp/renovate \
-e LOG_LEVEL="${LOG_LEVEL:-info}" \
"${{ steps.image.outputs.image }}"
"$RENOVATE_IMAGE"
+13
View File
@@ -0,0 +1,13 @@
# kubectl emits a List for files containing multiple resources.
(if .kind == "List" then .items[] else . end)
| (.metadata.namespace // "default") as $ns
| [
(.. | objects
| (.secretRef? // empty), (.secretKeyRef? // empty), (.secret? // empty)
| select(.optional != true)
| .name // .secretName // empty),
(.. | objects | .imagePullSecrets[]?.name)
]
| unique[]
| select(. != null and . != "")
| "\($ns) \(.)"
+62 -63
View File
@@ -1,71 +1,70 @@
#!/usr/bin/env bash
# usage: ssh-run.sh <stage>
# Runs one deploy-lib.sh stage on the workstation over SSH.
# The SSH client submits once and follows durable stages on workstation.
set -euo pipefail
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
: "${DEPLOY_USER:?missing DEPLOY_USER}"
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
deploy_port="${DEPLOY_PORT:-22}"
deploy_path="${DEPLOY_PATH:-/srv/homelab}"
deploy_path="$(printf '%s' "$deploy_path" | tr -d '\"' | tr -d '\r' | xargs)"
# The private key is written to a per-run directory that is removed on exit, so a
# failed or cancelled job cannot leave deploy credentials in the runner's temp
# directory. Do not use a fixed path: apply-k8s and apply-compose run in parallel.
: "${DEPLOY_KNOWN_HOSTS:?configure pinned DEPLOY_KNOWN_HOSTS}"
: "${DEPLOY_RUN_ID:?missing DEPLOY_RUN_ID}"
[[ "$DEPLOY_USER" =~ ^[A-Za-z_][A-Za-z0-9_.-]*$ ]] || exit 1
[[ "$DEPLOY_HOST" =~ ^[A-Za-z0-9_.:-]+$ ]] || exit 1
[[ "$DEPLOY_RUN_ID" =~ ^[0-9]+-[0-9]+$ ]] || exit 1
[[ "${DEPLOY_PORT:-22}" =~ ^[0-9]+$ ]] || exit 1
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")"
trap 'rm -rf "$key_dir"' EXIT INT TERM
ssh_key="$key_dir/deploy_key"
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
chmod 600 "$ssh_key"
# A connection that died silently used to hang until the job timeout, and the
# stage was never re-run: one flaky TCP session cost a whole 45-minute apply.
# ServerAlive* bounds how long a dead peer goes unnoticed, ConnectTimeout bounds
# setup. Only exit 255 - ssh's own transport failures - is retried. A stage that
# fails on its own merits exits with the remote's status, so a real failure
# still surfaces its own log instead of burning three attempts. The stages are
# declarative applies, so re-running one that had already committed is harmless.
ssh_opts=(
-i "$ssh_key" -p "$deploy_port"
-o BatchMode=yes -o StrictHostKeyChecking=accept-new
-o ConnectTimeout=15
-o ServerAliveInterval=15 -o ServerAliveCountMax=4
)
rc=0
# apply-k8s and apply-compose are separate workflow jobs so the graph stays
# intact for the verify job, but on a single node they must not run at once:
# host docker churn on top of cluster churn is what melts the node (load 40+,
# netbird/ssh die, helm is left pending-*). Serialize them on the workstation
# with a shared lock; whoever arrives second waits.
remote_cmd=(bash -se)
case "$1" in
apply-k8s | apply-compose)
remote_cmd=(flock -w 5400 /tmp/homelab-apply.lock bash -se)
trap 'rm -rf "$key_dir"' EXIT
chmod 700 "$key_dir"
printf '%s\n' "$DEPLOY_KEY" >"$key_dir/key"
printf '%s\n' "$DEPLOY_KNOWN_HOSTS" >"$key_dir/known_hosts"
chmod 600 "$key_dir/key" "$key_dir/known_hosts"
ssh_opts=(-i "$key_dir/key" -p "${DEPLOY_PORT:-22}" -o BatchMode=yes -o StrictHostKeyChecking=yes
-o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15
-o ServerAliveInterval=15 -o ServerAliveCountMax=4)
controller=.local/lib/homelab-deploy/controller.py
case "${1:?start, apply, verify, smoke or summary required}" in
start)
python3 - <<'PY' >"$key_dir/request.json"
import json
import os
from pathlib import Path
release = json.loads(Path('release.json').read_text())
print(json.dumps({'release': release, 'mode': os.environ.get('DEPLOY_MODE', 'changed'),
'refresh_images': os.environ.get('REFRESH_IMAGES', 'false') == 'true'}))
PY
for attempt in 1 2 3; do
rc=0
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" start "$DEPLOY_RUN_ID" <"$key_dir/request.json" || rc=$?
[ "$rc" -eq 0 ] && exit 0
[ "$rc" -eq 255 ] || exit "$rc"
sleep 5
done
exit "$rc"
;;
apply|verify|smoke)
result=0
for attempt in 1 2 3; do
rc=0
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$?
[ "$rc" -eq 0 ] && break
[ "$rc" -eq 255 ] || { result="$rc"; break; }
echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)"
if [ "$attempt" -eq 3 ]; then result=255; break; fi
sleep 5
done
exit "$result"
;;
summary)
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
rc=0
# shellcheck disable=SC2029 # The run ID is validated above.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" summary "$DEPLOY_RUN_ID" >"$key_dir/deploy-summary.md" || rc=$?
if [ "$rc" -eq 0 ]; then
cat "$key_dir/deploy-summary.md" >>"$GITHUB_STEP_SUMMARY" || echo "WARNING: cannot write the deploy summary"
else
echo 'Deploy summary is unavailable. The SSH connection failed or the controller did not respond. Check the job log.' >>"$GITHUB_STEP_SUMMARY" || true
fi
fi
;;
*) echo "Unknown SSH operation: $1" >&2; exit 1 ;;
esac
for attempt in 1 2 3; do
if [ "$attempt" -gt 1 ]; then
echo ":: warning::ssh transport failed, retrying (${attempt}/3)"
sleep $((attempt * 5))
fi
rc=0
# shellcheck disable=SC2029 # remote_cmd/ssh_opts expand on the client on purpose: they select the local ssh invocation, only the heredoc runs remotely.
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
env "REPO=$deploy_path" "APPLY_PRUNE=${APPLY_PRUNE:-false}" \
"DEPLOY_SHA=${DEPLOY_SHA:-}" "DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-}" \
"STAGE=$1" "${remote_cmd[@]}" <<'EOF' || rc=$?
source "$REPO/.gitea/workflows/deploy-lib.sh"
run_stage "$STAGE"
EOF
[ "$rc" -eq 0 ] && break
[ "$rc" -ne 255 ] && break
done
if [ "$rc" -ne 0 ]; then
echo ":: error::stage $1 failed over ssh (exit $rc)"
fi
exit "$rc"
+7 -1
View File
@@ -14,7 +14,7 @@ ACTIONLINT_VERSION="1.7.7"
SHELLCHECK_VERSION="0.11.0"
KUBECONFORM_VERSION="0.8.0"
PRETTIER_VERSION="3.8.1"
RUFF_VERSION="0.16.8"
RUFF_VERSION="0.16.10"
YAMLLINT_VERSION="1.38.0"
HADOLINT_VERSION="2.14.0"
# pip-audit reads the advisory database over the network, so a floating version
@@ -31,3 +31,9 @@ UV_VERSION="0.12.17"
# so the tree that gets tested is the tree that gets built. Renovate keeps this
# in step with the Dockerfile's node: tag via the "node runtime" group.
NODE_VERSION="22.23.3"
# Secret-reference regression tests parse rendered Kubernetes objects.
JQ_VERSION="1.8.1"
# BuildKit is the only auxiliary CI container; jobs themselves stay on the host.
BUILDKIT_IMAGE="moby/buildkit:v0.33.1"
-1
View File
@@ -94,7 +94,6 @@ replacements.txt
.idea
# Temp files
edu_master/temp/
temp/*
# Local-only tooling scratch space (pinned CI tools, verification scripts)
tmp/
+49 -49
View File
@@ -2,8 +2,8 @@
Configuration for my homelab: Kubernetes manifests, Docker Compose stacks, and the
Gitea Actions that build and deploy them. Most applications have both deployment
formats. Headscale, Nextcloud AIO, and the media stack run on Docker; Kubernetes
provides their ingress through Services and EndpointSlices.
formats. Headscale and Nextcloud AIO have Compose deployments with Kubernetes
ingress; the media stack has Compose and Kubernetes routing configuration.
These files contain this lab's domains, IP addresses, storage paths, and private
registry names. Running them on another machine takes some editing.
@@ -12,7 +12,8 @@ registry names. Running them on another machine takes some editing.
- [Service list](#services) — what each directory contains.
- [Deployment workflow](.gitea/README.md) — selection, validation, and recovery.
- [Repository review](docs/repository-review.md) — confirmed problems and fix branches.
- [Repository review](docs/repository-review.md) — findings from the 6 October baseline and their status.
- [EDU ownership handoff](.gitea/EDU_HANDOFF.md) — the EDU workloads now live in their own repository.
- [Shared PostgreSQL](postgres/README.md), [Traefik](traefik/README.md), and
[cert-manager](cert-manager/README.md) — common dependencies.
@@ -40,48 +41,47 @@ service is currently healthy or running.
## Services
| Service | Configuration | Selected by markers |
| ----------------------------------------------------------- | ---------------------------- | ------------------- |
| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes |
| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes |
| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual |
| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes |
| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual |
| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual |
| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes |
| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual |
| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual |
| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual |
| [EDU session keeper and Telegram bot](edu_master/README.md) | Kubernetes + Compose | Kubernetes |
| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes |
| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes |
| [Glance](glance/README.md) | Kubernetes + Compose | Manual |
| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual |
| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes |
| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes |
| [Kener](kener/README.md) | Kubernetes + Compose | Manual |
| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes |
| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes |
| [n8n](n8n/README.md) | Kubernetes + Compose | Manual |
| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes |
| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes |
| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes |
| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Penpot](penpot/README.md) | Compose | Manual |
| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual |
| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes |
| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes |
| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes |
| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes |
| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes |
| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual |
| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Termix](termix/README.md) | Kubernetes + Compose | Manual |
| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes |
| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes |
| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes |
| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes |
| Service | Configuration | Selected by markers |
| ---------------------------------------------- | ---------------------------- | ------------------- |
| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes |
| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes |
| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual |
| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes |
| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual |
| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual |
| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes |
| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual |
| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual |
| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual |
| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes |
| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes |
| [Glance](glance/README.md) | Kubernetes + Compose | Manual |
| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual |
| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes |
| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes |
| [Kener](kener/README.md) | Kubernetes + Compose | Manual |
| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes |
| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes |
| [n8n](n8n/README.md) | Kubernetes + Compose | Manual |
| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes |
| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes |
| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes |
| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Penpot](penpot/README.md) | Compose | Manual |
| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual |
| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes |
| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes |
| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes |
| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes |
| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes |
| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual |
| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Manual |
| [Termix](termix/README.md) | Kubernetes + Compose | Manual |
| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes |
| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes |
| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes |
| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes |
## Running a Compose stack
@@ -136,16 +136,16 @@ role's password; see the database README.
CI pins its tools in `.gitea/workflows/tool-versions.env`. Use the same versions:
```sh
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh)"
export PATH="$tools_dir:$PATH"
```fish
set tools_dir (bash .gitea/workflows/install-ci-tools.sh)
set -gx PATH $tools_dir $PATH
ruff check .
ruff format --check .
actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
.gitea/workflows/sync-renovate-configmap.sh --check
```
The [workflow README](.gitea/README.md#checks) lists the rest of the checks.
The [workflow README](.gitea/README.md#ci) lists the rest of the checks.
Structure checks do not establish that local Secrets, mounted files, storage,
or external services are ready.
+1 -1
View File
@@ -31,7 +31,7 @@ services:
- "traefik.http.routers.adguard-dev.entrypoints=websecure"
- "traefik.http.routers.adguard-dev.tls=true"
# DoH Router
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz` || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`))"
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz`) || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`)"
- "traefik.http.routers.dns-over-https.entrypoints=websecure"
- "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt"
+2 -2
View File
@@ -51,6 +51,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: adguard-deployment
namespace: adguard
spec:
@@ -64,8 +66,6 @@ spec:
metadata:
labels:
app: adguard
annotations:
reloader.stakater.com/auto: "true"
spec:
containers:
- name: adguard
+4
View File
@@ -27,6 +27,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: authentik-server-deployment
namespace: authentik
spec:
@@ -63,6 +65,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: authentik-worker-deployment
namespace: authentik
spec:
+2
View File
@@ -1,6 +1,8 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: cfddns
labels:
app: cfddns
+2
View File
@@ -17,6 +17,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: checkmk-deployment
namespace: checkmk
spec:
+3 -1
View File
@@ -1,6 +1,8 @@
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: cloudflared
labels:
app: cloudflared
@@ -18,7 +20,7 @@ spec:
spec:
containers:
- name: cloudflared
image: cloudflare/cloudflared:2026.9.3
image: cloudflare/cloudflared:2026.10.0
imagePullPolicy: IfNotPresent
args:
- tunnel
+2
View File
@@ -13,6 +13,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: convertx-deployment
namespace: converters
spec:
+59 -76
View File
@@ -1,39 +1,40 @@
# Repository review
# Repository review (6 October 2026 baseline)
Reviewed the tracked tree at `cc9c3de` and read the live workstation state on
6 October 2026. Changes are split into documentation and individual fix branches,
all based on that main commit. The original local checkout and its uncommitted
monitoring changes were preserved. No deployment was performed.
This records the tracked tree at `cc9c3de` and the workstation state observed on
6 October 2026. It is a historical review, not a current runtime inventory. The
listed code fixes have since merged into `main`; EDU ownership has moved to the
separate repository described in [the handoff record](../.gitea/EDU_HANDOFF.md).
See the [CI and deployment guide](../.gitea/README.md) and
[runner and recovery guide](../.gitea/runner/README.md) for the current workflow.
No deployment was performed during the original review.
## Confirmed problems with prepared fixes
## Findings at the baseline and current status
| Priority | Problem and consequence | Fix branch |
| -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------- |
| High | `APPLY_PRUNE=true` is passed to each individual manifest apply. Each invocation sees only that file's desired objects and can delete other resources selected by the shared label. | `fix/deploy-prune-guard` |
| High | Deploy validates Compose with interpolation and env/path resolution disabled. Required settings can pass validation and then fail during apply after other workloads have changed. | `fix/deploy-validation` |
| Medium | Secret validation is text-based and compares names across all namespaces. A Secret elsewhere can hide a missing local Secret; mounted Secrets are also missed. | `fix/deploy-validation` |
| Medium | Compose CI misses `postgres/shared-compose.yaml`, `netbird/client.compose.yaml`, and `renovate/renovate-compose.yaml`. | `fix/deploy-validation` |
| Medium | NetBird Compose mounts `entrypoint.sh`, but it is absent. Its README also calls a missing `setup.sh`; a fresh checkout cannot start this stack as documented. | `fix/netbird-compose-runtime` |
| Medium | Glance's CSS mount uses `glance-config`, whose keys do not include `user.css`. That key is in `glance-assets`; the pod's subPath mount cannot be prepared correctly. | `fix/glance-assets` |
| Medium | The shared PostgreSQL initializer requires `NETBOX_DB_PASSWORD`, but the Compose env example omits it. Following the example leaves first initialization incomplete. | `fix/postgres-env-example` |
| Medium | EDU's Compose env example uses old credential names and full URL variables, while the code reads `KEEPER_*` and paths under `EDU_URL_BASE`. | `fix/session-keeper-reliability` |
| Medium | Session keeper HTTP calls have no timeouts. Its Redis cookie never expires, probes only check existence, and its logs include cookies. A hung or failed refresh can leave a stale session appearing ready. | `fix/session-keeper-reliability` |
| Medium | AdGuard's DoH and SearXNG's Compose rules put Boolean expressions inside `Host(...)`. They are invalid router expressions despite valid YAML. | `fix/compose-router-rules` |
| Priority | Finding at the baseline | Current status |
| -------- | ---------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| High | Per-file `APPLY_PRUNE=true` could delete resources selected by a shared label. | The deploy workflow rejects unsafe pruning before applying resources. |
| High | Compose validation did not resolve the local configuration required at deploy time. | Preflight resolves the selected Compose configuration before apply. |
| Medium | Secret validation could miss namespace-specific and mounted Secret references. | Preflight checks rendered references in their namespaces, including mounted and projected Secrets. |
| Medium | Compose CI missed manual entry points such as `shared-compose.yaml` and `client.compose.yaml`. | CI checks all tracked Compose files. |
| Medium | NetBird Compose referenced missing setup and renderer files. | The setup and renderer files are now present; Compose remains a manual alternative to the active Kubernetes deployment. |
| Medium | Glance mounted its CSS from the wrong ConfigMap. | The mount now uses the ConfigMap that contains `user.css`. |
| Medium | The PostgreSQL env example omitted the required NetBox password. | The example now includes the required variable. |
| Medium | The former EDU code had stale Compose variable names and session reliability problems. | EDU workloads and their fixes moved out of this repository; see the handoff record. |
| Medium | AdGuard DoH and SearXNG Compose router expressions used invalid `Host(...)` syntax. | The router expressions now follow Traefik's rule syntax. |
Traefik matchers should be combined as `Host(a) || Host(b)`; the rule syntax is
described in the [Traefik rules documentation](https://doc.traefik.io/traefik/reference/routing-configuration/http/routing/rules-and-priority/).
The fix retains the DoH path constraint for both hostnames.
The prune fix deliberately rejects the unsafe option. It does not introduce
automatic deletion under a different implementation. Prune defaults to false,
and no tracked resource currently carries the selector label, so this is a
latent defect rather than evidence of a live deletion incident.
The current deploy workflow deliberately rejects the unsafe prune option. It
does not introduce automatic deletion under a different implementation. The
baseline finding was a configuration risk, not evidence of a live deletion
incident.
The session fix bounds HTTP and Redis calls, validates required credentials,
sets a cookie lifetime of two refresh intervals, and marks success only after
publishing the verified cookie. With the default ten-minute interval, an outage
longer than twenty minutes will make the existing Redis-key readiness checks fail.
That is an intentional change from indefinite apparent readiness.
The former session fix bounded HTTP and Redis calls, validated credentials, set
a cookie lifetime of two refresh intervals, and marked success only after
publishing the verified cookie. The service is now owned by the EDU repository;
see that repository for its current implementation.
The deployment fix extracts required pod Secret references from rendered JSON,
checks their namespaces, includes init containers, image-pull credentials, and
@@ -61,51 +62,34 @@ reviewed local commit. It has untracked host configuration and a separate
| Default `local-path` has reclaim policy Delete, while many existing PVs have been changed to Retain. | Current retention is partly live state. Recreating a claim can get a different policy from the old PV. |
| NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. | Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup. |
The monitoring files already modified in the user's local tree correspond to the
live migration. They are excluded from these branches. Reconcile that work before
using this review's baseline to deploy monitoring.
The VictoriaMetrics monitoring trial later merged into `main` in PR #95. The
first row above records the state before that change. Read
[`prometheus-stack/README.md`](../prometheus-stack/README.md) for the current
tracked monitoring configuration; the live observations in this section remain
a snapshot from 6 October.
## Remaining work
## Current recovery limits
These need recovery design or infrastructure decisions rather than a small
configuration correction:
The deployment controller and its recovery process changed after this review.
The current operator workflow is documented in the
[runner and recovery guide](../.gitea/runner/README.md). The remaining boundaries
are:
- **SSH apply retries can replace the rollback baseline.** `ssh-run.sh` retries
exit 255, including `apply-k8s`; every new invocation publishes a fresh snapshot.
If the first attempt already changed workloads, the retry snapshots that partial
state. Preserve a run-specific original baseline and verify it across retries.
- **Rollback can exceed the job budget.** Verification is parallel, but
`rollback_workloads` is serial with a five-minute limit per workload. The
thirty-minute job budget can expire before recovery finishes. Bound recovery
concurrency and account for both phases before choosing a new timeout.
- **Snapshot collection is allowed to fail.** Generation and workload snapshot
errors are warnings; verify can fall back to all workloads. A snapshot failure
must not permit unrelated workloads to be selected for automatic undo.
- **Rollback uses the previous revision, not the captured revision.** `rollout undo`
without an explicit revision cannot guarantee restoration to the snapshot after
retries or intervening rollouts. First deployments also have no previous revision.
- **Manual deploy dispatch bypasses the CI-success trigger.** Either validate the
target commit's successful CI run or document manual dispatch as an operator
override with its own required checks.
- **Direct Traefik API exposure is unauthenticated.** The latest local commit
explicitly added it for Homarr. Preserve that integration while choosing a
cluster-internal authenticated path or a verified network restriction; do not
simply disable an integration that is already in use.
- **Storage retention and backup are not reproducible as a whole.** Defaults and
several important PV policies are Delete. There is no repository-wide backup
schedule. Existing PVC StorageClass changes require migration rather than an
in-place YAML edit.
- **MeTube downloads are temporary on Kubernetes.** `/downloads` is a 20 GiB
emptyDir. Decide whether pod replacement should discard files or whether it
should use persistent storage. Compose uses a host directory instead.
- **First-time activation needs a bootstrap path.** Deploy validation dry-runs
namespaced resources before the apply stage creates namespaces and installs
selected charts. On a fresh cluster, missing namespaces and CRDs need separate
preparation; activation is not a complete installer.
- Kubernetes recovery can restore captured workload revisions. It does not
restore ConfigMaps, Secrets, database schemas, or persistent data.
- Compose recovery is manual. It uses saved resolved configuration, but it does
not restore volume data or reverse database migrations.
- Removed resources require manual review and removal; the deploy workflow does
not prune them automatically.
- Plan mode does not create namespaces. During apply, server validation for new
namespaces runs after namespace creation and chart installation; a failed
check can leave an empty namespace.
- Storage policy and backup coverage remain service-specific. Check the live PV,
PVC, and backup state before changing stateful workloads.
## Validation
Baseline lint checks passed for Python, shell, workflows, YAML, standard Compose
At the review baseline, lint checks passed for Python, shell, workflows, YAML, standard Compose
files, and Kubernetes resources with available schemas. Kubeconform found 347
resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources.
That skip count matters: passing schema validation does not validate Traefik rule
@@ -124,8 +108,8 @@ Fix validation covers:
documented Traefik grammar. They were not exercised on the live proxy.
- Prune rejection before any cluster invocation.
All seven fix branches and the documentation branch merged together without
conflicts in a disposable validation worktree. The combined tree passed the
At the time of review, all seven fix branches and the documentation branch
merged together in a disposable validation worktree. That combined tree passed the
CI-equivalent local checks, Markdown formatting/lint and link checks, all 35
Compose structure checks, and 11 Python regression tests plus the shell
validation regressions. CRD server-side validation and live rollout tests were
@@ -135,17 +119,16 @@ Runtime tests use fixtures and mocks, not production credentials. Live checks re
workload metadata, storage policies, chart versions, and container state only.
They did not read Secret contents or change services.
## Reloader follow-up
## Reloader follow-up (baseline)
`fix/reloader-integration` adds the active marker and opt-in annotations to 28
`fix/reloader-integration` added the active marker and opt-in annotations to
application Deployments/StatefulSets that consume runtime ConfigMaps or Secrets.
It corrects AdGuard's misplaced pod-template annotation. The Helm settings use
It corrected AdGuard's misplaced pod-template annotation. The Helm settings use
annotation-based reloads, keep global auto-reload disabled, and ignore Jobs and
CronJobs. PostgreSQL workloads are excluded because their credential variables
and init scripts are only effective on an empty data directory.
The controller was already running on workstation when inspected. Its live
configuration is unchanged by the branch: merge and deploy the integration to
apply the new policy and application annotations. Configuration reload behavior
was checked against the pinned chart, with Helm rendering and manifest validation;
no production configuration was changed to provoke a test restart.
The controller was running on the workstation when inspected. The original
review checked configuration against the pinned chart with Helm rendering and
manifest validation; it did not change production configuration to provoke a
test restart or confirm every application's live reload behavior.
-14
View File
@@ -1,14 +0,0 @@
EDU_LOGIN=your_edu_login_here
EDU_PASSWORD=your_edu_password_here
EDU_URL_LOGIN=https://edu.edu.vn.ua/user/login
EDU_URL_VERIFY=https://edu.edu.vn.ua/course/userlist
PHPSESSID_INTERVAL=10
USER_AGENT="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36"
WEBINAR_URL=https://edu.edu.vn.ua/webinar/useractive
WEBINAR_CHECK_INTERVAL=60
REDIS_HOST=redis
REDIS_PORT=6379
PLAYWRIGHT_WS=ws://playwright-service:3000/ws
TZ=Europe/Kyiv
WEBINAR_TELEGRAM_TOKEN=your_telegram_bot_token_here
WEBINAR_ADMIN_ID=123456789
-1
View File
@@ -1 +0,0 @@
1.56.0
-52
View File
@@ -1,52 +0,0 @@
# EDU session keeper and Telegram bot
Keeps an EDU login session in Redis and sends Telegram notifications for new webinars. The bot also serves diary and schedule commands.
`phpsessid-bot/` logs into EDU and publishes `EDU_PHPSESSID` in Redis.
`webinar-checker/` uses that cookie through a remote Playwright browser and stores
subscribers, language preferences, and webinar history in Redis.
Kubernetes runs in `edu-master`, with Redis data in `redis-data-pvc`.
`service.yaml`, `servicemonitor.yaml`, and `alerts.yaml` expose and monitor the
checker's metrics on port 8000. Its `/health` endpoint reflects recent checks.
## Configuration
Use the keys in `k8s/secrets.yaml.example` as the reference. The committed Compose
`.env.example` has stale names until `fix/session-keeper-reliability` is merged.
The code reads:
| Variable | Purpose |
| ----------------------------------------------------- | ------------------------------------------------ |
| `KEEPER_LOGIN`, `KEEPER_PASSWORD` | EDU login credentials. |
| `KEEPER_INTERVAL` | Session refresh interval in minutes; default 10. |
| `EDU_URL_BASE` | EDU site origin. |
| `EDU_URL_LOGIN`, `EDU_URL_COURSES`, `EDU_URL_WEBINAR` | Paths under that origin. |
| `WEBINAR_TELEGRAM_TOKEN`, `WEBINAR_ADMIN_ID` | Telegram bot and administrator. |
| `WEBINAR_CHECK_INTERVAL` | Checker interval in seconds; default 60. |
| `REDIS_HOST`, `REDIS_PORT` | Redis connection. |
| `PLAYWRIGHT_WS` | Remote browser WebSocket endpoint. |
Set the keeper keys explicitly in the Compose `.env`. Keep the Playwright Python
package, browser image, server command, and `PLAYWRIGHT_VERSION` file on matching
versions. The two Python images are built and published by CI.
## Bot use
Start a private chat with `/start` to subscribe. `/stop`, `/language`, `/diary`,
`/schedule`, and `/setclass` manage subscriptions and school views. The
administrator can manage the whitelist with `/adduser` and `/removeuser`.
Back up Redis if subscriber settings and notification history matter. Session
cookies and Telegram tokens are credentials; keep them out of shared logs.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n edu-master
kubectl get events -n edu-master --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
-49
View File
@@ -1,49 +0,0 @@
services:
redis:
image: redis:8.10.2-alpine
restart: unless-stopped
volumes:
- redis-data:/data
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 5s
timeout: 3s
retries: 5
playwright-service:
image: mcr.microsoft.com/playwright:v1.56.0-jammy
restart: unless-stopped
command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws
session-keeper:
build: ./phpsessid-bot
image: gcr.forust.xyz/forust/session-keeper:prod
pull_policy: build
env_file: .env
restart: unless-stopped
depends_on:
redis:
condition: service_healthy
healthcheck:
test: ["CMD-SHELL", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
interval: 30s
timeout: 5s
retries: 10
start_period: 60s
webinar-checker:
build: ./webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:prod
pull_policy: build
env_file: .env
restart: unless-stopped
depends_on:
redis:
condition: service_healthy
session-keeper:
condition: service_healthy
playwright-service:
condition: service_started
volumes:
redis-data:
-96
View File
@@ -1,96 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: edu-master-webinar
namespace: edu-master
labels:
release: prometheus-stack
spec:
groups:
- name: edu_master.webinar
rules:
# No successful webinar check for 5m (~2-3 missed 2-min checks).
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
# The last_success > 0 guard is mandatory: checker.py initialises
# last_success to 0, so without it `time() - 0` equals the current epoch
# and humanizeDuration renders ~20722d on every pod restart. Keep the
# duration expression on the left so $value stays the real gap.
- alert: WebinarCheckerNoSuccessfulCheck
expr: |
((time() - webinar_check_last_success_timestamp_seconds) > 300)
and (webinar_check_last_success_timestamp_seconds > 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 2m
labels:
severity: critical
annotations:
summary: "Webinar checker has no successful check for 5m"
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
# Checks are running but none has ever succeeded since pod start.
# Split out from the rule above so a zeroed gauge never feeds
# humanizeDuration.
- alert: WebinarCheckerNeverSucceeded
expr: |
(webinar_check_last_success_timestamp_seconds == 0)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker has never completed a successful check"
description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
- alert: WebinarCheckerConsecutiveFailures
expr: |
webinar_check_consecutive_failures >= 3
for: 5m
labels:
severity: critical
annotations:
summary: "Webinar checker failing consecutively"
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / playwright error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
- alert: WebinarCheckerScrapeDown
expr: |
absent(webinar_check_last_run_timestamp_seconds) == 1
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker metrics missing"
description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
- alert: EduPhpsessidMissing
expr: |
edu_phpsessid_present == 0
for: 10m
labels:
severity: critical
annotations:
summary: "EDU_PHPSESSID missing"
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
# Hard deps: checker and playwright deployments unavailable.
- alert: WebinarCheckerDeploymentDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker deployment unavailable"
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
- alert: PlaywrightServiceDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="playwright-service", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Playwright service unavailable"
description: "edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it."
-4
View File
@@ -1,4 +0,0 @@
apiVersion: v1
kind: Namespace
metadata:
name: edu-master
-69
View File
@@ -1,69 +0,0 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: playwright-service
namespace: edu-master
labels:
app: edu-master-playwright
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-playwright
strategy:
type: Recreate
template:
metadata:
labels:
app: edu-master-playwright
spec:
containers:
- name: playwright
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker
image: mcr.microsoft.com/playwright:v1.56.0-jammy
imagePullPolicy: IfNotPresent
# p95 412M, max 478M over 7 days, no limit before. Request is set at p95
# so the pod is not an eviction candidate; the limit stays above 2x the
# request because browser page lifetimes are unpredictable.
resources:
requests:
cpu: "200m"
memory: "416Mi"
limits:
memory: "1Gi"
command:
- npx
- -y
- playwright@1.56.0
- run-server
- --port
- "3000"
- --path
- /ws
ports:
- containerPort: 3000
readinessProbe:
tcpSocket:
port: 3000
initialDelaySeconds: 5
periodSeconds: 10
timeoutSeconds: 3
livenessProbe:
tcpSocket:
port: 3000
initialDelaySeconds: 15
periodSeconds: 20
timeoutSeconds: 3
---
apiVersion: v1
kind: Service
metadata:
name: playwright-service
namespace: edu-master
spec:
selector:
app: edu-master-playwright
ports:
- name: ws
port: 3000
targetPort: 3000
-75
View File
@@ -1,75 +0,0 @@
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: redis
namespace: edu-master
labels:
app: edu-master-redis
spec:
serviceName: redis
replicas: 1
selector:
matchLabels:
app: edu-master-redis
template:
metadata:
labels:
app: edu-master-redis
spec:
containers:
- name: redis
image: redis:8.10.2-alpine
imagePullPolicy: IfNotPresent
ports:
- containerPort: 6379
volumeMounts:
- name: redis-data
mountPath: /data
resources:
requests:
cpu: 25m
memory: 32Mi
limits:
cpu: 250m
memory: 128Mi
readinessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 3
livenessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 3
volumes:
- name: redis-data
persistentVolumeClaim:
claimName: redis-data-pvc
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: redis-data-pvc
namespace: edu-master
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 1Gi
---
apiVersion: v1
kind: Service
metadata:
name: redis
namespace: edu-master
spec:
selector:
app: edu-master-redis
ports:
- name: redis
port: 6379
targetPort: 6379
@@ -1,50 +0,0 @@
# One-time Job to migrate redis state from docker compose to k8s (maintenance window).
# The .example file is not applied by the deploy pipeline (mask *.example.yaml).
#
# Runbook:
# 1. docker compose -f <repo>/edu_master/compose.yaml stop # SIGTERM -> redis will flush dump.rdb
# 2. docker run --rm -v edu_master_redis-data:/data \
# -v /tmp/edu-master-backup:/backup \
# redis:alpine sh -c "cp /data/dump.rdb /backup/ && ls -la /backup"
# 3. kubectl apply -f edu_master/k8s/namespace.yaml
# 4. kubectl apply -f <only the PVC from redis.yaml> # seed must come BEFORE redis pod starts
# 5. kubectl apply -f edu_master/k8s/restore-seed-job.yaml.example
# kubectl wait --for=condition=complete job/redis-restore-seed -n edu-master --timeout=120s
# 6. kubectl delete job redis-restore-seed -n edu-master
# 7. kubectl apply -f edu_master/k8s/ -R # apply remaining manifests
apiVersion: batch/v1
kind: Job
metadata:
name: redis-restore-seed
namespace: edu-master
spec:
backoffLimit: 2
ttlSecondsAfterFinished: 3600
template:
spec:
restartPolicy: Never
containers:
- name: seed
image: redis:alpine
command:
- /bin/sh
- -ec
- |
ls -la /backup
cp /backup/dump.rdb /data/dump.rdb
chmod 644 /data/dump.rdb
ls -la /data
volumeMounts:
- name: redis-data
mountPath: /data
- name: backup
mountPath: /backup
readOnly: true
volumes:
- name: redis-data
persistentVolumeClaim:
claimName: redis-data-pvc
- name: backup
hostPath:
path: /tmp/edu-master-backup
type: DirectoryOrCreate
-29
View File
@@ -1,29 +0,0 @@
apiVersion: v1
kind: Secret
metadata:
name: edu-master-secrets
namespace: edu-master
type: Opaque
stringData:
# Session keeper credentials
KEEPER_LOGIN: ""
KEEPER_PASSWORD: ""
KEEPER_INTERVAL: "10"
# EDU links
EDU_URL_BASE: "https://edu.edu.vn.ua"
EDU_URL_LOGIN: "/user/login"
EDU_URL_COURSES: "/course/userlist"
EDU_URL_WEBINAR: "/webinar/useractive"
# Playwright
USER_AGENT: ""
PLAYWRIGHT_WS: "ws://playwright-service:3000/ws"
# Webinar-checker
WEBINAR_TELEGRAM_TOKEN: ""
WEBINAR_ADMIN_ID: ""
WEBINAR_CHECK_INTERVAL: "60"
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
METRICS_PORT: "8000"
# Database
REDIS_HOST: "redis"
REDIS_PORT: "6379"
TZ: "Europe/Kyiv"
-15
View File
@@ -1,15 +0,0 @@
apiVersion: v1
kind: Service
metadata:
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
selector:
app: edu-master-webinar-checker
ports:
- name: metrics
port: 8000
targetPort: metrics
protocol: TCP
-53
View File
@@ -1,53 +0,0 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: session-keeper
namespace: edu-master
labels:
app: edu-master-session-keeper
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-session-keeper
strategy:
type: Recreate
template:
metadata:
labels:
app: edu-master-session-keeper
spec:
initContainers:
- name: wait-redis
image: redis:8.10.2-alpine
command:
- /bin/sh
- -ec
- |
i=0
until redis-cli -h redis ping | grep -q PONG; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
sleep 2
done
echo "redis is ready"
containers:
- name: session-keeper
image: gcr.forust.xyz/forust/session-keeper:prod
envFrom:
- secretRef:
name: edu-master-secrets
resources:
requests:
cpu: 25m
memory: 32Mi
limits:
cpu: 250m
memory: 128Mi
readinessProbe:
exec:
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
initialDelaySeconds: 15
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 10
-75
View File
@@ -1,75 +0,0 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-webinar-checker
strategy:
type: Recreate
template:
metadata:
labels:
app: edu-master-webinar-checker
spec:
# Enforces dependency order like compose depends_on:
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) -> playwright started
initContainers:
- name: wait-deps
image: redis:8.10.2-alpine
command:
- /bin/sh
- -ec
- |
i=0
until redis-cli -h redis ping | grep -q PONG; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
sleep 2
done
echo "redis ok"
until [ "$(redis-cli -h redis EXISTS EDU_PHPSESSID)" = "1" ]; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: no PHPSESSID (session-keeper down?)" && exit 1
sleep 2
done
echo "PHPSESSID ok"
until nc -z playwright-service 3000; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: playwright-service not reachable" && exit 1
sleep 2
done
echo "playwright ok"
containers:
- name: webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:prod
ports:
- name: metrics
containerPort: 8000
protocol: TCP
readinessProbe:
httpGet:
path: /health
port: metrics
periodSeconds: 10
timeoutSeconds: 3
failureThreshold: 12
initialDelaySeconds: 10
envFrom:
- secretRef:
name: edu-master-secrets
env:
- name: TZ
value: "Europe/Kyiv"
resources:
requests:
cpu: "50m"
memory: "192Mi"
limits:
cpu: "600m"
memory: "384Mi"
-15
View File
@@ -1,15 +0,0 @@
FROM python:3.11-slim
WORKDIR /app
# Install system dependencies
RUN apt-get update && apt-get install -y --no-install-recommends redis-tools && rm -rf /var/lib/apt/lists/*
# Install dependencies
RUN pip install --no-cache-dir requests==2.32.3 redis==5.2.1
# Copy application code
COPY . .
# Run the bot
CMD ["python", "bot.py"]
-132
View File
@@ -1,132 +0,0 @@
import logging
import os
import time
from datetime import datetime
import redis
import requests
# Configure logging
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')
logger = logging.getLogger(__name__)
# Load configuration (adapted to .env keys)
def _env(key, default=None):
v = os.getenv(key, default)
if isinstance(v, str) and len(v) >= 2 and ((v[0] == '"' and v[-1] == '"') or (v[0] == "'" and v[-1] == "'")):
return v[1:-1]
return v
LOGIN = _env('KEEPER_LOGIN')
PASSWORD = _env('KEEPER_PASSWORD')
EDU_BASE = _env('EDU_URL_BASE', 'https://edu.edu.vn.ua')
EDU_LOGIN_PATH = _env('EDU_URL_LOGIN', '/user/login')
EDU_COURSES_PATH = _env('EDU_URL_COURSES', '/course/userlist')
URL_LOGIN = f'{EDU_BASE.rstrip("/")}/{EDU_LOGIN_PATH.lstrip("/")}'
URL_VERIFY = f'{EDU_BASE.rstrip("/")}/{EDU_COURSES_PATH.lstrip("/")}'
INTERVAL = int(_env('KEEPER_INTERVAL', 10))
USER_AGENT = _env(
'USER_AGENT',
'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36',
)
REDIS_HOST = _env('REDIS_HOST', 'redis')
REDIS_PORT = int(_env('REDIS_PORT', 6379))
SUCCESS_FILE = '/tmp/last_success' # noqa: S108
def touch_success_file():
"""Updates the timestamp of the success file for healthchecks."""
try:
with open(SUCCESS_FILE, 'w') as f:
f.write(str(datetime.now().timestamp()))
except Exception as e:
logger.error(f'Failed to touch success file: {e}')
def main():
logger.info('Starting Session Keeper Bot')
# Connect to Redis
try:
redis_client = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
redis_client.ping()
logger.info(f'Connected to Redis at {REDIS_HOST}:{REDIS_PORT}')
except Exception as e:
logger.error(f'Failed to connect to Redis: {e}')
return
session = requests.Session()
# Set headers
headers = {
'User-Agent': USER_AGENT,
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
'Accept-Language': 'en-US,en;q=0.9',
'Cache-Control': 'max-age=0',
'Upgrade-Insecure-Requests': '1',
'Sec-Fetch-Site': 'same-origin',
'Sec-Fetch-Mode': 'navigate',
'Sec-Fetch-User': '?1',
'Sec-Fetch-Dest': 'document',
'Sec-Ch-Ua': '"Not_A Brand";v="99", "Chromium";v="142"',
'Sec-Ch-Ua-Mobile': '?0',
'Sec-Ch-Ua-Platform': '"Linux"',
'Accept-Encoding': 'gzip, deflate, br',
'Priority': 'u=0, i',
}
session.headers.update(headers)
while True:
try:
logger.info('Attempting login...')
# Login payload
payload = {'login': LOGIN, 'password': PASSWORD}
# Perform Login
# Note: The user request shows a POST to /user/login with form data
# We need to make sure we handle the PHPSESSID correctly.
# If we already have a PHPSESSID, requests will send it.
login_response = session.post(URL_LOGIN, data=payload, allow_redirects=True)
logger.info(f'Login Response Status: {login_response.status_code}')
logger.info(f'Cookies after login: {session.cookies.get_dict()}')
# Verify Session
logger.info('Verifying session...')
verify_response = session.get(URL_VERIFY, allow_redirects=False)
logger.info(f'Verify Response Status: {verify_response.status_code}')
if verify_response.status_code == 200:
logger.info('Session verification SUCCESS (200 OK).')
touch_success_file()
# Save PHPSESSID to Redis
phpsessid = session.cookies.get('PHPSESSID')
if phpsessid:
try:
redis_client.set('EDU_PHPSESSID', phpsessid)
logger.info(f'Saved PHPSESSID to Redis: {phpsessid}')
except Exception as e:
logger.error(f'Failed to save PHPSESSID to Redis: {e}')
elif verify_response.status_code == 302:
logger.warning('Session verification FAILED (302 Redirect). Session might be invalid.')
else:
logger.warning(f'Session verification returned unexpected status: {verify_response.status_code}')
except Exception as e:
logger.error(f'An error occurred: {e}')
logger.info(f'Sleeping for {INTERVAL} minutes...')
time.sleep(INTERVAL * 60)
if __name__ == '__main__':
main()
-13
View File
@@ -1,13 +0,0 @@
FROM python:3.11-slim
WORKDIR /app
# renovate: datasource=pypi depName=playwright versioning=pep440
ARG PLAYWRIGHT_VERSION=1.56.0
# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
COPY checker.py .
CMD ["python", "checker.py"]
File diff suppressed because it is too large. Load diff
+2
View File
@@ -20,6 +20,8 @@ data:
GITEA__mailer__ENABLED: "false"
GITEA__metrics__ENABLED: "true"
# No code/issue search needed: bleve reindexes the whole issue index on
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
# the rotational disk for an hour. "db" serves issue search from postgres.
+4
View File
@@ -3,6 +3,8 @@ kind: Service
metadata:
name: gitea-service
namespace: gitea
labels:
app: gitea
spec:
selector:
app: gitea
@@ -17,6 +19,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: gitea-deployment
namespace: gitea
spec:
+2 -1
View File
@@ -7,7 +7,8 @@ spec:
entryPoints:
- websecure
routes:
- match: Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)
# Metrics are scraped directly through the cluster Service.
- match: (Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)) && !PathPrefix(`/metrics`)
kind: Rule
services:
- name: gitea-service
+16
View File
@@ -0,0 +1,16 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: gitea
namespace: gitea
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: gitea
endpoints:
- port: http
path: /metrics
interval: 30s
scrapeTimeout: 10s
+2 -2
View File
@@ -11,8 +11,8 @@ Docker socket. It references `glance-secrets` for `ADGUARD_PASSWORD`, but there
no tracked Secret example; create that Secret in `glance` before starting it.
Compose expects a local `.env` with the same password.
The CSS mount points at the wrong ConfigMap on the reviewed main commit;
`fix/glance-assets` corrects it.
The pod mounts `user.css` from `glance-assets`, which is the ConfigMap that
contains that key.
## Inspect
+3 -1
View File
@@ -13,6 +13,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: glance-deployment
namespace: glance
spec:
@@ -69,7 +71,7 @@ spec:
name: glance-config
- name: glance-assets
configMap:
name: glance-config
name: glance-assets
- name: docker-socket
hostPath:
path: /var/run/docker.sock
+1 -1
View File
@@ -7,7 +7,7 @@
# # Dev server_url
# server_url: https://hs.dev_internal_domain.internal
listen_addr: 0.0.0.0:8080
metrics_listen_addr: 127.0.0.1:9090
metrics_listen_addr: 0.0.0.0:9090
grpc_listen_addr: 127.0.0.1:50443
grpc_allow_insecure: false
noise:
@@ -3,6 +3,8 @@ kind: Service
metadata:
name: headscale-server-external
namespace: headscale
labels:
app: headscale
spec:
ports:
- port: 8080
+18
View File
@@ -0,0 +1,18 @@
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMServiceScrape
metadata:
name: headscale
namespace: headscale
labels:
release: prometheus-stack
spec:
# The external Service has a manually managed EndpointSlice, not Endpoints.
discoveryRole: endpointslice
selector:
matchLabels:
app: headscale
endpoints:
- port: metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
+1 -1
View File
@@ -1,7 +1,7 @@
services:
homarr:
container_name: homarr
image: ghcr.io/homarr-labs/homarr:v2.1.2
image: ghcr.io/homarr-labs/homarr:v2.2.0
restart: unless-stopped
volumes:
- ./appdata:/appdata
+3 -1
View File
@@ -13,6 +13,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: homarr-deployment
namespace: homarr
spec:
@@ -30,7 +32,7 @@ spec:
serviceAccountName: homarr
containers:
- name: homarr
image: ghcr.io/homarr-labs/homarr:v2.1.2
image: ghcr.io/homarr-labs/homarr:v2.2.0
envFrom:
- configMapRef:
name: homarr-config
+4
View File
@@ -6,6 +6,10 @@ metadata:
data:
TZ: "Europe/Bratislava"
IMMICH_TELEMETRY_INCLUDE: "all"
IMMICH_API_METRICS_PORT: "8081"
IMMICH_MICROSERVICES_METRICS_PORT: "8082"
# The database in this namespace, not the shared one in the database
# namespace: v3 needs VectorChord, and only the dedicated image carries it.
DB_HOSTNAME: "immich-postgres"
+14
View File
@@ -3,6 +3,8 @@ kind: Service
metadata:
name: immich-service
namespace: immich
labels:
app: immich
spec:
selector:
app: immich
@@ -10,10 +12,18 @@ spec:
- name: http
port: 2283
targetPort: 2283
- name: api-metrics
port: 8081
targetPort: api-metrics
- name: worker-metrics
port: 8082
targetPort: worker-metrics
---
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: immich-deployment
namespace: immich
labels:
@@ -39,6 +49,10 @@ spec:
ports:
- name: http
containerPort: 2283
- name: api-metrics
containerPort: 8081
- name: worker-metrics
containerPort: 8082
volumeMounts:
- name: immich-data
mountPath: /data
+2
View File
@@ -14,6 +14,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: immich-machine-learning-deployment
namespace: immich
labels:
+20
View File
@@ -0,0 +1,20 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: immich
namespace: immich
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: immich
endpoints:
- port: api-metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
- port: worker-metrics
path: /metrics
interval: 30s
scrapeTimeout: 10s
+2
View File
@@ -17,6 +17,8 @@ spec:
apiVersion: apps/v1
kind: StatefulSet
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: immich-valkey
namespace: immich
labels:
+1 -1
View File
@@ -1,6 +1,6 @@
services:
kener:
image: rajnandan1/kener:4.1.5
image: rajnandan1/kener:v4.1.7
container_name: kener
restart: unless-stopped
# ports:
+3 -1
View File
@@ -13,6 +13,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: kener-deployment
namespace: kener
spec:
@@ -29,7 +31,7 @@ spec:
spec:
containers:
- name: kener
image: rajnandan1/kener:4.1.5
image: rajnandan1/kener:v4.1.7
envFrom:
- configMapRef:
name: kener-config
+2
View File
@@ -13,6 +13,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: metube-deployment
namespace: metube
spec:
+1 -1
View File
@@ -1,6 +1,6 @@
services:
n8n:
image: docker.n8n.io/n8nio/n8n:2.42.3
image: docker.n8n.io/n8nio/n8n:2.43.2
container_name: n8n
restart: unless-stopped
environment:
+3 -1
View File
@@ -13,6 +13,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: n8n-deployment
namespace: n8n
spec:
@@ -29,7 +31,7 @@ spec:
spec:
containers:
- name: n8n
image: docker.n8n.io/n8nio/n8n:2.42.3
image: docker.n8n.io/n8nio/n8n:2.43.2
envFrom:
- configMapRef:
name: n8n-config
+113 -42
View File
@@ -1,64 +1,135 @@
# NetBird
Self-hosted NetBird with the combined management, signal, relay, and STUN server.
Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind Traefik. The Compose configuration uses the external Docker `proxy` network and publishes only STUN UDP `3478` directly.
Kubernetes runs the server and dashboard in `netbird`. The server uses SQLite
in `netbird-pvc`; `k8s/config.yaml` contains the template and runtime renderer.
Prepare `netbird-secrets` from `k8s/secrets.yaml.example` before the first start.
The single-instance server uses SQLite. Back up its data and datastore encryption key together.
## Routing and keys
## Kubernetes
The public hostname is set in the ConfigMap and ingress rules. Keep the issuer,
dashboard endpoints, and public routes consistent. HTTP, WebSocket, and gRPC
traffic go through Traefik; the STUN route uses UDP 3478. A CDN's HTTP proxy does
not provide that UDP listener.
`NETBIRD_PROXY_SUBNET` controls which forwarded client addresses are trusted.
Use the actual proxy network CIDR rather than assuming another lab's subnet.
Keep the datastore encryption key with every datastore backup. Regenerating it
can make stored credentials unreadable.
`k8s/active` selects the Kubernetes deployment. It runs the server and dashboard
in namespace `netbird`; the server stores SQLite data in `netbird-pvc`. The
configuration renderer and template are in `k8s/`. Prepare
`k8s/secrets.yaml` from `k8s/secrets.yaml.example` before the first deploy.
## Compose alternative
`compose.yaml` expects `entrypoint.sh`, a local `.env`, and two local files:
`secrets/relay-auth-secret` and `secrets/datastore-encryption-key`.
The reviewed main commit is missing the renderer and setup script.
`fix/netbird-compose-runtime` restores them. Merge that fix before following
these setup commands:
The Compose files are available for manual use. There is no root `active` marker,
so the automatic deploy workflow selects Kubernetes only.
```sh
cd netbird
## Files
- `compose.yaml`: dashboard and combined server; start it manually when using Compose.
- `config.template.yaml`: non-secret server configuration rendered at startup.
- `entrypoint.sh`: injects Docker secrets into an in-memory runtime configuration.
- `client.compose.yaml`: optional host-network peer using a dashboard-generated setup key.
- `.env`: ignored local hostnames, the detected Traefik Docker-network subnet, and optional client setup key.
- `secrets/`: ignored relay secret and datastore encryption key.
## First deployment
Run these commands on the Docker host before the first Compose start.
```bash
cd /srv/homelab/netbird
./setup.sh
$EDITOR .env
docker compose config --quiet
docker compose up -d
```
The setup script detects the IPv4 subnet of the external `proxy` network and
preserves existing secrets. Complete the initial owner setup through the public
TLS endpoint after starting the server.
Review the values in `.env` before starting. The example public hostname is `netbird.forust.xyz`; change it if a different public domain was selected. `setup.sh` replaces `NETBIRD_PROXY_SUBNET=auto` with the first IPv4 subnet of the external Docker `proxy` network. Keep that value synchronized with the network; set an explicit CIDR instead if the network is managed elsewhere.
`setup.sh` is idempotent and never replaces existing secrets. Do not delete or regenerate `secrets/datastore-encryption-key` after the first successful start unless all encrypted setup keys and API tokens are intentionally being invalidated.
## Network prerequisites
- Point the public hostname directly to the Docker host. Do not proxy UDP `3478` through Cloudflare or another CDN.
- Allow inbound TCP `80`, TCP `443`, and UDP `3478` through the host firewall and upstream router.
- Ensure the external `proxy` Docker network exists and Traefik uses its `websecure` entrypoint and `letsencrypt` resolver. `NETBIRD_PROXY_SUBNET` must describe that network; it is used to trust only forwarded client addresses from Traefik.
- Ensure the internal names in `.env` resolve where the local and development aliases are needed.
- Keep Traefik's `websecure` read timeout disabled for long-lived gRPC and WebSocket sessions. This repository configures `--entrypoints.websecure.transport.respondingTimeouts.readTimeout=0` in `traefik/compose.yaml`.
After startup, verify OIDC discovery through the public TLS endpoint:
```bash
curl -fsS "https://${NETBIRD_DOMAIN}/oauth2/.well-known/openid-configuration"
```
Open `https://${NETBIRD_DOMAIN}` immediately and complete the initial owner setup. Treat the initial setup flow as public until the owner exists.
## Optional host client
`client.compose.yaml` runs a host-network peer in a separate Compose project.
Set `NB_SETUP_KEY` and `NETBIRD_CLIENT_HOSTNAME` in the local `.env`, then run
`docker compose -f client.compose.yaml up -d`. It needs `/dev/net/tun` and elevated
network capabilities. The normal server deployment does not start this client.
The client intentionally lives in a separate Compose project. Normal server deploys use `--remove-orphans`, so keeping the client in the server project would cause it to be removed.
## Backup
1. Create a reusable or ephemeral setup key in the NetBird dashboard.
2. Put `NB_SETUP_KEY=<key>` in the ignored `netbird/.env` file.
3. Set `NETBIRD_CLIENT_HOSTNAME` to this machine's desired peer name.
4. Start and inspect the client:
Back up the SQLite data while the server is stopped, together with the encryption
key, relay secret, and local configuration. Test a restore on an isolated host.
For Compose, the datastore volume has the explicit name `netbird_data`.
Do not use `docker compose down -v` when keeping the installation.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n netbird
kubectl get events -n netbird --sort-by=.metadata.creationTimestamp
```bash
cd /srv/homelab/netbird
docker compose -f client.compose.yaml config --quiet
docker compose -f client.compose.yaml up -d
docker compose -f client.compose.yaml exec netbird-client netbird status
```
See the [repository README](../README.md) for deployment selection.
The client uses host networking and requires `NET_ADMIN`, `SYS_ADMIN`, `SYS_RESOURCE`, and `/dev/net/tun`. Remove it without affecting the server stack:
```bash
docker compose -f client.compose.yaml down
```
## Operations
Inspect status and logs:
```bash
docker compose ps
docker compose logs --tail=200 netbird-server dashboard
```
Stop or remove containers without deleting data:
```bash
docker compose down
```
Do not add `-v` to `docker compose down`; it would delete the NetBird datastore.
## Backup and restore
Back up both the persistent volume and the ignored secret files. For a consistent SQLite backup, briefly stop the server first and store the resulting archive and `datastore-encryption-key` in an encrypted backup:
```bash
cd /srv/homelab/netbird
mkdir -p backups
docker compose stop netbird-server
docker run --rm \
-v netbird_data:/data:ro \
-v "$PWD/backups:/backup" \
busybox:1.37.0 \
tar -C /data -czf "/backup/netbird-data-$(date -u +%Y%m%dT%H%M%SZ).tar.gz" .
docker compose start netbird-server
```
Also securely back up:
- `secrets/datastore-encryption-key` — required to decrypt stored secrets.
- `secrets/relay-auth-secret` — keeps issued relay credentials valid across restoration.
- `netbird/.env` — optional, but it records the public and internal hostnames.
Test a restore in an isolated Docker host before relying on a backup.
## Upgrade
1. Take and verify a backup.
2. Review NetBird release notes for server, client, and dashboard compatibility.
3. Update the pinned tags in `compose.yaml`; update `client.compose.yaml` separately when deploying the client.
4. Pull and recreate the selected services:
```bash
docker compose pull
docker compose up -d
```
The image tags are intentionally pinned instead of using `latest`, matching this repository's pull-on-deploy policy.
+109
View File
@@ -0,0 +1,109 @@
#!/bin/sh
set -eu
umask 077
TEMPLATE_PATH=/opt/netbird/config.template.yaml
RENDERED_PATH=/run/netbird/config.yaml
RELAY_SECRET_PATH=/run/secrets/relay_auth_secret
ENCRYPTION_KEY_PATH=/run/secrets/datastore_encryption_key
is_valid_proxy_subnet() {
candidate="$1"
case "$candidate" in
0.0.0.0/0)
return 1
;;
*/*)
address="${candidate%%/*}"
prefix="${candidate#*/}"
;;
*)
return 1
;;
esac
case "$prefix" in
0|[1-9]|[1-2][0-9]|3[0-2]) ;;
*)
return 1
;;
esac
old_ifs="$IFS"
IFS=.
# shellcheck disable=SC2086
set -- $address
IFS="$old_ifs"
[ "$#" -eq 4 ] || return 1
for octet do
case "$octet" in
0|[1-9]|[1-9][0-9]|1[0-9][0-9]|2[0-4][0-9]|25[0-5]) ;;
*)
return 1
;;
esac
done
}
read_secret() {
secret_path="$1"
if [ ! -r "$secret_path" ]; then
echo "Required secret is not readable: $secret_path" >&2
exit 1
fi
secret_value="$(cat "$secret_path")"
if [ -z "$secret_value" ]; then
echo "Required secret is empty: $secret_path" >&2
exit 1
fi
printf '%s' "$secret_value"
}
if [ -z "${NETBIRD_DOMAIN:-}" ]; then
echo "NETBIRD_DOMAIN must be set" >&2
exit 1
fi
case "$NETBIRD_DOMAIN" in
*[!A-Za-z0-9.-]*)
echo "NETBIRD_DOMAIN contains unsupported characters" >&2
exit 1
;;
esac
if [ -z "${NETBIRD_PROXY_SUBNET:-}" ] || [ "$NETBIRD_PROXY_SUBNET" = "auto" ]; then
echo "NETBIRD_PROXY_SUBNET must be an explicit IPv4 CIDR; run netbird/setup.sh first" >&2
exit 1
fi
if ! is_valid_proxy_subnet "$NETBIRD_PROXY_SUBNET"; then
echo "NETBIRD_PROXY_SUBNET must be a non-default IPv4 CIDR, for example 172.20.0.0/16" >&2
exit 1
fi
if [ "$#" -ne 2 ] || [ "$1" != "--config" ] || [ "$2" != "$RENDERED_PATH" ]; then
echo "Expected: --config $RENDERED_PATH" >&2
exit 1
fi
relay_secret="$(read_secret "$RELAY_SECRET_PATH")"
encryption_key="$(read_secret "$ENCRYPTION_KEY_PATH")"
mkdir -p "$(dirname "$RENDERED_PATH")"
sed \
-e "s|__NETBIRD_DOMAIN__|${NETBIRD_DOMAIN}|g" \
-e "s|__NETBIRD_AUTH_SECRET__|${relay_secret}|g" \
-e "s|__NETBIRD_ENCRYPTION_KEY__|${encryption_key}|g" \
-e "s|__NETBIRD_PROXY_SUBNET__|${NETBIRD_PROXY_SUBNET}|g" \
"$TEMPLATE_PATH" >"$RENDERED_PATH"
if grep -q '__NETBIRD_' "$RENDERED_PATH"; then
echo "Rendered NetBird configuration still contains unresolved placeholders" >&2
exit 1
fi
exec /go/bin/netbird-server "$@"
+13
View File
@@ -3,6 +3,8 @@ kind: Service
metadata:
name: netbird-server-service
namespace: netbird
labels:
app: netbird-server
spec:
selector:
app: netbird-server
@@ -11,6 +13,10 @@ spec:
name: http
targetPort: 80
protocol: TCP
- port: 9090
name: metrics
targetPort: metrics
protocol: TCP
- port: 3478
name: stun
targetPort: 3478
@@ -32,6 +38,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netbird-server-deployment
namespace: netbird
spec:
@@ -57,6 +65,9 @@ spec:
- containerPort: 80
name: http
protocol: TCP
- containerPort: 9090
name: metrics
protocol: TCP
- containerPort: 3478
name: stun
protocol: UDP
@@ -126,6 +137,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netbird-dashboard-deployment
namespace: netbird
spec:
@@ -1,14 +1,14 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: webinar-checker
namespace: edu-master
name: netbird-server
namespace: netbird
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: edu-master-webinar-checker
app: netbird-server
endpoints:
- port: metrics
path: /metrics
+38
View File
@@ -0,0 +1,38 @@
#!/usr/bin/env bash
# Prepare local Compose configuration without replacing existing credentials.
set -euo pipefail
cd "$(dirname "${BASH_SOURCE[0]}")"
umask 077
if [ ! -f .env ]; then
cp .env.example .env
fi
if grep -q '^NETBIRD_PROXY_SUBNET=auto$' .env; then
subnet="$(docker network inspect proxy --format '{{range .IPAM.Config}}{{println .Subnet}}{{end}}' | awk '/^[0-9]+\./ { print; exit }')"
if [ -z "$subnet" ]; then
echo "No IPv4 subnet found on the Docker proxy network. Set NETBIRD_PROXY_SUBNET in .env." >&2
exit 1
fi
# The detected value must be safe to substitute into the env file.
if [[ ! "$subnet" =~ ^[0-9.]+/[0-9]+$ ]]; then
echo "Unexpected Docker network subnet: $subnet" >&2
exit 1
fi
sed -i "s|^NETBIRD_PROXY_SUBNET=auto$|NETBIRD_PROXY_SUBNET=$subnet|" .env
fi
mkdir -p secrets
chmod 700 secrets
for name in relay-auth-secret datastore-encryption-key; do
path="secrets/$name"
if [ -e "$path" ]; then
if [ ! -s "$path" ]; then
echo "Existing secret is empty: $path. Restore it before continuing." >&2
exit 1
fi
else
openssl rand -base64 32 >"$path"
fi
chmod 600 "$path"
done
printf '%s\n' 'Local files are ready. Review .env, then run docker compose config --quiet.'
+4
View File
@@ -14,6 +14,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netbox-deployment
namespace: netbox
labels:
@@ -118,6 +120,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netbox-worker-deployment
namespace: netbox
labels:
+2
View File
@@ -17,6 +17,8 @@ spec:
apiVersion: apps/v1
kind: StatefulSet
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netbox-valkey
namespace: netbox
labels:
+1 -1
View File
@@ -1,6 +1,6 @@
services:
netronome:
image: ghcr.io/autobrr/netronome:v0.15.0
image: ghcr.io/autobrr/netronome:v0.16.0
restart: unless-stopped
container_name: netronome
ports:
+3 -1
View File
@@ -14,6 +14,8 @@ spec:
apiVersion: apps/v1
kind: Deployment
metadata:
annotations:
reloader.stakater.com/auto: "true"
name: netronome-deployment
namespace: netronome
labels:
@@ -32,7 +34,7 @@ spec:
spec:
containers:
- name: netronome
image: ghcr.io/autobrr/netronome:v0.15.0
image: ghcr.io/autobrr/netronome:v0.16.0
ports:
- name: netronome-port
protocol: TCP
+1 -1
View File
@@ -1,6 +1,6 @@
services:
portainer:
image: portainer/portainer-ce:2.45.1
image: portainer/portainer-ce:2.45.2
container_name: portainer
restart: always
volumes:
+1 -1
View File
@@ -29,7 +29,7 @@ spec:
spec:
containers:
- name: portainer
image: portainer/portainer-ce:2.45.1
image: portainer/portainer-ce:2.45.2
ports:
- containerPort: 9000
volumeMounts:
+1
View File
@@ -1,6 +1,7 @@
POSTGRES_ADMIN_PASSWORD=
AUTHENTIK_DB_PASSWORD=
GITEA_DB_PASSWORD=
NETBOX_DB_PASSWORD=
NETRONOME_DB_PASSWORD=
PENPOT_DB_PASSWORD=
STATUSPAGE_DB_PASSWORD=
+2 -2
View File
@@ -34,8 +34,8 @@ docker compose -f shared-compose.yaml config --quiet
docker compose -f shared-compose.yaml up -d
```
Add `NETBOX_DB_PASSWORD` to `.env` as well: the reviewed env example omits it;
`fix/postgres-env-example` restores the key. Fill every required password.
The example includes `NETBOX_DB_PASSWORD`; fill it and every other required
password before starting the stack.
This stack creates the `homelab-database` Docker network and the
`homelab-postgres` container. Compose applications need to join that network
explicitly to use it; several committed Compose stacks use their own databases.
+24 -18
View File
@@ -1,26 +1,32 @@
# Monitoring stack
Prometheus, Grafana, Alertmanager, and application alert rules.
The Kubernetes stack provides Grafana, Alertmanager, VictoriaMetrics, VMAgent,
and vmalert. The `k8s/active` marker selects the stack. The
`kube-prometheus-stack` Helm release installs Grafana, Alertmanager, the
Prometheus Operator, and related components. Its Prometheus server is configured
with zero replicas while VMAgent collects metrics and writes them to the
single-node VictoriaMetrics instance.
Kubernetes installs `kube-prometheus-stack` in `prometheus` through the deploy
library. Its chart version is pinned there; `k8s/grafana-values.yaml` contains the
values for the whole stack, despite the filename.
The `victoria-operator` Helm release converts selected Prometheus Operator
`ServiceMonitor` resources into `VMServiceScrape` resources. VMAgent selects
those scrapes across namespaces and writes to VictoriaMetrics. vmalert evaluates
the rule ConfigMap and sends alerts to the stack's Alertmanager. See the
[Kubernetes monitoring notes](k8s/README.md) for application metrics and
validation commands.
The values file is tracked. `.gitignore` also lists it, but that does not stop Git
tracking later edits. Keep local credentials in Secrets rather than treating
changes to this file as ignored.
The chart versions are pinned in `.gitea/workflows/deploy-lib.sh`. The tracked
`k8s/grafana-values.yaml` contains the Helm values for the stack. Create the
`grafana-admin` and `alertmanager-config` Secrets from the examples in `k8s/`;
keep their credentials out of the values file. Persistent volumes store data for
Prometheus, Grafana, Alertmanager, and VictoriaMetrics. Check the PVCs and
backups before changing storage. VictoriaMetrics currently retains 30 days of
data.
Prepare Grafana admin and Alertmanager Secrets separately. The Alertmanager
configuration example and Telegram template are in `k8s/`; the main deploy
selection excludes the example config. Certificates and ingress expose Grafana,
and the rule files add service-specific alerts.
A separate Compose configuration is present for manual use. There is no root
`active` marker, so the automatic deploy workflow does not select it.
The values use `local-path` PVCs for Grafana, Prometheus, and Alertmanager.
Retention is limited by both time and size. Back up Grafana state and any history
that must survive a storage failure.
The Compose stack has separate Prometheus and Alertmanager configuration files.
There is no root active marker. This README describes committed main files;
local VictoriaMetrics experiments are not part of that configuration.
The deploy workflow does not remove resources when manifests are deleted. For a
rollback of application-metrics changes, follow the explicit cleanup steps in
the [Kubernetes monitoring notes](k8s/README.md).
See the [repository README](../README.md) for deployment selection.
+43
View File
@@ -0,0 +1,43 @@
# VictoriaMetrics
The `victoria-operator` Helm release converts Prometheus Operator
`ServiceMonitor` resources into owned `VMServiceScrape` resources. The
`VMAgent` selects converted scrapes labeled `release: prometheus-stack` in all
namespaces and writes them to the existing single-node VictoriaMetrics
instance. Changes to selected `ServiceMonitor` resources are reconciled
automatically; there is no copied Prometheus scrape-config blob to regenerate.
The agent drops targets for the Prometheus server service to avoid duplicating
its self-scrape. `scraper: victoria` identifies the samples ingested by this
VMAgent.
The VictoriaMetrics Operator chart and its CRDs are installed before the
Kubernetes manifests by the normal deploy workflow. On a cluster where the
operator CRDs are not installed yet, CI skips the server-side dry-run of the
`VMAgent` resource; the deploy installs the chart before applying that resource.
## Application metrics
The application ServiceMonitors use a 30s interval and a 10s timeout:
- Headscale: the external Service points to the Compose host on port 19090.
A VMServiceScrape uses EndpointSlice discovery for this manually managed target.
The Compose configuration must bind metrics to `0.0.0.0:9090`.
- NetBird: the combined server exports `/metrics` on port 9090. The existing
`server.metricsPort` setting enables the listener.
- Gitea: `GITEA__metrics__ENABLED` enables `/metrics` on the HTTP port. The public
ingress excludes this path. The monitor uses the internal Service directly.
- Immich: `IMMICH_TELEMETRY_INCLUDE=all` enables API and worker metrics on ports
8081 and 8082. The monitor scrapes both ports on each server replica.
Deploy through the existing CI and deploy workflow. Gitea and Immich reload their
ConfigMap changes through Reloader. Check the VMAgent targets after deployment
and query `up{scraper="victoria",namespace=~"netbird|gitea|immich|headscale"}` in
VictoriaMetrics. All targets should report 1.
For rollback, revert the application metrics changes, run CI, and deploy the
revert. Remove the three application ServiceMonitors and the Headscale VMServiceScrape explicitly: the deployment
workflow applies manifests and does not prune removed resources.
For Headscale rollback, remove its VMServiceScrape and Service label, restore the
previous Compose metrics bind address, and restart only the Headscale service.
+11
View File
@@ -38,6 +38,8 @@ grafana:
# One block covers both the dashboards and datasources sidecars (p95 91M / 80M).
sidecar:
datasources:
defaultDatasourceEnabled: false
resources:
requests:
memory: "96Mi"
@@ -50,9 +52,18 @@ grafana:
type: loki
url: http://loki-gateway.prometheus.svc.cluster.local
access: proxy
- name: VictoriaMetrics
type: prometheus
url: http://victoria-metrics.prometheus.svc.cluster.local:8428
access: proxy
isDefault: true
prometheus:
prometheusSpec:
# VM trial: vmagent scrapes and remote-writes to VictoriaMetrics, so the
# Prometheus server itself stands down. Encoded here (not a kubectl patch)
# so helm keeps owning spec.replicas and upgrades do not conflict on it.
replicas: 0
retention: 60d
retentionSize: 32GB
storageSpec:
+102
View File
@@ -31,3 +31,105 @@ spec:
port: 80
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: prometheus-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`prom.workstation.internal`) || Host(`prom.gigaforust.internal`)
kind: Rule
services:
- name: prometheus-stack-kube-prom-prometheus
port: 9090
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: alertmanager-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`am.workstation.internal`) || Host(`am.gigaforust.internal`)
kind: Rule
services:
- name: prometheus-stack-kube-prom-alertmanager
port: 9093
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: loki-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`loki.workstation.internal`) || Host(`loki.gigaforust.internal`)
kind: Rule
services:
- name: loki-gateway
port: 80
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: alloy-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`alloy.workstation.internal`) || Host(`alloy.gigaforust.internal`)
kind: Rule
services:
- name: alloy
port: 12345
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: victoria-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`victoria.workstation.internal`) || Host(`victoria.gigaforust.internal`)
kind: Rule
services:
- name: victoria-metrics
port: 8428
tls:
secretName: internal-wildcard-tls
---
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: vmalert-local
namespace: prometheus
spec:
entryPoints:
- websecure
routes:
- match: Host(`vmalert.workstation.internal`) || Host(`vmalert.gigaforust.internal`)
kind: Rule
services:
- name: vmalert
port: 8880
tls:
secretName: internal-wildcard-tls
@@ -0,0 +1,12 @@
nameOverride: victoria-operator
operator:
enable_converter_ownership: true
resources:
requests:
cpu: 50m
memory: 96Mi
limits:
cpu: 200m
memory: 256Mi
+79
View File
@@ -0,0 +1,79 @@
apiVersion: v1
kind: Service
metadata:
name: victoria-metrics
namespace: prometheus
spec:
selector:
app: victoria-metrics
ports:
- port: 8428
targetPort: 8428
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: victoria-pvc
namespace: prometheus
spec:
resources:
requests:
storage: 10Gi
volumeMode: Filesystem
accessModes:
- ReadWriteOnce
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: victoria-deployment
namespace: prometheus
spec:
replicas: 1
selector:
matchLabels:
app: victoria-metrics
strategy:
type: Recreate
template:
metadata:
labels:
app: victoria-metrics
spec:
containers:
- name: victoria
image: victoriametrics/victoria-metrics:v1.153.0-scratch
args:
- -storageDataPath=/vmdata
- -retentionPeriod=30d
- -httpListenAddr=:8428
ports:
- containerPort: 8428
readinessProbe:
httpGet:
path: /health
port: 8428
initialDelaySeconds: 15
periodSeconds: 10
failureThreshold: 6
livenessProbe:
httpGet:
path: /health
port: 8428
initialDelaySeconds: 60
periodSeconds: 30
failureThreshold: 3
volumeMounts:
- name: vmdata
mountPath: /vmdata
resources:
requests:
cpu: "100m"
memory: "256Mi"
limits:
cpu: "1000m"
memory: "1Gi"
volumes:
- name: vmdata
persistentVolumeClaim:
claimName: victoria-pvc
+29
View File
@@ -0,0 +1,29 @@
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMAgent
metadata:
name: vmagent
namespace: prometheus
spec:
image:
tag: v1.153.0
scrapeInterval: 30s
externalLabels:
scraper: victoria
serviceScrapeNamespaceSelector: {}
serviceScrapeSelector:
matchLabels:
release: prometheus-stack
globalScrapeRelabelConfigs:
- action: drop
source_labels:
- __meta_kubernetes_service_name
regex: prometheus-stack-kube-prom-prometheus
remoteWrite:
- url: http://victoria-metrics.prometheus.svc.cluster.local:8428/api/v1/write
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: "1000m"
memory: 1Gi
+70
View File
@@ -0,0 +1,70 @@
apiVersion: v1
kind: Service
metadata:
name: vmalert
namespace: prometheus
spec:
selector:
app: vmalert
ports:
- port: 8880
targetPort: 8880
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: vmalert-deployment
namespace: prometheus
spec:
replicas: 1
selector:
matchLabels:
app: vmalert
strategy:
type: Recreate
template:
metadata:
labels:
app: vmalert
spec:
containers:
- name: vmalert
image: victoriametrics/vmalert:v1.153.0
args:
- -datasource.url=http://victoria-metrics.prometheus.svc.cluster.local:8428
- -remoteWrite.url=http://victoria-metrics.prometheus.svc.cluster.local:8428
- -notifier.url=http://prometheus-stack-kube-prom-alertmanager.prometheus.svc.cluster.local:9093
- -rule=/etc/vm/rules/*.yaml
- -evaluationInterval=60s
- -httpListenAddr=:8880
ports:
- containerPort: 8880
readinessProbe:
httpGet:
path: /metrics
port: 8880
initialDelaySeconds: 15
periodSeconds: 10
failureThreshold: 6
livenessProbe:
httpGet:
path: /metrics
port: 8880
initialDelaySeconds: 60
periodSeconds: 30
failureThreshold: 3
volumeMounts:
- name: rules
mountPath: /etc/vm/rules
readOnly: true
resources:
requests:
cpu: "50m"
memory: "64Mi"
limits:
cpu: "200m"
memory: "256Mi"
volumes:
- name: rules
configMap:
name: prometheus-prometheus-stack-kube-prom-prometheus-rulefiles-0
+2 -2
View File
@@ -19,8 +19,8 @@ Reloader discovers references in environment variables and mounted volumes.
This covers startup-only settings and ConfigMaps or Secrets mounted with `subPath`.
See the [upstream usage guide](https://github.com/stakater/Reloader/blob/v1.4.22/README.md#usage).
The application manifests opt in 28 workloads, including AdGuard's TLS files,
NetBird, both NetBox processes, EDU bots, and the password-protected Valkey servers.
The application manifests opt in workloads including AdGuard's TLS files,
NetBird, both NetBox processes, and the password-protected Valkey servers.
Inactive services have the same annotations ready for later activation.
## Controller policy
Loaded 100 of 126 files, more files were not shown because too many files have changed in this diff. Show more