Compare commits
2
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
5c8bc15e60 | ||
|
|
3c4732ae20 |
No files matched your search
@@ -1,50 +0,0 @@
|
||||
# EDU ownership handoff
|
||||
|
||||
## Status
|
||||
|
||||
The EDU ownership handoff is complete. The homelab repository no longer owns
|
||||
EDU workloads, images, routes, alerts, or deployment selection. The EDU
|
||||
repository is the only deployment owner: [forust/edu-master](https://git.forust.xyz/forust/edu-master).
|
||||
|
||||
Homelab PRs #99 and #105 are merged. PR #105 removed the EDU subtree and its
|
||||
build, deploy, rollback, verification, route-probe, and registry references.
|
||||
It also added the serial image build matrix for the homelab services. This
|
||||
handoff record is the only remaining EDU-specific file in homelab Git.
|
||||
|
||||
The dedicated workstation checkout is `/srv/edu-master`, at release
|
||||
`4f2b2a0e37dc11ac2c75441a15076c178e219d37`. It contains `k8s/active`; root
|
||||
`active` is absent. The old untracked `/srv/homelab/edu_master` checkout was
|
||||
moved outside the homelab repository to
|
||||
`/srv/edu-master-legacy-archive-20261007/edu_master`. Its private files remain
|
||||
mode `0600` inside an archive directory with mode `0700`. The homelab deploy
|
||||
checkout has no EDU marker or tracked EDU application/deployment files.
|
||||
`AUTODEPLOY=false` remains in place for homelab deployment.
|
||||
|
||||
## Release evidence
|
||||
|
||||
EDU PR #4 merged after its review and CI checks. Main-push CI run 1652 passed
|
||||
all validation and both image builds. Deploy run 1653 passed for the exact main
|
||||
SHA above.
|
||||
|
||||
The workstation rollout completed for both Deployments. The deployment
|
||||
verified `/health` and `/live` with HTTP 200, Redis AUTH, session TTL of 1058
|
||||
seconds, a delivery backlog of zero, and all nine EDU vmalert rules with
|
||||
matching expressions and healthy evaluation.
|
||||
|
||||
The images now run by digest:
|
||||
|
||||
- Session keeper: `sha256:998dea51aa3015fd9cabefb0f53b030157a650c3bef72e02fe84f17d5762613d`
|
||||
- Webinar checker: `sha256:92f3c1fa2bb7f9b4680a9fc76a5b33dfbea8ef3dd9c6490ebc45876fd4c54461`
|
||||
|
||||
Redis StatefulSet was unchanged. PVC `redis-data-pvc` remains bound to PV
|
||||
`pvc-a4f2a79a-363a-4c12-ae91-92cdfc2a0d2e` with capacity 1 GiB. The existing
|
||||
runtime Secret and Fernet key were preserved during the handoff. Notification
|
||||
delivery was verified before closeout, as confirmed by the operator. The
|
||||
deployment did not record downtime.
|
||||
|
||||
The release rollback snapshot is
|
||||
`/home/forust/.local/state/edu-master-deploy/20261007T180541Z-4f2b2a0e37dc11ac2c75441a15076c178e219d37`.
|
||||
The handoff data snapshot remains at
|
||||
`/home/forust/.local/state/edu-master-deploy/handoff-20261007T080838Z`.
|
||||
Both snapshots are outside Git. Do not restore old Redis data unless recovery
|
||||
requires it. Never delete or recreate the Redis PVC.
|
||||
@@ -0,0 +1,88 @@
|
||||
# Build and deployment workflows
|
||||
|
||||
Gitea Actions checks this repository, builds its custom images, and deploys
|
||||
selected services to the workstation. Workflows use the self-hosted runner labels
|
||||
`linux`, `arch`, and `homelab`; deployment jobs also require `prod`.
|
||||
|
||||
## Checks
|
||||
|
||||
`ci.yaml` runs Compose validation, actionlint, ShellCheck, Prettier, Ruff,
|
||||
yamllint, hadolint, and kubeconform. Tool versions are pinned in
|
||||
`workflows/tool-versions.env` and installed by `install-ci-tools.sh`.
|
||||
|
||||
Compose CI checks structure without resolving local environment files or paths.
|
||||
On the reviewed main commit it only discovers standard filenames; the
|
||||
`fix/deploy-validation` branch adds the manual Compose entry points too.
|
||||
|
||||
Kubeconform validates known resource schemas. Unknown CRDs are skipped. On main,
|
||||
CI also attempts server-side dry-runs for marked services; these require an
|
||||
existing namespace and contact the cluster's admission webhooks. A cluster that
|
||||
is unreachable produces a warning and skips that CI pass. Deploy validation has
|
||||
its own dry-run stage.
|
||||
|
||||
`renovate-ci.yaml` validates Renovate settings and checks that its generated
|
||||
ConfigMap matches `renovate/renovate.json`.
|
||||
|
||||
## Image builds
|
||||
|
||||
CI builds changed custom images for `errorpages`, both `homepages` variants, and
|
||||
the two `edu_master` Python services. Main builds publish `main`, `prod`, and a
|
||||
commit tag. Dev builds publish `dev`. Build jobs wait for the lint and manifest
|
||||
checks.
|
||||
|
||||
Kubernetes deployment resolves the lab's own registry images to digests, preferring
|
||||
commit-specific tags. Third-party image versions remain declared in the manifests.
|
||||
|
||||
## Deploy selection
|
||||
|
||||
`workflows/deploy-lib.sh` owns the stage logic; `ssh-run.sh` invokes it on the
|
||||
workstation through SSH. Kubernetes selection uses `k8s/active`; Compose selection
|
||||
uses an `active` file beside a standard `compose.yaml` or `compose.yml`.
|
||||
Kustomize overlays are supported, although the current tree primarily contains
|
||||
plain manifests.
|
||||
|
||||
Secret files, examples, Helm values, and patch files are excluded from plain
|
||||
manifest selection. Create local Kubernetes Secrets separately in their target
|
||||
namespaces. The Helm table lists Prometheus, Loki, Alloy, and Reloader, with each
|
||||
release controlled by its configured marker. Other charts need separate setup.
|
||||
|
||||
## Trigger and required settings
|
||||
|
||||
Automatic deployment follows a successful main CI run when the repository Actions
|
||||
variable `AUTODEPLOY` is `true`. The manual deploy workflow bypasses that switch
|
||||
and targets the fetched main branch when no validated commit SHA is provided.
|
||||
A manual dispatch does not prove that this commit passed CI.
|
||||
|
||||
Configure the Actions secrets `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_SSH_KEY`, and,
|
||||
where needed, `DEPLOY_PORT` and `DEPLOY_PATH`. Registry publishing uses
|
||||
`REGISTRY_USERNAME` and `REGISTRY_PASSWORD`. The remote user needs access to Git,
|
||||
Docker, kubectl, Helm, jq, and the state directory used for snapshots.
|
||||
|
||||
Keep `APPLY_PRUNE` false on the reviewed implementation: its per-file prune loop
|
||||
is unsafe. `fix/deploy-prune-guard` rejects that option before changes are applied.
|
||||
|
||||
Preflight fetches and resets the remote checkout. It refuses when tracked files
|
||||
have local changes; ignored local env and Secret files stay in place. Do not use
|
||||
a development checkout with uncommitted tracked changes as the deployment target.
|
||||
|
||||
## Stages and recovery
|
||||
|
||||
1. Preflight fetches the target commit and checks the remote working tree.
|
||||
2. Validate selects services, parses Compose, performs Kubernetes dry-runs, and
|
||||
checks referenced Secrets.
|
||||
3. Apply Kubernetes records a workload snapshot, upgrades selected Helm releases,
|
||||
applies resources, and refreshes owned custom images.
|
||||
4. Apply Compose recreates marked stacks and checks container state.
|
||||
5. Verify Kubernetes checks changed workloads and attempts rollback for failures.
|
||||
6. Smoke probes public routes after verification.
|
||||
|
||||
The two apply jobs share a remote lock. Workflow concurrency queues deployments
|
||||
rather than interrupting an older apply. Snapshots live under
|
||||
`$XDG_STATE_HOME/homelab-deploy`, or `~/.local/state/homelab-deploy` by default.
|
||||
They contain the pre-apply workload data and commit identifier.
|
||||
|
||||
Rollback uses workload revisions. It does not restore ConfigMaps, Secrets,
|
||||
database schemas, or data. Helm-owned workloads are handled through the Helm
|
||||
upgrade's rollback path; the generic rollback skips them. Compose has no automatic
|
||||
rollback. See the [review](../docs/repository-review.md) for remaining recovery
|
||||
limitations, including SSH retries and serial rollback timing.
|
||||
@@ -7,5 +7,4 @@ self-hosted-runner:
|
||||
labels:
|
||||
- arch
|
||||
- homelab
|
||||
- homelab-pr
|
||||
- prod
|
||||
@@ -1,3 +0,0 @@
|
||||
{
|
||||
"postgres": ["authentik", "gitea", "immich", "n8n", "netbox", "netronome"]
|
||||
}
|
||||
@@ -1,159 +0,0 @@
|
||||
# Homelab CI/CD
|
||||
|
||||
The native Gitea runners run on **vps**; production runs on **workstation**.
|
||||
Main-branch checks and image builds use `homelab:host`. Pull request and
|
||||
non-main checks use `homelab-pr:host` under a separate account without Docker
|
||||
access. The `homelab-pr` runner is registered at User scope for `forust`, so
|
||||
any repository under that account can schedule jobs that request this label.
|
||||
Each runner accepts one job at a time; the build waits for every check to pass.
|
||||
CI and deploy runs also show a summary with
|
||||
the release SHA, image build or reuse results, deploy mode, selected services,
|
||||
and image digests. Failed runs keep a summary of completed image builds, stage
|
||||
results, apply results, and recorded Kubernetes recovery. The final deploy
|
||||
summary is in the smoke job; earlier jobs show the state observed at that time.
|
||||
Apply success is separate from health and recovery. Update the installed
|
||||
workstation controller with `setup-workstation.sh` when no deploy is running.
|
||||
No job images or Kubernetes credentials are needed on the VPS. Builds use one
|
||||
pinned BuildKit helper container. CI and deploy are separate workflows.
|
||||
|
||||
## Runner installation
|
||||
|
||||
Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl,
|
||||
GNU tar/xz, flock and systemd using the host's package manager. Keep the existing
|
||||
Gitea runner 3.0.2 binary at `/usr/local/bin/gitea-runner`.
|
||||
|
||||
From this checkout on the VPS:
|
||||
|
||||
```sh
|
||||
sudo bash .gitea/runner/setup-runner.sh
|
||||
```
|
||||
|
||||
The installer reuses `/var/lib/gitea-runner/.runner` and the existing service.
|
||||
For a new host, install the same runner binary and register as `gitea-runner`
|
||||
using the registration token interactively, label `homelab:host`, and working
|
||||
directory `/var/lib/gitea-runner`; then rerun the installer. Tokens never belong
|
||||
in this repository or command-line examples.
|
||||
|
||||
Pinned tools live in the runner user's `~/.cache/homelab-ci`; CI repairs version
|
||||
drift there. Installations are locked. Buildx uses only the `homelab-ci` builder,
|
||||
pushes directly to the registry, and caps retained local cache at 1 GiB with a
|
||||
2 GiB free-space target. This is not a hard limit on peak build disk usage.
|
||||
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
|
||||
|
||||
### Pull request runner
|
||||
|
||||
Install the unprivileged host runner on the VPS:
|
||||
|
||||
```sh
|
||||
sudo bash .gitea/runner/setup-pr-runner.sh
|
||||
```
|
||||
|
||||
Get a registration token from the user Actions runner settings. Run the
|
||||
installer in a terminal. It asks for the token without echoing it, registers the
|
||||
runner as `homelab-pr` with label `homelab-pr:host`, then enables the service.
|
||||
The work directory is `/var/lib/gitea-pr-runner`. Confirm that Gitea lists the
|
||||
runner as User scope before merging the workflow change. An unmatched label can
|
||||
fall back to the default job image.
|
||||
|
||||
Renovate PR validation uses `pull_request_target`, which reads the workflow from
|
||||
the base branch. It checks out the PR head only after runner selection and runs
|
||||
that code on `homelab-pr`. Keep this workflow read-only and do not add secrets.
|
||||
|
||||
The PR runner has a separate home and tool cache. Do not add it to the `docker`
|
||||
group or give it access to `/var/run/docker.sock`. It runs repository code from
|
||||
pull requests, so keep its registration and permissions separate from the
|
||||
trusted `homelab` runner. This separates users and host permissions, but both
|
||||
runners still share the VPS kernel and network. Use a disposable VM if PRs from
|
||||
untrusted external authors must be fully isolated.
|
||||
|
||||
## Workstation setup
|
||||
|
||||
As the existing SSH deploy user on workstation:
|
||||
|
||||
```sh
|
||||
sudo loginctl enable-linger forust
|
||||
bash .gitea/runner/setup-workstation.sh
|
||||
```
|
||||
|
||||
The controller uses `/srv/homelab` as the persistent configuration tree and makes
|
||||
a detached source worktree for each SHA. It never resets `/srv/homelab`, moves
|
||||
local configuration, renames Compose projects, or changes volume names.
|
||||
The installer records the current Kubernetes context and cluster UID in
|
||||
`~/.config/homelab-deploy/environment`. Check these before installing.
|
||||
|
||||
Configure Gitea Actions Variables:
|
||||
|
||||
- `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`: the existing VPS-to-workstation SSH endpoint.
|
||||
- `DEPLOY_KNOWN_HOSTS`: workstation's verified SSH host key entry for that endpoint.
|
||||
- `AUTODEPLOY`: `false` initially; `true` enables deployment after successful main CI.
|
||||
|
||||
Keep `DEPLOY_SSH_KEY`, `REGISTRY_USERNAME` and `REGISTRY_PASSWORD` in Actions
|
||||
Secrets. Legacy endpoint secrets remain accepted during migration. The Actions
|
||||
token must have repository read and Actions read access for release downloads.
|
||||
The deploy user's existing Docker registry authentication remains necessary.
|
||||
|
||||
## Releases and deployment
|
||||
|
||||
CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image
|
||||
digests and build input fingerprints. Unchanged images are reused only from a
|
||||
successful main CI artifact, never from `:prod`. Expired artifacts cause CI to
|
||||
rebuild images; they block deployment until CI is rerun.
|
||||
|
||||
Run deploy from main with `deploy_ref=main` or a checked SHA:
|
||||
|
||||
- `full`: required for the first baseline; reconcile all active components.
|
||||
- `changed`: compare with the last fully successful production deploy.
|
||||
- `plan`: validate configuration and show selection without changing production resources.
|
||||
- `refresh_images=true`: explicitly refresh mutable third-party Compose tags.
|
||||
|
||||
The manual and automatic paths both require successful CI, a successful build
|
||||
job and the exact SHA's release artifact. PRs cannot publish images or deploy.
|
||||
Removed resources are reported and require explicit removal; no automatic prune.
|
||||
Service dependencies are listed in `.gitea/deploy-dependencies.json`.
|
||||
|
||||
A workstation user systemd service holds the deploy lock across validation,
|
||||
sequential apply, verification and smoke checks. SSH clients only submit/follow:
|
||||
disconnecting or cancelling the Actions client does not kill production apply.
|
||||
Retrying the same run ID does not start another apply. `ExecStopPost` recovers
|
||||
interrupted runs before the unit finishes. Kubernetes rolls back to captured
|
||||
revisions; configuration and persistent data are not reverted.
|
||||
|
||||
## Status and recovery
|
||||
|
||||
`--retry` repeats failed verification and smoke checks, never apply. Recovery
|
||||
keeps a failed deploy out of the successful baseline, even after rollback.
|
||||
|
||||
On workstation (replace the numeric ID with Actions run ID and attempt):
|
||||
|
||||
```sh
|
||||
python3 ~/.local/lib/homelab-deploy/controller.py status 123-1
|
||||
python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry
|
||||
journalctl --user -u homelab-deploy@123-1
|
||||
```
|
||||
|
||||
Runs live in `~/.local/state/homelab-deploy/runs`. Compose stores resolved configs
|
||||
with restricted permissions; these may contain credentials and must never be
|
||||
uploaded as CI artifacts. Stage logs print the exact manual recovery command
|
||||
using `compose-before/<stack>.json`, the original project directory and project
|
||||
name. Compose does not automatically roll back, and Nextcloud AIO's child
|
||||
containers remain managed by AIO. Preserve its own backups for data recovery.
|
||||
|
||||
The controller retains twenty successful/planned runs and preserves failures.
|
||||
Update the workstation dispatcher only when no deploy is running.
|
||||
|
||||
## Validation and migration rollback
|
||||
|
||||
```sh
|
||||
python3 -m unittest discover -s tests -v
|
||||
bash .gitea/tests/deploy-validation.sh
|
||||
```
|
||||
|
||||
Test on a separate namespace before the initial production `full` run. Check a
|
||||
failed rollout, interrupted SSH and repeated run ID, and verify that an isolated
|
||||
service change does not upgrade unrelated Helm releases or Compose stacks.
|
||||
|
||||
To roll back the migration, disable autodeploy and finish or recover the remote
|
||||
run first. Restore the runner config/unit from `.before-<timestamp>` backups,
|
||||
reload systemd and restart the runner. Restore the prior workflows from Git.
|
||||
Production data and persistent volumes stay where they were. Do not remove run
|
||||
state or Compose recovery files until recovery is confirmed.
|
||||
@@ -1,11 +0,0 @@
|
||||
[worker.oci]
|
||||
gc = true
|
||||
reservedSpace = "256MB"
|
||||
maxUsedSpace = "1GB"
|
||||
minFreeSpace = "2GB"
|
||||
|
||||
[[worker.oci.gcpolicy]]
|
||||
reservedSpace = "256MB"
|
||||
maxUsedSpace = "1GB"
|
||||
minFreeSpace = "2GB"
|
||||
all = true
|
||||
@@ -1,10 +0,0 @@
|
||||
runner:
|
||||
file: /var/lib/gitea-runner/.runner
|
||||
capacity: 1
|
||||
timeout: 5h
|
||||
labels:
|
||||
- homelab:host
|
||||
cache:
|
||||
enabled: false
|
||||
container:
|
||||
docker_host: unix:///var/run/docker.sock
|
||||
@@ -1,18 +0,0 @@
|
||||
[Unit]
|
||||
Description=Gitea Actions runner
|
||||
After=network-online.target docker.service
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
User=gitea-runner
|
||||
Group=gitea-runner
|
||||
SupplementaryGroups=docker
|
||||
WorkingDirectory=/var/lib/gitea-runner
|
||||
Environment=PATH=/var/lib/gitea-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
|
||||
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-runner/config.yaml
|
||||
Restart=on-failure
|
||||
RestartSec=5
|
||||
UMask=0077
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
@@ -1,12 +0,0 @@
|
||||
[Unit]
|
||||
Description=Homelab deploy %i
|
||||
|
||||
[Service]
|
||||
Type=exec
|
||||
EnvironmentFile=%h/.config/homelab-deploy/environment
|
||||
ExecStart=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py execute %i
|
||||
ExecStopPost=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py recover %i
|
||||
RuntimeMaxSec=5h
|
||||
TimeoutStopSec=135min
|
||||
KillMode=control-group
|
||||
UMask=0077
|
||||
@@ -1,8 +0,0 @@
|
||||
runner:
|
||||
file: /var/lib/gitea-pr-runner/.runner
|
||||
capacity: 1
|
||||
timeout: 5h
|
||||
labels:
|
||||
- homelab-pr:host
|
||||
cache:
|
||||
enabled: false
|
||||
@@ -1,27 +0,0 @@
|
||||
[Unit]
|
||||
Description=Gitea Actions untrusted pull request runner
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
User=gitea-pr-runner
|
||||
Group=gitea-pr-runner
|
||||
WorkingDirectory=/var/lib/gitea-pr-runner
|
||||
Environment=HOME=/var/lib/gitea-pr-runner
|
||||
Environment=PATH=/var/lib/gitea-pr-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
|
||||
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-pr-runner/config.yaml
|
||||
Restart=on-failure
|
||||
RestartSec=5
|
||||
NoNewPrivileges=yes
|
||||
PrivateTmp=yes
|
||||
ProtectSystem=full
|
||||
ProtectHome=yes
|
||||
ProtectKernelTunables=yes
|
||||
ProtectKernelModules=yes
|
||||
ProtectControlGroups=yes
|
||||
RestrictSUIDSGID=yes
|
||||
LockPersonality=yes
|
||||
UMask=0077
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
@@ -1,56 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Install a native runner for untrusted PR jobs without Docker access.
|
||||
set -euo pipefail
|
||||
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
|
||||
for tool in cp cut date getent id install runuser systemctl useradd; do
|
||||
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
|
||||
done
|
||||
command -v /usr/local/bin/gitea-runner >/dev/null || {
|
||||
echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
id gitea-pr-runner >/dev/null 2>&1 || \
|
||||
useradd --system --create-home --home-dir /var/lib/gitea-pr-runner --shell /usr/bin/bash gitea-pr-runner
|
||||
runner_home="$(getent passwd gitea-pr-runner | cut -d: -f6)"
|
||||
[ "$runner_home" = /var/lib/gitea-pr-runner ] || {
|
||||
echo 'Unexpected PR runner home; inspect the existing service first' >&2
|
||||
exit 1
|
||||
}
|
||||
case " $(id -nG gitea-pr-runner) " in
|
||||
*' docker '*)
|
||||
echo 'The PR runner account must not belong to the docker group' >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
|
||||
install -d -m 0755 /etc/gitea-pr-runner
|
||||
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||
for existing in /etc/gitea-pr-runner/config.yaml /etc/systemd/system/gitea-pr-runner.service; do
|
||||
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
|
||||
done
|
||||
install -m 0644 "$here/pr-config.yaml" /etc/gitea-pr-runner/config.yaml
|
||||
install -m 0644 "$here/pr-runner.service" /etc/systemd/system/gitea-pr-runner.service
|
||||
|
||||
if [ ! -f /var/lib/gitea-pr-runner/.runner ]; then
|
||||
read -r -s -p 'Enter the Gitea repository runner registration token: ' runner_token
|
||||
printf '\n'
|
||||
[ -n "$runner_token" ] || { echo 'Runner token is required' >&2; exit 1; }
|
||||
export GITEA_RUNNER_REGISTRATION_TOKEN="$runner_token"
|
||||
unset runner_token
|
||||
runuser --preserve-environment -u gitea-pr-runner -- \
|
||||
/usr/local/bin/gitea-runner register \
|
||||
--config /etc/gitea-pr-runner/config.yaml \
|
||||
--instance https://gitea.forust.xyz \
|
||||
--name homelab-pr \
|
||||
--labels homelab-pr:host \
|
||||
--no-interactive
|
||||
unset GITEA_RUNNER_REGISTRATION_TOKEN
|
||||
fi
|
||||
chmod 0600 /var/lib/gitea-pr-runner/.runner
|
||||
|
||||
systemctl daemon-reload
|
||||
systemctl enable --now gitea-pr-runner.service
|
||||
systemctl restart gitea-pr-runner.service
|
||||
echo "PR runner ready. Configuration backups: *.before-$stamp"
|
||||
@@ -1,48 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Native host runner, with pinned user-space tools and no extra CI images.
|
||||
set -euo pipefail
|
||||
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
|
||||
for tool in docker curl python3 git tar xz flock runuser systemctl; do
|
||||
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
|
||||
done
|
||||
docker info >/dev/null
|
||||
docker compose version >/dev/null
|
||||
docker buildx version >/dev/null
|
||||
id gitea-runner >/dev/null 2>&1 || useradd --system --create-home --home-dir /var/lib/gitea-runner --shell /usr/bin/bash gitea-runner
|
||||
# Reuse the established service account and runner registration.
|
||||
runner_home="$(getent passwd gitea-runner | cut -d: -f6)"
|
||||
[ "$runner_home" = /var/lib/gitea-runner ] || { echo 'Unexpected runner home; inspect the existing service first' >&2; exit 1; }
|
||||
runuser -u gitea-runner -- docker info >/dev/null || { echo "The runner user needs access to Docker before setup" >&2; exit 1; }
|
||||
command -v gitea-runner >/dev/null || { echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2; exit 1; }
|
||||
mkdir -p /etc/gitea-runner
|
||||
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
|
||||
for existing in /etc/gitea-runner/config.yaml /etc/systemd/system/gitea-runner.service; do
|
||||
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
|
||||
done
|
||||
scratch="$(mktemp -d)"
|
||||
trap 'rm -rf "$scratch"' EXIT
|
||||
chmod 755 "$scratch"
|
||||
install -m 0644 "$here/../workflows/install-ci-tools.sh" "$here/../workflows/tool-versions.env" "$scratch/"
|
||||
runuser -u gitea-runner -- bash "$scratch/install-ci-tools.sh"
|
||||
install -m 0644 "$here/config.yaml" /etc/gitea-runner/config.yaml
|
||||
python3 - <<'PYLABELS'
|
||||
import json
|
||||
from pathlib import Path
|
||||
registration = Path('/var/lib/gitea-runner/.runner')
|
||||
if registration.exists():
|
||||
labels = json.loads(registration.read_text()).get('labels', [])
|
||||
labels = [label for label in labels if isinstance(label, str) and label.split(':')[0] != 'homelab']
|
||||
labels.append('homelab:host')
|
||||
config = Path('/etc/gitea-runner/config.yaml')
|
||||
config.write_text(config.read_text().replace(' - homelab:host', '\n'.join(' - ' + json.dumps(label) for label in labels)))
|
||||
PYLABELS
|
||||
install -m 0644 "$here/gitea-runner.service" /etc/systemd/system/gitea-runner.service
|
||||
if [ ! -f /var/lib/gitea-runner/.runner ]; then
|
||||
echo 'Register once as gitea-runner with homelab:host before starting the service.'
|
||||
exit 0
|
||||
fi
|
||||
systemctl daemon-reload
|
||||
systemctl enable --now gitea-runner.service
|
||||
systemctl restart gitea-runner.service
|
||||
echo "Runner ready. Configuration backups: *.before-$stamp"
|
||||
@@ -1,32 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Run as the existing deploy user on workstation. Never resets the working tree.
|
||||
set -euo pipefail
|
||||
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
repo="${HOMELAB_REPO:-/srv/homelab}"
|
||||
for tool in python3 git kubectl helm docker flock timeout; do
|
||||
command -v "$tool" >/dev/null || { echo "Install missing dependency: $tool" >&2; exit 1; }
|
||||
done
|
||||
[ -d "$repo/.git" ] || { echo "Missing deploy checkout: $repo" >&2; exit 1; }
|
||||
[[ "$repo" =~ ^/[A-Za-z0-9_./-]+$ ]] || { echo 'Deploy path must be absolute and contain no whitespace' >&2; exit 1; }
|
||||
if [ "$(loginctl show-user "$USER" -p Linger --value)" != yes ]; then
|
||||
echo "Run once: sudo loginctl enable-linger $USER" >&2
|
||||
exit 1
|
||||
fi
|
||||
config="${XDG_CONFIG_HOME:-$HOME/.config}/homelab-deploy"
|
||||
mkdir -p "$config" "$HOME/.local/lib/homelab-deploy" "$HOME/.config/systemd/user"
|
||||
chmod 700 "$config"
|
||||
if [ ! -f "$config/environment" ]; then
|
||||
context="$(kubectl config current-context)"
|
||||
cluster_uid="$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')"
|
||||
printf 'HOMELAB_REPO=%s\nKUBE_CONTEXT=%s\nEXPECTED_CLUSTER_UID=%s\n' "$repo" "$context" "$cluster_uid" >"$config/environment"
|
||||
chmod 600 "$config/environment"
|
||||
fi
|
||||
# Do not replace a dispatcher while an existing deploy uses it.
|
||||
if systemctl --user list-units 'homelab-deploy@*' --state=running --no-legend | grep -q .; then
|
||||
echo 'An existing deploy is running; wait before updating the controller' >&2
|
||||
exit 1
|
||||
fi
|
||||
install -m 0755 "$here/../workflows/deploy-controller.py" "$HOME/.local/lib/homelab-deploy/controller.py"
|
||||
install -m 0644 "$here/homelab-deploy@.service" "$HOME/.config/systemd/user/homelab-deploy@.service"
|
||||
systemctl --user daemon-reload
|
||||
echo 'Controller ready. Run a checked main SHA in full mode for the initial baseline.'
|
||||
@@ -1,94 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# Local regressions only: kubectl is mocked and Docker is used for config parsing.
|
||||
set -euo pipefail
|
||||
repo="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
|
||||
scratch="$(mktemp -d)"
|
||||
trap 'rm -rf "$scratch"' EXIT
|
||||
|
||||
mkdir -p "$scratch/repo/app" "$scratch/repo/postgres" "$scratch/repo/netbird" "$scratch/repo/renovate"
|
||||
git -C "$scratch/repo" init -q
|
||||
for file in app/compose.yaml postgres/shared-compose.yaml netbird/client.compose.yaml renovate/renovate-compose.yaml; do
|
||||
touch "$scratch/repo/$file"
|
||||
done
|
||||
git -C "$scratch/repo" add .
|
||||
# shellcheck source=../workflows/compose-lint.sh
|
||||
source "$repo/.gitea/workflows/compose-lint.sh"
|
||||
actual="$(cd "$scratch/repo" && compose_files)"
|
||||
expected=$'app/compose.yaml\nnetbird/client.compose.yaml\npostgres/shared-compose.yaml\nrenovate/renovate-compose.yaml'
|
||||
[ "$actual" = "$expected" ] || { echo 'Compose discovery missed a file' >&2; exit 1; }
|
||||
|
||||
cat >"$scratch/compose.yaml" <<'YAML'
|
||||
services:
|
||||
example:
|
||||
image: busybox:1.37.0
|
||||
environment:
|
||||
REQUIRED: ${HOMELAB_TEST_REQUIRED:?required for this regression}
|
||||
YAML
|
||||
unset HOMELAB_TEST_REQUIRED
|
||||
if validate_compose_file "$scratch/compose.yaml" >"$scratch/config.log" 2>&1; then
|
||||
echo 'Full Compose validation accepted a missing variable' >&2
|
||||
exit 1
|
||||
fi
|
||||
grep -q 'required for this regression' "$scratch/config.log"
|
||||
HOMELAB_TEST_REQUIRED=present validate_compose_file "$scratch/compose.yaml"
|
||||
|
||||
cat >"$scratch/resources.json" <<'JSON'
|
||||
{"kind":"List","items":[
|
||||
{"kind":"Deployment","metadata":{"namespace":"app"},"spec":{"template":{"spec":{
|
||||
"containers":[{"envFrom":[{"secretRef":{"name":"credentials"}},{"secretRef":{"name":"optional","optional":true}}],"env":[{"valueFrom":{"secretKeyRef":{"name":"credentials","key":"password"}}}]}],
|
||||
"initContainers":[{"envFrom":[{"secretRef":{"name":"init"}}]}],
|
||||
"imagePullSecrets":[{"name":"registry"}],
|
||||
"volumes":[{"secret":{"secretName":"mounted"}},{"projected":{"sources":[{"secret":{"name":"projected"}},{"secret":{"name":"optional-projected","optional":true}}]}}]
|
||||
}}}},
|
||||
{"kind":"CronJob","metadata":{},"spec":{"jobTemplate":{"spec":{"template":{"spec":{"containers":[{"envFrom":[{"secretRef":{"name":"cron"}}]}]}}}}}},
|
||||
{"kind":"IngressRoute","metadata":{"namespace":"app"},"spec":{"tls":{"secretName":"controller-issued-tls"}}}
|
||||
]}
|
||||
JSON
|
||||
actual="$(jq -r -f "$repo/.gitea/workflows/secret-references.jq" "$scratch/resources.json" | sort)"
|
||||
expected=$'app credentials\napp init\napp mounted\napp projected\napp registry\ndefault cron'
|
||||
[ "$actual" = "$expected" ] || { echo "Unexpected Secret references: $actual" >&2; exit 1; }
|
||||
|
||||
REPO="$repo"
|
||||
# shellcheck source=../workflows/deploy-lib.sh
|
||||
source "$repo/.gitea/workflows/deploy-lib.sh"
|
||||
K8S_MANIFESTS=("$scratch/resources.json")
|
||||
KUSTOMIZE_APPS=()
|
||||
# No live cluster access. Reject credentials in app even if they exist elsewhere.
|
||||
kubectl() {
|
||||
case "$1" in
|
||||
create) cat "$scratch/resources.json" ;;
|
||||
get)
|
||||
if [ "$3" = credentials ] && [ "$5" = app ]; then
|
||||
return 1
|
||||
fi
|
||||
return 0
|
||||
;;
|
||||
*) echo "Unexpected kubectl invocation: $*" >&2; return 1 ;;
|
||||
esac
|
||||
}
|
||||
if check_referenced_secrets >"$scratch/secrets.log"; then
|
||||
echo 'Namespace-scoped Secret check accepted a missing Secret' >&2
|
||||
exit 1
|
||||
fi
|
||||
grep -q 'MISSING OR UNREADABLE: app/credentials' "$scratch/secrets.log"
|
||||
# API/rendering errors must not produce an empty reference list and pass.
|
||||
kubectl() { return 1; }
|
||||
if ! skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
|
||||
echo 'VMAgent preflight did not skip an uninstalled CRD' >&2
|
||||
exit 1
|
||||
fi
|
||||
kubectl() { return 0; }
|
||||
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
|
||||
echo 'VMAgent preflight skipped an installed CRD' >&2
|
||||
exit 1
|
||||
fi
|
||||
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/victoria.yaml"; then
|
||||
echo 'VMAgent preflight skipped an unrelated manifest' >&2
|
||||
exit 1
|
||||
fi
|
||||
kubectl() { return 1; }
|
||||
if check_referenced_secrets >"$scratch/secrets.log"; then
|
||||
echo 'Secret check accepted a failed manifest render' >&2
|
||||
exit 1
|
||||
fi
|
||||
printf '%s\n' 'Deploy validation regressions passed.'
|
||||
+386
-357
@@ -1,25 +1,38 @@
|
||||
name: ci
|
||||
"on":
|
||||
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
pull_request: null
|
||||
workflow_dispatch: null
|
||||
- "**"
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
|
||||
# Every job here is checkout plus local tools. The token needs to read the tree
|
||||
# and nothing else, and saying so keeps a future step that reaches for the API
|
||||
# from quietly holding a token that can write to the repository.
|
||||
permissions:
|
||||
contents: read
|
||||
actions: read
|
||||
|
||||
concurrency:
|
||||
group: ci-${{ github.ref }}
|
||||
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
|
||||
|
||||
env:
|
||||
REGISTRY: gcr.forust.xyz
|
||||
|
||||
jobs:
|
||||
compose:
|
||||
name: Compose
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
lint-compose:
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
id: source
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
# Structure check for every committed Compose file, active or not.
|
||||
# Interpolation, env-file and bind-mount resolution are all switched off,
|
||||
# because inactive stacks have no .env here and would only fail on their
|
||||
# ${VAR:?} guards. Active stacks get the full check with interpolation in
|
||||
# the deploy workflow, where the real .env files live.
|
||||
- name: Validate Compose files
|
||||
shell: bash
|
||||
run: |
|
||||
@@ -48,77 +61,35 @@ jobs:
|
||||
exit 1
|
||||
fi
|
||||
echo "checked ${#files[@]} Compose file(s)"
|
||||
id: check
|
||||
- name: Write the job result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: Compose
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP:
|
||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.source.conclusion == 'failure'
|
||||
&& 'Source checkout' || '' }}
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
fi
|
||||
workflows:
|
||||
name: Workflows
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
|
||||
lint-actionlint:
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
id: source
|
||||
- name: Prepare pinned tools
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint shellcheck)"
|
||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||
id: tools
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Lint Gitea Actions workflows with actionlint
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint)"
|
||||
export PATH="$tools_dir:$PATH"
|
||||
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
|
||||
id: check
|
||||
- name: Write the job result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: Workflows
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP:
|
||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
fi
|
||||
shell:
|
||||
name: Shell
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
|
||||
lint-shellcheck:
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
id: source
|
||||
- name: Prepare pinned tools
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck jq)"
|
||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||
id: tools
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Lint shell scripts with ShellCheck
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck)"
|
||||
export PATH="$tools_dir:$PATH"
|
||||
mapfile -t scripts < <(
|
||||
git ls-files '*.sh' ':(glob)**/*.bash'
|
||||
)
|
||||
@@ -127,42 +98,20 @@ jobs:
|
||||
exit 0
|
||||
fi
|
||||
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
|
||||
bash .gitea/tests/deploy-validation.sh
|
||||
id: check
|
||||
- name: Write the job result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: Shell
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP:
|
||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
fi
|
||||
formatting:
|
||||
name: Formatting
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
|
||||
lint-prettier:
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
id: source
|
||||
- name: Prepare pinned tools
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
|
||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||
id: tools
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Check formatting with Prettier
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
|
||||
export PATH="$tools_dir:$PATH"
|
||||
|
||||
mapfile -t prettier_files < <(
|
||||
git ls-files \
|
||||
@@ -176,79 +125,36 @@ jobs:
|
||||
fi
|
||||
|
||||
prettier --check --ignore-unknown "${prettier_files[@]}"
|
||||
id: check
|
||||
- name: Write the job result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: Formatting
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP:
|
||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
fi
|
||||
python:
|
||||
name: Python and tests
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
|
||||
lint-ruff:
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
id: source
|
||||
- name: Prepare pinned tools
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff jq)"
|
||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||
id: tools
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Lint and format-check Python with Ruff
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
ruff check . .gitea/workflows
|
||||
ruff format --check . .gitea/workflows
|
||||
python3 -m unittest discover -s tests -v
|
||||
id: check
|
||||
- name: Write the job result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: Python and tests
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP:
|
||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
fi
|
||||
yaml:
|
||||
name: YAML
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff)"
|
||||
export PATH="$tools_dir:$PATH"
|
||||
ruff check .
|
||||
ruff format --check .
|
||||
|
||||
lint-yaml:
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
id: source
|
||||
- name: Prepare pinned tools
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
|
||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||
id: tools
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Lint YAML syntax
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
|
||||
export PATH="$tools_dir:$PATH"
|
||||
|
||||
mapfile -t yaml_files < <(
|
||||
git ls-files '*.yaml' '*.yml' \
|
||||
@@ -262,41 +168,20 @@ jobs:
|
||||
fi
|
||||
|
||||
yamllint -c .yamllint "${yaml_files[@]}"
|
||||
id: check
|
||||
- name: Write the job result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: YAML
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP:
|
||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
fi
|
||||
dockerfiles:
|
||||
name: Dockerfiles
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
|
||||
lint-dockerfiles:
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
id: source
|
||||
- name: Prepare pinned tools
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
|
||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||
id: tools
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Lint Dockerfiles
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
|
||||
export PATH="$tools_dir:$PATH"
|
||||
|
||||
mapfile -t dockerfiles < <(
|
||||
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
|
||||
@@ -308,41 +193,20 @@ jobs:
|
||||
fi
|
||||
|
||||
hadolint -c .hadolint.yaml "${dockerfiles[@]}"
|
||||
id: check
|
||||
- name: Write the job result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: Dockerfiles
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP:
|
||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
fi
|
||||
kubernetes:
|
||||
name: Kubernetes
|
||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
timeout-minutes: 15
|
||||
|
||||
validate:
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 20
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
id: source
|
||||
- name: Prepare pinned tools
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
|
||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||
id: tools
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Validate Kubernetes manifests against JSON schemas
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
|
||||
export PATH="$tools_dir:$PATH"
|
||||
|
||||
mapfile -t manifests < <(
|
||||
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
|
||||
@@ -359,162 +223,327 @@ jobs:
|
||||
-ignore-missing-schemas \
|
||||
-summary \
|
||||
"${manifests[@]}"
|
||||
id: check
|
||||
- name: Write the job result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: Kubernetes
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP:
|
||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||
|
||||
# kubeconform has no schemas for CRDs, so every IngressRoute, Certificate,
|
||||
# PrometheusRule, Middleware, ServersTransport and ServiceMonitor is silently
|
||||
# skipped above. The live API server knows the real CRD schemas (and runs the
|
||||
# cert-manager / Traefik admission webhooks), so validate there too.
|
||||
#
|
||||
# Only services marked with a k8s/active marker are checked: server-side
|
||||
# dry-run needs the target namespace to exist, and inactive services are not
|
||||
# deployed. Services being enabled for the first time are still covered by
|
||||
# the JSON-schema pass above.
|
||||
#
|
||||
# Main pushes only. `--dry-run=server` persists nothing, but it does execute
|
||||
# the admission webhooks of the production API server, so anyone able to open
|
||||
# a pull request would be able to run arbitrary manifest content through
|
||||
# cert-manager and Traefik. A pull request has nothing to gain from it either:
|
||||
# only main is ever deployed, and this job runs to completion before the
|
||||
# deploy workflow is allowed to start, so a bad CRD is still caught before
|
||||
# anything reaches the cluster -- just on the push rather than on the PR.
|
||||
- name: Note the server-side check is not running here
|
||||
if: github.event_name == 'pull_request' || github.ref != 'refs/heads/main'
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
echo "::notice::Skipping the server-side dry-run. It executes the cert-manager and" \
|
||||
"Traefik admission webhooks against the production API server, so it is limited" \
|
||||
"to pushes to main. CRDs are still schema-checked by kubeconform above, and the" \
|
||||
"server-side pass still runs on main before the deploy."
|
||||
|
||||
- name: Validate active manifests against the live API server
|
||||
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
|
||||
if ! kubectl get --raw='/readyz' --request-timeout=10s >/dev/null 2>&1; then
|
||||
echo "::warning::Cluster unreachable — skipped server-side validation of CRDs (IngressRoute, Certificate, PrometheusRule). Review manifest changes manually."
|
||||
exit 0
|
||||
fi
|
||||
image-plan:
|
||||
needs: [compose, workflows, shell, formatting, python, yaml, dockerfiles, kubernetes]
|
||||
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
|
||||
runs-on: homelab
|
||||
timeout-minutes: 10
|
||||
|
||||
mapfile -t k8s_dirs < <(
|
||||
git ls-files '*.yaml' '*.yml' \
|
||||
| grep -E '(^|/)k8s/' \
|
||||
| sed -E 's#((^|.*/)k8s)/.*#\1#' \
|
||||
| sort -u
|
||||
)
|
||||
|
||||
manifests=()
|
||||
kustomize_apps=()
|
||||
for dir in "${k8s_dirs[@]}"; do
|
||||
if [ ! -f "${dir}/active" ]; then
|
||||
echo "skip (no k8s/active): ${dir}"
|
||||
continue
|
||||
fi
|
||||
if [ -f "${dir}/overlays/prod/kustomization.yaml" ]; then
|
||||
kustomize_apps+=("${dir}/overlays/prod")
|
||||
elif [ -f "${dir}/base/kustomization.yaml" ]; then
|
||||
kustomize_apps+=("${dir}/base")
|
||||
else
|
||||
while IFS= read -r f; do
|
||||
[ -n "$f" ] && manifests+=("$f")
|
||||
done < <(
|
||||
git ls-files "${dir}/*.yaml" "${dir}/*.yml" \
|
||||
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
|
||||
)
|
||||
fi
|
||||
done
|
||||
|
||||
echo "server-side dry-run: ${#manifests[@]} manifests, ${#kustomize_apps[@]} kustomize apps"
|
||||
failed=0
|
||||
for m in ${manifests[@]+"${manifests[@]}"}; do
|
||||
if ! out="$(kubectl apply --dry-run=server -f "$m" 2>&1)"; then
|
||||
failed=1
|
||||
echo "::error file=${m}::$(printf '%s' "$out" | head -1)"
|
||||
fi
|
||||
done
|
||||
for k in ${kustomize_apps[@]+"${kustomize_apps[@]}"}; do
|
||||
if ! out="$(kubectl apply -k "$k" --dry-run=server 2>&1)"; then
|
||||
failed=1
|
||||
echo "::error file=${k}::$(printf '%s' "$out" | head -1)"
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$failed" -ne 0 ]; then
|
||||
echo "Server-side validation failed. The API server (or an admission webhook) rejected these manifests."
|
||||
exit 1
|
||||
fi
|
||||
echo "server-side dry-run: all active manifests accepted by the API server"
|
||||
|
||||
build:
|
||||
needs:
|
||||
# The panel's scan-deps/test-backend/test-frontend jobs gated here until
|
||||
# userbot moved to its own repo; upstream's code is upstream's gate now.
|
||||
# The rule is unchanged: publishing and passing the checks are the same
|
||||
# gate, so a commit that fails any of these still cannot move :prod.
|
||||
[lint-actionlint, lint-shellcheck, lint-compose, lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
|
||||
if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev') && !startsWith(github.ref_name, 'renovate/')
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 60
|
||||
outputs:
|
||||
matrix: ${{ steps.plan.outputs.matrix }}
|
||||
services: ${{ steps.services.outputs.services }}
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
id: source
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
- name: Detect build inputs against successful CI
|
||||
id: plan
|
||||
env:
|
||||
GITEA_TOKEN: ${{ github.token }}
|
||||
run: python3 .gitea/workflows/release.py prepare --output build-plan.json
|
||||
- name: Store the image plan
|
||||
id: artifact
|
||||
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
||||
with:
|
||||
name: build-plan
|
||||
path: build-plan.json
|
||||
if-no-files-found: error
|
||||
retention-days: 30
|
||||
|
||||
- name: Write the plan result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: Image plan
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP: >-
|
||||
${{ steps.plan.conclusion == 'failure' && 'Build input detection' ||
|
||||
steps.artifact.conclusion == 'failure' && 'Plan upload' ||
|
||||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||
- name: Detect changed docker-built services
|
||||
id: services
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## Image plan\n\nResult: %s\n' "$SUMMARY_RESULT" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
set -euo pipefail
|
||||
base="${{ github.event.before }}"
|
||||
if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then
|
||||
base="$(git rev-list --max-parents=0 HEAD)"
|
||||
fi
|
||||
|
||||
images:
|
||||
name: Image (${{ matrix.name }})
|
||||
needs: [image-plan]
|
||||
if: needs.image-plan.result == 'success'
|
||||
runs-on: homelab
|
||||
timeout-minutes: 60
|
||||
strategy:
|
||||
max-parallel: 1
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJSON(needs.image-plan.outputs.matrix || '{"include":[{"name":"inactive"}]}') }}
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
id: source
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
- name: Download the checked image plan
|
||||
id: inputs
|
||||
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
|
||||
with:
|
||||
name: build-plan
|
||||
- name: Build or reuse this image
|
||||
id: check
|
||||
env:
|
||||
IMAGE_NAME: ${{ matrix.name }}
|
||||
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
|
||||
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
|
||||
run: python3 .gitea/workflows/release.py image --image "$IMAGE_NAME" --output image.json
|
||||
- name: Store the image result
|
||||
id: artifact
|
||||
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
||||
with:
|
||||
name: image-${{ matrix.name }}
|
||||
path: image.json
|
||||
if-no-files-found: error
|
||||
retention-days: 30
|
||||
- name: Write the job result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: Image (${{ matrix.name }})
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP: >-
|
||||
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
|
||||
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
|
||||
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
|
||||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||
# A failed diff used to leave changed_files empty, which reads exactly
|
||||
# like "nothing to build": the job went green having built nothing and
|
||||
# the tag never moved. The status is checked, not assumed.
|
||||
if ! changed="$(git diff --name-only "$base" "${GITHUB_SHA}")"; then
|
||||
echo "::error::cannot diff ${base}..${GITHUB_SHA}"
|
||||
exit 1
|
||||
fi
|
||||
mapfile -t changed_files <<<"$changed"
|
||||
|
||||
services=()
|
||||
|
||||
add_service() {
|
||||
local name="$1"
|
||||
local seen=0
|
||||
for existing in "${services[@]}"; do
|
||||
if [ "$existing" = "$name" ]; then
|
||||
seen=1
|
||||
break
|
||||
fi
|
||||
done
|
||||
if [ "$seen" -eq 0 ]; then
|
||||
services+=("$name")
|
||||
fi
|
||||
}
|
||||
|
||||
for file in "${changed_files[@]}"; do
|
||||
case "$file" in
|
||||
errorpages/*)
|
||||
add_service errorpages
|
||||
;;
|
||||
homepages/*)
|
||||
add_service homepages
|
||||
;;
|
||||
edu_master/phpsessid-bot/*|edu_master/webinar-checker/*|edu_master/compose.yaml)
|
||||
add_service edu_master
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [ "${#services[@]}" -eq 0 ]; then
|
||||
echo "No docker-built services changed."
|
||||
echo "services=" >> "$GITHUB_OUTPUT"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
printf '%s\n' "${services[@]}" | tee /tmp/services.txt
|
||||
echo "services=$(paste -sd, /tmp/services.txt)" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Log in to registry
|
||||
# The pin step below also writes (manifest PUTs), and it runs on every
|
||||
# main push — including manifest-only ones where services is empty. A
|
||||
# stale persistent login on the old runner used to mask this; a clean
|
||||
# runner pushes anonymously and gets 401.
|
||||
if: steps.services.outputs.services != '' || github.ref_name == 'main'
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
fi
|
||||
|
||||
# Retain the build job name required by the immutable release deployment gate.
|
||||
build:
|
||||
needs: [image-plan, images]
|
||||
runs-on: homelab
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
id: source
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||
- name: Download all image results
|
||||
id: inputs
|
||||
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
|
||||
with:
|
||||
path: artifacts
|
||||
- name: Pin SHA tags and write the complete release
|
||||
id: check
|
||||
# Through env, not by substitution into the script. A secret written
|
||||
# into a run: block is pasted into the shell source before bash parses
|
||||
# it, so a password containing a quote, a backtick or $(...) becomes
|
||||
# code that runs. Masking the value in the log does not prevent that.
|
||||
env:
|
||||
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
|
||||
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
|
||||
run: >-
|
||||
python3 .gitea/workflows/release.py finalize
|
||||
--plan artifacts/build-plan/build-plan.json
|
||||
- name: Store commit release
|
||||
id: artifact
|
||||
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
||||
with:
|
||||
name: release-${{ github.sha }}
|
||||
path: release.json
|
||||
if-no-files-found: error
|
||||
retention-days: 30
|
||||
- name: Write the job result
|
||||
if: always()
|
||||
env:
|
||||
SUMMARY_CHECK: Image release and SHA tags
|
||||
SUMMARY_RESULT: ${{ job.status }}
|
||||
SUMMARY_FAILED_STEP: >-
|
||||
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
|
||||
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
|
||||
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
|
||||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
printf '%s' "$REGISTRY_PASSWORD" | docker login "${REGISTRY}" \
|
||||
-u "$REGISTRY_USERNAME" \
|
||||
--password-stdin
|
||||
|
||||
- name: Build and push changed images
|
||||
if: steps.services.outputs.services != ''
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/release.py ]; then
|
||||
python3 .gitea/workflows/release.py check-summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
||||
# This step was the one run: block in the workflow without it, and it
|
||||
# is the one that cannot afford it: a docker push that failed partway
|
||||
# through the loop used to be followed by more pushes, the loop's exit
|
||||
# status came from the last one, and the job went green with half the
|
||||
# images missing from the registry.
|
||||
set -euo pipefail
|
||||
IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}"
|
||||
|
||||
# Tags for this push. The commit-pinned name is the point of this
|
||||
# step: the deploy resolves it in preference to :prod, so a deploy
|
||||
# that sat in the queue behind a later push still gets the build of
|
||||
# the commit CI validated, instead of whatever :prod points at by the
|
||||
# time it runs. See render_pinned in deploy-lib.sh.
|
||||
commit_tag=""
|
||||
if [ "${GITHUB_REF_NAME}" = "main" ]; then
|
||||
commit_tag="sha-${GITHUB_SHA:0:12}"
|
||||
fi
|
||||
|
||||
set_tags() {
|
||||
tags=()
|
||||
case "${GITHUB_REF_NAME}" in
|
||||
main) tags+=("main" "prod") ;;
|
||||
dev) tags+=("dev") ;;
|
||||
esac
|
||||
if [ -n "$commit_tag" ]; then
|
||||
tags+=("$commit_tag")
|
||||
fi
|
||||
}
|
||||
|
||||
for service in "${services[@]}"; do
|
||||
case "$service" in
|
||||
errorpages)
|
||||
image="${REGISTRY}/forust/error-pages"
|
||||
set_tags
|
||||
build_args=()
|
||||
for tag in "${tags[@]}"; do
|
||||
build_args+=(-t "${image}:${tag}")
|
||||
done
|
||||
docker build \
|
||||
--cache-from "type=registry,ref=${image}:buildcache" \
|
||||
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
|
||||
"${build_args[@]}" errorpages
|
||||
for tag in "${tags[@]}"; do
|
||||
docker push "${image}:${tag}"
|
||||
done
|
||||
;;
|
||||
homepages)
|
||||
for variant in forust xdfnx; do
|
||||
case "$variant" in
|
||||
forust)
|
||||
image="${REGISTRY}/forust/forust-homepage"
|
||||
;;
|
||||
xdfnx)
|
||||
image="${REGISTRY}/forust/xdfnx-homepage"
|
||||
;;
|
||||
esac
|
||||
set_tags
|
||||
build_args=()
|
||||
for tag in "${tags[@]}"; do
|
||||
build_args+=(-t "${image}:${tag}")
|
||||
done
|
||||
docker build \
|
||||
--cache-from "type=registry,ref=${image}:buildcache" \
|
||||
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
|
||||
"${build_args[@]}" -f "homepages/Dockerfile.${variant}" homepages
|
||||
for tag in "${tags[@]}"; do
|
||||
docker push "${image}:${tag}"
|
||||
done
|
||||
done
|
||||
;;
|
||||
edu_master)
|
||||
for variant in session-keeper webinar-checker; do
|
||||
case "$variant" in
|
||||
session-keeper)
|
||||
context="edu_master/phpsessid-bot"
|
||||
image="${REGISTRY}/forust/session-keeper"
|
||||
;;
|
||||
webinar-checker)
|
||||
context="edu_master/webinar-checker"
|
||||
image="${REGISTRY}/forust/webinar-checker"
|
||||
;;
|
||||
esac
|
||||
set_tags
|
||||
build_args=()
|
||||
for tag in "${tags[@]}"; do
|
||||
build_args+=(-t "${image}:${tag}")
|
||||
done
|
||||
docker build \
|
||||
--cache-from "type=registry,ref=${image}:buildcache" \
|
||||
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
|
||||
"${build_args[@]}" "$context"
|
||||
for tag in "${tags[@]}"; do
|
||||
docker push "${image}:${tag}"
|
||||
done
|
||||
done
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
# Every image the tree names has to carry the commit-pinned name, not only
|
||||
# the ones this push rebuilt. A push that touches nothing but manifests
|
||||
# builds nothing, and its deploy would then find no commit-pinned tag to
|
||||
# resolve and quietly fall back to the moving :prod - which is the whole
|
||||
# failure the commit-pinned name exists to remove.
|
||||
#
|
||||
# Re-tagging copies the manifest list and transfers no layers, so pinning
|
||||
# six images that already exist costs six registry writes.
|
||||
#
|
||||
# The list is derived from the tree rather than written out here, so an
|
||||
# image added to a manifest is covered without a second place to update.
|
||||
- name: Pin the commit name on the images this push did not rebuild
|
||||
if: github.ref_name == 'main'
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
commit_tag="sha-${GITHUB_SHA:0:12}"
|
||||
mapfile -t repos < <(
|
||||
git grep -hoE 'gcr\.forust\.xyz/forust/[A-Za-z0-9._-]+' -- '*.yaml' '*.yml' \
|
||||
| sort -u
|
||||
)
|
||||
if [ "${#repos[@]}" -eq 0 ]; then
|
||||
echo "No own images referenced by the tree."
|
||||
exit 0
|
||||
fi
|
||||
echo "pinning ${#repos[@]} image(s) to $commit_tag"
|
||||
for repo in "${repos[@]}"; do
|
||||
if docker buildx imagetools inspect "$repo:$commit_tag" >/dev/null 2>&1; then
|
||||
echo " already built by this push: ${repo##*/}"
|
||||
continue
|
||||
fi
|
||||
if ! docker buildx imagetools inspect "$repo:prod" >/dev/null 2>&1; then
|
||||
echo " WARNING: ${repo##*/} has no :prod to pin and no build produced it"
|
||||
continue
|
||||
fi
|
||||
docker buildx imagetools create --tag "$repo:$commit_tag" "$repo:prod"
|
||||
echo " pinned ${repo##*/}"
|
||||
done
|
||||
@@ -21,7 +21,8 @@
|
||||
# All committed Compose files, including the ones deploy never starts.
|
||||
compose_files() {
|
||||
git ls-files \
|
||||
'*compose.yaml' '*compose.yml'
|
||||
'*/compose.yaml' '*/compose.yml' 'compose.yaml' 'compose.yml' \
|
||||
'*/docker-compose.yaml' '*/docker-compose.yml'
|
||||
}
|
||||
|
||||
# Prints the flags that turn `docker compose config` into the general check.
|
||||
|
||||
@@ -1,103 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Resolve Compose images without changing project names or local bind paths."""
|
||||
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import subprocess
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def output(*args, **kwargs):
|
||||
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
|
||||
|
||||
|
||||
def resolve(reference):
|
||||
if '@sha256:' in reference:
|
||||
return reference
|
||||
descriptor = json.loads(
|
||||
output('docker', 'buildx', 'imagetools', 'inspect', reference, '--format', '{{json .Manifest}}')
|
||||
)
|
||||
digest = descriptor['digest']
|
||||
if not re.fullmatch(r'sha256:[0-9a-f]{64}', digest):
|
||||
raise ValueError(f'Invalid registry digest for {reference}')
|
||||
# Strip tag only from the final path segment (registry ports are preserved).
|
||||
repository = reference.rsplit('/', 1)
|
||||
repository[-1] = repository[-1].split(':')[0]
|
||||
return '/'.join(repository) + '@' + digest
|
||||
|
||||
|
||||
def prepare(source_file):
|
||||
config_repo = Path(os.environ['CONFIG_REPO'])
|
||||
source_repo = Path(os.environ['REPO'])
|
||||
directory = Path(os.environ['RUN_DIR'])
|
||||
relative = source_file.relative_to(source_repo)
|
||||
project_dir = config_repo / relative.parent
|
||||
base = ['docker', 'compose', '--project-directory', str(project_dir), '-f', str(source_file)]
|
||||
config = json.loads(output(*base, 'config', '--format', 'json', cwd=config_repo))
|
||||
project = config['name']
|
||||
previous_file = directory / 'previous.json'
|
||||
previous = json.loads(previous_file.read_text()) if previous_file.exists() else {}
|
||||
images_file = directory / 'compose-images.json'
|
||||
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
|
||||
release = json.loads((directory / 'release.json').read_text())
|
||||
before = json.loads(json.dumps(config))
|
||||
for service, settings in config['services'].items():
|
||||
reference = settings.get('image')
|
||||
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
|
||||
if not reference or settings.get('build'):
|
||||
raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
|
||||
image_repo = reference.split('@')[0].rsplit('/', 1)
|
||||
image_repo[-1] = image_repo[-1].split(':')[0]
|
||||
image_repo = '/'.join(image_repo)
|
||||
# Nextcloud AIO validates the mastercontainer image reference and rejects
|
||||
# a digest. Keep its configured tag so AIO can start and manage its stack.
|
||||
if nextcloud_aio_master:
|
||||
pinned = reference
|
||||
elif image_repo in release['images']:
|
||||
pinned = image_repo + '@' + release['images'][image_repo]
|
||||
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
|
||||
pinned = locks[reference]
|
||||
else:
|
||||
pinned = resolve(reference)
|
||||
settings['image'] = pinned
|
||||
locks[reference] = pinned
|
||||
# Capture what is running, not the current value of its mutable tag.
|
||||
ids = output(
|
||||
'docker',
|
||||
'ps',
|
||||
'-aq',
|
||||
'--filter',
|
||||
f'label=com.docker.compose.project={project}',
|
||||
'--filter',
|
||||
f'label=com.docker.compose.service={service}',
|
||||
).splitlines()
|
||||
actual = set()
|
||||
for container in ids:
|
||||
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
|
||||
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
|
||||
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
|
||||
if len(actual) > 1:
|
||||
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
|
||||
# AIO also rejects a digest in its recovery config. Preserve its tag in
|
||||
# both deploy and recovery files.
|
||||
if nextcloud_aio_master:
|
||||
before['services'][service]['image'] = reference
|
||||
else:
|
||||
before['services'][service]['image'] = next(iter(actual)) if actual else reference
|
||||
for name, data in (('compose', config), ('compose-before', before)):
|
||||
folder = directory / name
|
||||
folder.mkdir(mode=0o700, exist_ok=True)
|
||||
destination = folder / f'{relative.parent.name}.json'
|
||||
destination.write_text(json.dumps(data, indent=2) + '\n')
|
||||
destination.chmod(0o600)
|
||||
images_file.write_text(json.dumps(locks, indent=2) + '\n')
|
||||
print(f'Compose {project}: images pinned; local paths preserved')
|
||||
print(
|
||||
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never'
|
||||
)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
prepare(Path(sys.argv[1]))
|
||||
@@ -1,423 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Durable workstation deployment controller. Install with setup-workstation.sh."""
|
||||
|
||||
import argparse
|
||||
import contextlib
|
||||
import fcntl
|
||||
import importlib.util
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
STATE = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
|
||||
CONFIG_REPO = Path(os.environ.get('HOMELAB_REPO', '/srv/homelab'))
|
||||
RUN_ID = re.compile(r'[0-9]+-[0-9]+')
|
||||
|
||||
|
||||
def command(*args, **kwargs):
|
||||
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
|
||||
|
||||
|
||||
def atomic_json(path, data):
|
||||
temporary = path.with_suffix('.tmp')
|
||||
temporary.write_text(json.dumps(data, indent=2) + '\n')
|
||||
temporary.chmod(0o600)
|
||||
temporary.replace(path)
|
||||
|
||||
|
||||
@contextlib.contextmanager
|
||||
def lock(name):
|
||||
STATE.mkdir(mode=0o700, parents=True, exist_ok=True)
|
||||
with (STATE / name).open('a') as stream:
|
||||
fcntl.flock(stream, fcntl.LOCK_EX)
|
||||
yield
|
||||
|
||||
|
||||
def load_module(name, path):
|
||||
spec = importlib.util.spec_from_file_location(name, path)
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
def run_directory(run_id):
|
||||
if not RUN_ID.fullmatch(run_id):
|
||||
raise ValueError('Run ID must be numeric workflow-id and attempt')
|
||||
return STATE / 'runs' / run_id
|
||||
|
||||
|
||||
def start(run_id):
|
||||
payload = sys.stdin.buffer.read(256 * 1024 + 1)
|
||||
if len(payload) > 256 * 1024:
|
||||
raise ValueError('Deploy request exceeds 256 KiB')
|
||||
request = json.loads(payload)
|
||||
sha = request['release']['sha']
|
||||
if not re.fullmatch(r'[0-9a-f]{40}', sha) or request['mode'] not in ('changed', 'full', 'plan'):
|
||||
raise ValueError('Invalid deploy SHA or mode')
|
||||
if not isinstance(request['refresh_images'], bool):
|
||||
raise ValueError('refresh_images must be boolean')
|
||||
directory = run_directory(run_id)
|
||||
with lock('prepare.lock'):
|
||||
if (directory / 'request.json').exists():
|
||||
if json.loads((directory / 'request.json').read_text()) != request:
|
||||
raise ValueError('Run ID already belongs to a different request')
|
||||
else:
|
||||
directory.mkdir(mode=0o700, parents=True, exist_ok=True)
|
||||
command('git', '-C', str(CONFIG_REPO), 'fetch', '--quiet', 'origin', 'main')
|
||||
command('git', '-C', str(CONFIG_REPO), 'merge-base', '--is-ancestor', sha, 'origin/main')
|
||||
if not (directory / 'source').exists():
|
||||
command('git', '-C', str(CONFIG_REPO), 'worktree', 'add', '--detach', str(directory / 'source'), sha)
|
||||
if command('git', '-C', str(directory / 'source'), 'rev-parse', 'HEAD') != sha:
|
||||
raise ValueError('Prepared source does not match deploy SHA')
|
||||
release_module = load_module('release', directory / 'source/.gitea/workflows/release.py')
|
||||
release_module.validate_release(request['release'], sha)
|
||||
atomic_json(directory / 'release.json', request['release'])
|
||||
atomic_json(directory / 'request.json', request)
|
||||
if not (directory / 'status.json').exists():
|
||||
atomic_json(directory / 'status.json', {'state': 'queued', 'stages': {}})
|
||||
# Starting an existing active or finished ID is idempotent; never re-apply it.
|
||||
if json.loads((directory / 'status.json').read_text())['state'] == 'queued':
|
||||
command('systemctl', '--user', 'start', '--no-block', f'homelab-deploy@{run_id}.service')
|
||||
print(f'Accepted deploy {run_id} ({sha})')
|
||||
|
||||
|
||||
def environment(directory):
|
||||
request = json.loads((directory / 'request.json').read_text())
|
||||
return {
|
||||
**os.environ,
|
||||
'REPO': str(directory / 'source'),
|
||||
'CONFIG_REPO': str(CONFIG_REPO),
|
||||
'RUN_DIR': str(directory),
|
||||
'DEPLOY_SHA': request['release']['sha'],
|
||||
'RELEASE_FILE': str(directory / 'release.json'),
|
||||
'DEPLOY_PLAN': str(directory / 'plan.json'),
|
||||
'DEPLOY_SNAPSHOT_DIR': str(directory / 'snapshot'),
|
||||
'REFRESH_IMAGES': str(request['refresh_images']).lower(),
|
||||
'ROLLOUT_PARALLELISM': '4',
|
||||
}
|
||||
|
||||
|
||||
def stage(directory, name, budget):
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
if name in status['stages'] and status['stages'][name].get('result') in ('success', 'failure'):
|
||||
return status['stages'][name]['result'] == 'success'
|
||||
started = time.time()
|
||||
status['stages'][name] = {'result': 'running', 'started': started}
|
||||
atomic_json(directory / 'status.json', status)
|
||||
script = directory / 'source/.gitea/workflows/deploy-stage.sh'
|
||||
with (directory / f'{name}.log').open('a') as log:
|
||||
# timeout kills the whole stage process group, including children, before recovery.
|
||||
result = subprocess.run( # noqa: S603, S607
|
||||
[
|
||||
shutil.which('timeout') or '/usr/bin/timeout',
|
||||
'--signal=TERM',
|
||||
'--kill-after=30s',
|
||||
str(budget),
|
||||
'bash',
|
||||
str(script),
|
||||
name,
|
||||
],
|
||||
env=environment(directory),
|
||||
stdout=log,
|
||||
stderr=subprocess.STDOUT,
|
||||
check=False,
|
||||
).returncode
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
status['stages'][name].update(
|
||||
result='success' if result == 0 else 'failure', exit_code=result, seconds=round(time.time() - started)
|
||||
)
|
||||
atomic_json(directory / 'status.json', status)
|
||||
return result == 0
|
||||
|
||||
|
||||
def make_plan(directory):
|
||||
source = directory / 'source'
|
||||
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
|
||||
request = json.loads((directory / 'request.json').read_text())
|
||||
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
|
||||
# Helm 4 lists every release status by default and removed the --all flag.
|
||||
helm = json.loads(command('helm', 'list', '-A', '-o', 'json'))
|
||||
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
|
||||
if request['refresh_images']:
|
||||
plan['selected']['compose'] = plan['active']['compose']
|
||||
atomic_json(directory / 'plan.json', plan)
|
||||
if previous:
|
||||
atomic_json(directory / 'previous.json', previous)
|
||||
# Local config is deliberately separate from the immutable Git source.
|
||||
return plan
|
||||
|
||||
|
||||
def finish_success(directory, plan):
|
||||
# Repeating finalization after a crash is safe while holding deploy.lock.
|
||||
plan['run_id'] = directory.name
|
||||
path = directory / 'compose-images.json'
|
||||
previous = directory / 'previous.json'
|
||||
plan['compose-images'] = (
|
||||
json.loads(path.read_text())
|
||||
if path.exists()
|
||||
else json.loads(previous.read_text()).get('compose-images', {})
|
||||
if previous.exists()
|
||||
else {}
|
||||
)
|
||||
atomic_json(STATE / 'last-success.json', plan)
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
status['state'] = 'success'
|
||||
atomic_json(directory / 'status.json', status)
|
||||
try:
|
||||
retain_completed(directory)
|
||||
except (OSError, subprocess.CalledProcessError) as error:
|
||||
print(f'Retention deferred: {error}', flush=True)
|
||||
|
||||
|
||||
def recover(directory, retry=False):
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
if status['state'] in ('success', 'planned'):
|
||||
return
|
||||
completed = ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke')
|
||||
if all(status['stages'].get(name, {}).get('result') == 'success' for name in completed):
|
||||
finish_success(directory, json.loads((directory / 'plan.json').read_text()))
|
||||
return
|
||||
if retry:
|
||||
for name in ('verify-k8s', 'smoke'):
|
||||
if status['stages'].get(name, {}).get('result') == 'failure':
|
||||
del status['stages'][name]
|
||||
atomic_json(directory / 'status.json', status)
|
||||
snapshot = directory / 'snapshot/current'
|
||||
if snapshot.exists():
|
||||
stage(directory, 'verify-k8s', 7200)
|
||||
stage(directory, 'smoke', 600)
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
status['state'] = 'failure'
|
||||
atomic_json(directory / 'status.json', status)
|
||||
|
||||
|
||||
def execute(run_id):
|
||||
directory = run_directory(run_id)
|
||||
with lock('deploy.lock'):
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
if status['state'] != 'queued':
|
||||
return
|
||||
# A crashed predecessor must be recovered before another apply begins.
|
||||
for other in (STATE / 'runs').iterdir():
|
||||
if (
|
||||
other != directory
|
||||
and (other / 'status.json').exists()
|
||||
and json.loads((other / 'status.json').read_text())['state'] == 'running'
|
||||
):
|
||||
raise ValueError(f'Interrupted deploy {other.name}; run recover first')
|
||||
status['state'] = 'running'
|
||||
atomic_json(directory / 'status.json', status)
|
||||
phase = 'plan'
|
||||
try:
|
||||
plan = make_plan(directory)
|
||||
print(
|
||||
json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}),
|
||||
flush=True,
|
||||
)
|
||||
phase = 'doctor'
|
||||
if not stage(directory, 'doctor', 600):
|
||||
raise RuntimeError('Preflight failed')
|
||||
phase = 'validate'
|
||||
if not stage(directory, 'validate', 1200):
|
||||
raise RuntimeError('Validation failed')
|
||||
if json.loads((directory / 'request.json').read_text())['mode'] == 'plan':
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
status['state'] = 'planned'
|
||||
atomic_json(directory / 'status.json', status)
|
||||
return
|
||||
# Budget includes both rollout checks and rollback waves, plus API overhead.
|
||||
phase = 'Recovery budget'
|
||||
count = int(
|
||||
command(
|
||||
'bash',
|
||||
str(directory / 'source/.gitea/workflows/deploy-stage.sh'),
|
||||
'workload-count',
|
||||
env=environment(directory),
|
||||
)
|
||||
)
|
||||
verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120)
|
||||
if verify_budget > 7200:
|
||||
raise ValueError('More than two hours of recovery required; split this deploy')
|
||||
phase = 'apply-k8s'
|
||||
k8s_ok = stage(directory, 'apply-k8s', 2700)
|
||||
phase = 'apply-compose'
|
||||
compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False
|
||||
phase = 'verify-k8s'
|
||||
verify_ok = stage(directory, 'verify-k8s', verify_budget)
|
||||
phase = 'smoke'
|
||||
smoke_ok = stage(directory, 'smoke', 600)
|
||||
if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)):
|
||||
raise RuntimeError('Deploy failed; inspect stage logs and recovery report')
|
||||
phase = 'Save the successful baseline'
|
||||
finish_success(directory, plan)
|
||||
except Exception as error:
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
status['failure_stage'] = next(
|
||||
(name for name, result in status['stages'].items() if result.get('result') == 'failure'), phase
|
||||
)
|
||||
atomic_json(directory / 'status.json', status)
|
||||
with (directory / 'controller.log').open('a') as stream:
|
||||
stream.write(f'{error}\n')
|
||||
recover(directory)
|
||||
raise
|
||||
|
||||
|
||||
def retain_completed(current):
|
||||
finished = []
|
||||
for directory in (STATE / 'runs').iterdir():
|
||||
status_file = directory / 'status.json'
|
||||
if status_file.exists() and json.loads(status_file.read_text())['state'] in ('success', 'planned'):
|
||||
finished.append(directory)
|
||||
for directory in sorted(finished, key=lambda p: p.stat().st_mtime, reverse=True)[20:]:
|
||||
if directory == current:
|
||||
continue
|
||||
command('git', '-C', str(CONFIG_REPO), 'worktree', 'remove', '--force', str(directory / 'source'))
|
||||
shutil.rmtree(directory)
|
||||
|
||||
|
||||
def follow(run_id, phase):
|
||||
directory = run_directory(run_id)
|
||||
groups = {
|
||||
'apply': ('doctor', 'validate', 'apply-k8s', 'apply-compose'),
|
||||
'verify': ('verify-k8s',),
|
||||
'smoke': ('smoke',),
|
||||
}
|
||||
names = groups[phase]
|
||||
offsets = {}
|
||||
while True:
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
for name in (*names, 'controller'):
|
||||
path = directory / f'{name}.log'
|
||||
if path.exists():
|
||||
with path.open() as stream:
|
||||
stream.seek(offsets.get(name, 0))
|
||||
content = stream.read()
|
||||
if content:
|
||||
print(content, end='', flush=True)
|
||||
offsets[name] = stream.tell()
|
||||
stages = status['stages']
|
||||
if all(stages.get(name, {}).get('result') in ('success', 'failure') for name in names):
|
||||
return all(stages[name]['result'] == 'success' for name in names)
|
||||
if status['state'] in ('success', 'failure', 'planned'):
|
||||
return status['state'] in ('success', 'planned')
|
||||
time.sleep(3)
|
||||
|
||||
|
||||
def summary(run_id):
|
||||
directory = run_directory(run_id)
|
||||
request = json.loads((directory / 'request.json').read_text())
|
||||
release = request['release']
|
||||
plan_file = directory / 'plan.json'
|
||||
lines = [
|
||||
f'## Deploy `{release["sha"]}`',
|
||||
'',
|
||||
f'- Mode: `{request["mode"]}`',
|
||||
f'- Refresh third-party images: `{request["refresh_images"]}`',
|
||||
]
|
||||
status = json.loads((directory / 'status.json').read_text())
|
||||
if status.get('failure_stage'):
|
||||
lines.append(f'- Failed stage: **{status["failure_stage"]}**')
|
||||
lines.extend(
|
||||
[
|
||||
'',
|
||||
f'- Observed run state: **{status["state"]}**',
|
||||
'',
|
||||
'### Stage results',
|
||||
'| Stage | Result | Exit code |',
|
||||
'| --- | --- | --- |',
|
||||
]
|
||||
)
|
||||
for name in ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke'):
|
||||
stage_result = status['stages'].get(name, {})
|
||||
lines.append(f'| {name} | {stage_result.get("result", "not started")} | {stage_result.get("exit_code", "—")} |')
|
||||
lines.extend(['', '### Apply and Helm recovery results'])
|
||||
events_file = directory / 'apply-events.jsonl'
|
||||
events = []
|
||||
if events_file.exists():
|
||||
for line in events_file.read_text().splitlines():
|
||||
try:
|
||||
events.append(json.loads(line))
|
||||
except json.JSONDecodeError:
|
||||
lines.append('- An operation record is incomplete. Check the stage log.')
|
||||
latest = {(event['action'], event['target']): event['result'] for event in events}
|
||||
lines.extend(f'- `{action}` `{target}`: **{result}**' for (action, target), result in latest.items())
|
||||
if not latest:
|
||||
lines.append('- No apply results were recorded.')
|
||||
lines.append('- A completed apply does not confirm health. See verification and smoke results.')
|
||||
lines.extend(['', '### Kubernetes recovery'])
|
||||
pointer = directory / 'snapshot/current'
|
||||
failed = Path(pointer.read_text().strip()) / 'failed-workloads' if pointer.exists() else None
|
||||
if failed and failed.exists():
|
||||
contents = failed.read_text()
|
||||
counts = dict(re.findall(r'^(ROLLED_BACK|UNRECOVERED)=([0-9]+)$', contents, re.MULTILINE))
|
||||
if not contents.strip():
|
||||
lines.append('- No failed workloads were recorded. See the verification result above.')
|
||||
elif counts:
|
||||
lines.append(f'- Workloads restored: **{counts.get("ROLLED_BACK", "unknown")}**')
|
||||
lines.append(f'- Workloads that need manual recovery: **{counts.get("UNRECOVERED", "unknown")}**')
|
||||
else:
|
||||
lines.append('- Rollback has no recorded result yet. Check the verification log.')
|
||||
else:
|
||||
lines.append('- No workload rollback was recorded. This does not confirm health.')
|
||||
lines.append('- Compose requires manual recovery. Use the saved command in the apply log.')
|
||||
if not plan_file.exists():
|
||||
lines.extend(['', 'Plan was not created. Check the controller log.'])
|
||||
print('\n'.join(lines))
|
||||
return
|
||||
plan = json.loads(plan_file.read_text())
|
||||
lines.extend(['', '### Selected services'])
|
||||
count = 0
|
||||
for kind, services in plan['selected'].items():
|
||||
for service in services:
|
||||
lines.append(f'- `{kind}`: `{service}`')
|
||||
count += 1
|
||||
if not count:
|
||||
lines.append('- None')
|
||||
lines.extend(['', '### Selected Helm releases'])
|
||||
lines.extend(f'- `{release}`' for release in plan.get('helm', []))
|
||||
if not plan.get('helm'):
|
||||
lines.append('- None')
|
||||
lines.extend(['', '### Images pinned in the checked release'])
|
||||
lines.extend(f'- `{image}@{digest}`' for image, digest in sorted(release['images'].items()))
|
||||
lines.extend(['', '### Removed resources requiring manual review'])
|
||||
lines.extend(f'- `{item}`' for item in plan.get('removed', []))
|
||||
if not plan.get('removed'):
|
||||
lines.append('- None')
|
||||
print('\n'.join(lines))
|
||||
|
||||
|
||||
def main():
|
||||
os.umask(0o077)
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow', 'summary'))
|
||||
parser.add_argument('run_id')
|
||||
parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke'))
|
||||
parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply')
|
||||
args = parser.parse_args()
|
||||
directory = run_directory(args.run_id)
|
||||
if args.action == 'start':
|
||||
start(args.run_id)
|
||||
elif args.action == 'execute':
|
||||
execute(args.run_id)
|
||||
elif args.action == 'recover':
|
||||
with lock('deploy.lock'):
|
||||
recover(directory, retry=args.retry)
|
||||
elif args.action == 'status':
|
||||
print((directory / 'status.json').read_text())
|
||||
if (directory / 'plan.json').exists():
|
||||
plan = json.loads((directory / 'plan.json').read_text())
|
||||
print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2))
|
||||
elif args.action == 'summary':
|
||||
summary(args.run_id)
|
||||
elif not follow(args.run_id, args.phase):
|
||||
sys.exit(1)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
+444
-320
File diff suppressed because it is too large.
Load diff
@@ -1,124 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Calculate selected components against the last fully successful deploy."""
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import re
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def output(*args, **kwargs):
|
||||
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
|
||||
|
||||
|
||||
def tracked(repo):
|
||||
return output('git', '-C', str(repo), 'ls-files').splitlines()
|
||||
|
||||
|
||||
def helm_releases(repo):
|
||||
text = (repo / '.gitea/workflows/deploy-lib.sh').read_text()
|
||||
return [line.split('|') for line in re.findall(r'^ "([^"\n]+\|[^"\n]+)"$', text, re.MULTILINE)]
|
||||
|
||||
|
||||
def inventory(repo):
|
||||
files = tracked(repo)
|
||||
k8s = sorted(
|
||||
{f.split('/k8s/')[0] for f in files if '/k8s/' in f and (repo / f.split('/k8s/')[0] / 'k8s/active').is_file()}
|
||||
)
|
||||
compose = sorted(
|
||||
{
|
||||
str(Path(f).parent)
|
||||
for f in files
|
||||
if Path(f).name in ('compose.yaml', 'compose.yml') and (repo / Path(f).parent / 'active').is_file()
|
||||
}
|
||||
)
|
||||
return {'k8s': k8s, 'compose': compose}
|
||||
|
||||
|
||||
def file_hash(path):
|
||||
return hashlib.sha256(path.read_bytes()).hexdigest() if path.is_file() else 'missing'
|
||||
|
||||
|
||||
def make_plan(repo, config_repo, release, previous, mode, live_helm):
|
||||
active = inventory(repo)
|
||||
all_services = set(active['k8s'] + active['compose'])
|
||||
helm_inputs = {}
|
||||
helm_selected = []
|
||||
for name, chart, namespace, version, values, marker in helm_releases(repo):
|
||||
if not (repo / marker).is_file():
|
||||
continue
|
||||
value_path = repo / values if (repo / values).is_file() else config_repo / values
|
||||
if not value_path.is_file():
|
||||
raise ValueError(f'Missing Helm values: {values}')
|
||||
stamp = hashlib.sha256(f'{chart}|{version}|{file_hash(value_path)}'.encode()).hexdigest()
|
||||
helm_inputs[name] = stamp
|
||||
live = next((h for h in live_helm if h['name'] == name and h['namespace'] == namespace), None)
|
||||
if (
|
||||
mode == 'full'
|
||||
or previous is None
|
||||
or previous.get('helm_inputs', {}).get(name) != stamp
|
||||
or live is None
|
||||
or live.get('status') != 'deployed'
|
||||
or live.get('chart') != f'{chart.split("/")[-1]}-{version}'
|
||||
):
|
||||
helm_selected.append(name)
|
||||
local_inputs = {}
|
||||
for service in all_services:
|
||||
candidates = [config_repo / service / '.env']
|
||||
if service in active['compose']:
|
||||
candidates.append(config_repo / '.env')
|
||||
cfg = config_repo / service / 'config'
|
||||
if cfg.is_dir():
|
||||
candidates.extend(
|
||||
p for p in cfg.rglob('*') if p.is_file() and p.suffix in ('.yaml', '.yml', '.json', '.conf')
|
||||
)
|
||||
local_inputs[service] = hashlib.sha256(
|
||||
'\n'.join(f'{p.relative_to(config_repo)}:{file_hash(p)}' for p in sorted(candidates)).encode()
|
||||
).hexdigest()
|
||||
if previous is None:
|
||||
if mode == 'changed':
|
||||
raise ValueError('No successful baseline; run deploy in full mode first')
|
||||
changed = set(all_services)
|
||||
removed = []
|
||||
else:
|
||||
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
|
||||
changed = {path.split('/')[0] for path in paths}
|
||||
if any(path.startswith('.gitea/') for path in paths):
|
||||
changed |= all_services
|
||||
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
|
||||
for file in tracked(repo):
|
||||
service = file.split('/')[0]
|
||||
if service not in all_services or not file.endswith(('.yaml', '.yml')):
|
||||
continue
|
||||
text = (repo / file).read_text()
|
||||
if any(
|
||||
image in text and previous.get('images', {}).get(image) != digest
|
||||
for image, digest in release['images'].items()
|
||||
):
|
||||
changed.add(service)
|
||||
removed = sorted(
|
||||
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
|
||||
- all_services
|
||||
)
|
||||
removed += [path for path in paths if '/k8s/' in path and not (repo / path).exists()]
|
||||
if mode == 'full':
|
||||
changed = set(all_services)
|
||||
dependencies = json.loads((repo / '.gitea/deploy-dependencies.json').read_text())
|
||||
while True:
|
||||
expanded = changed | {dependent for service in changed for dependent in dependencies.get(service, [])}
|
||||
if expanded == changed:
|
||||
break
|
||||
changed = expanded
|
||||
return {
|
||||
'version': 1,
|
||||
'sha': release['sha'],
|
||||
'images': release['images'],
|
||||
'active': active,
|
||||
'selected': {kind: sorted(set(services) & changed) for kind, services in active.items()},
|
||||
'helm': helm_selected,
|
||||
'helm_inputs': helm_inputs,
|
||||
'local_inputs': local_inputs,
|
||||
'removed': sorted(set(removed)),
|
||||
'full_smoke': mode == 'full' or 'traefik' in changed,
|
||||
}
|
||||
@@ -1,10 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
source "${REPO:?}/.gitea/workflows/deploy-lib.sh"
|
||||
case "${1:?stage required}" in
|
||||
workload-count)
|
||||
select_manifests >/dev/null
|
||||
selected_workload_refs | sort -u | wc -l
|
||||
;;
|
||||
*) run_stage "$1" ;;
|
||||
esac
|
||||
+160
-103
@@ -1,140 +1,197 @@
|
||||
name: deploy
|
||||
|
||||
on:
|
||||
# Deploy only what CI already validated. workflow_run is used instead of
|
||||
# workflow_dispatch so a red lint/validate run can never reach the cluster.
|
||||
workflow_run:
|
||||
workflows: [ci]
|
||||
branches: [main]
|
||||
types: [completed]
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
deploy_ref:
|
||||
description: "Commit already checked by successful main CI (main or SHA)"
|
||||
default: main
|
||||
required: true
|
||||
deploy_mode:
|
||||
description: "First deploy requires full; plan changes no production resources"
|
||||
type: choice
|
||||
options: [changed, full, plan]
|
||||
default: changed
|
||||
refresh_images:
|
||||
description: "Explicitly refresh mutable third-party Compose tags"
|
||||
type: boolean
|
||||
default: false
|
||||
|
||||
# The deploy jobs read the tree, then reach the cluster over SSH with the
|
||||
# deploy key. The Actions token itself is not part of that path, so it gets
|
||||
# read-only contents and no more.
|
||||
permissions:
|
||||
contents: read
|
||||
actions: read
|
||||
|
||||
concurrency:
|
||||
group: deploy-main
|
||||
# Queue instead of cancelling. Cancelling a run kills the apply job mid-loop and
|
||||
# takes the verify job down with it, so a superseded deploy would leave the
|
||||
# cluster half-applied and unchecked — the exact failure the verify job exists
|
||||
# to catch. kubectl apply and docker compose up are both idempotent, so letting
|
||||
# the older run finish and then deploying the newer commit costs little.
|
||||
cancel-in-progress: false
|
||||
|
||||
env:
|
||||
DEPLOY_HOST: ${{ vars.DEPLOY_HOST || secrets.DEPLOY_HOST }}
|
||||
DEPLOY_PORT: ${{ vars.DEPLOY_PORT || secrets.DEPLOY_PORT }}
|
||||
DEPLOY_USER: ${{ vars.DEPLOY_USER || secrets.DEPLOY_USER }}
|
||||
DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
|
||||
DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
|
||||
DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
|
||||
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
|
||||
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
|
||||
DEPLOY_KNOWN_HOSTS: ${{ vars.DEPLOY_KNOWN_HOSTS }}
|
||||
DEPLOY_RUN_ID: ${{ github.run_id }}-${{ github.run_attempt || 1 }}
|
||||
DEPLOY_MODE: ${{ inputs.deploy_mode || 'changed' }}
|
||||
REFRESH_IMAGES: ${{ inputs.refresh_images && 'true' || 'false' }}
|
||||
APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
|
||||
# workflow_run's own GITHUB_SHA points at the branch head, not at the commit the
|
||||
# finished ci run checked. Pin the exact validated commit instead, so a push
|
||||
# landing mid-deploy cannot make the workstation deploy something else. Also
|
||||
# what the verify job checks the snapshot against. Empty for workflow_dispatch,
|
||||
# which falls back to the current origin/main.
|
||||
DEPLOY_SHA: ${{ github.event.workflow_run.head_sha }}
|
||||
|
||||
jobs:
|
||||
gate:
|
||||
preflight:
|
||||
# Autodeploy defaults to OFF: pushes deploy only when the AUTODEPLOY repo
|
||||
# variable is set to 'true' (Settings -> Actions -> Variables). A manual
|
||||
# Run workflow always bypasses the switch: dispatching it is the explicit
|
||||
# intent to deploy.
|
||||
if: >-
|
||||
github.ref == 'refs/heads/main' &&
|
||||
(vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') &&
|
||||
(github.event_name != 'workflow_run' ||
|
||||
(github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.head_branch == 'main'))
|
||||
runs-on: homelab
|
||||
(github.event.workflow_run.conclusion == 'success' &&
|
||||
github.event.workflow_run.head_branch == 'main'))
|
||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
||||
timeout-minutes: 10
|
||||
outputs:
|
||||
sha: ${{ steps.release.outputs.sha }}
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
with:
|
||||
fetch-depth: 0
|
||||
- name: Check successful CI and download the exact commit release
|
||||
id: release
|
||||
env:
|
||||
GITEA_TOKEN: ${{ github.token }}
|
||||
DEPLOY_REF: ${{ inputs.deploy_ref || 'main' }}
|
||||
EVENT_SHA: ${{ github.event.workflow_run.head_sha }}
|
||||
run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA"
|
||||
- name: Submit durable deploy to workstation
|
||||
run: bash .gitea/workflows/ssh-run.sh start
|
||||
- name: Write the request result
|
||||
if: always()
|
||||
env:
|
||||
REQUEST_RESULT: ${{ job.status }}
|
||||
CHECKED_SHA: ${{ steps.release.outputs.sha }}
|
||||
run: |
|
||||
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
printf '## Deploy request\n\n- Result: **%s**\n- Checked commit: %s\n- Mode: %s\n' "$REQUEST_RESULT" "${CHECKED_SHA:-not checked}" "$DEPLOY_MODE" >>"$GITHUB_STEP_SUMMARY"
|
||||
if [ "$REQUEST_RESULT" != success ]; then
|
||||
echo 'Open the failed step log. If SSH submission failed, check the remote controller state.' >>"$GITHUB_STEP_SUMMARY"
|
||||
fi
|
||||
fi
|
||||
|
||||
apply:
|
||||
needs: [gate]
|
||||
runs-on: homelab
|
||||
timeout-minutes: 120
|
||||
- name: Fetch and reset workstation
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
./.gitea/workflows/ssh-run.sh preflight
|
||||
|
||||
validate:
|
||||
needs: [preflight]
|
||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
||||
timeout-minutes: 20
|
||||
steps:
|
||||
- name: Checkout checked commit
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
with:
|
||||
ref: ${{ needs.gate.outputs.sha }}
|
||||
- name: Follow validation and sequential Kubernetes / Compose apply
|
||||
run: bash .gitea/workflows/ssh-run.sh apply
|
||||
- name: Write the deploy result
|
||||
if: always()
|
||||
run: |
|
||||
if [ -f .gitea/workflows/ssh-run.sh ]; then
|
||||
bash .gitea/workflows/ssh-run.sh summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
|
||||
fi
|
||||
|
||||
verify:
|
||||
needs: [gate, apply]
|
||||
if: always() && needs.gate.result == 'success'
|
||||
runs-on: homelab
|
||||
timeout-minutes: 130
|
||||
- name: Dry-run manifests and check Secrets
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
./.gitea/workflows/ssh-run.sh validate
|
||||
|
||||
apply-k8s:
|
||||
needs: [validate]
|
||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
||||
# Apply only, no verification, so this is just the work itself: snapshot,
|
||||
# then sequential `helm upgrade --install --wait --rollback-on-failure --timeout 10m`, then the apply loop.
|
||||
# Verification has its own job and its own budget.
|
||||
#
|
||||
# 45 is roughly four times the measured cost of the stage, which is
|
||||
# deliberately not raised on a theory:
|
||||
#
|
||||
# helm, healthy 3 no-op upgrades ~3-5 min
|
||||
# helm, one release bad rollback-on-failure spends its 10m, ~10-15 min
|
||||
# then rolls that one back
|
||||
# apply loop ~40 manifests, 4 of which ~1 min
|
||||
# resolve an image digest
|
||||
# restart_stale_images 7.6s to find 8 workloads, ~0.5 min
|
||||
# 9.8s to resolve their digests
|
||||
#
|
||||
# The helm figure is one release, not three: `set -e` aborts
|
||||
# upgrade_helm_releases on the first failure, so a broken release costs
|
||||
# 10m and the other two are never attempted. Multiplying 10m by three
|
||||
# overstates the worst case by 20 minutes.
|
||||
#
|
||||
# The 45 minutes this was last raised to 45 were still not enough, and the
|
||||
# job logs for those runs no longer exist, so what actually consumed the
|
||||
# budget is not known - the two measurable candidates above account for
|
||||
# ~15 of it. The unbounded `docker manifest inspect` against the registry's
|
||||
# known hang mode is now bounded inside registry_digest (25s timeout, 3
|
||||
# attempts): a dead registry fails each owned image after ~85s instead of
|
||||
# hanging the stage, and a blinking one is retried instead of failing the
|
||||
# whole apply file. Still open: make the stage announce which manifest it
|
||||
# is working on, so a killed run leaves a diagnosable last line.
|
||||
timeout-minutes: 45
|
||||
steps:
|
||||
- name: Checkout checked commit
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
with:
|
||||
ref: ${{ needs.gate.outputs.sha }}
|
||||
- name: Follow workload verification and recovery
|
||||
run: bash .gitea/workflows/ssh-run.sh verify
|
||||
- name: Write the deploy result
|
||||
if: always()
|
||||
run: |
|
||||
if [ -f .gitea/workflows/ssh-run.sh ]; then
|
||||
bash .gitea/workflows/ssh-run.sh summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
|
||||
fi
|
||||
|
||||
- name: Apply Kubernetes manifests
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
./.gitea/workflows/ssh-run.sh apply-k8s
|
||||
|
||||
apply-compose:
|
||||
needs: [validate]
|
||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
||||
timeout-minutes: 30
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Redeploy docker compose stacks
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
./.gitea/workflows/ssh-run.sh apply-compose
|
||||
|
||||
# Watches the workloads this deploy changed and rolls back the ones that never
|
||||
# became healthy. Runs even when the apply jobs failed, timed out or were
|
||||
# cancelled — that is the whole point of splitting it out. `always()` is what
|
||||
# lets it start after a failed dependency; the needs on apply-compose are a
|
||||
# barrier, so verification begins only once both applies are done.
|
||||
verify-k8s:
|
||||
needs: [apply-k8s, apply-compose]
|
||||
if: >-
|
||||
always() &&
|
||||
needs.apply-k8s.result != 'skipped' &&
|
||||
needs.apply-compose.result != 'skipped'
|
||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
||||
# Not raised, because the arithmetic does not close.
|
||||
#
|
||||
# 32 workloads are under management and the wave width is 8, so the verify
|
||||
# itself is 4 waves of ROLLOUT_TIMEOUT (300s) = 20 minutes worst case, when
|
||||
# every rollout times out rather than converging. That is already 20 of 30.
|
||||
#
|
||||
# The other 10 would have to absorb rollback, and rollback_workloads is a
|
||||
# serial `while read` loop at 300s per failed workload. 10 minutes buys two.
|
||||
# Any larger number is buying a bigger multiple of an unbounded term rather
|
||||
# than covering a known cost: 60 minutes buys eight, and 60 minutes is
|
||||
# therefore not a bound, it is a guess with two digits.
|
||||
#
|
||||
# The number becomes derivable the moment rollback uses the same wave width
|
||||
# as the verify: 32 failures then cost 4 waves = 20 minutes instead of 160,
|
||||
# and 45 covers verify plus rollback at full width. That change is to the
|
||||
# recovery path and is not folded into a timeout edit.
|
||||
timeout-minutes: 30
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
|
||||
- name: Verify workloads and roll back on failure
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
./.gitea/workflows/ssh-run.sh verify-k8s
|
||||
|
||||
# Asks the public route of every active service whether it is actually
|
||||
# serving, which the rollout check above structurally cannot: a pod can
|
||||
# converge and still be crash-looping, or be listening on a port no Service
|
||||
# points at, or answer 500.
|
||||
#
|
||||
# `always()` for the same reason verify-k8s has it, and it runs after that job
|
||||
# specifically because a rollback is when a route most needs re-checking. The
|
||||
# needs is a barrier, not a filter: whether verify-k8s passed, failed or was
|
||||
# cancelled, the probes are what say whether the cluster is serving, and
|
||||
# suppressing them on a rollback would hide the one run where the answer
|
||||
# matters most.
|
||||
smoke:
|
||||
needs: [gate, verify]
|
||||
if: always() && needs.gate.result == 'success'
|
||||
runs-on: homelab
|
||||
timeout-minutes: 15
|
||||
needs: [verify-k8s]
|
||||
if: always() && needs.verify-k8s.result != 'skipped'
|
||||
runs-on: [self-hosted, linux, arch, homelab, prod]
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout checked commit
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
with:
|
||||
ref: ${{ needs.gate.outputs.sha }}
|
||||
- name: Follow public route checks
|
||||
run: bash .gitea/workflows/ssh-run.sh smoke
|
||||
- name: Write the deploy result
|
||||
if: always()
|
||||
|
||||
- name: Probe the public route of every active service
|
||||
shell: bash
|
||||
run: |
|
||||
if [ -f .gitea/workflows/ssh-run.sh ]; then
|
||||
bash .gitea/workflows/ssh-run.sh summary
|
||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
|
||||
fi
|
||||
set -euo pipefail
|
||||
./.gitea/workflows/ssh-run.sh smoke
|
||||
@@ -13,14 +13,9 @@ here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
# shellcheck source=tool-versions.env
|
||||
. "$here/tool-versions.env"
|
||||
|
||||
TOOLS_DIR="${TOOLS_DIR:-${XDG_CACHE_HOME:-$HOME/.cache}/homelab-ci}"
|
||||
TOOLS_DIR="${TOOLS_DIR:-${RUNNER_TEMP:-/tmp}/homelab-tools}"
|
||||
BIN_DIR="$TOOLS_DIR/bin"
|
||||
mkdir -p "$BIN_DIR"
|
||||
# A runner may accept overlapping workflows even though each workflow is sequential.
|
||||
exec 9>"$TOOLS_DIR/install.lock"
|
||||
flock -w 300 9
|
||||
export UV_TOOL_DIR="$TOOLS_DIR/uv-tools"
|
||||
export UV_CACHE_DIR="$TOOLS_DIR/uv-cache"
|
||||
# The just-installed tools must resolve inside this script too: callers only
|
||||
# prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would
|
||||
# miss the binary install_uv just placed (exit 127 on a clean runner).
|
||||
@@ -53,7 +48,7 @@ esac
|
||||
fetch() {
|
||||
# fetch <url> <dest>
|
||||
if command -v curl >/dev/null 2>&1; then
|
||||
curl -sSLf --connect-timeout 15 --max-time 120 --retry 3 -o "$2" "$1"
|
||||
curl -sSLf --retry 3 -o "$2" "$1"
|
||||
elif command -v wget >/dev/null 2>&1; then
|
||||
wget -q -O "$2" "$1"
|
||||
else
|
||||
@@ -93,13 +88,10 @@ installed_version() {
|
||||
|
||||
# at_version <command> <expected>
|
||||
at_version() {
|
||||
local version expected="${2#v}"
|
||||
version="$(installed_version "$1")"
|
||||
if [[ "$version" =~ (^|[^0-9.])v?([0-9]+(\.[0-9]+)+) ]]; then
|
||||
[ "${BASH_REMATCH[2]}" = "$expected" ]
|
||||
else
|
||||
return 1
|
||||
fi
|
||||
case "$(installed_version "$1")" in
|
||||
*"$2"*) return 0 ;;
|
||||
*) return 1 ;;
|
||||
esac
|
||||
}
|
||||
|
||||
install_kubeconform() {
|
||||
@@ -128,15 +120,6 @@ install_shellcheck() {
|
||||
rm -rf "$tmp"
|
||||
}
|
||||
|
||||
install_jq() {
|
||||
if at_version jq "${JQ_VERSION}"; then
|
||||
return 0
|
||||
fi
|
||||
fetch "https://github.com/jqlang/jq/releases/download/jq-${JQ_VERSION}/jq-linux-${goarch}" \
|
||||
"$BIN_DIR/jq"
|
||||
chmod 0755 "$BIN_DIR/jq"
|
||||
}
|
||||
|
||||
install_uv() {
|
||||
if at_version uv "${UV_VERSION}"; then
|
||||
return 0
|
||||
@@ -184,7 +167,6 @@ install_pip_audit() {
|
||||
}
|
||||
|
||||
install_prettier() {
|
||||
install_node
|
||||
if at_version prettier "${PRETTIER_VERSION}"; then
|
||||
return 0
|
||||
fi
|
||||
@@ -245,41 +227,28 @@ install_actionlint() {
|
||||
rm -rf "$tmp"
|
||||
}
|
||||
|
||||
main() {
|
||||
wanted=("$@")
|
||||
if [ "${#wanted[@]}" -eq 0 ]; then
|
||||
wanted=(node jq kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
|
||||
fi
|
||||
wanted=("$@")
|
||||
if [ "${#wanted[@]}" -eq 0 ]; then
|
||||
wanted=(kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
|
||||
fi
|
||||
|
||||
for tool in "${wanted[@]}"; do
|
||||
case "$tool" in
|
||||
kubeconform) install_kubeconform ;;
|
||||
shellcheck) install_shellcheck ;;
|
||||
jq) install_jq ;;
|
||||
actionlint) install_actionlint ;;
|
||||
prettier) install_prettier ;;
|
||||
ruff) install_ruff ;;
|
||||
yamllint) install_yamllint ;;
|
||||
pip-audit) install_pip_audit ;;
|
||||
hadolint) install_hadolint ;;
|
||||
node) install_node ;;
|
||||
uv) install_uv ;;
|
||||
*)
|
||||
echo "install-ci-tools: unknown tool: $tool" >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
done
|
||||
for tool in "${wanted[@]}"; do
|
||||
case "$tool" in
|
||||
kubeconform) install_kubeconform ;;
|
||||
shellcheck) install_shellcheck ;;
|
||||
actionlint) install_actionlint ;;
|
||||
prettier) install_prettier ;;
|
||||
ruff) install_ruff ;;
|
||||
yamllint) install_yamllint ;;
|
||||
pip-audit) install_pip_audit ;;
|
||||
hadolint) install_hadolint ;;
|
||||
node) install_node ;;
|
||||
uv) install_uv ;;
|
||||
*)
|
||||
echo "install-ci-tools: unknown tool: $tool" >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
done
|
||||
|
||||
for old in "$BIN_DIR"/node-* "$BIN_DIR"/prettier-*; do
|
||||
[ -d "$old" ] || continue
|
||||
case "$(basename "$old")" in
|
||||
"node-$NODE_VERSION"|"prettier-$PRETTIER_VERSION") ;;
|
||||
*) rm -rf "$old" ;;
|
||||
esac
|
||||
done
|
||||
if [ -x "$BIN_DIR/uv" ]; then "$BIN_DIR/uv" cache prune >/dev/null; fi
|
||||
printf '%s\n' "$BIN_DIR"
|
||||
}
|
||||
|
||||
if [ "${BASH_SOURCE[0]}" = "$0" ]; then main "$@"; fi
|
||||
printf '%s\n' "$BIN_DIR"
|
||||
@@ -1,516 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""CI release artifacts and the SHA-specific Gitea deployment gate (stdlib only)."""
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import io
|
||||
import itertools
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import urllib.error
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
|
||||
SHA = re.compile(r'[0-9a-f]{40}')
|
||||
DIGEST = re.compile(r'sha256:[0-9a-f]{64}')
|
||||
IMAGES = {
|
||||
'error-pages': ('errorpages', 'errorpages/Dockerfile'),
|
||||
'forust-homepage': ('homepages', 'homepages/Dockerfile.forust'),
|
||||
'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'),
|
||||
}
|
||||
|
||||
|
||||
def command(*args, **kwargs):
|
||||
"""Arguments are passed directly to the executable, never to a shell."""
|
||||
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
|
||||
|
||||
|
||||
def validate_release(data, sha=None):
|
||||
if data.get('version') != 1 or not SHA.fullmatch(data.get('sha', '')):
|
||||
raise ValueError('Invalid release version or SHA')
|
||||
if sha is not None and data['sha'] != sha:
|
||||
raise ValueError('Release SHA does not match the checked CI commit')
|
||||
expected = {f'gcr.forust.xyz/forust/{name}' for name in IMAGES}
|
||||
if set(data.get('images', {})) != expected:
|
||||
raise ValueError('Release must contain all owned images')
|
||||
if not all(DIGEST.fullmatch(value) for value in data['images'].values()):
|
||||
raise ValueError('Release has an invalid image digest')
|
||||
if set(data.get('inputs', {})) != expected or not all(
|
||||
re.fullmatch(r'[0-9a-f]{64}', value) for value in data['inputs'].values()
|
||||
):
|
||||
raise ValueError('Release has invalid build input fingerprints')
|
||||
return data
|
||||
|
||||
|
||||
class NoRedirect(urllib.request.HTTPRedirectHandler):
|
||||
def redirect_request(self, _req, _fp, _code, _msg, _headers, _newurl):
|
||||
return None
|
||||
|
||||
|
||||
class Gitea:
|
||||
def __init__(self):
|
||||
self.origin = os.environ['GITHUB_SERVER_URL'].rstrip('/')
|
||||
if urllib.parse.urlsplit(self.origin).scheme != 'https':
|
||||
raise ValueError('Gitea API must use HTTPS')
|
||||
self.repository = os.environ['GITHUB_REPOSITORY']
|
||||
if not re.fullmatch(r'[\w.-]+/[\w.-]+', self.repository):
|
||||
raise ValueError('Invalid Gitea repository')
|
||||
self.token = os.environ['GITEA_TOKEN']
|
||||
self.base = f'{self.origin}/api/v1/repos/{self.repository}'
|
||||
|
||||
def request(self, url, *, archive=False):
|
||||
if not url.startswith(self.base + '/'):
|
||||
raise ValueError('Refusing to send the Actions token to another origin')
|
||||
req = urllib.request.Request(url, headers={'Authorization': f'token {self.token}'}) # noqa: S310 -- HTTPS origin validated above
|
||||
opener = urllib.request.build_opener(NoRedirect())
|
||||
try:
|
||||
response = opener.open(req, timeout=30) # noqa: S310
|
||||
except urllib.error.HTTPError as error:
|
||||
if not archive or error.code not in (301, 302, 303, 307, 308):
|
||||
raise RuntimeError(f'Gitea API returned HTTP {error.code}') from None
|
||||
target = urllib.parse.urljoin(url, error.headers['Location'])
|
||||
if urllib.parse.urlsplit(target).scheme != 'https':
|
||||
raise ValueError('Artifact redirect must use HTTPS') from None
|
||||
# Signed storage redirects must never receive the Gitea token.
|
||||
response = urllib.request.urlopen(target, timeout=30) # noqa: S310
|
||||
with response:
|
||||
payload = response.read(8 * 1024 * 1024 + 1)
|
||||
if len(payload) > 8 * 1024 * 1024:
|
||||
raise ValueError('Gitea response exceeds 8 MiB')
|
||||
return payload if archive else json.loads(payload)
|
||||
|
||||
def pages(self, path, key, **params):
|
||||
for page in range(1, 101):
|
||||
query = urllib.parse.urlencode({**params, 'page': page, 'limit': 50})
|
||||
data = self.request(f'{self.base}/{path}?{query}')
|
||||
entries = data[key]
|
||||
yield from entries
|
||||
if len(entries) < 50:
|
||||
return
|
||||
raise RuntimeError('Gitea pagination limit exceeded')
|
||||
|
||||
def successful_runs(self, sha=None):
|
||||
params = {'branch': 'main', 'status': 'success', 'exclude_pull_requests': 'true'}
|
||||
if sha:
|
||||
params['head_sha'] = sha
|
||||
for run in self.pages('actions/workflows/ci.yaml/runs', 'workflow_runs', **params):
|
||||
if (
|
||||
run.get('status') == 'completed'
|
||||
and run.get('conclusion') == 'success'
|
||||
and run.get('head_branch') == 'main'
|
||||
and run.get('event') in ('push', 'workflow_dispatch')
|
||||
and (run.get('repository') or {}).get('full_name') == self.repository
|
||||
and (run.get('head_repository') or run.get('repository') or {}).get('full_name') == self.repository
|
||||
and (sha is None or run.get('head_sha') == sha)
|
||||
):
|
||||
yield run
|
||||
|
||||
def release(self, run):
|
||||
sha = run['head_sha']
|
||||
jobs = list(self.pages(f'actions/runs/{run["id"]}/jobs', 'jobs'))
|
||||
# A green workflow with a skipped build must not authorize a deploy.
|
||||
if not any(job.get('name') == 'build' and job.get('conclusion') == 'success' for job in jobs):
|
||||
raise ValueError('CI build job did not succeed')
|
||||
artifacts = self.request(f'{self.base}/actions/runs/{run["id"]}/artifacts')['artifacts']
|
||||
matching = [a for a in artifacts if a['name'] == f'release-{sha}' and not a.get('expired')]
|
||||
if len(matching) != 1:
|
||||
raise ValueError('CI release artifact is missing, expired or ambiguous; rerun CI')
|
||||
blob = self.request(f'{self.base}/actions/artifacts/{matching[0]["id"]}/zip', archive=True)
|
||||
with zipfile.ZipFile(io.BytesIO(blob)) as archive:
|
||||
files = [entry for entry in archive.infolist() if not entry.is_dir()]
|
||||
if len(files) != 1 or files[0].filename != 'release.json' or files[0].file_size > 256 * 1024:
|
||||
raise ValueError('Unexpected release archive contents')
|
||||
return validate_release(json.loads(archive.read(files[0])), sha)
|
||||
|
||||
|
||||
def fingerprint(context, dockerfile):
|
||||
tree = command('git', 'ls-tree', '-r', 'HEAD', '--', context, dockerfile, '.gitea/workflows/release.py')
|
||||
return hashlib.sha256(tree.encode()).hexdigest()
|
||||
|
||||
|
||||
def gate(output, requested_ref, event_sha):
|
||||
command('git', 'fetch', '--quiet', 'origin', 'main')
|
||||
if event_sha:
|
||||
if not SHA.fullmatch(event_sha):
|
||||
raise ValueError('Invalid workflow_run SHA')
|
||||
sha = event_sha
|
||||
else:
|
||||
if requested_ref == 'main':
|
||||
requested_ref = 'origin/main'
|
||||
sha = command('git', 'rev-parse', '--verify', '--end-of-options', f'{requested_ref}^{{commit}}')
|
||||
if not SHA.fullmatch(sha):
|
||||
raise ValueError('Invalid deploy SHA')
|
||||
command('git', 'merge-base', '--is-ancestor', sha, 'origin/main')
|
||||
api = Gitea()
|
||||
runs = list(api.successful_runs(sha))
|
||||
if not runs:
|
||||
raise ValueError(f'No successful main CI for {sha}; run CI before deploying')
|
||||
release = api.release(max(runs, key=lambda run: run['id']))
|
||||
output.write_text(json.dumps(release, indent=2) + '\n')
|
||||
if os.environ.get('GITHUB_OUTPUT'):
|
||||
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
|
||||
stream.write(f'sha={sha}\n')
|
||||
print(f'CI gate accepted {sha}')
|
||||
|
||||
|
||||
def prepare_images(output):
|
||||
sha = command('git', 'rev-parse', 'HEAD')
|
||||
if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
|
||||
raise ValueError('Build checkout does not match GITHUB_SHA')
|
||||
api = Gitea()
|
||||
previous = None
|
||||
for run in sorted(itertools.islice(api.successful_runs(), 50), key=lambda item: item['id'], reverse=True):
|
||||
if str(run['id']) == os.environ.get('GITHUB_RUN_ID'):
|
||||
continue
|
||||
try:
|
||||
previous = api.release(run)
|
||||
break
|
||||
except ValueError:
|
||||
# Expired artifacts only cost a rebuild; mutable tags are never a fallback.
|
||||
continue
|
||||
targets = []
|
||||
for name, (context, dockerfile) in IMAGES.items():
|
||||
image = f'gcr.forust.xyz/forust/{name}'
|
||||
inputs = fingerprint(context, dockerfile)
|
||||
old_digest = (previous or {}).get('images', {}).get(image)
|
||||
targets.append(
|
||||
{
|
||||
'name': name,
|
||||
'image': image,
|
||||
'context': context,
|
||||
'dockerfile': dockerfile,
|
||||
'inputs': inputs,
|
||||
'reuse_digest': old_digest if (previous or {}).get('inputs', {}).get(image) == inputs else None,
|
||||
}
|
||||
)
|
||||
output.write_text(json.dumps({'sha': sha, 'targets': targets}, indent=2) + '\n')
|
||||
if os.environ.get('GITHUB_OUTPUT'):
|
||||
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
|
||||
stream.write('matrix=' + json.dumps({'include': targets}, separators=(',', ':')) + '\n')
|
||||
print(f'Prepared {len(targets)} image jobs; {sum(t["reuse_digest"] is None for t in targets)} require builds')
|
||||
|
||||
|
||||
def checked_plan(path):
|
||||
data = json.loads(path.read_text())
|
||||
sha = command('git', 'rev-parse', 'HEAD')
|
||||
if data.get('sha') != sha or sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
|
||||
raise ValueError('Image plan does not match the checked source commit')
|
||||
targets = data.get('targets', [])
|
||||
if sorted(t['name'] for t in targets) != sorted(IMAGES):
|
||||
raise ValueError('Image plan must contain each owned image once')
|
||||
for target in targets:
|
||||
name = target['name']
|
||||
context, dockerfile = IMAGES[name]
|
||||
if (target['context'], target['dockerfile'], target['image']) != (
|
||||
context,
|
||||
dockerfile,
|
||||
f'gcr.forust.xyz/forust/{name}',
|
||||
) or target['inputs'] != fingerprint(context, dockerfile):
|
||||
raise ValueError('Image plan has invalid build inputs')
|
||||
if target['reuse_digest'] is not None and not DIGEST.fullmatch(target['reuse_digest']):
|
||||
raise ValueError('Image plan has an invalid reuse digest')
|
||||
return data
|
||||
|
||||
|
||||
def build_images(output, report, name, plan):
|
||||
data = checked_plan(plan)
|
||||
sha = data['sha']
|
||||
target = next(t for t in data['targets'] if t['name'] == name)
|
||||
context, dockerfile = IMAGES[name]
|
||||
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
|
||||
builder_config = Path.home() / '.cache/homelab-ci/buildx'
|
||||
builder_config.mkdir(parents=True, exist_ok=True)
|
||||
env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)}
|
||||
try:
|
||||
report['phase'] = 'Registry login'
|
||||
subprocess.run( # noqa: S603, S607
|
||||
[
|
||||
shutil.which('docker') or '/usr/bin/docker',
|
||||
'login',
|
||||
'gcr.forust.xyz',
|
||||
'-u',
|
||||
os.environ['REGISTRY_USERNAME'],
|
||||
'--password-stdin',
|
||||
],
|
||||
input=os.environ['REGISTRY_PASSWORD'],
|
||||
text=True,
|
||||
check=True,
|
||||
env=env,
|
||||
)
|
||||
report['phase'] = 'Prepare the builder'
|
||||
builder = 'homelab-ci'
|
||||
versions = dict(
|
||||
re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE)
|
||||
)
|
||||
image = versions['BUILDKIT_IMAGE']
|
||||
signature = builder_config / 'homelab-ci-image'
|
||||
exists = (
|
||||
subprocess.run( # noqa: S603
|
||||
[shutil.which('docker') or '/usr/bin/docker', 'buildx', 'inspect', builder],
|
||||
capture_output=True,
|
||||
env=env,
|
||||
).returncode
|
||||
== 0
|
||||
)
|
||||
if exists and (not signature.exists() or signature.read_text().strip() != image):
|
||||
command('docker', 'buildx', 'rm', '--keep-state', builder, env=env)
|
||||
exists = False
|
||||
if not exists:
|
||||
command(
|
||||
'docker',
|
||||
'buildx',
|
||||
'create',
|
||||
'--name',
|
||||
builder,
|
||||
'--driver',
|
||||
'docker-container',
|
||||
'--driver-opt',
|
||||
f'image={image}',
|
||||
'--buildkitd-config',
|
||||
'.gitea/runner/buildkitd.toml',
|
||||
env=env,
|
||||
)
|
||||
signature.write_text(image + '\n')
|
||||
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
|
||||
report['images'] = release['images']
|
||||
report['phase'] = f'Build or reuse {name}'
|
||||
report['current'] = name
|
||||
image = f'gcr.forust.xyz/forust/{name}'
|
||||
inputs = target['inputs']
|
||||
old_digest = target['reuse_digest']
|
||||
exists = False
|
||||
if old_digest:
|
||||
exists = (
|
||||
subprocess.run( # noqa: S603, S607
|
||||
[
|
||||
shutil.which('docker') or '/usr/bin/docker',
|
||||
'buildx',
|
||||
'imagetools',
|
||||
'inspect',
|
||||
f'{image}@{old_digest}',
|
||||
],
|
||||
capture_output=True,
|
||||
env=env,
|
||||
timeout=60,
|
||||
).returncode
|
||||
== 0
|
||||
)
|
||||
if exists:
|
||||
print(f'Reuse {name}: inputs unchanged')
|
||||
digest = old_digest
|
||||
report['reused'].append(name)
|
||||
else:
|
||||
print(f'Build {name}', flush=True)
|
||||
metadata = Path(docker_config) / 'metadata.json'
|
||||
command(
|
||||
'docker',
|
||||
'buildx',
|
||||
'build',
|
||||
'--builder',
|
||||
builder,
|
||||
'--platform',
|
||||
'linux/amd64',
|
||||
'--provenance=false',
|
||||
'--cache-from',
|
||||
f'type=registry,ref={image}:buildcache',
|
||||
'--cache-to',
|
||||
f'type=registry,ref={image}:buildcache,mode=max',
|
||||
'--output',
|
||||
f'type=image,name={image},push-by-digest=true,name-canonical=true,push=true',
|
||||
'--metadata-file',
|
||||
str(metadata),
|
||||
'--file',
|
||||
dockerfile,
|
||||
context,
|
||||
env=env,
|
||||
)
|
||||
digest = json.loads(metadata.read_text())['containerimage.digest']
|
||||
report['built'].append(name)
|
||||
release['images'][image] = digest
|
||||
release['inputs'][image] = inputs
|
||||
if not DIGEST.fullmatch(digest):
|
||||
raise ValueError('Image job returned an invalid digest')
|
||||
output.write_text(json.dumps(release, indent=2) + '\n')
|
||||
report['current'] = None
|
||||
report['phase'] = 'Release file saved'
|
||||
finally:
|
||||
# Cleanup errors must neither leak credentials nor mask the original build error.
|
||||
try:
|
||||
subprocess.run( # noqa: S603
|
||||
[
|
||||
shutil.which('docker') or '/usr/bin/docker',
|
||||
'buildx',
|
||||
'prune',
|
||||
'--builder',
|
||||
'homelab-ci',
|
||||
'--force',
|
||||
'--max-used-space',
|
||||
'1gb',
|
||||
],
|
||||
env=env,
|
||||
timeout=60,
|
||||
)
|
||||
except (OSError, subprocess.TimeoutExpired):
|
||||
print('CI builder cache cleanup deferred', flush=True)
|
||||
finally:
|
||||
shutil.rmtree(docker_config)
|
||||
|
||||
|
||||
def write_summary(lines):
|
||||
path = os.environ.get('GITHUB_STEP_SUMMARY')
|
||||
if path:
|
||||
try:
|
||||
with Path(path).open('a') as stream:
|
||||
stream.write('\n'.join(lines) + '\n\n')
|
||||
except OSError:
|
||||
print('WARNING: cannot write the job summary')
|
||||
|
||||
|
||||
def check_summary():
|
||||
lines = [
|
||||
f'## {os.environ["SUMMARY_CHECK"]}',
|
||||
'',
|
||||
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
|
||||
f'- Result: **{os.environ["SUMMARY_RESULT"]}**',
|
||||
]
|
||||
if os.environ.get('SUMMARY_FAILED_STEP'):
|
||||
lines.append(f'- Failed step: {os.environ["SUMMARY_FAILED_STEP"]}')
|
||||
if os.environ['SUMMARY_RESULT'] != 'success':
|
||||
lines.append('- Open the failed step log for the error details.')
|
||||
write_summary(lines)
|
||||
|
||||
|
||||
def build(output, name, plan):
|
||||
report = {'phase': 'Check the source commit', 'current': None, 'built': [], 'reused': [], 'images': {}}
|
||||
result = 'failure'
|
||||
try:
|
||||
build_images(output, report, name, plan)
|
||||
result = 'success'
|
||||
finally:
|
||||
lines = [
|
||||
f'## Image release `{os.environ.get("GITHUB_SHA", "unknown")}`',
|
||||
'',
|
||||
f'- Result: **{result}**',
|
||||
f'- Last stage: {report["phase"]}',
|
||||
]
|
||||
if result == 'failure':
|
||||
lines.append('- No release from this build can be deployed. Open the failed step log.')
|
||||
if report['current']:
|
||||
lines.append(f'- Image at the failure: `{report["current"]}`')
|
||||
for title, key in (('Built', 'built'), ('Reused from successful CI', 'reused')):
|
||||
lines.extend(['', f'### {title}'])
|
||||
lines.extend(f'- `{name}`' for name in report[key])
|
||||
if not report[key]:
|
||||
lines.append('- None')
|
||||
lines.extend(['', '### Completed image digests'])
|
||||
lines.extend(f'- `{image}@{digest}`' for image, digest in report['images'].items())
|
||||
if not report['images']:
|
||||
lines.append('- None')
|
||||
write_summary(lines)
|
||||
|
||||
|
||||
def render(stream, destination):
|
||||
release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA'])
|
||||
image_line = re.compile(
|
||||
r"^(\s*(?:-\s*)?image:\s*)(['\"]?)(gcr\.forust\.xyz/forust/[\w.-]+)(?::[\w.-]+|@sha256:[0-9a-f]{64})\2(\s*(?:#.*)?)$"
|
||||
)
|
||||
rendered = []
|
||||
for line in stream:
|
||||
match = image_line.fullmatch(line.rstrip('\n'))
|
||||
if match:
|
||||
prefix, quote, image, tail = match.groups()
|
||||
if image not in release['images']:
|
||||
raise ValueError(f'Owned image missing from checked release: {image}')
|
||||
line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n'
|
||||
elif re.match(r'\s*(?:-\s*)?image:', line) and 'gcr.forust.xyz/forust/' in line:
|
||||
raise ValueError('Unsupported owned image syntax; refusing to apply a mutable tag')
|
||||
rendered.append(line)
|
||||
destination.writelines(rendered)
|
||||
|
||||
|
||||
def finalize_images(output, fragments, plan):
|
||||
data = checked_plan(plan)
|
||||
sha = data['sha']
|
||||
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
|
||||
for name in IMAGES:
|
||||
fragment = json.loads((fragments / f'image-{name}' / 'image.json').read_text())
|
||||
image = f'gcr.forust.xyz/forust/{name}'
|
||||
if fragment.get('sha') != sha or fragment.get('version') != 1 or set(fragment.get('images', {})) != {image}:
|
||||
raise ValueError('Image job artifact is missing or belongs to another commit')
|
||||
target = next(t for t in data['targets'] if t['name'] == name)
|
||||
if fragment.get('inputs') != {image: target['inputs']}:
|
||||
raise ValueError('Image artifact does not match the build plan')
|
||||
release['images'].update(fragment['images'])
|
||||
release['inputs'].update(fragment['inputs'])
|
||||
validate_release(release, sha)
|
||||
# Only a complete set of successful image jobs can publish the release tags.
|
||||
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
|
||||
env = {**os.environ, 'DOCKER_CONFIG': docker_config}
|
||||
try:
|
||||
subprocess.run( # noqa: S603, S607
|
||||
[
|
||||
shutil.which('docker') or '/usr/bin/docker',
|
||||
'login',
|
||||
'gcr.forust.xyz',
|
||||
'-u',
|
||||
os.environ['REGISTRY_USERNAME'],
|
||||
'--password-stdin',
|
||||
],
|
||||
input=os.environ['REGISTRY_PASSWORD'],
|
||||
text=True,
|
||||
check=True,
|
||||
env=env,
|
||||
)
|
||||
for image, digest in release['images'].items():
|
||||
command(
|
||||
'docker',
|
||||
'buildx',
|
||||
'imagetools',
|
||||
'create',
|
||||
'--prefer-index=false',
|
||||
'--tag',
|
||||
f'{image}:sha-{sha}',
|
||||
f'{image}@{digest}',
|
||||
env=env,
|
||||
timeout=90,
|
||||
)
|
||||
output.write_text(json.dumps(release, indent=2) + '\n')
|
||||
finally:
|
||||
shutil.rmtree(docker_config)
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument('action', choices=('prepare', 'image', 'finalize', 'gate', 'render', 'check-summary'))
|
||||
parser.add_argument('--output', type=Path, default=Path('release.json'))
|
||||
parser.add_argument('--ref', default='main')
|
||||
parser.add_argument('--event-sha', default='')
|
||||
parser.add_argument('--image', choices=IMAGES)
|
||||
parser.add_argument('--plan', type=Path, default=Path('build-plan.json'))
|
||||
parser.add_argument('--fragments', type=Path, default=Path('artifacts'))
|
||||
args = parser.parse_args()
|
||||
if args.action == 'check-summary':
|
||||
check_summary()
|
||||
elif args.action == 'render':
|
||||
render(sys.stdin, sys.stdout)
|
||||
elif args.action == 'gate':
|
||||
gate(args.output, args.ref, args.event_sha)
|
||||
elif args.action == 'prepare':
|
||||
prepare_images(args.output)
|
||||
elif args.action == 'image':
|
||||
if not args.image:
|
||||
parser.error('--image is required')
|
||||
build(args.output, args.image, args.plan)
|
||||
else:
|
||||
finalize_images(args.output, args.fragments, args.plan)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -1,26 +1,10 @@
|
||||
name: renovate-ci
|
||||
|
||||
on:
|
||||
# Read the workflow from the trusted base branch. PR code runs only on the
|
||||
# unprivileged runner selected below.
|
||||
pull_request_target:
|
||||
paths:
|
||||
- "renovate/**"
|
||||
- ".gitea/workflows/renovate-ci.yaml"
|
||||
- ".gitea/workflows/sync-renovate-configmap.sh"
|
||||
- ".gitea/workflows/compose-lint.sh"
|
||||
- ".gitea/workflows/install-ci-tools.sh"
|
||||
- ".gitea/workflows/tool-versions.env"
|
||||
pull_request:
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
paths:
|
||||
- "renovate/**"
|
||||
- ".gitea/workflows/renovate-ci.yaml"
|
||||
- ".gitea/workflows/sync-renovate-configmap.sh"
|
||||
- ".gitea/workflows/compose-lint.sh"
|
||||
- ".gitea/workflows/install-ci-tools.sh"
|
||||
- ".gitea/workflows/tool-versions.env"
|
||||
workflow_dispatch:
|
||||
|
||||
permissions:
|
||||
@@ -28,47 +12,37 @@ permissions:
|
||||
|
||||
jobs:
|
||||
validate-renovate:
|
||||
runs-on: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 20
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
with:
|
||||
ref: ${{ github.event_name == 'pull_request_target' && github.event.pull_request.head.sha || github.sha }}
|
||||
|
||||
# renovate/k8s/cronjob.yaml is the single source of truth for the version.
|
||||
- name: Resolve the deployed Renovate version
|
||||
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag,
|
||||
# so the same version that runs in the cluster is the one validated here.
|
||||
- name: Resolve the deployed Renovate image
|
||||
id: image
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
||||
renovate/k8s/cronjob.yaml | head -1)"
|
||||
if [[ ! "$image" =~ ^renovate/renovate:([0-9]+\.[0-9]+\.[0-9]+)$ ]]; then
|
||||
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
|
||||
if [ -z "$image" ]; then
|
||||
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
|
||||
exit 1
|
||||
fi
|
||||
version="${BASH_REMATCH[1]}"
|
||||
echo "using Renovate $version"
|
||||
printf 'version=%s\n' "$version" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Prepare pinned validation tools
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform node)"
|
||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||
echo "using $image"
|
||||
echo "image=$image" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Validate Renovate repository config
|
||||
shell: bash
|
||||
env:
|
||||
RENOVATE_VERSION: ${{ steps.image.outputs.version }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
npm_cache="$(mktemp -d "${RUNNER_TEMP:-/tmp}/renovate-npm-cache.XXXXXXXX")"
|
||||
trap 'rm -rf "$npm_cache"' EXIT
|
||||
NPM_CONFIG_CACHE="$npm_cache" RENOVATE_CONFIG_FILE="$PWD/renovate/renovate.json" \
|
||||
npm exec --yes --package="renovate@${RENOVATE_VERSION}" -- renovate-config-validator
|
||||
docker run --rm \
|
||||
-v "$PWD/renovate:/opt/renovate:ro" \
|
||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||
"${{ steps.image.outputs.image }}" \
|
||||
renovate-config-validator /opt/renovate/renovate.json
|
||||
|
||||
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
|
||||
# carries an inlined copy of the config. Fail if it no longer matches.
|
||||
@@ -82,6 +56,8 @@ jobs:
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
|
||||
export PATH="$tools_dir:$PATH"
|
||||
kubeconform \
|
||||
-strict \
|
||||
-ignore-missing-schemas \
|
||||
|
||||
@@ -32,14 +32,11 @@ concurrency:
|
||||
|
||||
jobs:
|
||||
run-renovate:
|
||||
if: github.ref == 'refs/heads/main'
|
||||
runs-on: homelab
|
||||
runs-on: [self-hosted, linux, arch, homelab]
|
||||
timeout-minutes: 60
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||
with:
|
||||
ref: refs/heads/main
|
||||
|
||||
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
|
||||
# Reading it here means this workflow validates and runs the exact version
|
||||
@@ -51,23 +48,21 @@ jobs:
|
||||
set -euo pipefail
|
||||
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
||||
renovate/k8s/cronjob.yaml | head -1)"
|
||||
if [[ ! "$image" =~ ^renovate/renovate:[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
|
||||
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
|
||||
if [ -z "$image" ]; then
|
||||
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
|
||||
exit 1
|
||||
fi
|
||||
echo "using $image"
|
||||
printf 'image=%s\n' "$image" >> "$GITHUB_OUTPUT"
|
||||
echo "image=$image" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Validate Renovate config
|
||||
shell: bash
|
||||
env:
|
||||
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
docker run --rm \
|
||||
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
|
||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||
"$RENOVATE_IMAGE" \
|
||||
"${{ steps.image.outputs.image }}" \
|
||||
renovate-config-validator
|
||||
|
||||
- name: Run Renovate
|
||||
@@ -78,7 +73,6 @@ jobs:
|
||||
RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
|
||||
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
|
||||
LOG_LEVEL: ${{ inputs.log_level }}
|
||||
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
|
||||
@@ -95,4 +89,4 @@ jobs:
|
||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||
-e RENOVATE_BASE_DIR=/tmp/renovate \
|
||||
-e LOG_LEVEL="${LOG_LEVEL:-info}" \
|
||||
"$RENOVATE_IMAGE"
|
||||
"${{ steps.image.outputs.image }}"
|
||||
@@ -1,13 +0,0 @@
|
||||
# kubectl emits a List for files containing multiple resources.
|
||||
(if .kind == "List" then .items[] else . end)
|
||||
| (.metadata.namespace // "default") as $ns
|
||||
| [
|
||||
(.. | objects
|
||||
| (.secretRef? // empty), (.secretKeyRef? // empty), (.secret? // empty)
|
||||
| select(.optional != true)
|
||||
| .name // .secretName // empty),
|
||||
(.. | objects | .imagePullSecrets[]?.name)
|
||||
]
|
||||
| unique[]
|
||||
| select(. != null and . != "")
|
||||
| "\($ns) \(.)"
|
||||
+63
-62
@@ -1,70 +1,71 @@
|
||||
#!/usr/bin/env bash
|
||||
# The SSH client submits once and follows durable stages on workstation.
|
||||
# usage: ssh-run.sh <stage>
|
||||
# Runs one deploy-lib.sh stage on the workstation over SSH.
|
||||
set -euo pipefail
|
||||
|
||||
: "${DEPLOY_HOST:?missing DEPLOY_HOST}"
|
||||
: "${DEPLOY_USER:?missing DEPLOY_USER}"
|
||||
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
|
||||
: "${DEPLOY_KNOWN_HOSTS:?configure pinned DEPLOY_KNOWN_HOSTS}"
|
||||
: "${DEPLOY_RUN_ID:?missing DEPLOY_RUN_ID}"
|
||||
[[ "$DEPLOY_USER" =~ ^[A-Za-z_][A-Za-z0-9_.-]*$ ]] || exit 1
|
||||
[[ "$DEPLOY_HOST" =~ ^[A-Za-z0-9_.:-]+$ ]] || exit 1
|
||||
[[ "$DEPLOY_RUN_ID" =~ ^[0-9]+-[0-9]+$ ]] || exit 1
|
||||
[[ "${DEPLOY_PORT:-22}" =~ ^[0-9]+$ ]] || exit 1
|
||||
|
||||
deploy_port="${DEPLOY_PORT:-22}"
|
||||
deploy_path="${DEPLOY_PATH:-/srv/homelab}"
|
||||
deploy_path="$(printf '%s' "$deploy_path" | tr -d '\"' | tr -d '\r' | xargs)"
|
||||
|
||||
# The private key is written to a per-run directory that is removed on exit, so a
|
||||
# failed or cancelled job cannot leave deploy credentials in the runner's temp
|
||||
# directory. Do not use a fixed path: apply-k8s and apply-compose run in parallel.
|
||||
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")"
|
||||
trap 'rm -rf "$key_dir"' EXIT
|
||||
chmod 700 "$key_dir"
|
||||
printf '%s\n' "$DEPLOY_KEY" >"$key_dir/key"
|
||||
printf '%s\n' "$DEPLOY_KNOWN_HOSTS" >"$key_dir/known_hosts"
|
||||
chmod 600 "$key_dir/key" "$key_dir/known_hosts"
|
||||
ssh_opts=(-i "$key_dir/key" -p "${DEPLOY_PORT:-22}" -o BatchMode=yes -o StrictHostKeyChecking=yes
|
||||
-o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15
|
||||
-o ServerAliveInterval=15 -o ServerAliveCountMax=4)
|
||||
controller=.local/lib/homelab-deploy/controller.py
|
||||
case "${1:?start, apply, verify, smoke or summary required}" in
|
||||
start)
|
||||
python3 - <<'PY' >"$key_dir/request.json"
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path
|
||||
release = json.loads(Path('release.json').read_text())
|
||||
print(json.dumps({'release': release, 'mode': os.environ.get('DEPLOY_MODE', 'changed'),
|
||||
'refresh_images': os.environ.get('REFRESH_IMAGES', 'false') == 'true'}))
|
||||
PY
|
||||
for attempt in 1 2 3; do
|
||||
rc=0
|
||||
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
|
||||
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" start "$DEPLOY_RUN_ID" <"$key_dir/request.json" || rc=$?
|
||||
[ "$rc" -eq 0 ] && exit 0
|
||||
[ "$rc" -eq 255 ] || exit "$rc"
|
||||
sleep 5
|
||||
done
|
||||
exit "$rc"
|
||||
trap 'rm -rf "$key_dir"' EXIT INT TERM
|
||||
|
||||
ssh_key="$key_dir/deploy_key"
|
||||
printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
|
||||
chmod 600 "$ssh_key"
|
||||
|
||||
# A connection that died silently used to hang until the job timeout, and the
|
||||
# stage was never re-run: one flaky TCP session cost a whole 45-minute apply.
|
||||
# ServerAlive* bounds how long a dead peer goes unnoticed, ConnectTimeout bounds
|
||||
# setup. Only exit 255 - ssh's own transport failures - is retried. A stage that
|
||||
# fails on its own merits exits with the remote's status, so a real failure
|
||||
# still surfaces its own log instead of burning three attempts. The stages are
|
||||
# declarative applies, so re-running one that had already committed is harmless.
|
||||
ssh_opts=(
|
||||
-i "$ssh_key" -p "$deploy_port"
|
||||
-o BatchMode=yes -o StrictHostKeyChecking=accept-new
|
||||
-o ConnectTimeout=15
|
||||
-o ServerAliveInterval=15 -o ServerAliveCountMax=4
|
||||
)
|
||||
|
||||
rc=0
|
||||
# apply-k8s and apply-compose are separate workflow jobs so the graph stays
|
||||
# intact for the verify job, but on a single node they must not run at once:
|
||||
# host docker churn on top of cluster churn is what melts the node (load 40+,
|
||||
# netbird/ssh die, helm is left pending-*). Serialize them on the workstation
|
||||
# with a shared lock; whoever arrives second waits.
|
||||
remote_cmd=(bash -se)
|
||||
case "$1" in
|
||||
apply-k8s | apply-compose)
|
||||
remote_cmd=(flock -w 5400 /tmp/homelab-apply.lock bash -se)
|
||||
;;
|
||||
apply|verify|smoke)
|
||||
result=0
|
||||
for attempt in 1 2 3; do
|
||||
rc=0
|
||||
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
|
||||
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$?
|
||||
[ "$rc" -eq 0 ] && break
|
||||
[ "$rc" -eq 255 ] || { result="$rc"; break; }
|
||||
echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)"
|
||||
if [ "$attempt" -eq 3 ]; then result=255; break; fi
|
||||
sleep 5
|
||||
done
|
||||
exit "$result"
|
||||
;;
|
||||
summary)
|
||||
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
||||
rc=0
|
||||
# shellcheck disable=SC2029 # The run ID is validated above.
|
||||
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" summary "$DEPLOY_RUN_ID" >"$key_dir/deploy-summary.md" || rc=$?
|
||||
if [ "$rc" -eq 0 ]; then
|
||||
cat "$key_dir/deploy-summary.md" >>"$GITHUB_STEP_SUMMARY" || echo "WARNING: cannot write the deploy summary"
|
||||
else
|
||||
echo 'Deploy summary is unavailable. The SSH connection failed or the controller did not respond. Check the job log.' >>"$GITHUB_STEP_SUMMARY" || true
|
||||
fi
|
||||
fi
|
||||
;;
|
||||
*) echo "Unknown SSH operation: $1" >&2; exit 1 ;;
|
||||
esac
|
||||
for attempt in 1 2 3; do
|
||||
if [ "$attempt" -gt 1 ]; then
|
||||
echo ":: warning::ssh transport failed, retrying (${attempt}/3)"
|
||||
sleep $((attempt * 5))
|
||||
fi
|
||||
rc=0
|
||||
# shellcheck disable=SC2029 # remote_cmd/ssh_opts expand on the client on purpose: they select the local ssh invocation, only the heredoc runs remotely.
|
||||
ssh "${ssh_opts[@]}" "${DEPLOY_USER}@${DEPLOY_HOST}" \
|
||||
env "REPO=$deploy_path" "APPLY_PRUNE=${APPLY_PRUNE:-false}" \
|
||||
"DEPLOY_SHA=${DEPLOY_SHA:-}" "DEPLOY_SNAPSHOT_DIR=${DEPLOY_SNAPSHOT_DIR:-}" \
|
||||
"STAGE=$1" "${remote_cmd[@]}" <<'EOF' || rc=$?
|
||||
source "$REPO/.gitea/workflows/deploy-lib.sh"
|
||||
run_stage "$STAGE"
|
||||
EOF
|
||||
[ "$rc" -eq 0 ] && break
|
||||
[ "$rc" -ne 255 ] && break
|
||||
done
|
||||
|
||||
if [ "$rc" -ne 0 ]; then
|
||||
echo ":: error::stage $1 failed over ssh (exit $rc)"
|
||||
fi
|
||||
exit "$rc"
|
||||
@@ -31,9 +31,3 @@ UV_VERSION="0.12.17"
|
||||
# so the tree that gets tested is the tree that gets built. Renovate keeps this
|
||||
# in step with the Dockerfile's node: tag via the "node runtime" group.
|
||||
NODE_VERSION="22.23.3"
|
||||
|
||||
# Secret-reference regression tests parse rendered Kubernetes objects.
|
||||
JQ_VERSION="1.8.1"
|
||||
|
||||
# BuildKit is the only auxiliary CI container; jobs themselves stay on the host.
|
||||
BUILDKIT_IMAGE="moby/buildkit:v0.33.1"
|
||||
@@ -94,6 +94,7 @@ replacements.txt
|
||||
.idea
|
||||
|
||||
# Temp files
|
||||
edu_master/temp/
|
||||
temp/*
|
||||
# Local-only tooling scratch space (pinned CI tools, verification scripts)
|
||||
tmp/
|
||||
|
||||
@@ -0,0 +1,163 @@
|
||||
# Homelab
|
||||
|
||||
Configuration for my homelab: Kubernetes manifests, Docker Compose stacks, and the
|
||||
Gitea Actions that build and deploy them. Most applications have both deployment
|
||||
formats. Headscale, Nextcloud AIO, and the media stack run on Docker; Kubernetes
|
||||
provides their ingress through Services and EndpointSlices.
|
||||
|
||||
These files contain this lab's domains, IP addresses, storage paths, and private
|
||||
registry names. Running them on another machine takes some editing.
|
||||
|
||||
## Start here
|
||||
|
||||
- [Service list](#services) — what each directory contains.
|
||||
- [Deployment workflow](.gitea/README.md) — selection, validation, and recovery.
|
||||
- [Repository review](docs/repository-review.md) — confirmed problems and fix branches.
|
||||
- [Shared PostgreSQL](postgres/README.md), [Traefik](traefik/README.md), and
|
||||
[cert-manager](cert-manager/README.md) — common dependencies.
|
||||
|
||||
## What gets deployed
|
||||
|
||||
The `active` files are switches for the deploy workflow, not health indicators.
|
||||
|
||||
| File | Effect |
|
||||
| ---------------------- | ----------------------------------------------------------- |
|
||||
| `<service>/active` | Include that directory's `compose.yaml` or `compose.yml`. |
|
||||
| `<service>/k8s/active` | Include its Kubernetes manifests or Kustomize overlay. |
|
||||
| Both | Run the Compose stack and apply the Kubernetes resources. |
|
||||
| Neither | Keep the configuration in Git without automatic deployment. |
|
||||
|
||||
`shared-compose.yaml`, `client.compose.yaml`, and `renovate-compose.yaml` are
|
||||
manual entry points. The deploy script does not discover them.
|
||||
|
||||
Kubernetes selection excludes secret files, examples, Helm values, and patches.
|
||||
Helm releases listed in `deploy-lib.sh` are upgraded separately. Traefik,
|
||||
cert-manager, and CrowdSec have additional bootstrap steps; an `active` marker
|
||||
does not install their charts.
|
||||
|
||||
The table below describes committed configuration. It does not claim that a
|
||||
service is currently healthy or running.
|
||||
|
||||
## Services
|
||||
|
||||
| Service | Configuration | Selected by markers |
|
||||
| ----------------------------------------------------------- | ---------------------------- | ------------------- |
|
||||
| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual |
|
||||
| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual |
|
||||
| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual |
|
||||
| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual |
|
||||
| [EDU session keeper and Telegram bot](edu_master/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Glance](glance/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
|
||||
| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Kener](kener/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes |
|
||||
| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [n8n](n8n/README.md) | Kubernetes + Compose | Manual |
|
||||
| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
|
||||
| [Penpot](penpot/README.md) | Compose | Manual |
|
||||
| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes |
|
||||
| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
|
||||
| [Termix](termix/README.md) | Kubernetes + Compose | Manual |
|
||||
| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes |
|
||||
| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes |
|
||||
|
||||
## Running a Compose stack
|
||||
|
||||
Use the service README first. Where a service has an env example, copy it inside
|
||||
that service's directory and replace the placeholders. The root `.env.example`
|
||||
is an older collection of variables, not a complete configuration for every stack.
|
||||
|
||||
For example, from the repository root:
|
||||
|
||||
```sh
|
||||
cd netbox
|
||||
cp .env.example .env
|
||||
$EDITOR .env
|
||||
docker compose config --quiet
|
||||
docker compose up -d
|
||||
docker compose ps
|
||||
```
|
||||
|
||||
Stacks that attach to `proxy` require an existing Docker network of that name and
|
||||
an appropriate reverse proxy. Published host ports still work independently of
|
||||
Traefik. Check port conflicts before starting an alternative to a Kubernetes
|
||||
service: DNS, STUN, and HTTP listeners can share the same host.
|
||||
|
||||
`docker compose down` keeps named volumes. Adding `-v` removes them.
|
||||
|
||||
## Preparing Kubernetes
|
||||
|
||||
The manifests assume Traefik CRDs, cert-manager, and a working storage provisioner.
|
||||
PrometheusRule and ServiceMonitor resources also need the Prometheus Operator.
|
||||
Replace the lab's hosts and addresses before using the configuration elsewhere.
|
||||
|
||||
Create a service's namespace, then prepare its ignored Secret from the example.
|
||||
For example:
|
||||
|
||||
```sh
|
||||
kubectl apply -f netbox/k8s/namespace.yaml
|
||||
cp netbox/k8s/secrets.yaml.example netbox/k8s/secrets.yaml
|
||||
$EDITOR netbox/k8s/secrets.yaml
|
||||
kubectl apply -f netbox/k8s/secrets.yaml
|
||||
```
|
||||
|
||||
The deploy workflow applies the tracked resources for marked services. Avoid
|
||||
applying an entire `k8s/` directory blindly: some directories contain Helm values,
|
||||
examples, and alternative routes. For a manual change, apply the selected manifest
|
||||
explicitly and check the resulting rollout.
|
||||
|
||||
Shared database passwords must agree between the `database` namespace and each
|
||||
application's Secret. Updating the PostgreSQL Secret does not change an existing
|
||||
role's password; see the database README.
|
||||
|
||||
## Local checks
|
||||
|
||||
CI pins its tools in `.gitea/workflows/tool-versions.env`. Use the same versions:
|
||||
|
||||
```sh
|
||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh)"
|
||||
export PATH="$tools_dir:$PATH"
|
||||
ruff check .
|
||||
ruff format --check .
|
||||
actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
|
||||
.gitea/workflows/sync-renovate-configmap.sh --check
|
||||
```
|
||||
|
||||
The [workflow README](.gitea/README.md#checks) lists the rest of the checks.
|
||||
Structure checks do not establish that local Secrets, mounted files, storage,
|
||||
or external services are ready.
|
||||
|
||||
## Data and recovery
|
||||
|
||||
State lives outside Git: PVCs, Docker volumes, bind mounts, databases, and ignored
|
||||
configuration. Keep backups of application data and the keys needed to read it.
|
||||
An image rollback does not roll back database migrations or ConfigMap contents.
|
||||
|
||||
Many PVCs use the cluster's default StorageClass; monitoring explicitly uses
|
||||
`local-path`. Check the PV reclaim policy before deleting a PVC or namespace.
|
||||
The manifests do not provide a repository-wide backup schedule.
|
||||
|
||||
`incident-archive/` contains past incident notes. `.docs/storage-audit-instruction.md`
|
||||
is a planning document, not evidence that NFS has been installed.
|
||||
@@ -0,0 +1,22 @@
|
||||
# AdGuard Home
|
||||
|
||||
DNS filtering with a web UI, DNS-over-TLS, and certificates from cert-manager.
|
||||
|
||||
The Kubernetes namespace is `adguard`. The workload uses `adguard-pvc` for
|
||||
configuration and working data, and mounts the `adguard-certs` TLS Secret.
|
||||
The LoadBalancer Service exposes DNS separately from the web ingress.
|
||||
|
||||
The Compose stack publishes TCP/UDP 53 and TCP 853 on the host. Prepare `conf/`
|
||||
and `certs/` before starting it. Starting both DNS deployments on the same address
|
||||
can cause a port conflict.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n adguard
|
||||
kubectl get events -n adguard --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -31,7 +31,7 @@ services:
|
||||
- "traefik.http.routers.adguard-dev.entrypoints=websecure"
|
||||
- "traefik.http.routers.adguard-dev.tls=true"
|
||||
# DoH Router
|
||||
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz`) || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`)"
|
||||
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz` || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`))"
|
||||
- "traefik.http.routers.dns-over-https.entrypoints=websecure"
|
||||
- "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt"
|
||||
|
||||
|
||||
@@ -51,8 +51,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: adguard-deployment
|
||||
namespace: adguard
|
||||
spec:
|
||||
@@ -66,6 +64,8 @@ spec:
|
||||
metadata:
|
||||
labels:
|
||||
app: adguard
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
spec:
|
||||
containers:
|
||||
- name: adguard
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
# Authentik
|
||||
|
||||
Identity provider with separate server and worker deployments.
|
||||
|
||||
Kubernetes connects to the shared PostgreSQL service in `database`. Set
|
||||
`AUTHENTIK_DB_PASSWORD` to the same value in both database and application Secrets.
|
||||
Keep `AUTHENTIK_SECRET_KEY` with the backups.
|
||||
|
||||
Compose uses its own PostgreSQL 15 container and bind-mounted media and templates.
|
||||
Its image defaults differ from Kubernetes; check both before an upgrade.
|
||||
The worker mounts the Docker socket for Docker outpost management.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n authentik
|
||||
kubectl get events -n authentik --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -27,8 +27,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: authentik-server-deployment
|
||||
namespace: authentik
|
||||
spec:
|
||||
@@ -65,8 +63,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: authentik-worker-deployment
|
||||
namespace: authentik
|
||||
spec:
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
# cert-manager
|
||||
|
||||
Public ACME issuers and an internal certificate authority.
|
||||
|
||||
This directory contains chart values and issuer resources, not the controller
|
||||
installation. Install the cert-manager chart with CRDs and the settings in
|
||||
`k8s/cert-manager-values.yaml` before applying the issuers.
|
||||
|
||||
`clusterissuer.yaml` defines staging and production Let's Encrypt issuers.
|
||||
They use HTTP-01 through the Traefik ingress class. Public DNS and inbound HTTP
|
||||
reachability must work for the requested names before issuance.
|
||||
`internal-ca.yaml` bootstraps the internal CA. Keep its private-key Secret backed
|
||||
up; the tracked `.crt` is only a public certificate.
|
||||
|
||||
This directory has no `k8s/active` marker. Apply the issuer files deliberately;
|
||||
`kubectl apply` does not interpret the Helm values file.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,22 @@
|
||||
# Cloudflare DDNS
|
||||
|
||||
Updates the lab DNS records when the public address changes.
|
||||
|
||||
Kubernetes runs in `default` with host networking and reads `cfddns-secrets`.
|
||||
The Compose stack also uses host networking. Configure the API token and domain
|
||||
list from the relevant example; keep DNS names consistent with the ingress rules.
|
||||
|
||||
`config.json.example` is a separate configuration example. The current Compose
|
||||
file does not mount a config.json file. Check configuration against the pinned
|
||||
DDNS image when changing between environment and file-based settings.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n default
|
||||
kubectl get events -n default --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -1,8 +1,6 @@
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: cfddns
|
||||
labels:
|
||||
app: cfddns
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
# Checkmk
|
||||
|
||||
Checkmk Raw monitoring site with web and agent-receiver ingress.
|
||||
|
||||
The site data lives in `checkmk-sites-pvc` on Kubernetes and the `sites` named
|
||||
volume on Compose. The agent receiver has a separate TCP route; enabling the
|
||||
web route alone does not expose it.
|
||||
|
||||
Prepare the password in the service env or Secret example. Inspect the Checkmk
|
||||
container logs during the first site creation.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n checkmk
|
||||
kubectl get events -n checkmk --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -17,8 +17,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: checkmk-deployment
|
||||
namespace: checkmk
|
||||
spec:
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
# Cloudflare Tunnel
|
||||
|
||||
A Kubernetes connector for an existing Cloudflare tunnel.
|
||||
|
||||
The Deployment runs in `default` and reads its token from the ignored Secret
|
||||
created from `k8s/secret.yaml.example`. Create the tunnel and its hostname rules
|
||||
in Cloudflare before starting the connector.
|
||||
|
||||
There is no Compose file or `k8s/active` marker. Apply the Secret first, then
|
||||
`k8s/deployment.yaml` when this tunnel is needed.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n default
|
||||
kubectl get events -n default --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -1,8 +1,6 @@
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: cloudflared
|
||||
labels:
|
||||
app: cloudflared
|
||||
@@ -20,7 +18,7 @@ spec:
|
||||
spec:
|
||||
containers:
|
||||
- name: cloudflared
|
||||
image: cloudflare/cloudflared:2026.10.0
|
||||
image: cloudflare/cloudflared:2026.9.3
|
||||
imagePullPolicy: IfNotPresent
|
||||
args:
|
||||
- tunnel
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
# File converters
|
||||
|
||||
ConvertX for server-side conversion and BentoPDF for PDF tools.
|
||||
|
||||
ConvertX persists files in `convertx-pvc`; BentoPDF has no persistent volume.
|
||||
Kubernetes configuration includes a local `config.yaml.example`, excluded from
|
||||
normal deployment. Copy and apply the real ConfigMap separately where required.
|
||||
|
||||
Compose publishes ConvertX on host port 9992 as well as attaching it to the
|
||||
proxy network. Replace the authentication settings from `.env.example` before
|
||||
exposing it outside the lab.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n converters
|
||||
kubectl get events -n converters --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -13,8 +13,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: convertx-deployment
|
||||
namespace: converters
|
||||
spec:
|
||||
|
||||
@@ -0,0 +1,25 @@
|
||||
# CrowdSec
|
||||
|
||||
Helm values, dashboards, network policy, and a maintenance CronJob.
|
||||
|
||||
Install CrowdSec separately using `k8s/crowdsec-values.yaml`; the deploy
|
||||
workflow does not have a CrowdSec Helm release entry. There is no `k8s/active`
|
||||
marker in this directory.
|
||||
|
||||
The LAPI policy and janitor run in `crowdsec`. The dashboard ConfigMaps are in
|
||||
`prometheus` for Grafana's sidecar. The janitor has its own ServiceAccount and
|
||||
namespace Role. Review its script and schedule before enabling cleanup.
|
||||
|
||||
Traefik's values state that enforcement moved to a host firewall bouncer. This
|
||||
repository does not install that host component.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n crowdsec
|
||||
kubectl get events -n crowdsec --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,21 @@
|
||||
# Dockmon
|
||||
|
||||
Docker management UI that talks to the host Docker daemon.
|
||||
|
||||
Both runtimes mount `/var/run/docker.sock`. On Kubernetes the socket belongs
|
||||
to the node hosting the pod, so this is not a cluster-wide container manager.
|
||||
|
||||
Compose stores application data in a named volume. Kubernetes uses a StatefulSet
|
||||
with a volume claim template. Its ServersTransport is specific to the upstream
|
||||
connection; keep it with the ingress resources.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n dockmon
|
||||
kubectl get events -n dockmon --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,151 @@
|
||||
# Repository review
|
||||
|
||||
Reviewed the tracked tree at `cc9c3de` and read the live workstation state on
|
||||
6 October 2026. Changes are split into documentation and individual fix branches,
|
||||
all based on that main commit. The original local checkout and its uncommitted
|
||||
monitoring changes were preserved. No deployment was performed.
|
||||
|
||||
## Confirmed problems with prepared fixes
|
||||
|
||||
| Priority | Problem and consequence | Fix branch |
|
||||
| -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------- |
|
||||
| High | `APPLY_PRUNE=true` is passed to each individual manifest apply. Each invocation sees only that file's desired objects and can delete other resources selected by the shared label. | `fix/deploy-prune-guard` |
|
||||
| High | Deploy validates Compose with interpolation and env/path resolution disabled. Required settings can pass validation and then fail during apply after other workloads have changed. | `fix/deploy-validation` |
|
||||
| Medium | Secret validation is text-based and compares names across all namespaces. A Secret elsewhere can hide a missing local Secret; mounted Secrets are also missed. | `fix/deploy-validation` |
|
||||
| Medium | Compose CI misses `postgres/shared-compose.yaml`, `netbird/client.compose.yaml`, and `renovate/renovate-compose.yaml`. | `fix/deploy-validation` |
|
||||
| Medium | NetBird Compose mounts `entrypoint.sh`, but it is absent. Its README also calls a missing `setup.sh`; a fresh checkout cannot start this stack as documented. | `fix/netbird-compose-runtime` |
|
||||
| Medium | Glance's CSS mount uses `glance-config`, whose keys do not include `user.css`. That key is in `glance-assets`; the pod's subPath mount cannot be prepared correctly. | `fix/glance-assets` |
|
||||
| Medium | The shared PostgreSQL initializer requires `NETBOX_DB_PASSWORD`, but the Compose env example omits it. Following the example leaves first initialization incomplete. | `fix/postgres-env-example` |
|
||||
| Medium | EDU's Compose env example uses old credential names and full URL variables, while the code reads `KEEPER_*` and paths under `EDU_URL_BASE`. | `fix/session-keeper-reliability` |
|
||||
| Medium | Session keeper HTTP calls have no timeouts. Its Redis cookie never expires, probes only check existence, and its logs include cookies. A hung or failed refresh can leave a stale session appearing ready. | `fix/session-keeper-reliability` |
|
||||
| Medium | AdGuard's DoH and SearXNG's Compose rules put Boolean expressions inside `Host(...)`. They are invalid router expressions despite valid YAML. | `fix/compose-router-rules` |
|
||||
|
||||
Traefik matchers should be combined as `Host(a) || Host(b)`; the rule syntax is
|
||||
described in the [Traefik rules documentation](https://doc.traefik.io/traefik/reference/routing-configuration/http/routing/rules-and-priority/).
|
||||
The fix retains the DoH path constraint for both hostnames.
|
||||
|
||||
The prune fix deliberately rejects the unsafe option. It does not introduce
|
||||
automatic deletion under a different implementation. Prune defaults to false,
|
||||
and no tracked resource currently carries the selector label, so this is a
|
||||
latent defect rather than evidence of a live deletion incident.
|
||||
|
||||
The session fix bounds HTTP and Redis calls, validates required credentials,
|
||||
sets a cookie lifetime of two refresh intervals, and marks success only after
|
||||
publishing the verified cookie. With the default ten-minute interval, an outage
|
||||
longer than twenty minutes will make the existing Redis-key readiness checks fail.
|
||||
That is an intentional change from indefinite apparent readiness.
|
||||
|
||||
The deployment fix extracts required pod Secret references from rendered JSON,
|
||||
checks their namespaces, includes init containers, image-pull credentials, and
|
||||
mounted/projected Secrets, and honors optional references. Ingress TLS Secrets
|
||||
issued by cert-manager are not treated as pre-existing pod prerequisites.
|
||||
It checks existence/access, not every key's contents or application validity.
|
||||
|
||||
## Live workstation observations
|
||||
|
||||
The SSH alias `workstation` is reachable. It has one Ready control-plane node,
|
||||
Kubernetes `v1.35.4+k0s`, and a Docker daemon alongside containerd. At inspection,
|
||||
no pods were Pending or in another non-running, non-completed phase. This is a
|
||||
point-in-time observation, not a complete application health test.
|
||||
|
||||
The deployment checkout at `/srv/homelab` is on main commit `2adf17c`, behind the
|
||||
reviewed local commit. It has untracked host configuration and a separate
|
||||
`userbot/` directory. It was not reset or cleaned.
|
||||
|
||||
| Observed difference | Implication |
|
||||
| ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| VictoriaMetrics and vmalert are running; the Prometheus StatefulSet has zero replicas. | A monitoring migration is already in progress outside committed main. Deploying the old Helm values can overwrite those settings. |
|
||||
| Homarr, Cloudflared, and Reloader are installed without their current Git active markers. | Installed services and marker-selected services are different inventories. Missing markers do not establish that a service is stopped. |
|
||||
| Cloudflare DDNS is running in both Docker and Kubernetes. | Confirm which instance should own DNS updates and whether their domain lists overlap before retiring either one. Secret values were not inspected. |
|
||||
| Traefik's LoadBalancer exposes port 8080 at `192.168.80.2`. | The direct API listener is deployed; its external reachability was not tested. |
|
||||
| Default `local-path` has reclaim policy Delete, while many existing PVs have been changed to Retain. | Current retention is partly live state. Recreating a claim can get a different policy from the old PV. |
|
||||
| NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. | Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup. |
|
||||
|
||||
The monitoring files already modified in the user's local tree correspond to the
|
||||
live migration. They are excluded from these branches. Reconcile that work before
|
||||
using this review's baseline to deploy monitoring.
|
||||
|
||||
## Remaining work
|
||||
|
||||
These need recovery design or infrastructure decisions rather than a small
|
||||
configuration correction:
|
||||
|
||||
- **SSH apply retries can replace the rollback baseline.** `ssh-run.sh` retries
|
||||
exit 255, including `apply-k8s`; every new invocation publishes a fresh snapshot.
|
||||
If the first attempt already changed workloads, the retry snapshots that partial
|
||||
state. Preserve a run-specific original baseline and verify it across retries.
|
||||
- **Rollback can exceed the job budget.** Verification is parallel, but
|
||||
`rollback_workloads` is serial with a five-minute limit per workload. The
|
||||
thirty-minute job budget can expire before recovery finishes. Bound recovery
|
||||
concurrency and account for both phases before choosing a new timeout.
|
||||
- **Snapshot collection is allowed to fail.** Generation and workload snapshot
|
||||
errors are warnings; verify can fall back to all workloads. A snapshot failure
|
||||
must not permit unrelated workloads to be selected for automatic undo.
|
||||
- **Rollback uses the previous revision, not the captured revision.** `rollout undo`
|
||||
without an explicit revision cannot guarantee restoration to the snapshot after
|
||||
retries or intervening rollouts. First deployments also have no previous revision.
|
||||
- **Manual deploy dispatch bypasses the CI-success trigger.** Either validate the
|
||||
target commit's successful CI run or document manual dispatch as an operator
|
||||
override with its own required checks.
|
||||
- **Direct Traefik API exposure is unauthenticated.** The latest local commit
|
||||
explicitly added it for Homarr. Preserve that integration while choosing a
|
||||
cluster-internal authenticated path or a verified network restriction; do not
|
||||
simply disable an integration that is already in use.
|
||||
- **Storage retention and backup are not reproducible as a whole.** Defaults and
|
||||
several important PV policies are Delete. There is no repository-wide backup
|
||||
schedule. Existing PVC StorageClass changes require migration rather than an
|
||||
in-place YAML edit.
|
||||
- **MeTube downloads are temporary on Kubernetes.** `/downloads` is a 20 GiB
|
||||
emptyDir. Decide whether pod replacement should discard files or whether it
|
||||
should use persistent storage. Compose uses a host directory instead.
|
||||
- **First-time activation needs a bootstrap path.** Deploy validation dry-runs
|
||||
namespaced resources before the apply stage creates namespaces and installs
|
||||
selected charts. On a fresh cluster, missing namespaces and CRDs need separate
|
||||
preparation; activation is not a complete installer.
|
||||
|
||||
## Validation
|
||||
|
||||
Baseline lint checks passed for Python, shell, workflows, YAML, standard Compose
|
||||
files, and Kubernetes resources with available schemas. Kubeconform found 347
|
||||
resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources.
|
||||
That skip count matters: passing schema validation does not validate Traefik rule
|
||||
strings or other controller-specific behavior.
|
||||
|
||||
Fix validation covers:
|
||||
|
||||
- Compose discovery of manual entry points, rejection of required-variable gaps,
|
||||
namespace-scoped and optional Secret references, and API/render failures.
|
||||
- NetBird setup idempotence, preservation of existing keys, file permissions,
|
||||
runtime rendering, and rejection of invalid trusted proxy CIDRs.
|
||||
- Session refresh success and failure paths, timeouts, cookie expiry, log redaction,
|
||||
missing credentials, and nonpositive refresh intervals.
|
||||
- Correct Glance ConfigMap key selection and PostgreSQL initializer/env alignment.
|
||||
- YAML and Compose structure for the corrected router rules, compared with the
|
||||
documented Traefik grammar. They were not exercised on the live proxy.
|
||||
- Prune rejection before any cluster invocation.
|
||||
|
||||
All seven fix branches and the documentation branch merged together without
|
||||
conflicts in a disposable validation worktree. The combined tree passed the
|
||||
CI-equivalent local checks, Markdown formatting/lint and link checks, all 35
|
||||
Compose structure checks, and 11 Python regression tests plus the shell
|
||||
validation regressions. CRD server-side validation and live rollout tests were
|
||||
not run.
|
||||
|
||||
Runtime tests use fixtures and mocks, not production credentials. Live checks read
|
||||
workload metadata, storage policies, chart versions, and container state only.
|
||||
They did not read Secret contents or change services.
|
||||
|
||||
## Reloader follow-up
|
||||
|
||||
`fix/reloader-integration` adds the active marker and opt-in annotations to 28
|
||||
application Deployments/StatefulSets that consume runtime ConfigMaps or Secrets.
|
||||
It corrects AdGuard's misplaced pod-template annotation. The Helm settings use
|
||||
annotation-based reloads, keep global auto-reload disabled, and ignore Jobs and
|
||||
CronJobs. PostgreSQL workloads are excluded because their credential variables
|
||||
and init scripts are only effective on an empty data directory.
|
||||
|
||||
The controller was already running on workstation when inspected. Its live
|
||||
configuration is unchanged by the branch: merge and deploy the integration to
|
||||
apply the new policy and application annotations. Configuration reload behavior
|
||||
was checked against the pinned chart, with Helm rendering and manifest validation;
|
||||
no production configuration was changed to provoke a test restart.
|
||||
@@ -0,0 +1,20 @@
|
||||
# Downtify
|
||||
|
||||
Download UI with a persistent downloads directory.
|
||||
|
||||
Compose stores downloads under `Downtify_downloads/`; Kubernetes uses
|
||||
`downtify-downloads-pvc`. The ingress manifests reference shared infrastructure,
|
||||
so check certificate and middleware availability before enabling them.
|
||||
|
||||
Back up downloads separately if they need to survive storage replacement.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n downtify
|
||||
kubectl get events -n downtify --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,14 @@
|
||||
EDU_LOGIN=your_edu_login_here
|
||||
EDU_PASSWORD=your_edu_password_here
|
||||
EDU_URL_LOGIN=https://edu.edu.vn.ua/user/login
|
||||
EDU_URL_VERIFY=https://edu.edu.vn.ua/course/userlist
|
||||
PHPSESSID_INTERVAL=10
|
||||
USER_AGENT="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36"
|
||||
WEBINAR_URL=https://edu.edu.vn.ua/webinar/useractive
|
||||
WEBINAR_CHECK_INTERVAL=60
|
||||
REDIS_HOST=redis
|
||||
REDIS_PORT=6379
|
||||
PLAYWRIGHT_WS=ws://playwright-service:3000/ws
|
||||
TZ=Europe/Kyiv
|
||||
WEBINAR_TELEGRAM_TOKEN=your_telegram_bot_token_here
|
||||
WEBINAR_ADMIN_ID=123456789
|
||||
@@ -0,0 +1 @@
|
||||
1.56.0
|
||||
@@ -0,0 +1,52 @@
|
||||
# EDU session keeper and Telegram bot
|
||||
|
||||
Keeps an EDU login session in Redis and sends Telegram notifications for new webinars. The bot also serves diary and schedule commands.
|
||||
|
||||
`phpsessid-bot/` logs into EDU and publishes `EDU_PHPSESSID` in Redis.
|
||||
`webinar-checker/` uses that cookie through a remote Playwright browser and stores
|
||||
subscribers, language preferences, and webinar history in Redis.
|
||||
|
||||
Kubernetes runs in `edu-master`, with Redis data in `redis-data-pvc`.
|
||||
`service.yaml`, `servicemonitor.yaml`, and `alerts.yaml` expose and monitor the
|
||||
checker's metrics on port 8000. Its `/health` endpoint reflects recent checks.
|
||||
|
||||
## Configuration
|
||||
|
||||
Use the keys in `k8s/secrets.yaml.example` as the reference. The committed Compose
|
||||
`.env.example` has stale names until `fix/session-keeper-reliability` is merged.
|
||||
The code reads:
|
||||
|
||||
| Variable | Purpose |
|
||||
| ----------------------------------------------------- | ------------------------------------------------ |
|
||||
| `KEEPER_LOGIN`, `KEEPER_PASSWORD` | EDU login credentials. |
|
||||
| `KEEPER_INTERVAL` | Session refresh interval in minutes; default 10. |
|
||||
| `EDU_URL_BASE` | EDU site origin. |
|
||||
| `EDU_URL_LOGIN`, `EDU_URL_COURSES`, `EDU_URL_WEBINAR` | Paths under that origin. |
|
||||
| `WEBINAR_TELEGRAM_TOKEN`, `WEBINAR_ADMIN_ID` | Telegram bot and administrator. |
|
||||
| `WEBINAR_CHECK_INTERVAL` | Checker interval in seconds; default 60. |
|
||||
| `REDIS_HOST`, `REDIS_PORT` | Redis connection. |
|
||||
| `PLAYWRIGHT_WS` | Remote browser WebSocket endpoint. |
|
||||
|
||||
Set the keeper keys explicitly in the Compose `.env`. Keep the Playwright Python
|
||||
package, browser image, server command, and `PLAYWRIGHT_VERSION` file on matching
|
||||
versions. The two Python images are built and published by CI.
|
||||
|
||||
## Bot use
|
||||
|
||||
Start a private chat with `/start` to subscribe. `/stop`, `/language`, `/diary`,
|
||||
`/schedule`, and `/setclass` manage subscriptions and school views. The
|
||||
administrator can manage the whitelist with `/adduser` and `/removeuser`.
|
||||
|
||||
Back up Redis if subscriber settings and notification history matter. Session
|
||||
cookies and Telegram tokens are credentials; keep them out of shared logs.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n edu-master
|
||||
kubectl get events -n edu-master --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,49 @@
|
||||
services:
|
||||
redis:
|
||||
image: redis:8.10.2-alpine
|
||||
restart: unless-stopped
|
||||
volumes:
|
||||
- redis-data:/data
|
||||
healthcheck:
|
||||
test: ["CMD", "redis-cli", "ping"]
|
||||
interval: 5s
|
||||
timeout: 3s
|
||||
retries: 5
|
||||
|
||||
playwright-service:
|
||||
image: mcr.microsoft.com/playwright:v1.56.0-jammy
|
||||
restart: unless-stopped
|
||||
command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws
|
||||
|
||||
session-keeper:
|
||||
build: ./phpsessid-bot
|
||||
image: gcr.forust.xyz/forust/session-keeper:prod
|
||||
pull_policy: build
|
||||
env_file: .env
|
||||
restart: unless-stopped
|
||||
depends_on:
|
||||
redis:
|
||||
condition: service_healthy
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
|
||||
interval: 30s
|
||||
timeout: 5s
|
||||
retries: 10
|
||||
start_period: 60s
|
||||
|
||||
webinar-checker:
|
||||
build: ./webinar-checker
|
||||
image: gcr.forust.xyz/forust/webinar-checker:prod
|
||||
pull_policy: build
|
||||
env_file: .env
|
||||
restart: unless-stopped
|
||||
depends_on:
|
||||
redis:
|
||||
condition: service_healthy
|
||||
session-keeper:
|
||||
condition: service_healthy
|
||||
playwright-service:
|
||||
condition: service_started
|
||||
|
||||
volumes:
|
||||
redis-data:
|
||||
File renamed without changes.
@@ -0,0 +1,96 @@
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: PrometheusRule
|
||||
metadata:
|
||||
name: edu-master-webinar
|
||||
namespace: edu-master
|
||||
labels:
|
||||
release: prometheus-stack
|
||||
spec:
|
||||
groups:
|
||||
- name: edu_master.webinar
|
||||
rules:
|
||||
# No successful webinar check for 5m (~2-3 missed 2-min checks).
|
||||
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
|
||||
# The last_success > 0 guard is mandatory: checker.py initialises
|
||||
# last_success to 0, so without it `time() - 0` equals the current epoch
|
||||
# and humanizeDuration renders ~20722d on every pod restart. Keep the
|
||||
# duration expression on the left so $value stays the real gap.
|
||||
- alert: WebinarCheckerNoSuccessfulCheck
|
||||
expr: |
|
||||
((time() - webinar_check_last_success_timestamp_seconds) > 300)
|
||||
and (webinar_check_last_success_timestamp_seconds > 0)
|
||||
and (webinar_check_last_run_timestamp_seconds > 0)
|
||||
for: 2m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Webinar checker has no successful check for 5m"
|
||||
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
|
||||
|
||||
# Checks are running but none has ever succeeded since pod start.
|
||||
# Split out from the rule above so a zeroed gauge never feeds
|
||||
# humanizeDuration.
|
||||
- alert: WebinarCheckerNeverSucceeded
|
||||
expr: |
|
||||
(webinar_check_last_success_timestamp_seconds == 0)
|
||||
and (webinar_check_last_run_timestamp_seconds > 0)
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Webinar checker has never completed a successful check"
|
||||
description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
|
||||
|
||||
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
|
||||
- alert: WebinarCheckerConsecutiveFailures
|
||||
expr: |
|
||||
webinar_check_consecutive_failures >= 3
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Webinar checker failing consecutively"
|
||||
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / playwright error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
|
||||
|
||||
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
|
||||
- alert: WebinarCheckerScrapeDown
|
||||
expr: |
|
||||
absent(webinar_check_last_run_timestamp_seconds) == 1
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Webinar checker metrics missing"
|
||||
description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
|
||||
|
||||
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
|
||||
- alert: EduPhpsessidMissing
|
||||
expr: |
|
||||
edu_phpsessid_present == 0
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "EDU_PHPSESSID missing"
|
||||
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
|
||||
|
||||
# Hard deps: checker and playwright deployments unavailable.
|
||||
- alert: WebinarCheckerDeploymentDown
|
||||
expr: |
|
||||
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Webinar checker deployment unavailable"
|
||||
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
|
||||
|
||||
- alert: PlaywrightServiceDown
|
||||
expr: |
|
||||
kube_deployment_status_replicas_unavailable{deployment="playwright-service", namespace="edu-master"} > 0
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Playwright service unavailable"
|
||||
description: "edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it."
|
||||
@@ -1,4 +1,4 @@
|
||||
apiVersion: v1
|
||||
kind: Namespace
|
||||
metadata:
|
||||
name: paperless
|
||||
name: edu-master
|
||||
@@ -0,0 +1,69 @@
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: playwright-service
|
||||
namespace: edu-master
|
||||
labels:
|
||||
app: edu-master-playwright
|
||||
spec:
|
||||
replicas: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app: edu-master-playwright
|
||||
strategy:
|
||||
type: Recreate
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: edu-master-playwright
|
||||
spec:
|
||||
containers:
|
||||
- name: playwright
|
||||
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker
|
||||
image: mcr.microsoft.com/playwright:v1.56.0-jammy
|
||||
imagePullPolicy: IfNotPresent
|
||||
# p95 412M, max 478M over 7 days, no limit before. Request is set at p95
|
||||
# so the pod is not an eviction candidate; the limit stays above 2x the
|
||||
# request because browser page lifetimes are unpredictable.
|
||||
resources:
|
||||
requests:
|
||||
cpu: "200m"
|
||||
memory: "416Mi"
|
||||
limits:
|
||||
memory: "1Gi"
|
||||
command:
|
||||
- npx
|
||||
- -y
|
||||
- playwright@1.56.0
|
||||
- run-server
|
||||
- --port
|
||||
- "3000"
|
||||
- --path
|
||||
- /ws
|
||||
ports:
|
||||
- containerPort: 3000
|
||||
readinessProbe:
|
||||
tcpSocket:
|
||||
port: 3000
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 10
|
||||
timeoutSeconds: 3
|
||||
livenessProbe:
|
||||
tcpSocket:
|
||||
port: 3000
|
||||
initialDelaySeconds: 15
|
||||
periodSeconds: 20
|
||||
timeoutSeconds: 3
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: playwright-service
|
||||
namespace: edu-master
|
||||
spec:
|
||||
selector:
|
||||
app: edu-master-playwright
|
||||
ports:
|
||||
- name: ws
|
||||
port: 3000
|
||||
targetPort: 3000
|
||||
@@ -0,0 +1,75 @@
|
||||
apiVersion: apps/v1
|
||||
kind: StatefulSet
|
||||
metadata:
|
||||
name: redis
|
||||
namespace: edu-master
|
||||
labels:
|
||||
app: edu-master-redis
|
||||
spec:
|
||||
serviceName: redis
|
||||
replicas: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app: edu-master-redis
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: edu-master-redis
|
||||
spec:
|
||||
containers:
|
||||
- name: redis
|
||||
image: redis:8.10.2-alpine
|
||||
imagePullPolicy: IfNotPresent
|
||||
ports:
|
||||
- containerPort: 6379
|
||||
volumeMounts:
|
||||
- name: redis-data
|
||||
mountPath: /data
|
||||
resources:
|
||||
requests:
|
||||
cpu: 25m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 250m
|
||||
memory: 128Mi
|
||||
readinessProbe:
|
||||
exec:
|
||||
command: ["redis-cli", "ping"]
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 5
|
||||
timeoutSeconds: 3
|
||||
livenessProbe:
|
||||
exec:
|
||||
command: ["redis-cli", "ping"]
|
||||
initialDelaySeconds: 10
|
||||
periodSeconds: 10
|
||||
timeoutSeconds: 3
|
||||
volumes:
|
||||
- name: redis-data
|
||||
persistentVolumeClaim:
|
||||
claimName: redis-data-pvc
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: redis-data-pvc
|
||||
namespace: edu-master
|
||||
spec:
|
||||
accessModes:
|
||||
- ReadWriteOnce
|
||||
resources:
|
||||
requests:
|
||||
storage: 1Gi
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: redis
|
||||
namespace: edu-master
|
||||
spec:
|
||||
selector:
|
||||
app: edu-master-redis
|
||||
ports:
|
||||
- name: redis
|
||||
port: 6379
|
||||
targetPort: 6379
|
||||
@@ -0,0 +1,50 @@
|
||||
# One-time Job to migrate redis state from docker compose to k8s (maintenance window).
|
||||
# The .example file is not applied by the deploy pipeline (mask *.example.yaml).
|
||||
#
|
||||
# Runbook:
|
||||
# 1. docker compose -f <repo>/edu_master/compose.yaml stop # SIGTERM -> redis will flush dump.rdb
|
||||
# 2. docker run --rm -v edu_master_redis-data:/data \
|
||||
# -v /tmp/edu-master-backup:/backup \
|
||||
# redis:alpine sh -c "cp /data/dump.rdb /backup/ && ls -la /backup"
|
||||
# 3. kubectl apply -f edu_master/k8s/namespace.yaml
|
||||
# 4. kubectl apply -f <only the PVC from redis.yaml> # seed must come BEFORE redis pod starts
|
||||
# 5. kubectl apply -f edu_master/k8s/restore-seed-job.yaml.example
|
||||
# kubectl wait --for=condition=complete job/redis-restore-seed -n edu-master --timeout=120s
|
||||
# 6. kubectl delete job redis-restore-seed -n edu-master
|
||||
# 7. kubectl apply -f edu_master/k8s/ -R # apply remaining manifests
|
||||
apiVersion: batch/v1
|
||||
kind: Job
|
||||
metadata:
|
||||
name: redis-restore-seed
|
||||
namespace: edu-master
|
||||
spec:
|
||||
backoffLimit: 2
|
||||
ttlSecondsAfterFinished: 3600
|
||||
template:
|
||||
spec:
|
||||
restartPolicy: Never
|
||||
containers:
|
||||
- name: seed
|
||||
image: redis:alpine
|
||||
command:
|
||||
- /bin/sh
|
||||
- -ec
|
||||
- |
|
||||
ls -la /backup
|
||||
cp /backup/dump.rdb /data/dump.rdb
|
||||
chmod 644 /data/dump.rdb
|
||||
ls -la /data
|
||||
volumeMounts:
|
||||
- name: redis-data
|
||||
mountPath: /data
|
||||
- name: backup
|
||||
mountPath: /backup
|
||||
readOnly: true
|
||||
volumes:
|
||||
- name: redis-data
|
||||
persistentVolumeClaim:
|
||||
claimName: redis-data-pvc
|
||||
- name: backup
|
||||
hostPath:
|
||||
path: /tmp/edu-master-backup
|
||||
type: DirectoryOrCreate
|
||||
@@ -0,0 +1,29 @@
|
||||
apiVersion: v1
|
||||
kind: Secret
|
||||
metadata:
|
||||
name: edu-master-secrets
|
||||
namespace: edu-master
|
||||
type: Opaque
|
||||
stringData:
|
||||
# Session keeper credentials
|
||||
KEEPER_LOGIN: ""
|
||||
KEEPER_PASSWORD: ""
|
||||
KEEPER_INTERVAL: "10"
|
||||
# EDU links
|
||||
EDU_URL_BASE: "https://edu.edu.vn.ua"
|
||||
EDU_URL_LOGIN: "/user/login"
|
||||
EDU_URL_COURSES: "/course/userlist"
|
||||
EDU_URL_WEBINAR: "/webinar/useractive"
|
||||
# Playwright
|
||||
USER_AGENT: ""
|
||||
PLAYWRIGHT_WS: "ws://playwright-service:3000/ws"
|
||||
# Webinar-checker
|
||||
WEBINAR_TELEGRAM_TOKEN: ""
|
||||
WEBINAR_ADMIN_ID: ""
|
||||
WEBINAR_CHECK_INTERVAL: "60"
|
||||
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
|
||||
METRICS_PORT: "8000"
|
||||
# Database
|
||||
REDIS_HOST: "redis"
|
||||
REDIS_PORT: "6379"
|
||||
TZ: "Europe/Kyiv"
|
||||
@@ -0,0 +1,15 @@
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: webinar-checker
|
||||
namespace: edu-master
|
||||
labels:
|
||||
app: edu-master-webinar-checker
|
||||
spec:
|
||||
selector:
|
||||
app: edu-master-webinar-checker
|
||||
ports:
|
||||
- name: metrics
|
||||
port: 8000
|
||||
targetPort: metrics
|
||||
protocol: TCP
|
||||
@@ -1,14 +1,14 @@
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: ServiceMonitor
|
||||
metadata:
|
||||
name: netbird-server
|
||||
namespace: netbird
|
||||
name: webinar-checker
|
||||
namespace: edu-master
|
||||
labels:
|
||||
release: prometheus-stack
|
||||
spec:
|
||||
selector:
|
||||
matchLabels:
|
||||
app: netbird-server
|
||||
app: edu-master-webinar-checker
|
||||
endpoints:
|
||||
- port: metrics
|
||||
path: /metrics
|
||||
@@ -0,0 +1,53 @@
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: session-keeper
|
||||
namespace: edu-master
|
||||
labels:
|
||||
app: edu-master-session-keeper
|
||||
spec:
|
||||
replicas: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app: edu-master-session-keeper
|
||||
strategy:
|
||||
type: Recreate
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: edu-master-session-keeper
|
||||
spec:
|
||||
initContainers:
|
||||
- name: wait-redis
|
||||
image: redis:8.10.2-alpine
|
||||
command:
|
||||
- /bin/sh
|
||||
- -ec
|
||||
- |
|
||||
i=0
|
||||
until redis-cli -h redis ping | grep -q PONG; do
|
||||
i=$((i+1))
|
||||
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
|
||||
sleep 2
|
||||
done
|
||||
echo "redis is ready"
|
||||
containers:
|
||||
- name: session-keeper
|
||||
image: gcr.forust.xyz/forust/session-keeper:prod
|
||||
envFrom:
|
||||
- secretRef:
|
||||
name: edu-master-secrets
|
||||
resources:
|
||||
requests:
|
||||
cpu: 25m
|
||||
memory: 32Mi
|
||||
limits:
|
||||
cpu: 250m
|
||||
memory: 128Mi
|
||||
readinessProbe:
|
||||
exec:
|
||||
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
|
||||
initialDelaySeconds: 15
|
||||
periodSeconds: 30
|
||||
timeoutSeconds: 5
|
||||
failureThreshold: 10
|
||||
@@ -0,0 +1,75 @@
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: webinar-checker
|
||||
namespace: edu-master
|
||||
labels:
|
||||
app: edu-master-webinar-checker
|
||||
spec:
|
||||
replicas: 1
|
||||
selector:
|
||||
matchLabels:
|
||||
app: edu-master-webinar-checker
|
||||
strategy:
|
||||
type: Recreate
|
||||
template:
|
||||
metadata:
|
||||
labels:
|
||||
app: edu-master-webinar-checker
|
||||
spec:
|
||||
# Enforces dependency order like compose depends_on:
|
||||
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) -> playwright started
|
||||
initContainers:
|
||||
- name: wait-deps
|
||||
image: redis:8.10.2-alpine
|
||||
command:
|
||||
- /bin/sh
|
||||
- -ec
|
||||
- |
|
||||
i=0
|
||||
until redis-cli -h redis ping | grep -q PONG; do
|
||||
i=$((i+1))
|
||||
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
|
||||
sleep 2
|
||||
done
|
||||
echo "redis ok"
|
||||
until [ "$(redis-cli -h redis EXISTS EDU_PHPSESSID)" = "1" ]; do
|
||||
i=$((i+1))
|
||||
[ "$i" -ge 300 ] && echo "TIMEOUT: no PHPSESSID (session-keeper down?)" && exit 1
|
||||
sleep 2
|
||||
done
|
||||
echo "PHPSESSID ok"
|
||||
until nc -z playwright-service 3000; do
|
||||
i=$((i+1))
|
||||
[ "$i" -ge 300 ] && echo "TIMEOUT: playwright-service not reachable" && exit 1
|
||||
sleep 2
|
||||
done
|
||||
echo "playwright ok"
|
||||
containers:
|
||||
- name: webinar-checker
|
||||
image: gcr.forust.xyz/forust/webinar-checker:prod
|
||||
ports:
|
||||
- name: metrics
|
||||
containerPort: 8000
|
||||
protocol: TCP
|
||||
readinessProbe:
|
||||
httpGet:
|
||||
path: /health
|
||||
port: metrics
|
||||
periodSeconds: 10
|
||||
timeoutSeconds: 3
|
||||
failureThreshold: 12
|
||||
initialDelaySeconds: 10
|
||||
envFrom:
|
||||
- secretRef:
|
||||
name: edu-master-secrets
|
||||
env:
|
||||
- name: TZ
|
||||
value: "Europe/Kyiv"
|
||||
resources:
|
||||
requests:
|
||||
cpu: "50m"
|
||||
memory: "192Mi"
|
||||
limits:
|
||||
cpu: "600m"
|
||||
memory: "384Mi"
|
||||
@@ -0,0 +1,15 @@
|
||||
FROM python:3.11-slim
|
||||
|
||||
WORKDIR /app
|
||||
|
||||
# Install system dependencies
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends redis-tools && rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# Install dependencies
|
||||
RUN pip install --no-cache-dir requests==2.32.3 redis==5.2.1
|
||||
|
||||
# Copy application code
|
||||
COPY . .
|
||||
|
||||
# Run the bot
|
||||
CMD ["python", "bot.py"]
|
||||
@@ -0,0 +1,132 @@
|
||||
import logging
|
||||
import os
|
||||
import time
|
||||
from datetime import datetime
|
||||
|
||||
import redis
|
||||
import requests
|
||||
|
||||
# Configure logging
|
||||
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
# Load configuration (adapted to .env keys)
|
||||
def _env(key, default=None):
|
||||
v = os.getenv(key, default)
|
||||
if isinstance(v, str) and len(v) >= 2 and ((v[0] == '"' and v[-1] == '"') or (v[0] == "'" and v[-1] == "'")):
|
||||
return v[1:-1]
|
||||
return v
|
||||
|
||||
|
||||
LOGIN = _env('KEEPER_LOGIN')
|
||||
PASSWORD = _env('KEEPER_PASSWORD')
|
||||
|
||||
EDU_BASE = _env('EDU_URL_BASE', 'https://edu.edu.vn.ua')
|
||||
EDU_LOGIN_PATH = _env('EDU_URL_LOGIN', '/user/login')
|
||||
EDU_COURSES_PATH = _env('EDU_URL_COURSES', '/course/userlist')
|
||||
URL_LOGIN = f'{EDU_BASE.rstrip("/")}/{EDU_LOGIN_PATH.lstrip("/")}'
|
||||
URL_VERIFY = f'{EDU_BASE.rstrip("/")}/{EDU_COURSES_PATH.lstrip("/")}'
|
||||
|
||||
INTERVAL = int(_env('KEEPER_INTERVAL', 10))
|
||||
USER_AGENT = _env(
|
||||
'USER_AGENT',
|
||||
'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36',
|
||||
)
|
||||
REDIS_HOST = _env('REDIS_HOST', 'redis')
|
||||
REDIS_PORT = int(_env('REDIS_PORT', 6379))
|
||||
|
||||
SUCCESS_FILE = '/tmp/last_success' # noqa: S108
|
||||
|
||||
|
||||
def touch_success_file():
|
||||
"""Updates the timestamp of the success file for healthchecks."""
|
||||
try:
|
||||
with open(SUCCESS_FILE, 'w') as f:
|
||||
f.write(str(datetime.now().timestamp()))
|
||||
except Exception as e:
|
||||
logger.error(f'Failed to touch success file: {e}')
|
||||
|
||||
|
||||
def main():
|
||||
logger.info('Starting Session Keeper Bot')
|
||||
|
||||
# Connect to Redis
|
||||
try:
|
||||
redis_client = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
|
||||
redis_client.ping()
|
||||
logger.info(f'Connected to Redis at {REDIS_HOST}:{REDIS_PORT}')
|
||||
except Exception as e:
|
||||
logger.error(f'Failed to connect to Redis: {e}')
|
||||
return
|
||||
|
||||
session = requests.Session()
|
||||
|
||||
# Set headers
|
||||
headers = {
|
||||
'User-Agent': USER_AGENT,
|
||||
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
|
||||
'Accept-Language': 'en-US,en;q=0.9',
|
||||
'Cache-Control': 'max-age=0',
|
||||
'Upgrade-Insecure-Requests': '1',
|
||||
'Sec-Fetch-Site': 'same-origin',
|
||||
'Sec-Fetch-Mode': 'navigate',
|
||||
'Sec-Fetch-User': '?1',
|
||||
'Sec-Fetch-Dest': 'document',
|
||||
'Sec-Ch-Ua': '"Not_A Brand";v="99", "Chromium";v="142"',
|
||||
'Sec-Ch-Ua-Mobile': '?0',
|
||||
'Sec-Ch-Ua-Platform': '"Linux"',
|
||||
'Accept-Encoding': 'gzip, deflate, br',
|
||||
'Priority': 'u=0, i',
|
||||
}
|
||||
session.headers.update(headers)
|
||||
|
||||
while True:
|
||||
try:
|
||||
logger.info('Attempting login...')
|
||||
|
||||
# Login payload
|
||||
payload = {'login': LOGIN, 'password': PASSWORD}
|
||||
|
||||
# Perform Login
|
||||
# Note: The user request shows a POST to /user/login with form data
|
||||
# We need to make sure we handle the PHPSESSID correctly.
|
||||
# If we already have a PHPSESSID, requests will send it.
|
||||
|
||||
login_response = session.post(URL_LOGIN, data=payload, allow_redirects=True)
|
||||
|
||||
logger.info(f'Login Response Status: {login_response.status_code}')
|
||||
logger.info(f'Cookies after login: {session.cookies.get_dict()}')
|
||||
|
||||
# Verify Session
|
||||
logger.info('Verifying session...')
|
||||
verify_response = session.get(URL_VERIFY, allow_redirects=False)
|
||||
|
||||
logger.info(f'Verify Response Status: {verify_response.status_code}')
|
||||
|
||||
if verify_response.status_code == 200:
|
||||
logger.info('Session verification SUCCESS (200 OK).')
|
||||
touch_success_file()
|
||||
|
||||
# Save PHPSESSID to Redis
|
||||
phpsessid = session.cookies.get('PHPSESSID')
|
||||
if phpsessid:
|
||||
try:
|
||||
redis_client.set('EDU_PHPSESSID', phpsessid)
|
||||
logger.info(f'Saved PHPSESSID to Redis: {phpsessid}')
|
||||
except Exception as e:
|
||||
logger.error(f'Failed to save PHPSESSID to Redis: {e}')
|
||||
elif verify_response.status_code == 302:
|
||||
logger.warning('Session verification FAILED (302 Redirect). Session might be invalid.')
|
||||
else:
|
||||
logger.warning(f'Session verification returned unexpected status: {verify_response.status_code}')
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f'An error occurred: {e}')
|
||||
|
||||
logger.info(f'Sleeping for {INTERVAL} minutes...')
|
||||
time.sleep(INTERVAL * 60)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,13 @@
|
||||
FROM python:3.11-slim
|
||||
|
||||
WORKDIR /app
|
||||
|
||||
# renovate: datasource=pypi depName=playwright versioning=pep440
|
||||
ARG PLAYWRIGHT_VERSION=1.56.0
|
||||
|
||||
# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers
|
||||
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
|
||||
|
||||
COPY checker.py .
|
||||
|
||||
CMD ["python", "checker.py"]
|
||||
File diff suppressed because it is too large.
Load diff
@@ -0,0 +1,21 @@
|
||||
# Error pages
|
||||
|
||||
Static HTTP error pages served by an Nginx image built in CI.
|
||||
|
||||
Edit the HTML in `html/`; the Dockerfile copies it into the image.
|
||||
Kubernetes exposes `error-pages-service` in `error-pages` for Traefik's error
|
||||
middleware. Keep the middleware's namespace and port aligned with that Service.
|
||||
|
||||
For a local build, run `docker build -t homelab-error-pages .` from this directory.
|
||||
Compose references the private registry image rather than a build context.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n error-pages
|
||||
kubectl get events -n error-pages --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Gitea
|
||||
|
||||
Git hosting with HTTP and a separate SSH route.
|
||||
|
||||
Kubernetes uses the shared PostgreSQL service and `gitea-pvc` for repositories
|
||||
and application data. Match the Gitea database password with the shared database
|
||||
Secret. SSH is routed through Traefik's TCP entrypoint on 2221.
|
||||
|
||||
Compose uses a separate PostgreSQL 14 database, bind mounts `gitea-data/` and
|
||||
`gitea-db/`, and publishes host port 2221. It is an alternative deployment with
|
||||
its own database, not a second frontend for the Kubernetes instance.
|
||||
|
||||
Back up repositories, application configuration, and a consistent database dump
|
||||
together. Gitea Actions definitions for this repository live in `../.gitea/`.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n gitea
|
||||
kubectl get events -n gitea --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -20,8 +20,6 @@ data:
|
||||
|
||||
GITEA__mailer__ENABLED: "false"
|
||||
|
||||
GITEA__metrics__ENABLED: "true"
|
||||
|
||||
# No code/issue search needed: bleve reindexes the whole issue index on
|
||||
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
|
||||
# the rotational disk for an hour. "db" serves issue search from postgres.
|
||||
|
||||
@@ -3,8 +3,6 @@ kind: Service
|
||||
metadata:
|
||||
name: gitea-service
|
||||
namespace: gitea
|
||||
labels:
|
||||
app: gitea
|
||||
spec:
|
||||
selector:
|
||||
app: gitea
|
||||
@@ -19,8 +17,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: gitea-deployment
|
||||
namespace: gitea
|
||||
spec:
|
||||
|
||||
@@ -7,8 +7,7 @@ spec:
|
||||
entryPoints:
|
||||
- websecure
|
||||
routes:
|
||||
# Metrics are scraped directly through the cluster Service.
|
||||
- match: (Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)) && !PathPrefix(`/metrics`)
|
||||
- match: Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)
|
||||
kind: Rule
|
||||
services:
|
||||
- name: gitea-service
|
||||
|
||||
@@ -1,16 +0,0 @@
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: ServiceMonitor
|
||||
metadata:
|
||||
name: gitea
|
||||
namespace: gitea
|
||||
labels:
|
||||
release: prometheus-stack
|
||||
spec:
|
||||
selector:
|
||||
matchLabels:
|
||||
app: gitea
|
||||
endpoints:
|
||||
- port: http
|
||||
path: /metrics
|
||||
interval: 30s
|
||||
scrapeTimeout: 10s
|
||||
@@ -0,0 +1,26 @@
|
||||
# Glance
|
||||
|
||||
Dashboard pages for links, service checks, and Docker containers.
|
||||
|
||||
Compose mounts `config/` and `assets/`. The Kubernetes equivalents are embedded
|
||||
in `k8s/glance-config.yaml`: `glance-config` holds pages and `glance-assets` holds
|
||||
`user.css`. Update both copies when changing shared content.
|
||||
|
||||
Kubernetes serves the dashboard under `/glance`. Its pod also mounts the node's
|
||||
Docker socket. It references `glance-secrets` for `ADGUARD_PASSWORD`, but there is
|
||||
no tracked Secret example; create that Secret in `glance` before starting it.
|
||||
Compose expects a local `.env` with the same password.
|
||||
|
||||
The CSS mount points at the wrong ConfigMap on the reviewed main commit;
|
||||
`fix/glance-assets` corrects it.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n glance
|
||||
kubectl get events -n glance --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -13,8 +13,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: glance-deployment
|
||||
namespace: glance
|
||||
spec:
|
||||
@@ -71,7 +69,7 @@ spec:
|
||||
name: glance-config
|
||||
- name: glance-assets
|
||||
configMap:
|
||||
name: glance-assets
|
||||
name: glance-config
|
||||
- name: docker-socket
|
||||
hostPath:
|
||||
path: /var/run/docker.sock
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
# Headscale
|
||||
|
||||
Headscale, Headplane, and a separate web administration UI on Docker.
|
||||
|
||||
Kubernetes only provides routes to the Docker host. Update the addresses in
|
||||
`k8s/routing/external-service.yaml` if the host moves.
|
||||
|
||||
Copy `config/headscale.yaml.example`, `config/headplane.yaml.example`, and
|
||||
`config/policy.json.example` to their names without `.example`. Set the public
|
||||
server URL, DNS settings, Headplane cookie secret, and Headscale public URL.
|
||||
The example URLs are placeholders.
|
||||
|
||||
Compose publishes Headscale on 18080, its metrics port on 19090, Headplane on
|
||||
13000, and the other UI on 10080. The data volumes store the Headscale database,
|
||||
keys, and Headplane state. The embedded DERP configuration needs reachable
|
||||
addresses; Compose does not publish its UDP 3478 listener.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -7,7 +7,7 @@
|
||||
# # Dev server_url
|
||||
# server_url: https://hs.dev_internal_domain.internal
|
||||
listen_addr: 0.0.0.0:8080
|
||||
metrics_listen_addr: 0.0.0.0:9090
|
||||
metrics_listen_addr: 127.0.0.1:9090
|
||||
grpc_listen_addr: 127.0.0.1:50443
|
||||
grpc_allow_insecure: false
|
||||
noise:
|
||||
|
||||
@@ -3,8 +3,6 @@ kind: Service
|
||||
metadata:
|
||||
name: headscale-server-external
|
||||
namespace: headscale
|
||||
labels:
|
||||
app: headscale
|
||||
spec:
|
||||
ports:
|
||||
- port: 8080
|
||||
|
||||
@@ -1,18 +0,0 @@
|
||||
apiVersion: operator.victoriametrics.com/v1beta1
|
||||
kind: VMServiceScrape
|
||||
metadata:
|
||||
name: headscale
|
||||
namespace: headscale
|
||||
labels:
|
||||
release: prometheus-stack
|
||||
spec:
|
||||
# The external Service has a manually managed EndpointSlice, not Endpoints.
|
||||
discoveryRole: endpointslice
|
||||
selector:
|
||||
matchLabels:
|
||||
app: headscale
|
||||
endpoints:
|
||||
- port: metrics
|
||||
path: /metrics
|
||||
interval: 30s
|
||||
scrapeTimeout: 10s
|
||||
@@ -0,0 +1,24 @@
|
||||
# Homarr
|
||||
|
||||
Dashboard with Kubernetes integration and persistent application state.
|
||||
|
||||
Kubernetes uses the `homarr` ServiceAccount and the read-only ClusterRole in
|
||||
`k8s/rbac.yaml`. Application data lives in `homarr-pvc`; supply the encryption key
|
||||
from `k8s/secrets.yaml.example` before the first start and retain it with backups.
|
||||
|
||||
The committed ingress is internal. There is no `k8s/active` marker even though
|
||||
manifests exist, so the workflow does not select Homarr automatically.
|
||||
|
||||
Compose publishes ports 80 and 81, mounts appdata and the Docker socket, and
|
||||
expects a local kubeconfig. Check these host ports against Traefik before use.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n homarr
|
||||
kubectl get events -n homarr --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
services:
|
||||
homarr:
|
||||
container_name: homarr
|
||||
image: ghcr.io/homarr-labs/homarr:v2.2.0
|
||||
image: ghcr.io/homarr-labs/homarr:v2.1.2
|
||||
restart: unless-stopped
|
||||
volumes:
|
||||
- ./appdata:/appdata
|
||||
|
||||
@@ -13,8 +13,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: homarr-deployment
|
||||
namespace: homarr
|
||||
spec:
|
||||
@@ -32,7 +30,7 @@ spec:
|
||||
serviceAccountName: homarr
|
||||
containers:
|
||||
- name: homarr
|
||||
image: ghcr.io/homarr-labs/homarr:v2.2.0
|
||||
image: ghcr.io/homarr-labs/homarr:v2.1.2
|
||||
envFrom:
|
||||
- configMapRef:
|
||||
name: homarr-config
|
||||
|
||||
@@ -0,0 +1,25 @@
|
||||
# Homepages
|
||||
|
||||
Two static sites: Forust and xdfnx.
|
||||
|
||||
The site sources are in `forust_files/` and `xdfnx_files/`. CI builds each with
|
||||
its own Dockerfile and publishes it to the private registry. Kubernetes serves
|
||||
the image contents; Compose overlays the source directories as bind mounts.
|
||||
|
||||
Both Traefik IngressRoute and Gateway API route manifests are committed.
|
||||
Keep their hostnames and backend Services aligned when changing routes.
|
||||
Certificate resources cover public and internal hostnames.
|
||||
|
||||
Build either site locally with `docker build -f Dockerfile.forust .` or
|
||||
`docker build -f Dockerfile.xdfnx .` from this directory.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n homepages
|
||||
kubectl get events -n homepages --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,29 @@
|
||||
# Immich
|
||||
|
||||
Photo library with its own vector-enabled PostgreSQL and machine-learning service.
|
||||
|
||||
This database is separate from the shared PostgreSQL instance. Keep the server
|
||||
and machine-learning versions aligned when upgrading.
|
||||
|
||||
Kubernetes bind-mounts `/mnt/immich/library` from the node. That directory must
|
||||
already exist and contain the intended library; moving the pod to a different
|
||||
node does not move the files. PostgreSQL and Valkey use StatefulSet storage, and
|
||||
the model cache has its own PVC.
|
||||
|
||||
Compose reads `UPLOAD_LOCATION` and `DB_DATA_LOCATION` from `.env`. The example
|
||||
uses the same library path as Kubernetes. Run one writer against that library;
|
||||
do not start both deployments as independent instances over the same files.
|
||||
|
||||
Back up the library and a consistent database dump together. The model cache
|
||||
can be rebuilt; the photo database cannot.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n immich
|
||||
kubectl get events -n immich --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -6,10 +6,6 @@ metadata:
|
||||
data:
|
||||
TZ: "Europe/Bratislava"
|
||||
|
||||
IMMICH_TELEMETRY_INCLUDE: "all"
|
||||
IMMICH_API_METRICS_PORT: "8081"
|
||||
IMMICH_MICROSERVICES_METRICS_PORT: "8082"
|
||||
|
||||
# The database in this namespace, not the shared one in the database
|
||||
# namespace: v3 needs VectorChord, and only the dedicated image carries it.
|
||||
DB_HOSTNAME: "immich-postgres"
|
||||
|
||||
@@ -3,8 +3,6 @@ kind: Service
|
||||
metadata:
|
||||
name: immich-service
|
||||
namespace: immich
|
||||
labels:
|
||||
app: immich
|
||||
spec:
|
||||
selector:
|
||||
app: immich
|
||||
@@ -12,18 +10,10 @@ spec:
|
||||
- name: http
|
||||
port: 2283
|
||||
targetPort: 2283
|
||||
- name: api-metrics
|
||||
port: 8081
|
||||
targetPort: api-metrics
|
||||
- name: worker-metrics
|
||||
port: 8082
|
||||
targetPort: worker-metrics
|
||||
---
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: immich-deployment
|
||||
namespace: immich
|
||||
labels:
|
||||
@@ -49,10 +39,6 @@ spec:
|
||||
ports:
|
||||
- name: http
|
||||
containerPort: 2283
|
||||
- name: api-metrics
|
||||
containerPort: 8081
|
||||
- name: worker-metrics
|
||||
containerPort: 8082
|
||||
volumeMounts:
|
||||
- name: immich-data
|
||||
mountPath: /data
|
||||
|
||||
@@ -14,8 +14,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: immich-machine-learning-deployment
|
||||
namespace: immich
|
||||
labels:
|
||||
|
||||
@@ -1,20 +0,0 @@
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: ServiceMonitor
|
||||
metadata:
|
||||
name: immich
|
||||
namespace: immich
|
||||
labels:
|
||||
release: prometheus-stack
|
||||
spec:
|
||||
selector:
|
||||
matchLabels:
|
||||
app: immich
|
||||
endpoints:
|
||||
- port: api-metrics
|
||||
path: /metrics
|
||||
interval: 30s
|
||||
scrapeTimeout: 10s
|
||||
- port: worker-metrics
|
||||
path: /metrics
|
||||
interval: 30s
|
||||
scrapeTimeout: 10s
|
||||
@@ -17,8 +17,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: StatefulSet
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: immich-valkey
|
||||
namespace: immich
|
||||
labels:
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
# Kener
|
||||
|
||||
Status page with Redis and persistent database and upload directories.
|
||||
|
||||
Kubernetes uses `kener-db-pvc`, `kener-uploads-pvc`, and a Redis StatefulSet.
|
||||
Compose keeps the corresponding directories in named volumes. Set the signing
|
||||
and other credentials from the env or Secret example.
|
||||
|
||||
The monitors and route settings live in `k8s/config.yaml` and `k8s/ingress.yaml`.
|
||||
There is no active marker for either runtime.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n kener
|
||||
kubectl get events -n kener --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -13,8 +13,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: kener-deployment
|
||||
namespace: kener
|
||||
spec:
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
# Loki and Alloy
|
||||
|
||||
Loki log storage and Alloy collection, both deployed through Helm.
|
||||
|
||||
The deploy library lists separate `loki` and `alloy` releases in `prometheus`,
|
||||
controlled by this directory's `k8s/active` marker. Chart versions are pinned in
|
||||
`deploy-lib.sh`; settings live in `loki-values.yaml` and `alloy-values.yaml`.
|
||||
|
||||
Alloy collects Kubernetes logs. Grafana's Loki datasource is configured in the
|
||||
monitoring stack. Review Loki retention and storage settings before enabling
|
||||
collection on a new cluster.
|
||||
|
||||
Check releases with `helm list -n prometheus` and inspect collector logs before
|
||||
assuming that an empty Grafana query means there were no events.
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -0,0 +1,21 @@
|
||||
# MeTube
|
||||
|
||||
Web downloader behind Traefik.
|
||||
|
||||
Compose bind-mounts `MeTube_downloads/` on the host. Kubernetes uses a 20 GiB
|
||||
`emptyDir` for `/downloads`: completed downloads disappear when the pod is
|
||||
replaced. Download files from the UI promptly if this temporary storage is intended.
|
||||
|
||||
Application settings are in `k8s/config.yaml`. Persisting downloads in Kubernetes
|
||||
would require changing the volume to a PVC and choosing a storage policy.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n metube
|
||||
kubectl get events -n metube --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
@@ -13,8 +13,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: metube-deployment
|
||||
namespace: metube
|
||||
spec:
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
# n8n
|
||||
|
||||
Workflow automation with persistent application and file storage.
|
||||
|
||||
Kubernetes keeps application state in `n8n-node-pvc` and files in
|
||||
`n8n-files-pvc`; Compose uses `node-data` and `files` named volumes.
|
||||
Webhook URLs and proxy settings are committed in the application config.
|
||||
|
||||
There is no active marker. Review the URLs before enabling the stack, and retain
|
||||
the credential encryption key with the database or application-data backup.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n n8n
|
||||
kubectl get events -n n8n --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
services:
|
||||
n8n:
|
||||
image: docker.n8n.io/n8nio/n8n:2.43.0
|
||||
image: docker.n8n.io/n8nio/n8n:2.42.3
|
||||
container_name: n8n
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
|
||||
+1
-3
@@ -13,8 +13,6 @@ spec:
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
annotations:
|
||||
reloader.stakater.com/auto: "true"
|
||||
name: n8n-deployment
|
||||
namespace: n8n
|
||||
spec:
|
||||
@@ -31,7 +29,7 @@ spec:
|
||||
spec:
|
||||
containers:
|
||||
- name: n8n
|
||||
image: docker.n8n.io/n8nio/n8n:2.43.0
|
||||
image: docker.n8n.io/n8nio/n8n:2.42.3
|
||||
envFrom:
|
||||
- configMapRef:
|
||||
name: n8n-config
|
||||
|
||||
+43
-102
@@ -1,123 +1,64 @@
|
||||
# NetBird
|
||||
|
||||
Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind the repository's existing external Traefik instance on the Docker `proxy` network. Only STUN UDP `3478` is published directly.
|
||||
Self-hosted NetBird with the combined management, signal, relay, and STUN server.
|
||||
|
||||
The deployment uses SQLite for a single-instance homelab server. The persistent `netbird_data` volume and the datastore encryption key are both required to recover the installation.
|
||||
Kubernetes runs the server and dashboard in `netbird`. The server uses SQLite
|
||||
in `netbird-pvc`; `k8s/config.yaml` contains the template and runtime renderer.
|
||||
Prepare `netbird-secrets` from `k8s/secrets.yaml.example` before the first start.
|
||||
|
||||
## Files
|
||||
## Routing and keys
|
||||
|
||||
- `compose.yaml`: dashboard and combined server; selected by the marker-driven deploy workflow through `active`.
|
||||
- `config.template.yaml`: non-secret server configuration rendered at startup.
|
||||
- `entrypoint.sh`: injects Docker secrets into an in-memory runtime configuration.
|
||||
- `client.compose.yaml`: optional host-network peer using a dashboard-generated setup key.
|
||||
- `.env`: ignored local hostnames, the detected Traefik Docker-network subnet, and optional client setup key.
|
||||
- `secrets/`: ignored relay secret and datastore encryption key.
|
||||
The public hostname is set in the ConfigMap and ingress rules. Keep the issuer,
|
||||
dashboard endpoints, and public routes consistent. HTTP, WebSocket, and gRPC
|
||||
traffic go through Traefik; the STUN route uses UDP 3478. A CDN's HTTP proxy does
|
||||
not provide that UDP listener.
|
||||
|
||||
## First deployment
|
||||
`NETBIRD_PROXY_SUBNET` controls which forwarded client addresses are trusted.
|
||||
Use the actual proxy network CIDR rather than assuming another lab's subnet.
|
||||
Keep the datastore encryption key with every datastore backup. Regenerating it
|
||||
can make stored credentials unreadable.
|
||||
|
||||
Run these commands on the Docker host before merging the activating branch. The deploy preflight resets tracked files but preserves ignored local state.
|
||||
## Compose alternative
|
||||
|
||||
```bash
|
||||
cd /srv/homelab/netbird
|
||||
`compose.yaml` expects `entrypoint.sh`, a local `.env`, and two local files:
|
||||
`secrets/relay-auth-secret` and `secrets/datastore-encryption-key`.
|
||||
The reviewed main commit is missing the renderer and setup script.
|
||||
`fix/netbird-compose-runtime` restores them. Merge that fix before following
|
||||
these setup commands:
|
||||
|
||||
```sh
|
||||
cd netbird
|
||||
./setup.sh
|
||||
$EDITOR .env
|
||||
docker compose config --quiet
|
||||
docker compose up -d
|
||||
```
|
||||
|
||||
Review the values in `.env` before starting. The example public hostname is `netbird.forust.xyz`; change it if a different public domain was selected. `setup.sh` replaces `NETBIRD_PROXY_SUBNET=auto` with the first IPv4 subnet of the external Docker `proxy` network. Keep that value synchronized with the network; set an explicit CIDR instead if the network is managed elsewhere.
|
||||
|
||||
`setup.sh` is idempotent and never replaces existing secrets. Do not delete or regenerate `secrets/datastore-encryption-key` after the first successful start unless all encrypted setup keys and API tokens are intentionally being invalidated.
|
||||
|
||||
## Network prerequisites
|
||||
|
||||
- Point the public hostname directly to the Docker host. Do not proxy UDP `3478` through Cloudflare or another CDN.
|
||||
- Allow inbound TCP `80`, TCP `443`, and UDP `3478` through the host firewall and upstream router.
|
||||
- Ensure the external `proxy` Docker network exists and Traefik uses its `websecure` entrypoint and `letsencrypt` resolver. `NETBIRD_PROXY_SUBNET` must describe that network; it is used to trust only forwarded client addresses from Traefik.
|
||||
- Ensure the internal names in `.env` resolve where the local and development aliases are needed.
|
||||
- Keep Traefik's `websecure` read timeout disabled for long-lived gRPC and WebSocket sessions. This repository configures `--entrypoints.websecure.transport.respondingTimeouts.readTimeout=0` in `traefik/compose.yaml`.
|
||||
|
||||
After startup, verify OIDC discovery through the public TLS endpoint:
|
||||
|
||||
```bash
|
||||
curl -fsS "https://${NETBIRD_DOMAIN}/oauth2/.well-known/openid-configuration"
|
||||
```
|
||||
|
||||
Open `https://${NETBIRD_DOMAIN}` immediately and complete the initial owner setup. Treat the initial setup flow as public until the owner exists.
|
||||
The setup script detects the IPv4 subnet of the external `proxy` network and
|
||||
preserves existing secrets. Complete the initial owner setup through the public
|
||||
TLS endpoint after starting the server.
|
||||
|
||||
## Optional host client
|
||||
|
||||
The client intentionally lives in a separate Compose project. Normal server deploys use `--remove-orphans`, so keeping the client in the server project would cause it to be removed.
|
||||
`client.compose.yaml` runs a host-network peer in a separate Compose project.
|
||||
Set `NB_SETUP_KEY` and `NETBIRD_CLIENT_HOSTNAME` in the local `.env`, then run
|
||||
`docker compose -f client.compose.yaml up -d`. It needs `/dev/net/tun` and elevated
|
||||
network capabilities. The normal server deployment does not start this client.
|
||||
|
||||
1. Create a reusable or ephemeral setup key in the NetBird dashboard.
|
||||
2. Put `NB_SETUP_KEY=<key>` in the ignored `netbird/.env` file.
|
||||
3. Set `NETBIRD_CLIENT_HOSTNAME` to this machine's desired peer name.
|
||||
4. Start and inspect the client:
|
||||
## Backup
|
||||
|
||||
```bash
|
||||
cd /srv/homelab/netbird
|
||||
docker compose -f client.compose.yaml config --quiet
|
||||
docker compose -f client.compose.yaml up -d
|
||||
docker compose -f client.compose.yaml exec netbird-client netbird status
|
||||
Back up the SQLite data while the server is stopped, together with the encryption
|
||||
key, relay secret, and local configuration. Test a restore on an isolated host.
|
||||
For Compose, the datastore volume has the explicit name `netbird_data`.
|
||||
Do not use `docker compose down -v` when keeping the installation.
|
||||
|
||||
## Inspect
|
||||
|
||||
From the repository root:
|
||||
|
||||
```sh
|
||||
kubectl get pods,svc,pvc -n netbird
|
||||
kubectl get events -n netbird --sort-by=.metadata.creationTimestamp
|
||||
```
|
||||
|
||||
The client uses host networking and requires `NET_ADMIN`, `SYS_ADMIN`, `SYS_RESOURCE`, and `/dev/net/tun`. Remove it without affecting the server stack:
|
||||
|
||||
```bash
|
||||
docker compose -f client.compose.yaml down
|
||||
```
|
||||
|
||||
## Operations
|
||||
|
||||
Inspect status and logs:
|
||||
|
||||
```bash
|
||||
docker compose ps
|
||||
docker compose logs --tail=200 netbird-server dashboard
|
||||
```
|
||||
|
||||
Stop or remove containers without deleting data:
|
||||
|
||||
```bash
|
||||
docker compose down
|
||||
```
|
||||
|
||||
Do not add `-v` to `docker compose down`; it would delete the NetBird datastore.
|
||||
|
||||
## Backup and restore
|
||||
|
||||
Back up both the persistent volume and the ignored secret files. For a consistent SQLite backup, briefly stop the server first and store the resulting archive and `datastore-encryption-key` in an encrypted backup:
|
||||
|
||||
```bash
|
||||
cd /srv/homelab/netbird
|
||||
mkdir -p backups
|
||||
docker compose stop netbird-server
|
||||
docker run --rm \
|
||||
-v netbird_data:/data:ro \
|
||||
-v "$PWD/backups:/backup" \
|
||||
busybox:1.37.0 \
|
||||
tar -C /data -czf "/backup/netbird-data-$(date -u +%Y%m%dT%H%M%SZ).tar.gz" .
|
||||
docker compose start netbird-server
|
||||
```
|
||||
|
||||
Also securely back up:
|
||||
|
||||
- `secrets/datastore-encryption-key` — required to decrypt stored secrets.
|
||||
- `secrets/relay-auth-secret` — keeps issued relay credentials valid across restoration.
|
||||
- `netbird/.env` — optional, but it records the public and internal hostnames.
|
||||
|
||||
Test a restore in an isolated Docker host before relying on a backup.
|
||||
|
||||
## Upgrade
|
||||
|
||||
1. Take and verify a backup.
|
||||
2. Review NetBird release notes for server, client, and dashboard compatibility.
|
||||
3. Update the pinned tags in `compose.yaml`; update `client.compose.yaml` separately when deploying the client.
|
||||
4. Pull and recreate the selected services:
|
||||
|
||||
```bash
|
||||
docker compose pull
|
||||
docker compose up -d
|
||||
```
|
||||
|
||||
The image tags are intentionally pinned instead of using `latest`, matching this repository's pull-on-deploy policy.
|
||||
See the [repository README](../README.md) for deployment selection.
|
||||
Loaded 100 of 163 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user