Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6a3eec0164 |
No files matched your search
@@ -1,50 +0,0 @@
|
|||||||
# EDU ownership handoff
|
|
||||||
|
|
||||||
## Status
|
|
||||||
|
|
||||||
The EDU ownership handoff is complete. The homelab repository no longer owns
|
|
||||||
EDU workloads, images, routes, alerts, or deployment selection. The EDU
|
|
||||||
repository is the only deployment owner: [forust/edu-master](https://git.forust.xyz/forust/edu-master).
|
|
||||||
|
|
||||||
Homelab PRs #99 and #105 are merged. PR #105 removed the EDU subtree and its
|
|
||||||
build, deploy, rollback, verification, route-probe, and registry references.
|
|
||||||
It also added the serial image build matrix for the homelab services. This
|
|
||||||
handoff record is the only remaining EDU-specific file in homelab Git.
|
|
||||||
|
|
||||||
The dedicated workstation checkout is `/srv/edu-master`, at release
|
|
||||||
`4f2b2a0e37dc11ac2c75441a15076c178e219d37`. It contains `k8s/active`; root
|
|
||||||
`active` is absent. The old untracked `/srv/homelab/edu_master` checkout was
|
|
||||||
moved outside the homelab repository to
|
|
||||||
`/srv/edu-master-legacy-archive-20261007/edu_master`. Its private files remain
|
|
||||||
mode `0600` inside an archive directory with mode `0700`. The homelab deploy
|
|
||||||
checkout has no EDU marker or tracked EDU application/deployment files.
|
|
||||||
`AUTODEPLOY=false` remains in place for homelab deployment.
|
|
||||||
|
|
||||||
## Release evidence
|
|
||||||
|
|
||||||
EDU PR #4 merged after its review and CI checks. Main-push CI run 1652 passed
|
|
||||||
all validation and both image builds. Deploy run 1653 passed for the exact main
|
|
||||||
SHA above.
|
|
||||||
|
|
||||||
The workstation rollout completed for both Deployments. The deployment
|
|
||||||
verified `/health` and `/live` with HTTP 200, Redis AUTH, session TTL of 1058
|
|
||||||
seconds, a delivery backlog of zero, and all nine EDU vmalert rules with
|
|
||||||
matching expressions and healthy evaluation.
|
|
||||||
|
|
||||||
The images now run by digest:
|
|
||||||
|
|
||||||
- Session keeper: `sha256:998dea51aa3015fd9cabefb0f53b030157a650c3bef72e02fe84f17d5762613d`
|
|
||||||
- Webinar checker: `sha256:92f3c1fa2bb7f9b4680a9fc76a5b33dfbea8ef3dd9c6490ebc45876fd4c54461`
|
|
||||||
|
|
||||||
Redis StatefulSet was unchanged. PVC `redis-data-pvc` remains bound to PV
|
|
||||||
`pvc-a4f2a79a-363a-4c12-ae91-92cdfc2a0d2e` with capacity 1 GiB. The existing
|
|
||||||
runtime Secret and Fernet key were preserved during the handoff. Notification
|
|
||||||
delivery was verified before closeout, as confirmed by the operator. The
|
|
||||||
deployment did not record downtime.
|
|
||||||
|
|
||||||
The release rollback snapshot is
|
|
||||||
`/home/forust/.local/state/edu-master-deploy/20261007T180541Z-4f2b2a0e37dc11ac2c75441a15076c178e219d37`.
|
|
||||||
The handoff data snapshot remains at
|
|
||||||
`/home/forust/.local/state/edu-master-deploy/handoff-20261007T080838Z`.
|
|
||||||
Both snapshots are outside Git. Do not restore old Redis data unless recovery
|
|
||||||
requires it. Never delete or recreate the Redis PVC.
|
|
||||||
@@ -1,64 +0,0 @@
|
|||||||
# CI and deployment
|
|
||||||
|
|
||||||
Gitea Actions validates changes, builds the repository's custom images, and can
|
|
||||||
deploy selected services to the workstation. CI and production deployment use
|
|
||||||
separate workflows. See the [runner and recovery guide](runner/README.md) for
|
|
||||||
installation, configuration, and operator commands.
|
|
||||||
|
|
||||||
## CI
|
|
||||||
|
|
||||||
`workflows/ci.yaml` runs Compose, workflow, shell, formatting, Python and unit
|
|
||||||
test, YAML, Dockerfile, and Kubernetes checks. Pull requests and non-main refs
|
|
||||||
use the unprivileged `homelab-pr` runner. Main-branch CI uses `homelab`. Tool
|
|
||||||
versions are pinned in `workflows/tool-versions.env`.
|
|
||||||
|
|
||||||
Compose CI checks every committed Compose file without requiring ignored `.env`
|
|
||||||
files. Kubernetes checks validate known schemas; unknown CRDs are skipped.
|
|
||||||
|
|
||||||
On main, CI plans builds for the three owned images: `error-pages`,
|
|
||||||
`forust-homepage`, and `xdfnx-homepage`. It builds changed inputs or reuses a
|
|
||||||
digest from a successful earlier main run. The successful build job publishes a
|
|
||||||
release artifact for the exact commit SHA. Pull requests do not publish images.
|
|
||||||
|
|
||||||
## Deployment gate
|
|
||||||
|
|
||||||
`workflows/deploy.yaml` starts a deployment after successful main CI when the
|
|
||||||
`AUTODEPLOY` Actions variable is `true`. Manual dispatch uses the same gate: the
|
|
||||||
requested `main` ref or commit must have successful main CI and its matching
|
|
||||||
release artifact. A manual dispatch does not bypass validation.
|
|
||||||
|
|
||||||
The workflow supports these modes:
|
|
||||||
|
|
||||||
- `changed`: select active services changed since the last successful deploy.
|
|
||||||
- `full`: select all active services; use this for the first baseline.
|
|
||||||
- `plan`: validate and show the selection without applying production resources.
|
|
||||||
|
|
||||||
`refresh_images=true` explicitly refreshes mutable third-party Compose tags.
|
|
||||||
|
|
||||||
## Selection and rollout
|
|
||||||
|
|
||||||
The active markers define automatic deployment. `<service>/active` selects a
|
|
||||||
standard Compose file; `<service>/k8s/active` selects Kubernetes resources. Helm
|
|
||||||
releases have their own markers in `workflows/deploy-lib.sh`. Service
|
|
||||||
dependencies are declared in `deploy-dependencies.json`. Removed resources are
|
|
||||||
reported for manual review; the workflow does not prune them automatically.
|
|
||||||
|
|
||||||
The workstation controller runs the checked source in a per-SHA worktree. It
|
|
||||||
validates configuration, applies Kubernetes and Compose changes in sequence,
|
|
||||||
verifies changed Kubernetes workloads, and checks public routes. A durable
|
|
||||||
systemd service continues the rollout if the Actions SSH client disconnects.
|
|
||||||
The workflow checks the exact CI release before it submits a deployment.
|
|
||||||
|
|
||||||
Kubernetes recovery uses captured workload revisions. It does not restore
|
|
||||||
ConfigMaps, Secrets, database schemas, or persistent data. Compose recovery is
|
|
||||||
manual and does not restore volume data or reverse migrations. Keep backups for
|
|
||||||
stateful services. The runner guide documents status, retry, logs, and recovery
|
|
||||||
commands.
|
|
||||||
|
|
||||||
## Settings
|
|
||||||
|
|
||||||
Configure `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`, and the verified
|
|
||||||
`DEPLOY_KNOWN_HOSTS` entry as Actions variables. Keep `DEPLOY_SSH_KEY`,
|
|
||||||
`REGISTRY_USERNAME`, and `REGISTRY_PASSWORD` in Actions secrets. The workstation
|
|
||||||
also needs its existing registry authentication. Set `AUTODEPLOY=false` until
|
|
||||||
automatic production deploys are intended.
|
|
||||||
@@ -7,5 +7,4 @@ self-hosted-runner:
|
|||||||
labels:
|
labels:
|
||||||
- arch
|
- arch
|
||||||
- homelab
|
- homelab
|
||||||
- homelab-pr
|
|
||||||
- prod
|
- prod
|
||||||
+5
-64
@@ -1,20 +1,8 @@
|
|||||||
# Homelab CI/CD
|
# Homelab CI/CD
|
||||||
|
|
||||||
The native Gitea runners run on **vps**; production runs on **workstation**.
|
The native Gitea runner runs on **vps**; production runs on **workstation**.
|
||||||
Main-branch checks and image builds use `homelab:host`. Pull request and
|
Jobs run on `homelab:host`, one at a time. No job images or Kubernetes credentials
|
||||||
non-main checks use `homelab-pr:host` under a separate account without Docker
|
are needed on the VPS. Builds use one pinned BuildKit helper container. CI and deploy are separate workflows.
|
||||||
access. The `homelab-pr` runner is registered at User scope for `forust`, so
|
|
||||||
any repository under that account can schedule jobs that request this label.
|
|
||||||
Each runner accepts one job at a time; the build waits for every check to pass.
|
|
||||||
CI and deploy runs also show a summary with
|
|
||||||
the release SHA, image build or reuse results, deploy mode, selected services,
|
|
||||||
and image digests. Failed runs keep a summary of completed image builds, stage
|
|
||||||
results, apply results, and recorded Kubernetes recovery. The final deploy
|
|
||||||
summary is in the smoke job; earlier jobs show the state observed at that time.
|
|
||||||
Apply success is separate from health and recovery. Update the installed
|
|
||||||
workstation controller with `setup-workstation.sh` when no deploy is running.
|
|
||||||
No job images or Kubernetes credentials are needed on the VPS. Builds use one
|
|
||||||
pinned BuildKit helper container. CI and deploy are separate workflows.
|
|
||||||
|
|
||||||
## Runner installation
|
## Runner installation
|
||||||
|
|
||||||
@@ -40,32 +28,6 @@ pushes directly to the registry, and caps retained local cache at 1 GiB with a
|
|||||||
2 GiB free-space target. This is not a hard limit on peak build disk usage.
|
2 GiB free-space target. This is not a hard limit on peak build disk usage.
|
||||||
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
|
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
|
||||||
|
|
||||||
### Pull request runner
|
|
||||||
|
|
||||||
Install the unprivileged host runner on the VPS:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
sudo bash .gitea/runner/setup-pr-runner.sh
|
|
||||||
```
|
|
||||||
|
|
||||||
Get a registration token from the user Actions runner settings. Run the
|
|
||||||
installer in a terminal. It asks for the token without echoing it, registers the
|
|
||||||
runner as `homelab-pr` with label `homelab-pr:host`, then enables the service.
|
|
||||||
The work directory is `/var/lib/gitea-pr-runner`. Confirm that Gitea lists the
|
|
||||||
runner as User scope before merging the workflow change. An unmatched label can
|
|
||||||
fall back to the default job image.
|
|
||||||
|
|
||||||
Renovate PR validation uses `pull_request_target`, which reads the workflow from
|
|
||||||
the base branch. It checks out the PR head only after runner selection and runs
|
|
||||||
that code on `homelab-pr`. Keep this workflow read-only and do not add secrets.
|
|
||||||
|
|
||||||
The PR runner has a separate home and tool cache. Do not add it to the `docker`
|
|
||||||
group or give it access to `/var/run/docker.sock`. It runs repository code from
|
|
||||||
pull requests, so keep its registration and permissions separate from the
|
|
||||||
trusted `homelab` runner. This separates users and host permissions, but both
|
|
||||||
runners still share the VPS kernel and network. Use a disposable VM if PRs from
|
|
||||||
untrusted external authors must be fully isolated.
|
|
||||||
|
|
||||||
## Workstation setup
|
## Workstation setup
|
||||||
|
|
||||||
As the existing SSH deploy user on workstation:
|
As the existing SSH deploy user on workstation:
|
||||||
@@ -96,7 +58,8 @@ The deploy user's existing Docker registry authentication remains necessary.
|
|||||||
|
|
||||||
CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image
|
CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image
|
||||||
digests and build input fingerprints. Unchanged images are reused only from a
|
digests and build input fingerprints. Unchanged images are reused only from a
|
||||||
successful main CI artifact, never from `:prod`. Expired artifacts cause CI to
|
successful main CI artifact, never from `:prod`. EDU images remain pinned to the
|
||||||
|
digests released by their application repository. Expired artifacts cause CI to
|
||||||
rebuild images; they block deployment until CI is rerun.
|
rebuild images; they block deployment until CI is rerun.
|
||||||
|
|
||||||
Run deploy from main with `deploy_ref=main` or a checked SHA:
|
Run deploy from main with `deploy_ref=main` or a checked SHA:
|
||||||
@@ -157,25 +120,3 @@ run first. Restore the runner config/unit from `.before-<timestamp>` backups,
|
|||||||
reload systemd and restart the runner. Restore the prior workflows from Git.
|
reload systemd and restart the runner. Restore the prior workflows from Git.
|
||||||
Production data and persistent volumes stay where they were. Do not remove run
|
Production data and persistent volumes stay where they were. Do not remove run
|
||||||
state or Compose recovery files until recovery is confirmed.
|
state or Compose recovery files until recovery is confirmed.
|
||||||
|
|
||||||
### Compose configuration recovery
|
|
||||||
|
|
||||||
Successful deploys save the complete resolved Compose configuration in
|
|
||||||
`~/.local/state/homelab-deploy/compose-configs/`. These files can contain secrets.
|
|
||||||
Keep them private and do not commit or upload them.
|
|
||||||
The next deploy uses this configuration for its recovery file, including old
|
|
||||||
commands, environment, mounts, ports, and removed services. The recovery command
|
|
||||||
uses `--remove-orphans` to remove services added by the failed deploy. It does
|
|
||||||
not restore volume data or reverse database migrations.
|
|
||||||
|
|
||||||
On the first run after this update, the controller can use the Compose file
|
|
||||||
from the previous successful run. If that file is absent, it reads the persistent
|
|
||||||
checkout and checks its service configuration hashes against existing containers.
|
|
||||||
A mismatch stops preflight. Restore the previous configuration before retrying.
|
|
||||||
Update the installed controller with `bash .gitea/runner/setup-workstation.sh`
|
|
||||||
from the reviewed checkout before using this change.
|
|
||||||
|
|
||||||
New namespaces are checked during preflight. Server validation of their resources
|
|
||||||
runs after namespace creation and before application resources are applied.
|
|
||||||
Plan mode does not create namespaces. A failed deferred check can leave an empty
|
|
||||||
namespace; inspect it before removing it.
|
|
||||||
@@ -1,8 +0,0 @@
|
|||||||
runner:
|
|
||||||
file: /var/lib/gitea-pr-runner/.runner
|
|
||||||
capacity: 1
|
|
||||||
timeout: 5h
|
|
||||||
labels:
|
|
||||||
- homelab-pr:host
|
|
||||||
cache:
|
|
||||||
enabled: false
|
|
||||||
@@ -1,27 +0,0 @@
|
|||||||
[Unit]
|
|
||||||
Description=Gitea Actions untrusted pull request runner
|
|
||||||
After=network-online.target
|
|
||||||
Wants=network-online.target
|
|
||||||
|
|
||||||
[Service]
|
|
||||||
User=gitea-pr-runner
|
|
||||||
Group=gitea-pr-runner
|
|
||||||
WorkingDirectory=/var/lib/gitea-pr-runner
|
|
||||||
Environment=HOME=/var/lib/gitea-pr-runner
|
|
||||||
Environment=PATH=/var/lib/gitea-pr-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
|
|
||||||
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-pr-runner/config.yaml
|
|
||||||
Restart=on-failure
|
|
||||||
RestartSec=5
|
|
||||||
NoNewPrivileges=yes
|
|
||||||
PrivateTmp=yes
|
|
||||||
ProtectSystem=full
|
|
||||||
ProtectHome=yes
|
|
||||||
ProtectKernelTunables=yes
|
|
||||||
ProtectKernelModules=yes
|
|
||||||
ProtectControlGroups=yes
|
|
||||||
RestrictSUIDSGID=yes
|
|
||||||
LockPersonality=yes
|
|
||||||
UMask=0077
|
|
||||||
|
|
||||||
[Install]
|
|
||||||
WantedBy=multi-user.target
|
|
||||||
@@ -1,56 +0,0 @@
|
|||||||
#!/usr/bin/env bash
|
|
||||||
# Install a native runner for untrusted PR jobs without Docker access.
|
|
||||||
set -euo pipefail
|
|
||||||
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
||||||
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
|
|
||||||
for tool in cp cut date getent id install runuser systemctl useradd; do
|
|
||||||
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
|
|
||||||
done
|
|
||||||
command -v /usr/local/bin/gitea-runner >/dev/null || {
|
|
||||||
echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2
|
|
||||||
exit 1
|
|
||||||
}
|
|
||||||
|
|
||||||
id gitea-pr-runner >/dev/null 2>&1 || \
|
|
||||||
useradd --system --create-home --home-dir /var/lib/gitea-pr-runner --shell /usr/bin/bash gitea-pr-runner
|
|
||||||
runner_home="$(getent passwd gitea-pr-runner | cut -d: -f6)"
|
|
||||||
[ "$runner_home" = /var/lib/gitea-pr-runner ] || {
|
|
||||||
echo 'Unexpected PR runner home; inspect the existing service first' >&2
|
|
||||||
exit 1
|
|
||||||
}
|
|
||||||
case " $(id -nG gitea-pr-runner) " in
|
|
||||||
*' docker '*)
|
|
||||||
echo 'The PR runner account must not belong to the docker group' >&2
|
|
||||||
exit 1
|
|
||||||
;;
|
|
||||||
esac
|
|
||||||
|
|
||||||
install -d -m 0755 /etc/gitea-pr-runner
|
|
||||||
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
|
|
||||||
for existing in /etc/gitea-pr-runner/config.yaml /etc/systemd/system/gitea-pr-runner.service; do
|
|
||||||
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
|
|
||||||
done
|
|
||||||
install -m 0644 "$here/pr-config.yaml" /etc/gitea-pr-runner/config.yaml
|
|
||||||
install -m 0644 "$here/pr-runner.service" /etc/systemd/system/gitea-pr-runner.service
|
|
||||||
|
|
||||||
if [ ! -f /var/lib/gitea-pr-runner/.runner ]; then
|
|
||||||
read -r -s -p 'Enter the Gitea repository runner registration token: ' runner_token
|
|
||||||
printf '\n'
|
|
||||||
[ -n "$runner_token" ] || { echo 'Runner token is required' >&2; exit 1; }
|
|
||||||
export GITEA_RUNNER_REGISTRATION_TOKEN="$runner_token"
|
|
||||||
unset runner_token
|
|
||||||
runuser --preserve-environment -u gitea-pr-runner -- \
|
|
||||||
/usr/local/bin/gitea-runner register \
|
|
||||||
--config /etc/gitea-pr-runner/config.yaml \
|
|
||||||
--instance https://gitea.forust.xyz \
|
|
||||||
--name homelab-pr \
|
|
||||||
--labels homelab-pr:host \
|
|
||||||
--no-interactive
|
|
||||||
unset GITEA_RUNNER_REGISTRATION_TOKEN
|
|
||||||
fi
|
|
||||||
chmod 0600 /var/lib/gitea-pr-runner/.runner
|
|
||||||
|
|
||||||
systemctl daemon-reload
|
|
||||||
systemctl enable --now gitea-pr-runner.service
|
|
||||||
systemctl restart gitea-pr-runner.service
|
|
||||||
echo "PR runner ready. Configuration backups: *.before-$stamp"
|
|
||||||
@@ -91,66 +91,4 @@ if check_referenced_secrets >"$scratch/secrets.log"; then
|
|||||||
echo 'Secret check accepted a failed manifest render' >&2
|
echo 'Secret check accepted a failed manifest render' >&2
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
|
|
||||||
# New declared namespaces defer only their own resources during preflight.
|
|
||||||
render_selected_resources() {
|
|
||||||
cat <<'JSON'
|
|
||||||
{"apiVersion":"v1","kind":"List","items":[
|
|
||||||
{"apiVersion":"v1","kind":"Namespace","metadata":{"name":"new"}},
|
|
||||||
{"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"new-config","namespace":"new"}},
|
|
||||||
{"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"existing-config","namespace":"default"}}
|
|
||||||
]}
|
|
||||||
JSON
|
|
||||||
}
|
|
||||||
kubectl() {
|
|
||||||
case "$1" in
|
|
||||||
get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}}]}' ;;
|
|
||||||
apply) cat >"$scratch/server-input.json" ;;
|
|
||||||
*) return 1 ;;
|
|
||||||
esac
|
|
||||||
}
|
|
||||||
validate_server_resources true
|
|
||||||
jq -e '.items | length == 2 and all(.metadata.name != "new-config")' "$scratch/server-input.json" >/dev/null
|
|
||||||
if validate_server_resources false 2>"$scratch/deferred.log"; then
|
|
||||||
echo 'Post-namespace validation accepted a missing namespace' >&2
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
kubectl() {
|
|
||||||
case "$1" in
|
|
||||||
get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}},{"metadata":{"name":"new"}}]}' ;;
|
|
||||||
apply) cat >"$scratch/server-input.json" ;;
|
|
||||||
*) return 1 ;;
|
|
||||||
esac
|
|
||||||
}
|
|
||||||
validate_server_resources false
|
|
||||||
jq -e '.items | length == 3' "$scratch/server-input.json" >/dev/null
|
|
||||||
render_selected_resources() {
|
|
||||||
printf '%s\n' '{"items":[{"kind":"ConfigMap","metadata":{"name":"bad","namespace":"undeclared"}}]}'
|
|
||||||
}
|
|
||||||
if validate_server_resources true 2>"$scratch/undeclared.log"; then
|
|
||||||
echo 'Preflight accepted an undeclared missing namespace' >&2
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
# Count services, not characters in the newline-separated service names.
|
|
||||||
compose() {
|
|
||||||
case "$*" in
|
|
||||||
*'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{},"init":{"restart":"no"}}}' ;;
|
|
||||||
*'ps --status running --services') printf '%s\n' headscale headplane web ;;
|
|
||||||
*) return 1 ;;
|
|
||||||
esac
|
|
||||||
}
|
|
||||||
verify_compose_stack example.yaml >"$scratch/compose-count.log"
|
|
||||||
grep -qF 'all 3 service(s) running' "$scratch/compose-count.log"
|
|
||||||
compose() {
|
|
||||||
case "$*" in
|
|
||||||
*'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{}}}' ;;
|
|
||||||
*'ps --status running --services') printf '%s\n' headscale headplane ;;
|
|
||||||
*) return 0 ;;
|
|
||||||
esac
|
|
||||||
}
|
|
||||||
if verify_compose_stack example.yaml >"$scratch/compose-missing.log"; then
|
|
||||||
echo 'Compose verification accepted a missing service' >&2
|
|
||||||
exit 1
|
|
||||||
fi
|
|
||||||
grep -qF 'NOT RUNNING: web' "$scratch/compose-missing.log"
|
|
||||||
printf '%s\n' 'Deploy validation regressions passed.'
|
printf '%s\n' 'Deploy validation regressions passed.'
|
||||||
+16
-364
@@ -12,14 +12,18 @@ concurrency:
|
|||||||
group: ci-${{ github.ref }}
|
group: ci-${{ github.ref }}
|
||||||
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
|
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
|
||||||
jobs:
|
jobs:
|
||||||
compose:
|
checks:
|
||||||
name: Compose
|
runs-on: homelab
|
||||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
timeout-minutes: 30
|
||||||
timeout-minutes: 15
|
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
id: source
|
- name: Prepare pinned tools
|
||||||
|
shell: bash
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh)"
|
||||||
|
echo "$tools_dir" >> "$GITHUB_PATH"
|
||||||
- name: Validate Compose files
|
- name: Validate Compose files
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
@@ -48,73 +52,11 @@ jobs:
|
|||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
echo "checked ${#files[@]} Compose file(s)"
|
echo "checked ${#files[@]} Compose file(s)"
|
||||||
id: check
|
|
||||||
- name: Write the job result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: Compose
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP:
|
|
||||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.source.conclusion == 'failure'
|
|
||||||
&& 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
workflows:
|
|
||||||
name: Workflows
|
|
||||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
|
||||||
timeout-minutes: 15
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
||||||
id: source
|
|
||||||
- name: Prepare pinned tools
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint shellcheck)"
|
|
||||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
|
||||||
id: tools
|
|
||||||
- name: Lint Gitea Actions workflows with actionlint
|
- name: Lint Gitea Actions workflows with actionlint
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
|
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
|
||||||
id: check
|
|
||||||
- name: Write the job result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: Workflows
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP:
|
|
||||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
|
||||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
shell:
|
|
||||||
name: Shell
|
|
||||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
|
||||||
timeout-minutes: 15
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
||||||
id: source
|
|
||||||
- name: Prepare pinned tools
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck jq)"
|
|
||||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
|
||||||
id: tools
|
|
||||||
- name: Lint shell scripts with ShellCheck
|
- name: Lint shell scripts with ShellCheck
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
@@ -128,37 +70,6 @@ jobs:
|
|||||||
fi
|
fi
|
||||||
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
|
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
|
||||||
bash .gitea/tests/deploy-validation.sh
|
bash .gitea/tests/deploy-validation.sh
|
||||||
id: check
|
|
||||||
- name: Write the job result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: Shell
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP:
|
|
||||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
|
||||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
formatting:
|
|
||||||
name: Formatting
|
|
||||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
|
||||||
timeout-minutes: 15
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
||||||
id: source
|
|
||||||
- name: Prepare pinned tools
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
|
|
||||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
|
||||||
id: tools
|
|
||||||
- name: Check formatting with Prettier
|
- name: Check formatting with Prettier
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
@@ -176,37 +87,6 @@ jobs:
|
|||||||
fi
|
fi
|
||||||
|
|
||||||
prettier --check --ignore-unknown "${prettier_files[@]}"
|
prettier --check --ignore-unknown "${prettier_files[@]}"
|
||||||
id: check
|
|
||||||
- name: Write the job result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: Formatting
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP:
|
|
||||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
|
||||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
python:
|
|
||||||
name: Python and tests
|
|
||||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
|
||||||
timeout-minutes: 15
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
||||||
id: source
|
|
||||||
- name: Prepare pinned tools
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff jq)"
|
|
||||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
|
||||||
id: tools
|
|
||||||
- name: Lint and format-check Python with Ruff
|
- name: Lint and format-check Python with Ruff
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
@@ -214,37 +94,6 @@ jobs:
|
|||||||
ruff check . .gitea/workflows
|
ruff check . .gitea/workflows
|
||||||
ruff format --check . .gitea/workflows
|
ruff format --check . .gitea/workflows
|
||||||
python3 -m unittest discover -s tests -v
|
python3 -m unittest discover -s tests -v
|
||||||
id: check
|
|
||||||
- name: Write the job result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: Python and tests
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP:
|
|
||||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
|
||||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
yaml:
|
|
||||||
name: YAML
|
|
||||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
|
||||||
timeout-minutes: 15
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
||||||
id: source
|
|
||||||
- name: Prepare pinned tools
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
|
|
||||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
|
||||||
id: tools
|
|
||||||
- name: Lint YAML syntax
|
- name: Lint YAML syntax
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
@@ -262,37 +111,6 @@ jobs:
|
|||||||
fi
|
fi
|
||||||
|
|
||||||
yamllint -c .yamllint "${yaml_files[@]}"
|
yamllint -c .yamllint "${yaml_files[@]}"
|
||||||
id: check
|
|
||||||
- name: Write the job result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: YAML
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP:
|
|
||||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
|
||||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
dockerfiles:
|
|
||||||
name: Dockerfiles
|
|
||||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
|
||||||
timeout-minutes: 15
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
||||||
id: source
|
|
||||||
- name: Prepare pinned tools
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
|
|
||||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
|
||||||
id: tools
|
|
||||||
- name: Lint Dockerfiles
|
- name: Lint Dockerfiles
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
@@ -308,37 +126,6 @@ jobs:
|
|||||||
fi
|
fi
|
||||||
|
|
||||||
hadolint -c .hadolint.yaml "${dockerfiles[@]}"
|
hadolint -c .hadolint.yaml "${dockerfiles[@]}"
|
||||||
id: check
|
|
||||||
- name: Write the job result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: Dockerfiles
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP:
|
|
||||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
|
||||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
kubernetes:
|
|
||||||
name: Kubernetes
|
|
||||||
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
|
||||||
timeout-minutes: 15
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
||||||
id: source
|
|
||||||
- name: Prepare pinned tools
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
|
|
||||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
|
||||||
id: tools
|
|
||||||
- name: Validate Kubernetes manifests against JSON schemas
|
- name: Validate Kubernetes manifests against JSON schemas
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
@@ -359,162 +146,27 @@ jobs:
|
|||||||
-ignore-missing-schemas \
|
-ignore-missing-schemas \
|
||||||
-summary \
|
-summary \
|
||||||
"${manifests[@]}"
|
"${manifests[@]}"
|
||||||
id: check
|
build:
|
||||||
- name: Write the job result
|
needs:
|
||||||
if: always()
|
- checks
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: Kubernetes
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP:
|
|
||||||
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
|
|
||||||
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
image-plan:
|
|
||||||
needs: [compose, workflows, shell, formatting, python, yaml, dockerfiles, kubernetes]
|
|
||||||
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
|
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
|
||||||
runs-on: homelab
|
runs-on: homelab
|
||||||
timeout-minutes: 10
|
timeout-minutes: 60
|
||||||
outputs:
|
|
||||||
matrix: ${{ steps.plan.outputs.matrix }}
|
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
id: source
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
||||||
with:
|
with:
|
||||||
fetch-depth: 0
|
fetch-depth: 0
|
||||||
- name: Detect build inputs against successful CI
|
- name: Build changed images and write release
|
||||||
id: plan
|
|
||||||
env:
|
env:
|
||||||
GITEA_TOKEN: ${{ github.token }}
|
GITEA_TOKEN: ${{ github.token }}
|
||||||
run: python3 .gitea/workflows/release.py prepare --output build-plan.json
|
|
||||||
- name: Store the image plan
|
|
||||||
id: artifact
|
|
||||||
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
|
||||||
with:
|
|
||||||
name: build-plan
|
|
||||||
path: build-plan.json
|
|
||||||
if-no-files-found: error
|
|
||||||
retention-days: 30
|
|
||||||
|
|
||||||
- name: Write the plan result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: Image plan
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP: >-
|
|
||||||
${{ steps.plan.conclusion == 'failure' && 'Build input detection' ||
|
|
||||||
steps.artifact.conclusion == 'failure' && 'Plan upload' ||
|
|
||||||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## Image plan\n\nResult: %s\n' "$SUMMARY_RESULT" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
|
|
||||||
images:
|
|
||||||
name: Image (${{ matrix.name }})
|
|
||||||
needs: [image-plan]
|
|
||||||
if: needs.image-plan.result == 'success'
|
|
||||||
runs-on: homelab
|
|
||||||
timeout-minutes: 60
|
|
||||||
strategy:
|
|
||||||
max-parallel: 1
|
|
||||||
fail-fast: false
|
|
||||||
matrix: ${{ fromJSON(needs.image-plan.outputs.matrix || '{"include":[{"name":"inactive"}]}') }}
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
id: source
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
||||||
- name: Download the checked image plan
|
|
||||||
id: inputs
|
|
||||||
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
|
|
||||||
with:
|
|
||||||
name: build-plan
|
|
||||||
- name: Build or reuse this image
|
|
||||||
id: check
|
|
||||||
env:
|
|
||||||
IMAGE_NAME: ${{ matrix.name }}
|
|
||||||
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
|
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
|
||||||
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
|
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
|
||||||
run: python3 .gitea/workflows/release.py image --image "$IMAGE_NAME" --output image.json
|
run: python3 .gitea/workflows/release.py build
|
||||||
- name: Store the image result
|
|
||||||
id: artifact
|
|
||||||
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
|
||||||
with:
|
|
||||||
name: image-${{ matrix.name }}
|
|
||||||
path: image.json
|
|
||||||
if-no-files-found: error
|
|
||||||
retention-days: 30
|
|
||||||
- name: Write the job result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: Image (${{ matrix.name }})
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP: >-
|
|
||||||
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
|
|
||||||
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
|
|
||||||
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
|
|
||||||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
|
|
||||||
# Retain the build job name required by the immutable release deployment gate.
|
|
||||||
build:
|
|
||||||
needs: [image-plan, images]
|
|
||||||
runs-on: homelab
|
|
||||||
timeout-minutes: 15
|
|
||||||
steps:
|
|
||||||
- name: Checkout repository
|
|
||||||
id: source
|
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
||||||
- name: Download all image results
|
|
||||||
id: inputs
|
|
||||||
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
|
|
||||||
with:
|
|
||||||
path: artifacts
|
|
||||||
- name: Pin SHA tags and write the complete release
|
|
||||||
id: check
|
|
||||||
env:
|
|
||||||
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
|
|
||||||
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
|
|
||||||
run: >-
|
|
||||||
python3 .gitea/workflows/release.py finalize
|
|
||||||
--plan artifacts/build-plan/build-plan.json
|
|
||||||
- name: Store commit release
|
- name: Store commit release
|
||||||
id: artifact
|
uses: actions/upload-artifact@c6a366c94c3e0affe28c06c8df20a878f24da3cf
|
||||||
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
|
||||||
with:
|
with:
|
||||||
name: release-${{ github.sha }}
|
name: release-${{ github.sha }}
|
||||||
path: release.json
|
path: release.json
|
||||||
if-no-files-found: error
|
if-no-files-found: error
|
||||||
retention-days: 30
|
retention-days: 30
|
||||||
- name: Write the job result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
SUMMARY_CHECK: Image release and SHA tags
|
|
||||||
SUMMARY_RESULT: ${{ job.status }}
|
|
||||||
SUMMARY_FAILED_STEP: >-
|
|
||||||
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
|
|
||||||
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
|
|
||||||
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
|
|
||||||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/release.py ]; then
|
|
||||||
python3 .gitea/workflows/release.py check-summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
@@ -42,50 +42,15 @@ def prepare(source_file):
|
|||||||
images_file = directory / 'compose-images.json'
|
images_file = directory / 'compose-images.json'
|
||||||
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
|
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
|
||||||
release = json.loads((directory / 'release.json').read_text())
|
release = json.loads((directory / 'release.json').read_text())
|
||||||
state = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
|
before = json.loads(json.dumps(config))
|
||||||
baseline = state / 'compose-configs' / f'{relative.parent.name}.json'
|
|
||||||
if not baseline.exists() and re.fullmatch(r'[0-9]+-[0-9]+', previous.get('run_id', '')):
|
|
||||||
baseline = state / 'runs' / previous['run_id'] / 'compose' / baseline.name
|
|
||||||
bootstrap = not baseline.exists()
|
|
||||||
if not bootstrap:
|
|
||||||
before = json.loads(baseline.read_text())
|
|
||||||
else:
|
|
||||||
# Bootstrap from the persistent configuration, never from the new source.
|
|
||||||
persistent_file = config_repo / relative
|
|
||||||
if persistent_file.exists():
|
|
||||||
before = json.loads(
|
|
||||||
output(
|
|
||||||
'docker',
|
|
||||||
'compose',
|
|
||||||
'--project-directory',
|
|
||||||
str(project_dir),
|
|
||||||
'-f',
|
|
||||||
str(persistent_file),
|
|
||||||
'config',
|
|
||||||
'--format',
|
|
||||||
'json',
|
|
||||||
cwd=config_repo,
|
|
||||||
)
|
|
||||||
)
|
|
||||||
elif output('docker', 'ps', '-aq', '--filter', f'label=com.docker.compose.project={project}'):
|
|
||||||
raise ValueError(f'{project}: no previous Compose configuration; restore it before deploy')
|
|
||||||
else:
|
|
||||||
before = {'name': project, 'services': {}}
|
|
||||||
if before['name'] != project:
|
|
||||||
raise ValueError('Compose project name changed; manual migration is required')
|
|
||||||
for service, settings in config['services'].items():
|
for service, settings in config['services'].items():
|
||||||
reference = settings.get('image')
|
reference = settings.get('image')
|
||||||
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
|
|
||||||
if not reference or settings.get('build'):
|
if not reference or settings.get('build'):
|
||||||
raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
|
raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
|
||||||
image_repo = reference.split('@')[0].rsplit('/', 1)
|
image_repo = reference.split('@')[0].rsplit('/', 1)
|
||||||
image_repo[-1] = image_repo[-1].split(':')[0]
|
image_repo[-1] = image_repo[-1].split(':')[0]
|
||||||
image_repo = '/'.join(image_repo)
|
image_repo = '/'.join(image_repo)
|
||||||
# Nextcloud AIO validates the mastercontainer image reference and rejects
|
if image_repo in release['images']:
|
||||||
# a digest. Keep its configured tag so AIO can start and manage its stack.
|
|
||||||
if nextcloud_aio_master:
|
|
||||||
pinned = reference
|
|
||||||
elif image_repo in release['images']:
|
|
||||||
pinned = image_repo + '@' + release['images'][image_repo]
|
pinned = image_repo + '@' + release['images'][image_repo]
|
||||||
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
|
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
|
||||||
pinned = locks[reference]
|
pinned = locks[reference]
|
||||||
@@ -93,12 +58,6 @@ def prepare(source_file):
|
|||||||
pinned = resolve(reference)
|
pinned = resolve(reference)
|
||||||
settings['image'] = pinned
|
settings['image'] = pinned
|
||||||
locks[reference] = pinned
|
locks[reference] = pinned
|
||||||
for service, settings in before['services'].items():
|
|
||||||
reference = settings['image']
|
|
||||||
image_repo = reference.split('@')[0].rsplit('/', 1)
|
|
||||||
image_repo[-1] = image_repo[-1].split(':')[0]
|
|
||||||
image_repo = '/'.join(image_repo)
|
|
||||||
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
|
|
||||||
# Capture what is running, not the current value of its mutable tag.
|
# Capture what is running, not the current value of its mutable tag.
|
||||||
ids = output(
|
ids = output(
|
||||||
'docker',
|
'docker',
|
||||||
@@ -110,43 +69,13 @@ def prepare(source_file):
|
|||||||
f'label=com.docker.compose.service={service}',
|
f'label=com.docker.compose.service={service}',
|
||||||
).splitlines()
|
).splitlines()
|
||||||
actual = set()
|
actual = set()
|
||||||
if bootstrap and ids:
|
|
||||||
expected_hash = output(
|
|
||||||
'docker',
|
|
||||||
'compose',
|
|
||||||
'--project-directory',
|
|
||||||
str(project_dir),
|
|
||||||
'-f',
|
|
||||||
str(persistent_file),
|
|
||||||
'config',
|
|
||||||
'--hash',
|
|
||||||
service,
|
|
||||||
cwd=config_repo,
|
|
||||||
).split()[-1]
|
|
||||||
for container in ids:
|
|
||||||
running_hash = output(
|
|
||||||
'docker',
|
|
||||||
'inspect',
|
|
||||||
container,
|
|
||||||
'--format',
|
|
||||||
'{{ index .Config.Labels "com.docker.compose.config-hash" }}',
|
|
||||||
)
|
|
||||||
if running_hash != expected_hash:
|
|
||||||
raise ValueError(
|
|
||||||
f'{project}/{service}: persistent config differs from running config; restore the previous config'
|
|
||||||
)
|
|
||||||
for container in ids:
|
for container in ids:
|
||||||
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
|
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
|
||||||
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
|
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
|
||||||
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
|
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
|
||||||
if len(actual) > 1:
|
if len(actual) > 1:
|
||||||
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
|
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
|
||||||
# AIO also rejects a digest in its recovery config. Preserve its tag in
|
before['services'][service]['image'] = next(iter(actual)) if actual else reference
|
||||||
# both deploy and recovery files.
|
|
||||||
if nextcloud_aio_master:
|
|
||||||
before['services'][service]['image'] = reference
|
|
||||||
else:
|
|
||||||
before['services'][service]['image'] = next(iter(actual)) if actual else reference
|
|
||||||
for name, data in (('compose', config), ('compose-before', before)):
|
for name, data in (('compose', config), ('compose-before', before)):
|
||||||
folder = directory / name
|
folder = directory / name
|
||||||
folder.mkdir(mode=0o700, exist_ok=True)
|
folder.mkdir(mode=0o700, exist_ok=True)
|
||||||
@@ -156,7 +85,7 @@ def prepare(source_file):
|
|||||||
images_file.write_text(json.dumps(locks, indent=2) + '\n')
|
images_file.write_text(json.dumps(locks, indent=2) + '\n')
|
||||||
print(f'Compose {project}: images pinned; local paths preserved')
|
print(f'Compose {project}: images pinned; local paths preserved')
|
||||||
print(
|
print(
|
||||||
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never --remove-orphans'
|
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never'
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -141,8 +141,7 @@ def make_plan(directory):
|
|||||||
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
|
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
|
||||||
request = json.loads((directory / 'request.json').read_text())
|
request = json.loads((directory / 'request.json').read_text())
|
||||||
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
|
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
|
||||||
# Helm 4 lists every release status by default and removed the --all flag.
|
helm = json.loads(command('helm', 'list', '--all', '-A', '-o', 'json'))
|
||||||
helm = json.loads(command('helm', 'list', '-A', '-o', 'json'))
|
|
||||||
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
|
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
|
||||||
if request['refresh_images']:
|
if request['refresh_images']:
|
||||||
plan['selected']['compose'] = plan['active']['compose']
|
plan['selected']['compose'] = plan['active']['compose']
|
||||||
@@ -165,10 +164,6 @@ def finish_success(directory, plan):
|
|||||||
if previous.exists()
|
if previous.exists()
|
||||||
else {}
|
else {}
|
||||||
)
|
)
|
||||||
configs = STATE / 'compose-configs'
|
|
||||||
configs.mkdir(mode=0o700, exist_ok=True)
|
|
||||||
for config in (directory / 'compose').glob('*.json'):
|
|
||||||
atomic_json(configs / config.name, json.loads(config.read_text()))
|
|
||||||
atomic_json(STATE / 'last-success.json', plan)
|
atomic_json(STATE / 'last-success.json', plan)
|
||||||
status = json.loads((directory / 'status.json').read_text())
|
status = json.loads((directory / 'status.json').read_text())
|
||||||
status['state'] = 'success'
|
status['state'] = 'success'
|
||||||
@@ -217,17 +212,14 @@ def execute(run_id):
|
|||||||
raise ValueError(f'Interrupted deploy {other.name}; run recover first')
|
raise ValueError(f'Interrupted deploy {other.name}; run recover first')
|
||||||
status['state'] = 'running'
|
status['state'] = 'running'
|
||||||
atomic_json(directory / 'status.json', status)
|
atomic_json(directory / 'status.json', status)
|
||||||
phase = 'plan'
|
|
||||||
try:
|
try:
|
||||||
plan = make_plan(directory)
|
plan = make_plan(directory)
|
||||||
print(
|
print(
|
||||||
json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}),
|
json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}),
|
||||||
flush=True,
|
flush=True,
|
||||||
)
|
)
|
||||||
phase = 'doctor'
|
|
||||||
if not stage(directory, 'doctor', 600):
|
if not stage(directory, 'doctor', 600):
|
||||||
raise RuntimeError('Preflight failed')
|
raise RuntimeError('Preflight failed')
|
||||||
phase = 'validate'
|
|
||||||
if not stage(directory, 'validate', 1200):
|
if not stage(directory, 'validate', 1200):
|
||||||
raise RuntimeError('Validation failed')
|
raise RuntimeError('Validation failed')
|
||||||
if json.loads((directory / 'request.json').read_text())['mode'] == 'plan':
|
if json.loads((directory / 'request.json').read_text())['mode'] == 'plan':
|
||||||
@@ -236,7 +228,6 @@ def execute(run_id):
|
|||||||
atomic_json(directory / 'status.json', status)
|
atomic_json(directory / 'status.json', status)
|
||||||
return
|
return
|
||||||
# Budget includes both rollout checks and rollback waves, plus API overhead.
|
# Budget includes both rollout checks and rollback waves, plus API overhead.
|
||||||
phase = 'Recovery budget'
|
|
||||||
count = int(
|
count = int(
|
||||||
command(
|
command(
|
||||||
'bash',
|
'bash',
|
||||||
@@ -248,24 +239,14 @@ def execute(run_id):
|
|||||||
verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120)
|
verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120)
|
||||||
if verify_budget > 7200:
|
if verify_budget > 7200:
|
||||||
raise ValueError('More than two hours of recovery required; split this deploy')
|
raise ValueError('More than two hours of recovery required; split this deploy')
|
||||||
phase = 'apply-k8s'
|
|
||||||
k8s_ok = stage(directory, 'apply-k8s', 2700)
|
k8s_ok = stage(directory, 'apply-k8s', 2700)
|
||||||
phase = 'apply-compose'
|
|
||||||
compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False
|
compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False
|
||||||
phase = 'verify-k8s'
|
|
||||||
verify_ok = stage(directory, 'verify-k8s', verify_budget)
|
verify_ok = stage(directory, 'verify-k8s', verify_budget)
|
||||||
phase = 'smoke'
|
|
||||||
smoke_ok = stage(directory, 'smoke', 600)
|
smoke_ok = stage(directory, 'smoke', 600)
|
||||||
if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)):
|
if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)):
|
||||||
raise RuntimeError('Deploy failed; inspect stage logs and recovery report')
|
raise RuntimeError('Deploy failed; inspect stage logs and recovery report')
|
||||||
phase = 'Save the successful baseline'
|
|
||||||
finish_success(directory, plan)
|
finish_success(directory, plan)
|
||||||
except Exception as error:
|
except Exception as error:
|
||||||
status = json.loads((directory / 'status.json').read_text())
|
|
||||||
status['failure_stage'] = next(
|
|
||||||
(name for name, result in status['stages'].items() if result.get('result') == 'failure'), phase
|
|
||||||
)
|
|
||||||
atomic_json(directory / 'status.json', status)
|
|
||||||
with (directory / 'controller.log').open('a') as stream:
|
with (directory / 'controller.log').open('a') as stream:
|
||||||
stream.write(f'{error}\n')
|
stream.write(f'{error}\n')
|
||||||
recover(directory)
|
recover(directory)
|
||||||
@@ -313,93 +294,10 @@ def follow(run_id, phase):
|
|||||||
time.sleep(3)
|
time.sleep(3)
|
||||||
|
|
||||||
|
|
||||||
def summary(run_id):
|
|
||||||
directory = run_directory(run_id)
|
|
||||||
request = json.loads((directory / 'request.json').read_text())
|
|
||||||
release = request['release']
|
|
||||||
plan_file = directory / 'plan.json'
|
|
||||||
lines = [
|
|
||||||
f'## Deploy `{release["sha"]}`',
|
|
||||||
'',
|
|
||||||
f'- Mode: `{request["mode"]}`',
|
|
||||||
f'- Refresh third-party images: `{request["refresh_images"]}`',
|
|
||||||
]
|
|
||||||
status = json.loads((directory / 'status.json').read_text())
|
|
||||||
if status.get('failure_stage'):
|
|
||||||
lines.append(f'- Failed stage: **{status["failure_stage"]}**')
|
|
||||||
lines.extend(
|
|
||||||
[
|
|
||||||
'',
|
|
||||||
f'- Observed run state: **{status["state"]}**',
|
|
||||||
'',
|
|
||||||
'### Stage results',
|
|
||||||
'| Stage | Result | Exit code |',
|
|
||||||
'| --- | --- | --- |',
|
|
||||||
]
|
|
||||||
)
|
|
||||||
for name in ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke'):
|
|
||||||
stage_result = status['stages'].get(name, {})
|
|
||||||
lines.append(f'| {name} | {stage_result.get("result", "not started")} | {stage_result.get("exit_code", "—")} |')
|
|
||||||
lines.extend(['', '### Apply and Helm recovery results'])
|
|
||||||
events_file = directory / 'apply-events.jsonl'
|
|
||||||
events = []
|
|
||||||
if events_file.exists():
|
|
||||||
for line in events_file.read_text().splitlines():
|
|
||||||
try:
|
|
||||||
events.append(json.loads(line))
|
|
||||||
except json.JSONDecodeError:
|
|
||||||
lines.append('- An operation record is incomplete. Check the stage log.')
|
|
||||||
latest = {(event['action'], event['target']): event['result'] for event in events}
|
|
||||||
lines.extend(f'- `{action}` `{target}`: **{result}**' for (action, target), result in latest.items())
|
|
||||||
if not latest:
|
|
||||||
lines.append('- No apply results were recorded.')
|
|
||||||
lines.append('- A completed apply does not confirm health. See verification and smoke results.')
|
|
||||||
lines.extend(['', '### Kubernetes recovery'])
|
|
||||||
pointer = directory / 'snapshot/current'
|
|
||||||
failed = Path(pointer.read_text().strip()) / 'failed-workloads' if pointer.exists() else None
|
|
||||||
if failed and failed.exists():
|
|
||||||
contents = failed.read_text()
|
|
||||||
counts = dict(re.findall(r'^(ROLLED_BACK|UNRECOVERED)=([0-9]+)$', contents, re.MULTILINE))
|
|
||||||
if not contents.strip():
|
|
||||||
lines.append('- No failed workloads were recorded. See the verification result above.')
|
|
||||||
elif counts:
|
|
||||||
lines.append(f'- Workloads restored: **{counts.get("ROLLED_BACK", "unknown")}**')
|
|
||||||
lines.append(f'- Workloads that need manual recovery: **{counts.get("UNRECOVERED", "unknown")}**')
|
|
||||||
else:
|
|
||||||
lines.append('- Rollback has no recorded result yet. Check the verification log.')
|
|
||||||
else:
|
|
||||||
lines.append('- No workload rollback was recorded. This does not confirm health.')
|
|
||||||
lines.append('- Compose requires manual recovery. Use the saved command in the apply log.')
|
|
||||||
if not plan_file.exists():
|
|
||||||
lines.extend(['', 'Plan was not created. Check the controller log.'])
|
|
||||||
print('\n'.join(lines))
|
|
||||||
return
|
|
||||||
plan = json.loads(plan_file.read_text())
|
|
||||||
lines.extend(['', '### Selected services'])
|
|
||||||
count = 0
|
|
||||||
for kind, services in plan['selected'].items():
|
|
||||||
for service in services:
|
|
||||||
lines.append(f'- `{kind}`: `{service}`')
|
|
||||||
count += 1
|
|
||||||
if not count:
|
|
||||||
lines.append('- None')
|
|
||||||
lines.extend(['', '### Selected Helm releases'])
|
|
||||||
lines.extend(f'- `{release}`' for release in plan.get('helm', []))
|
|
||||||
if not plan.get('helm'):
|
|
||||||
lines.append('- None')
|
|
||||||
lines.extend(['', '### Images pinned in the checked release'])
|
|
||||||
lines.extend(f'- `{image}@{digest}`' for image, digest in sorted(release['images'].items()))
|
|
||||||
lines.extend(['', '### Removed resources requiring manual review'])
|
|
||||||
lines.extend(f'- `{item}`' for item in plan.get('removed', []))
|
|
||||||
if not plan.get('removed'):
|
|
||||||
lines.append('- None')
|
|
||||||
print('\n'.join(lines))
|
|
||||||
|
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
os.umask(0o077)
|
os.umask(0o077)
|
||||||
parser = argparse.ArgumentParser(description=__doc__)
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow', 'summary'))
|
parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow'))
|
||||||
parser.add_argument('run_id')
|
parser.add_argument('run_id')
|
||||||
parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke'))
|
parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke'))
|
||||||
parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply')
|
parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply')
|
||||||
@@ -417,8 +315,6 @@ def main():
|
|||||||
if (directory / 'plan.json').exists():
|
if (directory / 'plan.json').exists():
|
||||||
plan = json.loads((directory / 'plan.json').read_text())
|
plan = json.loads((directory / 'plan.json').read_text())
|
||||||
print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2))
|
print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2))
|
||||||
elif args.action == 'summary':
|
|
||||||
summary(args.run_id)
|
|
||||||
elif not follow(args.run_id, args.phase):
|
elif not follow(args.run_id, args.phase):
|
||||||
sys.exit(1)
|
sys.exit(1)
|
||||||
|
|
||||||
|
|||||||
@@ -27,15 +27,6 @@ log() {
|
|||||||
echo "== $* =="
|
echo "== $* =="
|
||||||
}
|
}
|
||||||
|
|
||||||
# Store operation results without command output or local configuration values.
|
|
||||||
record_apply() {
|
|
||||||
[ -n "${RUN_DIR:-}" ] || return 0
|
|
||||||
jq -cn --arg action "$1" --arg target "$2" --arg result "$3" \
|
|
||||||
'{action: $action, target: $target, result: $result}' >>"$RUN_DIR/apply-events.jsonl" \
|
|
||||||
|| echo 'WARNING: cannot record an apply result' >&2
|
|
||||||
return 0
|
|
||||||
}
|
|
||||||
|
|
||||||
warn() {
|
warn() {
|
||||||
echo "WARNING: $*" >&2
|
echo "WARNING: $*" >&2
|
||||||
}
|
}
|
||||||
@@ -180,8 +171,7 @@ save_snapshot() {
|
|||||||
| select(any(.metadata.ownerReferences[]?; .uid == $w.metadata.uid))
|
| select(any(.metadata.ownerReferences[]?; .uid == $w.metadata.uid))
|
||||||
| select($w.kind != "StatefulSet" or .metadata.name == $w.status.currentRevision) | .revision] | max // 0) end)
|
| select($w.kind != "StatefulSet" or .metadata.name == $w.status.currentRevision) | .revision] | max // 0) end)
|
||||||
}]' "$dir/workloads.json" >"$dir/revisions.json" || return 1
|
}]' "$dir/workloads.json" >"$dir/revisions.json" || return 1
|
||||||
# Helm 4 lists every release status by default and removed the --all flag.
|
releases="$(helm list --all -A -o json)" || return 1
|
||||||
releases="$(helm list -A -o json)" || return 1
|
|
||||||
for entry in "${HELM_RELEASES[@]}"; do
|
for entry in "${HELM_RELEASES[@]}"; do
|
||||||
IFS='|' read -r release _ namespace _ _ _ <<<"$entry"
|
IFS='|' read -r release _ namespace _ _ _ <<<"$entry"
|
||||||
if ! jq -e --arg r "$release" --arg n "$namespace" \
|
if ! jq -e --arg r "$release" --arg n "$namespace" \
|
||||||
@@ -330,11 +320,11 @@ rollback_workloads() {
|
|||||||
# written straight into a `helm upgrade` command would never be updated: these
|
# written straight into a `helm upgrade` command would never be updated: these
|
||||||
# have to be declared as custom.regex managers in renovate/renovate.json.
|
# have to be declared as custom.regex managers in renovate/renovate.json.
|
||||||
HELM_RELEASES=(
|
HELM_RELEASES=(
|
||||||
"prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.3.2|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active"
|
"prometheus-stack|prometheus-community/kube-prometheus-stack|prometheus|86.2.3|prometheus-stack/k8s/grafana-values.yaml|prometheus-stack/k8s/active"
|
||||||
"victoria-operator|victoriametrics/victoria-metrics-operator|prometheus|0.68.1|prometheus-stack/k8s/victoria-operator-values.yaml|prometheus-stack/k8s/active"
|
"victoria-operator|victoriametrics/victoria-metrics-operator|prometheus|0.68.1|prometheus-stack/k8s/victoria-operator-values.yaml|prometheus-stack/k8s/active"
|
||||||
"loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active"
|
"loki|grafana/loki|prometheus|7.3.0|loki/k8s/loki-values.yaml|loki/k8s/active"
|
||||||
"alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active"
|
"alloy|grafana/alloy|prometheus|1.12.1|loki/k8s/alloy-values.yaml|loki/k8s/active"
|
||||||
"reloader|stakater/reloader|reloader|2.2.18|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
|
"reloader|stakater/reloader|reloader|2.2.17|reloader/k8s/reloader-values.yaml|reloader/k8s/active"
|
||||||
)
|
)
|
||||||
|
|
||||||
# "name url" for the Helm repository hosting a chart, empty if unknown.
|
# "name url" for the Helm repository hosting a chart, empty if unknown.
|
||||||
@@ -386,19 +376,15 @@ recover_pending_release() {
|
|||||||
echo "ERROR: no captured Helm revision for $release; manual recovery required"
|
echo "ERROR: no captured Helm revision for $release; manual recovery required"
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
record_apply helm-rollback "$namespace/$release" started
|
|
||||||
if ! helm rollback "$release" "$revision" -n "$namespace" --wait --timeout 10m; then
|
if ! helm rollback "$release" "$revision" -n "$namespace" --wait --timeout 10m; then
|
||||||
record_apply helm-rollback "$namespace/$release" failure
|
|
||||||
echo "WARN: helm rollback of $release did not complete"
|
echo "WARN: helm rollback of $release did not complete"
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
status="$(helm_release_status "$release" "$namespace")" || return 1
|
status="$(helm_release_status "$release" "$namespace")" || return 1
|
||||||
if [ "$status" != "deployed" ]; then
|
if [ "$status" != "deployed" ]; then
|
||||||
record_apply helm-rollback "$namespace/$release" failure
|
|
||||||
echo "WARN: $release is $status after rollback"
|
echo "WARN: $release is $status after rollback"
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
record_apply helm-rollback "$namespace/$release" success
|
|
||||||
;;
|
;;
|
||||||
esac
|
esac
|
||||||
return 0
|
return 0
|
||||||
@@ -462,13 +448,11 @@ upgrade_helm_releases() {
|
|||||||
# --rollback-on-failure (+ --wait) rolls the release back when the upgrade
|
# --rollback-on-failure (+ --wait) rolls the release back when the upgrade
|
||||||
# times out or the workloads it touches never become ready, so a bad chart
|
# times out or the workloads it touches never become ready, so a bad chart
|
||||||
# bump is not left half applied. (--atomic was this combo; deprecated.)
|
# bump is not left half applied. (--atomic was this combo; deprecated.)
|
||||||
record_apply helm-upgrade "$namespace/$release" started
|
|
||||||
if ! helm upgrade --install "$release" "$chart" \
|
if ! helm upgrade --install "$release" "$chart" \
|
||||||
--namespace "$namespace" \
|
--namespace "$namespace" \
|
||||||
--version "$version" \
|
--version "$version" \
|
||||||
--values "$values" \
|
--values "$values" \
|
||||||
--wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then
|
--wait --rollback-on-failure --cleanup-on-fail --timeout 10m; then
|
||||||
record_apply helm-upgrade "$namespace/$release" failure
|
|
||||||
echo "WARN: upgrade of $release failed, checking release state"
|
echo "WARN: upgrade of $release failed, checking release state"
|
||||||
# --rollback-on-failure already attempted its own rollback; finish the job when that
|
# --rollback-on-failure already attempted its own rollback; finish the job when that
|
||||||
# rollback never completed, otherwise the release stays pending-* and
|
# rollback never completed, otherwise the release stays pending-* and
|
||||||
@@ -478,10 +462,8 @@ upgrade_helm_releases() {
|
|||||||
else
|
else
|
||||||
echo "ERROR: upgrade of $release failed (release is back on its previous revision)."
|
echo "ERROR: upgrade of $release failed (release is back on its previous revision)."
|
||||||
fi
|
fi
|
||||||
record_apply helm-recovery-state "$namespace/$release" "$(helm_release_status "$release" "$namespace" || echo unknown)"
|
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
record_apply helm-upgrade "$namespace/$release" success
|
|
||||||
done
|
done
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -571,41 +553,6 @@ skip_uninstalled_vmagent_crd() {
|
|||||||
return 1
|
return 1
|
||||||
}
|
}
|
||||||
|
|
||||||
# Render one complete resource list so new namespaces can be identified across
|
|
||||||
# files and Kustomize apps. A missing undeclared namespace remains an error.
|
|
||||||
render_selected_resources() {
|
|
||||||
local m k
|
|
||||||
{
|
|
||||||
for m in "${K8S_MANIFESTS[@]}"; do
|
|
||||||
if skip_uninstalled_vmagent_crd "$m" >/dev/null; then continue; fi
|
|
||||||
kubectl create --dry-run=client --validate=false -f "$m" -o json || return 1
|
|
||||||
done
|
|
||||||
for k in "${KUSTOMIZE_APPS[@]}"; do
|
|
||||||
kubectl kustomize "$k" | kubectl create --dry-run=client --validate=false -f - -o json || return 1
|
|
||||||
done
|
|
||||||
} | jq -s '{apiVersion: "v1", kind: "List", items: [ .[] | if .kind == "List" then .items[] else . end ]}'
|
|
||||||
}
|
|
||||||
|
|
||||||
validate_server_resources() {
|
|
||||||
local defer_new="$1" resources existing filtered
|
|
||||||
resources="$(render_selected_resources)" || return 1
|
|
||||||
existing="$(kubectl get namespaces -o json)" || return 1
|
|
||||||
filtered="$(jq --argjson existing "$existing" --argjson defer "$defer_new" '
|
|
||||||
[.items[] | select(.kind == "Namespace") | .metadata.name] as $declared
|
|
||||||
| [$existing.items[].metadata.name] as $present
|
|
||||||
| .items |= map(
|
|
||||||
(.metadata.namespace // "default") as $ns
|
|
||||||
| if .kind == "Namespace" or ($present | index($ns)) != null then .
|
|
||||||
elif ($declared | index($ns)) == null then error("Undeclared missing namespace: " + $ns)
|
|
||||||
elif $defer then empty
|
|
||||||
else error("Namespace still missing after namespace apply: " + $ns)
|
|
||||||
end)
|
|
||||||
' <<<"$resources")" || return 1
|
|
||||||
if [ "$(jq '.items | length' <<<"$filtered")" -gt 0 ]; then
|
|
||||||
kubectl apply --dry-run=server -f - <<<"$filtered" >/dev/null
|
|
||||||
fi
|
|
||||||
}
|
|
||||||
|
|
||||||
stage_validate() {
|
stage_validate() {
|
||||||
check_prune_mode || return 1
|
check_prune_mode || return 1
|
||||||
cd "$REPO"
|
cd "$REPO"
|
||||||
@@ -632,7 +579,15 @@ stage_validate() {
|
|||||||
kubectl apply -k "$k" --dry-run=client >/dev/null
|
kubectl apply -k "$k" --dry-run=client >/dev/null
|
||||||
done
|
done
|
||||||
log "Validate k8s manifests (kubectl dry-run=server)"
|
log "Validate k8s manifests (kubectl dry-run=server)"
|
||||||
validate_server_resources true
|
for m in ${K8S_MANIFESTS[@]+"${K8S_MANIFESTS[@]}"}; do
|
||||||
|
if skip_uninstalled_vmagent_crd "$m"; then
|
||||||
|
continue
|
||||||
|
fi
|
||||||
|
kubectl apply --dry-run=server -f "$m" >/dev/null
|
||||||
|
done
|
||||||
|
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
|
||||||
|
kubectl apply -k "$k" --dry-run=server >/dev/null
|
||||||
|
done
|
||||||
log "Checking referenced Secrets exist"
|
log "Checking referenced Secrets exist"
|
||||||
echo " (deploy never applies *secret*.yaml; create missing ones manually)"
|
echo " (deploy never applies *secret*.yaml; create missing ones manually)"
|
||||||
check_referenced_secrets
|
check_referenced_secrets
|
||||||
@@ -676,22 +631,9 @@ stage_apply_k8s() {
|
|||||||
if [ "${#ns_files[@]}" -gt 0 ]; then
|
if [ "${#ns_files[@]}" -gt 0 ]; then
|
||||||
log "Applying namespaces (${#ns_files[@]} files)"
|
log "Applying namespaces (${#ns_files[@]} files)"
|
||||||
for m in "${ns_files[@]}"; do
|
for m in "${ns_files[@]}"; do
|
||||||
record_apply kubectl "${m#"$REPO"/}" started
|
kubectl apply -f "$m"
|
||||||
if ! kubectl apply -f "$m"; then
|
|
||||||
record_apply kubectl "${m#"$REPO"/}" failure
|
|
||||||
return 1
|
|
||||||
fi
|
|
||||||
record_apply kubectl "${m#"$REPO"/}" success
|
|
||||||
done
|
done
|
||||||
fi
|
fi
|
||||||
# Kustomize may declare namespaces inside its rendered resources too.
|
|
||||||
local namespace_resources
|
|
||||||
namespace_resources="$(render_selected_resources | jq '.items |= map(select(.kind == "Namespace"))')" || return 1
|
|
||||||
if [ "$(jq '.items | length' <<<"$namespace_resources")" -gt 0 ]; then
|
|
||||||
kubectl apply -f - <<<"$namespace_resources" || return 1
|
|
||||||
fi
|
|
||||||
# Complete the deferred server checks before Helm or application resources change.
|
|
||||||
validate_server_resources false || return 1
|
|
||||||
if selected_service k8s prometheus-stack && [ -f "$REPO/prometheus-stack/k8s/active" ]; then
|
if selected_service k8s prometheus-stack && [ -f "$REPO/prometheus-stack/k8s/active" ]; then
|
||||||
if [ ! -f "$CONFIG_REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then
|
if [ ! -f "$CONFIG_REPO/prometheus-stack/k8s/grafana-values.yaml" ]; then
|
||||||
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
|
echo "ERROR: prometheus-stack/k8s/grafana-values.yaml (gitignored) missing on workstation, restore it first."
|
||||||
@@ -704,24 +646,18 @@ stage_apply_k8s() {
|
|||||||
log "Applying resources (${#other_files[@]} files, our images pinned to digests)"
|
log "Applying resources (${#other_files[@]} files, our images pinned to digests)"
|
||||||
for m in "${other_files[@]}"; do
|
for m in "${other_files[@]}"; do
|
||||||
log "Applying ${m#"$REPO"/}"
|
log "Applying ${m#"$REPO"/}"
|
||||||
record_apply kubectl "${m#"$REPO"/}" started
|
|
||||||
if ! render_pinned <"$m" | kubectl apply -f -; then
|
if ! render_pinned <"$m" | kubectl apply -f -; then
|
||||||
record_apply kubectl "${m#"$REPO"/}" failure
|
|
||||||
echo "ERROR: apply failed for ${m#"$REPO"/}" >&2
|
echo "ERROR: apply failed for ${m#"$REPO"/}" >&2
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
record_apply kubectl "${m#"$REPO"/}" success
|
|
||||||
done
|
done
|
||||||
fi
|
fi
|
||||||
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
|
for k in ${KUSTOMIZE_APPS[@]+"${KUSTOMIZE_APPS[@]}"}; do
|
||||||
log "Applying kustomize app: ${k#"$REPO"/} (our images pinned to digests)"
|
log "Applying kustomize app: ${k#"$REPO"/} (our images pinned to digests)"
|
||||||
record_apply kustomize "${k#"$REPO"/}" started
|
|
||||||
if ! kubectl kustomize "$k" | render_pinned | kubectl apply -f -; then
|
if ! kubectl kustomize "$k" | render_pinned | kubectl apply -f -; then
|
||||||
record_apply kustomize "${k#"$REPO"/}" failure
|
|
||||||
echo "ERROR: apply failed for kustomize app ${k#"$REPO"/}" >&2
|
echo "ERROR: apply failed for kustomize app ${k#"$REPO"/}" >&2
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
record_apply kustomize "${k#"$REPO"/}" success
|
|
||||||
done
|
done
|
||||||
|
|
||||||
# No verification here on purpose. This stage may be killed at any point by
|
# No verification here on purpose. This stage may be killed at any point by
|
||||||
@@ -827,13 +763,12 @@ stage_verify_k8s() {
|
|||||||
# actually be running.
|
# actually be running.
|
||||||
verify_compose_stack() {
|
verify_compose_stack() {
|
||||||
local cf="$1"
|
local cf="$1"
|
||||||
local expected running svc missing=() service_count=0
|
local expected running missing=()
|
||||||
expected="$(compose "$cf" config --format json | jq -r ' .services | to_entries[] | select(.value.restart != "no") | .key' | sort)" || return 1
|
expected="$(compose "$cf" config --format json | jq -r ' .services | to_entries[] | select(.value.restart != "no") | .key' | sort)" || return 1
|
||||||
running="$(compose "$cf" ps --status running --services | sort)" || return 1
|
running="$(compose "$cf" ps --status running --services | sort)" || return 1
|
||||||
[ -n "$expected" ] || return 0
|
[ -n "$expected" ] || return 0
|
||||||
while IFS= read -r svc; do
|
while IFS= read -r svc; do
|
||||||
[ -n "$svc" ] || continue
|
[ -n "$svc" ] || continue
|
||||||
service_count=$((service_count + 1))
|
|
||||||
# restart:"no" services are allowed to have exited.
|
# restart:"no" services are allowed to have exited.
|
||||||
if ! printf '%s\n' "$running" | grep -qx "$svc"; then
|
if ! printf '%s\n' "$running" | grep -qx "$svc"; then
|
||||||
missing+=("$svc")
|
missing+=("$svc")
|
||||||
@@ -844,7 +779,7 @@ verify_compose_stack() {
|
|||||||
compose "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true
|
compose "$cf" ps --all 2>/dev/null | sed 's/^/ /' || true
|
||||||
return 1
|
return 1
|
||||||
fi
|
fi
|
||||||
echo " all $service_count service(s) running"
|
echo " all ${#expected} service(s) running"
|
||||||
return 0
|
return 0
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1024,12 +959,7 @@ stage_apply_compose() {
|
|||||||
local cf
|
local cf
|
||||||
for cf in "${COMPOSE_STACKS[@]}"; do
|
for cf in "${COMPOSE_STACKS[@]}"; do
|
||||||
log "Applying Compose ${cf#"$REPO"/}"
|
log "Applying Compose ${cf#"$REPO"/}"
|
||||||
record_apply compose "${cf#"$REPO"/}" started
|
compose "$cf" up -d --wait --wait-timeout 180 --pull missing --remove-orphans
|
||||||
if ! compose "$cf" up -d --wait --wait-timeout 180 --pull missing --remove-orphans; then
|
|
||||||
record_apply compose "${cf#"$REPO"/}" failure
|
|
||||||
return 1
|
|
||||||
fi
|
|
||||||
record_apply compose "${cf#"$REPO"/}" success
|
|
||||||
verify_compose_stack "$cf"
|
verify_compose_stack "$cf"
|
||||||
done
|
done
|
||||||
echo "Compose recovery files: $RUN_DIR/compose-before (manual recovery only)"
|
echo "Compose recovery files: $RUN_DIR/compose-before (manual recovery only)"
|
||||||
|
|||||||
@@ -83,20 +83,20 @@ def make_plan(repo, config_repo, release, previous, mode, live_helm):
|
|||||||
removed = []
|
removed = []
|
||||||
else:
|
else:
|
||||||
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
|
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
|
||||||
changed = {service for service in all_services for path in paths if path.startswith(service + '/')}
|
changed = {path.split('/')[0] for path in paths}
|
||||||
if any(path.startswith('.gitea/') for path in paths):
|
if any(path.startswith('.gitea/') for path in paths):
|
||||||
changed |= all_services
|
changed |= all_services
|
||||||
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
|
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
|
||||||
for file in tracked(repo):
|
for file in tracked(repo):
|
||||||
owners = {service for service in all_services if file.startswith(service + '/')}
|
service = file.split('/')[0]
|
||||||
if not owners or not file.endswith(('.yaml', '.yml')):
|
if service not in all_services or not file.endswith(('.yaml', '.yml')):
|
||||||
continue
|
continue
|
||||||
text = (repo / file).read_text()
|
text = (repo / file).read_text()
|
||||||
if any(
|
if any(
|
||||||
image in text and previous.get('images', {}).get(image) != digest
|
image in text and previous.get('images', {}).get(image) != digest
|
||||||
for image, digest in release['images'].items()
|
for image, digest in release['images'].items()
|
||||||
):
|
):
|
||||||
changed |= owners
|
changed.add(service)
|
||||||
removed = sorted(
|
removed = sorted(
|
||||||
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
|
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
|
||||||
- all_services
|
- all_services
|
||||||
|
|||||||
@@ -64,23 +64,11 @@ jobs:
|
|||||||
run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA"
|
run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA"
|
||||||
- name: Submit durable deploy to workstation
|
- name: Submit durable deploy to workstation
|
||||||
run: bash .gitea/workflows/ssh-run.sh start
|
run: bash .gitea/workflows/ssh-run.sh start
|
||||||
- name: Write the request result
|
|
||||||
if: always()
|
|
||||||
env:
|
|
||||||
REQUEST_RESULT: ${{ job.status }}
|
|
||||||
CHECKED_SHA: ${{ steps.release.outputs.sha }}
|
|
||||||
run: |
|
|
||||||
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
printf '## Deploy request\n\n- Result: **%s**\n- Checked commit: %s\n- Mode: %s\n' "$REQUEST_RESULT" "${CHECKED_SHA:-not checked}" "$DEPLOY_MODE" >>"$GITHUB_STEP_SUMMARY"
|
|
||||||
if [ "$REQUEST_RESULT" != success ]; then
|
|
||||||
echo 'Open the failed step log. If SSH submission failed, check the remote controller state.' >>"$GITHUB_STEP_SUMMARY"
|
|
||||||
fi
|
|
||||||
fi
|
|
||||||
|
|
||||||
apply:
|
apply:
|
||||||
needs: [gate]
|
needs: [gate]
|
||||||
runs-on: homelab
|
runs-on: homelab
|
||||||
timeout-minutes: 120
|
timeout-minutes: 100
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout checked commit
|
- name: Checkout checked commit
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||||
@@ -88,14 +76,6 @@ jobs:
|
|||||||
ref: ${{ needs.gate.outputs.sha }}
|
ref: ${{ needs.gate.outputs.sha }}
|
||||||
- name: Follow validation and sequential Kubernetes / Compose apply
|
- name: Follow validation and sequential Kubernetes / Compose apply
|
||||||
run: bash .gitea/workflows/ssh-run.sh apply
|
run: bash .gitea/workflows/ssh-run.sh apply
|
||||||
- name: Write the deploy result
|
|
||||||
if: always()
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/ssh-run.sh ]; then
|
|
||||||
bash .gitea/workflows/ssh-run.sh summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
|
|
||||||
fi
|
|
||||||
|
|
||||||
verify:
|
verify:
|
||||||
needs: [gate, apply]
|
needs: [gate, apply]
|
||||||
@@ -109,14 +89,6 @@ jobs:
|
|||||||
ref: ${{ needs.gate.outputs.sha }}
|
ref: ${{ needs.gate.outputs.sha }}
|
||||||
- name: Follow workload verification and recovery
|
- name: Follow workload verification and recovery
|
||||||
run: bash .gitea/workflows/ssh-run.sh verify
|
run: bash .gitea/workflows/ssh-run.sh verify
|
||||||
- name: Write the deploy result
|
|
||||||
if: always()
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/ssh-run.sh ]; then
|
|
||||||
bash .gitea/workflows/ssh-run.sh summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
|
|
||||||
fi
|
|
||||||
|
|
||||||
smoke:
|
smoke:
|
||||||
needs: [gate, verify]
|
needs: [gate, verify]
|
||||||
@@ -130,11 +102,3 @@ jobs:
|
|||||||
ref: ${{ needs.gate.outputs.sha }}
|
ref: ${{ needs.gate.outputs.sha }}
|
||||||
- name: Follow public route checks
|
- name: Follow public route checks
|
||||||
run: bash .gitea/workflows/ssh-run.sh smoke
|
run: bash .gitea/workflows/ssh-run.sh smoke
|
||||||
- name: Write the deploy result
|
|
||||||
if: always()
|
|
||||||
run: |
|
|
||||||
if [ -f .gitea/workflows/ssh-run.sh ]; then
|
|
||||||
bash .gitea/workflows/ssh-run.sh summary
|
|
||||||
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
|
|
||||||
fi
|
|
||||||
+63
-232
@@ -26,6 +26,9 @@ IMAGES = {
|
|||||||
'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'),
|
'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'),
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# These images are released by the EDU application repository.
|
||||||
|
EXTERNAL_IMAGES = {'gcr.forust.xyz/forust/session-keeper', 'gcr.forust.xyz/forust/webinar-checker'}
|
||||||
|
|
||||||
|
|
||||||
def command(*args, **kwargs):
|
def command(*args, **kwargs):
|
||||||
"""Arguments are passed directly to the executable, never to a shell."""
|
"""Arguments are passed directly to the executable, never to a shell."""
|
||||||
@@ -160,7 +163,7 @@ def gate(output, requested_ref, event_sha):
|
|||||||
print(f'CI gate accepted {sha}')
|
print(f'CI gate accepted {sha}')
|
||||||
|
|
||||||
|
|
||||||
def prepare_images(output):
|
def build(output):
|
||||||
sha = command('git', 'rev-parse', 'HEAD')
|
sha = command('git', 'rev-parse', 'HEAD')
|
||||||
if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
|
if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
|
||||||
raise ValueError('Build checkout does not match GITHUB_SHA')
|
raise ValueError('Build checkout does not match GITHUB_SHA')
|
||||||
@@ -175,61 +178,11 @@ def prepare_images(output):
|
|||||||
except ValueError:
|
except ValueError:
|
||||||
# Expired artifacts only cost a rebuild; mutable tags are never a fallback.
|
# Expired artifacts only cost a rebuild; mutable tags are never a fallback.
|
||||||
continue
|
continue
|
||||||
targets = []
|
|
||||||
for name, (context, dockerfile) in IMAGES.items():
|
|
||||||
image = f'gcr.forust.xyz/forust/{name}'
|
|
||||||
inputs = fingerprint(context, dockerfile)
|
|
||||||
old_digest = (previous or {}).get('images', {}).get(image)
|
|
||||||
targets.append(
|
|
||||||
{
|
|
||||||
'name': name,
|
|
||||||
'image': image,
|
|
||||||
'context': context,
|
|
||||||
'dockerfile': dockerfile,
|
|
||||||
'inputs': inputs,
|
|
||||||
'reuse_digest': old_digest if (previous or {}).get('inputs', {}).get(image) == inputs else None,
|
|
||||||
}
|
|
||||||
)
|
|
||||||
output.write_text(json.dumps({'sha': sha, 'targets': targets}, indent=2) + '\n')
|
|
||||||
if os.environ.get('GITHUB_OUTPUT'):
|
|
||||||
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
|
|
||||||
stream.write('matrix=' + json.dumps({'include': targets}, separators=(',', ':')) + '\n')
|
|
||||||
print(f'Prepared {len(targets)} image jobs; {sum(t["reuse_digest"] is None for t in targets)} require builds')
|
|
||||||
|
|
||||||
|
|
||||||
def checked_plan(path):
|
|
||||||
data = json.loads(path.read_text())
|
|
||||||
sha = command('git', 'rev-parse', 'HEAD')
|
|
||||||
if data.get('sha') != sha or sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
|
|
||||||
raise ValueError('Image plan does not match the checked source commit')
|
|
||||||
targets = data.get('targets', [])
|
|
||||||
if sorted(t['name'] for t in targets) != sorted(IMAGES):
|
|
||||||
raise ValueError('Image plan must contain each owned image once')
|
|
||||||
for target in targets:
|
|
||||||
name = target['name']
|
|
||||||
context, dockerfile = IMAGES[name]
|
|
||||||
if (target['context'], target['dockerfile'], target['image']) != (
|
|
||||||
context,
|
|
||||||
dockerfile,
|
|
||||||
f'gcr.forust.xyz/forust/{name}',
|
|
||||||
) or target['inputs'] != fingerprint(context, dockerfile):
|
|
||||||
raise ValueError('Image plan has invalid build inputs')
|
|
||||||
if target['reuse_digest'] is not None and not DIGEST.fullmatch(target['reuse_digest']):
|
|
||||||
raise ValueError('Image plan has an invalid reuse digest')
|
|
||||||
return data
|
|
||||||
|
|
||||||
|
|
||||||
def build_images(output, report, name, plan):
|
|
||||||
data = checked_plan(plan)
|
|
||||||
sha = data['sha']
|
|
||||||
target = next(t for t in data['targets'] if t['name'] == name)
|
|
||||||
context, dockerfile = IMAGES[name]
|
|
||||||
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
|
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
|
||||||
builder_config = Path.home() / '.cache/homelab-ci/buildx'
|
builder_config = Path.home() / '.cache/homelab-ci/buildx'
|
||||||
builder_config.mkdir(parents=True, exist_ok=True)
|
builder_config.mkdir(parents=True, exist_ok=True)
|
||||||
env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)}
|
env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)}
|
||||||
try:
|
try:
|
||||||
report['phase'] = 'Registry login'
|
|
||||||
subprocess.run( # noqa: S603, S607
|
subprocess.run( # noqa: S603, S607
|
||||||
[
|
[
|
||||||
shutil.which('docker') or '/usr/bin/docker',
|
shutil.which('docker') or '/usr/bin/docker',
|
||||||
@@ -244,7 +197,6 @@ def build_images(output, report, name, plan):
|
|||||||
check=True,
|
check=True,
|
||||||
env=env,
|
env=env,
|
||||||
)
|
)
|
||||||
report['phase'] = 'Prepare the builder'
|
|
||||||
builder = 'homelab-ci'
|
builder = 'homelab-ci'
|
||||||
versions = dict(
|
versions = dict(
|
||||||
re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE)
|
re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE)
|
||||||
@@ -279,66 +231,61 @@ def build_images(output, report, name, plan):
|
|||||||
)
|
)
|
||||||
signature.write_text(image + '\n')
|
signature.write_text(image + '\n')
|
||||||
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
|
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
|
||||||
report['images'] = release['images']
|
for name, (context, dockerfile) in IMAGES.items():
|
||||||
report['phase'] = f'Build or reuse {name}'
|
image = f'gcr.forust.xyz/forust/{name}'
|
||||||
report['current'] = name
|
inputs = fingerprint(context, dockerfile)
|
||||||
image = f'gcr.forust.xyz/forust/{name}'
|
old_digest = (previous or {}).get('images', {}).get(image)
|
||||||
inputs = target['inputs']
|
exists = False
|
||||||
old_digest = target['reuse_digest']
|
if old_digest and previous['inputs'].get(image) == inputs:
|
||||||
exists = False
|
exists = (
|
||||||
if old_digest:
|
subprocess.run( # noqa: S603, S607
|
||||||
exists = (
|
[
|
||||||
subprocess.run( # noqa: S603, S607
|
shutil.which('docker') or '/usr/bin/docker',
|
||||||
[
|
'buildx',
|
||||||
shutil.which('docker') or '/usr/bin/docker',
|
'imagetools',
|
||||||
'buildx',
|
'inspect',
|
||||||
'imagetools',
|
f'{image}@{old_digest}',
|
||||||
'inspect',
|
],
|
||||||
f'{image}@{old_digest}',
|
capture_output=True,
|
||||||
],
|
env=env,
|
||||||
capture_output=True,
|
timeout=60,
|
||||||
|
).returncode
|
||||||
|
== 0
|
||||||
|
)
|
||||||
|
if exists:
|
||||||
|
print(f'Reuse {name}: inputs unchanged')
|
||||||
|
digest = old_digest
|
||||||
|
else:
|
||||||
|
print(f'Build {name}', flush=True)
|
||||||
|
metadata = Path(docker_config) / 'metadata.json'
|
||||||
|
command(
|
||||||
|
'docker',
|
||||||
|
'buildx',
|
||||||
|
'build',
|
||||||
|
'--builder',
|
||||||
|
builder,
|
||||||
|
'--push',
|
||||||
|
'--platform',
|
||||||
|
'linux/amd64',
|
||||||
|
'--provenance=false',
|
||||||
|
'--cache-from',
|
||||||
|
f'type=registry,ref={image}:buildcache',
|
||||||
|
'--cache-to',
|
||||||
|
f'type=registry,ref={image}:buildcache,mode=max',
|
||||||
|
'--tag',
|
||||||
|
f'{image}:sha-{sha}',
|
||||||
|
'--metadata-file',
|
||||||
|
str(metadata),
|
||||||
|
'--file',
|
||||||
|
dockerfile,
|
||||||
|
context,
|
||||||
env=env,
|
env=env,
|
||||||
timeout=60,
|
)
|
||||||
).returncode
|
digest = json.loads(metadata.read_text())['containerimage.digest']
|
||||||
== 0
|
release['images'][image] = digest
|
||||||
)
|
release['inputs'][image] = inputs
|
||||||
if exists:
|
validate_release(release, sha)
|
||||||
print(f'Reuse {name}: inputs unchanged')
|
|
||||||
digest = old_digest
|
|
||||||
else:
|
|
||||||
print(f'Build {name}', flush=True)
|
|
||||||
metadata = Path(docker_config) / 'metadata.json'
|
|
||||||
command(
|
|
||||||
'docker',
|
|
||||||
'buildx',
|
|
||||||
'build',
|
|
||||||
'--builder',
|
|
||||||
builder,
|
|
||||||
'--platform',
|
|
||||||
'linux/amd64',
|
|
||||||
'--provenance=false',
|
|
||||||
'--cache-from',
|
|
||||||
f'type=registry,ref={image}:buildcache',
|
|
||||||
'--cache-to',
|
|
||||||
f'type=registry,ref={image}:buildcache,mode=max',
|
|
||||||
'--output',
|
|
||||||
f'type=image,name={image},push-by-digest=true,name-canonical=true,push=true',
|
|
||||||
'--metadata-file',
|
|
||||||
str(metadata),
|
|
||||||
'--file',
|
|
||||||
dockerfile,
|
|
||||||
context,
|
|
||||||
env=env,
|
|
||||||
)
|
|
||||||
digest = json.loads(metadata.read_text())['containerimage.digest']
|
|
||||||
if not isinstance(digest, str) or not DIGEST.fullmatch(digest):
|
|
||||||
raise ValueError('Image job returned an invalid digest')
|
|
||||||
release['images'][image] = digest
|
|
||||||
release['inputs'][image] = inputs
|
|
||||||
report['reused' if exists else 'built'].append(name)
|
|
||||||
output.write_text(json.dumps(release, indent=2) + '\n')
|
output.write_text(json.dumps(release, indent=2) + '\n')
|
||||||
report['current'] = None
|
|
||||||
report['phase'] = 'Image result file saved'
|
|
||||||
finally:
|
finally:
|
||||||
# Cleanup errors must neither leak credentials nor mask the original build error.
|
# Cleanup errors must neither leak credentials nor mask the original build error.
|
||||||
try:
|
try:
|
||||||
@@ -362,63 +309,6 @@ def build_images(output, report, name, plan):
|
|||||||
shutil.rmtree(docker_config)
|
shutil.rmtree(docker_config)
|
||||||
|
|
||||||
|
|
||||||
def write_summary(lines):
|
|
||||||
path = os.environ.get('GITHUB_STEP_SUMMARY')
|
|
||||||
if path:
|
|
||||||
try:
|
|
||||||
with Path(path).open('a') as stream:
|
|
||||||
stream.write('\n'.join(lines) + '\n\n')
|
|
||||||
except OSError:
|
|
||||||
print('WARNING: cannot write the job summary')
|
|
||||||
|
|
||||||
|
|
||||||
def check_summary():
|
|
||||||
lines = [
|
|
||||||
f'## {os.environ["SUMMARY_CHECK"]}',
|
|
||||||
'',
|
|
||||||
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
|
|
||||||
f'- Result: **{os.environ["SUMMARY_RESULT"]}**',
|
|
||||||
]
|
|
||||||
if os.environ.get('SUMMARY_FAILED_STEP'):
|
|
||||||
lines.append(f'- Failed step: {os.environ["SUMMARY_FAILED_STEP"]}')
|
|
||||||
if os.environ['SUMMARY_RESULT'] != 'success':
|
|
||||||
lines.append('- Open the failed step log for the error details.')
|
|
||||||
write_summary(lines)
|
|
||||||
|
|
||||||
|
|
||||||
def build(output, name, plan):
|
|
||||||
report = {'phase': 'Check the source commit', 'current': None, 'built': [], 'reused': [], 'images': {}}
|
|
||||||
result = 'failure'
|
|
||||||
try:
|
|
||||||
build_images(output, report, name, plan)
|
|
||||||
result = 'success'
|
|
||||||
finally:
|
|
||||||
lines = [
|
|
||||||
f'## Image build result `{name}`',
|
|
||||||
'',
|
|
||||||
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
|
|
||||||
'',
|
|
||||||
f'- Result: **{result}**',
|
|
||||||
f'- Last stage: {report["phase"]}',
|
|
||||||
]
|
|
||||||
if result == 'failure':
|
|
||||||
lines.append('- This image job failed. The complete release cannot be published. Open the failed step log.')
|
|
||||||
if result == 'success':
|
|
||||||
lines.append('- This is one image result. The final build job must publish the complete release.')
|
|
||||||
if report['current']:
|
|
||||||
lines.append(f'- Image at the failure: `{report["current"]}`')
|
|
||||||
for title, key in (('Built', 'built'), ('Reused from successful CI', 'reused')):
|
|
||||||
lines.extend(['', f'### {title}'])
|
|
||||||
lines.extend(f'- `{name}`' for name in report[key])
|
|
||||||
if not report[key]:
|
|
||||||
lines.append('- None')
|
|
||||||
lines.extend(['', '### Completed image digests'])
|
|
||||||
lines.extend(f'- `{image}@{digest}`' for image, digest in report['images'].items())
|
|
||||||
if not report['images']:
|
|
||||||
lines.append('- None')
|
|
||||||
write_summary(lines)
|
|
||||||
|
|
||||||
|
|
||||||
def render(stream, destination):
|
def render(stream, destination):
|
||||||
release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA'])
|
release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA'])
|
||||||
image_line = re.compile(
|
image_line = re.compile(
|
||||||
@@ -429,6 +319,9 @@ def render(stream, destination):
|
|||||||
match = image_line.fullmatch(line.rstrip('\n'))
|
match = image_line.fullmatch(line.rstrip('\n'))
|
||||||
if match:
|
if match:
|
||||||
prefix, quote, image, tail = match.groups()
|
prefix, quote, image, tail = match.groups()
|
||||||
|
if image in EXTERNAL_IMAGES and f'{image}@sha256:' in line:
|
||||||
|
rendered.append(line)
|
||||||
|
continue
|
||||||
if image not in release['images']:
|
if image not in release['images']:
|
||||||
raise ValueError(f'Owned image missing from checked release: {image}')
|
raise ValueError(f'Owned image missing from checked release: {image}')
|
||||||
line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n'
|
line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n'
|
||||||
@@ -438,81 +331,19 @@ def render(stream, destination):
|
|||||||
destination.writelines(rendered)
|
destination.writelines(rendered)
|
||||||
|
|
||||||
|
|
||||||
def finalize_images(output, fragments, plan):
|
|
||||||
data = checked_plan(plan)
|
|
||||||
sha = data['sha']
|
|
||||||
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
|
|
||||||
for name in IMAGES:
|
|
||||||
fragment = json.loads((fragments / f'image-{name}' / 'image.json').read_text())
|
|
||||||
image = f'gcr.forust.xyz/forust/{name}'
|
|
||||||
if fragment.get('sha') != sha or fragment.get('version') != 1 or set(fragment.get('images', {})) != {image}:
|
|
||||||
raise ValueError('Image job artifact is missing or belongs to another commit')
|
|
||||||
target = next(t for t in data['targets'] if t['name'] == name)
|
|
||||||
if fragment.get('inputs') != {image: target['inputs']}:
|
|
||||||
raise ValueError('Image artifact does not match the build plan')
|
|
||||||
release['images'].update(fragment['images'])
|
|
||||||
release['inputs'].update(fragment['inputs'])
|
|
||||||
validate_release(release, sha)
|
|
||||||
# Only a complete set of successful image jobs can publish the release tags.
|
|
||||||
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
|
|
||||||
env = {**os.environ, 'DOCKER_CONFIG': docker_config}
|
|
||||||
try:
|
|
||||||
subprocess.run( # noqa: S603, S607
|
|
||||||
[
|
|
||||||
shutil.which('docker') or '/usr/bin/docker',
|
|
||||||
'login',
|
|
||||||
'gcr.forust.xyz',
|
|
||||||
'-u',
|
|
||||||
os.environ['REGISTRY_USERNAME'],
|
|
||||||
'--password-stdin',
|
|
||||||
],
|
|
||||||
input=os.environ['REGISTRY_PASSWORD'],
|
|
||||||
text=True,
|
|
||||||
check=True,
|
|
||||||
env=env,
|
|
||||||
)
|
|
||||||
for image, digest in release['images'].items():
|
|
||||||
command(
|
|
||||||
'docker',
|
|
||||||
'buildx',
|
|
||||||
'imagetools',
|
|
||||||
'create',
|
|
||||||
'--prefer-index=false',
|
|
||||||
'--tag',
|
|
||||||
f'{image}:sha-{sha}',
|
|
||||||
f'{image}@{digest}',
|
|
||||||
env=env,
|
|
||||||
timeout=90,
|
|
||||||
)
|
|
||||||
output.write_text(json.dumps(release, indent=2) + '\n')
|
|
||||||
finally:
|
|
||||||
shutil.rmtree(docker_config)
|
|
||||||
|
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
parser = argparse.ArgumentParser(description=__doc__)
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
parser.add_argument('action', choices=('prepare', 'image', 'finalize', 'gate', 'render', 'check-summary'))
|
parser.add_argument('action', choices=('build', 'gate', 'render'))
|
||||||
parser.add_argument('--output', type=Path, default=Path('release.json'))
|
parser.add_argument('--output', type=Path, default=Path('release.json'))
|
||||||
parser.add_argument('--ref', default='main')
|
parser.add_argument('--ref', default='main')
|
||||||
parser.add_argument('--event-sha', default='')
|
parser.add_argument('--event-sha', default='')
|
||||||
parser.add_argument('--image', choices=IMAGES)
|
|
||||||
parser.add_argument('--plan', type=Path, default=Path('build-plan.json'))
|
|
||||||
parser.add_argument('--fragments', type=Path, default=Path('artifacts'))
|
|
||||||
args = parser.parse_args()
|
args = parser.parse_args()
|
||||||
if args.action == 'check-summary':
|
if args.action == 'render':
|
||||||
check_summary()
|
|
||||||
elif args.action == 'render':
|
|
||||||
render(sys.stdin, sys.stdout)
|
render(sys.stdin, sys.stdout)
|
||||||
elif args.action == 'gate':
|
elif args.action == 'gate':
|
||||||
gate(args.output, args.ref, args.event_sha)
|
gate(args.output, args.ref, args.event_sha)
|
||||||
elif args.action == 'prepare':
|
|
||||||
prepare_images(args.output)
|
|
||||||
elif args.action == 'image':
|
|
||||||
if not args.image:
|
|
||||||
parser.error('--image is required')
|
|
||||||
build(args.output, args.image, args.plan)
|
|
||||||
else:
|
else:
|
||||||
finalize_images(args.output, args.fragments, args.plan)
|
build(args.output)
|
||||||
|
|
||||||
|
|
||||||
if __name__ == '__main__':
|
if __name__ == '__main__':
|
||||||
|
|||||||
@@ -1,9 +1,7 @@
|
|||||||
name: renovate-ci
|
name: renovate-ci
|
||||||
|
|
||||||
on:
|
on:
|
||||||
# Read the workflow from the trusted base branch. PR code runs only on the
|
pull_request:
|
||||||
# unprivileged runner selected below.
|
|
||||||
pull_request_target:
|
|
||||||
paths:
|
paths:
|
||||||
- "renovate/**"
|
- "renovate/**"
|
||||||
- ".gitea/workflows/renovate-ci.yaml"
|
- ".gitea/workflows/renovate-ci.yaml"
|
||||||
@@ -28,47 +26,37 @@ permissions:
|
|||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
validate-renovate:
|
validate-renovate:
|
||||||
runs-on: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
|
runs-on: homelab
|
||||||
timeout-minutes: 20
|
timeout-minutes: 20
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||||
with:
|
|
||||||
ref: ${{ github.event_name == 'pull_request_target' && github.event.pull_request.head.sha || github.sha }}
|
|
||||||
|
|
||||||
# renovate/k8s/cronjob.yaml is the single source of truth for the version.
|
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag,
|
||||||
- name: Resolve the deployed Renovate version
|
# so the same version that runs in the cluster is the one validated here.
|
||||||
|
- name: Resolve the deployed Renovate image
|
||||||
id: image
|
id: image
|
||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
||||||
renovate/k8s/cronjob.yaml | head -1)"
|
renovate/k8s/cronjob.yaml | head -1)"
|
||||||
if [[ ! "$image" =~ ^renovate/renovate:([0-9]+\.[0-9]+\.[0-9]+)$ ]]; then
|
if [ -z "$image" ]; then
|
||||||
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
|
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
version="${BASH_REMATCH[1]}"
|
echo "using $image"
|
||||||
echo "using Renovate $version"
|
echo "image=$image" >> "$GITHUB_OUTPUT"
|
||||||
printf 'version=%s\n' "$version" >> "$GITHUB_OUTPUT"
|
|
||||||
|
|
||||||
- name: Prepare pinned validation tools
|
|
||||||
shell: bash
|
|
||||||
run: |
|
|
||||||
set -euo pipefail
|
|
||||||
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform node)"
|
|
||||||
echo "$tools_dir" >> "$GITHUB_PATH"
|
|
||||||
|
|
||||||
- name: Validate Renovate repository config
|
- name: Validate Renovate repository config
|
||||||
shell: bash
|
shell: bash
|
||||||
env:
|
|
||||||
RENOVATE_VERSION: ${{ steps.image.outputs.version }}
|
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
npm_cache="$(mktemp -d "${RUNNER_TEMP:-/tmp}/renovate-npm-cache.XXXXXXXX")"
|
docker run --rm \
|
||||||
trap 'rm -rf "$npm_cache"' EXIT
|
-v "$PWD/renovate:/opt/renovate:ro" \
|
||||||
NPM_CONFIG_CACHE="$npm_cache" RENOVATE_CONFIG_FILE="$PWD/renovate/renovate.json" \
|
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||||
npm exec --yes --package="renovate@${RENOVATE_VERSION}" -- renovate-config-validator
|
"${{ steps.image.outputs.image }}" \
|
||||||
|
renovate-config-validator /opt/renovate/renovate.json
|
||||||
|
|
||||||
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
|
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
|
||||||
# carries an inlined copy of the config. Fail if it no longer matches.
|
# carries an inlined copy of the config. Fail if it no longer matches.
|
||||||
@@ -82,6 +70,8 @@ jobs:
|
|||||||
shell: bash
|
shell: bash
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
|
||||||
|
export PATH="$tools_dir:$PATH"
|
||||||
kubeconform \
|
kubeconform \
|
||||||
-strict \
|
-strict \
|
||||||
-ignore-missing-schemas \
|
-ignore-missing-schemas \
|
||||||
|
|||||||
@@ -32,14 +32,11 @@ concurrency:
|
|||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
run-renovate:
|
run-renovate:
|
||||||
if: github.ref == 'refs/heads/main'
|
|
||||||
runs-on: homelab
|
runs-on: homelab
|
||||||
timeout-minutes: 60
|
timeout-minutes: 60
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout repository
|
- name: Checkout repository
|
||||||
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
||||||
with:
|
|
||||||
ref: refs/heads/main
|
|
||||||
|
|
||||||
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
|
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
|
||||||
# Reading it here means this workflow validates and runs the exact version
|
# Reading it here means this workflow validates and runs the exact version
|
||||||
@@ -51,23 +48,21 @@ jobs:
|
|||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
|
||||||
renovate/k8s/cronjob.yaml | head -1)"
|
renovate/k8s/cronjob.yaml | head -1)"
|
||||||
if [[ ! "$image" =~ ^renovate/renovate:[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
|
if [ -z "$image" ]; then
|
||||||
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
|
echo "::error::no renovate/renovate image found in renovate/k8s/cronjob.yaml"
|
||||||
exit 1
|
exit 1
|
||||||
fi
|
fi
|
||||||
echo "using $image"
|
echo "using $image"
|
||||||
printf 'image=%s\n' "$image" >> "$GITHUB_OUTPUT"
|
echo "image=$image" >> "$GITHUB_OUTPUT"
|
||||||
|
|
||||||
- name: Validate Renovate config
|
- name: Validate Renovate config
|
||||||
shell: bash
|
shell: bash
|
||||||
env:
|
|
||||||
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
|
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
docker run --rm \
|
docker run --rm \
|
||||||
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
|
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \
|
||||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||||
"$RENOVATE_IMAGE" \
|
"${{ steps.image.outputs.image }}" \
|
||||||
renovate-config-validator
|
renovate-config-validator
|
||||||
|
|
||||||
- name: Run Renovate
|
- name: Run Renovate
|
||||||
@@ -78,7 +73,6 @@ jobs:
|
|||||||
RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
|
RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
|
||||||
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
|
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
|
||||||
LOG_LEVEL: ${{ inputs.log_level }}
|
LOG_LEVEL: ${{ inputs.log_level }}
|
||||||
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
|
|
||||||
run: |
|
run: |
|
||||||
set -euo pipefail
|
set -euo pipefail
|
||||||
|
|
||||||
@@ -95,4 +89,4 @@ jobs:
|
|||||||
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \
|
||||||
-e RENOVATE_BASE_DIR=/tmp/renovate \
|
-e RENOVATE_BASE_DIR=/tmp/renovate \
|
||||||
-e LOG_LEVEL="${LOG_LEVEL:-info}" \
|
-e LOG_LEVEL="${LOG_LEVEL:-info}" \
|
||||||
"$RENOVATE_IMAGE"
|
"${{ steps.image.outputs.image }}"
|
||||||
@@ -20,7 +20,7 @@ ssh_opts=(-i "$key_dir/key" -p "${DEPLOY_PORT:-22}" -o BatchMode=yes -o StrictHo
|
|||||||
-o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15
|
-o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15
|
||||||
-o ServerAliveInterval=15 -o ServerAliveCountMax=4)
|
-o ServerAliveInterval=15 -o ServerAliveCountMax=4)
|
||||||
controller=.local/lib/homelab-deploy/controller.py
|
controller=.local/lib/homelab-deploy/controller.py
|
||||||
case "${1:?start, apply, verify, smoke or summary required}" in
|
case "${1:?start, apply, verify or smoke required}" in
|
||||||
start)
|
start)
|
||||||
python3 - <<'PY' >"$key_dir/request.json"
|
python3 - <<'PY' >"$key_dir/request.json"
|
||||||
import json
|
import json
|
||||||
@@ -41,30 +41,16 @@ PY
|
|||||||
exit "$rc"
|
exit "$rc"
|
||||||
;;
|
;;
|
||||||
apply|verify|smoke)
|
apply|verify|smoke)
|
||||||
result=0
|
|
||||||
for attempt in 1 2 3; do
|
for attempt in 1 2 3; do
|
||||||
rc=0
|
rc=0
|
||||||
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
|
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
|
||||||
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$?
|
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$?
|
||||||
[ "$rc" -eq 0 ] && break
|
[ "$rc" -eq 0 ] && exit 0
|
||||||
[ "$rc" -eq 255 ] || { result="$rc"; break; }
|
[ "$rc" -eq 255 ] || exit "$rc"
|
||||||
echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)"
|
echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)"
|
||||||
if [ "$attempt" -eq 3 ]; then result=255; break; fi
|
|
||||||
sleep 5
|
sleep 5
|
||||||
done
|
done
|
||||||
exit "$result"
|
exit "$rc"
|
||||||
;;
|
|
||||||
summary)
|
|
||||||
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
|
|
||||||
rc=0
|
|
||||||
# shellcheck disable=SC2029 # The run ID is validated above.
|
|
||||||
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" summary "$DEPLOY_RUN_ID" >"$key_dir/deploy-summary.md" || rc=$?
|
|
||||||
if [ "$rc" -eq 0 ]; then
|
|
||||||
cat "$key_dir/deploy-summary.md" >>"$GITHUB_STEP_SUMMARY" || echo "WARNING: cannot write the deploy summary"
|
|
||||||
else
|
|
||||||
echo 'Deploy summary is unavailable. The SSH connection failed or the controller did not respond. Check the job log.' >>"$GITHUB_STEP_SUMMARY" || true
|
|
||||||
fi
|
|
||||||
fi
|
|
||||||
;;
|
;;
|
||||||
*) echo "Unknown SSH operation: $1" >&2; exit 1 ;;
|
*) echo "Unknown SSH operation: $1" >&2; exit 1 ;;
|
||||||
esac
|
esac
|
||||||
@@ -94,6 +94,7 @@ replacements.txt
|
|||||||
.idea
|
.idea
|
||||||
|
|
||||||
# Temp files
|
# Temp files
|
||||||
|
edu_master/temp/
|
||||||
temp/*
|
temp/*
|
||||||
# Local-only tooling scratch space (pinned CI tools, verification scripts)
|
# Local-only tooling scratch space (pinned CI tools, verification scripts)
|
||||||
tmp/
|
tmp/
|
||||||
|
|||||||
@@ -1,163 +0,0 @@
|
|||||||
# Homelab
|
|
||||||
|
|
||||||
Configuration for my homelab: Kubernetes manifests, Docker Compose stacks, and the
|
|
||||||
Gitea Actions that build and deploy them. Most applications have both deployment
|
|
||||||
formats. Headscale and Nextcloud AIO have Compose deployments with Kubernetes
|
|
||||||
ingress; the media stack has Compose and Kubernetes routing configuration.
|
|
||||||
|
|
||||||
These files contain this lab's domains, IP addresses, storage paths, and private
|
|
||||||
registry names. Running them on another machine takes some editing.
|
|
||||||
|
|
||||||
## Start here
|
|
||||||
|
|
||||||
- [Service list](#services) — what each directory contains.
|
|
||||||
- [Deployment workflow](.gitea/README.md) — selection, validation, and recovery.
|
|
||||||
- [Repository review](docs/repository-review.md) — findings from the 6 October baseline and their status.
|
|
||||||
- [EDU ownership handoff](.gitea/EDU_HANDOFF.md) — the EDU workloads now live in their own repository.
|
|
||||||
- [Shared PostgreSQL](postgres/README.md), [Traefik](traefik/README.md), and
|
|
||||||
[cert-manager](cert-manager/README.md) — common dependencies.
|
|
||||||
|
|
||||||
## What gets deployed
|
|
||||||
|
|
||||||
The `active` files are switches for the deploy workflow, not health indicators.
|
|
||||||
|
|
||||||
| File | Effect |
|
|
||||||
| ---------------------- | ----------------------------------------------------------- |
|
|
||||||
| `<service>/active` | Include that directory's `compose.yaml` or `compose.yml`. |
|
|
||||||
| `<service>/k8s/active` | Include its Kubernetes manifests or Kustomize overlay. |
|
|
||||||
| Both | Run the Compose stack and apply the Kubernetes resources. |
|
|
||||||
| Neither | Keep the configuration in Git without automatic deployment. |
|
|
||||||
|
|
||||||
`shared-compose.yaml`, `client.compose.yaml`, and `renovate-compose.yaml` are
|
|
||||||
manual entry points. The deploy script does not discover them.
|
|
||||||
|
|
||||||
Kubernetes selection excludes secret files, examples, Helm values, and patches.
|
|
||||||
Helm releases listed in `deploy-lib.sh` are upgraded separately. Traefik,
|
|
||||||
cert-manager, and CrowdSec have additional bootstrap steps; an `active` marker
|
|
||||||
does not install their charts.
|
|
||||||
|
|
||||||
The table below describes committed configuration. It does not claim that a
|
|
||||||
service is currently healthy or running.
|
|
||||||
|
|
||||||
## Services
|
|
||||||
|
|
||||||
| Service | Configuration | Selected by markers |
|
|
||||||
| ---------------------------------------------- | ---------------------------- | ------------------- |
|
|
||||||
| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual |
|
|
||||||
| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual |
|
|
||||||
| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual |
|
|
||||||
| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual |
|
|
||||||
| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual |
|
|
||||||
| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual |
|
|
||||||
| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Glance](glance/README.md) | Kubernetes + Compose | Manual |
|
|
||||||
| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
|
|
||||||
| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual |
|
|
||||||
| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Kener](kener/README.md) | Kubernetes + Compose | Manual |
|
|
||||||
| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes |
|
|
||||||
| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [n8n](n8n/README.md) | Kubernetes + Compose | Manual |
|
|
||||||
| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
|
|
||||||
| [Penpot](penpot/README.md) | Compose | Manual |
|
|
||||||
| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual |
|
|
||||||
| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes |
|
|
||||||
| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual |
|
|
||||||
| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Manual |
|
|
||||||
| [Termix](termix/README.md) | Kubernetes + Compose | Manual |
|
|
||||||
| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes |
|
|
||||||
| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes |
|
|
||||||
|
|
||||||
## Running a Compose stack
|
|
||||||
|
|
||||||
Use the service README first. Where a service has an env example, copy it inside
|
|
||||||
that service's directory and replace the placeholders. The root `.env.example`
|
|
||||||
is an older collection of variables, not a complete configuration for every stack.
|
|
||||||
|
|
||||||
For example, from the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
cd netbox
|
|
||||||
cp .env.example .env
|
|
||||||
$EDITOR .env
|
|
||||||
docker compose config --quiet
|
|
||||||
docker compose up -d
|
|
||||||
docker compose ps
|
|
||||||
```
|
|
||||||
|
|
||||||
Stacks that attach to `proxy` require an existing Docker network of that name and
|
|
||||||
an appropriate reverse proxy. Published host ports still work independently of
|
|
||||||
Traefik. Check port conflicts before starting an alternative to a Kubernetes
|
|
||||||
service: DNS, STUN, and HTTP listeners can share the same host.
|
|
||||||
|
|
||||||
`docker compose down` keeps named volumes. Adding `-v` removes them.
|
|
||||||
|
|
||||||
## Preparing Kubernetes
|
|
||||||
|
|
||||||
The manifests assume Traefik CRDs, cert-manager, and a working storage provisioner.
|
|
||||||
PrometheusRule and ServiceMonitor resources also need the Prometheus Operator.
|
|
||||||
Replace the lab's hosts and addresses before using the configuration elsewhere.
|
|
||||||
|
|
||||||
Create a service's namespace, then prepare its ignored Secret from the example.
|
|
||||||
For example:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl apply -f netbox/k8s/namespace.yaml
|
|
||||||
cp netbox/k8s/secrets.yaml.example netbox/k8s/secrets.yaml
|
|
||||||
$EDITOR netbox/k8s/secrets.yaml
|
|
||||||
kubectl apply -f netbox/k8s/secrets.yaml
|
|
||||||
```
|
|
||||||
|
|
||||||
The deploy workflow applies the tracked resources for marked services. Avoid
|
|
||||||
applying an entire `k8s/` directory blindly: some directories contain Helm values,
|
|
||||||
examples, and alternative routes. For a manual change, apply the selected manifest
|
|
||||||
explicitly and check the resulting rollout.
|
|
||||||
|
|
||||||
Shared database passwords must agree between the `database` namespace and each
|
|
||||||
application's Secret. Updating the PostgreSQL Secret does not change an existing
|
|
||||||
role's password; see the database README.
|
|
||||||
|
|
||||||
## Local checks
|
|
||||||
|
|
||||||
CI pins its tools in `.gitea/workflows/tool-versions.env`. Use the same versions:
|
|
||||||
|
|
||||||
```fish
|
|
||||||
set tools_dir (bash .gitea/workflows/install-ci-tools.sh)
|
|
||||||
set -gx PATH $tools_dir $PATH
|
|
||||||
ruff check .
|
|
||||||
ruff format --check .
|
|
||||||
actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
|
|
||||||
.gitea/workflows/sync-renovate-configmap.sh --check
|
|
||||||
```
|
|
||||||
|
|
||||||
The [workflow README](.gitea/README.md#ci) lists the rest of the checks.
|
|
||||||
Structure checks do not establish that local Secrets, mounted files, storage,
|
|
||||||
or external services are ready.
|
|
||||||
|
|
||||||
## Data and recovery
|
|
||||||
|
|
||||||
State lives outside Git: PVCs, Docker volumes, bind mounts, databases, and ignored
|
|
||||||
configuration. Keep backups of application data and the keys needed to read it.
|
|
||||||
An image rollback does not roll back database migrations or ConfigMap contents.
|
|
||||||
|
|
||||||
Many PVCs use the cluster's default StorageClass; monitoring explicitly uses
|
|
||||||
`local-path`. Check the PV reclaim policy before deleting a PVC or namespace.
|
|
||||||
The manifests do not provide a repository-wide backup schedule.
|
|
||||||
|
|
||||||
`incident-archive/` contains past incident notes. `.docs/storage-audit-instruction.md`
|
|
||||||
is a planning document, not evidence that NFS has been installed.
|
|
||||||
@@ -1,22 +0,0 @@
|
|||||||
# AdGuard Home
|
|
||||||
|
|
||||||
DNS filtering with a web UI, DNS-over-TLS, and certificates from cert-manager.
|
|
||||||
|
|
||||||
The Kubernetes namespace is `adguard`. The workload uses `adguard-pvc` for
|
|
||||||
configuration and working data, and mounts the `adguard-certs` TLS Secret.
|
|
||||||
The LoadBalancer Service exposes DNS separately from the web ingress.
|
|
||||||
|
|
||||||
The Compose stack publishes TCP/UDP 53 and TCP 853 on the host. Prepare `conf/`
|
|
||||||
and `certs/` before starting it. Starting both DNS deployments on the same address
|
|
||||||
can cause a port conflict.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n adguard
|
|
||||||
kubectl get events -n adguard --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,22 +0,0 @@
|
|||||||
# Authentik
|
|
||||||
|
|
||||||
Identity provider with separate server and worker deployments.
|
|
||||||
|
|
||||||
Kubernetes connects to the shared PostgreSQL service in `database`. Set
|
|
||||||
`AUTHENTIK_DB_PASSWORD` to the same value in both database and application Secrets.
|
|
||||||
Keep `AUTHENTIK_SECRET_KEY` with the backups.
|
|
||||||
|
|
||||||
Compose uses its own PostgreSQL 15 container and bind-mounted media and templates.
|
|
||||||
Its image defaults differ from Kubernetes; check both before an upgrade.
|
|
||||||
The worker mounts the Docker socket for Docker outpost management.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n authentik
|
|
||||||
kubectl get events -n authentik --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,18 +0,0 @@
|
|||||||
# cert-manager
|
|
||||||
|
|
||||||
Public ACME issuers and an internal certificate authority.
|
|
||||||
|
|
||||||
This directory contains chart values and issuer resources, not the controller
|
|
||||||
installation. Install the cert-manager chart with CRDs and the settings in
|
|
||||||
`k8s/cert-manager-values.yaml` before applying the issuers.
|
|
||||||
|
|
||||||
`clusterissuer.yaml` defines staging and production Let's Encrypt issuers.
|
|
||||||
They use HTTP-01 through the Traefik ingress class. Public DNS and inbound HTTP
|
|
||||||
reachability must work for the requested names before issuance.
|
|
||||||
`internal-ca.yaml` bootstraps the internal CA. Keep its private-key Secret backed
|
|
||||||
up; the tracked `.crt` is only a public certificate.
|
|
||||||
|
|
||||||
This directory has no `k8s/active` marker. Apply the issuer files deliberately;
|
|
||||||
`kubectl apply` does not interpret the Helm values file.
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,22 +0,0 @@
|
|||||||
# Cloudflare DDNS
|
|
||||||
|
|
||||||
Updates the lab DNS records when the public address changes.
|
|
||||||
|
|
||||||
Kubernetes runs in `default` with host networking and reads `cfddns-secrets`.
|
|
||||||
The Compose stack also uses host networking. Configure the API token and domain
|
|
||||||
list from the relevant example; keep DNS names consistent with the ingress rules.
|
|
||||||
|
|
||||||
`config.json.example` is a separate configuration example. The current Compose
|
|
||||||
file does not mount a config.json file. Check configuration against the pinned
|
|
||||||
DDNS image when changing between environment and file-based settings.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n default
|
|
||||||
kubectl get events -n default --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,21 +0,0 @@
|
|||||||
# Checkmk
|
|
||||||
|
|
||||||
Checkmk Raw monitoring site with web and agent-receiver ingress.
|
|
||||||
|
|
||||||
The site data lives in `checkmk-sites-pvc` on Kubernetes and the `sites` named
|
|
||||||
volume on Compose. The agent receiver has a separate TCP route; enabling the
|
|
||||||
web route alone does not expose it.
|
|
||||||
|
|
||||||
Prepare the password in the service env or Secret example. Inspect the Checkmk
|
|
||||||
container logs during the first site creation.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n checkmk
|
|
||||||
kubectl get events -n checkmk --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,21 +0,0 @@
|
|||||||
# Cloudflare Tunnel
|
|
||||||
|
|
||||||
A Kubernetes connector for an existing Cloudflare tunnel.
|
|
||||||
|
|
||||||
The Deployment runs in `default` and reads its token from the ignored Secret
|
|
||||||
created from `k8s/secret.yaml.example`. Create the tunnel and its hostname rules
|
|
||||||
in Cloudflare before starting the connector.
|
|
||||||
|
|
||||||
There is no Compose file or `k8s/active` marker. Apply the Secret first, then
|
|
||||||
`k8s/deployment.yaml` when this tunnel is needed.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n default
|
|
||||||
kubectl get events -n default --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,22 +0,0 @@
|
|||||||
# File converters
|
|
||||||
|
|
||||||
ConvertX for server-side conversion and BentoPDF for PDF tools.
|
|
||||||
|
|
||||||
ConvertX persists files in `convertx-pvc`; BentoPDF has no persistent volume.
|
|
||||||
Kubernetes configuration includes a local `config.yaml.example`, excluded from
|
|
||||||
normal deployment. Copy and apply the real ConfigMap separately where required.
|
|
||||||
|
|
||||||
Compose publishes ConvertX on host port 9992 as well as attaching it to the
|
|
||||||
proxy network. Replace the authentication settings from `.env.example` before
|
|
||||||
exposing it outside the lab.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n converters
|
|
||||||
kubectl get events -n converters --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,25 +0,0 @@
|
|||||||
# CrowdSec
|
|
||||||
|
|
||||||
Helm values, dashboards, network policy, and a maintenance CronJob.
|
|
||||||
|
|
||||||
Install CrowdSec separately using `k8s/crowdsec-values.yaml`; the deploy
|
|
||||||
workflow does not have a CrowdSec Helm release entry. There is no `k8s/active`
|
|
||||||
marker in this directory.
|
|
||||||
|
|
||||||
The LAPI policy and janitor run in `crowdsec`. The dashboard ConfigMaps are in
|
|
||||||
`prometheus` for Grafana's sidecar. The janitor has its own ServiceAccount and
|
|
||||||
namespace Role. Review its script and schedule before enabling cleanup.
|
|
||||||
|
|
||||||
Traefik's values state that enforcement moved to a host firewall bouncer. This
|
|
||||||
repository does not install that host component.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n crowdsec
|
|
||||||
kubectl get events -n crowdsec --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,21 +0,0 @@
|
|||||||
# Dockmon
|
|
||||||
|
|
||||||
Docker management UI that talks to the host Docker daemon.
|
|
||||||
|
|
||||||
Both runtimes mount `/var/run/docker.sock`. On Kubernetes the socket belongs
|
|
||||||
to the node hosting the pod, so this is not a cluster-wide container manager.
|
|
||||||
|
|
||||||
Compose stores application data in a named volume. Kubernetes uses a StatefulSet
|
|
||||||
with a volume claim template. Its ServersTransport is specific to the upstream
|
|
||||||
connection; keep it with the ingress resources.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n dockmon
|
|
||||||
kubectl get events -n dockmon --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,134 +0,0 @@
|
|||||||
# Repository review (6 October 2026 baseline)
|
|
||||||
|
|
||||||
This records the tracked tree at `cc9c3de` and the workstation state observed on
|
|
||||||
6 October 2026. It is a historical review, not a current runtime inventory. The
|
|
||||||
listed code fixes have since merged into `main`; EDU ownership has moved to the
|
|
||||||
separate repository described in [the handoff record](../.gitea/EDU_HANDOFF.md).
|
|
||||||
See the [CI and deployment guide](../.gitea/README.md) and
|
|
||||||
[runner and recovery guide](../.gitea/runner/README.md) for the current workflow.
|
|
||||||
No deployment was performed during the original review.
|
|
||||||
|
|
||||||
## Findings at the baseline and current status
|
|
||||||
|
|
||||||
| Priority | Finding at the baseline | Current status |
|
|
||||||
| -------- | ---------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
|
|
||||||
| High | Per-file `APPLY_PRUNE=true` could delete resources selected by a shared label. | The deploy workflow rejects unsafe pruning before applying resources. |
|
|
||||||
| High | Compose validation did not resolve the local configuration required at deploy time. | Preflight resolves the selected Compose configuration before apply. |
|
|
||||||
| Medium | Secret validation could miss namespace-specific and mounted Secret references. | Preflight checks rendered references in their namespaces, including mounted and projected Secrets. |
|
|
||||||
| Medium | Compose CI missed manual entry points such as `shared-compose.yaml` and `client.compose.yaml`. | CI checks all tracked Compose files. |
|
|
||||||
| Medium | NetBird Compose referenced missing setup and renderer files. | The setup and renderer files are now present; Compose remains a manual alternative to the active Kubernetes deployment. |
|
|
||||||
| Medium | Glance mounted its CSS from the wrong ConfigMap. | The mount now uses the ConfigMap that contains `user.css`. |
|
|
||||||
| Medium | The PostgreSQL env example omitted the required NetBox password. | The example now includes the required variable. |
|
|
||||||
| Medium | The former EDU code had stale Compose variable names and session reliability problems. | EDU workloads and their fixes moved out of this repository; see the handoff record. |
|
|
||||||
| Medium | AdGuard DoH and SearXNG Compose router expressions used invalid `Host(...)` syntax. | The router expressions now follow Traefik's rule syntax. |
|
|
||||||
|
|
||||||
Traefik matchers should be combined as `Host(a) || Host(b)`; the rule syntax is
|
|
||||||
described in the [Traefik rules documentation](https://doc.traefik.io/traefik/reference/routing-configuration/http/routing/rules-and-priority/).
|
|
||||||
The fix retains the DoH path constraint for both hostnames.
|
|
||||||
|
|
||||||
The current deploy workflow deliberately rejects the unsafe prune option. It
|
|
||||||
does not introduce automatic deletion under a different implementation. The
|
|
||||||
baseline finding was a configuration risk, not evidence of a live deletion
|
|
||||||
incident.
|
|
||||||
|
|
||||||
The former session fix bounded HTTP and Redis calls, validated credentials, set
|
|
||||||
a cookie lifetime of two refresh intervals, and marked success only after
|
|
||||||
publishing the verified cookie. The service is now owned by the EDU repository;
|
|
||||||
see that repository for its current implementation.
|
|
||||||
|
|
||||||
The deployment fix extracts required pod Secret references from rendered JSON,
|
|
||||||
checks their namespaces, includes init containers, image-pull credentials, and
|
|
||||||
mounted/projected Secrets, and honors optional references. Ingress TLS Secrets
|
|
||||||
issued by cert-manager are not treated as pre-existing pod prerequisites.
|
|
||||||
It checks existence/access, not every key's contents or application validity.
|
|
||||||
|
|
||||||
## Live workstation observations
|
|
||||||
|
|
||||||
The SSH alias `workstation` is reachable. It has one Ready control-plane node,
|
|
||||||
Kubernetes `v1.35.4+k0s`, and a Docker daemon alongside containerd. At inspection,
|
|
||||||
no pods were Pending or in another non-running, non-completed phase. This is a
|
|
||||||
point-in-time observation, not a complete application health test.
|
|
||||||
|
|
||||||
The deployment checkout at `/srv/homelab` is on main commit `2adf17c`, behind the
|
|
||||||
reviewed local commit. It has untracked host configuration and a separate
|
|
||||||
`userbot/` directory. It was not reset or cleaned.
|
|
||||||
|
|
||||||
| Observed difference | Implication |
|
|
||||||
| ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
||||||
| VictoriaMetrics and vmalert are running; the Prometheus StatefulSet has zero replicas. | A monitoring migration is already in progress outside committed main. Deploying the old Helm values can overwrite those settings. |
|
|
||||||
| Homarr, Cloudflared, and Reloader are installed without their current Git active markers. | Installed services and marker-selected services are different inventories. Missing markers do not establish that a service is stopped. |
|
|
||||||
| Cloudflare DDNS is running in both Docker and Kubernetes. | Confirm which instance should own DNS updates and whether their domain lists overlap before retiring either one. Secret values were not inspected. |
|
|
||||||
| Traefik's LoadBalancer exposes port 8080 at `192.168.80.2`. | The direct API listener is deployed; its external reachability was not tested. |
|
|
||||||
| Default `local-path` has reclaim policy Delete, while many existing PVs have been changed to Retain. | Current retention is partly live state. Recreating a claim can get a different policy from the old PV. |
|
|
||||||
| NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. | Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup. |
|
|
||||||
|
|
||||||
The VictoriaMetrics monitoring trial later merged into `main` in PR #95. The
|
|
||||||
first row above records the state before that change. Read
|
|
||||||
[`prometheus-stack/README.md`](../prometheus-stack/README.md) for the current
|
|
||||||
tracked monitoring configuration; the live observations in this section remain
|
|
||||||
a snapshot from 6 October.
|
|
||||||
|
|
||||||
## Current recovery limits
|
|
||||||
|
|
||||||
The deployment controller and its recovery process changed after this review.
|
|
||||||
The current operator workflow is documented in the
|
|
||||||
[runner and recovery guide](../.gitea/runner/README.md). The remaining boundaries
|
|
||||||
are:
|
|
||||||
|
|
||||||
- Kubernetes recovery can restore captured workload revisions. It does not
|
|
||||||
restore ConfigMaps, Secrets, database schemas, or persistent data.
|
|
||||||
- Compose recovery is manual. It uses saved resolved configuration, but it does
|
|
||||||
not restore volume data or reverse database migrations.
|
|
||||||
- Removed resources require manual review and removal; the deploy workflow does
|
|
||||||
not prune them automatically.
|
|
||||||
- Plan mode does not create namespaces. During apply, server validation for new
|
|
||||||
namespaces runs after namespace creation and chart installation; a failed
|
|
||||||
check can leave an empty namespace.
|
|
||||||
- Storage policy and backup coverage remain service-specific. Check the live PV,
|
|
||||||
PVC, and backup state before changing stateful workloads.
|
|
||||||
|
|
||||||
## Validation
|
|
||||||
|
|
||||||
At the review baseline, lint checks passed for Python, shell, workflows, YAML, standard Compose
|
|
||||||
files, and Kubernetes resources with available schemas. Kubeconform found 347
|
|
||||||
resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources.
|
|
||||||
That skip count matters: passing schema validation does not validate Traefik rule
|
|
||||||
strings or other controller-specific behavior.
|
|
||||||
|
|
||||||
Fix validation covers:
|
|
||||||
|
|
||||||
- Compose discovery of manual entry points, rejection of required-variable gaps,
|
|
||||||
namespace-scoped and optional Secret references, and API/render failures.
|
|
||||||
- NetBird setup idempotence, preservation of existing keys, file permissions,
|
|
||||||
runtime rendering, and rejection of invalid trusted proxy CIDRs.
|
|
||||||
- Session refresh success and failure paths, timeouts, cookie expiry, log redaction,
|
|
||||||
missing credentials, and nonpositive refresh intervals.
|
|
||||||
- Correct Glance ConfigMap key selection and PostgreSQL initializer/env alignment.
|
|
||||||
- YAML and Compose structure for the corrected router rules, compared with the
|
|
||||||
documented Traefik grammar. They were not exercised on the live proxy.
|
|
||||||
- Prune rejection before any cluster invocation.
|
|
||||||
|
|
||||||
At the time of review, all seven fix branches and the documentation branch
|
|
||||||
merged together in a disposable validation worktree. That combined tree passed the
|
|
||||||
CI-equivalent local checks, Markdown formatting/lint and link checks, all 35
|
|
||||||
Compose structure checks, and 11 Python regression tests plus the shell
|
|
||||||
validation regressions. CRD server-side validation and live rollout tests were
|
|
||||||
not run.
|
|
||||||
|
|
||||||
Runtime tests use fixtures and mocks, not production credentials. Live checks read
|
|
||||||
workload metadata, storage policies, chart versions, and container state only.
|
|
||||||
They did not read Secret contents or change services.
|
|
||||||
|
|
||||||
## Reloader follow-up (baseline)
|
|
||||||
|
|
||||||
`fix/reloader-integration` added the active marker and opt-in annotations to
|
|
||||||
application Deployments/StatefulSets that consume runtime ConfigMaps or Secrets.
|
|
||||||
It corrected AdGuard's misplaced pod-template annotation. The Helm settings use
|
|
||||||
annotation-based reloads, keep global auto-reload disabled, and ignore Jobs and
|
|
||||||
CronJobs. PostgreSQL workloads are excluded because their credential variables
|
|
||||||
and init scripts are only effective on an empty data directory.
|
|
||||||
|
|
||||||
The controller was running on the workstation when inspected. The original
|
|
||||||
review checked configuration against the pinned chart with Helm rendering and
|
|
||||||
manifest validation; it did not change production configuration to provoke a
|
|
||||||
test restart or confirm every application's live reload behavior.
|
|
||||||
@@ -1,20 +0,0 @@
|
|||||||
# Downtify
|
|
||||||
|
|
||||||
Download UI with a persistent downloads directory.
|
|
||||||
|
|
||||||
Compose stores downloads under `Downtify_downloads/`; Kubernetes uses
|
|
||||||
`downtify-downloads-pvc`. The ingress manifests reference shared infrastructure,
|
|
||||||
so check certificate and middleware availability before enabling them.
|
|
||||||
|
|
||||||
Back up downloads separately if they need to survive storage replacement.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n downtify
|
|
||||||
kubectl get events -n downtify --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
EDU_LOGIN=your_edu_login_here
|
||||||
|
EDU_PASSWORD=your_edu_password_here
|
||||||
|
EDU_URL_LOGIN=https://edu.edu.vn.ua/user/login
|
||||||
|
EDU_URL_VERIFY=https://edu.edu.vn.ua/course/userlist
|
||||||
|
PHPSESSID_INTERVAL=10
|
||||||
|
USER_AGENT="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36"
|
||||||
|
WEBINAR_URL=https://edu.edu.vn.ua/webinar/useractive
|
||||||
|
WEBINAR_CHECK_INTERVAL=60
|
||||||
|
REDIS_HOST=redis
|
||||||
|
REDIS_PORT=6379
|
||||||
|
PLAYWRIGHT_WS=ws://playwright-service:3000/ws
|
||||||
|
TZ=Europe/Kyiv
|
||||||
|
WEBINAR_TELEGRAM_TOKEN=your_telegram_bot_token_here
|
||||||
|
WEBINAR_ADMIN_ID=123456789
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
1.56.0
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
# EDU deployment ownership
|
||||||
|
|
||||||
|
Application source and release builds: `forust/edu-master`.
|
||||||
|
The homelab pipeline deploys `edu_master/k8s` and preserves explicit image digests.
|
||||||
|
The application copies in this directory are legacy and are not build inputs.
|
||||||
|
Do not publish EDU `prod` images from homelab or resolve releases from moving tags.
|
||||||
|
|
||||||
|
For an EDU release, validate both images, select their digests in the keeper and
|
||||||
|
checker manifests, and run the existing homelab validation/apply/verification
|
||||||
|
helpers against this service. Keep the existing Secret and Redis PVC.
|
||||||
|
Coordinate Redis authentication changes with both clients and all init/probes;
|
||||||
|
keep a pre-rollout Redis backup and both previous compatible image references.
|
||||||
|
The current HTTP checker does not depend on Playwright; check other consumers
|
||||||
|
before removing the separate browser service.
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
services:
|
||||||
|
redis:
|
||||||
|
image: redis:8.10.2-alpine
|
||||||
|
restart: unless-stopped
|
||||||
|
volumes:
|
||||||
|
- redis-data:/data
|
||||||
|
healthcheck:
|
||||||
|
test: ["CMD", "redis-cli", "ping"]
|
||||||
|
interval: 5s
|
||||||
|
timeout: 3s
|
||||||
|
retries: 5
|
||||||
|
|
||||||
|
playwright-service:
|
||||||
|
image: mcr.microsoft.com/playwright:v1.56.0-jammy
|
||||||
|
restart: unless-stopped
|
||||||
|
command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws
|
||||||
|
|
||||||
|
session-keeper:
|
||||||
|
build: ./phpsessid-bot
|
||||||
|
image: gcr.forust.xyz/forust/session-keeper:prod
|
||||||
|
pull_policy: build
|
||||||
|
env_file: .env
|
||||||
|
restart: unless-stopped
|
||||||
|
depends_on:
|
||||||
|
redis:
|
||||||
|
condition: service_healthy
|
||||||
|
healthcheck:
|
||||||
|
test: ["CMD-SHELL", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
|
||||||
|
interval: 30s
|
||||||
|
timeout: 5s
|
||||||
|
retries: 10
|
||||||
|
start_period: 60s
|
||||||
|
|
||||||
|
webinar-checker:
|
||||||
|
build: ./webinar-checker
|
||||||
|
image: gcr.forust.xyz/forust/webinar-checker:prod
|
||||||
|
pull_policy: build
|
||||||
|
env_file: .env
|
||||||
|
restart: unless-stopped
|
||||||
|
depends_on:
|
||||||
|
redis:
|
||||||
|
condition: service_healthy
|
||||||
|
session-keeper:
|
||||||
|
condition: service_healthy
|
||||||
|
playwright-service:
|
||||||
|
condition: service_started
|
||||||
|
|
||||||
|
volumes:
|
||||||
|
redis-data:
|
||||||
Whitespace-only changes.
@@ -0,0 +1,115 @@
|
|||||||
|
apiVersion: monitoring.coreos.com/v1
|
||||||
|
kind: PrometheusRule
|
||||||
|
metadata:
|
||||||
|
name: edu-master-webinar
|
||||||
|
namespace: edu-master
|
||||||
|
labels:
|
||||||
|
release: prometheus-stack
|
||||||
|
spec:
|
||||||
|
groups:
|
||||||
|
- name: edu_master.webinar
|
||||||
|
rules:
|
||||||
|
# No successful webinar check for 5m (~2-3 missed 2-min checks).
|
||||||
|
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
|
||||||
|
# The last_success > 0 guard is mandatory: checker.py initialises
|
||||||
|
# last_success to 0, so without it `time() - 0` equals the current epoch
|
||||||
|
# and humanizeDuration renders ~20722d on every pod restart. Keep the
|
||||||
|
# duration expression on the left so $value stays the real gap.
|
||||||
|
- alert: WebinarCheckerNoSuccessfulCheck
|
||||||
|
expr: |
|
||||||
|
((time() - webinar_check_last_success_timestamp_seconds) > 300)
|
||||||
|
and (webinar_check_last_success_timestamp_seconds > 0)
|
||||||
|
and (webinar_check_last_run_timestamp_seconds > 0)
|
||||||
|
for: 2m
|
||||||
|
labels:
|
||||||
|
severity: critical
|
||||||
|
annotations:
|
||||||
|
summary: "Webinar checker has no successful check for 5m"
|
||||||
|
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
|
||||||
|
|
||||||
|
# Checks are running but none has ever succeeded since pod start.
|
||||||
|
# Split out from the rule above so a zeroed gauge never feeds
|
||||||
|
# humanizeDuration.
|
||||||
|
- alert: WebinarCheckerNeverSucceeded
|
||||||
|
expr: |
|
||||||
|
(webinar_check_last_success_timestamp_seconds == 0)
|
||||||
|
and (webinar_check_last_run_timestamp_seconds > 0)
|
||||||
|
for: 10m
|
||||||
|
labels:
|
||||||
|
severity: critical
|
||||||
|
annotations:
|
||||||
|
summary: "Webinar checker has never completed a successful check"
|
||||||
|
description: 'edu-master/webinar-checker: checks have been running for 10m but not one has ever succeeded since the pod started, so every check is failing. Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
|
||||||
|
|
||||||
|
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
|
||||||
|
- alert: WebinarCheckerConsecutiveFailures
|
||||||
|
expr: |
|
||||||
|
webinar_check_consecutive_failures >= 3
|
||||||
|
for: 5m
|
||||||
|
labels:
|
||||||
|
severity: critical
|
||||||
|
annotations:
|
||||||
|
summary: "Webinar checker failing consecutively"
|
||||||
|
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / http error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
|
||||||
|
|
||||||
|
- alert: WebinarCheckerNeverStarted
|
||||||
|
expr: |
|
||||||
|
(time() - edu_process_start > 120)
|
||||||
|
and (webinar_check_last_run_timestamp_seconds == 0)
|
||||||
|
for: 2m
|
||||||
|
labels:
|
||||||
|
severity: critical
|
||||||
|
annotations:
|
||||||
|
summary: "Webinar checker job has not started"
|
||||||
|
description: "The process exposes metrics but its webinar job has never started."
|
||||||
|
|
||||||
|
- alert: WebinarDeliveryPending
|
||||||
|
expr: edu_delivery_pending > 0
|
||||||
|
for: 5m
|
||||||
|
labels:
|
||||||
|
severity: warning
|
||||||
|
annotations:
|
||||||
|
summary: "Webinar notifications await delivery"
|
||||||
|
description: "Telegram delivery has pending recipients. Check delivery failures and retry status."
|
||||||
|
|
||||||
|
- alert: EduRedisUnavailable
|
||||||
|
expr: edu_redis_connected == 0
|
||||||
|
for: 2m
|
||||||
|
labels:
|
||||||
|
severity: critical
|
||||||
|
annotations:
|
||||||
|
summary: "EDU checker cannot reach Redis"
|
||||||
|
description: "Redis health checks are failing; checker commands and delivery may be unavailable."
|
||||||
|
|
||||||
|
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
|
||||||
|
- alert: WebinarCheckerScrapeDown
|
||||||
|
expr: |
|
||||||
|
absent(webinar_check_last_run_timestamp_seconds) == 1
|
||||||
|
for: 10m
|
||||||
|
labels:
|
||||||
|
severity: critical
|
||||||
|
annotations:
|
||||||
|
summary: "Webinar checker metrics missing"
|
||||||
|
description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
|
||||||
|
|
||||||
|
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
|
||||||
|
- alert: EduPhpsessidMissing
|
||||||
|
expr: |
|
||||||
|
edu_phpsessid_present == 0
|
||||||
|
for: 10m
|
||||||
|
labels:
|
||||||
|
severity: critical
|
||||||
|
annotations:
|
||||||
|
summary: "EDU_PHPSESSID missing"
|
||||||
|
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
|
||||||
|
|
||||||
|
# Hard deps: checker deployment unavailable.
|
||||||
|
- alert: WebinarCheckerDeploymentDown
|
||||||
|
expr: |
|
||||||
|
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
|
||||||
|
for: 10m
|
||||||
|
labels:
|
||||||
|
severity: critical
|
||||||
|
annotations:
|
||||||
|
summary: "Webinar checker deployment unavailable"
|
||||||
|
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
apiVersion: v1
|
||||||
|
kind: Namespace
|
||||||
|
metadata:
|
||||||
|
name: edu-master
|
||||||
@@ -0,0 +1,69 @@
|
|||||||
|
apiVersion: apps/v1
|
||||||
|
kind: Deployment
|
||||||
|
metadata:
|
||||||
|
name: playwright-service
|
||||||
|
namespace: edu-master
|
||||||
|
labels:
|
||||||
|
app: edu-master-playwright
|
||||||
|
spec:
|
||||||
|
replicas: 1
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app: edu-master-playwright
|
||||||
|
strategy:
|
||||||
|
type: Recreate
|
||||||
|
template:
|
||||||
|
metadata:
|
||||||
|
labels:
|
||||||
|
app: edu-master-playwright
|
||||||
|
spec:
|
||||||
|
containers:
|
||||||
|
- name: playwright
|
||||||
|
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker
|
||||||
|
image: mcr.microsoft.com/playwright:v1.56.0-jammy
|
||||||
|
imagePullPolicy: IfNotPresent
|
||||||
|
# p95 412M, max 478M over 7 days, no limit before. Request is set at p95
|
||||||
|
# so the pod is not an eviction candidate; the limit stays above 2x the
|
||||||
|
# request because browser page lifetimes are unpredictable.
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: "200m"
|
||||||
|
memory: "416Mi"
|
||||||
|
limits:
|
||||||
|
memory: "1Gi"
|
||||||
|
command:
|
||||||
|
- npx
|
||||||
|
- -y
|
||||||
|
- playwright@1.56.0
|
||||||
|
- run-server
|
||||||
|
- --port
|
||||||
|
- "3000"
|
||||||
|
- --path
|
||||||
|
- /ws
|
||||||
|
ports:
|
||||||
|
- containerPort: 3000
|
||||||
|
readinessProbe:
|
||||||
|
tcpSocket:
|
||||||
|
port: 3000
|
||||||
|
initialDelaySeconds: 5
|
||||||
|
periodSeconds: 10
|
||||||
|
timeoutSeconds: 3
|
||||||
|
livenessProbe:
|
||||||
|
tcpSocket:
|
||||||
|
port: 3000
|
||||||
|
initialDelaySeconds: 15
|
||||||
|
periodSeconds: 20
|
||||||
|
timeoutSeconds: 3
|
||||||
|
---
|
||||||
|
apiVersion: v1
|
||||||
|
kind: Service
|
||||||
|
metadata:
|
||||||
|
name: playwright-service
|
||||||
|
namespace: edu-master
|
||||||
|
spec:
|
||||||
|
selector:
|
||||||
|
app: edu-master-playwright
|
||||||
|
ports:
|
||||||
|
- name: ws
|
||||||
|
port: 3000
|
||||||
|
targetPort: 3000
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
apiVersion: networking.k8s.io/v1
|
||||||
|
kind: NetworkPolicy
|
||||||
|
metadata:
|
||||||
|
name: redis-clients-only
|
||||||
|
namespace: edu-master
|
||||||
|
spec:
|
||||||
|
podSelector:
|
||||||
|
matchLabels:
|
||||||
|
app: edu-master-redis
|
||||||
|
policyTypes:
|
||||||
|
- Ingress
|
||||||
|
ingress:
|
||||||
|
- from:
|
||||||
|
- podSelector:
|
||||||
|
matchLabels:
|
||||||
|
app: edu-master-session-keeper
|
||||||
|
- podSelector:
|
||||||
|
matchLabels:
|
||||||
|
app: edu-master-webinar-checker
|
||||||
|
ports:
|
||||||
|
- protocol: TCP
|
||||||
|
port: 6379
|
||||||
@@ -0,0 +1,96 @@
|
|||||||
|
apiVersion: apps/v1
|
||||||
|
kind: StatefulSet
|
||||||
|
metadata:
|
||||||
|
name: redis
|
||||||
|
namespace: edu-master
|
||||||
|
labels:
|
||||||
|
app: edu-master-redis
|
||||||
|
spec:
|
||||||
|
serviceName: redis
|
||||||
|
replicas: 1
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app: edu-master-redis
|
||||||
|
template:
|
||||||
|
metadata:
|
||||||
|
labels:
|
||||||
|
app: edu-master-redis
|
||||||
|
spec:
|
||||||
|
containers:
|
||||||
|
- name: redis
|
||||||
|
image: redis:8.10.2-alpine
|
||||||
|
imagePullPolicy: IfNotPresent
|
||||||
|
env:
|
||||||
|
- name: REDIS_PASSWORD
|
||||||
|
valueFrom:
|
||||||
|
secretKeyRef:
|
||||||
|
name: edu-master-secrets
|
||||||
|
key: REDIS_PASSWORD
|
||||||
|
- name: REDISCLI_AUTH
|
||||||
|
valueFrom:
|
||||||
|
secretKeyRef:
|
||||||
|
name: edu-master-secrets
|
||||||
|
key: REDIS_PASSWORD
|
||||||
|
command:
|
||||||
|
- /bin/sh
|
||||||
|
- -ec
|
||||||
|
- |
|
||||||
|
case "$REDIS_PASSWORD" in *[!0-9a-fA-F]*|'') echo 'REDIS_PASSWORD must be 64 hex characters' >&2; exit 1;; esac
|
||||||
|
[ "${#REDIS_PASSWORD}" -eq 64 ] || { echo 'REDIS_PASSWORD must be 64 hex characters' >&2; exit 1; }
|
||||||
|
umask 077
|
||||||
|
printf 'requirepass "%s"\n' "$REDIS_PASSWORD" > /tmp/redis-auth.conf
|
||||||
|
chown redis:redis /tmp/redis-auth.conf
|
||||||
|
exec docker-entrypoint.sh redis-server /tmp/redis-auth.conf
|
||||||
|
ports:
|
||||||
|
- containerPort: 6379
|
||||||
|
volumeMounts:
|
||||||
|
- name: redis-data
|
||||||
|
mountPath: /data
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: 25m
|
||||||
|
memory: 32Mi
|
||||||
|
limits:
|
||||||
|
cpu: 250m
|
||||||
|
memory: 128Mi
|
||||||
|
readinessProbe:
|
||||||
|
exec:
|
||||||
|
command: ["redis-cli", "ping"]
|
||||||
|
initialDelaySeconds: 5
|
||||||
|
periodSeconds: 5
|
||||||
|
timeoutSeconds: 3
|
||||||
|
livenessProbe:
|
||||||
|
exec:
|
||||||
|
command: ["redis-cli", "ping"]
|
||||||
|
initialDelaySeconds: 10
|
||||||
|
periodSeconds: 10
|
||||||
|
timeoutSeconds: 3
|
||||||
|
volumes:
|
||||||
|
- name: redis-data
|
||||||
|
persistentVolumeClaim:
|
||||||
|
claimName: redis-data-pvc
|
||||||
|
---
|
||||||
|
apiVersion: v1
|
||||||
|
kind: PersistentVolumeClaim
|
||||||
|
metadata:
|
||||||
|
name: redis-data-pvc
|
||||||
|
namespace: edu-master
|
||||||
|
spec:
|
||||||
|
accessModes:
|
||||||
|
- ReadWriteOnce
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
storage: 1Gi
|
||||||
|
---
|
||||||
|
apiVersion: v1
|
||||||
|
kind: Service
|
||||||
|
metadata:
|
||||||
|
name: redis
|
||||||
|
namespace: edu-master
|
||||||
|
spec:
|
||||||
|
selector:
|
||||||
|
app: edu-master-redis
|
||||||
|
ports:
|
||||||
|
- name: redis
|
||||||
|
port: 6379
|
||||||
|
targetPort: 6379
|
||||||
@@ -0,0 +1,50 @@
|
|||||||
|
# One-time Job to migrate redis state from docker compose to k8s (maintenance window).
|
||||||
|
# The .example file is not applied by the deploy pipeline (mask *.example.yaml).
|
||||||
|
#
|
||||||
|
# Runbook:
|
||||||
|
# 1. docker compose -f <repo>/edu_master/compose.yaml stop # SIGTERM -> redis will flush dump.rdb
|
||||||
|
# 2. docker run --rm -v edu_master_redis-data:/data \
|
||||||
|
# -v /tmp/edu-master-backup:/backup \
|
||||||
|
# redis:alpine sh -c "cp /data/dump.rdb /backup/ && ls -la /backup"
|
||||||
|
# 3. kubectl apply -f edu_master/k8s/namespace.yaml
|
||||||
|
# 4. kubectl apply -f <only the PVC from redis.yaml> # seed must come BEFORE redis pod starts
|
||||||
|
# 5. kubectl apply -f edu_master/k8s/restore-seed-job.yaml.example
|
||||||
|
# kubectl wait --for=condition=complete job/redis-restore-seed -n edu-master --timeout=120s
|
||||||
|
# 6. kubectl delete job redis-restore-seed -n edu-master
|
||||||
|
# 7. kubectl apply -f edu_master/k8s/ -R # apply remaining manifests
|
||||||
|
apiVersion: batch/v1
|
||||||
|
kind: Job
|
||||||
|
metadata:
|
||||||
|
name: redis-restore-seed
|
||||||
|
namespace: edu-master
|
||||||
|
spec:
|
||||||
|
backoffLimit: 2
|
||||||
|
ttlSecondsAfterFinished: 3600
|
||||||
|
template:
|
||||||
|
spec:
|
||||||
|
restartPolicy: Never
|
||||||
|
containers:
|
||||||
|
- name: seed
|
||||||
|
image: redis:alpine
|
||||||
|
command:
|
||||||
|
- /bin/sh
|
||||||
|
- -ec
|
||||||
|
- |
|
||||||
|
ls -la /backup
|
||||||
|
cp /backup/dump.rdb /data/dump.rdb
|
||||||
|
chmod 644 /data/dump.rdb
|
||||||
|
ls -la /data
|
||||||
|
volumeMounts:
|
||||||
|
- name: redis-data
|
||||||
|
mountPath: /data
|
||||||
|
- name: backup
|
||||||
|
mountPath: /backup
|
||||||
|
readOnly: true
|
||||||
|
volumes:
|
||||||
|
- name: redis-data
|
||||||
|
persistentVolumeClaim:
|
||||||
|
claimName: redis-data-pvc
|
||||||
|
- name: backup
|
||||||
|
hostPath:
|
||||||
|
path: /tmp/edu-master-backup
|
||||||
|
type: DirectoryOrCreate
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
apiVersion: v1
|
||||||
|
kind: Secret
|
||||||
|
metadata:
|
||||||
|
name: edu-master-secrets
|
||||||
|
namespace: edu-master
|
||||||
|
type: Opaque
|
||||||
|
stringData:
|
||||||
|
# Session keeper credentials
|
||||||
|
KEEPER_LOGIN: ""
|
||||||
|
KEEPER_PASSWORD: ""
|
||||||
|
KEEPER_INTERVAL: "10"
|
||||||
|
# EDU links
|
||||||
|
EDU_URL_BASE: "https://edu.edu.vn.ua"
|
||||||
|
EDU_URL_LOGIN: "/user/login"
|
||||||
|
EDU_URL_COURSES: "/course/userlist"
|
||||||
|
EDU_URL_WEBINAR: "/webinar/useractive"
|
||||||
|
# Playwright
|
||||||
|
USER_AGENT: ""
|
||||||
|
PLAYWRIGHT_WS: "ws://playwright-service:3000/ws"
|
||||||
|
# Webinar-checker
|
||||||
|
WEBINAR_TELEGRAM_TOKEN: ""
|
||||||
|
WEBINAR_ADMIN_ID: ""
|
||||||
|
WEBINAR_CHECK_INTERVAL: "60"
|
||||||
|
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
|
||||||
|
METRICS_PORT: "8000"
|
||||||
|
# Database
|
||||||
|
REDIS_HOST: "redis"
|
||||||
|
REDIS_PORT: "6379"
|
||||||
|
TZ: "Europe/Kyiv"
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
apiVersion: v1
|
||||||
|
kind: Service
|
||||||
|
metadata:
|
||||||
|
name: webinar-checker
|
||||||
|
namespace: edu-master
|
||||||
|
labels:
|
||||||
|
app: edu-master-webinar-checker
|
||||||
|
spec:
|
||||||
|
selector:
|
||||||
|
app: edu-master-webinar-checker
|
||||||
|
ports:
|
||||||
|
- name: metrics
|
||||||
|
port: 8000
|
||||||
|
targetPort: metrics
|
||||||
|
protocol: TCP
|
||||||
@@ -1,14 +1,14 @@
|
|||||||
apiVersion: monitoring.coreos.com/v1
|
apiVersion: monitoring.coreos.com/v1
|
||||||
kind: ServiceMonitor
|
kind: ServiceMonitor
|
||||||
metadata:
|
metadata:
|
||||||
name: netbird-server
|
name: webinar-checker
|
||||||
namespace: netbird
|
namespace: edu-master
|
||||||
labels:
|
labels:
|
||||||
release: prometheus-stack
|
release: prometheus-stack
|
||||||
spec:
|
spec:
|
||||||
selector:
|
selector:
|
||||||
matchLabels:
|
matchLabels:
|
||||||
app: netbird-server
|
app: edu-master-webinar-checker
|
||||||
endpoints:
|
endpoints:
|
||||||
- port: metrics
|
- port: metrics
|
||||||
path: /metrics
|
path: /metrics
|
||||||
@@ -0,0 +1,69 @@
|
|||||||
|
apiVersion: apps/v1
|
||||||
|
kind: Deployment
|
||||||
|
metadata:
|
||||||
|
annotations:
|
||||||
|
reloader.stakater.com/auto: "true"
|
||||||
|
name: session-keeper
|
||||||
|
namespace: edu-master
|
||||||
|
labels:
|
||||||
|
app: edu-master-session-keeper
|
||||||
|
spec:
|
||||||
|
replicas: 1
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app: edu-master-session-keeper
|
||||||
|
strategy:
|
||||||
|
type: Recreate
|
||||||
|
template:
|
||||||
|
metadata:
|
||||||
|
annotations:
|
||||||
|
edu.forust.xyz/source-commit: "90829d6c8080b9928f9da23587678e640939e10a"
|
||||||
|
labels:
|
||||||
|
app: edu-master-session-keeper
|
||||||
|
spec:
|
||||||
|
initContainers:
|
||||||
|
- name: wait-redis
|
||||||
|
image: redis:8.10.2-alpine
|
||||||
|
env:
|
||||||
|
- name: REDISCLI_AUTH
|
||||||
|
valueFrom:
|
||||||
|
secretKeyRef:
|
||||||
|
name: edu-master-secrets
|
||||||
|
key: REDIS_PASSWORD
|
||||||
|
command:
|
||||||
|
- /bin/sh
|
||||||
|
- -ec
|
||||||
|
- |
|
||||||
|
i=0
|
||||||
|
until redis-cli -h redis ping | grep -q PONG; do
|
||||||
|
i=$((i+1))
|
||||||
|
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
|
||||||
|
sleep 2
|
||||||
|
done
|
||||||
|
echo "redis is ready"
|
||||||
|
containers:
|
||||||
|
- name: session-keeper
|
||||||
|
image: gcr.forust.xyz/forust/session-keeper@sha256:49285e87cc5bc4cf4ffe190813d87927916c2df8a206daac0aeb7d227c636450
|
||||||
|
envFrom:
|
||||||
|
- secretRef:
|
||||||
|
name: edu-master-secrets
|
||||||
|
env:
|
||||||
|
- name: REDISCLI_AUTH
|
||||||
|
valueFrom:
|
||||||
|
secretKeyRef:
|
||||||
|
name: edu-master-secrets
|
||||||
|
key: REDIS_PASSWORD
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: 25m
|
||||||
|
memory: 32Mi
|
||||||
|
limits:
|
||||||
|
cpu: 250m
|
||||||
|
memory: 128Mi
|
||||||
|
readinessProbe:
|
||||||
|
exec:
|
||||||
|
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
|
||||||
|
initialDelaySeconds: 15
|
||||||
|
periodSeconds: 30
|
||||||
|
timeoutSeconds: 5
|
||||||
|
failureThreshold: 10
|
||||||
@@ -0,0 +1,87 @@
|
|||||||
|
apiVersion: apps/v1
|
||||||
|
kind: Deployment
|
||||||
|
metadata:
|
||||||
|
annotations:
|
||||||
|
reloader.stakater.com/auto: "true"
|
||||||
|
name: webinar-checker
|
||||||
|
namespace: edu-master
|
||||||
|
labels:
|
||||||
|
app: edu-master-webinar-checker
|
||||||
|
spec:
|
||||||
|
replicas: 1
|
||||||
|
selector:
|
||||||
|
matchLabels:
|
||||||
|
app: edu-master-webinar-checker
|
||||||
|
strategy:
|
||||||
|
type: Recreate
|
||||||
|
template:
|
||||||
|
metadata:
|
||||||
|
annotations:
|
||||||
|
edu.forust.xyz/source-commit: "90829d6c8080b9928f9da23587678e640939e10a"
|
||||||
|
labels:
|
||||||
|
app: edu-master-webinar-checker
|
||||||
|
spec:
|
||||||
|
# Enforces dependency order like compose depends_on:
|
||||||
|
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID)
|
||||||
|
initContainers:
|
||||||
|
- name: wait-deps
|
||||||
|
image: redis:8.10.2-alpine
|
||||||
|
env:
|
||||||
|
- name: REDISCLI_AUTH
|
||||||
|
valueFrom:
|
||||||
|
secretKeyRef:
|
||||||
|
name: edu-master-secrets
|
||||||
|
key: REDIS_PASSWORD
|
||||||
|
command:
|
||||||
|
- /bin/sh
|
||||||
|
- -ec
|
||||||
|
- |
|
||||||
|
i=0
|
||||||
|
until redis-cli -h redis ping | grep -q PONG; do
|
||||||
|
i=$((i+1))
|
||||||
|
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
|
||||||
|
sleep 2
|
||||||
|
done
|
||||||
|
echo "redis ok"
|
||||||
|
until [ "$(redis-cli -h redis EXISTS EDU_PHPSESSID)" = "1" ]; do
|
||||||
|
i=$((i+1))
|
||||||
|
[ "$i" -ge 300 ] && echo "TIMEOUT: no PHPSESSID (session-keeper down?)" && exit 1
|
||||||
|
sleep 2
|
||||||
|
done
|
||||||
|
echo "PHPSESSID ok"
|
||||||
|
containers:
|
||||||
|
- name: webinar-checker
|
||||||
|
image: gcr.forust.xyz/forust/webinar-checker@sha256:66c146f7b43cb9f0dc31ba9aa36d217e01df42ddafba5971b79c12ec215b2c01
|
||||||
|
ports:
|
||||||
|
- name: metrics
|
||||||
|
containerPort: 8000
|
||||||
|
protocol: TCP
|
||||||
|
readinessProbe:
|
||||||
|
httpGet:
|
||||||
|
path: /health
|
||||||
|
port: metrics
|
||||||
|
periodSeconds: 10
|
||||||
|
timeoutSeconds: 3
|
||||||
|
failureThreshold: 12
|
||||||
|
initialDelaySeconds: 10
|
||||||
|
livenessProbe:
|
||||||
|
httpGet:
|
||||||
|
path: /live
|
||||||
|
port: metrics
|
||||||
|
initialDelaySeconds: 60
|
||||||
|
periodSeconds: 15
|
||||||
|
timeoutSeconds: 3
|
||||||
|
failureThreshold: 4
|
||||||
|
envFrom:
|
||||||
|
- secretRef:
|
||||||
|
name: edu-master-secrets
|
||||||
|
env:
|
||||||
|
- name: TZ
|
||||||
|
value: "Europe/Kyiv"
|
||||||
|
resources:
|
||||||
|
requests:
|
||||||
|
cpu: "50m"
|
||||||
|
memory: "192Mi"
|
||||||
|
limits:
|
||||||
|
cpu: "600m"
|
||||||
|
memory: "384Mi"
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
FROM python:3.14-slim
|
||||||
|
|
||||||
|
WORKDIR /app
|
||||||
|
|
||||||
|
# Install system dependencies
|
||||||
|
RUN apt-get update && apt-get install -y --no-install-recommends redis-tools && rm -rf /var/lib/apt/lists/*
|
||||||
|
|
||||||
|
# Install dependencies
|
||||||
|
RUN pip install --no-cache-dir requests==2.32.3 redis==5.2.1
|
||||||
|
|
||||||
|
# Copy application code
|
||||||
|
COPY . .
|
||||||
|
|
||||||
|
# Run the bot
|
||||||
|
CMD ["python", "bot.py"]
|
||||||
@@ -0,0 +1,132 @@
|
|||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import time
|
||||||
|
from datetime import datetime
|
||||||
|
|
||||||
|
import redis
|
||||||
|
import requests
|
||||||
|
|
||||||
|
# Configure logging
|
||||||
|
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')
|
||||||
|
logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
|
||||||
|
# Load configuration (adapted to .env keys)
|
||||||
|
def _env(key, default=None):
|
||||||
|
v = os.getenv(key, default)
|
||||||
|
if isinstance(v, str) and len(v) >= 2 and ((v[0] == '"' and v[-1] == '"') or (v[0] == "'" and v[-1] == "'")):
|
||||||
|
return v[1:-1]
|
||||||
|
return v
|
||||||
|
|
||||||
|
|
||||||
|
LOGIN = _env('KEEPER_LOGIN')
|
||||||
|
PASSWORD = _env('KEEPER_PASSWORD')
|
||||||
|
|
||||||
|
EDU_BASE = _env('EDU_URL_BASE', 'https://edu.edu.vn.ua')
|
||||||
|
EDU_LOGIN_PATH = _env('EDU_URL_LOGIN', '/user/login')
|
||||||
|
EDU_COURSES_PATH = _env('EDU_URL_COURSES', '/course/userlist')
|
||||||
|
URL_LOGIN = f'{EDU_BASE.rstrip("/")}/{EDU_LOGIN_PATH.lstrip("/")}'
|
||||||
|
URL_VERIFY = f'{EDU_BASE.rstrip("/")}/{EDU_COURSES_PATH.lstrip("/")}'
|
||||||
|
|
||||||
|
INTERVAL = int(_env('KEEPER_INTERVAL', 10))
|
||||||
|
USER_AGENT = _env(
|
||||||
|
'USER_AGENT',
|
||||||
|
'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36',
|
||||||
|
)
|
||||||
|
REDIS_HOST = _env('REDIS_HOST', 'redis')
|
||||||
|
REDIS_PORT = int(_env('REDIS_PORT', 6379))
|
||||||
|
|
||||||
|
SUCCESS_FILE = '/tmp/last_success' # noqa: S108
|
||||||
|
|
||||||
|
|
||||||
|
def touch_success_file():
|
||||||
|
"""Updates the timestamp of the success file for healthchecks."""
|
||||||
|
try:
|
||||||
|
with open(SUCCESS_FILE, 'w') as f:
|
||||||
|
f.write(str(datetime.now().timestamp()))
|
||||||
|
except Exception as e:
|
||||||
|
logger.error(f'Failed to touch success file: {e}')
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
logger.info('Starting Session Keeper Bot')
|
||||||
|
|
||||||
|
# Connect to Redis
|
||||||
|
try:
|
||||||
|
redis_client = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
|
||||||
|
redis_client.ping()
|
||||||
|
logger.info(f'Connected to Redis at {REDIS_HOST}:{REDIS_PORT}')
|
||||||
|
except Exception as e:
|
||||||
|
logger.error(f'Failed to connect to Redis: {e}')
|
||||||
|
return
|
||||||
|
|
||||||
|
session = requests.Session()
|
||||||
|
|
||||||
|
# Set headers
|
||||||
|
headers = {
|
||||||
|
'User-Agent': USER_AGENT,
|
||||||
|
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
|
||||||
|
'Accept-Language': 'en-US,en;q=0.9',
|
||||||
|
'Cache-Control': 'max-age=0',
|
||||||
|
'Upgrade-Insecure-Requests': '1',
|
||||||
|
'Sec-Fetch-Site': 'same-origin',
|
||||||
|
'Sec-Fetch-Mode': 'navigate',
|
||||||
|
'Sec-Fetch-User': '?1',
|
||||||
|
'Sec-Fetch-Dest': 'document',
|
||||||
|
'Sec-Ch-Ua': '"Not_A Brand";v="99", "Chromium";v="142"',
|
||||||
|
'Sec-Ch-Ua-Mobile': '?0',
|
||||||
|
'Sec-Ch-Ua-Platform': '"Linux"',
|
||||||
|
'Accept-Encoding': 'gzip, deflate, br',
|
||||||
|
'Priority': 'u=0, i',
|
||||||
|
}
|
||||||
|
session.headers.update(headers)
|
||||||
|
|
||||||
|
while True:
|
||||||
|
try:
|
||||||
|
logger.info('Attempting login...')
|
||||||
|
|
||||||
|
# Login payload
|
||||||
|
payload = {'login': LOGIN, 'password': PASSWORD}
|
||||||
|
|
||||||
|
# Perform Login
|
||||||
|
# Note: The user request shows a POST to /user/login with form data
|
||||||
|
# We need to make sure we handle the PHPSESSID correctly.
|
||||||
|
# If we already have a PHPSESSID, requests will send it.
|
||||||
|
|
||||||
|
login_response = session.post(URL_LOGIN, data=payload, allow_redirects=True)
|
||||||
|
|
||||||
|
logger.info(f'Login Response Status: {login_response.status_code}')
|
||||||
|
logger.info(f'Cookies after login: {session.cookies.get_dict()}')
|
||||||
|
|
||||||
|
# Verify Session
|
||||||
|
logger.info('Verifying session...')
|
||||||
|
verify_response = session.get(URL_VERIFY, allow_redirects=False)
|
||||||
|
|
||||||
|
logger.info(f'Verify Response Status: {verify_response.status_code}')
|
||||||
|
|
||||||
|
if verify_response.status_code == 200:
|
||||||
|
logger.info('Session verification SUCCESS (200 OK).')
|
||||||
|
touch_success_file()
|
||||||
|
|
||||||
|
# Save PHPSESSID to Redis
|
||||||
|
phpsessid = session.cookies.get('PHPSESSID')
|
||||||
|
if phpsessid:
|
||||||
|
try:
|
||||||
|
redis_client.set('EDU_PHPSESSID', phpsessid)
|
||||||
|
logger.info(f'Saved PHPSESSID to Redis: {phpsessid}')
|
||||||
|
except Exception as e:
|
||||||
|
logger.error(f'Failed to save PHPSESSID to Redis: {e}')
|
||||||
|
elif verify_response.status_code == 302:
|
||||||
|
logger.warning('Session verification FAILED (302 Redirect). Session might be invalid.')
|
||||||
|
else:
|
||||||
|
logger.warning(f'Session verification returned unexpected status: {verify_response.status_code}')
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
logger.error(f'An error occurred: {e}')
|
||||||
|
|
||||||
|
logger.info(f'Sleeping for {INTERVAL} minutes...')
|
||||||
|
time.sleep(INTERVAL * 60)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
FROM python:3.14-slim
|
||||||
|
|
||||||
|
WORKDIR /app
|
||||||
|
|
||||||
|
# renovate: datasource=pypi depName=playwright versioning=pep440
|
||||||
|
ARG PLAYWRIGHT_VERSION=1.56.0
|
||||||
|
|
||||||
|
# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers
|
||||||
|
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
|
||||||
|
|
||||||
|
COPY checker.py .
|
||||||
|
|
||||||
|
CMD ["python", "checker.py"]
|
||||||
File diff suppressed because it is too large.
Load diff
@@ -1,21 +0,0 @@
|
|||||||
# Error pages
|
|
||||||
|
|
||||||
Static HTTP error pages served by an Nginx image built in CI.
|
|
||||||
|
|
||||||
Edit the HTML in `html/`; the Dockerfile copies it into the image.
|
|
||||||
Kubernetes exposes `error-pages-service` in `error-pages` for Traefik's error
|
|
||||||
middleware. Keep the middleware's namespace and port aligned with that Service.
|
|
||||||
|
|
||||||
For a local build, run `docker build -t homelab-error-pages .` from this directory.
|
|
||||||
Compose references the private registry image rather than a build context.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n error-pages
|
|
||||||
kubectl get events -n error-pages --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,25 +0,0 @@
|
|||||||
# Gitea
|
|
||||||
|
|
||||||
Git hosting with HTTP and a separate SSH route.
|
|
||||||
|
|
||||||
Kubernetes uses the shared PostgreSQL service and `gitea-pvc` for repositories
|
|
||||||
and application data. Match the Gitea database password with the shared database
|
|
||||||
Secret. SSH is routed through Traefik's TCP entrypoint on 2221.
|
|
||||||
|
|
||||||
Compose uses a separate PostgreSQL 14 database, bind mounts `gitea-data/` and
|
|
||||||
`gitea-db/`, and publishes host port 2221. It is an alternative deployment with
|
|
||||||
its own database, not a second frontend for the Kubernetes instance.
|
|
||||||
|
|
||||||
Back up repositories, application configuration, and a consistent database dump
|
|
||||||
together. Gitea Actions definitions for this repository live in `../.gitea/`.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n gitea
|
|
||||||
kubectl get events -n gitea --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -20,8 +20,6 @@ data:
|
|||||||
|
|
||||||
GITEA__mailer__ENABLED: "false"
|
GITEA__mailer__ENABLED: "false"
|
||||||
|
|
||||||
GITEA__metrics__ENABLED: "true"
|
|
||||||
|
|
||||||
# No code/issue search needed: bleve reindexes the whole issue index on
|
# No code/issue search needed: bleve reindexes the whole issue index on
|
||||||
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
|
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
|
||||||
# the rotational disk for an hour. "db" serves issue search from postgres.
|
# the rotational disk for an hour. "db" serves issue search from postgres.
|
||||||
|
|||||||
@@ -3,8 +3,6 @@ kind: Service
|
|||||||
metadata:
|
metadata:
|
||||||
name: gitea-service
|
name: gitea-service
|
||||||
namespace: gitea
|
namespace: gitea
|
||||||
labels:
|
|
||||||
app: gitea
|
|
||||||
spec:
|
spec:
|
||||||
selector:
|
selector:
|
||||||
app: gitea
|
app: gitea
|
||||||
|
|||||||
@@ -7,8 +7,7 @@ spec:
|
|||||||
entryPoints:
|
entryPoints:
|
||||||
- websecure
|
- websecure
|
||||||
routes:
|
routes:
|
||||||
# Metrics are scraped directly through the cluster Service.
|
- match: Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)
|
||||||
- match: (Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)) && !PathPrefix(`/metrics`)
|
|
||||||
kind: Rule
|
kind: Rule
|
||||||
services:
|
services:
|
||||||
- name: gitea-service
|
- name: gitea-service
|
||||||
|
|||||||
@@ -1,16 +0,0 @@
|
|||||||
apiVersion: monitoring.coreos.com/v1
|
|
||||||
kind: ServiceMonitor
|
|
||||||
metadata:
|
|
||||||
name: gitea
|
|
||||||
namespace: gitea
|
|
||||||
labels:
|
|
||||||
release: prometheus-stack
|
|
||||||
spec:
|
|
||||||
selector:
|
|
||||||
matchLabels:
|
|
||||||
app: gitea
|
|
||||||
endpoints:
|
|
||||||
- port: http
|
|
||||||
path: /metrics
|
|
||||||
interval: 30s
|
|
||||||
scrapeTimeout: 10s
|
|
||||||
@@ -1,26 +0,0 @@
|
|||||||
# Glance
|
|
||||||
|
|
||||||
Dashboard pages for links, service checks, and Docker containers.
|
|
||||||
|
|
||||||
Compose mounts `config/` and `assets/`. The Kubernetes equivalents are embedded
|
|
||||||
in `k8s/glance-config.yaml`: `glance-config` holds pages and `glance-assets` holds
|
|
||||||
`user.css`. Update both copies when changing shared content.
|
|
||||||
|
|
||||||
Kubernetes serves the dashboard under `/glance`. Its pod also mounts the node's
|
|
||||||
Docker socket. It references `glance-secrets` for `ADGUARD_PASSWORD`, but there is
|
|
||||||
no tracked Secret example; create that Secret in `glance` before starting it.
|
|
||||||
Compose expects a local `.env` with the same password.
|
|
||||||
|
|
||||||
The pod mounts `user.css` from `glance-assets`, which is the ConfigMap that
|
|
||||||
contains that key.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n glance
|
|
||||||
kubectl get events -n glance --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,18 +0,0 @@
|
|||||||
# Headscale
|
|
||||||
|
|
||||||
Headscale, Headplane, and a separate web administration UI on Docker.
|
|
||||||
|
|
||||||
Kubernetes only provides routes to the Docker host. Update the addresses in
|
|
||||||
`k8s/routing/external-service.yaml` if the host moves.
|
|
||||||
|
|
||||||
Copy `config/headscale.yaml.example`, `config/headplane.yaml.example`, and
|
|
||||||
`config/policy.json.example` to their names without `.example`. Set the public
|
|
||||||
server URL, DNS settings, Headplane cookie secret, and Headscale public URL.
|
|
||||||
The example URLs are placeholders.
|
|
||||||
|
|
||||||
Compose publishes Headscale on 18080, its metrics port on 19090, Headplane on
|
|
||||||
13000, and the other UI on 10080. The data volumes store the Headscale database,
|
|
||||||
keys, and Headplane state. The embedded DERP configuration needs reachable
|
|
||||||
addresses; Compose does not publish its UDP 3478 listener.
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -7,7 +7,7 @@
|
|||||||
# # Dev server_url
|
# # Dev server_url
|
||||||
# server_url: https://hs.dev_internal_domain.internal
|
# server_url: https://hs.dev_internal_domain.internal
|
||||||
listen_addr: 0.0.0.0:8080
|
listen_addr: 0.0.0.0:8080
|
||||||
metrics_listen_addr: 0.0.0.0:9090
|
metrics_listen_addr: 127.0.0.1:9090
|
||||||
grpc_listen_addr: 127.0.0.1:50443
|
grpc_listen_addr: 127.0.0.1:50443
|
||||||
grpc_allow_insecure: false
|
grpc_allow_insecure: false
|
||||||
noise:
|
noise:
|
||||||
|
|||||||
@@ -3,8 +3,6 @@ kind: Service
|
|||||||
metadata:
|
metadata:
|
||||||
name: headscale-server-external
|
name: headscale-server-external
|
||||||
namespace: headscale
|
namespace: headscale
|
||||||
labels:
|
|
||||||
app: headscale
|
|
||||||
spec:
|
spec:
|
||||||
ports:
|
ports:
|
||||||
- port: 8080
|
- port: 8080
|
||||||
|
|||||||
@@ -1,18 +0,0 @@
|
|||||||
apiVersion: operator.victoriametrics.com/v1beta1
|
|
||||||
kind: VMServiceScrape
|
|
||||||
metadata:
|
|
||||||
name: headscale
|
|
||||||
namespace: headscale
|
|
||||||
labels:
|
|
||||||
release: prometheus-stack
|
|
||||||
spec:
|
|
||||||
# The external Service has a manually managed EndpointSlice, not Endpoints.
|
|
||||||
discoveryRole: endpointslice
|
|
||||||
selector:
|
|
||||||
matchLabels:
|
|
||||||
app: headscale
|
|
||||||
endpoints:
|
|
||||||
- port: metrics
|
|
||||||
path: /metrics
|
|
||||||
interval: 30s
|
|
||||||
scrapeTimeout: 10s
|
|
||||||
@@ -1,24 +0,0 @@
|
|||||||
# Homarr
|
|
||||||
|
|
||||||
Dashboard with Kubernetes integration and persistent application state.
|
|
||||||
|
|
||||||
Kubernetes uses the `homarr` ServiceAccount and the read-only ClusterRole in
|
|
||||||
`k8s/rbac.yaml`. Application data lives in `homarr-pvc`; supply the encryption key
|
|
||||||
from `k8s/secrets.yaml.example` before the first start and retain it with backups.
|
|
||||||
|
|
||||||
The committed ingress is internal. There is no `k8s/active` marker even though
|
|
||||||
manifests exist, so the workflow does not select Homarr automatically.
|
|
||||||
|
|
||||||
Compose publishes ports 80 and 81, mounts appdata and the Docker socket, and
|
|
||||||
expects a local kubeconfig. Check these host ports against Traefik before use.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n homarr
|
|
||||||
kubectl get events -n homarr --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,25 +0,0 @@
|
|||||||
# Homepages
|
|
||||||
|
|
||||||
Two static sites: Forust and xdfnx.
|
|
||||||
|
|
||||||
The site sources are in `forust_files/` and `xdfnx_files/`. CI builds each with
|
|
||||||
its own Dockerfile and publishes it to the private registry. Kubernetes serves
|
|
||||||
the image contents; Compose overlays the source directories as bind mounts.
|
|
||||||
|
|
||||||
Both Traefik IngressRoute and Gateway API route manifests are committed.
|
|
||||||
Keep their hostnames and backend Services aligned when changing routes.
|
|
||||||
Certificate resources cover public and internal hostnames.
|
|
||||||
|
|
||||||
Build either site locally with `docker build -f Dockerfile.forust .` or
|
|
||||||
`docker build -f Dockerfile.xdfnx .` from this directory.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n homepages
|
|
||||||
kubectl get events -n homepages --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,29 +0,0 @@
|
|||||||
# Immich
|
|
||||||
|
|
||||||
Photo library with its own vector-enabled PostgreSQL and machine-learning service.
|
|
||||||
|
|
||||||
This database is separate from the shared PostgreSQL instance. Keep the server
|
|
||||||
and machine-learning versions aligned when upgrading.
|
|
||||||
|
|
||||||
Kubernetes bind-mounts `/mnt/immich/library` from the node. That directory must
|
|
||||||
already exist and contain the intended library; moving the pod to a different
|
|
||||||
node does not move the files. PostgreSQL and Valkey use StatefulSet storage, and
|
|
||||||
the model cache has its own PVC.
|
|
||||||
|
|
||||||
Compose reads `UPLOAD_LOCATION` and `DB_DATA_LOCATION` from `.env`. The example
|
|
||||||
uses the same library path as Kubernetes. Run one writer against that library;
|
|
||||||
do not start both deployments as independent instances over the same files.
|
|
||||||
|
|
||||||
Back up the library and a consistent database dump together. The model cache
|
|
||||||
can be rebuilt; the photo database cannot.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n immich
|
|
||||||
kubectl get events -n immich --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -6,10 +6,6 @@ metadata:
|
|||||||
data:
|
data:
|
||||||
TZ: "Europe/Bratislava"
|
TZ: "Europe/Bratislava"
|
||||||
|
|
||||||
IMMICH_TELEMETRY_INCLUDE: "all"
|
|
||||||
IMMICH_API_METRICS_PORT: "8081"
|
|
||||||
IMMICH_MICROSERVICES_METRICS_PORT: "8082"
|
|
||||||
|
|
||||||
# The database in this namespace, not the shared one in the database
|
# The database in this namespace, not the shared one in the database
|
||||||
# namespace: v3 needs VectorChord, and only the dedicated image carries it.
|
# namespace: v3 needs VectorChord, and only the dedicated image carries it.
|
||||||
DB_HOSTNAME: "immich-postgres"
|
DB_HOSTNAME: "immich-postgres"
|
||||||
|
|||||||
@@ -3,8 +3,6 @@ kind: Service
|
|||||||
metadata:
|
metadata:
|
||||||
name: immich-service
|
name: immich-service
|
||||||
namespace: immich
|
namespace: immich
|
||||||
labels:
|
|
||||||
app: immich
|
|
||||||
spec:
|
spec:
|
||||||
selector:
|
selector:
|
||||||
app: immich
|
app: immich
|
||||||
@@ -12,12 +10,6 @@ spec:
|
|||||||
- name: http
|
- name: http
|
||||||
port: 2283
|
port: 2283
|
||||||
targetPort: 2283
|
targetPort: 2283
|
||||||
- name: api-metrics
|
|
||||||
port: 8081
|
|
||||||
targetPort: api-metrics
|
|
||||||
- name: worker-metrics
|
|
||||||
port: 8082
|
|
||||||
targetPort: worker-metrics
|
|
||||||
---
|
---
|
||||||
apiVersion: apps/v1
|
apiVersion: apps/v1
|
||||||
kind: Deployment
|
kind: Deployment
|
||||||
@@ -49,10 +41,6 @@ spec:
|
|||||||
ports:
|
ports:
|
||||||
- name: http
|
- name: http
|
||||||
containerPort: 2283
|
containerPort: 2283
|
||||||
- name: api-metrics
|
|
||||||
containerPort: 8081
|
|
||||||
- name: worker-metrics
|
|
||||||
containerPort: 8082
|
|
||||||
volumeMounts:
|
volumeMounts:
|
||||||
- name: immich-data
|
- name: immich-data
|
||||||
mountPath: /data
|
mountPath: /data
|
||||||
|
|||||||
@@ -1,20 +0,0 @@
|
|||||||
apiVersion: monitoring.coreos.com/v1
|
|
||||||
kind: ServiceMonitor
|
|
||||||
metadata:
|
|
||||||
name: immich
|
|
||||||
namespace: immich
|
|
||||||
labels:
|
|
||||||
release: prometheus-stack
|
|
||||||
spec:
|
|
||||||
selector:
|
|
||||||
matchLabels:
|
|
||||||
app: immich
|
|
||||||
endpoints:
|
|
||||||
- port: api-metrics
|
|
||||||
path: /metrics
|
|
||||||
interval: 30s
|
|
||||||
scrapeTimeout: 10s
|
|
||||||
- port: worker-metrics
|
|
||||||
path: /metrics
|
|
||||||
interval: 30s
|
|
||||||
scrapeTimeout: 10s
|
|
||||||
@@ -1,21 +0,0 @@
|
|||||||
# Kener
|
|
||||||
|
|
||||||
Status page with Redis and persistent database and upload directories.
|
|
||||||
|
|
||||||
Kubernetes uses `kener-db-pvc`, `kener-uploads-pvc`, and a Redis StatefulSet.
|
|
||||||
Compose keeps the corresponding directories in named volumes. Set the signing
|
|
||||||
and other credentials from the env or Secret example.
|
|
||||||
|
|
||||||
The monitors and route settings live in `k8s/config.yaml` and `k8s/ingress.yaml`.
|
|
||||||
There is no active marker for either runtime.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n kener
|
|
||||||
kubectl get events -n kener --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
+1
-1
@@ -1,6 +1,6 @@
|
|||||||
services:
|
services:
|
||||||
kener:
|
kener:
|
||||||
image: rajnandan1/kener:v4.1.7
|
image: rajnandan1/kener:4.1.5
|
||||||
container_name: kener
|
container_name: kener
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
# ports:
|
# ports:
|
||||||
|
|||||||
@@ -31,7 +31,7 @@ spec:
|
|||||||
spec:
|
spec:
|
||||||
containers:
|
containers:
|
||||||
- name: kener
|
- name: kener
|
||||||
image: rajnandan1/kener:v4.1.7
|
image: rajnandan1/kener:4.1.5
|
||||||
envFrom:
|
envFrom:
|
||||||
- configMapRef:
|
- configMapRef:
|
||||||
name: kener-config
|
name: kener-config
|
||||||
|
|||||||
@@ -1,16 +0,0 @@
|
|||||||
# Loki and Alloy
|
|
||||||
|
|
||||||
Loki log storage and Alloy collection, both deployed through Helm.
|
|
||||||
|
|
||||||
The deploy library lists separate `loki` and `alloy` releases in `prometheus`,
|
|
||||||
controlled by this directory's `k8s/active` marker. Chart versions are pinned in
|
|
||||||
`deploy-lib.sh`; settings live in `loki-values.yaml` and `alloy-values.yaml`.
|
|
||||||
|
|
||||||
Alloy collects Kubernetes logs. Grafana's Loki datasource is configured in the
|
|
||||||
monitoring stack. Review Loki retention and storage settings before enabling
|
|
||||||
collection on a new cluster.
|
|
||||||
|
|
||||||
Check releases with `helm list -n prometheus` and inspect collector logs before
|
|
||||||
assuming that an empty Grafana query means there were no events.
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,21 +0,0 @@
|
|||||||
# MeTube
|
|
||||||
|
|
||||||
Web downloader behind Traefik.
|
|
||||||
|
|
||||||
Compose bind-mounts `MeTube_downloads/` on the host. Kubernetes uses a 20 GiB
|
|
||||||
`emptyDir` for `/downloads`: completed downloads disappear when the pod is
|
|
||||||
replaced. Download files from the UI promptly if this temporary storage is intended.
|
|
||||||
|
|
||||||
Application settings are in `k8s/config.yaml`. Persisting downloads in Kubernetes
|
|
||||||
would require changing the volume to a PVC and choosing a storage policy.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n metube
|
|
||||||
kubectl get events -n metube --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,21 +0,0 @@
|
|||||||
# n8n
|
|
||||||
|
|
||||||
Workflow automation with persistent application and file storage.
|
|
||||||
|
|
||||||
Kubernetes keeps application state in `n8n-node-pvc` and files in
|
|
||||||
`n8n-files-pvc`; Compose uses `node-data` and `files` named volumes.
|
|
||||||
Webhook URLs and proxy settings are committed in the application config.
|
|
||||||
|
|
||||||
There is no active marker. Review the URLs before enabling the stack, and retain
|
|
||||||
the credential encryption key with the database or application-data backup.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n n8n
|
|
||||||
kubectl get events -n n8n --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
+1
-1
@@ -1,6 +1,6 @@
|
|||||||
services:
|
services:
|
||||||
n8n:
|
n8n:
|
||||||
image: docker.n8n.io/n8nio/n8n:2.43.2
|
image: docker.n8n.io/n8nio/n8n:2.43.1
|
||||||
container_name: n8n
|
container_name: n8n
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
environment:
|
environment:
|
||||||
|
|||||||
+1
-1
@@ -31,7 +31,7 @@ spec:
|
|||||||
spec:
|
spec:
|
||||||
containers:
|
containers:
|
||||||
- name: n8n
|
- name: n8n
|
||||||
image: docker.n8n.io/n8nio/n8n:2.43.2
|
image: docker.n8n.io/n8nio/n8n:2.43.1
|
||||||
envFrom:
|
envFrom:
|
||||||
- configMapRef:
|
- configMapRef:
|
||||||
name: n8n-config
|
name: n8n-config
|
||||||
|
|||||||
+4
-16
@@ -1,24 +1,12 @@
|
|||||||
# NetBird
|
# NetBird
|
||||||
|
|
||||||
Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind Traefik. The Compose configuration uses the external Docker `proxy` network and publishes only STUN UDP `3478` directly.
|
Self-hosted NetBird with the combined management, signal, relay, and STUN server. The dashboard and server run behind the repository's existing external Traefik instance on the Docker `proxy` network. Only STUN UDP `3478` is published directly.
|
||||||
|
|
||||||
The single-instance server uses SQLite. Back up its data and datastore encryption key together.
|
The deployment uses SQLite for a single-instance homelab server. The persistent `netbird_data` volume and the datastore encryption key are both required to recover the installation.
|
||||||
|
|
||||||
## Kubernetes
|
|
||||||
|
|
||||||
`k8s/active` selects the Kubernetes deployment. It runs the server and dashboard
|
|
||||||
in namespace `netbird`; the server stores SQLite data in `netbird-pvc`. The
|
|
||||||
configuration renderer and template are in `k8s/`. Prepare
|
|
||||||
`k8s/secrets.yaml` from `k8s/secrets.yaml.example` before the first deploy.
|
|
||||||
|
|
||||||
## Compose alternative
|
|
||||||
|
|
||||||
The Compose files are available for manual use. There is no root `active` marker,
|
|
||||||
so the automatic deploy workflow selects Kubernetes only.
|
|
||||||
|
|
||||||
## Files
|
## Files
|
||||||
|
|
||||||
- `compose.yaml`: dashboard and combined server; start it manually when using Compose.
|
- `compose.yaml`: dashboard and combined server; selected by the marker-driven deploy workflow through `active`.
|
||||||
- `config.template.yaml`: non-secret server configuration rendered at startup.
|
- `config.template.yaml`: non-secret server configuration rendered at startup.
|
||||||
- `entrypoint.sh`: injects Docker secrets into an in-memory runtime configuration.
|
- `entrypoint.sh`: injects Docker secrets into an in-memory runtime configuration.
|
||||||
- `client.compose.yaml`: optional host-network peer using a dashboard-generated setup key.
|
- `client.compose.yaml`: optional host-network peer using a dashboard-generated setup key.
|
||||||
@@ -27,7 +15,7 @@ so the automatic deploy workflow selects Kubernetes only.
|
|||||||
|
|
||||||
## First deployment
|
## First deployment
|
||||||
|
|
||||||
Run these commands on the Docker host before the first Compose start.
|
Run these commands on the Docker host before merging the activating branch. The deploy preflight resets tracked files but preserves ignored local state.
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
cd /srv/homelab/netbird
|
cd /srv/homelab/netbird
|
||||||
|
|||||||
@@ -3,8 +3,6 @@ kind: Service
|
|||||||
metadata:
|
metadata:
|
||||||
name: netbird-server-service
|
name: netbird-server-service
|
||||||
namespace: netbird
|
namespace: netbird
|
||||||
labels:
|
|
||||||
app: netbird-server
|
|
||||||
spec:
|
spec:
|
||||||
selector:
|
selector:
|
||||||
app: netbird-server
|
app: netbird-server
|
||||||
@@ -13,10 +11,6 @@ spec:
|
|||||||
name: http
|
name: http
|
||||||
targetPort: 80
|
targetPort: 80
|
||||||
protocol: TCP
|
protocol: TCP
|
||||||
- port: 9090
|
|
||||||
name: metrics
|
|
||||||
targetPort: metrics
|
|
||||||
protocol: TCP
|
|
||||||
- port: 3478
|
- port: 3478
|
||||||
name: stun
|
name: stun
|
||||||
targetPort: 3478
|
targetPort: 3478
|
||||||
@@ -65,9 +59,6 @@ spec:
|
|||||||
- containerPort: 80
|
- containerPort: 80
|
||||||
name: http
|
name: http
|
||||||
protocol: TCP
|
protocol: TCP
|
||||||
- containerPort: 9090
|
|
||||||
name: metrics
|
|
||||||
protocol: TCP
|
|
||||||
- containerPort: 3478
|
- containerPort: 3478
|
||||||
name: stun
|
name: stun
|
||||||
protocol: UDP
|
protocol: UDP
|
||||||
|
|||||||
+88
-35
@@ -1,43 +1,96 @@
|
|||||||
# NetBox
|
# NetBox
|
||||||
|
|
||||||
Inventory and network documentation with a web process, worker, and Valkey.
|
NetBox for homelab documentation and visualization. Two runtimes are available:
|
||||||
|
|
||||||
Kubernetes uses the shared PostgreSQL service at
|
| Runtime | Manifest | Purpose |
|
||||||
`postgres.database.svc.cluster.local:5432`, database and role `netbox`.
|
| ------- | -------------- | -------------------------------------------------------------- |
|
||||||
The database and application Secrets must contain the same password.
|
| Docker | `compose.yaml` | Local stand on `127.0.0.1:8000` (no public exposure) |
|
||||||
Media, reports, scripts, and Valkey have persistent storage.
|
| k8s | `k8s/` | Homelab service on `netbox.forust.xyz` (and the internal name) |
|
||||||
|
|
||||||
Compose has its own PostgreSQL container and Valkey instances. It publishes the
|
Both use the same image (`netboxcommunity/netbox:v4.7-5.1.1`) and Valkey for tasks
|
||||||
web UI on `127.0.0.1:8000`; its Traefik labels can also expose it while a Docker
|
plus a second logical database for caching. The Docker stand keeps its own
|
||||||
proxy is running. Copy `.env.example` to `.env`, replace the credentials, and run
|
PostgreSQL container, while the k8s deployment uses the shared `database` cluster
|
||||||
`docker compose config --quiet` before starting it.
|
(`postgres.database.svc.cluster.local:5432`, role/database `netbox`); only Valkey
|
||||||
|
stays a per-service StatefulSet.
|
||||||
|
|
||||||
## First Kubernetes start
|
## Docker Compose
|
||||||
|
|
||||||
Create the namespace and application Secret. Provision the database through the
|
```bash
|
||||||
shared database initializer on a fresh instance, or create the role and database
|
cp .env.example .env
|
||||||
manually on an existing instance; see [PostgreSQL](../postgres/README.md).
|
# replace CHANGE_ME
|
||||||
The database NetworkPolicy already includes `netbox`.
|
docker compose up -d
|
||||||
|
|
||||||
Apply the selected application manifests after the database is ready. Startup
|
|
||||||
runs schema migrations, so the probes allow a longer first boot. Inspect web and
|
|
||||||
worker logs before retrying a slow migration.
|
|
||||||
|
|
||||||
## Settings and backup
|
|
||||||
|
|
||||||
`configuration/configuration.py` is the Compose settings file. Its Kubernetes
|
|
||||||
copy is embedded in `k8s/settings.yaml`; keep them aligned.
|
|
||||||
Back up the database and media together. Keep `SECRET_KEY` and
|
|
||||||
`API_TOKEN_PEPPER_1`: changing them invalidates sessions or API tokens.
|
|
||||||
A container rollback cannot undo a database migration.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n netbox
|
|
||||||
kubectl get events -n netbox --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
```
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
The UI is available at <http://localhost:8000>. The port is bound to `127.0.0.1`
|
||||||
|
intentionally, so this stand is not exposed on the LAN or public interfaces.
|
||||||
|
|
||||||
|
The `netbox` service is also attached to the external `proxy` network and carries
|
||||||
|
Traefik labels for `netbox.forust.xyz` and `netbox.workstation.internal`. Those
|
||||||
|
labels only take effect while the Docker Traefik stack is running; it is currently
|
||||||
|
stopped, and the live ingress path in this homelab is the k8s Traefik.
|
||||||
|
|
||||||
|
Inspect startup and health with:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker compose ps
|
||||||
|
docker compose logs -f netbox
|
||||||
|
```
|
||||||
|
|
||||||
|
Stop it with `docker compose down`; data is kept in the named volumes
|
||||||
|
`netbox-postgres`, `netbox-media-files`, `netbox-reports-files`,
|
||||||
|
`netbox-scripts-files` and `netbox-redis-data`.
|
||||||
|
|
||||||
|
## Kubernetes
|
||||||
|
|
||||||
|
`k8s/` is deployed in the homelab cluster and serves `netbox.forust.xyz` publicly
|
||||||
|
plus `netbox.workstation.internal` / `netbox.gigaforust.internal` internally. To
|
||||||
|
rebuild it from scratch:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. shared PostgreSQL: the password lives in the shared secret, NetBox keeps a copy
|
||||||
|
kubectl -n database patch secret postgres-shared-secrets \
|
||||||
|
--type merge -p '{"stringData":{"NETBOX_DB_PASSWORD":"<same value>"}}'
|
||||||
|
kubectl -n database exec postgres17-0 -- psql -U postgres -d postgres \
|
||||||
|
-c 'CREATE ROLE netbox LOGIN PASSWORD ...' -c 'CREATE DATABASE netbox OWNER netbox'
|
||||||
|
|
||||||
|
# 2. secrets first: the deploy workflow never applies *secret*.yaml
|
||||||
|
cp k8s/secrets.yaml.example k8s/secrets.yaml # replace CHANGE_ME
|
||||||
|
kubectl apply -f k8s/secrets.yaml
|
||||||
|
|
||||||
|
# 3. manifests
|
||||||
|
kubectl apply -f k8s/
|
||||||
|
```
|
||||||
|
|
||||||
|
The shared cluster is reached at `postgres.database.svc.cluster.local:5432`. Its
|
||||||
|
NetworkPolicy (`postgres/k8s/network-policy.yaml`) must list the `netbox` namespace
|
||||||
|
or connections are dropped, and `postgres/initdb/01-create-databases.sh` already
|
||||||
|
creates the role and database on a fresh data directory. NetBox has no PostgreSQL
|
||||||
|
StatefulSet of its own — only `netbox-valkey`.
|
||||||
|
|
||||||
|
`netbox.forust.xyz` resolves to this host (`78.98.72.122`) through the `DOMAINS`
|
||||||
|
list in the `default/cfddns` secret. cert-manager issues `netbox-prod-tls` with the
|
||||||
|
`letsencrypt-prod` issuer, the internal route uses `internal-wildcard-tls`.
|
||||||
|
|
||||||
|
Resources are permanent again now that the first-boot migrations are complete:
|
||||||
|
the web container reserves `100m`/`512Mi` and is capped at `2` CPU/`2Gi`, the
|
||||||
|
worker reserves `50m`/`256Mi` and is capped at `1` CPU/`1Gi`, and Valkey reserves
|
||||||
|
`25m`/`64Mi` and is capped at `250m`/`256Mi`. The deliberately generous CPU caps
|
||||||
|
leave enough headroom for future schema migrations without letting one process
|
||||||
|
consume the whole node.
|
||||||
|
|
||||||
|
The first start applies ~810 migrations, each in its own transaction with DDL and
|
||||||
|
a commit; every later start is a no-op. The startup probe allows 15 minutes and
|
||||||
|
`progressDeadlineSeconds` is 1800 for the same reason. Probes run inside the pod
|
||||||
|
and explicitly set `Host: netbox.forust.xyz`; a kubelet `httpGet.host` field would
|
||||||
|
replace the probe destination with that public hostname and bypass the pod.
|
||||||
|
|
||||||
|
## Secrets
|
||||||
|
|
||||||
|
- `netbox/.env` (compose) and `netbox/k8s/secrets.yaml` (k8s) are gitignored. Only
|
||||||
|
`.env.example` and `k8s/secrets.yaml.example` are committed.
|
||||||
|
- `netbox/configuration/configuration.py` is env-driven: hosts, database, Redis and
|
||||||
|
the Django keys all come from the environment, so the same settings file works in
|
||||||
|
both runtimes. The k8s copy lives in the `netbox-settings` ConfigMap
|
||||||
|
(`k8s/settings.yaml`) and must be kept in sync with the file.
|
||||||
|
- Rotating `SECRET_KEY` invalidates all sessions; rotating `API_TOKEN_PEPPER_1`
|
||||||
|
invalidates every API token.
|
||||||
@@ -1,22 +0,0 @@
|
|||||||
# Netronome
|
|
||||||
|
|
||||||
Network monitoring application using the shared PostgreSQL instance on Kubernetes.
|
|
||||||
|
|
||||||
Kubernetes reads application settings from its ConfigMap and Secret. Match the
|
|
||||||
Netronome role password with `NETRONOME_DB_PASSWORD` in the shared database Secret.
|
|
||||||
Its namespace is included in the PostgreSQL NetworkPolicy.
|
|
||||||
|
|
||||||
The Compose configuration is a separate deployment; review its local database
|
|
||||||
settings and env example before starting it. Keep monitoring history in the
|
|
||||||
database backup.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n netronome
|
|
||||||
kubectl get events -n netronome --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,17 +0,0 @@
|
|||||||
# Nextcloud AIO
|
|
||||||
|
|
||||||
Nextcloud All-in-One on Docker, with Kubernetes routes to the Docker host.
|
|
||||||
|
|
||||||
The master container manages its own child containers through the Docker
|
|
||||||
socket. Kubernetes does not run the Nextcloud application; the EndpointSlices
|
|
||||||
under `k8s/routing/` point to host services.
|
|
||||||
|
|
||||||
Compose publishes the AIO administration interface on 8888. The Apache frontend
|
|
||||||
uses host port 11000. `NEXTCLOUD_DATADIR` is `/mnt/nextcloud/ncdata`; prepare that
|
|
||||||
storage before first setup and do not change the path casually afterwards.
|
|
||||||
|
|
||||||
Use AIO's backup and restore tools for the managed application. Keep the master
|
|
||||||
configuration volume and the data directory with the recovery plan. Do not
|
|
||||||
remove child containers just because they do not appear as Compose services.
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,12 +0,0 @@
|
|||||||
# Penpot
|
|
||||||
|
|
||||||
A Compose-only design application with frontend, backend, exporter, database, and cache.
|
|
||||||
|
|
||||||
There is no active marker or Kubernetes deployment here. Configure the public
|
|
||||||
URL and credentials from `.env.example` before starting `compose.yaml`.
|
|
||||||
|
|
||||||
Penpot has its own PostgreSQL container. The shared database initializer still
|
|
||||||
contains a Penpot role, but this Compose stack does not use it.
|
|
||||||
Back up the application assets and database together.
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,21 +0,0 @@
|
|||||||
# Portainer
|
|
||||||
|
|
||||||
Container management UI backed by the host Docker socket.
|
|
||||||
|
|
||||||
Kubernetes mounts the node's Docker socket and persists application data in
|
|
||||||
`portainer-data-pvc`. This targets Docker on that node, not Kubernetes workloads.
|
|
||||||
Compose uses the `portainer_data` volume for its state.
|
|
||||||
|
|
||||||
Review initial administrator setup and route access before exposing the UI.
|
|
||||||
Neither deployment has an active marker.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n portainer
|
|
||||||
kubectl get events -n portainer --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
services:
|
services:
|
||||||
portainer:
|
portainer:
|
||||||
image: portainer/portainer-ce:2.45.2
|
image: portainer/portainer-ce:2.45.1
|
||||||
container_name: portainer
|
container_name: portainer
|
||||||
restart: always
|
restart: always
|
||||||
volumes:
|
volumes:
|
||||||
|
|||||||
@@ -29,7 +29,7 @@ spec:
|
|||||||
spec:
|
spec:
|
||||||
containers:
|
containers:
|
||||||
- name: portainer
|
- name: portainer
|
||||||
image: portainer/portainer-ce:2.45.2
|
image: portainer/portainer-ce:2.45.1
|
||||||
ports:
|
ports:
|
||||||
- containerPort: 9000
|
- containerPort: 9000
|
||||||
volumeMounts:
|
volumeMounts:
|
||||||
|
|||||||
+28
-56
@@ -1,63 +1,35 @@
|
|||||||
# Shared PostgreSQL
|
# Shared PostgreSQL
|
||||||
|
|
||||||
PostgreSQL 17 for the Kubernetes deployments of Authentik, Gitea, NetBox, and Netronome.
|
This directory contains the shared PostgreSQL 17 deployment for Authentik,
|
||||||
|
Gitea, NetBox, Netronome, and Statuspage. It creates one database and one login role
|
||||||
|
per service. Per-service standalone databases were removed after the
|
||||||
|
migration (Sep 2026); Penpot stays on its own compose PostgreSQL (archived,
|
||||||
|
not part of the shared instance).
|
||||||
|
|
||||||
The server runs in `database` as StatefulSet `postgres17`, with data in
|
## Compatibility baseline
|
||||||
`postgres17-data`. Applications connect to
|
|
||||||
`postgres.database.svc.cluster.local:5432`. The NetworkPolicy allows only the
|
|
||||||
listed application namespaces; add a new consumer there as well as provisioning
|
|
||||||
its database.
|
|
||||||
|
|
||||||
## Initialization
|
| Service | Current application | Shared PostgreSQL 17 |
|
||||||
|
| ---------- | ------------------- | -------------------------------------- |
|
||||||
|
| Authentik | 2025.10.x | Supported (Authentik requires 14+) |
|
||||||
|
| Gitea | 1.27.3 | Supported (Gitea requires 12+) |
|
||||||
|
| NetBox | 4.7.x | Supported (NetBox 4.x requires 13+) |
|
||||||
|
| Netronome | 0.14.0 | Supported (upstream's example uses 17) |
|
||||||
|
| Statuspage | custom | Supported |
|
||||||
|
|
||||||
`initdb/01-create-databases.sh` creates roles and databases on an empty data
|
A major-version change must use a logical dump/restore; changing only the
|
||||||
directory. The Kubernetes copy is embedded in `k8s/postgres.yaml`.
|
image tag while keeping a data directory is not supported.
|
||||||
It also provisions Penpot and Statuspage roles, even though those are not active
|
|
||||||
consumers in the current Kubernetes manifests.
|
|
||||||
|
|
||||||
The initializer requires every listed password. Prepare `k8s/secrets.yaml` from
|
For Compose, copy `.env.example` to `.env`, set all passwords, and start it with
|
||||||
the example before applying the StatefulSet. Existing application Secrets keep
|
`docker compose -f shared-compose.yaml up -d`. This file is intentionally not
|
||||||
copies of their own database passwords; they must match the corresponding role.
|
named `compose.yaml`, so the repository deploy workflow does not start a second
|
||||||
|
database accidentally.
|
||||||
|
Applications that use this database must also join that external network and use
|
||||||
|
`homelab-postgres:5432`.
|
||||||
|
|
||||||
The init scripts do not run again when an existing data directory is mounted.
|
For Kubernetes, create `k8s/secrets.yaml` from the example before applying the
|
||||||
Changing a Secret does not rotate the PostgreSQL role password. Rotate the role
|
manifests. The `k8s/active` marker makes the normal deploy workflow include the
|
||||||
with SQL and update the application Secret together.
|
namespace, StatefulSet, ConfigMap, and NetworkPolicy. Applications use
|
||||||
|
`postgres.database.svc.cluster.local:5432`.
|
||||||
## Compose alternative
|
Migrate each existing database with a tested logical dump/restore before
|
||||||
|
switching an application. Do not reuse a PostgreSQL 14 or 17 data directory
|
||||||
From this directory:
|
with PostgreSQL 15.
|
||||||
|
|
||||||
```sh
|
|
||||||
cp .env.example .env
|
|
||||||
$EDITOR .env
|
|
||||||
docker compose -f shared-compose.yaml config --quiet
|
|
||||||
docker compose -f shared-compose.yaml up -d
|
|
||||||
```
|
|
||||||
|
|
||||||
The example includes `NETBOX_DB_PASSWORD`; fill it and every other required
|
|
||||||
password before starting the stack.
|
|
||||||
This stack creates the `homelab-database` Docker network and the
|
|
||||||
`homelab-postgres` container. Compose applications need to join that network
|
|
||||||
explicitly to use it; several committed Compose stacks use their own databases.
|
|
||||||
|
|
||||||
The filename is intentional: the automatic deploy discovery does not start this
|
|
||||||
stack just because the Kubernetes database is active.
|
|
||||||
|
|
||||||
## Backup and upgrades
|
|
||||||
|
|
||||||
Keep database dumps and role definitions, including ownership and grants.
|
|
||||||
Take a logical backup before changing a major PostgreSQL version. A new image
|
|
||||||
tag over the existing data directory is not a major-version migration.
|
|
||||||
Test restores separately before changing application connection settings.
|
|
||||||
Immich uses its own vector-enabled database and is outside this shared instance.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n database
|
|
||||||
kubectl get events -n database --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,32 +0,0 @@
|
|||||||
# Monitoring stack
|
|
||||||
|
|
||||||
The Kubernetes stack provides Grafana, Alertmanager, VictoriaMetrics, VMAgent,
|
|
||||||
and vmalert. The `k8s/active` marker selects the stack. The
|
|
||||||
`kube-prometheus-stack` Helm release installs Grafana, Alertmanager, the
|
|
||||||
Prometheus Operator, and related components. Its Prometheus server is configured
|
|
||||||
with zero replicas while VMAgent collects metrics and writes them to the
|
|
||||||
single-node VictoriaMetrics instance.
|
|
||||||
|
|
||||||
The `victoria-operator` Helm release converts selected Prometheus Operator
|
|
||||||
`ServiceMonitor` resources into `VMServiceScrape` resources. VMAgent selects
|
|
||||||
those scrapes across namespaces and writes to VictoriaMetrics. vmalert evaluates
|
|
||||||
the rule ConfigMap and sends alerts to the stack's Alertmanager. See the
|
|
||||||
[Kubernetes monitoring notes](k8s/README.md) for application metrics and
|
|
||||||
validation commands.
|
|
||||||
|
|
||||||
The chart versions are pinned in `.gitea/workflows/deploy-lib.sh`. The tracked
|
|
||||||
`k8s/grafana-values.yaml` contains the Helm values for the stack. Create the
|
|
||||||
`grafana-admin` and `alertmanager-config` Secrets from the examples in `k8s/`;
|
|
||||||
keep their credentials out of the values file. Persistent volumes store data for
|
|
||||||
Prometheus, Grafana, Alertmanager, and VictoriaMetrics. Check the PVCs and
|
|
||||||
backups before changing storage. VictoriaMetrics currently retains 30 days of
|
|
||||||
data.
|
|
||||||
|
|
||||||
A separate Compose configuration is present for manual use. There is no root
|
|
||||||
`active` marker, so the automatic deploy workflow does not select it.
|
|
||||||
|
|
||||||
The deploy workflow does not remove resources when manifests are deleted. For a
|
|
||||||
rollback of application-metrics changes, follow the explicit cleanup steps in
|
|
||||||
the [Kubernetes monitoring notes](k8s/README.md).
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -15,29 +15,3 @@ The VictoriaMetrics Operator chart and its CRDs are installed before the
|
|||||||
Kubernetes manifests by the normal deploy workflow. On a cluster where the
|
Kubernetes manifests by the normal deploy workflow. On a cluster where the
|
||||||
operator CRDs are not installed yet, CI skips the server-side dry-run of the
|
operator CRDs are not installed yet, CI skips the server-side dry-run of the
|
||||||
`VMAgent` resource; the deploy installs the chart before applying that resource.
|
`VMAgent` resource; the deploy installs the chart before applying that resource.
|
||||||
|
|
||||||
## Application metrics
|
|
||||||
|
|
||||||
The application ServiceMonitors use a 30s interval and a 10s timeout:
|
|
||||||
|
|
||||||
- Headscale: the external Service points to the Compose host on port 19090.
|
|
||||||
A VMServiceScrape uses EndpointSlice discovery for this manually managed target.
|
|
||||||
The Compose configuration must bind metrics to `0.0.0.0:9090`.
|
|
||||||
- NetBird: the combined server exports `/metrics` on port 9090. The existing
|
|
||||||
`server.metricsPort` setting enables the listener.
|
|
||||||
- Gitea: `GITEA__metrics__ENABLED` enables `/metrics` on the HTTP port. The public
|
|
||||||
ingress excludes this path. The monitor uses the internal Service directly.
|
|
||||||
- Immich: `IMMICH_TELEMETRY_INCLUDE=all` enables API and worker metrics on ports
|
|
||||||
8081 and 8082. The monitor scrapes both ports on each server replica.
|
|
||||||
|
|
||||||
Deploy through the existing CI and deploy workflow. Gitea and Immich reload their
|
|
||||||
ConfigMap changes through Reloader. Check the VMAgent targets after deployment
|
|
||||||
and query `up{scraper="victoria",namespace=~"netbird|gitea|immich|headscale"}` in
|
|
||||||
VictoriaMetrics. All targets should report 1.
|
|
||||||
|
|
||||||
For rollback, revert the application metrics changes, run CI, and deploy the
|
|
||||||
revert. Remove the three application ServiceMonitors and the Headscale VMServiceScrape explicitly: the deployment
|
|
||||||
workflow applies manifests and does not prune removed resources.
|
|
||||||
|
|
||||||
For Headscale rollback, remove its VMServiceScrape and Service label, restore the
|
|
||||||
previous Compose metrics bind address, and restart only the Headscale service.
|
|
||||||
@@ -1,20 +0,0 @@
|
|||||||
# RackPeek
|
|
||||||
|
|
||||||
Rack inventory UI behind Traefik.
|
|
||||||
|
|
||||||
Kubernetes stores configuration in `rackpeek-pvc`. The Compose alternative uses
|
|
||||||
its own data mount. Keep rack descriptions and inventory data in the backup.
|
|
||||||
|
|
||||||
Public and internal certificates and routes are in `k8s/`. There are no tracked
|
|
||||||
Secret examples for this service.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n rackpeek
|
|
||||||
kubectl get events -n rackpeek --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,56 +0,0 @@
|
|||||||
# Reloader
|
|
||||||
|
|
||||||
Restarts opted-in workloads when the ConfigMaps or Secrets they consume change.
|
|
||||||
The deploy workflow upgrades the `reloader` Helm release in namespace `reloader`;
|
|
||||||
`k8s/active` enables it. The chart version is pinned in `deploy-lib.sh`.
|
|
||||||
|
|
||||||
## Workload integration
|
|
||||||
|
|
||||||
Put this annotation on the Deployment or StatefulSet metadata:
|
|
||||||
|
|
||||||
```yaml
|
|
||||||
metadata:
|
|
||||||
annotations:
|
|
||||||
reloader.stakater.com/auto: "true"
|
|
||||||
```
|
|
||||||
|
|
||||||
The annotation belongs to the workload, not `spec.template.metadata`.
|
|
||||||
Reloader discovers references in environment variables and mounted volumes.
|
|
||||||
This covers startup-only settings and ConfigMaps or Secrets mounted with `subPath`.
|
|
||||||
See the [upstream usage guide](https://github.com/stakater/Reloader/blob/v1.4.22/README.md#usage).
|
|
||||||
|
|
||||||
The application manifests opt in workloads including AdGuard's TLS files,
|
|
||||||
NetBird, both NetBox processes, and the password-protected Valkey servers.
|
|
||||||
Inactive services have the same annotations ready for later activation.
|
|
||||||
|
|
||||||
## Controller policy
|
|
||||||
|
|
||||||
The controller watches all namespaces but only restarts annotated workloads.
|
|
||||||
It uses the `annotations` reload strategy, so changes trigger a pod-template
|
|
||||||
annotation rather than injecting extra environment variables.
|
|
||||||
|
|
||||||
Jobs and CronJobs are excluded: their next execution reads current configuration.
|
|
||||||
PostgreSQL is intentionally not opted in. Its password variables and init scripts
|
|
||||||
apply to first initialization; restarting an existing database does not rotate
|
|
||||||
roles or rerun those scripts. Rotate database credentials with SQL and update the
|
|
||||||
clients' Secrets together.
|
|
||||||
|
|
||||||
Helm-managed monitoring components already have their own configuration reload
|
|
||||||
paths; Traefik watches its file-provider configuration. They are not globally
|
|
||||||
opted in. The controller does not react to files in PVCs or changes to external
|
|
||||||
services unless a watched ConfigMap or Secret changes.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl -n reloader rollout status deployment/reloader-reloader
|
|
||||||
kubectl -n reloader logs deployment/reloader-reloader --since=10m
|
|
||||||
kubectl -n netbird get deployment netbird-server-deployment \
|
|
||||||
-o jsonpath='{.metadata.annotations.reloader\.stakater\.com/auto}'
|
|
||||||
```
|
|
||||||
|
|
||||||
A changed configuration can briefly interrupt a single-replica service, especially
|
|
||||||
one using `Recreate`. Installing annotations does not validate the configuration
|
|
||||||
or migrate database data. Keep changes to shared Secrets coordinated across consumers.
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
+88
-38
@@ -1,51 +1,101 @@
|
|||||||
# Renovate
|
# Renovate for Gitea
|
||||||
|
|
||||||
Container and chart dependency updates for the Gitea repository.
|
Renovate runs as a Kubernetes CronJob and creates container image update pull
|
||||||
|
requests in Gitea. It does not deploy changes itself.
|
||||||
|
|
||||||
The Kubernetes CronJob runs in `renovate` every six hours with overlapping
|
## Kubernetes
|
||||||
CronJob executions forbidden. Prepare the bot PAT from the Secret example.
|
|
||||||
Give the dedicated Gitea user access to the repositories it should update.
|
|
||||||
|
|
||||||
`renovate.json` is the source configuration. The ConfigMap is a generated copy:
|
Create a dedicated Gitea user named `renovate-bot`, create a repository access
|
||||||
|
token, and grant it repository read/write plus issue read/write permissions.
|
||||||
|
Add `read:packages` if Renovate must inspect private Gitea registry images.
|
||||||
|
|
||||||
|
Create the ignored Secret locally; never commit the PAT:
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
.gitea/workflows/sync-renovate-configmap.sh
|
cp renovate/k8s/secrets.yaml.example renovate/k8s/secrets.yaml
|
||||||
.gitea/workflows/sync-renovate-configmap.sh --check
|
$EDITOR renovate/k8s/secrets.yaml
|
||||||
|
kubectl apply -f renovate/k8s/namespace.yaml
|
||||||
|
kubectl apply -f renovate/k8s/secrets.yaml
|
||||||
|
kubectl apply -f renovate/k8s/configmap.yaml
|
||||||
|
kubectl apply -f renovate/k8s/cronjob.yaml
|
||||||
```
|
```
|
||||||
|
|
||||||
Run those commands from the repository root. The `renovate-ci` workflow checks
|
The `renovate/k8s/active` marker makes the normal deployment workflow include
|
||||||
that the generated configuration agrees with the source.
|
the namespace, ConfigMap, and CronJob. The Secret is intentionally excluded
|
||||||
|
from Git and must be applied separately after every new cluster.
|
||||||
|
|
||||||
## Run manually
|
Run it immediately instead of waiting for the six-hour schedule.
|
||||||
|
|
||||||
From the repository root:
|
Two options, both use the same `renovate/renovate.json`:
|
||||||
|
|
||||||
```fish
|
|
||||||
kubectl create job --from=cronjob/renovate renovate-manual-(date +%s) -n renovate
|
|
||||||
kubectl get jobs,pods -n renovate
|
|
||||||
```
|
|
||||||
|
|
||||||
Alternatively use the `renovate-run` Actions workflow. It reads the image tag
|
|
||||||
from the CronJob and accepts repository, log-level, and dry-run inputs. Actions
|
|
||||||
requires `RENOVATE_TOKEN`; `RENOVATE_GITHUB_COM_TOKEN` is optional.
|
|
||||||
The Actions concurrency group and the CronJob policy are separate, so avoid
|
|
||||||
starting both against the same repository at once.
|
|
||||||
|
|
||||||
For Compose, copy `.env.example` to `.env` in this directory and run
|
|
||||||
`docker compose -f renovate-compose.yaml run --rm renovate`. That file is a
|
|
||||||
manual entry point and is not selected by the deploy workflow.
|
|
||||||
|
|
||||||
The config also tracks chart versions in `deploy-lib.sh` and tool versions in
|
|
||||||
`.gitea/workflows/tool-versions.env`. Renovate opens pull requests; the normal CI and deploy
|
|
||||||
workflows handle changes after merge.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
kubectl get pods,svc,pvc -n renovate
|
kubectl create job --from=cronjob/renovate renovate-manual-$(date +%s) -n renovate
|
||||||
kubectl get events -n renovate --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
```
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
or the `renovate-run` Actions workflow (Actions tab → `renovate-run` →
|
||||||
|
Run workflow). It runs the same image as the CronJob on the self-hosted runner
|
||||||
|
via Docker — the tag is read out of `renovate/k8s/cronjob.yaml` at run time
|
||||||
|
rather than hardcoded, so the two cannot drift apart. Required Actions secrets
|
||||||
|
(repo or org settings):
|
||||||
|
|
||||||
|
- `RENOVATE_TOKEN` — renovate-bot PAT (repository + issue read/write).
|
||||||
|
- `RENOVATE_GITHUB_COM_TOKEN` — optional, for changelogs and GitHub rate limits.
|
||||||
|
|
||||||
|
Inputs: `repositories` (default `forust/homelab`), `log_level`
|
||||||
|
(`info`/`debug`). Only one run at a time (concurrency group
|
||||||
|
`renovate-run`), same as the CronJob `Forbid` policy.
|
||||||
|
|
||||||
|
Inspect runs with:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
kubectl get cronjob,jobs,pods -n renovate
|
||||||
|
kubectl logs -n renovate job/<job-name>
|
||||||
|
```
|
||||||
|
|
||||||
|
`RENOVATE_GITHUB_COM_TOKEN` is optional but recommended for changelogs and
|
||||||
|
GitHub API rate limits. Set it in the Kubernetes Secret if available.
|
||||||
|
|
||||||
|
## Compose
|
||||||
|
|
||||||
|
Copy `.env.example` to `.env`, set the PAT, and run:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
docker compose -f renovate-compose.yaml run --rm renovate
|
||||||
|
```
|
||||||
|
|
||||||
|
The Compose file is intentionally named `renovate-compose.yaml`, so the
|
||||||
|
repository's automatic deployment discovery does not start it accidentally.
|
||||||
|
|
||||||
|
## Configuration
|
||||||
|
|
||||||
|
`renovate/renovate.json` is the single source of truth. The Compose file and the
|
||||||
|
`renovate-run` workflow mount that file directly.
|
||||||
|
|
||||||
|
A ConfigMap cannot read from the repository, so the CronJob needs the config
|
||||||
|
inlined. `renovate/k8s/configmap.yaml` is therefore a **generated** copy:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
.gitea/workflows/sync-renovate-configmap.sh # regenerate after editing
|
||||||
|
.gitea/workflows/sync-renovate-configmap.sh --check # fail if out of date
|
||||||
|
```
|
||||||
|
|
||||||
|
The `renovate-ci` workflow runs the `--check` form on every PR and push, so a
|
||||||
|
config edit that forgets to regenerate the ConfigMap cannot be merged.
|
||||||
|
|
||||||
|
Beyond images, `customManagers` in the config track:
|
||||||
|
|
||||||
|
- Helm chart versions pinned in `.gitea/workflows/deploy-lib.sh`. The built-in
|
||||||
|
`helmv3` manager only reads `Chart.yaml` and `helm-values` only reads values
|
||||||
|
files, so neither sees a version written into a `helm upgrade` command —
|
||||||
|
these are declared as `custom.regex` managers against the `helm` datasource.
|
||||||
|
- CI linter versions in `.gitea/workflows/tool-versions.env`.
|
||||||
|
|
||||||
|
The Renovate image tag is deliberately _not_ in `tool-versions.env`:
|
||||||
|
`renovate/k8s/cronjob.yaml` owns it, and the workflows read it from there.
|
||||||
|
|
||||||
|
## How updates flow
|
||||||
|
|
||||||
|
Renovate scans both `compose.yaml` files and Kubernetes manifests, opens a
|
||||||
|
branch and PR with image tag changes, and waits for CI. After merge, the
|
||||||
|
existing deployment workflow applies Kubernetes changes or redeploys Compose
|
||||||
|
stacks. Renovate never updates running workloads directly.
|
||||||
@@ -29,6 +29,31 @@ data:
|
|||||||
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
|
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
|
||||||
},
|
},
|
||||||
"customManagers": [
|
"customManagers": [
|
||||||
|
{
|
||||||
|
"customType": "regex",
|
||||||
|
"description": "singlesource: playwright npm version pinned in npx command (k8s + compose)",
|
||||||
|
"managerFilePatterns": ["edu_master/k8s/playwright.yaml", "edu_master/compose.yaml"],
|
||||||
|
"matchStrings": ["playwright@(?<currentValue>\\d+\\.\\d+\\.\\d+)"],
|
||||||
|
"datasourceTemplate": "npm",
|
||||||
|
"depNameTemplate": "playwright"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"customType": "regex",
|
||||||
|
"description": "singlesource: PLAYWRIGHT_VERSION file",
|
||||||
|
"managerFilePatterns": ["edu_master/PLAYWRIGHT_VERSION"],
|
||||||
|
"matchStrings": ["^(?<currentValue>\\d+\\.\\d+\\.\\d+)(?:\\r?\\n)?$"],
|
||||||
|
"datasourceTemplate": "pypi",
|
||||||
|
"depNameTemplate": "playwright"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"customType": "regex",
|
||||||
|
"description": "singlesource: playwright Python client version pinned in Dockerfile ARG",
|
||||||
|
"managerFilePatterns": ["edu_master/webinar-checker/Dockerfile"],
|
||||||
|
"matchStrings": ["(?:^|\\n)ARG PLAYWRIGHT_VERSION=(?<currentValue>\\d+\\.\\d+\\.\\d+)(?:\\r?\\n|$)"],
|
||||||
|
"datasourceTemplate": "pypi",
|
||||||
|
"depNameTemplate": "playwright",
|
||||||
|
"versioningTemplate": "pep440"
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "kube-prometheus-stack chart version pinned in the deploy workflow",
|
"description": "kube-prometheus-stack chart version pinned in the deploy workflow",
|
||||||
@@ -187,6 +212,17 @@ data:
|
|||||||
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
|
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
|
||||||
"enabled": false
|
"enabled": false
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"description": "singlesource playwright - use whichever version is found, keep docker+pypi+npm in sync",
|
||||||
|
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
|
||||||
|
"groupName": "playwright singlesource",
|
||||||
|
"groupSlug": "playwright"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"description": "playwright must not automerge - version skew breaks the WS handshake (checker.py:1523 vs playwright.yaml:20)",
|
||||||
|
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
|
||||||
|
"automerge": false
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
|
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
|
||||||
"matchPackageNames": ["renovate/renovate"],
|
"matchPackageNames": ["renovate/renovate"],
|
||||||
|
|||||||
@@ -19,7 +19,7 @@ spec:
|
|||||||
restartPolicy: Never
|
restartPolicy: Never
|
||||||
containers:
|
containers:
|
||||||
- name: renovate
|
- name: renovate
|
||||||
image: renovate/renovate:44.147.0
|
image: renovate/renovate:44.140.0
|
||||||
env:
|
env:
|
||||||
- name: RENOVATE_PLATFORM
|
- name: RENOVATE_PLATFORM
|
||||||
value: gitea
|
value: gitea
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ services:
|
|||||||
renovate:
|
renovate:
|
||||||
# Kept in step with renovate/k8s/cronjob.yaml by the "renovate self-update"
|
# Kept in step with renovate/k8s/cronjob.yaml by the "renovate self-update"
|
||||||
# package rule in renovate/renovate.json.
|
# package rule in renovate/renovate.json.
|
||||||
image: renovate/renovate:44.147.0
|
image: renovate/renovate:44.136.0
|
||||||
container_name: renovate
|
container_name: renovate
|
||||||
restart: "no"
|
restart: "no"
|
||||||
env_file:
|
env_file:
|
||||||
|
|||||||
@@ -18,6 +18,31 @@
|
|||||||
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
|
"managerFilePatterns": ["/k8s/.+\\.ya?ml$/"]
|
||||||
},
|
},
|
||||||
"customManagers": [
|
"customManagers": [
|
||||||
|
{
|
||||||
|
"customType": "regex",
|
||||||
|
"description": "singlesource: playwright npm version pinned in npx command (k8s + compose)",
|
||||||
|
"managerFilePatterns": ["edu_master/k8s/playwright.yaml", "edu_master/compose.yaml"],
|
||||||
|
"matchStrings": ["playwright@(?<currentValue>\\d+\\.\\d+\\.\\d+)"],
|
||||||
|
"datasourceTemplate": "npm",
|
||||||
|
"depNameTemplate": "playwright"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"customType": "regex",
|
||||||
|
"description": "singlesource: PLAYWRIGHT_VERSION file",
|
||||||
|
"managerFilePatterns": ["edu_master/PLAYWRIGHT_VERSION"],
|
||||||
|
"matchStrings": ["^(?<currentValue>\\d+\\.\\d+\\.\\d+)(?:\\r?\\n)?$"],
|
||||||
|
"datasourceTemplate": "pypi",
|
||||||
|
"depNameTemplate": "playwright"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"customType": "regex",
|
||||||
|
"description": "singlesource: playwright Python client version pinned in Dockerfile ARG",
|
||||||
|
"managerFilePatterns": ["edu_master/webinar-checker/Dockerfile"],
|
||||||
|
"matchStrings": ["(?:^|\\n)ARG PLAYWRIGHT_VERSION=(?<currentValue>\\d+\\.\\d+\\.\\d+)(?:\\r?\\n|$)"],
|
||||||
|
"datasourceTemplate": "pypi",
|
||||||
|
"depNameTemplate": "playwright",
|
||||||
|
"versioningTemplate": "pep440"
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"customType": "regex",
|
"customType": "regex",
|
||||||
"description": "kube-prometheus-stack chart version pinned in the deploy workflow",
|
"description": "kube-prometheus-stack chart version pinned in the deploy workflow",
|
||||||
@@ -176,6 +201,17 @@
|
|||||||
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
|
"matchPackageNames": ["/gcr\\.forust\\.xyz\\/forust\\/.+/"],
|
||||||
"enabled": false
|
"enabled": false
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"description": "singlesource playwright - use whichever version is found, keep docker+pypi+npm in sync",
|
||||||
|
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
|
||||||
|
"groupName": "playwright singlesource",
|
||||||
|
"groupSlug": "playwright"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"description": "playwright must not automerge - version skew breaks the WS handshake (checker.py:1523 vs playwright.yaml:20)",
|
||||||
|
"matchPackageNames": ["playwright", "mcr.microsoft.com/playwright"],
|
||||||
|
"automerge": false
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
|
"description": "Renovate updates itself in lockstep across the CronJob and the Compose file",
|
||||||
"matchPackageNames": ["renovate/renovate"],
|
"matchPackageNames": ["renovate/renovate"],
|
||||||
|
|||||||
@@ -1,22 +0,0 @@
|
|||||||
# SearXNG
|
|
||||||
|
|
||||||
Search frontend with a separate Valkey cache.
|
|
||||||
|
|
||||||
Kubernetes keeps the application settings in a ConfigMap and starts Valkey as a
|
|
||||||
StatefulSet. Set the secret from the example before exposing the search endpoint.
|
|
||||||
There is no active marker.
|
|
||||||
|
|
||||||
Compose expects local configuration under `core-config/`, which is ignored.
|
|
||||||
Prepare it before starting the stack; a container image alone does not supply
|
|
||||||
this lab's settings.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n searxng
|
|
||||||
kubectl get events -n searxng --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
@@ -1,18 +0,0 @@
|
|||||||
# Media stack
|
|
||||||
|
|
||||||
Docker services for playback, requests, library management, and downloads.
|
|
||||||
|
|
||||||
Compose runs Jellyfin, Jellyseerr, Sonarr, Radarr, Prowlarr, qBittorrent, and the
|
|
||||||
other services declared in the file. Kubernetes only routes to host endpoints;
|
|
||||||
update `k8s/routing/external-service.yaml` when the Docker host or ports change.
|
|
||||||
|
|
||||||
Prepare the paths, user/group IDs, and credentials from `.env.example`. Service
|
|
||||||
configuration and media/download directories are bind mounts. Preserve their
|
|
||||||
permissions when moving data, and keep the application databases with backups.
|
|
||||||
|
|
||||||
Review device mounts for hardware acceleration before starting on another host.
|
|
||||||
The Compose and Kubernetes routing files have no active markers, so automatic
|
|
||||||
deploys do not select this stack. Start the Compose project or apply its routing
|
|
||||||
resources manually when needed.
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
Whitespace-only changes.
Whitespace-only changes.
@@ -1,20 +0,0 @@
|
|||||||
# Termix
|
|
||||||
|
|
||||||
Terminal and SSH connection manager with persistent application data.
|
|
||||||
|
|
||||||
Kubernetes stores state in `termix-pvc`; Compose mounts `termix-data/`.
|
|
||||||
The application config and routes are committed separately under `k8s/`.
|
|
||||||
|
|
||||||
There is no active marker. Review access control and retain the application data
|
|
||||||
needed to recover saved connections before enabling it.
|
|
||||||
|
|
||||||
## Inspect
|
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```sh
|
|
||||||
kubectl get pods,svc,pvc -n termix
|
|
||||||
kubectl get events -n termix --sort-by=.metadata.creationTimestamp
|
|
||||||
```
|
|
||||||
|
|
||||||
See the [repository README](../README.md) for deployment selection.
|
|
||||||
Loaded 100 of 111 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user