Compare commits

..
Author SHA1 Message Date
renovate-bot Bot fc42670cfb chore(deps): update renovate/renovate docker tag to v44.111.4
deploy / validate (push) Skipped
renovate-ci / validate-renovate (push) Skipped
ci / validate (push) Successful in 2s
ci / build (push) Skipped
ci / lint-prettier (pull_request) Successful in 2s
ci / lint-ruff (pull_request) Successful in 0s
ci / lint-yaml (pull_request) Successful in 2s
renovate-ci / validate-renovate (pull_request) Successful in 13s
ci / lint-prettier (push) Successful in 4s
ci / lint-ruff (push) Successful in 1s
ci / lint-yaml (push) Successful in 2s
ci / lint-dockerfiles (push) Successful in 1s
ci / lint-dockerfiles (pull_request) Successful in 1s
ci / validate (pull_request) Successful in 1s
ci / build (pull_request) Skipped
2026-09-24 04:18:33 +00:00
439 changed files with 28429 additions and 11228 deletions

No files matched your search

-50
View File
@@ -1,50 +0,0 @@
# EDU ownership handoff
## Status
The EDU ownership handoff is complete. The homelab repository no longer owns
EDU workloads, images, routes, alerts, or deployment selection. The EDU
repository is the only deployment owner: [forust/edu-master](https://git.forust.xyz/forust/edu-master).
Homelab PRs #99 and #105 are merged. PR #105 removed the EDU subtree and its
build, deploy, rollback, verification, route-probe, and registry references.
It also added the serial image build matrix for the homelab services. This
handoff record is the only remaining EDU-specific file in homelab Git.
The dedicated workstation checkout is `/srv/edu-master`, at release
`4f2b2a0e37dc11ac2c75441a15076c178e219d37`. It contains `k8s/active`; root
`active` is absent. The old untracked `/srv/homelab/edu_master` checkout was
moved outside the homelab repository to
`/srv/edu-master-legacy-archive-20261007/edu_master`. Its private files remain
mode `0600` inside an archive directory with mode `0700`. The homelab deploy
checkout has no EDU marker or tracked EDU application/deployment files.
`AUTODEPLOY=false` remains in place for homelab deployment.
## Release evidence
EDU PR #4 merged after its review and CI checks. Main-push CI run 1652 passed
all validation and both image builds. Deploy run 1653 passed for the exact main
SHA above.
The workstation rollout completed for both Deployments. The deployment
verified `/health` and `/live` with HTTP 200, Redis AUTH, session TTL of 1058
seconds, a delivery backlog of zero, and all nine EDU vmalert rules with
matching expressions and healthy evaluation.
The images now run by digest:
- Session keeper: `sha256:998dea51aa3015fd9cabefb0f53b030157a650c3bef72e02fe84f17d5762613d`
- Webinar checker: `sha256:92f3c1fa2bb7f9b4680a9fc76a5b33dfbea8ef3dd9c6490ebc45876fd4c54461`
Redis StatefulSet was unchanged. PVC `redis-data-pvc` remains bound to PV
`pvc-a4f2a79a-363a-4c12-ae91-92cdfc2a0d2e` with capacity 1 GiB. The existing
runtime Secret and Fernet key were preserved during the handoff. Notification
delivery was verified before closeout, as confirmed by the operator. The
deployment did not record downtime.
The release rollback snapshot is
`/home/forust/.local/state/edu-master-deploy/20261007T180541Z-4f2b2a0e37dc11ac2c75441a15076c178e219d37`.
The handoff data snapshot remains at
`/home/forust/.local/state/edu-master-deploy/handoff-20261007T080838Z`.
Both snapshots are outside Git. Do not restore old Redis data unless recovery
requires it. Never delete or recreate the Redis PVC.
-64
View File
@@ -1,64 +0,0 @@
# CI and deployment
Gitea Actions validates changes, builds the repository's custom images, and can
deploy selected services to the workstation. CI and production deployment use
separate workflows. See the [runner and recovery guide](runner/README.md) for
installation, configuration, and operator commands.
## CI
`workflows/ci.yaml` runs Compose, workflow, shell, formatting, Python and unit
test, YAML, Dockerfile, and Kubernetes checks. Pull requests and non-main refs
use the unprivileged `homelab-pr` runner. Main-branch CI uses `homelab`. Tool
versions are pinned in `workflows/tool-versions.env`.
Compose CI checks every committed Compose file without requiring ignored `.env`
files. Kubernetes checks validate known schemas; unknown CRDs are skipped.
On main, CI plans builds for the three owned images: `error-pages`,
`forust-homepage`, and `xdfnx-homepage`. It builds changed inputs or reuses a
digest from a successful earlier main run. The successful build job publishes a
release artifact for the exact commit SHA. Pull requests do not publish images.
## Deployment gate
`workflows/deploy.yaml` starts a deployment after successful main CI when the
`AUTODEPLOY` Actions variable is `true`. Manual dispatch uses the same gate: the
requested `main` ref or commit must have successful main CI and its matching
release artifact. A manual dispatch does not bypass validation.
The workflow supports these modes:
- `changed`: select active services changed since the last successful deploy.
- `full`: select all active services; use this for the first baseline.
- `plan`: validate and show the selection without applying production resources.
`refresh_images=true` explicitly refreshes mutable third-party Compose tags.
## Selection and rollout
The active markers define automatic deployment. `<service>/active` selects a
standard Compose file; `<service>/k8s/active` selects Kubernetes resources. Helm
releases have their own markers in `workflows/deploy-lib.sh`. Service
dependencies are declared in `deploy-dependencies.json`. Removed resources are
reported for manual review; the workflow does not prune them automatically.
The workstation controller runs the checked source in a per-SHA worktree. It
validates configuration, applies Kubernetes and Compose changes in sequence,
verifies changed Kubernetes workloads, and checks public routes. A durable
systemd service continues the rollout if the Actions SSH client disconnects.
The workflow checks the exact CI release before it submits a deployment.
Kubernetes recovery uses captured workload revisions. It does not restore
ConfigMaps, Secrets, database schemas, or persistent data. Compose recovery is
manual and does not restore volume data or reverse migrations. Keep backups for
stateful services. The runner guide documents status, retry, logs, and recovery
commands.
## Settings
Configure `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`, and the verified
`DEPLOY_KNOWN_HOSTS` entry as Actions variables. Keep `DEPLOY_SSH_KEY`,
`REGISTRY_USERNAME`, and `REGISTRY_PASSWORD` in Actions secrets. The workstation
also needs its existing registry authentication. Set `AUTODEPLOY=false` until
automatic production deploys are intended.
-11
View File
@@ -1,11 +0,0 @@
# actionlint configuration. Passed explicitly from the ci workflow:
# actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
#
# The self-hosted act_runner registers custom labels that actionlint cannot know
# about, so declare them here instead of silencing the whole runner-label check.
self-hosted-runner:
labels:
- arch
- homelab
- homelab-pr
- prod
-3
View File
@@ -1,3 +0,0 @@
{
"postgres": ["authentik", "gitea", "immich", "n8n", "netbox", "netronome"]
}
-181
View File
@@ -1,181 +0,0 @@
# Homelab CI/CD
The native Gitea runners run on **vps**; production runs on **workstation**.
Main-branch checks and image builds use `homelab:host`. Pull request and
non-main checks use `homelab-pr:host` under a separate account without Docker
access. The `homelab-pr` runner is registered at User scope for `forust`, so
any repository under that account can schedule jobs that request this label.
Each runner accepts one job at a time; the build waits for every check to pass.
CI and deploy runs also show a summary with
the release SHA, image build or reuse results, deploy mode, selected services,
and image digests. Failed runs keep a summary of completed image builds, stage
results, apply results, and recorded Kubernetes recovery. The final deploy
summary is in the smoke job; earlier jobs show the state observed at that time.
Apply success is separate from health and recovery. Update the installed
workstation controller with `setup-workstation.sh` when no deploy is running.
No job images or Kubernetes credentials are needed on the VPS. Builds use one
pinned BuildKit helper container. CI and deploy are separate workflows.
## Runner installation
Install Docker Engine with Compose and Buildx, Git, Python 3.11+, Bash, curl,
GNU tar/xz, flock and systemd using the host's package manager. Keep the existing
Gitea runner 3.0.2 binary at `/usr/local/bin/gitea-runner`.
From this checkout on the VPS:
```sh
sudo bash .gitea/runner/setup-runner.sh
```
The installer reuses `/var/lib/gitea-runner/.runner` and the existing service.
For a new host, install the same runner binary and register as `gitea-runner`
using the registration token interactively, label `homelab:host`, and working
directory `/var/lib/gitea-runner`; then rerun the installer. Tokens never belong
in this repository or command-line examples.
Pinned tools live in the runner user's `~/.cache/homelab-ci`; CI repairs version
drift there. Installations are locked. Buildx uses only the `homelab-ci` builder,
pushes directly to the registry, and caps retained local cache at 1 GiB with a
2 GiB free-space target. This is not a hard limit on peak build disk usage.
Nothing runs `docker system prune`, removes unrelated images, or deletes volumes.
### Pull request runner
Install the unprivileged host runner on the VPS:
```sh
sudo bash .gitea/runner/setup-pr-runner.sh
```
Get a registration token from the user Actions runner settings. Run the
installer in a terminal. It asks for the token without echoing it, registers the
runner as `homelab-pr` with label `homelab-pr:host`, then enables the service.
The work directory is `/var/lib/gitea-pr-runner`. Confirm that Gitea lists the
runner as User scope before merging the workflow change. An unmatched label can
fall back to the default job image.
Renovate PR validation uses `pull_request_target`, which reads the workflow from
the base branch. It checks out the PR head only after runner selection and runs
that code on `homelab-pr`. Keep this workflow read-only and do not add secrets.
The PR runner has a separate home and tool cache. Do not add it to the `docker`
group or give it access to `/var/run/docker.sock`. It runs repository code from
pull requests, so keep its registration and permissions separate from the
trusted `homelab` runner. This separates users and host permissions, but both
runners still share the VPS kernel and network. Use a disposable VM if PRs from
untrusted external authors must be fully isolated.
## Workstation setup
As the existing SSH deploy user on workstation:
```sh
sudo loginctl enable-linger forust
bash .gitea/runner/setup-workstation.sh
```
The controller uses `/srv/homelab` as the persistent configuration tree and makes
a detached source worktree for each SHA. It never resets `/srv/homelab`, moves
local configuration, renames Compose projects, or changes volume names.
The installer records the current Kubernetes context and cluster UID in
`~/.config/homelab-deploy/environment`. Check these before installing.
Configure Gitea Actions Variables:
- `DEPLOY_HOST`, `DEPLOY_USER`, `DEPLOY_PORT`: the existing VPS-to-workstation SSH endpoint.
- `DEPLOY_KNOWN_HOSTS`: workstation's verified SSH host key entry for that endpoint.
- `AUTODEPLOY`: `false` initially; `true` enables deployment after successful main CI.
Keep `DEPLOY_SSH_KEY`, `REGISTRY_USERNAME` and `REGISTRY_PASSWORD` in Actions
Secrets. Legacy endpoint secrets remain accepted during migration. The Actions
token must have repository read and Actions read access for release downloads.
The deploy user's existing Docker registry authentication remains necessary.
## Releases and deployment
CI publishes `release-<full SHA>` as a Gitea artifact with all three owned image
digests and build input fingerprints. Unchanged images are reused only from a
successful main CI artifact, never from `:prod`. Expired artifacts cause CI to
rebuild images; they block deployment until CI is rerun.
Run deploy from main with `deploy_ref=main` or a checked SHA:
- `full`: required for the first baseline; reconcile all active components.
- `changed`: compare with the last fully successful production deploy.
- `plan`: validate configuration and show selection without changing production resources.
- `refresh_images=true`: explicitly refresh mutable third-party Compose tags.
The manual and automatic paths both require successful CI, a successful build
job and the exact SHA's release artifact. PRs cannot publish images or deploy.
Removed resources are reported and require explicit removal; no automatic prune.
Service dependencies are listed in `.gitea/deploy-dependencies.json`.
A workstation user systemd service holds the deploy lock across validation,
sequential apply, verification and smoke checks. SSH clients only submit/follow:
disconnecting or cancelling the Actions client does not kill production apply.
Retrying the same run ID does not start another apply. `ExecStopPost` recovers
interrupted runs before the unit finishes. Kubernetes rolls back to captured
revisions; configuration and persistent data are not reverted.
## Status and recovery
`--retry` repeats failed verification and smoke checks, never apply. Recovery
keeps a failed deploy out of the successful baseline, even after rollback.
On workstation (replace the numeric ID with Actions run ID and attempt):
```sh
python3 ~/.local/lib/homelab-deploy/controller.py status 123-1
python3 ~/.local/lib/homelab-deploy/controller.py recover 123-1 --retry
journalctl --user -u homelab-deploy@123-1
```
Runs live in `~/.local/state/homelab-deploy/runs`. Compose stores resolved configs
with restricted permissions; these may contain credentials and must never be
uploaded as CI artifacts. Stage logs print the exact manual recovery command
using `compose-before/<stack>.json`, the original project directory and project
name. Compose does not automatically roll back, and Nextcloud AIO's child
containers remain managed by AIO. Preserve its own backups for data recovery.
The controller retains twenty successful/planned runs and preserves failures.
Update the workstation dispatcher only when no deploy is running.
## Validation and migration rollback
```sh
python3 -m unittest discover -s tests -v
bash .gitea/tests/deploy-validation.sh
```
Test on a separate namespace before the initial production `full` run. Check a
failed rollout, interrupted SSH and repeated run ID, and verify that an isolated
service change does not upgrade unrelated Helm releases or Compose stacks.
To roll back the migration, disable autodeploy and finish or recover the remote
run first. Restore the runner config/unit from `.before-<timestamp>` backups,
reload systemd and restart the runner. Restore the prior workflows from Git.
Production data and persistent volumes stay where they were. Do not remove run
state or Compose recovery files until recovery is confirmed.
### Compose configuration recovery
Successful deploys save the complete resolved Compose configuration in
`~/.local/state/homelab-deploy/compose-configs/`. These files can contain secrets.
Keep them private and do not commit or upload them.
The next deploy uses this configuration for its recovery file, including old
commands, environment, mounts, ports, and removed services. The recovery command
uses `--remove-orphans` to remove services added by the failed deploy. It does
not restore volume data or reverse database migrations.
On the first run after this update, the controller can use the Compose file
from the previous successful run. If that file is absent, it reads the persistent
checkout and checks its service configuration hashes against existing containers.
A mismatch stops preflight. Restore the previous configuration before retrying.
Update the installed controller with `bash .gitea/runner/setup-workstation.sh`
from the reviewed checkout before using this change.
New namespaces are checked during preflight. Server validation of their resources
runs after namespace creation and before application resources are applied.
Plan mode does not create namespaces. A failed deferred check can leave an empty
namespace; inspect it before removing it.
-11
View File
@@ -1,11 +0,0 @@
[worker.oci]
gc = true
reservedSpace = "256MB"
maxUsedSpace = "1GB"
minFreeSpace = "2GB"
[[worker.oci.gcpolicy]]
reservedSpace = "256MB"
maxUsedSpace = "1GB"
minFreeSpace = "2GB"
all = true
-10
View File
@@ -1,10 +0,0 @@
runner:
file: /var/lib/gitea-runner/.runner
capacity: 1
timeout: 5h
labels:
- homelab:host
cache:
enabled: false
container:
docker_host: unix:///var/run/docker.sock
-18
View File
@@ -1,18 +0,0 @@
[Unit]
Description=Gitea Actions runner
After=network-online.target docker.service
Wants=network-online.target
[Service]
User=gitea-runner
Group=gitea-runner
SupplementaryGroups=docker
WorkingDirectory=/var/lib/gitea-runner
Environment=PATH=/var/lib/gitea-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-runner/config.yaml
Restart=on-failure
RestartSec=5
UMask=0077
[Install]
WantedBy=multi-user.target
-12
View File
@@ -1,12 +0,0 @@
[Unit]
Description=Homelab deploy %i
[Service]
Type=exec
EnvironmentFile=%h/.config/homelab-deploy/environment
ExecStart=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py execute %i
ExecStopPost=/usr/bin/python3 %h/.local/lib/homelab-deploy/controller.py recover %i
RuntimeMaxSec=5h
TimeoutStopSec=135min
KillMode=control-group
UMask=0077
-8
View File
@@ -1,8 +0,0 @@
runner:
file: /var/lib/gitea-pr-runner/.runner
capacity: 1
timeout: 5h
labels:
- homelab-pr:host
cache:
enabled: false
-27
View File
@@ -1,27 +0,0 @@
[Unit]
Description=Gitea Actions untrusted pull request runner
After=network-online.target
Wants=network-online.target
[Service]
User=gitea-pr-runner
Group=gitea-pr-runner
WorkingDirectory=/var/lib/gitea-pr-runner
Environment=HOME=/var/lib/gitea-pr-runner
Environment=PATH=/var/lib/gitea-pr-runner/.cache/homelab-ci/bin:/usr/local/bin:/usr/bin:/bin
ExecStart=/usr/local/bin/gitea-runner daemon --config /etc/gitea-pr-runner/config.yaml
Restart=on-failure
RestartSec=5
NoNewPrivileges=yes
PrivateTmp=yes
ProtectSystem=full
ProtectHome=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
RestrictSUIDSGID=yes
LockPersonality=yes
UMask=0077
[Install]
WantedBy=multi-user.target
-56
View File
@@ -1,56 +0,0 @@
#!/usr/bin/env bash
# Install a native runner for untrusted PR jobs without Docker access.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
for tool in cp cut date getent id install runuser systemctl useradd; do
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
done
command -v /usr/local/bin/gitea-runner >/dev/null || {
echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2
exit 1
}
id gitea-pr-runner >/dev/null 2>&1 || \
useradd --system --create-home --home-dir /var/lib/gitea-pr-runner --shell /usr/bin/bash gitea-pr-runner
runner_home="$(getent passwd gitea-pr-runner | cut -d: -f6)"
[ "$runner_home" = /var/lib/gitea-pr-runner ] || {
echo 'Unexpected PR runner home; inspect the existing service first' >&2
exit 1
}
case " $(id -nG gitea-pr-runner) " in
*' docker '*)
echo 'The PR runner account must not belong to the docker group' >&2
exit 1
;;
esac
install -d -m 0755 /etc/gitea-pr-runner
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
for existing in /etc/gitea-pr-runner/config.yaml /etc/systemd/system/gitea-pr-runner.service; do
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
done
install -m 0644 "$here/pr-config.yaml" /etc/gitea-pr-runner/config.yaml
install -m 0644 "$here/pr-runner.service" /etc/systemd/system/gitea-pr-runner.service
if [ ! -f /var/lib/gitea-pr-runner/.runner ]; then
read -r -s -p 'Enter the Gitea repository runner registration token: ' runner_token
printf '\n'
[ -n "$runner_token" ] || { echo 'Runner token is required' >&2; exit 1; }
export GITEA_RUNNER_REGISTRATION_TOKEN="$runner_token"
unset runner_token
runuser --preserve-environment -u gitea-pr-runner -- \
/usr/local/bin/gitea-runner register \
--config /etc/gitea-pr-runner/config.yaml \
--instance https://gitea.forust.xyz \
--name homelab-pr \
--labels homelab-pr:host \
--no-interactive
unset GITEA_RUNNER_REGISTRATION_TOKEN
fi
chmod 0600 /var/lib/gitea-pr-runner/.runner
systemctl daemon-reload
systemctl enable --now gitea-pr-runner.service
systemctl restart gitea-pr-runner.service
echo "PR runner ready. Configuration backups: *.before-$stamp"
-48
View File
@@ -1,48 +0,0 @@
#!/usr/bin/env bash
# Native host runner, with pinned user-space tools and no extra CI images.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
[ "$(id -u)" -eq 0 ] || { echo 'Run with sudo on the runner host' >&2; exit 1; }
for tool in docker curl python3 git tar xz flock runuser systemctl; do
command -v "$tool" >/dev/null || { echo "Install missing prerequisite: $tool" >&2; exit 1; }
done
docker info >/dev/null
docker compose version >/dev/null
docker buildx version >/dev/null
id gitea-runner >/dev/null 2>&1 || useradd --system --create-home --home-dir /var/lib/gitea-runner --shell /usr/bin/bash gitea-runner
# Reuse the established service account and runner registration.
runner_home="$(getent passwd gitea-runner | cut -d: -f6)"
[ "$runner_home" = /var/lib/gitea-runner ] || { echo 'Unexpected runner home; inspect the existing service first' >&2; exit 1; }
runuser -u gitea-runner -- docker info >/dev/null || { echo "The runner user needs access to Docker before setup" >&2; exit 1; }
command -v gitea-runner >/dev/null || { echo 'Install gitea-runner 3.0.2 at /usr/local/bin/gitea-runner first' >&2; exit 1; }
mkdir -p /etc/gitea-runner
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
for existing in /etc/gitea-runner/config.yaml /etc/systemd/system/gitea-runner.service; do
[ ! -f "$existing" ] || cp -p "$existing" "$existing.before-$stamp"
done
scratch="$(mktemp -d)"
trap 'rm -rf "$scratch"' EXIT
chmod 755 "$scratch"
install -m 0644 "$here/../workflows/install-ci-tools.sh" "$here/../workflows/tool-versions.env" "$scratch/"
runuser -u gitea-runner -- bash "$scratch/install-ci-tools.sh"
install -m 0644 "$here/config.yaml" /etc/gitea-runner/config.yaml
python3 - <<'PYLABELS'
import json
from pathlib import Path
registration = Path('/var/lib/gitea-runner/.runner')
if registration.exists():
labels = json.loads(registration.read_text()).get('labels', [])
labels = [label for label in labels if isinstance(label, str) and label.split(':')[0] != 'homelab']
labels.append('homelab:host')
config = Path('/etc/gitea-runner/config.yaml')
config.write_text(config.read_text().replace(' - homelab:host', '\n'.join(' - ' + json.dumps(label) for label in labels)))
PYLABELS
install -m 0644 "$here/gitea-runner.service" /etc/systemd/system/gitea-runner.service
if [ ! -f /var/lib/gitea-runner/.runner ]; then
echo 'Register once as gitea-runner with homelab:host before starting the service.'
exit 0
fi
systemctl daemon-reload
systemctl enable --now gitea-runner.service
systemctl restart gitea-runner.service
echo "Runner ready. Configuration backups: *.before-$stamp"
-32
View File
@@ -1,32 +0,0 @@
#!/usr/bin/env bash
# Run as the existing deploy user on workstation. Never resets the working tree.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
repo="${HOMELAB_REPO:-/srv/homelab}"
for tool in python3 git kubectl helm docker flock timeout; do
command -v "$tool" >/dev/null || { echo "Install missing dependency: $tool" >&2; exit 1; }
done
[ -d "$repo/.git" ] || { echo "Missing deploy checkout: $repo" >&2; exit 1; }
[[ "$repo" =~ ^/[A-Za-z0-9_./-]+$ ]] || { echo 'Deploy path must be absolute and contain no whitespace' >&2; exit 1; }
if [ "$(loginctl show-user "$USER" -p Linger --value)" != yes ]; then
echo "Run once: sudo loginctl enable-linger $USER" >&2
exit 1
fi
config="${XDG_CONFIG_HOME:-$HOME/.config}/homelab-deploy"
mkdir -p "$config" "$HOME/.local/lib/homelab-deploy" "$HOME/.config/systemd/user"
chmod 700 "$config"
if [ ! -f "$config/environment" ]; then
context="$(kubectl config current-context)"
cluster_uid="$(kubectl get namespace kube-system -o jsonpath='{.metadata.uid}')"
printf 'HOMELAB_REPO=%s\nKUBE_CONTEXT=%s\nEXPECTED_CLUSTER_UID=%s\n' "$repo" "$context" "$cluster_uid" >"$config/environment"
chmod 600 "$config/environment"
fi
# Do not replace a dispatcher while an existing deploy uses it.
if systemctl --user list-units 'homelab-deploy@*' --state=running --no-legend | grep -q .; then
echo 'An existing deploy is running; wait before updating the controller' >&2
exit 1
fi
install -m 0755 "$here/../workflows/deploy-controller.py" "$HOME/.local/lib/homelab-deploy/controller.py"
install -m 0644 "$here/homelab-deploy@.service" "$HOME/.config/systemd/user/homelab-deploy@.service"
systemctl --user daemon-reload
echo 'Controller ready. Run a checked main SHA in full mode for the initial baseline.'
-156
View File
@@ -1,156 +0,0 @@
#!/usr/bin/env bash
# Local regressions only: kubectl is mocked and Docker is used for config parsing.
set -euo pipefail
repo="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
scratch="$(mktemp -d)"
trap 'rm -rf "$scratch"' EXIT
mkdir -p "$scratch/repo/app" "$scratch/repo/postgres" "$scratch/repo/netbird" "$scratch/repo/renovate"
git -C "$scratch/repo" init -q
for file in app/compose.yaml postgres/shared-compose.yaml netbird/client.compose.yaml renovate/renovate-compose.yaml; do
touch "$scratch/repo/$file"
done
git -C "$scratch/repo" add .
# shellcheck source=../workflows/compose-lint.sh
source "$repo/.gitea/workflows/compose-lint.sh"
actual="$(cd "$scratch/repo" && compose_files)"
expected=$'app/compose.yaml\nnetbird/client.compose.yaml\npostgres/shared-compose.yaml\nrenovate/renovate-compose.yaml'
[ "$actual" = "$expected" ] || { echo 'Compose discovery missed a file' >&2; exit 1; }
cat >"$scratch/compose.yaml" <<'YAML'
services:
example:
image: busybox:1.37.0
environment:
REQUIRED: ${HOMELAB_TEST_REQUIRED:?required for this regression}
YAML
unset HOMELAB_TEST_REQUIRED
if validate_compose_file "$scratch/compose.yaml" >"$scratch/config.log" 2>&1; then
echo 'Full Compose validation accepted a missing variable' >&2
exit 1
fi
grep -q 'required for this regression' "$scratch/config.log"
HOMELAB_TEST_REQUIRED=present validate_compose_file "$scratch/compose.yaml"
cat >"$scratch/resources.json" <<'JSON'
{"kind":"List","items":[
{"kind":"Deployment","metadata":{"namespace":"app"},"spec":{"template":{"spec":{
"containers":[{"envFrom":[{"secretRef":{"name":"credentials"}},{"secretRef":{"name":"optional","optional":true}}],"env":[{"valueFrom":{"secretKeyRef":{"name":"credentials","key":"password"}}}]}],
"initContainers":[{"envFrom":[{"secretRef":{"name":"init"}}]}],
"imagePullSecrets":[{"name":"registry"}],
"volumes":[{"secret":{"secretName":"mounted"}},{"projected":{"sources":[{"secret":{"name":"projected"}},{"secret":{"name":"optional-projected","optional":true}}]}}]
}}}},
{"kind":"CronJob","metadata":{},"spec":{"jobTemplate":{"spec":{"template":{"spec":{"containers":[{"envFrom":[{"secretRef":{"name":"cron"}}]}]}}}}}},
{"kind":"IngressRoute","metadata":{"namespace":"app"},"spec":{"tls":{"secretName":"controller-issued-tls"}}}
]}
JSON
actual="$(jq -r -f "$repo/.gitea/workflows/secret-references.jq" "$scratch/resources.json" | sort)"
expected=$'app credentials\napp init\napp mounted\napp projected\napp registry\ndefault cron'
[ "$actual" = "$expected" ] || { echo "Unexpected Secret references: $actual" >&2; exit 1; }
REPO="$repo"
# shellcheck source=../workflows/deploy-lib.sh
source "$repo/.gitea/workflows/deploy-lib.sh"
K8S_MANIFESTS=("$scratch/resources.json")
KUSTOMIZE_APPS=()
# No live cluster access. Reject credentials in app even if they exist elsewhere.
kubectl() {
case "$1" in
create) cat "$scratch/resources.json" ;;
get)
if [ "$3" = credentials ] && [ "$5" = app ]; then
return 1
fi
return 0
;;
*) echo "Unexpected kubectl invocation: $*" >&2; return 1 ;;
esac
}
if check_referenced_secrets >"$scratch/secrets.log"; then
echo 'Namespace-scoped Secret check accepted a missing Secret' >&2
exit 1
fi
grep -q 'MISSING OR UNREADABLE: app/credentials' "$scratch/secrets.log"
# API/rendering errors must not produce an empty reference list and pass.
kubectl() { return 1; }
if ! skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
echo 'VMAgent preflight did not skip an uninstalled CRD' >&2
exit 1
fi
kubectl() { return 0; }
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/vmagent.yaml"; then
echo 'VMAgent preflight skipped an installed CRD' >&2
exit 1
fi
if skip_uninstalled_vmagent_crd "$REPO/prometheus-stack/k8s/victoria.yaml"; then
echo 'VMAgent preflight skipped an unrelated manifest' >&2
exit 1
fi
kubectl() { return 1; }
if check_referenced_secrets >"$scratch/secrets.log"; then
echo 'Secret check accepted a failed manifest render' >&2
exit 1
fi
# New declared namespaces defer only their own resources during preflight.
render_selected_resources() {
cat <<'JSON'
{"apiVersion":"v1","kind":"List","items":[
{"apiVersion":"v1","kind":"Namespace","metadata":{"name":"new"}},
{"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"new-config","namespace":"new"}},
{"apiVersion":"v1","kind":"ConfigMap","metadata":{"name":"existing-config","namespace":"default"}}
]}
JSON
}
kubectl() {
case "$1" in
get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}}]}' ;;
apply) cat >"$scratch/server-input.json" ;;
*) return 1 ;;
esac
}
validate_server_resources true
jq -e '.items | length == 2 and all(.metadata.name != "new-config")' "$scratch/server-input.json" >/dev/null
if validate_server_resources false 2>"$scratch/deferred.log"; then
echo 'Post-namespace validation accepted a missing namespace' >&2
exit 1
fi
kubectl() {
case "$1" in
get) printf '%s\n' '{"items":[{"metadata":{"name":"default"}},{"metadata":{"name":"new"}}]}' ;;
apply) cat >"$scratch/server-input.json" ;;
*) return 1 ;;
esac
}
validate_server_resources false
jq -e '.items | length == 3' "$scratch/server-input.json" >/dev/null
render_selected_resources() {
printf '%s\n' '{"items":[{"kind":"ConfigMap","metadata":{"name":"bad","namespace":"undeclared"}}]}'
}
if validate_server_resources true 2>"$scratch/undeclared.log"; then
echo 'Preflight accepted an undeclared missing namespace' >&2
exit 1
fi
# Count services, not characters in the newline-separated service names.
compose() {
case "$*" in
*'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{},"init":{"restart":"no"}}}' ;;
*'ps --status running --services') printf '%s\n' headscale headplane web ;;
*) return 1 ;;
esac
}
verify_compose_stack example.yaml >"$scratch/compose-count.log"
grep -qF 'all 3 service(s) running' "$scratch/compose-count.log"
compose() {
case "$*" in
*'config --format json') printf '%s\n' '{"services":{"headscale":{},"headplane":{},"web":{}}}' ;;
*'ps --status running --services') printf '%s\n' headscale headplane ;;
*) return 0 ;;
esac
}
if verify_compose_stack example.yaml >"$scratch/compose-missing.log"; then
echo 'Compose verification accepted a missing service' >&2
exit 1
fi
grep -qF 'NOT RUNNING: web' "$scratch/compose-missing.log"
printf '%s\n' 'Deploy validation regressions passed.'
+256 -427
View File
@@ -1,169 +1,29 @@
name: ci name: ci
"on":
on:
push: push:
branches: branches:
- main - "**"
pull_request: null pull_request:
workflow_dispatch: null workflow_dispatch:
permissions:
contents: read
actions: read
concurrency: concurrency:
group: ci-${{ github.ref }} group: ci-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }} cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
env:
REGISTRY: gcr.forust.xyz
jobs: jobs:
compose: lint-prettier:
name: Compose runs-on: [self-hosted, linux, arch, homelab]
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
id: source
- name: Validate Compose files
shell: bash
run: |
set -euo pipefail
source .gitea/workflows/compose-lint.sh
mapfile -t safe_flags < <(compose_safe_flags)
echo "docker compose config ${safe_flags[*]-}"
mapfile -t files < <(compose_files)
if [ "${#files[@]}" -eq 0 ]; then
echo "No Compose files found."
exit 0
fi
failed=0
for f in "${files[@]}"; do
if ! out="$(validate_compose_file "$f" ${safe_flags[@]+"${safe_flags[@]}"} 2>&1)"; then
failed=1
echo "::error file=${f}::$(printf '%s' "$out" | head -1)"
fi
done
if [ "$failed" -ne 0 ]; then
echo "Compose validation failed."
exit 1
fi
echo "checked ${#files[@]} Compose file(s)"
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Compose
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.source.conclusion == 'failure'
&& 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
workflows:
name: Workflows
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh actionlint shellcheck)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint Gitea Actions workflows with actionlint
shell: bash
run: |
set -euo pipefail
actionlint -config-file .gitea/actionlint.yaml -color .gitea/workflows/*.yaml
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Workflows
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
shell:
name: Shell
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh shellcheck jq)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint shell scripts with ShellCheck
shell: bash
run: |
set -euo pipefail
mapfile -t scripts < <(
git ls-files '*.sh' ':(glob)**/*.bash'
)
if [ "${#scripts[@]}" -eq 0 ]; then
echo "No shell scripts found."
exit 0
fi
shellcheck --external-sources --source-path=SCRIPTDIR --severity=style "${scripts[@]}"
bash .gitea/tests/deploy-validation.sh
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Shell
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
formatting:
name: Formatting
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh prettier)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Check formatting with Prettier - name: Check formatting with Prettier
shell: bash shell: bash
run: | run: |
set -euo pipefail
mapfile -t prettier_files < <( mapfile -t prettier_files < <(
git ls-files \ git ls-files \
| grep -E '\.(md|json|ya?ml|html|css)$' \ | grep -E '\.(md|json|ya?ml|html|css)$' \
@@ -176,80 +36,27 @@ jobs:
fi fi
prettier --check --ignore-unknown "${prettier_files[@]}" prettier --check --ignore-unknown "${prettier_files[@]}"
id: check
- name: Write the job result lint-ruff:
if: always() runs-on: [self-hosted, linux, arch, homelab]
env:
SUMMARY_CHECK: Formatting
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
python:
name: Python and tests
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
id: source
- name: Prepare pinned tools - name: Lint Python with Ruff
shell: bash shell: bash
run: | run: |
set -euo pipefail ruff check .
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh ruff jq)"
echo "$tools_dir" >> "$GITHUB_PATH" lint-yaml:
id: tools runs-on: [self-hosted, linux, arch, homelab]
- name: Lint and format-check Python with Ruff
shell: bash
run: |
set -euo pipefail
ruff check . .gitea/workflows
ruff format --check . .gitea/workflows
python3 -m unittest discover -s tests -v
id: check
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Python and tests
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
yaml:
name: YAML
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh yamllint)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint YAML syntax - name: Lint YAML syntax
shell: bash shell: bash
run: | run: |
set -euo pipefail
mapfile -t yaml_files < <( mapfile -t yaml_files < <(
git ls-files '*.yaml' '*.yml' \ git ls-files '*.yaml' '*.yml' \
':!node_modules/**' \ ':!node_modules/**' \
@@ -262,42 +69,16 @@ jobs:
fi fi
yamllint -c .yamllint "${yaml_files[@]}" yamllint -c .yamllint "${yaml_files[@]}"
id: check
- name: Write the job result lint-dockerfiles:
if: always() runs-on: [self-hosted, linux, arch, homelab]
env:
SUMMARY_CHECK: YAML
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
dockerfiles:
name: Dockerfiles
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh hadolint)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Lint Dockerfiles - name: Lint Dockerfiles
shell: bash shell: bash
run: | run: |
set -euo pipefail
mapfile -t dockerfiles < <( mapfile -t dockerfiles < <(
git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*' git ls-files ':(glob)**/Dockerfile' ':(glob)**/Dockerfile.*'
) )
@@ -308,42 +89,16 @@ jobs:
fi fi
hadolint -c .hadolint.yaml "${dockerfiles[@]}" hadolint -c .hadolint.yaml "${dockerfiles[@]}"
id: check
- name: Write the job result validate:
if: always() runs-on: [self-hosted, linux, arch, homelab]
env:
SUMMARY_CHECK: Dockerfiles
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
kubernetes:
name: Kubernetes
runs-on: ${{ github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }}
timeout-minutes: 15
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
id: source
- name: Prepare pinned tools
shell: bash
run: |
set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform)"
echo "$tools_dir" >> "$GITHUB_PATH"
id: tools
- name: Validate Kubernetes manifests against JSON schemas
shell: bash
run: |
set -euo pipefail
- name: Validate Kubernetes manifests
shell: bash
run: |
mapfile -t manifests < <( mapfile -t manifests < <(
git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \ git ls-files ':(glob)**/k8s/**/*.yaml' ':(glob)**/k8s/**/*.yml' \
| grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$' | grep -Ev '(^|/)(kustomization\.ya?ml|.*\.example\.ya?ml|.*values\.ya?ml|patch-.*\.ya?ml)$'
@@ -359,162 +114,236 @@ jobs:
-ignore-missing-schemas \ -ignore-missing-schemas \
-summary \ -summary \
"${manifests[@]}" "${manifests[@]}"
id: check
- name: Write the job result build:
if: always() needs: [lint-prettier, lint-ruff, lint-yaml, lint-dockerfiles, validate]
env: if: github.event_name != 'pull_request' && (github.ref_name == 'main' || github.ref_name == 'dev')
SUMMARY_CHECK: Kubernetes runs-on: [self-hosted, linux, arch, homelab]
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP:
${{ steps.check.conclusion == 'failure' && 'Check or image build' || steps.tools.conclusion == 'failure'
&& 'Tool setup' || steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi
image-plan:
needs: [compose, workflows, shell, formatting, python, yaml, dockerfiles, kubernetes]
if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main'
runs-on: homelab
timeout-minutes: 10
outputs: outputs:
matrix: ${{ steps.plan.outputs.matrix }} services: ${{ steps.services.outputs.services }}
steps: steps:
- name: Checkout repository - name: Checkout repository
id: source uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
with: with:
fetch-depth: 0 fetch-depth: 0
- name: Detect build inputs against successful CI
id: plan
env:
GITEA_TOKEN: ${{ github.token }}
run: python3 .gitea/workflows/release.py prepare --output build-plan.json
- name: Store the image plan
id: artifact
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: build-plan
path: build-plan.json
if-no-files-found: error
retention-days: 30
- name: Write the plan result - name: Detect changed docker-built services
if: always() id: services
env:
SUMMARY_CHECK: Image plan
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP: >-
${{ steps.plan.conclusion == 'failure' && 'Build input detection' ||
steps.artifact.conclusion == 'failure' && 'Plan upload' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash shell: bash
run: | run: |
if [ -f .gitea/workflows/release.py ]; then base="${{ github.event.before }}"
python3 .gitea/workflows/release.py check-summary if [ -z "$base" ] || [ "$base" = "0000000000000000000000000000000000000000" ]; then
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then base="$(git rev-list --max-parents=0 HEAD)"
printf '## Image plan\n\nResult: %s\n' "$SUMMARY_RESULT" >>"$GITHUB_STEP_SUMMARY" || true
fi fi
images: mapfile -t changed_files < <(git diff --name-only "$base" "${GITHUB_SHA}")
name: Image (${{ matrix.name }})
needs: [image-plan] services=()
if: needs.image-plan.result == 'success'
runs-on: homelab add_service() {
timeout-minutes: 60 local name="$1"
strategy: local seen=0
max-parallel: 1 for existing in "${services[@]}"; do
fail-fast: false if [ "$existing" = "$name" ]; then
matrix: ${{ fromJSON(needs.image-plan.outputs.matrix || '{"include":[{"name":"inactive"}]}') }} seen=1
steps: break
- name: Checkout repository fi
id: source done
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 if [ "$seen" -eq 0 ]; then
- name: Download the checked image plan services+=("$name")
id: inputs fi
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0 }
with:
name: build-plan for file in "${changed_files[@]}"; do
- name: Build or reuse this image case "$file" in
id: check dtek_notif/*)
env: add_service dtek_notif
IMAGE_NAME: ${{ matrix.name }} ;;
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }} errorpages/*)
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }} add_service errorpages
run: python3 .gitea/workflows/release.py image --image "$IMAGE_NAME" --output image.json ;;
- name: Store the image result userbot/*)
id: artifact add_service userbot
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2 ;;
with: homepages/*)
name: image-${{ matrix.name }} add_service homepages
path: image.json ;;
if-no-files-found: error edu_master/phpsessid-bot/*|edu_master/webinar-checker/*|edu_master/compose.yaml)
retention-days: 30 add_service edu_master
- name: Write the job result ;;
if: always() esac
env: done
SUMMARY_CHECK: Image (${{ matrix.name }})
SUMMARY_RESULT: ${{ job.status }} if [ "${#services[@]}" -eq 0 ]; then
SUMMARY_FAILED_STEP: >- echo "No docker-built services changed."
${{ steps.check.conclusion == 'failure' && 'Build or tag images' || echo "services=" >> "$GITHUB_OUTPUT"
steps.artifact.conclusion == 'failure' && 'Artifact upload' || exit 0
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash
run: |
if [ -f .gitea/workflows/release.py ]; then
python3 .gitea/workflows/release.py check-summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi fi
# Retain the build job name required by the immutable release deployment gate. printf '%s\n' "${services[@]}" | tee /tmp/services.txt
build: echo "services=$(paste -sd, /tmp/services.txt)" >> "$GITHUB_OUTPUT"
needs: [image-plan, images]
runs-on: homelab - name: Log in to registry
timeout-minutes: 15 if: steps.services.outputs.services != ''
steps:
- name: Checkout repository
id: source
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
- name: Download all image results
id: inputs
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
path: artifacts
- name: Pin SHA tags and write the complete release
id: check
env:
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: >-
python3 .gitea/workflows/release.py finalize
--plan artifacts/build-plan/build-plan.json
- name: Store commit release
id: artifact
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: release-${{ github.sha }}
path: release.json
if-no-files-found: error
retention-days: 30
- name: Write the job result
if: always()
env:
SUMMARY_CHECK: Image release and SHA tags
SUMMARY_RESULT: ${{ job.status }}
SUMMARY_FAILED_STEP: >-
${{ steps.check.conclusion == 'failure' && 'Build or tag images' ||
steps.artifact.conclusion == 'failure' && 'Artifact upload' ||
steps.inputs.conclusion == 'failure' && 'Artifact download' ||
steps.source.conclusion == 'failure' && 'Source checkout' || '' }}
shell: bash shell: bash
run: | run: |
if [ -f .gitea/workflows/release.py ]; then echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login "${REGISTRY}" \
python3 .gitea/workflows/release.py check-summary -u "${{ secrets.REGISTRY_USERNAME }}" \
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then --password-stdin
printf '## %s\n\n- Result: **%s**\n- Failed step: %s\n' "$SUMMARY_CHECK" "$SUMMARY_RESULT" "$SUMMARY_FAILED_STEP" >>"$GITHUB_STEP_SUMMARY" || true
fi - name: Build and push changed images
if: steps.services.outputs.services != ''
shell: bash
run: |
IFS=, read -r -a services <<< "${{ steps.services.outputs.services }}"
for service in "${services[@]}"; do
case "$service" in
dtek_notif)
image="${REGISTRY}/forust/dtek-notif"
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" dtek_notif
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
;;
errorpages)
image="${REGISTRY}/forust/error-pages"
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" errorpages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
;;
userbot)
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
for target in runtime panel; do
case "$target" in
runtime)
context="userbot"
image="${REGISTRY}/forust/userbot"
;;
panel)
context="userbot/panel"
image="${REGISTRY}/forust/userbot-panel"
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" "$context"
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
homepages)
for service in forust xdfnx; do
case "$service" in
forust)
image="${REGISTRY}/forust/forust-homepage"
;;
xdfnx)
image="${REGISTRY}/forust/xdfnx-homepage"
;;
esac
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" -f "homepages/Dockerfile.${service}" homepages
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
edu_master)
for service in session-keeper webinar-checker; do
case "$service" in
session-keeper)
context="edu_master/phpsessid-bot"
image="${REGISTRY}/forust/session-keeper"
;;
webinar-checker)
context="edu_master/webinar-checker"
image="${REGISTRY}/forust/webinar-checker"
;;
esac
tags=("latest")
case "${GITHUB_REF_NAME}" in
main)
tags+=("main" "prod")
;;
dev)
tags+=("dev")
;;
esac
build_args=()
for tag in "${tags[@]}"; do
build_args+=(-t "${image}:${tag}")
done
docker build \
--cache-from "type=registry,ref=${image}:buildcache" \
--cache-to "type=registry,ref=${image}:buildcache,mode=max" \
"${build_args[@]}" "$context"
for tag in "${tags[@]}"; do
docker push "${image}:${tag}"
done
done
;;
esac
done
-45
View File
@@ -1,45 +0,0 @@
#!/usr/bin/env bash
# Shared helpers for validating Compose files. Sourced both by steps in
# .gitea/workflows/ci.yaml and by deploy-lib.sh on the workstation.
#
# Two levels of checking, matching how the repo is structured:
#
# general every committed Compose file, active or not. Pure structure check:
# no ${VAR} interpolation, no .env lookup, no bind-mount path
# resolution. Disabled stacks deliberately have no .env in the repo
# and no values on the CI runner, so a full `config` run would fail on
# their `${VAR:?}` guards for reasons that have nothing to do with the
# change under review.
#
# full active stacks only, with interpolation and env-file resolution, so
# required variables and referenced files are actually resolved. Needs
# the gitignored .env files, so this only runs in the deploy workflow
# on the workstation.
#
# This file is meant to be sourced, not executed.
# All committed Compose files, including the ones deploy never starts.
compose_files() {
git ls-files \
'*compose.yaml' '*compose.yml'
}
# Prints the flags that turn `docker compose config` into the general check.
# Probed rather than hardcoded so an older Compose without --no-env-resolution
# still gets the flags it does support.
compose_safe_flags() {
local help flag
help="$(docker compose config --help 2>/dev/null || true)"
for flag in --no-interpolate --no-env-resolution --no-path-resolution; do
if printf '%s' "$help" | grep -q -- "$flag"; then
printf '%s\n' "$flag"
fi
done
}
# validate_compose_file <file> [extra docker compose config flags...]
validate_compose_file() {
local file="$1"
shift
docker compose -f "$file" config --quiet "$@"
}
-164
View File
@@ -1,164 +0,0 @@
#!/usr/bin/env python3
"""Resolve Compose images without changing project names or local bind paths."""
import json
import os
import re
import subprocess
import sys
from pathlib import Path
def output(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
def resolve(reference):
if '@sha256:' in reference:
return reference
descriptor = json.loads(
output('docker', 'buildx', 'imagetools', 'inspect', reference, '--format', '{{json .Manifest}}')
)
digest = descriptor['digest']
if not re.fullmatch(r'sha256:[0-9a-f]{64}', digest):
raise ValueError(f'Invalid registry digest for {reference}')
# Strip tag only from the final path segment (registry ports are preserved).
repository = reference.rsplit('/', 1)
repository[-1] = repository[-1].split(':')[0]
return '/'.join(repository) + '@' + digest
def prepare(source_file):
config_repo = Path(os.environ['CONFIG_REPO'])
source_repo = Path(os.environ['REPO'])
directory = Path(os.environ['RUN_DIR'])
relative = source_file.relative_to(source_repo)
project_dir = config_repo / relative.parent
base = ['docker', 'compose', '--project-directory', str(project_dir), '-f', str(source_file)]
config = json.loads(output(*base, 'config', '--format', 'json', cwd=config_repo))
project = config['name']
previous_file = directory / 'previous.json'
previous = json.loads(previous_file.read_text()) if previous_file.exists() else {}
images_file = directory / 'compose-images.json'
locks = json.loads(images_file.read_text()) if images_file.exists() else previous.get('compose-images', {})
release = json.loads((directory / 'release.json').read_text())
state = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
baseline = state / 'compose-configs' / f'{relative.parent.name}.json'
if not baseline.exists() and re.fullmatch(r'[0-9]+-[0-9]+', previous.get('run_id', '')):
baseline = state / 'runs' / previous['run_id'] / 'compose' / baseline.name
bootstrap = not baseline.exists()
if not bootstrap:
before = json.loads(baseline.read_text())
else:
# Bootstrap from the persistent configuration, never from the new source.
persistent_file = config_repo / relative
if persistent_file.exists():
before = json.loads(
output(
'docker',
'compose',
'--project-directory',
str(project_dir),
'-f',
str(persistent_file),
'config',
'--format',
'json',
cwd=config_repo,
)
)
elif output('docker', 'ps', '-aq', '--filter', f'label=com.docker.compose.project={project}'):
raise ValueError(f'{project}: no previous Compose configuration; restore it before deploy')
else:
before = {'name': project, 'services': {}}
if before['name'] != project:
raise ValueError('Compose project name changed; manual migration is required')
for service, settings in config['services'].items():
reference = settings.get('image')
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
if not reference or settings.get('build'):
raise ValueError(f'{project}/{service}: Compose deploy requires a published image')
image_repo = reference.split('@')[0].rsplit('/', 1)
image_repo[-1] = image_repo[-1].split(':')[0]
image_repo = '/'.join(image_repo)
# Nextcloud AIO validates the mastercontainer image reference and rejects
# a digest. Keep its configured tag so AIO can start and manage its stack.
if nextcloud_aio_master:
pinned = reference
elif image_repo in release['images']:
pinned = image_repo + '@' + release['images'][image_repo]
elif os.environ.get('REFRESH_IMAGES') != 'true' and reference in locks:
pinned = locks[reference]
else:
pinned = resolve(reference)
settings['image'] = pinned
locks[reference] = pinned
for service, settings in before['services'].items():
reference = settings['image']
image_repo = reference.split('@')[0].rsplit('/', 1)
image_repo[-1] = image_repo[-1].split(':')[0]
image_repo = '/'.join(image_repo)
nextcloud_aio_master = project == 'nextcloud' and service == 'nextcloud-aio-mastercontainer'
# Capture what is running, not the current value of its mutable tag.
ids = output(
'docker',
'ps',
'-aq',
'--filter',
f'label=com.docker.compose.project={project}',
'--filter',
f'label=com.docker.compose.service={service}',
).splitlines()
actual = set()
if bootstrap and ids:
expected_hash = output(
'docker',
'compose',
'--project-directory',
str(project_dir),
'-f',
str(persistent_file),
'config',
'--hash',
service,
cwd=config_repo,
).split()[-1]
for container in ids:
running_hash = output(
'docker',
'inspect',
container,
'--format',
'{{ index .Config.Labels "com.docker.compose.config-hash" }}',
)
if running_hash != expected_hash:
raise ValueError(
f'{project}/{service}: persistent config differs from running config; restore the previous config'
)
for container in ids:
image_id = output('docker', 'inspect', container, '--format', '{{.Image}}')
digests = json.loads(output('docker', 'image', 'inspect', image_id, '--format', '{{json .RepoDigests}}'))
actual.add(next((d for d in digests or [] if d.split('@')[0] == image_repo), image_id))
if len(actual) > 1:
raise ValueError(f'{project}/{service}: mixed running images, cannot capture one recovery config')
# AIO also rejects a digest in its recovery config. Preserve its tag in
# both deploy and recovery files.
if nextcloud_aio_master:
before['services'][service]['image'] = reference
else:
before['services'][service]['image'] = next(iter(actual)) if actual else reference
for name, data in (('compose', config), ('compose-before', before)):
folder = directory / name
folder.mkdir(mode=0o700, exist_ok=True)
destination = folder / f'{relative.parent.name}.json'
destination.write_text(json.dumps(data, indent=2) + '\n')
destination.chmod(0o600)
images_file.write_text(json.dumps(locks, indent=2) + '\n')
print(f'Compose {project}: images pinned; local paths preserved')
print(
f'Recovery: docker compose --project-directory {project_dir} -p {project} -f {directory}/compose-before/{relative.parent.name}.json up -d --pull never --remove-orphans'
)
if __name__ == '__main__':
prepare(Path(sys.argv[1]))
-427
View File
@@ -1,427 +0,0 @@
#!/usr/bin/env python3
"""Durable workstation deployment controller. Install with setup-workstation.sh."""
import argparse
import contextlib
import fcntl
import importlib.util
import json
import math
import os
import re
import shutil
import subprocess
import sys
import time
from pathlib import Path
STATE = Path(os.environ.get('HOMELAB_STATE', Path.home() / '.local/state/homelab-deploy'))
CONFIG_REPO = Path(os.environ.get('HOMELAB_REPO', '/srv/homelab'))
RUN_ID = re.compile(r'[0-9]+-[0-9]+')
def command(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
def atomic_json(path, data):
temporary = path.with_suffix('.tmp')
temporary.write_text(json.dumps(data, indent=2) + '\n')
temporary.chmod(0o600)
temporary.replace(path)
@contextlib.contextmanager
def lock(name):
STATE.mkdir(mode=0o700, parents=True, exist_ok=True)
with (STATE / name).open('a') as stream:
fcntl.flock(stream, fcntl.LOCK_EX)
yield
def load_module(name, path):
spec = importlib.util.spec_from_file_location(name, path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def run_directory(run_id):
if not RUN_ID.fullmatch(run_id):
raise ValueError('Run ID must be numeric workflow-id and attempt')
return STATE / 'runs' / run_id
def start(run_id):
payload = sys.stdin.buffer.read(256 * 1024 + 1)
if len(payload) > 256 * 1024:
raise ValueError('Deploy request exceeds 256 KiB')
request = json.loads(payload)
sha = request['release']['sha']
if not re.fullmatch(r'[0-9a-f]{40}', sha) or request['mode'] not in ('changed', 'full', 'plan'):
raise ValueError('Invalid deploy SHA or mode')
if not isinstance(request['refresh_images'], bool):
raise ValueError('refresh_images must be boolean')
directory = run_directory(run_id)
with lock('prepare.lock'):
if (directory / 'request.json').exists():
if json.loads((directory / 'request.json').read_text()) != request:
raise ValueError('Run ID already belongs to a different request')
else:
directory.mkdir(mode=0o700, parents=True, exist_ok=True)
command('git', '-C', str(CONFIG_REPO), 'fetch', '--quiet', 'origin', 'main')
command('git', '-C', str(CONFIG_REPO), 'merge-base', '--is-ancestor', sha, 'origin/main')
if not (directory / 'source').exists():
command('git', '-C', str(CONFIG_REPO), 'worktree', 'add', '--detach', str(directory / 'source'), sha)
if command('git', '-C', str(directory / 'source'), 'rev-parse', 'HEAD') != sha:
raise ValueError('Prepared source does not match deploy SHA')
release_module = load_module('release', directory / 'source/.gitea/workflows/release.py')
release_module.validate_release(request['release'], sha)
atomic_json(directory / 'release.json', request['release'])
atomic_json(directory / 'request.json', request)
if not (directory / 'status.json').exists():
atomic_json(directory / 'status.json', {'state': 'queued', 'stages': {}})
# Starting an existing active or finished ID is idempotent; never re-apply it.
if json.loads((directory / 'status.json').read_text())['state'] == 'queued':
command('systemctl', '--user', 'start', '--no-block', f'homelab-deploy@{run_id}.service')
print(f'Accepted deploy {run_id} ({sha})')
def environment(directory):
request = json.loads((directory / 'request.json').read_text())
return {
**os.environ,
'REPO': str(directory / 'source'),
'CONFIG_REPO': str(CONFIG_REPO),
'RUN_DIR': str(directory),
'DEPLOY_SHA': request['release']['sha'],
'RELEASE_FILE': str(directory / 'release.json'),
'DEPLOY_PLAN': str(directory / 'plan.json'),
'DEPLOY_SNAPSHOT_DIR': str(directory / 'snapshot'),
'REFRESH_IMAGES': str(request['refresh_images']).lower(),
'ROLLOUT_PARALLELISM': '4',
}
def stage(directory, name, budget):
status = json.loads((directory / 'status.json').read_text())
if name in status['stages'] and status['stages'][name].get('result') in ('success', 'failure'):
return status['stages'][name]['result'] == 'success'
started = time.time()
status['stages'][name] = {'result': 'running', 'started': started}
atomic_json(directory / 'status.json', status)
script = directory / 'source/.gitea/workflows/deploy-stage.sh'
with (directory / f'{name}.log').open('a') as log:
# timeout kills the whole stage process group, including children, before recovery.
result = subprocess.run( # noqa: S603, S607
[
shutil.which('timeout') or '/usr/bin/timeout',
'--signal=TERM',
'--kill-after=30s',
str(budget),
'bash',
str(script),
name,
],
env=environment(directory),
stdout=log,
stderr=subprocess.STDOUT,
check=False,
).returncode
status = json.loads((directory / 'status.json').read_text())
status['stages'][name].update(
result='success' if result == 0 else 'failure', exit_code=result, seconds=round(time.time() - started)
)
atomic_json(directory / 'status.json', status)
return result == 0
def make_plan(directory):
source = directory / 'source'
planner = load_module('deploy_plan', source / '.gitea/workflows/deploy-plan.py')
request = json.loads((directory / 'request.json').read_text())
previous = json.loads((STATE / 'last-success.json').read_text()) if (STATE / 'last-success.json').exists() else None
# Helm 4 lists every release status by default and removed the --all flag.
helm = json.loads(command('helm', 'list', '-A', '-o', 'json'))
plan = planner.make_plan(source, CONFIG_REPO, request['release'], previous, request['mode'], helm)
if request['refresh_images']:
plan['selected']['compose'] = plan['active']['compose']
atomic_json(directory / 'plan.json', plan)
if previous:
atomic_json(directory / 'previous.json', previous)
# Local config is deliberately separate from the immutable Git source.
return plan
def finish_success(directory, plan):
# Repeating finalization after a crash is safe while holding deploy.lock.
plan['run_id'] = directory.name
path = directory / 'compose-images.json'
previous = directory / 'previous.json'
plan['compose-images'] = (
json.loads(path.read_text())
if path.exists()
else json.loads(previous.read_text()).get('compose-images', {})
if previous.exists()
else {}
)
configs = STATE / 'compose-configs'
configs.mkdir(mode=0o700, exist_ok=True)
for config in (directory / 'compose').glob('*.json'):
atomic_json(configs / config.name, json.loads(config.read_text()))
atomic_json(STATE / 'last-success.json', plan)
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'success'
atomic_json(directory / 'status.json', status)
try:
retain_completed(directory)
except (OSError, subprocess.CalledProcessError) as error:
print(f'Retention deferred: {error}', flush=True)
def recover(directory, retry=False):
status = json.loads((directory / 'status.json').read_text())
if status['state'] in ('success', 'planned'):
return
completed = ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke')
if all(status['stages'].get(name, {}).get('result') == 'success' for name in completed):
finish_success(directory, json.loads((directory / 'plan.json').read_text()))
return
if retry:
for name in ('verify-k8s', 'smoke'):
if status['stages'].get(name, {}).get('result') == 'failure':
del status['stages'][name]
atomic_json(directory / 'status.json', status)
snapshot = directory / 'snapshot/current'
if snapshot.exists():
stage(directory, 'verify-k8s', 7200)
stage(directory, 'smoke', 600)
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'failure'
atomic_json(directory / 'status.json', status)
def execute(run_id):
directory = run_directory(run_id)
with lock('deploy.lock'):
status = json.loads((directory / 'status.json').read_text())
if status['state'] != 'queued':
return
# A crashed predecessor must be recovered before another apply begins.
for other in (STATE / 'runs').iterdir():
if (
other != directory
and (other / 'status.json').exists()
and json.loads((other / 'status.json').read_text())['state'] == 'running'
):
raise ValueError(f'Interrupted deploy {other.name}; run recover first')
status['state'] = 'running'
atomic_json(directory / 'status.json', status)
phase = 'plan'
try:
plan = make_plan(directory)
print(
json.dumps({'selected': plan['selected'], 'helm': plan['helm'], 'manual_removals': plan['removed']}),
flush=True,
)
phase = 'doctor'
if not stage(directory, 'doctor', 600):
raise RuntimeError('Preflight failed')
phase = 'validate'
if not stage(directory, 'validate', 1200):
raise RuntimeError('Validation failed')
if json.loads((directory / 'request.json').read_text())['mode'] == 'plan':
status = json.loads((directory / 'status.json').read_text())
status['state'] = 'planned'
atomic_json(directory / 'status.json', status)
return
# Budget includes both rollout checks and rollback waves, plus API overhead.
phase = 'Recovery budget'
count = int(
command(
'bash',
str(directory / 'source/.gitea/workflows/deploy-stage.sh'),
'workload-count',
env=environment(directory),
)
)
verify_budget = max(600, 2 * math.ceil(count / 4) * 300 + 120)
if verify_budget > 7200:
raise ValueError('More than two hours of recovery required; split this deploy')
phase = 'apply-k8s'
k8s_ok = stage(directory, 'apply-k8s', 2700)
phase = 'apply-compose'
compose_ok = stage(directory, 'apply-compose', 1800) if k8s_ok else False
phase = 'verify-k8s'
verify_ok = stage(directory, 'verify-k8s', verify_budget)
phase = 'smoke'
smoke_ok = stage(directory, 'smoke', 600)
if not all((k8s_ok, compose_ok, verify_ok, smoke_ok)):
raise RuntimeError('Deploy failed; inspect stage logs and recovery report')
phase = 'Save the successful baseline'
finish_success(directory, plan)
except Exception as error:
status = json.loads((directory / 'status.json').read_text())
status['failure_stage'] = next(
(name for name, result in status['stages'].items() if result.get('result') == 'failure'), phase
)
atomic_json(directory / 'status.json', status)
with (directory / 'controller.log').open('a') as stream:
stream.write(f'{error}\n')
recover(directory)
raise
def retain_completed(current):
finished = []
for directory in (STATE / 'runs').iterdir():
status_file = directory / 'status.json'
if status_file.exists() and json.loads(status_file.read_text())['state'] in ('success', 'planned'):
finished.append(directory)
for directory in sorted(finished, key=lambda p: p.stat().st_mtime, reverse=True)[20:]:
if directory == current:
continue
command('git', '-C', str(CONFIG_REPO), 'worktree', 'remove', '--force', str(directory / 'source'))
shutil.rmtree(directory)
def follow(run_id, phase):
directory = run_directory(run_id)
groups = {
'apply': ('doctor', 'validate', 'apply-k8s', 'apply-compose'),
'verify': ('verify-k8s',),
'smoke': ('smoke',),
}
names = groups[phase]
offsets = {}
while True:
status = json.loads((directory / 'status.json').read_text())
for name in (*names, 'controller'):
path = directory / f'{name}.log'
if path.exists():
with path.open() as stream:
stream.seek(offsets.get(name, 0))
content = stream.read()
if content:
print(content, end='', flush=True)
offsets[name] = stream.tell()
stages = status['stages']
if all(stages.get(name, {}).get('result') in ('success', 'failure') for name in names):
return all(stages[name]['result'] == 'success' for name in names)
if status['state'] in ('success', 'failure', 'planned'):
return status['state'] in ('success', 'planned')
time.sleep(3)
def summary(run_id):
directory = run_directory(run_id)
request = json.loads((directory / 'request.json').read_text())
release = request['release']
plan_file = directory / 'plan.json'
lines = [
f'## Deploy `{release["sha"]}`',
'',
f'- Mode: `{request["mode"]}`',
f'- Refresh third-party images: `{request["refresh_images"]}`',
]
status = json.loads((directory / 'status.json').read_text())
if status.get('failure_stage'):
lines.append(f'- Failed stage: **{status["failure_stage"]}**')
lines.extend(
[
'',
f'- Observed run state: **{status["state"]}**',
'',
'### Stage results',
'| Stage | Result | Exit code |',
'| --- | --- | --- |',
]
)
for name in ('doctor', 'validate', 'apply-k8s', 'apply-compose', 'verify-k8s', 'smoke'):
stage_result = status['stages'].get(name, {})
lines.append(f'| {name} | {stage_result.get("result", "not started")} | {stage_result.get("exit_code", "—")} |')
lines.extend(['', '### Apply and Helm recovery results'])
events_file = directory / 'apply-events.jsonl'
events = []
if events_file.exists():
for line in events_file.read_text().splitlines():
try:
events.append(json.loads(line))
except json.JSONDecodeError:
lines.append('- An operation record is incomplete. Check the stage log.')
latest = {(event['action'], event['target']): event['result'] for event in events}
lines.extend(f'- `{action}` `{target}`: **{result}**' for (action, target), result in latest.items())
if not latest:
lines.append('- No apply results were recorded.')
lines.append('- A completed apply does not confirm health. See verification and smoke results.')
lines.extend(['', '### Kubernetes recovery'])
pointer = directory / 'snapshot/current'
failed = Path(pointer.read_text().strip()) / 'failed-workloads' if pointer.exists() else None
if failed and failed.exists():
contents = failed.read_text()
counts = dict(re.findall(r'^(ROLLED_BACK|UNRECOVERED)=([0-9]+)$', contents, re.MULTILINE))
if not contents.strip():
lines.append('- No failed workloads were recorded. See the verification result above.')
elif counts:
lines.append(f'- Workloads restored: **{counts.get("ROLLED_BACK", "unknown")}**')
lines.append(f'- Workloads that need manual recovery: **{counts.get("UNRECOVERED", "unknown")}**')
else:
lines.append('- Rollback has no recorded result yet. Check the verification log.')
else:
lines.append('- No workload rollback was recorded. This does not confirm health.')
lines.append('- Compose requires manual recovery. Use the saved command in the apply log.')
if not plan_file.exists():
lines.extend(['', 'Plan was not created. Check the controller log.'])
print('\n'.join(lines))
return
plan = json.loads(plan_file.read_text())
lines.extend(['', '### Selected services'])
count = 0
for kind, services in plan['selected'].items():
for service in services:
lines.append(f'- `{kind}`: `{service}`')
count += 1
if not count:
lines.append('- None')
lines.extend(['', '### Selected Helm releases'])
lines.extend(f'- `{release}`' for release in plan.get('helm', []))
if not plan.get('helm'):
lines.append('- None')
lines.extend(['', '### Images pinned in the checked release'])
lines.extend(f'- `{image}@{digest}`' for image, digest in sorted(release['images'].items()))
lines.extend(['', '### Removed resources requiring manual review'])
lines.extend(f'- `{item}`' for item in plan.get('removed', []))
if not plan.get('removed'):
lines.append('- None')
print('\n'.join(lines))
def main():
os.umask(0o077)
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('start', 'execute', 'recover', 'status', 'follow', 'summary'))
parser.add_argument('run_id')
parser.add_argument('phase', nargs='?', choices=('apply', 'verify', 'smoke'))
parser.add_argument('--retry', action='store_true', help='Retry failed recovery checks; never repeat apply')
args = parser.parse_args()
directory = run_directory(args.run_id)
if args.action == 'start':
start(args.run_id)
elif args.action == 'execute':
execute(args.run_id)
elif args.action == 'recover':
with lock('deploy.lock'):
recover(directory, retry=args.retry)
elif args.action == 'status':
print((directory / 'status.json').read_text())
if (directory / 'plan.json').exists():
plan = json.loads((directory / 'plan.json').read_text())
print(json.dumps({k: plan[k] for k in ('sha', 'selected', 'helm', 'removed')}, indent=2))
elif args.action == 'summary':
summary(args.run_id)
elif not follow(args.run_id, args.phase):
sys.exit(1)
if __name__ == '__main__':
main()
File diff suppressed because it is too large. Load diff
-124
View File
@@ -1,124 +0,0 @@
#!/usr/bin/env python3
"""Calculate selected components against the last fully successful deploy."""
import hashlib
import json
import re
import subprocess
from pathlib import Path
def output(*args, **kwargs):
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603
def tracked(repo):
return output('git', '-C', str(repo), 'ls-files').splitlines()
def helm_releases(repo):
text = (repo / '.gitea/workflows/deploy-lib.sh').read_text()
return [line.split('|') for line in re.findall(r'^ "([^"\n]+\|[^"\n]+)"$', text, re.MULTILINE)]
def inventory(repo):
files = tracked(repo)
k8s = sorted(
{f.split('/k8s/')[0] for f in files if '/k8s/' in f and (repo / f.split('/k8s/')[0] / 'k8s/active').is_file()}
)
compose = sorted(
{
str(Path(f).parent)
for f in files
if Path(f).name in ('compose.yaml', 'compose.yml') and (repo / Path(f).parent / 'active').is_file()
}
)
return {'k8s': k8s, 'compose': compose}
def file_hash(path):
return hashlib.sha256(path.read_bytes()).hexdigest() if path.is_file() else 'missing'
def make_plan(repo, config_repo, release, previous, mode, live_helm):
active = inventory(repo)
all_services = set(active['k8s'] + active['compose'])
helm_inputs = {}
helm_selected = []
for name, chart, namespace, version, values, marker in helm_releases(repo):
if not (repo / marker).is_file():
continue
value_path = repo / values if (repo / values).is_file() else config_repo / values
if not value_path.is_file():
raise ValueError(f'Missing Helm values: {values}')
stamp = hashlib.sha256(f'{chart}|{version}|{file_hash(value_path)}'.encode()).hexdigest()
helm_inputs[name] = stamp
live = next((h for h in live_helm if h['name'] == name and h['namespace'] == namespace), None)
if (
mode == 'full'
or previous is None
or previous.get('helm_inputs', {}).get(name) != stamp
or live is None
or live.get('status') != 'deployed'
or live.get('chart') != f'{chart.split("/")[-1]}-{version}'
):
helm_selected.append(name)
local_inputs = {}
for service in all_services:
candidates = [config_repo / service / '.env']
if service in active['compose']:
candidates.append(config_repo / '.env')
cfg = config_repo / service / 'config'
if cfg.is_dir():
candidates.extend(
p for p in cfg.rglob('*') if p.is_file() and p.suffix in ('.yaml', '.yml', '.json', '.conf')
)
local_inputs[service] = hashlib.sha256(
'\n'.join(f'{p.relative_to(config_repo)}:{file_hash(p)}' for p in sorted(candidates)).encode()
).hexdigest()
if previous is None:
if mode == 'changed':
raise ValueError('No successful baseline; run deploy in full mode first')
changed = set(all_services)
removed = []
else:
paths = output('git', '-C', str(repo), 'diff', '--name-only', previous['sha'], release['sha']).splitlines()
changed = {service for service in all_services for path in paths if path.startswith(service + '/')}
if any(path.startswith('.gitea/') for path in paths):
changed |= all_services
changed |= {s for s in all_services if previous.get('local_inputs', {}).get(s) != local_inputs[s]}
for file in tracked(repo):
owners = {service for service in all_services if file.startswith(service + '/')}
if not owners or not file.endswith(('.yaml', '.yml')):
continue
text = (repo / file).read_text()
if any(
image in text and previous.get('images', {}).get(image) != digest
for image, digest in release['images'].items()
):
changed |= owners
removed = sorted(
set(previous.get('active', {}).get('k8s', []) + previous.get('active', {}).get('compose', []))
- all_services
)
removed += [path for path in paths if '/k8s/' in path and not (repo / path).exists()]
if mode == 'full':
changed = set(all_services)
dependencies = json.loads((repo / '.gitea/deploy-dependencies.json').read_text())
while True:
expanded = changed | {dependent for service in changed for dependent in dependencies.get(service, [])}
if expanded == changed:
break
changed = expanded
return {
'version': 1,
'sha': release['sha'],
'images': release['images'],
'active': active,
'selected': {kind: sorted(set(services) & changed) for kind, services in active.items()},
'helm': helm_selected,
'helm_inputs': helm_inputs,
'local_inputs': local_inputs,
'removed': sorted(set(removed)),
'full_smoke': mode == 'full' or 'traefik' in changed,
}
-10
View File
@@ -1,10 +0,0 @@
#!/usr/bin/env bash
set -euo pipefail
source "${REPO:?}/.gitea/workflows/deploy-lib.sh"
case "${1:?stage required}" in
workload-count)
select_manifests >/dev/null
selected_workload_refs | sort -u | wc -l
;;
*) run_stage "$1" ;;
esac
+48 -117
View File
@@ -1,140 +1,71 @@
name: deploy name: deploy
on: on:
workflow_run: push:
workflows: [ci] branches:
branches: [main] - main
types: [completed]
workflow_dispatch: workflow_dispatch:
inputs:
deploy_ref:
description: "Commit already checked by successful main CI (main or SHA)"
default: main
required: true
deploy_mode:
description: "First deploy requires full; plan changes no production resources"
type: choice
options: [changed, full, plan]
default: changed
refresh_images:
description: "Explicitly refresh mutable third-party Compose tags"
type: boolean
default: false
permissions:
contents: read
actions: read
concurrency: concurrency:
group: deploy-main group: deploy-main
cancel-in-progress: false cancel-in-progress: false
env: env:
DEPLOY_HOST: ${{ vars.DEPLOY_HOST || secrets.DEPLOY_HOST }} DEPLOY_HOST: ${{ secrets.DEPLOY_HOST }}
DEPLOY_PORT: ${{ vars.DEPLOY_PORT || secrets.DEPLOY_PORT }} DEPLOY_PORT: ${{ secrets.DEPLOY_PORT }}
DEPLOY_USER: ${{ vars.DEPLOY_USER || secrets.DEPLOY_USER }} DEPLOY_USER: ${{ secrets.DEPLOY_USER }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }} DEPLOY_KEY: ${{ secrets.DEPLOY_SSH_KEY }}
DEPLOY_KNOWN_HOSTS: ${{ vars.DEPLOY_KNOWN_HOSTS }} APPLY_PRUNE: ${{ vars.APPLY_PRUNE }}
DEPLOY_RUN_ID: ${{ github.run_id }}-${{ github.run_attempt || 1 }}
DEPLOY_MODE: ${{ inputs.deploy_mode || 'changed' }}
REFRESH_IMAGES: ${{ inputs.refresh_images && 'true' || 'false' }}
jobs: jobs:
gate: preflight:
if: >- runs-on: [self-hosted, linux, arch, homelab, prod]
github.ref == 'refs/heads/main' &&
(vars.AUTODEPLOY == 'true' || github.event_name == 'workflow_dispatch') &&
(github.event_name != 'workflow_run' ||
(github.event.workflow_run.conclusion == 'success' && github.event.workflow_run.head_branch == 'main'))
runs-on: homelab
timeout-minutes: 10
outputs:
sha: ${{ steps.release.outputs.sha }}
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
fetch-depth: 0
- name: Check successful CI and download the exact commit release
id: release
env:
GITEA_TOKEN: ${{ github.token }}
DEPLOY_REF: ${{ inputs.deploy_ref || 'main' }}
EVENT_SHA: ${{ github.event.workflow_run.head_sha }}
run: python3 .gitea/workflows/release.py gate --ref "$DEPLOY_REF" --event-sha "$EVENT_SHA"
- name: Submit durable deploy to workstation
run: bash .gitea/workflows/ssh-run.sh start
- name: Write the request result
if: always()
env:
REQUEST_RESULT: ${{ job.status }}
CHECKED_SHA: ${{ steps.release.outputs.sha }}
run: |
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
printf '## Deploy request\n\n- Result: **%s**\n- Checked commit: %s\n- Mode: %s\n' "$REQUEST_RESULT" "${CHECKED_SHA:-not checked}" "$DEPLOY_MODE" >>"$GITHUB_STEP_SUMMARY"
if [ "$REQUEST_RESULT" != success ]; then
echo 'Open the failed step log. If SSH submission failed, check the remote controller state.' >>"$GITHUB_STEP_SUMMARY"
fi
fi
apply: - name: Fetch and reset workstation
needs: [gate] shell: bash
runs-on: homelab
timeout-minutes: 120
steps:
- name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow validation and sequential Kubernetes / Compose apply
run: bash .gitea/workflows/ssh-run.sh apply
- name: Write the deploy result
if: always()
run: | run: |
if [ -f .gitea/workflows/ssh-run.sh ]; then set -euo pipefail
bash .gitea/workflows/ssh-run.sh summary ./.gitea/workflows/ssh-run.sh preflight
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
verify: validate:
needs: [gate, apply] needs: [preflight]
if: always() && needs.gate.result == 'success' runs-on: [self-hosted, linux, arch, homelab, prod]
runs-on: homelab
timeout-minutes: 130
steps: steps:
- name: Checkout checked commit - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow workload verification and recovery
run: bash .gitea/workflows/ssh-run.sh verify
- name: Write the deploy result
if: always()
run: |
if [ -f .gitea/workflows/ssh-run.sh ]; then
bash .gitea/workflows/ssh-run.sh summary
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY"
fi
smoke: - name: Dry-run manifests and check Secrets
needs: [gate, verify] shell: bash
if: always() && needs.gate.result == 'success'
runs-on: homelab
timeout-minutes: 15
steps:
- name: Checkout checked commit
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ needs.gate.outputs.sha }}
- name: Follow public route checks
run: bash .gitea/workflows/ssh-run.sh smoke
- name: Write the deploy result
if: always()
run: | run: |
if [ -f .gitea/workflows/ssh-run.sh ]; then set -euo pipefail
bash .gitea/workflows/ssh-run.sh summary ./.gitea/workflows/ssh-run.sh validate
elif [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
echo 'Source checkout failed. The remote deploy state is unknown. Check the job log.' >>"$GITHUB_STEP_SUMMARY" apply-k8s:
fi needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Apply Kubernetes manifests
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-k8s
apply-compose:
needs: [validate]
runs-on: [self-hosted, linux, arch, homelab, prod]
steps:
- name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
- name: Redeploy docker compose stacks
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/ssh-run.sh apply-compose
-285
View File
@@ -1,285 +0,0 @@
#!/usr/bin/env bash
# Installs the pinned CI tools into "$TOOLS_DIR/bin" and echoes that directory
# on stdout, so callers can do:
#
# export PATH="$(bash .gitea/workflows/install-ci-tools.sh kubeconform shellcheck):$PATH"
#
# Versions come from tool-versions.env next to this script and are kept fresh by
# Renovate. Re-running is cheap: an already-installed tool at the pinned version
# is left alone.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=tool-versions.env
. "$here/tool-versions.env"
TOOLS_DIR="${TOOLS_DIR:-${XDG_CACHE_HOME:-$HOME/.cache}/homelab-ci}"
BIN_DIR="$TOOLS_DIR/bin"
mkdir -p "$BIN_DIR"
# A runner may accept overlapping workflows even though each workflow is sequential.
exec 9>"$TOOLS_DIR/install.lock"
flock -w 300 9
export UV_TOOL_DIR="$TOOLS_DIR/uv-tools"
export UV_CACHE_DIR="$TOOLS_DIR/uv-cache"
# The just-installed tools must resolve inside this script too: callers only
# prepend BIN_DIR to PATH after the script exits, so a bare `uv` below would
# miss the binary install_uv just placed (exit 127 on a clean runner).
export PATH="$BIN_DIR:$PATH"
arch="$(uname -m)"
# Upstream projects disagree on arch spelling: kubeconform and actionlint use
# Go names (amd64/arm64), shellcheck uses uname names (x86_64/aarch64), node
# uses neither (x64/arm64), and hadolint mixes the two in a single release
# (x86_64 but arm64).
case "$arch" in
x86_64 | amd64)
goarch=amd64
sharch=x86_64
nodearch=x64
hadolintarch=x86_64
;;
aarch64 | arm64)
goarch=arm64
sharch=aarch64
nodearch=arm64
hadolintarch=arm64
;;
*)
echo "install-ci-tools: unsupported architecture: $arch" >&2
exit 1
;;
esac
fetch() {
# fetch <url> <dest>
if command -v curl >/dev/null 2>&1; then
curl -sSLf --connect-timeout 15 --max-time 120 --retry 3 -o "$2" "$1"
elif command -v wget >/dev/null 2>&1; then
wget -q -O "$2" "$1"
else
echo "install-ci-tools: neither curl nor wget is available" >&2
exit 1
fi
}
# resolve <command>
# Absolute path to use for invoking a tool: the copy in BIN_DIR when present,
# otherwise the name for PATH lookup. Every version check and every in-script
# invocation goes through this, so a tool missing from both places reads as
# "not installed" instead of dying with 127 under `set -e`.
resolve() {
if [ -x "$BIN_DIR/$1" ]; then
printf '%s' "$BIN_DIR/$1"
else
printf '%s' "$1"
fi
}
# installed_version <command>
# Prints the version of an already-installed tool, or nothing. Each tool spells
# its version flag differently, hence the case.
installed_version() {
local bin out
bin="$(resolve "$1")"
if ! command -v "$bin" >/dev/null 2>&1; then
return 0
fi
case "$1" in
kubeconform) out="$("$bin" -v 2>/dev/null | head -1 || true)" ;;
*) out="$("$bin" --version 2>/dev/null | head -1 || true)" ;;
esac
printf '%s' "$out"
}
# at_version <command> <expected>
at_version() {
local version expected="${2#v}"
version="$(installed_version "$1")"
if [[ "$version" =~ (^|[^0-9.])v?([0-9]+(\.[0-9]+)+) ]]; then
[ "${BASH_REMATCH[2]}" = "$expected" ]
else
return 1
fi
}
install_kubeconform() {
if at_version kubeconform "v${KUBECONFORM_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/yannh/kubeconform/releases/download/v${KUBECONFORM_VERSION}/kubeconform-linux-${goarch}.tar.gz" \
"$tmp/kubeconform.tar.gz"
tar -xzf "$tmp/kubeconform.tar.gz" -C "$tmp" kubeconform
install -m 0755 "$tmp/kubeconform" "$BIN_DIR/kubeconform"
rm -rf "$tmp"
}
install_shellcheck() {
if at_version shellcheck "${SHELLCHECK_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/koalaman/shellcheck/releases/download/v${SHELLCHECK_VERSION}/shellcheck-v${SHELLCHECK_VERSION}.linux.${sharch}.tar.xz" \
"$tmp/shellcheck.tar.xz"
tar -xJf "$tmp/shellcheck.tar.xz" -C "$tmp" --strip-components=1 "shellcheck-v${SHELLCHECK_VERSION}/shellcheck"
install -m 0755 "$tmp/shellcheck" "$BIN_DIR/shellcheck"
rm -rf "$tmp"
}
install_jq() {
if at_version jq "${JQ_VERSION}"; then
return 0
fi
fetch "https://github.com/jqlang/jq/releases/download/jq-${JQ_VERSION}/jq-linux-${goarch}" \
"$BIN_DIR/jq"
chmod 0755 "$BIN_DIR/jq"
}
install_uv() {
if at_version uv "${UV_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
# uv release tags carry no leading v, unlike every other tool installed here.
fetch "https://github.com/astral-sh/uv/releases/download/${UV_VERSION}/uv-${sharch}-unknown-linux-gnu.tar.gz" \
"$tmp/uv.tar.gz"
tar -xzf "$tmp/uv.tar.gz" -C "$tmp" --strip-components=1 "uv-${sharch}-unknown-linux-gnu/uv"
install -m 0755 "$tmp/uv" "$BIN_DIR/uv"
rm -rf "$tmp"
}
install_hadolint() {
if at_version hadolint "${HADOLINT_VERSION}"; then
return 0
fi
# A bare binary, no archive: hadolint ships one file per platform.
fetch "https://github.com/hadolint/hadolint/releases/download/v${HADOLINT_VERSION}/hadolint-linux-${hadolintarch}" \
"$BIN_DIR/hadolint"
chmod 0755 "$BIN_DIR/hadolint"
}
# ruff and yamllint both come from PyPI as wheels, which uv unpacks for us.
install_uv_tool() {
# <package> <pinned version>
if at_version "$1" "$2"; then
return 0
fi
install_uv
UV_TOOL_BIN_DIR="$BIN_DIR" "$BIN_DIR/uv" tool install --force "$1==$2" >/dev/null
}
install_ruff() {
install_uv_tool ruff "${RUFF_VERSION}"
}
install_yamllint() {
install_uv_tool yamllint "${YAMLLINT_VERSION}"
}
install_pip_audit() {
install_uv_tool pip-audit "${PIP_AUDIT_VERSION}"
}
install_prettier() {
install_node
if at_version prettier "${PRETTIER_VERSION}"; then
return 0
fi
# Not a standalone binary: prettier's entry point requires ../package.json
# relative to its own real path, so the package directory has to survive
# next to it. Hence a versioned directory plus a relative symlink, rather
# than copying the one file out as the other installers do.
local dir="$BIN_DIR/prettier-${PRETTIER_VERSION}"
if [ ! -f "$dir/package/package.json" ]; then
rm -rf "$dir"
mkdir -p "$dir"
fetch "https://registry.npmjs.org/prettier/-/prettier-${PRETTIER_VERSION}.tgz" "$dir/prettier.tgz"
tar -xzf "$dir/prettier.tgz" -C "$dir"
rm -f "$dir/prettier.tgz"
# npm strips the exec bit from bin/ on the way into the tarball.
chmod 0755 "$dir/package/bin/prettier.cjs"
fi
# Relative, so the whole tree stays valid if TOOLS_DIR is relocated.
ln -sfn "prettier-${PRETTIER_VERSION}/package/bin/prettier.cjs" "$BIN_DIR/prettier"
}
install_node() {
# npm gets checked by running it, not by looking it up: what matters is that
# it answers, so a stub, a half-removed Arch package or a name that resolves
# to something broken all have to read as "not installed". The runner's npm
# is a symlink into /usr/lib/node_modules/npm, which is exactly the kind of
# thing that disappears between runs.
if at_version node "v${NODE_VERSION}" && [ -n "$(installed_version npm)" ]; then
return 0
fi
# Same shape as prettier above: the tarball's bin/npm and bin/npx are links
# into lib/node_modules, so the whole tree has to survive next to them.
local dir="$BIN_DIR/node-${NODE_VERSION}"
if [ ! -x "$dir/bin/node" ]; then
rm -rf "$dir"
mkdir -p "$dir"
fetch "https://nodejs.org/dist/v${NODE_VERSION}/node-v${NODE_VERSION}-linux-${nodearch}.tar.xz" \
"$dir/node.tar.xz"
tar -xJf "$dir/node.tar.xz" -C "$dir" --strip-components=1 "node-v${NODE_VERSION}-linux-${nodearch}"
rm -f "$dir/node.tar.xz"
fi
# Relative, so the whole tree stays valid if TOOLS_DIR is relocated.
for bin in node npm npx; do
ln -sfn "node-${NODE_VERSION}/bin/${bin}" "$BIN_DIR/${bin}"
done
}
install_actionlint() {
if at_version actionlint "${ACTIONLINT_VERSION}"; then
return 0
fi
local tmp
tmp="$(mktemp -d)"
fetch "https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/actionlint_${ACTIONLINT_VERSION}_linux_${goarch}.tar.gz" \
"$tmp/actionlint.tar.gz"
tar -xzf "$tmp/actionlint.tar.gz" -C "$tmp" actionlint
install -m 0755 "$tmp/actionlint" "$BIN_DIR/actionlint"
rm -rf "$tmp"
}
main() {
wanted=("$@")
if [ "${#wanted[@]}" -eq 0 ]; then
wanted=(node jq kubeconform shellcheck actionlint prettier ruff yamllint hadolint)
fi
for tool in "${wanted[@]}"; do
case "$tool" in
kubeconform) install_kubeconform ;;
shellcheck) install_shellcheck ;;
jq) install_jq ;;
actionlint) install_actionlint ;;
prettier) install_prettier ;;
ruff) install_ruff ;;
yamllint) install_yamllint ;;
pip-audit) install_pip_audit ;;
hadolint) install_hadolint ;;
node) install_node ;;
uv) install_uv ;;
*)
echo "install-ci-tools: unknown tool: $tool" >&2
exit 1
;;
esac
done
for old in "$BIN_DIR"/node-* "$BIN_DIR"/prettier-*; do
[ -d "$old" ] || continue
case "$(basename "$old")" in
"node-$NODE_VERSION"|"prettier-$PRETTIER_VERSION") ;;
*) rm -rf "$old" ;;
esac
done
if [ -x "$BIN_DIR/uv" ]; then "$BIN_DIR/uv" cache prune >/dev/null; fi
printf '%s\n' "$BIN_DIR"
}
if [ "${BASH_SOURCE[0]}" = "$0" ]; then main "$@"; fi
-519
View File
@@ -1,519 +0,0 @@
#!/usr/bin/env python3
"""CI release artifacts and the SHA-specific Gitea deployment gate (stdlib only)."""
import argparse
import hashlib
import io
import itertools
import json
import os
import re
import shutil
import subprocess
import sys
import tempfile
import urllib.error
import urllib.parse
import urllib.request
import zipfile
from pathlib import Path
SHA = re.compile(r'[0-9a-f]{40}')
DIGEST = re.compile(r'sha256:[0-9a-f]{64}')
IMAGES = {
'error-pages': ('errorpages', 'errorpages/Dockerfile'),
'forust-homepage': ('homepages', 'homepages/Dockerfile.forust'),
'xdfnx-homepage': ('homepages', 'homepages/Dockerfile.xdfnx'),
}
def command(*args, **kwargs):
"""Arguments are passed directly to the executable, never to a shell."""
return subprocess.check_output(args, text=True, **kwargs).strip() # noqa: S603, S607
def validate_release(data, sha=None):
if data.get('version') != 1 or not SHA.fullmatch(data.get('sha', '')):
raise ValueError('Invalid release version or SHA')
if sha is not None and data['sha'] != sha:
raise ValueError('Release SHA does not match the checked CI commit')
expected = {f'gcr.forust.xyz/forust/{name}' for name in IMAGES}
if set(data.get('images', {})) != expected:
raise ValueError('Release must contain all owned images')
if not all(DIGEST.fullmatch(value) for value in data['images'].values()):
raise ValueError('Release has an invalid image digest')
if set(data.get('inputs', {})) != expected or not all(
re.fullmatch(r'[0-9a-f]{64}', value) for value in data['inputs'].values()
):
raise ValueError('Release has invalid build input fingerprints')
return data
class NoRedirect(urllib.request.HTTPRedirectHandler):
def redirect_request(self, _req, _fp, _code, _msg, _headers, _newurl):
return None
class Gitea:
def __init__(self):
self.origin = os.environ['GITHUB_SERVER_URL'].rstrip('/')
if urllib.parse.urlsplit(self.origin).scheme != 'https':
raise ValueError('Gitea API must use HTTPS')
self.repository = os.environ['GITHUB_REPOSITORY']
if not re.fullmatch(r'[\w.-]+/[\w.-]+', self.repository):
raise ValueError('Invalid Gitea repository')
self.token = os.environ['GITEA_TOKEN']
self.base = f'{self.origin}/api/v1/repos/{self.repository}'
def request(self, url, *, archive=False):
if not url.startswith(self.base + '/'):
raise ValueError('Refusing to send the Actions token to another origin')
req = urllib.request.Request(url, headers={'Authorization': f'token {self.token}'}) # noqa: S310 -- HTTPS origin validated above
opener = urllib.request.build_opener(NoRedirect())
try:
response = opener.open(req, timeout=30) # noqa: S310
except urllib.error.HTTPError as error:
if not archive or error.code not in (301, 302, 303, 307, 308):
raise RuntimeError(f'Gitea API returned HTTP {error.code}') from None
target = urllib.parse.urljoin(url, error.headers['Location'])
if urllib.parse.urlsplit(target).scheme != 'https':
raise ValueError('Artifact redirect must use HTTPS') from None
# Signed storage redirects must never receive the Gitea token.
response = urllib.request.urlopen(target, timeout=30) # noqa: S310
with response:
payload = response.read(8 * 1024 * 1024 + 1)
if len(payload) > 8 * 1024 * 1024:
raise ValueError('Gitea response exceeds 8 MiB')
return payload if archive else json.loads(payload)
def pages(self, path, key, **params):
for page in range(1, 101):
query = urllib.parse.urlencode({**params, 'page': page, 'limit': 50})
data = self.request(f'{self.base}/{path}?{query}')
entries = data[key]
yield from entries
if len(entries) < 50:
return
raise RuntimeError('Gitea pagination limit exceeded')
def successful_runs(self, sha=None):
params = {'branch': 'main', 'status': 'success', 'exclude_pull_requests': 'true'}
if sha:
params['head_sha'] = sha
for run in self.pages('actions/workflows/ci.yaml/runs', 'workflow_runs', **params):
if (
run.get('status') == 'completed'
and run.get('conclusion') == 'success'
and run.get('head_branch') == 'main'
and run.get('event') in ('push', 'workflow_dispatch')
and (run.get('repository') or {}).get('full_name') == self.repository
and (run.get('head_repository') or run.get('repository') or {}).get('full_name') == self.repository
and (sha is None or run.get('head_sha') == sha)
):
yield run
def release(self, run):
sha = run['head_sha']
jobs = list(self.pages(f'actions/runs/{run["id"]}/jobs', 'jobs'))
# A green workflow with a skipped build must not authorize a deploy.
if not any(job.get('name') == 'build' and job.get('conclusion') == 'success' for job in jobs):
raise ValueError('CI build job did not succeed')
artifacts = self.request(f'{self.base}/actions/runs/{run["id"]}/artifacts')['artifacts']
matching = [a for a in artifacts if a['name'] == f'release-{sha}' and not a.get('expired')]
if len(matching) != 1:
raise ValueError('CI release artifact is missing, expired or ambiguous; rerun CI')
blob = self.request(f'{self.base}/actions/artifacts/{matching[0]["id"]}/zip', archive=True)
with zipfile.ZipFile(io.BytesIO(blob)) as archive:
files = [entry for entry in archive.infolist() if not entry.is_dir()]
if len(files) != 1 or files[0].filename != 'release.json' or files[0].file_size > 256 * 1024:
raise ValueError('Unexpected release archive contents')
return validate_release(json.loads(archive.read(files[0])), sha)
def fingerprint(context, dockerfile):
tree = command('git', 'ls-tree', '-r', 'HEAD', '--', context, dockerfile, '.gitea/workflows/release.py')
return hashlib.sha256(tree.encode()).hexdigest()
def gate(output, requested_ref, event_sha):
command('git', 'fetch', '--quiet', 'origin', 'main')
if event_sha:
if not SHA.fullmatch(event_sha):
raise ValueError('Invalid workflow_run SHA')
sha = event_sha
else:
if requested_ref == 'main':
requested_ref = 'origin/main'
sha = command('git', 'rev-parse', '--verify', '--end-of-options', f'{requested_ref}^{{commit}}')
if not SHA.fullmatch(sha):
raise ValueError('Invalid deploy SHA')
command('git', 'merge-base', '--is-ancestor', sha, 'origin/main')
api = Gitea()
runs = list(api.successful_runs(sha))
if not runs:
raise ValueError(f'No successful main CI for {sha}; run CI before deploying')
release = api.release(max(runs, key=lambda run: run['id']))
output.write_text(json.dumps(release, indent=2) + '\n')
if os.environ.get('GITHUB_OUTPUT'):
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
stream.write(f'sha={sha}\n')
print(f'CI gate accepted {sha}')
def prepare_images(output):
sha = command('git', 'rev-parse', 'HEAD')
if sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
raise ValueError('Build checkout does not match GITHUB_SHA')
api = Gitea()
previous = None
for run in sorted(itertools.islice(api.successful_runs(), 50), key=lambda item: item['id'], reverse=True):
if str(run['id']) == os.environ.get('GITHUB_RUN_ID'):
continue
try:
previous = api.release(run)
break
except ValueError:
# Expired artifacts only cost a rebuild; mutable tags are never a fallback.
continue
targets = []
for name, (context, dockerfile) in IMAGES.items():
image = f'gcr.forust.xyz/forust/{name}'
inputs = fingerprint(context, dockerfile)
old_digest = (previous or {}).get('images', {}).get(image)
targets.append(
{
'name': name,
'image': image,
'context': context,
'dockerfile': dockerfile,
'inputs': inputs,
'reuse_digest': old_digest if (previous or {}).get('inputs', {}).get(image) == inputs else None,
}
)
output.write_text(json.dumps({'sha': sha, 'targets': targets}, indent=2) + '\n')
if os.environ.get('GITHUB_OUTPUT'):
with Path(os.environ['GITHUB_OUTPUT']).open('a') as stream:
stream.write('matrix=' + json.dumps({'include': targets}, separators=(',', ':')) + '\n')
print(f'Prepared {len(targets)} image jobs; {sum(t["reuse_digest"] is None for t in targets)} require builds')
def checked_plan(path):
data = json.loads(path.read_text())
sha = command('git', 'rev-parse', 'HEAD')
if data.get('sha') != sha or sha != os.environ['GITHUB_SHA'] or not SHA.fullmatch(sha):
raise ValueError('Image plan does not match the checked source commit')
targets = data.get('targets', [])
if sorted(t['name'] for t in targets) != sorted(IMAGES):
raise ValueError('Image plan must contain each owned image once')
for target in targets:
name = target['name']
context, dockerfile = IMAGES[name]
if (target['context'], target['dockerfile'], target['image']) != (
context,
dockerfile,
f'gcr.forust.xyz/forust/{name}',
) or target['inputs'] != fingerprint(context, dockerfile):
raise ValueError('Image plan has invalid build inputs')
if target['reuse_digest'] is not None and not DIGEST.fullmatch(target['reuse_digest']):
raise ValueError('Image plan has an invalid reuse digest')
return data
def build_images(output, report, name, plan):
data = checked_plan(plan)
sha = data['sha']
target = next(t for t in data['targets'] if t['name'] == name)
context, dockerfile = IMAGES[name]
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
builder_config = Path.home() / '.cache/homelab-ci/buildx'
builder_config.mkdir(parents=True, exist_ok=True)
env = {**os.environ, 'DOCKER_CONFIG': docker_config, 'BUILDX_CONFIG': str(builder_config)}
try:
report['phase'] = 'Registry login'
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'login',
'gcr.forust.xyz',
'-u',
os.environ['REGISTRY_USERNAME'],
'--password-stdin',
],
input=os.environ['REGISTRY_PASSWORD'],
text=True,
check=True,
env=env,
)
report['phase'] = 'Prepare the builder'
builder = 'homelab-ci'
versions = dict(
re.findall(r'^([A-Z_]+)="([^"\n]+)"$', Path('.gitea/workflows/tool-versions.env').read_text(), re.MULTILINE)
)
image = versions['BUILDKIT_IMAGE']
signature = builder_config / 'homelab-ci-image'
exists = (
subprocess.run( # noqa: S603
[shutil.which('docker') or '/usr/bin/docker', 'buildx', 'inspect', builder],
capture_output=True,
env=env,
).returncode
== 0
)
if exists and (not signature.exists() or signature.read_text().strip() != image):
command('docker', 'buildx', 'rm', '--keep-state', builder, env=env)
exists = False
if not exists:
command(
'docker',
'buildx',
'create',
'--name',
builder,
'--driver',
'docker-container',
'--driver-opt',
f'image={image}',
'--buildkitd-config',
'.gitea/runner/buildkitd.toml',
env=env,
)
signature.write_text(image + '\n')
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
report['images'] = release['images']
report['phase'] = f'Build or reuse {name}'
report['current'] = name
image = f'gcr.forust.xyz/forust/{name}'
inputs = target['inputs']
old_digest = target['reuse_digest']
exists = False
if old_digest:
exists = (
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'buildx',
'imagetools',
'inspect',
f'{image}@{old_digest}',
],
capture_output=True,
env=env,
timeout=60,
).returncode
== 0
)
if exists:
print(f'Reuse {name}: inputs unchanged')
digest = old_digest
else:
print(f'Build {name}', flush=True)
metadata = Path(docker_config) / 'metadata.json'
command(
'docker',
'buildx',
'build',
'--builder',
builder,
'--platform',
'linux/amd64',
'--provenance=false',
'--cache-from',
f'type=registry,ref={image}:buildcache',
'--cache-to',
f'type=registry,ref={image}:buildcache,mode=max',
'--output',
f'type=image,name={image},push-by-digest=true,name-canonical=true,push=true',
'--metadata-file',
str(metadata),
'--file',
dockerfile,
context,
env=env,
)
digest = json.loads(metadata.read_text())['containerimage.digest']
if not isinstance(digest, str) or not DIGEST.fullmatch(digest):
raise ValueError('Image job returned an invalid digest')
release['images'][image] = digest
release['inputs'][image] = inputs
report['reused' if exists else 'built'].append(name)
output.write_text(json.dumps(release, indent=2) + '\n')
report['current'] = None
report['phase'] = 'Image result file saved'
finally:
# Cleanup errors must neither leak credentials nor mask the original build error.
try:
subprocess.run( # noqa: S603
[
shutil.which('docker') or '/usr/bin/docker',
'buildx',
'prune',
'--builder',
'homelab-ci',
'--force',
'--max-used-space',
'1gb',
],
env=env,
timeout=60,
)
except (OSError, subprocess.TimeoutExpired):
print('CI builder cache cleanup deferred', flush=True)
finally:
shutil.rmtree(docker_config)
def write_summary(lines):
path = os.environ.get('GITHUB_STEP_SUMMARY')
if path:
try:
with Path(path).open('a') as stream:
stream.write('\n'.join(lines) + '\n\n')
except OSError:
print('WARNING: cannot write the job summary')
def check_summary():
lines = [
f'## {os.environ["SUMMARY_CHECK"]}',
'',
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
f'- Result: **{os.environ["SUMMARY_RESULT"]}**',
]
if os.environ.get('SUMMARY_FAILED_STEP'):
lines.append(f'- Failed step: {os.environ["SUMMARY_FAILED_STEP"]}')
if os.environ['SUMMARY_RESULT'] != 'success':
lines.append('- Open the failed step log for the error details.')
write_summary(lines)
def build(output, name, plan):
report = {'phase': 'Check the source commit', 'current': None, 'built': [], 'reused': [], 'images': {}}
result = 'failure'
try:
build_images(output, report, name, plan)
result = 'success'
finally:
lines = [
f'## Image build result `{name}`',
'',
f'- Commit: `{os.environ.get("GITHUB_SHA", "unknown")}`',
'',
f'- Result: **{result}**',
f'- Last stage: {report["phase"]}',
]
if result == 'failure':
lines.append('- This image job failed. The complete release cannot be published. Open the failed step log.')
if result == 'success':
lines.append('- This is one image result. The final build job must publish the complete release.')
if report['current']:
lines.append(f'- Image at the failure: `{report["current"]}`')
for title, key in (('Built', 'built'), ('Reused from successful CI', 'reused')):
lines.extend(['', f'### {title}'])
lines.extend(f'- `{name}`' for name in report[key])
if not report[key]:
lines.append('- None')
lines.extend(['', '### Completed image digests'])
lines.extend(f'- `{image}@{digest}`' for image, digest in report['images'].items())
if not report['images']:
lines.append('- None')
write_summary(lines)
def render(stream, destination):
release = validate_release(json.loads(Path(os.environ['RELEASE_FILE']).read_text()), os.environ['DEPLOY_SHA'])
image_line = re.compile(
r"^(\s*(?:-\s*)?image:\s*)(['\"]?)(gcr\.forust\.xyz/forust/[\w.-]+)(?::[\w.-]+|@sha256:[0-9a-f]{64})\2(\s*(?:#.*)?)$"
)
rendered = []
for line in stream:
match = image_line.fullmatch(line.rstrip('\n'))
if match:
prefix, quote, image, tail = match.groups()
if image not in release['images']:
raise ValueError(f'Owned image missing from checked release: {image}')
line = f'{prefix}{quote}{image}@{release["images"][image]}{quote}{tail}\n'
elif re.match(r'\s*(?:-\s*)?image:', line) and 'gcr.forust.xyz/forust/' in line:
raise ValueError('Unsupported owned image syntax; refusing to apply a mutable tag')
rendered.append(line)
destination.writelines(rendered)
def finalize_images(output, fragments, plan):
data = checked_plan(plan)
sha = data['sha']
release = {'version': 1, 'sha': sha, 'images': {}, 'inputs': {}}
for name in IMAGES:
fragment = json.loads((fragments / f'image-{name}' / 'image.json').read_text())
image = f'gcr.forust.xyz/forust/{name}'
if fragment.get('sha') != sha or fragment.get('version') != 1 or set(fragment.get('images', {})) != {image}:
raise ValueError('Image job artifact is missing or belongs to another commit')
target = next(t for t in data['targets'] if t['name'] == name)
if fragment.get('inputs') != {image: target['inputs']}:
raise ValueError('Image artifact does not match the build plan')
release['images'].update(fragment['images'])
release['inputs'].update(fragment['inputs'])
validate_release(release, sha)
# Only a complete set of successful image jobs can publish the release tags.
docker_config = tempfile.mkdtemp(prefix='homelab-registry-')
env = {**os.environ, 'DOCKER_CONFIG': docker_config}
try:
subprocess.run( # noqa: S603, S607
[
shutil.which('docker') or '/usr/bin/docker',
'login',
'gcr.forust.xyz',
'-u',
os.environ['REGISTRY_USERNAME'],
'--password-stdin',
],
input=os.environ['REGISTRY_PASSWORD'],
text=True,
check=True,
env=env,
)
for image, digest in release['images'].items():
command(
'docker',
'buildx',
'imagetools',
'create',
'--prefer-index=false',
'--tag',
f'{image}:sha-{sha}',
f'{image}@{digest}',
env=env,
timeout=90,
)
output.write_text(json.dumps(release, indent=2) + '\n')
finally:
shutil.rmtree(docker_config)
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=('prepare', 'image', 'finalize', 'gate', 'render', 'check-summary'))
parser.add_argument('--output', type=Path, default=Path('release.json'))
parser.add_argument('--ref', default='main')
parser.add_argument('--event-sha', default='')
parser.add_argument('--image', choices=IMAGES)
parser.add_argument('--plan', type=Path, default=Path('build-plan.json'))
parser.add_argument('--fragments', type=Path, default=Path('artifacts'))
args = parser.parse_args()
if args.action == 'check-summary':
check_summary()
elif args.action == 'render':
render(sys.stdin, sys.stdout)
elif args.action == 'gate':
gate(args.output, args.ref, args.event_sha)
elif args.action == 'prepare':
prepare_images(args.output)
elif args.action == 'image':
if not args.image:
parser.error('--image is required')
build(args.output, args.image, args.plan)
else:
finalize_images(args.output, args.fragments, args.plan)
if __name__ == '__main__':
main()
+21 -69
View File
@@ -1,88 +1,39 @@
name: renovate-ci name: renovate-ci
on: on:
# Read the workflow from the trusted base branch. PR code runs only on the pull_request:
# unprivileged runner selected below.
pull_request_target:
paths:
- "renovate/**"
- ".gitea/workflows/renovate-ci.yaml"
- ".gitea/workflows/sync-renovate-configmap.sh"
- ".gitea/workflows/compose-lint.sh"
- ".gitea/workflows/install-ci-tools.sh"
- ".gitea/workflows/tool-versions.env"
push: push:
branches: branches:
- main - main
paths:
- "renovate/**"
- ".gitea/workflows/renovate-ci.yaml"
- ".gitea/workflows/sync-renovate-configmap.sh"
- ".gitea/workflows/compose-lint.sh"
- ".gitea/workflows/install-ci-tools.sh"
- ".gitea/workflows/tool-versions.env"
workflow_dispatch: workflow_dispatch:
permissions:
contents: read
jobs: jobs:
validate-renovate: validate-renovate:
runs-on: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' && 'homelab' || 'homelab-pr' }} runs-on: [self-hosted, linux, arch, homelab]
timeout-minutes: 20
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: ${{ github.event_name == 'pull_request_target' && github.event.pull_request.head.sha || github.sha }}
# renovate/k8s/cronjob.yaml is the single source of truth for the version. - name: Validate Renovate Compose draft
- name: Resolve the deployed Renovate version
id: image
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \ trap 'rm -f renovate/.env' EXIT
renovate/k8s/cronjob.yaml | head -1)" printf '%s\n' \
if [[ ! "$image" =~ ^renovate/renovate:([0-9]+\.[0-9]+\.[0-9]+)$ ]]; then 'RENOVATE_ENDPOINT=https://gitea.example/api/v1' \
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml" 'RENOVATE_TOKEN=test-token' \
exit 1 'RENOVATE_REPOSITORIES=forust/homelab' \
fi > renovate/.env
version="${BASH_REMATCH[1]}" docker compose -f renovate/renovate-compose.yaml config --quiet
echo "using Renovate $version"
printf 'version=%s\n' "$version" >> "$GITHUB_OUTPUT"
- name: Prepare pinned validation tools - name: Validate Kubernetes manifests
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
tools_dir="$(bash .gitea/workflows/install-ci-tools.sh kubeconform node)" docker run --rm \
echo "$tools_dir" >> "$GITHUB_PATH" -v "$PWD:/work" \
-w /work \
- name: Validate Renovate repository config ghcr.io/yannh/kubeconform:latest \
shell: bash
env:
RENOVATE_VERSION: ${{ steps.image.outputs.version }}
run: |
set -euo pipefail
npm_cache="$(mktemp -d "${RUNNER_TEMP:-/tmp}/renovate-npm-cache.XXXXXXXX")"
trap 'rm -rf "$npm_cache"' EXIT
NPM_CONFIG_CACHE="$npm_cache" RENOVATE_CONFIG_FILE="$PWD/renovate/renovate.json" \
npm exec --yes --package="renovate@${RENOVATE_VERSION}" -- renovate-config-validator
# The CronJob cannot read the repository, so renovate/k8s/configmap.yaml
# carries an inlined copy of the config. Fail if it no longer matches.
- name: Check the generated Renovate ConfigMap
shell: bash
run: |
set -euo pipefail
./.gitea/workflows/sync-renovate-configmap.sh --check
- name: Validate Renovate Kubernetes manifests
shell: bash
run: |
set -euo pipefail
kubeconform \
-strict \ -strict \
-ignore-missing-schemas \ -ignore-missing-schemas \
-summary \ -summary \
@@ -90,11 +41,12 @@ jobs:
renovate/k8s/configmap.yaml \ renovate/k8s/configmap.yaml \
renovate/k8s/cronjob.yaml renovate/k8s/cronjob.yaml
- name: Validate Renovate Compose file - name: Validate Renovate repository config
shell: bash shell: bash
run: | run: |
set -euo pipefail set -euo pipefail
source .gitea/workflows/compose-lint.sh docker run --rm \
mapfile -t safe_flags < <(compose_safe_flags) -v "$PWD:/work" \
validate_compose_file renovate/renovate-compose.yaml \ -w /work \
${safe_flags[@]+"${safe_flags[@]}"} renovate/renovate:44.103.0 \
renovate-config-validator renovate.json
+8 -37
View File
@@ -21,53 +21,25 @@ on:
default: false default: false
type: boolean type: boolean
# Renovate writes through its own bot PAT, passed in as RENOVATE_TOKEN, so the
# Actions token is only ever used to read the checkout.
permissions:
contents: read
concurrency: concurrency:
group: renovate-run group: renovate-run
cancel-in-progress: false cancel-in-progress: false
jobs: jobs:
run-renovate: run-renovate:
if: github.ref == 'refs/heads/main' runs-on: [self-hosted, linux, arch, homelab]
runs-on: homelab
timeout-minutes: 60
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4 uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
with:
ref: refs/heads/main
# renovate/k8s/cronjob.yaml is the single source of truth for the image tag.
# Reading it here means this workflow validates and runs the exact version
# that is deployed, instead of a copy that silently goes stale.
- name: Resolve the deployed Renovate image
id: image
shell: bash
run: |
set -euo pipefail
image="$(sed -n 's|.*image:[[:space:]]*\(renovate/renovate:[^[:space:]]*\).*|\1|p' \
renovate/k8s/cronjob.yaml | head -1)"
if [[ ! "$image" =~ ^renovate/renovate:[0-9]+\.[0-9]+\.[0-9]+$ ]]; then
echo "::error::expected a pinned renovate/renovate semantic version in renovate/k8s/cronjob.yaml"
exit 1
fi
echo "using $image"
printf 'image=%s\n' "$image" >> "$GITHUB_OUTPUT"
- name: Validate Renovate config - name: Validate Renovate config
shell: bash shell: bash
env:
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
run: | run: |
set -euo pipefail set -euo pipefail
docker run --rm \ docker run --rm \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \ -v "$PWD/renovate/config.js:/opt/renovate/config.js:ro" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \ -e RENOVATE_CONFIG_FILE=/opt/renovate/config.js \
"$RENOVATE_IMAGE" \ renovate/renovate:44.103.0 \
renovate-config-validator renovate-config-validator
- name: Run Renovate - name: Run Renovate
@@ -78,21 +50,20 @@ jobs:
RENOVATE_REPOSITORIES: ${{ inputs.repositories }} RENOVATE_REPOSITORIES: ${{ inputs.repositories }}
RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }} RENOVATE_DRY_RUN: ${{ inputs.dry_run && 'full' || '' }}
LOG_LEVEL: ${{ inputs.log_level }} LOG_LEVEL: ${{ inputs.log_level }}
RENOVATE_IMAGE: ${{ steps.image.outputs.image }}
run: | run: |
set -euo pipefail set -euo pipefail
: "${RENOVATE_TOKEN:?missing RENOVATE_TOKEN secret — add a renovate-bot PAT in repo/org Actions secrets}" : "${RENOVATE_TOKEN:?missing RENOVATE_TOKEN secret — add a renovate-bot PAT in repo/org Actions secrets}"
docker run --rm \ docker run --rm \
-v "$PWD/renovate/renovate.json:/opt/renovate/renovate.json:ro" \ -v "$PWD/renovate/config.js:/opt/renovate/config.js:ro" \
-e RENOVATE_PLATFORM=gitea \ -e RENOVATE_PLATFORM=gitea \
-e RENOVATE_ENDPOINT=https://git.forust.xyz/api/v1 \ -e RENOVATE_ENDPOINT=https://gitea.forust.xyz/api/v1 \
-e RENOVATE_TOKEN="$RENOVATE_TOKEN" \ -e RENOVATE_TOKEN="$RENOVATE_TOKEN" \
-e RENOVATE_GITHUB_COM_TOKEN="${RENOVATE_GITHUB_COM_TOKEN:-}" \ -e RENOVATE_GITHUB_COM_TOKEN="${RENOVATE_GITHUB_COM_TOKEN:-}" \
-e RENOVATE_REPOSITORIES="${RENOVATE_REPOSITORIES:-forust/homelab}" \ -e RENOVATE_REPOSITORIES="${RENOVATE_REPOSITORIES:-forust/homelab}" \
-e RENOVATE_DRY_RUN="${RENOVATE_DRY_RUN:-}" \ -e RENOVATE_DRY_RUN="${RENOVATE_DRY_RUN:-}" \
-e RENOVATE_CONFIG_FILE=/opt/renovate/renovate.json \ -e RENOVATE_CONFIG_FILE=/opt/renovate/config.js \
-e RENOVATE_BASE_DIR=/tmp/renovate \ -e RENOVATE_BASE_DIR=/tmp/renovate \
-e LOG_LEVEL="${LOG_LEVEL:-info}" \ -e LOG_LEVEL="${LOG_LEVEL:-info}" \
"$RENOVATE_IMAGE" renovate/renovate:44.103.0
-13
View File
@@ -1,13 +0,0 @@
# kubectl emits a List for files containing multiple resources.
(if .kind == "List" then .items[] else . end)
| (.metadata.namespace // "default") as $ns
| [
(.. | objects
| (.secretRef? // empty), (.secretKeyRef? // empty), (.secret? // empty)
| select(.optional != true)
| .name // .secretName // empty),
(.. | objects | .imagePullSecrets[]?.name)
]
| unique[]
| select(. != null and . != "")
| "\($ns) \(.)"
+20 -65
View File
@@ -1,70 +1,25 @@
#!/usr/bin/env bash #!/usr/bin/env bash
# The SSH client submits once and follows durable stages on workstation. # usage: ssh-run.sh <stage>
# Runs one deploy-lib.sh stage on the workstation over SSH.
set -euo pipefail set -euo pipefail
: "${DEPLOY_HOST:?missing DEPLOY_HOST}" : "${DEPLOY_HOST:?missing DEPLOY_HOST}"
: "${DEPLOY_USER:?missing DEPLOY_USER}" : "${DEPLOY_USER:?missing DEPLOY_USER}"
: "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}" : "${DEPLOY_KEY:?missing DEPLOY_SSH_KEY}"
: "${DEPLOY_KNOWN_HOSTS:?configure pinned DEPLOY_KNOWN_HOSTS}"
: "${DEPLOY_RUN_ID:?missing DEPLOY_RUN_ID}" deploy_port="${DEPLOY_PORT:-22}"
[[ "$DEPLOY_USER" =~ ^[A-Za-z_][A-Za-z0-9_.-]*$ ]] || exit 1 deploy_path="${DEPLOY_PATH:-/srv/homelab}"
[[ "$DEPLOY_HOST" =~ ^[A-Za-z0-9_.:-]+$ ]] || exit 1 deploy_path="$(printf '%s' "$deploy_path" | tr -d '\"' | tr -d '\r' | xargs)"
[[ "$DEPLOY_RUN_ID" =~ ^[0-9]+-[0-9]+$ ]] || exit 1
[[ "${DEPLOY_PORT:-22}" =~ ^[0-9]+$ ]] || exit 1 ssh_key="$RUNNER_TEMP/deploy_key"
key_dir="$(mktemp -d "${RUNNER_TEMP:-/tmp}/homelab-deploy-key.XXXXXXXX")" mkdir -p "$RUNNER_TEMP"
trap 'rm -rf "$key_dir"' EXIT printf '%s\n' "$DEPLOY_KEY" > "$ssh_key"
chmod 700 "$key_dir" chmod 600 "$ssh_key"
printf '%s\n' "$DEPLOY_KEY" >"$key_dir/key"
printf '%s\n' "$DEPLOY_KNOWN_HOSTS" >"$key_dir/known_hosts" ssh -i "$ssh_key" -p "$deploy_port" \
chmod 600 "$key_dir/key" "$key_dir/known_hosts" -o BatchMode=yes -o StrictHostKeyChecking=accept-new \
ssh_opts=(-i "$key_dir/key" -p "${DEPLOY_PORT:-22}" -o BatchMode=yes -o StrictHostKeyChecking=yes "${DEPLOY_USER}@${DEPLOY_HOST}" \
-o "UserKnownHostsFile=$key_dir/known_hosts" -o ConnectTimeout=15 "REPO=$deploy_path APPLY_PRUNE=${APPLY_PRUNE:-false} STAGE=$1 bash -se" <<'EOF'
-o ServerAliveInterval=15 -o ServerAliveCountMax=4) source "$REPO/.gitea/workflows/deploy-lib.sh"
controller=.local/lib/homelab-deploy/controller.py run_stage "$STAGE"
case "${1:?start, apply, verify, smoke or summary required}" in EOF
start)
python3 - <<'PY' >"$key_dir/request.json"
import json
import os
from pathlib import Path
release = json.loads(Path('release.json').read_text())
print(json.dumps({'release': release, 'mode': os.environ.get('DEPLOY_MODE', 'changed'),
'refresh_images': os.environ.get('REFRESH_IMAGES', 'false') == 'true'}))
PY
for attempt in 1 2 3; do
rc=0
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" start "$DEPLOY_RUN_ID" <"$key_dir/request.json" || rc=$?
[ "$rc" -eq 0 ] && exit 0
[ "$rc" -eq 255 ] || exit "$rc"
sleep 5
done
exit "$rc"
;;
apply|verify|smoke)
result=0
for attempt in 1 2 3; do
rc=0
# shellcheck disable=SC2029 # The run ID and operation are validated local arguments, not remote variables.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" follow "$DEPLOY_RUN_ID" "$1" || rc=$?
[ "$rc" -eq 0 ] && break
[ "$rc" -eq 255 ] || { result="$rc"; break; }
echo "SSH disconnected; reconnecting to the existing deploy ($attempt/3)"
if [ "$attempt" -eq 3 ]; then result=255; break; fi
sleep 5
done
exit "$result"
;;
summary)
if [ -n "${GITHUB_STEP_SUMMARY:-}" ]; then
rc=0
# shellcheck disable=SC2029 # The run ID is validated above.
ssh "${ssh_opts[@]}" "$DEPLOY_USER@$DEPLOY_HOST" python3 "$controller" summary "$DEPLOY_RUN_ID" >"$key_dir/deploy-summary.md" || rc=$?
if [ "$rc" -eq 0 ]; then
cat "$key_dir/deploy-summary.md" >>"$GITHUB_STEP_SUMMARY" || echo "WARNING: cannot write the deploy summary"
else
echo 'Deploy summary is unavailable. The SSH connection failed or the controller did not respond. Check the job log.' >>"$GITHUB_STEP_SUMMARY" || true
fi
fi
;;
*) echo "Unknown SSH operation: $1" >&2; exit 1 ;;
esac
@@ -1,55 +0,0 @@
#!/usr/bin/env bash
# Regenerates renovate/k8s/configmap.yaml from renovate/renovate.json.
#
# renovate/renovate.json is the single source of truth: the CronJob, the Compose
# file and the renovate-run workflow all mount that exact file. A ConfigMap cannot
# read a file from the repository, so the same bytes are inlined here as a literal
# block. This script keeps the copy honest:
#
# .gitea/workflows/sync-renovate-configmap.sh # rewrite in place
# .gitea/workflows/sync-renovate-configmap.sh --check # fail if out of date
#
# renovate-ci runs the --check form on every PR and push, so a config change that
# forgets to regenerate the ConfigMap cannot be merged.
set -euo pipefail
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
repo="$(git -C "$here" rev-parse --show-toplevel)"
src="$repo/renovate/renovate.json"
dst="$repo/renovate/k8s/configmap.yaml"
[ -f "$src" ] || {
echo "missing $src" >&2
exit 1
}
render() {
cat <<'HEADER'
# GENERATED FILE - do not edit by hand.
# Source: renovate/renovate.json
# Regenerate: .gitea/workflows/sync-renovate-configmap.sh
# Verify: .gitea/workflows/sync-renovate-configmap.sh --check
apiVersion: v1
kind: ConfigMap
metadata:
name: renovate-config
namespace: renovate
data:
renovate.json: |
HEADER
sed 's/^/ /' "$src"
}
if [ "${1:-}" = "--check" ]; then
if ! diff -u "$dst" <(render) >/dev/null 2>&1; then
echo "ERROR: $dst is out of sync with renovate/renovate.json"
echo "Run: .gitea/workflows/sync-renovate-configmap.sh"
diff -u "$dst" <(render) || true
exit 1
fi
echo "renovate/k8s/configmap.yaml is in sync with renovate/renovate.json"
exit 0
fi
render >"$dst"
echo "wrote $dst"
-39
View File
@@ -1,39 +0,0 @@
# Pinned versions of the CI tools installed by install-ci-tools.sh.
# Renovate keeps these up to date (see customManagers in renovate/renovate.json).
#
# Every version here except NODE_VERSION matches what was already installed on
# the runner, so pinning them changes what CI does not at all. It changes what
# CI does when the runner is rebuilt with something else: today
# install-ci-tools.sh finds the pinned version already on PATH and installs
# nothing, and a runner that drifts gets the pinned one installed over it.
#
# The renovate image version is NOT pinned here: renovate/k8s/cronjob.yaml is the
# single source of truth and the workflows read the tag from it, so there is
# nothing to drift.
ACTIONLINT_VERSION="1.7.7"
SHELLCHECK_VERSION="0.11.0"
KUBECONFORM_VERSION="0.8.0"
PRETTIER_VERSION="3.8.1"
RUFF_VERSION="0.16.10"
YAMLLINT_VERSION="1.38.0"
HADOLINT_VERSION="2.14.0"
# pip-audit reads the advisory database over the network, so a floating version
# would make the same commit report different things on different days. Pin it
# like the rest: the advisories themselves are the moving part, not the tool.
PIP_AUDIT_VERSION="2.10.1"
# uv builds the throwaway venv the pytest job runs in, and unpacks the PyPI
# wheels for ruff, yamllint and pip-audit.
UV_VERSION="0.12.17"
# node runs `npm ci` for the frontend tests and the npm audit, and it is the one
# pin here that does NOT come from the runner: the runner's system node is a
# rolling Arch package (it was node 26 with no npm at all when this was pinned),
# and the panel image is node:22-alpine. Pinned to the image's major on purpose,
# so the tree that gets tested is the tree that gets built. Renovate keeps this
# in step with the Dockerfile's node: tag via the "node runtime" group.
NODE_VERSION="22.23.3"
# Secret-reference regression tests parse rendered Kubernetes objects.
JQ_VERSION="1.8.1"
# BuildKit is the only auxiliary CI container; jobs themselves stay on the host.
BUILDKIT_IMAGE="moby/buildkit:v0.33.1"
+1 -5
View File
@@ -21,9 +21,6 @@ checkmk/checkmk/*
downtify/Downtify_downloads downtify/Downtify_downloads
headscale/config/* headscale/config/*
headscale/data/* headscale/data/*
# NetBird local hostnames and generated secrets
netbird/.env
netbird/secrets/
searxng/core-config/* searxng/core-config/*
# Steaming services files # Steaming services files
@@ -94,9 +91,8 @@ replacements.txt
.idea .idea
# Temp files # Temp files
edu_master/temp/
temp/* temp/*
# Local-only tooling scratch space (pinned CI tools, verification scripts)
tmp/
# Environment # Environment
.env .env
-163
View File
@@ -1,163 +0,0 @@
# Homelab
Configuration for my homelab: Kubernetes manifests, Docker Compose stacks, and the
Gitea Actions that build and deploy them. Most applications have both deployment
formats. Headscale and Nextcloud AIO have Compose deployments with Kubernetes
ingress; the media stack has Compose and Kubernetes routing configuration.
These files contain this lab's domains, IP addresses, storage paths, and private
registry names. Running them on another machine takes some editing.
## Start here
- [Service list](#services) — what each directory contains.
- [Deployment workflow](.gitea/README.md) — selection, validation, and recovery.
- [Repository review](docs/repository-review.md) — findings from the 6 October baseline and their status.
- [EDU ownership handoff](.gitea/EDU_HANDOFF.md) — the EDU workloads now live in their own repository.
- [Shared PostgreSQL](postgres/README.md), [Traefik](traefik/README.md), and
[cert-manager](cert-manager/README.md) — common dependencies.
## What gets deployed
The `active` files are switches for the deploy workflow, not health indicators.
| File | Effect |
| ---------------------- | ----------------------------------------------------------- |
| `<service>/active` | Include that directory's `compose.yaml` or `compose.yml`. |
| `<service>/k8s/active` | Include its Kubernetes manifests or Kustomize overlay. |
| Both | Run the Compose stack and apply the Kubernetes resources. |
| Neither | Keep the configuration in Git without automatic deployment. |
`shared-compose.yaml`, `client.compose.yaml`, and `renovate-compose.yaml` are
manual entry points. The deploy script does not discover them.
Kubernetes selection excludes secret files, examples, Helm values, and patches.
Helm releases listed in `deploy-lib.sh` are upgraded separately. Traefik,
cert-manager, and CrowdSec have additional bootstrap steps; an `active` marker
does not install their charts.
The table below describes committed configuration. It does not claim that a
service is currently healthy or running.
## Services
| Service | Configuration | Selected by markers |
| ---------------------------------------------- | ---------------------------- | ------------------- |
| [AdGuard Home](adguardhome/README.md) | Kubernetes + Compose | Kubernetes |
| [Authentik](authentik/README.md) | Kubernetes + Compose | Kubernetes |
| [cert-manager](cert-manager/README.md) | Kubernetes / Helm | Manual |
| [Cloudflare DDNS](cfddns/README.md) | Kubernetes + Compose | Kubernetes |
| [Checkmk](checkmk/README.md) | Kubernetes + Compose | Manual |
| [Cloudflare Tunnel](cloudflared/README.md) | Kubernetes / Helm | Manual |
| [File converters](converters/README.md) | Kubernetes + Compose | Kubernetes |
| [CrowdSec](crowdsec/README.md) | Kubernetes / Helm | Manual |
| [Dockmon](dockmon/README.md) | Kubernetes + Compose | Manual |
| [Downtify](downtify/README.md) | Kubernetes + Compose | Manual |
| [Error pages](errorpages/README.md) | Kubernetes + Compose | Kubernetes |
| [Gitea](gitea/README.md) | Kubernetes + Compose | Kubernetes |
| [Glance](glance/README.md) | Kubernetes + Compose | Manual |
| [Headscale](headscale/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Homarr](homarr/README.md) | Kubernetes + Compose | Manual |
| [Homepages](homepages/README.md) | Kubernetes + Compose | Kubernetes |
| [Immich](immich/README.md) | Kubernetes + Compose | Kubernetes |
| [Kener](kener/README.md) | Kubernetes + Compose | Manual |
| [Loki and Alloy](loki/README.md) | Kubernetes / Helm | Kubernetes |
| [MeTube](metube/README.md) | Kubernetes + Compose | Kubernetes |
| [n8n](n8n/README.md) | Kubernetes + Compose | Manual |
| [NetBird](netbird/README.md) | Kubernetes + Compose | Kubernetes |
| [NetBox](netbox/README.md) | Kubernetes + Compose | Kubernetes |
| [Netronome](netronome/README.md) | Kubernetes + Compose | Kubernetes |
| [Nextcloud AIO](nextcloud/README.md) | Compose + Kubernetes routing | Compose, Kubernetes |
| [Penpot](penpot/README.md) | Compose | Manual |
| [Portainer](portainer/README.md) | Kubernetes + Compose | Manual |
| [Shared PostgreSQL](postgres/README.md) | Kubernetes + Compose | Kubernetes |
| [Monitoring stack](prometheus-stack/README.md) | Kubernetes + Compose | Kubernetes |
| [RackPeek](rackpeek/README.md) | Kubernetes + Compose | Kubernetes |
| [Reloader](reloader/README.md) | Kubernetes / Helm | Kubernetes |
| [Renovate](renovate/README.md) | Kubernetes + Compose | Kubernetes |
| [SearXNG](searxng/README.md) | Kubernetes + Compose | Manual |
| [Media stack](streaming/README.md) | Compose + Kubernetes routing | Manual |
| [Termix](termix/README.md) | Kubernetes + Compose | Manual |
| [Traefik](traefik/README.md) | Kubernetes + Compose | Kubernetes |
| [Uptime Kuma](uptime-kuma/README.md) | Kubernetes + Compose | Kubernetes |
| [Vaultwarden](vaultwarden/README.md) | Kubernetes + Compose | Kubernetes |
| [3x-ui](vpn/xui/README.md) | Kubernetes | Kubernetes |
## Running a Compose stack
Use the service README first. Where a service has an env example, copy it inside
that service's directory and replace the placeholders. The root `.env.example`
is an older collection of variables, not a complete configuration for every stack.
For example, from the repository root:
```sh
cd netbox
cp .env.example .env
$EDITOR .env
docker compose config --quiet
docker compose up -d
docker compose ps
```
Stacks that attach to `proxy` require an existing Docker network of that name and
an appropriate reverse proxy. Published host ports still work independently of
Traefik. Check port conflicts before starting an alternative to a Kubernetes
service: DNS, STUN, and HTTP listeners can share the same host.
`docker compose down` keeps named volumes. Adding `-v` removes them.
## Preparing Kubernetes
The manifests assume Traefik CRDs, cert-manager, and a working storage provisioner.
PrometheusRule and ServiceMonitor resources also need the Prometheus Operator.
Replace the lab's hosts and addresses before using the configuration elsewhere.
Create a service's namespace, then prepare its ignored Secret from the example.
For example:
```sh
kubectl apply -f netbox/k8s/namespace.yaml
cp netbox/k8s/secrets.yaml.example netbox/k8s/secrets.yaml
$EDITOR netbox/k8s/secrets.yaml
kubectl apply -f netbox/k8s/secrets.yaml
```
The deploy workflow applies the tracked resources for marked services. Avoid
applying an entire `k8s/` directory blindly: some directories contain Helm values,
examples, and alternative routes. For a manual change, apply the selected manifest
explicitly and check the resulting rollout.
Shared database passwords must agree between the `database` namespace and each
application's Secret. Updating the PostgreSQL Secret does not change an existing
role's password; see the database README.
## Local checks
CI pins its tools in `.gitea/workflows/tool-versions.env`. Use the same versions:
```fish
set tools_dir (bash .gitea/workflows/install-ci-tools.sh)
set -gx PATH $tools_dir $PATH
ruff check .
ruff format --check .
actionlint -config-file .gitea/actionlint.yaml .gitea/workflows/*.yaml
.gitea/workflows/sync-renovate-configmap.sh --check
```
The [workflow README](.gitea/README.md#ci) lists the rest of the checks.
Structure checks do not establish that local Secrets, mounted files, storage,
or external services are ready.
## Data and recovery
State lives outside Git: PVCs, Docker volumes, bind mounts, databases, and ignored
configuration. Keep backups of application data and the keys needed to read it.
An image rollback does not roll back database migrations or ConfigMap contents.
Many PVCs use the cluster's default StorageClass; monitoring explicitly uses
`local-path`. Check the PV reclaim policy before deleting a PVC or namespace.
The manifests do not provide a repository-wide backup schedule.
`incident-archive/` contains past incident notes. `.docs/storage-audit-instruction.md`
is a planning document, not evidence that NFS has been installed.
-22
View File
@@ -1,22 +0,0 @@
# AdGuard Home
DNS filtering with a web UI, DNS-over-TLS, and certificates from cert-manager.
The Kubernetes namespace is `adguard`. The workload uses `adguard-pvc` for
configuration and working data, and mounts the `adguard-certs` TLS Secret.
The LoadBalancer Service exposes DNS separately from the web ingress.
The Compose stack publishes TCP/UDP 53 and TCP 853 on the host. Prepare `conf/`
and `certs/` before starting it. Starting both DNS deployments on the same address
can cause a port conflict.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n adguard
kubectl get events -n adguard --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -31,7 +31,7 @@ services:
- "traefik.http.routers.adguard-dev.entrypoints=websecure" - "traefik.http.routers.adguard-dev.entrypoints=websecure"
- "traefik.http.routers.adguard-dev.tls=true" - "traefik.http.routers.adguard-dev.tls=true"
# DoH Router # DoH Router
- "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz`) || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`)" - "traefik.http.routers.dns-over-https.rule=(Host(`dns.forust.xyz` || Host(`adguard.forust.xyz`)) && PathPrefix(`/dns-query`))"
- "traefik.http.routers.dns-over-https.entrypoints=websecure" - "traefik.http.routers.dns-over-https.entrypoints=websecure"
- "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt" - "traefik.http.routers.dns-over-https.tls.certresolver=letsencrypt"
+3 -12
View File
@@ -51,8 +51,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: adguard-deployment name: adguard-deployment
namespace: adguard namespace: adguard
spec: spec:
@@ -60,12 +58,12 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: adguard app: adguard
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
app: adguard app: adguard
annotations:
reloader.stakater.com/auto: "true"
spec: spec:
containers: containers:
- name: adguard - name: adguard
@@ -75,7 +73,7 @@ spec:
memory: "1.5Gi" memory: "1.5Gi"
cpu: "300m" cpu: "300m"
requests: requests:
memory: "512Mi" memory: "500Mi"
cpu: "50m" cpu: "50m"
ports: ports:
- containerPort: 3000 - containerPort: 3000
@@ -84,13 +82,6 @@ spec:
name: dns name: dns
- containerPort: 853 - containerPort: 853
name: dot name: dot
readinessProbe:
tcpSocket:
port: dns
initialDelaySeconds: 5
periodSeconds: 5
successThreshold: 1
failureThreshold: 3
volumeMounts: volumeMounts:
- name: adguard-data - name: adguard-data
mountPath: /opt/adguardhome/work mountPath: /opt/adguardhome/work
+3
View File
@@ -9,6 +9,9 @@ spec:
routes: routes:
- match: Host(`dns.forust.xyz`) - match: Host(`dns.forust.xyz`)
kind: Rule kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: adguard-service - name: adguard-service
port: 3000 port: 3000
-22
View File
@@ -1,22 +0,0 @@
# Authentik
Identity provider with separate server and worker deployments.
Kubernetes connects to the shared PostgreSQL service in `database`. Set
`AUTHENTIK_DB_PASSWORD` to the same value in both database and application Secrets.
Keep `AUTHENTIK_SECRET_KEY` with the backups.
Compose uses its own PostgreSQL 15 container and bind-mounted media and templates.
Its image defaults differ from Kubernetes; check both before an upgrade.
The worker mounts the Docker socket for Docker outpost management.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n authentik
kubectl get events -n authentik --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+5 -13
View File
@@ -27,8 +27,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: authentik-server-deployment name: authentik-server-deployment
namespace: authentik namespace: authentik
spec: spec:
@@ -36,8 +34,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: authentik-server app: authentik-server
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -56,8 +52,8 @@ spec:
- containerPort: 9000 - containerPort: 9000
resources: resources:
requests: requests:
memory: "768Mi" memory: "700Mi"
cpu: "100m" cpu: "300m"
limits: limits:
memory: "1.5Gi" memory: "1.5Gi"
cpu: "1000m" cpu: "1000m"
@@ -65,8 +61,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: authentik-worker-deployment name: authentik-worker-deployment
namespace: authentik namespace: authentik
spec: spec:
@@ -74,8 +68,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: authentik-worker app: authentik-worker
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -94,8 +86,8 @@ spec:
name: authentik-secrets name: authentik-secrets
resources: resources:
requests: requests:
memory: "320Mi" memory: "512Mi"
cpu: "100m" cpu: "300m"
limits: limits:
memory: "768Mi" memory: "1Gi"
cpu: "700m" cpu: "700m"
+3
View File
@@ -9,6 +9,9 @@ spec:
routes: routes:
- match: Host(`auth.forust.xyz`) - match: Host(`auth.forust.xyz`)
kind: Rule kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: authentik-server-service - name: authentik-server-service
port: 9000 port: 9000
-18
View File
@@ -1,18 +0,0 @@
# cert-manager
Public ACME issuers and an internal certificate authority.
This directory contains chart values and issuer resources, not the controller
installation. Install the cert-manager chart with CRDs and the settings in
`k8s/cert-manager-values.yaml` before applying the issuers.
`clusterissuer.yaml` defines staging and production Let's Encrypt issuers.
They use HTTP-01 through the Traefik ingress class. Public DNS and inbound HTTP
reachability must work for the requested names before issuance.
`internal-ca.yaml` bootstraps the internal CA. Keep its private-key Secret backed
up; the tracked `.crt` is only a public certificate.
This directory has no `k8s/active` marker. Apply the issuer files deliberately;
`kubectl apply` does not interpret the Helm values file.
See the [repository README](../README.md) for deployment selection.
-22
View File
@@ -1,22 +0,0 @@
# Cloudflare DDNS
Updates the lab DNS records when the public address changes.
Kubernetes runs in `default` with host networking and reads `cfddns-secrets`.
The Compose stack also uses host networking. Configure the API token and domain
list from the relevant example; keep DNS names consistent with the ingress rules.
`config.json.example` is a separate configuration example. The current Compose
file does not mount a config.json file. Check configuration against the pinned
DDNS image when changing between environment and file-based settings.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n default
kubectl get events -n default --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+2 -6
View File
@@ -1,8 +1,6 @@
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: cfddns name: cfddns
labels: labels:
app: cfddns app: cfddns
@@ -11,8 +9,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: cfddns app: cfddns
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -26,10 +22,10 @@ spec:
imagePullPolicy: Always imagePullPolicy: Always
resources: resources:
requests: requests:
memory: "32Mi" memory: "20Mi"
cpu: "30m" cpu: "30m"
limits: limits:
memory: "128Mi" memory: "64Mi"
cpu: "50m" cpu: "50m"
envFrom: envFrom:
- secretRef: - secretRef:
-21
View File
@@ -1,21 +0,0 @@
# Checkmk
Checkmk Raw monitoring site with web and agent-receiver ingress.
The site data lives in `checkmk-sites-pvc` on Kubernetes and the `sites` named
volume on Compose. The agent receiver has a separate TCP route; enabling the
web route alone does not expose it.
Prepare the password in the service env or Secret example. Inspect the Checkmk
container logs during the first site creation.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n checkmk
kubectl get events -n checkmk --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
File renamed without changes.
-4
View File
@@ -17,8 +17,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: checkmk-deployment name: checkmk-deployment
namespace: checkmk namespace: checkmk
spec: spec:
@@ -26,8 +24,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: checkmk app: checkmk
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
+3
View File
@@ -9,6 +9,9 @@ spec:
routes: routes:
- match: Host(`cmk.forust.xyz`) - match: Host(`cmk.forust.xyz`)
kind: Rule kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: checkmk-service - name: checkmk-service
port: 5000 port: 5000
-21
View File
@@ -1,21 +0,0 @@
# Cloudflare Tunnel
A Kubernetes connector for an existing Cloudflare tunnel.
The Deployment runs in `default` and reads its token from the ignored Secret
created from `k8s/secret.yaml.example`. Create the tunnel and its hostname rules
in Cloudflare before starting the connector.
There is no Compose file or `k8s/active` marker. Apply the Secret first, then
`k8s/deployment.yaml` when this tunnel is needed.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n default
kubectl get events -n default --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
File renamed without changes.
+3 -7
View File
@@ -1,8 +1,6 @@
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: cloudflared name: cloudflared
labels: labels:
app: cloudflared app: cloudflared
@@ -11,8 +9,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: cloudflared app: cloudflared
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -20,7 +16,7 @@ spec:
spec: spec:
containers: containers:
- name: cloudflared - name: cloudflared
image: cloudflare/cloudflared:2026.10.0 image: cloudflare/cloudflared:2026.9.1
imagePullPolicy: IfNotPresent imagePullPolicy: IfNotPresent
args: args:
- tunnel - tunnel
@@ -34,8 +30,8 @@ spec:
key: TUNNEL_TOKEN key: TUNNEL_TOKEN
resources: resources:
requests: requests:
memory: "128Mi" memory: "32Mi"
cpu: "30m" cpu: "30m"
limits: limits:
memory: "256Mi" memory: "128Mi"
cpu: "200m" cpu: "200m"
-22
View File
@@ -1,22 +0,0 @@
# File converters
ConvertX for server-side conversion and BentoPDF for PDF tools.
ConvertX persists files in `convertx-pvc`; BentoPDF has no persistent volume.
Kubernetes configuration includes a local `config.yaml.example`, excluded from
normal deployment. Copy and apply the real ConfigMap separately where required.
Compose publishes ConvertX on host port 9992 as well as attaching it to the
proxy network. Replace the authentication settings from `.env.example` before
exposing it outside the lab.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n converters
kubectl get events -n converters --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+2 -2
View File
@@ -1,7 +1,7 @@
services: services:
convertx: convertx:
container_name: convertx container_name: convertx
image: ghcr.io/c4illin/convertx:v0.19.0 image: ghcr.io/c4illin/convertx:v0.18.0
restart: unless-stopped restart: unless-stopped
ports: ports:
- "9992:3000" - "9992:3000"
@@ -54,7 +54,7 @@ services:
- "traefik.http.routers.bentopdf.tls.certresolver=letsencrypt" - "traefik.http.routers.bentopdf.tls.certresolver=letsencrypt"
- "traefik.http.routers.bentopdf.tls=true" - "traefik.http.routers.bentopdf.tls=true"
# Local router # Local router
- "traefik.http.routers.bentopdf-local.rule=Host(`pdf.workstation.internal`)" - "traefik.http.routers.bentopdf-local.rule=Host(`pdf.wokstation.internal`)"
- "traefik.http.routers.bentopdf-local.entrypoints=websecure" - "traefik.http.routers.bentopdf-local.entrypoints=websecure"
- "traefik.http.routers.bentopdf-local.tls=true" - "traefik.http.routers.bentopdf-local.tls=true"
# Dev router # Dev router
+2 -3
View File
@@ -31,13 +31,12 @@ spec:
name: bentopdf name: bentopdf
ports: ports:
- containerPort: 8080 - containerPort: 8080
# p95 4M, max 11M over 7 days. Was 50Mi/700Mi.
resources: resources:
requests: requests:
memory: "32Mi" memory: "50Mi"
cpu: "50m" cpu: "50m"
ephemeral-storage: "100Mi" ephemeral-storage: "100Mi"
limits: limits:
memory: "128Mi" memory: "700Mi"
cpu: "700m" cpu: "700m"
ephemeral-storage: "5Gi" ephemeral-storage: "5Gi"
+3 -8
View File
@@ -13,8 +13,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: convertx-deployment name: convertx-deployment
namespace: converters namespace: converters
spec: spec:
@@ -22,15 +20,13 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: convertx app: convertx
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
app: convertx app: convertx
spec: spec:
containers: containers:
- image: ghcr.io/c4illin/convertx:v0.19.0 - image: ghcr.io/c4illin/convertx:v0.18.0
name: convertx name: convertx
envFrom: envFrom:
- configMapRef: - configMapRef:
@@ -42,14 +38,13 @@ spec:
volumeMounts: volumeMounts:
- mountPath: /data - mountPath: /data
name: data name: data
# p95 85M, max 136M over 7 days, spikes while converting. Was 250Mi/1.5Gi.
resources: resources:
requests: requests:
memory: "128Mi" memory: "250Mi"
cpu: "100m" cpu: "100m"
limits: limits:
cpu: "1500m" cpu: "1500m"
memory: "512Mi" memory: "1.5Gi"
volumes: volumes:
- name: data - name: data
persistentVolumeClaim: persistentVolumeClaim:
-25
View File
@@ -1,25 +0,0 @@
# CrowdSec
Helm values, dashboards, network policy, and a maintenance CronJob.
Install CrowdSec separately using `k8s/crowdsec-values.yaml`; the deploy
workflow does not have a CrowdSec Helm release entry. There is no `k8s/active`
marker in this directory.
The LAPI policy and janitor run in `crowdsec`. The dashboard ConfigMaps are in
`prometheus` for Grafana's sidecar. The janitor has its own ServiceAccount and
namespace Role. Review its script and schedule before enabling cleanup.
Traefik's values state that enforcement moved to a host firewall bouncer. This
repository does not install that host component.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n crowdsec
kubectl get events -n crowdsec --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+14
View File
@@ -0,0 +1,14 @@
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: crowdsec-bouncer
namespace: crowdsec
spec:
plugin:
crowdsec-bouncer:
enabled: true
LogLevel: INFO
CrowdsecMode: live
CrowdsecLapiScheme: http
CrowdsecLapiHost: crowdsec-service.crowdsec.svc.cluster.local:8080
CrowdsecLapiKeyFile: "/etc/traefik/secrets/traefik-api-key"
+4 -105
View File
@@ -16,12 +16,6 @@ agent:
value: crowdsecurity/traefik crowdsecurity/base-http-scenarios value: crowdsecurity/traefik crowdsecurity/base-http-scenarios
- name: DISABLE_COLLECTIONS - name: DISABLE_COLLECTIONS
value: crowdsecurity/sshd value: crowdsecurity/sshd
# Bans on 401/403 bursts hurt more than they protect: with L3 enforcement
# a false positive cuts the IP off everything (SSH included), and past
# incidents show legit automation (deploy runner, mesh peers, registry
# pulls) tripping this probe. Probing/XSS/SQLi/CVE scenarios stay.
- name: DISABLE_SCENARIOS
value: crowdsecurity/http-generic-bf
metrics: metrics:
enabled: true enabled: true
serviceMonitor: serviceMonitor:
@@ -64,44 +58,9 @@ config:
reason: "Mobile IP whitelist" reason: "Mobile IP whitelist"
cidr: cidr:
- "84.245.64.0/18" - "84.245.64.0/18"
# CrowdSec's own guidance: CIDR allowlisting belongs at the parser stage.
# A parser whitelist discards the event before it reaches a bucket, so
# these addresses never produce an overflow and never become a decision.
# A postoverflow whitelist is checked only *after* the ban exists, and
# the bouncer answers 403 for as long as it does - which is a window we
# do not want the deploy sitting in.
local-network.yaml: |
name: forust/local-network
description: "Whitelist loopback, private and VPN networks"
whitelist:
reason: "Local network"
cidr:
- "127.0.0.0/8"
- "10.0.0.0/8"
- "172.16.0.0/12"
- "192.168.0.0/16"
# CGNAT range (RFC 6598). The workstation and the k0s node live
# here on WireGuard, and 100.64.0.0/10 is not covered by the
# RFC 1918 blocks above.
- "100.64.0.0/10"
- "169.254.0.0/16"
- "fc00::/7"
- "fe80::/10"
vps-whitelist.yaml: |
name: forust/vps-whitelist
description: "Whitelist static VPS"
whitelist:
reason: "VPS"
ip:
- "193.181.211.79"
postoverflows: postoverflows:
s01-whitelist: s01-whitelist:
# The one whitelist that has to stay here: resolving a hostname is a
# network call, and the docs put expensive lookups in postoverflows on
# purpose - it runs only when a bucket actually overflows.
# ddns.forust.xyz is the public home address, not a private one, so
# forust/local-network does not cover it.
home-dynamic-ip.yaml: | home-dynamic-ip.yaml: |
name: forust/home-dynamic-ip name: forust/home-dynamic-ip
description: "Whitelist home dynamic IP" description: "Whitelist home dynamic IP"
@@ -110,59 +69,6 @@ config:
expression: expression:
- evt.Overflow.Alert.Source.IP in LookupHost("ddns.forust.xyz") - evt.Overflow.Alert.Source.IP in LookupHost("ddns.forust.xyz")
# LAPI-only main config override, merged over config.yaml. NOTE: the
# chart's own default for this key is REPLACED, not merged, so its
# auto_registration block is repeated verbatim below - drop it and the
# agent can no longer register itself.
config.yaml.local: |
api:
server:
auto_registration: # Activate if not using TLS for authentication
enabled: true
token: "${REGISTRATION_TOKEN}" # /!\ Do not modify this variable (auto-generated and handled by the chart)
allowed_ranges: # /!\ Make sure to adapt to the pod IP ranges used by your cluster
- "127.0.0.1/32"
- "192.168.0.0/16"
- "10.0.0.0/8"
- "172.16.0.0/12"
# This homelab has no egress to console.crowdsec.cloud: DNS does
# not resolve. The LAPI kept trying anyway ("Signal push: N
# signals to push", "capi metrics: sending" every 10s) and each
# attempt sat on a resolver timeout WHILE HOLDING A WRITE
# TRANSACTION, which is what kept stalling per-request decision
# lookups even with WAL enabled. Nothing to share and nothing to
# pull - turn the Central API off instead of letting it block the
# only database writer we have.
online_client:
sharing: false
pull:
community: false
blocklists: false
disable_usage_metrics_export: true
db_config:
# SQLite without WAL serialises every reader behind the writer's
# rollback journal, and the LAPI writes constantly: the agent pushes
# Traefik alerts read from Loki, the metrics collector counts
# decisions, the bouncer touches "last pull" on every request.
# Symptom: decision lookups taking 10-30s (and a second connection
# that could not even open the database) while the LAPI sat at 28m
# CPU - the process was blocked in fsync, not computing. Every
# bouncer-protected request then blew through the plugin timeout and
# fail-closed with 403, on every site at once.
# The PVC is local-path-retain (hostPath), not a network share, so
# WAL is safe here; the crowdsec docs recommend it for exactly this
# ("allowing more concurrency in SQLite that will improve
# performances in most scenarios").
use_wal: true
# Keeps the alert table bounded. At the 5000/7d default the file
# reached 54MB in 15 days off the Traefik access log alone, and the
# metrics collector counts decisions on a timer; a smaller working
# set means fewer full scans. Crowdsec only prunes - SQLite never
# shrinks the file, so the size stays until a manual VACUUM.
flush:
max_items: 1000
max_age: 24h
lapi: lapi:
env: env:
- name: COLLECTIONS - name: COLLECTIONS
@@ -184,20 +90,13 @@ lapi:
enabled: true enabled: true
size: 1Gi size: 1Gi
storageClassName: local-path-retain storageClassName: local-path-retain
# LAPI answers a blocking /v1/decisions lookup for EVERY bouncer-protected
# request (whole Traefik front door), so it is the hot path of the proxy.
# At 400m/500Mi it went CPU-throttled and idle lookups measured 1.3-7.4s,
# which pushed requests into the bouncer's fail-closed 403.
# Single replica on purpose: LAPI is stateful (BoltDB on the `data` PVC,
# credentials on the `config` PVC) - two replicas sharing those RWO
# volumes would corrupt the decision store. Scale up CPU, not replicas.
resources: resources:
limits: limits:
cpu: 1500m cpu: 400m
memory: 1Gi
requests:
cpu: 250m
memory: 500Mi memory: 500Mi
requests:
cpu: 50m
memory: 150Mi
service: service:
type: ClusterIP type: ClusterIP
storeLAPICscliCredentialsInSecret: true storeLAPICscliCredentialsInSecret: true
-7
View File
@@ -32,13 +32,6 @@
# on their own - same name + same password); # on their own - same name + same password);
# 4. prune bouncer entries idle for 30d. # 4. prune bouncer entries idle for 30d.
# #
# It used to also delete LePresidente/http-generic-403-bf decisions hourly.
# That was a workaround for the bouncer failing closed on a slow LAPI and
# 403-ing the deploy runner into a 4h ban. The bouncer now polls decisions
# into a cache and never blocks on an unreachable LAPI, so it cannot
# manufacture those 403s any more, and the scenario only fires against real
# scanners - deleting their decisions hourly was undoing a working ban.
#
# Manual apply (crowdsec/k8s is NOT managed by deploy.yaml): # Manual apply (crowdsec/k8s is NOT managed by deploy.yaml):
# kubectl apply -f crowdsec/k8s/janitor-cronjob.yaml # kubectl apply -f crowdsec/k8s/janitor-cronjob.yaml
# Force a run: # Force a run:
-21
View File
@@ -1,21 +0,0 @@
# Dockmon
Docker management UI that talks to the host Docker daemon.
Both runtimes mount `/var/run/docker.sock`. On Kubernetes the socket belongs
to the node hosting the pod, so this is not a cluster-wide container manager.
Compose stores application data in a named volume. Kubernetes uses a StatefulSet
with a volume claim template. Its ServersTransport is specific to the upstream
connection; keep it with the ingress resources.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n dockmon
kubectl get events -n dockmon --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+2
View File
@@ -18,6 +18,8 @@ spec:
- match: Host(`dockmon.forust.xyz`) - match: Host(`dockmon.forust.xyz`)
kind: Rule kind: Rule
middlewares: middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
- name: security-headers@file - name: security-headers@file
services: services:
- name: dockmon-service - name: dockmon-service
-134
View File
@@ -1,134 +0,0 @@
# Repository review (6 October 2026 baseline)
This records the tracked tree at `cc9c3de` and the workstation state observed on
6 October 2026. It is a historical review, not a current runtime inventory. The
listed code fixes have since merged into `main`; EDU ownership has moved to the
separate repository described in [the handoff record](../.gitea/EDU_HANDOFF.md).
See the [CI and deployment guide](../.gitea/README.md) and
[runner and recovery guide](../.gitea/runner/README.md) for the current workflow.
No deployment was performed during the original review.
## Findings at the baseline and current status
| Priority | Finding at the baseline | Current status |
| -------- | ---------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| High | Per-file `APPLY_PRUNE=true` could delete resources selected by a shared label. | The deploy workflow rejects unsafe pruning before applying resources. |
| High | Compose validation did not resolve the local configuration required at deploy time. | Preflight resolves the selected Compose configuration before apply. |
| Medium | Secret validation could miss namespace-specific and mounted Secret references. | Preflight checks rendered references in their namespaces, including mounted and projected Secrets. |
| Medium | Compose CI missed manual entry points such as `shared-compose.yaml` and `client.compose.yaml`. | CI checks all tracked Compose files. |
| Medium | NetBird Compose referenced missing setup and renderer files. | The setup and renderer files are now present; Compose remains a manual alternative to the active Kubernetes deployment. |
| Medium | Glance mounted its CSS from the wrong ConfigMap. | The mount now uses the ConfigMap that contains `user.css`. |
| Medium | The PostgreSQL env example omitted the required NetBox password. | The example now includes the required variable. |
| Medium | The former EDU code had stale Compose variable names and session reliability problems. | EDU workloads and their fixes moved out of this repository; see the handoff record. |
| Medium | AdGuard DoH and SearXNG Compose router expressions used invalid `Host(...)` syntax. | The router expressions now follow Traefik's rule syntax. |
Traefik matchers should be combined as `Host(a) || Host(b)`; the rule syntax is
described in the [Traefik rules documentation](https://doc.traefik.io/traefik/reference/routing-configuration/http/routing/rules-and-priority/).
The fix retains the DoH path constraint for both hostnames.
The current deploy workflow deliberately rejects the unsafe prune option. It
does not introduce automatic deletion under a different implementation. The
baseline finding was a configuration risk, not evidence of a live deletion
incident.
The former session fix bounded HTTP and Redis calls, validated credentials, set
a cookie lifetime of two refresh intervals, and marked success only after
publishing the verified cookie. The service is now owned by the EDU repository;
see that repository for its current implementation.
The deployment fix extracts required pod Secret references from rendered JSON,
checks their namespaces, includes init containers, image-pull credentials, and
mounted/projected Secrets, and honors optional references. Ingress TLS Secrets
issued by cert-manager are not treated as pre-existing pod prerequisites.
It checks existence/access, not every key's contents or application validity.
## Live workstation observations
The SSH alias `workstation` is reachable. It has one Ready control-plane node,
Kubernetes `v1.35.4+k0s`, and a Docker daemon alongside containerd. At inspection,
no pods were Pending or in another non-running, non-completed phase. This is a
point-in-time observation, not a complete application health test.
The deployment checkout at `/srv/homelab` is on main commit `2adf17c`, behind the
reviewed local commit. It has untracked host configuration and a separate
`userbot/` directory. It was not reset or cleaned.
| Observed difference | Implication |
| ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| VictoriaMetrics and vmalert are running; the Prometheus StatefulSet has zero replicas. | A monitoring migration is already in progress outside committed main. Deploying the old Helm values can overwrite those settings. |
| Homarr, Cloudflared, and Reloader are installed without their current Git active markers. | Installed services and marker-selected services are different inventories. Missing markers do not establish that a service is stopped. |
| Cloudflare DDNS is running in both Docker and Kubernetes. | Confirm which instance should own DNS updates and whether their domain lists overlap before retiring either one. Secret values were not inspected. |
| Traefik's LoadBalancer exposes port 8080 at `192.168.80.2`. | The direct API listener is deployed; its external reachability was not tested. |
| Default `local-path` has reclaim policy Delete, while many existing PVs have been changed to Retain. | Current retention is partly live state. Recreating a claim can get a different policy from the old PV. |
| NetBird, NetBox media/reports/scripts, EDU Redis, Homarr, and VictoriaMetrics have Delete-policy PVs. | Deleting their claims can delete important state. Plan backup and retention changes before namespace cleanup. |
The VictoriaMetrics monitoring trial later merged into `main` in PR #95. The
first row above records the state before that change. Read
[`prometheus-stack/README.md`](../prometheus-stack/README.md) for the current
tracked monitoring configuration; the live observations in this section remain
a snapshot from 6 October.
## Current recovery limits
The deployment controller and its recovery process changed after this review.
The current operator workflow is documented in the
[runner and recovery guide](../.gitea/runner/README.md). The remaining boundaries
are:
- Kubernetes recovery can restore captured workload revisions. It does not
restore ConfigMaps, Secrets, database schemas, or persistent data.
- Compose recovery is manual. It uses saved resolved configuration, but it does
not restore volume data or reverse database migrations.
- Removed resources require manual review and removal; the deploy workflow does
not prune them automatically.
- Plan mode does not create namespaces. During apply, server validation for new
namespaces runs after namespace creation and chart installation; a failed
check can leave an empty namespace.
- Storage policy and backup coverage remain service-specific. Check the live PV,
PVC, and backup state before changing stateful workloads.
## Validation
At the review baseline, lint checks passed for Python, shell, workflows, YAML, standard Compose
files, and Kubernetes resources with available schemas. Kubeconform found 347
resources in 174 files: 201 valid, 146 skipped CRDs, zero invalid resources.
That skip count matters: passing schema validation does not validate Traefik rule
strings or other controller-specific behavior.
Fix validation covers:
- Compose discovery of manual entry points, rejection of required-variable gaps,
namespace-scoped and optional Secret references, and API/render failures.
- NetBird setup idempotence, preservation of existing keys, file permissions,
runtime rendering, and rejection of invalid trusted proxy CIDRs.
- Session refresh success and failure paths, timeouts, cookie expiry, log redaction,
missing credentials, and nonpositive refresh intervals.
- Correct Glance ConfigMap key selection and PostgreSQL initializer/env alignment.
- YAML and Compose structure for the corrected router rules, compared with the
documented Traefik grammar. They were not exercised on the live proxy.
- Prune rejection before any cluster invocation.
At the time of review, all seven fix branches and the documentation branch
merged together in a disposable validation worktree. That combined tree passed the
CI-equivalent local checks, Markdown formatting/lint and link checks, all 35
Compose structure checks, and 11 Python regression tests plus the shell
validation regressions. CRD server-side validation and live rollout tests were
not run.
Runtime tests use fixtures and mocks, not production credentials. Live checks read
workload metadata, storage policies, chart versions, and container state only.
They did not read Secret contents or change services.
## Reloader follow-up (baseline)
`fix/reloader-integration` added the active marker and opt-in annotations to
application Deployments/StatefulSets that consume runtime ConfigMaps or Secrets.
It corrected AdGuard's misplaced pod-template annotation. The Helm settings use
annotation-based reloads, keep global auto-reload disabled, and ignore Jobs and
CronJobs. PostgreSQL workloads are excluded because their credential variables
and init scripts are only effective on an empty data directory.
The controller was running on the workstation when inspected. The original
review checked configuration against the pinned chart with Helm rendering and
manifest validation; it did not change production configuration to provoke a
test restart or confirm every application's live reload behavior.
-20
View File
@@ -1,20 +0,0 @@
# Downtify
Download UI with a persistent downloads directory.
Compose stores downloads under `Downtify_downloads/`; Kubernetes uses
`downtify-downloads-pvc`. The ingress manifests reference shared infrastructure,
so check certificate and middleware availability before enabling them.
Back up downloads separately if they need to survive storage replacement.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n downtify
kubectl get events -n downtify --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,7 +1,7 @@
services: services:
downtify: downtify:
container_name: downtify container_name: downtify
image: ghcr.io/henriquesebastiao/downtify:3.4.0 image: ghcr.io/henriquesebastiao/downtify:3.1.0
restart: unless-stopped restart: unless-stopped
# ports: # ports:
# - '7077:8000' # - '7077:8000'
+1 -3
View File
@@ -20,8 +20,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: downtify app: downtify
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -29,7 +27,7 @@ spec:
spec: spec:
containers: containers:
- name: downtify - name: downtify
image: ghcr.io/henriquesebastiao/downtify:3.4.0 image: ghcr.io/henriquesebastiao/downtify:3.1.0
ports: ports:
- containerPort: 8000 - containerPort: 8000
volumeMounts: volumeMounts:
+2
View File
@@ -10,6 +10,8 @@ spec:
- match: Host(`downtify.forust.xyz`) - match: Host(`downtify.forust.xyz`)
kind: Rule kind: Rule
middlewares: middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
- name: security-chain@file - name: security-chain@file
services: services:
- name: downtify-service - name: downtify-service
+13
View File
@@ -0,0 +1,13 @@
FROM python:3.9-alpine
WORKDIR /app
# Установка зависимостей
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Копирование кода
COPY main.py .
COPY .env .
# Запуск бота
CMD ["python", "-u", "main.py"]
+373
View File
@@ -0,0 +1,373 @@
Mozilla Public License Version 2.0
==================================
1. Definitions
--------------
1.1. "Contributor"
means each individual or legal entity that creates, contributes to
the creation of, or owns Covered Software.
1.2. "Contributor Version"
means the combination of the Contributions of others (if any) used
by a Contributor and that particular Contributor's Contribution.
1.3. "Contribution"
means Covered Software of a particular Contributor.
1.4. "Covered Software"
means Source Code Form to which the initial Contributor has attached
the notice in Exhibit A, the Executable Form of such Source Code
Form, and Modifications of such Source Code Form, in each case
including portions thereof.
1.5. "Incompatible With Secondary Licenses"
means
(a) that the initial Contributor has attached the notice described
in Exhibit B to the Covered Software; or
(b) that the Covered Software was made available under the terms of
version 1.1 or earlier of the License, but not also under the
terms of a Secondary License.
1.6. "Executable Form"
means any form of the work other than Source Code Form.
1.7. "Larger Work"
means a work that combines Covered Software with other material, in
a separate file or files, that is not Covered Software.
1.8. "License"
means this document.
1.9. "Licensable"
means having the right to grant, to the maximum extent possible,
whether at the time of the initial grant or subsequently, any and
all of the rights conveyed by this License.
1.10. "Modifications"
means any of the following:
(a) any file in Source Code Form that results from an addition to,
deletion from, or modification of the contents of Covered
Software; or
(b) any new file in Source Code Form that contains any Covered
Software.
1.11. "Patent Claims" of a Contributor
means any patent claim(s), including without limitation, method,
process, and apparatus claims, in any patent Licensable by such
Contributor that would be infringed, but for the grant of the
License, by the making, using, selling, offering for sale, having
made, import, or transfer of either its Contributions or its
Contributor Version.
1.12. "Secondary License"
means either the GNU General Public License, Version 2.0, the GNU
Lesser General Public License, Version 2.1, the GNU Affero General
Public License, Version 3.0, or any later versions of those
licenses.
1.13. "Source Code Form"
means the form of the work preferred for making modifications.
1.14. "You" (or "Your")
means an individual or a legal entity exercising rights under this
License. For legal entities, "You" includes any entity that
controls, is controlled by, or is under common control with You. For
purposes of this definition, "control" means (a) the power, direct
or indirect, to cause the direction or management of such entity,
whether by contract or otherwise, or (b) ownership of more than
fifty percent (50%) of the outstanding shares or beneficial
ownership of such entity.
2. License Grants and Conditions
--------------------------------
2.1. Grants
Each Contributor hereby grants You a world-wide, royalty-free,
non-exclusive license:
(a) under intellectual property rights (other than patent or trademark)
Licensable by such Contributor to use, reproduce, make available,
modify, display, perform, distribute, and otherwise exploit its
Contributions, either on an unmodified basis, with Modifications, or
as part of a Larger Work; and
(b) under Patent Claims of such Contributor to make, use, sell, offer
for sale, have made, import, and otherwise transfer either its
Contributions or its Contributor Version.
2.2. Effective Date
The licenses granted in Section 2.1 with respect to any Contribution
become effective for each Contribution on the date the Contributor first
distributes such Contribution.
2.3. Limitations on Grant Scope
The licenses granted in this Section 2 are the only rights granted under
this License. No additional rights or licenses will be implied from the
distribution or licensing of Covered Software under this License.
Notwithstanding Section 2.1(b) above, no patent license is granted by a
Contributor:
(a) for any code that a Contributor has removed from Covered Software;
or
(b) for infringements caused by: (i) Your and any other third party's
modifications of Covered Software, or (ii) the combination of its
Contributions with other software (except as part of its Contributor
Version); or
(c) under Patent Claims infringed by Covered Software in the absence of
its Contributions.
This License does not grant any rights in the trademarks, service marks,
or logos of any Contributor (except as may be necessary to comply with
the notice requirements in Section 3.4).
2.4. Subsequent Licenses
No Contributor makes additional grants as a result of Your choice to
distribute the Covered Software under a subsequent version of this
License (see Section 10.2) or under the terms of a Secondary License (if
permitted under the terms of Section 3.3).
2.5. Representation
Each Contributor represents that the Contributor believes its
Contributions are its original creation(s) or it has sufficient rights
to grant the rights to its Contributions conveyed by this License.
2.6. Fair Use
This License is not intended to limit any rights You have under
applicable copyright doctrines of fair use, fair dealing, or other
equivalents.
2.7. Conditions
Sections 3.1, 3.2, 3.3, and 3.4 are conditions of the licenses granted
in Section 2.1.
3. Responsibilities
-------------------
3.1. Distribution of Source Form
All distribution of Covered Software in Source Code Form, including any
Modifications that You create or to which You contribute, must be under
the terms of this License. You must inform recipients that the Source
Code Form of the Covered Software is governed by the terms of this
License, and how they can obtain a copy of this License. You may not
attempt to alter or restrict the recipients' rights in the Source Code
Form.
3.2. Distribution of Executable Form
If You distribute Covered Software in Executable Form then:
(a) such Covered Software must also be made available in Source Code
Form, as described in Section 3.1, and You must inform recipients of
the Executable Form how they can obtain a copy of such Source Code
Form by reasonable means in a timely manner, at a charge no more
than the cost of distribution to the recipient; and
(b) You may distribute such Executable Form under the terms of this
License, or sublicense it under different terms, provided that the
license for the Executable Form does not attempt to limit or alter
the recipients' rights in the Source Code Form under this License.
3.3. Distribution of a Larger Work
You may create and distribute a Larger Work under terms of Your choice,
provided that You also comply with the requirements of this License for
the Covered Software. If the Larger Work is a combination of Covered
Software with a work governed by one or more Secondary Licenses, and the
Covered Software is not Incompatible With Secondary Licenses, this
License permits You to additionally distribute such Covered Software
under the terms of such Secondary License(s), so that the recipient of
the Larger Work may, at their option, further distribute the Covered
Software under the terms of either this License or such Secondary
License(s).
3.4. Notices
You may not remove or alter the substance of any license notices
(including copyright notices, patent notices, disclaimers of warranty,
or limitations of liability) contained within the Source Code Form of
the Covered Software, except that You may alter any license notices to
the extent required to remedy known factual inaccuracies.
3.5. Application of Additional Terms
You may choose to offer, and to charge a fee for, warranty, support,
indemnity or liability obligations to one or more recipients of Covered
Software. However, You may do so only on Your own behalf, and not on
behalf of any Contributor. You must make it absolutely clear that any
such warranty, support, indemnity, or liability obligation is offered by
You alone, and You hereby agree to indemnify every Contributor for any
liability incurred by such Contributor as a result of warranty, support,
indemnity or liability terms You offer. You may include additional
disclaimers of warranty and limitations of liability specific to any
jurisdiction.
4. Inability to Comply Due to Statute or Regulation
---------------------------------------------------
If it is impossible for You to comply with any of the terms of this
License with respect to some or all of the Covered Software due to
statute, judicial order, or regulation then You must: (a) comply with
the terms of this License to the maximum extent possible; and (b)
describe the limitations and the code they affect. Such description must
be placed in a text file included with all distributions of the Covered
Software under this License. Except to the extent prohibited by statute
or regulation, such description must be sufficiently detailed for a
recipient of ordinary skill to be able to understand it.
5. Termination
--------------
5.1. The rights granted under this License will terminate automatically
if You fail to comply with any of its terms. However, if You become
compliant, then the rights granted under this License from a particular
Contributor are reinstated (a) provisionally, unless and until such
Contributor explicitly and finally terminates Your grants, and (b) on an
ongoing basis, if such Contributor fails to notify You of the
non-compliance by some reasonable means prior to 60 days after You have
come back into compliance. Moreover, Your grants from a particular
Contributor are reinstated on an ongoing basis if such Contributor
notifies You of the non-compliance by some reasonable means, this is the
first time You have received notice of non-compliance with this License
from such Contributor, and You become compliant prior to 30 days after
Your receipt of the notice.
5.2. If You initiate litigation against any entity by asserting a patent
infringement claim (excluding declaratory judgment actions,
counter-claims, and cross-claims) alleging that a Contributor Version
directly or indirectly infringes any patent, then the rights granted to
You by any and all Contributors for the Covered Software under Section
2.1 of this License shall terminate.
5.3. In the event of termination under Sections 5.1 or 5.2 above, all
end user license agreements (excluding distributors and resellers) which
have been validly granted by You or Your distributors under this License
prior to termination shall survive termination.
************************************************************************
* *
* 6. Disclaimer of Warranty *
* ------------------------- *
* *
* Covered Software is provided under this License on an "as is" *
* basis, without warranty of any kind, either expressed, implied, or *
* statutory, including, without limitation, warranties that the *
* Covered Software is free of defects, merchantable, fit for a *
* particular purpose or non-infringing. The entire risk as to the *
* quality and performance of the Covered Software is with You. *
* Should any Covered Software prove defective in any respect, You *
* (not any Contributor) assume the cost of any necessary servicing, *
* repair, or correction. This disclaimer of warranty constitutes an *
* essential part of this License. No use of any Covered Software is *
* authorized under this License except under this disclaimer. *
* *
************************************************************************
************************************************************************
* *
* 7. Limitation of Liability *
* -------------------------- *
* *
* Under no circumstances and under no legal theory, whether tort *
* (including negligence), contract, or otherwise, shall any *
* Contributor, or anyone who distributes Covered Software as *
* permitted above, be liable to You for any direct, indirect, *
* special, incidental, or consequential damages of any character *
* including, without limitation, damages for lost profits, loss of *
* goodwill, work stoppage, computer failure or malfunction, or any *
* and all other commercial damages or losses, even if such party *
* shall have been informed of the possibility of such damages. This *
* limitation of liability shall not apply to liability for death or *
* personal injury resulting from such party's negligence to the *
* extent applicable law prohibits such limitation. Some *
* jurisdictions do not allow the exclusion or limitation of *
* incidental or consequential damages, so this exclusion and *
* limitation may not apply to You. *
* *
************************************************************************
8. Litigation
-------------
Any litigation relating to this License may be brought only in the
courts of a jurisdiction where the defendant maintains its principal
place of business and such litigation shall be governed by laws of that
jurisdiction, without reference to its conflict-of-law provisions.
Nothing in this Section shall prevent a party's ability to bring
cross-claims or counter-claims.
9. Miscellaneous
----------------
This License represents the complete agreement concerning the subject
matter hereof. If any provision of this License is held to be
unenforceable, such provision shall be reformed only to the extent
necessary to make it enforceable. Any law or regulation which provides
that the language of a contract shall be construed against the drafter
shall not be used to construe this License against a Contributor.
10. Versions of the License
---------------------------
10.1. New Versions
Mozilla Foundation is the license steward. Except as provided in Section
10.3, no one other than the license steward has the right to modify or
publish new versions of this License. Each version will be given a
distinguishing version number.
10.2. Effect of New Versions
You may distribute the Covered Software under the terms of the version
of the License under which You originally received the Covered Software,
or under the terms of any subsequent version published by the license
steward.
10.3. Modified Versions
If you create software not governed by this License, and you want to
create a new license for such software, you may create and use a
modified version of this License if you rename the license and remove
any references to the name of the license steward (except to note that
such modified license differs from this License).
10.4. Distributing Source Code Form that is Incompatible With Secondary
Licenses
If You choose to distribute Source Code Form that is Incompatible With
Secondary Licenses under the terms of this version of the License, the
notice described in Exhibit B of this License must be attached.
Exhibit A - Source Code Form License Notice
-------------------------------------------
This Source Code Form is subject to the terms of the Mozilla Public
License, v. 2.0. If a copy of the MPL was not distributed with this
file, You can obtain one at https://mozilla.org/MPL/2.0/.
If it is not possible or desirable to put the notice in a particular
file, then You may include the notice in a location (such as a LICENSE
file in a relevant directory) where a recipient would be likely to look
for such a notice.
You may add additional accurate notices of copyright ownership.
Exhibit B - "Incompatible With Secondary Licenses" Notice
---------------------------------------------------------
This Source Code Form is "Incompatible With Secondary Licenses", as
defined by the Mozilla Public License, v. 2.0.
+15
View File
@@ -0,0 +1,15 @@
services:
dtek_notif:
build:
context: .
dockerfile: Dockerfile
image: gcr.forust.xyz/forust/dtek-notif:latest
pull_policy: build
restart: unless-stopped
environment:
- TZ=Europe/Kyiv
dns:
- 1.1.1.1
- 8.8.8.8
networks:
- default
+748
View File
@@ -0,0 +1,748 @@
import asyncio
import contextlib
import logging
import os
from datetime import datetime, timedelta
import requests
from aiogram import Bot, Dispatcher
from aiogram.filters import Command
from aiogram.types import KeyboardButton, Message
from aiogram.utils.keyboard import ReplyKeyboardBuilder
from bs4 import BeautifulSoup
from dotenv import load_dotenv
# Загрузка переменных окружения
load_dotenv()
# Настройки
TELEGRAM_TOKEN = os.getenv('TELEGRAM_TOKEN', 'YOUR_TOKEN_HERE')
ALLOWED_CHAT_IDS = list(map(int, os.getenv('ALLOWED_CHAT_IDS', '').split(','))) if os.getenv('ALLOWED_CHAT_IDS') else []
CHECK_INTERVAL = int(os.getenv('CHECK_INTERVAL', '120'))
# Параметры для запроса
VOE_CITY_ID = int(os.getenv('VOE_CITY_ID', 'VOE_CITY_ID'))
VOE_STREET_ID = int(os.getenv('VOE_STREET_ID', 'VOE_STREET_ID'))
VOE_HOUSE_ID = int(os.getenv('VOE_HOUSE_ID', 'VOE_HOUSE_ID'))
# Настройка логирования
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(name)s - %(levelname)s - %(message)s')
logger = logging.getLogger(__name__)
# Глобальные переменные
bot = Bot(token=TELEGRAM_TOKEN)
dp = Dispatcher()
last_schedule: list[dict] | None = None
last_notification_time: dict[str, datetime] = {}
# ============================================================================
# УТИЛИТЫ
# ============================================================================
def format_time_duration(minutes: int) -> str:
"""Форматирует время из минут в часы и минуты"""
hours = minutes // 60
mins = minutes % 60
if hours == 0:
return f'{mins}м'
elif mins == 0:
return f'{hours}ч'
return f'{hours}ч {mins}м'
def get_day_statistics(day_blocks: list[dict]) -> dict[str, int]:
"""Получает статистику по дню"""
total_minutes = 0
confirmed_minutes = 0
possible_minutes = 0
for block in day_blocks:
for half in [block['first_half'], block['second_half']]:
if half['status'] == 'off':
total_minutes += 30
if half['confirmed']:
confirmed_minutes += 30
else:
possible_minutes += 30
return {'total': total_minutes, 'confirmed': confirmed_minutes, 'possible': possible_minutes}
# ============================================================================
# ПАРСИНГ ДАННЫХ
# ============================================================================
def parse_html(html: str) -> list[dict]:
"""Парсит HTML с графиком отключений (логика от 15.11.2024)"""
soup = BeautifulSoup(html, 'html.parser')
cells = soup.select('.disconnection-detailed-table-cell.cell')
schedule = []
current_hour = 0
current_day = 0
for cell in cells:
if 'legend' in cell.get('class', []) or 'head' in cell.get('class', []):
continue
cell_classes = cell.get('class', [])
# ПРоверка статуса отключения на весь час
full_hour_off = 'has_disconnection' in cell_classes and 'full_hour' in cell_classes
hour_block = cell.select_one('.hour_block')
if not hour_block:
continue
# Проверка подтверждённости отключения для всего часа
cell_confirmed = None
if 'confirm_1' in cell_classes:
cell_confirmed = True
elif 'confirm_0' in cell_classes:
cell_confirmed = False
# Проверка половин часа
left = hour_block.select_one('.half.left')
right = hour_block.select_one('.half.right')
def parse_half(half, is_full_hour_off: bool, cell_confirmed: bool | None = None) -> dict:
"""Парсит половину часа"""
if not half:
return {'status': 'on', 'queue': None, 'confirmed': None}
half_classes = half.get('class', [])
# Если вся ячейка full_hour - используем статус ячейки
if is_full_hour_off:
return {'status': 'off', 'queue': None, 'confirmed': cell_confirmed}
# Определяем статус половины
if 'has_disconnection' in half_classes:
status = 'off'
elif 'no_disconnection' in half_classes:
status = 'on'
else:
status = 'on' # По умолчанию считаем включенным
# Если выключено - ищем подробности
queue = None
confirmed = None
if status == 'off':
disconnection_div = half.select_one('.disconnection')
if disconnection_div:
# Ищем номер черги в title
if disconnection_div.has_attr('title'):
title = disconnection_div['title']
if 'Номер черги' in title or 'Номер черги:' in title:
with contextlib.suppress(BaseException):
queue = title.split(':')[-1].strip()
# Определяем подтверждение
disc_classes = disconnection_div.get('class', [])
if 'disconnection_confirm_1' in disc_classes:
confirmed = True
elif 'disconnection_confirm_0' in disc_classes:
confirmed = False
return {'status': status, 'queue': queue, 'confirmed': confirmed}
first_half_data = parse_half(left, full_hour_off, cell_confirmed)
second_half_data = parse_half(right, full_hour_off, cell_confirmed)
schedule.append(
{
'hour': current_hour,
'day': current_day,
'first_half': first_half_data,
'second_half': second_half_data,
}
)
current_hour += 1
if current_hour >= 24:
current_hour = 0
current_day += 1
return schedule
def get_voe_html(city_id: int, street_id: int, house_id: int) -> str:
"""Получает HTML с сайта VOE"""
url = 'https://www.voe.com.ua/disconnection/detailed?ajax_form=1&_wrapper_format=drupal_ajax'
headers = {
'Content-Type': 'application/x-www-form-urlencoded; charset=UTF-8',
'X-Requested-With': 'XMLHttpRequest',
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36',
}
data = {
'search_type': 0,
'city_id': city_id,
'street_id': street_id,
'house_id': house_id,
'form_build_id': 'form-Irv5aHw1R2FT_Ik2apyHOZ47hTH5xPNH_LQnBrmpSTc',
'form_id': 'disconnection_detailed_search_form',
'_triggering_element_name': 'search',
'_triggering_element_value': 'Показати',
'_drupal_ajax': 1,
}
try:
response = requests.post(url, headers=headers, data=data, timeout=10)
response.raise_for_status()
resp_json = response.json()
insert_html = next((item['data'] for item in resp_json if item.get('command') == 'insert'), None)
if not insert_html:
raise ValueError('HTML не найден в ответе')
return insert_html
except requests.exceptions.RequestException as e:
logger.error(f'Ошибка запроса VOE: {e}')
raise
# ============================================================================
# ФОРМАТИРОВАНИЕ СООБЩЕНИЙ
# ============================================================================
def get_main_keyboard():
"""Создает главную клавиатуру"""
builder = ReplyKeyboardBuilder()
builder.row(KeyboardButton(text='📊 Графік'), KeyboardButton(text='🔄 Оновити'))
builder.row(KeyboardButton(text='📅 Сьогодні'), KeyboardButton(text='📅 Завтра'))
builder.row(KeyboardButton(text='ℹ️ Про бота'))
return builder.as_markup(resize_keyboard=True)
def format_schedule_message(schedule: list[dict], days_to_show: int = 2) -> str:
"""Форматирует полный график на несколько дней"""
lines = [
'⚡️ <b>Графік відключень світла</b>',
f'🕐 Оновлено: {datetime.now().strftime("%d.%m.%Y %H:%M:%S")}',
'─' * 30,
'',
]
start_date = datetime.now()
for day in range(min(days_to_show, 2)):
day_blocks = [b for b in schedule if b['day'] == day]
if not day_blocks:
continue
date_str = (start_date + timedelta(days=day)).strftime('%d.%m.%Y')
day_name = '🌅 <b>Сьогодні</b>' if day == 0 else '🌄 <b>Завтра</b>'
lines.append(f'{day_name} ({date_str})')
# Статистика
stats = get_day_statistics(day_blocks)
if stats['total'] > 0:
lines.append(f'⏱ Всього: <code>{format_time_duration(stats["total"])}</code>')
if stats['confirmed'] > 0:
lines.append(f'🔴 Підтверджено: <code>{format_time_duration(stats["confirmed"])}</code>')
if stats['possible'] > 0:
lines.append(f'🟠 Можливо: <code>{format_time_duration(stats["possible"])}</code>')
else:
lines.append('🟢 <b>Відключень немає!</b>')
lines.append('')
# Детальный список отключений
disconnections = []
current_status = None
start_time = None
current_confirmed = None
current_queue = None
for block in day_blocks:
hour = block['hour']
for half_idx, half in enumerate([block['first_half'], block['second_half']]):
time_str = f'{hour:02d}:00' if half_idx == 0 else f'{hour:02d}:30'
if half['status'] == 'off':
if current_status != 'off':
start_time = time_str
current_confirmed = half['confirmed']
current_queue = half['queue']
current_status = 'off'
else:
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч{current_queue})' if current_queue else ''
disconnections.append(f'{icon} <code>{start_time} - {time_str}</code>{queue_text}')
current_status = half['status']
# Если день закончился на отключении
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч{current_queue})' if current_queue else ''
next_hour = (day_blocks[-1]['hour'] + 1) % 24
end_time = f'{next_hour:02d}:00'
disconnections.append(f'{icon} <code>{start_time} - {end_time}</code>{queue_text}')
if disconnections:
for idx, disc in enumerate(disconnections, 1):
lines.append(f'{idx}. {disc}')
lines.append('')
lines.append('<i>🔴 = підтверджено • 🟠 = можливо • 🟢 = світло</i>')
return '\n'.join(lines)
def format_single_day_schedule(schedule: list[dict], day: int) -> str:
"""Форматирует график на один день"""
day_blocks = [b for b in schedule if b['day'] == day]
if not day_blocks:
return '❌ Немає даних для цього дня'
start_date = datetime.now()
date_str = (start_date + timedelta(days=day)).strftime('%d.%m.%Y')
day_name = '🟠 <b>Сьогодні</b>' if day == 0 else '🔶 <b>Завтра</b>'
lines = [f'{day_name} • {date_str}', '']
# Статистика
lines.append('<b>📊 Статистика</b>')
stats = get_day_statistics(day_blocks)
if stats['total'] == 0:
lines.append('└ 🟢 <b>Відключень немає!</b>')
else:
total_time = format_time_duration(stats['total'])
lines.append(f'├ ⏱ Всього: <code>{total_time}</code>')
if stats['confirmed'] > 0:
confirmed_time = format_time_duration(stats['confirmed'])
lines.append(f'├ 🔴 Підтверджено: <code>{confirmed_time}</code>')
if stats['possible'] > 0:
possible_time = format_time_duration(stats['possible'])
lines.append(f'└ 🟠 Можливо: <code>{possible_time}</code>')
else:
lines.append('└ 🟢 Решта часу світло')
lines.append('')
# Детальный список отключений
disconnections = []
current_status = None
start_time = None
current_confirmed = None
current_queue = None
for block in day_blocks:
hour = block['hour']
for half_idx, half in enumerate([block['first_half'], block['second_half']]):
time_str = f'{hour:02d}:00' if half_idx == 0 else f'{hour:02d}:30'
if half['status'] == 'off':
if current_status != 'off':
start_time = time_str
current_confirmed = half['confirmed']
current_queue = half['queue']
current_status = 'off'
else:
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч.{current_queue})' if current_queue else ''
disconnections.append(f'{icon} <code>{start_time} - {time_str}</code>{queue_text}')
current_status = half['status']
# Если день закончился на отключении
if current_status == 'off':
icon = '🔴' if current_confirmed else '🟠'
queue_text = f' (Ч.{current_queue})' if current_queue else ''
next_hour = (day_blocks[-1]['hour'] + 1) % 24
end_time = f'{next_hour:02d}:00'
disconnections.append(f'{icon} <code>{start_time} - {end_time}</code>{queue_text}')
if disconnections:
lines.append('<b>⚡️ Розклад відключень</b>')
for idx, disc in enumerate(disconnections, 1):
lines.append(f'{idx}. {disc}')
lines.append('')
lines.append('<i>🔴 підтверджено • 🟠 можливо • 🟢 світло</i>')
return '\n'.join(lines)
def schedules_differ(old_schedule: list[dict] | None, new_schedule: list[dict] | None) -> bool:
"""Проверяет отличия между графиками"""
if old_schedule is None or new_schedule is None:
return True
if len(old_schedule) != len(new_schedule):
return True
for old, new in zip(old_schedule, new_schedule, strict=False):
if old['day'] >= 2:
break
if old['first_half'] != new['first_half'] or old['second_half'] != new['second_half']:
return True
return False
# ============================================================================
# УВЕДОМЛЕНИЯ
# ============================================================================
async def send_to_all_users(message_text: str, parse_mode: str = 'HTML'):
"""Отправляет сообщение всем пользователям"""
if not ALLOWED_CHAT_IDS:
logger.warning('Нет допущенных ID чатов для отправки уведомлений')
return
for chat_id in ALLOWED_CHAT_IDS:
try:
await bot.send_message(chat_id, message_text, parse_mode=parse_mode)
logger.info(f'✅ Сообщение отправлено пользователю {chat_id}')
except Exception as e:
logger.error(f'❌ Ошибка отправки пользователю {chat_id}: {e}')
await asyncio.sleep(0.5)
async def check_schedule():
"""Проверяет график и отправляет уведомления"""
global last_schedule
try:
logger.info('🔍 Проверка графика...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
new_schedule = parse_html(html)
if schedules_differ(last_schedule, new_schedule):
logger.info('✨ Обнаружены изменения!')
message = format_schedule_message(new_schedule, days_to_show=2)
if last_schedule is not None:
await send_to_all_users(f'🔄 <b>Графік оновлено!</b>\n\n{message}')
last_schedule = new_schedule
else:
logger.info('✓ Графік без змін')
except Exception as e:
logger.error(f'❌ Ошибка при проверке графика: {e}')
async def check_upcoming_disconnections():
"""Проверяет предстоящие события и отправляет предупреждения за 5 минут"""
global last_notification_time
if last_schedule is None:
return
now = datetime.now()
today_blocks = [b for b in last_schedule if b['day'] == 0]
# Создаем список всех переходов (off -> on или on -> off)
transitions = []
prev_status = None
for block in today_blocks:
hour = block['hour']
for half_idx, half in enumerate([block['first_half'], block['second_half']]):
minute = 0 if half_idx == 0 else 30
time_str = f'{hour:02d}:{minute:02d}'
current_status = half['status']
# Если статус изменился - это переход
if prev_status is not None and prev_status != current_status:
transitions.append(
{
'hour': hour,
'minute': minute,
'time_str': time_str,
'from_status': prev_status,
'to_status': current_status,
'confirmed': half.get('confirmed'),
'queue': half.get('queue'),
}
)
prev_status = current_status
# Проверяем переходы
for transition in transitions:
event_time = now.replace(hour=transition['hour'], minute=transition['minute'], second=0, microsecond=0)
time_until = (event_time - now).total_seconds() / 60
notification_key = f'{transition["hour"]}:{transition["minute"]}_{transition["to_status"]}'
# Если за 5 минут до события (±1 минута) и еще не отправляли
if 4 <= time_until <= 6:
# Проверяем, не отправляли ли уже уведомление сегодня
if notification_key in last_notification_time:
last_notif_time = last_notification_time[notification_key]
if last_notif_time.date() == now.date():
continue # Уже отправляли сегодня
# Переход на ОТКЛЮЧЕНИЕ (on -> off)
if transition['from_status'] == 'on' and transition['to_status'] == 'off':
icon = '🔴' if transition['confirmed'] else '🟠'
status = 'підтверджено' if transition['confirmed'] else 'можливе'
queue_info = f' (Черга {transition["queue"]})' if transition['queue'] else ''
warning = (
f'⚠️ <b>УВАГА! ВІДКЛЮЧЕННЯ</b>\n\n'
f'Через ~5 хвилин\n'
f'Час: <code>{transition["time_str"]}</code>\n'
f'Статус: {icon} {status}{queue_info}'
)
await send_to_all_users(warning)
last_notification_time[notification_key] = now
logger.info(f'📢 Відправлено попередження про ВІДКЛЮЧЕННЯ в {transition["time_str"]}')
# Переход на ВКЛЮЧЕНИЕ (off -> on)
elif transition['from_status'] == 'off' and transition['to_status'] == 'on':
warning = (
f'✅ <b>УВАГА! ВКЛЮЧЕННЯ</b>\n\n'
f'Через ~5 хвилин буде світло\n'
f'Час: <code>{transition["time_str"]}</code>'
)
await send_to_all_users(warning)
last_notification_time[notification_key] = now
logger.info(f'📢 Відправлено попередження про ВКЛЮЧЕННЯ в {transition["time_str"]}')
async def monitoring_loop():
"""Основной цикл мониторинга"""
await check_schedule()
while True:
try:
await asyncio.sleep(CHECK_INTERVAL)
await check_schedule()
await check_upcoming_disconnections()
except Exception as e:
logger.error(f'Ошибка в цикле мониторинга: {e}')
await asyncio.sleep(5)
# ============================================================================
# ОБРАБОТЧИКИ КОМАНД
# ============================================================================
@dp.message(Command('start'))
async def cmd_start(message: Message):
"""Обработчик /start"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу до цього бота.')
return
await message.answer(
'👋 <b>Ласкаво просимо!</b>\n\n'
'🤖 <b>Бот для моніторингу графіку відключень світла</b>\n\n'
'✨ <b>Можливості:</b>\n'
'• 📊 Перегляд графіку на сьогодні і завтра\n'
'• 🔔 Автоматичні сповіщення за 5 хвилин до подій\n'
'• 🔄 Моніторинг змін графіку\n\n'
'Використовуйте кнопки нижче 👇',
parse_mode='HTML',
reply_markup=get_main_keyboard(),
)
@dp.message(lambda msg: msg.text == 'ℹ️ Про бота')
async def cmd_info(message: Message):
"""Показывает информацию о боте"""
if message.chat.id not in ALLOWED_CHAT_IDS:
return
await message.answer(
'<b>ℹ️ Про бота</b>\n\n'
'🚀 <b>Версія:</b> 2.2 (Стабільна)\n\n'
'📝 <b>Реліз-ноути:</b>\n'
'├ 15.11.2024: Адаптація під оновлену логіку сайту VOE\n'
'├ Виправлено парсинг half.left та half.right\n'
'├ Покращено визначення підтвердження відключень\n'
'└ Оптимізовано обробку статусу для всієї години\n\n'
'⚡ <b>Функціональність:</b>\n'
'├ Моніторинг графіку 24/7\n'
'├ Сповіщення за 5 хвилин\n'
'├ Детальна статистика дня\n'
'└ Красива візуалізація\n\n'
'🔐 <b>Безпека:</b> Використовуються .env файли\n'
'💾 <b>Джерело:</b> voe.com.ua',
parse_mode='HTML',
reply_markup=get_main_keyboard(),
)
@dp.message(Command('schedule'))
async def cmd_schedule(message: Message):
"""Показывает полный график"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('⏳ Завантаження графіку...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
schedule = parse_html(html)
text = format_schedule_message(schedule, days_to_show=2)
await message.answer(text, parse_mode='HTML', reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message(Command('today'))
async def cmd_today(message: Message):
"""Показывает график на сегодня"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('⏳ Завантаження графіку сьогодні...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
schedule = parse_html(html)
text = format_single_day_schedule(schedule, 0)
await message.answer(text, parse_mode='HTML', reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message(Command('tomorrow'))
async def cmd_tomorrow(message: Message):
"""Показывает график на завтра"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('⏳ Завантаження графіку завтра...')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
schedule = parse_html(html)
text = format_single_day_schedule(schedule, 1)
await message.answer(text, parse_mode='HTML', reply_markup=get_main_keyboard())
# await message.answer("❌ Функція тимчасово недоступна. Чекаємо на оновлення сайту", parse_mode="HTML", reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message(Command('check'))
async def cmd_check(message: Message):
"""Принудительная проверка графика"""
if message.chat.id not in ALLOWED_CHAT_IDS:
await message.answer('❌ У вас немає доступу.')
return
try:
await message.answer('🔄 <b>Перевіряю графік...</b>', parse_mode='HTML')
html = get_voe_html(VOE_CITY_ID, VOE_STREET_ID, VOE_HOUSE_ID)
new_schedule = parse_html(html)
prefix = (
'✅ <b>Знайдено зміни!</b>\n\n'
if schedules_differ(last_schedule, new_schedule)
else '✓ <b>Графік без змін</b>\n\n'
)
result = prefix + format_schedule_message(new_schedule, days_to_show=2)
await message.answer(result, parse_mode='HTML', reply_markup=get_main_keyboard())
except Exception as e:
await message.answer(f'❌ <b>Помилка:</b> {str(e)}', parse_mode='HTML', reply_markup=get_main_keyboard())
@dp.message()
async def handle_text(message: Message):
"""Обработчик текстовых сообщений и кнопок"""
if message.chat.id not in ALLOWED_CHAT_IDS:
return
text = message.text
# Кнопка "Графік"
if text == '📊 Графік':
await cmd_schedule(message)
# Кнопка "Сьогодні"
elif text == '📅 Сьогодні':
await cmd_today(message)
# Кнопка "Завтра"
elif text == '📅 Завтра':
await cmd_tomorrow(message)
# Кнопка "Оновити"
elif text == '🔄 Оновити':
await cmd_check(message)
# Кнопка "Про бота"
elif text == 'ℹ️ Про бота':
await cmd_info(message)
# Неизвестная команда
else:
await message.answer(
'❓ <b>Команда не розпізнана</b>\n\n'
'Використовуйте кнопки на клавіатурі або команди:\n'
'/start • /today • /tomorrow • /schedule • /check',
parse_mode='HTML',
reply_markup=get_main_keyboard(),
)
# ============================================================================
# ГЛАВНАЯ ФУНКЦИЯ
# ============================================================================
async def main():
"""Главная функция"""
logger.info('=' * 50)
logger.info('ЗАПУСК БОТА V2.2 (stable 2.2, 15.11.2025)')
logger.info('=' * 50)
if not TELEGRAM_TOKEN or os.getenv('TELEGRAM_TOKEN', 'YOUR_TOKEN_HERE') == TELEGRAM_TOKEN:
logger.error('❌ TELEGRAM_TOKEN не конфігурований! Напишіть токен в .env файл')
return
if not ALLOWED_CHAT_IDS:
logger.error('❌ ALLOWED_CHAT_IDS не конфігуровані! Напишіть ID в .env файл')
return
logger.info(f'📌 Allowed chat ids: {ALLOWED_CHAT_IDS}')
logger.info(f'⏱ Інтервал перевірки: {CHECK_INTERVAL} сек')
logger.info('=' * 50)
# Запускаем мониторинг
monitoring_task = asyncio.create_task(monitoring_loop())
try:
await dp.start_polling(bot)
except KeyboardInterrupt:
logger.info('⏹ Бот зупинений користувачем')
finally:
monitoring_task.cancel()
await bot.session.close()
logger.info('✓ Підключення закрито')
if __name__ == '__main__':
try:
asyncio.run(main())
except KeyboardInterrupt:
logger.info('⏹ Завершено')
+7
View File
@@ -0,0 +1,7 @@
[project]
name = "dtek-notif"
version = "0.1.0"
description = "Add your description here"
readme = "README.md"
requires-python = ">=3.13"
dependencies = []
+5
View File
@@ -0,0 +1,5 @@
requests>=2.31.0
beautifulsoup4>=4.12.0
aiogram>=3.3.0
python-dotenv>=1.0.0
aiohttp>=3.9.0
+14
View File
@@ -0,0 +1,14 @@
EDU_LOGIN=your_edu_login_here
EDU_PASSWORD=your_edu_password_here
EDU_URL_LOGIN=https://edu.edu.vn.ua/user/login
EDU_URL_VERIFY=https://edu.edu.vn.ua/course/userlist
PHPSESSID_INTERVAL=10
USER_AGENT="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36"
WEBINAR_URL=https://edu.edu.vn.ua/webinar/useractive
WEBINAR_CHECK_INTERVAL=60
REDIS_HOST=redis
REDIS_PORT=6379
PLAYWRIGHT_WS=ws://playwright-service:3000/ws
TZ=Europe/Kyiv
WEBINAR_TELEGRAM_TOKEN=your_telegram_bot_token_here
WEBINAR_ADMIN_ID=123456789
+1
View File
@@ -0,0 +1 @@
1.56.0
+49
View File
@@ -0,0 +1,49 @@
services:
redis:
image: redis:8.10.1-alpine
restart: unless-stopped
volumes:
- redis-data:/data
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 5s
timeout: 3s
retries: 5
playwright-service:
image: mcr.microsoft.com/playwright:v1.56.0-jammy
restart: unless-stopped
command: npx -y playwright@1.56.0 run-server --port 3000 --path /ws
session-keeper:
build: ./phpsessid-bot
image: gcr.forust.xyz/forust/session-keeper:latest
pull_policy: build
env_file: .env
restart: unless-stopped
depends_on:
redis:
condition: service_healthy
healthcheck:
test: ["CMD-SHELL", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
interval: 30s
timeout: 5s
retries: 10
start_period: 60s
webinar-checker:
build: ./webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:latest
pull_policy: build
env_file: .env
restart: unless-stopped
depends_on:
redis:
condition: service_healthy
session-keeper:
condition: service_healthy
playwright-service:
condition: service_started
volumes:
redis-data:
File renamed without changes.
+77
View File
@@ -0,0 +1,77 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: edu-master-webinar
namespace: edu-master
labels:
release: prometheus-stack
spec:
groups:
- name: edu_master.webinar
rules:
# No successful webinar check for 5m (~2-3 missed 2-min checks).
# Catches: playwright hangs/timeouts, version skew, site changes, hung job.
- alert: WebinarCheckerNoSuccessfulCheck
expr: |
(time() - webinar_check_last_success_timestamp_seconds > 300)
and (webinar_check_last_run_timestamp_seconds > 0)
for: 2m
labels:
severity: critical
annotations:
summary: "Webinar checker has no successful check for 5m"
description: "edu-master/webinar-checker: last successful webinar check was {{ $value | humanizeDuration }} ago. Checks are failing or hanging (see consecutive failures alert). Notifications about new webinars are NOT being sent."
# Fast path: 3 consecutive failures (~6+ min at 2-min interval).
- alert: WebinarCheckerConsecutiveFailures
expr: |
webinar_check_consecutive_failures >= 3
for: 5m
labels:
severity: critical
annotations:
summary: "Webinar checker failing consecutively"
description: 'edu-master/webinar-checker: {{ $value }} consecutive webinar check failures (timeout / playwright error / page error). Check pod logs (Loki: {namespace="edu-master", container="webinar-checker"}).'
# Metrics endpoint not scraped for 10m: pod down, metrics server dead, or ServiceMonitor broken.
- alert: WebinarCheckerScrapeDown
expr: |
absent(webinar_check_last_run_timestamp_seconds) == 1
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker metrics missing"
description: "edu-master/webinar-checker: no metrics series for 10m. Pod may be down, metrics server dead, or ServiceMonitor/Service broken. Webinar checks are unobserved."
# EDU session lost: session-keeper down or credentials expired. Without PHPSESSID every check is skipped.
- alert: EduPhpsessidMissing
expr: |
edu_phpsessid_present == 0
for: 10m
labels:
severity: critical
annotations:
summary: "EDU_PHPSESSID missing"
description: "edu-master: EDU_PHPSESSID absent from redis for 10m. Webinar/diari/schedule checks are all skipped. Check session-keeper logs and EDU credentials."
# Hard deps: checker and playwright deployments unavailable.
- alert: WebinarCheckerDeploymentDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="webinar-checker", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Webinar checker deployment unavailable"
description: "edu-master/webinar-checker deployment has {{ $value }} unavailable replica(s) for 10m."
- alert: PlaywrightServiceDown
expr: |
kube_deployment_status_replicas_unavailable{deployment="playwright-service", namespace="edu-master"} > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Playwright service unavailable"
description: "edu-master/playwright-service deployment has {{ $value }} unavailable replica(s) for 10m. All webinar/diari/schedule checks fail without it."
@@ -1,4 +1,4 @@
apiVersion: v1 apiVersion: v1
kind: Namespace kind: Namespace
metadata: metadata:
name: homarr name: edu-master
+58
View File
@@ -0,0 +1,58 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: playwright-service
namespace: edu-master
labels:
app: edu-master-playwright
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-playwright
template:
metadata:
labels:
app: edu-master-playwright
spec:
containers:
- name: playwright
# renovate: datasource=docker depName=mcr.microsoft.com/playwright versioning=docker
image: mcr.microsoft.com/playwright:v1.56.0-jammy
imagePullPolicy: IfNotPresent
command:
- npx
- -y
- playwright@1.56.0
- run-server
- --port
- "3000"
- --path
- /ws
ports:
- containerPort: 3000
readinessProbe:
tcpSocket:
port: 3000
initialDelaySeconds: 5
periodSeconds: 10
timeoutSeconds: 3
livenessProbe:
tcpSocket:
port: 3000
initialDelaySeconds: 15
periodSeconds: 20
timeoutSeconds: 3
---
apiVersion: v1
kind: Service
metadata:
name: playwright-service
namespace: edu-master
spec:
selector:
app: edu-master-playwright
ports:
- name: ws
port: 3000
targetPort: 3000
+75
View File
@@ -0,0 +1,75 @@
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: redis
namespace: edu-master
labels:
app: edu-master-redis
spec:
serviceName: redis
replicas: 1
selector:
matchLabels:
app: edu-master-redis
template:
metadata:
labels:
app: edu-master-redis
spec:
containers:
- name: redis
image: redis:8.10.1-alpine
imagePullPolicy: IfNotPresent
ports:
- containerPort: 6379
volumeMounts:
- name: redis-data
mountPath: /data
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 250m
memory: 256Mi
readinessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 3
livenessProbe:
exec:
command: ["redis-cli", "ping"]
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 3
volumes:
- name: redis-data
persistentVolumeClaim:
claimName: redis-data-pvc
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: redis-data-pvc
namespace: edu-master
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 1Gi
---
apiVersion: v1
kind: Service
metadata:
name: redis
namespace: edu-master
spec:
selector:
app: edu-master-redis
ports:
- name: redis
port: 6379
targetPort: 6379
@@ -0,0 +1,50 @@
# One-time Job to migrate redis state from docker compose to k8s (maintenance window).
# The .example file is not applied by the deploy pipeline (mask *.example.yaml).
#
# Runbook:
# 1. docker compose -f <repo>/edu_master/compose.yaml stop # SIGTERM -> redis will flush dump.rdb
# 2. docker run --rm -v edu_master_redis-data:/data \
# -v /tmp/edu-master-backup:/backup \
# redis:alpine sh -c "cp /data/dump.rdb /backup/ && ls -la /backup"
# 3. kubectl apply -f edu_master/k8s/namespace.yaml
# 4. kubectl apply -f <only the PVC from redis.yaml> # seed must come BEFORE redis pod starts
# 5. kubectl apply -f edu_master/k8s/restore-seed-job.yaml.example
# kubectl wait --for=condition=complete job/redis-restore-seed -n edu-master --timeout=120s
# 6. kubectl delete job redis-restore-seed -n edu-master
# 7. kubectl apply -f edu_master/k8s/ -R # apply remaining manifests
apiVersion: batch/v1
kind: Job
metadata:
name: redis-restore-seed
namespace: edu-master
spec:
backoffLimit: 2
ttlSecondsAfterFinished: 3600
template:
spec:
restartPolicy: Never
containers:
- name: seed
image: redis:alpine
command:
- /bin/sh
- -ec
- |
ls -la /backup
cp /backup/dump.rdb /data/dump.rdb
chmod 644 /data/dump.rdb
ls -la /data
volumeMounts:
- name: redis-data
mountPath: /data
- name: backup
mountPath: /backup
readOnly: true
volumes:
- name: redis-data
persistentVolumeClaim:
claimName: redis-data-pvc
- name: backup
hostPath:
path: /tmp/edu-master-backup
type: DirectoryOrCreate
+29
View File
@@ -0,0 +1,29 @@
apiVersion: v1
kind: Secret
metadata:
name: edu-master-secrets
namespace: edu-master
type: Opaque
stringData:
# Session keeper credentials
KEEPER_LOGIN: ""
KEEPER_PASSWORD: ""
KEEPER_INTERVAL: "10"
# EDU links
EDU_URL_BASE: "https://edu.edu.vn.ua"
EDU_URL_LOGIN: "/user/login"
EDU_URL_COURSES: "/course/userlist"
EDU_URL_WEBINAR: "/webinar/useractive"
# Playwright
USER_AGENT: ""
PLAYWRIGHT_WS: "ws://playwright-service:3000/ws"
# Webinar-checker
WEBINAR_TELEGRAM_TOKEN: ""
WEBINAR_ADMIN_ID: ""
WEBINAR_CHECK_INTERVAL: "60"
# Prometheus metrics endpoint (scraped via ServiceMonitor, alerts in k8s/alerts.yaml)
METRICS_PORT: "8000"
# Database
REDIS_HOST: "redis"
REDIS_PORT: "6379"
TZ: "Europe/Kyiv"
+15
View File
@@ -0,0 +1,15 @@
apiVersion: v1
kind: Service
metadata:
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
selector:
app: edu-master-webinar-checker
ports:
- name: metrics
port: 8000
targetPort: metrics
protocol: TCP
@@ -1,14 +1,14 @@
apiVersion: monitoring.coreos.com/v1 apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor kind: ServiceMonitor
metadata: metadata:
name: netbird-server name: webinar-checker
namespace: netbird namespace: edu-master
labels: labels:
release: prometheus-stack release: prometheus-stack
spec: spec:
selector: selector:
matchLabels: matchLabels:
app: netbird-server app: edu-master-webinar-checker
endpoints: endpoints:
- port: metrics - port: metrics
path: /metrics path: /metrics
+52
View File
@@ -0,0 +1,52 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: session-keeper
namespace: edu-master
labels:
app: edu-master-session-keeper
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-session-keeper
template:
metadata:
labels:
app: edu-master-session-keeper
spec:
initContainers:
- name: wait-redis
image: redis:8.10.1-alpine
command:
- /bin/sh
- -ec
- |
i=0
until redis-cli -h redis ping | grep -q PONG; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
sleep 2
done
echo "redis is ready"
containers:
- name: session-keeper
image: gcr.forust.xyz/forust/session-keeper:latest
imagePullPolicy: Always
envFrom:
- secretRef:
name: edu-master-secrets
resources:
requests:
cpu: 25m
memory: 96Mi
limits:
cpu: 250m
memory: 256Mi
readinessProbe:
exec:
command: ["/bin/sh", "-ec", "redis-cli -h redis EXISTS EDU_PHPSESSID | grep -q 1"]
initialDelaySeconds: 15
periodSeconds: 30
timeoutSeconds: 5
failureThreshold: 10
+66
View File
@@ -0,0 +1,66 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: webinar-checker
namespace: edu-master
labels:
app: edu-master-webinar-checker
spec:
replicas: 1
selector:
matchLabels:
app: edu-master-webinar-checker
template:
metadata:
labels:
app: edu-master-webinar-checker
spec:
# Enforces dependency order like compose depends_on:
# redis healthy -> session-keeper healthy (EXISTS EDU_PHPSESSID) -> playwright started
initContainers:
- name: wait-deps
image: redis:8.10.1-alpine
command:
- /bin/sh
- -ec
- |
i=0
until redis-cli -h redis ping | grep -q PONG; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: redis not ready" && exit 1
sleep 2
done
echo "redis ok"
until [ "$(redis-cli -h redis EXISTS EDU_PHPSESSID)" = "1" ]; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: no PHPSESSID (session-keeper down?)" && exit 1
sleep 2
done
echo "PHPSESSID ok"
until nc -z playwright-service 3000; do
i=$((i+1))
[ "$i" -ge 300 ] && echo "TIMEOUT: playwright-service not reachable" && exit 1
sleep 2
done
echo "playwright ok"
containers:
- name: webinar-checker
image: gcr.forust.xyz/forust/webinar-checker:latest
imagePullPolicy: Always
ports:
- name: metrics
containerPort: 8000
protocol: TCP
envFrom:
- secretRef:
name: edu-master-secrets
env:
- name: TZ
value: "Europe/Kyiv"
resources:
requests:
cpu: "50m"
memory: "128Mi"
limits:
cpu: "600m"
memory: "512Mi"
+15
View File
@@ -0,0 +1,15 @@
FROM python:3.11-slim
WORKDIR /app
# Install system dependencies
RUN apt-get update && apt-get install -y --no-install-recommends redis-tools && rm -rf /var/lib/apt/lists/*
# Install dependencies
RUN pip install --no-cache-dir requests==2.32.3 redis==5.2.1
# Copy application code
COPY . .
# Run the bot
CMD ["python", "bot.py"]
+132
View File
@@ -0,0 +1,132 @@
import logging
import os
import time
from datetime import datetime
import redis
import requests
# Configure logging
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s')
logger = logging.getLogger(__name__)
# Load configuration (adapted to .env keys)
def _env(key, default=None):
v = os.getenv(key, default)
if isinstance(v, str) and len(v) >= 2 and ((v[0] == '"' and v[-1] == '"') or (v[0] == "'" and v[-1] == "'")):
return v[1:-1]
return v
LOGIN = _env('KEEPER_LOGIN')
PASSWORD = _env('KEEPER_PASSWORD')
EDU_BASE = _env('EDU_URL_BASE', 'https://edu.edu.vn.ua')
EDU_LOGIN_PATH = _env('EDU_URL_LOGIN', '/user/login')
EDU_COURSES_PATH = _env('EDU_URL_COURSES', '/course/userlist')
URL_LOGIN = f'{EDU_BASE.rstrip("/")}/{EDU_LOGIN_PATH.lstrip("/")}'
URL_VERIFY = f'{EDU_BASE.rstrip("/")}/{EDU_COURSES_PATH.lstrip("/")}'
INTERVAL = int(_env('KEEPER_INTERVAL', 10))
USER_AGENT = _env(
'USER_AGENT',
'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/142.0.0.0 Safari/537.36',
)
REDIS_HOST = _env('REDIS_HOST', 'redis')
REDIS_PORT = int(_env('REDIS_PORT', 6379))
SUCCESS_FILE = '/tmp/last_success' # noqa: S108
def touch_success_file():
"""Updates the timestamp of the success file for healthchecks."""
try:
with open(SUCCESS_FILE, 'w') as f:
f.write(str(datetime.now().timestamp()))
except Exception as e:
logger.error(f'Failed to touch success file: {e}')
def main():
logger.info('Starting Session Keeper Bot')
# Connect to Redis
try:
redis_client = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
redis_client.ping()
logger.info(f'Connected to Redis at {REDIS_HOST}:{REDIS_PORT}')
except Exception as e:
logger.error(f'Failed to connect to Redis: {e}')
return
session = requests.Session()
# Set headers
headers = {
'User-Agent': USER_AGENT,
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7',
'Accept-Language': 'en-US,en;q=0.9',
'Cache-Control': 'max-age=0',
'Upgrade-Insecure-Requests': '1',
'Sec-Fetch-Site': 'same-origin',
'Sec-Fetch-Mode': 'navigate',
'Sec-Fetch-User': '?1',
'Sec-Fetch-Dest': 'document',
'Sec-Ch-Ua': '"Not_A Brand";v="99", "Chromium";v="142"',
'Sec-Ch-Ua-Mobile': '?0',
'Sec-Ch-Ua-Platform': '"Linux"',
'Accept-Encoding': 'gzip, deflate, br',
'Priority': 'u=0, i',
}
session.headers.update(headers)
while True:
try:
logger.info('Attempting login...')
# Login payload
payload = {'login': LOGIN, 'password': PASSWORD}
# Perform Login
# Note: The user request shows a POST to /user/login with form data
# We need to make sure we handle the PHPSESSID correctly.
# If we already have a PHPSESSID, requests will send it.
login_response = session.post(URL_LOGIN, data=payload, allow_redirects=True)
logger.info(f'Login Response Status: {login_response.status_code}')
logger.info(f'Cookies after login: {session.cookies.get_dict()}')
# Verify Session
logger.info('Verifying session...')
verify_response = session.get(URL_VERIFY, allow_redirects=False)
logger.info(f'Verify Response Status: {verify_response.status_code}')
if verify_response.status_code == 200:
logger.info('Session verification SUCCESS (200 OK).')
touch_success_file()
# Save PHPSESSID to Redis
phpsessid = session.cookies.get('PHPSESSID')
if phpsessid:
try:
redis_client.set('EDU_PHPSESSID', phpsessid)
logger.info(f'Saved PHPSESSID to Redis: {phpsessid}')
except Exception as e:
logger.error(f'Failed to save PHPSESSID to Redis: {e}')
elif verify_response.status_code == 302:
logger.warning('Session verification FAILED (302 Redirect). Session might be invalid.')
else:
logger.warning(f'Session verification returned unexpected status: {verify_response.status_code}')
except Exception as e:
logger.error(f'An error occurred: {e}')
logger.info(f'Sleeping for {INTERVAL} minutes...')
time.sleep(INTERVAL * 60)
if __name__ == '__main__':
main()
+13
View File
@@ -0,0 +1,13 @@
FROM python:3.11-slim
WORKDIR /app
# renovate: datasource=pypi depName=playwright versioning=pep440
ARG PLAYWRIGHT_VERSION=1.56.0
# Install dependencies - PLAYWRIGHT_VERSION is single-source, renovate updates ARG above and all other places via regexManagers
RUN pip install --no-cache-dir pip==25.0.1 && pip install --no-cache-dir playwright==${PLAYWRIGHT_VERSION} redis==5.2.1 requests==2.32.3 "python-telegram-bot[job-queue]==21.10"
COPY checker.py .
CMD ["python", "checker.py"]
File diff suppressed because it is too large. Load diff
-21
View File
@@ -1,21 +0,0 @@
# Error pages
Static HTTP error pages served by an Nginx image built in CI.
Edit the HTML in `html/`; the Dockerfile copies it into the image.
Kubernetes exposes `error-pages-service` in `error-pages` for Traefik's error
middleware. Keep the middleware's namespace and port aligned with that Service.
For a local build, run `docker build -t homelab-error-pages .` from this directory.
Compose references the private registry image rather than a build context.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n error-pages
kubectl get events -n error-pages --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+1 -1
View File
@@ -1,7 +1,7 @@
services: services:
errorpage: errorpage:
build: . build: .
image: gcr.forust.xyz/forust/error-pages:prod image: gcr.forust.xyz/forust/error-pages:latest
pull_policy: build pull_policy: build
container_name: error-pages container_name: error-pages
restart: unless-stopped restart: unless-stopped
+1 -17
View File
@@ -20,8 +20,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: error-pages app: error-pages
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -29,21 +27,7 @@ spec:
spec: spec:
containers: containers:
- name: error-pages - name: error-pages
image: gcr.forust.xyz/forust/error-pages:prod image: gcr.forust.xyz/forust/error-pages:latest
# p95 6M, max 10M, no limit before.
resources:
requests:
cpu: "10m"
memory: "32Mi"
limits:
memory: "128Mi"
ports: ports:
- containerPort: 80 - containerPort: 80
readinessProbe:
httpGet:
path: /404.html
port: 80
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
--- ---
-25
View File
@@ -1,25 +0,0 @@
# Gitea
Git hosting with HTTP and a separate SSH route.
Kubernetes uses the shared PostgreSQL service and `gitea-pvc` for repositories
and application data. Match the Gitea database password with the shared database
Secret. SSH is routed through Traefik's TCP entrypoint on 2221.
Compose uses a separate PostgreSQL 14 database, bind mounts `gitea-data/` and
`gitea-db/`, and publishes host port 2221. It is an alternative deployment with
its own database, not a second frontend for the Kubernetes instance.
Back up repositories, application configuration, and a consistent database dump
together. Gitea Actions definitions for this repository live in `../.gitea/`.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n gitea
kubectl get events -n gitea --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
+3 -6
View File
@@ -1,6 +1,6 @@
services: services:
server: server:
image: docker.gitea.com/gitea:28.0.0 image: docker.gitea.com/gitea:1.27.3
container_name: gitea container_name: gitea
restart: always restart: always
environment: environment:
@@ -13,12 +13,9 @@ services:
- GITEA__database__PASSWD=gitea - GITEA__database__PASSWD=gitea
- GITEA__database__NAME=gitea - GITEA__database__NAME=gitea
# Server # Server
- GITEA__server__ROOT_URL=https://git.forust.xyz - GITEA__server__ROOT_URL=https://gitea.forust.xyz
- GITEA__server__SSH_DOMAIN=gitssh.forust.xyz - GITEA__server__SSH_DOMAIN=gitssh.forust.xyz
- GITEA__server__SSH_PORT=2221 - GITEA__server__SSH_PORT=2221
# Pin 28.0 defaults explicitly (see k8s/config.yaml for rationale)
- GITEA__service__DISABLE_REGISTRATION=true
- GITEA__actions__RUN_RETENTION_DAYS=90
# Mailer # Mailer
- GITEA__mailer__ENABLED=true - GITEA__mailer__ENABLED=true
- GITEA__mailer__FROM=${SERVICE_EMAIL} - GITEA__mailer__FROM=${SERVICE_EMAIL}
@@ -37,7 +34,7 @@ services:
- "traefik.http.services.gitea.loadbalancer.server.port=3000" - "traefik.http.services.gitea.loadbalancer.server.port=3000"
# Prod Router # Prod Router
- "traefik.http.routers.gitea.rule=Host(`git.forust.xyz`) || Host(`gitea.forust.xyz`)" - "traefik.http.routers.gitea.rule=Host(`gitea.forust.xyz`)"
- "traefik.http.routers.gitea.entrypoints=websecure" - "traefik.http.routers.gitea.entrypoints=websecure"
- "traefik.http.routers.gitea.tls.certresolver" - "traefik.http.routers.gitea.tls.certresolver"
# Local Router # Local Router
-1
View File
@@ -8,7 +8,6 @@ spec:
dnsNames: dnsNames:
- gcr.forust.xyz - gcr.forust.xyz
- gitea.forust.xyz - gitea.forust.xyz
- git.forust.xyz
issuerRef: issuerRef:
name: letsencrypt-prod name: letsencrypt-prod
kind: ClusterIssuer kind: ClusterIssuer
+2 -13
View File
@@ -4,14 +4,11 @@ metadata:
name: gitea-config name: gitea-config
namespace: gitea namespace: gitea
data: data:
GITEA__server__ROOT_URL: "https://git.forust.xyz" GITEA__server__DOMAIN: "gitea.forust.xyz"
GITEA__server__ROOT_URL: "https://gitea.forust.xyz"
GITEA__server__SSH_DOMAIN: "gitssh.forust.xyz" GITEA__server__SSH_DOMAIN: "gitssh.forust.xyz"
GITEA__server__SSH_PORT: "2221" GITEA__server__SSH_PORT: "2221"
GITEA__service__DISABLE_REGISTRATION: "true"
GITEA__actions__RUN_RETENTION_DAYS: "90"
GITEA__database__DB_TYPE: "postgres" GITEA__database__DB_TYPE: "postgres"
GITEA__database__HOST: "postgres.database.svc.cluster.local:5432" GITEA__database__HOST: "postgres.database.svc.cluster.local:5432"
GITEA__database__NAME: "gitea" GITEA__database__NAME: "gitea"
@@ -20,14 +17,6 @@ data:
GITEA__mailer__ENABLED: "false" GITEA__mailer__ENABLED: "false"
GITEA__metrics__ENABLED: "true"
# No code/issue search needed: bleve reindexes the whole issue index on
# every pod restart (cron.rebuild_issue_indexer RUN_AT_START) and hammers
# the rotational disk for an hour. "db" serves issue search from postgres.
GITEA__indexer__ISSUE_INDEXER_TYPE: "db"
GITEA__indexer__REPO_INDEXER_ENABLED: "false"
GITEA__log__logger.access.MODE: "console, file" GITEA__log__logger.access.MODE: "console, file"
USER_UID: "1000" USER_UID: "1000"
USER_GID: "1000" USER_GID: "1000"
+4 -10
View File
@@ -3,8 +3,6 @@ kind: Service
metadata: metadata:
name: gitea-service name: gitea-service
namespace: gitea namespace: gitea
labels:
app: gitea
spec: spec:
selector: selector:
app: gitea app: gitea
@@ -19,8 +17,6 @@ spec:
apiVersion: apps/v1 apiVersion: apps/v1
kind: Deployment kind: Deployment
metadata: metadata:
annotations:
reloader.stakater.com/auto: "true"
name: gitea-deployment name: gitea-deployment
namespace: gitea namespace: gitea
spec: spec:
@@ -28,8 +24,6 @@ spec:
selector: selector:
matchLabels: matchLabels:
app: gitea app: gitea
strategy:
type: Recreate
template: template:
metadata: metadata:
labels: labels:
@@ -37,7 +31,7 @@ spec:
spec: spec:
containers: containers:
- name: gitea - name: gitea
image: gitea/gitea:28.0.0 image: gitea/gitea:1.27.3
envFrom: envFrom:
- configMapRef: - configMapRef:
name: gitea-config name: gitea-config
@@ -53,10 +47,10 @@ spec:
mountPath: /data mountPath: /data
resources: resources:
requests: requests:
memory: "320Mi" memory: "512Mi"
cpu: "100m" cpu: "300m"
limits: limits:
memory: "1Gi" memory: "1.5Gi"
cpu: "1300m" cpu: "1300m"
volumes: volumes:
- name: gitea-data - name: gitea-data
+8 -3
View File
@@ -7,14 +7,19 @@ spec:
entryPoints: entryPoints:
- websecure - websecure
routes: routes:
# Metrics are scraped directly through the cluster Service. - match: Host(`gitea.forust.xyz`)
- match: (Host(`gitea.forust.xyz`) || Host(`git.forust.xyz`)) && !PathPrefix(`/metrics`)
kind: Rule kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: gitea-service - name: gitea-service
port: 3000 port: 3000
- match: Host(`gcr.forust.xyz`) && PathPrefix(`/v2`) - match: Host(`gcr.forust.xyz`) && PathPrefix(`/v2`)
kind: Rule kind: Rule
middlewares:
- name: crowdsec-bouncer
namespace: crowdsec
services: services:
- name: gitea-service - name: gitea-service
port: 3000 port: 3000
@@ -30,7 +35,7 @@ spec:
entryPoints: entryPoints:
- websecure - websecure
routes: routes:
- match: (Host(`gitea.workstation.internal`) || Host(`gitea.gigaforust.internal`)) || (Host(`git.workstation.internal`) || Host(`git.gigaforust.internal`)) - match: Host(`gitea.workstation.internal`) || Host(`gitea.gigaforust.internal`)
kind: Rule kind: Rule
services: services:
- name: gitea-service - name: gitea-service
-16
View File
@@ -1,16 +0,0 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: gitea
namespace: gitea
labels:
release: prometheus-stack
spec:
selector:
matchLabels:
app: gitea
endpoints:
- port: http
path: /metrics
interval: 30s
scrapeTimeout: 10s
-26
View File
@@ -1,26 +0,0 @@
# Glance
Dashboard pages for links, service checks, and Docker containers.
Compose mounts `config/` and `assets/`. The Kubernetes equivalents are embedded
in `k8s/glance-config.yaml`: `glance-config` holds pages and `glance-assets` holds
`user.css`. Update both copies when changing shared content.
Kubernetes serves the dashboard under `/glance`. Its pod also mounts the node's
Docker socket. It references `glance-secrets` for `ADGUARD_PASSWORD`, but there is
no tracked Secret example; create that Secret in `glance` before starting it.
Compose expects a local `.env` with the same password.
The pod mounts `user.css` from `glance-assets`, which is the ConfigMap that
contains that key.
## Inspect
From the repository root:
```sh
kubectl get pods,svc,pvc -n glance
kubectl get events -n glance --sort-by=.metadata.creationTimestamp
```
See the [repository README](../README.md) for deployment selection.
Loaded 100 of 439 files, more files were not shown because too many files have changed in this diff. Show more